{
  "id": 238630,
  "title": "[LB 0.91] Let's talk about RNNs...",
  "url": "/competitions/seti-breakthrough-listen/discussion/238630",
  "author_name": "xhlulu",
  "post_date": "2021-05-12T21:03:58.407000",
  "votes": 24,
  "comment_count": 19,
  "views": 0,
  "content": "<p>I just released a notebook that establishes a baseline for RNN - <a href=\"https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras\" target=\"_blank\">Bidirectional GRU in Keras</a>. </p>\n<p>At time of writing this, the score sits at LB 0.85, which is lower than the top CNN models. Moreover, there's a few things that could be improved.</p>\n<ol>\n<li>How would you deal with the 6 positions? At the moment I'm apply a shared RNN layer on each position separately, then combining them using an average layer. However, would it make more sense to combine them before applying the RNN, or after? Is there some way to share the information? and would it make sense to concatenate instead of averaging?</li>\n<li>What are the best hyperparameters? How many layers &amp; hidden units, what learning rate, and what optimizer would you use?</li>\n<li>Is GRU the best model for this task? Would LSTM, or some other sequential architecture (e.g. transformers) make more sense?</li>\n<li>How could we speed up the data loading so it doesn't slow down the actually forward/backward pass on the GPU?</li>\n</ol>\n<p>I'm happy to hear your thoughts on what is the best way to approach them, and I'm looking forward incorporating the feedback into the notebook.</p>\n<p>EDIT: I tried concatenating the pooled output of the RNNs instead of averaging. I also used the entire dataset instead of just a sampled subset of the negative results. With those two changes I got an improve LB 0.85 -&gt; LB 0.9</p>",
  "messages": [
    {
      "id": 1304761,
      "postDate": "2021-05-12T21:03:58.407Z",
      "content": "<p>I just released a notebook that establishes a baseline for RNN - <a href=\"https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras\" target=\"_blank\">Bidirectional GRU in Keras</a>. </p>\n<p>At time of writing this, the score sits at LB 0.85, which is lower than the top CNN models. Moreover, there's a few things that could be improved.</p>\n<ol>\n<li>How would you deal with the 6 positions? At the moment I'm apply a shared RNN layer on each position separately, then combining them using an average layer. However, would it make more sense to combine them before applying the RNN, or after? Is there some way to share the information? and would it make sense to concatenate instead of averaging?</li>\n<li>What are the best hyperparameters? How many layers &amp; hidden units, what learning rate, and what optimizer would you use?</li>\n<li>Is GRU the best model for this task? Would LSTM, or some other sequential architecture (e.g. transformers) make more sense?</li>\n<li>How could we speed up the data loading so it doesn't slow down the actually forward/backward pass on the GPU?</li>\n</ol>\n<p>I'm happy to hear your thoughts on what is the best way to approach them, and I'm looking forward incorporating the feedback into the notebook.</p>\n<p>EDIT: I tried concatenating the pooled output of the RNNs instead of averaging. I also used the entire dataset instead of just a sampled subset of the negative results. With those two changes I got an improve LB 0.85 -&gt; LB 0.9</p>",
      "rawMarkdown": "I just released a notebook that establishes a baseline for RNN - [Bidirectional GRU in Keras](https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras). \n\nAt time of writing this, the score sits at LB 0.85, which is lower than the top CNN models. Moreover, there's a few things that could be improved.\n\n1. How would you deal with the 6 positions? At the moment I'm apply a shared RNN layer on each position separately, then combining them using an average layer. However, would it make more sense to combine them before applying the RNN, or after? Is there some way to share the information? and would it make sense to concatenate instead of averaging?\n2. What are the best hyperparameters? How many layers & hidden units, what learning rate, and what optimizer would you use?\n3. Is GRU the best model for this task? Would LSTM, or some other sequential architecture (e.g. transformers) make more sense?\n4. How could we speed up the data loading so it doesn't slow down the actually forward/backward pass on the GPU?\n\nI'm happy to hear your thoughts on what is the best way to approach them, and I'm looking forward incorporating the feedback into the notebook.\n\nEDIT: I tried concatenating the pooled output of the RNNs instead of averaging. I also used the entire dataset instead of just a sampled subset of the negative results. With those two changes I got an improve LB 0.85 -> LB 0.9",
      "votes": 24
    },
    {
      "id": 1306574,
      "postDate": "2021-05-13T22:13:52.577Z",
      "content": "<p>Thanks for starting this interesting thread and creating a starting nb! I'm really curious to see if RNN would make difference to this problem - however it is still not very clear to me how we are trying to exploit the temporal information ? I guess we speak about the \"intra-cadence\" relationships ? ie the sixtuple sequence ABACAD = 6 channel info   </p>\n<p>If that is the case then I would try first to combine them - or better not to split by input head - and then passing to RNN.  </p>\n<p>For points 2,3 need to make some experiments fist but I'm out of gpu quota :(</p>\n<p>PS: Still remember your famous GRU kernel in OpenVaccine comp. ;)</p>",
      "rawMarkdown": "Thanks for starting this interesting thread and creating a starting nb! I'm really curious to see if RNN would make difference to this problem - however it is still not very clear to me how we are trying to exploit the temporal information ? I guess we speak about the \"intra-cadence\" relationships ? ie the sixtuple sequence ABACAD = 6 channel info   \n\nIf that is the case then I would try first to combine them - or better not to split by input head - and then passing to RNN.  \n\nFor points 2,3 need to make some experiments fist but I'm out of gpu quota :(\n\nPS: Still remember your famous GRU kernel in OpenVaccine comp. ;)",
      "votes": 1,
      "replies": [
        {
          "id": 1307773,
          "postDate": "2021-05-14T16:43:30.900Z",
          "content": "<p>Hahaha at this rate i'll be known as the GRU guy ;)</p>\n<p>In my notebook, the sequence in this case is the 1st dimension of the 2d spectrogram. I'm running the RNN on each of the 6 channels separately, and afterwards I concatenate them.</p>\n<p>However I can't say if there's a \"temporal\" significance to the 1st dimension of the spectrogram; happy if someone can correct my approach here!</p>",
          "rawMarkdown": "Hahaha at this rate i'll be known as the GRU guy ;)\n\nIn my notebook, the sequence in this case is the 1st dimension of the 2d spectrogram. I'm running the RNN on each of the 6 channels separately, and afterwards I concatenate them.\n\nHowever I can't say if there's a \"temporal\" significance to the 1st dimension of the spectrogram; happy if someone can correct my approach here!",
          "votes": 3
        },
        {
          "id": 1320506,
          "postDate": "2021-05-24T06:20:10.067Z",
          "content": "<p>I think RNN time series combined with CNN model can also get better results, which can be used as a branch of ensemble. But at present, how to effectively use the 273 time dimension is a problem. I saw the information about GRU on the published notebook, but there is no in-depth study. I just used the code of that notebook and changed it to nfnet_ L0 model, got 0.96 +</p>",
          "rawMarkdown": "I think RNN time series combined with CNN model can also get better results, which can be used as a branch of ensemble. But at present, how to effectively use the 273 time dimension is a problem. I saw the information about GRU on the published notebook, but there is no in-depth study. I just used the code of that notebook and changed it to nfnet_ L0 model, got 0.96 +"
        },
        {
          "id": 1324707,
          "postDate": "2021-05-27T07:09:42.810Z",
          "content": "<p>You put nfnet_l0 before RNN?</p>",
          "rawMarkdown": "You put nfnet_l0 before RNN?"
        }
      ]
    },
    {
      "id": 1320219,
      "postDate": "2021-05-23T21:48:16.953Z",
      "content": "<p>Hello, </p>\n<p>For the first point, I'm currently trying another approach with two input in my model. I try to separate the A signals and the B, C and D. So, one input it will have the three A signal and for the other input the three others. I don't really know if it's possible to do something like that, I'm afraid I maybe lose some sense in the data. </p>\n<p>Currently, in my model, I used LSTM. There are more time consuming than GRU, but as the model is simple, I prefer to use them. However, I didn't take the time to use the same architecture as yours to truly see if there is a noticeable difference between GRU and LSTM. I will maybe try it in the next days and put an update on this post. </p>\n<p>At the moment, I more interest in the transposition you've made in your kernel. I need to maybe more understand the data ^^</p>\n<p>For the second point, in order to optimize your model, if think Dropout can be useful. For the learning rate, I personally use the ReduceLROnPlateau callback from the Keras library. It allows you to reduce the learning rate when, for example, the loss in the validation data doesn't decrease.</p>",
      "rawMarkdown": "Hello, \n\nFor the first point, I'm currently trying another approach with two input in my model. I try to separate the A signals and the B, C and D. So, one input it will have the three A signal and for the other input the three others. I don't really know if it's possible to do something like that, I'm afraid I maybe lose some sense in the data. \n\nCurrently, in my model, I used LSTM. There are more time consuming than GRU, but as the model is simple, I prefer to use them. However, I didn't take the time to use the same architecture as yours to truly see if there is a noticeable difference between GRU and LSTM. I will maybe try it in the next days and put an update on this post. \n\nAt the moment, I more interest in the transposition you've made in your kernel. I need to maybe more understand the data ^^\n\nFor the second point, in order to optimize your model, if think Dropout can be useful. For the learning rate, I personally use the ReduceLROnPlateau callback from the Keras library. It allows you to reduce the learning rate when, for example, the loss in the validation data doesn't decrease.\n",
      "votes": 2,
      "replies": [
        {
          "id": 1320325,
          "postDate": "2021-05-24T02:22:27.137Z",
          "content": "<p>Building a tf.keras model with two inputs is pretty easy stuff - the very hard part for me is building the data generator :)</p>\n<p>Let us know if you have success getting it to run.</p>",
          "rawMarkdown": "Building a tf.keras model with two inputs is pretty easy stuff - the very hard part for me is building the data generator :)\n\nLet us know if you have success getting it to run."
        },
        {
          "id": 1320511,
          "postDate": "2021-05-24T06:24:47.267Z",
          "content": "<p>I think it's also a good choice to take apart, but it may be necessary to merge and identify useful information, and the generator mentioned above is also a problem</p>",
          "rawMarkdown": "I think it's also a good choice to take apart, but it may be necessary to merge and identify useful information, and the generator mentioned above is also a problem"
        },
        {
          "id": 1320663,
          "postDate": "2021-05-24T08:27:13.817Z",
          "content": "<p>For the generator, I simply create two np array (one for the signal A and another for the rest). I fill them with the correct signal. </p>\n<pre><code> def __getitem__(self, index):\n\n    # Create the batch\n    batch = self.df[index * self.batch_size:(index + 1) * self.batch_size]\n\n    # Create the two input\n    signals_1 = np.empty((len(batch), 3, 273, 256), dtype=np.float32)\n    signals_2 = np.empty((len(batch), 3, 273, 256), dtype=np.float32)\n\n    i = 0\n\n    for filename in batch.id:\n        # Get the path : directory/first_char/filename.npy\n        path = os.path.join(self.directory, filename[0], filename + \".npy\")\n        data = np.load(path)\n\n        # Separate each signal\n        signals_1[i, 0,] = data[0]\n        signals_1[i, 1,] = data[2]\n        signals_1[i, 2,] = data[4]\n\n        signals_2[i, 0,] = data[1]\n        signals_2[i, 1,] = data[3]\n        signals_2[i, 2,] = data[5]\n\n        i += 1\n\n    inp = [signals_1, signals_2]\n\n    if self.training:\n        return inp, batch.target.values\n    else:\n        return inp\n</code></pre>\n<p>And then, for the neural network, I just create two input layer : </p>\n<pre><code>inp_1 = Input((3, 273, 256))\nx = TimeDistributed(Bidirectional(LSTM(128, return_sequences=True)))(inp_1)\nx = Dropout(0.5)(x)\n\ninp_2 = Input((3, 273, 256))\ny = TimeDistributed(Bidirectional(LSTM(128, return_sequences=True)))(inp_2)\ny = Dropout(0.5)(y)\n</code></pre>\n<p>Finally, when I build the model, I precise the two inputs as :</p>\n<pre><code>model = Model(inputs=[inp_1, inp_2], outputs=output)\n</code></pre>\n<p>If you want, you can see the <a href=\"https://www.kaggle.com/rerere/signal-search-keras-two-input-lstm\" target=\"_blank\">full code here.</a></p>\n<p>I don't know if my way of doing is the right thing to do or not haha ^^ <br>\nDon't hesitate, if you have questions :)</p>",
          "rawMarkdown": "For the generator, I simply create two np array (one for the signal A and another for the rest). I fill them with the correct signal. \n\n```\n def __getitem__(self, index):\n\n    # Create the batch\n    batch = self.df[index * self.batch_size:(index + 1) * self.batch_size]\n        \n    # Create the two input\n    signals_1 = np.empty((len(batch), 3, 273, 256), dtype=np.float32)\n    signals_2 = np.empty((len(batch), 3, 273, 256), dtype=np.float32)\n        \n    i = 0\n        \n    for filename in batch.id:\n        # Get the path : directory/first_char/filename.npy\n        path = os.path.join(self.directory, filename[0], filename + \".npy\")\n        data = np.load(path)\n            \n        # Separate each signal\n        signals_1[i, 0,] = data[0]\n        signals_1[i, 1,] = data[2]\n        signals_1[i, 2,] = data[4]\n        \n        signals_2[i, 0,] = data[1]\n        signals_2[i, 1,] = data[3]\n        signals_2[i, 2,] = data[5]\n            \n        i += 1\n        \n    inp = [signals_1, signals_2]\n        \n    if self.training:\n        return inp, batch.target.values\n    else:\n        return inp\n```\n\nAnd then, for the neural network, I just create two input layer : \n\n```\ninp_1 = Input((3, 273, 256))\nx = TimeDistributed(Bidirectional(LSTM(128, return_sequences=True)))(inp_1)\nx = Dropout(0.5)(x)\n    \ninp_2 = Input((3, 273, 256))\ny = TimeDistributed(Bidirectional(LSTM(128, return_sequences=True)))(inp_2)\ny = Dropout(0.5)(y)\n```\n\nFinally, when I build the model, I precise the two inputs as :\n```\nmodel = Model(inputs=[inp_1, inp_2], outputs=output)\n```\n\nIf you want, you can see the [full code here.](https://www.kaggle.com/rerere/signal-search-keras-two-input-lstm)\n\nI don't know if my way of doing is the right thing to do or not haha ^^ \nDon't hesitate, if you have questions :)"
        },
        {
          "id": 1320687,
          "postDate": "2021-05-24T08:44:42.657Z",
          "content": "<p>Interesting. How was your LB or best AUC?</p>",
          "rawMarkdown": "Interesting. How was your LB or best AUC?"
        },
        {
          "id": 1320698,
          "postDate": "2021-05-24T08:51:28.843Z",
          "content": "<p>Very bad at the moment haha ^^<br>\nI yet don't use the transposition in my data. Also, I just update my kernel in order to use the AUC metric (at the moment, I used accuracy, but it's not really relevant here ^^' ) . My update should be done in the afternoon. I will update this message when it's done :)</p>\n<p>Edit : <br>\nSo, I trained my network on 10 epochs. For the simple LSTM (only 2 bidirectional LSTM stacked) and with the 6 signals, I got :</p>\n<ul>\n<li>loss: 0.1482 - auc: 0.9195 - val_loss: 0.2853 - val_auc: 0.7562 (* on the 10th epochs)<br>\nWith 0.72 on LB </li>\n</ul>\n<p>With the other approach, with the decomposition of the signals (2 inputs LSTM), I got :</p>\n<ul>\n<li>loss: 0.2059 - auc: 0.8913 - val_loss: 0.3435 - val_auc: 0.6740 (* on the 10th epochs)<br>\nWith 0.61 on LB</li>\n</ul>\n<p>Based on the loss, I think the second network with the 2 inputs has some difficulties to learn. <br>\nSo it's maybe not a good way to separate the A signals from the rest. I was thinking that maybe I should try different approach when I'm merging the output from the two LSTM layers ? </p>\n<p>Also, LSTM layers take much more times than GRU for training. I will maybe adjust the number of units on my LSTM or maybe change them. </p>",
          "rawMarkdown": "Very bad at the moment haha ^^\nI yet don't use the transposition in my data. Also, I just update my kernel in order to use the AUC metric (at the moment, I used accuracy, but it's not really relevant here ^^' ) . My update should be done in the afternoon. I will update this message when it's done :)\n\nEdit : \nSo, I trained my network on 10 epochs. For the simple LSTM (only 2 bidirectional LSTM stacked) and with the 6 signals, I got :\n- loss: 0.1482 - auc: 0.9195 - val_loss: 0.2853 - val_auc: 0.7562 (* on the 10th epochs)\nWith 0.72 on LB \n\nWith the other approach, with the decomposition of the signals (2 inputs LSTM), I got :\n- loss: 0.2059 - auc: 0.8913 - val_loss: 0.3435 - val_auc: 0.6740 (* on the 10th epochs)\nWith 0.61 on LB\n\nBased on the loss, I think the second network with the 2 inputs has some difficulties to learn. \nSo it's maybe not a good way to separate the A signals from the rest. I was thinking that maybe I should try different approach when I'm merging the output from the two LSTM layers ? \n\nAlso, LSTM layers take much more times than GRU for training. I will maybe adjust the number of units on my LSTM or maybe change them. ",
          "votes": 1
        },
        {
          "id": 1325251,
          "postDate": "2021-05-27T15:42:01.090Z",
          "content": "<p><a href=\"https://www.kaggle.com/rerere\" target=\"_blank\">Regis</a></p>\n<p>Forking your code on local machine this afternoon.  Have a couple of changes in mind - will let you know if any success with the two inputs approach.</p>",
          "rawMarkdown": "[Regis](https://www.kaggle.com/rerere)\n\nForking your code on local machine this afternoon.  Have a couple of changes in mind - will let you know if any success with the two inputs approach."
        },
        {
          "id": 1325465,
          "postDate": "2021-05-27T19:10:27.063Z",
          "content": "<p>Ok no problem, I hope your ideas will work. Keep me posted, I'm curious about your results :)</p>\n<p>Also, I tried an approach with three LSTM. I was thinking that maybe if the three models can separately analyze the good signal A with one of the other signals and merge all the result after, it could help them. With this approach, I got similar result as 2 LSTM. Unfortunately, I used all my GPU time available, so it will have to wait before I can improve them haha ^^</p>",
          "rawMarkdown": "Ok no problem, I hope your ideas will work. Keep me posted, I'm curious about your results :)\n\nAlso, I tried an approach with three LSTM. I was thinking that maybe if the three models can separately analyze the good signal A with one of the other signals and merge all the result after, it could help them. With this approach, I got similar result as 2 LSTM. Unfortunately, I used all my GPU time available, so it will have to wait before I can improve them haha ^^\n "
        }
      ]
    },
    {
      "id": 1355536,
      "postDate": "2021-06-18T10:58:03.333Z",
      "content": "<p>So a quick check - any breakthrough in using LSTM? My initial impression was it is the natural candidate but looks like no one is speaking about that outside this thread. Even here also the results are not very encouraging. </p>",
      "rawMarkdown": "So a quick check - any breakthrough in using LSTM? My initial impression was it is the natural candidate but looks like no one is speaking about that outside this thread. Even here also the results are not very encouraging. "
    },
    {
      "id": 1319321,
      "postDate": "2021-05-23T05:36:51.073Z",
      "content": "<p>I'm also going to study this direction, but I don't have time to study it at present, but it may change in a few days</p>",
      "rawMarkdown": "I'm also going to study this direction, but I don't have time to study it at present, but it may change in a few days"
    },
    {
      "id": 1306687,
      "postDate": "2021-05-14T02:32:36.377Z",
      "content": "<p>my poor transformer doesn't work (auc=0.5)…</p>",
      "rawMarkdown": "my poor transformer doesn't work (auc=0.5)...",
      "replies": [
        {
          "id": 1307753,
          "postDate": "2021-05-14T16:38:13.737Z",
          "content": "<p>Have you tried swapping out the RNN in my notebook with the <code>layers.MultiHeadAttention</code> (<a href=\"https://keras.io/api/layers/attention_layers/multi_head_attention/\" target=\"_blank\">docs</a>)?</p>",
          "rawMarkdown": "Have you tried swapping out the RNN in my notebook with the `layers.MultiHeadAttention` ([docs](https://keras.io/api/layers/attention_layers/multi_head_attention/))?"
        },
        {
          "id": 1309357,
          "postDate": "2021-05-15T21:10:28.483Z",
          "content": "<p>My poor transformer (0.6) emphasizes with yours… </p>",
          "rawMarkdown": "My poor transformer (0.6) emphasizes with yours... "
        },
        {
          "id": 1309583,
          "postDate": "2021-05-16T05:45:02.453Z",
          "content": "<p>Vit can`t go over 0.6 in my testing</p>",
          "rawMarkdown": "Vit can`t go over 0.6 in my testing"
        }
      ]
    },
    {
      "id": 1305012,
      "postDate": "2021-05-13T04:12:24.033Z",
      "content": "<p>Visual Transformer would be interesting as well.</p>",
      "rawMarkdown": "Visual Transformer would be interesting as well."
    }
  ],
  "comments": [
    {
      "id": 1306574,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-05-13T22:13:52.577000",
      "content": "<p>Thanks for starting this interesting thread and creating a starting nb! I'm really curious to see if RNN would make difference to this problem - however it is still not very clear to me how we are trying to exploit the temporal information ? I guess we speak about the \"intra-cadence\" relationships ? ie the sixtuple sequence ABACAD = 6 channel info   </p>\n<p>If that is the case then I would try first to combine them - or better not to split by input head - and then passing to RNN.  </p>\n<p>For points 2,3 need to make some experiments fist but I'm out of gpu quota :(</p>\n<p>PS: Still remember your famous GRU kernel in OpenVaccine comp. ;)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1307773,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2021-05-14T16:43:30.900000",
          "content": "<p>Hahaha at this rate i'll be known as the GRU guy ;)</p>\n<p>In my notebook, the sequence in this case is the 1st dimension of the 2d spectrogram. I'm running the RNN on each of the 6 channels separately, and afterwards I concatenate them.</p>\n<p>However I can't say if there's a \"temporal\" significance to the 1st dimension of the spectrogram; happy if someone can correct my approach here!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1320506,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-05-24T06:20:10.067000",
          "content": "<p>I think RNN time series combined with CNN model can also get better results, which can be used as a branch of ensemble. But at present, how to effectively use the 273 time dimension is a problem. I saw the information about GRU on the published notebook, but there is no in-depth study. I just used the code of that notebook and changed it to nfnet_ L0 model, got 0.96 +</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1324707,
          "author_name": "Baran Hashemi",
          "author_url": "",
          "post_date": "2021-05-27T07:09:42.810000",
          "content": "<p>You put nfnet_l0 before RNN?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1320219,
      "author_name": "Régis",
      "author_url": "",
      "post_date": "2021-05-23T21:48:16.953000",
      "content": "<p>Hello, </p>\n<p>For the first point, I'm currently trying another approach with two input in my model. I try to separate the A signals and the B, C and D. So, one input it will have the three A signal and for the other input the three others. I don't really know if it's possible to do something like that, I'm afraid I maybe lose some sense in the data. </p>\n<p>Currently, in my model, I used LSTM. There are more time consuming than GRU, but as the model is simple, I prefer to use them. However, I didn't take the time to use the same architecture as yours to truly see if there is a noticeable difference between GRU and LSTM. I will maybe try it in the next days and put an update on this post. </p>\n<p>At the moment, I more interest in the transposition you've made in your kernel. I need to maybe more understand the data ^^</p>\n<p>For the second point, in order to optimize your model, if think Dropout can be useful. For the learning rate, I personally use the ReduceLROnPlateau callback from the Keras library. It allows you to reduce the learning rate when, for example, the loss in the validation data doesn't decrease.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1320325,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-05-24T02:22:27.137000",
          "content": "<p>Building a tf.keras model with two inputs is pretty easy stuff - the very hard part for me is building the data generator :)</p>\n<p>Let us know if you have success getting it to run.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320511,
          "author_name": "zhangeng",
          "author_url": "",
          "post_date": "2021-05-24T06:24:47.267000",
          "content": "<p>I think it's also a good choice to take apart, but it may be necessary to merge and identify useful information, and the generator mentioned above is also a problem</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320663,
          "author_name": "Régis",
          "author_url": "",
          "post_date": "2021-05-24T08:27:13.817000",
          "content": "<p>For the generator, I simply create two np array (one for the signal A and another for the rest). I fill them with the correct signal. </p>\n<pre><code> def __getitem__(self, index):\n\n    # Create the batch\n    batch = self.df[index * self.batch_size:(index + 1) * self.batch_size]\n\n    # Create the two input\n    signals_1 = np.empty((len(batch), 3, 273, 256), dtype=np.float32)\n    signals_2 = np.empty((len(batch), 3, 273, 256), dtype=np.float32)\n\n    i = 0\n\n    for filename in batch.id:\n        # Get the path : directory/first_char/filename.npy\n        path = os.path.join(self.directory, filename[0], filename + \".npy\")\n        data = np.load(path)\n\n        # Separate each signal\n        signals_1[i, 0,] = data[0]\n        signals_1[i, 1,] = data[2]\n        signals_1[i, 2,] = data[4]\n\n        signals_2[i, 0,] = data[1]\n        signals_2[i, 1,] = data[3]\n        signals_2[i, 2,] = data[5]\n\n        i += 1\n\n    inp = [signals_1, signals_2]\n\n    if self.training:\n        return inp, batch.target.values\n    else:\n        return inp\n</code></pre>\n<p>And then, for the neural network, I just create two input layer : </p>\n<pre><code>inp_1 = Input((3, 273, 256))\nx = TimeDistributed(Bidirectional(LSTM(128, return_sequences=True)))(inp_1)\nx = Dropout(0.5)(x)\n\ninp_2 = Input((3, 273, 256))\ny = TimeDistributed(Bidirectional(LSTM(128, return_sequences=True)))(inp_2)\ny = Dropout(0.5)(y)\n</code></pre>\n<p>Finally, when I build the model, I precise the two inputs as :</p>\n<pre><code>model = Model(inputs=[inp_1, inp_2], outputs=output)\n</code></pre>\n<p>If you want, you can see the <a href=\"https://www.kaggle.com/rerere/signal-search-keras-two-input-lstm\" target=\"_blank\">full code here.</a></p>\n<p>I don't know if my way of doing is the right thing to do or not haha ^^ <br>\nDon't hesitate, if you have questions :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320687,
          "author_name": "Baran Hashemi",
          "author_url": "",
          "post_date": "2021-05-24T08:44:42.657000",
          "content": "<p>Interesting. How was your LB or best AUC?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1320698,
          "author_name": "Régis",
          "author_url": "",
          "post_date": "2021-05-24T08:51:28.843000",
          "content": "<p>Very bad at the moment haha ^^<br>\nI yet don't use the transposition in my data. Also, I just update my kernel in order to use the AUC metric (at the moment, I used accuracy, but it's not really relevant here ^^' ) . My update should be done in the afternoon. I will update this message when it's done :)</p>\n<p>Edit : <br>\nSo, I trained my network on 10 epochs. For the simple LSTM (only 2 bidirectional LSTM stacked) and with the 6 signals, I got :</p>\n<ul>\n<li>loss: 0.1482 - auc: 0.9195 - val_loss: 0.2853 - val_auc: 0.7562 (* on the 10th epochs)<br>\nWith 0.72 on LB </li>\n</ul>\n<p>With the other approach, with the decomposition of the signals (2 inputs LSTM), I got :</p>\n<ul>\n<li>loss: 0.2059 - auc: 0.8913 - val_loss: 0.3435 - val_auc: 0.6740 (* on the 10th epochs)<br>\nWith 0.61 on LB</li>\n</ul>\n<p>Based on the loss, I think the second network with the 2 inputs has some difficulties to learn. <br>\nSo it's maybe not a good way to separate the A signals from the rest. I was thinking that maybe I should try different approach when I'm merging the output from the two LSTM layers ? </p>\n<p>Also, LSTM layers take much more times than GRU for training. I will maybe adjust the number of units on my LSTM or maybe change them. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1325251,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2021-05-27T15:42:01.090000",
          "content": "<p><a href=\"https://www.kaggle.com/rerere\" target=\"_blank\">Regis</a></p>\n<p>Forking your code on local machine this afternoon.  Have a couple of changes in mind - will let you know if any success with the two inputs approach.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1325465,
          "author_name": "Régis",
          "author_url": "",
          "post_date": "2021-05-27T19:10:27.063000",
          "content": "<p>Ok no problem, I hope your ideas will work. Keep me posted, I'm curious about your results :)</p>\n<p>Also, I tried an approach with three LSTM. I was thinking that maybe if the three models can separately analyze the good signal A with one of the other signals and merge all the result after, it could help them. With this approach, I got similar result as 2 LSTM. Unfortunately, I used all my GPU time available, so it will have to wait before I can improve them haha ^^</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1355536,
      "author_name": "Saurav",
      "author_url": "",
      "post_date": "2021-06-18T10:58:03.333000",
      "content": "<p>So a quick check - any breakthrough in using LSTM? My initial impression was it is the natural candidate but looks like no one is speaking about that outside this thread. Even here also the results are not very encouraging. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1319321,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "2021-05-23T05:36:51.073000",
      "content": "<p>I'm also going to study this direction, but I don't have time to study it at present, but it may change in a few days</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1306687,
      "author_name": "patriot",
      "author_url": "",
      "post_date": "2021-05-14T02:32:36.377000",
      "content": "<p>my poor transformer doesn't work (auc=0.5)…</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1307753,
          "author_name": "xhlulu",
          "author_url": "",
          "post_date": "2021-05-14T16:38:13.737000",
          "content": "<p>Have you tried swapping out the RNN in my notebook with the <code>layers.MultiHeadAttention</code> (<a href=\"https://keras.io/api/layers/attention_layers/multi_head_attention/\" target=\"_blank\">docs</a>)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1309357,
          "author_name": "Hannes Öhler",
          "author_url": "",
          "post_date": "2021-05-15T21:10:28.483000",
          "content": "<p>My poor transformer (0.6) emphasizes with yours… </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1309583,
          "author_name": "Mithil Salunkhe",
          "author_url": "",
          "post_date": "2021-05-16T05:45:02.453000",
          "content": "<p>Vit can`t go over 0.6 in my testing</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1305012,
      "author_name": "abcd",
      "author_url": "",
      "post_date": "2021-05-13T04:12:24.033000",
      "content": "<p>Visual Transformer would be interesting as well.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1304761": "I just released a notebook that establishes a baseline for RNN - [Bidirectional GRU in Keras](https://www.kaggle.com/xhlulu/seti-bidirectional-gru-in-keras). \n\nAt time of writing this, the score sits at LB 0.85, which is lower than the top CNN models. Moreover, there's a few things that could be improved.\n\n1. How would you deal with the 6 positions? At the moment I'm apply a shared RNN layer on each position separately, then combining them using an average layer. However, would it make more sense to combine them before applying the RNN, or after? Is there some way to share the information? and would it make sense to concatenate instead of averaging?\n2. What are the best hyperparameters? How many layers & hidden units, what learning rate, and what optimizer would you use?\n3. Is GRU the best model for this task? Would LSTM, or some other sequential architecture (e.g. transformers) make more sense?\n4. How could we speed up the data loading so it doesn't slow down the actually forward/backward pass on the GPU?\n\nI'm happy to hear your thoughts on what is the best way to approach them, and I'm looking forward incorporating the feedback into the notebook.\n\nEDIT: I tried concatenating the pooled output of the RNNs instead of averaging. I also used the entire dataset instead of just a sampled subset of the negative results. With those two changes I got an improve LB 0.85 -> LB 0.9",
    "1306574": "Thanks for starting this interesting thread and creating a starting nb! I'm really curious to see if RNN would make difference to this problem - however it is still not very clear to me how we are trying to exploit the temporal information ? I guess we speak about the \"intra-cadence\" relationships ? ie the sixtuple sequence ABACAD = 6 channel info   \n\nIf that is the case then I would try first to combine them - or better not to split by input head - and then passing to RNN.  \n\nFor points 2,3 need to make some experiments fist but I'm out of gpu quota :(\n\nPS: Still remember your famous GRU kernel in OpenVaccine comp. ;)",
    "1320219": "Hello, \n\nFor the first point, I'm currently trying another approach with two input in my model. I try to separate the A signals and the B, C and D. So, one input it will have the three A signal and for the other input the three others. I don't really know if it's possible to do something like that, I'm afraid I maybe lose some sense in the data. \n\nCurrently, in my model, I used LSTM. There are more time consuming than GRU, but as the model is simple, I prefer to use them. However, I didn't take the time to use the same architecture as yours to truly see if there is a noticeable difference between GRU and LSTM. I will maybe try it in the next days and put an update on this post. \n\nAt the moment, I more interest in the transposition you've made in your kernel. I need to maybe more understand the data ^^\n\nFor the second point, in order to optimize your model, if think Dropout can be useful. For the learning rate, I personally use the ReduceLROnPlateau callback from the Keras library. It allows you to reduce the learning rate when, for example, the loss in the validation data doesn't decrease.\n",
    "1355536": "So a quick check - any breakthrough in using LSTM? My initial impression was it is the natural candidate but looks like no one is speaking about that outside this thread. Even here also the results are not very encouraging. ",
    "1319321": "I'm also going to study this direction, but I don't have time to study it at present, but it may change in a few days",
    "1306687": "my poor transformer doesn't work (auc=0.5)...",
    "1305012": "Visual Transformer would be interesting as well."
  }
}