{
  "id": 473592,
  "title": "Use of RNNs?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/473592",
  "author_name": "",
  "post_date": "2024-02-05T13:18:46.604687800Z",
  "votes": 14,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hey all!  Has anyone had success with RNNs on this data?  My intuition says that they should be an ok approach, but I am not that experienced with them and was wondering if anyone smarter than me has gotten performance on par with the fully convolutional models so graciously shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a>, and many others later. To date, I have tried the following:</p>\n<ul>\n<li><p>Network architectures such as <a href=\"https://arxiv.org/pdf/1412.5567.pdf\" target=\"_blank\">DeepSpeech</a> and <a href=\"https://arxiv.org/pdf/1512.02595.pdf\" target=\"_blank\">DeepSpeech2</a> used for speech recognition.  I could not get these to work well (CV 0.88, LB 0.57 with this <a href=\"https://www.kaggle.com/code/robbob62287/top-features-with-effnet-ds1-ensemble\" target=\"_blank\">notebook</a> I hacked together from the <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">EfficentNet</a>starter and trying different variations of CNNs feeding RNNs.  To my surprise, this shallow one had better training CV than ones that more closely resembled what is described in the papers.  Does anyone have any thoughts on why that may be? </p></li>\n<li><p>Changing out the convolutional layers for <a href=\"https://arxiv.org/pdf/1409.4842.pdf\" target=\"_blank\">Inception</a> inspired convolutional layers so that different filter sizes would be used since it is not known which filter size would be optimal because I don't know what frequencies would necessarily work well.  Even with the change to the convolutional layers,  I only achieved similar so-so results as the shared notebook.  </p></li>\n<li><p>I recently stumbled upon a network architecture named <a href=\"https://www.semanticscholar.org/reader/631caf3219766d1a23df89fa69b1bebbaa9d6959\" target=\"_blank\">ChronoNet</a>, and this took the Inception inspired convolutional layer and then also added <a href=\"https://arxiv.org/pdf/1608.06993.pdf\" target=\"_blank\">DenseNet</a> ideas of fully-connecting the RNN layers, and applied this to EEG data.  I thought that was a pretty interesting idea and made a <a href=\"https://www.kaggle.com/robbob62287/chrononet-for-eeg-data-starter\" target=\"_blank\">notebook </a>to test out this idea, but I am still not getting CV scores (~1.2 range) that would improve upon what I have already submitted.  </p></li>\n</ul>\n<p>I am attempting to replace the RNN layers with transformers, will post in the comments if this works.  Please let me know if anything sticks out as being done poorly on my end that could be contributing to this just not.</p>",
  "messages": [
    {
      "id": "2637041",
      "postDate": "02/05/2024 13:18:46",
      "content": "<p>Hey all!  Has anyone had success with RNNs on this data?  My intuition says that they should be an ok approach, but I am not that experienced with them and was wondering if anyone smarter than me has gotten performance on par with the fully convolutional models so graciously shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/yunsuxiaozi\" target=\"_blank\">@yunsuxiaozi</a>, and many others later. To date, I have tried the following:</p>\n<ul>\n<li><p>Network architectures such as <a href=\"https://arxiv.org/pdf/1412.5567.pdf\" target=\"_blank\">DeepSpeech</a> and <a href=\"https://arxiv.org/pdf/1512.02595.pdf\" target=\"_blank\">DeepSpeech2</a> used for speech recognition.  I could not get these to work well (CV 0.88, LB 0.57 with this <a href=\"https://www.kaggle.com/code/robbob62287/top-features-with-effnet-ds1-ensemble\" target=\"_blank\">notebook</a> I hacked together from the <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">EfficentNet</a>starter and trying different variations of CNNs feeding RNNs.  To my surprise, this shallow one had better training CV than ones that more closely resembled what is described in the papers.  Does anyone have any thoughts on why that may be? </p></li>\n<li><p>Changing out the convolutional layers for <a href=\"https://arxiv.org/pdf/1409.4842.pdf\" target=\"_blank\">Inception</a> inspired convolutional layers so that different filter sizes would be used since it is not known which filter size would be optimal because I don't know what frequencies would necessarily work well.  Even with the change to the convolutional layers,  I only achieved similar so-so results as the shared notebook.  </p></li>\n<li><p>I recently stumbled upon a network architecture named <a href=\"https://www.semanticscholar.org/reader/631caf3219766d1a23df89fa69b1bebbaa9d6959\" target=\"_blank\">ChronoNet</a>, and this took the Inception inspired convolutional layer and then also added <a href=\"https://arxiv.org/pdf/1608.06993.pdf\" target=\"_blank\">DenseNet</a> ideas of fully-connecting the RNN layers, and applied this to EEG data.  I thought that was a pretty interesting idea and made a <a href=\"https://www.kaggle.com/robbob62287/chrononet-for-eeg-data-starter\" target=\"_blank\">notebook </a>to test out this idea, but I am still not getting CV scores (~1.2 range) that would improve upon what I have already submitted.  </p></li>\n</ul>\n<p>I am attempting to replace the RNN layers with transformers, will post in the comments if this works.  Please let me know if anything sticks out as being done poorly on my end that could be contributing to this just not.</p>",
      "rawMarkdown": "Hey all!  Has anyone had success with RNNs on this data?  My intuition says that they should be an ok approach, but I am not that experienced with them and was wondering if anyone smarter than me has gotten performance on par with the fully convolutional models so graciously shared by @cdeotte, @yunsuxiaozi, and many others later. To date, I have tried the following:\n\n-  Network architectures such as [DeepSpeech](https://arxiv.org/pdf/1412.5567.pdf) and [DeepSpeech2](https://arxiv.org/pdf/1512.02595.pdf) used for speech recognition.  I could not get these to work well (CV 0.88, LB 0.57 with this [notebook](https://www.kaggle.com/code/robbob62287/top-features-with-effnet-ds1-ensemble) I hacked together from the [EfficentNet] (https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43)starter and trying different variations of CNNs feeding RNNs.  To my surprise, this shallow one had better training CV than ones that more closely resembled what is described in the papers.  Does anyone have any thoughts on why that may be? \n\n- Changing out the convolutional layers for [Inception](https://arxiv.org/pdf/1409.4842.pdf) inspired convolutional layers so that different filter sizes would be used since it is not known which filter size would be optimal because I don't know what frequencies would necessarily work well.  Even with the change to the convolutional layers,  I only achieved similar so-so results as the shared notebook.  \n\n- I recently stumbled upon a network architecture named [ChronoNet](https://www.semanticscholar.org/reader/631caf3219766d1a23df89fa69b1bebbaa9d6959), and this took the Inception inspired convolutional layer and then also added [DenseNet](https://arxiv.org/pdf/1608.06993.pdf) ideas of fully-connecting the RNN layers, and applied this to EEG data.  I thought that was a pretty interesting idea and made a [notebook ](https://www.kaggle.com/robbob62287/chrononet-for-eeg-data-starter)to test out this idea, but I am still not getting CV scores (~1.2 range) that would improve upon what I have already submitted.  \n\nI am attempting to replace the RNN layers with transformers, will post in the comments if this works.  Please let me know if anything sticks out as being done poorly on my end that could be contributing to this just not.",
      "votes": null
    },
    {
      "id": "2638188",
      "postDate": "02/06/2024 06:14:00",
      "content": "<p>I think time series analysis is a promising direction, but I think it should not be in the form of spectrograms but in the form of complete time series. I tried to make a Transformer-based time series framework where loss becomes NAN during runtime, which I am still exploring</p>",
      "rawMarkdown": "I think time series analysis is a promising direction, but I think it should not be in the form of spectrograms but in the form of complete time series. I tried to make a Transformer-based time series framework where loss becomes NAN during runtime, which I am still exploring",
      "votes": null
    },
    {
      "id": "2638246",
      "postDate": "02/06/2024 06:45:28",
      "content": "<p>I was hoping to be able to pull it off without spectrograms like the ChronoNet paper/notebook, but DS2 uses RNNs on spectrograms so it could work either way.  I was able to get a transformer-based network to work and also experienced the nan problem.  I found gradient clipping resolved it.  In TensorFlow you just set the \"clipvalue\" argument in the loss function definition to a number like this:</p>\n<p><code>tf.keras.optimizers.Adam(\n    learning_rate=0.001,\n    clipvalue=10,\n)</code></p>\n<p>Please share if you find anything!  I have kind of given up on this approach for now.</p>",
      "rawMarkdown": "I was hoping to be able to pull it off without spectrograms like the ChronoNet paper/notebook, but DS2 uses RNNs on spectrograms so it could work either way.  I was able to get a transformer-based network to work and also experienced the nan problem.  I found gradient clipping resolved it.  In TensorFlow you just set the \"clipvalue\" argument in the loss function definition to a number like this:\n\n`tf.keras.optimizers.Adam(\n    learning_rate=0.001,\n    clipvalue=10,\n)`\n\nPlease share if you find anything!  I have kind of given up on this approach for now.",
      "votes": null
    },
    {
      "id": "2638298",
      "postDate": "02/06/2024 08:04:43",
      "content": "<p>How well do your Transformer-based models work? My framework is based on pytorch but I will improve it based on your hints</p>",
      "rawMarkdown": "How well do your Transformer-based models work? My framework is based on pytorch but I will improve it based on your hints",
      "votes": null
    },
    {
      "id": "2640467",
      "postDate": "02/07/2024 00:31:34",
      "content": "<p>Not great. The KL-Div scores for the cross-validation data (CV) hovers around 1.0 and the area under the cover is around 0.75.  For reference, I have convolutional-only models that are &lt; 0.60 CV and &gt; 0.88 AUC on the CV sets and receive LB &lt; 0.42.  This has been achieved using just the signals and the spectrograms.</p>",
      "rawMarkdown": "Not great. The KL-Div scores for the cross-validation data (CV) hovers around 1.0 and the area under the cover is around 0.75.  For reference, I have convolutional-only models that are < 0.60 CV and > 0.88 AUC on the CV sets and receive LB < 0.42.  This has been achieved using just the signals and the spectrograms.",
      "votes": null
    },
    {
      "id": "2640741",
      "postDate": "02/07/2024 04:14:06",
      "content": "<p>i got LB 1.16</p>",
      "rawMarkdown": "i got LB 1.16",
      "votes": null
    },
    {
      "id": "2641460",
      "postDate": "02/07/2024 13:50:02",
      "content": "<p>I tried LSTM and got 1.09 LB</p>",
      "rawMarkdown": "I tried LSTM and got 1.09 LB",
      "votes": null
    },
    {
      "id": "2642047",
      "postDate": "02/07/2024 21:29:48",
      "content": "<p>Thanks for sharing, Rob. I think Chononet is such a good find. It may not yield he best results by itself, but could add a lot of diversity to some established ensembles.</p>",
      "rawMarkdown": "Thanks for sharing, Rob. I think Chononet is such a good find. It may not yield he best results by itself, but could add a lot of diversity to some established ensembles.",
      "votes": null
    },
    {
      "id": "2642105",
      "postDate": "02/07/2024 22:41:26",
      "content": "<p>For anyone wanting to try transformers to classify time series data here is the example I worked off of:</p>\n<p><a href=\"https://keras.io/examples/timeseries/timeseries_classification_transformer/\" target=\"_blank\">https://keras.io/examples/timeseries/timeseries_classification_transformer/</a></p>",
      "rawMarkdown": "For anyone wanting to try transformers to classify time series data here is the example I worked off of:\n\nhttps://keras.io/examples/timeseries/timeseries_classification_transformer/",
      "votes": null
    },
    {
      "id": "2642110",
      "postDate": "02/07/2024 22:59:57",
      "content": "<p>I publish some Transformer Keras TF starter code in Sleep comp <a href=\"https://www.kaggle.com/code/cdeotte/11th-place-gold-cv-835-public-lb-788\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "I publish some Transformer Keras TF starter code in Sleep comp [here][1]\n\n[1]: https://www.kaggle.com/code/cdeotte/11th-place-gold-cv-835-public-lb-788",
      "votes": null
    },
    {
      "id": "2642136",
      "postDate": "02/08/2024 00:06:44",
      "content": "<p>Thank you Chris!  I'm sure you get this all the time, but the notebook you provided is enlightening!</p>",
      "rawMarkdown": "Thank you Chris!  I'm sure you get this all the time, but the notebook you provided is enlightening!",
      "votes": null
    },
    {
      "id": "2642620",
      "postDate": "02/08/2024 09:42:30",
      "content": "<p>Hi! i've been trying to get RNN's to work and i got a weird LB-CV gap… Did you normalize the signal? if so, how?</p>",
      "rawMarkdown": "Hi! i've been trying to get RNN's to work and i got a weird LB-CV gap... Did you normalize the signal? if so, how?",
      "votes": null
    },
    {
      "id": "2652283",
      "postDate": "02/14/2024 17:00:42",
      "content": "<p>I don't think they'll work as good as 1D CNNs.</p>",
      "rawMarkdown": "I don't think they'll work as good as 1D CNNs.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2638188,
      "author_name": "chenboluo",
      "author_url": "",
      "post_date": "02/06/2024 06:14:00",
      "content": "<p>I think time series analysis is a promising direction, but I think it should not be in the form of spectrograms but in the form of complete time series. I tried to make a Transformer-based time series framework where loss becomes NAN during runtime, which I am still exploring</p>",
      "votes": null,
      "replies": [
        {
          "id": 2638246,
          "author_name": "robbob62287",
          "author_url": "",
          "post_date": "02/06/2024 06:45:28",
          "content": "<p>I was hoping to be able to pull it off without spectrograms like the ChronoNet paper/notebook, but DS2 uses RNNs on spectrograms so it could work either way.  I was able to get a transformer-based network to work and also experienced the nan problem.  I found gradient clipping resolved it.  In TensorFlow you just set the \"clipvalue\" argument in the loss function definition to a number like this:</p>\n<p><code>tf.keras.optimizers.Adam(\n    learning_rate=0.001,\n    clipvalue=10,\n)</code></p>\n<p>Please share if you find anything!  I have kind of given up on this approach for now.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2638298,
              "author_name": "chenboluo",
              "author_url": "",
              "post_date": "02/06/2024 08:04:43",
              "content": "<p>How well do your Transformer-based models work? My framework is based on pytorch but I will improve it based on your hints</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2640467,
                  "author_name": "robbob62287",
                  "author_url": "",
                  "post_date": "02/07/2024 00:31:34",
                  "content": "<p>Not great. The KL-Div scores for the cross-validation data (CV) hovers around 1.0 and the area under the cover is around 0.75.  For reference, I have convolutional-only models that are &lt; 0.60 CV and &gt; 0.88 AUC on the CV sets and receive LB &lt; 0.42.  This has been achieved using just the signals and the spectrograms.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2640741,
                      "author_name": "chenboluo",
                      "author_url": "",
                      "post_date": "02/07/2024 04:14:06",
                      "content": "<p>i got LB 1.16</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2641460,
                          "author_name": "varunmanojgupta",
                          "author_url": "",
                          "post_date": "02/07/2024 13:50:02",
                          "content": "<p>I tried LSTM and got 1.09 LB</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2642047,
      "author_name": "m000sey",
      "author_url": "",
      "post_date": "02/07/2024 21:29:48",
      "content": "<p>Thanks for sharing, Rob. I think Chononet is such a good find. It may not yield he best results by itself, but could add a lot of diversity to some established ensembles.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2642105,
      "author_name": "robbob62287",
      "author_url": "",
      "post_date": "02/07/2024 22:41:26",
      "content": "<p>For anyone wanting to try transformers to classify time series data here is the example I worked off of:</p>\n<p><a href=\"https://keras.io/examples/timeseries/timeseries_classification_transformer/\" target=\"_blank\">https://keras.io/examples/timeseries/timeseries_classification_transformer/</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2642110,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/07/2024 22:59:57",
          "content": "<p>I publish some Transformer Keras TF starter code in Sleep comp <a href=\"https://www.kaggle.com/code/cdeotte/11th-place-gold-cv-835-public-lb-788\" target=\"_blank\">here</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 2642136,
              "author_name": "robbob62287",
              "author_url": "",
              "post_date": "02/08/2024 00:06:44",
              "content": "<p>Thank you Chris!  I'm sure you get this all the time, but the notebook you provided is enlightening!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2642620,
      "author_name": "roysegalz",
      "author_url": "",
      "post_date": "02/08/2024 09:42:30",
      "content": "<p>Hi! i've been trying to get RNN's to work and i got a weird LB-CV gap… Did you normalize the signal? if so, how?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2652283,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "02/14/2024 17:00:42",
      "content": "<p>I don't think they'll work as good as 1D CNNs.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2637041": "Hey all!  Has anyone had success with RNNs on this data?  My intuition says that they should be an ok approach, but I am not that experienced with them and was wondering if anyone smarter than me has gotten performance on par with the fully convolutional models so graciously shared by @cdeotte, @yunsuxiaozi, and many others later. To date, I have tried the following:\n\n-  Network architectures such as [DeepSpeech](https://arxiv.org/pdf/1412.5567.pdf) and [DeepSpeech2](https://arxiv.org/pdf/1512.02595.pdf) used for speech recognition.  I could not get these to work well (CV 0.88, LB 0.57 with this [notebook](https://www.kaggle.com/code/robbob62287/top-features-with-effnet-ds1-ensemble) I hacked together from the [EfficentNet] (https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43)starter and trying different variations of CNNs feeding RNNs.  To my surprise, this shallow one had better training CV than ones that more closely resembled what is described in the papers.  Does anyone have any thoughts on why that may be? \n\n- Changing out the convolutional layers for [Inception](https://arxiv.org/pdf/1409.4842.pdf) inspired convolutional layers so that different filter sizes would be used since it is not known which filter size would be optimal because I don't know what frequencies would necessarily work well.  Even with the change to the convolutional layers,  I only achieved similar so-so results as the shared notebook.  \n\n- I recently stumbled upon a network architecture named [ChronoNet](https://www.semanticscholar.org/reader/631caf3219766d1a23df89fa69b1bebbaa9d6959), and this took the Inception inspired convolutional layer and then also added [DenseNet](https://arxiv.org/pdf/1608.06993.pdf) ideas of fully-connecting the RNN layers, and applied this to EEG data.  I thought that was a pretty interesting idea and made a [notebook ](https://www.kaggle.com/robbob62287/chrononet-for-eeg-data-starter)to test out this idea, but I am still not getting CV scores (~1.2 range) that would improve upon what I have already submitted.  \n\nI am attempting to replace the RNN layers with transformers, will post in the comments if this works.  Please let me know if anything sticks out as being done poorly on my end that could be contributing to this just not.",
    "2638188": "I think time series analysis is a promising direction, but I think it should not be in the form of spectrograms but in the form of complete time series. I tried to make a Transformer-based time series framework where loss becomes NAN during runtime, which I am still exploring",
    "2638246": "I was hoping to be able to pull it off without spectrograms like the ChronoNet paper/notebook, but DS2 uses RNNs on spectrograms so it could work either way.  I was able to get a transformer-based network to work and also experienced the nan problem.  I found gradient clipping resolved it.  In TensorFlow you just set the \"clipvalue\" argument in the loss function definition to a number like this:\n\n`tf.keras.optimizers.Adam(\n    learning_rate=0.001,\n    clipvalue=10,\n)`\n\nPlease share if you find anything!  I have kind of given up on this approach for now.",
    "2638298": "How well do your Transformer-based models work? My framework is based on pytorch but I will improve it based on your hints",
    "2640467": "Not great. The KL-Div scores for the cross-validation data (CV) hovers around 1.0 and the area under the cover is around 0.75.  For reference, I have convolutional-only models that are < 0.60 CV and > 0.88 AUC on the CV sets and receive LB < 0.42.  This has been achieved using just the signals and the spectrograms.",
    "2640741": "i got LB 1.16",
    "2641460": "I tried LSTM and got 1.09 LB",
    "2642047": "Thanks for sharing, Rob. I think Chononet is such a good find. It may not yield he best results by itself, but could add a lot of diversity to some established ensembles.",
    "2642105": "For anyone wanting to try transformers to classify time series data here is the example I worked off of:\n\nhttps://keras.io/examples/timeseries/timeseries_classification_transformer/",
    "2642110": "I publish some Transformer Keras TF starter code in Sleep comp [here][1]\n\n[1]: https://www.kaggle.com/code/cdeotte/11th-place-gold-cv-835-public-lb-788",
    "2642136": "Thank you Chris!  I'm sure you get this all the time, but the notebook you provided is enlightening!",
    "2642620": "Hi! i've been trying to get RNN's to work and i got a weird LB-CV gap... Did you normalize the signal? if so, how?",
    "2652283": "I don't think they'll work as good as 1D CNNs."
  },
  "source": "meta"
}