{
  "id": 478195,
  "title": "[CV 0.68 | LB 0.46] with Only Raw EEG Signals in PyTorch",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/478195",
  "author_name": "AbaoJiang",
  "post_date": "2024-02-19T15:29:45.691000",
  "votes": 66,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Hello everyone,</p>\n<p>After reading this forum, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/473821\" target=\"_blank\">Raw EEG CV vs LB</a>, I'm also interested in finding out how far raw EEG model can go without using spectrogram information. Through many experiments, I finally build a model architecture to get <strong>CV 0.67 and LB 0.46</strong>. Let's see the model architecture first,<br>\n<a href=\"https://postimg.cc/7bdRnCBZ\" target=\"_blank\"><img src=\"https://i.postimg.cc/MKp8xVVV/Screenshot-2024-02-19-at-1-11-40-PM.png\" alt=\"Screenshot-2024-02-19-at-1-11-40-PM.png\"></a></p>\n<p>This model architecture can be seen as an extension of Chris' version. I modify the original dilated convolution to <strong>dilated inception convolution</strong>. Concretely speaking, instead of using only one kernel size per layer in <code>WaveBlock</code>, this model tries to capture enriched temporal patterns by considering different kernel sizes (2, 3, 6, and 7 here) at the same time. Furthermore, to restrain model parameters from growing too much, the model <strong>narrows down the output channels</strong> of convolutions with different kernel sizes. Let <code>out_channels = 64</code>, each dilated convolution with a specific kernel size only outputs 16 channels, then the concatenation along channel dimension is applied to match <code>out_channels = 64</code>.</p>\n<p>Adding this auxiliary block helps boost my CV about 0.15 and LB about 0.06 (compared with Chris'). If you're interested in training your own models and submitting in pure PyTorch way, you can refer to these two notebooks,</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-training\" target=\"_blank\">[LB 0.46] DilatedInception WaveNet - Training</a></li>\n<li><a href=\"https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-inference?scriptVersionId=163448688\" target=\"_blank\">[LB 0.46] DilatedInception WaveNet - Inference</a></li>\n</ul>\n<p>Thanks for reading!</p>",
  "messages": [
    {
      "id": 2659070,
      "postDate": "2024-02-19T15:29:45.690Z",
      "content": "<p>Hello everyone,</p>\n<p>After reading this forum, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/473821\" target=\"_blank\">Raw EEG CV vs LB</a>, I'm also interested in finding out how far raw EEG model can go without using spectrogram information. Through many experiments, I finally build a model architecture to get <strong>CV 0.67 and LB 0.46</strong>. Let's see the model architecture first,<br>\n<a href=\"https://postimg.cc/7bdRnCBZ\" target=\"_blank\"><img src=\"https://i.postimg.cc/MKp8xVVV/Screenshot-2024-02-19-at-1-11-40-PM.png\" alt=\"Screenshot-2024-02-19-at-1-11-40-PM.png\"></a></p>\n<p>This model architecture can be seen as an extension of Chris' version. I modify the original dilated convolution to <strong>dilated inception convolution</strong>. Concretely speaking, instead of using only one kernel size per layer in <code>WaveBlock</code>, this model tries to capture enriched temporal patterns by considering different kernel sizes (2, 3, 6, and 7 here) at the same time. Furthermore, to restrain model parameters from growing too much, the model <strong>narrows down the output channels</strong> of convolutions with different kernel sizes. Let <code>out_channels = 64</code>, each dilated convolution with a specific kernel size only outputs 16 channels, then the concatenation along channel dimension is applied to match <code>out_channels = 64</code>.</p>\n<p>Adding this auxiliary block helps boost my CV about 0.15 and LB about 0.06 (compared with Chris'). If you're interested in training your own models and submitting in pure PyTorch way, you can refer to these two notebooks,</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-training\" target=\"_blank\">[LB 0.46] DilatedInception WaveNet - Training</a></li>\n<li><a href=\"https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-inference?scriptVersionId=163448688\" target=\"_blank\">[LB 0.46] DilatedInception WaveNet - Inference</a></li>\n</ul>\n<p>Thanks for reading!</p>",
      "rawMarkdown": "Hello everyone,\n\nAfter reading this forum, [Raw EEG CV vs LB](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/473821), I'm also interested in finding out how far raw EEG model can go without using spectrogram information. Through many experiments, I finally build a model architecture to get **CV 0.67 and LB 0.46**. Let's see the model architecture first,\n[![Screenshot-2024-02-19-at-1-11-40-PM.png](https://i.postimg.cc/MKp8xVVV/Screenshot-2024-02-19-at-1-11-40-PM.png)](https://postimg.cc/7bdRnCBZ)\n\nThis model architecture can be seen as an extension of Chris' version. I modify the original dilated convolution to **dilated inception convolution**. Concretely speaking, instead of using only one kernel size per layer in `WaveBlock`, this model tries to capture enriched temporal patterns by considering different kernel sizes (2, 3, 6, and 7 here) at the same time. Furthermore, to restrain model parameters from growing too much, the model **narrows down the output channels** of convolutions with different kernel sizes. Let `out_channels = 64`, each dilated convolution with a specific kernel size only outputs 16 channels, then the concatenation along channel dimension is applied to match `out_channels = 64`.\n\nAdding this auxiliary block helps boost my CV about 0.15 and LB about 0.06 (compared with Chris'). If you're interested in training your own models and submitting in pure PyTorch way, you can refer to these two notebooks,\n* [[LB 0.46] DilatedInception WaveNet - Training](https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-training)\n* [[LB 0.46] DilatedInception WaveNet - Inference](https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-inference?scriptVersionId=163448688)\n\nThanks for reading!",
      "votes": 66
    },
    {
      "id": 2661599,
      "postDate": "2024-02-21T11:18:58.060Z",
      "content": "<p>Good work! with some modifications I was able to get <strong>CV 0.6 and LB 0.38</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3350264%2Fad9e55523487c56ceecae0d3a719163f%2FScreenshot%20from%202024-02-21%2016-17-24.png?generation=1708514299642251&amp;alt=media\"></p>",
      "rawMarkdown": "Good work! with some modifications I was able to get **CV 0.6 and LB 0.38**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3350264%2Fad9e55523487c56ceecae0d3a719163f%2FScreenshot%20from%202024-02-21%2016-17-24.png?generation=1708514299642251&alt=media)",
      "votes": 9,
      "replies": [
        {
          "id": 2661834,
          "postDate": "2024-02-21T14:33:50.180Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/muhammad4hmed\" target=\"_blank\">@muhammad4hmed</a>,</p>\n<p>Thanks very much for your appreciation! <br>\nIt's so glad to hear that. Your CV and LB scores look pretty good. Keep fighting!!</p>",
          "rawMarkdown": "Hi @muhammad4hmed,\n\nThanks very much for your appreciation! \nIt's so glad to hear that. Your CV and LB scores look pretty good. Keep fighting!!",
          "votes": 1
        },
        {
          "id": 2662566,
          "postDate": "2024-02-22T02:02:15.137Z",
          "content": "<p><a href=\"https://www.kaggle.com/muhammad4hmed\" target=\"_blank\">@muhammad4hmed</a>  Could you please specify which aspects are primarily changed?</p>",
          "rawMarkdown": "@muhammad4hmed  Could you please specify which aspects are primarily changed?",
          "votes": 1
        },
        {
          "id": 2664144,
          "postDate": "2024-02-22T20:24:04.763Z",
          "content": "<p><a href=\"https://www.kaggle.com/muhammad4hmed\" target=\"_blank\">@muhammad4hmed</a> Reducing from 0.46 to 0.38 is amazing! I imagine you had to make several modifications</p>",
          "rawMarkdown": "@muhammad4hmed Reducing from 0.46 to 0.38 is amazing! I imagine you had to make several modifications"
        }
      ]
    },
    {
      "id": 2659559,
      "postDate": "2024-02-20T01:01:07.753Z",
      "content": "<p>Thank you. I even got below 0.5 by using lightgbm with massive feature engineering for raw EEG data. You just use basic features.</p>\n<pre><code>    x_tmp = np.zeros((, self.n_feats), dtype=)\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n</code></pre>",
      "rawMarkdown": "Thank you. I even got below 0.5 by using lightgbm with massive feature engineering for raw EEG data. You just use basic features.\n        \n        x_tmp = np.zeros((EEG_PTS, self.n_feats), dtype=\"float32\")\n        x_tmp[:, 0] = x[:, self.FEAT2CODE[\"Fp1\"]] - x[:, self.FEAT2CODE[\"T3\"]]\n        x_tmp[:, 1] = x[:, self.FEAT2CODE[\"T3\"]] - x[:, self.FEAT2CODE[\"O1\"]]\n        x_tmp[:, 2] = x[:, self.FEAT2CODE[\"Fp1\"]] - x[:, self.FEAT2CODE[\"C3\"]]\n        x_tmp[:, 3] = x[:, self.FEAT2CODE[\"C3\"]] - x[:, self.FEAT2CODE[\"O1\"]]\n        x_tmp[:, 4] = x[:, self.FEAT2CODE[\"Fp2\"]] - x[:, self.FEAT2CODE[\"C4\"]]\n        x_tmp[:, 5] = x[:, self.FEAT2CODE[\"C4\"]] - x[:, self.FEAT2CODE[\"O2\"]]\n        x_tmp[:, 6] = x[:, self.FEAT2CODE[\"Fp2\"]] - x[:, self.FEAT2CODE[\"T4\"]]\n        x_tmp[:, 7] = x[:, self.FEAT2CODE[\"T4\"]] - x[:, self.FEAT2CODE[\"O2\"]]\n\n",
      "votes": 8,
      "replies": [
        {
          "id": 2660088,
          "postDate": "2024-02-20T10:35:52.113Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>,</p>\n<p>Thanks for the comment.<br>\nThe credits of generating these features should be given to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. He's the one kindly sharing these formula out!</p>",
          "rawMarkdown": "Hi @sweetyheehee,\n\nThanks for the comment.\nThe credits of generating these features should be given to @cdeotte. He's the one kindly sharing these formula out!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2702519,
      "postDate": "2024-03-17T17:09:36.083Z",
      "content": "<p>thanks for sharing. I got LB 0.35 without any change of model architecture.<br>\nUPDATE: LB0.34</p>",
      "rawMarkdown": "thanks for sharing. I got LB 0.35 without any change of model architecture.\nUPDATE: LB0.34",
      "votes": 1,
      "replies": [
        {
          "id": 2703034,
          "postDate": "2024-03-18T00:57:57.470Z",
          "content": "<p>Amazing scores, may I ask what improvements you have made?</p>",
          "rawMarkdown": "Amazing scores, may I ask what improvements you have made?"
        },
        {
          "id": 2704534,
          "postDate": "2024-03-18T19:26:45.327Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>,</p>\n<p>That sounds amazing. Really glad to hear that, keep going!</p>",
          "rawMarkdown": "Hi @clearwaterkzk,\n\nThat sounds amazing. Really glad to hear that, keep going!",
          "replies": [
            {
              "id": 2705660,
              "postDate": "2024-03-19T12:58:02.510Z",
              "content": "<p>About kernel, how did you find <code>kernel_size = [2, 3, 6, 7]</code> ?</p>",
              "rawMarkdown": "About kernel, how did you find `kernel_size = [2, 3, 6, 7]` ?"
            },
            {
              "id": 2742916,
              "postDate": "2024-04-09T06:21:41.167Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>.<br>\nThanks for sharing nice model.<br>\nI got  0.36 in public LB, 0.345 in CV(vote&gt;=10)<br>\nThe change is following.</p>\n<ul>\n<li>2stage learning, <code>20 epoch</code> in each stage.</li>\n<li><code>lr=1e-5</code> in 2nd stage.</li>\n<li><code>kernel = [3, 6, 7, 9]</code> got best CV(vote). just changing kernel results in better score. I got the idea from <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468684\" target=\"_blank\">chris' discussion.</a></li>\n</ul>",
              "rawMarkdown": "Hi @abaojiang.\nThanks for sharing nice model.\nI got ~~0.34~~ 0.36 in public LB, 0.345 in CV(vote>=10)\nThe change is following.\n- 2stage learning, `20 epoch` in each stage.\n- `lr=1e-5` in 2nd stage.\n- `kernel = [3, 6, 7, 9]` got best CV(vote). just changing kernel results in better score. I got the idea from [chris' discussion.](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468684)\n",
              "votes": 1
            },
            {
              "id": 2743698,
              "postDate": "2024-04-09T15:22:57.687Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>,</p>\n<p>It's great to hear that! Tbh, I start my sprint about two weeks ago and didn't observe significant performance boost with this 1D model. So, I switch to 2D model with spectrogram. I'll go back and try again this week. Thanks for sharing your result!<br>\nBtw, may I ask how your 1D model perform on private LB?</p>",
              "rawMarkdown": "Hi @clearwaterkzk,\n\nIt's great to hear that! Tbh, I start my sprint about two weeks ago and didn't observe significant performance boost with this 1D model. So, I switch to 2D model with spectrogram. I'll go back and try again this week. Thanks for sharing your result!\nBtw, may I ask how your 1D model perform on private LB?"
            },
            {
              "id": 2743747,
              "postDate": "2024-04-09T15:48:59.313Z",
              "content": "<p>Here is the details.<br>\nThe training took 360 minutes in 1st stage and 120 minutes in 2nd stage with my GPU(A10G)</p>\n<ul>\n<li><p><strong>best CV(vote&gt;=10) 0.345, private 0.46, public 0.36</strong><br>\nkernel=[3, 6, 7, 9], mixup alpha=2.0, 20 epoch, flip augmentation(time reversal, p=0.5) in train and test time, train_batch=32, validation with vote&gt;=10 in eeg-unique-df, downsample=5<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fc9cbfb6835e684b3c00bd0839b214105%2Fwavenet_bestcv.png?generation=1712676777340886&amp;alt=media\"></p></li>\n<li><p><strong>all data(no val score), private 0.44, public 0.34</strong><br>\nkernel=[3, 5, 7, 9], mixup alpha=2.0, 20 epoch, flip augmentation(time reversal, p=0.5) in train and test time, train_batch=32, train all data, seed ensemble, downsample=5</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F3f6e56ab15cbcb9b41d13dea11708529%2Fwavenet_bestlb.png?generation=1712676748677925&amp;alt=media\" alt=\"img1\"></p>\n<p>This is flip augmentation code.<br>\nI also tried scaling and jittering augmentation but got worse result.</p>\n<pre><code> ():\n     np.flip(X, axis=)\n\n ():\n    scalingFactor = np.random.normal(loc=, scale=sigma, size=(,X.shape[])) \n    myNoise = np.matmul(np.ones((X.shape[],)), scalingFactor)\n     X*myNoise\n\n ():\n    myNoise = np.random.normal(loc=, scale=sigma, size=X.shape)\n     X+myNoise\n</code></pre>\n<p>If you want my full code, I will be happy to share !!</p>",
              "rawMarkdown": "Here is the details.\nThe training took 360 minutes in 1st stage and 120 minutes in 2nd stage with my GPU(A10G)\n\n- **best CV(vote>=10) 0.345, private 0.46, public 0.36**\nkernel=[3, 6, 7, 9], mixup alpha=2.0, 20 epoch, flip augmentation(time reversal, p=0.5) in train and test time, train_batch=32, validation with vote>=10 in eeg-unique-df, downsample=5\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fc9cbfb6835e684b3c00bd0839b214105%2Fwavenet_bestcv.png?generation=1712676777340886&alt=media)\n\n\n- **all data(no val score), private 0.44, public 0.34**\nkernel=[3, 5, 7, 9], mixup alpha=2.0, 20 epoch, flip augmentation(time reversal, p=0.5) in train and test time, train_batch=32, train all data, seed ensemble, downsample=5\n\n\n![img1](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F3f6e56ab15cbcb9b41d13dea11708529%2Fwavenet_bestlb.png?generation=1712676748677925&alt=media)\n\nThis is flip augmentation code.\nI also tried scaling and jittering augmentation but got worse result.\n```python\ndef DA_Flip(X):\n    return np.flip(X, axis=0)\n\ndef DA_Scaling(X, sigma=0.1):\n    scalingFactor = np.random.normal(loc=1.0, scale=sigma, size=(1,X.shape[1])) # shape=(1,3)\n    myNoise = np.matmul(np.ones((X.shape[0],1)), scalingFactor)\n    return X*myNoise\n\ndef DA_Jitter(X, sigma=0.05):\n    myNoise = np.random.normal(loc=0, scale=sigma, size=X.shape)\n    return X+myNoise\n```\n\nIf you want my full code, I will be happy to share !!",
              "votes": 1
            },
            {
              "id": 2744596,
              "postDate": "2024-04-10T02:45:41.607Z",
              "content": "<p>Here is my latest 1D model submission during the last week,<br>\n<a href=\"https://postimg.cc/R3HrP0cX\" target=\"_blank\"><img src=\"https://i.postimg.cc/5tPx2Yjd/Screenshot-2024-04-10-at-10-37-37-AM.png\" alt=\"Screenshot-2024-04-10-at-10-37-37-AM.png\"></a></p>\n<p>The difference with my original released version in the post is that I downsample the sequence along time dimension in the second and third <code>WaveBlock</code>, which accelerates the training speed without compromising too much performance.<br>\nSome more observations from my experiments,</p>\n<ol>\n<li>Augmentations didn't work for my 1D model (<em>e.g.,</em> hflip, vflip, gaussian noise, mixup, time and frequency mask).</li>\n<li>The single 1D model performs not that good in private LB but does contribute to ensembling model due to diversity.</li>\n</ol>\n<p>I think the second point coincide with yours. Thanks for sharing your experiments!</p>",
              "rawMarkdown": "Here is my latest 1D model submission during the last week,\n[![Screenshot-2024-04-10-at-10-37-37-AM.png](https://i.postimg.cc/5tPx2Yjd/Screenshot-2024-04-10-at-10-37-37-AM.png)](https://postimg.cc/R3HrP0cX)\n\nThe difference with my original released version in the post is that I downsample the sequence along time dimension in the second and third `WaveBlock`, which accelerates the training speed without compromising too much performance.\nSome more observations from my experiments,\n1. Augmentations didn't work for my 1D model (*e.g.,* hflip, vflip, gaussian noise, mixup, time and frequency mask).\n2. The single 1D model performs not that good in private LB but does contribute to ensembling model due to diversity.\n\nI think the second point coincide with yours. Thanks for sharing your experiments!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2659685,
      "postDate": "2024-02-20T04:48:26.983Z",
      "content": "<p>hello - Do we get .680 by just running this training script ? I remember running it for fold 1 and result was not comparable ?</p>",
      "rawMarkdown": "hello - Do we get .680 by just running this training script ? I remember running it for fold 1 and result was not comparable ?",
      "votes": 1,
      "replies": [
        {
          "id": 2659848,
          "postDate": "2024-02-20T07:11:13.723Z",
          "content": "<pre><code>  Load model  /kaggle//0217---/model-last_seed0_fold0.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold1.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold2.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold3.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold4.pth\n  Sum of row   y_preds .\n</code></pre>\n<p>He loaded all.</p>",
          "rawMarkdown": "```python\n  Load model from /kaggle/input/0217-15-11-37/model-last_seed0_fold0.pth\n  Load model from /kaggle/input/0217-15-11-37/model-last_seed0_fold1.pth\n  Load model from /kaggle/input/0217-15-11-37/model-last_seed0_fold2.pth\n  Load model from /kaggle/input/0217-15-11-37/model-last_seed0_fold3.pth\n  Load model from /kaggle/input/0217-15-11-37/model-last_seed0_fold4.pth\n  Sum of row 0 in y_preds 1.0.\n```\n\nHe loaded all.",
          "votes": 1,
          "replies": [
            {
              "id": 2659894,
              "postDate": "2024-02-20T07:45:42.083Z",
              "content": "<p>My question was for the CV .. it says 5 fold oof CV is .680 .. But for fold1 CV was around .900+ .. thats why I was thinking if there is some tweak required from baseline</p>",
              "rawMarkdown": "My question was for the CV .. it says 5 fold oof CV is .680 .. But for fold1 CV was around .900+ .. thats why I was thinking if there is some tweak required from baseline\n",
              "votes": 1
            },
            {
              "id": 2660101,
              "postDate": "2024-02-20T10:44:20.283Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a>,</p>\n<p>The setting is exactly what I used to obtain the CV score. If you want to train 5-fold models, you need to set <code>one_fold_only=False</code>, which is just a debug option for convenience.<br>\nSome observations to note are,</p>\n<ol>\n<li>Different seeds can make CV scores different, which is within the range of 0.02 when I take 3 seeds.</li>\n<li>Different folds have quite different scores (<em>e.g.,</em> fold0 is much worse than fold4).</li>\n</ol>\n<p>For your question, did you finish the 5-epoch training? 0.900+ is observed exactly on epoch 0. I just add a seeding functionality, can you try version2 again? If there's any problem, please let me know. Thanks!</p>",
              "rawMarkdown": "Hi @phoenix9032,\n\nThe setting is exactly what I used to obtain the CV score. If you want to train 5-fold models, you need to set `one_fold_only=False`, which is just a debug option for convenience.\nSome observations to note are,\n1. Different seeds can make CV scores different, which is within the range of 0.02 when I take 3 seeds.\n2. Different folds have quite different scores (*e.g.,* fold0 is much worse than fold4).\n\nFor your question, did you finish the 5-epoch training? 0.900+ is observed exactly on epoch 0. I just add a seeding functionality, can you try version2 again? If there's any problem, please let me know. Thanks!",
              "votes": 1
            },
            {
              "id": 2661451,
              "postDate": "2024-02-21T08:58:23.890Z",
              "content": "<p>add seed and retrain still get diff score</p>",
              "rawMarkdown": "add seed and retrain still get diff score",
              "votes": 1
            },
            {
              "id": 2661831,
              "postDate": "2024-02-21T14:30:06.130Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a>,</p>\n<p>May I ask what CV score you obtain by rerunning version2? <br>\nAs explained above, score can differ fold-by-fold and also seed-by-seed (<em>i.e.,</em> the randomness matters). My CV score 0.68 is reported for only one seed, which is also randomly picked (<em>i.e.,</em> I don't fix seed here). For experiments with more seeds, I get 0.67 or even 0.69 sometimes. </p>",
              "rawMarkdown": "Hi @tonymarkchris,\n\nMay I ask what CV score you obtain by rerunning version2? \nAs explained above, score can differ fold-by-fold and also seed-by-seed (*i.e.,* the randomness matters). My CV score 0.68 is reported for only one seed, which is also randomly picked (*i.e.,* I don't fix seed here). For experiments with more seeds, I get 0.67 or even 0.69 sometimes. "
            },
            {
              "id": 2661862,
              "postDate": "2024-02-21T14:50:13.833Z",
              "content": "<p>0.6762 ~ 0.69+</p>",
              "rawMarkdown": "0.6762 ~ 0.69+"
            },
            {
              "id": 2661904,
              "postDate": "2024-02-21T15:21:23.297Z",
              "content": "<p>So, it sounds reasonable. For the previous post of mine, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/476899\" target=\"_blank\">Reproducibility of Chris' WaveNet Baseline in torch and Randomness Study</a>, I do simple randomness study and show CV scores along two dimensions, fold-dim and seed-dim. As can be observed, the score can fluctuate within a wide range. Hence, randomness factor should be taken into consideration when trying to improve CV scores. For me, cosine decay with only picking the last epoch helps stabilize my CV scores a little bit.</p>",
              "rawMarkdown": "So, it sounds reasonable. For the previous post of mine, [Reproducibility of Chris' WaveNet Baseline in torch and Randomness Study](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/476899), I do simple randomness study and show CV scores along two dimensions, fold-dim and seed-dim. As can be observed, the score can fluctuate within a wide range. Hence, randomness factor should be taken into consideration when trying to improve CV scores. For me, cosine decay with only picking the last epoch helps stabilize my CV scores a little bit."
            },
            {
              "id": 2665885,
              "postDate": "2024-02-23T23:05:25.753Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> ，how to avoid diff score even when retrain?</p>",
              "rawMarkdown": "Hi @tonymarkchris ，how to avoid diff score even when retrain?"
            }
          ]
        }
      ]
    },
    {
      "id": 2671595,
      "postDate": "2024-02-27T15:44:52.040Z",
      "content": "<p>Nice work abaojiang, I see wavenet as a great way to give diversity to the ensembles used in the competition, I will use your starter model, keep sharing the good work</p>",
      "rawMarkdown": "Nice work abaojiang, I see wavenet as a great way to give diversity to the ensembles used in the competition, I will use your starter model, keep sharing the good work",
      "votes": 2,
      "replies": [
        {
          "id": 2671901,
          "postDate": "2024-02-27T18:57:54.337Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>,</p>\n<p>Thanks for your appreciation, I'm glad to hear that it works for you!</p>",
          "rawMarkdown": "Hi @rafaelzimmermann1,\n\nThanks for your appreciation, I'm glad to hear that it works for you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 2664143,
      "postDate": "2024-02-22T20:23:10.623Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2661599,
      "author_name": "Muhammad Ahmed",
      "author_url": "",
      "post_date": "2024-02-21T11:18:58.060000",
      "content": "<p>Good work! with some modifications I was able to get <strong>CV 0.6 and LB 0.38</strong><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3350264%2Fad9e55523487c56ceecae0d3a719163f%2FScreenshot%20from%202024-02-21%2016-17-24.png?generation=1708514299642251&amp;alt=media\"></p>",
      "votes": 9,
      "replies": [
        {
          "id": 2661834,
          "author_name": "AbaoJiang",
          "author_url": "",
          "post_date": "2024-02-21T14:33:50.180000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/muhammad4hmed\" target=\"_blank\">@muhammad4hmed</a>,</p>\n<p>Thanks very much for your appreciation! <br>\nIt's so glad to hear that. Your CV and LB scores look pretty good. Keep fighting!!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2662566,
          "author_name": "lhwcv",
          "author_url": "",
          "post_date": "2024-02-22T02:02:15.137000",
          "content": "<p><a href=\"https://www.kaggle.com/muhammad4hmed\" target=\"_blank\">@muhammad4hmed</a>  Could you please specify which aspects are primarily changed?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2664144,
          "author_name": "Yan Teixeira",
          "author_url": "",
          "post_date": "2024-02-22T20:24:04.763000",
          "content": "<p><a href=\"https://www.kaggle.com/muhammad4hmed\" target=\"_blank\">@muhammad4hmed</a> Reducing from 0.46 to 0.38 is amazing! I imagine you had to make several modifications</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2659559,
      "author_name": "Timmy Juicehouse",
      "author_url": "",
      "post_date": "2024-02-20T01:01:07.753000",
      "content": "<p>Thank you. I even got below 0.5 by using lightgbm with massive feature engineering for raw EEG data. You just use basic features.</p>\n<pre><code>    x_tmp = np.zeros((, self.n_feats), dtype=)\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n    x_tmp[:, ] = x[:, self.[]] - x[:, self.[]]\n</code></pre>",
      "votes": 8,
      "replies": [
        {
          "id": 2660088,
          "author_name": "AbaoJiang",
          "author_url": "",
          "post_date": "2024-02-20T10:35:52.113000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sweetyheehee\" target=\"_blank\">@sweetyheehee</a>,</p>\n<p>Thanks for the comment.<br>\nThe credits of generating these features should be given to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. He's the one kindly sharing these formula out!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2702519,
      "author_name": "Aurora_blue",
      "author_url": "",
      "post_date": "2024-03-17T17:09:36.083000",
      "content": "<p>thanks for sharing. I got LB 0.35 without any change of model architecture.<br>\nUPDATE: LB0.34</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2703034,
          "author_name": "Wisp Vale",
          "author_url": "",
          "post_date": "2024-03-18T00:57:57.470000",
          "content": "<p>Amazing scores, may I ask what improvements you have made?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2704534,
          "author_name": "AbaoJiang",
          "author_url": "",
          "post_date": "2024-03-18T19:26:45.327000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>,</p>\n<p>That sounds amazing. Really glad to hear that, keep going!</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2705660,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-03-19T12:58:02.510000",
              "content": "<p>About kernel, how did you find <code>kernel_size = [2, 3, 6, 7]</code> ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2742916,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-09T06:21:41.167000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/abaojiang\" target=\"_blank\">@abaojiang</a>.<br>\nThanks for sharing nice model.<br>\nI got  0.36 in public LB, 0.345 in CV(vote&gt;=10)<br>\nThe change is following.</p>\n<ul>\n<li>2stage learning, <code>20 epoch</code> in each stage.</li>\n<li><code>lr=1e-5</code> in 2nd stage.</li>\n<li><code>kernel = [3, 6, 7, 9]</code> got best CV(vote). just changing kernel results in better score. I got the idea from <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/468684\" target=\"_blank\">chris' discussion.</a></li>\n</ul>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2743698,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2024-04-09T15:22:57.687000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/clearwaterkzk\" target=\"_blank\">@clearwaterkzk</a>,</p>\n<p>It's great to hear that! Tbh, I start my sprint about two weeks ago and didn't observe significant performance boost with this 1D model. So, I switch to 2D model with spectrogram. I'll go back and try again this week. Thanks for sharing your result!<br>\nBtw, may I ask how your 1D model perform on private LB?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2743747,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-09T15:48:59.313000",
              "content": "<p>Here is the details.<br>\nThe training took 360 minutes in 1st stage and 120 minutes in 2nd stage with my GPU(A10G)</p>\n<ul>\n<li><p><strong>best CV(vote&gt;=10) 0.345, private 0.46, public 0.36</strong><br>\nkernel=[3, 6, 7, 9], mixup alpha=2.0, 20 epoch, flip augmentation(time reversal, p=0.5) in train and test time, train_batch=32, validation with vote&gt;=10 in eeg-unique-df, downsample=5<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2Fc9cbfb6835e684b3c00bd0839b214105%2Fwavenet_bestcv.png?generation=1712676777340886&amp;alt=media\"></p></li>\n<li><p><strong>all data(no val score), private 0.44, public 0.34</strong><br>\nkernel=[3, 5, 7, 9], mixup alpha=2.0, 20 epoch, flip augmentation(time reversal, p=0.5) in train and test time, train_batch=32, train all data, seed ensemble, downsample=5</p></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4250230%2F3f6e56ab15cbcb9b41d13dea11708529%2Fwavenet_bestlb.png?generation=1712676748677925&amp;alt=media\" alt=\"img1\"></p>\n<p>This is flip augmentation code.<br>\nI also tried scaling and jittering augmentation but got worse result.</p>\n<pre><code> ():\n     np.flip(X, axis=)\n\n ():\n    scalingFactor = np.random.normal(loc=, scale=sigma, size=(,X.shape[])) \n    myNoise = np.matmul(np.ones((X.shape[],)), scalingFactor)\n     X*myNoise\n\n ():\n    myNoise = np.random.normal(loc=, scale=sigma, size=X.shape)\n     X+myNoise\n</code></pre>\n<p>If you want my full code, I will be happy to share !!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2744596,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2024-04-10T02:45:41.607000",
              "content": "<p>Here is my latest 1D model submission during the last week,<br>\n<a href=\"https://postimg.cc/R3HrP0cX\" target=\"_blank\"><img src=\"https://i.postimg.cc/5tPx2Yjd/Screenshot-2024-04-10-at-10-37-37-AM.png\" alt=\"Screenshot-2024-04-10-at-10-37-37-AM.png\"></a></p>\n<p>The difference with my original released version in the post is that I downsample the sequence along time dimension in the second and third <code>WaveBlock</code>, which accelerates the training speed without compromising too much performance.<br>\nSome more observations from my experiments,</p>\n<ol>\n<li>Augmentations didn't work for my 1D model (<em>e.g.,</em> hflip, vflip, gaussian noise, mixup, time and frequency mask).</li>\n<li>The single 1D model performs not that good in private LB but does contribute to ensembling model due to diversity.</li>\n</ol>\n<p>I think the second point coincide with yours. Thanks for sharing your experiments!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2659685,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2024-02-20T04:48:26.983000",
      "content": "<p>hello - Do we get .680 by just running this training script ? I remember running it for fold 1 and result was not comparable ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2659848,
          "author_name": "Timmy Juicehouse",
          "author_url": "",
          "post_date": "2024-02-20T07:11:13.723000",
          "content": "<pre><code>  Load model  /kaggle//0217---/model-last_seed0_fold0.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold1.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold2.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold3.pth\n  Load model  /kaggle//0217---/model-last_seed0_fold4.pth\n  Sum of row   y_preds .\n</code></pre>\n<p>He loaded all.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2659894,
              "author_name": "Nirjhar Roy",
              "author_url": "",
              "post_date": "2024-02-20T07:45:42.083000",
              "content": "<p>My question was for the CV .. it says 5 fold oof CV is .680 .. But for fold1 CV was around .900+ .. thats why I was thinking if there is some tweak required from baseline</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2660101,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2024-02-20T10:44:20.283000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/phoenix9032\" target=\"_blank\">@phoenix9032</a>,</p>\n<p>The setting is exactly what I used to obtain the CV score. If you want to train 5-fold models, you need to set <code>one_fold_only=False</code>, which is just a debug option for convenience.<br>\nSome observations to note are,</p>\n<ol>\n<li>Different seeds can make CV scores different, which is within the range of 0.02 when I take 3 seeds.</li>\n<li>Different folds have quite different scores (<em>e.g.,</em> fold0 is much worse than fold4).</li>\n</ol>\n<p>For your question, did you finish the 5-epoch training? 0.900+ is observed exactly on epoch 0. I just add a seeding functionality, can you try version2 again? If there's any problem, please let me know. Thanks!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2661451,
              "author_name": "DJ_Xia",
              "author_url": "",
              "post_date": "2024-02-21T08:58:23.890000",
              "content": "<p>add seed and retrain still get diff score</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2661831,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2024-02-21T14:30:06.130000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a>,</p>\n<p>May I ask what CV score you obtain by rerunning version2? <br>\nAs explained above, score can differ fold-by-fold and also seed-by-seed (<em>i.e.,</em> the randomness matters). My CV score 0.68 is reported for only one seed, which is also randomly picked (<em>i.e.,</em> I don't fix seed here). For experiments with more seeds, I get 0.67 or even 0.69 sometimes. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2661862,
              "author_name": "DJ_Xia",
              "author_url": "",
              "post_date": "2024-02-21T14:50:13.833000",
              "content": "<p>0.6762 ~ 0.69+</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2661904,
              "author_name": "AbaoJiang",
              "author_url": "",
              "post_date": "2024-02-21T15:21:23.297000",
              "content": "<p>So, it sounds reasonable. For the previous post of mine, <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/476899\" target=\"_blank\">Reproducibility of Chris' WaveNet Baseline in torch and Randomness Study</a>, I do simple randomness study and show CV scores along two dimensions, fold-dim and seed-dim. As can be observed, the score can fluctuate within a wide range. Hence, randomness factor should be taken into consideration when trying to improve CV scores. For me, cosine decay with only picking the last epoch helps stabilize my CV scores a little bit.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2665885,
              "author_name": "lhwcv",
              "author_url": "",
              "post_date": "2024-02-23T23:05:25.753000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/tonymarkchris\" target=\"_blank\">@tonymarkchris</a> ，how to avoid diff score even when retrain?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2671595,
      "author_name": "Rafael Zimmermann",
      "author_url": "",
      "post_date": "2024-02-27T15:44:52.040000",
      "content": "<p>Nice work abaojiang, I see wavenet as a great way to give diversity to the ensembles used in the competition, I will use your starter model, keep sharing the good work</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2671901,
          "author_name": "AbaoJiang",
          "author_url": "",
          "post_date": "2024-02-27T18:57:54.337000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rafaelzimmermann1\" target=\"_blank\">@rafaelzimmermann1</a>,</p>\n<p>Thanks for your appreciation, I'm glad to hear that it works for you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2664143,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-22T20:23:10.623000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2659070": "Hello everyone,\n\nAfter reading this forum, [Raw EEG CV vs LB](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/473821), I'm also interested in finding out how far raw EEG model can go without using spectrogram information. Through many experiments, I finally build a model architecture to get **CV 0.67 and LB 0.46**. Let's see the model architecture first,\n[![Screenshot-2024-02-19-at-1-11-40-PM.png](https://i.postimg.cc/MKp8xVVV/Screenshot-2024-02-19-at-1-11-40-PM.png)](https://postimg.cc/7bdRnCBZ)\n\nThis model architecture can be seen as an extension of Chris' version. I modify the original dilated convolution to **dilated inception convolution**. Concretely speaking, instead of using only one kernel size per layer in `WaveBlock`, this model tries to capture enriched temporal patterns by considering different kernel sizes (2, 3, 6, and 7 here) at the same time. Furthermore, to restrain model parameters from growing too much, the model **narrows down the output channels** of convolutions with different kernel sizes. Let `out_channels = 64`, each dilated convolution with a specific kernel size only outputs 16 channels, then the concatenation along channel dimension is applied to match `out_channels = 64`.\n\nAdding this auxiliary block helps boost my CV about 0.15 and LB about 0.06 (compared with Chris'). If you're interested in training your own models and submitting in pure PyTorch way, you can refer to these two notebooks,\n* [[LB 0.46] DilatedInception WaveNet - Training](https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-training)\n* [[LB 0.46] DilatedInception WaveNet - Inference](https://www.kaggle.com/code/abaojiang/lb-0-46-dilatedinception-wavenet-inference?scriptVersionId=163448688)\n\nThanks for reading!",
    "2661599": "Good work! with some modifications I was able to get **CV 0.6 and LB 0.38**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3350264%2Fad9e55523487c56ceecae0d3a719163f%2FScreenshot%20from%202024-02-21%2016-17-24.png?generation=1708514299642251&alt=media)",
    "2659559": "Thank you. I even got below 0.5 by using lightgbm with massive feature engineering for raw EEG data. You just use basic features.\n        \n        x_tmp = np.zeros((EEG_PTS, self.n_feats), dtype=\"float32\")\n        x_tmp[:, 0] = x[:, self.FEAT2CODE[\"Fp1\"]] - x[:, self.FEAT2CODE[\"T3\"]]\n        x_tmp[:, 1] = x[:, self.FEAT2CODE[\"T3\"]] - x[:, self.FEAT2CODE[\"O1\"]]\n        x_tmp[:, 2] = x[:, self.FEAT2CODE[\"Fp1\"]] - x[:, self.FEAT2CODE[\"C3\"]]\n        x_tmp[:, 3] = x[:, self.FEAT2CODE[\"C3\"]] - x[:, self.FEAT2CODE[\"O1\"]]\n        x_tmp[:, 4] = x[:, self.FEAT2CODE[\"Fp2\"]] - x[:, self.FEAT2CODE[\"C4\"]]\n        x_tmp[:, 5] = x[:, self.FEAT2CODE[\"C4\"]] - x[:, self.FEAT2CODE[\"O2\"]]\n        x_tmp[:, 6] = x[:, self.FEAT2CODE[\"Fp2\"]] - x[:, self.FEAT2CODE[\"T4\"]]\n        x_tmp[:, 7] = x[:, self.FEAT2CODE[\"T4\"]] - x[:, self.FEAT2CODE[\"O2\"]]\n\n",
    "2702519": "thanks for sharing. I got LB 0.35 without any change of model architecture.\nUPDATE: LB0.34",
    "2659685": "hello - Do we get .680 by just running this training script ? I remember running it for fold 1 and result was not comparable ?",
    "2671595": "Nice work abaojiang, I see wavenet as a great way to give diversity to the ensembles used in the competition, I will use your starter model, keep sharing the good work",
    "2664143": ""
  }
}