{
  "id": 492211,
  "title": "33th solution : Crazy 10 days sprint with 2D pipeline",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492211",
  "author_name": "CPMP",
  "post_date": "2024-04-09T00:59:02.523000",
  "votes": 62,
  "comment_count": 23,
  "views": 0,
  "content": "<p>I joined quite late (10 days before end) and have lukewarm feelings with my result. It is good enough to not have regrets for entering so late. At the same time, I am not sure that few more days would have helped me. One more month maybe.</p>\n<p>Anyway, entering late was only made possible because of the high quality sharing from all the community, and its distillation by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . A remark by <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a> on 10+ votes being a reliable CV was the other major sharing for me. <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> post on 10+vote mean being better than all data mean was also one of the key insights for me. There is more, and I can't cite them all. But I thank them warmly. Small data was also key as it enabled fast iterations.</p>\n<p>Given the short time frame I could not invest a lot of time on things that did not pay off immediately. It is why I did not use 1D models at all, I could not make them work well enough.</p>\n<p>My pipeline is similar to Chris efficientnet starter with these differences:</p>\n<ul>\n<li>Both mel spectrograms and Q transforms (QT) were used. I used QT successfully in g2net competition 3 years ago (see <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433</a>)</li>\n<li>All data processing is done in GPU, torchaudio for STFT and Mel Spectrograms, a modified nnAudio code for QT.</li>\n<li>Data normalization was important, parameters of the log transform were chosen to yield an average mean close to 1 and values between -1 and 1.</li>\n<li>The 16 montage differences were directly used instead of averaging spectrograms over the 4 lines.</li>\n<li>Each of the 16 was used to create a small 32 x250 spectrogram  or Q transform. For spectrogram it meant 32 mels, and  n_fft being as low as 192. for Q transform it meant 5 octaves with 32 bins.</li>\n<li>50 second spectrograms or QT were stacked into a 512 x 250 image. Middle 10 seconds spectrograms or QT were stacked into a 512 x 250 image and the two were concatenated frequency side by side. Original Kaggle 10 mins spectrograms were added as in Chris' pipeline.</li>\n<li>fmax was set to 20 Hz or 25 Hz depending on the runs. 25 Hz was a bit better in general.</li>\n<li>I worked all the competition till last day with tf-efficientnet-b0 for the speed. It enabled me to experiment quickly. One epoch ran in 1 minute or less on a single V100 GPU. The last day I used larger efficientnets, b1 to b3, and did not have time to tune for more recent ones.</li>\n<li>targets were aggregated by eeg_id as in Chris' pipeline</li>\n<li>Kaggle 10 mins spectrograms were sampled randomly from min and max offsets during training. This was quite effective. </li>\n<li>I tried ways to sample EEGs from other places than the middle and did not see much difference. I stayed with the simplicity.</li>\n<li>I used pseudo labeling (aka teach student). First trained models on 10+ votes data. Then used them to predict on the low votes data. Then used that as additional training data. This worked well, it yields optimistic CV scores but LB always improved.</li>\n<li>I think that the issue with low votes is not so much the votes. It is rather the vote distribution. For instance seizure happens in about 30% of low votes while it is only few percents in high vote. To avoid this distribution shift I sample pseudo labelled data by target, to keep the same distribution as training data.</li>\n</ul>\n<p>There are many things I wish I could have tried but overall it went well, with my LB score improving at every submission. Correlation between public, score, private score, and 10 votes CV was great. I did not select my subs, best public was also best private and best CV.</p>\n<p>Now I will rest a bit and read all about the 1D models I should have used if I had one more month! Happy kaggling!</p>",
  "messages": [
    {
      "id": 2742561,
      "postDate": "2024-04-09T00:59:02.523Z",
      "content": "<p>I joined quite late (10 days before end) and have lukewarm feelings with my result. It is good enough to not have regrets for entering so late. At the same time, I am not sure that few more days would have helped me. One more month maybe.</p>\n<p>Anyway, entering late was only made possible because of the high quality sharing from all the community, and its distillation by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> . A remark by <a href=\"https://www.kaggle.com/goldenlock\" target=\"_blank\">@goldenlock</a> on 10+ votes being a reliable CV was the other major sharing for me. <a href=\"https://www.kaggle.com/pcjimmmy\" target=\"_blank\">@pcjimmmy</a> post on 10+vote mean being better than all data mean was also one of the key insights for me. There is more, and I can't cite them all. But I thank them warmly. Small data was also key as it enabled fast iterations.</p>\n<p>Given the short time frame I could not invest a lot of time on things that did not pay off immediately. It is why I did not use 1D models at all, I could not make them work well enough.</p>\n<p>My pipeline is similar to Chris efficientnet starter with these differences:</p>\n<ul>\n<li>Both mel spectrograms and Q transforms (QT) were used. I used QT successfully in g2net competition 3 years ago (see <a href=\"https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433\" target=\"_blank\">https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433</a>)</li>\n<li>All data processing is done in GPU, torchaudio for STFT and Mel Spectrograms, a modified nnAudio code for QT.</li>\n<li>Data normalization was important, parameters of the log transform were chosen to yield an average mean close to 1 and values between -1 and 1.</li>\n<li>The 16 montage differences were directly used instead of averaging spectrograms over the 4 lines.</li>\n<li>Each of the 16 was used to create a small 32 x250 spectrogram  or Q transform. For spectrogram it meant 32 mels, and  n_fft being as low as 192. for Q transform it meant 5 octaves with 32 bins.</li>\n<li>50 second spectrograms or QT were stacked into a 512 x 250 image. Middle 10 seconds spectrograms or QT were stacked into a 512 x 250 image and the two were concatenated frequency side by side. Original Kaggle 10 mins spectrograms were added as in Chris' pipeline.</li>\n<li>fmax was set to 20 Hz or 25 Hz depending on the runs. 25 Hz was a bit better in general.</li>\n<li>I worked all the competition till last day with tf-efficientnet-b0 for the speed. It enabled me to experiment quickly. One epoch ran in 1 minute or less on a single V100 GPU. The last day I used larger efficientnets, b1 to b3, and did not have time to tune for more recent ones.</li>\n<li>targets were aggregated by eeg_id as in Chris' pipeline</li>\n<li>Kaggle 10 mins spectrograms were sampled randomly from min and max offsets during training. This was quite effective. </li>\n<li>I tried ways to sample EEGs from other places than the middle and did not see much difference. I stayed with the simplicity.</li>\n<li>I used pseudo labeling (aka teach student). First trained models on 10+ votes data. Then used them to predict on the low votes data. Then used that as additional training data. This worked well, it yields optimistic CV scores but LB always improved.</li>\n<li>I think that the issue with low votes is not so much the votes. It is rather the vote distribution. For instance seizure happens in about 30% of low votes while it is only few percents in high vote. To avoid this distribution shift I sample pseudo labelled data by target, to keep the same distribution as training data.</li>\n</ul>\n<p>There are many things I wish I could have tried but overall it went well, with my LB score improving at every submission. Correlation between public, score, private score, and 10 votes CV was great. I did not select my subs, best public was also best private and best CV.</p>\n<p>Now I will rest a bit and read all about the 1D models I should have used if I had one more month! Happy kaggling!</p>",
      "rawMarkdown": "I joined quite late (10 days before end) and have lukewarm feelings with my result. It is good enough to not have regrets for entering so late. At the same time, I am not sure that few more days would have helped me. One more month maybe.\n\nAnyway, entering late was only made possible because of the high quality sharing from all the community, and its distillation by @cdeotte . A remark by @goldenlock on 10+ votes being a reliable CV was the other major sharing for me. @pcjimmmy post on 10+vote mean being better than all data mean was also one of the key insights for me. There is more, and I can't cite them all. But I thank them warmly. Small data was also key as it enabled fast iterations.\n\nGiven the short time frame I could not invest a lot of time on things that did not pay off immediately. It is why I did not use 1D models at all, I could not make them work well enough.\n\nMy pipeline is similar to Chris efficientnet starter with these differences:\n\n- Both mel spectrograms and Q transforms (QT) were used. I used QT successfully in g2net competition 3 years ago (see https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433)\n- All data processing is done in GPU, torchaudio for STFT and Mel Spectrograms, a modified nnAudio code for QT.\n- Data normalization was important, parameters of the log transform were chosen to yield an average mean close to 1 and values between -1 and 1.\n- The 16 montage differences were directly used instead of averaging spectrograms over the 4 lines.\n- Each of the 16 was used to create a small 32 x250 spectrogram  or Q transform. For spectrogram it meant 32 mels, and  n_fft being as low as 192. for Q transform it meant 5 octaves with 32 bins.\n- 50 second spectrograms or QT were stacked into a 512 x 250 image. Middle 10 seconds spectrograms or QT were stacked into a 512 x 250 image and the two were concatenated frequency side by side. Original Kaggle 10 mins spectrograms were added as in Chris' pipeline.\n- fmax was set to 20 Hz or 25 Hz depending on the runs. 25 Hz was a bit better in general.\n- I worked all the competition till last day with tf-efficientnet-b0 for the speed. It enabled me to experiment quickly. One epoch ran in 1 minute or less on a single V100 GPU. The last day I used larger efficientnets, b1 to b3, and did not have time to tune for more recent ones.\n- targets were aggregated by eeg_id as in Chris' pipeline\n- Kaggle 10 mins spectrograms were sampled randomly from min and max offsets during training. This was quite effective. \n- I tried ways to sample EEGs from other places than the middle and did not see much difference. I stayed with the simplicity.\n- I used pseudo labeling (aka teach student). First trained models on 10+ votes data. Then used them to predict on the low votes data. Then used that as additional training data. This worked well, it yields optimistic CV scores but LB always improved.\n- I think that the issue with low votes is not so much the votes. It is rather the vote distribution. For instance seizure happens in about 30% of low votes while it is only few percents in high vote. To avoid this distribution shift I sample pseudo labelled data by target, to keep the same distribution as training data.\n\nThere are many things I wish I could have tried but overall it went well, with my LB score improving at every submission. Correlation between public, score, private score, and 10 votes CV was great. I did not select my subs, best public was also best private and best CV.\n\nNow I will rest a bit and read all about the 1D models I should have used if I had one more month! Happy kaggling!",
      "votes": 62
    },
    {
      "id": 2743408,
      "postDate": "2024-04-09T12:47:27.600Z",
      "content": "<p>Good work!</p>\n<p>\"At the same time, I am not sure that few more days would have helped me. One more month maybe.\"<br>\ngone are the days when you can win a kaggle competition within one week or two.<br>\ni also face the same problems that modern kaggle competitions need more efforts : more resources, more experiments, more …</p>",
      "rawMarkdown": "Good work!\n\n\"At the same time, I am not sure that few more days would have helped me. One more month maybe.\"\ngone are the days when you can win a kaggle competition within one week or two.\ni also face the same problems that modern kaggle competitions need more efforts : more resources, more experiments, more ...\n",
      "votes": 5,
      "replies": [
        {
          "id": 2743424,
          "postDate": "2024-04-09T12:57:07.833Z",
          "content": "<p>Yes, it was crazy to assume gold could be reached in 10 days. But I prefer to try and fail than not try and regret!</p>",
          "rawMarkdown": "Yes, it was crazy to assume gold could be reached in 10 days. But I prefer to try and fail than not try and regret!",
          "votes": 1
        },
        {
          "id": 2743486,
          "postDate": "2024-04-09T13:27:50.187Z",
          "content": "<blockquote>\n  <p>i also face the same problems that modern kaggle competitions need more efforts : more resources, more experiments, more …</p>\n</blockquote>\n<p>I have also notice this trend on Kaggle. In this comp, I performed 1000 experiments to find my solution. My original models after dozens of experiments are good but won't win gold.</p>",
          "rawMarkdown": ">i also face the same problems that modern kaggle competitions need more efforts : more resources, more experiments, more …\n\nI have also notice this trend on Kaggle. In this comp, I performed 1000 experiments to find my solution. My original models after dozens of experiments are good but won't win gold.",
          "votes": 8,
          "replies": [
            {
              "id": 2743509,
              "postDate": "2024-04-09T13:49:40.803Z",
              "content": "<p>I have 414 experiments…</p>",
              "rawMarkdown": "I have 414 experiments...\n",
              "votes": 1
            },
            {
              "id": 2743584,
              "postDate": "2024-04-09T14:26:28.373Z",
              "content": "<p>414/10 = 41 experiments per day<br>\n41/24 = 2 experiments per hour</p>\n<p>many people do not understand the meaning of \"experiences\" in datascience/AI/ML<br>\nIt is not something that can be gain from online courses or a few projects from github, etc. you need to repeat paper's results or join competitions …</p>\n<p>I use to ask my job interviewees:<br>\n\"for software engineering we always asked how many lines of codes you have written.<br>\nNow what is your experience level in ML?  how many models have you trained last year and on average, how many experiments you have done per day?\"</p>",
              "rawMarkdown": "414/10 = 41 experiments per day\n41/24 = 2 experiments per hour\n\nmany people do not understand the meaning of \"experiences\" in datascience/AI/ML\nIt is not something that can be gain from online courses or a few projects from github, etc. you need to repeat paper's results or join competitions ...\n\nI use to ask my job interviewees:\n\"for software engineering we always asked how many lines of codes you have written.\nNow what is your experience level in ML?  how many models have you trained last year and on average, how many experiments you have done per day?\"\n",
              "votes": 5
            },
            {
              "id": 2743672,
              "postDate": "2024-04-09T15:12:00.410Z",
              "content": "<p>The number of experiment per day is even higher as there is ramp up in producing a pipeline.</p>",
              "rawMarkdown": "The number of experiment per day is even higher as there is ramp up in producing a pipeline.",
              "votes": 1
            },
            {
              "id": 2743755,
              "postDate": "2024-04-09T15:54:49.253Z",
              "content": "<blockquote>\n  <p>many people do not understand the meaning of \"experiences\" in datascience/AI/ML<br>\n  It is not something that can be gain from online courses or a few projects from github, etc. you need to repeat paper's results or join competitions …</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you made a very good point. I always explain that ML is an applied science like physics, chemistry or biology (sure there are theoretical physics, but most of it is experimental).  one needs to make hypothesis, design ways to test them, perform tests in a rigorous manner, and analyze results. It is why former physicists become good ML people in general.</p>",
              "rawMarkdown": "> many people do not understand the meaning of \"experiences\" in datascience/AI/ML\nIt is not something that can be gain from online courses or a few projects from github, etc. you need to repeat paper's results or join competitions …\n\n@hengck23 you made a very good point. I always explain that ML is an applied science like physics, chemistry or biology (sure there are theoretical physics, but most of it is experimental).  one needs to make hypothesis, design ways to test them, perform tests in a rigorous manner, and analyze results. It is why former physicists become good ML people in general.\n",
              "votes": 7
            },
            {
              "id": 2745166,
              "postDate": "2024-04-10T12:33:04.493Z",
              "content": "<p>Does your one experiment include various hyper parameters ? </p>\n<p>It's unimaginable speed  for deep learning beginner, me, to gather information and do coding, debugging and computing.<br>\nMoreover, we need tough mind to keep changing things even if the experiment that was expected to boost CV does not go well.. </p>",
              "rawMarkdown": "Does your one experiment include various hyper parameters ? \n\nIt's unimaginable speed  for deep learning beginner, me, to gather information and do coding, debugging and computing.\nMoreover, we need tough mind to keep changing things even if the experiment that was expected to boost CV does not go well.. "
            },
            {
              "id": 2745359,
              "postDate": "2024-04-10T15:12:22.203Z",
              "content": "<blockquote>\n  <p>Does your one experiment include various hyper parameters ? </p>\n</blockquote>\n<p>Yes. I use largest possible batch size, and I tune max learning rate and number of epochs, that's it. I don't use early stopping.</p>",
              "rawMarkdown": "> Does your one experiment include various hyper parameters ? \n\nYes. I use largest possible batch size, and I tune max learning rate and number of epochs, that's it. I don't use early stopping.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2743410,
      "postDate": "2024-04-09T12:48:47.390Z",
      "content": "<blockquote>\n  <p>I joined quite late (10 days before end) </p>\n</blockquote>\n<p>It is so surprising for me that you reached the very good score in such a short period!<br>\nI'm curious about how to speed up modeling experiments.<br>\nDo you check the cv score for any single change to your model when you test your ideas?</p>",
      "rawMarkdown": ">I joined quite late (10 days before end) \n\nIt is so surprising for me that you reached the very good score in such a short period!\nI'm curious about how to speed up modeling experiments.\nDo you check the cv score for any single change to your model when you test your ideas?",
      "votes": 1,
      "replies": [
        {
          "id": 2743422,
          "postDate": "2024-04-09T12:56:27.157Z",
          "content": "<p>Main thing was to move all computation to GPU. Once I did that I could experiment with different spectrogram generation in real time, running one fold was taking me 5 minutes or so.</p>",
          "rawMarkdown": "Main thing was to move all computation to GPU. Once I did that I could experiment with different spectrogram generation in real time, running one fold was taking me 5 minutes or so.",
          "votes": 1,
          "replies": [
            {
              "id": 2743479,
              "postDate": "2024-04-09T13:23:43.937Z",
              "content": "<p>I see… I will check public notebooks to learn how to accelerate the whole process on GPU including spectrogram generation and so on. Thanks for the answer!</p>",
              "rawMarkdown": "I see... I will check public notebooks to learn how to accelerate the whole process on GPU including spectrogram generation and so on. Thanks for the answer!"
            },
            {
              "id": 2743484,
              "postDate": "2024-04-09T13:26:20.487Z",
              "content": "<blockquote>\n  <p>Once I did that I could experiment with different spectrogram generation in real time, running one fold was taking me 5 minutes or so.</p>\n</blockquote>\n<p>Absolutely. Furthermore we can do data augmentation on the raw waveforms instead of data augmentation on the spectrograms. This allows for more augmentation techniques.</p>",
              "rawMarkdown": ">Once I did that I could experiment with different spectrogram generation in real time, running one fold was taking me 5 minutes or so.\n\nAbsolutely. Furthermore we can do data augmentation on the raw waveforms instead of data augmentation on the spectrograms. This allows for more augmentation techniques.",
              "votes": 3
            },
            {
              "id": 2744004,
              "postDate": "2024-04-09T17:48:47.363Z",
              "content": "<p><a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> shared a pipeline with torchaudio, look for his notebooks.</p>",
              "rawMarkdown": "@iglovikov shared a pipeline with torchaudio, look for his notebooks.",
              "votes": 3
            },
            {
              "id": 2744586,
              "postDate": "2024-04-10T02:18:08.337Z",
              "content": "<p>Thanks! I'll check it!</p>",
              "rawMarkdown": "Thanks! I'll check it!"
            }
          ]
        }
      ]
    },
    {
      "id": 2743488,
      "postDate": "2024-04-09T13:28:27.287Z",
      "content": "<p>Congratulations CPMP. This is great work in 10 days. I saw you jump upward on the leaderboard in the last week and was really impressed.</p>",
      "rawMarkdown": "Congratulations CPMP. This is great work in 10 days. I saw you jump upward on the leaderboard in the last week and was really impressed.",
      "votes": 2,
      "replies": [
        {
          "id": 2743518,
          "postDate": "2024-04-09T13:54:30.953Z",
          "content": "<p>Thanks. I owe a lot to your sharing.</p>\n<p>I am amazed by how effective the aggregation of target by eeg_id is. I am not sure I would have found it myself.</p>",
          "rawMarkdown": "Thanks. I owe a lot to your sharing.\n\nI am amazed by how effective the aggregation of target by eeg_id is. I am not sure I would have found it myself.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2743318,
      "postDate": "2024-04-09T11:31:59.953Z",
      "content": "<p>Maybe few more days would have helped. My next todo was to try other backbones than the good old efficientnets.</p>\n<p>Here are my last subs:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F151105d05cddd344bffa0504c1ca5b63%2Fimage%20(33).png?generation=1712662301487223&amp;alt=media\"></p>",
      "rawMarkdown": "Maybe few more days would have helped. My next todo was to try other backbones than the good old efficientnets.\n\nHere are my last subs:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F151105d05cddd344bffa0504c1ca5b63%2Fimage%20(33).png?generation=1712662301487223&alt=media)",
      "votes": 2,
      "replies": [
        {
          "id": 2743480,
          "postDate": "2024-04-09T13:23:44.633Z",
          "content": "<p>I'm impressed that you achieved such a good score with EfficientNet. For me all vision transformers outperform CNN, specifically using <code>tiny_vit_21m_512</code> gave <code>+0.02</code> boost. You should try this backbone in your model. Note you will probably need to retune your learning rate and schedule to get the most out of VIT.</p>",
          "rawMarkdown": "I'm impressed that you achieved such a good score with EfficientNet. For me all vision transformers outperform CNN, specifically using `tiny_vit_21m_512` gave `+0.02` boost. You should try this backbone in your model. Note you will probably need to retune your learning rate and schedule to get the most out of VIT.",
          "votes": 2,
          "replies": [
            {
              "id": 2743515,
              "postDate": "2024-04-09T13:52:38.390Z",
              "content": "<p>I did not even use efficientnet v2!</p>\n<p>I now thinks that using more recent backbones like you did was the way to go. But it is easy to say in hindsight.</p>",
              "rawMarkdown": "I did not even use efficientnet v2!\n\nI now thinks that using more recent backbones like you did was the way to go. But it is easy to say in hindsight.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2742575,
      "postDate": "2024-04-09T01:08:47.933Z",
      "content": "<p>Good idea to resmaple votes &lt;=10 data, I used lowe weight for them, but I could see my model would produce higher seizure rate due to distribution diff. Finetune on votes &gt;= 10 data yet would hurt cv a bit but seems it helps on LB and PB.</p>",
      "rawMarkdown": "Good idea to resmaple votes <=10 data, I used lowe weight for them, but I could see my model would produce higher seizure rate due to distribution diff. Finetune on votes >= 10 data yet would hurt cv a bit but seems it helps on LB and PB.",
      "votes": 2
    },
    {
      "id": 2746720,
      "postDate": "2024-04-11T12:20:59.977Z",
      "content": "<p>Great Work</p>",
      "rawMarkdown": "Great Work\n"
    },
    {
      "id": 2747388,
      "postDate": "2024-04-11T20:50:24.803Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2743408,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-09T12:47:27.600000",
      "content": "<p>Good work!</p>\n<p>\"At the same time, I am not sure that few more days would have helped me. One more month maybe.\"<br>\ngone are the days when you can win a kaggle competition within one week or two.<br>\ni also face the same problems that modern kaggle competitions need more efforts : more resources, more experiments, more …</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2743424,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2024-04-09T12:57:07.833000",
          "content": "<p>Yes, it was crazy to assume gold could be reached in 10 days. But I prefer to try and fail than not try and regret!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2743486,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T13:27:50.187000",
          "content": "<blockquote>\n  <p>i also face the same problems that modern kaggle competitions need more efforts : more resources, more experiments, more …</p>\n</blockquote>\n<p>I have also notice this trend on Kaggle. In this comp, I performed 1000 experiments to find my solution. My original models after dozens of experiments are good but won't win gold.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2743509,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-09T13:49:40.803000",
              "content": "<p>I have 414 experiments…</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2743584,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-04-09T14:26:28.373000",
              "content": "<p>414/10 = 41 experiments per day<br>\n41/24 = 2 experiments per hour</p>\n<p>many people do not understand the meaning of \"experiences\" in datascience/AI/ML<br>\nIt is not something that can be gain from online courses or a few projects from github, etc. you need to repeat paper's results or join competitions …</p>\n<p>I use to ask my job interviewees:<br>\n\"for software engineering we always asked how many lines of codes you have written.<br>\nNow what is your experience level in ML?  how many models have you trained last year and on average, how many experiments you have done per day?\"</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2743672,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-09T15:12:00.410000",
              "content": "<p>The number of experiment per day is even higher as there is ramp up in producing a pipeline.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2743755,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-09T15:54:49.253000",
              "content": "<blockquote>\n  <p>many people do not understand the meaning of \"experiences\" in datascience/AI/ML<br>\n  It is not something that can be gain from online courses or a few projects from github, etc. you need to repeat paper's results or join competitions …</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> you made a very good point. I always explain that ML is an applied science like physics, chemistry or biology (sure there are theoretical physics, but most of it is experimental).  one needs to make hypothesis, design ways to test them, perform tests in a rigorous manner, and analyze results. It is why former physicists become good ML people in general.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 2745166,
              "author_name": "Aurora_blue",
              "author_url": "",
              "post_date": "2024-04-10T12:33:04.493000",
              "content": "<p>Does your one experiment include various hyper parameters ? </p>\n<p>It's unimaginable speed  for deep learning beginner, me, to gather information and do coding, debugging and computing.<br>\nMoreover, we need tough mind to keep changing things even if the experiment that was expected to boost CV does not go well.. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2745359,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-10T15:12:22.203000",
              "content": "<blockquote>\n  <p>Does your one experiment include various hyper parameters ? </p>\n</blockquote>\n<p>Yes. I use largest possible batch size, and I tune max learning rate and number of epochs, that's it. I don't use early stopping.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2743410,
      "author_name": "nynyny67",
      "author_url": "",
      "post_date": "2024-04-09T12:48:47.390000",
      "content": "<blockquote>\n  <p>I joined quite late (10 days before end) </p>\n</blockquote>\n<p>It is so surprising for me that you reached the very good score in such a short period!<br>\nI'm curious about how to speed up modeling experiments.<br>\nDo you check the cv score for any single change to your model when you test your ideas?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2743422,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2024-04-09T12:56:27.157000",
          "content": "<p>Main thing was to move all computation to GPU. Once I did that I could experiment with different spectrogram generation in real time, running one fold was taking me 5 minutes or so.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2743479,
              "author_name": "nynyny67",
              "author_url": "",
              "post_date": "2024-04-09T13:23:43.937000",
              "content": "<p>I see… I will check public notebooks to learn how to accelerate the whole process on GPU including spectrogram generation and so on. Thanks for the answer!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2743484,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-04-09T13:26:20.487000",
              "content": "<blockquote>\n  <p>Once I did that I could experiment with different spectrogram generation in real time, running one fold was taking me 5 minutes or so.</p>\n</blockquote>\n<p>Absolutely. Furthermore we can do data augmentation on the raw waveforms instead of data augmentation on the spectrograms. This allows for more augmentation techniques.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2744004,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-09T17:48:47.363000",
              "content": "<p><a href=\"https://www.kaggle.com/iglovikov\" target=\"_blank\">@iglovikov</a> shared a pipeline with torchaudio, look for his notebooks.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2744586,
              "author_name": "nynyny67",
              "author_url": "",
              "post_date": "2024-04-10T02:18:08.337000",
              "content": "<p>Thanks! I'll check it!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2743488,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-04-09T13:28:27.287000",
      "content": "<p>Congratulations CPMP. This is great work in 10 days. I saw you jump upward on the leaderboard in the last week and was really impressed.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2743518,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2024-04-09T13:54:30.953000",
          "content": "<p>Thanks. I owe a lot to your sharing.</p>\n<p>I am amazed by how effective the aggregation of target by eeg_id is. I am not sure I would have found it myself.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2743318,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-04-09T11:31:59.953000",
      "content": "<p>Maybe few more days would have helped. My next todo was to try other backbones than the good old efficientnets.</p>\n<p>Here are my last subs:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F151105d05cddd344bffa0504c1ca5b63%2Fimage%20(33).png?generation=1712662301487223&amp;alt=media\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 2743480,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T13:23:44.633000",
          "content": "<p>I'm impressed that you achieved such a good score with EfficientNet. For me all vision transformers outperform CNN, specifically using <code>tiny_vit_21m_512</code> gave <code>+0.02</code> boost. You should try this backbone in your model. Note you will probably need to retune your learning rate and schedule to get the most out of VIT.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2743515,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-09T13:52:38.390000",
              "content": "<p>I did not even use efficientnet v2!</p>\n<p>I now thinks that using more recent backbones like you did was the way to go. But it is easy to say in hindsight.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2742575,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2024-04-09T01:08:47.933000",
      "content": "<p>Good idea to resmaple votes &lt;=10 data, I used lowe weight for them, but I could see my model would produce higher seizure rate due to distribution diff. Finetune on votes &gt;= 10 data yet would hurt cv a bit but seems it helps on LB and PB.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2746720,
      "author_name": "Ganesh Talwar",
      "author_url": "",
      "post_date": "2024-04-11T12:20:59.977000",
      "content": "<p>Great Work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2747388,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-11T20:50:24.803000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742561": "I joined quite late (10 days before end) and have lukewarm feelings with my result. It is good enough to not have regrets for entering so late. At the same time, I am not sure that few more days would have helped me. One more month maybe.\n\nAnyway, entering late was only made possible because of the high quality sharing from all the community, and its distillation by @cdeotte . A remark by @goldenlock on 10+ votes being a reliable CV was the other major sharing for me. @pcjimmmy post on 10+vote mean being better than all data mean was also one of the key insights for me. There is more, and I can't cite them all. But I thank them warmly. Small data was also key as it enabled fast iterations.\n\nGiven the short time frame I could not invest a lot of time on things that did not pay off immediately. It is why I did not use 1D models at all, I could not make them work well enough.\n\nMy pipeline is similar to Chris efficientnet starter with these differences:\n\n- Both mel spectrograms and Q transforms (QT) were used. I used QT successfully in g2net competition 3 years ago (see https://www.kaggle.com/competitions/g2net-gravitational-wave-detection/discussion/275433)\n- All data processing is done in GPU, torchaudio for STFT and Mel Spectrograms, a modified nnAudio code for QT.\n- Data normalization was important, parameters of the log transform were chosen to yield an average mean close to 1 and values between -1 and 1.\n- The 16 montage differences were directly used instead of averaging spectrograms over the 4 lines.\n- Each of the 16 was used to create a small 32 x250 spectrogram  or Q transform. For spectrogram it meant 32 mels, and  n_fft being as low as 192. for Q transform it meant 5 octaves with 32 bins.\n- 50 second spectrograms or QT were stacked into a 512 x 250 image. Middle 10 seconds spectrograms or QT were stacked into a 512 x 250 image and the two were concatenated frequency side by side. Original Kaggle 10 mins spectrograms were added as in Chris' pipeline.\n- fmax was set to 20 Hz or 25 Hz depending on the runs. 25 Hz was a bit better in general.\n- I worked all the competition till last day with tf-efficientnet-b0 for the speed. It enabled me to experiment quickly. One epoch ran in 1 minute or less on a single V100 GPU. The last day I used larger efficientnets, b1 to b3, and did not have time to tune for more recent ones.\n- targets were aggregated by eeg_id as in Chris' pipeline\n- Kaggle 10 mins spectrograms were sampled randomly from min and max offsets during training. This was quite effective. \n- I tried ways to sample EEGs from other places than the middle and did not see much difference. I stayed with the simplicity.\n- I used pseudo labeling (aka teach student). First trained models on 10+ votes data. Then used them to predict on the low votes data. Then used that as additional training data. This worked well, it yields optimistic CV scores but LB always improved.\n- I think that the issue with low votes is not so much the votes. It is rather the vote distribution. For instance seizure happens in about 30% of low votes while it is only few percents in high vote. To avoid this distribution shift I sample pseudo labelled data by target, to keep the same distribution as training data.\n\nThere are many things I wish I could have tried but overall it went well, with my LB score improving at every submission. Correlation between public, score, private score, and 10 votes CV was great. I did not select my subs, best public was also best private and best CV.\n\nNow I will rest a bit and read all about the 1D models I should have used if I had one more month! Happy kaggling!",
    "2743408": "Good work!\n\n\"At the same time, I am not sure that few more days would have helped me. One more month maybe.\"\ngone are the days when you can win a kaggle competition within one week or two.\ni also face the same problems that modern kaggle competitions need more efforts : more resources, more experiments, more ...\n",
    "2743410": ">I joined quite late (10 days before end) \n\nIt is so surprising for me that you reached the very good score in such a short period!\nI'm curious about how to speed up modeling experiments.\nDo you check the cv score for any single change to your model when you test your ideas?",
    "2743488": "Congratulations CPMP. This is great work in 10 days. I saw you jump upward on the leaderboard in the last week and was really impressed.",
    "2743318": "Maybe few more days would have helped. My next todo was to try other backbones than the good old efficientnets.\n\nHere are my last subs:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F75976%2F151105d05cddd344bffa0504c1ca5b63%2Fimage%20(33).png?generation=1712662301487223&alt=media)",
    "2742575": "Good idea to resmaple votes <=10 data, I used lowe weight for them, but I could see my model would produce higher seizure rate due to distribution diff. Finetune on votes >= 10 data yet would hurt cv a bit but seems it helps on LB and PB.",
    "2746720": "Great Work\n",
    "2747388": ""
  }
}