{
  "id": 492482,
  "title": "8th Place Gold - One Mega Model",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/492482",
  "author_name": "Chris Deotte",
  "post_date": "2024-04-09T18:08:31.937000",
  "votes": 159,
  "comment_count": 62,
  "views": 0,
  "content": "<h1>Thank You Kaggle and Hosts!</h1>\n<p>Very fun competition! I love competitions with opportunities to be creative, explore, and make discoveries! In the past 3 months, I enthusiastically conducted 1000+ experiments!</p>\n<p>I also love competitions where a single model can achieve Gold medal (instead of a large ensemble).</p>\n<p>Thank you Kagglers for all the great discussion shares and notebook shares!</p>\n<h1>Understand Test Data - What is Competition Task?</h1>\n<p>The task in this competition is to <strong>predict annotator opinions</strong>. By reading the host's paper <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf\" target=\"_blank\">here</a>, we learn that there are <strong>119 overall annotators</strong> in train data and <strong>20 expert annotators</strong> in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%. </p>\n<p>Wow, this is a big difference therefore we need to train our models with <strong>expert opinions</strong> instead of <strong>overall opinions</strong>. We can locate the expert annotators in the train data with the filter <code>expert annotators = train.loc[train.vote_count&gt;=10]</code>.</p>\n<h1>Which Features Do Annotators See?</h1>\n<p>From the paper we see that the annotators see 2D spectrograms and 2D plots of waveforms. So let's input 2D spectrograms and 2D plots of waveforms into our models. (Also inputting 1D raw waveform helped too).</p>\n<h1>Train 3 Separate Models with Vote&gt;=10</h1>\n<p>First we train 3 separate models using train data with <code>vote_count &gt;= 10</code> (i.e. the <strong>expert annotator opinions</strong>). </p>\n<ul>\n<li>One model receives 2D spectrograms (created from EEG) as input and uses <code>tiny_vit_21m_512</code> to classify into 6 targets. </li>\n<li>Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image and uses <code>tiny_vit_21m_224</code> to classify into 6 targets. </li>\n<li>And third model receives 1D raw time series waveform and uses <code>EegNet-1D</code> <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\" target=\"_blank\">here</a> to classify into 6 targets.</li>\n</ul>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single1.png\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single2.png\"></p>\n<h1>Pseudo Label Vote&lt;10</h1>\n<p>Using our 3 single models above, we pseudo label the train data with <code>vote&lt;10</code>. We set the new labels as <code>new_label = 0.1 * old_label + 0.3 * model1 + 0.3 * model2 + 0.3 * model3</code>.</p>\n<h1>Build One Mega Model and Train with Vote&gt;=10 plus Pseudo Vote&lt;10</h1>\n<p>We load our 3 pretrained models above and remove the 3 heads. We combine them into one model, concatenate the final embeddings and add 1 new head. Then we train with 2x copies of <code>vote&gt;=10</code> data concatenated with 1x copy of <code>vote&lt;10 with pseudo</code> data. (Note the 2D spectrogram model above actually trains with both 10 sec spectrograms and 50 spectrograms in one image as depicted below):</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/mega_model.png\"></p>\n<h1>Single Model Performance</h1>\n<h2>==&gt; 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉</h2>\n<h1>More Details</h1>\n<p>To successfully accomplish the above and achieve <strong>single model Gold Medal</strong> Hooray!, there are many small details. Here are some:</p>\n<ul>\n<li><p>Preprocess all raw waveform with MNE library </p>\n<pre><code>     mne.filter import filter_data, notch_filter\n    sample = data.T[[0,4,5,6, 11,15,16,17, 0,1,2,3, 11,12,13,14]]\\\n             - data.T[[4,5,6,7, 15,16,17,18, 1,2,3,7, 12,13,14,18]]\n    sample = notch_filter(sample.astype(), 200, 60, =32, =)\n    sample = filter_data(sample.astype(), 200, 0.5, 40, =32, =) \n    sample = np.clip(sample,-500,500)\n    sample = np.nan_to_num(sample, =0)\n</code></pre></li>\n<li><p>Use all 107k rows of train data (instead of only 1 row per eeg_id).</p></li>\n<li><p>Read all eeg parquets into CPU RAM and create spectrograms inside dataloader with <code>torchaudio.transforms.MelSpectrogram</code> and crop with <code>eeg_label_offset_seconds</code>.</p>\n<pre><code>import torchaudio.transforms as T\n\nmake_spec1 = T.MelSpectrogram\n\nmake_spec2 = T.MelSpectrogram\n</code></pre></li>\n<li><p>Each epoch train with 1 sample from each unique <code>eeg_id</code>. But randomly select 1 sample from all <code>eeg_id</code> samples with <code>train = train.loc[train.eeg_id == EEG_ID].sample(1)</code>.</p></li>\n<li><p>Apply data augmentations to raw waveform (not spectrogram) inside dataloader before creating spectrogram.</p></li>\n<li><p>Use data augmentation \"brain-right-left flip\", \"brain-temporal-parasagittal-flip\", \"waveform-crop-rescale\", \"waveform-invert\", \"frequency-dropout\"</p></li>\n<li><p>Use <code>tiny_vit_21m_512</code> for 2D spectrograms and <code>tiny_vit_21m_224</code> for 2D wave plots.</p></li>\n<li><p>Use Nischay EegNet-1D <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\" target=\"_blank\">here</a> for 1D raw wave data model.</p></li>\n<li><p>Tune LR and learning schedule of each model individually using my discussion post <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/488083\" target=\"_blank\">here</a></p></li>\n<li><p>Cross validate with train <code>vote&gt;=10</code> and observe perfect correlation between CV and LB.</p></li>\n</ul>\n<h1>Inference Code Published</h1>\n<p>My single model inference notebook is published <a href=\"https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24\" target=\"_blank\">here</a>. This single model achieves <strong>Gold 14th place</strong>. My final submission is an ensemble of this model with some other models and achieves <strong>Gold 8th place</strong>.</p>\n<p>The code describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!</p>",
  "messages": [
    {
      "id": 2744038,
      "postDate": "2024-04-09T18:08:31.937Z",
      "content": "<h1>Thank You Kaggle and Hosts!</h1>\n<p>Very fun competition! I love competitions with opportunities to be creative, explore, and make discoveries! In the past 3 months, I enthusiastically conducted 1000+ experiments!</p>\n<p>I also love competitions where a single model can achieve Gold medal (instead of a large ensemble).</p>\n<p>Thank you Kagglers for all the great discussion shares and notebook shares!</p>\n<h1>Understand Test Data - What is Competition Task?</h1>\n<p>The task in this competition is to <strong>predict annotator opinions</strong>. By reading the host's paper <a href=\"https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf\" target=\"_blank\">here</a>, we learn that there are <strong>119 overall annotators</strong> in train data and <strong>20 expert annotators</strong> in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%. </p>\n<p>Wow, this is a big difference therefore we need to train our models with <strong>expert opinions</strong> instead of <strong>overall opinions</strong>. We can locate the expert annotators in the train data with the filter <code>expert annotators = train.loc[train.vote_count&gt;=10]</code>.</p>\n<h1>Which Features Do Annotators See?</h1>\n<p>From the paper we see that the annotators see 2D spectrograms and 2D plots of waveforms. So let's input 2D spectrograms and 2D plots of waveforms into our models. (Also inputting 1D raw waveform helped too).</p>\n<h1>Train 3 Separate Models with Vote&gt;=10</h1>\n<p>First we train 3 separate models using train data with <code>vote_count &gt;= 10</code> (i.e. the <strong>expert annotator opinions</strong>). </p>\n<ul>\n<li>One model receives 2D spectrograms (created from EEG) as input and uses <code>tiny_vit_21m_512</code> to classify into 6 targets. </li>\n<li>Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image and uses <code>tiny_vit_21m_224</code> to classify into 6 targets. </li>\n<li>And third model receives 1D raw time series waveform and uses <code>EegNet-1D</code> <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\" target=\"_blank\">here</a> to classify into 6 targets.</li>\n</ul>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single1.png\"></p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single2.png\"></p>\n<h1>Pseudo Label Vote&lt;10</h1>\n<p>Using our 3 single models above, we pseudo label the train data with <code>vote&lt;10</code>. We set the new labels as <code>new_label = 0.1 * old_label + 0.3 * model1 + 0.3 * model2 + 0.3 * model3</code>.</p>\n<h1>Build One Mega Model and Train with Vote&gt;=10 plus Pseudo Vote&lt;10</h1>\n<p>We load our 3 pretrained models above and remove the 3 heads. We combine them into one model, concatenate the final embeddings and add 1 new head. Then we train with 2x copies of <code>vote&gt;=10</code> data concatenated with 1x copy of <code>vote&lt;10 with pseudo</code> data. (Note the 2D spectrogram model above actually trains with both 10 sec spectrograms and 50 spectrograms in one image as depicted below):</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/mega_model.png\"></p>\n<h1>Single Model Performance</h1>\n<h2>==&gt; 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉</h2>\n<h1>More Details</h1>\n<p>To successfully accomplish the above and achieve <strong>single model Gold Medal</strong> Hooray!, there are many small details. Here are some:</p>\n<ul>\n<li><p>Preprocess all raw waveform with MNE library </p>\n<pre><code>     mne.filter import filter_data, notch_filter\n    sample = data.T[[0,4,5,6, 11,15,16,17, 0,1,2,3, 11,12,13,14]]\\\n             - data.T[[4,5,6,7, 15,16,17,18, 1,2,3,7, 12,13,14,18]]\n    sample = notch_filter(sample.astype(), 200, 60, =32, =)\n    sample = filter_data(sample.astype(), 200, 0.5, 40, =32, =) \n    sample = np.clip(sample,-500,500)\n    sample = np.nan_to_num(sample, =0)\n</code></pre></li>\n<li><p>Use all 107k rows of train data (instead of only 1 row per eeg_id).</p></li>\n<li><p>Read all eeg parquets into CPU RAM and create spectrograms inside dataloader with <code>torchaudio.transforms.MelSpectrogram</code> and crop with <code>eeg_label_offset_seconds</code>.</p>\n<pre><code>import torchaudio.transforms as T\n\nmake_spec1 = T.MelSpectrogram\n\nmake_spec2 = T.MelSpectrogram\n</code></pre></li>\n<li><p>Each epoch train with 1 sample from each unique <code>eeg_id</code>. But randomly select 1 sample from all <code>eeg_id</code> samples with <code>train = train.loc[train.eeg_id == EEG_ID].sample(1)</code>.</p></li>\n<li><p>Apply data augmentations to raw waveform (not spectrogram) inside dataloader before creating spectrogram.</p></li>\n<li><p>Use data augmentation \"brain-right-left flip\", \"brain-temporal-parasagittal-flip\", \"waveform-crop-rescale\", \"waveform-invert\", \"frequency-dropout\"</p></li>\n<li><p>Use <code>tiny_vit_21m_512</code> for 2D spectrograms and <code>tiny_vit_21m_224</code> for 2D wave plots.</p></li>\n<li><p>Use Nischay EegNet-1D <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\" target=\"_blank\">here</a> for 1D raw wave data model.</p></li>\n<li><p>Tune LR and learning schedule of each model individually using my discussion post <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/488083\" target=\"_blank\">here</a></p></li>\n<li><p>Cross validate with train <code>vote&gt;=10</code> and observe perfect correlation between CV and LB.</p></li>\n</ul>\n<h1>Inference Code Published</h1>\n<p>My single model inference notebook is published <a href=\"https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24\" target=\"_blank\">here</a>. This single model achieves <strong>Gold 14th place</strong>. My final submission is an ensemble of this model with some other models and achieves <strong>Gold 8th place</strong>.</p>\n<p>The code describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!</p>",
      "rawMarkdown": "# Thank You Kaggle and Hosts!\nVery fun competition! I love competitions with opportunities to be creative, explore, and make discoveries! In the past 3 months, I enthusiastically conducted 1000+ experiments!\n\nI also love competitions where a single model can achieve Gold medal (instead of a large ensemble).\n\nThank you Kagglers for all the great discussion shares and notebook shares!\n\n# Understand Test Data - What is Competition Task?\nThe task in this competition is to **predict annotator opinions**. By reading the host's paper [here][1], we learn that there are **119 overall annotators** in train data and **20 expert annotators** in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%. \n\nWow, this is a big difference therefore we need to train our models with **expert opinions** instead of **overall opinions**. We can locate the expert annotators in the train data with the filter `expert annotators = train.loc[train.vote_count>=10]`.\n\n# Which Features Do Annotators See?\nFrom the paper we see that the annotators see 2D spectrograms and 2D plots of waveforms. So let's input 2D spectrograms and 2D plots of waveforms into our models. (Also inputting 1D raw waveform helped too).\n\n# Train 3 Separate Models with Vote>=10\nFirst we train 3 separate models using train data with `vote_count >= 10` (i.e. the **expert annotator opinions**). \n* One model receives 2D spectrograms (created from EEG) as input and uses `tiny_vit_21m_512` to classify into 6 targets. \n* Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image and uses `tiny_vit_21m_224` to classify into 6 targets. \n* And third model receives 1D raw time series waveform and uses `EegNet-1D` [here][2] to classify into 6 targets.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single1.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single2.png)\n\n# Pseudo Label Vote<10\nUsing our 3 single models above, we pseudo label the train data with `vote<10`. We set the new labels as `new_label = 0.1 * old_label + 0.3 * model1 + 0.3 * model2 + 0.3 * model3`.\n\n# Build One Mega Model and Train with Vote>=10 plus Pseudo Vote<10\nWe load our 3 pretrained models above and remove the 3 heads. We combine them into one model, concatenate the final embeddings and add 1 new head. Then we train with 2x copies of `vote>=10` data concatenated with 1x copy of `vote<10 with pseudo` data. (Note the 2D spectrogram model above actually trains with both 10 sec spectrograms and 50 spectrograms in one image as depicted below):\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/mega_model.png)\n\n# Single Model Performance\n## ==> 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉\n\n# More Details\nTo successfully accomplish the above and achieve **single model Gold Medal** Hooray!, there are many small details. Here are some:\n* Preprocess all raw waveform with MNE library \n\n            from mne.filter import filter_data, notch_filter\n            sample = data.T[[0,4,5,6, 11,15,16,17, 0,1,2,3, 11,12,13,14]]\\\n                     - data.T[[4,5,6,7, 15,16,17,18, 1,2,3,7, 12,13,14,18]]\n            sample = notch_filter(sample.astype('float64'), 200, 60, n_jobs=32, verbose='ERROR')\n            sample = filter_data(sample.astype('float64'), 200, 0.5, 40, n_jobs=32, verbose='ERROR') \n            sample = np.clip(sample,-500,500)\n            sample = np.nan_to_num(sample, nan=0)\n\n* Use all 107k rows of train data (instead of only 1 row per eeg_id).\n* Read all eeg parquets into CPU RAM and create spectrograms inside dataloader with `torchaudio.transforms.MelSpectrogram` and crop with `eeg_label_offset_seconds`.\n\n        import torchaudio.transforms as T\n        # 50 SECOND SPEC\n        make_spec1 = T.MelSpectrogram(n_fft=2048, win_length=1280, hop_length=19,\n                                        f_min=0,f_max=20,sample_rate=200,n_mels=64).to(device)\n        # MIDDLE 10 SECOND SPEC\n        make_spec2 = T.MelSpectrogram(n_fft=1280, win_length=32, hop_length=4,\n                                        f_min=0,f_max=20,sample_rate=200,n_mels=64).to(device)\n\n* Each epoch train with 1 sample from each unique `eeg_id`. But randomly select 1 sample from all `eeg_id` samples with `train = train.loc[train.eeg_id == EEG_ID].sample(1)`.\n\n* Apply data augmentations to raw waveform (not spectrogram) inside dataloader before creating spectrogram.\n* Use data augmentation \"brain-right-left flip\", \"brain-temporal-parasagittal-flip\", \"waveform-crop-rescale\", \"waveform-invert\", \"frequency-dropout\"\n\n* Use `tiny_vit_21m_512` for 2D spectrograms and `tiny_vit_21m_224` for 2D wave plots.\n* Use Nischay EegNet-1D [here][2] for 1D raw wave data model.\n* Tune LR and learning schedule of each model individually using my discussion post [here][3]\n* Cross validate with train `vote>=10` and observe perfect correlation between CV and LB.\n\n# Inference Code Published\nMy single model inference notebook is published [here][4]. This single model achieves **Gold 14th place**. My final submission is an ensemble of this model with some other models and achieves **Gold 8th place**.\n\nThe code describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!\n\n[1]: https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf \n[2]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\n[3]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/488083\n[4]: https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24",
      "votes": 159
    },
    {
      "id": 2744045,
      "postDate": "2024-04-09T18:15:33.080Z",
      "content": "<p>Using plots is an intriguing idea. I never saw it before.</p>\n<p>Glad you ended in top 10 despite sharing so much. Congrats!</p>",
      "rawMarkdown": "Using plots is an intriguing idea. I never saw it before.\n\nGlad you ended in top 10 despite sharing so much. Congrats!\n",
      "votes": 11,
      "replies": [
        {
          "id": 2744054,
          "postDate": "2024-04-09T18:21:59.520Z",
          "content": "<p>Thanks. I thought about using plots a month ago but didn't try it because it seemed silly. I tried it in the last week and it boost CV LB <code>0.01</code>. I'm surprised that it works.</p>",
          "rawMarkdown": "Thanks. I thought about using plots a month ago but didn't try it because it seemed silly. I tried it in the last week and it boost CV LB `0.01`. I'm surprised that it works.",
          "votes": 5,
          "replies": [
            {
              "id": 2744095,
              "postDate": "2024-04-09T18:42:51.900Z",
              "content": "<p>Solution 20th also uses plots! I tagged you there.</p>",
              "rawMarkdown": "Solution 20th also uses plots! I tagged you there.\n",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2747902,
      "postDate": "2024-04-12T06:28:13.390Z",
      "content": "<p>Hi Chris! 🎉<strong>Congratulations</strong> on your gold medal, and thank you for your selfless sharing.</p>\n<p>As a novice, I am now confused about choosing PyTorch or TensorFlow. Many baselines were produced during the beginning of competition based on TensorFlow (including your <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter</a>, <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">EfficientNetB0 Starter</a>), but in the end most gold medals' solutions used PyTorch. I wanna ask if there is any reason?</p>\n<p>It seems that PyTorch has a more complete API for implementing many custom aspects: </p>\n<ul>\n<li>lr schedule (Cosine annealing is manually implemented in the <em>Train Scheduler</em> section of your <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43#Train-Scheduler\" target=\"_blank\">EfficientNetB0 Starter</a>, but it can be simply implemented through <code>torch.optim.lr_scheduler.CosineAnnealingLR</code> in PyTorch)</li>\n<li>Adding sample weights (PyTorch is easier than the TF)</li>\n<li>Use <code>entmax</code> instead of <code>softmax</code> (referring to the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560\" target=\"_blank\">1st place solution</a>, I did not find a direct way to implement it in TF)</li>\n<li>And many more…</li>\n</ul>\n<p>In this competition, I also spent much time solving the problem of different TF versions. Version 2.13 in Kaggle Notebook and my local version 2.15 are incompatible with <em>.h5</em> weight files🥲. </p>\n<p>To sum up, TF seems far less useful than PyTorch. Is this the current situation or is it just my bias haha. Based on this, can you give any suggestions on the use of these two tools?</p>",
      "rawMarkdown": "Hi Chris! 🎉**Congratulations** on your gold medal, and thank you for your selfless sharing.\n\nAs a novice, I am now confused about choosing PyTorch or TensorFlow. Many baselines were produced during the beginning of competition based on TensorFlow (including your [WaveNet Starter](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52), [EfficientNetB0 Starter](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43)), but in the end most gold medals' solutions used PyTorch. I wanna ask if there is any reason?\n\nIt seems that PyTorch has a more complete API for implementing many custom aspects: \n\n- lr schedule (Cosine annealing is manually implemented in the *Train Scheduler* section of your [EfficientNetB0 Starter](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43#Train-Scheduler), but it can be simply implemented through `torch.optim.lr_scheduler.CosineAnnealingLR` in PyTorch)\n- Adding sample weights (PyTorch is easier than the TF)\n- Use `entmax` instead of `softmax` (referring to the [1st place solution](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560), I did not find a direct way to implement it in TF)\n- And many more...\n\nIn this competition, I also spent much time solving the problem of different TF versions. Version 2.13 in Kaggle Notebook and my local version 2.15 are incompatible with *.h5* weight files🥲. \n\nTo sum up, TF seems far less useful than PyTorch. Is this the current situation or is it just my bias haha. Based on this, can you give any suggestions on the use of these two tools?",
      "votes": 3,
      "replies": [
        {
          "id": 2748718,
          "postDate": "2024-04-12T14:50:18.047Z",
          "content": "<p>Hi. I still like to use both TensorFlow and PyTorch. My final ensemble (which improves my single model 14th place Gold to 8th place Gold) has models from both libraries. The main reason that I often use PyTorch is because we have access to over 1000+ different backbones. To display all possible backbones use</p>\n<pre><code>import timm\nm = timm()\n     ,k  (m):\n    (,k)\n</code></pre>\n<p>This prints out over 1000+ models, and we can try each one in PyTorch! It's like being a kid in a candy store! :-) With TensorFlow we cannot access all these models.</p>\n<p>(For example using <code>tiny_vit_21m_512</code> gave <code>+0.02</code> or more boost compared with <code>EffientNetB0</code> but I could not find this model for TensorFlow. In TensorFlow I was able to find and use <code>ConvNext</code> <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/convnext\" target=\"_blank\">here</a> and that did better than EfficientNet but not as good as <code>TinyVIT</code>.)</p>",
          "rawMarkdown": "Hi. I still like to use both TensorFlow and PyTorch. My final ensemble (which improves my single model 14th place Gold to 8th place Gold) has models from both libraries. The main reason that I often use PyTorch is because we have access to over 1000+ different backbones. To display all possible backbones use\n\n    import timm\n    m = timm.list_models()\n        for i,k in enumerate(m):\n        print(i,k)\n\nThis prints out over 1000+ models, and we can try each one in PyTorch! It's like being a kid in a candy store! :-) With TensorFlow we cannot access all these models.\n\n(For example using `tiny_vit_21m_512` gave `+0.02` or more boost compared with `EffientNetB0` but I could not find this model for TensorFlow. In TensorFlow I was able to find and use `ConvNext` [here][1] and that did better than EfficientNet but not as good as `TinyVIT`.)\n\n[1]: https://www.tensorflow.org/api_docs/python/tf/keras/applications/convnext",
          "votes": 4
        }
      ]
    },
    {
      "id": 2744226,
      "postDate": "2024-04-09T19:59:24.500Z",
      "content": "<p>\"Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image\" 🤯 mind blown</p>",
      "rawMarkdown": "\"Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image\" 🤯 mind blown",
      "votes": 3,
      "replies": [
        {
          "id": 2744913,
          "postDate": "2024-04-10T07:54:38.703Z",
          "content": "<p>Same here. Congratulations Chris. you should get another gold medal just for sharing alone. I have to admit that I only joined this competition because you were in it and I knew that I would learn a lot! </p>",
          "rawMarkdown": "Same here. Congratulations Chris. you should get another gold medal just for sharing alone. I have to admit that I only joined this competition because you were in it and I knew that I would learn a lot! ",
          "votes": 4,
          "replies": [
            {
              "id": 2751435,
              "postDate": "2024-04-14T08:42:51.957Z",
              "content": "<p>Thank you Dennis and Nymfree!</p>",
              "rawMarkdown": "Thank you Dennis and Nymfree!"
            }
          ]
        }
      ]
    },
    {
      "id": 2745920,
      "postDate": "2024-04-10T23:27:25.713Z",
      "content": "<p><strong>UPDATE</strong> I published my single model inference notebook <a href=\"https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24\" target=\"_blank\">here</a>. It achieves 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉 It describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!</p>",
      "rawMarkdown": "**UPDATE** I published my single model inference notebook [here][1]. It achieves 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉 It describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24",
      "votes": 4
    },
    {
      "id": 2744085,
      "postDate": "2024-04-09T18:39:12.913Z",
      "content": "<p>Hi Chris, Congratulations!  I do have to say that my brian has exploded :  from this \"The task in this competition is to predict annotator opinions. By reading the host's paper here, we learn that there are 119 overall annotators in train data and 20 expert annotators in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%.</p>\n<p>Wow, this is a big difference therefore we need to train our models with expert opinions instead of overall opinions. We can locate the expert annotators in the train data with the filter expert annotators = train.loc[train.vote_count&gt;=10].\"</p>\n<p>I tried reading the original paper and I totally missed that there was a difference amongst annotators - I thought it was all radiology fellows.  Furthermore, I had no idea that the experts were those with a vote count of more than 10.  Why isn't this on the main page ?  Is this what all Kaggle competitions are like? </p>",
      "rawMarkdown": "Hi Chris, Congratulations!  I do have to say that my brian has exploded :  from this \"The task in this competition is to predict annotator opinions. By reading the host's paper here, we learn that there are 119 overall annotators in train data and 20 expert annotators in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%.\n\nWow, this is a big difference therefore we need to train our models with expert opinions instead of overall opinions. We can locate the expert annotators in the train data with the filter expert annotators = train.loc[train.vote_count>=10].\"\n\nI tried reading the original paper and I totally missed that there was a difference amongst annotators - I thought it was all radiology fellows.  Furthermore, I had no idea that the experts were those with a vote count of more than 10.  Why isn't this on the main page ?  Is this what all Kaggle competitions are like? ",
      "votes": 4,
      "replies": [
        {
          "id": 2744092,
          "postDate": "2024-04-09T18:42:05.470Z",
          "content": "<p>Personally, this is my biggest complaint about Kaggle. They never tell us what the real task is. In other words, the real task is about predicting the test data but they never tell us what the test data is.</p>\n<p>In every competition it is our first task to discover what the test data is and hence what is our real task.</p>",
          "rawMarkdown": "Personally, this is my biggest complaint about Kaggle. They never tell us what the real task is. In other words, the real task is about predicting the test data but they never tell us what the test data is.\n\nIn every competition it is our first task to discover what the test data is and hence what is our real task.",
          "votes": 6,
          "replies": [
            {
              "id": 2744275,
              "postDate": "2024-04-09T20:16:43.360Z",
              "content": "<p>Let's look from the host point of view. What they want is to evaluate ML pipeline quality. Kaggle makes that evaluation through the use of a test dataset. Host designed the test dataset that will provide them what they believe is a good estimate of ML pipelines quality.</p>\n<p>Disclosing how the test set is constructed would defeat the purpose.  Reverse engineering test data creation defeats the purpose.</p>\n<p>TL;DR we have two conflicting goals at play: host goal is to get an unbiased evaluation of models and pipelines, kagglers goal is to get the best possible score. Hiding test data is a way to align these two goals.</p>",
              "rawMarkdown": "Let's look from the host point of view. What they want is to evaluate ML pipeline quality. Kaggle makes that evaluation through the use of a test dataset. Host designed the test dataset that will provide them what they believe is a good estimate of ML pipelines quality.\n\n\nDisclosing how the test set is constructed would defeat the purpose.  Reverse engineering test data creation defeats the purpose.\n\nTL;DR we have two conflicting goals at play: host goal is to get an unbiased evaluation of models and pipelines, kagglers goal is to get the best possible score. Hiding test data is a way to align these two goals.\n\n\n\n",
              "votes": 6
            },
            {
              "id": 2744359,
              "postDate": "2024-04-09T21:00:00.613Z",
              "content": "<p>In the Coursera couses I took, Andrew Ng emphasized the the test and validation/dev set have to come from the same distribution.  This makes sense to me from a theoretical perspective as well.  So we were all spending hours fitting to the wrong target (as Andrew Ng describes the validation set) and the results could have been even better.   I guess I should have asked in the discussions why the CV scores were different from LB scores  (!!) but this was my first competition.   It's also here (<a href=\"https://cs230.stanford.edu/blog/split/#:~:text=It%20is%20important%20to%20choose,to%20get%20in%20the%20future.)\" target=\"_blank\">https://cs230.stanford.edu/blog/split/#:~:text=It%20is%20important%20to%20choose,to%20get%20in%20the%20future.)</a></p>",
              "rawMarkdown": " In the Coursera couses I took, Andrew Ng emphasized the the test and validation/dev set have to come from the same distribution.  This makes sense to me from a theoretical perspective as well.  So we were all spending hours fitting to the wrong target (as Andrew Ng describes the validation set) and the results could have been even better.   I guess I should have asked in the discussions why the CV scores were different from LB scores  (!!) but this was my first competition.   It's also here (https://cs230.stanford.edu/blog/split/#:~:text=It%20is%20important%20to%20choose,to%20get%20in%20the%20future.)",
              "votes": 2
            }
          ]
        },
        {
          "id": 2744419,
          "postDate": "2024-04-09T22:42:39.847Z",
          "content": "<p>Agree, actually though I splitted data by nvotes durting the game, but the same as CPMP I tought the novtes &lt; 10 dataset just means data distribtuion diff not by mean of lower quality…</p>",
          "rawMarkdown": "Agree, actually though I splitted data by nvotes durting the game, but the same as CPMP I tought the novtes < 10 dataset just means data distribtuion diff not by mean of lower quality...",
          "votes": 2
        }
      ]
    },
    {
      "id": 2796609,
      "postDate": "2024-05-06T10:29:54.577Z",
      "content": "<p>Thank you for your sharings throughout the competition.</p>",
      "rawMarkdown": "Thank you for your sharings throughout the competition.",
      "votes": 1
    },
    {
      "id": 2782804,
      "postDate": "2024-04-29T13:06:57.963Z",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi, thanks a lot for this post, I highly appreciate the work you have done here.<br>\nI have the following question.<br>\nAs I understand from this post <code>tiny_vit_21m_224</code><br>\nuses that 10seconds data as an input, but it is already pretrained model. How can I find that model that only uses 10seconds data and train separately ? </p>",
      "rawMarkdown": "@cdeotte Hi, thanks a lot for this post, I highly appreciate the work you have done here.\nI have the following question.\nAs I understand from this post `tiny_vit_21m_224 `\nuses that 10seconds data as an input, but it is already pretrained model. How can I find that model that only uses 10seconds data and train separately ? ",
      "votes": 1,
      "replies": [
        {
          "id": 2782828,
          "postDate": "2024-04-29T13:13:26.023Z",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18779914%2F962bd0f42b665d27da4f3ee3c10dc906%2FScreenshot%202024-04-29%20170855.png?generation=1714396404499433&amp;alt=media\"></p>",
          "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18779914%2F962bd0f42b665d27da4f3ee3c10dc906%2FScreenshot%202024-04-29%20170855.png?generation=1714396404499433&alt=media)",
          "votes": 1,
          "replies": [
            {
              "id": 2783110,
              "postDate": "2024-04-29T15:22:57.157Z",
              "content": "<p>We download the <strong>pretrained</strong> <code>tiny_vit_21m_224</code> model from Timm. Then we <strong>finetune</strong> it by itself in stage 1. Then we use it to make pseudo label. Then we <strong>finetune</strong> it again as part of the Mega model in stage 2.</p>",
              "rawMarkdown": "We download the **pretrained** `tiny_vit_21m_224` model from Timm. Then we **finetune** it by itself in stage 1. Then we use it to make pseudo label. Then we **finetune** it again as part of the Mega model in stage 2."
            },
            {
              "id": 2783119,
              "postDate": "2024-04-29T15:28:06.593Z",
              "content": "<p>ok, got it, I want to create an additional model, I want to fine tune this model based on tiny_vit_21m_224, and the difference is that this model is gonna be trained on different slice of eeg, not the central 10seconds, then I want to combine those 4 models into one and see the results, So am i in right direction ? could you please a supervise here a little :) I want to use the exact same model as  the central 10 seconds data is inputted to, but trained with different dataset.</p>",
              "rawMarkdown": "ok, got it, I want to create an additional model, I want to fine tune this model based on tiny_vit_21m_224, and the difference is that this model is gonna be trained on different slice of eeg, not the central 10seconds, then I want to combine those 4 models into one and see the results, So am i in right direction ? could you please a supervise here a little :) I want to use the exact same model as  the central 10 seconds data is inputted to, but trained with different dataset.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2758405,
      "postDate": "2024-04-18T06:36:22.543Z",
      "content": "<p>I ran into some issues where data enhancement became a bottleneck for the training network. I used some data enhancement methods to increase network generalization, which resulted in my cpu running at 100% while my gpu was only using 10%. How can data enhancement also be accelerated using Gpus?</p>",
      "rawMarkdown": "I ran into some issues where data enhancement became a bottleneck for the training network. I used some data enhancement methods to increase network generalization, which resulted in my cpu running at 100% while my gpu was only using 10%. How can data enhancement also be accelerated using Gpus?",
      "votes": 1
    },
    {
      "id": 2755002,
      "postDate": "2024-04-16T10:23:17.137Z",
      "content": "<p>As described comment <a href=\"https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs/comments#2643122\" target=\"_blank\">here</a>, the eeg net 1D model seems to have a GRU layer which treats the batch dimension as sequence dimension because the layer is instantiated with the default setting batch_first=False.<br>\nSo there are interactions between samples in batches and I'm confused to interpret this model architecture.<br>\nThis seemed bug for me at first glance but not modified here.</p>\n<p>Here is the result of the one mega model from the <a href=\"https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24\" target=\"_blank\">solution notebook</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F13f2be8355bd04d98a13c68e9ea04ecf%2F2024-04-16%2018.40.01.png?generation=1713262210701006&amp;alt=media\" alt=\"one sample\"></p>\n<p>To test there are interactions with samples in a batch, I duplicated the test record and got slightly different result from above.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F7eb4de8dfdff629a3c7f70fc1a3d1b0e%2F2024-04-16%2018.47.36.png?generation=1713262237564915&amp;alt=media\" alt=\"two samples\"></p>\n<p>Next I set the parameter batch_first=True, and get the same output for the duplicated row.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2Ff169e7e6215eecfb278b927d51b72eb7%2F2024-04-16%2019.06.12.png?generation=1713262278157549&amp;alt=media\" alt=\"two samples batch first\"></p>\n<p>I wonder there is an intent behind not setting batch_first=True or this is just a bug.</p>",
      "rawMarkdown": "As described comment [here](https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs/comments#2643122), the eeg net 1D model seems to have a GRU layer which treats the batch dimension as sequence dimension because the layer is instantiated with the default setting batch_first=False.\nSo there are interactions between samples in batches and I'm confused to interpret this model architecture.\nThis seemed bug for me at first glance but not modified here.\n\nHere is the result of the one mega model from the [solution notebook](https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24)\n\n![one sample](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F13f2be8355bd04d98a13c68e9ea04ecf%2F2024-04-16%2018.40.01.png?generation=1713262210701006&alt=media)\n\nTo test there are interactions with samples in a batch, I duplicated the test record and got slightly different result from above.\n\n![two samples](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F7eb4de8dfdff629a3c7f70fc1a3d1b0e%2F2024-04-16%2018.47.36.png?generation=1713262237564915&alt=media)\n\nNext I set the parameter batch_first=True, and get the same output for the duplicated row.\n\n![two samples batch first](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2Ff169e7e6215eecfb278b927d51b72eb7%2F2024-04-16%2019.06.12.png?generation=1713262278157549&alt=media)\n\nI wonder there is an intent behind not setting batch_first=True or this is just a bug.",
      "votes": 1,
      "replies": [
        {
          "id": 2756130,
          "postDate": "2024-04-16T22:16:33.503Z",
          "content": "<p>Interesting. I was not aware of this. I will need to think about this and perform some tests to fully understand this. Thanks for sharing.</p>",
          "rawMarkdown": "Interesting. I was not aware of this. I will need to think about this and perform some tests to fully understand this. Thanks for sharing.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2754948,
      "postDate": "2024-04-16T09:31:33.417Z",
      "content": "<p>thank you for sharing this, very helpful</p>",
      "rawMarkdown": "thank you for sharing this, very helpful",
      "votes": 1
    },
    {
      "id": 2752717,
      "postDate": "2024-04-15T05:44:30.063Z",
      "content": "<p>Thank you for your sharings.<br>\nI have a question about \"Then we train with 2x copies of vote&gt;=10 data concatenated with 1x copy of vote&lt;10 with pseudo data.\".<br>\nwhat is the meaning of copies ?? <br>\nI don't understand what exactly it is.</p>",
      "rawMarkdown": "Thank you for your sharings.\nI have a question about \"Then we train with 2x copies of vote>=10 data concatenated with 1x copy of vote<10 with pseudo data.\".\nwhat is the meaning of copies ?? \nI don't understand what exactly it is.",
      "votes": 1,
      "replies": [
        {
          "id": 2752749,
          "postDate": "2024-04-15T06:14:36.583Z",
          "content": "<p>We give <code>vote&gt;=10</code> weight=2 and <code>vote&lt;10</code> weight=1 like this. (And for vote&gt;=10 we use original labels and vote&lt;10 we use pseudo labels):</p>\n<pre><code>train = pd()\ntrain = train(axis=)\ndf1 = train\ndf2 = train\ntrain_data = pd(,axis=)\n</code></pre>",
          "rawMarkdown": "We give `vote>=10` weight=2 and `vote<10` weight=1 like this. (And for vote>=10 we use original labels and vote<10 we use pseudo labels):\n\n    train = pd.read_csv('train.csv')\n    train['ct'] = train.iloc[:,-6:].sum(axis=1)\n    df1 = train.loc[train.ct >= 10]\n    df2 = train.loc[train.ct < 10]\n    train_data = pd.concat([df1, df1, df2],axis=0)",
          "votes": 1
        }
      ]
    },
    {
      "id": 2751453,
      "postDate": "2024-04-14T09:02:05.577Z",
      "content": "<p>thank you for sharing this, very helpful </p>",
      "rawMarkdown": "thank you for sharing this, very helpful ",
      "votes": 1
    },
    {
      "id": 2748258,
      "postDate": "2024-04-12T10:44:36.703Z",
      "content": "<p>Very helpful</p>",
      "rawMarkdown": "Very helpful",
      "votes": 1
    },
    {
      "id": 2746120,
      "postDate": "2024-04-11T04:03:59.177Z",
      "content": "<p>the graphs is looking little dangerous and attracting me to know more!!<br>\ncongratulations for the gold <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !!</p>",
      "rawMarkdown": "the graphs is looking little dangerous and attracting me to know more!!\ncongratulations for the gold @cdeotte !!",
      "votes": 1,
      "replies": [
        {
          "id": 2747588,
          "postDate": "2024-04-12T01:29:25.420Z",
          "content": "<p>haha thank you. </p>",
          "rawMarkdown": "haha thank you. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2745937,
      "postDate": "2024-04-10T23:55:37.827Z",
      "content": "<p>Very helpful!!</p>",
      "rawMarkdown": "Very helpful!!",
      "votes": 1
    },
    {
      "id": 2745732,
      "postDate": "2024-04-10T19:09:06.373Z",
      "content": "<p>Thanks for sharing this solution it will be very helpful to us.</p>",
      "rawMarkdown": "Thanks for sharing this solution it will be very helpful to us.",
      "votes": 1
    },
    {
      "id": 2745654,
      "postDate": "2024-04-10T18:25:07.657Z",
      "content": "<p>Great and helpful work👌🏻 congrats☘️</p>",
      "rawMarkdown": "Great and helpful work👌🏻 congrats☘️",
      "votes": 1
    },
    {
      "id": 2744907,
      "postDate": "2024-04-10T07:46:52.740Z",
      "content": "<p>Great work on HMS - Harmful Brain Activity Classification! Your approach was clear and effective. Have you considered [specific suggestion] for even better results?</p>",
      "rawMarkdown": "Great work on HMS - Harmful Brain Activity Classification! Your approach was clear and effective. Have you considered [specific suggestion] for even better results?",
      "votes": 1
    },
    {
      "id": 2744788,
      "postDate": "2024-04-10T06:04:12.897Z",
      "content": "<p>Hi Chris! Congrats with the gold, and thank you for your early stages posts and especially 50s spectrogramms, those were very handy! </p>\n<p>I also used 1D -&gt; 2D plot with the same motivation to show the model ~ same stuff as were shown to annotators, and I also was surprised it works, a nice trick indeed.</p>\n<p>Have you measured the pseudolabels impact? I found PLs to improve the public score just a bit (on ~0.001 order) on multiple occasions and decided to trust it, but did not do any hold out validation so could not tell more precisely, that would be interesting to see if it helped you more.</p>",
      "rawMarkdown": "Hi Chris! Congrats with the gold, and thank you for your early stages posts and especially 50s spectrogramms, those were very handy! \n\nI also used 1D -> 2D plot with the same motivation to show the model ~ same stuff as were shown to annotators, and I also was surprised it works, a nice trick indeed.\n\nHave you measured the pseudolabels impact? I found PLs to improve the public score just a bit (on ~0.001 order) on multiple occasions and decided to trust it, but did not do any hold out validation so could not tell more precisely, that would be interesting to see if it helped you more.",
      "votes": 1,
      "replies": [
        {
          "id": 2744817,
          "postDate": "2024-04-10T06:27:37.483Z",
          "content": "<p>Pseudo labeling gave large boosts like <code>+0.03</code>. Here is how i performed clean (leakfree) CV. First train 5 KFold models on <code>vote&gt;=10</code>. Then we pseudo label each fold separately. So we take fold 1 model and use it to pseudo label the fold 1 train data <code>vote&lt;10</code>. Next we take the fold 2 model and use it to pseudo label the fold 2 train data, etc etc. (So we have 5 sets of pseudo labels). </p>\n<p>Afterward we train another 5 KFold using the <code>vote&gt;=10</code> data concatenated with <code>vote&lt;10</code> pseudo label data. When training the new fold 1 model we use the fold 1 pseudo labels. When training the new fold 2 model, we use the fold 2 pseudo labels, etc. Lastly we perform a stage 3 where we finetune each fold model using only <code>vote&gt;=10</code>. </p>\n<p>What I describe here is a 3 stage approach. This is a different model than what I describe in my solution post above. We see below that stage 1 cannot score better than 0.31 CV no matter what LR and train schedule we use. However stage 2 (which uses pseudo labels created from stage 1 models) improves the CV score more. And lastly stage 3 improves the CV score even more. This model uses spectrograms and raw waves and has clean CV score 0.27 and LB 0.28. Each stages' LR and train schedule was carefully discovered. Below show the validation KL Divergence loss for fold 1:</p>\n<p>Note that if we don't create a clean (leakfree) KFold we will never know the correct LR and train schedule because a leaky KFold will just prefer the largest LR possible to overfit the leak.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/3-stage.png\"></p>",
          "rawMarkdown": "Pseudo labeling gave large boosts like `+0.03`. Here is how i performed clean (leakfree) CV. First train 5 KFold models on `vote>=10`. Then we pseudo label each fold separately. So we take fold 1 model and use it to pseudo label the fold 1 train data `vote<10`. Next we take the fold 2 model and use it to pseudo label the fold 2 train data, etc etc. (So we have 5 sets of pseudo labels). \n\nAfterward we train another 5 KFold using the `vote>=10` data concatenated with `vote<10` pseudo label data. When training the new fold 1 model we use the fold 1 pseudo labels. When training the new fold 2 model, we use the fold 2 pseudo labels, etc. Lastly we perform a stage 3 where we finetune each fold model using only `vote>=10`. \n\nWhat I describe here is a 3 stage approach. This is a different model than what I describe in my solution post above. We see below that stage 1 cannot score better than 0.31 CV no matter what LR and train schedule we use. However stage 2 (which uses pseudo labels created from stage 1 models) improves the CV score more. And lastly stage 3 improves the CV score even more. This model uses spectrograms and raw waves and has clean CV score 0.27 and LB 0.28. Each stages' LR and train schedule was carefully discovered. Below show the validation KL Divergence loss for fold 1:\n\nNote that if we don't create a clean (leakfree) KFold we will never know the correct LR and train schedule because a leaky KFold will just prefer the largest LR possible to overfit the leak.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/3-stage.png)",
          "votes": 5
        },
        {
          "id": 2744826,
          "postDate": "2024-04-10T06:34:11.353Z",
          "content": "<p>Also note that pseudo labels allows us to perform knowledge distillation which gives even more boost than simple pseudo labels.</p>\n<p>With knowledge distillation, we use pseudo labels from one model (like a 1D EegNet) and transfer that intelligence into another type of model (like 2D VIT) using pseudo labels. Or we can create an ensemble of lots of models, average all the pseudo labels, and transfer that combined knowledge into a new different model.</p>\n<p>The reason this helps is because when a teacher trains a student, then the student adds is own \"personality\" to the knowledge. Afterward if we ensemble the new student with the old teacher, the pair will be smarter than the both the teacher and the student alone.</p>\n<p>I used this technique here in Brain comp and I used it previously in AMEX comp <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641\" target=\"_blank\">here</a>. In AMEX comp, I transferred the knowledge of LGBM model into NN model. The NN model received the learnings from LGBM and added its own personality thus creating a new representation of the knowledge.</p>",
          "rawMarkdown": "Also note that pseudo labels allows us to perform knowledge distillation which gives even more boost than simple pseudo labels.\n\nWith knowledge distillation, we use pseudo labels from one model (like a 1D EegNet) and transfer that intelligence into another type of model (like 2D VIT) using pseudo labels. Or we can create an ensemble of lots of models, average all the pseudo labels, and transfer that combined knowledge into a new different model.\n\nThe reason this helps is because when a teacher trains a student, then the student adds is own \"personality\" to the knowledge. Afterward if we ensemble the new student with the old teacher, the pair will be smarter than the both the teacher and the student alone.\n\nI used this technique here in Brain comp and I used it previously in AMEX comp [here][1]. In AMEX comp, I transferred the knowledge of LGBM model into NN model. The NN model received the learnings from LGBM and added its own personality thus creating a new representation of the knowledge.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641",
          "votes": 8,
          "replies": [
            {
              "id": 2752352,
              "postDate": "2024-04-14T23:26:22.273Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thank you for your detailed explanation about pseudo labeling!<br>\nI have a question about this part. Why is it necessary to pseudo label each fold separately?<br>\nWhen you train 5 KFold models on vote &gt;= 10, it seems that vote &lt; 10 records don't belong to any of the folds and you can create pseudo labels for these records using all 5 KFold models (average of all models or something). There seems no leak and pseudo labels could be slightly accurate.<br>\nAm I missing something? <br>\nI would appreciate if you could answer to my question. Thank you!</p>\n<blockquote>\n  <p>Here is how i performed clean (leakfree) CV. First train 5 KFold models on vote&gt;=10. Then we pseudo label each fold separately. So we take fold 1 model and use it to pseudo label the fold 1 train data vote&lt;10. Next we take the fold 2 model and use it to pseudo label the fold 2 train data, etc etc. (So we have 5 sets of pseudo labels).</p>\n</blockquote>",
              "rawMarkdown": "\nHi @cdeotte , thank you for your detailed explanation about pseudo labeling!\n\n\nI have a question about this part. Why is it necessary to pseudo label each fold separately?\nWhen you train 5 KFold models on vote >= 10, it seems that vote < 10 records don't belong to any of the folds and you can create pseudo labels for these records using all 5 KFold models (average of all models or something). There seems no leak and pseudo labels could be slightly accurate.\nAm I missing something? \nI would appreciate if you could answer to my question. Thank you!\n\n> Here is how i performed clean (leakfree) CV. First train 5 KFold models on vote>=10. Then we pseudo label each fold separately. So we take fold 1 model and use it to pseudo label the fold 1 train data vote<10. Next we take the fold 2 model and use it to pseudo label the fold 2 train data, etc etc. (So we have 5 sets of pseudo labels).\n\n\n",
              "votes": 2
            },
            {
              "id": 2752500,
              "postDate": "2024-04-15T02:22:56.820Z",
              "content": "<p><a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a> For stage 1, let's imagine that we have Folds A, B, C, D, E. When we train <code>fold 1</code> model, we train with Folds B, C, D, E and validation with Fold A. Now let's pseudo label ANY DATA with this <code>fold 1</code> model and call the result <code>pseudo 1</code>. The process of making pseudo label transfers the ground truth of Folds B, C, D, E into the ANY DATA (via knowledge distillation principle even though ANY DATA is different data than Folds B, C, D, E data). So <code>pseudo 1</code> contains the ground truth from Folds B, C, D, E (from knowledge transfer).</p>\n<p>Imagine now that we are training stage 2 <code>fold 2</code> model, we train with Folds A, C, D, E and validation with Fold B. If we add <code>pseudo 1</code> to the training of stage 2 <code>fold 2</code> model, then we leak. Because <code>pseudo 1</code> contains the ground truth of Folds B, C, D, E and afterward we want to validate <code>fold 2</code> model with Fold B.</p>\n<p>In conclusion, whenever we do pseudo label or knowledge transfer with KFold models, we need to keep K separate pseudo label data for usage in stage 2 to prevent leak. If we have a leak then the stage 2 models will perform poorly on LB because as their <strong>fake</strong> leaked CV score improve, their <strong>real</strong> leak-free CV score will be getting worse without us knowing.</p>",
              "rawMarkdown": "@nynyny67 For stage 1, let's imagine that we have Folds A, B, C, D, E. When we train `fold 1` model, we train with Folds B, C, D, E and validation with Fold A. Now let's pseudo label ANY DATA with this `fold 1` model and call the result `pseudo 1`. The process of making pseudo label transfers the ground truth of Folds B, C, D, E into the ANY DATA (via knowledge distillation principle even though ANY DATA is different data than Folds B, C, D, E data). So `pseudo 1` contains the ground truth from Folds B, C, D, E (from knowledge transfer).\n\nImagine now that we are training stage 2 `fold 2` model, we train with Folds A, C, D, E and validation with Fold B. If we add `pseudo 1` to the training of stage 2 `fold 2` model, then we leak. Because `pseudo 1` contains the ground truth of Folds B, C, D, E and afterward we want to validate `fold 2` model with Fold B.\n\nIn conclusion, whenever we do pseudo label or knowledge transfer with KFold models, we need to keep K separate pseudo label data for usage in stage 2 to prevent leak. If we have a leak then the stage 2 models will perform poorly on LB because as their **fake** leaked CV score improve, their **real** leak-free CV score will be getting worse without us knowing.",
              "votes": 5
            },
            {
              "id": 2753326,
              "postDate": "2024-04-15T12:27:56.873Z",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nThank you so much for answering my question!<br>\nI understood why it's necessary to keep K separate pseudo label!<br>\nOnce we use a pseudo label from a model trained with some ground truth information, we can never use the ground truth data to evaluate the models that used the pseudo label to train because it has a leak!<br>\nSo if we create pseudo label from all folds, the model trained with it cannot be evaluated locally.</p>",
              "rawMarkdown": "@cdeotte\nThank you so much for answering my question!\nI understood why it's necessary to keep K separate pseudo label!\nOnce we use a pseudo label from a model trained with some ground truth information, we can never use the ground truth data to evaluate the models that used the pseudo label to train because it has a leak!\nSo if we create pseudo label from all folds, the model trained with it cannot be evaluated locally.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2744412,
      "postDate": "2024-04-09T22:36:51.187Z",
      "content": "<p>I believe ViT here refers to visual transformers? I had started training one from scratch, should have known better to rely on pre-trained solutions, very nice!</p>",
      "rawMarkdown": "I believe ViT here refers to visual transformers? I had started training one from scratch, should have known better to rely on pre-trained solutions, very nice!",
      "votes": 1,
      "replies": [
        {
          "id": 2744420,
          "postDate": "2024-04-09T22:44:59.817Z",
          "content": "<p>Yes VIT is Vision Transformer. We can print out a full list of the 1000 pretrained models available in timm library with:</p>\n<pre><code>import timm\nm = timm()\n     ,k  (m):\n    (,k)\n</code></pre>\n<p>Then we can try lots of these models to see what works best. For each model we should use our previous LR train schedule but also try <code>LR = LR * 10</code> and <code>LR = LR/10</code> because some of these different models require stronger or weaker learning rates.</p>",
          "rawMarkdown": "Yes VIT is Vision Transformer. We can print out a full list of the 1000 pretrained models available in timm library with:\n\n    import timm\n    m = timm.list_models()\n        for i,k in enumerate(m):\n        print(i,k)\n\nThen we can try lots of these models to see what works best. For each model we should use our previous LR train schedule but also try `LR = LR * 10` and `LR = LR/10` because some of these different models require stronger or weaker learning rates.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2744254,
      "postDate": "2024-04-09T20:08:18.267Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> on achieve 8th Place.</p>",
      "rawMarkdown": "Congrats @cdeotte on achieve 8th Place.",
      "votes": 1,
      "replies": [
        {
          "id": 2745393,
          "postDate": "2024-04-10T15:40:35.260Z",
          "content": "<p>Thank you Adnan</p>",
          "rawMarkdown": "Thank you Adnan",
          "votes": 1
        }
      ]
    },
    {
      "id": 2752510,
      "postDate": "2024-04-15T02:38:09.147Z",
      "content": "<p>Thank you for your sharings throughout the competition.</p>\n<p>Did you use larger models instead of tiny-vit and get worse results, or is there a specific reason for using smaller models?</p>",
      "rawMarkdown": "Thank you for your sharings throughout the competition.\n\nDid you use larger models instead of tiny-vit and get worse results, or is there a specific reason for using smaller models?",
      "votes": 2,
      "replies": [
        {
          "id": 2752513,
          "postDate": "2024-04-15T02:43:18.823Z",
          "content": "<p>Yes. I tried many sized models and always got the best performance with <strong>tiny</strong> vision transformers. Furthermore, <code>tiny-vit</code> is special (among tiny models). It is pretrained with knowledge transfer from larger model into the tiny model.</p>",
          "rawMarkdown": "Yes. I tried many sized models and always got the best performance with **tiny** vision transformers. Furthermore, `tiny-vit` is special (among tiny models). It is pretrained with knowledge transfer from larger model into the tiny model.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2745444,
      "postDate": "2024-04-10T16:28:24.923Z",
      "content": "<p>Thank you for sharing your process selflessly from beginning to end.<br>\nYou are a true benevolent person who is admired by kagglers.<br>\nTHX ! ! </p>",
      "rawMarkdown": "Thank you for sharing your process selflessly from beginning to end.\nYou are a true benevolent person who is admired by kagglers.\nTHX ! ! ",
      "votes": 2,
      "replies": [
        {
          "id": 2749183,
          "postDate": "2024-04-12T22:16:30.300Z",
          "content": "<p>Thank you Warrenkuo!</p>",
          "rawMarkdown": "Thank you Warrenkuo!",
          "votes": 2
        }
      ]
    },
    {
      "id": 2744670,
      "postDate": "2024-04-10T04:19:05.450Z",
      "content": "<p>Congratulations on securing 8th place in this competition. Thanks for sharing your solution details. </p>",
      "rawMarkdown": "Congratulations on securing 8th place in this competition. Thanks for sharing your solution details. ",
      "votes": 2,
      "replies": [
        {
          "id": 2744754,
          "postDate": "2024-04-10T05:19:58.700Z",
          "content": "<p>Thank you.    </p>",
          "rawMarkdown": "Thank you.    ",
          "votes": 1
        }
      ]
    },
    {
      "id": 2744421,
      "postDate": "2024-04-09T22:45:03.073Z",
      "content": "<p>Great work, your sharing early in the game make this game much more interesting and helped so much people to join.<br>\n Could you help share your cv results of your first step 3 models? </p>",
      "rawMarkdown": "Great work, your sharing early in the game make this game much more interesting and helped so much people to join.\n Could you help share your cv results of your first step 3 models? ",
      "votes": 2,
      "replies": [
        {
          "id": 2744436,
          "postDate": "2024-04-09T23:14:59.050Z",
          "content": "<p>Thanks. I enjoyed sharing and engaging in discussions with the community. It did make this competition more challenging.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2D Eeg Plot Image</td>\n<td>0.280</td>\n<td>0.256</td>\n<td>0.317</td>\n</tr>\n<tr>\n<td>2D Eeg Spectrogram Image</td>\n<td>0.281</td>\n<td>0.275</td>\n<td>0.326</td>\n</tr>\n<tr>\n<td>1D Raw Eeg Waveform</td>\n<td>0.303</td>\n<td>0.305</td>\n<td>0.363</td>\n</tr>\n<tr>\n<td>Ordinary Ensemble</td>\n<td>0.253</td>\n<td>???</td>\n<td>???</td>\n</tr>\n<tr>\n<td>One Mega Model</td>\n<td>0.246</td>\n<td>0.242</td>\n<td>0.290</td>\n</tr>\n</tbody>\n</table>\n<p>.<br>\nSurprisingly the 2D Eeg Plot model achieves LB 0.25 by itself wow! We also see using a combined mega model achieves better CV than simple ensemble of 3 models. Note a mega model is like <strong>stacking</strong> we are training a new model to make predictions from the output of 3 previous models.</p>\n<p>Note that these are not my best single modality input models. I have another 1D Raw Eeg Waveform model that achieves CV 0.28 LB 0.28 by modifying the public notebook architecture. And I probably have better versions of the other two single models too. But for my mega model the table above shows which models are utilized for it.</p>",
          "rawMarkdown": "Thanks. I enjoyed sharing and engaging in discussions with the community. It did make this competition more challenging.\n\n| Model | CV | Public LB | Private LB |\n| --- | --- | --- | --- |\n| 2D Eeg Plot Image | 0.280 | 0.256 | 0.317 |\n| 2D Eeg Spectrogram Image | 0.281 | 0.275 | 0.326 |\n| 1D Raw Eeg Waveform | 0.303 | 0.305 | 0.363 |\n| Ordinary Ensemble | 0.253 | ??? | ??? |\n| One Mega Model | 0.246 | 0.242 | 0.290 |\n\n.\nSurprisingly the 2D Eeg Plot model achieves LB 0.25 by itself wow! We also see using a combined mega model achieves better CV than simple ensemble of 3 models. Note a mega model is like **stacking** we are training a new model to make predictions from the output of 3 previous models.\n\nNote that these are not my best single modality input models. I have another 1D Raw Eeg Waveform model that achieves CV 0.28 LB 0.28 by modifying the public notebook architecture. And I probably have better versions of the other two single models too. But for my mega model the table above shows which models are utilized for it.\n",
          "votes": 3
        }
      ]
    },
    {
      "id": 2744344,
      "postDate": "2024-04-09T20:51:57.857Z",
      "content": "<p>Great solution Chris! I also tried to use similar graphs implemented through matplotlib at the very beginning, but gave up on this idea because I couldn't get good results, and all the articles I saw said that this is not a good solution.</p>",
      "rawMarkdown": "Great solution Chris! I also tried to use similar graphs implemented through matplotlib at the very beginning, but gave up on this idea because I couldn't get good results, and all the articles I saw said that this is not a good solution.",
      "votes": 2
    },
    {
      "id": 2744261,
      "postDate": "2024-04-09T20:11:10.023Z",
      "content": "<p>What post processing did you use? You aid you used some in a past yesterday.</p>",
      "rawMarkdown": "What post processing did you use? You aid you used some in a past yesterday.",
      "votes": 2,
      "replies": [
        {
          "id": 2744414,
          "postDate": "2024-04-09T22:40:24.320Z",
          "content": "<p>My final solution is an ensemble of multiple diverse models (and increases my private rank from 15th to 8th). The model shown in this discussion post by itself achieves gold medal and does not require any post process. Similar to you when I used pseudo label for some models I adjusted the target distribution of pseudo labeled train vote&lt;10 with:</p>\n<blockquote>\n  <p>To avoid this distribution shift I sample pseudo labelled data by target, to keep the same distribution as training data. (Quote from CPMP <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492211\" target=\"_blank\">here</a>)</p>\n</blockquote>\n<p>Some of my other models were trained with different subsets (i.e. different target distributions) of train data and consequently don't have the right prediction distribution of targets. </p>\n<p>For example, we can train equally with all <code>vote&lt;10</code> original targets and <code>vote&gt;=10</code> original targets then the model will predict too many <code>Seizures</code> and too many <code>Other</code>. Afterward we correct the predictions with the following post process:</p>\n<pre><code>\npred = submission[TARGETS].copy()\nfor k,v in [(,),(,)]: \n    pred[:,k] = pred[:,k]/(-pred[:,k]) \n    pred[:,k] *= v \n    pred[:,k] = pred[:,k]/(pred[:,k]+) \npred = pred / pred.sum(axis=,keepdims=) \nsubmission[TARGETS] = pred\n</code></pre>",
          "rawMarkdown": "My final solution is an ensemble of multiple diverse models (and increases my private rank from 15th to 8th). The model shown in this discussion post by itself achieves gold medal and does not require any post process. Similar to you when I used pseudo label for some models I adjusted the target distribution of pseudo labeled train vote<10 with:\n>To avoid this distribution shift I sample pseudo labelled data by target, to keep the same distribution as training data. (Quote from CPMP [here][1])\n\nSome of my other models were trained with different subsets (i.e. different target distributions) of train data and consequently don't have the right prediction distribution of targets. \n\nFor example, we can train equally with all `vote<10` original targets and `vote>=10` original targets then the model will predict too many `Seizures` and too many `Other`. Afterward we correct the predictions with the following post process:\n\n    # REDUCE SEIZURE AND OTHER PREDICTIONS\n    pred = submission[TARGETS].copy()\n    for k,v in [(0,0.1),(5,0.3)]: # TARGETS SEIZURE AND OTHER\n        pred[:,k] = pred[:,k]/(1-pred[:,k]) # CONVERT PROB TO ODDS\n        pred[:,k] *= v # BAYESIAN UPDATE\n        pred[:,k] = pred[:,k]/(pred[:,k]+1) # CONVERT ODDS TO PROBS\n    pred = pred / pred.sum(axis=1,keepdims=True) # ADJUST SUM TO ONE\n    submission[TARGETS] = pred\n\n[1]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492211",
          "votes": 5,
          "replies": [
            {
              "id": 2744854,
              "postDate": "2024-04-10T07:00:34.250Z",
              "content": "<p>ok, I did not have the distribution shift issue indeed.</p>",
              "rawMarkdown": "ok, I did not have the distribution shift issue indeed.",
              "votes": 1
            },
            {
              "id": 2753945,
              "postDate": "2024-04-15T18:14:55.937Z",
              "content": "<p>Very Nice!!</p>",
              "rawMarkdown": "Very Nice!!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2824203,
      "postDate": "2024-05-19T16:24:54.180Z",
      "content": "<p>Hi Chris … Congratulations on your goldmedal! </p>",
      "rawMarkdown": "Hi Chris ... Congratulations on your goldmedal! "
    },
    {
      "id": 2782824,
      "postDate": "2024-04-29T13:12:27.030Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 2756059,
      "postDate": "2024-04-16T19:59:23.567Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 2753171,
      "postDate": "2024-04-15T10:36:46.610Z",
      "content": "<p>Thank you for sharing</p>",
      "rawMarkdown": "Thank you for sharing",
      "votes": 1
    },
    {
      "id": 2746716,
      "postDate": "2024-04-11T12:19:20.697Z",
      "content": "<p>Thank you for sharing your process</p>",
      "rawMarkdown": "Thank you for sharing your process",
      "votes": 1
    },
    {
      "id": 2768884,
      "postDate": "2024-04-23T04:40:32.933Z",
      "content": "<p>Thank you for your sharings.</p>",
      "rawMarkdown": "Thank you for your sharings."
    }
  ],
  "comments": [
    {
      "id": 2744045,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-04-09T18:15:33.080000",
      "content": "<p>Using plots is an intriguing idea. I never saw it before.</p>\n<p>Glad you ended in top 10 despite sharing so much. Congrats!</p>",
      "votes": 11,
      "replies": [
        {
          "id": 2744054,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T18:21:59.520000",
          "content": "<p>Thanks. I thought about using plots a month ago but didn't try it because it seemed silly. I tried it in the last week and it boost CV LB <code>0.01</code>. I'm surprised that it works.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 2744095,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-09T18:42:51.900000",
              "content": "<p>Solution 20th also uses plots! I tagged you there.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2747902,
      "author_name": "Fang Zitao",
      "author_url": "",
      "post_date": "2024-04-12T06:28:13.390000",
      "content": "<p>Hi Chris! 🎉<strong>Congratulations</strong> on your gold medal, and thank you for your selfless sharing.</p>\n<p>As a novice, I am now confused about choosing PyTorch or TensorFlow. Many baselines were produced during the beginning of competition based on TensorFlow (including your <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter</a>, <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43\" target=\"_blank\">EfficientNetB0 Starter</a>), but in the end most gold medals' solutions used PyTorch. I wanna ask if there is any reason?</p>\n<p>It seems that PyTorch has a more complete API for implementing many custom aspects: </p>\n<ul>\n<li>lr schedule (Cosine annealing is manually implemented in the <em>Train Scheduler</em> section of your <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43#Train-Scheduler\" target=\"_blank\">EfficientNetB0 Starter</a>, but it can be simply implemented through <code>torch.optim.lr_scheduler.CosineAnnealingLR</code> in PyTorch)</li>\n<li>Adding sample weights (PyTorch is easier than the TF)</li>\n<li>Use <code>entmax</code> instead of <code>softmax</code> (referring to the <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560\" target=\"_blank\">1st place solution</a>, I did not find a direct way to implement it in TF)</li>\n<li>And many more…</li>\n</ul>\n<p>In this competition, I also spent much time solving the problem of different TF versions. Version 2.13 in Kaggle Notebook and my local version 2.15 are incompatible with <em>.h5</em> weight files🥲. </p>\n<p>To sum up, TF seems far less useful than PyTorch. Is this the current situation or is it just my bias haha. Based on this, can you give any suggestions on the use of these two tools?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2748718,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-12T14:50:18.047000",
          "content": "<p>Hi. I still like to use both TensorFlow and PyTorch. My final ensemble (which improves my single model 14th place Gold to 8th place Gold) has models from both libraries. The main reason that I often use PyTorch is because we have access to over 1000+ different backbones. To display all possible backbones use</p>\n<pre><code>import timm\nm = timm()\n     ,k  (m):\n    (,k)\n</code></pre>\n<p>This prints out over 1000+ models, and we can try each one in PyTorch! It's like being a kid in a candy store! :-) With TensorFlow we cannot access all these models.</p>\n<p>(For example using <code>tiny_vit_21m_512</code> gave <code>+0.02</code> or more boost compared with <code>EffientNetB0</code> but I could not find this model for TensorFlow. In TensorFlow I was able to find and use <code>ConvNext</code> <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/convnext\" target=\"_blank\">here</a> and that did better than EfficientNet but not as good as <code>TinyVIT</code>.)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2744226,
      "author_name": "DennisSakva",
      "author_url": "",
      "post_date": "2024-04-09T19:59:24.500000",
      "content": "<p>\"Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image\" 🤯 mind blown</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2744913,
          "author_name": "nymfree",
          "author_url": "",
          "post_date": "2024-04-10T07:54:38.703000",
          "content": "<p>Same here. Congratulations Chris. you should get another gold medal just for sharing alone. I have to admit that I only joined this competition because you were in it and I knew that I would learn a lot! </p>",
          "votes": 4,
          "replies": [
            {
              "id": 2751435,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-04-14T08:42:51.957000",
              "content": "<p>Thank you Dennis and Nymfree!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2745920,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2024-04-10T23:27:25.713000",
      "content": "<p><strong>UPDATE</strong> I published my single model inference notebook <a href=\"https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24\" target=\"_blank\">here</a>. It achieves 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉 It describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 2744085,
      "author_name": "Idith Haber",
      "author_url": "",
      "post_date": "2024-04-09T18:39:12.913000",
      "content": "<p>Hi Chris, Congratulations!  I do have to say that my brian has exploded :  from this \"The task in this competition is to predict annotator opinions. By reading the host's paper here, we learn that there are 119 overall annotators in train data and 20 expert annotators in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%.</p>\n<p>Wow, this is a big difference therefore we need to train our models with expert opinions instead of overall opinions. We can locate the expert annotators in the train data with the filter expert annotators = train.loc[train.vote_count&gt;=10].\"</p>\n<p>I tried reading the original paper and I totally missed that there was a difference amongst annotators - I thought it was all radiology fellows.  Furthermore, I had no idea that the experts were those with a vote count of more than 10.  Why isn't this on the main page ?  Is this what all Kaggle competitions are like? </p>",
      "votes": 4,
      "replies": [
        {
          "id": 2744092,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T18:42:05.470000",
          "content": "<p>Personally, this is my biggest complaint about Kaggle. They never tell us what the real task is. In other words, the real task is about predicting the test data but they never tell us what the test data is.</p>\n<p>In every competition it is our first task to discover what the test data is and hence what is our real task.</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2744275,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-09T20:16:43.360000",
              "content": "<p>Let's look from the host point of view. What they want is to evaluate ML pipeline quality. Kaggle makes that evaluation through the use of a test dataset. Host designed the test dataset that will provide them what they believe is a good estimate of ML pipelines quality.</p>\n<p>Disclosing how the test set is constructed would defeat the purpose.  Reverse engineering test data creation defeats the purpose.</p>\n<p>TL;DR we have two conflicting goals at play: host goal is to get an unbiased evaluation of models and pipelines, kagglers goal is to get the best possible score. Hiding test data is a way to align these two goals.</p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2744359,
              "author_name": "Idith Haber",
              "author_url": "",
              "post_date": "2024-04-09T21:00:00.613000",
              "content": "<p>In the Coursera couses I took, Andrew Ng emphasized the the test and validation/dev set have to come from the same distribution.  This makes sense to me from a theoretical perspective as well.  So we were all spending hours fitting to the wrong target (as Andrew Ng describes the validation set) and the results could have been even better.   I guess I should have asked in the discussions why the CV scores were different from LB scores  (!!) but this was my first competition.   It's also here (<a href=\"https://cs230.stanford.edu/blog/split/#:~:text=It%20is%20important%20to%20choose,to%20get%20in%20the%20future.)\" target=\"_blank\">https://cs230.stanford.edu/blog/split/#:~:text=It%20is%20important%20to%20choose,to%20get%20in%20the%20future.)</a></p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 2744419,
          "author_name": "gezi",
          "author_url": "",
          "post_date": "2024-04-09T22:42:39.847000",
          "content": "<p>Agree, actually though I splitted data by nvotes durting the game, but the same as CPMP I tought the novtes &lt; 10 dataset just means data distribtuion diff not by mean of lower quality…</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2796609,
      "author_name": "Superman",
      "author_url": "",
      "post_date": "2024-05-06T10:29:54.577000",
      "content": "<p>Thank you for your sharings throughout the competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2782804,
      "author_name": "Datevik",
      "author_url": "",
      "post_date": "2024-04-29T13:06:57.963000",
      "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> Hi, thanks a lot for this post, I highly appreciate the work you have done here.<br>\nI have the following question.<br>\nAs I understand from this post <code>tiny_vit_21m_224</code><br>\nuses that 10seconds data as an input, but it is already pretrained model. How can I find that model that only uses 10seconds data and train separately ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2782828,
          "author_name": "Datevik",
          "author_url": "",
          "post_date": "2024-04-29T13:13:26.023000",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F18779914%2F962bd0f42b665d27da4f3ee3c10dc906%2FScreenshot%202024-04-29%20170855.png?generation=1714396404499433&amp;alt=media\"></p>",
          "votes": 1,
          "replies": [
            {
              "id": 2783110,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-04-29T15:22:57.157000",
              "content": "<p>We download the <strong>pretrained</strong> <code>tiny_vit_21m_224</code> model from Timm. Then we <strong>finetune</strong> it by itself in stage 1. Then we use it to make pseudo label. Then we <strong>finetune</strong> it again as part of the Mega model in stage 2.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2783119,
              "author_name": "Datevik",
              "author_url": "",
              "post_date": "2024-04-29T15:28:06.593000",
              "content": "<p>ok, got it, I want to create an additional model, I want to fine tune this model based on tiny_vit_21m_224, and the difference is that this model is gonna be trained on different slice of eeg, not the central 10seconds, then I want to combine those 4 models into one and see the results, So am i in right direction ? could you please a supervise here a little :) I want to use the exact same model as  the central 10 seconds data is inputted to, but trained with different dataset.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2758405,
      "author_name": "franticXu",
      "author_url": "",
      "post_date": "2024-04-18T06:36:22.543000",
      "content": "<p>I ran into some issues where data enhancement became a bottleneck for the training network. I used some data enhancement methods to increase network generalization, which resulted in my cpu running at 100% while my gpu was only using 10%. How can data enhancement also be accelerated using Gpus?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2755002,
      "author_name": "nynyny67",
      "author_url": "",
      "post_date": "2024-04-16T10:23:17.137000",
      "content": "<p>As described comment <a href=\"https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs/comments#2643122\" target=\"_blank\">here</a>, the eeg net 1D model seems to have a GRU layer which treats the batch dimension as sequence dimension because the layer is instantiated with the default setting batch_first=False.<br>\nSo there are interactions between samples in batches and I'm confused to interpret this model architecture.<br>\nThis seemed bug for me at first glance but not modified here.</p>\n<p>Here is the result of the one mega model from the <a href=\"https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24\" target=\"_blank\">solution notebook</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F13f2be8355bd04d98a13c68e9ea04ecf%2F2024-04-16%2018.40.01.png?generation=1713262210701006&amp;alt=media\" alt=\"one sample\"></p>\n<p>To test there are interactions with samples in a batch, I duplicated the test record and got slightly different result from above.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F7eb4de8dfdff629a3c7f70fc1a3d1b0e%2F2024-04-16%2018.47.36.png?generation=1713262237564915&amp;alt=media\" alt=\"two samples\"></p>\n<p>Next I set the parameter batch_first=True, and get the same output for the duplicated row.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2Ff169e7e6215eecfb278b927d51b72eb7%2F2024-04-16%2019.06.12.png?generation=1713262278157549&amp;alt=media\" alt=\"two samples batch first\"></p>\n<p>I wonder there is an intent behind not setting batch_first=True or this is just a bug.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2756130,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-16T22:16:33.503000",
          "content": "<p>Interesting. I was not aware of this. I will need to think about this and perform some tests to fully understand this. Thanks for sharing.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2754948,
      "author_name": "Chenghao Peng1013",
      "author_url": "",
      "post_date": "2024-04-16T09:31:33.417000",
      "content": "<p>thank you for sharing this, very helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2752717,
      "author_name": "SeongwookLee",
      "author_url": "",
      "post_date": "2024-04-15T05:44:30.063000",
      "content": "<p>Thank you for your sharings.<br>\nI have a question about \"Then we train with 2x copies of vote&gt;=10 data concatenated with 1x copy of vote&lt;10 with pseudo data.\".<br>\nwhat is the meaning of copies ?? <br>\nI don't understand what exactly it is.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2752749,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-15T06:14:36.583000",
          "content": "<p>We give <code>vote&gt;=10</code> weight=2 and <code>vote&lt;10</code> weight=1 like this. (And for vote&gt;=10 we use original labels and vote&lt;10 we use pseudo labels):</p>\n<pre><code>train = pd()\ntrain = train(axis=)\ndf1 = train\ndf2 = train\ntrain_data = pd(,axis=)\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2751453,
      "author_name": "Ki Kwang Min",
      "author_url": "",
      "post_date": "2024-04-14T09:02:05.577000",
      "content": "<p>thank you for sharing this, very helpful </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2748258,
      "author_name": "Aditya Srivastava",
      "author_url": "",
      "post_date": "2024-04-12T10:44:36.703000",
      "content": "<p>Very helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2746120,
      "author_name": "Anmol gupta",
      "author_url": "",
      "post_date": "2024-04-11T04:03:59.177000",
      "content": "<p>the graphs is looking little dangerous and attracting me to know more!!<br>\ncongratulations for the gold <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2747588,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-12T01:29:25.420000",
          "content": "<p>haha thank you. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2745937,
      "author_name": "Hoonsuk Bae",
      "author_url": "",
      "post_date": "2024-04-10T23:55:37.827000",
      "content": "<p>Very helpful!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2745732,
      "author_name": "Aaditya Porwal",
      "author_url": "",
      "post_date": "2024-04-10T19:09:06.373000",
      "content": "<p>Thanks for sharing this solution it will be very helpful to us.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2745654,
      "author_name": "Atefeh Mirnaseri",
      "author_url": "",
      "post_date": "2024-04-10T18:25:07.657000",
      "content": "<p>Great and helpful work👌🏻 congrats☘️</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2744907,
      "author_name": "Yugandhar Dasari",
      "author_url": "",
      "post_date": "2024-04-10T07:46:52.740000",
      "content": "<p>Great work on HMS - Harmful Brain Activity Classification! Your approach was clear and effective. Have you considered [specific suggestion] for even better results?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2744788,
      "author_name": "Mikhail Kotyushev",
      "author_url": "",
      "post_date": "2024-04-10T06:04:12.897000",
      "content": "<p>Hi Chris! Congrats with the gold, and thank you for your early stages posts and especially 50s spectrogramms, those were very handy! </p>\n<p>I also used 1D -&gt; 2D plot with the same motivation to show the model ~ same stuff as were shown to annotators, and I also was surprised it works, a nice trick indeed.</p>\n<p>Have you measured the pseudolabels impact? I found PLs to improve the public score just a bit (on ~0.001 order) on multiple occasions and decided to trust it, but did not do any hold out validation so could not tell more precisely, that would be interesting to see if it helped you more.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2744817,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-10T06:27:37.483000",
          "content": "<p>Pseudo labeling gave large boosts like <code>+0.03</code>. Here is how i performed clean (leakfree) CV. First train 5 KFold models on <code>vote&gt;=10</code>. Then we pseudo label each fold separately. So we take fold 1 model and use it to pseudo label the fold 1 train data <code>vote&lt;10</code>. Next we take the fold 2 model and use it to pseudo label the fold 2 train data, etc etc. (So we have 5 sets of pseudo labels). </p>\n<p>Afterward we train another 5 KFold using the <code>vote&gt;=10</code> data concatenated with <code>vote&lt;10</code> pseudo label data. When training the new fold 1 model we use the fold 1 pseudo labels. When training the new fold 2 model, we use the fold 2 pseudo labels, etc. Lastly we perform a stage 3 where we finetune each fold model using only <code>vote&gt;=10</code>. </p>\n<p>What I describe here is a 3 stage approach. This is a different model than what I describe in my solution post above. We see below that stage 1 cannot score better than 0.31 CV no matter what LR and train schedule we use. However stage 2 (which uses pseudo labels created from stage 1 models) improves the CV score more. And lastly stage 3 improves the CV score even more. This model uses spectrograms and raw waves and has clean CV score 0.27 and LB 0.28. Each stages' LR and train schedule was carefully discovered. Below show the validation KL Divergence loss for fold 1:</p>\n<p>Note that if we don't create a clean (leakfree) KFold we will never know the correct LR and train schedule because a leaky KFold will just prefer the largest LR possible to overfit the leak.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/3-stage.png\"></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 2744826,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-10T06:34:11.353000",
          "content": "<p>Also note that pseudo labels allows us to perform knowledge distillation which gives even more boost than simple pseudo labels.</p>\n<p>With knowledge distillation, we use pseudo labels from one model (like a 1D EegNet) and transfer that intelligence into another type of model (like 2D VIT) using pseudo labels. Or we can create an ensemble of lots of models, average all the pseudo labels, and transfer that combined knowledge into a new different model.</p>\n<p>The reason this helps is because when a teacher trains a student, then the student adds is own \"personality\" to the knowledge. Afterward if we ensemble the new student with the old teacher, the pair will be smarter than the both the teacher and the student alone.</p>\n<p>I used this technique here in Brain comp and I used it previously in AMEX comp <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/347641\" target=\"_blank\">here</a>. In AMEX comp, I transferred the knowledge of LGBM model into NN model. The NN model received the learnings from LGBM and added its own personality thus creating a new representation of the knowledge.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 2752352,
              "author_name": "nynyny67",
              "author_url": "",
              "post_date": "2024-04-14T23:26:22.273000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , thank you for your detailed explanation about pseudo labeling!<br>\nI have a question about this part. Why is it necessary to pseudo label each fold separately?<br>\nWhen you train 5 KFold models on vote &gt;= 10, it seems that vote &lt; 10 records don't belong to any of the folds and you can create pseudo labels for these records using all 5 KFold models (average of all models or something). There seems no leak and pseudo labels could be slightly accurate.<br>\nAm I missing something? <br>\nI would appreciate if you could answer to my question. Thank you!</p>\n<blockquote>\n  <p>Here is how i performed clean (leakfree) CV. First train 5 KFold models on vote&gt;=10. Then we pseudo label each fold separately. So we take fold 1 model and use it to pseudo label the fold 1 train data vote&lt;10. Next we take the fold 2 model and use it to pseudo label the fold 2 train data, etc etc. (So we have 5 sets of pseudo labels).</p>\n</blockquote>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2752500,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2024-04-15T02:22:56.820000",
              "content": "<p><a href=\"https://www.kaggle.com/nynyny67\" target=\"_blank\">@nynyny67</a> For stage 1, let's imagine that we have Folds A, B, C, D, E. When we train <code>fold 1</code> model, we train with Folds B, C, D, E and validation with Fold A. Now let's pseudo label ANY DATA with this <code>fold 1</code> model and call the result <code>pseudo 1</code>. The process of making pseudo label transfers the ground truth of Folds B, C, D, E into the ANY DATA (via knowledge distillation principle even though ANY DATA is different data than Folds B, C, D, E data). So <code>pseudo 1</code> contains the ground truth from Folds B, C, D, E (from knowledge transfer).</p>\n<p>Imagine now that we are training stage 2 <code>fold 2</code> model, we train with Folds A, C, D, E and validation with Fold B. If we add <code>pseudo 1</code> to the training of stage 2 <code>fold 2</code> model, then we leak. Because <code>pseudo 1</code> contains the ground truth of Folds B, C, D, E and afterward we want to validate <code>fold 2</code> model with Fold B.</p>\n<p>In conclusion, whenever we do pseudo label or knowledge transfer with KFold models, we need to keep K separate pseudo label data for usage in stage 2 to prevent leak. If we have a leak then the stage 2 models will perform poorly on LB because as their <strong>fake</strong> leaked CV score improve, their <strong>real</strong> leak-free CV score will be getting worse without us knowing.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2753326,
              "author_name": "nynyny67",
              "author_url": "",
              "post_date": "2024-04-15T12:27:56.873000",
              "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a><br>\nThank you so much for answering my question!<br>\nI understood why it's necessary to keep K separate pseudo label!<br>\nOnce we use a pseudo label from a model trained with some ground truth information, we can never use the ground truth data to evaluate the models that used the pseudo label to train because it has a leak!<br>\nSo if we create pseudo label from all folds, the model trained with it cannot be evaluated locally.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2744412,
      "author_name": "Yijie Xu",
      "author_url": "",
      "post_date": "2024-04-09T22:36:51.187000",
      "content": "<p>I believe ViT here refers to visual transformers? I had started training one from scratch, should have known better to rely on pre-trained solutions, very nice!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2744420,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T22:44:59.817000",
          "content": "<p>Yes VIT is Vision Transformer. We can print out a full list of the 1000 pretrained models available in timm library with:</p>\n<pre><code>import timm\nm = timm()\n     ,k  (m):\n    (,k)\n</code></pre>\n<p>Then we can try lots of these models to see what works best. For each model we should use our previous LR train schedule but also try <code>LR = LR * 10</code> and <code>LR = LR/10</code> because some of these different models require stronger or weaker learning rates.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2744254,
      "author_name": "Adnan Alaref",
      "author_url": "",
      "post_date": "2024-04-09T20:08:18.267000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> on achieve 8th Place.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2745393,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-10T15:40:35.260000",
          "content": "<p>Thank you Adnan</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2752510,
      "author_name": "Bartu Hacılar",
      "author_url": "",
      "post_date": "2024-04-15T02:38:09.147000",
      "content": "<p>Thank you for your sharings throughout the competition.</p>\n<p>Did you use larger models instead of tiny-vit and get worse results, or is there a specific reason for using smaller models?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2752513,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-15T02:43:18.823000",
          "content": "<p>Yes. I tried many sized models and always got the best performance with <strong>tiny</strong> vision transformers. Furthermore, <code>tiny-vit</code> is special (among tiny models). It is pretrained with knowledge transfer from larger model into the tiny model.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2745444,
      "author_name": "warrenkuo",
      "author_url": "",
      "post_date": "2024-04-10T16:28:24.923000",
      "content": "<p>Thank you for sharing your process selflessly from beginning to end.<br>\nYou are a true benevolent person who is admired by kagglers.<br>\nTHX ! ! </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2749183,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-12T22:16:30.300000",
          "content": "<p>Thank you Warrenkuo!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2744670,
      "author_name": "C R Suthikshn Kumar",
      "author_url": "",
      "post_date": "2024-04-10T04:19:05.450000",
      "content": "<p>Congratulations on securing 8th place in this competition. Thanks for sharing your solution details. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2744754,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-10T05:19:58.700000",
          "content": "<p>Thank you.    </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2744421,
      "author_name": "gezi",
      "author_url": "",
      "post_date": "2024-04-09T22:45:03.073000",
      "content": "<p>Great work, your sharing early in the game make this game much more interesting and helped so much people to join.<br>\n Could you help share your cv results of your first step 3 models? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 2744436,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T23:14:59.050000",
          "content": "<p>Thanks. I enjoyed sharing and engaging in discussions with the community. It did make this competition more challenging.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>CV</th>\n<th>Public LB</th>\n<th>Private LB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2D Eeg Plot Image</td>\n<td>0.280</td>\n<td>0.256</td>\n<td>0.317</td>\n</tr>\n<tr>\n<td>2D Eeg Spectrogram Image</td>\n<td>0.281</td>\n<td>0.275</td>\n<td>0.326</td>\n</tr>\n<tr>\n<td>1D Raw Eeg Waveform</td>\n<td>0.303</td>\n<td>0.305</td>\n<td>0.363</td>\n</tr>\n<tr>\n<td>Ordinary Ensemble</td>\n<td>0.253</td>\n<td>???</td>\n<td>???</td>\n</tr>\n<tr>\n<td>One Mega Model</td>\n<td>0.246</td>\n<td>0.242</td>\n<td>0.290</td>\n</tr>\n</tbody>\n</table>\n<p>.<br>\nSurprisingly the 2D Eeg Plot model achieves LB 0.25 by itself wow! We also see using a combined mega model achieves better CV than simple ensemble of 3 models. Note a mega model is like <strong>stacking</strong> we are training a new model to make predictions from the output of 3 previous models.</p>\n<p>Note that these are not my best single modality input models. I have another 1D Raw Eeg Waveform model that achieves CV 0.28 LB 0.28 by modifying the public notebook architecture. And I probably have better versions of the other two single models too. But for my mega model the table above shows which models are utilized for it.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2744344,
      "author_name": "Andrij",
      "author_url": "",
      "post_date": "2024-04-09T20:51:57.857000",
      "content": "<p>Great solution Chris! I also tried to use similar graphs implemented through matplotlib at the very beginning, but gave up on this idea because I couldn't get good results, and all the articles I saw said that this is not a good solution.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2744261,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2024-04-09T20:11:10.023000",
      "content": "<p>What post processing did you use? You aid you used some in a past yesterday.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2744414,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2024-04-09T22:40:24.320000",
          "content": "<p>My final solution is an ensemble of multiple diverse models (and increases my private rank from 15th to 8th). The model shown in this discussion post by itself achieves gold medal and does not require any post process. Similar to you when I used pseudo label for some models I adjusted the target distribution of pseudo labeled train vote&lt;10 with:</p>\n<blockquote>\n  <p>To avoid this distribution shift I sample pseudo labelled data by target, to keep the same distribution as training data. (Quote from CPMP <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492211\" target=\"_blank\">here</a>)</p>\n</blockquote>\n<p>Some of my other models were trained with different subsets (i.e. different target distributions) of train data and consequently don't have the right prediction distribution of targets. </p>\n<p>For example, we can train equally with all <code>vote&lt;10</code> original targets and <code>vote&gt;=10</code> original targets then the model will predict too many <code>Seizures</code> and too many <code>Other</code>. Afterward we correct the predictions with the following post process:</p>\n<pre><code>\npred = submission[TARGETS].copy()\nfor k,v in [(,),(,)]: \n    pred[:,k] = pred[:,k]/(-pred[:,k]) \n    pred[:,k] *= v \n    pred[:,k] = pred[:,k]/(pred[:,k]+) \npred = pred / pred.sum(axis=,keepdims=) \nsubmission[TARGETS] = pred\n</code></pre>",
          "votes": 5,
          "replies": [
            {
              "id": 2744854,
              "author_name": "CPMP",
              "author_url": "",
              "post_date": "2024-04-10T07:00:34.250000",
              "content": "<p>ok, I did not have the distribution shift issue indeed.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2753945,
              "author_name": "Kazuma143",
              "author_url": "",
              "post_date": "2024-04-15T18:14:55.937000",
              "content": "<p>Very Nice!!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2824203,
      "author_name": "Vaishno Kumar Tiwari",
      "author_url": "",
      "post_date": "2024-05-19T16:24:54.180000",
      "content": "<p>Hi Chris … Congratulations on your goldmedal! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2782824,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-29T13:12:27.030000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2756059,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-16T19:59:23.567000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2753171,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-15T10:36:46.610000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2746716,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-11T12:19:20.697000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2768884,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-23T04:40:32.933000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2744038": "# Thank You Kaggle and Hosts!\nVery fun competition! I love competitions with opportunities to be creative, explore, and make discoveries! In the past 3 months, I enthusiastically conducted 1000+ experiments!\n\nI also love competitions where a single model can achieve Gold medal (instead of a large ensemble).\n\nThank you Kagglers for all the great discussion shares and notebook shares!\n\n# Understand Test Data - What is Competition Task?\nThe task in this competition is to **predict annotator opinions**. By reading the host's paper [here][1], we learn that there are **119 overall annotators** in train data and **20 expert annotators** in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%. \n\nWow, this is a big difference therefore we need to train our models with **expert opinions** instead of **overall opinions**. We can locate the expert annotators in the train data with the filter `expert annotators = train.loc[train.vote_count>=10]`.\n\n# Which Features Do Annotators See?\nFrom the paper we see that the annotators see 2D spectrograms and 2D plots of waveforms. So let's input 2D spectrograms and 2D plots of waveforms into our models. (Also inputting 1D raw waveform helped too).\n\n# Train 3 Separate Models with Vote>=10\nFirst we train 3 separate models using train data with `vote_count >= 10` (i.e. the **expert annotator opinions**). \n* One model receives 2D spectrograms (created from EEG) as input and uses `tiny_vit_21m_512` to classify into 6 targets. \n* Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image and uses `tiny_vit_21m_224` to classify into 6 targets. \n* And third model receives 1D raw time series waveform and uses `EegNet-1D` [here][2] to classify into 6 targets.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single1.png)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/single2.png)\n\n# Pseudo Label Vote<10\nUsing our 3 single models above, we pseudo label the train data with `vote<10`. We set the new labels as `new_label = 0.1 * old_label + 0.3 * model1 + 0.3 * model2 + 0.3 * model3`.\n\n# Build One Mega Model and Train with Vote>=10 plus Pseudo Vote<10\nWe load our 3 pretrained models above and remove the 3 heads. We combine them into one model, concatenate the final embeddings and add 1 new head. Then we train with 2x copies of `vote>=10` data concatenated with 1x copy of `vote<10 with pseudo` data. (Note the 2D spectrogram model above actually trains with both 10 sec spectrograms and 50 spectrograms in one image as depicted below):\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2024/mega_model.png)\n\n# Single Model Performance\n## ==> 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉\n\n# More Details\nTo successfully accomplish the above and achieve **single model Gold Medal** Hooray!, there are many small details. Here are some:\n* Preprocess all raw waveform with MNE library \n\n            from mne.filter import filter_data, notch_filter\n            sample = data.T[[0,4,5,6, 11,15,16,17, 0,1,2,3, 11,12,13,14]]\\\n                     - data.T[[4,5,6,7, 15,16,17,18, 1,2,3,7, 12,13,14,18]]\n            sample = notch_filter(sample.astype('float64'), 200, 60, n_jobs=32, verbose='ERROR')\n            sample = filter_data(sample.astype('float64'), 200, 0.5, 40, n_jobs=32, verbose='ERROR') \n            sample = np.clip(sample,-500,500)\n            sample = np.nan_to_num(sample, nan=0)\n\n* Use all 107k rows of train data (instead of only 1 row per eeg_id).\n* Read all eeg parquets into CPU RAM and create spectrograms inside dataloader with `torchaudio.transforms.MelSpectrogram` and crop with `eeg_label_offset_seconds`.\n\n        import torchaudio.transforms as T\n        # 50 SECOND SPEC\n        make_spec1 = T.MelSpectrogram(n_fft=2048, win_length=1280, hop_length=19,\n                                        f_min=0,f_max=20,sample_rate=200,n_mels=64).to(device)\n        # MIDDLE 10 SECOND SPEC\n        make_spec2 = T.MelSpectrogram(n_fft=1280, win_length=32, hop_length=4,\n                                        f_min=0,f_max=20,sample_rate=200,n_mels=64).to(device)\n\n* Each epoch train with 1 sample from each unique `eeg_id`. But randomly select 1 sample from all `eeg_id` samples with `train = train.loc[train.eeg_id == EEG_ID].sample(1)`.\n\n* Apply data augmentations to raw waveform (not spectrogram) inside dataloader before creating spectrogram.\n* Use data augmentation \"brain-right-left flip\", \"brain-temporal-parasagittal-flip\", \"waveform-crop-rescale\", \"waveform-invert\", \"frequency-dropout\"\n\n* Use `tiny_vit_21m_512` for 2D spectrograms and `tiny_vit_21m_224` for 2D wave plots.\n* Use Nischay EegNet-1D [here][2] for 1D raw wave data model.\n* Tune LR and learning schedule of each model individually using my discussion post [here][3]\n* Cross validate with train `vote>=10` and observe perfect correlation between CV and LB.\n\n# Inference Code Published\nMy single model inference notebook is published [here][4]. This single model achieves **Gold 14th place**. My final submission is an ensemble of this model with some other models and achieves **Gold 8th place**.\n\nThe code describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!\n\n[1]: https://github.com/bdsp-core/IIIC-SPaRCNet/blob/main/IIIC_Classification-Supplemental.pdf \n[2]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/471666\n[3]: https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/488083\n[4]: https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24",
    "2744045": "Using plots is an intriguing idea. I never saw it before.\n\nGlad you ended in top 10 despite sharing so much. Congrats!\n",
    "2747902": "Hi Chris! 🎉**Congratulations** on your gold medal, and thank you for your selfless sharing.\n\nAs a novice, I am now confused about choosing PyTorch or TensorFlow. Many baselines were produced during the beginning of competition based on TensorFlow (including your [WaveNet Starter](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52), [EfficientNetB0 Starter](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43)), but in the end most gold medals' solutions used PyTorch. I wanna ask if there is any reason?\n\nIt seems that PyTorch has a more complete API for implementing many custom aspects: \n\n- lr schedule (Cosine annealing is manually implemented in the *Train Scheduler* section of your [EfficientNetB0 Starter](https://www.kaggle.com/code/cdeotte/efficientnetb0-starter-lb-0-43#Train-Scheduler), but it can be simply implemented through `torch.optim.lr_scheduler.CosineAnnealingLR` in PyTorch)\n- Adding sample weights (PyTorch is easier than the TF)\n- Use `entmax` instead of `softmax` (referring to the [1st place solution](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/492560), I did not find a direct way to implement it in TF)\n- And many more...\n\nIn this competition, I also spent much time solving the problem of different TF versions. Version 2.13 in Kaggle Notebook and my local version 2.15 are incompatible with *.h5* weight files🥲. \n\nTo sum up, TF seems far less useful than PyTorch. Is this the current situation or is it just my bias haha. Based on this, can you give any suggestions on the use of these two tools?",
    "2744226": "\"Second model receives 2D matplotlib plots of the montage waveforms converted into a 2D image\" 🤯 mind blown",
    "2745920": "**UPDATE** I published my single model inference notebook [here][1]. It achieves 🎉 CV = 0.246, Public LB = 0.243, Private LB = 0.290 🎉 It describes the PyTorch model details, the preprocess details, PyTorch data loader details, and data augmentation details. Enjoy!\n\n[1]: https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24",
    "2744085": "Hi Chris, Congratulations!  I do have to say that my brian has exploded :  from this \"The task in this competition is to predict annotator opinions. By reading the host's paper here, we learn that there are 119 overall annotators in train data and 20 expert annotators in test data. The paper shows the average Seizure prediction from 119 annotators is 18.8% whereas the 20 experts are less likely to classify as Seizure and their average is 1.5%.\n\nWow, this is a big difference therefore we need to train our models with expert opinions instead of overall opinions. We can locate the expert annotators in the train data with the filter expert annotators = train.loc[train.vote_count>=10].\"\n\nI tried reading the original paper and I totally missed that there was a difference amongst annotators - I thought it was all radiology fellows.  Furthermore, I had no idea that the experts were those with a vote count of more than 10.  Why isn't this on the main page ?  Is this what all Kaggle competitions are like? ",
    "2796609": "Thank you for your sharings throughout the competition.",
    "2782804": "@cdeotte Hi, thanks a lot for this post, I highly appreciate the work you have done here.\nI have the following question.\nAs I understand from this post `tiny_vit_21m_224 `\nuses that 10seconds data as an input, but it is already pretrained model. How can I find that model that only uses 10seconds data and train separately ? ",
    "2758405": "I ran into some issues where data enhancement became a bottleneck for the training network. I used some data enhancement methods to increase network generalization, which resulted in my cpu running at 100% while my gpu was only using 10%. How can data enhancement also be accelerated using Gpus?",
    "2755002": "As described comment [here](https://www.kaggle.com/code/nischaydnk/lightning-1d-eegnet-training-pipeline-hbs/comments#2643122), the eeg net 1D model seems to have a GRU layer which treats the batch dimension as sequence dimension because the layer is instantiated with the default setting batch_first=False.\nSo there are interactions between samples in batches and I'm confused to interpret this model architecture.\nThis seemed bug for me at first glance but not modified here.\n\nHere is the result of the one mega model from the [solution notebook](https://www.kaggle.com/code/cdeotte/single-model-gold-solution-cv-0-24-lb-0-24)\n\n![one sample](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F13f2be8355bd04d98a13c68e9ea04ecf%2F2024-04-16%2018.40.01.png?generation=1713262210701006&alt=media)\n\nTo test there are interactions with samples in a batch, I duplicated the test record and got slightly different result from above.\n\n![two samples](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2F7eb4de8dfdff629a3c7f70fc1a3d1b0e%2F2024-04-16%2018.47.36.png?generation=1713262237564915&alt=media)\n\nNext I set the parameter batch_first=True, and get the same output for the duplicated row.\n\n![two samples batch first](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8363906%2Ff169e7e6215eecfb278b927d51b72eb7%2F2024-04-16%2019.06.12.png?generation=1713262278157549&alt=media)\n\nI wonder there is an intent behind not setting batch_first=True or this is just a bug.",
    "2754948": "thank you for sharing this, very helpful",
    "2752717": "Thank you for your sharings.\nI have a question about \"Then we train with 2x copies of vote>=10 data concatenated with 1x copy of vote<10 with pseudo data.\".\nwhat is the meaning of copies ?? \nI don't understand what exactly it is.",
    "2751453": "thank you for sharing this, very helpful ",
    "2748258": "Very helpful",
    "2746120": "the graphs is looking little dangerous and attracting me to know more!!\ncongratulations for the gold @cdeotte !!",
    "2745937": "Very helpful!!",
    "2745732": "Thanks for sharing this solution it will be very helpful to us.",
    "2745654": "Great and helpful work👌🏻 congrats☘️",
    "2744907": "Great work on HMS - Harmful Brain Activity Classification! Your approach was clear and effective. Have you considered [specific suggestion] for even better results?",
    "2744788": "Hi Chris! Congrats with the gold, and thank you for your early stages posts and especially 50s spectrogramms, those were very handy! \n\nI also used 1D -> 2D plot with the same motivation to show the model ~ same stuff as were shown to annotators, and I also was surprised it works, a nice trick indeed.\n\nHave you measured the pseudolabels impact? I found PLs to improve the public score just a bit (on ~0.001 order) on multiple occasions and decided to trust it, but did not do any hold out validation so could not tell more precisely, that would be interesting to see if it helped you more.",
    "2744412": "I believe ViT here refers to visual transformers? I had started training one from scratch, should have known better to rely on pre-trained solutions, very nice!",
    "2744254": "Congrats @cdeotte on achieve 8th Place.",
    "2752510": "Thank you for your sharings throughout the competition.\n\nDid you use larger models instead of tiny-vit and get worse results, or is there a specific reason for using smaller models?",
    "2745444": "Thank you for sharing your process selflessly from beginning to end.\nYou are a true benevolent person who is admired by kagglers.\nTHX ! ! ",
    "2744670": "Congratulations on securing 8th place in this competition. Thanks for sharing your solution details. ",
    "2744421": "Great work, your sharing early in the game make this game much more interesting and helped so much people to join.\n Could you help share your cv results of your first step 3 models? ",
    "2744344": "Great solution Chris! I also tried to use similar graphs implemented through matplotlib at the very beginning, but gave up on this idea because I couldn't get good results, and all the articles I saw said that this is not a good solution.",
    "2744261": "What post processing did you use? You aid you used some in a past yesterday.",
    "2824203": "Hi Chris ... Congratulations on your goldmedal! ",
    "2782824": "",
    "2756059": "Thanks for sharing!",
    "2753171": "Thank you for sharing",
    "2746716": "Thank you for sharing your process",
    "2768884": "Thank you for your sharings."
  }
}