{
  "id": 492234,
  "title": "38th solution: 1d models and 1d+2d models",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/27-38th-solution-1d-models-and-1d-2d-models",
  "author_name": "",
  "post_date": "2024-04-10T04:55:21.963Z",
  "votes": 28,
  "comment_count": 3,
  "views": 0,
  "content": "<h1>Summary</h1>\n<ul>\n<li>1d models (public LB ~0.26)</li>\n<li>1d+2d models (public LB ~0.25)</li>\n<li>2stages training</li>\n<li>training 2d models' backbone and 1d+2d models separately.</li>\n</ul>\n<h1>TmT&amp;Ken_Ken_pa Part : <strong>1d Models</strong></h1>\n<p>The 1d model is primarily based on Wavenet and uses the attention pool.<br>\nFirst train on all votes, then train on the part where vote&gt;=10.<br>\nAnd after a nice model is obtained, the pseudo label pre-training is carried out.</p>\n<p><strong>Data Augumentation</strong>: </p>\n<ul>\n<li>Flipping the left and right brain.  Inference with both the flipped and the unflipped brain.</li>\n<li>Use upsampling to replicate samples from fewer classes.</li>\n</ul>\n<h1>BladeRunner Part: <strong>2d Models</strong></h1>\n<p>The 2D spectrogram solution uses a two-stage training.</p>\n<ul>\n<li>Stage 1: training on all data.</li>\n<li>Stage 2: training on only data with vote &gt; 9.\nData preparation and pre-processing\nUsing kaggle spectrogram + EEG spectrogram\nOnly one sample is generated for an eeg_id, and the data is taken right in the middle.\nkaggle_spec shape (512, 256, 1)\neeg_spec shape (512, 256, 1)\ncombine (512, 512, 1)\nEEG  spectrogram generated method: mel spectrogram<ul>\n<li>height = 128</li>\n<li>width = 256</li>\n<li>n_fft = 1024</li>\n<li>win_length = 128</li>\n<li>fmin = 0.5</li>\n<li>fmax = 20</li>\n<li>USE_WAVELET = None<br>\nkfold:  5fold group by patients<br>\nData Aug</li></ul></li>\n<li>xymask</li>\n<li>random brightness</li>\n<li>mixup<br>\nModel</li>\n<li>tf_efficientnet_b1 cv:0.27 LB:0.29</li>\n<li>tiny_vit_21m_512 cv:0.27 LB:0.29<br>\nTraining Detail</li>\n</ul>\n<ol>\n<li>Loss: nn.KLDivLoss(reduction=\"batchmean\")</li>\n<li>LR: 3e-3</li>\n<li>optimizer: AdamW</li>\n<li>scheduler: OneCycleLR</li>\n</ol>\n<h1>Horikita_Saku&amp;HB Part: <strong>1d+2d Models</strong></h1>\n<ul>\n<li>First, the 2d model is trained with 2stage.</li>\n<li>The backbone is then extracted and fused using the attention pool in the 1d model.<br>\nThe reason for this is that I noticed that the gradient changes in the 1d model and the 2d model are very different(The maximum difference is 1e10 times). At the same time, the changes in learning rate required for 1d and 2d model fitting are also very different (Again, the difference is on the order of 1e10-1e100). So I came up with the idea of training separately.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F25fe25ca47cee24592ef075604247b3a%2F6(99HLPSRAD8DH7JPT.png?generation=1712627676420178&amp;alt=media\"><br>\nAt the same time, some modifications were made to the 1d model. We slice the EEG wave. For example, 10000→2000*5. Then instead of Conv over 10,000, we Conv over every 2000. That gave us a bit of a boost.</p>\n<p><strong>TTA</strong><br>\nWe used TTA for a large number of models to improve the performance of each model as much as possible.</p>\n<p><strong>Sorted Kaggle Spec</strong><br>\nThe order of the four channels of the kaggle spec is ['LL','RL','LP','RP']. Whereas EEG is ['LL','LP','RP','RR '], so in some of these models we try to reorder the Kaggle Spec, sorted it in the same channel order as EEG.</p>",
  "messages": [
    {
      "id": "2742623",
      "postDate": "04/09/2024 02:01:08",
      "content": "<h1>Summary</h1>\n<ul>\n<li>1d models (public LB ~0.26)</li>\n<li>1d+2d models (public LB ~0.25)</li>\n<li>2stages training</li>\n<li>training 2d models' backbone and 1d+2d models separately.</li>\n</ul>\n<h1>TmT&amp;Ken_Ken_pa Part : <strong>1d Models</strong></h1>\n<p>The 1d model is primarily based on Wavenet and uses the attention pool.<br>\nFirst train on all votes, then train on the part where vote&gt;=10.<br>\nAnd after a nice model is obtained, the pseudo label pre-training is carried out.</p>\n<p><strong>Data Augumentation</strong>: </p>\n<ul>\n<li>Flipping the left and right brain.  Inference with both the flipped and the unflipped brain.</li>\n<li>Use upsampling to replicate samples from fewer classes.</li>\n</ul>\n<h1>BladeRunner Part: <strong>2d Models</strong></h1>\n<p>The 2D spectrogram solution uses a two-stage training.</p>\n<ul>\n<li>Stage 1: training on all data.</li>\n<li>Stage 2: training on only data with vote &gt; 9.\nData preparation and pre-processing\nUsing kaggle spectrogram + EEG spectrogram\nOnly one sample is generated for an eeg_id, and the data is taken right in the middle.\nkaggle_spec shape (512, 256, 1)\neeg_spec shape (512, 256, 1)\ncombine (512, 512, 1)\nEEG  spectrogram generated method: mel spectrogram<ul>\n<li>height = 128</li>\n<li>width = 256</li>\n<li>n_fft = 1024</li>\n<li>win_length = 128</li>\n<li>fmin = 0.5</li>\n<li>fmax = 20</li>\n<li>USE_WAVELET = None<br>\nkfold:  5fold group by patients<br>\nData Aug</li></ul></li>\n<li>xymask</li>\n<li>random brightness</li>\n<li>mixup<br>\nModel</li>\n<li>tf_efficientnet_b1 cv:0.27 LB:0.29</li>\n<li>tiny_vit_21m_512 cv:0.27 LB:0.29<br>\nTraining Detail</li>\n</ul>\n<ol>\n<li>Loss: nn.KLDivLoss(reduction=\"batchmean\")</li>\n<li>LR: 3e-3</li>\n<li>optimizer: AdamW</li>\n<li>scheduler: OneCycleLR</li>\n</ol>\n<h1>Horikita_Saku&amp;HB Part: <strong>1d+2d Models</strong></h1>\n<ul>\n<li>First, the 2d model is trained with 2stage.</li>\n<li>The backbone is then extracted and fused using the attention pool in the 1d model.<br>\nThe reason for this is that I noticed that the gradient changes in the 1d model and the 2d model are very different(The maximum difference is 1e10 times). At the same time, the changes in learning rate required for 1d and 2d model fitting are also very different (Again, the difference is on the order of 1e10-1e100). So I came up with the idea of training separately.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F25fe25ca47cee24592ef075604247b3a%2F6(99HLPSRAD8DH7JPT.png?generation=1712627676420178&amp;alt=media\"><br>\nAt the same time, some modifications were made to the 1d model. We slice the EEG wave. For example, 10000→2000*5. Then instead of Conv over 10,000, we Conv over every 2000. That gave us a bit of a boost.</p>\n<p><strong>TTA</strong><br>\nWe used TTA for a large number of models to improve the performance of each model as much as possible.</p>\n<p><strong>Sorted Kaggle Spec</strong><br>\nThe order of the four channels of the kaggle spec is ['LL','RL','LP','RP']. Whereas EEG is ['LL','LP','RP','RR '], so in some of these models we try to reorder the Kaggle Spec, sorted it in the same channel order as EEG.</p>",
      "rawMarkdown": "# Summary\n\n- 1d models (public LB ~0.26)\n- 1d+2d models (public LB ~0.25)\n- 2stages training\n- training 2d models' backbone and 1d+2d models separately.\n\n# TmT&Ken_Ken_pa Part : **1d Models**\nThe 1d model is primarily based on Wavenet and uses the attention pool.\nFirst train on all votes, then train on the part where vote>=10.\nAnd after a nice model is obtained, the pseudo label pre-training is carried out.\n\n**Data Augumentation**: \n  - Flipping the left and right brain.  Inference with both the flipped and the unflipped brain.\n  - Use upsampling to replicate samples from fewer classes.\n\n# BladeRunner Part: **2d Models**\nThe 2D spectrogram solution uses a two-stage training.\n- Stage 1: training on all data.\n- Stage 2: training on only data with vote > 9.\nData preparation and pre-processing\nUsing kaggle spectrogram + EEG spectrogram\nOnly one sample is generated for an eeg_id, and the data is taken right in the middle.\nkaggle_spec shape (512, 256, 1)\neeg_spec shape (512, 256, 1)\ncombine (512, 512, 1)\nEEG  spectrogram generated method: mel spectrogram\n  - height = 128\n  - width = 256\n  - n_fft = 1024\n  - win_length = 128\n  - fmin = 0.5\n  - fmax = 20\n  - USE_WAVELET = None\nkfold:  5fold group by patients\nData Aug\n- xymask\n- random brightness\n- mixup\nModel\n- tf_efficientnet_b1 cv:0.27 LB:0.29\n- tiny_vit_21m_512 cv:0.27 LB:0.29\nTraining Detail\n1. Loss: nn.KLDivLoss(reduction=\"batchmean\")\n2. LR: 3e-3\n3. optimizer: AdamW\n4. scheduler: OneCycleLR\n\n# Horikita_Saku&HB Part: **1d+2d Models**\n- First, the 2d model is trained with 2stage.\n- The backbone is then extracted and fused using the attention pool in the 1d model.\nThe reason for this is that I noticed that the gradient changes in the 1d model and the 2d model are very different(The maximum difference is 1e10 times). At the same time, the changes in learning rate required for 1d and 2d model fitting are also very different (Again, the difference is on the order of 1e10-1e100). So I came up with the idea of training separately.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F25fe25ca47cee24592ef075604247b3a%2F6(99HLPSRAD8DH7JPT.png?generation=1712627676420178&alt=media)\nAt the same time, some modifications were made to the 1d model. We slice the EEG wave. For example, 10000→2000*5. Then instead of Conv over 10,000, we Conv over every 2000. That gave us a bit of a boost.\n\n**TTA**\nWe used TTA for a large number of models to improve the performance of each model as much as possible.\n\n**Sorted Kaggle Spec**\nThe order of the four channels of the kaggle spec is ['LL','RL','LP','RP']. Whereas EEG is ['LL','LP','RP','RR '], so in some of these models we try to reorder the Kaggle Spec, sorted it in the same channel order as EEG.",
      "votes": null
    },
    {
      "id": "2742806",
      "postDate": "04/09/2024 05:18:53",
      "content": "<p><a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a> Thanks for sharing, may I know what is the important part to get 1d model work well? I don't think there are public notes which has cv lower then 0.3 or single model LB lower then 0.3.</p>",
      "rawMarkdown": "horikitasaku Thanks for sharing, may I know what is the important part to get 1d model work well? I don't think there are public notes which has cv lower then 0.3 or single model LB lower then 0.3.",
      "votes": null
    },
    {
      "id": "2742836",
      "postDate": "04/09/2024 05:38:13",
      "content": "<p>I think the upsampling is doing a good job. Direct replication of the samples from less classes. I tried not to upsample but cv stayed at 0.33.<br>\nIt's also interesting to note that upsampling seems to work poorly on 2d models</p>",
      "rawMarkdown": "I think the upsampling is doing a good job. Direct replication of the samples from less classes. I tried not to upsample but cv stayed at 0.33.\nIt's also interesting to note that upsampling seems to work poorly on 2d models",
      "votes": null
    },
    {
      "id": "2742905",
      "postDate": "04/09/2024 06:17:20",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a>, your team and CPMP all use sampling on dataset, good finding!</p>",
      "rawMarkdown": "Thanks @horikitasaku, your team and CPMP all use sampling on dataset, good finding!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2742806,
      "author_name": "goldenlock",
      "author_url": "",
      "post_date": "04/09/2024 05:18:53",
      "content": "<p><a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a> Thanks for sharing, may I know what is the important part to get 1d model work well? I don't think there are public notes which has cv lower then 0.3 or single model LB lower then 0.3.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2742836,
          "author_name": "horikitasaku",
          "author_url": "",
          "post_date": "04/09/2024 05:38:13",
          "content": "<p>I think the upsampling is doing a good job. Direct replication of the samples from less classes. I tried not to upsample but cv stayed at 0.33.<br>\nIt's also interesting to note that upsampling seems to work poorly on 2d models</p>",
          "votes": null,
          "replies": [
            {
              "id": 2742905,
              "author_name": "goldenlock",
              "author_url": "",
              "post_date": "04/09/2024 06:17:20",
              "content": "<p>Thanks <a href=\"https://www.kaggle.com/horikitasaku\" target=\"_blank\">@horikitasaku</a>, your team and CPMP all use sampling on dataset, good finding!</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2742623": "# Summary\n\n- 1d models (public LB ~0.26)\n- 1d+2d models (public LB ~0.25)\n- 2stages training\n- training 2d models' backbone and 1d+2d models separately.\n\n# TmT&Ken_Ken_pa Part : **1d Models**\nThe 1d model is primarily based on Wavenet and uses the attention pool.\nFirst train on all votes, then train on the part where vote>=10.\nAnd after a nice model is obtained, the pseudo label pre-training is carried out.\n\n**Data Augumentation**: \n  - Flipping the left and right brain.  Inference with both the flipped and the unflipped brain.\n  - Use upsampling to replicate samples from fewer classes.\n\n# BladeRunner Part: **2d Models**\nThe 2D spectrogram solution uses a two-stage training.\n- Stage 1: training on all data.\n- Stage 2: training on only data with vote > 9.\nData preparation and pre-processing\nUsing kaggle spectrogram + EEG spectrogram\nOnly one sample is generated for an eeg_id, and the data is taken right in the middle.\nkaggle_spec shape (512, 256, 1)\neeg_spec shape (512, 256, 1)\ncombine (512, 512, 1)\nEEG  spectrogram generated method: mel spectrogram\n  - height = 128\n  - width = 256\n  - n_fft = 1024\n  - win_length = 128\n  - fmin = 0.5\n  - fmax = 20\n  - USE_WAVELET = None\nkfold:  5fold group by patients\nData Aug\n- xymask\n- random brightness\n- mixup\nModel\n- tf_efficientnet_b1 cv:0.27 LB:0.29\n- tiny_vit_21m_512 cv:0.27 LB:0.29\nTraining Detail\n1. Loss: nn.KLDivLoss(reduction=\"batchmean\")\n2. LR: 3e-3\n3. optimizer: AdamW\n4. scheduler: OneCycleLR\n\n# Horikita_Saku&HB Part: **1d+2d Models**\n- First, the 2d model is trained with 2stage.\n- The backbone is then extracted and fused using the attention pool in the 1d model.\nThe reason for this is that I noticed that the gradient changes in the 1d model and the 2d model are very different(The maximum difference is 1e10 times). At the same time, the changes in learning rate required for 1d and 2d model fitting are also very different (Again, the difference is on the order of 1e10-1e100). So I came up with the idea of training separately.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F11676771%2F25fe25ca47cee24592ef075604247b3a%2F6(99HLPSRAD8DH7JPT.png?generation=1712627676420178&alt=media)\nAt the same time, some modifications were made to the 1d model. We slice the EEG wave. For example, 10000→2000*5. Then instead of Conv over 10,000, we Conv over every 2000. That gave us a bit of a boost.\n\n**TTA**\nWe used TTA for a large number of models to improve the performance of each model as much as possible.\n\n**Sorted Kaggle Spec**\nThe order of the four channels of the kaggle spec is ['LL','RL','LP','RP']. Whereas EEG is ['LL','LP','RP','RR '], so in some of these models we try to reorder the Kaggle Spec, sorted it in the same channel order as EEG.",
    "2742806": "horikitasaku Thanks for sharing, may I know what is the important part to get 1d model work well? I don't think there are public notes which has cv lower then 0.3 or single model LB lower then 0.3.",
    "2742836": "I think the upsampling is doing a good job. Direct replication of the samples from less classes. I tried not to upsample but cv stayed at 0.33.\nIt's also interesting to note that upsampling seems to work poorly on 2d models",
    "2742905": "Thanks @horikitasaku, your team and CPMP all use sampling on dataset, good finding!"
  },
  "source": "meta"
}