{
  "id": 492251,
  "title": "16th place write-up: EEG montages -> 2D image + specs",
  "url": "/competitions/hms-harmful-brain-activity-classification/writeups/mikhail-kotyushev-16th-place-write-up-eeg-montages",
  "author_name": "",
  "post_date": "2024-04-09T03:15:09.833Z",
  "votes": 42,
  "comment_count": 5,
  "views": 0,
  "content": "<h1>Overview</h1>\n<p>First, I would like to thank the organizers and the community of the competition for a great challenge and very interesting time spent in the notebooks &amp; discussions. Unfortunately, this time I was only a reader and will try to compensate it with this write-up. Please do not hesitate to comment or ask questions if you have any!</p>\n<p>Here is the high level description of my solution for 16th place in the HMS - Harmful Brain Activity Classification competition, along with a bit more detailed dive in the most important topics.</p>\n<ul>\n<li><strong>Data:</strong> competition's data 10min &amp; <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s (big thanks to him!) 50s spectrograms concatenated with plots of normalized EEG bipolar montages</li>\n<li><strong>Model:</strong> timm's tinyViT as a backbone (the best quality / speed ratio)</li>\n<li><strong>Training:</strong> stage 1 training with only &gt; 7 voters, stage 2 by pseudolabeling (PL) the rest of the data &amp; re-training</li>\n<li><strong>Augmentations:</strong> random selection of EEG-sub-id, channel flipping of left-side and right-side electrodes in raw EEGs and spectrograms, image augs: coarse dropout + brightness / contrast</li>\n<li><strong>Other:</strong> cosine LR schedule w warm-up, EMA with decay 0.8, TTA by flipping</li>\n</ul>\n<p>The final solution is 6 seeds ensemble on all the data (&gt; 7 labels, &lt;= 7 PLs), PLs for &lt;= 7 voters part was also obtained by 10-seeds ensemble.</p>\n<p>The most helpful parts were selecting only quality labels with &gt; 7 voters, plotting 1D as 2D, using global bipolar montages' normalization instead of per-object min-max, luck of the backbone selection and using large image size. PLs, usage of EMA and TTA introduced only minor (~0.001 order of magnitude) improvements. I have not properly measured the augmentations impact and only could state that adding them greatly reduced overfitting.</p>\n<h1>Some insights from the papers</h1>\n<p>The following supplementary materials <a href=\"https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\" target=\"_blank\">paper</a> helped with two things: </p>\n<ol>\n<li>Answer to the question: does test data contain only objects with high number of voters?</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3ca1a241c9369e5a0aedab3830118532%2Ftrain_test_split.png?generation=1712629787176792&amp;alt=media\" alt=\"Train-test split\"></p>\n<ol>\n<li>Channel flipping augmentation.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F635d896eace302770b690d96ac764b7c%2Faug.png?generation=1712629811426093&amp;alt=media\" alt=\"Channel flipping aug\"></p>\n<p>Another thing covered here was train-test splitting strategy based on errors' equalization that could help to have more equal distribution of the hard samples in train and test, but I have not used it, sticking to the familiar <code>StratifiedGroupKFold</code>.</p>\n<h1>Data</h1>\n<p>Here, the idea was to leverage 2D vision models and weights (esp. transformers architecture) and simultaneously show the model just the same picture as which was shown to the voters when they were labeling the data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fb76379e629771a99ba4d6a6604ab0f1f%2Fvoter_gui.png?generation=1712629446043533&amp;alt=media\" alt=\"Voter's GUI\"></p>\n<p>That was achieved by converting everything in 2D as follows:</p>\n<p>Competition's data spectrograms with the following pre-processing:</p>\n<pre><code>    [np.isnan(spectrogram)] = \n     = np.log10(spectrogram + e-)\n    , max_ = np.quantile(spectrogram, .), np.quantile(spectrogram, .)\n     = np.clip(spectrogram, min_, max_)\n     = (spectrogram - min_) / (max_ - min_)\n</code></pre>\n<p>And <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s spectrograms with no additional pre-processing other than converting to [0, 1] range. Both specs were resized to 320 x 320 and concatenated vertically.</p>\n<p>Next, EEG bipolar montages of central 20s were calculated, filtered in 1-70 Hz range excluding possible power line noise and normalized. Pre-calculated <code>EEG_DIFF_ABS_MAX[i]</code> values are mostly around 100 as 0.05 and 0.95 quantiles of the corresponding montages.</p>\n<pre><code>    #  50 Hz &amp; 60 Hz noise\n    b, a = butter(5, (59, 61), =, =, =EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, =0)\n\n    # Bandpass 1-70 Hz\n    b, a = butter(5, (1, 70), =, =, =EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, =0)\n\n    # Normalize\n    min_, max_ = \\\n        -2 * EEG_DIFF_ABS_MAX[i], \\\n        2 * EEG_DIFF_ABS_MAX[i]\n    y = np.clip(y, min_, max_)\n    y = (y - min_) / (max_ - min_)\n    y[0] = 0\n    y[-1] = 1\n</code></pre>\n<p>The 20 resulting montages were plotted to 2D array of size 640 x 1600 by <code>justpyplot</code> library, each covers up to 64 x 1600 area with overlaps and concatenated to the spectrograms. The total image size was 640 x 1920, find an example of the image below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F702b42f5142c89598a429e28b7a4bba2%2Fimage_example.png?generation=1712629513515473&amp;alt=media\" alt=\"Example of the model's input image\"></p>\n<h1>Model</h1>\n<p>Initially I have selected the <code>tiny_vit_21m_512.dist_in22k_ft_in1k</code> model to experiment with as indeed tiny and fast to train. At some moment I have tried to use some other models, but it appears that it was distilled to be very powerful, so to reach its performance I needed e.g. <code>hf_hub:timm/eva02_base_patch14_448.mim_in22k_ft_in22k_in1k</code> model, which is ~ 5 times larger and much slower. </p>\n<p>So, I could only recommend trying the mentioned TinyViT model in your problems!</p>\n<p><strong>Edit:</strong> here is the <a href=\"https://github.com/mkotyushev/hms\" target=\"_blank\">code</a> of the solution in Pytorch / Lightning / Docker.</p>",
  "messages": [
    {
      "id": "2742662",
      "postDate": "04/09/2024 02:49:33",
      "content": "<h1>Overview</h1>\n<p>First, I would like to thank the organizers and the community of the competition for a great challenge and very interesting time spent in the notebooks &amp; discussions. Unfortunately, this time I was only a reader and will try to compensate it with this write-up. Please do not hesitate to comment or ask questions if you have any!</p>\n<p>Here is the high level description of my solution for 16th place in the HMS - Harmful Brain Activity Classification competition, along with a bit more detailed dive in the most important topics.</p>\n<ul>\n<li><strong>Data:</strong> competition's data 10min &amp; <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s (big thanks to him!) 50s spectrograms concatenated with plots of normalized EEG bipolar montages</li>\n<li><strong>Model:</strong> timm's tinyViT as a backbone (the best quality / speed ratio)</li>\n<li><strong>Training:</strong> stage 1 training with only &gt; 7 voters, stage 2 by pseudolabeling (PL) the rest of the data &amp; re-training</li>\n<li><strong>Augmentations:</strong> random selection of EEG-sub-id, channel flipping of left-side and right-side electrodes in raw EEGs and spectrograms, image augs: coarse dropout + brightness / contrast</li>\n<li><strong>Other:</strong> cosine LR schedule w warm-up, EMA with decay 0.8, TTA by flipping</li>\n</ul>\n<p>The final solution is 6 seeds ensemble on all the data (&gt; 7 labels, &lt;= 7 PLs), PLs for &lt;= 7 voters part was also obtained by 10-seeds ensemble.</p>\n<p>The most helpful parts were selecting only quality labels with &gt; 7 voters, plotting 1D as 2D, using global bipolar montages' normalization instead of per-object min-max, luck of the backbone selection and using large image size. PLs, usage of EMA and TTA introduced only minor (~0.001 order of magnitude) improvements. I have not properly measured the augmentations impact and only could state that adding them greatly reduced overfitting.</p>\n<h1>Some insights from the papers</h1>\n<p>The following supplementary materials <a href=\"https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf\" target=\"_blank\">paper</a> helped with two things: </p>\n<ol>\n<li>Answer to the question: does test data contain only objects with high number of voters?</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3ca1a241c9369e5a0aedab3830118532%2Ftrain_test_split.png?generation=1712629787176792&amp;alt=media\" alt=\"Train-test split\"></p>\n<ol>\n<li>Channel flipping augmentation.</li>\n</ol>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F635d896eace302770b690d96ac764b7c%2Faug.png?generation=1712629811426093&amp;alt=media\" alt=\"Channel flipping aug\"></p>\n<p>Another thing covered here was train-test splitting strategy based on errors' equalization that could help to have more equal distribution of the hard samples in train and test, but I have not used it, sticking to the familiar <code>StratifiedGroupKFold</code>.</p>\n<h1>Data</h1>\n<p>Here, the idea was to leverage 2D vision models and weights (esp. transformers architecture) and simultaneously show the model just the same picture as which was shown to the voters when they were labeling the data. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fb76379e629771a99ba4d6a6604ab0f1f%2Fvoter_gui.png?generation=1712629446043533&amp;alt=media\" alt=\"Voter's GUI\"></p>\n<p>That was achieved by converting everything in 2D as follows:</p>\n<p>Competition's data spectrograms with the following pre-processing:</p>\n<pre><code>    [np.isnan(spectrogram)] = \n     = np.log10(spectrogram + e-)\n    , max_ = np.quantile(spectrogram, .), np.quantile(spectrogram, .)\n     = np.clip(spectrogram, min_, max_)\n     = (spectrogram - min_) / (max_ - min_)\n</code></pre>\n<p>And <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s spectrograms with no additional pre-processing other than converting to [0, 1] range. Both specs were resized to 320 x 320 and concatenated vertically.</p>\n<p>Next, EEG bipolar montages of central 20s were calculated, filtered in 1-70 Hz range excluding possible power line noise and normalized. Pre-calculated <code>EEG_DIFF_ABS_MAX[i]</code> values are mostly around 100 as 0.05 and 0.95 quantiles of the corresponding montages.</p>\n<pre><code>    #  50 Hz &amp; 60 Hz noise\n    b, a = butter(5, (59, 61), =, =, =EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, =0)\n\n    # Bandpass 1-70 Hz\n    b, a = butter(5, (1, 70), =, =, =EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, =0)\n\n    # Normalize\n    min_, max_ = \\\n        -2 * EEG_DIFF_ABS_MAX[i], \\\n        2 * EEG_DIFF_ABS_MAX[i]\n    y = np.clip(y, min_, max_)\n    y = (y - min_) / (max_ - min_)\n    y[0] = 0\n    y[-1] = 1\n</code></pre>\n<p>The 20 resulting montages were plotted to 2D array of size 640 x 1600 by <code>justpyplot</code> library, each covers up to 64 x 1600 area with overlaps and concatenated to the spectrograms. The total image size was 640 x 1920, find an example of the image below.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F702b42f5142c89598a429e28b7a4bba2%2Fimage_example.png?generation=1712629513515473&amp;alt=media\" alt=\"Example of the model's input image\"></p>\n<h1>Model</h1>\n<p>Initially I have selected the <code>tiny_vit_21m_512.dist_in22k_ft_in1k</code> model to experiment with as indeed tiny and fast to train. At some moment I have tried to use some other models, but it appears that it was distilled to be very powerful, so to reach its performance I needed e.g. <code>hf_hub:timm/eva02_base_patch14_448.mim_in22k_ft_in22k_in1k</code> model, which is ~ 5 times larger and much slower. </p>\n<p>So, I could only recommend trying the mentioned TinyViT model in your problems!</p>\n<p><strong>Edit:</strong> here is the <a href=\"https://github.com/mkotyushev/hms\" target=\"_blank\">code</a> of the solution in Pytorch / Lightning / Docker.</p>",
      "rawMarkdown": "# Overview\n\nFirst, I would like to thank the organizers and the community of the competition for a great challenge and very interesting time spent in the notebooks & discussions. Unfortunately, this time I was only a reader and will try to compensate it with this write-up. Please do not hesitate to comment or ask questions if you have any!\n\nHere is the high level description of my solution for 16th place in the HMS - Harmful Brain Activity Classification competition, along with a bit more detailed dive in the most important topics.\n\n- **Data:** competition's data 10min & @cdeotte's (big thanks to him!) 50s spectrograms concatenated with plots of normalized EEG bipolar montages\n- **Model:** timm's tinyViT as a backbone (the best quality / speed ratio)\n- **Training:** stage 1 training with only > 7 voters, stage 2 by pseudolabeling (PL) the rest of the data & re-training\n- **Augmentations:** random selection of EEG-sub-id, channel flipping of left-side and right-side electrodes in raw EEGs and spectrograms, image augs: coarse dropout + brightness / contrast\n- **Other:** cosine LR schedule w warm-up, EMA with decay 0.8, TTA by flipping\n\nThe final solution is 6 seeds ensemble on all the data (> 7 labels, <= 7 PLs), PLs for <= 7 voters part was also obtained by 10-seeds ensemble.\n\nThe most helpful parts were selecting only quality labels with > 7 voters, plotting 1D as 2D, using global bipolar montages' normalization instead of per-object min-max, luck of the backbone selection and using large image size. PLs, usage of EMA and TTA introduced only minor (~0.001 order of magnitude) improvements. I have not properly measured the augmentations impact and only could state that adding them greatly reduced overfitting.\n\n# Some insights from the papers\n\nThe following supplementary materials [paper](https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf) helped with two things: \n\n1. Answer to the question: does test data contain only objects with high number of voters?\n\n![Train-test split](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3ca1a241c9369e5a0aedab3830118532%2Ftrain_test_split.png?generation=1712629787176792&alt=media)\n \n2. Channel flipping augmentation.\n\n![Channel flipping aug](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F635d896eace302770b690d96ac764b7c%2Faug.png?generation=1712629811426093&alt=media)\n\nAnother thing covered here was train-test splitting strategy based on errors' equalization that could help to have more equal distribution of the hard samples in train and test, but I have not used it, sticking to the familiar `StratifiedGroupKFold`.\n\n# Data\n\nHere, the idea was to leverage 2D vision models and weights (esp. transformers architecture) and simultaneously show the model just the same picture as which was shown to the voters when they were labeling the data. \n\n![Voter's GUI](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fb76379e629771a99ba4d6a6604ab0f1f%2Fvoter_gui.png?generation=1712629446043533&alt=media)\n\nThat was achieved by converting everything in 2D as follows:\n\nCompetition's data spectrograms with the following pre-processing:\n\n```\n    spectrogram[np.isnan(spectrogram)] = 0\n    spectrogram = np.log10(spectrogram + 1e-6)\n    min_, max_ = np.quantile(spectrogram, 0.01), np.quantile(spectrogram, 0.99)\n    spectrogram = np.clip(spectrogram, min_, max_)\n    spectrogram = (spectrogram - min_) / (max_ - min_)\n```\n\nAnd @cdeotte's spectrograms with no additional pre-processing other than converting to [0, 1] range. Both specs were resized to 320 x 320 and concatenated vertically.\n\nNext, EEG bipolar montages of central 20s were calculated, filtered in 1-70 Hz range excluding possible power line noise and normalized. Pre-calculated `EEG_DIFF_ABS_MAX[i]` values are mostly around 100 as 0.05 and 0.95 quantiles of the corresponding montages.\n\n```\n    # Remove 50 Hz & 60 Hz noise\n    b, a = butter(5, (59, 61), btype='bandstop', analog=False, fs=EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, axis=0)\n\n    # Bandpass 1-70 Hz\n    b, a = butter(5, (1, 70), btype='bandpass', analog=False, fs=EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, axis=0)\n\n    # Normalize\n    min_, max_ = \\\n        -2 * EEG_DIFF_ABS_MAX[i], \\\n        2 * EEG_DIFF_ABS_MAX[i]\n    y = np.clip(y, min_, max_)\n    y = (y - min_) / (max_ - min_)\n    y[0] = 0\n    y[-1] = 1\n```\n\nThe 20 resulting montages were plotted to 2D array of size 640 x 1600 by `justpyplot` library, each covers up to 64 x 1600 area with overlaps and concatenated to the spectrograms. The total image size was 640 x 1920, find an example of the image below.\n\n![Example of the model's input image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F702b42f5142c89598a429e28b7a4bba2%2Fimage_example.png?generation=1712629513515473&alt=media)\n\n# Model\n\nInitially I have selected the `tiny_vit_21m_512.dist_in22k_ft_in1k` model to experiment with as indeed tiny and fast to train. At some moment I have tried to use some other models, but it appears that it was distilled to be very powerful, so to reach its performance I needed e.g. `hf_hub:timm/eva02_base_patch14_448.mim_in22k_ft_in22k_in1k` model, which is ~ 5 times larger and much slower. \n\nSo, I could only recommend trying the mentioned TinyViT model in your problems!\n\n**Edit:** here is the [code](https://github.com/mkotyushev/hms) of the solution in Pytorch / Lightning / Docker.",
      "votes": null
    },
    {
      "id": "2742682",
      "postDate": "04/09/2024 03:00:04",
      "content": "<p>Thank you for really good solution and congratulation!<br>\nSo did you input image of 640 x 1920 into model <code>tiny_vit_21m_512.dist_in22k_ft_in1k</code>?<br>\nAnd Can I ask your Best CV?</p>",
      "rawMarkdown": "Thank you for really good solution and congratulation!\nSo did you input image of 640 x 1920 into model `tiny_vit_21m_512.dist_in22k_ft_in1k`?\nAnd Can I ask your Best CV?",
      "votes": null
    },
    {
      "id": "2742781",
      "postDate": "04/09/2024 04:39:42",
      "content": "<p>thank you for sharing your code, there's a lot must to learn especially signal processing…</p>",
      "rawMarkdown": "thank you for sharing your code, there's a lot must to learn especially signal processing...",
      "votes": null
    },
    {
      "id": "2743600",
      "postDate": "04/09/2024 14:38:19",
      "content": "<p>Cool idea to try and match what the voters saw. Nice solution!</p>",
      "rawMarkdown": "Cool idea to try and match what the voters saw. Nice solution!",
      "votes": null
    },
    {
      "id": "2743636",
      "postDate": "04/09/2024 15:00:19",
      "content": "<p>Hi, thanks! Yes, the model is <code>tiny_vit_21m_512.dist_in22k_ft_in1k</code> adapted for 1 channel input (first convolution weight summed) and the 640 x 1920 image is fed into it. </p>\n<p>For CV: 5-fold average KL is ~0.235 (see fig), but the final solution is trained on the full data, so no validation done, and I trusted public LB here. Public LB was 0.24 for stage 1 (10 seeds, no PLs) and 0.23 for stage 2 (6 seeds with PLs generated by stage 1 models), and private was 0.29 for both stages.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F63185ad83c1b06fb9306761b82f3a3f7%2FScreenshot%202024-04-09%20194720.png?generation=1712674767764667&amp;alt=media\" alt=\"5-fold CV\"></p>",
      "rawMarkdown": "Hi, thanks! Yes, the model is `tiny_vit_21m_512.dist_in22k_ft_in1k` adapted for 1 channel input (first convolution weight summed) and the 640 x 1920 image is fed into it. \n\nFor CV: 5-fold average KL is ~0.235 (see fig), but the final solution is trained on the full data, so no validation done, and I trusted public LB here. Public LB was 0.24 for stage 1 (10 seeds, no PLs) and 0.23 for stage 2 (6 seeds with PLs generated by stage 1 models), and private was 0.29 for both stages.\n\n![5-fold CV](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F63185ad83c1b06fb9306761b82f3a3f7%2FScreenshot%202024-04-09%20194720.png?generation=1712674767764667&alt=media)",
      "votes": null
    },
    {
      "id": "2971370",
      "postDate": "08/27/2024 05:47:16",
      "content": "<p>Could you please share which web-based UI is in the picture you shared in this post?</p>",
      "rawMarkdown": "Could you please share which web-based UI is in the picture you shared in this post?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2742682,
      "author_name": "haruki741",
      "author_url": "",
      "post_date": "04/09/2024 03:00:04",
      "content": "<p>Thank you for really good solution and congratulation!<br>\nSo did you input image of 640 x 1920 into model <code>tiny_vit_21m_512.dist_in22k_ft_in1k</code>?<br>\nAnd Can I ask your Best CV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2743636,
          "author_name": "mkotyushev",
          "author_url": "",
          "post_date": "04/09/2024 15:00:19",
          "content": "<p>Hi, thanks! Yes, the model is <code>tiny_vit_21m_512.dist_in22k_ft_in1k</code> adapted for 1 channel input (first convolution weight summed) and the 640 x 1920 image is fed into it. </p>\n<p>For CV: 5-fold average KL is ~0.235 (see fig), but the final solution is trained on the full data, so no validation done, and I trusted public LB here. Public LB was 0.24 for stage 1 (10 seeds, no PLs) and 0.23 for stage 2 (6 seeds with PLs generated by stage 1 models), and private was 0.29 for both stages.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F63185ad83c1b06fb9306761b82f3a3f7%2FScreenshot%202024-04-09%20194720.png?generation=1712674767764667&amp;alt=media\" alt=\"5-fold CV\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2742781,
      "author_name": "saidkoussi",
      "author_url": "",
      "post_date": "04/09/2024 04:39:42",
      "content": "<p>thank you for sharing your code, there's a lot must to learn especially signal processing…</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2743600,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "04/09/2024 14:38:19",
      "content": "<p>Cool idea to try and match what the voters saw. Nice solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2971370,
      "author_name": "deepak412",
      "author_url": "",
      "post_date": "08/27/2024 05:47:16",
      "content": "<p>Could you please share which web-based UI is in the picture you shared in this post?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2742662": "# Overview\n\nFirst, I would like to thank the organizers and the community of the competition for a great challenge and very interesting time spent in the notebooks & discussions. Unfortunately, this time I was only a reader and will try to compensate it with this write-up. Please do not hesitate to comment or ask questions if you have any!\n\nHere is the high level description of my solution for 16th place in the HMS - Harmful Brain Activity Classification competition, along with a bit more detailed dive in the most important topics.\n\n- **Data:** competition's data 10min & @cdeotte's (big thanks to him!) 50s spectrograms concatenated with plots of normalized EEG bipolar montages\n- **Model:** timm's tinyViT as a backbone (the best quality / speed ratio)\n- **Training:** stage 1 training with only > 7 voters, stage 2 by pseudolabeling (PL) the rest of the data & re-training\n- **Augmentations:** random selection of EEG-sub-id, channel flipping of left-side and right-side electrodes in raw EEGs and spectrograms, image augs: coarse dropout + brightness / contrast\n- **Other:** cosine LR schedule w warm-up, EMA with decay 0.8, TTA by flipping\n\nThe final solution is 6 seeds ensemble on all the data (> 7 labels, <= 7 PLs), PLs for <= 7 voters part was also obtained by 10-seeds ensemble.\n\nThe most helpful parts were selecting only quality labels with > 7 voters, plotting 1D as 2D, using global bipolar montages' normalization instead of per-object min-max, luck of the backbone selection and using large image size. PLs, usage of EMA and TTA introduced only minor (~0.001 order of magnitude) improvements. I have not properly measured the augmentations impact and only could state that adding them greatly reduced overfitting.\n\n# Some insights from the papers\n\nThe following supplementary materials [paper](https://cdn-links.lww.com/permalink/wnl/c/wnl_2023_02_26_westover_1_sdc1.pdf) helped with two things: \n\n1. Answer to the question: does test data contain only objects with high number of voters?\n\n![Train-test split](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F3ca1a241c9369e5a0aedab3830118532%2Ftrain_test_split.png?generation=1712629787176792&alt=media)\n \n2. Channel flipping augmentation.\n\n![Channel flipping aug](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F635d896eace302770b690d96ac764b7c%2Faug.png?generation=1712629811426093&alt=media)\n\nAnother thing covered here was train-test splitting strategy based on errors' equalization that could help to have more equal distribution of the hard samples in train and test, but I have not used it, sticking to the familiar `StratifiedGroupKFold`.\n\n# Data\n\nHere, the idea was to leverage 2D vision models and weights (esp. transformers architecture) and simultaneously show the model just the same picture as which was shown to the voters when they were labeling the data. \n\n![Voter's GUI](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2Fb76379e629771a99ba4d6a6604ab0f1f%2Fvoter_gui.png?generation=1712629446043533&alt=media)\n\nThat was achieved by converting everything in 2D as follows:\n\nCompetition's data spectrograms with the following pre-processing:\n\n```\n    spectrogram[np.isnan(spectrogram)] = 0\n    spectrogram = np.log10(spectrogram + 1e-6)\n    min_, max_ = np.quantile(spectrogram, 0.01), np.quantile(spectrogram, 0.99)\n    spectrogram = np.clip(spectrogram, min_, max_)\n    spectrogram = (spectrogram - min_) / (max_ - min_)\n```\n\nAnd @cdeotte's spectrograms with no additional pre-processing other than converting to [0, 1] range. Both specs were resized to 320 x 320 and concatenated vertically.\n\nNext, EEG bipolar montages of central 20s were calculated, filtered in 1-70 Hz range excluding possible power line noise and normalized. Pre-calculated `EEG_DIFF_ABS_MAX[i]` values are mostly around 100 as 0.05 and 0.95 quantiles of the corresponding montages.\n\n```\n    # Remove 50 Hz & 60 Hz noise\n    b, a = butter(5, (59, 61), btype='bandstop', analog=False, fs=EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, axis=0)\n\n    # Bandpass 1-70 Hz\n    b, a = butter(5, (1, 70), btype='bandpass', analog=False, fs=EED_SAMPLING_RATE_HZ)\n    y = lfilter(b, a, y, axis=0)\n\n    # Normalize\n    min_, max_ = \\\n        -2 * EEG_DIFF_ABS_MAX[i], \\\n        2 * EEG_DIFF_ABS_MAX[i]\n    y = np.clip(y, min_, max_)\n    y = (y - min_) / (max_ - min_)\n    y[0] = 0\n    y[-1] = 1\n```\n\nThe 20 resulting montages were plotted to 2D array of size 640 x 1600 by `justpyplot` library, each covers up to 64 x 1600 area with overlaps and concatenated to the spectrograms. The total image size was 640 x 1920, find an example of the image below.\n\n![Example of the model's input image](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F702b42f5142c89598a429e28b7a4bba2%2Fimage_example.png?generation=1712629513515473&alt=media)\n\n# Model\n\nInitially I have selected the `tiny_vit_21m_512.dist_in22k_ft_in1k` model to experiment with as indeed tiny and fast to train. At some moment I have tried to use some other models, but it appears that it was distilled to be very powerful, so to reach its performance I needed e.g. `hf_hub:timm/eva02_base_patch14_448.mim_in22k_ft_in22k_in1k` model, which is ~ 5 times larger and much slower. \n\nSo, I could only recommend trying the mentioned TinyViT model in your problems!\n\n**Edit:** here is the [code](https://github.com/mkotyushev/hms) of the solution in Pytorch / Lightning / Docker.",
    "2742682": "Thank you for really good solution and congratulation!\nSo did you input image of 640 x 1920 into model `tiny_vit_21m_512.dist_in22k_ft_in1k`?\nAnd Can I ask your Best CV?",
    "2742781": "thank you for sharing your code, there's a lot must to learn especially signal processing...",
    "2743600": "Cool idea to try and match what the voters saw. Nice solution!",
    "2743636": "Hi, thanks! Yes, the model is `tiny_vit_21m_512.dist_in22k_ft_in1k` adapted for 1 channel input (first convolution weight summed) and the 640 x 1920 image is fed into it. \n\nFor CV: 5-fold average KL is ~0.235 (see fig), but the final solution is trained on the full data, so no validation done, and I trusted public LB here. Public LB was 0.24 for stage 1 (10 seeds, no PLs) and 0.23 for stage 2 (6 seeds with PLs generated by stage 1 models), and private was 0.29 for both stages.\n\n![5-fold CV](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2190976%2F63185ad83c1b06fb9306761b82f3a3f7%2FScreenshot%202024-04-09%20194720.png?generation=1712674767764667&alt=media)",
    "2971370": "Could you please share which web-based UI is in the picture you shared in this post?"
  },
  "source": "meta"
}