{
  "id": 479776,
  "title": "Proper Augmentations is a Key!",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/479776",
  "author_name": "",
  "post_date": "2024-02-25T23:59:43.675860400Z",
  "votes": 63,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I originally got little gain from augmentations and started to double their usefulness for this competition. After some further experimentation, this is absolutely not the case. I thought I could share some of my research results here. </p>\n<p>What worked:</p>\n<ul>\n<li>A few augmentations that worked was the mentioned XYMask which is described in detail with an example notebook here<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477825\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477825</a></li>\n<li>Vertical flips</li>\n</ul>\n<p>What didn't work:</p>\n<ul>\n<li>Horizontal flips</li>\n<li>Brightness contrast</li>\n<li>rotation</li>\n<li>cropping</li>\n</ul>\n<p>Conclusion: It may be due to a knowledge gap in EEG and spectrogram data but I was really surprised by some of the results. I really expected value in horizontal flips and none with vertical flips. However, after extensive testing that just wasn't the case. I also believe some are focusing too much on LB, which is common of course. So I would like to remind people to pay attention to and trust the CV a bit! </p>",
  "messages": [
    {
      "id": "2668827",
      "postDate": "02/25/2024 23:59:43",
      "content": "<p>I originally got little gain from augmentations and started to double their usefulness for this competition. After some further experimentation, this is absolutely not the case. I thought I could share some of my research results here. </p>\n<p>What worked:</p>\n<ul>\n<li>A few augmentations that worked was the mentioned XYMask which is described in detail with an example notebook here<br>\n<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477825\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477825</a></li>\n<li>Vertical flips</li>\n</ul>\n<p>What didn't work:</p>\n<ul>\n<li>Horizontal flips</li>\n<li>Brightness contrast</li>\n<li>rotation</li>\n<li>cropping</li>\n</ul>\n<p>Conclusion: It may be due to a knowledge gap in EEG and spectrogram data but I was really surprised by some of the results. I really expected value in horizontal flips and none with vertical flips. However, after extensive testing that just wasn't the case. I also believe some are focusing too much on LB, which is common of course. So I would like to remind people to pay attention to and trust the CV a bit! </p>",
      "rawMarkdown": "I originally got little gain from augmentations and started to double their usefulness for this competition. After some further experimentation, this is absolutely not the case. I thought I could share some of my research results here. \n\nWhat worked:\n* A few augmentations that worked was the mentioned XYMask which is described in detail with an example notebook here\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477825\n* Vertical flips\n\nWhat didn't work:\n* Horizontal flips\n* Brightness contrast\n* rotation\n* cropping\n\nConclusion: It may be due to a knowledge gap in EEG and spectrogram data but I was really surprised by some of the results. I really expected value in horizontal flips and none with vertical flips. However, after extensive testing that just wasn't the case. I also believe some are focusing too much on LB, which is common of course. So I would like to remind people to pay attention to and trust the CV a bit!",
      "votes": null
    },
    {
      "id": "2668856",
      "postDate": "02/26/2024 00:37:07",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2668906",
      "postDate": "02/26/2024 01:33:01",
      "content": "<p>Sure!  I used some augmentation and improved my CV score, but my LB score unchanged.</p>",
      "rawMarkdown": "Sure!  I used some augmentation and improved my CV score, but my LB score unchanged.",
      "votes": null
    },
    {
      "id": "2668912",
      "postDate": "02/26/2024 02:09:30",
      "content": "<p>Yes I have a feeling LB will only mean so much in this competition I have suspicions that many of the top LB scores are using data that is not available in test. And usually in competition CV is king</p>",
      "rawMarkdown": "Yes I have a feeling LB will only mean so much in this competition I have suspicions that many of the top LB scores are using data that is not available in test. And usually in competition CV is king",
      "votes": null
    },
    {
      "id": "2669595",
      "postDate": "02/26/2024 11:07:42",
      "content": "<p>I am surprised about vertical flips. It would change the place of every frequency of spectrogram. Some of the patterns are frequency specific (lateralized rhythmic <strong>delta</strong> activity (LRDA), generalized rhythmic <strong>delta</strong> activity (GRDA). My intuition is that this would worsen the classification of these patterns.</p>",
      "rawMarkdown": "I am surprised about vertical flips. It would change the place of every frequency of spectrogram. Some of the patterns are frequency specific (lateralized rhythmic **delta** activity (LRDA), generalized rhythmic **delta** activity (GRDA). My intuition is that this would worsen the classification of these patterns.",
      "votes": null
    },
    {
      "id": "2670307",
      "postDate": "02/26/2024 19:04:08",
      "content": "<p>I think the same way on this, perhaps it’s overfitting but it does improve my CV score so it’s tough to tell if it is legitimate or not</p>",
      "rawMarkdown": "I think the same way on this, perhaps it’s overfitting but it does improve my CV score so it’s tough to tell if it is legitimate or not",
      "votes": null
    },
    {
      "id": "2670370",
      "postDate": "02/26/2024 20:16:16",
      "content": "<p>XYMasking is working well for me, but it's quite sensitive to parameter choices, I find that generally multiple smaller masks work better than a few larger ones.</p>",
      "rawMarkdown": "XYMasking is working well for me, but it's quite sensitive to parameter choices, I find that generally multiple smaller masks work better than a few larger ones.",
      "votes": null
    },
    {
      "id": "2671166",
      "postDate": "02/27/2024 11:10:59",
      "content": "<p>It seems that finetuning on samples with a higher number of votes will boost lb but will make cv worse, this what probably explains the jump in the two stage approach (0.4* -&gt; 0.36) .<br>\nThis won't be overfitting if the private test has the same distribution as the public one.<br>\nI can think of a good reason for this to be the case : simply the organisers thought that samples with more votes are more trusthworthy and they should be used in both public and private test.<br>\nUnfortunately, if this is true, we should redefine our cv or finetune all our models </p>",
      "rawMarkdown": "It seems that finetuning on samples with a higher number of votes will boost lb but will make cv worse, this what probably explains the jump in the two stage approach (0.4* -> 0.36) .\nThis won't be overfitting if the private test has the same distribution as the public one.\nI can think of a good reason for this to be the case : simply the organisers thought that samples with more votes are more trusthworthy and they should be used in both public and private test.\nUnfortunately, if this is true, we should redefine our cv or finetune all our models",
      "votes": null
    },
    {
      "id": "2675822",
      "postDate": "03/01/2024 05:53:50",
      "content": "<p>what are your experiences on mixup? I have seen both positive and negative responses about mixup, it didn't work for me. can anyone share their experiences with mixup?</p>",
      "rawMarkdown": "what are your experiences on mixup? I have seen both positive and negative responses about mixup, it didn't work for me. can anyone share their experiences with mixup?",
      "votes": null
    },
    {
      "id": "2676033",
      "postDate": "03/01/2024 08:51:35",
      "content": "<p>For me, my CV has boost ~0.02 when using mix up, but LB has not changes</p>",
      "rawMarkdown": "For me, my CV has boost ~0.02 when using mix up, but LB has not changes",
      "votes": null
    },
    {
      "id": "2676150",
      "postDate": "03/01/2024 10:12:30",
      "content": "<p>Thank you, same with me, my LB had no changes with mixup so far, but I am trying some few tricks with mixup after that I will move on to other augs</p>",
      "rawMarkdown": "Thank you, same with me, my LB had no changes with mixup so far, but I am trying some few tricks with mixup after that I will move on to other augs",
      "votes": null
    },
    {
      "id": "2676311",
      "postDate": "03/01/2024 12:12:22",
      "content": "<p>My experience:</p>\n<ul>\n<li>The parameters (alpha and probability) should be tuned. Optimal parameters seem to differ between kaggle spectrograms and eeg spectrograms. </li>\n<li>Using different augmentation methods together with mixup possibly make mixup effects invisible. For example, when I used both horizontal-flip and mixup, which can reduce LB 0.01 respectively, I got only 0.01 reduction (not 0.01+0.01=0.02). I tried to tune probabilities, but I couldn't find the way to make both methods effective. <ul>\n<li>Note: My experiments were based on kaggle spectrograms. An effective combination of mixup and other methods is reported in a different situation (<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/480674)\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/480674)</a>, so additional experiments still seem to be valuable.</li></ul></li>\n<li>Mixup effect was 0.02 at most. </li>\n</ul>",
      "rawMarkdown": "My experience:\n- The parameters (alpha and probability) should be tuned. Optimal parameters seem to differ between kaggle spectrograms and eeg spectrograms. \n- Using different augmentation methods together with mixup possibly make mixup effects invisible. For example, when I used both horizontal-flip and mixup, which can reduce LB 0.01 respectively, I got only 0.01 reduction (not 0.01+0.01=0.02). I tried to tune probabilities, but I couldn't find the way to make both methods effective. \n  - Note: My experiments were based on kaggle spectrograms. An effective combination of mixup and other methods is reported in a different situation (https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/480674), so additional experiments still seem to be valuable.\n- Mixup effect was 0.02 at most.",
      "votes": null
    },
    {
      "id": "2677534",
      "postDate": "03/02/2024 09:10:44",
      "content": "<p>It's a true <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>, various transpose techniques making key to proper argumentation and to good CV score.</p>",
      "rawMarkdown": "It's a true @cody11null, various transpose techniques making key to proper argumentation and to good CV score.",
      "votes": null
    },
    {
      "id": "2688342",
      "postDate": "03/09/2024 07:06:04",
      "content": "<p>Hello, I have two inquiries:</p>\n<ol>\n<li><p>Could you explain the integration of both EEG data and spectrogram into the processes of model training?</p></li>\n<li><p>Is it imperative to apply filtering, denoising, or normalization techniques to the provided readings?</p></li>\n</ol>",
      "rawMarkdown": "Hello, I have two inquiries:\n\n1. Could you explain the integration of both EEG data and spectrogram into the processes of model training?\n\n2. Is it imperative to apply filtering, denoising, or normalization techniques to the provided readings?",
      "votes": null
    },
    {
      "id": "3159311",
      "postDate": "03/25/2025 12:52:27",
      "content": "<p>Thanks! Interesting to know. I should have discovered this post before uploading my data augmentation package haha</p>",
      "rawMarkdown": "Thanks! Interesting to know. I should have discovered this post before uploading my data augmentation package haha",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2668856,
      "author_name": "akulvaishnavi",
      "author_url": "",
      "post_date": "02/26/2024 00:37:07",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2668906,
      "author_name": "sweetyheehee",
      "author_url": "",
      "post_date": "02/26/2024 01:33:01",
      "content": "<p>Sure!  I used some augmentation and improved my CV score, but my LB score unchanged.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2668912,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "02/26/2024 02:09:30",
          "content": "<p>Yes I have a feeling LB will only mean so much in this competition I have suspicions that many of the top LB scores are using data that is not available in test. And usually in competition CV is king</p>",
          "votes": null,
          "replies": [
            {
              "id": 2671166,
              "author_name": "ahmedelfazouan",
              "author_url": "",
              "post_date": "02/27/2024 11:10:59",
              "content": "<p>It seems that finetuning on samples with a higher number of votes will boost lb but will make cv worse, this what probably explains the jump in the two stage approach (0.4* -&gt; 0.36) .<br>\nThis won't be overfitting if the private test has the same distribution as the public one.<br>\nI can think of a good reason for this to be the case : simply the organisers thought that samples with more votes are more trusthworthy and they should be used in both public and private test.<br>\nUnfortunately, if this is true, we should redefine our cv or finetune all our models </p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2669595,
      "author_name": "serjhenrique",
      "author_url": "",
      "post_date": "02/26/2024 11:07:42",
      "content": "<p>I am surprised about vertical flips. It would change the place of every frequency of spectrogram. Some of the patterns are frequency specific (lateralized rhythmic <strong>delta</strong> activity (LRDA), generalized rhythmic <strong>delta</strong> activity (GRDA). My intuition is that this would worsen the classification of these patterns.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2670307,
          "author_name": "cody11null",
          "author_url": "",
          "post_date": "02/26/2024 19:04:08",
          "content": "<p>I think the same way on this, perhaps it’s overfitting but it does improve my CV score so it’s tough to tell if it is legitimate or not</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2670370,
      "author_name": "kionce",
      "author_url": "",
      "post_date": "02/26/2024 20:16:16",
      "content": "<p>XYMasking is working well for me, but it's quite sensitive to parameter choices, I find that generally multiple smaller masks work better than a few larger ones.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2675822,
      "author_name": "arunsensei",
      "author_url": "",
      "post_date": "03/01/2024 05:53:50",
      "content": "<p>what are your experiences on mixup? I have seen both positive and negative responses about mixup, it didn't work for me. can anyone share their experiences with mixup?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2676033,
          "author_name": "haruki741",
          "author_url": "",
          "post_date": "03/01/2024 08:51:35",
          "content": "<p>For me, my CV has boost ~0.02 when using mix up, but LB has not changes</p>",
          "votes": null,
          "replies": [
            {
              "id": 2676150,
              "author_name": "arunsensei",
              "author_url": "",
              "post_date": "03/01/2024 10:12:30",
              "content": "<p>Thank you, same with me, my LB had no changes with mixup so far, but I am trying some few tricks with mixup after that I will move on to other augs</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2676311,
          "author_name": "symz00",
          "author_url": "",
          "post_date": "03/01/2024 12:12:22",
          "content": "<p>My experience:</p>\n<ul>\n<li>The parameters (alpha and probability) should be tuned. Optimal parameters seem to differ between kaggle spectrograms and eeg spectrograms. </li>\n<li>Using different augmentation methods together with mixup possibly make mixup effects invisible. For example, when I used both horizontal-flip and mixup, which can reduce LB 0.01 respectively, I got only 0.01 reduction (not 0.01+0.01=0.02). I tried to tune probabilities, but I couldn't find the way to make both methods effective. <ul>\n<li>Note: My experiments were based on kaggle spectrograms. An effective combination of mixup and other methods is reported in a different situation (<a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/480674)\" target=\"_blank\">https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/480674)</a>, so additional experiments still seem to be valuable.</li></ul></li>\n<li>Mixup effect was 0.02 at most. </li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2677534,
      "author_name": "ravindragaikar",
      "author_url": "",
      "post_date": "03/02/2024 09:10:44",
      "content": "<p>It's a true <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a>, various transpose techniques making key to proper argumentation and to good CV score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2688342,
      "author_name": "farwa99",
      "author_url": "",
      "post_date": "03/09/2024 07:06:04",
      "content": "<p>Hello, I have two inquiries:</p>\n<ol>\n<li><p>Could you explain the integration of both EEG data and spectrogram into the processes of model training?</p></li>\n<li><p>Is it imperative to apply filtering, denoising, or normalization techniques to the provided readings?</p></li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3159311,
      "author_name": "alexandrelemercier",
      "author_url": "",
      "post_date": "03/25/2025 12:52:27",
      "content": "<p>Thanks! Interesting to know. I should have discovered this post before uploading my data augmentation package haha</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2668827": "I originally got little gain from augmentations and started to double their usefulness for this competition. After some further experimentation, this is absolutely not the case. I thought I could share some of my research results here. \n\nWhat worked:\n* A few augmentations that worked was the mentioned XYMask which is described in detail with an example notebook here\nhttps://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477825\n* Vertical flips\n\nWhat didn't work:\n* Horizontal flips\n* Brightness contrast\n* rotation\n* cropping\n\nConclusion: It may be due to a knowledge gap in EEG and spectrogram data but I was really surprised by some of the results. I really expected value in horizontal flips and none with vertical flips. However, after extensive testing that just wasn't the case. I also believe some are focusing too much on LB, which is common of course. So I would like to remind people to pay attention to and trust the CV a bit!",
    "2668856": "Thanks for sharing!",
    "2668906": "Sure!  I used some augmentation and improved my CV score, but my LB score unchanged.",
    "2668912": "Yes I have a feeling LB will only mean so much in this competition I have suspicions that many of the top LB scores are using data that is not available in test. And usually in competition CV is king",
    "2669595": "I am surprised about vertical flips. It would change the place of every frequency of spectrogram. Some of the patterns are frequency specific (lateralized rhythmic **delta** activity (LRDA), generalized rhythmic **delta** activity (GRDA). My intuition is that this would worsen the classification of these patterns.",
    "2670307": "I think the same way on this, perhaps it’s overfitting but it does improve my CV score so it’s tough to tell if it is legitimate or not",
    "2670370": "XYMasking is working well for me, but it's quite sensitive to parameter choices, I find that generally multiple smaller masks work better than a few larger ones.",
    "2671166": "It seems that finetuning on samples with a higher number of votes will boost lb but will make cv worse, this what probably explains the jump in the two stage approach (0.4* -> 0.36) .\nThis won't be overfitting if the private test has the same distribution as the public one.\nI can think of a good reason for this to be the case : simply the organisers thought that samples with more votes are more trusthworthy and they should be used in both public and private test.\nUnfortunately, if this is true, we should redefine our cv or finetune all our models",
    "2675822": "what are your experiences on mixup? I have seen both positive and negative responses about mixup, it didn't work for me. can anyone share their experiences with mixup?",
    "2676033": "For me, my CV has boost ~0.02 when using mix up, but LB has not changes",
    "2676150": "Thank you, same with me, my LB had no changes with mixup so far, but I am trying some few tricks with mixup after that I will move on to other augs",
    "2676311": "My experience:\n- The parameters (alpha and probability) should be tuned. Optimal parameters seem to differ between kaggle spectrograms and eeg spectrograms. \n- Using different augmentation methods together with mixup possibly make mixup effects invisible. For example, when I used both horizontal-flip and mixup, which can reduce LB 0.01 respectively, I got only 0.01 reduction (not 0.01+0.01=0.02). I tried to tune probabilities, but I couldn't find the way to make both methods effective. \n  - Note: My experiments were based on kaggle spectrograms. An effective combination of mixup and other methods is reported in a different situation (https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/480674), so additional experiments still seem to be valuable.\n- Mixup effect was 0.02 at most.",
    "2677534": "It's a true @cody11null, various transpose techniques making key to proper argumentation and to good CV score.",
    "2688342": "Hello, I have two inquiries:\n\n1. Could you explain the integration of both EEG data and spectrogram into the processes of model training?\n\n2. Is it imperative to apply filtering, denoising, or normalization techniques to the provided readings?",
    "3159311": "Thanks! Interesting to know. I should have discovered this post before uploading my data augmentation package haha"
  },
  "source": "meta"
}