{
  "id": 479446,
  "title": "Same-Class-CutMix: The Only Data Aug that worked for me",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/479446",
  "author_name": "Vikram Sandu",
  "post_date": "2024-02-24T17:48:58.841000",
  "votes": 42,
  "comment_count": 10,
  "views": 0,
  "content": "<p>CutMixing the samples from the SAME CLASSES where Inputs are proportionally cut and mixed but not labels. (Experimental)</p>\n<p>To be clear, we are cutting a whole time window from image-1 and replacing it randomly in image-2 except the middle portion of the image.</p>\n<pre><code>\n random\n\n ():\n\n    \n    cutmix_data = data.clone()\n\n     label_idx  (target.size()):\n        \n        indices = torch.nonzero((target[:, label_idx] &gt;= cutmix_thr), as_tuple=)\n\n        \n         (indices) &lt; :\n            \n\n        \n        data_orig = data[indices]\n\n        \n        shuffled_indices = torch.randperm((indices))\n        data_shuffled = data_orig[shuffled_indices]\n\n        \n        start = random.randint(, margin)  random.choice([, ])  random.randint(-max_size-margin, -max_size)\n        size = random.randint(min_size, max_size)\n\n        \n        cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n\n        \n         cut_eeg_spec:\n            start =  +  + start \n            cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n\n     cutmix_data, target\n</code></pre>\n<p>Model | OOF-CV | LB<br>\nEffB0  | 0.6642    | 0.41<br>\nEffB0 + CutMix (p=0.4)  | 0.6530    | 0.39</p>\n<p>Here are the list of augmentations that didn't work for ME:</p>\n<ol>\n<li>TimeAxis Cutout</li>\n<li>FrequencyAxis Cutout</li>\n<li>HFlip</li>\n<li>SpecTranslation </li>\n<li>Translate and Mix</li>\n<li>MixUP</li>\n<li>Same Class Mixup</li>\n<li>Spec Brightness Adjustment</li>\n<li>Random Noise</li>\n</ol>\n<p>Thank you all for sharing your amazing work and insights, especially Chris and Tawara :)</p>",
  "messages": [
    {
      "id": 2666878,
      "postDate": "2024-02-24T17:48:58.840Z",
      "content": "<p>CutMixing the samples from the SAME CLASSES where Inputs are proportionally cut and mixed but not labels. (Experimental)</p>\n<p>To be clear, we are cutting a whole time window from image-1 and replacing it randomly in image-2 except the middle portion of the image.</p>\n<pre><code>\n random\n\n ():\n\n    \n    cutmix_data = data.clone()\n\n     label_idx  (target.size()):\n        \n        indices = torch.nonzero((target[:, label_idx] &gt;= cutmix_thr), as_tuple=)\n\n        \n         (indices) &lt; :\n            \n\n        \n        data_orig = data[indices]\n\n        \n        shuffled_indices = torch.randperm((indices))\n        data_shuffled = data_orig[shuffled_indices]\n\n        \n        start = random.randint(, margin)  random.choice([, ])  random.randint(-max_size-margin, -max_size)\n        size = random.randint(min_size, max_size)\n\n        \n        cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n\n        \n         cut_eeg_spec:\n            start =  +  + start \n            cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n\n     cutmix_data, target\n</code></pre>\n<p>Model | OOF-CV | LB<br>\nEffB0  | 0.6642    | 0.41<br>\nEffB0 + CutMix (p=0.4)  | 0.6530    | 0.39</p>\n<p>Here are the list of augmentations that didn't work for ME:</p>\n<ol>\n<li>TimeAxis Cutout</li>\n<li>FrequencyAxis Cutout</li>\n<li>HFlip</li>\n<li>SpecTranslation </li>\n<li>Translate and Mix</li>\n<li>MixUP</li>\n<li>Same Class Mixup</li>\n<li>Spec Brightness Adjustment</li>\n<li>Random Noise</li>\n</ol>\n<p>Thank you all for sharing your amazing work and insights, especially Chris and Tawara :)</p>",
      "rawMarkdown": "CutMixing the samples from the SAME CLASSES where Inputs are proportionally cut and mixed but not labels. (Experimental)\n\nTo be clear, we are cutting a whole time window from image-1 and replacing it randomly in image-2 except the middle portion of the image.\n\n```python\n'''\n  This applies on the Batch\n'''\nimport random\n\ndef hbac_cutmix(data, \n                target,\n                cutmix_thr = 0.75, # Threshold to consider samples\n                margin=50, # Either Cut 50 from begining  or end\n                min_size=25, # Min Cutout Size\n                max_size=75, # Max Cutout Size\n                cut_eeg_spec = True # CutMix in EEG Specs\n               ):\n    \n    # CutMix Data\n    cutmix_data = data.clone()\n    \n    for label_idx in range(target.size(1)):\n        # Indices with a confidence score greater than cutmix_thr for particular target\n        indices = torch.nonzero((target[:, label_idx] >= cutmix_thr), as_tuple=False)\n        \n        # Skip if less than 2 samples with coconfidence score 1.0\n        if len(indices) < 2:\n            continue\n            \n        # Original Data\n        data_orig = data[indices]\n        \n        # Shuffle\n        shuffled_indices = torch.randperm(len(indices))\n        data_shuffled = data_orig[shuffled_indices]\n        \n        # CutMix augmentation logic\n        start = random.randint(0, margin) if random.choice([True, False]) else random.randint(300-max_size-margin, 300-max_size)\n        size = random.randint(min_size, max_size)\n        \n        # CutMix in Specs\n        cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n        \n        # CutMix in EEG Specs\n        if cut_eeg_spec:\n            start = 300 + 40 + start # Size + Padding + Start\n            cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n            \n    return cutmix_data, target\n```\n\nModel | OOF-CV | LB\nEffB0  | 0.6642    | 0.41\nEffB0 + CutMix (p=0.4)  | 0.6530    | 0.39\n\nHere are the list of augmentations that didn't work for ME:\n\n1. TimeAxis Cutout\n2. FrequencyAxis Cutout\n3. HFlip\n4. SpecTranslation \n5. Translate and Mix\n6. MixUP\n7. Same Class Mixup\n8. Spec Brightness Adjustment\n9. Random Noise\n\nThank you all for sharing your amazing work and insights, especially Chris and Tawara :)",
      "votes": 42
    },
    {
      "id": 2671683,
      "postDate": "2024-02-27T16:44:02.193Z",
      "content": "<p>Do you think the key is the cutmix_thr filter logic? Simple in-class MixUp doesn't work for me either, so it seems that the filter is necessary.</p>\n<p>I wonder if it has something to do with the apparent <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">dual-modality of training data</a>, and by mixing with a particular subset of the data, the data is less noisy overall.</p>",
      "rawMarkdown": "Do you think the key is the cutmix_thr filter logic? Simple in-class MixUp doesn't work for me either, so it seems that the filter is necessary.\n\nI wonder if it has something to do with the apparent [dual-modality of training data](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135), and by mixing with a particular subset of the data, the data is less noisy overall.",
      "votes": 3,
      "replies": [
        {
          "id": 2671802,
          "postDate": "2024-02-27T17:48:35.207Z",
          "content": "<p>According to my experiments;</p>\n<ol>\n<li>Cutmix Thr = 1 ; LB = 0.40</li>\n<li>Cutmix Thr= 0.75; LB = 0.39</li>\n<li>Cutmix Thr = 0.50; LB = 0.39 better</li>\n</ol>\n<p>About the dual modality of training data; Maybe 🤔 IDK </p>",
          "rawMarkdown": "According to my experiments;\n1. Cutmix Thr = 1 ; LB = 0.40\n2. Cutmix Thr= 0.75; LB = 0.39\n3. Cutmix Thr = 0.50; LB = 0.39 better\n\nAbout the dual modality of training data; Maybe 🤔 IDK ",
          "votes": 4
        }
      ]
    },
    {
      "id": 2667273,
      "postDate": "2024-02-25T03:20:40.797Z",
      "content": "<p>Very helpful！</p>",
      "rawMarkdown": "Very helpful！",
      "votes": 1
    },
    {
      "id": 2667198,
      "postDate": "2024-02-25T00:20:45.757Z",
      "content": "<p>Nice find!</p>",
      "rawMarkdown": "Nice find!",
      "votes": 1
    },
    {
      "id": 2667099,
      "postDate": "2024-02-24T21:04:09.790Z",
      "content": "<p>Cool. An interesting idea. It will be necessary to check how it works.</p>",
      "rawMarkdown": "Cool. An interesting idea. It will be necessary to check how it works.",
      "votes": 1
    },
    {
      "id": 2666907,
      "postDate": "2024-02-24T18:16:04.263Z",
      "content": "<p>Thanks for share. I've seen many people suggesting Hflip. You never know till you try but… is not flipping time kind  a weird from beggining?</p>",
      "rawMarkdown": "Thanks for share. I've seen many people suggesting Hflip. You never know till you try but... is not flipping time kind  a weird from beggining?",
      "votes": 1,
      "replies": [
        {
          "id": 2667514,
          "postDate": "2024-02-25T06:38:11.763Z",
          "content": "<p>I think Hfilp implies a prior: reverse time order does not affect classification. This sounds more like a type of anomaly detection (for example, there may be an abnormal peak). I'm not sure if this is correct. Anyway, in my experiments, Hfilp is indeed useful.</p>",
          "rawMarkdown": "I think Hfilp implies a prior: reverse time order does not affect classification. This sounds more like a type of anomaly detection (for example, there may be an abnormal peak). I'm not sure if this is correct. Anyway, in my experiments, Hfilp is indeed useful.",
          "votes": 4,
          "replies": [
            {
              "id": 2667801,
              "postDate": "2024-02-25T11:02:18.547Z",
              "content": "<p>As autor commented. For pure CNN it's allrigth. But for sequences… I'm actually still DL newbe but reversing time doesn't feel very reasonable.</p>",
              "rawMarkdown": "As autor commented. For pure CNN it's allrigth. But for sequences... I'm actually still DL newbe but reversing time doesn't feel very reasonable.",
              "votes": 3
            },
            {
              "id": 2669792,
              "postDate": "2024-02-26T13:46:39.487Z",
              "content": "<p>Oh, I haven't tried models trained with sequence data. I agree with you about this type.</p>",
              "rawMarkdown": "Oh, I haven't tried models trained with sequence data. I agree with you about this type."
            }
          ]
        }
      ]
    },
    {
      "id": 2666928,
      "postDate": "2024-02-24T18:28:01.657Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2671683,
      "author_name": "Ry",
      "author_url": "",
      "post_date": "2024-02-27T16:44:02.193000",
      "content": "<p>Do you think the key is the cutmix_thr filter logic? Simple in-class MixUp doesn't work for me either, so it seems that the filter is necessary.</p>\n<p>I wonder if it has something to do with the apparent <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135\" target=\"_blank\">dual-modality of training data</a>, and by mixing with a particular subset of the data, the data is less noisy overall.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2671802,
          "author_name": "Vikram Sandu",
          "author_url": "",
          "post_date": "2024-02-27T17:48:35.207000",
          "content": "<p>According to my experiments;</p>\n<ol>\n<li>Cutmix Thr = 1 ; LB = 0.40</li>\n<li>Cutmix Thr= 0.75; LB = 0.39</li>\n<li>Cutmix Thr = 0.50; LB = 0.39 better</li>\n</ol>\n<p>About the dual modality of training data; Maybe 🤔 IDK </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 2667273,
      "author_name": "Yuri Sun",
      "author_url": "",
      "post_date": "2024-02-25T03:20:40.797000",
      "content": "<p>Very helpful！</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2667198,
      "author_name": "Yan Teixeira",
      "author_url": "",
      "post_date": "2024-02-25T00:20:45.757000",
      "content": "<p>Nice find!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2667099,
      "author_name": "Zaakcii Ru",
      "author_url": "",
      "post_date": "2024-02-24T21:04:09.790000",
      "content": "<p>Cool. An interesting idea. It will be necessary to check how it works.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2666907,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-02-24T18:16:04.263000",
      "content": "<p>Thanks for share. I've seen many people suggesting Hflip. You never know till you try but… is not flipping time kind  a weird from beggining?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2667514,
          "author_name": "Koolo",
          "author_url": "",
          "post_date": "2024-02-25T06:38:11.763000",
          "content": "<p>I think Hfilp implies a prior: reverse time order does not affect classification. This sounds more like a type of anomaly detection (for example, there may be an abnormal peak). I'm not sure if this is correct. Anyway, in my experiments, Hfilp is indeed useful.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 2667801,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-02-25T11:02:18.547000",
              "content": "<p>As autor commented. For pure CNN it's allrigth. But for sequences… I'm actually still DL newbe but reversing time doesn't feel very reasonable.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2669792,
              "author_name": "Koolo",
              "author_url": "",
              "post_date": "2024-02-26T13:46:39.487000",
              "content": "<p>Oh, I haven't tried models trained with sequence data. I agree with you about this type.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2666928,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-02-24T18:28:01.657000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2666878": "CutMixing the samples from the SAME CLASSES where Inputs are proportionally cut and mixed but not labels. (Experimental)\n\nTo be clear, we are cutting a whole time window from image-1 and replacing it randomly in image-2 except the middle portion of the image.\n\n```python\n'''\n  This applies on the Batch\n'''\nimport random\n\ndef hbac_cutmix(data, \n                target,\n                cutmix_thr = 0.75, # Threshold to consider samples\n                margin=50, # Either Cut 50 from begining  or end\n                min_size=25, # Min Cutout Size\n                max_size=75, # Max Cutout Size\n                cut_eeg_spec = True # CutMix in EEG Specs\n               ):\n    \n    # CutMix Data\n    cutmix_data = data.clone()\n    \n    for label_idx in range(target.size(1)):\n        # Indices with a confidence score greater than cutmix_thr for particular target\n        indices = torch.nonzero((target[:, label_idx] >= cutmix_thr), as_tuple=False)\n        \n        # Skip if less than 2 samples with coconfidence score 1.0\n        if len(indices) < 2:\n            continue\n            \n        # Original Data\n        data_orig = data[indices]\n        \n        # Shuffle\n        shuffled_indices = torch.randperm(len(indices))\n        data_shuffled = data_orig[shuffled_indices]\n        \n        # CutMix augmentation logic\n        start = random.randint(0, margin) if random.choice([True, False]) else random.randint(300-max_size-margin, 300-max_size)\n        size = random.randint(min_size, max_size)\n        \n        # CutMix in Specs\n        cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n        \n        # CutMix in EEG Specs\n        if cut_eeg_spec:\n            start = 300 + 40 + start # Size + Padding + Start\n            cutmix_data[indices, :, start:start+size] = data_shuffled[:, :, :,start:start+size]\n            \n    return cutmix_data, target\n```\n\nModel | OOF-CV | LB\nEffB0  | 0.6642    | 0.41\nEffB0 + CutMix (p=0.4)  | 0.6530    | 0.39\n\nHere are the list of augmentations that didn't work for ME:\n\n1. TimeAxis Cutout\n2. FrequencyAxis Cutout\n3. HFlip\n4. SpecTranslation \n5. Translate and Mix\n6. MixUP\n7. Same Class Mixup\n8. Spec Brightness Adjustment\n9. Random Noise\n\nThank you all for sharing your amazing work and insights, especially Chris and Tawara :)",
    "2671683": "Do you think the key is the cutmix_thr filter logic? Simple in-class MixUp doesn't work for me either, so it seems that the filter is necessary.\n\nI wonder if it has something to do with the apparent [dual-modality of training data](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/477135), and by mixing with a particular subset of the data, the data is less noisy overall.",
    "2667273": "Very helpful！",
    "2667198": "Nice find!",
    "2667099": "Cool. An interesting idea. It will be necessary to check how it works.",
    "2666907": "Thanks for share. I've seen many people suggesting Hflip. You never know till you try but... is not flipping time kind  a weird from beggining?",
    "2666928": ""
  }
}