{
  "id": 266621,
  "title": "Does Mixup Help with New Data?",
  "url": "/competitions/seti-breakthrough-listen/discussion/266621",
  "author_name": "Chris Deotte",
  "post_date": "2021-08-19T18:33:02.381000",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<h2>Question: Does Mixup Help?</h2>\n<p>I started the comp after the data reset. Using only the new data, I did not see a CV improvement using mixup versus not using mixup. (Update: i didn't do enough experiments during the comp and only used low values of alpha) Therefore, i did not use mixup. However now, i'm curious if mixup was an important technique to decrease the CV LB gap? (My model without mixup had a huge 0.125 CV LB gap described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266534\" target=\"_blank\">here</a>).</p>\n<p>Reading old discussions, it appears that mixup did improve CV score before competition reset (i.e. with the original old train data) by a huge 7% described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245152\" target=\"_blank\">here</a>. So perhaps teams that began this comp before reset used mixup and continued using mixup after reset and did not check the CV score with and without?</p>\n<p>Does anyone have statistics on how much mixup helped CV LB before competition reset? And how much mixup helps CV LB after competition reset (i.e. using only new data)? Also did anyone have a high scoring LB model that did not use mixup?</p>\n<h2>Answer: Yes, Mixup is the Magic #0!</h2>\n<p>Mixup is the Magic #0 to achieve Silver medal. And Magic's #1 and #2 to achieve Gold medal are described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a> (in first place solution)<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. </p>\n<p>Fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
  "messages": [
    {
      "id": 1481930,
      "postDate": "2021-08-19T18:33:02.380Z",
      "content": "<h2>Question: Does Mixup Help?</h2>\n<p>I started the comp after the data reset. Using only the new data, I did not see a CV improvement using mixup versus not using mixup. (Update: i didn't do enough experiments during the comp and only used low values of alpha) Therefore, i did not use mixup. However now, i'm curious if mixup was an important technique to decrease the CV LB gap? (My model without mixup had a huge 0.125 CV LB gap described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266534\" target=\"_blank\">here</a>).</p>\n<p>Reading old discussions, it appears that mixup did improve CV score before competition reset (i.e. with the original old train data) by a huge 7% described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245152\" target=\"_blank\">here</a>. So perhaps teams that began this comp before reset used mixup and continued using mixup after reset and did not check the CV score with and without?</p>\n<p>Does anyone have statistics on how much mixup helped CV LB before competition reset? And how much mixup helps CV LB after competition reset (i.e. using only new data)? Also did anyone have a high scoring LB model that did not use mixup?</p>\n<h2>Answer: Yes, Mixup is the Magic #0!</h2>\n<p>Mixup is the Magic #0 to achieve Silver medal. And Magic's #1 and #2 to achieve Gold medal are described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a> (in first place solution)<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. </p>\n<p>Fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
      "rawMarkdown": "## Question: Does Mixup Help?\nI started the comp after the data reset. Using only the new data, I did not see a CV improvement using mixup versus not using mixup. (Update: i didn't do enough experiments during the comp and only used low values of alpha) Therefore, i did not use mixup. However now, i'm curious if mixup was an important technique to decrease the CV LB gap? (My model without mixup had a huge 0.125 CV LB gap described [here][1]).\n\nReading old discussions, it appears that mixup did improve CV score before competition reset (i.e. with the original old train data) by a huge 7% described [here][2]. So perhaps teams that began this comp before reset used mixup and continued using mixup after reset and did not check the CV score with and without?\n\nDoes anyone have statistics on how much mixup helped CV LB before competition reset? And how much mixup helps CV LB after competition reset (i.e. using only new data)? Also did anyone have a high scoring LB model that did not use mixup?\n\n## Answer: Yes, Mixup is the Magic #0! \nMixup is the Magic #0 to achieve Silver medal. And Magic's #1 and #2 to achieve Gold medal are described [here][3] (in first place solution)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. \n\nFold 0 CV is 0.894. Therefore CV LB gap is 0.008!\n\n[3]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266534\n[2]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245152",
      "votes": 9
    },
    {
      "id": 1483654,
      "postDate": "2021-08-20T18:47:33.183Z",
      "content": "<p>I cant remember exactly, but it also gave us very significant boosts on CV. Once I tried removing it, and the scores dropped a lot. Try to mix images very equally (around 50-50), always apply mixup, take the maximum of the target, and optionally re-normalize after mixing and you should see good results.</p>",
      "rawMarkdown": "I cant remember exactly, but it also gave us very significant boosts on CV. Once I tried removing it, and the scores dropped a lot. Try to mix images very equally (around 50-50), always apply mixup, take the maximum of the target, and optionally re-normalize after mixing and you should see good results.",
      "votes": 3
    },
    {
      "id": 1487171,
      "postDate": "2021-08-23T13:04:38.193Z",
      "content": "<p>UPDATE: <strong>Mixup is the Magic #0 to achieve Silver medal</strong>. And Magic's #1 and #2 to achieve Gold medal are described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a> (in 1st place solution)<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make to beat LB 780 and achieve Silver medal. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. </p>\n<p>Fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
      "rawMarkdown": "UPDATE: **Mixup is the Magic #0 to achieve Silver medal**. And Magic's #1 and #2 to achieve Gold medal are described [here][1] (in 1st place solution)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make to beat LB 780 and achieve Silver medal. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. \n\nFold 0 CV is 0.894. Therefore CV LB gap is 0.008!\n\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385",
      "votes": 1
    },
    {
      "id": 1483184,
      "postDate": "2021-08-20T13:47:07.690Z",
      "content": "<p>Every experiment i did to lower or alternate down from 100% mixup with alpha=2. failed (lower CV). My mixup was kind of 'OR' and with respect of crop with ABACAD (i was not mixing between ON and OFF channels). Crop was done 'batchwise' </p>",
      "rawMarkdown": "Every experiment i did to lower or alternate down from 100% mixup with alpha=2. failed (lower CV). My mixup was kind of 'OR' and with respect of crop with ABACAD (i was not mixing between ON and OFF channels). Crop was done 'batchwise' ",
      "votes": 1,
      "replies": [
        {
          "id": 1483222,
          "postDate": "2021-08-20T14:10:05.433Z",
          "content": "<p>Thanks Gleb. I'm running experiments now. I want to learn what my models were missing.</p>\n<p>You say \"crop was done batchwise\", did you use cutmix or mixup?</p>",
          "rawMarkdown": "Thanks Gleb. I'm running experiments now. I want to learn what my models were missing.\n\nYou say \"crop was done batchwise\", did you use cutmix or mixup?"
        },
        {
          "id": 1483299,
          "postDate": "2021-08-20T15:04:46.557Z",
          "content": "<p>Simplified version of my data pipeline:</p>\n<ol>\n<li>Load batch Nx1xHxW</li>\n<li>Add guide layer for ABACAD frames: Nx2xHxW</li>\n<li>Batch-wise crop N x 2 x h x w</li>\n<li>separate anomaly cutmix on n x 1 x h x w and (N-n) x 1 x h x w subbatches ([n] == num of anomaly samples). cutmix crop was also done batch-wise. </li>\n<li>'or' mixup </li>\n</ol>\n<p>I use both (at first it was one of them, but without mixup on every batch CV was lower)</p>",
          "rawMarkdown": "Simplified version of my data pipeline:\n1. Load batch Nx1xHxW\n2. Add guide layer for ABACAD frames: Nx2xHxW\n3. Batch-wise crop N x 2 x h x w\n4. separate anomaly cutmix on n x 1 x h x w and (N-n) x 1 x h x w subbatches ([n] == num of anomaly samples). cutmix crop was also done batch-wise. \n5. 'or' mixup \n\nI use both (at first it was one of them, but without mixup on every batch CV was lower)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1482078,
      "postDate": "2021-08-19T21:30:26.837Z",
      "content": "<p>I didn't bother submitting without mixup since the reset, but to give you a feel, here is an otherwise identically trained tf_efficientnet_b5_ns model trained at 512 * 512 resolution (640*640 test) without (purple) and with (yellow) mixup: <a href=\"https://i.ibb.co/q1D8NH8/mixup.png\" target=\"_blank\">image</a></p>\n<p>I <em>DID</em> have a very high ranking pre-reset (maybe 3rd place at one point?) and I found at that stage I was able to get CVs of 0.97 without mixup, and mixup didn't help. But because of the leak I don't think it's easy to drawn any conclusions from, though.</p>",
      "rawMarkdown": "I didn't bother submitting without mixup since the reset, but to give you a feel, here is an otherwise identically trained tf_efficientnet_b5_ns model trained at 512 * 512 resolution (640*640 test) without (purple) and with (yellow) mixup: [image](https://i.ibb.co/q1D8NH8/mixup.png)\n\nI _DID_ have a very high ranking pre-reset (maybe 3rd place at one point?) and I found at that stage I was able to get CVs of 0.97 without mixup, and mixup didn't help. But because of the leak I don't think it's easy to drawn any conclusions from, though.",
      "votes": 1,
      "replies": [
        {
          "id": 1482155,
          "postDate": "2021-08-19T22:43:45.660Z",
          "content": "<p>Thanks for the image. For reference, i show it below. I'm confused what I'm looking at. I think the axis are labeled wrong. Maybe the bottom two plots are loss and the top two are AUC? What are the values on the y axis? Are you saying that mixup helped your local CV score? Perhaps i did not implement mixup correctly.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup.png\" alt=\"\"></p>",
          "rawMarkdown": "Thanks for the image. For reference, i show it below. I'm confused what I'm looking at. I think the axis are labeled wrong. Maybe the bottom two plots are loss and the top two are AUC? What are the values on the y axis? Are you saying that mixup helped your local CV score? Perhaps i did not implement mixup correctly.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup.png)",
          "votes": 1
        },
        {
          "id": 1482157,
          "postDate": "2021-08-19T22:44:32.127Z",
          "content": "<p>I'm trying to implement mixup correctly now with my models because i want to see how much it will close the CV LB gap.</p>",
          "rawMarkdown": "I'm trying to implement mixup correctly now with my models because i want to see how much it will close the CV LB gap."
        },
        {
          "id": 1482626,
          "postDate": "2021-08-20T07:10:24.590Z",
          "content": "<p>The axes are correct, but maybe it's confusing because I plot \"1-AUC\" (see the title above) so that I can plot it on a log scale to see continual benefits better.</p>\n<p>The top 2 plots are the training data (loss left, 1-AUC right), and bottom 2 validation (loss left, 1-AUC right).</p>\n<p>You can see after 10-12 epochs, without mixup (purple) the validation loss and AUC start rising as it overfits significantly.</p>\n<p>Adding mixup results in slower training (yellow), but after 12 epochs it's caught up and continues to improve substantially.</p>",
          "rawMarkdown": "The axes are correct, but maybe it's confusing because I plot \"1-AUC\" (see the title above) so that I can plot it on a log scale to see continual benefits better.\n\nThe top 2 plots are the training data (loss left, 1-AUC right), and bottom 2 validation (loss left, 1-AUC right).\n\nYou can see after 10-12 epochs, without mixup (purple) the validation loss and AUC start rising as it overfits significantly.\n\nAdding mixup results in slower training (yellow), but after 12 epochs it's caught up and continues to improve substantially.",
          "votes": 1
        },
        {
          "id": 1482755,
          "postDate": "2021-08-20T08:51:34.590Z",
          "content": "<p>After the reset i gave up on mixup just because i noticed the slow learning curve (I was also at the top of the PB at the time using it), I thinks that's why ended up 50+ position lower ahaha. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> let us know your findings! I'm short gpus resources rn :(</p>",
          "rawMarkdown": "After the reset i gave up on mixup just because i noticed the slow learning curve (I was also at the top of the PB at the time using it), I thinks that's why ended up 50+ position lower ahaha. @cdeotte let us know your findings! I'm short gpus resources rn :(",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1483654,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2021-08-20T18:47:33.183000",
      "content": "<p>I cant remember exactly, but it also gave us very significant boosts on CV. Once I tried removing it, and the scores dropped a lot. Try to mix images very equally (around 50-50), always apply mixup, take the maximum of the target, and optionally re-normalize after mixing and you should see good results.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1487171,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2021-08-23T13:04:38.193000",
      "content": "<p>UPDATE: <strong>Mixup is the Magic #0 to achieve Silver medal</strong>. And Magic's #1 and #2 to achieve Gold medal are described <a href=\"https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\" target=\"_blank\">here</a> (in 1st place solution)<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png\" alt=\"\"><br>\nThe above score of <strong>Public LB 0.786 and Private LB 781</strong> is the simplest model you can make to beat LB 780 and achieve Silver medal. It just uses new train data with <code>np.vstack( img[::2] )</code>, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. </p>\n<p>Fold 0 CV is 0.894. Therefore CV LB gap is 0.008!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1483184,
      "author_name": "Gleb",
      "author_url": "",
      "post_date": "2021-08-20T13:47:07.690000",
      "content": "<p>Every experiment i did to lower or alternate down from 100% mixup with alpha=2. failed (lower CV). My mixup was kind of 'OR' and with respect of crop with ABACAD (i was not mixing between ON and OFF channels). Crop was done 'batchwise' </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1483222,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-20T14:10:05.433000",
          "content": "<p>Thanks Gleb. I'm running experiments now. I want to learn what my models were missing.</p>\n<p>You say \"crop was done batchwise\", did you use cutmix or mixup?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1483299,
          "author_name": "Gleb",
          "author_url": "",
          "post_date": "2021-08-20T15:04:46.557000",
          "content": "<p>Simplified version of my data pipeline:</p>\n<ol>\n<li>Load batch Nx1xHxW</li>\n<li>Add guide layer for ABACAD frames: Nx2xHxW</li>\n<li>Batch-wise crop N x 2 x h x w</li>\n<li>separate anomaly cutmix on n x 1 x h x w and (N-n) x 1 x h x w subbatches ([n] == num of anomaly samples). cutmix crop was also done batch-wise. </li>\n<li>'or' mixup </li>\n</ol>\n<p>I use both (at first it was one of them, but without mixup on every batch CV was lower)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1482078,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2021-08-19T21:30:26.837000",
      "content": "<p>I didn't bother submitting without mixup since the reset, but to give you a feel, here is an otherwise identically trained tf_efficientnet_b5_ns model trained at 512 * 512 resolution (640*640 test) without (purple) and with (yellow) mixup: <a href=\"https://i.ibb.co/q1D8NH8/mixup.png\" target=\"_blank\">image</a></p>\n<p>I <em>DID</em> have a very high ranking pre-reset (maybe 3rd place at one point?) and I found at that stage I was able to get CVs of 0.97 without mixup, and mixup didn't help. But because of the leak I don't think it's easy to drawn any conclusions from, though.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1482155,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T22:43:45.660000",
          "content": "<p>Thanks for the image. For reference, i show it below. I'm confused what I'm looking at. I think the axis are labeled wrong. Maybe the bottom two plots are loss and the top two are AUC? What are the values on the y axis? Are you saying that mixup helped your local CV score? Perhaps i did not implement mixup correctly.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup.png\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482157,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2021-08-19T22:44:32.127000",
          "content": "<p>I'm trying to implement mixup correctly now with my models because i want to see how much it will close the CV LB gap.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1482626,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2021-08-20T07:10:24.590000",
          "content": "<p>The axes are correct, but maybe it's confusing because I plot \"1-AUC\" (see the title above) so that I can plot it on a log scale to see continual benefits better.</p>\n<p>The top 2 plots are the training data (loss left, 1-AUC right), and bottom 2 validation (loss left, 1-AUC right).</p>\n<p>You can see after 10-12 epochs, without mixup (purple) the validation loss and AUC start rising as it overfits significantly.</p>\n<p>Adding mixup results in slower training (yellow), but after 12 epochs it's caught up and continues to improve substantially.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1482755,
          "author_name": "Riccardo",
          "author_url": "",
          "post_date": "2021-08-20T08:51:34.590000",
          "content": "<p>After the reset i gave up on mixup just because i noticed the slow learning curve (I was also at the top of the PB at the time using it), I thinks that's why ended up 50+ position lower ahaha. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> let us know your findings! I'm short gpus resources rn :(</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1481930": "## Question: Does Mixup Help?\nI started the comp after the data reset. Using only the new data, I did not see a CV improvement using mixup versus not using mixup. (Update: i didn't do enough experiments during the comp and only used low values of alpha) Therefore, i did not use mixup. However now, i'm curious if mixup was an important technique to decrease the CV LB gap? (My model without mixup had a huge 0.125 CV LB gap described [here][1]).\n\nReading old discussions, it appears that mixup did improve CV score before competition reset (i.e. with the original old train data) by a huge 7% described [here][2]. So perhaps teams that began this comp before reset used mixup and continued using mixup after reset and did not check the CV score with and without?\n\nDoes anyone have statistics on how much mixup helped CV LB before competition reset? And how much mixup helps CV LB after competition reset (i.e. using only new data)? Also did anyone have a high scoring LB model that did not use mixup?\n\n## Answer: Yes, Mixup is the Magic #0! \nMixup is the Magic #0 to achieve Silver medal. And Magic's #1 and #2 to achieve Gold medal are described [here][3] (in first place solution)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. \n\nFold 0 CV is 0.894. Therefore CV LB gap is 0.008!\n\n[3]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266534\n[2]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/245152",
    "1483654": "I cant remember exactly, but it also gave us very significant boosts on CV. Once I tried removing it, and the scores dropped a lot. Try to mix images very equally (around 50-50), always apply mixup, take the maximum of the target, and optionally re-normalize after mixing and you should see good results.",
    "1487171": "UPDATE: **Mixup is the Magic #0 to achieve Silver medal**. And Magic's #1 and #2 to achieve Gold medal are described [here][1] (in 1st place solution)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Aug-2021/mixup3.png)\nThe above score of **Public LB 0.786 and Private LB 781** is the simplest model you can make to beat LB 780 and achieve Silver medal. It just uses new train data with `np.vstack( img[::2] )`, and trains 40 epochs with cosine schedule with warmup. It only uses Hflip, Vflip, and Mixup (alpha=3, max target) augmentation. That's it nothing else, nothing fancy!! The score above is just 1 fold of image size 768x768 and EfficientNetB4. \n\nFold 0 CV is 0.894. Therefore CV LB gap is 0.008!\n\n[1]: https://www.kaggle.com/c/seti-breakthrough-listen/discussion/266385",
    "1483184": "Every experiment i did to lower or alternate down from 100% mixup with alpha=2. failed (lower CV). My mixup was kind of 'OR' and with respect of crop with ABACAD (i was not mixing between ON and OFF channels). Crop was done 'batchwise' ",
    "1482078": "I didn't bother submitting without mixup since the reset, but to give you a feel, here is an otherwise identically trained tf_efficientnet_b5_ns model trained at 512 * 512 resolution (640*640 test) without (purple) and with (yellow) mixup: [image](https://i.ibb.co/q1D8NH8/mixup.png)\n\nI _DID_ have a very high ranking pre-reset (maybe 3rd place at one point?) and I found at that stage I was able to get CVs of 0.97 without mixup, and mixup didn't help. But because of the leak I don't think it's easy to drawn any conclusions from, though."
  }
}