{
  "id": 127486,
  "title": "Cutmix-mixup instability--shakup?",
  "url": "/competitions/bengaliai-cv19/discussion/127486",
  "author_name": "",
  "post_date": "2020-01-24T06:37:49.516338Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>From my experiments it is shown that using cutmix &amp; mixup at (.5-.5 or .7-.3) probability as opposed to using normal training results in fluctuating validation metric &amp; loss; indeed much more unstable than normal training alone. With similar setup I am seeing ~.005 difference in LB.</p>\n\n<p>I wonder if u guys are also observing this? Do we increase the risk and magnitude of shakup when using cutmix &amp; mixup? How is your CV stability when using cutmix &amp; mixup?</p>\n\n<p>Any comments welcome, thx</p>",
  "messages": [
    {
      "id": "727880",
      "postDate": "01/24/2020 06:37:49",
      "content": "<p>From my experiments it is shown that using cutmix &amp; mixup at (.5-.5 or .7-.3) probability as opposed to using normal training results in fluctuating validation metric &amp; loss; indeed much more unstable than normal training alone. With similar setup I am seeing ~.005 difference in LB.</p>\n\n<p>I wonder if u guys are also observing this? Do we increase the risk and magnitude of shakup when using cutmix &amp; mixup? How is your CV stability when using cutmix &amp; mixup?</p>\n\n<p>Any comments welcome, thx</p>",
      "rawMarkdown": "From my experiments it is shown that using cutmix &amp; mixup at (.5-.5 or .7-.3) probability as opposed to using normal training results in fluctuating validation metric &amp; loss; indeed much more unstable than normal training alone. With similar setup I am seeing ~.005 difference in LB.\n\nI wonder if u guys are also observing this? Do we increase the risk and magnitude of shakup when using cutmix &amp; mixup? How is your CV stability when using cutmix &amp; mixup?\n\nAny comments welcome, thx",
      "votes": null
    },
    {
      "id": "728169",
      "postDate": "01/24/2020 13:28:24",
      "content": "<p>For me only cut mix works best. Combining cut mix with mixup results in poor performance. Not sure what is the reason for this maybe what you said... maybe I need to train more... =)</p>",
      "rawMarkdown": "For me only cut mix works best. Combining cut mix with mixup results in poor performance. Not sure what is the reason for this maybe what you said... maybe I need to train more... =)",
      "votes": null
    },
    {
      "id": "728238",
      "postDate": "01/24/2020 14:20:59",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Is there a specific implementation of cutmix are you using? I have a feeling a more custom type of cutmix would work well since we are dealing with written text.</p>",
      "rawMarkdown": "drhabib Is there a specific implementation of cutmix are you using? I have a feeling a more custom type of cutmix would work well since we are dealing with written text.",
      "votes": null
    },
    {
      "id": "728253",
      "postDate": "01/24/2020 14:39:33",
      "content": "<p>Hi <a href=\"/robikscube\">@robikscube</a>  =)\nI use very similar cutmix implementation from this excellent post <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\">https://www.kaggle.com/c/bengaliai-cv19/discussion/126504</a>. The only thing I modified is my <code>alpha</code> is not fix to <code>1.0</code> but its random from <code>0.8</code> to <code>1.0</code>. I got inspiration from <code>fastai</code> when they use mixup training they use random range for <code>alpha</code>. </p>\n\n<p>Again for <code>cutmix</code> I am not sure if varying <code>alpha</code> performs better than just simple fix <code>alpha = 1.0</code>. if I will run out of ideas perhaps I will do comparison =)   </p>",
      "rawMarkdown": "Hi @robikscube  =)\nI use very similar cutmix implementation from this excellent post https://www.kaggle.com/c/bengaliai-cv19/discussion/126504. The only thing I modified is my `alpha` is not fix to `1.0` but its random from `0.8` to `1.0`. I got inspiration from `fastai` when they use mixup training they use random range for `alpha`. \n\nAgain for `cutmix` I am not sure if varying `alpha` performs better than just simple fix `alpha = 1.0`. if I will run out of ideas perhaps I will do comparison =)",
      "votes": null
    },
    {
      "id": "728659",
      "postDate": "01/25/2020 03:29:19",
      "content": "<p>Thanks for pointing me back to that thread! Trying it out now. Super helpful as always.</p>",
      "rawMarkdown": "Thanks for pointing me back to that thread! Trying it out now. Super helpful as always.",
      "votes": null
    },
    {
      "id": "729415",
      "postDate": "01/26/2020 07:07:17",
      "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> how many epochs did you train for using Cutmix , and both Cutmix and mixup? </p>",
      "rawMarkdown": "Hi @drhabib how many epochs did you train for using Cutmix , and both Cutmix and mixup?",
      "votes": null
    },
    {
      "id": "729693",
      "postDate": "01/26/2020 14:31:47",
      "content": "<p>Hi <a href=\"/p4rallax\">@p4rallax</a> </p>\n\n<p>I first tried 100 epoch, but later reduced to 80 =) \nGood luck </p>",
      "rawMarkdown": "Hi @p4rallax \n\nI first tried 100 epoch, but later reduced to 80 =) \nGood luck",
      "votes": null
    },
    {
      "id": "1084421",
      "postDate": "11/20/2020 02:51:44",
      "content": "<p>now I have the instability phenomenon ,</p>",
      "rawMarkdown": "now I have the instability phenomenon ,",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1084421,
      "author_name": "xiangzidai",
      "author_url": "",
      "post_date": "11/20/2020 02:51:44",
      "content": "<p>now I have the instability phenomenon ,</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 728169,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "01/24/2020 13:28:24",
      "content": "<p>For me only cut mix works best. Combining cut mix with mixup results in poor performance. Not sure what is the reason for this maybe what you said... maybe I need to train more... =)</p>",
      "votes": null,
      "replies": [
        {
          "id": 728238,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "01/24/2020 14:20:59",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Is there a specific implementation of cutmix are you using? I have a feeling a more custom type of cutmix would work well since we are dealing with written text.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728253,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "01/24/2020 14:39:33",
          "content": "<p>Hi <a href=\"/robikscube\">@robikscube</a>  =)\nI use very similar cutmix implementation from this excellent post <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/126504\">https://www.kaggle.com/c/bengaliai-cv19/discussion/126504</a>. The only thing I modified is my <code>alpha</code> is not fix to <code>1.0</code> but its random from <code>0.8</code> to <code>1.0</code>. I got inspiration from <code>fastai</code> when they use mixup training they use random range for <code>alpha</code>. </p>\n\n<p>Again for <code>cutmix</code> I am not sure if varying <code>alpha</code> performs better than just simple fix <code>alpha = 1.0</code>. if I will run out of ideas perhaps I will do comparison =)   </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728659,
          "author_name": "robikscube",
          "author_url": "",
          "post_date": "01/25/2020 03:29:19",
          "content": "<p>Thanks for pointing me back to that thread! Trying it out now. Super helpful as always.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729415,
          "author_name": "p4rallax",
          "author_url": "",
          "post_date": "01/26/2020 07:07:17",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> how many epochs did you train for using Cutmix , and both Cutmix and mixup? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 729693,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "01/26/2020 14:31:47",
          "content": "<p>Hi <a href=\"/p4rallax\">@p4rallax</a> </p>\n\n<p>I first tried 100 epoch, but later reduced to 80 =) \nGood luck </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "727880": "From my experiments it is shown that using cutmix &amp; mixup at (.5-.5 or .7-.3) probability as opposed to using normal training results in fluctuating validation metric &amp; loss; indeed much more unstable than normal training alone. With similar setup I am seeing ~.005 difference in LB.\n\nI wonder if u guys are also observing this? Do we increase the risk and magnitude of shakup when using cutmix &amp; mixup? How is your CV stability when using cutmix &amp; mixup?\n\nAny comments welcome, thx",
    "728169": "For me only cut mix works best. Combining cut mix with mixup results in poor performance. Not sure what is the reason for this maybe what you said... maybe I need to train more... =)",
    "728238": "drhabib Is there a specific implementation of cutmix are you using? I have a feeling a more custom type of cutmix would work well since we are dealing with written text.",
    "728253": "Hi @robikscube  =)\nI use very similar cutmix implementation from this excellent post https://www.kaggle.com/c/bengaliai-cv19/discussion/126504. The only thing I modified is my `alpha` is not fix to `1.0` but its random from `0.8` to `1.0`. I got inspiration from `fastai` when they use mixup training they use random range for `alpha`. \n\nAgain for `cutmix` I am not sure if varying `alpha` performs better than just simple fix `alpha = 1.0`. if I will run out of ideas perhaps I will do comparison =)",
    "728659": "Thanks for pointing me back to that thread! Trying it out now. Super helpful as always.",
    "729415": "Hi @drhabib how many epochs did you train for using Cutmix , and both Cutmix and mixup?",
    "729693": "Hi @p4rallax \n\nI first tried 100 epoch, but later reduced to 80 =) \nGood luck",
    "1084421": "now I have the instability phenomenon ,"
  },
  "source": "meta"
}