{
  "id": 132119,
  "title": "confusion about mixup/cutmix",
  "url": "/competitions/bengaliai-cv19/discussion/132119",
  "author_name": "",
  "post_date": "2020-02-24T09:17:48.293997100Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>when I apply mixup and cutmix to my model ,my cv score became very low(ablou 0.5),even though training for lots of epochs(70),I was wandering if there are something I should pay attention to when apply mixup or cutmix.The main code are as follows,hope someone could help me handle it.THX.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fb1bbdf1777b34a49ca7a16005965021b%2F1.PNG?generation=1582535667598838&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fd1ae11dfa28c3c6801fa0536366ec041%2F2.PNG?generation=1582535670112624&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2F71a38f4d082379804552005b5470b466%2F4.PNG?generation=1582535669233412&amp;alt=media\" alt=\"\">\nclassifer code:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fa5f93d560eeb093617eeade1ad9463aa%2F1.PNG?generation=1582538026560022&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "754959",
      "postDate": "02/24/2020 09:17:48",
      "content": "<p>when I apply mixup and cutmix to my model ,my cv score became very low(ablou 0.5),even though training for lots of epochs(70),I was wandering if there are something I should pay attention to when apply mixup or cutmix.The main code are as follows,hope someone could help me handle it.THX.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fb1bbdf1777b34a49ca7a16005965021b%2F1.PNG?generation=1582535667598838&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fd1ae11dfa28c3c6801fa0536366ec041%2F2.PNG?generation=1582535670112624&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2F71a38f4d082379804552005b5470b466%2F4.PNG?generation=1582535669233412&amp;alt=media\" alt=\"\">\nclassifer code:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fa5f93d560eeb093617eeade1ad9463aa%2F1.PNG?generation=1582538026560022&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "when I apply mixup and cutmix to my model ,my cv score became very low(ablou 0.5),even though training for lots of epochs(70),I was wandering if there are something I should pay attention to when apply mixup or cutmix.The main code are as follows,hope someone could help me handle it.THX.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fb1bbdf1777b34a49ca7a16005965021b%2F1.PNG?generation=1582535667598838&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fd1ae11dfa28c3c6801fa0536366ec041%2F2.PNG?generation=1582535670112624&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2F71a38f4d082379804552005b5470b466%2F4.PNG?generation=1582535669233412&amp;alt=media)\nclassifer code:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fa5f93d560eeb093617eeade1ad9463aa%2F1.PNG?generation=1582538026560022&amp;alt=media)",
      "votes": null
    },
    {
      "id": "754972",
      "postDate": "02/24/2020 09:37:16",
      "content": "<p>bug #1: your consonant loss seems a bit wrong. In the 2nd part (1-lam) CE y[4] should be y[5].\nWhat is the shape of <code>x</code> before you call cutmix/mixup?</p>",
      "rawMarkdown": "bug #1: your consonant loss seems a bit wrong. In the 2nd part (1-lam) CE y[4] should be y[5].\nWhat is the shape of `x` before you call cutmix/mixup?",
      "votes": null
    },
    {
      "id": "754982",
      "postDate": "02/24/2020 09:54:13",
      "content": "<p>Oh yes ,it's a clerical error,I have changed the screenshot.The batch_size I used is 768  and the shape of x is torch.Size([768, 1, 64, 64])</p>",
      "rawMarkdown": "Oh yes ,it's a clerical error,I have changed the screenshot.The batch_size I used is 768  and the shape of x is torch.Size([768, 1, 64, 64])",
      "votes": null
    },
    {
      "id": "754993",
      "postDate": "02/24/2020 10:13:26",
      "content": "<p>If the shape is 768x1x64x64, then there is a bug in the rand_bbox too:</p>\n\n<p><code>\nW = size[1]\nH = size[2]\n</code>\nshould be:\n<code>\nW = size[2]\nH = size[3]\n</code></p>",
      "rawMarkdown": "If the shape is 768x1x64x64, then there is a bug in the rand_bbox too:\n\n```\nW = size[1]\nH = size[2]\n```\nshould be:\n```\nW = size[2]\nH = size[3]\n```",
      "votes": null
    },
    {
      "id": "755011",
      "postDate": "02/24/2020 10:39:49",
      "content": "<p>seems like I'm not the only one who faced the same problem. acc &lt; 0.65 after 100 epochs.</p>",
      "rawMarkdown": "seems like I'm not the only one who faced the same problem. acc &lt; 0.65 after 100 epochs.",
      "votes": null
    },
    {
      "id": "755024",
      "postDate": "02/24/2020 10:58:49",
      "content": "<p>wow,I didn't notice it.Thank you ,I will change it and retry</p>",
      "rawMarkdown": "wow,I didn't notice it.Thank you ,I will change it and retry",
      "votes": null
    },
    {
      "id": "755075",
      "postDate": "02/24/2020 12:01:52",
      "content": "<p><a href=\"/tiandaye\">@tiandaye</a> <a href=\"/thefatcat\">@thefatcat</a> The problem is that you are also calling cutmix &amp; mixup on validation data. While cutmixup should be called on training data, it should not be used during validation for a meaningful CV. Hope this helps</p>\n\n<p>Also, according to the definition of beta distribution, assigned lambdas are randomly &gt;.5. Taken the best case that your model perfectly fits cutmixup, this only corresponds to a maximum accuracy of .5 and marginally better macro-recall. Therefore the training loss is a better indicator of overfitting than the training metric (though the training loss is not too informative, too). The training metric is basically rendered useless by use of cutmixup.</p>",
      "rawMarkdown": "tiandaye @thefatcat The problem is that you are also calling cutmix &amp; mixup on validation data. While cutmixup should be called on training data, it should not be used during validation for a meaningful CV. Hope this helps\n\nAlso, according to the definition of beta distribution, assigned lambdas are randomly &gt;.5. Taken the best case that your model perfectly fits cutmixup, this only corresponds to a maximum accuracy of .5 and marginally better macro-recall. Therefore the training loss is a better indicator of overfitting than the training metric (though the training loss is not too informative, too). The training metric is basically rendered useless by use of cutmixup.",
      "votes": null
    },
    {
      "id": "755141",
      "postDate": "02/24/2020 13:41:15",
      "content": "<p>Thanks for your advice! My cv score improve significantly after I cancel mixup and cutmix on validation data.  </p>",
      "rawMarkdown": "Thanks for your advice! My cv score improve significantly after I cancel mixup and cutmix on validation data.",
      "votes": null
    },
    {
      "id": "755145",
      "postDate": "02/24/2020 13:42:51",
      "content": "<p>haha,try a few more times</p>",
      "rawMarkdown": "haha,try a few more times",
      "votes": null
    },
    {
      "id": "755159",
      "postDate": "02/24/2020 14:03:03",
      "content": "<p>Glad it helps👍</p>",
      "rawMarkdown": "Glad it helps👍",
      "votes": null
    },
    {
      "id": "755166",
      "postDate": "02/24/2020 14:08:41",
      "content": "<p>training losses make no sense, you should see the cv score without cutmix/mixup.</p>",
      "rawMarkdown": "training losses make no sense, you should see the cv score without cutmix/mixup.",
      "votes": null
    },
    {
      "id": "755172",
      "postDate": "02/24/2020 14:15:47",
      "content": "<p>yean,I'm trying to train my model by the way you said,thx:)</p>",
      "rawMarkdown": "yean,I'm trying to train my model by the way you said,thx:)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 754972,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "02/24/2020 09:37:16",
      "content": "<p>bug #1: your consonant loss seems a bit wrong. In the 2nd part (1-lam) CE y[4] should be y[5].\nWhat is the shape of <code>x</code> before you call cutmix/mixup?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 754982,
      "author_name": "thefatcat",
      "author_url": "",
      "post_date": "02/24/2020 09:54:13",
      "content": "<p>Oh yes ,it's a clerical error,I have changed the screenshot.The batch_size I used is 768  and the shape of x is torch.Size([768, 1, 64, 64])</p>",
      "votes": null,
      "replies": [
        {
          "id": 754993,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/24/2020 10:13:26",
          "content": "<p>If the shape is 768x1x64x64, then there is a bug in the rand_bbox too:</p>\n\n<p><code>\nW = size[1]\nH = size[2]\n</code>\nshould be:\n<code>\nW = size[2]\nH = size[3]\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755024,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "02/24/2020 10:58:49",
          "content": "<p>wow,I didn't notice it.Thank you ,I will change it and retry</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755011,
      "author_name": "tiandaye",
      "author_url": "",
      "post_date": "02/24/2020 10:39:49",
      "content": "<p>seems like I'm not the only one who faced the same problem. acc &lt; 0.65 after 100 epochs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 755145,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "02/24/2020 13:42:51",
          "content": "<p>haha,try a few more times</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755075,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "02/24/2020 12:01:52",
      "content": "<p><a href=\"/tiandaye\">@tiandaye</a> <a href=\"/thefatcat\">@thefatcat</a> The problem is that you are also calling cutmix &amp; mixup on validation data. While cutmixup should be called on training data, it should not be used during validation for a meaningful CV. Hope this helps</p>\n\n<p>Also, according to the definition of beta distribution, assigned lambdas are randomly &gt;.5. Taken the best case that your model perfectly fits cutmixup, this only corresponds to a maximum accuracy of .5 and marginally better macro-recall. Therefore the training loss is a better indicator of overfitting than the training metric (though the training loss is not too informative, too). The training metric is basically rendered useless by use of cutmixup.</p>",
      "votes": null,
      "replies": [
        {
          "id": 755141,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "02/24/2020 13:41:15",
          "content": "<p>Thanks for your advice! My cv score improve significantly after I cancel mixup and cutmix on validation data.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755159,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/24/2020 14:03:03",
          "content": "<p>Glad it helps👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 755166,
      "author_name": "masterzj",
      "author_url": "",
      "post_date": "02/24/2020 14:08:41",
      "content": "<p>training losses make no sense, you should see the cv score without cutmix/mixup.</p>",
      "votes": null,
      "replies": [
        {
          "id": 755172,
          "author_name": "thefatcat",
          "author_url": "",
          "post_date": "02/24/2020 14:15:47",
          "content": "<p>yean,I'm trying to train my model by the way you said,thx:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "754959": "when I apply mixup and cutmix to my model ,my cv score became very low(ablou 0.5),even though training for lots of epochs(70),I was wandering if there are something I should pay attention to when apply mixup or cutmix.The main code are as follows,hope someone could help me handle it.THX.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fb1bbdf1777b34a49ca7a16005965021b%2F1.PNG?generation=1582535667598838&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fd1ae11dfa28c3c6801fa0536366ec041%2F2.PNG?generation=1582535670112624&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2F71a38f4d082379804552005b5470b466%2F4.PNG?generation=1582535669233412&amp;alt=media)\nclassifer code:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2969049%2Fa5f93d560eeb093617eeade1ad9463aa%2F1.PNG?generation=1582538026560022&amp;alt=media)",
    "754972": "bug #1: your consonant loss seems a bit wrong. In the 2nd part (1-lam) CE y[4] should be y[5].\nWhat is the shape of `x` before you call cutmix/mixup?",
    "754982": "Oh yes ,it's a clerical error,I have changed the screenshot.The batch_size I used is 768  and the shape of x is torch.Size([768, 1, 64, 64])",
    "754993": "If the shape is 768x1x64x64, then there is a bug in the rand_bbox too:\n\n```\nW = size[1]\nH = size[2]\n```\nshould be:\n```\nW = size[2]\nH = size[3]\n```",
    "755011": "seems like I'm not the only one who faced the same problem. acc &lt; 0.65 after 100 epochs.",
    "755024": "wow,I didn't notice it.Thank you ,I will change it and retry",
    "755075": "tiandaye @thefatcat The problem is that you are also calling cutmix &amp; mixup on validation data. While cutmixup should be called on training data, it should not be used during validation for a meaningful CV. Hope this helps\n\nAlso, according to the definition of beta distribution, assigned lambdas are randomly &gt;.5. Taken the best case that your model perfectly fits cutmixup, this only corresponds to a maximum accuracy of .5 and marginally better macro-recall. Therefore the training loss is a better indicator of overfitting than the training metric (though the training loss is not too informative, too). The training metric is basically rendered useless by use of cutmixup.",
    "755141": "Thanks for your advice! My cv score improve significantly after I cancel mixup and cutmix on validation data.",
    "755145": "haha,try a few more times",
    "755159": "Glad it helps👍",
    "755166": "training losses make no sense, you should see the cv score without cutmix/mixup.",
    "755172": "yean,I'm trying to train my model by the way you said,thx:)"
  },
  "source": "meta"
}