{
  "id": 126504,
  "title": "mixup/cutmix is all you need",
  "url": "/competitions/bengaliai-cv19/discussion/126504",
  "author_name": "MachineLP",
  "post_date": "2020-01-18T01:51:40.844000",
  "votes": 134,
  "comment_count": 57,
  "views": 0,
  "content": "<p>```python\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p>def cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\nbbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\ndata[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n# adjust lambda to exactly match pixel ratio\nlam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\nreturn data, targets\n</code></pre>\n\n<p>def mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\ndata = data * lam + shuffled_data * (1 - lam)\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\nreturn data, targets\n</code></pre>\n\n<p>def cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)</p>\n\n<p>def mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)\n```</p>",
  "messages": [
    {
      "id": 722013,
      "postDate": "2020-01-18T01:51:40.843Z",
      "content": "<p>```python\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p>def cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\nbbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\ndata[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n# adjust lambda to exactly match pixel ratio\nlam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\nreturn data, targets\n</code></pre>\n\n<p>def mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\ndata = data * lam + shuffled_data * (1 - lam)\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\nreturn data, targets\n</code></pre>\n\n<p>def cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)</p>\n\n<p>def mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)\n```</p>",
      "rawMarkdown": "```python\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n\n    return bbx1, bby1, bbx2, bby2\ndef cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    data[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n    return data, targets\n\ndef mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    data = data * lam + shuffled_data * (1 - lam)\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\n    return data, targets\n\n\ndef cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)\n\ndef mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)\n```",
      "votes": 133
    },
    {
      "id": 730097,
      "postDate": "2020-01-27T05:23:21.007Z",
      "content": "<p>@DrHB</p>\n\n<p>\" ... you should try to train on half precision \"</p>\n\n<p>if you are using half-precision for training, remember to switch back to full precision for the last finetunning. it may improve results</p>",
      "rawMarkdown": "@DrHB\n \n\" ... you should try to train on half precision \"\n\nif you are using half-precision for training, remember to switch back to full precision for the last finetunning. it may improve results",
      "votes": 20,
      "replies": [
        {
          "id": 730389,
          "postDate": "2020-01-27T12:24:17.623Z",
          "content": "<p>Ahhh!!! That’s smart! Thanks for the advice!</p>",
          "rawMarkdown": "Ahhh!!! That’s smart! Thanks for the advice!"
        },
        {
          "id": 754521,
          "postDate": "2020-02-23T17:23:11.187Z",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> <a href=\"/drhabib\">@drhabib</a>   what's the meaning of training on half precision?👀 </p>",
          "rawMarkdown": "Hi @hengck23 @drhabib   what's the meaning of training on half precision?👀 "
        },
        {
          "id": 754553,
          "postDate": "2020-02-23T18:13:11.890Z",
          "content": "<p>fp16 I guess. int 16??</p>",
          "rawMarkdown": "fp16 I guess. int 16??"
        }
      ]
    },
    {
      "id": 722191,
      "postDate": "2020-01-18T08:48:17.823Z",
      "content": "<p>Update:\n```\n......\nfor i, (image_id, images, label1, label2, label3) in enumerate(data_loader_train):\n            images = images.to(device)\n            label1 = label1.to(device)\n            label2 = label2.to(device)\n            label3 = label3.to(device)\n            # print (image_id, label1, label2, label3)</p>\n\n<pre><code>        if np.random.rand()&lt;0.5:\n            images, targets = mixup(images, label1, label2, label3, 0.4)\n            output1, output2, output3 = model(images)\n            loss = mixup_criterion(output1,output2,output3, targets) \n        else:\n            images, targets = cutmix(images, label1, label2, label3, 0.4)\n            output1, output2, output3 = model(images)\n            loss = cutmix_criterion(output1,output2,output3, targets) \n</code></pre>\n\n<p>......\n```</p>",
      "rawMarkdown": "Update:\n```\n......\nfor i, (image_id, images, label1, label2, label3) in enumerate(data_loader_train):\n            images = images.to(device)\n            label1 = label1.to(device)\n            label2 = label2.to(device)\n            label3 = label3.to(device)\n            # print (image_id, label1, label2, label3)\n           \n            if np.random.rand()&lt;0.5:\n                images, targets = mixup(images, label1, label2, label3, 0.4)\n                output1, output2, output3 = model(images)\n                loss = mixup_criterion(output1,output2,output3, targets) \n            else:\n                images, targets = cutmix(images, label1, label2, label3, 0.4)\n                output1, output2, output3 = model(images)\n                loss = cutmix_criterion(output1,output2,output3, targets) \n......\n```",
      "votes": 19,
      "replies": [
        {
          "id": 722242,
          "postDate": "2020-01-18T09:57:14.630Z",
          "content": "<p>Thanks for the implementation. I'm not too familiar with torch but keras user. I didn’t try out your code yet but would you please point out to me is there any changes I need to make of this code while using it keras. A pseudo code would be helpful. Thanks.</p>",
          "rawMarkdown": "Thanks for the implementation. I'm not too familiar with torch but keras user. I didn’t try out your code yet but would you please point out to me is there any changes I need to make of this code while using it keras. A pseudo code would be helpful. Thanks.",
          "votes": 1
        },
        {
          "id": 722268,
          "postDate": "2020-01-18T10:47:46.150Z",
          "content": "<p>Please check this repo friend \n<a href=\"https://github.com/DevBruce/CutMixImageDataGenerator_For_Keras\">https://github.com/DevBruce/CutMixImageDataGenerator_For_Keras</a></p>",
          "rawMarkdown": "Please check this repo friend \nhttps://github.com/DevBruce/CutMixImageDataGenerator_For_Keras",
          "votes": 3
        },
        {
          "id": 722275,
          "postDate": "2020-01-18T10:52:47.253Z",
          "content": "<p>Oh, Thank you 🙂</p>",
          "rawMarkdown": "Oh, Thank you 🙂"
        },
        {
          "id": 724298,
          "postDate": "2020-01-21T02:50:09.287Z",
          "content": "<p>I found that mixup is hard to train. even after 40 epochs, still low acc.</p>",
          "rawMarkdown": "I found that mixup is hard to train. even after 40 epochs, still low acc."
        },
        {
          "id": 724301,
          "postDate": "2020-01-21T03:00:27.400Z",
          "content": "<p>you have to try up to 100 epochs =) </p>",
          "rawMarkdown": "you have to try up to 100 epochs =) ",
          "votes": 2
        },
        {
          "id": 724547,
          "postDate": "2020-01-21T08:28:47.237Z",
          "content": "<p>I am only running models with mixup for up to 20 epochs, where each epoch takes 1.25 hrs to train. Within the 20 epochs, the score starts plateauing after like 14-15 epochs and changes very slowly. So, <a href=\"/drhabib\">@drhabib</a> can you please let me know if this behavior is similar to the models of yours? I would just like some confirmation that my model would eventually have a better score if I train it for 100 epochs, sacrificing 120 hours of my life :p My model architecture is based on se_resnext-50. </p>",
          "rawMarkdown": "I am only running models with mixup for up to 20 epochs, where each epoch takes 1.25 hrs to train. Within the 20 epochs, the score starts plateauing after like 14-15 epochs and changes very slowly. So, @drhabib can you please let me know if this behavior is similar to the models of yours? I would just like some confirmation that my model would eventually have a better score if I train it for 100 epochs, sacrificing 120 hours of my life :p My model architecture is based on se_resnext-50. ",
          "votes": 1
        },
        {
          "id": 724746,
          "postDate": "2020-01-21T12:56:28.400Z",
          "content": "<p>rather 120 hrs of your GPUs' life? :p</p>",
          "rawMarkdown": "rather 120 hrs of your GPUs' life? :p"
        },
        {
          "id": 724797,
          "postDate": "2020-01-21T13:46:33.580Z",
          "content": "<p><a href=\"/mahtabshaan\">@mahtabshaan</a>  you should try to train on <code>half precision</code>, you will able to fit twice as much images in the batch, (see <code>apex</code> library). Since we are faraway from deadline you can also do all your experiments on smaller model and lower resolution images. Once you are confident with the results and competition deadline is close you can scale your model with the image resolution =) </p>",
          "rawMarkdown": "@mahtabshaan  you should try to train on `half precision`, you will able to fit twice as much images in the batch, (see `apex` library). Since we are faraway from deadline you can also do all your experiments on smaller model and lower resolution images. Once you are confident with the results and competition deadline is close you can scale your model with the image resolution =) ",
          "votes": 9
        },
        {
          "id": 725904,
          "postDate": "2020-01-22T15:36:11.817Z",
          "rawMarkdown": ""
        },
        {
          "id": 728975,
          "postDate": "2020-01-25T14:24:30.357Z",
          "content": "<p>Sorry if this an the obvious question. I'm working through trying to implement this cutmix/mixup into my current pipeline and I get errors for cutmux:\n```\n      4 def rand_bbox(size, lam):\n      5     W = size[2]\n----&gt; 6     H = size[3]\n      7     cut_rat = np.sqrt(1. - lam)\n      8     cut_w = np.int(W * cut_rat)</p>\n\n<p>IndexError: tuple index out of range\n```</p>\n\n<p>and this error for mixup:\n```\n    144     def forward(self, x):\n    145         x = self.static_padding(x)\n--&gt; 146         x = F.conv2d(x, self.weight, self.bias, self.stride, self.padding, self.dilation, self.groups)\n    147         return x\n    148 </p>\n\n<p>RuntimeError: Expected 4-dimensional input for 4-dimensional weight 32 1 3 3, but got 3-dimensional input of size [256, 129, 129] instead\n```</p>\n\n<p>I at least can check that the mixup function is working because the resulting images look like a mix between two samples:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F644036%2F52c78c908dc282b169a6684612ef4b5a%2Ftest.png?generation=1579962027461251&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Sorry if this an the obvious question. I'm working through trying to implement this cutmix/mixup into my current pipeline and I get errors for cutmux:\n```\n      4 def rand_bbox(size, lam):\n      5     W = size[2]\n----&gt; 6     H = size[3]\n      7     cut_rat = np.sqrt(1. - lam)\n      8     cut_w = np.int(W * cut_rat)\n\nIndexError: tuple index out of range\n```\n\nand this error for mixup:\n```\n    144     def forward(self, x):\n    145         x = self.static_padding(x)\n--&gt; 146         x = F.conv2d(x, self.weight, self.bias, self.stride, self.padding, self.dilation, self.groups)\n    147         return x\n    148 \n\nRuntimeError: Expected 4-dimensional input for 4-dimensional weight 32 1 3 3, but got 3-dimensional input of size [256, 129, 129] instead\n```\n\nI at least can check that the mixup function is working because the resulting images look like a mix between two samples:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F644036%2F52c78c908dc282b169a6684612ef4b5a%2Ftest.png?generation=1579962027461251&amp;alt=media)\n"
        },
        {
          "id": 729047,
          "postDate": "2020-01-25T16:36:46.630Z",
          "content": "<p>It look like you want to use 1-channel input to feed CNN, in this case you also need to add a dimension indicate that it's a 1-channel input.\nTo be specific, your input size now is [256, 129, 129] but your CNN needs it to be [256, 1, 129, 129]</p>",
          "rawMarkdown": "It look like you want to use 1-channel input to feed CNN, in this case you also need to add a dimension indicate that it's a 1-channel input.\nTo be specific, your input size now is [256, 129, 129] but your CNN needs it to be [256, 1, 129, 129]",
          "votes": 1
        },
        {
          "id": 730047,
          "postDate": "2020-01-27T03:09:24.400Z",
          "content": "<p>Thanks for your advice. I had a similar problem, too. \nIn this case, it is easier to modify the mixup or cutmix function instead of the CNN. In other words, it is calculated on 3-dimentions\nDo you think so?</p>",
          "rawMarkdown": "Thanks for your advice. I had a similar problem, too. \nIn this case, it is easier to modify the mixup or cutmix function instead of the CNN. In other words, it is calculated on 3-dimentions\nDo you think so?"
        },
        {
          "id": 732688,
          "postDate": "2020-01-30T05:36:14.580Z",
          "content": "<p><a href=\"/machinelp\">@machinelp</a> , quick question - In your loop, you are always doing either cutmix or mixup to train in every batch. Does it not make sense to do these augmentation only for x% of batches and not for all of them. </p>\n\n<p>Ideally i would have imagined something like this.</p>\n\n<p>```\nif np.rand.rand()&lt;0.25:\n   do mixup</p>\n\n<p>if (np.rand.rand()&lt;0.5) and (np.rand.rand()&gt;0.25):\n   do cutmix</p>\n\n<p>else:\n    normal training with usual augmentations (rotate etc)\n```</p>\n\n<p>What are your thought on this.</p>",
          "rawMarkdown": "@machinelp , quick question - In your loop, you are always doing either cutmix or mixup to train in every batch. Does it not make sense to do these augmentation only for x% of batches and not for all of them. \n\nIdeally i would have imagined something like this.\n\n```\nif np.rand.rand()&lt;0.25:\n   do mixup\n\nif (np.rand.rand()&lt;0.5) and (np.rand.rand()&gt;0.25):\n   do cutmix\n\nelse:\n    normal training with usual augmentations (rotate etc)\n```\n\nWhat are your thought on this.",
          "votes": 3
        },
        {
          "id": 732706,
          "postDate": "2020-01-30T06:09:12.573Z",
          "content": "<p>You can have a try</p>",
          "rawMarkdown": "You can have a try",
          "votes": 2
        },
        {
          "id": 747372,
          "postDate": "2020-02-16T10:54:47.180Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 750478,
          "postDate": "2020-02-19T12:27:09.770Z",
          "content": "<p>Hey <a href=\"/robikscube\">@robikscube</a> Facing with the same issue. Have you solved it?Can you tell me how to solve this? Thanks in advance 😄 </p>",
          "rawMarkdown": "Hey @robikscube Facing with the same issue. Have you solved it?Can you tell me how to solve this? Thanks in advance 😄 "
        },
        {
          "id": 752696,
          "postDate": "2020-02-21T10:07:00.643Z",
          "content": "<p><a href=\"/machinelp\">@machinelp</a> \nafter doing : loss = mixup_criterion(outputs1,outputs2,outputs3, targets) </p>\n\n<p>when i try loss.backward() i get this error : \nAttributeError: 'list' object has no attribute 'backward'</p>\n\n<p>we don't need pytorch autograd Backward() function here?\nor how can i solve this problem?</p>",
          "rawMarkdown": "@machinelp \nafter doing : loss = mixup_criterion(outputs1,outputs2,outputs3, targets) \n\nwhen i try loss.backward() i get this error : \nAttributeError: 'list' object has no attribute 'backward'\n\nwe don't need pytorch autograd Backward() function here?\nor how can i solve this problem?"
        }
      ]
    },
    {
      "id": 724078,
      "postDate": "2020-01-20T19:17:29.987Z",
      "content": "<p>Great! If people are curios what <code>alpha</code> should be used in cutmix its (<code>1.0</code>). Below is the image from original paper. (Lower top 1% error is better) \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F9d76a95113937205c68ec10681e86b53%2FScreen%20Shot%202020-01-20%20at%202.16.37%20PM.png?generation=1579547824680882&amp;alt=media\" alt=\"\"></p>\n\n<p>paper: \n<a href=\"https://arxiv.org/pdf/1905.04899.pdf\">https://arxiv.org/pdf/1905.04899.pdf</a></p>",
      "rawMarkdown": "Great! If people are curios what `alpha` should be used in cutmix its (`1.0`). Below is the image from original paper. (Lower top 1% error is better) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F9d76a95113937205c68ec10681e86b53%2FScreen%20Shot%202020-01-20%20at%202.16.37%20PM.png?generation=1579547824680882&amp;alt=media)\n\npaper: \nhttps://arxiv.org/pdf/1905.04899.pdf\n ",
      "votes": 13,
      "replies": [
        {
          "id": 724368,
          "postDate": "2020-01-21T05:13:54.307Z",
          "content": "<p>Setting <code>alpha</code> to 1.0 ( Beta(1.0, 1.0) ) is just using Uniform Distribution, as we can see from the definition of <a href=\"https://en.wikipedia.org/wiki/Beta_distribution\"><strong>Beta Distribution</strong></a>. <br>\nBelow shows how distribution changes with parameter <code>alpha</code>. Small alpha( ~ 0) attributes to original images, while enough large alpha mixes images half and half. So the optimal alpha will depend on the image property.  </p>\n\n<p><br>\n<br></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F61543332570dd26485e1a8e9f76eeced%2Ffig.png?generation=1579583175899649&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Setting `alpha` to 1.0 ( Beta(1.0, 1.0) ) is just using Uniform Distribution, as we can see from the definition of [**Beta Distribution**](https://en.wikipedia.org/wiki/Beta_distribution).  \nBelow shows how distribution changes with parameter `alpha`. Small alpha( ~ 0) attributes to original images, while enough large alpha mixes images half and half. So the optimal alpha will depend on the image property.  \n  \n<br>\n<br>\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F61543332570dd26485e1a8e9f76eeced%2Ffig.png?generation=1579583175899649&amp;alt=media)\n",
          "votes": 10
        },
        {
          "id": 769741,
          "postDate": "2020-03-12T08:00:32.580Z",
          "content": "<p>Did you use an alpha equals to 1?</p>",
          "rawMarkdown": "Did you use an alpha equals to 1?"
        }
      ]
    },
    {
      "id": 722255,
      "postDate": "2020-01-18T10:19:23.893Z",
      "content": "<p>I like your implementation!\nEDIT: I've tried your implementation, it helped a lot!</p>",
      "rawMarkdown": "I like your implementation!\nEDIT: I've tried your implementation, it helped a lot!",
      "votes": 11,
      "replies": [
        {
          "id": 722262,
          "postDate": "2020-01-18T10:26:33.437Z",
          "content": "<p>thanks😄</p>",
          "rawMarkdown": "thanks😄",
          "votes": 2
        }
      ]
    },
    {
      "id": 724956,
      "postDate": "2020-01-21T16:36:11.930Z",
      "content": "<p>the full power of mixup/cutout will not be realized unless your network has the \"right capacity and structure\" to learn the augmentation. </p>\n\n<p>i.e. you need to improve your network design and augmentation at the same time. just doing any one alone may only give limited improvement</p>",
      "rawMarkdown": "the full power of mixup/cutout will not be realized unless your network has the \"right capacity and structure\" to learn the augmentation. \n\ni.e. you need to improve your network design and augmentation at the same time. just doing any one alone may only give limited improvement",
      "votes": 9,
      "replies": [
        {
          "id": 725247,
          "postDate": "2020-01-21T23:58:13.403Z",
          "content": "<p>What kind of design should I do specifically?</p>",
          "rawMarkdown": "What kind of design should I do specifically?"
        },
        {
          "id": 725455,
          "postDate": "2020-01-22T05:54:09.613Z",
          "content": "<p>for handwriting recognition, there target object is made up of strokes. they are quite \"thin\" (we need to detect high frequency)</p>\n\n<p>in network tweaking, the usual modifications should work, e.g\n- number of filters, wide? pyramid?\n- size of filters (3x3,5x5)\n- scaling (downsampling/upsampling)\n- pooling (max vs avg vs generalised, etc)\n- activation\n- convolution: grouped? dilated?\n- connection: residual, densely connect, multi scale (inception style)\n- deeply supervised, etc\n- attention: channel vs spatial</p>",
          "rawMarkdown": "for handwriting recognition, there target object is made up of strokes. they are quite \"thin\" (we need to detect high frequency)\n\nin network tweaking, the usual modifications should work, e.g\n- number of filters, wide? pyramid?\n- size of filters (3x3,5x5)\n- scaling (downsampling/upsampling)\n- pooling (max vs avg vs generalised, etc)\n- activation\n- convolution: grouped? dilated?\n- connection: residual, densely connect, multi scale (inception style)\n- deeply supervised, etc\n- attention: channel vs spatial\n",
          "votes": 20
        },
        {
          "id": 725712,
          "postDate": "2020-01-22T12:25:47.477Z",
          "content": "<p>I see. Thank you for reply.</p>",
          "rawMarkdown": "I see. Thank you for reply."
        },
        {
          "id": 747365,
          "postDate": "2020-02-16T10:43:22.660Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 763161,
      "postDate": "2020-03-04T07:44:53.920Z",
      "content": "<p>FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">https://www.kaggle.com/c/bengaliai-cv19/discussion/133322</a></p>",
      "rawMarkdown": "FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：https://www.kaggle.com/c/bengaliai-cv19/discussion/133322",
      "votes": 3
    },
    {
      "id": 734767,
      "postDate": "2020-02-02T01:32:51.033Z",
      "content": "<p>mixup/cutmix with ohem loss : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
      "rawMarkdown": "mixup/cutmix with ohem loss : https://www.kaggle.com/c/bengaliai-cv19/discussion/128637",
      "votes": 3
    },
    {
      "id": 722775,
      "postDate": "2020-01-19T04:57:15.690Z",
      "content": "<p>If you use/refer opensource, please let others know the original implementation. It seems to be similar to source codes that I'm familiar with. BTW, thanks for sharing and it helps a lot for many people!</p>",
      "rawMarkdown": "If you use/refer opensource, please let others know the original implementation. It seems to be similar to source codes that I'm familiar with. BTW, thanks for sharing and it helps a lot for many people!",
      "votes": 2
    },
    {
      "id": 722731,
      "postDate": "2020-01-19T02:24:58.913Z",
      "content": "<p>cutmix_criterion == mixup_criterion.</p>",
      "rawMarkdown": "cutmix_criterion == mixup_criterion.",
      "votes": 2
    },
    {
      "id": 723773,
      "postDate": "2020-01-20T12:47:35.930Z",
      "content": "<p>Thanks for the share. Very informative. \nHave a deserving upvote.\n:) </p>",
      "rawMarkdown": "Thanks for the share. Very informative. \nHave a deserving upvote.\n:) ",
      "votes": 1,
      "replies": [
        {
          "id": 723941,
          "postDate": "2020-01-20T16:07:02.337Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!",
          "votes": 2
        }
      ]
    },
    {
      "id": 723424,
      "postDate": "2020-01-20T02:39:51.483Z",
      "content": "<p>Do you mind sharing you train loss and recall with cutmix and mixup <a href=\"/machinelp\">@machinelp</a> </p>",
      "rawMarkdown": "Do you mind sharing you train loss and recall with cutmix and mixup @machinelp ",
      "votes": 1
    },
    {
      "id": 743517,
      "postDate": "2020-02-12T05:53:41.733Z",
      "content": "<p>Hey <a href=\"/machinelp\">@machinelp</a>, Thanks for posting this! I'm trying to implement CutMix on CIFAR-10 dataset instead of directly applying on this competition's data as I'm new to this type of implementations. Now when I train the model on CIFAR-10 , the training loss/accuracy fluctuates way too much but the validation accuracy is stable. Here is the accuracy curve:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1905996%2Fabf72a67151d5abf78daa8ee069d8cfc%2FScreenshot%20(126\" alt=\"\">.png?generation=1581487322942081&amp;alt=media)</p>\n\n<p>Is this normal behavior or I messed up the implementation? I also feel that this might be because of CIFAR-10 image resolution (32 x 32). Here is my complete implementation: <a href=\"https://www.kaggle.com/kaushal2896/cifar-10-simple-cnn-with-cutmix-using-pytorch\">https://www.kaggle.com/kaushal2896/cifar-10-simple-cnn-with-cutmix-using-pytorch</a>.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hey @machinelp, Thanks for posting this! I'm trying to implement CutMix on CIFAR-10 dataset instead of directly applying on this competition's data as I'm new to this type of implementations. Now when I train the model on CIFAR-10 , the training loss/accuracy fluctuates way too much but the validation accuracy is stable. Here is the accuracy curve:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1905996%2Fabf72a67151d5abf78daa8ee069d8cfc%2FScreenshot%20(126).png?generation=1581487322942081&amp;alt=media)\n\nIs this normal behavior or I messed up the implementation? I also feel that this might be because of CIFAR-10 image resolution (32 x 32). Here is my complete implementation: https://www.kaggle.com/kaushal2896/cifar-10-simple-cnn-with-cutmix-using-pytorch.\n\nThanks!",
      "votes": 2,
      "replies": [
        {
          "id": 755080,
          "postDate": "2020-02-24T12:07:37.990Z",
          "content": "<p><a href=\"/kaushal2896\">@kaushal2896</a> This behavior is explainable if you are using single random lambda for the whole batch cutmix/mixup. When lambda &gt;.5 (original sample given weights), your model classifies correctly with an accuracy reasonably close to the validation accuracy. When lambda &lt;.5, the model performs very terrible on the original \"labels\" (they should really be replaced by the shuffled labels in this case). Because lambda &gt;.5 with a chance of 1/2, this is reflected in your logs by the model performing miserably on training metrics half the time. Try changing the labels for evaluation to shuffled labels when batch lambda &lt;.5 and observe the training metric. I would suppose that there is a bug if the logs still fluctuate this much; but there should be no problem :) Hope this helps</p>",
          "rawMarkdown": "@kaushal2896 This behavior is explainable if you are using single random lambda for the whole batch cutmix/mixup. When lambda &gt;.5 (original sample given weights), your model classifies correctly with an accuracy reasonably close to the validation accuracy. When lambda &lt;.5, the model performs very terrible on the original \"labels\" (they should really be replaced by the shuffled labels in this case). Because lambda &gt;.5 with a chance of 1/2, this is reflected in your logs by the model performing miserably on training metrics half the time. Try changing the labels for evaluation to shuffled labels when batch lambda &lt;.5 and observe the training metric. I would suppose that there is a bug if the logs still fluctuate this much; but there should be no problem :) Hope this helps",
          "votes": 1
        }
      ]
    },
    {
      "id": 740610,
      "postDate": "2020-02-09T15:30:10.053Z",
      "content": "<p>Hi MachineLP thanks for the code. It's really helpful.</p>\n\n<p>However may I ask for some tip.\nI have used the above code in my own training implementation, however I am getting a lot of fluctuation and slower convergence, compared with this fastdai implementation: <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964</a></p>\n\n<p>I have inspected the mixup and loss in this kernel and it does not seem to differ much, but it has much stable loss and converge much quicker. </p>\n\n<p>I am just wondering if you could shed some light into this.</p>\n\n<p>Many thanks!</p>",
      "rawMarkdown": "Hi MachineLP thanks for the code. It's really helpful.\n\nHowever may I ask for some tip.\nI have used the above code in my own training implementation, however I am getting a lot of fluctuation and slower convergence, compared with this fastdai implementation: https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\n\nI have inspected the mixup and loss in this kernel and it does not seem to differ much, but it has much stable loss and converge much quicker. \n\nI am just wondering if you could shed some light into this.\n\nMany thanks!",
      "votes": 2
    },
    {
      "id": 731419,
      "postDate": "2020-01-28T15:59:41.613Z",
      "content": "<p>Does anyone try above CutMix/MixUp with Keras? Unfortunately, I didn't find any robust keras implementation that could fit such a problem. 🙁 </p>",
      "rawMarkdown": "Does anyone try above CutMix/MixUp with Keras? Unfortunately, I didn't find any robust keras implementation that could fit such a problem. 🙁 ",
      "votes": 1,
      "replies": [
        {
          "id": 731457,
          "postDate": "2020-01-28T17:01:33.037Z",
          "content": "<p>There is a kernel .....</p>\n\n<p><a href=\"https://www.kaggle.com/code1110/mixup-cutmix-in-keras\">https://www.kaggle.com/code1110/mixup-cutmix-in-keras</a></p>",
          "rawMarkdown": "There is a kernel .....\n\nhttps://www.kaggle.com/code1110/mixup-cutmix-in-keras",
          "votes": 2
        },
        {
          "id": 731464,
          "postDate": "2020-01-28T17:09:05.067Z",
          "content": "<p>Ops didn't notice. Actually, it published two hours ago before asking! Thanks <a href=\"/drhabib\">@drhabib</a> 😛 😅 </p>",
          "rawMarkdown": "Ops didn't notice. Actually, it published two hours ago before asking! Thanks @drhabib 😛 😅 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 767202,
      "postDate": "2020-03-09T09:49:52.407Z",
      "content": "<p>Just interesting, can we use mixup/cutmix simultaneously or we need to pick one</p>",
      "rawMarkdown": "Just interesting, can we use mixup/cutmix simultaneously or we need to pick one"
    },
    {
      "id": 754416,
      "postDate": "2020-02-23T14:22:06.180Z",
      "content": "<p>I have tried the cutmix implementation for FastAI. But it throws a 'tensors of different dimensions while concatenating error'. When I tried to debug, I was able to see that 2 [128,2,3] -sized tensors were concatenated with a [128,1] tensor. How can I rectify the error? The link to the FastAI callback: \n<a href=\"https://github.com/oguiza/fastai_extensions/blob/master/fastai_extensions/exp/nb_NewDataAugmentation.py\">https://github.com/oguiza/fastai_extensions/blob/master/fastai_extensions/exp/nb_NewDataAugmentation.py</a>\nPS: I don't know his Kaggle handle.</p>",
      "rawMarkdown": "I have tried the cutmix implementation for FastAI. But it throws a 'tensors of different dimensions while concatenating error'. When I tried to debug, I was able to see that 2 [128,2,3] -sized tensors were concatenated with a [128,1] tensor. How can I rectify the error? The link to the FastAI callback: \nhttps://github.com/oguiza/fastai_extensions/blob/master/fastai_extensions/exp/nb_NewDataAugmentation.py\nPS: I don't know his Kaggle handle."
    },
    {
      "id": 747426,
      "postDate": "2020-02-16T12:23:30.507Z",
      "content": "<p><a href=\"/machinelp\">@machinelp</a> <a href=\"/haqishen\">@haqishen</a> Curious that this implementation of mixup does not weigh grapheme/vowel/conso by 2:1:1? Wondering if any light can be shed onto this</p>",
      "rawMarkdown": "@machinelp @haqishen Curious that this implementation of mixup does not weigh grapheme/vowel/conso by 2:1:1? Wondering if any light can be shed onto this"
    },
    {
      "id": 733043,
      "postDate": "2020-01-30T15:21:39.610Z",
      "content": "<p>When you guys are using mixup/cutmix, are you using them separately, i.e. randomly select whether a batch is mixup or cutmix? Or are you combining mixup with cutmix, i.e. first cutmix inside the batch, and then mix up the cutmixed samples?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "When you guys are using mixup/cutmix, are you using them separately, i.e. randomly select whether a batch is mixup or cutmix? Or are you combining mixup with cutmix, i.e. first cutmix inside the batch, and then mix up the cutmixed samples?\n\nThanks!"
    },
    {
      "id": 728974,
      "postDate": "2020-01-25T14:24:17.787Z",
      "content": "<p>I noticed this specific implementation uses the same lambda for the whole batch compared to image specific lambda's. Even though it might be only a minor detail, wouldn't it be beneficial to use different cut-ratios, which are determined by the lambda, for each image in the batch?</p>",
      "rawMarkdown": "I noticed this specific implementation uses the same lambda for the whole batch compared to image specific lambda's. Even though it might be only a minor detail, wouldn't it be beneficial to use different cut-ratios, which are determined by the lambda, for each image in the batch?"
    },
    {
      "id": 739948,
      "postDate": "2020-02-08T15:45:29.210Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 724679,
      "postDate": "2020-01-21T11:31:38.080Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 724866,
      "postDate": "2020-01-21T15:23:10.623Z",
      "content": "<p><a href=\"/machinelp\">@machinelp</a>   nice work thanks!!</p>",
      "rawMarkdown": "@machinelp   nice work thanks!!",
      "votes": 1
    },
    {
      "id": 723696,
      "postDate": "2020-01-20T10:44:37.103Z",
      "content": "<p>Thanks!😃  </p>",
      "rawMarkdown": "Thanks!😃  ",
      "votes": 1
    },
    {
      "id": 723430,
      "postDate": "2020-01-20T02:59:59.977Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1
    },
    {
      "id": 722837,
      "postDate": "2020-01-19T07:03:45.120Z",
      "content": "<p>Nice Thanks</p>",
      "rawMarkdown": "Nice Thanks",
      "votes": 1
    },
    {
      "id": 722272,
      "postDate": "2020-01-18T10:49:26.363Z",
      "content": "<p>Nice . Thanks !</p>",
      "rawMarkdown": "Nice . Thanks !",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 730097,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-27T05:23:21.007000",
      "content": "<p>@DrHB</p>\n\n<p>\" ... you should try to train on half precision \"</p>\n\n<p>if you are using half-precision for training, remember to switch back to full precision for the last finetunning. it may improve results</p>",
      "votes": 20,
      "replies": [
        {
          "id": 730389,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-27T12:24:17.623000",
          "content": "<p>Ahhh!!! That’s smart! Thanks for the advice!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754521,
          "author_name": "cswwp",
          "author_url": "",
          "post_date": "2020-02-23T17:23:11.187000",
          "content": "<p>Hi <a href=\"/hengck23\">@hengck23</a> <a href=\"/drhabib\">@drhabib</a>   what's the meaning of training on half precision?👀 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754553,
          "author_name": "Son of Anton v3.0",
          "author_url": "",
          "post_date": "2020-02-23T18:13:11.890000",
          "content": "<p>fp16 I guess. int 16??</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 722191,
      "author_name": "MachineLP",
      "author_url": "",
      "post_date": "2020-01-18T08:48:17.823000",
      "content": "<p>Update:\n```\n......\nfor i, (image_id, images, label1, label2, label3) in enumerate(data_loader_train):\n            images = images.to(device)\n            label1 = label1.to(device)\n            label2 = label2.to(device)\n            label3 = label3.to(device)\n            # print (image_id, label1, label2, label3)</p>\n\n<pre><code>        if np.random.rand()&lt;0.5:\n            images, targets = mixup(images, label1, label2, label3, 0.4)\n            output1, output2, output3 = model(images)\n            loss = mixup_criterion(output1,output2,output3, targets) \n        else:\n            images, targets = cutmix(images, label1, label2, label3, 0.4)\n            output1, output2, output3 = model(images)\n            loss = cutmix_criterion(output1,output2,output3, targets) \n</code></pre>\n\n<p>......\n```</p>",
      "votes": 19,
      "replies": [
        {
          "id": 722242,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-18T09:57:14.630000",
          "content": "<p>Thanks for the implementation. I'm not too familiar with torch but keras user. I didn’t try out your code yet but would you please point out to me is there any changes I need to make of this code while using it keras. A pseudo code would be helpful. Thanks.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 722268,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-01-18T10:47:46.150000",
          "content": "<p>Please check this repo friend \n<a href=\"https://github.com/DevBruce/CutMixImageDataGenerator_For_Keras\">https://github.com/DevBruce/CutMixImageDataGenerator_For_Keras</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 722275,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-18T10:52:47.253000",
          "content": "<p>Oh, Thank you 🙂</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 724298,
          "author_name": "Tian",
          "author_url": "",
          "post_date": "2020-01-21T02:50:09.287000",
          "content": "<p>I found that mixup is hard to train. even after 40 epochs, still low acc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 724301,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-21T03:00:27.400000",
          "content": "<p>you have to try up to 100 epochs =) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 724547,
          "author_name": "Mahtab Noor Shaan",
          "author_url": "",
          "post_date": "2020-01-21T08:28:47.237000",
          "content": "<p>I am only running models with mixup for up to 20 epochs, where each epoch takes 1.25 hrs to train. Within the 20 epochs, the score starts plateauing after like 14-15 epochs and changes very slowly. So, <a href=\"/drhabib\">@drhabib</a> can you please let me know if this behavior is similar to the models of yours? I would just like some confirmation that my model would eventually have a better score if I train it for 100 epochs, sacrificing 120 hours of my life :p My model architecture is based on se_resnext-50. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 724746,
          "author_name": "kobi2000",
          "author_url": "",
          "post_date": "2020-01-21T12:56:28.400000",
          "content": "<p>rather 120 hrs of your GPUs' life? :p</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 724797,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-21T13:46:33.580000",
          "content": "<p><a href=\"/mahtabshaan\">@mahtabshaan</a>  you should try to train on <code>half precision</code>, you will able to fit twice as much images in the batch, (see <code>apex</code> library). Since we are faraway from deadline you can also do all your experiments on smaller model and lower resolution images. Once you are confident with the results and competition deadline is close you can scale your model with the image resolution =) </p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 725904,
          "author_name": "Opfer-Ab0",
          "author_url": "",
          "post_date": "2020-01-22T15:36:11.817000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728975,
          "author_name": "Rob Mulla",
          "author_url": "",
          "post_date": "2020-01-25T14:24:30.357000",
          "content": "<p>Sorry if this an the obvious question. I'm working through trying to implement this cutmix/mixup into my current pipeline and I get errors for cutmux:\n```\n      4 def rand_bbox(size, lam):\n      5     W = size[2]\n----&gt; 6     H = size[3]\n      7     cut_rat = np.sqrt(1. - lam)\n      8     cut_w = np.int(W * cut_rat)</p>\n\n<p>IndexError: tuple index out of range\n```</p>\n\n<p>and this error for mixup:\n```\n    144     def forward(self, x):\n    145         x = self.static_padding(x)\n--&gt; 146         x = F.conv2d(x, self.weight, self.bias, self.stride, self.padding, self.dilation, self.groups)\n    147         return x\n    148 </p>\n\n<p>RuntimeError: Expected 4-dimensional input for 4-dimensional weight 32 1 3 3, but got 3-dimensional input of size [256, 129, 129] instead\n```</p>\n\n<p>I at least can check that the mixup function is working because the resulting images look like a mix between two samples:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F644036%2F52c78c908dc282b169a6684612ef4b5a%2Ftest.png?generation=1579962027461251&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 729047,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-01-25T16:36:46.630000",
          "content": "<p>It look like you want to use 1-channel input to feed CNN, in this case you also need to add a dimension indicate that it's a 1-channel input.\nTo be specific, your input size now is [256, 129, 129] but your CNN needs it to be [256, 1, 129, 129]</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 730047,
          "author_name": "Syn",
          "author_url": "",
          "post_date": "2020-01-27T03:09:24.400000",
          "content": "<p>Thanks for your advice. I had a similar problem, too. \nIn this case, it is easier to modify the mixup or cutmix function instead of the CNN. In other words, it is calculated on 3-dimentions\nDo you think so?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 732688,
          "author_name": "Phaedrus",
          "author_url": "",
          "post_date": "2020-01-30T05:36:14.580000",
          "content": "<p><a href=\"/machinelp\">@machinelp</a> , quick question - In your loop, you are always doing either cutmix or mixup to train in every batch. Does it not make sense to do these augmentation only for x% of batches and not for all of them. </p>\n\n<p>Ideally i would have imagined something like this.</p>\n\n<p>```\nif np.rand.rand()&lt;0.25:\n   do mixup</p>\n\n<p>if (np.rand.rand()&lt;0.5) and (np.rand.rand()&gt;0.25):\n   do cutmix</p>\n\n<p>else:\n    normal training with usual augmentations (rotate etc)\n```</p>\n\n<p>What are your thought on this.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 732706,
          "author_name": "MachineLP",
          "author_url": "",
          "post_date": "2020-01-30T06:09:12.573000",
          "content": "<p>You can have a try</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 747372,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-16T10:54:47.180000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750478,
          "author_name": "ratan rohith",
          "author_url": "",
          "post_date": "2020-02-19T12:27:09.770000",
          "content": "<p>Hey <a href=\"/robikscube\">@robikscube</a> Facing with the same issue. Have you solved it?Can you tell me how to solve this? Thanks in advance 😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752696,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2020-02-21T10:07:00.643000",
          "content": "<p><a href=\"/machinelp\">@machinelp</a> \nafter doing : loss = mixup_criterion(outputs1,outputs2,outputs3, targets) </p>\n\n<p>when i try loss.backward() i get this error : \nAttributeError: 'list' object has no attribute 'backward'</p>\n\n<p>we don't need pytorch autograd Backward() function here?\nor how can i solve this problem?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 724078,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2020-01-20T19:17:29.987000",
      "content": "<p>Great! If people are curios what <code>alpha</code> should be used in cutmix its (<code>1.0</code>). Below is the image from original paper. (Lower top 1% error is better) \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F9d76a95113937205c68ec10681e86b53%2FScreen%20Shot%202020-01-20%20at%202.16.37%20PM.png?generation=1579547824680882&amp;alt=media\" alt=\"\"></p>\n\n<p>paper: \n<a href=\"https://arxiv.org/pdf/1905.04899.pdf\">https://arxiv.org/pdf/1905.04899.pdf</a></p>",
      "votes": 13,
      "replies": [
        {
          "id": 724368,
          "author_name": "Maxwell",
          "author_url": "",
          "post_date": "2020-01-21T05:13:54.307000",
          "content": "<p>Setting <code>alpha</code> to 1.0 ( Beta(1.0, 1.0) ) is just using Uniform Distribution, as we can see from the definition of <a href=\"https://en.wikipedia.org/wiki/Beta_distribution\"><strong>Beta Distribution</strong></a>. <br>\nBelow shows how distribution changes with parameter <code>alpha</code>. Small alpha( ~ 0) attributes to original images, while enough large alpha mixes images half and half. So the optimal alpha will depend on the image property.  </p>\n\n<p><br>\n<br></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F479538%2F61543332570dd26485e1a8e9f76eeced%2Ffig.png?generation=1579583175899649&amp;alt=media\" alt=\"\"></p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 769741,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-12T08:00:32.580000",
          "content": "<p>Did you use an alpha equals to 1?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 722255,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-01-18T10:19:23.893000",
      "content": "<p>I like your implementation!\nEDIT: I've tried your implementation, it helped a lot!</p>",
      "votes": 11,
      "replies": [
        {
          "id": 722262,
          "author_name": "MachineLP",
          "author_url": "",
          "post_date": "2020-01-18T10:26:33.437000",
          "content": "<p>thanks😄</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 724956,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-21T16:36:11.930000",
      "content": "<p>the full power of mixup/cutout will not be realized unless your network has the \"right capacity and structure\" to learn the augmentation. </p>\n\n<p>i.e. you need to improve your network design and augmentation at the same time. just doing any one alone may only give limited improvement</p>",
      "votes": 9,
      "replies": [
        {
          "id": 725247,
          "author_name": "Syn",
          "author_url": "",
          "post_date": "2020-01-21T23:58:13.403000",
          "content": "<p>What kind of design should I do specifically?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 725455,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-01-22T05:54:09.613000",
          "content": "<p>for handwriting recognition, there target object is made up of strokes. they are quite \"thin\" (we need to detect high frequency)</p>\n\n<p>in network tweaking, the usual modifications should work, e.g\n- number of filters, wide? pyramid?\n- size of filters (3x3,5x5)\n- scaling (downsampling/upsampling)\n- pooling (max vs avg vs generalised, etc)\n- activation\n- convolution: grouped? dilated?\n- connection: residual, densely connect, multi scale (inception style)\n- deeply supervised, etc\n- attention: channel vs spatial</p>",
          "votes": 20,
          "replies": []
        },
        {
          "id": 725712,
          "author_name": "Syn",
          "author_url": "",
          "post_date": "2020-01-22T12:25:47.477000",
          "content": "<p>I see. Thank you for reply.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747365,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-16T10:43:22.660000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 763161,
      "author_name": "MachineLP",
      "author_url": "",
      "post_date": "2020-03-04T07:44:53.920000",
      "content": "<p>FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">https://www.kaggle.com/c/bengaliai-cv19/discussion/133322</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 734767,
      "author_name": "MachineLP",
      "author_url": "",
      "post_date": "2020-02-02T01:32:51.033000",
      "content": "<p>mixup/cutmix with ohem loss : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 722775,
      "author_name": "Ildoo Kim",
      "author_url": "",
      "post_date": "2020-01-19T04:57:15.690000",
      "content": "<p>If you use/refer opensource, please let others know the original implementation. It seems to be similar to source codes that I'm familiar with. BTW, thanks for sharing and it helps a lot for many people!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 722731,
      "author_name": "Tian",
      "author_url": "",
      "post_date": "2020-01-19T02:24:58.913000",
      "content": "<p>cutmix_criterion == mixup_criterion.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 723773,
      "author_name": "Shahbaz Khan",
      "author_url": "",
      "post_date": "2020-01-20T12:47:35.930000",
      "content": "<p>Thanks for the share. Very informative. \nHave a deserving upvote.\n:) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 723941,
          "author_name": "MachineLP",
          "author_url": "",
          "post_date": "2020-01-20T16:07:02.337000",
          "content": "<p>Thanks!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 723424,
      "author_name": "cswwp",
      "author_url": "",
      "post_date": "2020-01-20T02:39:51.483000",
      "content": "<p>Do you mind sharing you train loss and recall with cutmix and mixup <a href=\"/machinelp\">@machinelp</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 743517,
      "author_name": "Kaushal Shah",
      "author_url": "",
      "post_date": "2020-02-12T05:53:41.733000",
      "content": "<p>Hey <a href=\"/machinelp\">@machinelp</a>, Thanks for posting this! I'm trying to implement CutMix on CIFAR-10 dataset instead of directly applying on this competition's data as I'm new to this type of implementations. Now when I train the model on CIFAR-10 , the training loss/accuracy fluctuates way too much but the validation accuracy is stable. Here is the accuracy curve:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1905996%2Fabf72a67151d5abf78daa8ee069d8cfc%2FScreenshot%20(126\" alt=\"\">.png?generation=1581487322942081&amp;alt=media)</p>\n\n<p>Is this normal behavior or I messed up the implementation? I also feel that this might be because of CIFAR-10 image resolution (32 x 32). Here is my complete implementation: <a href=\"https://www.kaggle.com/kaushal2896/cifar-10-simple-cnn-with-cutmix-using-pytorch\">https://www.kaggle.com/kaushal2896/cifar-10-simple-cnn-with-cutmix-using-pytorch</a>.</p>\n\n<p>Thanks!</p>",
      "votes": 2,
      "replies": [
        {
          "id": 755080,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-02-24T12:07:37.990000",
          "content": "<p><a href=\"/kaushal2896\">@kaushal2896</a> This behavior is explainable if you are using single random lambda for the whole batch cutmix/mixup. When lambda &gt;.5 (original sample given weights), your model classifies correctly with an accuracy reasonably close to the validation accuracy. When lambda &lt;.5, the model performs very terrible on the original \"labels\" (they should really be replaced by the shuffled labels in this case). Because lambda &gt;.5 with a chance of 1/2, this is reflected in your logs by the model performing miserably on training metrics half the time. Try changing the labels for evaluation to shuffled labels when batch lambda &lt;.5 and observe the training metric. I would suppose that there is a bug if the logs still fluctuate this much; but there should be no problem :) Hope this helps</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 740610,
      "author_name": "YL",
      "author_url": "",
      "post_date": "2020-02-09T15:30:10.053000",
      "content": "<p>Hi MachineLP thanks for the code. It's really helpful.</p>\n\n<p>However may I ask for some tip.\nI have used the above code in my own training implementation, however I am getting a lot of fluctuation and slower convergence, compared with this fastdai implementation: <a href=\"https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\">https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964</a></p>\n\n<p>I have inspected the mixup and loss in this kernel and it does not seem to differ much, but it has much stable loss and converge much quicker. </p>\n\n<p>I am just wondering if you could shed some light into this.</p>\n\n<p>Many thanks!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 731419,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-01-28T15:59:41.613000",
      "content": "<p>Does anyone try above CutMix/MixUp with Keras? Unfortunately, I didn't find any robust keras implementation that could fit such a problem. 🙁 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 731457,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2020-01-28T17:01:33.037000",
          "content": "<p>There is a kernel .....</p>\n\n<p><a href=\"https://www.kaggle.com/code1110/mixup-cutmix-in-keras\">https://www.kaggle.com/code1110/mixup-cutmix-in-keras</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 731464,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-01-28T17:09:05.067000",
          "content": "<p>Ops didn't notice. Actually, it published two hours ago before asking! Thanks <a href=\"/drhabib\">@drhabib</a> 😛 😅 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 767202,
      "author_name": "Michael Kalinin",
      "author_url": "",
      "post_date": "2020-03-09T09:49:52.407000",
      "content": "<p>Just interesting, can we use mixup/cutmix simultaneously or we need to pick one</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 754416,
      "author_name": "Son of Anton v3.0",
      "author_url": "",
      "post_date": "2020-02-23T14:22:06.180000",
      "content": "<p>I have tried the cutmix implementation for FastAI. But it throws a 'tensors of different dimensions while concatenating error'. When I tried to debug, I was able to see that 2 [128,2,3] -sized tensors were concatenated with a [128,1] tensor. How can I rectify the error? The link to the FastAI callback: \n<a href=\"https://github.com/oguiza/fastai_extensions/blob/master/fastai_extensions/exp/nb_NewDataAugmentation.py\">https://github.com/oguiza/fastai_extensions/blob/master/fastai_extensions/exp/nb_NewDataAugmentation.py</a>\nPS: I don't know his Kaggle handle.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 747426,
      "author_name": "Nicholas Lyu",
      "author_url": "",
      "post_date": "2020-02-16T12:23:30.507000",
      "content": "<p><a href=\"/machinelp\">@machinelp</a> <a href=\"/haqishen\">@haqishen</a> Curious that this implementation of mixup does not weigh grapheme/vowel/conso by 2:1:1? Wondering if any light can be shed onto this</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 733043,
      "author_name": "PY",
      "author_url": "",
      "post_date": "2020-01-30T15:21:39.610000",
      "content": "<p>When you guys are using mixup/cutmix, are you using them separately, i.e. randomly select whether a batch is mixup or cutmix? Or are you combining mixup with cutmix, i.e. first cutmix inside the batch, and then mix up the cutmixed samples?</p>\n\n<p>Thanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728974,
      "author_name": "Morris",
      "author_url": "",
      "post_date": "2020-01-25T14:24:17.787000",
      "content": "<p>I noticed this specific implementation uses the same lambda for the whole batch compared to image specific lambda's. Even though it might be only a minor detail, wouldn't it be beneficial to use different cut-ratios, which are determined by the lambda, for each image in the batch?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 739948,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-08T15:45:29.210000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 724679,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-21T11:31:38.080000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 724866,
      "author_name": "A/C",
      "author_url": "",
      "post_date": "2020-01-21T15:23:10.623000",
      "content": "<p><a href=\"/machinelp\">@machinelp</a>   nice work thanks!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 723696,
      "author_name": "Vivek Mittal",
      "author_url": "",
      "post_date": "2020-01-20T10:44:37.103000",
      "content": "<p>Thanks!😃  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 723430,
      "author_name": "kaihenguestc",
      "author_url": "",
      "post_date": "2020-01-20T02:59:59.977000",
      "content": "<p>Thanks!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 722837,
      "author_name": "A/C",
      "author_url": "",
      "post_date": "2020-01-19T07:03:45.120000",
      "content": "<p>Nice Thanks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 722272,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2020-01-18T10:49:26.363000",
      "content": "<p>Nice . Thanks !</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "722013": "```python\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n\n    return bbx1, bby1, bbx2, bby2\ndef cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    data[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n    return data, targets\n\ndef mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    data = data * lam + shuffled_data * (1 - lam)\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\n    return data, targets\n\n\ndef cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)\n\ndef mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = nn.CrossEntropyLoss(reduction='mean')\n    return lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2) + lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4) + lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6)\n```",
    "730097": "@DrHB\n \n\" ... you should try to train on half precision \"\n\nif you are using half-precision for training, remember to switch back to full precision for the last finetunning. it may improve results",
    "722191": "Update:\n```\n......\nfor i, (image_id, images, label1, label2, label3) in enumerate(data_loader_train):\n            images = images.to(device)\n            label1 = label1.to(device)\n            label2 = label2.to(device)\n            label3 = label3.to(device)\n            # print (image_id, label1, label2, label3)\n           \n            if np.random.rand()&lt;0.5:\n                images, targets = mixup(images, label1, label2, label3, 0.4)\n                output1, output2, output3 = model(images)\n                loss = mixup_criterion(output1,output2,output3, targets) \n            else:\n                images, targets = cutmix(images, label1, label2, label3, 0.4)\n                output1, output2, output3 = model(images)\n                loss = cutmix_criterion(output1,output2,output3, targets) \n......\n```",
    "724078": "Great! If people are curios what `alpha` should be used in cutmix its (`1.0`). Below is the image from original paper. (Lower top 1% error is better) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F9d76a95113937205c68ec10681e86b53%2FScreen%20Shot%202020-01-20%20at%202.16.37%20PM.png?generation=1579547824680882&amp;alt=media)\n\npaper: \nhttps://arxiv.org/pdf/1905.04899.pdf\n ",
    "722255": "I like your implementation!\nEDIT: I've tried your implementation, it helped a lot!",
    "724956": "the full power of mixup/cutout will not be realized unless your network has the \"right capacity and structure\" to learn the augmentation. \n\ni.e. you need to improve your network design and augmentation at the same time. just doing any one alone may only give limited improvement",
    "763161": "FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：https://www.kaggle.com/c/bengaliai-cv19/discussion/133322",
    "734767": "mixup/cutmix with ohem loss : https://www.kaggle.com/c/bengaliai-cv19/discussion/128637",
    "722775": "If you use/refer opensource, please let others know the original implementation. It seems to be similar to source codes that I'm familiar with. BTW, thanks for sharing and it helps a lot for many people!",
    "722731": "cutmix_criterion == mixup_criterion.",
    "723773": "Thanks for the share. Very informative. \nHave a deserving upvote.\n:) ",
    "723424": "Do you mind sharing you train loss and recall with cutmix and mixup @machinelp ",
    "743517": "Hey @machinelp, Thanks for posting this! I'm trying to implement CutMix on CIFAR-10 dataset instead of directly applying on this competition's data as I'm new to this type of implementations. Now when I train the model on CIFAR-10 , the training loss/accuracy fluctuates way too much but the validation accuracy is stable. Here is the accuracy curve:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1905996%2Fabf72a67151d5abf78daa8ee069d8cfc%2FScreenshot%20(126).png?generation=1581487322942081&amp;alt=media)\n\nIs this normal behavior or I messed up the implementation? I also feel that this might be because of CIFAR-10 image resolution (32 x 32). Here is my complete implementation: https://www.kaggle.com/kaushal2896/cifar-10-simple-cnn-with-cutmix-using-pytorch.\n\nThanks!",
    "740610": "Hi MachineLP thanks for the code. It's really helpful.\n\nHowever may I ask for some tip.\nI have used the above code in my own training implementation, however I am getting a lot of fluctuation and slower convergence, compared with this fastdai implementation: https://www.kaggle.com/iafoss/grapheme-fast-ai-starter-lb-0-964\n\nI have inspected the mixup and loss in this kernel and it does not seem to differ much, but it has much stable loss and converge much quicker. \n\nI am just wondering if you could shed some light into this.\n\nMany thanks!",
    "731419": "Does anyone try above CutMix/MixUp with Keras? Unfortunately, I didn't find any robust keras implementation that could fit such a problem. 🙁 ",
    "767202": "Just interesting, can we use mixup/cutmix simultaneously or we need to pick one",
    "754416": "I have tried the cutmix implementation for FastAI. But it throws a 'tensors of different dimensions while concatenating error'. When I tried to debug, I was able to see that 2 [128,2,3] -sized tensors were concatenated with a [128,1] tensor. How can I rectify the error? The link to the FastAI callback: \nhttps://github.com/oguiza/fastai_extensions/blob/master/fastai_extensions/exp/nb_NewDataAugmentation.py\nPS: I don't know his Kaggle handle.",
    "747426": "@machinelp @haqishen Curious that this implementation of mixup does not weigh grapheme/vowel/conso by 2:1:1? Wondering if any light can be shed onto this",
    "733043": "When you guys are using mixup/cutmix, are you using them separately, i.e. randomly select whether a batch is mixup or cutmix? Or are you combining mixup with cutmix, i.e. first cutmix inside the batch, and then mix up the cutmixed samples?\n\nThanks!",
    "728974": "I noticed this specific implementation uses the same lambda for the whole batch compared to image specific lambda's. Even though it might be only a minor detail, wouldn't it be beneficial to use different cut-ratios, which are determined by the lambda, for each image in the batch?",
    "739948": "",
    "724679": "",
    "724866": "@machinelp   nice work thanks!!",
    "723696": "Thanks!😃  ",
    "723430": "Thanks!",
    "722837": "Nice Thanks",
    "722272": "Nice . Thanks !"
  }
}