{
  "id": 128115,
  "title": "mixup/cutmix with label smoothing",
  "url": "/competitions/bengaliai-cv19/discussion/128115",
  "author_name": "",
  "post_date": "2020-01-29T02:52:35.758407500Z",
  "votes": 34,
  "comment_count": 5,
  "views": 0,
  "content": "<p><code>``python\ndef onehot_encoding(label, n_classes):\n    return torch.zeros(label.size(0), n_classes).to(label.device).scatter_(\n        1, label.view(-1, 1), 1)\ndef cross_entropy_loss(input, target, reduction):\n    logp = F.log_softmax(input, dim=1)\n    loss = torch.sum(-logp * target, dim=1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'mean':\n        return loss.mean()\n    elif reduction == 'sum':\n        return loss.sum()\n    else:\n        raise ValueError(\n            '</code>reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\ndef label_smoothing_criterion(epsilon=0.1, reduction='mean'):\n    def _label_smoothing_criterion(preds, targets):\n        n_classes = preds.size(1)\n        device = preds.device</p>\n\n<pre><code>    onehot = onehot_encoding(targets, n_classes).float().to(device)\n    targets = onehot * (1 - epsilon) + torch.ones_like(onehot).to(\n        device) * epsilon / n_classes\n    loss = cross_entropy_loss(preds, targets, reduction)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'mean':\n        return loss.mean()\n    elif reduction == 'sum':\n        return loss.sum()\n    else:\n        raise ValueError(\n            '`reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\n\nreturn _label_smoothing_criterion\n</code></pre>\n\n<h1>criterion1 = label_smoothing_criterion()</h1>\n\n<h1>criterion2 = label_smoothing_criterion()</h1>\n\n<h1>criterion3 = label_smoothing_criterion()</h1>\n\n<p>def rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p>def cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\nbbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\ndata[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n# adjust lambda to exactly match pixel ratio\nlam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\nreturn data, targets\n</code></pre>\n\n<h1>loss 变化</h1>\n\n<p>def cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [ lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]</p>\n\n<p>def mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\ndata = data * lam + shuffled_data * (1 - lam)\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\nreturn data, targets\n</code></pre>\n\n<p>def mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]</p>\n\n<p>```</p>",
  "messages": [
    {
      "id": "731765",
      "postDate": "01/29/2020 02:52:35",
      "content": "<p><code>``python\ndef onehot_encoding(label, n_classes):\n    return torch.zeros(label.size(0), n_classes).to(label.device).scatter_(\n        1, label.view(-1, 1), 1)\ndef cross_entropy_loss(input, target, reduction):\n    logp = F.log_softmax(input, dim=1)\n    loss = torch.sum(-logp * target, dim=1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'mean':\n        return loss.mean()\n    elif reduction == 'sum':\n        return loss.sum()\n    else:\n        raise ValueError(\n            '</code>reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\ndef label_smoothing_criterion(epsilon=0.1, reduction='mean'):\n    def _label_smoothing_criterion(preds, targets):\n        n_classes = preds.size(1)\n        device = preds.device</p>\n\n<pre><code>    onehot = onehot_encoding(targets, n_classes).float().to(device)\n    targets = onehot * (1 - epsilon) + torch.ones_like(onehot).to(\n        device) * epsilon / n_classes\n    loss = cross_entropy_loss(preds, targets, reduction)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'mean':\n        return loss.mean()\n    elif reduction == 'sum':\n        return loss.sum()\n    else:\n        raise ValueError(\n            '`reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\n\nreturn _label_smoothing_criterion\n</code></pre>\n\n<h1>criterion1 = label_smoothing_criterion()</h1>\n\n<h1>criterion2 = label_smoothing_criterion()</h1>\n\n<h1>criterion3 = label_smoothing_criterion()</h1>\n\n<p>def rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)</p>\n\n<pre><code># uniform\ncx = np.random.randint(W)\ncy = np.random.randint(H)\n\nbbx1 = np.clip(cx - cut_w // 2, 0, W)\nbby1 = np.clip(cy - cut_h // 2, 0, H)\nbbx2 = np.clip(cx + cut_w // 2, 0, W)\nbby2 = np.clip(cy + cut_h // 2, 0, H)\n\nreturn bbx1, bby1, bbx2, bby2\n</code></pre>\n\n<p>def cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\nbbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\ndata[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n# adjust lambda to exactly match pixel ratio\nlam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\nreturn data, targets\n</code></pre>\n\n<h1>loss 变化</h1>\n\n<p>def cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [ lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]</p>\n\n<p>def mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]</p>\n\n<pre><code>lam = np.random.beta(alpha, alpha)\ndata = data * lam + shuffled_data * (1 - lam)\ntargets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\nreturn data, targets\n</code></pre>\n\n<p>def mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]</p>\n\n<p>```</p>",
      "rawMarkdown": "```python\ndef onehot_encoding(label, n_classes):\n    return torch.zeros(label.size(0), n_classes).to(label.device).scatter_(\n        1, label.view(-1, 1), 1)\ndef cross_entropy_loss(input, target, reduction):\n    logp = F.log_softmax(input, dim=1)\n    loss = torch.sum(-logp * target, dim=1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'mean':\n        return loss.mean()\n    elif reduction == 'sum':\n        return loss.sum()\n    else:\n        raise ValueError(\n            '`reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\ndef label_smoothing_criterion(epsilon=0.1, reduction='mean'):\n    def _label_smoothing_criterion(preds, targets):\n        n_classes = preds.size(1)\n        device = preds.device\n\n        onehot = onehot_encoding(targets, n_classes).float().to(device)\n        targets = onehot * (1 - epsilon) + torch.ones_like(onehot).to(\n            device) * epsilon / n_classes\n        loss = cross_entropy_loss(preds, targets, reduction)\n        if reduction == 'none':\n            return loss\n        elif reduction == 'mean':\n            return loss.mean()\n        elif reduction == 'sum':\n            return loss.sum()\n        else:\n            raise ValueError(\n                '`reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\n\n    return _label_smoothing_criterion\n#criterion1 = label_smoothing_criterion()\n#criterion2 = label_smoothing_criterion()\n#criterion3 = label_smoothing_criterion()\n\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n\n    return bbx1, bby1, bbx2, bby2\ndef cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    data[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n    return data, targets\n# loss 变化\ndef cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [ lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]\n\n\ndef mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    data = data * lam + shuffled_data * (1 - lam)\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\n    return data, targets\n\n\ndef mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]\n\n```",
      "votes": null
    },
    {
      "id": "734770",
      "postDate": "02/02/2020 01:33:41",
      "content": "<p>mixup/cutmix with ohem loss : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
      "rawMarkdown": "mixup/cutmix with ohem loss : https://www.kaggle.com/c/bengaliai-cv19/discussion/128637",
      "votes": null
    },
    {
      "id": "755085",
      "postDate": "02/24/2020 12:10:37",
      "content": "<p><a href=\"/machinelp\">@machinelp</a> I wonder if you have observed better results when using cutmix/mixup with label smoothing? Seems really counter-intuitive to me though on my to-do list</p>",
      "rawMarkdown": "machinelp I wonder if you have observed better results when using cutmix/mixup with label smoothing? Seems really counter-intuitive to me though on my to-do list",
      "votes": null
    },
    {
      "id": "763163",
      "postDate": "03/04/2020 07:45:46",
      "content": "<p>FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">https://www.kaggle.com/c/bengaliai-cv19/discussion/133322</a></p>",
      "rawMarkdown": "FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：https://www.kaggle.com/c/bengaliai-cv19/discussion/133322",
      "votes": null
    },
    {
      "id": "768605",
      "postDate": "03/11/2020 02:31:16",
      "content": "<p>Thanks for your great code. But I am trying mixup for the first time. Can you give me some example about how to apply it? </p>\n\n<p>As for my understanding, mixup or cutmix are applied in data augmentation such as transforms. And cutmix_loss or mixup_loss are applied loss computation. </p>\n\n<p>Is my undertanding correct?  Thanks for your time : )</p>",
      "rawMarkdown": "Thanks for your great code. But I am trying mixup for the first time. Can you give me some example about how to apply it? \n\nAs for my understanding, mixup or cutmix are applied in data augmentation such as transforms. And cutmix_loss or mixup_loss are applied loss computation. \n\nIs my undertanding correct?  Thanks for your time : )",
      "votes": null
    },
    {
      "id": "1090444",
      "postDate": "11/25/2020 10:34:49",
      "content": "<p><a href=\"https://www.kaggle.com/machinelp\" target=\"_blank\">@machinelp</a> Thank you so much for your posts. At the time of this competition, I was so novice that I didn't understand any of these but after almost a year, in another CV competition (cassava leaf disease classification) these posts are so useful !</p>",
      "rawMarkdown": "machinelp Thank you so much for your posts. At the time of this competition, I was so novice that I didn't understand any of these but after almost a year, in another CV competition (cassava leaf disease classification) these posts are so useful !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1090444,
      "author_name": "kaushal2896",
      "author_url": "",
      "post_date": "11/25/2020 10:34:49",
      "content": "<p><a href=\"https://www.kaggle.com/machinelp\" target=\"_blank\">@machinelp</a> Thank you so much for your posts. At the time of this competition, I was so novice that I didn't understand any of these but after almost a year, in another CV competition (cassava leaf disease classification) these posts are so useful !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 734770,
      "author_name": "machinelp",
      "author_url": "",
      "post_date": "02/02/2020 01:33:41",
      "content": "<p>mixup/cutmix with ohem loss : <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128637\">https://www.kaggle.com/c/bengaliai-cv19/discussion/128637</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 755085,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "02/24/2020 12:10:37",
      "content": "<p><a href=\"/machinelp\">@machinelp</a> I wonder if you have observed better results when using cutmix/mixup with label smoothing? Seems really counter-intuitive to me though on my to-do list</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 763163,
      "author_name": "machinelp",
      "author_url": "",
      "post_date": "03/04/2020 07:45:46",
      "content": "<p>FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/133322\">https://www.kaggle.com/c/bengaliai-cv19/discussion/133322</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 768605,
      "author_name": "bryce1010",
      "author_url": "",
      "post_date": "03/11/2020 02:31:16",
      "content": "<p>Thanks for your great code. But I am trying mixup for the first time. Can you give me some example about how to apply it? </p>\n\n<p>As for my understanding, mixup or cutmix are applied in data augmentation such as transforms. And cutmix_loss or mixup_loss are applied loss computation. </p>\n\n<p>Is my undertanding correct?  Thanks for your time : )</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "731765": "```python\ndef onehot_encoding(label, n_classes):\n    return torch.zeros(label.size(0), n_classes).to(label.device).scatter_(\n        1, label.view(-1, 1), 1)\ndef cross_entropy_loss(input, target, reduction):\n    logp = F.log_softmax(input, dim=1)\n    loss = torch.sum(-logp * target, dim=1)\n    if reduction == 'none':\n        return loss\n    elif reduction == 'mean':\n        return loss.mean()\n    elif reduction == 'sum':\n        return loss.sum()\n    else:\n        raise ValueError(\n            '`reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\ndef label_smoothing_criterion(epsilon=0.1, reduction='mean'):\n    def _label_smoothing_criterion(preds, targets):\n        n_classes = preds.size(1)\n        device = preds.device\n\n        onehot = onehot_encoding(targets, n_classes).float().to(device)\n        targets = onehot * (1 - epsilon) + torch.ones_like(onehot).to(\n            device) * epsilon / n_classes\n        loss = cross_entropy_loss(preds, targets, reduction)\n        if reduction == 'none':\n            return loss\n        elif reduction == 'mean':\n            return loss.mean()\n        elif reduction == 'sum':\n            return loss.sum()\n        else:\n            raise ValueError(\n                '`reduction` must be one of \\'none\\', \\'mean\\', or \\'sum\\'.')\n\n    return _label_smoothing_criterion\n#criterion1 = label_smoothing_criterion()\n#criterion2 = label_smoothing_criterion()\n#criterion3 = label_smoothing_criterion()\n\ndef rand_bbox(size, lam):\n    W = size[2]\n    H = size[3]\n    cut_rat = np.sqrt(1. - lam)\n    cut_w = np.int(W * cut_rat)\n    cut_h = np.int(H * cut_rat)\n\n    # uniform\n    cx = np.random.randint(W)\n    cy = np.random.randint(H)\n\n    bbx1 = np.clip(cx - cut_w // 2, 0, W)\n    bby1 = np.clip(cy - cut_h // 2, 0, H)\n    bbx2 = np.clip(cx + cut_w // 2, 0, W)\n    bby2 = np.clip(cy + cut_h // 2, 0, H)\n\n    return bbx1, bby1, bbx2, bby2\ndef cutmix(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    bbx1, bby1, bbx2, bby2 = rand_bbox(data.size(), lam)\n    data[:, :, bbx1:bbx2, bby1:bby2] = data[indices, :, bbx1:bbx2, bby1:bby2]\n    # adjust lambda to exactly match pixel ratio\n    lam = 1 - ((bbx2 - bbx1) * (bby2 - bby1) / (data.size()[-1] * data.size()[-2]))\n\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n    return data, targets\n# loss 变化\ndef cutmix_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [ lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]\n\n\ndef mixup(data, targets1, targets2, targets3, alpha):\n    indices = torch.randperm(data.size(0))\n    shuffled_data = data[indices]\n    shuffled_targets1 = targets1[indices]\n    shuffled_targets2 = targets2[indices]\n    shuffled_targets3 = targets3[indices]\n    \n    lam = np.random.beta(alpha, alpha)\n    data = data * lam + shuffled_data * (1 - lam)\n    targets = [targets1, shuffled_targets1, targets2, shuffled_targets2, targets3, shuffled_targets3, lam]\n\n    return data, targets\n\n\ndef mixup_criterion(preds1,preds2,preds3, targets):\n    targets1, targets2,targets3, targets4,targets5, targets6, lam = targets[0], targets[1], targets[2], targets[3], targets[4], targets[5], targets[6]\n    criterion = label_smoothing_criterion()\n    return [lam * criterion(preds1, targets1) + (1 - lam) * criterion(preds1, targets2), lam * criterion(preds2, targets3) + (1 - lam) * criterion(preds2, targets4), lam * criterion(preds3, targets5) + (1 - lam) * criterion(preds3, targets6) ]\n\n```",
    "734770": "mixup/cutmix with ohem loss : https://www.kaggle.com/c/bengaliai-cv19/discussion/128637",
    "755085": "machinelp I wonder if you have observed better results when using cutmix/mixup with label smoothing? Seems really counter-intuitive to me though on my to-do list",
    "763163": "FMix: FMix improves performance over MixUp and CutMix for a number of state-of-the- art models across a range of data sets and problem settings：https://www.kaggle.com/c/bengaliai-cv19/discussion/133322",
    "768605": "Thanks for your great code. But I am trying mixup for the first time. Can you give me some example about how to apply it? \n\nAs for my understanding, mixup or cutmix are applied in data augmentation such as transforms. And cutmix_loss or mixup_loss are applied loss computation. \n\nIs my undertanding correct?  Thanks for your time : )",
    "1090444": "machinelp Thank you so much for your posts. At the time of this competition, I was so novice that I didn't understand any of these but after almost a year, in another CV competition (cassava leaf disease classification) these posts are so useful !"
  },
  "source": "meta"
}