{
  "id": 174503,
  "title": "MixUp implementation for Tensorflow datasets",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/174503",
  "author_name": "",
  "post_date": "2020-08-13T19:34:55.235476500Z",
  "votes": 6,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hey everyone, I was going to give a try to MixUp on this competition, so I made an adaptation from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu#Display-MixUp-Augmentation\" target=\"_blank\">MixUp implementations</a>, the thing is when I use it the regular way my epochs gets 5x slower.</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>\n<p>If I remove the unbatch and shuffle, epochs run a lot faster but I get bad results as if something is messing the data.</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\n#ds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>\n<p>Any one have a working implementation that runs on normal time and give the expected results?</p>\n<p>This is my current implementation:</p>\n<pre><code>def mixup(image, label, alpha=0.1, height=256, width=256, channels=3, batch_size,  classes=1):\n    # input image - is a batch of images of size [batch_size, height, width, channels] not a single image of [height, width, channels]\n    # output - a batch of images with mixup applied\n\n    imgs = []; labs = []\n    for j in range(batch_size):\n        # random chose images\n        k = tf.cast(tf.random.uniform([], 0, batch_size), tf.int32)\n        # MixUp image\n        img1 = image[j,]\n        img2 = image[k,]\n        imgs.append((1-alpha)*img1 + alpha*img2)\n        # MixUp label\n        if classes &gt; 1: # multi-class\n            lab1 = tf.one_hot(label[j], classes)\n            lab2 = tf.one_hot(label[k], classes)\n        else:           # binary\n            lab1 = label[j,]\n            lab2 = label[k,]\n        labs.append((1-alpha)*lab1 + alpha*lab2)\n\n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR\n    image2 = tf.reshape(tf.stack(imgs), (batch_size, height, width, 3))\n    label2 = tf.reshape(tf.stack(labs), (batch_size, classes))\n\n    return image2, label2\n</code></pre>",
  "messages": [
    {
      "id": "969588",
      "postDate": "08/13/2020 19:34:55",
      "content": "<p>Hey everyone, I was going to give a try to MixUp on this competition, so I made an adaptation from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu#Display-MixUp-Augmentation\" target=\"_blank\">MixUp implementations</a>, the thing is when I use it the regular way my epochs gets 5x slower.</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>\n<p>If I remove the unbatch and shuffle, epochs run a lot faster but I get bad results as if something is messing the data.</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\n#ds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>\n<p>Any one have a working implementation that runs on normal time and give the expected results?</p>\n<p>This is my current implementation:</p>\n<pre><code>def mixup(image, label, alpha=0.1, height=256, width=256, channels=3, batch_size,  classes=1):\n    # input image - is a batch of images of size [batch_size, height, width, channels] not a single image of [height, width, channels]\n    # output - a batch of images with mixup applied\n\n    imgs = []; labs = []\n    for j in range(batch_size):\n        # random chose images\n        k = tf.cast(tf.random.uniform([], 0, batch_size), tf.int32)\n        # MixUp image\n        img1 = image[j,]\n        img2 = image[k,]\n        imgs.append((1-alpha)*img1 + alpha*img2)\n        # MixUp label\n        if classes &gt; 1: # multi-class\n            lab1 = tf.one_hot(label[j], classes)\n            lab2 = tf.one_hot(label[k], classes)\n        else:           # binary\n            lab1 = label[j,]\n            lab2 = label[k,]\n        labs.append((1-alpha)*lab1 + alpha*lab2)\n\n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR\n    image2 = tf.reshape(tf.stack(imgs), (batch_size, height, width, 3))\n    label2 = tf.reshape(tf.stack(labs), (batch_size, classes))\n\n    return image2, label2\n</code></pre>",
      "rawMarkdown": "Hey everyone, I was going to give a try to MixUp on this competition, so I made an adaptation from @cdeotte [MixUp implementations](https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu#Display-MixUp-Augmentation), the thing is when I use it the regular way my epochs gets 5x slower.\n\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```\n\nIf I remove the unbatch and shuffle, epochs run a lot faster but I get bad results as if something is messing the data.\n\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\n#ds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```\n\nAny one have a working implementation that runs on normal time and give the expected results?\n\nThis is my current implementation:\n```\ndef mixup(image, label, alpha=0.1, height=256, width=256, channels=3, batch_size,  classes=1):\n    # input image - is a batch of images of size [batch_size, height, width, channels] not a single image of [height, width, channels]\n    # output - a batch of images with mixup applied\n    \n    imgs = []; labs = []\n    for j in range(batch_size):\n        # random chose images\n        k = tf.cast(tf.random.uniform([], 0, batch_size), tf.int32)\n        # MixUp image\n        img1 = image[j,]\n        img2 = image[k,]\n        imgs.append((1-alpha)*img1 + alpha*img2)\n        # MixUp label\n        if classes > 1: # multi-class\n            lab1 = tf.one_hot(label[j], classes)\n            lab2 = tf.one_hot(label[k], classes)\n        else:           # binary\n            lab1 = label[j,]\n            lab2 = label[k,]\n        labs.append((1-alpha)*lab1 + alpha*lab2)\n            \n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR\n    image2 = tf.reshape(tf.stack(imgs), (batch_size, height, width, 3))\n    label2 = tf.reshape(tf.stack(labs), (batch_size, classes))\n\n    return image2, label2\n```",
      "votes": null
    },
    {
      "id": "969600",
      "postDate": "08/13/2020 19:46:11",
      "content": "<p>In the notebook <a href=\"https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations\" target=\"_blank\">here</a>, user Yih-Dar SHIEH converted my MixUp to be 5-10x faster doing entire batch at once. (Note this will require more CPU RAM which may be a problem for Kaggle's GPU but probably not Kaggle's TPU). </p>",
      "rawMarkdown": "In the notebook [here][1], user Yih-Dar SHIEH converted my MixUp to be 5-10x faster doing entire batch at once. (Note this will require more CPU RAM which may be a problem for Kaggle's GPU but probably not Kaggle's TPU). \n\n[1]: https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations",
      "votes": null
    },
    {
      "id": "969642",
      "postDate": "08/13/2020 20:16:03",
      "content": "<p>Thanks you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I will give it a try.</p>",
      "rawMarkdown": "Thanks you @cdeotte I will give it a try.",
      "votes": null
    },
    {
      "id": "969671",
      "postDate": "08/13/2020 20:55:31",
      "content": "<p>I am no expert but 'for loop' seems like main issue here, I also wonder how's the mixup works on this competition, if you get any results would be great if you share them!</p>",
      "rawMarkdown": "I am no expert but 'for loop' seems like main issue here, I also wonder how's the mixup works on this competition, if you get any results would be great if you share them!",
      "votes": null
    },
    {
      "id": "969733",
      "postDate": "08/13/2020 22:25:20",
      "content": "<p>Yeah <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> , probably using loops is not ideal, but I also use loops on my implementations of Cutout augmentation and it works with good speed, I think the issue is because MixUp works on batch level, so the suggestion by Chris may improve things.</p>\n<p>Some people reported that it does not improve results, but I would like to try anyway.</p>",
      "rawMarkdown": "Yeah @datafan07 , probably using loops is not ideal, but I also use loops on my implementations of Cutout augmentation and it works with good speed, I think the issue is because MixUp works on batch level, so the suggestion by Chris may improve things.\n\nSome people reported that it does not improve results, but I would like to try anyway.",
      "votes": null
    },
    {
      "id": "970612",
      "postDate": "08/14/2020 15:56:42",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>, I have not used it in this competition, but if you share a stripped down version of it in a kernel, it will probably be easier to get help.</p>",
      "rawMarkdown": "dimitreoliveira, I have not used it in this competition, but if you share a stripped down version of it in a kernel, it will probably be easier to get help.",
      "votes": null
    },
    {
      "id": "970908",
      "postDate": "08/15/2020 01:59:53",
      "content": "<p>Hey guys, I found what the problem was, in case someone is interested in trying MixUp, the implementation above works, the problem was that I was using the following code:</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>\n<p>but the correct way was this one:</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE']) # batch before augment # not multiply by REPLICAS(8)\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>",
      "rawMarkdown": "Hey guys, I found what the problem was, in case someone is interested in trying MixUp, the implementation above works, the problem was that I was using the following code:\n\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```\n\nbut the correct way was this one:\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE']) # batch before augment # not multiply by REPLICAS(8)\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```",
      "votes": null
    },
    {
      "id": "970909",
      "postDate": "08/15/2020 02:00:22",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> I got it to work on a normal speed.</p>",
      "rawMarkdown": "Thanks @sheriytm I got it to work on a normal speed.",
      "votes": null
    },
    {
      "id": "970911",
      "postDate": "08/15/2020 02:03:14",
      "content": "<p>Thanks for the update</p>",
      "rawMarkdown": "Thanks for the update",
      "votes": null
    },
    {
      "id": "970926",
      "postDate": "08/15/2020 02:33:16",
      "content": "<p>Thanks for the update <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>. If possible, could you report back if you see any improvement or not? If you rather wait until after cmpetition close, I will understand.</p>",
      "rawMarkdown": "Thanks for the update @dimitreoliveira. If possible, could you report back if you see any improvement or not? If you rather wait until after cmpetition close, I will understand.",
      "votes": null
    },
    {
      "id": "971382",
      "postDate": "08/15/2020 13:22:05",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> , for my first experiments I got very similar results, but I was using MixUp with alpha = 0.1, for my next experiments I will try to use MixUp with alpha varying in the range of [0, 1] and also I will remove label smoothing, I think makes more sense to remove label smoothing with MixUp since the labels will already be smoothed a bit.</p>",
      "rawMarkdown": "Hi @sheriytm , for my first experiments I got very similar results, but I was using MixUp with alpha = 0.1, for my next experiments I will try to use MixUp with alpha varying in the range of [0, 1] and also I will remove label smoothing, I think makes more sense to remove label smoothing with MixUp since the labels will already be smoothed a bit.",
      "votes": null
    },
    {
      "id": "971394",
      "postDate": "08/15/2020 13:29:30",
      "content": "<p>I noted something interesting while using MixUp, it seems that MixUp makes the training set harder, but give a similar score, you can see that for the experiment with MixUp the train loss is higher and the train AUC is lower.</p>\n<p>Without MixUp:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F143a923062b512646babf1cb96b1e101%2FScreenshot%20from%202020-08-15%2010-22-41.png?generation=1597497795572770&amp;alt=media\" alt=\"\"></p>\n<p>With MixUp (alpha 0.1):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F2be6ec5e0b8f206f2afdd6152b30c5d1%2FScreenshot%20from%202020-08-15%2010-22-53.png?generation=1597497809272257&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I noted something interesting while using MixUp, it seems that MixUp makes the training set harder, but give a similar score, you can see that for the experiment with MixUp the train loss is higher and the train AUC is lower.\n\nWithout MixUp:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F143a923062b512646babf1cb96b1e101%2FScreenshot%20from%202020-08-15%2010-22-41.png?generation=1597497795572770&alt=media)\n\n\nWith MixUp (alpha 0.1):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F2be6ec5e0b8f206f2afdd6152b30c5d1%2FScreenshot%20from%202020-08-15%2010-22-53.png?generation=1597497809272257&alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969600,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "08/13/2020 19:46:11",
      "content": "<p>In the notebook <a href=\"https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations\" target=\"_blank\">here</a>, user Yih-Dar SHIEH converted my MixUp to be 5-10x faster doing entire batch at once. (Note this will require more CPU RAM which may be a problem for Kaggle's GPU but probably not Kaggle's TPU). </p>",
      "votes": null,
      "replies": [
        {
          "id": 969642,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "08/13/2020 20:16:03",
          "content": "<p>Thanks you <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I will give it a try.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 969671,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "08/13/2020 20:55:31",
      "content": "<p>I am no expert but 'for loop' seems like main issue here, I also wonder how's the mixup works on this competition, if you get any results would be great if you share them!</p>",
      "votes": null,
      "replies": [
        {
          "id": 969733,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "08/13/2020 22:25:20",
          "content": "<p>Yeah <a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> , probably using loops is not ideal, but I also use loops on my implementations of Cutout augmentation and it works with good speed, I think the issue is because MixUp works on batch level, so the suggestion by Chris may improve things.</p>\n<p>Some people reported that it does not improve results, but I would like to try anyway.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970612,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "08/14/2020 15:56:42",
      "content": "<p><a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>, I have not used it in this competition, but if you share a stripped down version of it in a kernel, it will probably be easier to get help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 970909,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "08/15/2020 02:00:22",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> I got it to work on a normal speed.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 970926,
          "author_name": "sheriytm",
          "author_url": "",
          "post_date": "08/15/2020 02:33:16",
          "content": "<p>Thanks for the update <a href=\"https://www.kaggle.com/dimitreoliveira\" target=\"_blank\">@dimitreoliveira</a>. If possible, could you report back if you see any improvement or not? If you rather wait until after cmpetition close, I will understand.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 971382,
          "author_name": "dimitreoliveira",
          "author_url": "",
          "post_date": "08/15/2020 13:22:05",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sheriytm\" target=\"_blank\">@sheriytm</a> , for my first experiments I got very similar results, but I was using MixUp with alpha = 0.1, for my next experiments I will try to use MixUp with alpha varying in the range of [0, 1] and also I will remove label smoothing, I think makes more sense to remove label smoothing with MixUp since the labels will already be smoothed a bit.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 970908,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "08/15/2020 01:59:53",
      "content": "<p>Hey guys, I found what the problem was, in case someone is interested in trying MixUp, the implementation above works, the problem was that I was using the following code:</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>\n<p>but the correct way was this one:</p>\n<pre><code>... regular stuff here\nds = ds.batch(config['BATCH_SIZE']) # batch before augment # not multiply by REPLICAS(8)\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 970911,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "08/15/2020 02:03:14",
          "content": "<p>Thanks for the update</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 971394,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "08/15/2020 13:29:30",
      "content": "<p>I noted something interesting while using MixUp, it seems that MixUp makes the training set harder, but give a similar score, you can see that for the experiment with MixUp the train loss is higher and the train AUC is lower.</p>\n<p>Without MixUp:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F143a923062b512646babf1cb96b1e101%2FScreenshot%20from%202020-08-15%2010-22-41.png?generation=1597497795572770&amp;alt=media\" alt=\"\"></p>\n<p>With MixUp (alpha 0.1):<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F2be6ec5e0b8f206f2afdd6152b30c5d1%2FScreenshot%20from%202020-08-15%2010-22-53.png?generation=1597497809272257&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "969588": "Hey everyone, I was going to give a try to MixUp on this competition, so I made an adaptation from @cdeotte [MixUp implementations](https://www.kaggle.com/cdeotte/cutmix-and-mixup-on-gpu-tpu#Display-MixUp-Augmentation), the thing is when I use it the regular way my epochs gets 5x slower.\n\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```\n\nIf I remove the unbatch and shuffle, epochs run a lot faster but I get bad results as if something is messing the data.\n\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\n#ds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```\n\nAny one have a working implementation that runs on normal time and give the expected results?\n\nThis is my current implementation:\n```\ndef mixup(image, label, alpha=0.1, height=256, width=256, channels=3, batch_size,  classes=1):\n    # input image - is a batch of images of size [batch_size, height, width, channels] not a single image of [height, width, channels]\n    # output - a batch of images with mixup applied\n    \n    imgs = []; labs = []\n    for j in range(batch_size):\n        # random chose images\n        k = tf.cast(tf.random.uniform([], 0, batch_size), tf.int32)\n        # MixUp image\n        img1 = image[j,]\n        img2 = image[k,]\n        imgs.append((1-alpha)*img1 + alpha*img2)\n        # MixUp label\n        if classes > 1: # multi-class\n            lab1 = tf.one_hot(label[j], classes)\n            lab2 = tf.one_hot(label[k], classes)\n        else:           # binary\n            lab1 = label[j,]\n            lab2 = label[k,]\n        labs.append((1-alpha)*lab1 + alpha*lab2)\n            \n    # RESHAPE HACK SO TPU COMPILER KNOWS SHAPE OF OUTPUT TENSOR\n    image2 = tf.reshape(tf.stack(imgs), (batch_size, height, width, 3))\n    label2 = tf.reshape(tf.stack(labs), (batch_size, classes))\n\n    return image2, label2\n```",
    "969600": "In the notebook [here][1], user Yih-Dar SHIEH converted my MixUp to be 5-10x faster doing entire batch at once. (Note this will require more CPU RAM which may be a problem for Kaggle's GPU but probably not Kaggle's TPU). \n\n[1]: https://www.kaggle.com/yihdarshieh/batch-implementation-of-more-data-augmentations",
    "969642": "Thanks you @cdeotte I will give it a try.",
    "969671": "I am no expert but 'for loop' seems like main issue here, I also wonder how's the mixup works on this competition, if you get any results would be great if you share them!",
    "969733": "Yeah @datafan07 , probably using loops is not ideal, but I also use loops on my implementations of Cutout augmentation and it works with good speed, I think the issue is because MixUp works on batch level, so the suggestion by Chris may improve things.\n\nSome people reported that it does not improve results, but I would like to try anyway.",
    "970612": "dimitreoliveira, I have not used it in this competition, but if you share a stripped down version of it in a kernel, it will probably be easier to get help.",
    "970908": "Hey guys, I found what the problem was, in case someone is interested in trying MixUp, the implementation above works, the problem was that I was using the following code:\n\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS) # batch before augment\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```\n\nbut the correct way was this one:\n```\n... regular stuff here\nds = ds.batch(config['BATCH_SIZE']) # batch before augment # not multiply by REPLICAS(8)\nds = ds.map(mixup, num_parallel_calls=AUTO)\nds = ds.unbatch().shuffle(1024*8) # unbatch to shuffle\nds = ds.batch(config['BATCH_SIZE'] * REPLICAS).prefetch(AUTO)\n```",
    "970909": "Thanks @sheriytm I got it to work on a normal speed.",
    "970911": "Thanks for the update",
    "970926": "Thanks for the update @dimitreoliveira. If possible, could you report back if you see any improvement or not? If you rather wait until after cmpetition close, I will understand.",
    "971382": "Hi @sheriytm , for my first experiments I got very similar results, but I was using MixUp with alpha = 0.1, for my next experiments I will try to use MixUp with alpha varying in the range of [0, 1] and also I will remove label smoothing, I think makes more sense to remove label smoothing with MixUp since the labels will already be smoothed a bit.",
    "971394": "I noted something interesting while using MixUp, it seems that MixUp makes the training set harder, but give a similar score, you can see that for the experiment with MixUp the train loss is higher and the train AUC is lower.\n\nWithout MixUp:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F143a923062b512646babf1cb96b1e101%2FScreenshot%20from%202020-08-15%2010-22-41.png?generation=1597497795572770&alt=media)\n\n\nWith MixUp (alpha 0.1):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1182060%2F2be6ec5e0b8f206f2afdd6152b30c5d1%2FScreenshot%20from%202020-08-15%2010-22-53.png?generation=1597497809272257&alt=media)"
  },
  "source": "meta"
}