{
  "id": 102152,
  "title": "MixUp Visualization : Does it make sense?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102152",
  "author_name": "",
  "post_date": "2019-07-31T10:15:20.504497100Z",
  "votes": 13,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi everybody, there are some mentions on unsuccessful application of <a href=\"https://arxiv.org/abs/1710.09412\"><strong>MixUp Augmentation</strong></a> on the current problem. </p>\n\n<p>To help us better understand mixup in the current context, I just plot a simple mixup example of two images (the first is class 3, and the second is class 0), on various weights (lambdas), and the mixup (combined) label, just to help you see whether the combined label makes sense or not.</p>\n\n<p>In the case of using pure regression, suppose a weight is 0.5, you will get mixup class (0+3)/2=1.5 (not sure if it sensible)</p>\n\n<p><img src=\"https://i.ibb.co/BzptDkq/Mix-Up-Examples.jpg\" alt=\"MixUp Viz\"></p>\n\n<p>Perhaps in the case of (ordinal) regression, we should come up with better non-linear label combinations?</p>",
  "messages": [
    {
      "id": "588988",
      "postDate": "07/31/2019 10:15:20",
      "content": "<p>Hi everybody, there are some mentions on unsuccessful application of <a href=\"https://arxiv.org/abs/1710.09412\"><strong>MixUp Augmentation</strong></a> on the current problem. </p>\n\n<p>To help us better understand mixup in the current context, I just plot a simple mixup example of two images (the first is class 3, and the second is class 0), on various weights (lambdas), and the mixup (combined) label, just to help you see whether the combined label makes sense or not.</p>\n\n<p>In the case of using pure regression, suppose a weight is 0.5, you will get mixup class (0+3)/2=1.5 (not sure if it sensible)</p>\n\n<p><img src=\"https://i.ibb.co/BzptDkq/Mix-Up-Examples.jpg\" alt=\"MixUp Viz\"></p>\n\n<p>Perhaps in the case of (ordinal) regression, we should come up with better non-linear label combinations?</p>",
      "rawMarkdown": "Hi everybody, there are some mentions on unsuccessful application of [**MixUp Augmentation**](https://arxiv.org/abs/1710.09412) on the current problem. \n\nTo help us better understand mixup in the current context, I just plot a simple mixup example of two images (the first is class 3, and the second is class 0), on various weights (lambdas), and the mixup (combined) label, just to help you see whether the combined label makes sense or not.\n\nIn the case of using pure regression, suppose a weight is 0.5, you will get mixup class (0+3)/2=1.5 (not sure if it sensible)\n\n![MixUp Viz](https://i.ibb.co/BzptDkq/Mix-Up-Examples.jpg)\n\nPerhaps in the case of (ordinal) regression, we should come up with better non-linear label combinations?",
      "votes": null
    },
    {
      "id": "588995",
      "postDate": "07/31/2019 10:25:03",
      "content": "<p>Another example of classes 0 &amp; 4 combined: \n<img src=\"https://i.ibb.co/NSD9BRp/Mix-Up-Examples2.jpg\" alt=\"MixUp2\"></p>",
      "rawMarkdown": "Another example of classes 0 &amp; 4 combined: \n![MixUp2](https://i.ibb.co/NSD9BRp/Mix-Up-Examples2.jpg)",
      "votes": null
    },
    {
      "id": "589007",
      "postDate": "07/31/2019 10:36:06",
      "content": "<p>I've tried mixup, it makes the training process slower and more unstable. But LB only went down a little bit, maybe I don't train it well. I don't know it's useful or not, but I do believe mixup can alleviate overfitting. If I will make the reduction of LB  within acceptable limits, I would choose to use it.</p>",
      "rawMarkdown": "I've tried mixup, it makes the training process slower and more unstable. But LB only went down a little bit, maybe I don't train it well. I don't know it's useful or not, but I do believe mixup can alleviate overfitting. If I will make the reduction of LB  within acceptable limits, I would choose to use it.",
      "votes": null
    },
    {
      "id": "589069",
      "postDate": "07/31/2019 12:27:45",
      "content": "<p>I also implemented it, but it brought down LB score a bit.\nMight work better on the private test set...</p>",
      "rawMarkdown": "I also implemented it, but it brought down LB score a bit.\nMight work better on the private test set...",
      "votes": null
    },
    {
      "id": "589518",
      "postDate": "08/01/2019 03:15:17",
      "content": "<p>At first, I appreciate your all works for the preprocessing. I got much helps from your scripts(kernel). So, Thanks!</p>\n\n<p>I thought the mixup, but I didn't. As you can see above pictures, If we choose two images and the direction of the pupil are different, the mixed-up images have two pupils. It doesn't seem to be reasonable. </p>\n\n<p>And, I think mixing the grading is also dangerous. Is this reasonable to say that if we mix two images with 2:8 ratio, (one is 1 grade, other is 4 grade),  the new mixed-up eye is 3.4 (0.2 +1 + 0.8 * 4 = 3.4)? </p>\n\n<p>I don't think the symptom could be represented in the form of a continuous manner. </p>\n\n<p>So, for now, I think mixing technique is not easy to this competition.</p>\n\n<p>But, I expect an amazing kaggler can solve it! As we know, Kaggler is very creative.</p>",
      "rawMarkdown": "At first, I appreciate your all works for the preprocessing. I got much helps from your scripts(kernel). So, Thanks!\n\nI thought the mixup, but I didn't. As you can see above pictures, If we choose two images and the direction of the pupil are different, the mixed-up images have two pupils. It doesn't seem to be reasonable. \n\nAnd, I think mixing the grading is also dangerous. Is this reasonable to say that if we mix two images with 2:8 ratio, (one is 1 grade, other is 4 grade),  the new mixed-up eye is 3.4 (0.2 +1 + 0.8 * 4 = 3.4)? \n\nI don't think the symptom could be represented in the form of a continuous manner. \n\nSo, for now, I think mixing technique is not easy to this competition.\n\nBut, I expect an amazing kaggler can solve it! As we know, Kaggler is very creative.",
      "votes": null
    },
    {
      "id": "589534",
      "postDate": "08/01/2019 03:32:25",
      "content": "<p><a href=\"/youhanlee\">@youhanlee</a> Thanks for your thought!  You also did very well here, great jobs YouHan !</p>",
      "rawMarkdown": "youhanlee Thanks for your thought!  You also did very well here, great jobs YouHan !",
      "votes": null
    },
    {
      "id": "589901",
      "postDate": "08/01/2019 14:36:03",
      "content": "<p>Thanks for creating this topic! I have been thinking about this lately. I agree that it doesn't work well if you use just mixup training. But there is  one trick which helped me a bit (still very experimental need more tests):</p>\n\n<p>1) Train on Mixup,\n2) Freeze everything until last layer and fine tune for 1-2 epoch model with no Mixup </p>\n\n<p>I got some boost from this method. They only explanation which I was able to come that maybe network generalize better on the first step. And  on the second step instead of predicting average labels you are predicting actual label...  </p>\n\n<p>Hope it helps. </p>\n\n<p>EDIT:  Made a quick picture \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F0a6f9e597739b66ff62c322e40ae1115%2FScreen%20Shot%202019-08-01%20at%2010.53.54%20AM.png?generation=1564671342362318&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for creating this topic! I have been thinking about this lately. I agree that it doesn't work well if you use just mixup training. But there is  one trick which helped me a bit (still very experimental need more tests):\n\n1) Train on Mixup,\n2) Freeze everything until last layer and fine tune for 1-2 epoch model with no Mixup \n\nI got some boost from this method. They only explanation which I was able to come that maybe network generalize better on the first step. And  on the second step instead of predicting average labels you are predicting actual label...  \n\nHope it helps. \n\nEDIT:  Made a quick picture \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F0a6f9e597739b66ff62c322e40ae1115%2FScreen%20Shot%202019-08-01%20at%2010.53.54%20AM.png?generation=1564671342362318&amp;alt=media)",
      "votes": null
    },
    {
      "id": "590235",
      "postDate": "08/01/2019 23:08:38",
      "content": "<p>I am a fan of your quick picture :)</p>",
      "rawMarkdown": "I am a fan of your quick picture :)",
      "votes": null
    },
    {
      "id": "590236",
      "postDate": "08/01/2019 23:09:52",
      "content": "<p>Nice thought <a href=\"/buaazijian\">@buaazijian</a> .</p>",
      "rawMarkdown": "Nice thought @buaazijian .",
      "votes": null
    },
    {
      "id": "590238",
      "postDate": "08/01/2019 23:14:11",
      "content": "<p>Thanks. Just recently learned Microsoft Paint =)</p>",
      "rawMarkdown": "Thanks. Just recently learned Microsoft Paint =)",
      "votes": null
    },
    {
      "id": "591282",
      "postDate": "08/03/2019 12:52:52",
      "content": "<p>Thanks for the nice inputs.  Agree that mix-up in this case  should take in to account the progressive nature of the diabetic retinopathy.   May be mixing adjacent grades  (0,1) or (1,2) or (3,4) should be more meaningful and helps the model to capture the progression? Just a random thought</p>",
      "rawMarkdown": "Thanks for the nice inputs.  Agree that mix-up in this case  should take in to account the progressive nature of the diabetic retinopathy.   May be mixing adjacent grades  (0,1) or (1,2) or (3,4) should be more meaningful and helps the model to capture the progression? Just a random thought",
      "votes": null
    },
    {
      "id": "592233",
      "postDate": "08/05/2019 03:36:40",
      "content": "<p>so do you use avg pool in your custom tail?</p>",
      "rawMarkdown": "so do you use avg pool in your custom tail?",
      "votes": null
    },
    {
      "id": "592575",
      "postDate": "08/05/2019 13:52:53",
      "content": "<p>Hi its up to you. But I had equal success  with  <code>AdaptiveAvgPool2d</code>  or <code>AdaptiveMaxPool2d</code> =) </p>",
      "rawMarkdown": "Hi its up to you. But I had equal success  with  `AdaptiveAvgPool2d`  or `AdaptiveMaxPool2d` =)",
      "votes": null
    },
    {
      "id": "593558",
      "postDate": "08/06/2019 19:45:19",
      "content": "<p>So I've been playing with mixup a bit and my idea was to do mixup with slight modifications:\n1) Enforcing lambda to stay somewhere in [0.4-0.6] \n2) Computing target as <code>max(a,b)</code> instead of <code>a * lambda + b * (1-lambda)</code> </p>\n\n<p>My line of thinking was the following - suppose we have images graded as <code>0</code> and <code>4</code>. If we mix them together, there is still visible signs of Proliferative DR, which does not go anywhere. So it should me max grade of two.</p>",
      "rawMarkdown": "So I've been playing with mixup a bit and my idea was to do mixup with slight modifications:\n1) Enforcing lambda to stay somewhere in [0.4-0.6] \n2) Computing target as `max(a,b)` instead of `a * lambda + b * (1-lambda)` \n\nMy line of thinking was the following - suppose we have images graded as `0` and `4`. If we mix them together, there is still visible signs of Proliferative DR, which does not go anywhere. So it should me max grade of two.",
      "votes": null
    },
    {
      "id": "597162",
      "postDate": "08/12/2019 00:09:03",
      "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> Thanks Eugene. I had a similar line of thoughts. Let us try what should work best!!</p>",
      "rawMarkdown": "bloodaxe Thanks Eugene. I had a similar line of thoughts. Let us try what should work best!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 588995,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "07/31/2019 10:25:03",
      "content": "<p>Another example of classes 0 &amp; 4 combined: \n<img src=\"https://i.ibb.co/NSD9BRp/Mix-Up-Examples2.jpg\" alt=\"MixUp2\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 589007,
      "author_name": "buaazijian",
      "author_url": "",
      "post_date": "07/31/2019 10:36:06",
      "content": "<p>I've tried mixup, it makes the training process slower and more unstable. But LB only went down a little bit, maybe I don't train it well. I don't know it's useful or not, but I do believe mixup can alleviate overfitting. If I will make the reduction of LB  within acceptable limits, I would choose to use it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 590236,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "08/01/2019 23:09:52",
          "content": "<p>Nice thought <a href=\"/buaazijian\">@buaazijian</a> .</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 589069,
      "author_name": "nemethpeti",
      "author_url": "",
      "post_date": "07/31/2019 12:27:45",
      "content": "<p>I also implemented it, but it brought down LB score a bit.\nMight work better on the private test set...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 589518,
      "author_name": "youhanlee",
      "author_url": "",
      "post_date": "08/01/2019 03:15:17",
      "content": "<p>At first, I appreciate your all works for the preprocessing. I got much helps from your scripts(kernel). So, Thanks!</p>\n\n<p>I thought the mixup, but I didn't. As you can see above pictures, If we choose two images and the direction of the pupil are different, the mixed-up images have two pupils. It doesn't seem to be reasonable. </p>\n\n<p>And, I think mixing the grading is also dangerous. Is this reasonable to say that if we mix two images with 2:8 ratio, (one is 1 grade, other is 4 grade),  the new mixed-up eye is 3.4 (0.2 +1 + 0.8 * 4 = 3.4)? </p>\n\n<p>I don't think the symptom could be represented in the form of a continuous manner. </p>\n\n<p>So, for now, I think mixing technique is not easy to this competition.</p>\n\n<p>But, I expect an amazing kaggler can solve it! As we know, Kaggler is very creative.</p>",
      "votes": null,
      "replies": [
        {
          "id": 589534,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "08/01/2019 03:32:25",
          "content": "<p><a href=\"/youhanlee\">@youhanlee</a> Thanks for your thought!  You also did very well here, great jobs YouHan !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 589901,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "08/01/2019 14:36:03",
      "content": "<p>Thanks for creating this topic! I have been thinking about this lately. I agree that it doesn't work well if you use just mixup training. But there is  one trick which helped me a bit (still very experimental need more tests):</p>\n\n<p>1) Train on Mixup,\n2) Freeze everything until last layer and fine tune for 1-2 epoch model with no Mixup </p>\n\n<p>I got some boost from this method. They only explanation which I was able to come that maybe network generalize better on the first step. And  on the second step instead of predicting average labels you are predicting actual label...  </p>\n\n<p>Hope it helps. </p>\n\n<p>EDIT:  Made a quick picture \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F0a6f9e597739b66ff62c322e40ae1115%2FScreen%20Shot%202019-08-01%20at%2010.53.54%20AM.png?generation=1564671342362318&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 590235,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "08/01/2019 23:08:38",
          "content": "<p>I am a fan of your quick picture :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 590238,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "08/01/2019 23:14:11",
          "content": "<p>Thanks. Just recently learned Microsoft Paint =)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592233,
          "author_name": "sidhanthholalkere",
          "author_url": "",
          "post_date": "08/05/2019 03:36:40",
          "content": "<p>so do you use avg pool in your custom tail?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 592575,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "08/05/2019 13:52:53",
          "content": "<p>Hi its up to you. But I had equal success  with  <code>AdaptiveAvgPool2d</code>  or <code>AdaptiveMaxPool2d</code> =) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 591282,
      "author_name": "riyasvk",
      "author_url": "",
      "post_date": "08/03/2019 12:52:52",
      "content": "<p>Thanks for the nice inputs.  Agree that mix-up in this case  should take in to account the progressive nature of the diabetic retinopathy.   May be mixing adjacent grades  (0,1) or (1,2) or (3,4) should be more meaningful and helps the model to capture the progression? Just a random thought</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 593558,
      "author_name": "bloodaxe",
      "author_url": "",
      "post_date": "08/06/2019 19:45:19",
      "content": "<p>So I've been playing with mixup a bit and my idea was to do mixup with slight modifications:\n1) Enforcing lambda to stay somewhere in [0.4-0.6] \n2) Computing target as <code>max(a,b)</code> instead of <code>a * lambda + b * (1-lambda)</code> </p>\n\n<p>My line of thinking was the following - suppose we have images graded as <code>0</code> and <code>4</code>. If we mix them together, there is still visible signs of Proliferative DR, which does not go anywhere. So it should me max grade of two.</p>",
      "votes": null,
      "replies": [
        {
          "id": 597162,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "08/12/2019 00:09:03",
          "content": "<p><a href=\"/bloodaxe\">@bloodaxe</a> Thanks Eugene. I had a similar line of thoughts. Let us try what should work best!!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "588988": "Hi everybody, there are some mentions on unsuccessful application of [**MixUp Augmentation**](https://arxiv.org/abs/1710.09412) on the current problem. \n\nTo help us better understand mixup in the current context, I just plot a simple mixup example of two images (the first is class 3, and the second is class 0), on various weights (lambdas), and the mixup (combined) label, just to help you see whether the combined label makes sense or not.\n\nIn the case of using pure regression, suppose a weight is 0.5, you will get mixup class (0+3)/2=1.5 (not sure if it sensible)\n\n![MixUp Viz](https://i.ibb.co/BzptDkq/Mix-Up-Examples.jpg)\n\nPerhaps in the case of (ordinal) regression, we should come up with better non-linear label combinations?",
    "588995": "Another example of classes 0 &amp; 4 combined: \n![MixUp2](https://i.ibb.co/NSD9BRp/Mix-Up-Examples2.jpg)",
    "589007": "I've tried mixup, it makes the training process slower and more unstable. But LB only went down a little bit, maybe I don't train it well. I don't know it's useful or not, but I do believe mixup can alleviate overfitting. If I will make the reduction of LB  within acceptable limits, I would choose to use it.",
    "589069": "I also implemented it, but it brought down LB score a bit.\nMight work better on the private test set...",
    "589518": "At first, I appreciate your all works for the preprocessing. I got much helps from your scripts(kernel). So, Thanks!\n\nI thought the mixup, but I didn't. As you can see above pictures, If we choose two images and the direction of the pupil are different, the mixed-up images have two pupils. It doesn't seem to be reasonable. \n\nAnd, I think mixing the grading is also dangerous. Is this reasonable to say that if we mix two images with 2:8 ratio, (one is 1 grade, other is 4 grade),  the new mixed-up eye is 3.4 (0.2 +1 + 0.8 * 4 = 3.4)? \n\nI don't think the symptom could be represented in the form of a continuous manner. \n\nSo, for now, I think mixing technique is not easy to this competition.\n\nBut, I expect an amazing kaggler can solve it! As we know, Kaggler is very creative.",
    "589534": "youhanlee Thanks for your thought!  You also did very well here, great jobs YouHan !",
    "589901": "Thanks for creating this topic! I have been thinking about this lately. I agree that it doesn't work well if you use just mixup training. But there is  one trick which helped me a bit (still very experimental need more tests):\n\n1) Train on Mixup,\n2) Freeze everything until last layer and fine tune for 1-2 epoch model with no Mixup \n\nI got some boost from this method. They only explanation which I was able to come that maybe network generalize better on the first step. And  on the second step instead of predicting average labels you are predicting actual label...  \n\nHope it helps. \n\nEDIT:  Made a quick picture \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F0a6f9e597739b66ff62c322e40ae1115%2FScreen%20Shot%202019-08-01%20at%2010.53.54%20AM.png?generation=1564671342362318&amp;alt=media)",
    "590235": "I am a fan of your quick picture :)",
    "590236": "Nice thought @buaazijian .",
    "590238": "Thanks. Just recently learned Microsoft Paint =)",
    "591282": "Thanks for the nice inputs.  Agree that mix-up in this case  should take in to account the progressive nature of the diabetic retinopathy.   May be mixing adjacent grades  (0,1) or (1,2) or (3,4) should be more meaningful and helps the model to capture the progression? Just a random thought",
    "592233": "so do you use avg pool in your custom tail?",
    "592575": "Hi its up to you. But I had equal success  with  `AdaptiveAvgPool2d`  or `AdaptiveMaxPool2d` =)",
    "593558": "So I've been playing with mixup a bit and my idea was to do mixup with slight modifications:\n1) Enforcing lambda to stay somewhere in [0.4-0.6] \n2) Computing target as `max(a,b)` instead of `a * lambda + b * (1-lambda)` \n\nMy line of thinking was the following - suppose we have images graded as `0` and `4`. If we mix them together, there is still visible signs of Proliferative DR, which does not go anywhere. So it should me max grade of two.",
    "597162": "bloodaxe Thanks Eugene. I had a similar line of thoughts. Let us try what should work best!!"
  },
  "source": "meta"
}