{
  "id": 126557,
  "title": "Isn't Cutmix a bit risky for this competition?",
  "url": "/competitions/bengaliai-cv19/discussion/126557",
  "author_name": "",
  "post_date": "2020-01-18T10:43:12.199506700Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I've seen many high public LB scoring kagglers talk about the use of Cutmix in this competition which has seemed to improve their results. I just wanted to ask if this wasn't maybe prone to give worse results on the private test set, because of the nature of the test set and what cutmix does. Let me explain:</p>\n\n<p>First of all, I'm pretty new to these more advanced techniques, and still learning a lot, so my understanding might be flawed as well. In that case, please let me know, I'd be happy to be wrong and learn!</p>\n\n<p>If I understand correctly, Cutmix crops random parts of different images, pastes them together and then attributes a weighted label to the newly created image based on how much of the original images are in the newly created one, as follows (from the <a href=\"https://github.com/clovaai/CutMix-PyTorch\">github repo of the project</a>:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2Ff29492171d83dfa6b6fcae2af414fcf8%2FCutmix_exmaple.png?generation=1579343294489994&amp;alt=media\" alt=\"\"></p>\n\n<p>I get that this could be great for cats and dogs, but when it comes to multi-label characters like this, I am a little bit less convinced. Some of the vowels or consonants are in just one part of the image and not all of it, but the label only tells you that the image contains that label, not where it is. So from there, we could image a random crop that would take part of the image NOT containing the vowel or consonant, and stitching it in a new image and still labeling it as containing it.</p>\n\n<p>I tried to make a simple illustrated example here (number and labels are purely examples) where the resulting image has some label of a vowel not even present anymore in the final image.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F776f19eb12023e37d9b980096f197692%2FBengali.aiMixup.png?generation=1579344075293328&amp;alt=media\" alt=\"\"></p>\n\n<p>Wouldn't this be particularly bad, since the private test set contains images with unseen combinations of roots, vowels and consonants?</p>\n\n<p>Just a thought that crossed my mind, curious to hear what those who have implemented it think.</p>",
  "messages": [
    {
      "id": "722266",
      "postDate": "01/18/2020 10:43:12",
      "content": "<p>I've seen many high public LB scoring kagglers talk about the use of Cutmix in this competition which has seemed to improve their results. I just wanted to ask if this wasn't maybe prone to give worse results on the private test set, because of the nature of the test set and what cutmix does. Let me explain:</p>\n\n<p>First of all, I'm pretty new to these more advanced techniques, and still learning a lot, so my understanding might be flawed as well. In that case, please let me know, I'd be happy to be wrong and learn!</p>\n\n<p>If I understand correctly, Cutmix crops random parts of different images, pastes them together and then attributes a weighted label to the newly created image based on how much of the original images are in the newly created one, as follows (from the <a href=\"https://github.com/clovaai/CutMix-PyTorch\">github repo of the project</a>:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2Ff29492171d83dfa6b6fcae2af414fcf8%2FCutmix_exmaple.png?generation=1579343294489994&amp;alt=media\" alt=\"\"></p>\n\n<p>I get that this could be great for cats and dogs, but when it comes to multi-label characters like this, I am a little bit less convinced. Some of the vowels or consonants are in just one part of the image and not all of it, but the label only tells you that the image contains that label, not where it is. So from there, we could image a random crop that would take part of the image NOT containing the vowel or consonant, and stitching it in a new image and still labeling it as containing it.</p>\n\n<p>I tried to make a simple illustrated example here (number and labels are purely examples) where the resulting image has some label of a vowel not even present anymore in the final image.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F776f19eb12023e37d9b980096f197692%2FBengali.aiMixup.png?generation=1579344075293328&amp;alt=media\" alt=\"\"></p>\n\n<p>Wouldn't this be particularly bad, since the private test set contains images with unseen combinations of roots, vowels and consonants?</p>\n\n<p>Just a thought that crossed my mind, curious to hear what those who have implemented it think.</p>",
      "rawMarkdown": "I've seen many high public LB scoring kagglers talk about the use of Cutmix in this competition which has seemed to improve their results. I just wanted to ask if this wasn't maybe prone to give worse results on the private test set, because of the nature of the test set and what cutmix does. Let me explain:\n\nFirst of all, I'm pretty new to these more advanced techniques, and still learning a lot, so my understanding might be flawed as well. In that case, please let me know, I'd be happy to be wrong and learn!\n\nIf I understand correctly, Cutmix crops random parts of different images, pastes them together and then attributes a weighted label to the newly created image based on how much of the original images are in the newly created one, as follows (from the [github repo of the project](https://github.com/clovaai/CutMix-PyTorch):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2Ff29492171d83dfa6b6fcae2af414fcf8%2FCutmix_exmaple.png?generation=1579343294489994&amp;alt=media)\n\n I get that this could be great for cats and dogs, but when it comes to multi-label characters like this, I am a little bit less convinced. Some of the vowels or consonants are in just one part of the image and not all of it, but the label only tells you that the image contains that label, not where it is. So from there, we could image a random crop that would take part of the image NOT containing the vowel or consonant, and stitching it in a new image and still labeling it as containing it.\n\nI tried to make a simple illustrated example here (number and labels are purely examples) where the resulting image has some label of a vowel not even present anymore in the final image.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F776f19eb12023e37d9b980096f197692%2FBengali.aiMixup.png?generation=1579344075293328&amp;alt=media)\n\nWouldn't this be particularly bad, since the private test set contains images with unseen combinations of roots, vowels and consonants?\n\nJust a thought that crossed my mind, curious to hear what those who have implemented it think.",
      "votes": null
    },
    {
      "id": "722310",
      "postDate": "01/18/2020 11:39:32",
      "content": "<p>For me both cutout and cutmix does not improve my score (maybe i did something wrong) but i didn't tried to train my model for more than 5 epochs, while mixup seems effective (i got an improvement of ~0.5% LB score in 10 epochs).</p>",
      "rawMarkdown": "For me both cutout and cutmix does not improve my score (maybe i did something wrong) but i didn't tried to train my model for more than 5 epochs, while mixup seems effective (i got an improvement of ~0.5% LB score in 10 epochs).",
      "votes": null
    },
    {
      "id": "722321",
      "postDate": "01/18/2020 11:49:31",
      "content": "<p>how about cut and mix like the following\n1) grapheme_root without constant or vowel\n2) grapheme_root with constant </p>\n\n<p>hence mixing (1)+(2) will only have one constant</p>\n\n<p>do not cut and mix randomly. choose carefully from the graphemes will mix \"correctly\"</p>\n\n<p>further, the crop need not to be random, e.g. vowel is always at bottom, or right, left etc</p>",
      "rawMarkdown": "how about cut and mix like the following\n1) grapheme\\_root without constant or vowel\n2) grapheme\\_root with constant \n\nhence mixing (1)+(2) will only have one constant\n\ndo not cut and mix randomly. choose carefully from the graphemes will mix \"correctly\"\n\nfurther, the crop need not to be random, e.g. vowel is always at bottom, or right, left etc",
      "votes": null
    },
    {
      "id": "722557",
      "postDate": "01/18/2020 17:58:12",
      "content": "<p>That already sounds like a more interesting idea!</p>\n\n<p>Hopefully I'll have enough time and GPU time to try those things out.</p>",
      "rawMarkdown": "That already sounds like a more interesting idea!\n\nHopefully I'll have enough time and GPU time to try those things out.",
      "votes": null
    },
    {
      "id": "722562",
      "postDate": "01/18/2020 18:05:53",
      "content": "<p>I'm guessing these techniques provide diminishing returns in times of \"bang for your buck\". I think at first trying to find better models, training them with a bit more epoch maybe, fine-tuning them, those things have shown to work across the board; no matter what type of images.</p>\n\n<p>But I'd still be interested in having more feedback on the topic!</p>",
      "rawMarkdown": "I'm guessing these techniques provide diminishing returns in times of \"bang for your buck\". I think at first trying to find better models, training them with a bit more epoch maybe, fine-tuning them, those things have shown to work across the board; no matter what type of images.\n\nBut I'd still be interested in having more feedback on the topic!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 722310,
      "author_name": "lazcoder",
      "author_url": "",
      "post_date": "01/18/2020 11:39:32",
      "content": "<p>For me both cutout and cutmix does not improve my score (maybe i did something wrong) but i didn't tried to train my model for more than 5 epochs, while mixup seems effective (i got an improvement of ~0.5% LB score in 10 epochs).</p>",
      "votes": null,
      "replies": [
        {
          "id": 722562,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "01/18/2020 18:05:53",
          "content": "<p>I'm guessing these techniques provide diminishing returns in times of \"bang for your buck\". I think at first trying to find better models, training them with a bit more epoch maybe, fine-tuning them, those things have shown to work across the board; no matter what type of images.</p>\n\n<p>But I'd still be interested in having more feedback on the topic!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 722321,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/18/2020 11:49:31",
      "content": "<p>how about cut and mix like the following\n1) grapheme_root without constant or vowel\n2) grapheme_root with constant </p>\n\n<p>hence mixing (1)+(2) will only have one constant</p>\n\n<p>do not cut and mix randomly. choose carefully from the graphemes will mix \"correctly\"</p>\n\n<p>further, the crop need not to be random, e.g. vowel is always at bottom, or right, left etc</p>",
      "votes": null,
      "replies": [
        {
          "id": 722557,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "01/18/2020 17:58:12",
          "content": "<p>That already sounds like a more interesting idea!</p>\n\n<p>Hopefully I'll have enough time and GPU time to try those things out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "722266": "I've seen many high public LB scoring kagglers talk about the use of Cutmix in this competition which has seemed to improve their results. I just wanted to ask if this wasn't maybe prone to give worse results on the private test set, because of the nature of the test set and what cutmix does. Let me explain:\n\nFirst of all, I'm pretty new to these more advanced techniques, and still learning a lot, so my understanding might be flawed as well. In that case, please let me know, I'd be happy to be wrong and learn!\n\nIf I understand correctly, Cutmix crops random parts of different images, pastes them together and then attributes a weighted label to the newly created image based on how much of the original images are in the newly created one, as follows (from the [github repo of the project](https://github.com/clovaai/CutMix-PyTorch):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2Ff29492171d83dfa6b6fcae2af414fcf8%2FCutmix_exmaple.png?generation=1579343294489994&amp;alt=media)\n\n I get that this could be great for cats and dogs, but when it comes to multi-label characters like this, I am a little bit less convinced. Some of the vowels or consonants are in just one part of the image and not all of it, but the label only tells you that the image contains that label, not where it is. So from there, we could image a random crop that would take part of the image NOT containing the vowel or consonant, and stitching it in a new image and still labeling it as containing it.\n\nI tried to make a simple illustrated example here (number and labels are purely examples) where the resulting image has some label of a vowel not even present anymore in the final image.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2605845%2F776f19eb12023e37d9b980096f197692%2FBengali.aiMixup.png?generation=1579344075293328&amp;alt=media)\n\nWouldn't this be particularly bad, since the private test set contains images with unseen combinations of roots, vowels and consonants?\n\nJust a thought that crossed my mind, curious to hear what those who have implemented it think.",
    "722310": "For me both cutout and cutmix does not improve my score (maybe i did something wrong) but i didn't tried to train my model for more than 5 epochs, while mixup seems effective (i got an improvement of ~0.5% LB score in 10 epochs).",
    "722321": "how about cut and mix like the following\n1) grapheme\\_root without constant or vowel\n2) grapheme\\_root with constant \n\nhence mixing (1)+(2) will only have one constant\n\ndo not cut and mix randomly. choose carefully from the graphemes will mix \"correctly\"\n\nfurther, the crop need not to be random, e.g. vowel is always at bottom, or right, left etc",
    "722557": "That already sounds like a more interesting idea!\n\nHopefully I'll have enough time and GPU time to try those things out.",
    "722562": "I'm guessing these techniques provide diminishing returns in times of \"bang for your buck\". I think at first trying to find better models, training them with a bit more epoch maybe, fine-tuning them, those things have shown to work across the board; no matter what type of images.\n\nBut I'd still be interested in having more feedback on the topic!"
  },
  "source": "meta"
}