{
  "id": 127902,
  "title": "why cutout/mixup works",
  "url": "/competitions/bengaliai-cv19/discussion/127902",
  "author_name": "",
  "post_date": "2020-01-27T14:47:22.805585300Z",
  "votes": 35,
  "comment_count": 13,
  "views": 0,
  "content": "<p>this is the reason ....and we can make it works better via better \"mixing\"</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F856828837128e8e77e54def004736f96%2FSelection_130.png?generation=1580136440020750&amp;alt=media\" alt=\"\"></p>\n\n<p>a crude way is to just e.g. take 10% of the left side of the bounding box to remove vowel. a better way is to add vowel label (pixel label) by hand. yet another way is to use heatmap (e.g. CAM response) as pesudo label</p>",
  "messages": [
    {
      "id": "730488",
      "postDate": "01/27/2020 14:47:22",
      "content": "<p>this is the reason ....and we can make it works better via better \"mixing\"</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F856828837128e8e77e54def004736f96%2FSelection_130.png?generation=1580136440020750&amp;alt=media\" alt=\"\"></p>\n\n<p>a crude way is to just e.g. take 10% of the left side of the bounding box to remove vowel. a better way is to add vowel label (pixel label) by hand. yet another way is to use heatmap (e.g. CAM response) as pesudo label</p>",
      "rawMarkdown": "this is the reason ....and we can make it works better via better \"mixing\"\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F856828837128e8e77e54def004736f96%2FSelection_130.png?generation=1580136440020750&amp;alt=media)\n\n\na crude way is to just e.g. take 10% of the left side of the bounding box to remove vowel. a better way is to add vowel label (pixel label) by hand. yet another way is to use heatmap (e.g. CAM response) as pesudo label",
      "votes": null
    },
    {
      "id": "730490",
      "postDate": "01/27/2020 14:49:17",
      "content": "<p>if there is any GAN (adversarial net) that can learns parts , please let me know.\nboth with and without part labeling are welcome</p>",
      "rawMarkdown": "if there is any GAN (adversarial net) that can learns parts , please let me know.\nboth with and without part labeling are welcome",
      "votes": null
    },
    {
      "id": "730519",
      "postDate": "01/27/2020 15:15:56",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa7e13e9974dd1da673aa07a0750caf1a%2FSelection_052.png?generation=1580138133875470&amp;alt=media\" alt=\"\"></p>\n\n<p>similar network can be made for vowel removal</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa7e13e9974dd1da673aa07a0750caf1a%2FSelection_052.png?generation=1580138133875470&amp;alt=media)\n\n\nsimilar network can be made for vowel removal",
      "votes": null
    },
    {
      "id": "730528",
      "postDate": "01/27/2020 15:22:38",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F315cc2b43553ac7b8e39d78706b1dd14%2FSelection_053.png?generation=1580138554793208&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F315cc2b43553ac7b8e39d78706b1dd14%2FSelection_053.png?generation=1580138554793208&amp;alt=media)",
      "votes": null
    },
    {
      "id": "730529",
      "postDate": "01/27/2020 15:31:53",
      "content": "<p>You can split the unicode characters into \"parts\"</p>\n\n<p><code>\nin [1]: list('ক্ট্রো')\nout [1]: ['ক', '্', 'ট', '্', 'র', 'ো'])\n</code></p>\n\n<p>In this example the labels are:\n- root: 15 (<code>'ক্ট'</code>) -&gt; <code>list('ক্ট') = ['ক', '্', 'ট']</code>\n- vovel_d: 9 (<code>'ো'</code>) -&gt; <code>list('ো') = ['ো']</code>\n- consonant_d: 5 (<code>'্র'</code>) -&gt; <code>list('্র') = ['্', 'র']</code></p>\n\n<p>I used these as a target (multi-label; part of the overall loss). What do you think? Is this could work or I just wasting my time? (my baseline improved, but I am still far from .98LB)  </p>",
      "rawMarkdown": "You can split the unicode characters into \"parts\"\n\n```\nin [1]: list('ক্ট্রো')\nout [1]: ['ক', '্', 'ট', '্', 'র', 'ো'])\n```\n\nIn this example the labels are:\n- root: 15 (`'ক্ট'`) -&gt; `list('ক্ট') = ['ক', '্', 'ট']`\n- vovel_d: 9 (`'ো'`) -&gt; `list('ো') = ['ো']`\n- consonant_d: 5 (`'্র'`) -&gt; `list('্র') = ['্', 'র']`\n\nI used these as a target (multi-label; part of the overall loss). What do you think? Is this could work or I just wasting my time? (my baseline improved, but I am still far from .98LB)",
      "votes": null
    },
    {
      "id": "730548",
      "postDate": "01/27/2020 16:04:19",
      "content": "<p>Hi, Peter, May I ask how to split it to part? Thanks</p>",
      "rawMarkdown": "Hi, Peter, May I ask how to split it to part? Thanks",
      "votes": null
    },
    {
      "id": "730555",
      "postDate": "01/27/2020 16:12:38",
      "content": "<p>simple python <code>list</code> function, like <code>list('grapheme_char')</code>\n(I am using python 3.7; not sure this works with older versions..)</p>",
      "rawMarkdown": "simple python `list` function, like `list('grapheme_char')`\n(I am using python 3.7; not sure this works with older versions..)",
      "votes": null
    },
    {
      "id": "730741",
      "postDate": "01/27/2020 20:53:16",
      "content": "<p>structural code to encode location of vowel (consonant)\n<img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/code.png\" alt=\"\">\n<img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/structure.png\" alt=\"\"><img src=\"https://www.researchgate.net/profile/Cheng-Lin_Liu/publication/220412438/figure/download/fig4/AS:277377025363970@1443143244172/10-types-of-Chinese-character-structures-single-radical-left-right-up-down-up-left.png\" alt=\"\">\n<img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/07f6301809e45582d8330ce9d55f52dda4e42611/2-Figure2-1.png\" alt=\"\"></p>",
      "rawMarkdown": "structural code to encode location of vowel (consonant)\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/code.png)\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/structure.png)![](https://www.researchgate.net/profile/Cheng-Lin_Liu/publication/220412438/figure/download/fig4/AS:277377025363970@1443143244172/10-types-of-Chinese-character-structures-single-radical-left-right-up-down-up-left.png)\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/07f6301809e45582d8330ce9d55f52dda4e42611/2-Figure2-1.png)",
      "votes": null
    },
    {
      "id": "730795",
      "postDate": "01/27/2020 23:45:31",
      "content": "<p>A really slow but good way to do this is to train a unet type model that outputs 3 channels - root, consonant, vowel. Pre-train a classifier on the whole dataset then use that model to tell the unet if it split it well. You can then sum up the different part of the channel - root + cons, root + vowel, root alone and use the pre-trained model for the loss.I also added an intersection loss for the 3 channels and a loss so that the sum of the 3 channels made the same image as the input. I then convolution over the ouput to make the final decision of where each pixel goes. Samples with root 13:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F404040%2F69aab1e1429167b7637d35322cd78ed7%2Fsample_split_output.PNG?generation=1580168716653253&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "A really slow but good way to do this is to train a unet type model that outputs 3 channels - root, consonant, vowel. Pre-train a classifier on the whole dataset then use that model to tell the unet if it split it well. You can then sum up the different part of the channel - root + cons, root + vowel, root alone and use the pre-trained model for the loss.I also added an intersection loss for the 3 channels and a loss so that the sum of the 3 channels made the same image as the input. I then convolution over the ouput to make the final decision of where each pixel goes. Samples with root 13:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F404040%2F69aab1e1429167b7637d35322cd78ed7%2Fsample_split_output.PNG?generation=1580168716653253&amp;alt=media)",
      "votes": null
    },
    {
      "id": "730883",
      "postDate": "01/28/2020 03:56:56",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "731612",
      "postDate": "01/28/2020 20:19:35",
      "content": "<p>Here is <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128059\">one more good thread</a> on generating more data by DrHB </p>\n\n<p>Looks like making the model to generalize well on unseen combinations is the key. Apparently, the data set is <a href=\"https://www.kaggle.com/sibmike/pivots-hint-why-cutmix-works\">constructed in such a way</a> to make us work on this problem. Can we avoid data generation? Maybe if we hide full combination sets for validation we can make the model generalize well on unseen data. Or a combination of these might work.</p>",
      "rawMarkdown": "Here is [one more good thread](https://www.kaggle.com/c/bengaliai-cv19/discussion/128059) on generating more data by DrHB \n\nLooks like making the model to generalize well on unseen combinations is the key. Apparently, the data set is [constructed in such a way](https://www.kaggle.com/sibmike/pivots-hint-why-cutmix-works) to make us work on this problem. Can we avoid data generation? Maybe if we hide full combination sets for validation we can make the model generalize well on unseen data. Or a combination of these might work.",
      "votes": null
    },
    {
      "id": "731757",
      "postDate": "01/29/2020 02:25:59",
      "content": "<p>it is weird, the mixups/cutmix can't improve the cv from my experiments😨 </p>",
      "rawMarkdown": "it is weird, the mixups/cutmix can't improve the cv from my experiments😨",
      "votes": null
    },
    {
      "id": "737104",
      "postDate": "02/04/2020 23:21:51",
      "content": "<p>This seems to be a very good idea <a href=\"/zero00\">@zero00</a>. Have you been able to try it? </p>",
      "rawMarkdown": "This seems to be a very good idea @zero00. Have you been able to try it?",
      "votes": null
    },
    {
      "id": "740791",
      "postDate": "02/09/2020 19:45:20",
      "content": "<p><a href=\"/pestipeti\">@pestipeti</a> This is an amazing idea! I've been working on something similar myself. In fact, I think it will help to create more labels (finer labels) which will, in turn, increase the performance. \nI recently found a related paper on a similar phenomenon of \"fine labels\": \n<a href=\"https://arxiv.org/abs/1901.07012\">Understanding the Impact of Label Granularity on CNN-based Image Classification</a></p>",
      "rawMarkdown": "pestipeti This is an amazing idea! I've been working on something similar myself. In fact, I think it will help to create more labels (finer labels) which will, in turn, increase the performance. \nI recently found a related paper on a similar phenomenon of \"fine labels\": \n[Understanding the Impact of Label Granularity on CNN-based Image Classification](https://arxiv.org/abs/1901.07012)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 730490,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/27/2020 14:49:17",
      "content": "<p>if there is any GAN (adversarial net) that can learns parts , please let me know.\nboth with and without part labeling are welcome</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 730519,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/27/2020 15:15:56",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa7e13e9974dd1da673aa07a0750caf1a%2FSelection_052.png?generation=1580138133875470&amp;alt=media\" alt=\"\"></p>\n\n<p>similar network can be made for vowel removal</p>",
      "votes": null,
      "replies": [
        {
          "id": 730528,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/27/2020 15:22:38",
          "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F315cc2b43553ac7b8e39d78706b1dd14%2FSelection_053.png?generation=1580138554793208&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 730529,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "01/27/2020 15:31:53",
      "content": "<p>You can split the unicode characters into \"parts\"</p>\n\n<p><code>\nin [1]: list('ক্ট্রো')\nout [1]: ['ক', '্', 'ট', '্', 'র', 'ো'])\n</code></p>\n\n<p>In this example the labels are:\n- root: 15 (<code>'ক্ট'</code>) -&gt; <code>list('ক্ট') = ['ক', '্', 'ট']</code>\n- vovel_d: 9 (<code>'ো'</code>) -&gt; <code>list('ো') = ['ো']</code>\n- consonant_d: 5 (<code>'্র'</code>) -&gt; <code>list('্র') = ['্', 'র']</code></p>\n\n<p>I used these as a target (multi-label; part of the overall loss). What do you think? Is this could work or I just wasting my time? (my baseline improved, but I am still far from .98LB)  </p>",
      "votes": null,
      "replies": [
        {
          "id": 730548,
          "author_name": "hesene",
          "author_url": "",
          "post_date": "01/27/2020 16:04:19",
          "content": "<p>Hi, Peter, May I ask how to split it to part? Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 730555,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "01/27/2020 16:12:38",
          "content": "<p>simple python <code>list</code> function, like <code>list('grapheme_char')</code>\n(I am using python 3.7; not sure this works with older versions..)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 730883,
          "author_name": "hesene",
          "author_url": "",
          "post_date": "01/28/2020 03:56:56",
          "content": "<p>Thanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 740791,
          "author_name": "timetraveller98",
          "author_url": "",
          "post_date": "02/09/2020 19:45:20",
          "content": "<p><a href=\"/pestipeti\">@pestipeti</a> This is an amazing idea! I've been working on something similar myself. In fact, I think it will help to create more labels (finer labels) which will, in turn, increase the performance. \nI recently found a related paper on a similar phenomenon of \"fine labels\": \n<a href=\"https://arxiv.org/abs/1901.07012\">Understanding the Impact of Label Granularity on CNN-based Image Classification</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 730741,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/27/2020 20:53:16",
      "content": "<p>structural code to encode location of vowel (consonant)\n<img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/code.png\" alt=\"\">\n<img src=\"https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/structure.png\" alt=\"\"><img src=\"https://www.researchgate.net/profile/Cheng-Lin_Liu/publication/220412438/figure/download/fig4/AS:277377025363970@1443143244172/10-types-of-Chinese-character-structures-single-radical-left-right-up-down-up-left.png\" alt=\"\">\n<img src=\"https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/07f6301809e45582d8330ce9d55f52dda4e42611/2-Figure2-1.png\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 730795,
      "author_name": "zero00",
      "author_url": "",
      "post_date": "01/27/2020 23:45:31",
      "content": "<p>A really slow but good way to do this is to train a unet type model that outputs 3 channels - root, consonant, vowel. Pre-train a classifier on the whole dataset then use that model to tell the unet if it split it well. You can then sum up the different part of the channel - root + cons, root + vowel, root alone and use the pre-trained model for the loss.I also added an intersection loss for the 3 channels and a loss so that the sum of the 3 channels made the same image as the input. I then convolution over the ouput to make the final decision of where each pixel goes. Samples with root 13:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F404040%2F69aab1e1429167b7637d35322cd78ed7%2Fsample_split_output.PNG?generation=1580168716653253&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 737104,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "02/04/2020 23:21:51",
          "content": "<p>This seems to be a very good idea <a href=\"/zero00\">@zero00</a>. Have you been able to try it? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 731612,
      "author_name": "sibmike",
      "author_url": "",
      "post_date": "01/28/2020 20:19:35",
      "content": "<p>Here is <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/128059\">one more good thread</a> on generating more data by DrHB </p>\n\n<p>Looks like making the model to generalize well on unseen combinations is the key. Apparently, the data set is <a href=\"https://www.kaggle.com/sibmike/pivots-hint-why-cutmix-works\">constructed in such a way</a> to make us work on this problem. Can we avoid data generation? Maybe if we hide full combination sets for validation we can make the model generalize well on unseen data. Or a combination of these might work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 731757,
      "author_name": "garybios",
      "author_url": "",
      "post_date": "01/29/2020 02:25:59",
      "content": "<p>it is weird, the mixups/cutmix can't improve the cv from my experiments😨 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "730488": "this is the reason ....and we can make it works better via better \"mixing\"\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F856828837128e8e77e54def004736f96%2FSelection_130.png?generation=1580136440020750&amp;alt=media)\n\n\na crude way is to just e.g. take 10% of the left side of the bounding box to remove vowel. a better way is to add vowel label (pixel label) by hand. yet another way is to use heatmap (e.g. CAM response) as pesudo label",
    "730490": "if there is any GAN (adversarial net) that can learns parts , please let me know.\nboth with and without part labeling are welcome",
    "730519": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fa7e13e9974dd1da673aa07a0750caf1a%2FSelection_052.png?generation=1580138133875470&amp;alt=media)\n\n\nsimilar network can be made for vowel removal",
    "730528": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2F315cc2b43553ac7b8e39d78706b1dd14%2FSelection_053.png?generation=1580138554793208&amp;alt=media)",
    "730529": "You can split the unicode characters into \"parts\"\n\n```\nin [1]: list('ক্ট্রো')\nout [1]: ['ক', '্', 'ট', '্', 'র', 'ো'])\n```\n\nIn this example the labels are:\n- root: 15 (`'ক্ট'`) -&gt; `list('ক্ট') = ['ক', '্', 'ট']`\n- vovel_d: 9 (`'ো'`) -&gt; `list('ো') = ['ো']`\n- consonant_d: 5 (`'্র'`) -&gt; `list('্র') = ['্', 'র']`\n\nI used these as a target (multi-label; part of the overall loss). What do you think? Is this could work or I just wasting my time? (my baseline improved, but I am still far from .98LB)",
    "730548": "Hi, Peter, May I ask how to split it to part? Thanks",
    "730555": "simple python `list` function, like `list('grapheme_char')`\n(I am using python 3.7; not sure this works with older versions..)",
    "730741": "structural code to encode location of vowel (consonant)\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/code.png)\n![](https://storage.googleapis.com/groundai-web-prod/media/users/user_13647/project_16001/images/structure.png)![](https://www.researchgate.net/profile/Cheng-Lin_Liu/publication/220412438/figure/download/fig4/AS:277377025363970@1443143244172/10-types-of-Chinese-character-structures-single-radical-left-right-up-down-up-left.png)\n![](https://ai2-s2-public.s3.amazonaws.com/figures/2017-08-08/07f6301809e45582d8330ce9d55f52dda4e42611/2-Figure2-1.png)",
    "730795": "A really slow but good way to do this is to train a unet type model that outputs 3 channels - root, consonant, vowel. Pre-train a classifier on the whole dataset then use that model to tell the unet if it split it well. You can then sum up the different part of the channel - root + cons, root + vowel, root alone and use the pre-trained model for the loss.I also added an intersection loss for the 3 channels and a loss so that the sum of the 3 channels made the same image as the input. I then convolution over the ouput to make the final decision of where each pixel goes. Samples with root 13:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F404040%2F69aab1e1429167b7637d35322cd78ed7%2Fsample_split_output.PNG?generation=1580168716653253&amp;alt=media)",
    "730883": "Thanks",
    "731612": "Here is [one more good thread](https://www.kaggle.com/c/bengaliai-cv19/discussion/128059) on generating more data by DrHB \n\nLooks like making the model to generalize well on unseen combinations is the key. Apparently, the data set is [constructed in such a way](https://www.kaggle.com/sibmike/pivots-hint-why-cutmix-works) to make us work on this problem. Can we avoid data generation? Maybe if we hide full combination sets for validation we can make the model generalize well on unseen data. Or a combination of these might work.",
    "731757": "it is weird, the mixups/cutmix can't improve the cv from my experiments😨",
    "737104": "This seems to be a very good idea @zero00. Have you been able to try it?",
    "740791": "pestipeti This is an amazing idea! I've been working on something similar myself. In fact, I think it will help to create more labels (finer labels) which will, in turn, increase the performance. \nI recently found a related paper on a similar phenomenon of \"fine labels\": \n[Understanding the Impact of Label Granularity on CNN-based Image Classification](https://arxiv.org/abs/1901.07012)"
  },
  "source": "meta"
}