{
  "id": 134035,
  "title": "Are your models robust enough for unseen graphemes?",
  "url": "/competitions/bengaliai-cv19/discussion/134035",
  "author_name": "",
  "post_date": "2020-03-05T14:25:32.349354900Z",
  "votes": 24,
  "comment_count": 16,
  "views": 0,
  "content": "<p>It's not a secret that there are some unseen graphemes in the test set, making the gap between our CV score and LB score.</p>\n\n<p>According to the LB score, I guess the number of unseen graphemes in public LB test set is not that much, maybe not more than 2%.</p>\n\n<p>But what if there are around 5%~10% unseen graphemes in private LB test set? If this comes true, I can imagine there will be a BIG SHAKE. I think improve the robustness of our models for unseen graphemes is a key for avoiding the shake.</p>\n\n<p>So back to the topic. Are your models robust enough for unseen graphemes?</p>",
  "messages": [
    {
      "id": "764493",
      "postDate": "03/05/2020 14:25:32",
      "content": "<p>It's not a secret that there are some unseen graphemes in the test set, making the gap between our CV score and LB score.</p>\n\n<p>According to the LB score, I guess the number of unseen graphemes in public LB test set is not that much, maybe not more than 2%.</p>\n\n<p>But what if there are around 5%~10% unseen graphemes in private LB test set? If this comes true, I can imagine there will be a BIG SHAKE. I think improve the robustness of our models for unseen graphemes is a key for avoiding the shake.</p>\n\n<p>So back to the topic. Are your models robust enough for unseen graphemes?</p>",
      "rawMarkdown": "It's not a secret that there are some unseen graphemes in the test set, making the gap between our CV score and LB score.\n\nAccording to the LB score, I guess the number of unseen graphemes in public LB test set is not that much, maybe not more than 2%.\n\nBut what if there are around 5%~10% unseen graphemes in private LB test set? If this comes true, I can imagine there will be a BIG SHAKE. I think improve the robustness of our models for unseen graphemes is a key for avoiding the shake.\n\nSo back to the topic. Are your models robust enough for unseen graphemes?",
      "votes": null
    },
    {
      "id": "764506",
      "postDate": "03/05/2020 14:38:44",
      "content": "<blockquote>\n  <p>Are your models robust enough for unseen graphemes?</p>\n</blockquote>\n\n<p>Is there a metric to measure how robust our models are to unseen graphemes?</p>",
      "rawMarkdown": "&gt; Are your models robust enough for unseen graphemes?\n\nIs there a metric to measure how robust our models are to unseen graphemes?",
      "votes": null
    },
    {
      "id": "764511",
      "postDate": "03/05/2020 14:46:25",
      "content": "<p>It's easy, for example, hold out 5% of the graphemes in training set (while keep all kind of components remains) as validation set, then train a model and doing validating on it.</p>",
      "rawMarkdown": "It's easy, for example, hold out 5% of the graphemes in training set (while keep all kind of components remains) as validation set, then train a model and doing validating on it.",
      "votes": null
    },
    {
      "id": "764513",
      "postDate": "03/05/2020 14:46:43",
      "content": "<p>It can be evaluated somehow by splitting dataset into train and val dataset so that some graphemes are not included in train dataset, while keeping all grapheme root, vowel diacritics, and consonant diacritics included in both train and val dataset. Did someone try this?</p>",
      "rawMarkdown": "It can be evaluated somehow by splitting dataset into train and val dataset so that some graphemes are not included in train dataset, while keeping all grapheme root, vowel diacritics, and consonant diacritics included in both train and val dataset. Did someone try this?",
      "votes": null
    },
    {
      "id": "764516",
      "postDate": "03/05/2020 14:48:24",
      "content": "<p>Indeed, I've done what you said.\nThe score is quite low.... Like, 0.6~0.7 or something, and unstable.</p>",
      "rawMarkdown": "Indeed, I've done what you said.\nThe score is quite low.... Like, 0.6~0.7 or something, and unstable.",
      "votes": null
    },
    {
      "id": "764548",
      "postDate": "03/05/2020 15:27:33",
      "content": "<p>I too expect a huge shakeup..today our best model achieved validation recall 0.9824 but lb 0.9757\nI see large shakeup is coming.</p>",
      "rawMarkdown": "I too expect a huge shakeup..today our best model achieved validation recall 0.9824 but lb 0.9757\nI see large shakeup is coming.",
      "votes": null
    },
    {
      "id": "764553",
      "postDate": "03/05/2020 15:37:53",
      "content": "<p>Wow, I thought the public test &amp; private test should be homogenous. Never thought of it! I took a wild guess that your magic is post-processing.....</p>",
      "rawMarkdown": "Wow, I thought the public test &amp; private test should be homogenous. Never thought of it! I took a wild guess that your magic is post-processing.....",
      "votes": null
    },
    {
      "id": "764585",
      "postDate": "03/05/2020 16:12:28",
      "content": "<p>It's not impossible... So we have to be prepared for it.</p>",
      "rawMarkdown": "It's not impossible... So we have to be prepared for it.",
      "votes": null
    },
    {
      "id": "764594",
      "postDate": "03/05/2020 16:22:09",
      "content": "<p>Good job but terrible results...</p>",
      "rawMarkdown": "Good job but terrible results...",
      "votes": null
    },
    {
      "id": "764613",
      "postDate": "03/05/2020 16:42:29",
      "content": "<p>Which means, our score is mostly come from seen graphemes.</p>",
      "rawMarkdown": "Which means, our score is mostly come from seen graphemes.",
      "votes": null
    },
    {
      "id": "764625",
      "postDate": "03/05/2020 16:58:05",
      "content": "<p>since the model is trained on cutmix and cutout, i feel that it should be robust against \"such noises like unseen graphemes\"</p>\n\n<p>test graphemes = some part of train graphemes + some part of noise</p>",
      "rawMarkdown": "since the model is trained on cutmix and cutout, i feel that it should be robust against \"such noises like unseen graphemes\"\n\ntest graphemes = some part of train graphemes + some part of noise",
      "votes": null
    },
    {
      "id": "764685",
      "postDate": "03/05/2020 18:26:36",
      "content": "<p>I feel cutout may have opposite effect, as the model may learn to the \"fill the gap\". In the training set there are lots of grapheme roots that are associated with a certain few vowels and consonants.</p>",
      "rawMarkdown": "I feel cutout may have opposite effect, as the model may learn to the \"fill the gap\". In the training set there are lots of grapheme roots that are associated with a certain few vowels and consonants.",
      "votes": null
    },
    {
      "id": "764686",
      "postDate": "03/05/2020 18:27:51",
      "content": "<p>but cutmix shoud be the opposit.</p>",
      "rawMarkdown": "but cutmix shoud be the opposit.",
      "votes": null
    },
    {
      "id": "764895",
      "postDate": "03/06/2020 03:06:57",
      "content": "<p>I'm not sure what augmentations are doing good impact on unseen graphemes and what are not. But the fact is that, if I only doing validation on unseen graphemes, the score is very low 😨 </p>",
      "rawMarkdown": "I'm not sure what augmentations are doing good impact on unseen graphemes and what are not. But the fact is that, if I only doing validation on unseen graphemes, the score is very low 😨",
      "votes": null
    },
    {
      "id": "765064",
      "postDate": "03/06/2020 08:08:28",
      "content": "<p><a href=\"/haqishen\">@haqishen</a> it is not very low for me, it's motherf*cking low, like 99xx -&gt; 5xxx</p>",
      "rawMarkdown": "haqishen it is not very low for me, it's motherf*cking low, like 99xx -&gt; 5xxx",
      "votes": null
    },
    {
      "id": "765066",
      "postDate": "03/06/2020 08:09:44",
      "content": "<p>yea i tried and yes, it's a motherf*cking hell of a mess like <a href=\"/haqishen\">@haqishen</a> just said. 30-40% drop</p>",
      "rawMarkdown": "yea i tried and yes, it's a motherf*cking hell of a mess like @haqishen just said. 30-40% drop",
      "votes": null
    },
    {
      "id": "766509",
      "postDate": "03/08/2020 09:15:55",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> Rather using cutmix and cutout is probably making it worse, </p>\n\n<p>Suppose the situation when using cutout/cutmix a component might completely get removed, but the lambda value is only the percentage of pixels removed for cutmix and no change in labels for cutout. This means that the model learns the combinations as a whole by seeing the rest of the image instead of learning the separate components. </p>\n\n<p>To test this I used only mixup and the result for unseen graphemes is 10% better (though LB score is lower for only mixup). More experimentation is needed.</p>",
      "rawMarkdown": "hengck23 Rather using cutmix and cutout is probably making it worse, \n\nSuppose the situation when using cutout/cutmix a component might completely get removed, but the lambda value is only the percentage of pixels removed for cutmix and no change in labels for cutout. This means that the model learns the combinations as a whole by seeing the rest of the image instead of learning the separate components. \n\nTo test this I used only mixup and the result for unseen graphemes is 10% better (though LB score is lower for only mixup). More experimentation is needed.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 764506,
      "author_name": "bibek777",
      "author_url": "",
      "post_date": "03/05/2020 14:38:44",
      "content": "<blockquote>\n  <p>Are your models robust enough for unseen graphemes?</p>\n</blockquote>\n\n<p>Is there a metric to measure how robust our models are to unseen graphemes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 764511,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/05/2020 14:46:25",
          "content": "<p>It's easy, for example, hold out 5% of the graphemes in training set (while keep all kind of components remains) as validation set, then train a model and doing validating on it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 764513,
      "author_name": "ren4yu",
      "author_url": "",
      "post_date": "03/05/2020 14:46:43",
      "content": "<p>It can be evaluated somehow by splitting dataset into train and val dataset so that some graphemes are not included in train dataset, while keeping all grapheme root, vowel diacritics, and consonant diacritics included in both train and val dataset. Did someone try this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 764516,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/05/2020 14:48:24",
          "content": "<p>Indeed, I've done what you said.\nThe score is quite low.... Like, 0.6~0.7 or something, and unstable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764594,
          "author_name": "ren4yu",
          "author_url": "",
          "post_date": "03/05/2020 16:22:09",
          "content": "<p>Good job but terrible results...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764613,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/05/2020 16:42:29",
          "content": "<p>Which means, our score is mostly come from seen graphemes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765066,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/06/2020 08:09:44",
          "content": "<p>yea i tried and yes, it's a motherf*cking hell of a mess like <a href=\"/haqishen\">@haqishen</a> just said. 30-40% drop</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 764548,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "03/05/2020 15:27:33",
      "content": "<p>I too expect a huge shakeup..today our best model achieved validation recall 0.9824 but lb 0.9757\nI see large shakeup is coming.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 764553,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "03/05/2020 15:37:53",
      "content": "<p>Wow, I thought the public test &amp; private test should be homogenous. Never thought of it! I took a wild guess that your magic is post-processing.....</p>",
      "votes": null,
      "replies": [
        {
          "id": 764585,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/05/2020 16:12:28",
          "content": "<p>It's not impossible... So we have to be prepared for it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 764625,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/05/2020 16:58:05",
      "content": "<p>since the model is trained on cutmix and cutout, i feel that it should be robust against \"such noises like unseen graphemes\"</p>\n\n<p>test graphemes = some part of train graphemes + some part of noise</p>",
      "votes": null,
      "replies": [
        {
          "id": 764685,
          "author_name": "yl1202",
          "author_url": "",
          "post_date": "03/05/2020 18:26:36",
          "content": "<p>I feel cutout may have opposite effect, as the model may learn to the \"fill the gap\". In the training set there are lots of grapheme roots that are associated with a certain few vowels and consonants.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764686,
          "author_name": "yl1202",
          "author_url": "",
          "post_date": "03/05/2020 18:27:51",
          "content": "<p>but cutmix shoud be the opposit.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 764895,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "03/06/2020 03:06:57",
          "content": "<p>I'm not sure what augmentations are doing good impact on unseen graphemes and what are not. But the fact is that, if I only doing validation on unseen graphemes, the score is very low 😨 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765064,
          "author_name": "moewie94",
          "author_url": "",
          "post_date": "03/06/2020 08:08:28",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> it is not very low for me, it's motherf*cking low, like 99xx -&gt; 5xxx</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766509,
          "author_name": "dipamc77",
          "author_url": "",
          "post_date": "03/08/2020 09:15:55",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Rather using cutmix and cutout is probably making it worse, </p>\n\n<p>Suppose the situation when using cutout/cutmix a component might completely get removed, but the lambda value is only the percentage of pixels removed for cutmix and no change in labels for cutout. This means that the model learns the combinations as a whole by seeing the rest of the image instead of learning the separate components. </p>\n\n<p>To test this I used only mixup and the result for unseen graphemes is 10% better (though LB score is lower for only mixup). More experimentation is needed.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "764493": "It's not a secret that there are some unseen graphemes in the test set, making the gap between our CV score and LB score.\n\nAccording to the LB score, I guess the number of unseen graphemes in public LB test set is not that much, maybe not more than 2%.\n\nBut what if there are around 5%~10% unseen graphemes in private LB test set? If this comes true, I can imagine there will be a BIG SHAKE. I think improve the robustness of our models for unseen graphemes is a key for avoiding the shake.\n\nSo back to the topic. Are your models robust enough for unseen graphemes?",
    "764506": "&gt; Are your models robust enough for unseen graphemes?\n\nIs there a metric to measure how robust our models are to unseen graphemes?",
    "764511": "It's easy, for example, hold out 5% of the graphemes in training set (while keep all kind of components remains) as validation set, then train a model and doing validating on it.",
    "764513": "It can be evaluated somehow by splitting dataset into train and val dataset so that some graphemes are not included in train dataset, while keeping all grapheme root, vowel diacritics, and consonant diacritics included in both train and val dataset. Did someone try this?",
    "764516": "Indeed, I've done what you said.\nThe score is quite low.... Like, 0.6~0.7 or something, and unstable.",
    "764548": "I too expect a huge shakeup..today our best model achieved validation recall 0.9824 but lb 0.9757\nI see large shakeup is coming.",
    "764553": "Wow, I thought the public test &amp; private test should be homogenous. Never thought of it! I took a wild guess that your magic is post-processing.....",
    "764585": "It's not impossible... So we have to be prepared for it.",
    "764594": "Good job but terrible results...",
    "764613": "Which means, our score is mostly come from seen graphemes.",
    "764625": "since the model is trained on cutmix and cutout, i feel that it should be robust against \"such noises like unseen graphemes\"\n\ntest graphemes = some part of train graphemes + some part of noise",
    "764685": "I feel cutout may have opposite effect, as the model may learn to the \"fill the gap\". In the training set there are lots of grapheme roots that are associated with a certain few vowels and consonants.",
    "764686": "but cutmix shoud be the opposit.",
    "764895": "I'm not sure what augmentations are doing good impact on unseen graphemes and what are not. But the fact is that, if I only doing validation on unseen graphemes, the score is very low 😨",
    "765064": "haqishen it is not very low for me, it's motherf*cking low, like 99xx -&gt; 5xxx",
    "765066": "yea i tried and yes, it's a motherf*cking hell of a mess like @haqishen just said. 30-40% drop",
    "766509": "hengck23 Rather using cutmix and cutout is probably making it worse, \n\nSuppose the situation when using cutout/cutmix a component might completely get removed, but the lambda value is only the percentage of pixels removed for cutmix and no change in labels for cutout. This means that the model learns the combinations as a whole by seeing the rest of the image instead of learning the separate components. \n\nTo test this I used only mixup and the result for unseen graphemes is 10% better (though LB score is lower for only mixup). More experimentation is needed."
  },
  "source": "meta"
}