{
  "id": 130503,
  "title": "Thinking out of the box",
  "url": "/competitions/bengaliai-cv19/discussion/130503",
  "author_name": "",
  "post_date": "2020-02-14T14:40:39.030099100Z",
  "votes": 40,
  "comment_count": 16,
  "views": 0,
  "content": "<p>So far, most discussions I see are focused on model selection or \"tricks\" (like replacing Adam with RAdam).</p>\n\n<p>This could work but typically gives mediocre results compared to more problem-specific methods. I propose you to start thinking differently.</p>\n\n<p>For instance, you could ask yourself the following questions:</p>\n\n<ol>\n<li>The task seems to be quite easy for practically any model except for a tiny portion of images. What is so specific for these images? Could you figure out a correct class by eye? Should you be really worried about bad performance on those?</li>\n<li>Image class is a combination of 3 subclasses. A possible number of combinations is quite large, while there is only 1300 classes in the training set. Is there is a method to make a model not to output unrealistic triplets?</li>\n<li>Graphemes have a more or less defined structure. A classifier is forced to learn this structure implicitly by finding the correlation between grapheme components and classes. Is there a method to supply this structure information to the classifier?</li>\n</ol>",
  "messages": [
    {
      "id": "746046",
      "postDate": "02/14/2020 14:40:39",
      "content": "<p>So far, most discussions I see are focused on model selection or \"tricks\" (like replacing Adam with RAdam).</p>\n\n<p>This could work but typically gives mediocre results compared to more problem-specific methods. I propose you to start thinking differently.</p>\n\n<p>For instance, you could ask yourself the following questions:</p>\n\n<ol>\n<li>The task seems to be quite easy for practically any model except for a tiny portion of images. What is so specific for these images? Could you figure out a correct class by eye? Should you be really worried about bad performance on those?</li>\n<li>Image class is a combination of 3 subclasses. A possible number of combinations is quite large, while there is only 1300 classes in the training set. Is there is a method to make a model not to output unrealistic triplets?</li>\n<li>Graphemes have a more or less defined structure. A classifier is forced to learn this structure implicitly by finding the correlation between grapheme components and classes. Is there a method to supply this structure information to the classifier?</li>\n</ol>",
      "rawMarkdown": "So far, most discussions I see are focused on model selection or \"tricks\" (like replacing Adam with RAdam).\n\nThis could work but typically gives mediocre results compared to more problem-specific methods. I propose you to start thinking differently.\n\nFor instance, you could ask yourself the following questions:\n\n1. The task seems to be quite easy for practically any model except for a tiny portion of images. What is so specific for these images? Could you figure out a correct class by eye? Should you be really worried about bad performance on those?\n2. Image class is a combination of 3 subclasses. A possible number of combinations is quite large, while there is only 1300 classes in the training set. Is there is a method to make a model not to output unrealistic triplets?\n3. Graphemes have a more or less defined structure. A classifier is forced to learn this structure implicitly by finding the correlation between grapheme components and classes. Is there a method to supply this structure information to the classifier?",
      "votes": null
    },
    {
      "id": "746079",
      "postDate": "02/14/2020 15:42:47",
      "content": "<p>I agree. Replacing Adam with RAdam won't give us a 0.02 boost on the LB.</p>\n\n<p>I analyzed my model's predictions (<code>grapheme_root</code>), and I found that ~30-40% of the error can be traced back to two graphemes: <code>ণ</code> (class: 59) and <code>ন</code> (class: 81) and their conjunctions with something.</p>\n\n<p>Here is the list (misclassifications between <code>graphemer_root</code> pairs):\n- 59 (<code>ণ</code>) - 81 (<code>ন</code>)\n- 60 (<code>ণ্ট</code>; conjunction parts: [<code>ণ</code>, <code>ট</code>]) - 83 (<code>ন্ট</code>; parts: [<code>ন</code>, <code>ট</code>])\n- 61 (<code>ণ্ঠ</code>; parts: [<code>ণ</code>, <code>ঠ</code>]) - 84 (<code>ন্ঠ</code>; parts: [<code>ন</code>, <code>ঠ</code>])\n- 62 (<code>ণ্ড</code>; parts: [<code>ণ</code>, <code>ড</code>]) - 85 (<code>ন্ড</code>; parts: [<code>ন</code>, <code>ড</code>])</p>\n\n<h2>True label: 59 (<code>ণ</code>) | Predicted (mostly) 81 (<code>ন</code>)</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F864684%2F8e487a8cdaf9ec4f0ef7acb831e0a07c%2F59_81_missclassification.png?generation=1581694744622331&amp;alt=media\" alt=\"\"></p>\n\n<p>I think some of these misclassifications can be fixed by postprocessing.</p>",
      "rawMarkdown": "I agree. Replacing Adam with RAdam won't give us a 0.02 boost on the LB.\n\nI analyzed my model's predictions (`grapheme_root`), and I found that ~30-40% of the error can be traced back to two graphemes: `ণ` (class: 59) and `ন` (class: 81) and their conjunctions with something.\n\nHere is the list (misclassifications between `graphemer_root` pairs):\n- 59 (`ণ`) - 81 (`ন`)\n- 60 (`ণ্ট`; conjunction parts: [`ণ`, `ট`]) - 83 (`ন্ট`; parts: [`ন`, `ট`])\n- 61 (`ণ্ঠ`; parts: [`ণ`, `ঠ`]) - 84 (`ন্ঠ`; parts: [`ন`, `ঠ`])\n- 62 (`ণ্ড`; parts: [`ণ`, `ড`]) - 85 (`ন্ড`; parts: [`ন`, `ড`])\n\n## True label: 59 (`ণ`) | Predicted (mostly) 81 (`ন`)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F864684%2F8e487a8cdaf9ec4f0ef7acb831e0a07c%2F59_81_missclassification.png?generation=1581694744622331&amp;alt=media)\n\n\nI think some of these misclassifications can be fixed by postprocessing.",
      "votes": null
    },
    {
      "id": "746102",
      "postDate": "02/14/2020 16:11:38",
      "content": "<p>Are you sure those are not labeling errors? </p>",
      "rawMarkdown": "Are you sure those are not labeling errors?",
      "votes": null
    },
    {
      "id": "746106",
      "postDate": "02/14/2020 16:27:40",
      "content": "<p>No. Honestly, I've never seen Bengali characters before. I can't tell the difference.</p>",
      "rawMarkdown": "No. Honestly, I've never seen Bengali characters before. I can't tell the difference.",
      "votes": null
    },
    {
      "id": "746107",
      "postDate": "02/14/2020 16:27:43",
      "content": "<p>At least some of them (like the bottom right corner one) look like they are.</p>",
      "rawMarkdown": "At least some of them (like the bottom right corner one) look like they are.",
      "votes": null
    },
    {
      "id": "746108",
      "postDate": "02/14/2020 16:28:46",
      "content": "<p>Yeah, me neither. But you still can figure out how the individual grapheme components look like by examining multiple images.</p>",
      "rawMarkdown": "Yeah, me neither. But you still can figure out how the individual grapheme components look like by examining multiple images.",
      "votes": null
    },
    {
      "id": "746124",
      "postDate": "02/14/2020 16:45:23",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a>  what do you think?</p>",
      "rawMarkdown": "phoenix9032  what do you think?",
      "votes": null
    },
    {
      "id": "746133",
      "postDate": "02/14/2020 16:59:53",
      "content": "<p>My analysis also same . I had really good model run through the training data I.e. mostly seen data and I have found that model is struggling to find difference between retroflex na (59) and dental na (81)  and other graphemes which is made up of these two . \nThey have very small difference between them visually (and today's time phonetically ) . When they go together with other consonants like 62 and 85 the difference  become even smaller . Most of the misinterpretation is between 62 and 85 . Unfortunately there is no thought currently in my mind to differenciate them based on rule . (Except for the fact that the 59 does not have matra or horizontal line on top of it and 81 has that . This is true for the Graphemes made based on these consonants as well . \nI was thinking probably another model or loss to separate these hard examples more clearly as a refinement would be good  .</p>",
      "rawMarkdown": "My analysis also same . I had really good model run through the training data I.e. mostly seen data and I have found that model is struggling to find difference between retroflex na (59) and dental na (81)  and other graphemes which is made up of these two . \nThey have very small difference between them visually (and today's time phonetically ) . When they go together with other consonants like 62 and 85 the difference  become even smaller . Most of the misinterpretation is between 62 and 85 . Unfortunately there is no thought currently in my mind to differenciate them based on rule . (Except for the fact that the 59 does not have matra or horizontal line on top of it and 81 has that . This is true for the Graphemes made based on these consonants as well . \nI was thinking probably another model or loss to separate these hard examples more clearly as a refinement would be good  .",
      "votes": null
    },
    {
      "id": "746136",
      "postDate": "02/14/2020 17:03:20",
      "content": "<p>mb that`s why cutmix or mixups works fine in this competition</p>",
      "rawMarkdown": "mb that`s why cutmix or mixups works fine in this competition",
      "votes": null
    },
    {
      "id": "746146",
      "postDate": "02/14/2020 17:18:21",
      "content": "<p>I am not active in the competition. But as a native Bengali speaker, the prediction seems to be perfect in the image.</p>",
      "rawMarkdown": "I am not active in the competition. But as a native Bengali speaker, the prediction seems to be perfect in the image.",
      "votes": null
    },
    {
      "id": "746154",
      "postDate": "02/14/2020 17:22:32",
      "content": "<p>Ooops ..I misread the True and prediction  .. I thought g = ground truth .. The prediction is completely perfect . The  Truth label is wrong ..</p>",
      "rawMarkdown": "Ooops ..I misread the True and prediction  .. I thought g = ground truth .. The prediction is completely perfect . The  Truth label is wrong ..",
      "votes": null
    },
    {
      "id": "746159",
      "postDate": "02/14/2020 17:26:04",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> thanks for your feedback. Yes, 'g' means the predicted grapheme_root</p>",
      "rawMarkdown": "phoenix9032 thanks for your feedback. Yes, 'g' means the predicted grapheme_root",
      "votes": null
    },
    {
      "id": "746409",
      "postDate": "02/15/2020 01:20:06",
      "content": "<p>Yeah, I second Bukun's observation.</p>",
      "rawMarkdown": "Yeah, I second Bukun's observation.",
      "votes": null
    },
    {
      "id": "746584",
      "postDate": "02/15/2020 08:37:51",
      "content": "<p>So, do we need to reproduce the mislabeling to win this competition ?</p>",
      "rawMarkdown": "So, do we need to reproduce the mislabeling to win this competition ?",
      "votes": null
    },
    {
      "id": "746585",
      "postDate": "02/15/2020 08:41:33",
      "content": "<p>This was one of my concern and I have asked the organizers to check for the test-set . Because I have seen quite a few cases model is predicting correctly but the gt is wrong in train_set . if its the  same in validation set , it might  swing the top slots ..</p>",
      "rawMarkdown": "This was one of my concern and I have asked the organizers to check for the test-set . Because I have seen quite a few cases model is predicting correctly but the gt is wrong in train_set . if its the  same in validation set , it might  swing the top slots ..",
      "votes": null
    },
    {
      "id": "746608",
      "postDate": "02/15/2020 09:10:41",
      "content": "<p>Mislabeled cases have been parts of many, many kaggle competitions before. There is not much that can be done about it. Usually, the scores are not that perfect, so it plays a much lower role.</p>\n\n<p>Here it is of course more critical, and luck can be a deciding factor. A funny comparison is actually \"Instant Gratification\" where data was generated and a small percentage was on purpose mislabeled. Solutions were perfect except for those mislabeled ones. Top spots were then purely by chance, and some guys even submitted two random versions.</p>",
      "rawMarkdown": "Mislabeled cases have been parts of many, many kaggle competitions before. There is not much that can be done about it. Usually, the scores are not that perfect, so it plays a much lower role.\n\nHere it is of course more critical, and luck can be a deciding factor. A funny comparison is actually \"Instant Gratification\" where data was generated and a small percentage was on purpose mislabeled. Solutions were perfect except for those mislabeled ones. Top spots were then purely by chance, and some guys even submitted two random versions.",
      "votes": null
    },
    {
      "id": "746634",
      "postDate": "02/15/2020 09:43:46",
      "content": "<p>Thank you guys for the information.\nI think it somehow explains the observation that decompose 3 components from predicted grapheme always get higher local score, but lower LB score than predict 3 components separately.</p>",
      "rawMarkdown": "Thank you guys for the information.\nI think it somehow explains the observation that decompose 3 components from predicted grapheme always get higher local score, but lower LB score than predict 3 components separately.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 746079,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "02/14/2020 15:42:47",
      "content": "<p>I agree. Replacing Adam with RAdam won't give us a 0.02 boost on the LB.</p>\n\n<p>I analyzed my model's predictions (<code>grapheme_root</code>), and I found that ~30-40% of the error can be traced back to two graphemes: <code>ণ</code> (class: 59) and <code>ন</code> (class: 81) and their conjunctions with something.</p>\n\n<p>Here is the list (misclassifications between <code>graphemer_root</code> pairs):\n- 59 (<code>ণ</code>) - 81 (<code>ন</code>)\n- 60 (<code>ণ্ট</code>; conjunction parts: [<code>ণ</code>, <code>ট</code>]) - 83 (<code>ন্ট</code>; parts: [<code>ন</code>, <code>ট</code>])\n- 61 (<code>ণ্ঠ</code>; parts: [<code>ণ</code>, <code>ঠ</code>]) - 84 (<code>ন্ঠ</code>; parts: [<code>ন</code>, <code>ঠ</code>])\n- 62 (<code>ণ্ড</code>; parts: [<code>ণ</code>, <code>ড</code>]) - 85 (<code>ন্ড</code>; parts: [<code>ন</code>, <code>ড</code>])</p>\n\n<h2>True label: 59 (<code>ণ</code>) | Predicted (mostly) 81 (<code>ন</code>)</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F864684%2F8e487a8cdaf9ec4f0ef7acb831e0a07c%2F59_81_missclassification.png?generation=1581694744622331&amp;alt=media\" alt=\"\"></p>\n\n<p>I think some of these misclassifications can be fixed by postprocessing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 746102,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "02/14/2020 16:11:38",
          "content": "<p>Are you sure those are not labeling errors? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746106,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/14/2020 16:27:40",
          "content": "<p>No. Honestly, I've never seen Bengali characters before. I can't tell the difference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746107,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "02/14/2020 16:27:43",
          "content": "<p>At least some of them (like the bottom right corner one) look like they are.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746108,
          "author_name": "ddanevskyi",
          "author_url": "",
          "post_date": "02/14/2020 16:28:46",
          "content": "<p>Yeah, me neither. But you still can figure out how the individual grapheme components look like by examining multiple images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746124,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/14/2020 16:45:23",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a>  what do you think?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746133,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "02/14/2020 16:59:53",
          "content": "<p>My analysis also same . I had really good model run through the training data I.e. mostly seen data and I have found that model is struggling to find difference between retroflex na (59) and dental na (81)  and other graphemes which is made up of these two . \nThey have very small difference between them visually (and today's time phonetically ) . When they go together with other consonants like 62 and 85 the difference  become even smaller . Most of the misinterpretation is between 62 and 85 . Unfortunately there is no thought currently in my mind to differenciate them based on rule . (Except for the fact that the 59 does not have matra or horizontal line on top of it and 81 has that . This is true for the Graphemes made based on these consonants as well . \nI was thinking probably another model or loss to separate these hard examples more clearly as a refinement would be good  .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746136,
          "author_name": "kupchanski",
          "author_url": "",
          "post_date": "02/14/2020 17:03:20",
          "content": "<p>mb that`s why cutmix or mixups works fine in this competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746146,
          "author_name": "ambarish",
          "author_url": "",
          "post_date": "02/14/2020 17:18:21",
          "content": "<p>I am not active in the competition. But as a native Bengali speaker, the prediction seems to be perfect in the image.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746154,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "02/14/2020 17:22:32",
          "content": "<p>Ooops ..I misread the True and prediction  .. I thought g = ground truth .. The prediction is completely perfect . The  Truth label is wrong ..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746159,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/14/2020 17:26:04",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> thanks for your feedback. Yes, 'g' means the predicted grapheme_root</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746584,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/15/2020 08:37:51",
          "content": "<p>So, do we need to reproduce the mislabeling to win this competition ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746585,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "02/15/2020 08:41:33",
          "content": "<p>This was one of my concern and I have asked the organizers to check for the test-set . Because I have seen quite a few cases model is predicting correctly but the gt is wrong in train_set . if its the  same in validation set , it might  swing the top slots ..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746608,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/15/2020 09:10:41",
          "content": "<p>Mislabeled cases have been parts of many, many kaggle competitions before. There is not much that can be done about it. Usually, the scores are not that perfect, so it plays a much lower role.</p>\n\n<p>Here it is of course more critical, and luck can be a deciding factor. A funny comparison is actually \"Instant Gratification\" where data was generated and a small percentage was on purpose mislabeled. Solutions were perfect except for those mislabeled ones. Top spots were then purely by chance, and some guys even submitted two random versions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746634,
          "author_name": "haqishen",
          "author_url": "",
          "post_date": "02/15/2020 09:43:46",
          "content": "<p>Thank you guys for the information.\nI think it somehow explains the observation that decompose 3 components from predicted grapheme always get higher local score, but lower LB score than predict 3 components separately.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 746409,
      "author_name": "amaity0",
      "author_url": "",
      "post_date": "02/15/2020 01:20:06",
      "content": "<p>Yeah, I second Bukun's observation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "746046": "So far, most discussions I see are focused on model selection or \"tricks\" (like replacing Adam with RAdam).\n\nThis could work but typically gives mediocre results compared to more problem-specific methods. I propose you to start thinking differently.\n\nFor instance, you could ask yourself the following questions:\n\n1. The task seems to be quite easy for practically any model except for a tiny portion of images. What is so specific for these images? Could you figure out a correct class by eye? Should you be really worried about bad performance on those?\n2. Image class is a combination of 3 subclasses. A possible number of combinations is quite large, while there is only 1300 classes in the training set. Is there is a method to make a model not to output unrealistic triplets?\n3. Graphemes have a more or less defined structure. A classifier is forced to learn this structure implicitly by finding the correlation between grapheme components and classes. Is there a method to supply this structure information to the classifier?",
    "746079": "I agree. Replacing Adam with RAdam won't give us a 0.02 boost on the LB.\n\nI analyzed my model's predictions (`grapheme_root`), and I found that ~30-40% of the error can be traced back to two graphemes: `ণ` (class: 59) and `ন` (class: 81) and their conjunctions with something.\n\nHere is the list (misclassifications between `graphemer_root` pairs):\n- 59 (`ণ`) - 81 (`ন`)\n- 60 (`ণ্ট`; conjunction parts: [`ণ`, `ট`]) - 83 (`ন্ট`; parts: [`ন`, `ট`])\n- 61 (`ণ্ঠ`; parts: [`ণ`, `ঠ`]) - 84 (`ন্ঠ`; parts: [`ন`, `ঠ`])\n- 62 (`ণ্ড`; parts: [`ণ`, `ড`]) - 85 (`ন্ড`; parts: [`ন`, `ড`])\n\n## True label: 59 (`ণ`) | Predicted (mostly) 81 (`ন`)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F864684%2F8e487a8cdaf9ec4f0ef7acb831e0a07c%2F59_81_missclassification.png?generation=1581694744622331&amp;alt=media)\n\n\nI think some of these misclassifications can be fixed by postprocessing.",
    "746102": "Are you sure those are not labeling errors?",
    "746106": "No. Honestly, I've never seen Bengali characters before. I can't tell the difference.",
    "746107": "At least some of them (like the bottom right corner one) look like they are.",
    "746108": "Yeah, me neither. But you still can figure out how the individual grapheme components look like by examining multiple images.",
    "746124": "phoenix9032  what do you think?",
    "746133": "My analysis also same . I had really good model run through the training data I.e. mostly seen data and I have found that model is struggling to find difference between retroflex na (59) and dental na (81)  and other graphemes which is made up of these two . \nThey have very small difference between them visually (and today's time phonetically ) . When they go together with other consonants like 62 and 85 the difference  become even smaller . Most of the misinterpretation is between 62 and 85 . Unfortunately there is no thought currently in my mind to differenciate them based on rule . (Except for the fact that the 59 does not have matra or horizontal line on top of it and 81 has that . This is true for the Graphemes made based on these consonants as well . \nI was thinking probably another model or loss to separate these hard examples more clearly as a refinement would be good  .",
    "746136": "mb that`s why cutmix or mixups works fine in this competition",
    "746146": "I am not active in the competition. But as a native Bengali speaker, the prediction seems to be perfect in the image.",
    "746154": "Ooops ..I misread the True and prediction  .. I thought g = ground truth .. The prediction is completely perfect . The  Truth label is wrong ..",
    "746159": "phoenix9032 thanks for your feedback. Yes, 'g' means the predicted grapheme_root",
    "746409": "Yeah, I second Bukun's observation.",
    "746584": "So, do we need to reproduce the mislabeling to win this competition ?",
    "746585": "This was one of my concern and I have asked the organizers to check for the test-set . Because I have seen quite a few cases model is predicting correctly but the gt is wrong in train_set . if its the  same in validation set , it might  swing the top slots ..",
    "746608": "Mislabeled cases have been parts of many, many kaggle competitions before. There is not much that can be done about it. Usually, the scores are not that perfect, so it plays a much lower role.\n\nHere it is of course more critical, and luck can be a deciding factor. A funny comparison is actually \"Instant Gratification\" where data was generated and a small percentage was on purpose mislabeled. Solutions were perfect except for those mislabeled ones. Top spots were then purely by chance, and some guys even submitted two random versions.",
    "746634": "Thank you guys for the information.\nI think it somehow explains the observation that decompose 3 components from predicted grapheme always get higher local score, but lower LB score than predict 3 components separately."
  },
  "source": "meta"
}