{
  "id": 131078,
  "title": "Train Test Distribution",
  "url": "/competitions/bengaliai-cv19/discussion/131078",
  "author_name": "",
  "post_date": "2020-02-18T05:04:34.817745800Z",
  "votes": 31,
  "comment_count": 18,
  "views": 0,
  "content": "<p>Thank you Bengali.AI for sharing your dataset and hosting this interesting competition. I think I finally figured out what's going on :-)</p>\n\n<h3>Description</h3>\n\n<p>The training set contains approximately 150 each of 1292 common words. At the bottom of this post are 12 examples each of 4 common words.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7096ae8ae94b5bf462562d443f6b88fe%2Fwords.png?generation=1582001513517840&amp;alt=media\" alt=\"\"></p>\n\n<p>Each word is composed of three pieces; a (1) Grapheme root (which confusingly are actually vowels, consonants, and consonant conjuncts themselves), a (2) vowel diacritic, and a (3) consonant diacritic. There are 168 different roots, 11 different v-diacritics, and 7 different c-diacritics total over train and test. Our job is to find and detect these pieces in words. The test set will contain words not present in the train's 1292 words (but uses the same pieces).</p>\n\n<h3>Questions</h3>\n\n<ol>\n<li>Will all the test data words be new? Or will some be the same as training words?</li>\n<li>Are the 1292 common words representative of the language. i.e. is the frequency distribution of grapheme roots in train data the same as test data? Are rare graphemes rare in both train and test? And what about distribution of vowel diacritics and consonant diacritics?</li>\n<li>How many ways can the same word be drawn? I notice images that look different below among the same word?</li>\n</ol>\n\n<h3>Word 1</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fff591a018c0d602ddc6c8040db2e6a9a%2FScreen%20Shot%202020-02-17%20at%208.47.33%20PM.png?generation=1582001726970413&amp;alt=media\" alt=\"\"></p>\n\n<h3>Word 2</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F22330c5d5c2c5acea1db961e84804bfa%2FScreen%20Shot%202020-02-17%20at%208.47.53%20PM.png?generation=1582001743502117&amp;alt=media\" alt=\"\"></p>\n\n<h3>Word 3</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb94d29b6470e26665a35499a4012704d%2FScreen%20Shot%202020-02-17%20at%208.48.17%20PM.png?generation=1582001756393939&amp;alt=media\" alt=\"\"></p>\n\n<h3>Word 4</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F78378bf870f2205a2acaad08df6b1b2a%2FScreen%20Shot%202020-02-17%20at%208.45.37%20PM.png?generation=1582001770729552&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "748885",
      "postDate": "02/18/2020 05:04:34",
      "content": "<p>Thank you Bengali.AI for sharing your dataset and hosting this interesting competition. I think I finally figured out what's going on :-)</p>\n\n<h3>Description</h3>\n\n<p>The training set contains approximately 150 each of 1292 common words. At the bottom of this post are 12 examples each of 4 common words.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7096ae8ae94b5bf462562d443f6b88fe%2Fwords.png?generation=1582001513517840&amp;alt=media\" alt=\"\"></p>\n\n<p>Each word is composed of three pieces; a (1) Grapheme root (which confusingly are actually vowels, consonants, and consonant conjuncts themselves), a (2) vowel diacritic, and a (3) consonant diacritic. There are 168 different roots, 11 different v-diacritics, and 7 different c-diacritics total over train and test. Our job is to find and detect these pieces in words. The test set will contain words not present in the train's 1292 words (but uses the same pieces).</p>\n\n<h3>Questions</h3>\n\n<ol>\n<li>Will all the test data words be new? Or will some be the same as training words?</li>\n<li>Are the 1292 common words representative of the language. i.e. is the frequency distribution of grapheme roots in train data the same as test data? Are rare graphemes rare in both train and test? And what about distribution of vowel diacritics and consonant diacritics?</li>\n<li>How many ways can the same word be drawn? I notice images that look different below among the same word?</li>\n</ol>\n\n<h3>Word 1</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fff591a018c0d602ddc6c8040db2e6a9a%2FScreen%20Shot%202020-02-17%20at%208.47.33%20PM.png?generation=1582001726970413&amp;alt=media\" alt=\"\"></p>\n\n<h3>Word 2</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F22330c5d5c2c5acea1db961e84804bfa%2FScreen%20Shot%202020-02-17%20at%208.47.53%20PM.png?generation=1582001743502117&amp;alt=media\" alt=\"\"></p>\n\n<h3>Word 3</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb94d29b6470e26665a35499a4012704d%2FScreen%20Shot%202020-02-17%20at%208.48.17%20PM.png?generation=1582001756393939&amp;alt=media\" alt=\"\"></p>\n\n<h3>Word 4</h3>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F78378bf870f2205a2acaad08df6b1b2a%2FScreen%20Shot%202020-02-17%20at%208.45.37%20PM.png?generation=1582001770729552&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thank you Bengali.AI for sharing your dataset and hosting this interesting competition. I think I finally figured out what's going on :-)\n\n### Description\n\nThe training set contains approximately 150 each of 1292 common words. At the bottom of this post are 12 examples each of 4 common words.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7096ae8ae94b5bf462562d443f6b88fe%2Fwords.png?generation=1582001513517840&amp;alt=media)\n\nEach word is composed of three pieces; a (1) Grapheme root (which confusingly are actually vowels, consonants, and consonant conjuncts themselves), a (2) vowel diacritic, and a (3) consonant diacritic. There are 168 different roots, 11 different v-diacritics, and 7 different c-diacritics total over train and test. Our job is to find and detect these pieces in words. The test set will contain words not present in the train's 1292 words (but uses the same pieces).\n\n### Questions\n1. Will all the test data words be new? Or will some be the same as training words?\n2. Are the 1292 common words representative of the language. i.e. is the frequency distribution of grapheme roots in train data the same as test data? Are rare graphemes rare in both train and test? And what about distribution of vowel diacritics and consonant diacritics?\n3. How many ways can the same word be drawn? I notice images that look different below among the same word?\n\n### Word 1\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fff591a018c0d602ddc6c8040db2e6a9a%2FScreen%20Shot%202020-02-17%20at%208.47.33%20PM.png?generation=1582001726970413&amp;alt=media)\n\n### Word 2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F22330c5d5c2c5acea1db961e84804bfa%2FScreen%20Shot%202020-02-17%20at%208.47.53%20PM.png?generation=1582001743502117&amp;alt=media)\n\n### Word 3\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb94d29b6470e26665a35499a4012704d%2FScreen%20Shot%202020-02-17%20at%208.48.17%20PM.png?generation=1582001756393939&amp;alt=media)\n\n### Word 4\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F78378bf870f2205a2acaad08df6b1b2a%2FScreen%20Shot%202020-02-17%20at%208.45.37%20PM.png?generation=1582001770729552&amp;alt=media)",
      "votes": null
    },
    {
      "id": "749083",
      "postDate": "02/18/2020 09:32:55",
      "content": "<p>Hence why I really think we need to take some time to push for methods like Class Activation Maps to \"see\" which parts of the image are useful for a model. It feels like for now a lot of the development of a model is very trial &amp; error like: we try something and we see how that influences our local CV. I have the feeling that we´re not putting much effort (at least publicly, here, on Kaggle, I don't see that much) on error analysis, activation maps, etc.</p>\n\n<p>I feel like this would go a long way in helping answer some of those questions: are the models we're making robust to new combinations of symbols, or are there finding information in the whole character at once?</p>",
      "rawMarkdown": "Hence why I really think we need to take some time to push for methods like Class Activation Maps to \"see\" which parts of the image are useful for a model. It feels like for now a lot of the development of a model is very trial &amp; error like: we try something and we see how that influences our local CV. I have the feeling that we´re not putting much effort (at least publicly, here, on Kaggle, I don't see that much) on error analysis, activation maps, etc.\n\nI feel like this would go a long way in helping answer some of those questions: are the models we're making robust to new combinations of symbols, or are there finding information in the whole character at once?",
      "votes": null
    },
    {
      "id": "749126",
      "postDate": "02/18/2020 10:28:54",
      "content": "<p>As far as I know, all Grapheme roots, vowel diacritics, consonant diacritics between train and test are identical, but test set has more new connections from Grapheme roots, vowel diacritics and consonant diacritics.</p>",
      "rawMarkdown": "As far as I know, all Grapheme roots, vowel diacritics, consonant diacritics between train and test are identical, but test set has more new connections from Grapheme roots, vowel diacritics and consonant diacritics.",
      "votes": null
    },
    {
      "id": "749480",
      "postDate": "02/18/2020 18:20:38",
      "content": "<p>Can any naive speakers comment on the distribution (relative frequency counts) of Grapheme roots and diacritics of train versus test? This would be helpful to all of us. Is the distribution among the 1292 common words the same distribution among the rest of the language (and the test data)?</p>",
      "rawMarkdown": "Can any naive speakers comment on the distribution (relative frequency counts) of Grapheme roots and diacritics of train versus test? This would be helpful to all of us. Is the distribution among the 1292 common words the same distribution among the rest of the language (and the test data)?",
      "votes": null
    },
    {
      "id": "749486",
      "postDate": "02/18/2020 18:22:47",
      "content": "<p>Thanks. Yes. But i wonder is the most popular Grapheme root among the 1292 common words also the most popular among less common words (the test set)? What about the least common Grapheme root? Participants who don't know Bengla have no idea.</p>",
      "rawMarkdown": "Thanks. Yes. But i wonder is the most popular Grapheme root among the 1292 common words also the most popular among less common words (the test set)? What about the least common Grapheme root? Participants who don't know Bengla have no idea.",
      "votes": null
    },
    {
      "id": "749488",
      "postDate": "02/18/2020 18:24:43",
      "content": "<p>Yes. Knowing where models are directing their attention could help us build better models. I like your notebook <a href=\"https://www.kaggle.com/maxlenormand/multi-class-activation-map-with-efficientnetb0\">here</a> showing this. Great job.</p>",
      "rawMarkdown": "Yes. Knowing where models are directing their attention could help us build better models. I like your notebook [here][1] showing this. Great job.\n\n[1]: https://www.kaggle.com/maxlenormand/multi-class-activation-map-with-efficientnetb0",
      "votes": null
    },
    {
      "id": "749502",
      "postDate": "02/18/2020 18:29:48",
      "content": "<p>In the test set, we have the same 168-11-7 classes, but there are unseen combinations. The host confirmed this somewhere.</p>\n\n<p>Take a look at this doc (published by the host):\n<a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf</a></p>",
      "rawMarkdown": "In the test set, we have the same 168-11-7 classes, but there are unseen combinations. The host confirmed this somewhere.\n\nTake a look at this doc (published by the host):\nhttps://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf",
      "votes": null
    },
    {
      "id": "749507",
      "postDate": "02/18/2020 18:34:19",
      "content": "<p>Yes thanks. But the Grapheme's are unbalanced in the train set. Are they unbalanced in the same way in the test set?</p>",
      "rawMarkdown": "Yes thanks. But the Grapheme's are unbalanced in the train set. Are they unbalanced in the same way in the test set?",
      "votes": null
    },
    {
      "id": "749514",
      "postDate": "02/18/2020 18:39:31",
      "content": "<p>Noone knows and this is also no information hosts / kaggle will share. </p>",
      "rawMarkdown": "Noone knows and this is also no information hosts / kaggle will share.",
      "votes": null
    },
    {
      "id": "749548",
      "postDate": "02/18/2020 19:15:09",
      "content": "<p>In many experiments of the participants (in mine, too), a strict correlation was observed between cross-validation score and the public test score - CV = PL + 0.007..0.009.  In my opinion, this indicates that the distribution in the train and public test is the same.  You can only guess about the distribution in a private test. I bet that distribution in the private test is the same and therefore the shake up will be minimal.</p>",
      "rawMarkdown": "In many experiments of the participants (in mine, too), a strict correlation was observed between cross-validation score and the public test score - CV = PL + 0.007..0.009.  In my opinion, this indicates that the distribution in the train and public test is the same.  You can only guess about the distribution in a private test. I bet that distribution in the private test is the same and therefore the shake up will be minimal.",
      "votes": null
    },
    {
      "id": "749600",
      "postDate": "02/18/2020 19:51:00",
      "content": "<p>Thanks for the info. My 20% validation holdout is 0.977 but my LB is only 0.967, so I think something is off with my validation set up. (Perhaps I need to isolate words or stratify Grapheme roots). Those scores are significantly different implying that there is something different about the relationship of test to train when compared to my validation set to train set.</p>",
      "rawMarkdown": "Thanks for the info. My 20% validation holdout is 0.977 but my LB is only 0.967, so I think something is off with my validation set up. (Perhaps I need to isolate words or stratify Grapheme roots). Those scores are significantly different implying that there is something different about the relationship of test to train when compared to my validation set to train set.",
      "votes": null
    },
    {
      "id": "749713",
      "postDate": "02/18/2020 21:42:36",
      "content": "<p>using a random 20% validation isnt good idea, remember you said there are 150 of each 3-unit-combo. you want all of that 150 in your validation to represent test</p>",
      "rawMarkdown": "using a random 20% validation isnt good idea, remember you said there are 150 of each 3-unit-combo. you want all of that 150 in your validation to represent test",
      "votes": null
    },
    {
      "id": "749722",
      "postDate": "02/18/2020 21:49:00",
      "content": "<p>Yeah, my \"hold out set\" is just my first fold of 5-Fold. I'm just too lazy to train all 5 folds. Currently it is plain KFold.</p>\n\n<p>There are 1292 words. I'm not sure where I want each word (the sets of 150). Perhaps GroupKFold where the 150 is either all in train or all in validation is best. Or perhaps StratifiedKFold where we stratify on the 168 Grapheme roots is best (and let words go where they go). Still trying different CVs. What's your CV?</p>",
      "rawMarkdown": "Yeah, my \"hold out set\" is just my first fold of 5-Fold. I'm just too lazy to train all 5 folds. Currently it is plain KFold.\n\nThere are 1292 words. I'm not sure where I want each word (the sets of 150). Perhaps GroupKFold where the 150 is either all in train or all in validation is best. Or perhaps StratifiedKFold where we stratify on the 168 Grapheme roots is best (and let words go where they go). Still trying different CVs. What's your CV?",
      "votes": null
    },
    {
      "id": "749757",
      "postDate": "02/18/2020 22:20:18",
      "content": "<p>We are using iterative stratification and retain the 0.007..0.009 gap between CV and public LB. I am not very bothered by this gap because it is small &amp; consistent. Personally i think groupkfold on each word would be the best; however, we are doing experiments on just 1 fold, and if you are using groupkfold on each word, then you run the risk of overfitting to that single fold, so we would honestly need to run experiments on all folds to get something reliable which would take waaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaay too long.</p>\n\n<p>Anyway, to get the best results you want to retrain on full train dataset without using folds, so I am really calm with the gap.</p>",
      "rawMarkdown": "We are using iterative stratification and retain the 0.007..0.009 gap between CV and public LB. I am not very bothered by this gap because it is small &amp; consistent. Personally i think groupkfold on each word would be the best; however, we are doing experiments on just 1 fold, and if you are using groupkfold on each word, then you run the risk of overfitting to that single fold, so we would honestly need to run experiments on all folds to get something reliable which would take waaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaay too long.\n\nAnyway, to get the best results you want to retrain on full train dataset without using folds, so I am really calm with the gap.",
      "votes": null
    },
    {
      "id": "749801",
      "postDate": "02/18/2020 22:56:11",
      "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I have a bigger CV/LB gap of 0.02. I am thinking that the different distribution might be the problem, too.</p>",
      "rawMarkdown": "cdeotte I have a bigger CV/LB gap of 0.02. I am thinking that the different distribution might be the problem, too.",
      "votes": null
    },
    {
      "id": "749829",
      "postDate": "02/18/2020 23:29:48",
      "content": "<p>Same thing for me, i can have 1% diff between my cv and lb if i dont balance de train data. One test i did was training imbalanced and then i only balanced consonant classes and it didn't really give me a noticeable boost so i guess consonants are not the problem because it's the most imbalanced part of the train data. I assume grapheme root is the problem??</p>",
      "rawMarkdown": "Same thing for me, i can have 1% diff between my cv and lb if i dont balance de train data. One test i did was training imbalanced and then i only balanced consonant classes and it didn't really give me a noticeable boost so i guess consonants are not the problem because it's the most imbalanced part of the train data. I assume grapheme root is the problem??",
      "votes": null
    },
    {
      "id": "750126",
      "postDate": "02/19/2020 06:41:18",
      "content": "<p>I see most error between the grapheme root 59 to 92 , in the validation set . </p>",
      "rawMarkdown": "I see most error between the grapheme root 59 to 92 , in the validation set .",
      "votes": null
    },
    {
      "id": "750299",
      "postDate": "02/19/2020 09:12:41",
      "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a>  I am not sure I completely understand . Actually even if the person knows Bangla (Like me) would not be able to understand the distribution . Because</p>\n\n<p>a) These are not words - Words are some meaningful combination of characters . So there is no such word in the set . These are like compund-characters , made up of three different characters.</p>\n\n<p>b) The characters crowdsourced and people are asked to write them as they want . So there is no way we can identify anything that are popular . Yes , there are some characters which are less used  in our day to day life , but that does not mean that test set might have less number of them . As an example word E and word Z in english , you know that E is used a lot more than Z , that does not mean if its crowdsourced to write E and Z , Z will have lesser population in the set . </p>\n\n<p>I am sorry , if I completely misunderstood your inference . Happy to help with little knowledge i have in this language . </p>",
      "rawMarkdown": "Hi @cdeotte  I am not sure I completely understand . Actually even if the person knows Bangla (Like me) would not be able to understand the distribution . Because\n\na) These are not words - Words are some meaningful combination of characters . So there is no such word in the set . These are like compund-characters , made up of three different characters.\n\nb) The characters crowdsourced and people are asked to write them as they want . So there is no way we can identify anything that are popular . Yes , there are some characters which are less used  in our day to day life , but that does not mean that test set might have less number of them . As an example word E and word Z in english , you know that E is used a lot more than Z , that does not mean if its crowdsourced to write E and Z , Z will have lesser population in the set . \n\nI am sorry , if I completely misunderstood your inference . Happy to help with little knowledge i have in this language .",
      "votes": null
    },
    {
      "id": "750384",
      "postDate": "02/19/2020 10:35:18",
      "content": "<p>Thanks, glad you like it!</p>",
      "rawMarkdown": "Thanks, glad you like it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 749083,
      "author_name": "maxlenormand",
      "author_url": "",
      "post_date": "02/18/2020 09:32:55",
      "content": "<p>Hence why I really think we need to take some time to push for methods like Class Activation Maps to \"see\" which parts of the image are useful for a model. It feels like for now a lot of the development of a model is very trial &amp; error like: we try something and we see how that influences our local CV. I have the feeling that we´re not putting much effort (at least publicly, here, on Kaggle, I don't see that much) on error analysis, activation maps, etc.</p>\n\n<p>I feel like this would go a long way in helping answer some of those questions: are the models we're making robust to new combinations of symbols, or are there finding information in the whole character at once?</p>",
      "votes": null,
      "replies": [
        {
          "id": 749488,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2020 18:24:43",
          "content": "<p>Yes. Knowing where models are directing their attention could help us build better models. I like your notebook <a href=\"https://www.kaggle.com/maxlenormand/multi-class-activation-map-with-efficientnetb0\">here</a> showing this. Great job.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750384,
          "author_name": "maxlenormand",
          "author_url": "",
          "post_date": "02/19/2020 10:35:18",
          "content": "<p>Thanks, glad you like it!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749126,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "02/18/2020 10:28:54",
      "content": "<p>As far as I know, all Grapheme roots, vowel diacritics, consonant diacritics between train and test are identical, but test set has more new connections from Grapheme roots, vowel diacritics and consonant diacritics.</p>",
      "votes": null,
      "replies": [
        {
          "id": 749486,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2020 18:22:47",
          "content": "<p>Thanks. Yes. But i wonder is the most popular Grapheme root among the 1292 common words also the most popular among less common words (the test set)? What about the least common Grapheme root? Participants who don't know Bengla have no idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750299,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "02/19/2020 09:12:41",
          "content": "<p>Hi <a href=\"/cdeotte\">@cdeotte</a>  I am not sure I completely understand . Actually even if the person knows Bangla (Like me) would not be able to understand the distribution . Because</p>\n\n<p>a) These are not words - Words are some meaningful combination of characters . So there is no such word in the set . These are like compund-characters , made up of three different characters.</p>\n\n<p>b) The characters crowdsourced and people are asked to write them as they want . So there is no way we can identify anything that are popular . Yes , there are some characters which are less used  in our day to day life , but that does not mean that test set might have less number of them . As an example word E and word Z in english , you know that E is used a lot more than Z , that does not mean if its crowdsourced to write E and Z , Z will have lesser population in the set . </p>\n\n<p>I am sorry , if I completely misunderstood your inference . Happy to help with little knowledge i have in this language . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 749480,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/18/2020 18:20:38",
      "content": "<p>Can any naive speakers comment on the distribution (relative frequency counts) of Grapheme roots and diacritics of train versus test? This would be helpful to all of us. Is the distribution among the 1292 common words the same distribution among the rest of the language (and the test data)?</p>",
      "votes": null,
      "replies": [
        {
          "id": 749502,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/18/2020 18:29:48",
          "content": "<p>In the test set, we have the same 168-11-7 classes, but there are unseen combinations. The host confirmed this somewhere.</p>\n\n<p>Take a look at this doc (published by the host):\n<a href=\"https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf\">https://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749507,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2020 18:34:19",
          "content": "<p>Yes thanks. But the Grapheme's are unbalanced in the train set. Are they unbalanced in the same way in the test set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749514,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "02/18/2020 18:39:31",
          "content": "<p>Noone knows and this is also no information hosts / kaggle will share. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749548,
          "author_name": "andreyzotov",
          "author_url": "",
          "post_date": "02/18/2020 19:15:09",
          "content": "<p>In many experiments of the participants (in mine, too), a strict correlation was observed between cross-validation score and the public test score - CV = PL + 0.007..0.009.  In my opinion, this indicates that the distribution in the train and public test is the same.  You can only guess about the distribution in a private test. I bet that distribution in the private test is the same and therefore the shake up will be minimal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749600,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2020 19:51:00",
          "content": "<p>Thanks for the info. My 20% validation holdout is 0.977 but my LB is only 0.967, so I think something is off with my validation set up. (Perhaps I need to isolate words or stratify Grapheme roots). Those scores are significantly different implying that there is something different about the relationship of test to train when compared to my validation set to train set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749713,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "02/18/2020 21:42:36",
          "content": "<p>using a random 20% validation isnt good idea, remember you said there are 150 of each 3-unit-combo. you want all of that 150 in your validation to represent test</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749722,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "02/18/2020 21:49:00",
          "content": "<p>Yeah, my \"hold out set\" is just my first fold of 5-Fold. I'm just too lazy to train all 5 folds. Currently it is plain KFold.</p>\n\n<p>There are 1292 words. I'm not sure where I want each word (the sets of 150). Perhaps GroupKFold where the 150 is either all in train or all in validation is best. Or perhaps StratifiedKFold where we stratify on the 168 Grapheme roots is best (and let words go where they go). Still trying different CVs. What's your CV?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749757,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "02/18/2020 22:20:18",
          "content": "<p>We are using iterative stratification and retain the 0.007..0.009 gap between CV and public LB. I am not very bothered by this gap because it is small &amp; consistent. Personally i think groupkfold on each word would be the best; however, we are doing experiments on just 1 fold, and if you are using groupkfold on each word, then you run the risk of overfitting to that single fold, so we would honestly need to run experiments on all folds to get something reliable which would take waaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaay too long.</p>\n\n<p>Anyway, to get the best results you want to retrain on full train dataset without using folds, so I am really calm with the gap.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749801,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/18/2020 22:56:11",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> I have a bigger CV/LB gap of 0.02. I am thinking that the different distribution might be the problem, too.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749829,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "02/18/2020 23:29:48",
          "content": "<p>Same thing for me, i can have 1% diff between my cv and lb if i dont balance de train data. One test i did was training imbalanced and then i only balanced consonant classes and it didn't really give me a noticeable boost so i guess consonants are not the problem because it's the most imbalanced part of the train data. I assume grapheme root is the problem??</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750126,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "02/19/2020 06:41:18",
          "content": "<p>I see most error between the grapheme root 59 to 92 , in the validation set . </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "748885": "Thank you Bengali.AI for sharing your dataset and hosting this interesting competition. I think I finally figured out what's going on :-)\n\n### Description\n\nThe training set contains approximately 150 each of 1292 common words. At the bottom of this post are 12 examples each of 4 common words.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F7096ae8ae94b5bf462562d443f6b88fe%2Fwords.png?generation=1582001513517840&amp;alt=media)\n\nEach word is composed of three pieces; a (1) Grapheme root (which confusingly are actually vowels, consonants, and consonant conjuncts themselves), a (2) vowel diacritic, and a (3) consonant diacritic. There are 168 different roots, 11 different v-diacritics, and 7 different c-diacritics total over train and test. Our job is to find and detect these pieces in words. The test set will contain words not present in the train's 1292 words (but uses the same pieces).\n\n### Questions\n1. Will all the test data words be new? Or will some be the same as training words?\n2. Are the 1292 common words representative of the language. i.e. is the frequency distribution of grapheme roots in train data the same as test data? Are rare graphemes rare in both train and test? And what about distribution of vowel diacritics and consonant diacritics?\n3. How many ways can the same word be drawn? I notice images that look different below among the same word?\n\n### Word 1\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fff591a018c0d602ddc6c8040db2e6a9a%2FScreen%20Shot%202020-02-17%20at%208.47.33%20PM.png?generation=1582001726970413&amp;alt=media)\n\n### Word 2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F22330c5d5c2c5acea1db961e84804bfa%2FScreen%20Shot%202020-02-17%20at%208.47.53%20PM.png?generation=1582001743502117&amp;alt=media)\n\n### Word 3\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb94d29b6470e26665a35499a4012704d%2FScreen%20Shot%202020-02-17%20at%208.48.17%20PM.png?generation=1582001756393939&amp;alt=media)\n\n### Word 4\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F78378bf870f2205a2acaad08df6b1b2a%2FScreen%20Shot%202020-02-17%20at%208.45.37%20PM.png?generation=1582001770729552&amp;alt=media)",
    "749083": "Hence why I really think we need to take some time to push for methods like Class Activation Maps to \"see\" which parts of the image are useful for a model. It feels like for now a lot of the development of a model is very trial &amp; error like: we try something and we see how that influences our local CV. I have the feeling that we´re not putting much effort (at least publicly, here, on Kaggle, I don't see that much) on error analysis, activation maps, etc.\n\nI feel like this would go a long way in helping answer some of those questions: are the models we're making robust to new combinations of symbols, or are there finding information in the whole character at once?",
    "749126": "As far as I know, all Grapheme roots, vowel diacritics, consonant diacritics between train and test are identical, but test set has more new connections from Grapheme roots, vowel diacritics and consonant diacritics.",
    "749480": "Can any naive speakers comment on the distribution (relative frequency counts) of Grapheme roots and diacritics of train versus test? This would be helpful to all of us. Is the distribution among the 1292 common words the same distribution among the rest of the language (and the test data)?",
    "749486": "Thanks. Yes. But i wonder is the most popular Grapheme root among the 1292 common words also the most popular among less common words (the test set)? What about the least common Grapheme root? Participants who don't know Bengla have no idea.",
    "749488": "Yes. Knowing where models are directing their attention could help us build better models. I like your notebook [here][1] showing this. Great job.\n\n[1]: https://www.kaggle.com/maxlenormand/multi-class-activation-map-with-efficientnetb0",
    "749502": "In the test set, we have the same 168-11-7 classes, but there are unseen combinations. The host confirmed this somewhere.\n\nTake a look at this doc (published by the host):\nhttps://bengali.ai/wp-content/uploads/CV19-COCO-Grapheme.pdf",
    "749507": "Yes thanks. But the Grapheme's are unbalanced in the train set. Are they unbalanced in the same way in the test set?",
    "749514": "Noone knows and this is also no information hosts / kaggle will share.",
    "749548": "In many experiments of the participants (in mine, too), a strict correlation was observed between cross-validation score and the public test score - CV = PL + 0.007..0.009.  In my opinion, this indicates that the distribution in the train and public test is the same.  You can only guess about the distribution in a private test. I bet that distribution in the private test is the same and therefore the shake up will be minimal.",
    "749600": "Thanks for the info. My 20% validation holdout is 0.977 but my LB is only 0.967, so I think something is off with my validation set up. (Perhaps I need to isolate words or stratify Grapheme roots). Those scores are significantly different implying that there is something different about the relationship of test to train when compared to my validation set to train set.",
    "749713": "using a random 20% validation isnt good idea, remember you said there are 150 of each 3-unit-combo. you want all of that 150 in your validation to represent test",
    "749722": "Yeah, my \"hold out set\" is just my first fold of 5-Fold. I'm just too lazy to train all 5 folds. Currently it is plain KFold.\n\nThere are 1292 words. I'm not sure where I want each word (the sets of 150). Perhaps GroupKFold where the 150 is either all in train or all in validation is best. Or perhaps StratifiedKFold where we stratify on the 168 Grapheme roots is best (and let words go where they go). Still trying different CVs. What's your CV?",
    "749757": "We are using iterative stratification and retain the 0.007..0.009 gap between CV and public LB. I am not very bothered by this gap because it is small &amp; consistent. Personally i think groupkfold on each word would be the best; however, we are doing experiments on just 1 fold, and if you are using groupkfold on each word, then you run the risk of overfitting to that single fold, so we would honestly need to run experiments on all folds to get something reliable which would take waaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaaay too long.\n\nAnyway, to get the best results you want to retrain on full train dataset without using folds, so I am really calm with the gap.",
    "749801": "cdeotte I have a bigger CV/LB gap of 0.02. I am thinking that the different distribution might be the problem, too.",
    "749829": "Same thing for me, i can have 1% diff between my cv and lb if i dont balance de train data. One test i did was training imbalanced and then i only balanced consonant classes and it didn't really give me a noticeable boost so i guess consonants are not the problem because it's the most imbalanced part of the train data. I assume grapheme root is the problem??",
    "750126": "I see most error between the grapheme root 59 to 92 , in the validation set .",
    "750299": "Hi @cdeotte  I am not sure I completely understand . Actually even if the person knows Bangla (Like me) would not be able to understand the distribution . Because\n\na) These are not words - Words are some meaningful combination of characters . So there is no such word in the set . These are like compund-characters , made up of three different characters.\n\nb) The characters crowdsourced and people are asked to write them as they want . So there is no way we can identify anything that are popular . Yes , there are some characters which are less used  in our day to day life , but that does not mean that test set might have less number of them . As an example word E and word Z in english , you know that E is used a lot more than Z , that does not mean if its crowdsourced to write E and Z , Z will have lesser population in the set . \n\nI am sorry , if I completely misunderstood your inference . Happy to help with little knowledge i have in this language .",
    "750384": "Thanks, glad you like it!"
  },
  "source": "meta"
}