{
  "id": 134434,
  "title": " Have a Try on Validating With Unseen?",
  "url": "/competitions/bengaliai-cv19/discussion/134434",
  "author_name": "Qishen Ha",
  "post_date": "2020-03-08T04:29:51.575000",
  "votes": 34,
  "comment_count": 20,
  "views": 0,
  "content": "<p>There is a risk of shake-up for this competition, which I have mentioned here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134035</a></p>\n\n<p>If we make the unseen validation splits separately, it is difficult to compare the ROBUSTNESS of models with each other.</p>\n\n<p>So I've made a split for you guys ;) <a href=\"https://www.kaggle.com/haqishen/validation-with-unseen\">https://www.kaggle.com/haqishen/validation-with-unseen</a></p>\n\n<p>Using my split, we will have approximately 38k SEEN samples, and exacly 7578 UNSEEN samples for validation in each fold. (16.4% unseen)</p>\n\n<p>Download the file <code>train_v2.csv</code> generated by this script and have a try on yourself!</p>",
  "messages": [
    {
      "id": 766380,
      "postDate": "2020-03-08T04:29:51.577Z",
      "content": "<p>There is a risk of shake-up for this competition, which I have mentioned here: <a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\">https://www.kaggle.com/c/bengaliai-cv19/discussion/134035</a></p>\n\n<p>If we make the unseen validation splits separately, it is difficult to compare the ROBUSTNESS of models with each other.</p>\n\n<p>So I've made a split for you guys ;) <a href=\"https://www.kaggle.com/haqishen/validation-with-unseen\">https://www.kaggle.com/haqishen/validation-with-unseen</a></p>\n\n<p>Using my split, we will have approximately 38k SEEN samples, and exacly 7578 UNSEEN samples for validation in each fold. (16.4% unseen)</p>\n\n<p>Download the file <code>train_v2.csv</code> generated by this script and have a try on yourself!</p>",
      "rawMarkdown": "There is a risk of shake-up for this competition, which I have mentioned here: https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\n\nIf we make the unseen validation splits separately, it is difficult to compare the ROBUSTNESS of models with each other.\n\nSo I've made a split for you guys ;) https://www.kaggle.com/haqishen/validation-with-unseen\n\nUsing my split, we will have approximately 38k SEEN samples, and exacly 7578 UNSEEN samples for validation in each fold. (16.4% unseen)\n\nDownload the file `train_v2.csv` generated by this script and have a try on yourself!",
      "votes": 34
    },
    {
      "id": 766525,
      "postDate": "2020-03-08T09:55:10.603Z",
      "content": "<p>From my first fold experiment, I dont see much difference between random split and your unseen split. 🙁 . \nIn fact, I did split by grapheme (100% unseen validation) and could not surpass 0.7 CV</p>",
      "rawMarkdown": "From my first fold experiment, I dont see much difference between random split and your unseen split. 🙁 . \nIn fact, I did split by grapheme (100% unseen validation) and could not surpass 0.7 CV",
      "votes": 1,
      "replies": [
        {
          "id": 766575,
          "postDate": "2020-03-08T11:57:44.510Z",
          "content": "<p>Yes, it's hard to generalize to unseen graphemes</p>",
          "rawMarkdown": "Yes, it's hard to generalize to unseen graphemes"
        }
      ]
    },
    {
      "id": 766708,
      "postDate": "2020-03-08T15:43:48.033Z",
      "content": "<p>My experiment on this validation split:\n<code>\nLocal Score: 0.9888  (my fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n</code></p>",
      "rawMarkdown": "My experiment on this validation split:\n```\nLocal Score: 0.9888  (my fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n```",
      "votes": 2,
      "replies": [
        {
          "id": 766714,
          "postDate": "2020-03-08T15:50:11.300Z",
          "content": "<p>Do you compute score on unseen only?. \nMy CV:\nRandom split: 0.9756.\nUnseen split: 0.9774. Unseen score: 0.83</p>",
          "rawMarkdown": "Do you compute score on unseen only?. \nMy CV:\nRandom split: 0.9756.\nUnseen split: 0.9774. Unseen score: 0.83"
        },
        {
          "id": 766716,
          "postDate": "2020-03-08T15:54:00.497Z",
          "content": "<p>My unseen cannot surpass 0.8😲 I think your 0.83 is very good score.\nAnd what's your <code>Random split</code> setting? How much unseen in that split?</p>",
          "rawMarkdown": "My unseen cannot surpass 0.8😲 I think your 0.83 is very good score.\nAnd what's your `Random split` setting? How much unseen in that split?"
        },
        {
          "id": 766721,
          "postDate": "2020-03-08T15:59:23.190Z",
          "content": "<p>Random split doesnot have unseen, so I could not score it. Here I just want to state that, unseen split does not matter to my old random split one. </p>",
          "rawMarkdown": "Random split doesnot have unseen, so I could not score it. Here I just want to state that, unseen split does not matter to my old random split one. "
        },
        {
          "id": 766730,
          "postDate": "2020-03-08T16:08:45.957Z",
          "content": "<p>Oh, I see.</p>\n\n<p>&gt; Unseen split: 0.9774</p>\n\n<p>Is this score comes from seen graphemes only? Or is it computed from seen + unseen samples?</p>",
          "rawMarkdown": "Oh, I see.\n\n&gt; Unseen split: 0.9774\n\nIs this score comes from seen graphemes only? Or is it computed from seen + unseen samples?"
        },
        {
          "id": 766736,
          "postDate": "2020-03-08T16:12:47.340Z",
          "content": "<p>it is computed on both seen and unseen graphems</p>",
          "rawMarkdown": "it is computed on both seen and unseen graphems"
        },
        {
          "id": 766739,
          "postDate": "2020-03-08T16:18:13.950Z",
          "content": "<p>That's strange, you got a 0.83 on unseen only (16.4% samples),  but have a similar score to old split when computed on both seen (83.6%) and unseen (16.4%) 😲</p>\n\n<p>It seems that... you got a 0.99+ on seen graphemes only?</p>",
          "rawMarkdown": "That's strange, you got a 0.83 on unseen only (16.4% samples),  but have a similar score to old split when computed on both seen (83.6%) and unseen (16.4%) 😲\n\nIt seems that... you got a 0.99+ on seen graphemes only?"
        },
        {
          "id": 766755,
          "postDate": "2020-03-08T16:39:15.967Z",
          "content": "<p>I downloaded your unseen split. There are only 1515 unseen graphemes (~3.7%) in fold 0. Did you change something?</p>",
          "rawMarkdown": "I downloaded your unseen split. There are only 1515 unseen graphemes (~3.7%) in fold 0. Did you change something?"
        },
        {
          "id": 766766,
          "postDate": "2020-03-08T16:54:35.917Z",
          "content": "<p>Sorry, that's my fault, I've updated the notebook and made all unseen graphemes to <code>fold = -1</code> to ensure no one use it to train.... When doing validation, please use all unseen graphemes in each folds</p>",
          "rawMarkdown": "Sorry, that's my fault, I've updated the notebook and made all unseen graphemes to `fold = -1` to ensure no one use it to train.... When doing validation, please use all unseen graphemes in each folds"
        },
        {
          "id": 766972,
          "postDate": "2020-03-09T01:52:06.203Z",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> Wondering what is the single-model LB for the same approach using the old validation scheme only?</p>",
          "rawMarkdown": "@haqishen Wondering what is the single-model LB for the same approach using the old validation scheme only?"
        }
      ]
    },
    {
      "id": 770215,
      "postDate": "2020-03-12T17:31:53.930Z",
      "content": "<p>how about training external dataset?\ne.g.\nIII. The BSU Bangla Dataset\n<a href=\"https://scholarworks.boisestate.edu/cgi/viewcontent.cgi?article=1392&amp;context=electrical_facpubs\">https://scholarworks.boisestate.edu/cgi/viewcontent.cgi?article=1392&amp;context=electrical_facpubs</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc0d7efb6ac607af4e82b7d3fc61ec66c%2FSelection_059.png?generation=1584034311341730&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "how about training external dataset?\ne.g.\nIII. The BSU Bangla Dataset\nhttps://scholarworks.boisestate.edu/cgi/viewcontent.cgi?article=1392&amp;context=electrical_facpubs\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc0d7efb6ac607af4e82b7d3fc61ec66c%2FSelection_059.png?generation=1584034311341730&amp;alt=media)\n"
    },
    {
      "id": 766971,
      "postDate": "2020-03-09T01:49:20.570Z",
      "content": "<p>So do you recommend to use this split data for the training process or using random split? Thank you very much.</p>",
      "rawMarkdown": "So do you recommend to use this split data for the training process or using random split? Thank you very much.",
      "replies": [
        {
          "id": 767014,
          "postDate": "2020-03-09T03:18:17.350Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#767010\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#767010</a>\nThis is what I found and what I have done, but it's not necessary for you guys to follow.\nYou have to make your own choice base on the information.</p>",
          "rawMarkdown": "https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#767010\nThis is what I found and what I have done, but it's not necessary for you guys to follow.\nYou have to make your own choice base on the information.",
          "votes": 1
        }
      ]
    },
    {
      "id": 766461,
      "postDate": "2020-03-08T07:01:09.947Z",
      "content": "<p>May I ask why use 16.4% this ratio for unseen in each fold? </p>",
      "rawMarkdown": "May I ask why use 16.4% this ratio for unseen in each fold? ",
      "replies": [
        {
          "id": 766463,
          "postDate": "2020-03-08T07:04:49.600Z",
          "content": "<p>all unseen samples are used in all fold for validation, so yes, approximately 16.4% unseen in each fold (Though they are exactly the same samples). </p>",
          "rawMarkdown": "all unseen samples are used in all fold for validation, so yes, approximately 16.4% unseen in each fold (Though they are exactly the same samples). "
        },
        {
          "id": 766467,
          "postDate": "2020-03-08T07:13:46.517Z",
          "content": "<p>I mean that in your notebook, you set grapheme &gt;= 1245 as unseen. So why do you split like this? </p>",
          "rawMarkdown": "I mean that in your notebook, you set grapheme &gt;= 1245 as unseen. So why do you split like this? "
        },
        {
          "id": 766468,
          "postDate": "2020-03-08T07:18:31.923Z",
          "content": "<p>Just for convenience.</p>\n\n<p>It maintained all components in training set, while we don't need to modify that much code for getting it start training.</p>",
          "rawMarkdown": "Just for convenience.\n\nIt maintained all components in training set, while we don't need to modify that much code for getting it start training."
        }
      ]
    },
    {
      "id": 766383,
      "postDate": "2020-03-08T04:31:58.100Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 766525,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2020-03-08T09:55:10.603000",
      "content": "<p>From my first fold experiment, I dont see much difference between random split and your unseen split. 🙁 . \nIn fact, I did split by grapheme (100% unseen validation) and could not surpass 0.7 CV</p>",
      "votes": 1,
      "replies": [
        {
          "id": 766575,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T11:57:44.510000",
          "content": "<p>Yes, it's hard to generalize to unseen graphemes</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 766708,
      "author_name": "Qishen Ha",
      "author_url": "",
      "post_date": "2020-03-08T15:43:48.033000",
      "content": "<p>My experiment on this validation split:\n<code>\nLocal Score: 0.9888  (my fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 766714,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2020-03-08T15:50:11.300000",
          "content": "<p>Do you compute score on unseen only?. \nMy CV:\nRandom split: 0.9756.\nUnseen split: 0.9774. Unseen score: 0.83</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766716,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T15:54:00.497000",
          "content": "<p>My unseen cannot surpass 0.8😲 I think your 0.83 is very good score.\nAnd what's your <code>Random split</code> setting? How much unseen in that split?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766721,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2020-03-08T15:59:23.190000",
          "content": "<p>Random split doesnot have unseen, so I could not score it. Here I just want to state that, unseen split does not matter to my old random split one. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766730,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T16:08:45.957000",
          "content": "<p>Oh, I see.</p>\n\n<p>&gt; Unseen split: 0.9774</p>\n\n<p>Is this score comes from seen graphemes only? Or is it computed from seen + unseen samples?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766736,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2020-03-08T16:12:47.340000",
          "content": "<p>it is computed on both seen and unseen graphems</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766739,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T16:18:13.950000",
          "content": "<p>That's strange, you got a 0.83 on unseen only (16.4% samples),  but have a similar score to old split when computed on both seen (83.6%) and unseen (16.4%) 😲</p>\n\n<p>It seems that... you got a 0.99+ on seen graphemes only?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766755,
          "author_name": "cab",
          "author_url": "",
          "post_date": "2020-03-08T16:39:15.967000",
          "content": "<p>I downloaded your unseen split. There are only 1515 unseen graphemes (~3.7%) in fold 0. Did you change something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766766,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T16:54:35.917000",
          "content": "<p>Sorry, that's my fault, I've updated the notebook and made all unseen graphemes to <code>fold = -1</code> to ensure no one use it to train.... When doing validation, please use all unseen graphemes in each folds</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766972,
          "author_name": "Nicholas Lyu",
          "author_url": "",
          "post_date": "2020-03-09T01:52:06.203000",
          "content": "<p><a href=\"/haqishen\">@haqishen</a> Wondering what is the single-model LB for the same approach using the old validation scheme only?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 770215,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-03-12T17:31:53.930000",
      "content": "<p>how about training external dataset?\ne.g.\nIII. The BSU Bangla Dataset\n<a href=\"https://scholarworks.boisestate.edu/cgi/viewcontent.cgi?article=1392&amp;context=electrical_facpubs\">https://scholarworks.boisestate.edu/cgi/viewcontent.cgi?article=1392&amp;context=electrical_facpubs</a></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc0d7efb6ac607af4e82b7d3fc61ec66c%2FSelection_059.png?generation=1584034311341730&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 766971,
      "author_name": "Bo Peng",
      "author_url": "",
      "post_date": "2020-03-09T01:49:20.570000",
      "content": "<p>So do you recommend to use this split data for the training process or using random split? Thank you very much.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 767014,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-09T03:18:17.350000",
          "content": "<p><a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#767010\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123198#767010</a>\nThis is what I found and what I have done, but it's not necessary for you guys to follow.\nYou have to make your own choice base on the information.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 766461,
      "author_name": "Morphy",
      "author_url": "",
      "post_date": "2020-03-08T07:01:09.947000",
      "content": "<p>May I ask why use 16.4% this ratio for unseen in each fold? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 766463,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T07:04:49.600000",
          "content": "<p>all unseen samples are used in all fold for validation, so yes, approximately 16.4% unseen in each fold (Though they are exactly the same samples). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766467,
          "author_name": "Morphy",
          "author_url": "",
          "post_date": "2020-03-08T07:13:46.517000",
          "content": "<p>I mean that in your notebook, you set grapheme &gt;= 1245 as unseen. So why do you split like this? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766468,
          "author_name": "Qishen Ha",
          "author_url": "",
          "post_date": "2020-03-08T07:18:31.923000",
          "content": "<p>Just for convenience.</p>\n\n<p>It maintained all components in training set, while we don't need to modify that much code for getting it start training.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 766383,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-08T04:31:58.100000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "766380": "There is a risk of shake-up for this competition, which I have mentioned here: https://www.kaggle.com/c/bengaliai-cv19/discussion/134035\n\nIf we make the unseen validation splits separately, it is difficult to compare the ROBUSTNESS of models with each other.\n\nSo I've made a split for you guys ;) https://www.kaggle.com/haqishen/validation-with-unseen\n\nUsing my split, we will have approximately 38k SEEN samples, and exacly 7578 UNSEEN samples for validation in each fold. (16.4% unseen)\n\nDownload the file `train_v2.csv` generated by this script and have a try on yourself!",
    "766525": "From my first fold experiment, I dont see much difference between random split and your unseen split. 🙁 . \nIn fact, I did split by grapheme (100% unseen validation) and could not surpass 0.7 CV",
    "766708": "My experiment on this validation split:\n```\nLocal Score: 0.9888  (my fold_0: 38653 seen samples and 7578 unseen samples)\nLocal Score: 0.998~ (old split, single fold)\nPublic LB: 0.9923 (single fold)\n```",
    "770215": "how about training external dataset?\ne.g.\nIII. The BSU Bangla Dataset\nhttps://scholarworks.boisestate.edu/cgi/viewcontent.cgi?article=1392&amp;context=electrical_facpubs\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fc0d7efb6ac607af4e82b7d3fc61ec66c%2FSelection_059.png?generation=1584034311341730&amp;alt=media)\n",
    "766971": "So do you recommend to use this split data for the training process or using random split? Thank you very much.",
    "766461": "May I ask why use 16.4% this ratio for unseen in each fold? ",
    "766383": ""
  }
}