{
  "id": 135900,
  "title": "Different Data Distribution between Train/Test?",
  "url": "/competitions/bengaliai-cv19/discussion/135900",
  "author_name": "",
  "post_date": "2020-03-16T16:31:13.477310200Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Top ranking solutions perform very well.</p>\n\n<p>In my case, I got around 0.998 in my CV and 0.990 in LB.</p>\n\n<p>As you can see my CV score, there is almost no room to improve. Also my training score is not that differ from the score. </p>\n\n<hr>\n\n<p>This leads me to guess that there is a data distribution shift between training set and LB test set. </p>\n\n<p>That means, top rankers have figured out how to match their training distribution with data processing method. </p>\n\n<hr>\n\n<p>If private LB have also different distribution than public LB, (this is highly unlikely), big shake-ups can be anticipated. </p>\n\n<p>If current top ranking solutions will stay at the top of the private LB, I will look forward to hear their solutions to make their models work on LB data distribution.</p>\n\n<p>:) </p>",
  "messages": [
    {
      "id": "775393",
      "postDate": "03/16/2020 16:31:13",
      "content": "<p>Top ranking solutions perform very well.</p>\n\n<p>In my case, I got around 0.998 in my CV and 0.990 in LB.</p>\n\n<p>As you can see my CV score, there is almost no room to improve. Also my training score is not that differ from the score. </p>\n\n<hr>\n\n<p>This leads me to guess that there is a data distribution shift between training set and LB test set. </p>\n\n<p>That means, top rankers have figured out how to match their training distribution with data processing method. </p>\n\n<hr>\n\n<p>If private LB have also different distribution than public LB, (this is highly unlikely), big shake-ups can be anticipated. </p>\n\n<p>If current top ranking solutions will stay at the top of the private LB, I will look forward to hear their solutions to make their models work on LB data distribution.</p>\n\n<p>:) </p>",
      "rawMarkdown": "Top ranking solutions perform very well.\n\nIn my case, I got around 0.998 in my CV and 0.990 in LB.\n\nAs you can see my CV score, there is almost no room to improve. Also my training score is not that differ from the score. \n\n----\n\nThis leads me to guess that there is a data distribution shift between training set and LB test set. \n\nThat means, top rankers have figured out how to match their training distribution with data processing method. \n\n----\n\nIf private LB have also different distribution than public LB, (this is highly unlikely), big shake-ups can be anticipated. \n\nIf current top ranking solutions will stay at the top of the private LB, I will look forward to hear their solutions to make their models work on LB data distribution.\n\n:)",
      "votes": null
    },
    {
      "id": "775421",
      "postDate": "03/16/2020 16:57:55",
      "content": "<p>Well it is too late even if someone found out the way haha...</p>",
      "rawMarkdown": "Well it is too late even if someone found out the way haha...",
      "votes": null
    },
    {
      "id": "775428",
      "postDate": "03/16/2020 17:01:31",
      "content": "<p>So let's stay back and relax! I guess everyone's up hahaha</p>",
      "rawMarkdown": "So let's stay back and relax! I guess everyone's up hahaha",
      "votes": null
    },
    {
      "id": "775453",
      "postDate": "03/16/2020 17:39:24",
      "content": "<p>\"That means, top rankers have figured out how to match their training distribution with data processing method.\"</p>\n\n<p>i don't think so. \nthey are probably doing something very different, e.. different supervisory signal or loss, post-processing or error correction,etc</p>",
      "rawMarkdown": "\"That means, top rankers have figured out how to match their training distribution with data processing method.\"\n\ni don't think so. \nthey are probably doing something very different, e.. different supervisory signal or loss, post-processing or error correction,etc",
      "votes": null
    },
    {
      "id": "775462",
      "postDate": "03/16/2020 17:45:44",
      "content": "<p>well maybe you're right but i have a hunch that it might be not that too complicated.</p>",
      "rawMarkdown": "well maybe you're right but i have a hunch that it might be not that too complicated.",
      "votes": null
    },
    {
      "id": "775494",
      "postDate": "03/16/2020 18:29:26",
      "content": "<p>My hypothesis is that it’s some data trick😃 well we’ll know very shortly. Good solutions are often surprisingly simple</p>",
      "rawMarkdown": "My hypothesis is that it’s some data trick😃 well we’ll know very shortly. Good solutions are often surprisingly simple",
      "votes": null
    },
    {
      "id": "775543",
      "postDate": "03/16/2020 19:41:38",
      "content": "<p>My current LB doesn't involve any post processing of predictions. Just very boring ensembling of 4 folds ... and I am really curious what those tricks are </p>",
      "rawMarkdown": "My current LB doesn't involve any post processing of predictions. Just very boring ensembling of 4 folds ... and I am really curious what those tricks are",
      "votes": null
    },
    {
      "id": "775544",
      "postDate": "03/16/2020 19:42:56",
      "content": "<p>its not =) our team jumped from 20th to 7th within 5 hours =) so don't give up and good luck! =)</p>",
      "rawMarkdown": "its not =) our team jumped from 20th to 7th within 5 hours =) so don't give up and good luck! =)",
      "votes": null
    },
    {
      "id": "775680",
      "postDate": "03/16/2020 23:57:56",
      "content": "<p>in the end, our solution are just plain ensemble... we will know what the magic is really soon ;)</p>",
      "rawMarkdown": "in the end, our solution are just plain ensemble... we will know what the magic is really soon ;)",
      "votes": null
    },
    {
      "id": "775760",
      "postDate": "03/17/2020 00:44:04",
      "content": "<p>Public LB cannot be trust at all</p>",
      "rawMarkdown": "Public LB cannot be trust at all",
      "votes": null
    },
    {
      "id": "775778",
      "postDate": "03/17/2020 00:53:55",
      "content": "<p>It is not surprising to me to have 8% new graphemes in privprivate dataset ;)</p>",
      "rawMarkdown": "It is not surprising to me to have 8% new graphemes in privprivate dataset ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 775421,
      "author_name": "soonhwankwon",
      "author_url": "",
      "post_date": "03/16/2020 16:57:55",
      "content": "<p>Well it is too late even if someone found out the way haha...</p>",
      "votes": null,
      "replies": [
        {
          "id": 775544,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "03/16/2020 19:42:56",
          "content": "<p>its not =) our team jumped from 20th to 7th within 5 hours =) so don't give up and good luck! =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775428,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "03/16/2020 17:01:31",
      "content": "<p>So let's stay back and relax! I guess everyone's up hahaha</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775453,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/16/2020 17:39:24",
      "content": "<p>\"That means, top rankers have figured out how to match their training distribution with data processing method.\"</p>\n\n<p>i don't think so. \nthey are probably doing something very different, e.. different supervisory signal or loss, post-processing or error correction,etc</p>",
      "votes": null,
      "replies": [
        {
          "id": 775462,
          "author_name": "soonhwankwon",
          "author_url": "",
          "post_date": "03/16/2020 17:45:44",
          "content": "<p>well maybe you're right but i have a hunch that it might be not that too complicated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 775494,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "03/16/2020 18:29:26",
          "content": "<p>My hypothesis is that it’s some data trick😃 well we’ll know very shortly. Good solutions are often surprisingly simple</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 775543,
      "author_name": "weimin",
      "author_url": "",
      "post_date": "03/16/2020 19:41:38",
      "content": "<p>My current LB doesn't involve any post processing of predictions. Just very boring ensembling of 4 folds ... and I am really curious what those tricks are </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775680,
      "author_name": "moewie94",
      "author_url": "",
      "post_date": "03/16/2020 23:57:56",
      "content": "<p>in the end, our solution are just plain ensemble... we will know what the magic is really soon ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775760,
      "author_name": "kasim0226",
      "author_url": "",
      "post_date": "03/17/2020 00:44:04",
      "content": "<p>Public LB cannot be trust at all</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 775778,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "03/17/2020 00:53:55",
      "content": "<p>It is not surprising to me to have 8% new graphemes in privprivate dataset ;)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "775393": "Top ranking solutions perform very well.\n\nIn my case, I got around 0.998 in my CV and 0.990 in LB.\n\nAs you can see my CV score, there is almost no room to improve. Also my training score is not that differ from the score. \n\n----\n\nThis leads me to guess that there is a data distribution shift between training set and LB test set. \n\nThat means, top rankers have figured out how to match their training distribution with data processing method. \n\n----\n\nIf private LB have also different distribution than public LB, (this is highly unlikely), big shake-ups can be anticipated. \n\nIf current top ranking solutions will stay at the top of the private LB, I will look forward to hear their solutions to make their models work on LB data distribution.\n\n:)",
    "775421": "Well it is too late even if someone found out the way haha...",
    "775428": "So let's stay back and relax! I guess everyone's up hahaha",
    "775453": "\"That means, top rankers have figured out how to match their training distribution with data processing method.\"\n\ni don't think so. \nthey are probably doing something very different, e.. different supervisory signal or loss, post-processing or error correction,etc",
    "775462": "well maybe you're right but i have a hunch that it might be not that too complicated.",
    "775494": "My hypothesis is that it’s some data trick😃 well we’ll know very shortly. Good solutions are often surprisingly simple",
    "775543": "My current LB doesn't involve any post processing of predictions. Just very boring ensembling of 4 folds ... and I am really curious what those tricks are",
    "775544": "its not =) our team jumped from 20th to 7th within 5 hours =) so don't give up and good luck! =)",
    "775680": "in the end, our solution are just plain ensemble... we will know what the magic is really soon ;)",
    "775760": "Public LB cannot be trust at all",
    "775778": "It is not surprising to me to have 8% new graphemes in privprivate dataset ;)"
  },
  "source": "meta"
}