{
  "id": 99397,
  "title": "Can we rely on the public LB?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/99397",
  "author_name": "",
  "post_date": "2019-07-11T01:24:21.364291200Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>CV is too high comparing to the LB score. I have ran a model that scores 0.93 CV, 0.699 on LB and 0.42 on the train old data.\nMoreover, training a model on the old dataset would give 0.5-0.6 score on LB with 0.8+ CV.\nBoth datasets (the new and the old one) have the same nature of targets and  same images but models failed to predict a dataset based on the other (Train on 2015 data and predict the new one and vice versa).</p>\n\n<p>Here are some experiments I did trying to understand the gap between CV and LB:</p>\n\n<blockquote>\n  <p>Train on 2015 data and submit for this competition's test : 0.8+ cv and 0.5- LB\n  =&gt; 2015 data is different from our test\n  Train on the actual train data and submit : 0.92 CV and 0.6+ LB\n  =&gt; train data is different from the test\n  Train on old train and submit for the old test : CV and LB are correlated\n  =&gt; the old train is similar to this old test\n  Train on the actual train and predict for the old train : 0.92CV and 0.42 on the old train\n  =&gt; new train is different from the old train, but here we have (0.92+0.42)/2 ~ LB score\n  So I thought to train on both datasets old train + actual train: got 0.83 CV and 0.55LB</p>\n</blockquote>\n\n<p>After running all these experiments I think public LB is useless and there would be a big shape up at the end of the competition unless someone would discover a new way to validate our models.</p>",
  "messages": [
    {
      "id": "572468",
      "postDate": "07/11/2019 01:24:21",
      "content": "<p>CV is too high comparing to the LB score. I have ran a model that scores 0.93 CV, 0.699 on LB and 0.42 on the train old data.\nMoreover, training a model on the old dataset would give 0.5-0.6 score on LB with 0.8+ CV.\nBoth datasets (the new and the old one) have the same nature of targets and  same images but models failed to predict a dataset based on the other (Train on 2015 data and predict the new one and vice versa).</p>\n\n<p>Here are some experiments I did trying to understand the gap between CV and LB:</p>\n\n<blockquote>\n  <p>Train on 2015 data and submit for this competition's test : 0.8+ cv and 0.5- LB\n  =&gt; 2015 data is different from our test\n  Train on the actual train data and submit : 0.92 CV and 0.6+ LB\n  =&gt; train data is different from the test\n  Train on old train and submit for the old test : CV and LB are correlated\n  =&gt; the old train is similar to this old test\n  Train on the actual train and predict for the old train : 0.92CV and 0.42 on the old train\n  =&gt; new train is different from the old train, but here we have (0.92+0.42)/2 ~ LB score\n  So I thought to train on both datasets old train + actual train: got 0.83 CV and 0.55LB</p>\n</blockquote>\n\n<p>After running all these experiments I think public LB is useless and there would be a big shape up at the end of the competition unless someone would discover a new way to validate our models.</p>",
      "rawMarkdown": "CV is too high comparing to the LB score. I have ran a model that scores 0.93 CV, 0.699 on LB and 0.42 on the train old data.\nMoreover, training a model on the old dataset would give 0.5-0.6 score on LB with 0.8+ CV.\nBoth datasets (the new and the old one) have the same nature of targets and  same images but models failed to predict a dataset based on the other (Train on 2015 data and predict the new one and vice versa).\n\nHere are some experiments I did trying to understand the gap between CV and LB:\n&gt;Train on 2015 data and submit for this competition's test : 0.8+ cv and 0.5- LB\n=&gt; 2015 data is different from our test\n&gt;Train on the actual train data and submit : 0.92 CV and 0.6+ LB\n=&gt; train data is different from the test\n&gt;Train on old train and submit for the old test : CV and LB are correlated\n=&gt; the old train is similar to this old test\n&gt;Train on the actual train and predict for the old train : 0.92CV and 0.42 on the old train\n=&gt; new train is different from the old train, but here we have (0.92+0.42)/2 ~ LB score\n&gt;So I thought to train on both datasets old train + actual train: got 0.83 CV and 0.55LB\n\nAfter running all these experiments I think public LB is useless and there would be a big shape up at the end of the competition unless someone would discover a new way to validate our models.",
      "votes": null
    },
    {
      "id": "572472",
      "postDate": "07/11/2019 01:27:56",
      "content": "<p>Unfortunately, not. See <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361\">here</a> for more information.</p>",
      "rawMarkdown": "Unfortunately, not. See [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361) for more information.",
      "votes": null
    },
    {
      "id": "572533",
      "postDate": "07/11/2019 03:27:39",
      "content": "<p>Trust your CV :)</p>",
      "rawMarkdown": "Trust your CV :)",
      "votes": null
    },
    {
      "id": "572575",
      "postDate": "07/11/2019 05:52:46",
      "content": "<p>nop</p>",
      "rawMarkdown": "nop",
      "votes": null
    },
    {
      "id": "572811",
      "postDate": "07/11/2019 12:31:42",
      "content": "<p>Should I trust my stratified train test split. It's giving 90 on CV but 66 on LB.</p>",
      "rawMarkdown": "Should I trust my stratified train test split. It's giving 90 on CV but 66 on LB.",
      "votes": null
    },
    {
      "id": "572939",
      "postDate": "07/11/2019 15:31:53",
      "content": "<p>I can recommend you read this <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-572069\">discussion</a> for understanding why you obtain this result.</p>",
      "rawMarkdown": "I can recommend you read this [discussion](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-572069) for understanding why you obtain this result.",
      "votes": null
    },
    {
      "id": "572948",
      "postDate": "07/11/2019 15:37:41",
      "content": "<p>In the old competition, they didnt have this problem... Why we re having it now?</p>",
      "rawMarkdown": "In the old competition, they didnt have this problem... Why we re having it now?",
      "votes": null
    },
    {
      "id": "572963",
      "postDate": "07/11/2019 15:46:09",
      "content": "<p>In this competition now used fake test dataset with 8% leaks and very different from train dataset data. Also current leaderboard calculated on 15% -&gt; 300 img. But it's very short.</p>",
      "rawMarkdown": "In this competition now used fake test dataset with 8% leaks and very different from train dataset data. Also current leaderboard calculated on 15% -&gt; 300 img. But it's very short.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 572472,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "07/11/2019 01:27:56",
      "content": "<p>Unfortunately, not. See <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361\">here</a> for more information.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 572533,
      "author_name": "snakayama",
      "author_url": "",
      "post_date": "07/11/2019 03:27:39",
      "content": "<p>Trust your CV :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 572575,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "07/11/2019 05:52:46",
      "content": "<p>nop</p>",
      "votes": null,
      "replies": [
        {
          "id": 572811,
          "author_name": "nitin29",
          "author_url": "",
          "post_date": "07/11/2019 12:31:42",
          "content": "<p>Should I trust my stratified train test split. It's giving 90 on CV but 66 on LB.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 572939,
      "author_name": "miklgr500",
      "author_url": "",
      "post_date": "07/11/2019 15:31:53",
      "content": "<p>I can recommend you read this <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-572069\">discussion</a> for understanding why you obtain this result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 572948,
          "author_name": "rinnqd",
          "author_url": "",
          "post_date": "07/11/2019 15:37:41",
          "content": "<p>In the old competition, they didnt have this problem... Why we re having it now?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 572963,
          "author_name": "miklgr500",
          "author_url": "",
          "post_date": "07/11/2019 15:46:09",
          "content": "<p>In this competition now used fake test dataset with 8% leaks and very different from train dataset data. Also current leaderboard calculated on 15% -&gt; 300 img. But it's very short.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "572468": "CV is too high comparing to the LB score. I have ran a model that scores 0.93 CV, 0.699 on LB and 0.42 on the train old data.\nMoreover, training a model on the old dataset would give 0.5-0.6 score on LB with 0.8+ CV.\nBoth datasets (the new and the old one) have the same nature of targets and  same images but models failed to predict a dataset based on the other (Train on 2015 data and predict the new one and vice versa).\n\nHere are some experiments I did trying to understand the gap between CV and LB:\n&gt;Train on 2015 data and submit for this competition's test : 0.8+ cv and 0.5- LB\n=&gt; 2015 data is different from our test\n&gt;Train on the actual train data and submit : 0.92 CV and 0.6+ LB\n=&gt; train data is different from the test\n&gt;Train on old train and submit for the old test : CV and LB are correlated\n=&gt; the old train is similar to this old test\n&gt;Train on the actual train and predict for the old train : 0.92CV and 0.42 on the old train\n=&gt; new train is different from the old train, but here we have (0.92+0.42)/2 ~ LB score\n&gt;So I thought to train on both datasets old train + actual train: got 0.83 CV and 0.55LB\n\nAfter running all these experiments I think public LB is useless and there would be a big shape up at the end of the competition unless someone would discover a new way to validate our models.",
    "572472": "Unfortunately, not. See [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493#latest-568361) for more information.",
    "572533": "Trust your CV :)",
    "572575": "nop",
    "572811": "Should I trust my stratified train test split. It's giving 90 on CV but 66 on LB.",
    "572939": "I can recommend you read this [discussion](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98856#latest-572069) for understanding why you obtain this result.",
    "572948": "In the old competition, they didnt have this problem... Why we re having it now?",
    "572963": "In this competition now used fake test dataset with 8% leaks and very different from train dataset data. Also current leaderboard calculated on 15% -&gt; 300 img. But it's very short."
  },
  "source": "meta"
}