{
  "id": 102827,
  "title": "how to do cv correctly?",
  "url": "/competitions/aptos2019-blindness-detection/discussion/102827",
  "author_name": "Abhishek Thakur",
  "post_date": "2019-08-05T11:19:25.797000",
  "votes": 31,
  "comment_count": 16,
  "views": 0,
  "content": "<p>So far, we have been doing kfolds in a stratified manner. However, it seems CV and LB vary a lot. With a CV of 0.92, we get 0.79. I would, thus, like to know what your CV approach is and if we are doing something wrong in the cross validation? </p>",
  "messages": [
    {
      "id": 592484,
      "postDate": "2019-08-05T11:19:25.797Z",
      "content": "<p>So far, we have been doing kfolds in a stratified manner. However, it seems CV and LB vary a lot. With a CV of 0.92, we get 0.79. I would, thus, like to know what your CV approach is and if we are doing something wrong in the cross validation? </p>",
      "rawMarkdown": "So far, we have been doing kfolds in a stratified manner. However, it seems CV and LB vary a lot. With a CV of 0.92, we get 0.79. I would, thus, like to know what your CV approach is and if we are doing something wrong in the cross validation? ",
      "votes": 31
    },
    {
      "id": 593276,
      "postDate": "2019-08-06T11:52:12.697Z",
      "content": "<p>Not sure how reliable is LB as well. In my opinion it is not too indicative. So we can (probably) expect quite serious shakeup :D </p>",
      "rawMarkdown": "Not sure how reliable is LB as well. In my opinion it is not too indicative. So we can (probably) expect quite serious shakeup :D ",
      "votes": 4
    },
    {
      "id": 596394,
      "postDate": "2019-08-10T15:29:27.783Z",
      "content": "<p>It seems like the test set has a very different distribution than the training set. Could it be the case that the test set has more worse cases (3's and 4's) than the training set, which consists mostly of 0's and 2's? </p>\n\n<p>For me the score on LB is consistently ± 0.15 lower than local validation score. 😬 </p>",
      "rawMarkdown": "It seems like the test set has a very different distribution than the training set. Could it be the case that the test set has more worse cases (3's and 4's) than the training set, which consists mostly of 0's and 2's? \n\nFor me the score on LB is consistently ± 0.15 lower than local validation score. 😬 ",
      "votes": 1,
      "replies": [
        {
          "id": 596433,
          "postDate": "2019-08-10T16:24:19.603Z",
          "content": "<p>You could check to see ur models predicted distribution by changing ur lb score based on conditions</p>",
          "rawMarkdown": "You could check to see ur models predicted distribution by changing ur lb score based on conditions",
          "votes": 2
        }
      ]
    },
    {
      "id": 595171,
      "postDate": "2019-08-08T23:42:32.527Z",
      "content": "<p>My hypothesis:\n- Our models do really well classifying class 0\n-  Train data set has a higher proportion of class 0 data than test data set</p>\n\n<p>Therefore CV is consistently better than LB</p>",
      "rawMarkdown": "My hypothesis:\n- Our models do really well classifying class 0\n-  Train data set has a higher proportion of class 0 data than test data set\n\nTherefore CV is consistently better than LB",
      "votes": 1
    },
    {
      "id": 594703,
      "postDate": "2019-08-08T10:03:33.013Z",
      "content": "<p>Talk at kaggle days on <code>Building the K-fold cross validation strategy</code> by Dmytro <a href=\"https://www.youtube.com/watch?v=HfQjDfJ44Fs&amp;t=34s\">https://www.youtube.com/watch?v=HfQjDfJ44Fs&amp;t=34s</a> . Might trigger some ideas.</p>\n\n<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\">https://www.kaggle.com/dmytropoplavskiy</a></p>",
      "rawMarkdown": "Talk at kaggle days on `Building the K-fold cross validation strategy` by Dmytro https://www.youtube.com/watch?v=HfQjDfJ44Fs&amp;t=34s . Might trigger some ideas.\n\nhttps://www.kaggle.com/dmytropoplavskiy",
      "votes": 1
    },
    {
      "id": 593002,
      "postDate": "2019-08-06T05:14:47.210Z",
      "content": "<p>I've just been using a specific seed(for val split) that has an ok correlation between cv and lb. No kfold cv yet</p>",
      "rawMarkdown": "I've just been using a specific seed(for val split) that has an ok correlation between cv and lb. No kfold cv yet",
      "votes": 1,
      "replies": [
        {
          "id": 593134,
          "postDate": "2019-08-06T07:58:50.683Z",
          "content": "<p>This is it, a true kaggler always tunes his random seed.</p>",
          "rawMarkdown": "This is it, a true kaggler always tunes his random seed.",
          "votes": 18
        },
        {
          "id": 594595,
          "postDate": "2019-08-08T07:42:13.537Z",
          "content": "<p><a href=\"/sidhanthholalkere\">@sidhanthholalkere</a> hehe okay, do you mind sharing your <code>seed</code> number?</p>",
          "rawMarkdown": "@sidhanthholalkere hehe okay, do you mind sharing your `seed` number?"
        },
        {
          "id": 595046,
          "postDate": "2019-08-08T19:30:16.897Z",
          "content": "<p>use  <strong>42</strong>. <a href=\"https://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy\">42</a> is the answer to all your cv and lb problems =)  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F69f6324e3cee76b4a665b17ddb7765fa%2F49141_0.jpg?generation=1565292608513763&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "use  **42**. [42](https://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy) is the answer to all your cv and lb problems =)  \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F69f6324e3cee76b4a665b17ddb7765fa%2F49141_0.jpg?generation=1565292608513763&amp;alt=media)\n",
          "votes": 9
        },
        {
          "id": 596544,
          "postDate": "2019-08-10T20:12:39.387Z",
          "content": "<p>Ah thank you! That is probably the problem. I've been using 1234. 😁 </p>",
          "rawMarkdown": "Ah thank you! That is probably the problem. I've been using 1234. 😁 ",
          "votes": 1
        },
        {
          "id": 596965,
          "postDate": "2019-08-11T15:36:36.253Z",
          "content": "<p>thanks</p>",
          "rawMarkdown": "thanks",
          "votes": 1
        }
      ]
    },
    {
      "id": 592581,
      "postDate": "2019-08-05T14:02:33.367Z",
      "content": "<p>same with you.</p>",
      "rawMarkdown": "same with you.",
      "votes": 1
    },
    {
      "id": 594047,
      "postDate": "2019-08-07T13:45:42.170Z",
      "content": "<p>Same here, sometimes the scores are anti-correlated. I think this will be a battle of who can make their model generalize best.</p>",
      "rawMarkdown": "Same here, sometimes the scores are anti-correlated. I think this will be a battle of who can make their model generalize best."
    },
    {
      "id": 593271,
      "postDate": "2019-08-06T11:48:04.563Z",
      "content": "<p>CV doesn't work for me. Looks like properly built validation set shows better correlation.</p>",
      "rawMarkdown": "CV doesn't work for me. Looks like properly built validation set shows better correlation."
    },
    {
      "id": 593175,
      "postDate": "2019-08-06T08:51:09.847Z",
      "content": "<p>From what i understand, it seems to be correlated to the test set sizes images too. A usual preprocessing seems to be insufficient. I didn't pay too much attention on it tho</p>",
      "rawMarkdown": "From what i understand, it seems to be correlated to the test set sizes images too. A usual preprocessing seems to be insufficient. I didn't pay too much attention on it tho"
    },
    {
      "id": 592569,
      "postDate": "2019-08-05T13:42:32.197Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 593276,
      "author_name": "Georgi Pamukov",
      "author_url": "",
      "post_date": "2019-08-06T11:52:12.697000",
      "content": "<p>Not sure how reliable is LB as well. In my opinion it is not too indicative. So we can (probably) expect quite serious shakeup :D </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 596394,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2019-08-10T15:29:27.783000",
      "content": "<p>It seems like the test set has a very different distribution than the training set. Could it be the case that the test set has more worse cases (3's and 4's) than the training set, which consists mostly of 0's and 2's? </p>\n\n<p>For me the score on LB is consistently ± 0.15 lower than local validation score. 😬 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 596433,
          "author_name": "sh",
          "author_url": "",
          "post_date": "2019-08-10T16:24:19.603000",
          "content": "<p>You could check to see ur models predicted distribution by changing ur lb score based on conditions</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 595171,
      "author_name": "S D",
      "author_url": "",
      "post_date": "2019-08-08T23:42:32.527000",
      "content": "<p>My hypothesis:\n- Our models do really well classifying class 0\n-  Train data set has a higher proportion of class 0 data than test data set</p>\n\n<p>Therefore CV is consistently better than LB</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 594703,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "2019-08-08T10:03:33.013000",
      "content": "<p>Talk at kaggle days on <code>Building the K-fold cross validation strategy</code> by Dmytro <a href=\"https://www.youtube.com/watch?v=HfQjDfJ44Fs&amp;t=34s\">https://www.youtube.com/watch?v=HfQjDfJ44Fs&amp;t=34s</a> . Might trigger some ideas.</p>\n\n<p><a href=\"https://www.kaggle.com/dmytropoplavskiy\">https://www.kaggle.com/dmytropoplavskiy</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 593002,
      "author_name": "sh",
      "author_url": "",
      "post_date": "2019-08-06T05:14:47.210000",
      "content": "<p>I've just been using a specific seed(for val split) that has an ok correlation between cv and lb. No kfold cv yet</p>",
      "votes": 1,
      "replies": [
        {
          "id": 593134,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2019-08-06T07:58:50.683000",
          "content": "<p>This is it, a true kaggler always tunes his random seed.</p>",
          "votes": 18,
          "replies": []
        },
        {
          "id": 594595,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-08-08T07:42:13.537000",
          "content": "<p><a href=\"/sidhanthholalkere\">@sidhanthholalkere</a> hehe okay, do you mind sharing your <code>seed</code> number?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 595046,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-08-08T19:30:16.897000",
          "content": "<p>use  <strong>42</strong>. <a href=\"https://en.wikipedia.org/wiki/Phrases_from_The_Hitchhiker%27s_Guide_to_the_Galaxy\">42</a> is the answer to all your cv and lb problems =)  </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F69f6324e3cee76b4a665b17ddb7765fa%2F49141_0.jpg?generation=1565292608513763&amp;alt=media\" alt=\"\"></p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 596544,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-10T20:12:39.387000",
          "content": "<p>Ah thank you! That is probably the problem. I've been using 1234. 😁 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 596965,
          "author_name": "Deepak Dhaka",
          "author_url": "",
          "post_date": "2019-08-11T15:36:36.253000",
          "content": "<p>thanks</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 592581,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-08-05T14:02:33.367000",
      "content": "<p>same with you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 594047,
      "author_name": "shatz",
      "author_url": "",
      "post_date": "2019-08-07T13:45:42.170000",
      "content": "<p>Same here, sometimes the scores are anti-correlated. I think this will be a battle of who can make their model generalize best.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 593271,
      "author_name": "Georgi Pamukov",
      "author_url": "",
      "post_date": "2019-08-06T11:48:04.563000",
      "content": "<p>CV doesn't work for me. Looks like properly built validation set shows better correlation.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 593175,
      "author_name": "Samuel Amoyal",
      "author_url": "",
      "post_date": "2019-08-06T08:51:09.847000",
      "content": "<p>From what i understand, it seems to be correlated to the test set sizes images too. A usual preprocessing seems to be insufficient. I didn't pay too much attention on it tho</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 592569,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-08-05T13:42:32.197000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "592484": "So far, we have been doing kfolds in a stratified manner. However, it seems CV and LB vary a lot. With a CV of 0.92, we get 0.79. I would, thus, like to know what your CV approach is and if we are doing something wrong in the cross validation? ",
    "593276": "Not sure how reliable is LB as well. In my opinion it is not too indicative. So we can (probably) expect quite serious shakeup :D ",
    "596394": "It seems like the test set has a very different distribution than the training set. Could it be the case that the test set has more worse cases (3's and 4's) than the training set, which consists mostly of 0's and 2's? \n\nFor me the score on LB is consistently ± 0.15 lower than local validation score. 😬 ",
    "595171": "My hypothesis:\n- Our models do really well classifying class 0\n-  Train data set has a higher proportion of class 0 data than test data set\n\nTherefore CV is consistently better than LB",
    "594703": "Talk at kaggle days on `Building the K-fold cross validation strategy` by Dmytro https://www.youtube.com/watch?v=HfQjDfJ44Fs&amp;t=34s . Might trigger some ideas.\n\nhttps://www.kaggle.com/dmytropoplavskiy",
    "593002": "I've just been using a specific seed(for val split) that has an ok correlation between cv and lb. No kfold cv yet",
    "592581": "same with you.",
    "594047": "Same here, sometimes the scores are anti-correlated. I think this will be a battle of who can make their model generalize best.",
    "593271": "CV doesn't work for me. Looks like properly built validation set shows better correlation.",
    "593175": "From what i understand, it seems to be correlated to the test set sizes images too. A usual preprocessing seems to be insufficient. I didn't pay too much attention on it tho",
    "592569": ""
  }
}