{
  "id": 77803,
  "title": "Cross Validation Strategy",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77803",
  "author_name": "",
  "post_date": "2019-01-16T19:20:28.226410Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Can someone suggest me a good cross validation strategy here for NN</p>",
  "messages": [
    {
      "id": "456945",
      "postDate": "01/16/2019 19:20:28",
      "content": "<p>Can someone suggest me a good cross validation strategy here for NN</p>",
      "rawMarkdown": "Can someone suggest me a good cross validation strategy here for NN",
      "votes": null
    },
    {
      "id": "457120",
      "postDate": "01/17/2019 01:17:02",
      "content": "<p>It's pretty frustrating to find that CV and LB scores usually go opposite directions, which makes CV score not a very good indicator for LB (or maybe I am doing something wrong...). It would be very helpful if fellow kagglers could share about their CV strategies and how to overcome the inconsistency between CV and LB scores. </p>\n\n<p>Should we adjust our CV method to make it more closely reflect the LB score ? \nOr maybe we should just trust CV and believe that a higher CV will lead to a higher position on the private LB ?</p>",
      "rawMarkdown": "It's pretty frustrating to find that CV and LB scores usually go opposite directions, which makes CV score not a very good indicator for LB (or maybe I am doing something wrong...). It would be very helpful if fellow kagglers could share about their CV strategies and how to overcome the inconsistency between CV and LB scores. \n\nShould we adjust our CV method to make it more closely reflect the LB score ? \nOr maybe we should just trust CV and believe that a higher CV will lead to a higher position on the private LB ?",
      "votes": null
    },
    {
      "id": "457289",
      "postDate": "01/17/2019 07:29:19",
      "content": "<p>The idea of cross validation is having a hold out(unseen) set for your model to learn on seen data and test it then and there itself. So that it performs good on new(Test) set when it arrives.\nI think trusting our <code>CV</code> in hope that higher the CV, higher will be the private LB is a good strategy.</p>\n\n<p>As this is a huge dataset and we will be given 2 hours time in the second stage, i am concerned how cross validation strategy should be.</p>",
      "rawMarkdown": "The idea of cross validation is having a hold out(unseen) set for your model to learn on seen data and test it then and there itself. So that it performs good on new(Test) set when it arrives.\nI think trusting our `CV` in hope that higher the CV, higher will be the private LB is a good strategy.\n\nAs this is a huge dataset and we will be given 2 hours time in the second stage, i am concerned how cross validation strategy should be.",
      "votes": null
    },
    {
      "id": "457394",
      "postDate": "01/17/2019 10:45:53",
      "content": "<p>@rbakshee Have a look here <a href=\"https://www.kaggle.com/c/home-depot-product-search-relevance/discussion/19648\">https://www.kaggle.com/c/home-depot-product-search-relevance/discussion/19648</a>\nand <a href=\"https://www.kaggle.com/c/telstra-recruiting-network/discussion/19277\">https://www.kaggle.com/c/telstra-recruiting-network/discussion/19277</a> and <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44593\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44593</a>. I am sure if you go through the entire threads you will find interesting points as some are by grandmasters.</p>",
      "rawMarkdown": "rbakshee Have a look here https://www.kaggle.com/c/home-depot-product-search-relevance/discussion/19648\nand https://www.kaggle.com/c/telstra-recruiting-network/discussion/19277 and https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44593. I am sure if you go through the entire threads you will find interesting points as some are by grandmasters.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 457120,
      "author_name": "m7catsue",
      "author_url": "",
      "post_date": "01/17/2019 01:17:02",
      "content": "<p>It's pretty frustrating to find that CV and LB scores usually go opposite directions, which makes CV score not a very good indicator for LB (or maybe I am doing something wrong...). It would be very helpful if fellow kagglers could share about their CV strategies and how to overcome the inconsistency between CV and LB scores. </p>\n\n<p>Should we adjust our CV method to make it more closely reflect the LB score ? \nOr maybe we should just trust CV and believe that a higher CV will lead to a higher position on the private LB ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457289,
      "author_name": "rahulbakshee",
      "author_url": "",
      "post_date": "01/17/2019 07:29:19",
      "content": "<p>The idea of cross validation is having a hold out(unseen) set for your model to learn on seen data and test it then and there itself. So that it performs good on new(Test) set when it arrives.\nI think trusting our <code>CV</code> in hope that higher the CV, higher will be the private LB is a good strategy.</p>\n\n<p>As this is a huge dataset and we will be given 2 hours time in the second stage, i am concerned how cross validation strategy should be.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 457394,
      "author_name": "cyberia",
      "author_url": "",
      "post_date": "01/17/2019 10:45:53",
      "content": "<p>@rbakshee Have a look here <a href=\"https://www.kaggle.com/c/home-depot-product-search-relevance/discussion/19648\">https://www.kaggle.com/c/home-depot-product-search-relevance/discussion/19648</a>\nand <a href=\"https://www.kaggle.com/c/telstra-recruiting-network/discussion/19277\">https://www.kaggle.com/c/telstra-recruiting-network/discussion/19277</a> and <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44593\">https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44593</a>. I am sure if you go through the entire threads you will find interesting points as some are by grandmasters.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "456945": "Can someone suggest me a good cross validation strategy here for NN",
    "457120": "It's pretty frustrating to find that CV and LB scores usually go opposite directions, which makes CV score not a very good indicator for LB (or maybe I am doing something wrong...). It would be very helpful if fellow kagglers could share about their CV strategies and how to overcome the inconsistency between CV and LB scores. \n\nShould we adjust our CV method to make it more closely reflect the LB score ? \nOr maybe we should just trust CV and believe that a higher CV will lead to a higher position on the private LB ?",
    "457289": "The idea of cross validation is having a hold out(unseen) set for your model to learn on seen data and test it then and there itself. So that it performs good on new(Test) set when it arrives.\nI think trusting our `CV` in hope that higher the CV, higher will be the private LB is a good strategy.\n\nAs this is a huge dataset and we will be given 2 hours time in the second stage, i am concerned how cross validation strategy should be.",
    "457394": "rbakshee Have a look here https://www.kaggle.com/c/home-depot-product-search-relevance/discussion/19648\nand https://www.kaggle.com/c/telstra-recruiting-network/discussion/19277 and https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/44593. I am sure if you go through the entire threads you will find interesting points as some are by grandmasters."
  },
  "source": "meta"
}