{
  "id": 373722,
  "title": "Do we really need to use cross-validation over holdout in large datasets to win a competition?",
  "url": "/competitions/otto-recommender-system/discussion/373722",
  "author_name": "",
  "post_date": "2022-12-23T01:48:16.944898800Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>As the dataset is larger do we really need to use cross-validation instead holdout to win a competition? </p>",
  "messages": [
    {
      "id": "2073370",
      "postDate": "12/23/2022 01:48:16",
      "content": "<p>As the dataset is larger do we really need to use cross-validation instead holdout to win a competition? </p>",
      "rawMarkdown": "As the dataset is larger do we really need to use cross-validation instead holdout to win a competition?",
      "votes": null
    },
    {
      "id": "2073402",
      "postDate": "12/23/2022 02:55:45",
      "content": "<p>In this competition, I suspect everyone is using holdout (because it is difficult to cross validate with time data and 25% holdout is already large enough with 1.8M users). Note: often on Kaggle people use the word \"CV\" to refer to both \"cross validation\" and \"holdout\" (even though it technically means cross validation). So in this competition when people say \"CV\", i believe everyone means \"holdout\".</p>\n<p>The most common holdout (in this comp) is using the last 1 week of the 4 weeks of train data to compute validation score. This is what i am doing and my validation score perfects correlates with LB. Radek posted discussion about this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "In this competition, I suspect everyone is using holdout (because it is difficult to cross validate with time data and 25% holdout is already large enough with 1.8M users). Note: often on Kaggle people use the word \"CV\" to refer to both \"cross validation\" and \"holdout\" (even though it technically means cross validation). So in this competition when people say \"CV\", i believe everyone means \"holdout\".\n\nThe most common holdout (in this comp) is using the last 1 week of the 4 weeks of train data to compute validation score. This is what i am doing and my validation score perfects correlates with LB. Radek posted discussion about this [here][1]\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991",
      "votes": null
    },
    {
      "id": "2073868",
      "postDate": "12/23/2022 13:49:10",
      "content": "<p>However Chris, as you mentioned in the training step of your thread <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">How To Build a GBT Ranker Model</a>, you are training a GroupKFold (CV) on your training dataset (which is probably the holdout dataset you just mentioned). Am I right?</p>",
      "rawMarkdown": "However Chris, as you mentioned in the training step of your thread [How To Build a GBT Ranker Model](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210), you are training a GroupKFold (CV) on your training dataset (which is probably the holdout dataset you just mentioned). Am I right?",
      "votes": null
    },
    {
      "id": "2073871",
      "postDate": "12/23/2022 13:55:39",
      "content": "<p>Great point haha, i am using both cross validation and holdout. However, for the purpose of local model evaluation, i ignore the cross validation metric score. I am only using the holdout metric score to tune and modify my model.</p>",
      "rawMarkdown": "Great point haha, i am using both cross validation and holdout. However, for the purpose of local model evaluation, i ignore the cross validation metric score. I am only using the holdout metric score to tune and modify my model.",
      "votes": null
    },
    {
      "id": "2073873",
      "postDate": "12/23/2022 14:03:36",
      "content": "<p>Hmm… now that i think about it, i don't know what it is called what i am doing. In normal holdout (or normal cross validation plus holdout), we do not use anything from holdout during training. However in this competition, i use the targets from the holdout data. So perhaps i am using cross validation and i am not using any holdout. I will need to think about this.</p>\n<p>Earlier i said i ignore cross validation metric score. But actually my cross validation metric score and \"holdout\" metric are the same. So i am using both metric scores…</p>",
      "rawMarkdown": "Hmm... now that i think about it, i don't know what it is called what i am doing. In normal holdout (or normal cross validation plus holdout), we do not use anything from holdout during training. However in this competition, i use the targets from the holdout data. So perhaps i am using cross validation and i am not using any holdout. I will need to think about this.\n\nEarlier i said i ignore cross validation metric score. But actually my cross validation metric score and \"holdout\" metric are the same. So i am using both metric scores...",
      "votes": null
    },
    {
      "id": "2076373",
      "postDate": "12/26/2022 12:22:11",
      "content": "<p>This is actually interesting, when you experiment with the local CV you measure the score of the holdout, correct? (You said you use the target of the holdout? <br>\nSo as for the GroupedKFold, do you simply use it only on the training set?</p>",
      "rawMarkdown": "This is actually interesting, when you experiment with the local CV you measure the score of the holdout, correct? (You said you use the target of the holdout? \nSo as for the GroupedKFold, do you simply use it only on the training set?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2073402,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "12/23/2022 02:55:45",
      "content": "<p>In this competition, I suspect everyone is using holdout (because it is difficult to cross validate with time data and 25% holdout is already large enough with 1.8M users). Note: often on Kaggle people use the word \"CV\" to refer to both \"cross validation\" and \"holdout\" (even though it technically means cross validation). So in this competition when people say \"CV\", i believe everyone means \"holdout\".</p>\n<p>The most common holdout (in this comp) is using the last 1 week of the 4 weeks of train data to compute validation score. This is what i am doing and my validation score perfects correlates with LB. Radek posted discussion about this <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991\" target=\"_blank\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2073868,
          "author_name": "joaopmpeinado",
          "author_url": "",
          "post_date": "12/23/2022 13:49:10",
          "content": "<p>However Chris, as you mentioned in the training step of your thread <a href=\"https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210\" target=\"_blank\">How To Build a GBT Ranker Model</a>, you are training a GroupKFold (CV) on your training dataset (which is probably the holdout dataset you just mentioned). Am I right?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2073871,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "12/23/2022 13:55:39",
              "content": "<p>Great point haha, i am using both cross validation and holdout. However, for the purpose of local model evaluation, i ignore the cross validation metric score. I am only using the holdout metric score to tune and modify my model.</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2073873,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "12/23/2022 14:03:36",
              "content": "<p>Hmm… now that i think about it, i don't know what it is called what i am doing. In normal holdout (or normal cross validation plus holdout), we do not use anything from holdout during training. However in this competition, i use the targets from the holdout data. So perhaps i am using cross validation and i am not using any holdout. I will need to think about this.</p>\n<p>Earlier i said i ignore cross validation metric score. But actually my cross validation metric score and \"holdout\" metric are the same. So i am using both metric scores…</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2076373,
                  "author_name": "thedevastator",
                  "author_url": "",
                  "post_date": "12/26/2022 12:22:11",
                  "content": "<p>This is actually interesting, when you experiment with the local CV you measure the score of the holdout, correct? (You said you use the target of the holdout? <br>\nSo as for the GroupedKFold, do you simply use it only on the training set?</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2073370": "As the dataset is larger do we really need to use cross-validation instead holdout to win a competition?",
    "2073402": "In this competition, I suspect everyone is using holdout (because it is difficult to cross validate with time data and 25% holdout is already large enough with 1.8M users). Note: often on Kaggle people use the word \"CV\" to refer to both \"cross validation\" and \"holdout\" (even though it technically means cross validation). So in this competition when people say \"CV\", i believe everyone means \"holdout\".\n\nThe most common holdout (in this comp) is using the last 1 week of the 4 weeks of train data to compute validation score. This is what i am doing and my validation score perfects correlates with LB. Radek posted discussion about this [here][1]\n\n[1]: https://www.kaggle.com/competitions/otto-recommender-system/discussion/364991",
    "2073868": "However Chris, as you mentioned in the training step of your thread [How To Build a GBT Ranker Model](https://www.kaggle.com/competitions/otto-recommender-system/discussion/370210), you are training a GroupKFold (CV) on your training dataset (which is probably the holdout dataset you just mentioned). Am I right?",
    "2073871": "Great point haha, i am using both cross validation and holdout. However, for the purpose of local model evaluation, i ignore the cross validation metric score. I am only using the holdout metric score to tune and modify my model.",
    "2073873": "Hmm... now that i think about it, i don't know what it is called what i am doing. In normal holdout (or normal cross validation plus holdout), we do not use anything from holdout during training. However in this competition, i use the targets from the holdout data. So perhaps i am using cross validation and i am not using any holdout. I will need to think about this.\n\nEarlier i said i ignore cross validation metric score. But actually my cross validation metric score and \"holdout\" metric are the same. So i am using both metric scores...",
    "2076373": "This is actually interesting, when you experiment with the local CV you measure the score of the holdout, correct? (You said you use the target of the holdout? \nSo as for the GroupedKFold, do you simply use it only on the training set?"
  },
  "source": "meta"
}