{
  "id": 346024,
  "title": "Purpose of validation data",
  "url": "/competitions/amex-default-prediction/discussion/346024",
  "author_name": "",
  "post_date": "2022-08-17T15:24:23.853108400Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Excuse me for asking like a newbie.</p>\n<p>I believe the purpose of cross-validation is to know without having to submit CV scores. On the other hand, I think there are quite a few notebooks that also set validation data when modeling for submission. (Like as follows)</p>\n<p>**    bst = xgb.train(params, dtrain=dtrain,<br>\n                num_boost_round=20500,evals=watchlist,<br>\n                early_stopping_rounds=500, feval=xgb_amex, maximize=True,<br>\n                verbose_eval=100) **<br>\n(cited from <a href=\"https://www.kaggle.com/code/zb1373/xgb-lgbm-catboost-cnn-stacking\" target=\"_blank\">https://www.kaggle.com/code/zb1373/xgb-lgbm-catboost-cnn-stacking</a>)</p>\n<p>My personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?</p>",
  "messages": [
    {
      "id": "1903693",
      "postDate": "08/17/2022 15:24:23",
      "content": "<p>Excuse me for asking like a newbie.</p>\n<p>I believe the purpose of cross-validation is to know without having to submit CV scores. On the other hand, I think there are quite a few notebooks that also set validation data when modeling for submission. (Like as follows)</p>\n<p>**    bst = xgb.train(params, dtrain=dtrain,<br>\n                num_boost_round=20500,evals=watchlist,<br>\n                early_stopping_rounds=500, feval=xgb_amex, maximize=True,<br>\n                verbose_eval=100) **<br>\n(cited from <a href=\"https://www.kaggle.com/code/zb1373/xgb-lgbm-catboost-cnn-stacking\" target=\"_blank\">https://www.kaggle.com/code/zb1373/xgb-lgbm-catboost-cnn-stacking</a>)</p>\n<p>My personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?</p>",
      "rawMarkdown": "Excuse me for asking like a newbie.\n\nI believe the purpose of cross-validation is to know without having to submit CV scores. On the other hand, I think there are quite a few notebooks that also set validation data when modeling for submission. (Like as follows)\n\n**    bst = xgb.train(params, dtrain=dtrain,\n                num_boost_round=20500,evals=watchlist,\n                early_stopping_rounds=500, feval=xgb_amex, maximize=True,\n                verbose_eval=100) **\n(cited from https://www.kaggle.com/code/zb1373/xgb-lgbm-catboost-cnn-stacking)\n\nMy personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?",
      "votes": null
    },
    {
      "id": "1903722",
      "postDate": "08/17/2022 15:44:49",
      "content": "<p>Usually your point is fine, but one may want to augment ones score in this competition leading to this approach. In the industry, one may leave out the out of sample test set in totality during model training.</p>",
      "rawMarkdown": "Usually your point is fine, but one may want to augment ones score in this competition leading to this approach. In the industry, one may leave out the out of sample test set in totality during model training.",
      "votes": null
    },
    {
      "id": "1903987",
      "postDate": "08/17/2022 20:28:36",
      "content": "<p>As I know, K fold cross validation allows all dataset to be used to train, and validate. <br>\nFor sub model, K-1 folds was used to train, 1 fold was used to validate.<br>\nMaybe you refer to train-test split. For this case, yes, you are right. But you dont know when the model overfitted without checking validation.</p>",
      "rawMarkdown": "As I know, K fold cross validation allows all dataset to be used to train, and validate. \nFor sub model, K-1 folds was used to train, 1 fold was used to validate.\nMaybe you refer to train-test split. For this case, yes, you are right. But you dont know when the model overfitted without checking validation.",
      "votes": null
    },
    {
      "id": "1906683",
      "postDate": "08/20/2022 06:03:53",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/meisa0\" target=\"_blank\">@meisa0</a> </p>\n<p>In the above code snippet it seems that early stopping is used. In order for that to work then the estimator requires a chunk of validation data to be set aside so as to monitor progress. I suggest that this is part of the hyperparameter selection process; once one knows roughly how many boosting rounds to use (it may vary somewhat over folds), this can then be 'hard coded' and the early stopping removed. I agree with you in that once one has a set of optimal hyperparameters, then one can proceed to setting those hyperparameters, and re-running the estimator on the whole dataset.</p>\n<p>Note however that in advanced notebooks, people <em>still</em> perform cross-validation even after having found the optimal hyperparameters. This is in order to save the OOF predictions produced by the estimator for the test data, and those predictions are saved and go on to become a feature column in a stacking ensemble.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @meisa0 \n\nIn the above code snippet it seems that early stopping is used. In order for that to work then the estimator requires a chunk of validation data to be set aside so as to monitor progress. I suggest that this is part of the hyperparameter selection process; once one knows roughly how many boosting rounds to use (it may vary somewhat over folds), this can then be 'hard coded' and the early stopping removed. I agree with you in that once one has a set of optimal hyperparameters, then one can proceed to setting those hyperparameters, and re-running the estimator on the whole dataset.\n\nNote however that in advanced notebooks, people *still* perform cross-validation even after having found the optimal hyperparameters. This is in order to save the OOF predictions produced by the estimator for the test data, and those predictions are saved and go on to become a feature column in a stacking ensemble.\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1908941",
      "postDate": "08/22/2022 06:42:09",
      "content": "<p>\"My personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?\"</p>\n<p>I would like to disagree somewhat, because in general if you have more training data it would not necessarily imply that more patterns are picked up, thus it can lead to underestimation of the parameters involved. Validation data in such cases becomes useful when the training data is large yet it does not pick up any unforeseen pattern. <br>\nOne can of-course tweak the model after checking the accuracy w.r.t the validation data and find any flaws. <br>\nIf there is a very large training set and no validation set one has effectively lost control over his model. </p>",
      "rawMarkdown": "\"My personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?\"\n\nI would like to disagree somewhat, because in general if you have more training data it would not necessarily imply that more patterns are picked up, thus it can lead to underestimation of the parameters involved. Validation data in such cases becomes useful when the training data is large yet it does not pick up any unforeseen pattern. \nOne can of-course tweak the model after checking the accuracy w.r.t the validation data and find any flaws. \nIf there is a very large training set and no validation set one has effectively lost control over his model.",
      "votes": null
    },
    {
      "id": "1908953",
      "postDate": "08/22/2022 06:54:13",
      "content": "<p>Totally agree with you <a href=\"https://www.kaggle.com/amoghbajpai\" target=\"_blank\">@amoghbajpai</a> <br>\nBut this is only when u have a large dataset </p>",
      "rawMarkdown": "Totally agree with you @amoghbajpai \nBut this is only when u have a large dataset",
      "votes": null
    },
    {
      "id": "1911892",
      "postDate": "08/24/2022 11:28:36",
      "content": "<p>Thank you for your comment. My understanding is as follows:<br>\n\"If there is enough data and all patterns are considered complete, the disadvantage of a possible shift in hyperparameters outweighs the advantage of slightly more data\"</p>",
      "rawMarkdown": "Thank you for your comment. My understanding is as follows:\n\"If there is enough data and all patterns are considered complete, the disadvantage of a possible shift in hyperparameters outweighs the advantage of slightly more data\"",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1903722,
      "author_name": "ravi20076",
      "author_url": "",
      "post_date": "08/17/2022 15:44:49",
      "content": "<p>Usually your point is fine, but one may want to augment ones score in this competition leading to this approach. In the industry, one may leave out the out of sample test set in totality during model training.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1903987,
      "author_name": "zhehaoliang",
      "author_url": "",
      "post_date": "08/17/2022 20:28:36",
      "content": "<p>As I know, K fold cross validation allows all dataset to be used to train, and validate. <br>\nFor sub model, K-1 folds was used to train, 1 fold was used to validate.<br>\nMaybe you refer to train-test split. For this case, yes, you are right. But you dont know when the model overfitted without checking validation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1906683,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "08/20/2022 06:03:53",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/meisa0\" target=\"_blank\">@meisa0</a> </p>\n<p>In the above code snippet it seems that early stopping is used. In order for that to work then the estimator requires a chunk of validation data to be set aside so as to monitor progress. I suggest that this is part of the hyperparameter selection process; once one knows roughly how many boosting rounds to use (it may vary somewhat over folds), this can then be 'hard coded' and the early stopping removed. I agree with you in that once one has a set of optimal hyperparameters, then one can proceed to setting those hyperparameters, and re-running the estimator on the whole dataset.</p>\n<p>Note however that in advanced notebooks, people <em>still</em> perform cross-validation even after having found the optimal hyperparameters. This is in order to save the OOF predictions produced by the estimator for the test data, and those predictions are saved and go on to become a feature column in a stacking ensemble.</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1908941,
      "author_name": "amoghbajpai",
      "author_url": "",
      "post_date": "08/22/2022 06:42:09",
      "content": "<p>\"My personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?\"</p>\n<p>I would like to disagree somewhat, because in general if you have more training data it would not necessarily imply that more patterns are picked up, thus it can lead to underestimation of the parameters involved. Validation data in such cases becomes useful when the training data is large yet it does not pick up any unforeseen pattern. <br>\nOne can of-course tweak the model after checking the accuracy w.r.t the validation data and find any flaws. <br>\nIf there is a very large training set and no validation set one has effectively lost control over his model. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1908953,
          "author_name": "subhamjain",
          "author_url": "",
          "post_date": "08/22/2022 06:54:13",
          "content": "<p>Totally agree with you <a href=\"https://www.kaggle.com/amoghbajpai\" target=\"_blank\">@amoghbajpai</a> <br>\nBut this is only when u have a large dataset </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1911892,
          "author_name": "meisa0",
          "author_url": "",
          "post_date": "08/24/2022 11:28:36",
          "content": "<p>Thank you for your comment. My understanding is as follows:<br>\n\"If there is enough data and all patterns are considered complete, the disadvantage of a possible shift in hyperparameters outweighs the advantage of slightly more data\"</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1903693": "Excuse me for asking like a newbie.\n\nI believe the purpose of cross-validation is to know without having to submit CV scores. On the other hand, I think there are quite a few notebooks that also set validation data when modeling for submission. (Like as follows)\n\n**    bst = xgb.train(params, dtrain=dtrain,\n                num_boost_round=20500,evals=watchlist,\n                early_stopping_rounds=500, feval=xgb_amex, maximize=True,\n                verbose_eval=100) **\n(cited from https://www.kaggle.com/code/zb1373/xgb-lgbm-catboost-cnn-stacking)\n\nMy personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?",
    "1903722": "Usually your point is fine, but one may want to augment ones score in this competition leading to this approach. In the industry, one may leave out the out of sample test set in totality during model training.",
    "1903987": "As I know, K fold cross validation allows all dataset to be used to train, and validate. \nFor sub model, K-1 folds was used to train, 1 fold was used to validate.\nMaybe you refer to train-test split. For this case, yes, you are right. But you dont know when the model overfitted without checking validation.",
    "1906683": "Dear @meisa0 \n\nIn the above code snippet it seems that early stopping is used. In order for that to work then the estimator requires a chunk of validation data to be set aside so as to monitor progress. I suggest that this is part of the hyperparameter selection process; once one knows roughly how many boosting rounds to use (it may vary somewhat over folds), this can then be 'hard coded' and the early stopping removed. I agree with you in that once one has a set of optimal hyperparameters, then one can proceed to setting those hyperparameters, and re-running the estimator on the whole dataset.\n\nNote however that in advanced notebooks, people *still* perform cross-validation even after having found the optimal hyperparameters. This is in order to save the OOF predictions produced by the estimator for the test data, and those predictions are saved and go on to become a feature column in a stacking ensemble.\n\nAll the best,\ncarl",
    "1908941": "\"My personal opinion is that in general, the more teacher data you have, the better the accuracy, so I think it is better not to set validation data in the actual model to be submitted, but can someone please tell me if my idea is good or not?\"\n\nI would like to disagree somewhat, because in general if you have more training data it would not necessarily imply that more patterns are picked up, thus it can lead to underestimation of the parameters involved. Validation data in such cases becomes useful when the training data is large yet it does not pick up any unforeseen pattern. \nOne can of-course tweak the model after checking the accuracy w.r.t the validation data and find any flaws. \nIf there is a very large training set and no validation set one has effectively lost control over his model.",
    "1908953": "Totally agree with you @amoghbajpai \nBut this is only when u have a large dataset",
    "1911892": "Thank you for your comment. My understanding is as follows:\n\"If there is enough data and all patterns are considered complete, the disadvantage of a possible shift in hyperparameters outweighs the advantage of slightly more data\""
  },
  "source": "meta"
}