{
  "id": 75310,
  "title": "About cross validation",
  "url": "/competitions/quora-insincere-questions-classification/discussion/75310",
  "author_name": "Bai",
  "post_date": "2018-12-20T13:01:26.648000",
  "votes": 3,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I saw a lot of people using cross-validation, but another set of validation sets, maybe 0.1, maybe 0.04. This surprised me. I never found such a division in the knowledge I learned. Maybe they want to simulate a test set, but this is not accurate. Because the data distribution of the test set is unknown.</p>\n\n<p>So, I want to send some knowledge about cross-validation to everyone.</p>\n\n<p>The method we often use is called K-fold cross-validation.</p>\n\n<p>Assume K=5，our steps are as follows：\n1、Divide the data into five；\n2、Not repeat one of the test sets each time and use the other four to do the training set training model，then calculate the predicted score of the model on the test set.</p>\n\n<p>You can use the best one of the models to predict the test score，you can also average all the predicted scores for the test data. </p>\n\n<p>You can use the best model to pick the threshold, or you can use all the training data to pick the threshold. But don't set up another validation set because the validation set already exists. This is meaningless.</p>\n\n<p>If I have any mistakes, welcome everyone to criticize me.</p>",
  "messages": [
    {
      "id": 442757,
      "postDate": "2018-12-20T13:01:26.650Z",
      "content": "<p>I saw a lot of people using cross-validation, but another set of validation sets, maybe 0.1, maybe 0.04. This surprised me. I never found such a division in the knowledge I learned. Maybe they want to simulate a test set, but this is not accurate. Because the data distribution of the test set is unknown.</p>\n\n<p>So, I want to send some knowledge about cross-validation to everyone.</p>\n\n<p>The method we often use is called K-fold cross-validation.</p>\n\n<p>Assume K=5，our steps are as follows：\n1、Divide the data into five；\n2、Not repeat one of the test sets each time and use the other four to do the training set training model，then calculate the predicted score of the model on the test set.</p>\n\n<p>You can use the best one of the models to predict the test score，you can also average all the predicted scores for the test data. </p>\n\n<p>You can use the best model to pick the threshold, or you can use all the training data to pick the threshold. But don't set up another validation set because the validation set already exists. This is meaningless.</p>\n\n<p>If I have any mistakes, welcome everyone to criticize me.</p>",
      "rawMarkdown": "I saw a lot of people using cross-validation, but another set of validation sets, maybe 0.1, maybe 0.04. This surprised me. I never found such a division in the knowledge I learned. Maybe they want to simulate a test set, but this is not accurate. Because the data distribution of the test set is unknown.\n\nSo, I want to send some knowledge about cross-validation to everyone.\n\nThe method we often use is called K-fold cross-validation.\n\nAssume K=5，our steps are as follows：\n1、Divide the data into five；\n2、Not repeat one of the test sets each time and use the other four to do the training set training model，then calculate the predicted score of the model on the test set.\n\nYou can use the best one of the models to predict the test score，you can also average all the predicted scores for the test data. \n\nYou can use the best model to pick the threshold, or you can use all the training data to pick the threshold. But don't set up another validation set because the validation set already exists. This is meaningless.\n\nIf I have any mistakes, welcome everyone to criticize me.",
      "votes": 3
    },
    {
      "id": 445994,
      "postDate": "2018-12-27T10:26:06.353Z",
      "content": "<p>If your model takes some hours to train, you would need to find other strategies, given that it is a Kernel competition.</p>",
      "rawMarkdown": "If your model takes some hours to train, you would need to find other strategies, given that it is a Kernel competition."
    },
    {
      "id": 445248,
      "postDate": "2018-12-26T03:52:15.833Z",
      "content": "<p>k-fold validation is good, but local cv  show negative relationship with LB, some kaggler  found  this division(with k-fold train) can reflect public lb </p>",
      "rawMarkdown": "k-fold validation is good, but local cv  show negative relationship with LB, some kaggler  found  this division(with k-fold train) can reflect public lb ",
      "replies": [
        {
          "id": 445256,
          "postDate": "2018-12-26T04:07:59.593Z",
          "content": "<p><a href=\"/luckyboyde\">@luckyboyde</a>\nbut it reduces the number of training data.</p>",
          "rawMarkdown": "@luckyboyde\nbut it reduces the number of training data.",
          "votes": 1
        },
        {
          "id": 445275,
          "postDate": "2018-12-26T05:29:21.687Z",
          "content": "<p>but it is the reason they do that, quora competition did not have a convinced metrics up to now, both local cv and public lb</p>",
          "rawMarkdown": "but it is the reason they do that, quora competition did not have a convinced metrics up to now, both local cv and public lb"
        },
        {
          "id": 445732,
          "postDate": "2018-12-27T01:59:29.747Z",
          "content": "<p><a href=\"/luckyboyde\">@luckyboyde</a>\nSorry, I don't know what you mean.</p>",
          "rawMarkdown": "@luckyboyde\nSorry, I don't know what you mean."
        },
        {
          "id": 445737,
          "postDate": "2018-12-27T02:18:02.323Z",
          "content": "<p><a href=\"/luckyboyde\">@luckyboyde</a>\nI'm not really convinced that this method is resulting in convincing validation scores. From what I've read thus far, a lot of people using this method are getting optimistic public scores (~0.700), but their local f1 score is like 0.675. That's hard for me to stomach in terms of the delta.</p>\n\n<p>With that said, I agree with @Bai. I don't think holding out 10% of the data prior to performing k-fold makes sense to me at all. </p>",
          "rawMarkdown": "@luckyboyde\nI'm not really convinced that this method is resulting in convincing validation scores. From what I've read thus far, a lot of people using this method are getting optimistic public scores (~0.700), but their local f1 score is like 0.675. That's hard for me to stomach in terms of the delta.\n\nWith that said, I agree with @Bai. I don't think holding out 10% of the data prior to performing k-fold makes sense to me at all. "
        },
        {
          "id": 445925,
          "postDate": "2018-12-27T08:23:12.417Z",
          "content": "<p>Oh, actually i agree with @Bai, i just account for the reason kagglers they do that. </p>",
          "rawMarkdown": "Oh, actually i agree with @Bai, i just account for the reason kagglers they do that. "
        }
      ]
    },
    {
      "id": 443255,
      "postDate": "2018-12-21T09:57:15.383Z",
      "content": "<p>I totally agree with your point bai. You cant use a separate validation set. The whole point of validation is to generalise your model to give you best results. You cant create your validation set. Use something which is already there like create a split on your train data.</p>",
      "rawMarkdown": "I totally agree with your point bai. You cant use a separate validation set. The whole point of validation is to generalise your model to give you best results. You cant create your validation set. Use something which is already there like create a split on your train data."
    },
    {
      "id": 443678,
      "postDate": "2018-12-22T05:56:36.023Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 445994,
      "author_name": "ManuelSH",
      "author_url": "",
      "post_date": "2018-12-27T10:26:06.353000",
      "content": "<p>If your model takes some hours to train, you would need to find other strategies, given that it is a Kernel competition.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 445248,
      "author_name": "hey!2019",
      "author_url": "",
      "post_date": "2018-12-26T03:52:15.833000",
      "content": "<p>k-fold validation is good, but local cv  show negative relationship with LB, some kaggler  found  this division(with k-fold train) can reflect public lb </p>",
      "votes": 0,
      "replies": [
        {
          "id": 445256,
          "author_name": "Bai",
          "author_url": "",
          "post_date": "2018-12-26T04:07:59.593000",
          "content": "<p><a href=\"/luckyboyde\">@luckyboyde</a>\nbut it reduces the number of training data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 445275,
          "author_name": "hey!2019",
          "author_url": "",
          "post_date": "2018-12-26T05:29:21.687000",
          "content": "<p>but it is the reason they do that, quora competition did not have a convinced metrics up to now, both local cv and public lb</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445732,
          "author_name": "Bai",
          "author_url": "",
          "post_date": "2018-12-27T01:59:29.747000",
          "content": "<p><a href=\"/luckyboyde\">@luckyboyde</a>\nSorry, I don't know what you mean.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445737,
          "author_name": "Thomas Yokota",
          "author_url": "",
          "post_date": "2018-12-27T02:18:02.323000",
          "content": "<p><a href=\"/luckyboyde\">@luckyboyde</a>\nI'm not really convinced that this method is resulting in convincing validation scores. From what I've read thus far, a lot of people using this method are getting optimistic public scores (~0.700), but their local f1 score is like 0.675. That's hard for me to stomach in terms of the delta.</p>\n\n<p>With that said, I agree with @Bai. I don't think holding out 10% of the data prior to performing k-fold makes sense to me at all. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 445925,
          "author_name": "hey!2019",
          "author_url": "",
          "post_date": "2018-12-27T08:23:12.417000",
          "content": "<p>Oh, actually i agree with @Bai, i just account for the reason kagglers they do that. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 443255,
      "author_name": "Raghav Agrawal",
      "author_url": "",
      "post_date": "2018-12-21T09:57:15.383000",
      "content": "<p>I totally agree with your point bai. You cant use a separate validation set. The whole point of validation is to generalise your model to give you best results. You cant create your validation set. Use something which is already there like create a split on your train data.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 443678,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-12-22T05:56:36.023000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "442757": "I saw a lot of people using cross-validation, but another set of validation sets, maybe 0.1, maybe 0.04. This surprised me. I never found such a division in the knowledge I learned. Maybe they want to simulate a test set, but this is not accurate. Because the data distribution of the test set is unknown.\n\nSo, I want to send some knowledge about cross-validation to everyone.\n\nThe method we often use is called K-fold cross-validation.\n\nAssume K=5，our steps are as follows：\n1、Divide the data into five；\n2、Not repeat one of the test sets each time and use the other four to do the training set training model，then calculate the predicted score of the model on the test set.\n\nYou can use the best one of the models to predict the test score，you can also average all the predicted scores for the test data. \n\nYou can use the best model to pick the threshold, or you can use all the training data to pick the threshold. But don't set up another validation set because the validation set already exists. This is meaningless.\n\nIf I have any mistakes, welcome everyone to criticize me.",
    "445994": "If your model takes some hours to train, you would need to find other strategies, given that it is a Kernel competition.",
    "445248": "k-fold validation is good, but local cv  show negative relationship with LB, some kaggler  found  this division(with k-fold train) can reflect public lb ",
    "443255": "I totally agree with your point bai. You cant use a separate validation set. The whole point of validation is to generalise your model to give you best results. You cant create your validation set. Use something which is already there like create a split on your train data.",
    "443678": ""
  }
}