{
  "id": 53458,
  "title": "Which validation set to use?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53458",
  "author_name": "",
  "post_date": "2018-03-30T20:44:42.800764500Z",
  "votes": 2,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone! At the beginning I will say only that this is my first competition on Kaggle, I'm still learning and I'm very sorry if my question seems obvious to anyone. I am very curious which validation set do you use? This topic is currently sparing my dreams. I searched the forum for information about what fragments of the training set people use for validation. Of course, I have found a lot of helpful information <a href=\"https://www.kaggle.com/konradb/validation-set/code\">here</a> and <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492\">here</a>. </p>\n\n<p>I know it is usefull to take data from November 9th between 4 and 5 am because they are probably part of the test set on Kaggle. I still wonder if this is a good practice. I also noticed that many people simply take the last 5% of the training set. </p>\n\n<p>Does any of you use cross validation? If so, how do you propose it?</p>\n\n<p>Thank you in advance for your answer.</p>",
  "messages": [
    {
      "id": "306695",
      "postDate": "03/30/2018 20:44:42",
      "content": "<p>Hi everyone! At the beginning I will say only that this is my first competition on Kaggle, I'm still learning and I'm very sorry if my question seems obvious to anyone. I am very curious which validation set do you use? This topic is currently sparing my dreams. I searched the forum for information about what fragments of the training set people use for validation. Of course, I have found a lot of helpful information <a href=\"https://www.kaggle.com/konradb/validation-set/code\">here</a> and <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492\">here</a>. </p>\n\n<p>I know it is usefull to take data from November 9th between 4 and 5 am because they are probably part of the test set on Kaggle. I still wonder if this is a good practice. I also noticed that many people simply take the last 5% of the training set. </p>\n\n<p>Does any of you use cross validation? If so, how do you propose it?</p>\n\n<p>Thank you in advance for your answer.</p>",
      "rawMarkdown": "Hi everyone! At the beginning I will say only that this is my first competition on Kaggle, I'm still learning and I'm very sorry if my question seems obvious to anyone. I am very curious which validation set do you use? This topic is currently sparing my dreams. I searched the forum for information about what fragments of the training set people use for validation. Of course, I have found a lot of helpful information [here][1] and [here][2]. \n\nI know it is usefull to take data from November 9th between 4 and 5 am because they are probably part of the test set on Kaggle. I still wonder if this is a good practice. I also noticed that many people simply take the last 5% of the training set. \n\nDoes any of you use cross validation? If so, how do you propose it?\n\nThank you in advance for your answer.\n\n  [1]: https://www.kaggle.com/konradb/validation-set/code\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492",
      "votes": null
    },
    {
      "id": "306710",
      "postDate": "03/30/2018 21:18:02",
      "content": "<p>I use whole day 9 as a validation set.</p>",
      "rawMarkdown": "I use whole day 9 as a validation set.",
      "votes": null
    },
    {
      "id": "306717",
      "postDate": "03/30/2018 21:30:12",
      "content": "<p>Do you notice a big difference between the score of your validation and public score on Kaggle on submit? I am also curious if you train the model on the entire training set?</p>",
      "rawMarkdown": "Do you notice a big difference between the score of your validation and public score on Kaggle on submit? I am also curious if you train the model on the entire training set?",
      "votes": null
    },
    {
      "id": "306724",
      "postDate": "03/30/2018 21:50:47",
      "content": "<p>I use a <a href=\"https://www.kaggle.com/aharless/training-and-validation-data-pickle\">validation set from day 9</a> that matches the <a href=\"https://www.kaggle.com/aharless/critical-times\">times for the test data</a> (6 full hours in separate 2 hour segments plus 1 additional second at the end of each of the segments). Sometimes I divide it into \"public\" and \"private\" <a href=\"https://www.kaggle.com/aharless/validation-of-kireev-style-nn/log\">subsets</a> corresponding to the public and private test data.  (The public data are the first hour, which corresponds to the first 4032690 records of my validation data for day 9.)</p>",
      "rawMarkdown": "I use a [validation set from day 9][1] that matches the [times for the test data][2] (6 full hours in separate 2 hour segments plus 1 additional second at the end of each of the segments). Sometimes I divide it into \"public\" and \"private\" [subsets][3] corresponding to the public and private test data.  (The public data are the first hour, which corresponds to the first 4032690 records of my validation data for day 9.)\n\n\n  [1]: https://www.kaggle.com/aharless/training-and-validation-data-pickle\n  [2]: https://www.kaggle.com/aharless/critical-times\n  [3]: https://www.kaggle.com/aharless/validation-of-kireev-style-nn/log",
      "votes": null
    },
    {
      "id": "306872",
      "postDate": "03/31/2018 07:55:06",
      "content": "<p>Thank you very much for your answer. A lot of valuable information has appeared <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325\">here</a>  too.</p>",
      "rawMarkdown": "Thank you very much for your answer. A lot of valuable information has appeared [here][1]  too.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325",
      "votes": null
    },
    {
      "id": "306873",
      "postDate": "03/31/2018 07:56:08",
      "content": "<p>Does anyone use some kind of cross validation?</p>",
      "rawMarkdown": "Does anyone use some kind of cross validation?",
      "votes": null
    },
    {
      "id": "306879",
      "postDate": "03/31/2018 08:19:32",
      "content": "<p>Difference between my CV and LB is consistent, in 0.002 - 0.003 range. And I do not use all training set. I train on day 8, validate on 9, if validation score increases, the retrain on day 8+9 and predict on test.</p>",
      "rawMarkdown": "Difference between my CV and LB is consistent, in 0.002 - 0.003 range. And I do not use all training set. I train on day 8, validate on 9, if validation score increases, the retrain on day 8+9 and predict on test.",
      "votes": null
    },
    {
      "id": "306943",
      "postDate": "03/31/2018 11:38:01",
      "content": "<p>For me the difference at this point is around 0.015 ... I certainly do something wrong and I try to understand what the problem is. I will try to modify my current script, bringing it closer to your validation. Maybe it's just my impression, but from what I read the forum people try to deliberately overfit...</p>\n\n<p>Out of curiosity, your result 0.9704 - is it one model or stacking?</p>",
      "rawMarkdown": "For me the difference at this point is around 0.015 ... I certainly do something wrong and I try to understand what the problem is. I will try to modify my current script, bringing it closer to your validation. Maybe it's just my impression, but from what I read the forum people try to deliberately overfit...\n\nOut of curiosity, your result 0.9704 - is it one model or stacking?",
      "votes": null
    },
    {
      "id": "306948",
      "postDate": "03/31/2018 11:57:05",
      "content": "<p>It's a single model.</p>",
      "rawMarkdown": "It's a single model.",
      "votes": null
    },
    {
      "id": "307080",
      "postDate": "03/31/2018 19:09:13",
      "content": "<p>I am using the kaggle kernel to do this competition because my local machine only has 4gb ram. could someone tell me which subset of the train data should I pick? </p>",
      "rawMarkdown": "I am using the kaggle kernel to do this competition because my local machine only has 4gb ram. could someone tell me which subset of the train data should I pick?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 306710,
      "author_name": "sohaibomar",
      "author_url": "",
      "post_date": "03/30/2018 21:18:02",
      "content": "<p>I use whole day 9 as a validation set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 306717,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "03/30/2018 21:30:12",
          "content": "<p>Do you notice a big difference between the score of your validation and public score on Kaggle on submit? I am also curious if you train the model on the entire training set?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306879,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/31/2018 08:19:32",
          "content": "<p>Difference between my CV and LB is consistent, in 0.002 - 0.003 range. And I do not use all training set. I train on day 8, validate on 9, if validation score increases, the retrain on day 8+9 and predict on test.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306943,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "03/31/2018 11:38:01",
          "content": "<p>For me the difference at this point is around 0.015 ... I certainly do something wrong and I try to understand what the problem is. I will try to modify my current script, bringing it closer to your validation. Maybe it's just my impression, but from what I read the forum people try to deliberately overfit...</p>\n\n<p>Out of curiosity, your result 0.9704 - is it one model or stacking?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 306948,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "03/31/2018 11:57:05",
          "content": "<p>It's a single model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306724,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/30/2018 21:50:47",
      "content": "<p>I use a <a href=\"https://www.kaggle.com/aharless/training-and-validation-data-pickle\">validation set from day 9</a> that matches the <a href=\"https://www.kaggle.com/aharless/critical-times\">times for the test data</a> (6 full hours in separate 2 hour segments plus 1 additional second at the end of each of the segments). Sometimes I divide it into \"public\" and \"private\" <a href=\"https://www.kaggle.com/aharless/validation-of-kireev-style-nn/log\">subsets</a> corresponding to the public and private test data.  (The public data are the first hour, which corresponds to the first 4032690 records of my validation data for day 9.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 306872,
          "author_name": "skalskip",
          "author_url": "",
          "post_date": "03/31/2018 07:55:06",
          "content": "<p>Thank you very much for your answer. A lot of valuable information has appeared <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325\">here</a>  too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 306873,
      "author_name": "skalskip",
      "author_url": "",
      "post_date": "03/31/2018 07:56:08",
      "content": "<p>Does anyone use some kind of cross validation?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 307080,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "03/31/2018 19:09:13",
      "content": "<p>I am using the kaggle kernel to do this competition because my local machine only has 4gb ram. could someone tell me which subset of the train data should I pick? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "306695": "Hi everyone! At the beginning I will say only that this is my first competition on Kaggle, I'm still learning and I'm very sorry if my question seems obvious to anyone. I am very curious which validation set do you use? This topic is currently sparing my dreams. I searched the forum for information about what fragments of the training set people use for validation. Of course, I have found a lot of helpful information [here][1] and [here][2]. \n\nI know it is usefull to take data from November 9th between 4 and 5 am because they are probably part of the test set on Kaggle. I still wonder if this is a good practice. I also noticed that many people simply take the last 5% of the training set. \n\nDoes any of you use cross validation? If so, how do you propose it?\n\nThank you in advance for your answer.\n\n  [1]: https://www.kaggle.com/konradb/validation-set/code\n  [2]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/51492",
    "306710": "I use whole day 9 as a validation set.",
    "306717": "Do you notice a big difference between the score of your validation and public score on Kaggle on submit? I am also curious if you train the model on the entire training set?",
    "306724": "I use a [validation set from day 9][1] that matches the [times for the test data][2] (6 full hours in separate 2 hour segments plus 1 additional second at the end of each of the segments). Sometimes I divide it into \"public\" and \"private\" [subsets][3] corresponding to the public and private test data.  (The public data are the first hour, which corresponds to the first 4032690 records of my validation data for day 9.)\n\n\n  [1]: https://www.kaggle.com/aharless/training-and-validation-data-pickle\n  [2]: https://www.kaggle.com/aharless/critical-times\n  [3]: https://www.kaggle.com/aharless/validation-of-kireev-style-nn/log",
    "306872": "Thank you very much for your answer. A lot of valuable information has appeared [here][1]  too.\n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53325",
    "306873": "Does anyone use some kind of cross validation?",
    "306879": "Difference between my CV and LB is consistent, in 0.002 - 0.003 range. And I do not use all training set. I train on day 8, validate on 9, if validation score increases, the retrain on day 8+9 and predict on test.",
    "306943": "For me the difference at this point is around 0.015 ... I certainly do something wrong and I try to understand what the problem is. I will try to modify my current script, bringing it closer to your validation. Maybe it's just my impression, but from what I read the forum people try to deliberately overfit...\n\nOut of curiosity, your result 0.9704 - is it one model or stacking?",
    "306948": "It's a single model.",
    "307080": "I am using the kaggle kernel to do this competition because my local machine only has 4gb ram. could someone tell me which subset of the train data should I pick?"
  },
  "source": "meta"
}