{
  "id": 53989,
  "title": "Sample Training Based on the Testing Time Distribution",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53989",
  "author_name": "",
  "post_date": "2018-04-08T01:50:36.732601800Z",
  "votes": null,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hi Kagglers: </p>\n\n<p>Does anybody have any experience on sampling the training data based on the testing dataset's time distribution? This helped me by shrinking the dataset down to about 60m records. I was able to train a simple lgbm on it and get testing result about 0.96. I would love to hear any experience on how much value does the rest of the training set have. </p>",
  "messages": [
    {
      "id": "310568",
      "postDate": "04/08/2018 01:50:36",
      "content": "<p>Hi Kagglers: </p>\n\n<p>Does anybody have any experience on sampling the training data based on the testing dataset's time distribution? This helped me by shrinking the dataset down to about 60m records. I was able to train a simple lgbm on it and get testing result about 0.96. I would love to hear any experience on how much value does the rest of the training set have. </p>",
      "rawMarkdown": "Hi Kagglers: \n\nDoes anybody have any experience on sampling the training data based on the testing dataset's time distribution? This helped me by shrinking the dataset down to about 60m records. I was able to train a simple lgbm on it and get testing result about 0.96. I would love to hear any experience on how much value does the rest of the training set have.",
      "votes": null
    },
    {
      "id": "310664",
      "postDate": "04/08/2018 07:43:03",
      "content": "<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\">This discussion</a>  is really worth reading. Strategies shared by fellow Kagglers will certainly save a lot time if you're planning to go with more data. </p>\n\n<p>@<strong>Kaggle</strong> , I really feel that search functionality needs to be improved (<strong>fixed</strong>  - to be honest)  at least in discussion forum. </p>",
      "rawMarkdown": "[This discussion][1]  is really worth reading. Strategies shared by fellow Kagglers will certainly save a lot time if you're planning to go with more data. \n\n@**Kaggle** , I really feel that search functionality needs to be improved (**fixed**  - to be honest)  at least in discussion forum. \n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 310664,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/08/2018 07:43:03",
      "content": "<p><a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634\">This discussion</a>  is really worth reading. Strategies shared by fellow Kagglers will certainly save a lot time if you're planning to go with more data. </p>\n\n<p>@<strong>Kaggle</strong> , I really feel that search functionality needs to be improved (<strong>fixed</strong>  - to be honest)  at least in discussion forum. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "310568": "Hi Kagglers: \n\nDoes anybody have any experience on sampling the training data based on the testing dataset's time distribution? This helped me by shrinking the dataset down to about 60m records. I was able to train a simple lgbm on it and get testing result about 0.96. I would love to hear any experience on how much value does the rest of the training set have.",
    "310664": "[This discussion][1]  is really worth reading. Strategies shared by fellow Kagglers will certainly save a lot time if you're planning to go with more data. \n\n@**Kaggle** , I really feel that search functionality needs to be improved (**fixed**  - to be honest)  at least in discussion forum. \n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634"
  },
  "source": "meta"
}