{
  "id": 52890,
  "title": "Using all of the data",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/52890",
  "author_name": "",
  "post_date": "2018-03-24T17:49:11.305408200Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I see many kernels obtaining pretty nice LB scores from just using a subset of the training data. Doesn't that mean they can usually be improved if the entire training data is used? More data is always better right?</p>\n\n<p>I can't think of a reason why using the full data might hurt your score (besides computational resources).</p>",
  "messages": [
    {
      "id": "302750",
      "postDate": "03/24/2018 17:49:11",
      "content": "<p>I see many kernels obtaining pretty nice LB scores from just using a subset of the training data. Doesn't that mean they can usually be improved if the entire training data is used? More data is always better right?</p>\n\n<p>I can't think of a reason why using the full data might hurt your score (besides computational resources).</p>",
      "rawMarkdown": "I see many kernels obtaining pretty nice LB scores from just using a subset of the training data. Doesn't that mean they can usually be improved if the entire training data is used? More data is always better right?\n\nI can't think of a reason why using the full data might hurt your score (besides computational resources).",
      "votes": null
    },
    {
      "id": "302878",
      "postDate": "03/24/2018 22:35:30",
      "content": "<p>I trained the same model for 100k, 1 million, 10 million, and 100 million rows and of course the score increases.\nI wrote down the scores <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52927\">here</a> </p>",
      "rawMarkdown": "I trained the same model for 100k, 1 million, 10 million, and 100 million rows and of course the score increases.\nI wrote down the scores [here][1] \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52927",
      "votes": null
    },
    {
      "id": "302942",
      "postDate": "03/25/2018 03:10:16",
      "content": "<p>I've generally obtained higher scores by running public kernels on the full data set at home (where I have 32MB RAM plus swap space, so don't have the 16MB limitation that requires the kernels on Kaggle to use limited data).  In some cases the scores are much higher, although there are a few where I didn't get any improvement.</p>",
      "rawMarkdown": "I've generally obtained higher scores by running public kernels on the full data set at home (where I have 32MB RAM plus swap space, so don't have the 16MB limitation that requires the kernels on Kaggle to use limited data).  In some cases the scores are much higher, although there are a few where I didn't get any improvement.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 302878,
      "author_name": "araksstepanyan",
      "author_url": "",
      "post_date": "03/24/2018 22:35:30",
      "content": "<p>I trained the same model for 100k, 1 million, 10 million, and 100 million rows and of course the score increases.\nI wrote down the scores <a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52927\">here</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 302942,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "03/25/2018 03:10:16",
      "content": "<p>I've generally obtained higher scores by running public kernels on the full data set at home (where I have 32MB RAM plus swap space, so don't have the 16MB limitation that requires the kernels on Kaggle to use limited data).  In some cases the scores are much higher, although there are a few where I didn't get any improvement.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "302750": "I see many kernels obtaining pretty nice LB scores from just using a subset of the training data. Doesn't that mean they can usually be improved if the entire training data is used? More data is always better right?\n\nI can't think of a reason why using the full data might hurt your score (besides computational resources).",
    "302878": "I trained the same model for 100k, 1 million, 10 million, and 100 million rows and of course the score increases.\nI wrote down the scores [here][1] \n\n\n  [1]: https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/52927",
    "302942": "I've generally obtained higher scores by running public kernels on the full data set at home (where I have 32MB RAM plus swap space, so don't have the 16MB limitation that requires the kernels on Kaggle to use limited data).  In some cases the scores are much higher, although there are a few where I didn't get any improvement."
  },
  "source": "meta"
}