{
  "id": 77595,
  "title": "Dont load (all) of the data",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77595",
  "author_name": "",
  "post_date": "2019-01-14T17:24:39.367964300Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Just a tought:</p>\n\n<p>If you have not found a way to subsample the data from the train dataset (since its a one continous measurament) in a way that it represents correctly samples from the test set ( possibly independent from each other measuraments) than training the model on ALL of the train data may actually hurt predictions on the test dataset. In other words if you just sub-sample a lot of 150 000rows samples, that are not representative of the test data set your score will be lower (mine was atleast)</p>\n\n<p>Conclusion: use nrows for now and dont overfit on training data...</p>",
  "messages": [
    {
      "id": "455847",
      "postDate": "01/14/2019 17:24:39",
      "content": "<p>Just a tought:</p>\n\n<p>If you have not found a way to subsample the data from the train dataset (since its a one continous measurament) in a way that it represents correctly samples from the test set ( possibly independent from each other measuraments) than training the model on ALL of the train data may actually hurt predictions on the test dataset. In other words if you just sub-sample a lot of 150 000rows samples, that are not representative of the test data set your score will be lower (mine was atleast)</p>\n\n<p>Conclusion: use nrows for now and dont overfit on training data...</p>",
      "rawMarkdown": "Just a tought:\n\n\nIf you have not found a way to subsample the data from the train dataset (since its a one continous measurament) in a way that it represents correctly samples from the test set ( possibly independent from each other measuraments) than training the model on ALL of the train data may actually hurt predictions on the test dataset. In other words if you just sub-sample a lot of 150 000rows samples, that are not representative of the test data set your score will be lower (mine was atleast)\n\nConclusion: use nrows for now and dont overfit on training data...",
      "votes": null
    },
    {
      "id": "456353",
      "postDate": "01/15/2019 16:52:08",
      "content": "<p>What do you mean by nrows?</p>",
      "rawMarkdown": "What do you mean by nrows?",
      "votes": null
    },
    {
      "id": "456356",
      "postDate": "01/15/2019 16:56:22",
      "content": "<p>nrows is a parameter in pandas’ read_csv. i think thats what was meant by OP</p>",
      "rawMarkdown": "nrows is a parameter in pandas’ read_csv. i think thats what was meant by OP",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 456353,
      "author_name": "horizonpicking2k18",
      "author_url": "",
      "post_date": "01/15/2019 16:52:08",
      "content": "<p>What do you mean by nrows?</p>",
      "votes": null,
      "replies": [
        {
          "id": 456356,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "01/15/2019 16:56:22",
          "content": "<p>nrows is a parameter in pandas’ read_csv. i think thats what was meant by OP</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "455847": "Just a tought:\n\n\nIf you have not found a way to subsample the data from the train dataset (since its a one continous measurament) in a way that it represents correctly samples from the test set ( possibly independent from each other measuraments) than training the model on ALL of the train data may actually hurt predictions on the test dataset. In other words if you just sub-sample a lot of 150 000rows samples, that are not representative of the test data set your score will be lower (mine was atleast)\n\nConclusion: use nrows for now and dont overfit on training data...",
    "456353": "What do you mean by nrows?",
    "456356": "nrows is a parameter in pandas’ read_csv. i think thats what was meant by OP"
  },
  "source": "meta"
}