{
  "id": 77275,
  "title": "Reduce data size",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77275",
  "author_name": "",
  "post_date": "2019-01-11T03:51:33.386738800Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>How are you guys thinking about reducing the size of the training data? With nearly 630 million lines, it will be near impossible for me to use this data without making some kind of reduction.</p>",
  "messages": [
    {
      "id": "454039",
      "postDate": "01/11/2019 03:51:33",
      "content": "<p>How are you guys thinking about reducing the size of the training data? With nearly 630 million lines, it will be near impossible for me to use this data without making some kind of reduction.</p>",
      "rawMarkdown": "How are you guys thinking about reducing the size of the training data? With nearly 630 million lines, it will be near impossible for me to use this data without making some kind of reduction.",
      "votes": null
    },
    {
      "id": "454134",
      "postDate": "01/11/2019 07:00:07",
      "content": "<p><a href=\"https://www.kaggle.com/gemartin/load-data-reduce-memory-usage\">https://www.kaggle.com/gemartin/load-data-reduce-memory-usage</a></p>",
      "rawMarkdown": "https://www.kaggle.com/gemartin/load-data-reduce-memory-usage",
      "votes": null
    },
    {
      "id": "454637",
      "postDate": "01/11/2019 22:59:31",
      "content": "<p>You can use int16 for acoustic data (only 2 bytes) and maybe float32 for time, although the precision might be important (have to check more carefully). Another option is working with samples, especially for feature engineering.</p>",
      "rawMarkdown": "You can use int16 for acoustic data (only 2 bytes) and maybe float32 for time, although the precision might be important (have to check more carefully). Another option is working with samples, especially for feature engineering.",
      "votes": null
    },
    {
      "id": "455147",
      "postDate": "01/13/2019 04:23:11",
      "content": "<p>You can use the Pandas resample functionality\n<a href=\"https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.resample.html\">https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.resample.html</a></p>",
      "rawMarkdown": "You can use the Pandas resample functionality\nhttps://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.resample.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454134,
      "author_name": "timon88",
      "author_url": "",
      "post_date": "01/11/2019 07:00:07",
      "content": "<p><a href=\"https://www.kaggle.com/gemartin/load-data-reduce-memory-usage\">https://www.kaggle.com/gemartin/load-data-reduce-memory-usage</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 454637,
      "author_name": "jsaguiar",
      "author_url": "",
      "post_date": "01/11/2019 22:59:31",
      "content": "<p>You can use int16 for acoustic data (only 2 bytes) and maybe float32 for time, although the precision might be important (have to check more carefully). Another option is working with samples, especially for feature engineering.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 455147,
      "author_name": "cv13j0",
      "author_url": "",
      "post_date": "01/13/2019 04:23:11",
      "content": "<p>You can use the Pandas resample functionality\n<a href=\"https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.resample.html\">https://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.resample.html</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454039": "How are you guys thinking about reducing the size of the training data? With nearly 630 million lines, it will be near impossible for me to use this data without making some kind of reduction.",
    "454134": "https://www.kaggle.com/gemartin/load-data-reduce-memory-usage",
    "454637": "You can use int16 for acoustic data (only 2 bytes) and maybe float32 for time, although the precision might be important (have to check more carefully). Another option is working with samples, especially for feature engineering.",
    "455147": "You can use the Pandas resample functionality\nhttps://pandas.pydata.org/pandas-docs/stable/generated/pandas.DataFrame.resample.html"
  },
  "source": "meta"
}