{
  "id": 91572,
  "title": "Use msgpack to load/save data.",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91572",
  "author_name": "",
  "post_date": "2019-05-06T16:23:27.892700800Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I just started in this competition, but from what I have seen all the people are loading the data as csv. I know that this is normal in the kernels but not in your local experiments. Probably many of you know it, but just in case.</p>\n\n<p>On your local, to improve timing, just save the raw data as <a href=\"https://msgpack.org/index.html\">msgpack</a></p>\n\n<p>In my i7 it took only 8.5 secs to load the full dataset (comparing with the 15 minutes with csv)</p>\n\n<p><code>\ndf_train_csv = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\ndf_train_csv.to_msgpack('../output/train.msg')\n</code></p>\n\n<p><code>\ndf_train_msg = pd.read_msgpack('../output/train.msg')\n</code></p>",
  "messages": [
    {
      "id": "527922",
      "postDate": "05/06/2019 16:23:27",
      "content": "<p>I just started in this competition, but from what I have seen all the people are loading the data as csv. I know that this is normal in the kernels but not in your local experiments. Probably many of you know it, but just in case.</p>\n\n<p>On your local, to improve timing, just save the raw data as <a href=\"https://msgpack.org/index.html\">msgpack</a></p>\n\n<p>In my i7 it took only 8.5 secs to load the full dataset (comparing with the 15 minutes with csv)</p>\n\n<p><code>\ndf_train_csv = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\ndf_train_csv.to_msgpack('../output/train.msg')\n</code></p>\n\n<p><code>\ndf_train_msg = pd.read_msgpack('../output/train.msg')\n</code></p>",
      "rawMarkdown": "I just started in this competition, but from what I have seen all the people are loading the data as csv. I know that this is normal in the kernels but not in your local experiments. Probably many of you know it, but just in case.\n\nOn your local, to improve timing, just save the raw data as [msgpack](https://msgpack.org/index.html)\n\nIn my i7 it took only 8.5 secs to load the full dataset (comparing with the 15 minutes with csv)\n\n```\ndf_train_csv = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\ndf_train_csv.to_msgpack('../output/train.msg')\n```\n\n```\ndf_train_msg = pd.read_msgpack('../output/train.msg')\n```",
      "votes": null
    },
    {
      "id": "527986",
      "postDate": "05/06/2019 18:46:37",
      "content": "<p>Please note as per pandas' documentation: \"THIS IS AN EXPERIMENTAL LIBRARY and the storage format may not be stable until a future release.\"</p>",
      "rawMarkdown": "Please note as per pandas' documentation: \"THIS IS AN EXPERIMENTAL LIBRARY and the storage format may not be stable until a future release.\"",
      "votes": null
    },
    {
      "id": "528172",
      "postDate": "05/07/2019 07:26:52",
      "content": "<p>You can use other implementations <a href=\"https://github.com/msgpack/msgpack-python\">https://github.com/msgpack/msgpack-python</a></p>\n\n<p>or use other formats such as hdf5 <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to_hdf.html\">https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to_hdf.html</a></p>\n\n<p>The key is not to use csv.</p>",
      "rawMarkdown": "You can use other implementations https://github.com/msgpack/msgpack-python\n\nor use other formats such as hdf5 https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to_hdf.html\n\nThe key is not to use csv.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 527986,
      "author_name": "mhviraf",
      "author_url": "",
      "post_date": "05/06/2019 18:46:37",
      "content": "<p>Please note as per pandas' documentation: \"THIS IS AN EXPERIMENTAL LIBRARY and the storage format may not be stable until a future release.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 528172,
          "author_name": "rhortelanos",
          "author_url": "",
          "post_date": "05/07/2019 07:26:52",
          "content": "<p>You can use other implementations <a href=\"https://github.com/msgpack/msgpack-python\">https://github.com/msgpack/msgpack-python</a></p>\n\n<p>or use other formats such as hdf5 <a href=\"https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to_hdf.html\">https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to_hdf.html</a></p>\n\n<p>The key is not to use csv.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "527922": "I just started in this competition, but from what I have seen all the people are loading the data as csv. I know that this is normal in the kernels but not in your local experiments. Probably many of you know it, but just in case.\n\nOn your local, to improve timing, just save the raw data as [msgpack](https://msgpack.org/index.html)\n\nIn my i7 it took only 8.5 secs to load the full dataset (comparing with the 15 minutes with csv)\n\n```\ndf_train_csv = pd.read_csv('../input/train.csv', dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\ndf_train_csv.to_msgpack('../output/train.msg')\n```\n\n```\ndf_train_msg = pd.read_msgpack('../output/train.msg')\n```",
    "527986": "Please note as per pandas' documentation: \"THIS IS AN EXPERIMENTAL LIBRARY and the storage format may not be stable until a future release.\"",
    "528172": "You can use other implementations https://github.com/msgpack/msgpack-python\n\nor use other formats such as hdf5 https://pandas.pydata.org/pandas-docs/stable/reference/api/pandas.DataFrame.to_hdf.html\n\nThe key is not to use csv."
  },
  "source": "meta"
}