{
  "id": 77292,
  "title": "Using dask to load chunks of the data into memory",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/77292",
  "author_name": "",
  "post_date": "2019-01-11T07:04:27.461234Z",
  "votes": 15,
  "comment_count": 2,
  "views": 0,
  "content": "<p>If you're like me and don't have a large enough RAM in your system to load the train data all at once, you might find it useful to use something like <a href=\"https://dask.readthedocs.io/en/latest/dataframe.html\">dask</a> to load the data in chunks.</p>\n\n<p>Example usage:</p>\n\n<pre><code>import numpy as np\nimport dask.dataframe as dd\n\ndf = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\n\n# returns the first \"partition\" of the dataframe\npart_acoustic_data = np.array(df.acoustic_data.partitions[0]) \npart_time = np.array(df.time_to_failure.partitions[0])\n\n# print total number of partitions\nprint(df.npartitions)\n</code></pre>",
  "messages": [
    {
      "id": "454141",
      "postDate": "01/11/2019 07:04:27",
      "content": "<p>If you're like me and don't have a large enough RAM in your system to load the train data all at once, you might find it useful to use something like <a href=\"https://dask.readthedocs.io/en/latest/dataframe.html\">dask</a> to load the data in chunks.</p>\n\n<p>Example usage:</p>\n\n<pre><code>import numpy as np\nimport dask.dataframe as dd\n\ndf = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\n\n# returns the first \"partition\" of the dataframe\npart_acoustic_data = np.array(df.acoustic_data.partitions[0]) \npart_time = np.array(df.time_to_failure.partitions[0])\n\n# print total number of partitions\nprint(df.npartitions)\n</code></pre>",
      "rawMarkdown": "If you're like me and don't have a large enough RAM in your system to load the train data all at once, you might find it useful to use something like [dask][1] to load the data in chunks.\n\nExample usage:\n\n    import numpy as np\n    import dask.dataframe as dd\n    \n    df = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\n    \n    # returns the first \"partition\" of the dataframe\n    part_acoustic_data = np.array(df.acoustic_data.partitions[0]) \n    part_time = np.array(df.time_to_failure.partitions[0])\n    \n    # print total number of partitions\n    print(df.npartitions)\n\n  [1]: https://dask.readthedocs.io/en/latest/dataframe.html",
      "votes": null
    },
    {
      "id": "454228",
      "postDate": "01/11/2019 09:18:33",
      "content": "<p>You can do the same with pandas using 'nrows' and 'chunksize' arguments of read_csv method.</p>",
      "rawMarkdown": "You can do the same with pandas using 'nrows' and 'chunksize' arguments of read_csv method.",
      "votes": null
    },
    {
      "id": "455359",
      "postDate": "01/13/2019 17:00:08",
      "content": "<p>If you are computing features of data and then feed them to your model you can try compute features in kaggle-kernel, save them, and then download to your personal machine and work with them locally.  </p>",
      "rawMarkdown": "If you are computing features of data and then feed them to your model you can try compute features in kaggle-kernel, save them, and then download to your personal machine and work with them locally.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 454228,
      "author_name": "nroman",
      "author_url": "",
      "post_date": "01/11/2019 09:18:33",
      "content": "<p>You can do the same with pandas using 'nrows' and 'chunksize' arguments of read_csv method.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 455359,
      "author_name": "kostyaatarik",
      "author_url": "",
      "post_date": "01/13/2019 17:00:08",
      "content": "<p>If you are computing features of data and then feed them to your model you can try compute features in kaggle-kernel, save them, and then download to your personal machine and work with them locally.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454141": "If you're like me and don't have a large enough RAM in your system to load the train data all at once, you might find it useful to use something like [dask][1] to load the data in chunks.\n\nExample usage:\n\n    import numpy as np\n    import dask.dataframe as dd\n    \n    df = dd.read_csv(\"train.csv\", dtype={'acoustic_data': np.int16, 'time_to_failure': np.float64})\n    \n    # returns the first \"partition\" of the dataframe\n    part_acoustic_data = np.array(df.acoustic_data.partitions[0]) \n    part_time = np.array(df.time_to_failure.partitions[0])\n    \n    # print total number of partitions\n    print(df.npartitions)\n\n  [1]: https://dask.readthedocs.io/en/latest/dataframe.html",
    "454228": "You can do the same with pandas using 'nrows' and 'chunksize' arguments of read_csv method.",
    "455359": "If you are computing features of data and then feed them to your model you can try compute features in kaggle-kernel, save them, and then download to your personal machine and work with them locally."
  },
  "source": "meta"
}