{
  "id": 189499,
  "title": "Feather files",
  "url": "/competitions/riiid-test-answer-prediction/discussion/189499",
  "author_name": "",
  "post_date": "2020-10-07T19:24:40.974389200Z",
  "votes": 15,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I made a dataset with feather format. I tried to preserve the original data format as much as possible. Loading this data goes from 5 minutes to 6 seconds.</p>\n<p>dataset: <a href=\"https://www.kaggle.com/ryati131457/riiid-feather-files\" target=\"_blank\">https://www.kaggle.com/ryati131457/riiid-feather-files</a></p>\n<p>notebook: <a href=\"https://www.kaggle.com/ryati131457/riiid-file-convert-feather\" target=\"_blank\">https://www.kaggle.com/ryati131457/riiid-file-convert-feather</a></p>\n<h2>Parquet Format</h2>\n<p>After the suggestion for <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I experimented more with the parquet format. You can get the filesize to around 670 MB and only sacrifice 1-2 seconds of load time. Take a look at the data here:<br>\n<a href=\"https://www.kaggle.com/ryati131457/riiid-parquet-files\" target=\"_blank\">https://www.kaggle.com/ryati131457/riiid-parquet-files</a></p>",
  "messages": [
    {
      "id": "1041499",
      "postDate": "10/07/2020 19:24:40",
      "content": "<p>I made a dataset with feather format. I tried to preserve the original data format as much as possible. Loading this data goes from 5 minutes to 6 seconds.</p>\n<p>dataset: <a href=\"https://www.kaggle.com/ryati131457/riiid-feather-files\" target=\"_blank\">https://www.kaggle.com/ryati131457/riiid-feather-files</a></p>\n<p>notebook: <a href=\"https://www.kaggle.com/ryati131457/riiid-file-convert-feather\" target=\"_blank\">https://www.kaggle.com/ryati131457/riiid-file-convert-feather</a></p>\n<h2>Parquet Format</h2>\n<p>After the suggestion for <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> I experimented more with the parquet format. You can get the filesize to around 670 MB and only sacrifice 1-2 seconds of load time. Take a look at the data here:<br>\n<a href=\"https://www.kaggle.com/ryati131457/riiid-parquet-files\" target=\"_blank\">https://www.kaggle.com/ryati131457/riiid-parquet-files</a></p>",
      "rawMarkdown": "I made a dataset with feather format. I tried to preserve the original data format as much as possible. Loading this data goes from 5 minutes to 6 seconds.\n\ndataset: https://www.kaggle.com/ryati131457/riiid-feather-files\n\nnotebook: https://www.kaggle.com/ryati131457/riiid-file-convert-feather\n\n## Parquet Format\n\nAfter the suggestion for @sohier I experimented more with the parquet format. You can get the filesize to around 670 MB and only sacrifice 1-2 seconds of load time. Take a look at the data here:\nhttps://www.kaggle.com/ryati131457/riiid-parquet-files",
      "votes": null
    },
    {
      "id": "1041506",
      "postDate": "10/07/2020 19:28:53",
      "content": "<p>I've never heard of this filetype before. Very interesting! Thank you for sharing</p>",
      "rawMarkdown": "I've never heard of this filetype before. Very interesting! Thank you for sharing",
      "votes": null
    },
    {
      "id": "1041591",
      "postDate": "10/07/2020 20:36:24",
      "content": "<p>There is also parquet, but it was a bit slower than feather when I tested it. CSV is universal, but super slow.</p>",
      "rawMarkdown": "There is also parquet, but it was a bit slower than feather when I tested it. CSV is universal, but super slow.",
      "votes": null
    },
    {
      "id": "1041600",
      "postDate": "10/07/2020 20:43:39",
      "content": "<p>Cool! Any idea how it compares to say hd5py or numpy loading?</p>",
      "rawMarkdown": "Cool! Any idea how it compares to say hd5py or numpy loading?",
      "votes": null
    },
    {
      "id": "1041606",
      "postDate": "10/07/2020 20:48:07",
      "content": "<p><a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> Parquet should be the same speed if you turn off compression, no?</p>",
      "rawMarkdown": "ryati131457 Parquet should be the same speed if you turn off compression, no?",
      "votes": null
    },
    {
      "id": "1041634",
      "postDate": "10/07/2020 21:07:55",
      "content": "<p>This would be one of the things I learned from this competition :)</p>",
      "rawMarkdown": "This would be one of the things I learned from this competition :)",
      "votes": null
    },
    {
      "id": "1041636",
      "postDate": "10/07/2020 21:10:54",
      "content": "<p><a href=\"https://www.kaggle.com/matthewmasters\" target=\"_blank\">@matthewmasters</a> I am not sure about hd5py. Might be worth checking out. I think the npy format is really fast, but as far as I know, you don't have preservation of column names.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> It looks like feather is still faster in this case. However, there may be other cases where parquet wins:<br>\n<a href=\"https://www.kaggle.com/ryati131457/feather-vs-parquet\" target=\"_blank\">https://www.kaggle.com/ryati131457/feather-vs-parquet</a></p>",
      "rawMarkdown": "matthewmasters I am not sure about hd5py. Might be worth checking out. I think the npy format is really fast, but as far as I know, you don't have preservation of column names.\n\n@sohier It looks like feather is still faster in this case. However, there may be other cases where parquet wins:\nhttps://www.kaggle.com/ryati131457/feather-vs-parquet",
      "votes": null
    },
    {
      "id": "1041637",
      "postDate": "10/07/2020 21:12:26",
      "content": "<p>Thanks for you input! 👍</p>",
      "rawMarkdown": "Thanks for you input! 👍",
      "votes": null
    },
    {
      "id": "1041640",
      "postDate": "10/07/2020 21:19:23",
      "content": "<p><a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> Would you mind to tell me why you are able to load the whole train.csv without memory error? Probably I missed some discussion in this forum. Thanks in advance.</p>",
      "rawMarkdown": "ryati131457 Would you mind to tell me why you are able to load the whole train.csv without memory error? Probably I missed some discussion in this forum. Thanks in advance.",
      "votes": null
    },
    {
      "id": "1041665",
      "postDate": "10/07/2020 21:35:16",
      "content": "<p>Edit: I guess you can't load the whole file as is. Use a dtype dictionary to decrease file size..</p>\n<p>I had no issues loading it. But I am also not loading anything else. If you are loading tons of python modules, it may take up a bit of memory, leaving you with less.</p>\n<p>The key is to reducing the file size is to specify a dtype dictionary when you <code>read_csv</code>. Pandas has to scan the file and make a guess as to the data type. But the guess is really conservative, so it will choose large formats, like 'float64' and 'int64' or 'object' for everything. Loading with 'int8' on this column and 'float32' on that column really adds up. You can see that the fully loaded feather file is around 3.5 GB. I think if you leave it with defaults (no dtype dictionary) its over 7GB. </p>\n<p>Also, remember that if you re-ran your notebook a few times, you might still be holding on to some unused memory. If you look at the notebook linked above, you see i use <code>del train</code> and <code>gc.collect()</code>. <code>del</code> deletes the python object <code>train</code>. <code>gc.collect()</code> tells the garbage collector to free up memory now. If you just run delete, it can still be holding on to that memory because the garbage collector has not run yet.</p>",
      "rawMarkdown": "Edit: I guess you can't load the whole file as is. Use a dtype dictionary to decrease file size..\n\nI had no issues loading it. But I am also not loading anything else. If you are loading tons of python modules, it may take up a bit of memory, leaving you with less.\n\nThe key is to reducing the file size is to specify a dtype dictionary when you `read_csv`. Pandas has to scan the file and make a guess as to the data type. But the guess is really conservative, so it will choose large formats, like 'float64' and 'int64' or 'object' for everything. Loading with 'int8' on this column and 'float32' on that column really adds up. You can see that the fully loaded feather file is around 3.5 GB. I think if you leave it with defaults (no dtype dictionary) its over 7GB. \n\nAlso, remember that if you re-ran your notebook a few times, you might still be holding on to some unused memory. If you look at the notebook linked above, you see i use `del train` and `gc.collect()`. `del` deletes the python object `train`. `gc.collect()` tells the garbage collector to free up memory now. If you just run delete, it can still be holding on to that memory because the garbage collector has not run yet.",
      "votes": null
    },
    {
      "id": "1042152",
      "postDate": "10/08/2020 05:09:11",
      "content": "<p><a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> You <strong>can</strong> load the entire original train data set in a Kaggle kernel.</p>\n<p>This question was <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908\" target=\"_blank\">already answered</a> along with a <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> to demonstrate this.</p>",
      "rawMarkdown": "ryati131457 You **can** load the entire original train data set in a Kaggle kernel.\n\nThis question was [already answered](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908) along with a [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) to demonstrate this.",
      "votes": null
    },
    {
      "id": "1042354",
      "postDate": "10/08/2020 07:28:00",
      "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> , Thanks. But in the notebook of <a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> , he didn't need to use chunk, and he was able to load the whole csv once.</p>",
      "rawMarkdown": "sirishks , Thanks. But in the notebook of @ryati131457 , he didn't need to use chunk, and he was able to load the whole csv once.",
      "votes": null
    },
    {
      "id": "1043119",
      "postDate": "10/08/2020 17:49:34",
      "content": "<p><a href=\"https://datatable.readthedocs.io/en/latest/index.html\" target=\"_blank\">Python datatable</a> can read the entire train data from binary format in less than a second: <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid</a></p>\n<p>You might still want to convert it to pandas ultimately, which takes less than a minute with datatable from the raw csv file, so it's a good option to use if you want a one-line solution 🙂</p>",
      "rawMarkdown": "[Python datatable](https://datatable.readthedocs.io/en/latest/index.html) can read the entire train data from binary format in less than a second: https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\n\nYou might still want to convert it to pandas ultimately, which takes less than a minute with datatable from the raw csv file, so it's a good option to use if you want a one-line solution 🙂",
      "votes": null
    },
    {
      "id": "1043144",
      "postDate": "10/08/2020 18:08:45",
      "content": "<p>very cool. I have never seen Python datatable. </p>",
      "rawMarkdown": "very cool. I have never seen Python datatable.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1041506,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "10/07/2020 19:28:53",
      "content": "<p>I've never heard of this filetype before. Very interesting! Thank you for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041591,
          "author_name": "ryati131457",
          "author_url": "",
          "post_date": "10/07/2020 20:36:24",
          "content": "<p>There is also parquet, but it was a bit slower than feather when I tested it. CSV is universal, but super slow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041600,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "10/07/2020 20:43:39",
          "content": "<p>Cool! Any idea how it compares to say hd5py or numpy loading?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041606,
          "author_name": "sohier",
          "author_url": "",
          "post_date": "10/07/2020 20:48:07",
          "content": "<p><a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> Parquet should be the same speed if you turn off compression, no?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041636,
          "author_name": "ryati131457",
          "author_url": "",
          "post_date": "10/07/2020 21:10:54",
          "content": "<p><a href=\"https://www.kaggle.com/matthewmasters\" target=\"_blank\">@matthewmasters</a> I am not sure about hd5py. Might be worth checking out. I think the npy format is really fast, but as far as I know, you don't have preservation of column names.</p>\n<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> It looks like feather is still faster in this case. However, there may be other cases where parquet wins:<br>\n<a href=\"https://www.kaggle.com/ryati131457/feather-vs-parquet\" target=\"_blank\">https://www.kaggle.com/ryati131457/feather-vs-parquet</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1041637,
          "author_name": "matthewmasters",
          "author_url": "",
          "post_date": "10/07/2020 21:12:26",
          "content": "<p>Thanks for you input! 👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1041634,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "10/07/2020 21:07:55",
      "content": "<p>This would be one of the things I learned from this competition :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1041640,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "10/07/2020 21:19:23",
      "content": "<p><a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> Would you mind to tell me why you are able to load the whole train.csv without memory error? Probably I missed some discussion in this forum. Thanks in advance.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1041665,
          "author_name": "ryati131457",
          "author_url": "",
          "post_date": "10/07/2020 21:35:16",
          "content": "<p>Edit: I guess you can't load the whole file as is. Use a dtype dictionary to decrease file size..</p>\n<p>I had no issues loading it. But I am also not loading anything else. If you are loading tons of python modules, it may take up a bit of memory, leaving you with less.</p>\n<p>The key is to reducing the file size is to specify a dtype dictionary when you <code>read_csv</code>. Pandas has to scan the file and make a guess as to the data type. But the guess is really conservative, so it will choose large formats, like 'float64' and 'int64' or 'object' for everything. Loading with 'int8' on this column and 'float32' on that column really adds up. You can see that the fully loaded feather file is around 3.5 GB. I think if you leave it with defaults (no dtype dictionary) its over 7GB. </p>\n<p>Also, remember that if you re-ran your notebook a few times, you might still be holding on to some unused memory. If you look at the notebook linked above, you see i use <code>del train</code> and <code>gc.collect()</code>. <code>del</code> deletes the python object <code>train</code>. <code>gc.collect()</code> tells the garbage collector to free up memory now. If you just run delete, it can still be holding on to that memory because the garbage collector has not run yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1042152,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "10/08/2020 05:09:11",
          "content": "<p><a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> You <strong>can</strong> load the entire original train data set in a Kaggle kernel.</p>\n<p>This question was <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908\" target=\"_blank\">already answered</a> along with a <a href=\"https://www.kaggle.com/sirishks/cant-load-train-data\" target=\"_blank\">kernel</a> to demonstrate this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1042354,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "10/08/2020 07:28:00",
          "content": "<p><a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> , Thanks. But in the notebook of <a href=\"https://www.kaggle.com/ryati131457\" target=\"_blank\">@ryati131457</a> , he didn't need to use chunk, and he was able to load the whole csv once.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1043119,
      "author_name": "rohanrao",
      "author_url": "",
      "post_date": "10/08/2020 17:49:34",
      "content": "<p><a href=\"https://datatable.readthedocs.io/en/latest/index.html\" target=\"_blank\">Python datatable</a> can read the entire train data from binary format in less than a second: <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid</a></p>\n<p>You might still want to convert it to pandas ultimately, which takes less than a minute with datatable from the raw csv file, so it's a good option to use if you want a one-line solution 🙂</p>",
      "votes": null,
      "replies": [
        {
          "id": 1043144,
          "author_name": "ryati131457",
          "author_url": "",
          "post_date": "10/08/2020 18:08:45",
          "content": "<p>very cool. I have never seen Python datatable. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1041499": "I made a dataset with feather format. I tried to preserve the original data format as much as possible. Loading this data goes from 5 minutes to 6 seconds.\n\ndataset: https://www.kaggle.com/ryati131457/riiid-feather-files\n\nnotebook: https://www.kaggle.com/ryati131457/riiid-file-convert-feather\n\n## Parquet Format\n\nAfter the suggestion for @sohier I experimented more with the parquet format. You can get the filesize to around 670 MB and only sacrifice 1-2 seconds of load time. Take a look at the data here:\nhttps://www.kaggle.com/ryati131457/riiid-parquet-files",
    "1041506": "I've never heard of this filetype before. Very interesting! Thank you for sharing",
    "1041591": "There is also parquet, but it was a bit slower than feather when I tested it. CSV is universal, but super slow.",
    "1041600": "Cool! Any idea how it compares to say hd5py or numpy loading?",
    "1041606": "ryati131457 Parquet should be the same speed if you turn off compression, no?",
    "1041634": "This would be one of the things I learned from this competition :)",
    "1041636": "matthewmasters I am not sure about hd5py. Might be worth checking out. I think the npy format is really fast, but as far as I know, you don't have preservation of column names.\n\n@sohier It looks like feather is still faster in this case. However, there may be other cases where parquet wins:\nhttps://www.kaggle.com/ryati131457/feather-vs-parquet",
    "1041637": "Thanks for you input! 👍",
    "1041640": "ryati131457 Would you mind to tell me why you are able to load the whole train.csv without memory error? Probably I missed some discussion in this forum. Thanks in advance.",
    "1041665": "Edit: I guess you can't load the whole file as is. Use a dtype dictionary to decrease file size..\n\nI had no issues loading it. But I am also not loading anything else. If you are loading tons of python modules, it may take up a bit of memory, leaving you with less.\n\nThe key is to reducing the file size is to specify a dtype dictionary when you `read_csv`. Pandas has to scan the file and make a guess as to the data type. But the guess is really conservative, so it will choose large formats, like 'float64' and 'int64' or 'object' for everything. Loading with 'int8' on this column and 'float32' on that column really adds up. You can see that the fully loaded feather file is around 3.5 GB. I think if you leave it with defaults (no dtype dictionary) its over 7GB. \n\nAlso, remember that if you re-ran your notebook a few times, you might still be holding on to some unused memory. If you look at the notebook linked above, you see i use `del train` and `gc.collect()`. `del` deletes the python object `train`. `gc.collect()` tells the garbage collector to free up memory now. If you just run delete, it can still be holding on to that memory because the garbage collector has not run yet.",
    "1042152": "ryati131457 You **can** load the entire original train data set in a Kaggle kernel.\n\nThis question was [already answered](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/188908) along with a [kernel](https://www.kaggle.com/sirishks/cant-load-train-data) to demonstrate this.",
    "1042354": "sirishks , Thanks. But in the notebook of @ryati131457 , he didn't need to use chunk, and he was able to load the whole csv once.",
    "1043119": "[Python datatable](https://datatable.readthedocs.io/en/latest/index.html) can read the entire train data from binary format in less than a second: https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\n\nYou might still want to convert it to pandas ultimately, which takes less than a minute with datatable from the raw csv file, so it's a good option to use if you want a one-line solution 🙂",
    "1043144": "very cool. I have never seen Python datatable."
  },
  "source": "meta"
}