{
  "id": 54948,
  "title": "Load 240MM rows in... 8 seconds",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54948",
  "author_name": "",
  "post_date": "2018-04-20T04:10:08.144808100Z",
  "votes": 46,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Been away from the competition for 6 days and my LB drops 254 spots? Brutal lol. Time to read every forum post and kernel to see what dark magix everyone's discovered this past week. Anyhow, onto the goods:</p>\n\n<p>There is a new format on the block called feather. It's more optimal than pickling, or--God forbid--raw .csvs. It essentially writes files to disk in exactly the format it's stored in RAM, so no conversions need occur. Wickedly fast and supported natively by Pandas (v0.3.1+)! Simply <code>conda install feather-format -c conda-forge</code>; then, to write a dataframe to feather:</p>\n\n<p><code>your_pandas_datafram.to_feather('./better_than_csv.ftr')</code></p>\n\n<p>And when you wanna read it back in,</p>\n\n<p><code>your_pandas_datafram = pd.read_feather('./better_than_csv.ftr')</code></p>\n\n<p>No need to mess with index=False, or whatever other params you usually throw in there since it's mirroring the DF as stored in memory. For comprison, loading 240MM rows from CSV takes more than than I'm willing to time it with, what with the parsing of datetimes, and float64s, etc. After doing all that stuff and getting things in your desired data types, reading the 240MM rows on my i8700K / SSD takes just 7-8 secs. Other languages and frameworks also have father support as well. You can read more about it -- <a href=\"https://duckduckgo.com/?q=pandas+feather+install\">https://duckduckgo.com/?q=pandas+feather+install</a> Hope this helps speed up your testing =)!</p>",
  "messages": [
    {
      "id": "316846",
      "postDate": "04/20/2018 04:10:08",
      "content": "<p>Been away from the competition for 6 days and my LB drops 254 spots? Brutal lol. Time to read every forum post and kernel to see what dark magix everyone's discovered this past week. Anyhow, onto the goods:</p>\n\n<p>There is a new format on the block called feather. It's more optimal than pickling, or--God forbid--raw .csvs. It essentially writes files to disk in exactly the format it's stored in RAM, so no conversions need occur. Wickedly fast and supported natively by Pandas (v0.3.1+)! Simply <code>conda install feather-format -c conda-forge</code>; then, to write a dataframe to feather:</p>\n\n<p><code>your_pandas_datafram.to_feather('./better_than_csv.ftr')</code></p>\n\n<p>And when you wanna read it back in,</p>\n\n<p><code>your_pandas_datafram = pd.read_feather('./better_than_csv.ftr')</code></p>\n\n<p>No need to mess with index=False, or whatever other params you usually throw in there since it's mirroring the DF as stored in memory. For comprison, loading 240MM rows from CSV takes more than than I'm willing to time it with, what with the parsing of datetimes, and float64s, etc. After doing all that stuff and getting things in your desired data types, reading the 240MM rows on my i8700K / SSD takes just 7-8 secs. Other languages and frameworks also have father support as well. You can read more about it -- <a href=\"https://duckduckgo.com/?q=pandas+feather+install\">https://duckduckgo.com/?q=pandas+feather+install</a> Hope this helps speed up your testing =)!</p>",
      "rawMarkdown": "Been away from the competition for 6 days and my LB drops 254 spots? Brutal lol. Time to read every forum post and kernel to see what dark magix everyone's discovered this past week. Anyhow, onto the goods:\n\nThere is a new format on the block called feather. It's more optimal than pickling, or--God forbid--raw .csvs. It essentially writes files to disk in exactly the format it's stored in RAM, so no conversions need occur. Wickedly fast and supported natively by Pandas (v0.3.1+)! Simply `conda install feather-format -c conda-forge`; then, to write a dataframe to feather:\n\n`your_pandas_datafram.to_feather('./better_than_csv.ftr')`\n\nAnd when you wanna read it back in,\n\n`your_pandas_datafram = pd.read_feather('./better_than_csv.ftr')`\n\nNo need to mess with index=False, or whatever other params you usually throw in there since it's mirroring the DF as stored in memory. For comprison, loading 240MM rows from CSV takes more than than I'm willing to time it with, what with the parsing of datetimes, and float64s, etc. After doing all that stuff and getting things in your desired data types, reading the 240MM rows on my i8700K / SSD takes just 7-8 secs. Other languages and frameworks also have father support as well. You can read more about it -- https://duckduckgo.com/?q=pandas+feather+install Hope this helps speed up your testing =)!",
      "votes": null
    },
    {
      "id": "316867",
      "postDate": "04/20/2018 05:12:34",
      "content": "<p>I will try. thanks</p>",
      "rawMarkdown": "I will try. thanks",
      "votes": null
    },
    {
      "id": "316905",
      "postDate": "04/20/2018 07:02:59",
      "content": "<p>Cool, sounds even better than hdf!</p>",
      "rawMarkdown": "Cool, sounds even better than hdf!",
      "votes": null
    },
    {
      "id": "316917",
      "postDate": "04/20/2018 07:25:48",
      "content": "<p>Awsome! Life is easier with this.</p>",
      "rawMarkdown": "Awsome! Life is easier with this.",
      "votes": null
    },
    {
      "id": "316939",
      "postDate": "04/20/2018 08:04:58",
      "content": "<p>OK, think feather beats hdf.</p>\n\n<p>On my machine, hdf takes 18.3 s to write the training dataframe, while feather takes 12.6s.</p>",
      "rawMarkdown": "OK, think feather beats hdf.\n\nOn my machine, hdf takes 18.3 s to write the training dataframe, while feather takes 12.6s.",
      "votes": null
    },
    {
      "id": "316965",
      "postDate": "04/20/2018 09:20:44",
      "content": "<p>I'm using feather via the mlcrate library from@anokas: <a href=\"https://github.com/mxbi/mlcrate\">https://github.com/mxbi/mlcrate</a></p>\n\n<p>Feather support was added at my request ;)</p>",
      "rawMarkdown": "I'm using feather via the mlcrate library from@anokas: https://github.com/mxbi/mlcrate\n\nFeather support was added at my request ;)",
      "votes": null
    },
    {
      "id": "317015",
      "postDate": "04/20/2018 13:24:09",
      "content": "<p>What about read time?  Feather is supposed to be optimized for read, not for write.</p>",
      "rawMarkdown": "What about read time?  Feather is supposed to be optimized for read, not for write.",
      "votes": null
    },
    {
      "id": "317277",
      "postDate": "04/21/2018 05:16:01",
      "content": "<p>One also has to care about storage, the size on disk. The feather will be 3x bigger ;-)\nI am using feather already, just pointing out the size matter for ppl that might use it in low storage environments.</p>",
      "rawMarkdown": "One also has to care about storage, the size on disk. The feather will be 3x bigger ;-)\nI am using feather already, just pointing out the size matter for ppl that might use it in low storage environments.",
      "votes": null
    },
    {
      "id": "317534",
      "postDate": "04/21/2018 20:06:23",
      "content": "<p>Thanks for sharing !</p>",
      "rawMarkdown": "Thanks for sharing !",
      "votes": null
    },
    {
      "id": "317656",
      "postDate": "04/22/2018 06:47:02",
      "content": "<p>FYI,  my .ftr file size is the same as the size showed in the end of dataframe.info()</p>",
      "rawMarkdown": "FYI,  my .ftr file size is the same as the size showed in the end of dataframe.info()",
      "votes": null
    },
    {
      "id": "317914",
      "postDate": "04/22/2018 20:27:47",
      "content": "<p>Thanks, you saved a lot of my time :)</p>",
      "rawMarkdown": "Thanks, you saved a lot of my time :)",
      "votes": null
    },
    {
      "id": "319652",
      "postDate": "04/26/2018 14:25:56",
      "content": "<p>This is a lifesaver. Thanks :)</p>",
      "rawMarkdown": "This is a lifesaver. Thanks :)",
      "votes": null
    },
    {
      "id": "319834",
      "postDate": "04/26/2018 23:27:27",
      "content": "<p>I wish I had known about this at the start of the competition! Thanks for sharing, very helpful.</p>",
      "rawMarkdown": "I wish I had known about this at the start of the competition! Thanks for sharing, very helpful.",
      "votes": null
    },
    {
      "id": "319988",
      "postDate": "04/27/2018 08:44:29",
      "content": "<p>Thanks for your share. </p>",
      "rawMarkdown": "Thanks for your share.",
      "votes": null
    },
    {
      "id": "320623",
      "postDate": "04/29/2018 10:03:27",
      "content": "<p>Is someone getting this error while writing dataframe to feather format?</p>\n\n<blockquote>\n  <p>Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array</p>\n</blockquote>",
      "rawMarkdown": "Is someone getting this error while writing dataframe to feather format?\n\n&gt; Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array",
      "votes": null
    },
    {
      "id": "320709",
      "postDate": "04/29/2018 15:08:43",
      "content": "<p>try resetting your df's index and lmk</p>",
      "rawMarkdown": "try resetting your df's index and lmk",
      "votes": null
    },
    {
      "id": "322018",
      "postDate": "05/02/2018 09:56:41",
      "content": "<p>which one is better?</p>",
      "rawMarkdown": "which one is better?",
      "votes": null
    },
    {
      "id": "322019",
      "postDate": "05/02/2018 09:57:52",
      "content": "<p>i see, i think this is very awsome</p>",
      "rawMarkdown": "i see, i think this is very awsome",
      "votes": null
    },
    {
      "id": "323199",
      "postDate": "05/04/2018 15:41:44",
      "content": "<p>Hi <a href=\"/authman\">@authman</a>, sorry but what does 'lmk' mean?</p>",
      "rawMarkdown": "Hi @authman, sorry but what does 'lmk' mean?",
      "votes": null
    },
    {
      "id": "323305",
      "postDate": "05/04/2018 19:58:26",
      "content": "<p>lmk = \"let me know\"</p>",
      "rawMarkdown": "lmk = \"let me know\"",
      "votes": null
    },
    {
      "id": "323328",
      "postDate": "05/04/2018 20:55:34",
      "content": "<p><a href=\"/inversion\">@inversion</a> - thank you, one more acronym in my bag )</p>",
      "rawMarkdown": "inversion - thank you, one more acronym in my bag )",
      "votes": null
    },
    {
      "id": "323582",
      "postDate": "05/05/2018 15:54:15",
      "content": "<p>tl;dr:</p>\n\n<ul>\n<li>Python: feather (Apache Arrow) does a great job</li>\n<li>R: fst does a great job, feather is \"useless\" due to not being updated for R / Julia</li>\n</ul>\n\n<p>Quoting for KaggleNoobs Slack (with edits) because this might be very interesting for anyone, we are using train (test_supplement added optionally) to benchmark:</p>\n\n<blockquote>\n  <p>[7d] Laurae: How fast are people loading the TalkingData training data after preprocessing (using any format)? Here it takes approx 3.5s to load the full 185M dataset (with date format on dates)</p>\n  \n  <p>[7d] miguel_perez: Just timed on laptop, 21 seconds to load those same 185M observations, also POSIXct format on dates, R.data on i7-6700HQ, so...  much less fast. Will try later with better hardware on AWS, and see how better... but 3.5 seconds sounds hard to beat</p>\n  \n  <p>[6d] cpmp: 3.11s wall time here, using mlcrate and feather.</p>\n  \n  <p>[6d] cpmp: This is for the concatenation of train and test_supplement.  Also the size is smaller than original csv.</p>\n  \n  <p>[6d] Laurae: I use fst format. <a href=\"/miguel\">@miguel</a>_perez has less than 1.5s for loading training data with it now (we are using regular SSDs, less than 600 MB/s, train fst less than 1GB)</p>\n  \n  <p>[6d] miguel_perez: Yup, 21 secs + @Laurae advice = 1.3 secs :-) huge difference between plain R binaries and fst in this data.</p>\n  \n  <p>[6d] Laurae: @cpmp R feather and Python feather are not fully interoperable, also the R/Julia feather package current status is frozen which means no more updates for nearly 1 year. btw my drive hardware is outdated (goes through a slow hardware RAID card), you should take <a href=\"/miguel\">@miguel</a>_perez timings to compare as he got a more realistic hardware scenario.\n  With another local RAID 0 of 4x 250GB 960 Pro it takes about 0.4s using R fst, 0.9s for Python feather (no RAM caching), and R feather takes over 6 seconds (reading the CSV takes 17 seconds). preprocessing was very simple: load CSV, then transform all date strings to date format.</p>\n</blockquote>\n\n<p>In my case / miguel_perez: we use the raw sets with dates converted to date format.</p>\n\n<p><a href=\"https://github.com/mxbi/mlcrate\">mlcrate</a> is <a href=\"/anokas\">@anokas</a> ' package.</p>\n\n<p><a href=\"https://cran.r-project.org/web/packages/fst/index.html\">fst</a> is R package for fast tabular data I/O using a special compressed multithreaded format:</p>\n\n<ul>\n<li>train: 7,360,986 KB =&gt; 857,218 KB</li>\n<li>test: 843,039 KB =&gt; 88,103 KB</li>\n<li>test_supplement: 2,603,064 KB =&gt; 254,563 KB</li>\n</ul>",
      "rawMarkdown": "tl;dr:\n\n* Python: feather (Apache Arrow) does a great job\n* R: fst does a great job, feather is \"useless\" due to not being updated for R / Julia\n\nQuoting for KaggleNoobs Slack (with edits) because this might be very interesting for anyone, we are using train (test_supplement added optionally) to benchmark:\n\n&gt; [7d] Laurae: How fast are people loading the TalkingData training data after preprocessing (using any format)? Here it takes approx 3.5s to load the full 185M dataset (with date format on dates)\n\n&gt; [7d] miguel_perez: Just timed on laptop, 21 seconds to load those same 185M observations, also POSIXct format on dates, R.data on i7-6700HQ, so...  much less fast. Will try later with better hardware on AWS, and see how better... but 3.5 seconds sounds hard to beat\n\n&gt; [6d] cpmp: 3.11s wall time here, using mlcrate and feather.\n\n&gt; [6d] cpmp: This is for the concatenation of train and test_supplement.  Also the size is smaller than original csv.\n\n&gt; [6d] Laurae: I use fst format. @miguel_perez has less than 1.5s for loading training data with it now (we are using regular SSDs, less than 600 MB/s, train fst less than 1GB)\n\n&gt; [6d] miguel_perez: Yup, 21 secs + @Laurae advice = 1.3 secs :-) huge difference between plain R binaries and fst in this data.\n\n&gt; [6d] Laurae: @cpmp R feather and Python feather are not fully interoperable, also the R/Julia feather package current status is frozen which means no more updates for nearly 1 year. btw my drive hardware is outdated (goes through a slow hardware RAID card), you should take @miguel_perez timings to compare as he got a more realistic hardware scenario.\nWith another local RAID 0 of 4x 250GB 960 Pro it takes about 0.4s using R fst, 0.9s for Python feather (no RAM caching), and R feather takes over 6 seconds (reading the CSV takes 17 seconds). preprocessing was very simple: load CSV, then transform all date strings to date format.\n\nIn my case / miguel_perez: we use the raw sets with dates converted to date format.\n\n[mlcrate](https://github.com/mxbi/mlcrate) is @anokas ' package.\n\n[fst](https://cran.r-project.org/web/packages/fst/index.html) is R package for fast tabular data I/O using a special compressed multithreaded format:\n\n* train: 7,360,986 KB =&gt; 857,218 KB\n* test: 843,039 KB =&gt; 88,103 KB\n* test_supplement: 2,603,064 KB =&gt; 254,563 KB",
      "votes": null
    },
    {
      "id": "323850",
      "postDate": "05/06/2018 12:06:49",
      "content": "<p>Unfortunatley I am getting the same error as Burhan.</p>\n\n<p>\"Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array\"</p>\n\n<p>Resetting the index didnt help.\nany other ideas?</p>",
      "rawMarkdown": "Unfortunatley I am getting the same error as Burhan.\n\n\"Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array\"\n\nResetting the index didnt help.\nany other ideas?",
      "votes": null
    },
    {
      "id": "324423",
      "postDate": "05/07/2018 17:25:13",
      "content": "<p>Hey <a href=\"/authman\">@authman</a>, \nI'm really sorry for not responding to that. I never tried that fix.\nIn fact, i had rather used the mlcrate package suggested by CPMP.</p>",
      "rawMarkdown": "Hey @authman, \nI'm really sorry for not responding to that. I never tried that fix.\nIn fact, i had rather used the mlcrate package suggested by CPMP.",
      "votes": null
    },
    {
      "id": "324426",
      "postDate": "05/07/2018 17:27:42",
      "content": "<p>You could try using the mlcrate package for python. It has methods for saving and loading data as feather files. That worked for me.</p>",
      "rawMarkdown": "You could try using the mlcrate package for python. It has methods for saving and loading data as feather files. That worked for me.",
      "votes": null
    },
    {
      "id": "1505788",
      "postDate": "09/07/2021 14:50:21",
      "content": "<p>I am following you👍👍</p>",
      "rawMarkdown": "I am following you👍👍",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1505788,
      "author_name": "saliblue",
      "author_url": "",
      "post_date": "09/07/2021 14:50:21",
      "content": "<p>I am following you👍👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 316867,
      "author_name": "liuhdsgoal",
      "author_url": "",
      "post_date": "04/20/2018 05:12:34",
      "content": "<p>I will try. thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 316905,
      "author_name": "enfeizhan",
      "author_url": "",
      "post_date": "04/20/2018 07:02:59",
      "content": "<p>Cool, sounds even better than hdf!</p>",
      "votes": null,
      "replies": [
        {
          "id": 316939,
          "author_name": "enfeizhan",
          "author_url": "",
          "post_date": "04/20/2018 08:04:58",
          "content": "<p>OK, think feather beats hdf.</p>\n\n<p>On my machine, hdf takes 18.3 s to write the training dataframe, while feather takes 12.6s.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 317015,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/20/2018 13:24:09",
          "content": "<p>What about read time?  Feather is supposed to be optimized for read, not for write.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 316917,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "04/20/2018 07:25:48",
      "content": "<p>Awsome! Life is easier with this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 316965,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/20/2018 09:20:44",
      "content": "<p>I'm using feather via the mlcrate library from@anokas: <a href=\"https://github.com/mxbi/mlcrate\">https://github.com/mxbi/mlcrate</a></p>\n\n<p>Feather support was added at my request ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 322018,
          "author_name": "marvinxu",
          "author_url": "",
          "post_date": "05/02/2018 09:56:41",
          "content": "<p>which one is better?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 322019,
          "author_name": "marvinxu",
          "author_url": "",
          "post_date": "05/02/2018 09:57:52",
          "content": "<p>i see, i think this is very awsome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 317277,
      "author_name": "profetul",
      "author_url": "",
      "post_date": "04/21/2018 05:16:01",
      "content": "<p>One also has to care about storage, the size on disk. The feather will be 3x bigger ;-)\nI am using feather already, just pointing out the size matter for ppl that might use it in low storage environments.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 317534,
      "author_name": "nathanlauga",
      "author_url": "",
      "post_date": "04/21/2018 20:06:23",
      "content": "<p>Thanks for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 317656,
      "author_name": "yyqing",
      "author_url": "",
      "post_date": "04/22/2018 06:47:02",
      "content": "<p>FYI,  my .ftr file size is the same as the size showed in the end of dataframe.info()</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 317914,
      "author_name": "walkerous",
      "author_url": "",
      "post_date": "04/22/2018 20:27:47",
      "content": "<p>Thanks, you saved a lot of my time :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 319652,
      "author_name": "combalgorythm",
      "author_url": "",
      "post_date": "04/26/2018 14:25:56",
      "content": "<p>This is a lifesaver. Thanks :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 319834,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/26/2018 23:27:27",
      "content": "<p>I wish I had known about this at the start of the competition! Thanks for sharing, very helpful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 319988,
      "author_name": "mcggood",
      "author_url": "",
      "post_date": "04/27/2018 08:44:29",
      "content": "<p>Thanks for your share. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320623,
      "author_name": "burhanusm",
      "author_url": "",
      "post_date": "04/29/2018 10:03:27",
      "content": "<p>Is someone getting this error while writing dataframe to feather format?</p>\n\n<blockquote>\n  <p>Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 320709,
          "author_name": "authman",
          "author_url": "",
          "post_date": "04/29/2018 15:08:43",
          "content": "<p>try resetting your df's index and lmk</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323199,
          "author_name": "kruegger",
          "author_url": "",
          "post_date": "05/04/2018 15:41:44",
          "content": "<p>Hi <a href=\"/authman\">@authman</a>, sorry but what does 'lmk' mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323305,
          "author_name": "inversion",
          "author_url": "",
          "post_date": "05/04/2018 19:58:26",
          "content": "<p>lmk = \"let me know\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 323328,
          "author_name": "kruegger",
          "author_url": "",
          "post_date": "05/04/2018 20:55:34",
          "content": "<p><a href=\"/inversion\">@inversion</a> - thank you, one more acronym in my bag )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324423,
          "author_name": "burhanusm",
          "author_url": "",
          "post_date": "05/07/2018 17:25:13",
          "content": "<p>Hey <a href=\"/authman\">@authman</a>, \nI'm really sorry for not responding to that. I never tried that fix.\nIn fact, i had rather used the mlcrate package suggested by CPMP.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 323582,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "05/05/2018 15:54:15",
      "content": "<p>tl;dr:</p>\n\n<ul>\n<li>Python: feather (Apache Arrow) does a great job</li>\n<li>R: fst does a great job, feather is \"useless\" due to not being updated for R / Julia</li>\n</ul>\n\n<p>Quoting for KaggleNoobs Slack (with edits) because this might be very interesting for anyone, we are using train (test_supplement added optionally) to benchmark:</p>\n\n<blockquote>\n  <p>[7d] Laurae: How fast are people loading the TalkingData training data after preprocessing (using any format)? Here it takes approx 3.5s to load the full 185M dataset (with date format on dates)</p>\n  \n  <p>[7d] miguel_perez: Just timed on laptop, 21 seconds to load those same 185M observations, also POSIXct format on dates, R.data on i7-6700HQ, so...  much less fast. Will try later with better hardware on AWS, and see how better... but 3.5 seconds sounds hard to beat</p>\n  \n  <p>[6d] cpmp: 3.11s wall time here, using mlcrate and feather.</p>\n  \n  <p>[6d] cpmp: This is for the concatenation of train and test_supplement.  Also the size is smaller than original csv.</p>\n  \n  <p>[6d] Laurae: I use fst format. <a href=\"/miguel\">@miguel</a>_perez has less than 1.5s for loading training data with it now (we are using regular SSDs, less than 600 MB/s, train fst less than 1GB)</p>\n  \n  <p>[6d] miguel_perez: Yup, 21 secs + @Laurae advice = 1.3 secs :-) huge difference between plain R binaries and fst in this data.</p>\n  \n  <p>[6d] Laurae: @cpmp R feather and Python feather are not fully interoperable, also the R/Julia feather package current status is frozen which means no more updates for nearly 1 year. btw my drive hardware is outdated (goes through a slow hardware RAID card), you should take <a href=\"/miguel\">@miguel</a>_perez timings to compare as he got a more realistic hardware scenario.\n  With another local RAID 0 of 4x 250GB 960 Pro it takes about 0.4s using R fst, 0.9s for Python feather (no RAM caching), and R feather takes over 6 seconds (reading the CSV takes 17 seconds). preprocessing was very simple: load CSV, then transform all date strings to date format.</p>\n</blockquote>\n\n<p>In my case / miguel_perez: we use the raw sets with dates converted to date format.</p>\n\n<p><a href=\"https://github.com/mxbi/mlcrate\">mlcrate</a> is <a href=\"/anokas\">@anokas</a> ' package.</p>\n\n<p><a href=\"https://cran.r-project.org/web/packages/fst/index.html\">fst</a> is R package for fast tabular data I/O using a special compressed multithreaded format:</p>\n\n<ul>\n<li>train: 7,360,986 KB =&gt; 857,218 KB</li>\n<li>test: 843,039 KB =&gt; 88,103 KB</li>\n<li>test_supplement: 2,603,064 KB =&gt; 254,563 KB</li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 323850,
      "author_name": "martykuentzel",
      "author_url": "",
      "post_date": "05/06/2018 12:06:49",
      "content": "<p>Unfortunatley I am getting the same error as Burhan.</p>\n\n<p>\"Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array\"</p>\n\n<p>Resetting the index didnt help.\nany other ideas?</p>",
      "votes": null,
      "replies": [
        {
          "id": 324426,
          "author_name": "burhanusm",
          "author_url": "",
          "post_date": "05/07/2018 17:27:42",
          "content": "<p>You could try using the mlcrate package for python. It has methods for saving and loading data as feather files. That worked for me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "316846": "Been away from the competition for 6 days and my LB drops 254 spots? Brutal lol. Time to read every forum post and kernel to see what dark magix everyone's discovered this past week. Anyhow, onto the goods:\n\nThere is a new format on the block called feather. It's more optimal than pickling, or--God forbid--raw .csvs. It essentially writes files to disk in exactly the format it's stored in RAM, so no conversions need occur. Wickedly fast and supported natively by Pandas (v0.3.1+)! Simply `conda install feather-format -c conda-forge`; then, to write a dataframe to feather:\n\n`your_pandas_datafram.to_feather('./better_than_csv.ftr')`\n\nAnd when you wanna read it back in,\n\n`your_pandas_datafram = pd.read_feather('./better_than_csv.ftr')`\n\nNo need to mess with index=False, or whatever other params you usually throw in there since it's mirroring the DF as stored in memory. For comprison, loading 240MM rows from CSV takes more than than I'm willing to time it with, what with the parsing of datetimes, and float64s, etc. After doing all that stuff and getting things in your desired data types, reading the 240MM rows on my i8700K / SSD takes just 7-8 secs. Other languages and frameworks also have father support as well. You can read more about it -- https://duckduckgo.com/?q=pandas+feather+install Hope this helps speed up your testing =)!",
    "316867": "I will try. thanks",
    "316905": "Cool, sounds even better than hdf!",
    "316917": "Awsome! Life is easier with this.",
    "316939": "OK, think feather beats hdf.\n\nOn my machine, hdf takes 18.3 s to write the training dataframe, while feather takes 12.6s.",
    "316965": "I'm using feather via the mlcrate library from@anokas: https://github.com/mxbi/mlcrate\n\nFeather support was added at my request ;)",
    "317015": "What about read time?  Feather is supposed to be optimized for read, not for write.",
    "317277": "One also has to care about storage, the size on disk. The feather will be 3x bigger ;-)\nI am using feather already, just pointing out the size matter for ppl that might use it in low storage environments.",
    "317534": "Thanks for sharing !",
    "317656": "FYI,  my .ftr file size is the same as the size showed in the end of dataframe.info()",
    "317914": "Thanks, you saved a lot of my time :)",
    "319652": "This is a lifesaver. Thanks :)",
    "319834": "I wish I had known about this at the start of the competition! Thanks for sharing, very helpful.",
    "319988": "Thanks for your share.",
    "320623": "Is someone getting this error while writing dataframe to feather format?\n\n&gt; Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array",
    "320709": "try resetting your df's index and lmk",
    "322018": "which one is better?",
    "322019": "i see, i think this is very awsome",
    "323199": "Hi @authman, sorry but what does 'lmk' mean?",
    "323305": "lmk = \"let me know\"",
    "323328": "inversion - thank you, one more acronym in my bag )",
    "323582": "tl;dr:\n\n* Python: feather (Apache Arrow) does a great job\n* R: fst does a great job, feather is \"useless\" due to not being updated for R / Julia\n\nQuoting for KaggleNoobs Slack (with edits) because this might be very interesting for anyone, we are using train (test_supplement added optionally) to benchmark:\n\n&gt; [7d] Laurae: How fast are people loading the TalkingData training data after preprocessing (using any format)? Here it takes approx 3.5s to load the full 185M dataset (with date format on dates)\n\n&gt; [7d] miguel_perez: Just timed on laptop, 21 seconds to load those same 185M observations, also POSIXct format on dates, R.data on i7-6700HQ, so...  much less fast. Will try later with better hardware on AWS, and see how better... but 3.5 seconds sounds hard to beat\n\n&gt; [6d] cpmp: 3.11s wall time here, using mlcrate and feather.\n\n&gt; [6d] cpmp: This is for the concatenation of train and test_supplement.  Also the size is smaller than original csv.\n\n&gt; [6d] Laurae: I use fst format. @miguel_perez has less than 1.5s for loading training data with it now (we are using regular SSDs, less than 600 MB/s, train fst less than 1GB)\n\n&gt; [6d] miguel_perez: Yup, 21 secs + @Laurae advice = 1.3 secs :-) huge difference between plain R binaries and fst in this data.\n\n&gt; [6d] Laurae: @cpmp R feather and Python feather are not fully interoperable, also the R/Julia feather package current status is frozen which means no more updates for nearly 1 year. btw my drive hardware is outdated (goes through a slow hardware RAID card), you should take @miguel_perez timings to compare as he got a more realistic hardware scenario.\nWith another local RAID 0 of 4x 250GB 960 Pro it takes about 0.4s using R fst, 0.9s for Python feather (no RAM caching), and R feather takes over 6 seconds (reading the CSV takes 17 seconds). preprocessing was very simple: load CSV, then transform all date strings to date format.\n\nIn my case / miguel_perez: we use the raw sets with dates converted to date format.\n\n[mlcrate](https://github.com/mxbi/mlcrate) is @anokas ' package.\n\n[fst](https://cran.r-project.org/web/packages/fst/index.html) is R package for fast tabular data I/O using a special compressed multithreaded format:\n\n* train: 7,360,986 KB =&gt; 857,218 KB\n* test: 843,039 KB =&gt; 88,103 KB\n* test_supplement: 2,603,064 KB =&gt; 254,563 KB",
    "323850": "Unfortunatley I am getting the same error as Burhan.\n\n\"Cannot convert pyarrow.lib.ChunkedArray to pyarrow.lib.Array\"\n\nResetting the index didnt help.\nany other ideas?",
    "324423": "Hey @authman, \nI'm really sorry for not responding to that. I never tried that fix.\nIn fact, i had rather used the mlcrate package suggested by CPMP.",
    "324426": "You could try using the mlcrate package for python. It has methods for saving and loading data as feather files. That worked for me.",
    "1505788": "I am following you👍👍"
  },
  "source": "meta"
}