{
  "id": 190288,
  "title": "Read train.csv in 10 sec",
  "url": "/competitions/riiid-test-answer-prediction/discussion/190288",
  "author_name": "ONODERA",
  "post_date": "2020-10-11T04:17:48.873000",
  "votes": 167,
  "comment_count": 54,
  "views": 0,
  "content": "<p>Using <a href=\"https://docs.rapids.ai/api/cudf/stable/\" target=\"_blank\">cuDF</a> you can read train.csv in <strong>10 seconds</strong> in <a href=\"https://www.kaggle.com/onodera/riiid-read-csv-in-cudf\" target=\"_blank\">kaggle notebook</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F374bc1dee5812ff949677eb07e5ef8b5%2F2020-10-11%2013.13.38.png?generation=1602389639041354&amp;alt=media\" alt=\"\"></p>\n<p>In my environment, pandas takes <strong>1 minutes</strong>, but cuDF only takes <strong>3 seconds</strong>!!! <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F9796020e8bc09be7d359cb83176b65d7%2FEkAJfR_U0AEQhAa.jpeg?generation=1602389509691312&amp;alt=media\" alt=\"\"></p>\n<p>Life is short, use cuDF</p>",
  "messages": [
    {
      "id": 1045804,
      "postDate": "2020-10-11T04:17:48.873Z",
      "content": "<p>Using <a href=\"https://docs.rapids.ai/api/cudf/stable/\" target=\"_blank\">cuDF</a> you can read train.csv in <strong>10 seconds</strong> in <a href=\"https://www.kaggle.com/onodera/riiid-read-csv-in-cudf\" target=\"_blank\">kaggle notebook</a>.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F374bc1dee5812ff949677eb07e5ef8b5%2F2020-10-11%2013.13.38.png?generation=1602389639041354&amp;alt=media\" alt=\"\"></p>\n<p>In my environment, pandas takes <strong>1 minutes</strong>, but cuDF only takes <strong>3 seconds</strong>!!! <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F9796020e8bc09be7d359cb83176b65d7%2FEkAJfR_U0AEQhAa.jpeg?generation=1602389509691312&amp;alt=media\" alt=\"\"></p>\n<p>Life is short, use cuDF</p>",
      "rawMarkdown": "Using [cuDF](https://docs.rapids.ai/api/cudf/stable/) you can read train.csv in **10 seconds** in [kaggle notebook](https://www.kaggle.com/onodera/riiid-read-csv-in-cudf).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F374bc1dee5812ff949677eb07e5ef8b5%2F2020-10-11%2013.13.38.png?generation=1602389639041354&alt=media)\n\n\n\nIn my environment, pandas takes **1 minutes**, but cuDF only takes **3 seconds**!!! \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F9796020e8bc09be7d359cb83176b65d7%2FEkAJfR_U0AEQhAa.jpeg?generation=1602389509691312&alt=media)\n\n\nLife is short, use cuDF",
      "votes": 166
    },
    {
      "id": 1049661,
      "postDate": "2020-10-14T16:11:59.487Z",
      "content": "<p>Thank you for your help.  Using cudf may be a good solution especially when having a good gpu.<br>\nHowever I think that getting the data into jay format and then converting it to pandas using datatable as shown in <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">this wonderful notebook</a> is more effective.<br>\n<a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">This notebook </a>is very good resource of knowledge regarding the different ways to import a large datasets. I recommend to visit it.</p>",
      "rawMarkdown": "Thank you for your help.  Using cudf may be a good solution especially when having a good gpu.\nHowever I think that getting the data into jay format and then converting it to pandas using datatable as shown in [this wonderful notebook](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid) is more effective.\n[This notebook ](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets)is very good resource of knowledge regarding the different ways to import a large datasets. I recommend to visit it.",
      "votes": 9
    },
    {
      "id": 1046464,
      "postDate": "2020-10-11T17:18:43.937Z",
      "content": "<p>Yassine - not all is lost (though I agree a GPU would be nice.)</p>\n<pre><code>data_types_dict = ...\ndf = pd.read_csv(\"train.csv\", dtype=data_types_dict, index_col=0)\ndf.to_pickle(\"train.pkl\", protocol=pickle.HIGHEST_PROTOCOL)\n</code></pre>\n<pre><code>%%time\ntrain_df = pd.read_pickle(\"train.pkl\")\n</code></pre>\n<pre><code>CPU times: user 0 ns, sys: 1.17 s, total: 1.17 s Wall time: 1.17 s\n</code></pre>",
      "rawMarkdown": "Yassine - not all is lost (though I agree a GPU would be nice.)\n\n\n```\ndata_types_dict = ...\ndf = pd.read_csv(\"train.csv\", dtype=data_types_dict, index_col=0)\ndf.to_pickle(\"train.pkl\", protocol=pickle.HIGHEST_PROTOCOL)\n\n```\n```\n%%time\ntrain_df = pd.read_pickle(\"train.pkl\")\n\n```\n```\nCPU times: user 0 ns, sys: 1.17 s, total: 1.17 s Wall time: 1.17 s\n\n```\n",
      "votes": 3,
      "replies": [
        {
          "id": 1046488,
          "postDate": "2020-10-11T17:49:45.063Z",
          "content": "<p>I recommend you to replace pd with cudf<br>\n<code>df = cudf.read_csv(\"train.csv\")</code><br>\nand then write it into pkl or feather</p>",
          "rawMarkdown": "I recommend you to replace pd with cudf\n`df = cudf.read_csv(\"train.csv\")`\nand then write it into pkl or feather",
          "votes": 4
        },
        {
          "id": 1063679,
          "postDate": "2020-10-29T07:49:40.637Z",
          "content": "<p><a href=\"https://www.kaggle.com/roman99\" target=\"_blank\">@roman99</a>  are we able to read all data into memory  during read csv  and read pickle?</p>",
          "rawMarkdown": "@roman99  are we able to read all data into memory  during read csv  and read pickle?"
        }
      ]
    },
    {
      "id": 1047127,
      "postDate": "2020-10-12T09:33:02.243Z",
      "content": "<p>Could you please tell in your own words why cudf is so much faster than pandas?:) Thanx</p>",
      "rawMarkdown": "Could you please tell in your own words why cudf is so much faster than pandas?:) Thanx",
      "votes": 1,
      "replies": [
        {
          "id": 1047164,
          "postDate": "2020-10-12T10:31:46.427Z",
          "content": "<p>Because cudf is a GPU accelerated library build on top of CUDA with very similar API to pandas. The following link will help you get a better understanding of cudf and the Rapids ecosystem <br>\n<a href=\"url\" target=\"_blank\">https://rapids.ai/about.html</a></p>",
          "rawMarkdown": "Because cudf is a GPU accelerated library build on top of CUDA with very similar API to pandas. The following link will help you get a better understanding of cudf and the Rapids ecosystem \n[https://rapids.ai/about.html](url)",
          "votes": 3
        },
        {
          "id": 1047182,
          "postDate": "2020-10-12T10:56:02.680Z",
          "content": "<p>Hello, Aravind! Thanks for your answer! I will check provided resource:)</p>\n<p>By the way, does the using of cudf will decrease my Kaggle's GPU limit?:)</p>",
          "rawMarkdown": "Hello, Aravind! Thanks for your answer! I will check provided resource:)\n\nBy the way, does the using of cudf will decrease my Kaggle's GPU limit?:)"
        },
        {
          "id": 1047215,
          "postDate": "2020-10-12T11:27:28.440Z",
          "content": "<blockquote>\n  <p>does the using of cudf will decrease my Kaggle's GPU limit?</p>\n</blockquote>\n<p>Yes</p>",
          "rawMarkdown": "> does the using of cudf will decrease my Kaggle's GPU limit?\n\nYes",
          "votes": 1
        }
      ]
    },
    {
      "id": 1046537,
      "postDate": "2020-10-11T18:46:26.420Z",
      "content": "<p>sadly, converting <code>train</code> back to pandas gives an out of memory error.<br>\nBut great library, thanks for introducing it. Will explore more.</p>",
      "rawMarkdown": "sadly, converting `train` back to pandas gives an out of memory error.\nBut great library, thanks for introducing it. Will explore more.",
      "votes": 1
    },
    {
      "id": 1045934,
      "postDate": "2020-10-11T06:54:58.800Z",
      "content": "<p>Adding to your point, XGBoost integrates with cuDF, Dask, and the entire RAPIDS ecosystem. Which will enable us to buid faster models by harvesting the power of GPU</p>",
      "rawMarkdown": "Adding to your point, XGBoost integrates with cuDF, Dask, and the entire RAPIDS ecosystem. Which will enable us to buid faster models by harvesting the power of GPU",
      "votes": 1
    },
    {
      "id": 1083876,
      "postDate": "2020-11-19T13:40:21.470Z",
      "content": "<p>On my desktop computer, with pandas, reading the whole dataset in <strong>feather format</strong> takes <strong>less than 5 seconds</strong> and writing it to feather format takes about 23 seconds :<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F277616%2F149912a60fc0a349f539bace6ed6a828%2F2020-11-19%2014_39_44-Window.png?generation=1605793212318961&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "On my desktop computer, with pandas, reading the whole dataset in **feather format** takes **less than 5 seconds** and writing it to feather format takes about 23 seconds :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F277616%2F149912a60fc0a349f539bace6ed6a828%2F2020-11-19%2014_39_44-Window.png?generation=1605793212318961&alt=media)",
      "votes": 2
    },
    {
      "id": 1045890,
      "postDate": "2020-10-11T06:13:28.037Z",
      "content": "<p>I wish cudf came with a GPU attached. ;)<br>\nIndeed, life is short, so only waste it in what matters. </p>",
      "rawMarkdown": "I wish cudf came with a GPU attached. ;)\nIndeed, life is short, so only waste it in what matters. ",
      "votes": 2,
      "replies": [
        {
          "id": 1053912,
          "postDate": "2020-10-19T13:32:49.257Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1054131,
      "postDate": "2020-10-19T17:50:17.467Z",
      "content": "<p>Installation takes forever. Is that only me? I'm referring <a href=\"https://rapids.ai/start.html\" target=\"_blank\">https://rapids.ai/start.html</a> with conda command.</p>",
      "rawMarkdown": "Installation takes forever. Is that only me? I'm referring https://rapids.ai/start.html with conda command.",
      "replies": [
        {
          "id": 1054428,
          "postDate": "2020-10-19T23:33:20.293Z",
          "content": "<p>Did you create environment? and run the command on the environment?<br>\nYou can create and activate a environment with following commands;</p>\n<pre><code>conda create -n foo python=3.7\nconda activate foo\n</code></pre>",
          "rawMarkdown": "Did you create environment? and run the command on the environment?\nYou can create and activate a environment with following commands;\n```\nconda create -n foo python=3.7\nconda activate foo\n```"
        }
      ]
    },
    {
      "id": 1054031,
      "postDate": "2020-10-19T15:45:04.427Z",
      "content": "<p>Hi how do you install cudf with pip on Kaggle notebook?</p>",
      "rawMarkdown": "Hi how do you install cudf with pip on Kaggle notebook?",
      "replies": [
        {
          "id": 1054039,
          "postDate": "2020-10-19T15:55:27.887Z",
          "content": "<p>I found thank !</p>",
          "rawMarkdown": "I found thank !"
        }
      ]
    },
    {
      "id": 1052469,
      "postDate": "2020-10-17T18:54:20.873Z",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>, It's the firs time to hear about that library </p>",
      "rawMarkdown": "thanks for sharing @onodera, It's the firs time to hear about that library "
    },
    {
      "id": 1052140,
      "postDate": "2020-10-17T11:23:49.363Z",
      "content": "<p>That's a great way. Thanks for sharing. <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> </p>",
      "rawMarkdown": "That's a great way. Thanks for sharing. @onodera "
    },
    {
      "id": 1052021,
      "postDate": "2020-10-17T07:43:48.610Z",
      "content": "<p><a href=\"https://www.kaggle.com/onodere\" target=\"_blank\">@onodere</a> <br>\nLike you said loading the data in cudf makes in faster than pandas.<br>\nBut even in cudf the time taken for one feature engineering is 1 minute.<br>\nHow are you overcoming this??</p>",
      "rawMarkdown": "@onodere \nLike you said loading the data in cudf makes in faster than pandas.\nBut even in cudf the time taken for one feature engineering is 1 minute.\nHow are you overcoming this??",
      "replies": [
        {
          "id": 1052596,
          "postDate": "2020-10-18T01:40:27.237Z",
          "content": "<p>Depends on the feature</p>",
          "rawMarkdown": "Depends on the feature"
        }
      ]
    },
    {
      "id": 1050913,
      "postDate": "2020-10-15T22:10:25.013Z",
      "content": "<p>Awesome work buddy</p>",
      "rawMarkdown": "Awesome work buddy"
    },
    {
      "id": 1050688,
      "postDate": "2020-10-15T16:19:47.767Z",
      "content": "<p>I wonder what Radeon has to offer in terms of ML</p>",
      "rawMarkdown": "I wonder what Radeon has to offer in terms of ML"
    },
    {
      "id": 1050617,
      "postDate": "2020-10-15T15:31:01.280Z",
      "content": "<p>Thanks for sharing sir, now I go search for cudf, kaggle give something new every single day,</p>",
      "rawMarkdown": "Thanks for sharing sir, now I go search for cudf, kaggle give something new every single day,"
    },
    {
      "id": 1050551,
      "postDate": "2020-10-15T14:30:40.653Z",
      "content": "<p>Thanks a lot for sharing. Should give it a try. </p>",
      "rawMarkdown": "Thanks a lot for sharing. Should give it a try. "
    },
    {
      "id": 1048791,
      "postDate": "2020-10-13T19:31:37.953Z",
      "content": "<p>Fantastic!</p>",
      "rawMarkdown": "Fantastic!"
    },
    {
      "id": 1047869,
      "postDate": "2020-10-13T02:44:23.717Z",
      "content": "<p>Nice! Can cudf be installed with pip ?</p>",
      "rawMarkdown": "Nice! Can cudf be installed with pip ?",
      "replies": [
        {
          "id": 1048001,
          "postDate": "2020-10-13T05:48:39.970Z",
          "content": "<p>Refer to <a href=\"https://rapids.ai/start.html\" target=\"_blank\">here</a></p>",
          "rawMarkdown": "Refer to [here](https://rapids.ai/start.html)",
          "votes": 2
        },
        {
          "id": 1048373,
          "postDate": "2020-10-13T12:32:12.347Z",
          "content": "<p>When I use <code>XXX.apply()</code>, I got <code>AttributeError: 'Series' object has no attribute 'apply'</code><br>\nAre there any alternatives to the <code>.apply</code> method? Or it's just not being implemented yet.</p>",
          "rawMarkdown": "When I use `XXX.apply()`, I got `AttributeError: 'Series' object has no attribute 'apply'`\nAre there any alternatives to the `.apply` method? Or it's just not being implemented yet."
        },
        {
          "id": 1048408,
          "postDate": "2020-10-13T13:11:09.400Z",
          "content": "<p>How about \"<strong>applymap</strong>\" ?</p>",
          "rawMarkdown": "How about \"**applymap**\" ?",
          "votes": 1
        },
        {
          "id": 1048423,
          "postDate": "2020-10-13T13:32:57.240Z",
          "content": "<p>Thanks, that's what I want.</p>",
          "rawMarkdown": "Thanks, that's what I want."
        }
      ]
    },
    {
      "id": 1047784,
      "postDate": "2020-10-12T23:48:54.703Z",
      "content": "<p>Сompare the capabilities of the libraries, not just download the file.</p>",
      "rawMarkdown": "Сompare the capabilities of the libraries, not just download the file."
    },
    {
      "id": 1047748,
      "postDate": "2020-10-12T22:30:19.607Z",
      "content": "<p>Are you using Colab or a local environment (if so, what GPU are you using?)</p>",
      "rawMarkdown": "Are you using Colab or a local environment (if so, what GPU are you using?)",
      "replies": [
        {
          "id": 1047775,
          "postDate": "2020-10-12T23:17:37.303Z",
          "content": "<p>local, V100</p>",
          "rawMarkdown": "local, V100",
          "votes": 1
        }
      ]
    },
    {
      "id": 1047605,
      "postDate": "2020-10-12T18:37:07.837Z",
      "content": "<p>Thanks for sharing nice info. Is there a way to converting from cudf to dataframe faster ?</p>",
      "rawMarkdown": "Thanks for sharing nice info. Is there a way to converting from cudf to dataframe faster ?"
    },
    {
      "id": 1047238,
      "postDate": "2020-10-12T11:56:18.713Z",
      "content": "<p>Thanks for sharing…Will explore more on this.</p>",
      "rawMarkdown": "Thanks for sharing...Will explore more on this."
    },
    {
      "id": 1045812,
      "postDate": "2020-10-11T04:31:10.073Z",
      "content": "<p>I liked the quote btw😂</p>",
      "rawMarkdown": "I liked the quote btw😂"
    },
    {
      "id": 1092307,
      "postDate": "2020-11-26T17:26:28.790Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1053698,
      "postDate": "2020-10-19T08:48:30.997Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1048075,
      "postDate": "2020-10-13T06:57:27.937Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1047541,
      "postDate": "2020-10-12T17:40:41.600Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1050431,
      "postDate": "2020-10-15T12:14:16.233Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> </p>",
      "rawMarkdown": "Thanks for sharing @onodera ",
      "votes": 1
    },
    {
      "id": 1056382,
      "postDate": "2020-10-21T16:37:07.160Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1054261,
      "postDate": "2020-10-19T19:30:49.747Z",
      "content": "<p>Thanks a lot for sharing this.</p>",
      "rawMarkdown": "Thanks a lot for sharing this."
    },
    {
      "id": 1053940,
      "postDate": "2020-10-19T14:01:06.927Z",
      "content": "<p>Thanks for sharing Man 👍</p>",
      "rawMarkdown": "Thanks for sharing Man 👍"
    },
    {
      "id": 1053742,
      "postDate": "2020-10-19T09:54:28.847Z",
      "content": "<p>thanks for sharing </p>",
      "rawMarkdown": "thanks for sharing "
    },
    {
      "id": 1052758,
      "postDate": "2020-10-18T08:09:09.163Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a></p>",
      "rawMarkdown": "Thanks for sharing @onodera"
    },
    {
      "id": 1052346,
      "postDate": "2020-10-17T16:09:17.123Z",
      "content": "<p>Thank you for sharing , very helpful</p>",
      "rawMarkdown": "Thank you for sharing , very helpful"
    },
    {
      "id": 1052105,
      "postDate": "2020-10-17T10:13:43.270Z",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you."
    },
    {
      "id": 1051985,
      "postDate": "2020-10-17T06:29:16.240Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 1050015,
      "postDate": "2020-10-15T01:34:01.370Z",
      "content": "<p>Thanks for sharing this..</p>",
      "rawMarkdown": "Thanks for sharing this.."
    },
    {
      "id": 1049750,
      "postDate": "2020-10-14T18:47:50.140Z",
      "content": "<p>Thank you for your help.</p>",
      "rawMarkdown": "Thank you for your help."
    },
    {
      "id": 1048470,
      "postDate": "2020-10-13T14:17:42.840Z",
      "content": "<p>Thanks for mentioning this</p>",
      "rawMarkdown": "Thanks for mentioning this"
    },
    {
      "id": 1045872,
      "postDate": "2020-10-11T05:58:57.317Z",
      "content": "<p>thank you for sharing:)</p>",
      "rawMarkdown": "thank you for sharing:)"
    }
  ],
  "comments": [
    {
      "id": 1049661,
      "author_name": "sami",
      "author_url": "",
      "post_date": "2020-10-14T16:11:59.487000",
      "content": "<p>Thank you for your help.  Using cudf may be a good solution especially when having a good gpu.<br>\nHowever I think that getting the data into jay format and then converting it to pandas using datatable as shown in <a href=\"https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid\" target=\"_blank\">this wonderful notebook</a> is more effective.<br>\n<a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">This notebook </a>is very good resource of knowledge regarding the different ways to import a large datasets. I recommend to visit it.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1046464,
      "author_name": "R Puttkammer",
      "author_url": "",
      "post_date": "2020-10-11T17:18:43.937000",
      "content": "<p>Yassine - not all is lost (though I agree a GPU would be nice.)</p>\n<pre><code>data_types_dict = ...\ndf = pd.read_csv(\"train.csv\", dtype=data_types_dict, index_col=0)\ndf.to_pickle(\"train.pkl\", protocol=pickle.HIGHEST_PROTOCOL)\n</code></pre>\n<pre><code>%%time\ntrain_df = pd.read_pickle(\"train.pkl\")\n</code></pre>\n<pre><code>CPU times: user 0 ns, sys: 1.17 s, total: 1.17 s Wall time: 1.17 s\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 1046488,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-11T17:49:45.063000",
          "content": "<p>I recommend you to replace pd with cudf<br>\n<code>df = cudf.read_csv(\"train.csv\")</code><br>\nand then write it into pkl or feather</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1063679,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-10-29T07:49:40.637000",
          "content": "<p><a href=\"https://www.kaggle.com/roman99\" target=\"_blank\">@roman99</a>  are we able to read all data into memory  during read csv  and read pickle?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1047127,
      "author_name": "Arkadiy Synovets",
      "author_url": "",
      "post_date": "2020-10-12T09:33:02.243000",
      "content": "<p>Could you please tell in your own words why cudf is so much faster than pandas?:) Thanx</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1047164,
          "author_name": "Aravind P",
          "author_url": "",
          "post_date": "2020-10-12T10:31:46.427000",
          "content": "<p>Because cudf is a GPU accelerated library build on top of CUDA with very similar API to pandas. The following link will help you get a better understanding of cudf and the Rapids ecosystem <br>\n<a href=\"url\" target=\"_blank\">https://rapids.ai/about.html</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1047182,
          "author_name": "Arkadiy Synovets",
          "author_url": "",
          "post_date": "2020-10-12T10:56:02.680000",
          "content": "<p>Hello, Aravind! Thanks for your answer! I will check provided resource:)</p>\n<p>By the way, does the using of cudf will decrease my Kaggle's GPU limit?:)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1047215,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-12T11:27:28.440000",
          "content": "<blockquote>\n  <p>does the using of cudf will decrease my Kaggle's GPU limit?</p>\n</blockquote>\n<p>Yes</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1046537,
      "author_name": "GauthamKumaran",
      "author_url": "",
      "post_date": "2020-10-11T18:46:26.420000",
      "content": "<p>sadly, converting <code>train</code> back to pandas gives an out of memory error.<br>\nBut great library, thanks for introducing it. Will explore more.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1045934,
      "author_name": "Aravind P",
      "author_url": "",
      "post_date": "2020-10-11T06:54:58.800000",
      "content": "<p>Adding to your point, XGBoost integrates with cuDF, Dask, and the entire RAPIDS ecosystem. Which will enable us to buid faster models by harvesting the power of GPU</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1083876,
      "author_name": "ISMAX",
      "author_url": "",
      "post_date": "2020-11-19T13:40:21.470000",
      "content": "<p>On my desktop computer, with pandas, reading the whole dataset in <strong>feather format</strong> takes <strong>less than 5 seconds</strong> and writing it to feather format takes about 23 seconds :<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F277616%2F149912a60fc0a349f539bace6ed6a828%2F2020-11-19%2014_39_44-Window.png?generation=1605793212318961&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1045890,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-10-11T06:13:28.037000",
      "content": "<p>I wish cudf came with a GPU attached. ;)<br>\nIndeed, life is short, so only waste it in what matters. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1053912,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-19T13:32:49.257000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1054131,
      "author_name": "Kaushal Shah",
      "author_url": "",
      "post_date": "2020-10-19T17:50:17.467000",
      "content": "<p>Installation takes forever. Is that only me? I'm referring <a href=\"https://rapids.ai/start.html\" target=\"_blank\">https://rapids.ai/start.html</a> with conda command.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1054428,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-19T23:33:20.293000",
          "content": "<p>Did you create environment? and run the command on the environment?<br>\nYou can create and activate a environment with following commands;</p>\n<pre><code>conda create -n foo python=3.7\nconda activate foo\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1054031,
      "author_name": "Karl BINA",
      "author_url": "",
      "post_date": "2020-10-19T15:45:04.427000",
      "content": "<p>Hi how do you install cudf with pip on Kaggle notebook?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1054039,
          "author_name": "Karl BINA",
          "author_url": "",
          "post_date": "2020-10-19T15:55:27.887000",
          "content": "<p>I found thank !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1052469,
      "author_name": "Mohamed LazyBob",
      "author_url": "",
      "post_date": "2020-10-17T18:54:20.873000",
      "content": "<p>thanks for sharing <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a>, It's the firs time to hear about that library </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052140,
      "author_name": "Rajat Kumar",
      "author_url": "",
      "post_date": "2020-10-17T11:23:49.363000",
      "content": "<p>That's a great way. Thanks for sharing. <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052021,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-10-17T07:43:48.610000",
      "content": "<p><a href=\"https://www.kaggle.com/onodere\" target=\"_blank\">@onodere</a> <br>\nLike you said loading the data in cudf makes in faster than pandas.<br>\nBut even in cudf the time taken for one feature engineering is 1 minute.<br>\nHow are you overcoming this??</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1052596,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-18T01:40:27.237000",
          "content": "<p>Depends on the feature</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1050913,
      "author_name": "Muhammad Saim Alam Khan",
      "author_url": "",
      "post_date": "2020-10-15T22:10:25.013000",
      "content": "<p>Awesome work buddy</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050688,
      "author_name": "Gaurav Chauhan",
      "author_url": "",
      "post_date": "2020-10-15T16:19:47.767000",
      "content": "<p>I wonder what Radeon has to offer in terms of ML</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050617,
      "author_name": "Tanagool Yenjai",
      "author_url": "",
      "post_date": "2020-10-15T15:31:01.280000",
      "content": "<p>Thanks for sharing sir, now I go search for cudf, kaggle give something new every single day,</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050551,
      "author_name": "Ahmed Lashin",
      "author_url": "",
      "post_date": "2020-10-15T14:30:40.653000",
      "content": "<p>Thanks a lot for sharing. Should give it a try. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1048791,
      "author_name": "Gabriel Atkin",
      "author_url": "",
      "post_date": "2020-10-13T19:31:37.953000",
      "content": "<p>Fantastic!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1047869,
      "author_name": "Yu Kang",
      "author_url": "",
      "post_date": "2020-10-13T02:44:23.717000",
      "content": "<p>Nice! Can cudf be installed with pip ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1048001,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-13T05:48:39.970000",
          "content": "<p>Refer to <a href=\"https://rapids.ai/start.html\" target=\"_blank\">here</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1048373,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-10-13T12:32:12.347000",
          "content": "<p>When I use <code>XXX.apply()</code>, I got <code>AttributeError: 'Series' object has no attribute 'apply'</code><br>\nAre there any alternatives to the <code>.apply</code> method? Or it's just not being implemented yet.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1048408,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-13T13:11:09.400000",
          "content": "<p>How about \"<strong>applymap</strong>\" ?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1048423,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-10-13T13:32:57.240000",
          "content": "<p>Thanks, that's what I want.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1047784,
      "author_name": "Timur V",
      "author_url": "",
      "post_date": "2020-10-12T23:48:54.703000",
      "content": "<p>Сompare the capabilities of the libraries, not just download the file.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1047748,
      "author_name": "Tarun Kumar",
      "author_url": "",
      "post_date": "2020-10-12T22:30:19.607000",
      "content": "<p>Are you using Colab or a local environment (if so, what GPU are you using?)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1047775,
          "author_name": "ONODERA",
          "author_url": "",
          "post_date": "2020-10-12T23:17:37.303000",
          "content": "<p>local, V100</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1047605,
      "author_name": "Sri Polu",
      "author_url": "",
      "post_date": "2020-10-12T18:37:07.837000",
      "content": "<p>Thanks for sharing nice info. Is there a way to converting from cudf to dataframe faster ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1047238,
      "author_name": "Gryffindor",
      "author_url": "",
      "post_date": "2020-10-12T11:56:18.713000",
      "content": "<p>Thanks for sharing…Will explore more on this.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1045812,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-10-11T04:31:10.073000",
      "content": "<p>I liked the quote btw😂</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1092307,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-11-26T17:26:28.790000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1053698,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-19T08:48:30.997000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1048075,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-13T06:57:27.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1047541,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-12T17:40:41.600000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050431,
      "author_name": "Devesh Mishra",
      "author_url": "",
      "post_date": "2020-10-15T12:14:16.233000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/onodera\" target=\"_blank\">@onodera</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1056382,
      "author_name": "Sudhir",
      "author_url": "",
      "post_date": "2020-10-21T16:37:07.160000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1054261,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-19T19:30:49.747000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1053940,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-19T14:01:06.927000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1053742,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-19T09:54:28.847000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052758,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-18T08:09:09.163000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052346,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-17T16:09:17.123000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052105,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-17T10:13:43.270000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1051985,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-17T06:29:16.240000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1050015,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-15T01:34:01.370000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1049750,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-14T18:47:50.140000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1048470,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-13T14:17:42.840000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1045872,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-10-11T05:58:57.317000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1045804": "Using [cuDF](https://docs.rapids.ai/api/cudf/stable/) you can read train.csv in **10 seconds** in [kaggle notebook](https://www.kaggle.com/onodera/riiid-read-csv-in-cudf).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F374bc1dee5812ff949677eb07e5ef8b5%2F2020-10-11%2013.13.38.png?generation=1602389639041354&alt=media)\n\n\n\nIn my environment, pandas takes **1 minutes**, but cuDF only takes **3 seconds**!!! \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F317344%2F9796020e8bc09be7d359cb83176b65d7%2FEkAJfR_U0AEQhAa.jpeg?generation=1602389509691312&alt=media)\n\n\nLife is short, use cuDF",
    "1049661": "Thank you for your help.  Using cudf may be a good solution especially when having a good gpu.\nHowever I think that getting the data into jay format and then converting it to pandas using datatable as shown in [this wonderful notebook](https://www.kaggle.com/rohanrao/riiid-with-blazing-fast-rid) is more effective.\n[This notebook ](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets)is very good resource of knowledge regarding the different ways to import a large datasets. I recommend to visit it.",
    "1046464": "Yassine - not all is lost (though I agree a GPU would be nice.)\n\n\n```\ndata_types_dict = ...\ndf = pd.read_csv(\"train.csv\", dtype=data_types_dict, index_col=0)\ndf.to_pickle(\"train.pkl\", protocol=pickle.HIGHEST_PROTOCOL)\n\n```\n```\n%%time\ntrain_df = pd.read_pickle(\"train.pkl\")\n\n```\n```\nCPU times: user 0 ns, sys: 1.17 s, total: 1.17 s Wall time: 1.17 s\n\n```\n",
    "1047127": "Could you please tell in your own words why cudf is so much faster than pandas?:) Thanx",
    "1046537": "sadly, converting `train` back to pandas gives an out of memory error.\nBut great library, thanks for introducing it. Will explore more.",
    "1045934": "Adding to your point, XGBoost integrates with cuDF, Dask, and the entire RAPIDS ecosystem. Which will enable us to buid faster models by harvesting the power of GPU",
    "1083876": "On my desktop computer, with pandas, reading the whole dataset in **feather format** takes **less than 5 seconds** and writing it to feather format takes about 23 seconds :\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F277616%2F149912a60fc0a349f539bace6ed6a828%2F2020-11-19%2014_39_44-Window.png?generation=1605793212318961&alt=media)",
    "1045890": "I wish cudf came with a GPU attached. ;)\nIndeed, life is short, so only waste it in what matters. ",
    "1054131": "Installation takes forever. Is that only me? I'm referring https://rapids.ai/start.html with conda command.",
    "1054031": "Hi how do you install cudf with pip on Kaggle notebook?",
    "1052469": "thanks for sharing @onodera, It's the firs time to hear about that library ",
    "1052140": "That's a great way. Thanks for sharing. @onodera ",
    "1052021": "@onodere \nLike you said loading the data in cudf makes in faster than pandas.\nBut even in cudf the time taken for one feature engineering is 1 minute.\nHow are you overcoming this??",
    "1050913": "Awesome work buddy",
    "1050688": "I wonder what Radeon has to offer in terms of ML",
    "1050617": "Thanks for sharing sir, now I go search for cudf, kaggle give something new every single day,",
    "1050551": "Thanks a lot for sharing. Should give it a try. ",
    "1048791": "Fantastic!",
    "1047869": "Nice! Can cudf be installed with pip ?",
    "1047784": "Сompare the capabilities of the libraries, not just download the file.",
    "1047748": "Are you using Colab or a local environment (if so, what GPU are you using?)",
    "1047605": "Thanks for sharing nice info. Is there a way to converting from cudf to dataframe faster ?",
    "1047238": "Thanks for sharing...Will explore more on this.",
    "1045812": "I liked the quote btw😂",
    "1092307": "",
    "1053698": "",
    "1048075": "",
    "1047541": "",
    "1050431": "Thanks for sharing @onodera ",
    "1056382": "Thanks for sharing",
    "1054261": "Thanks a lot for sharing this.",
    "1053940": "Thanks for sharing Man 👍",
    "1053742": "thanks for sharing ",
    "1052758": "Thanks for sharing @onodera",
    "1052346": "Thank you for sharing , very helpful",
    "1052105": "Thank you.",
    "1051985": "Thanks for sharing",
    "1050015": "Thanks for sharing this..",
    "1049750": "Thank you for your help.",
    "1048470": "Thanks for mentioning this",
    "1045872": "thank you for sharing:)"
  }
}