{
  "id": 210239,
  "title": "My first competition using only RAPIDS",
  "url": "/competitions/riiid-test-answer-prediction/discussion/210239",
  "author_name": "",
  "post_date": "2021-01-10T08:13:00.738238100Z",
  "votes": 30,
  "comment_count": 3,
  "views": 0,
  "content": "<p>This competition was the perfect oppotunity to learn and practice for the very first time the RAPIDS suite of libraries. </p>\n<p>As the data was <strong>100mil+</strong> rows, usual libraries like Numpy, Pandas and Scikit-learn couldn't get the job done, or if they could, the process would take a very long time. Issues with insufficient memory would pop up as well, especially if you were using the Kaggle environment.</p>\n<p>This happens because these particular libraries are created to run on CPU, which becomes slower and slower as the volume of data increases.</p>\n<p>Hence, RAPIDS offers a suite of libraries very similar to the usual Numpy, Pandas, Scikit-learn etc., but combined with the power of GPU, which speeds up the process of analysis:</p>\n<p><img src=\"https://www.researchgate.net/ii/hosted.content.attachment/AS:682987469426688@1539848310063_xl\"></p>\n<p>I worked on a Z8 Workstation that has an NVIDIA Quadro RTX 8000 on it, but if you don't have local GPUs you can always go ahead and use the Kaggle environment, which is already equiped with a pretty powerful NVIDIA Tesla P-100 GPU.</p>\n<h2>What I've learned</h2>\n<ul>\n<li><code>cudf</code> is the replacer of <code>pandas</code>. It has implemented lots of helpful functions, such as <code>.read_csv()</code>, <code>.isna()</code>, <code>.fillna()</code>, <code>.groupby()</code> or <code>.agg()</code>.</li>\n<li><code>cudf</code> reads 100 mil rows in less than 10 seconds.</li>\n<li>saving the dataframe as <code>.parquet</code> instead of <code>.csv</code> is making the process 6x faster.</li>\n<li><code>cupy</code> is the replacer for <code>numpy</code>. It can be used to easily transfer data from <code>GPU</code> array to <code>CPU</code> array, so you can perform visualizations using the entire dataset.</li>\n<li><code>cuml</code> is the replacer for <code>sklearn</code>.</li>\n<li>the <code>roc_auc_score</code> from <code>cuml</code> is 16 times faster than the one from <code>sklearn</code>.</li>\n<li>model training was fast like the wind: 10 seconds for 10 million rows using <code>xgboost</code> on GPU</li>\n<li>helpful references are <a href=\"https://docs.rapids.ai/api/cudf/stable/10min.html#Dask-Performance-Tips\" target=\"_blank\">10 mins to cudf</a> and <a href=\"https://towardsdatascience.com/lightning-fast-xgboost-on-multiple-gpus-32710815c7c3\" target=\"_blank\">xgboost on GPU tutorial</a>.</li>\n</ul>\n<h2>Struggles I've encountered</h2>\n<ul>\n<li>RAPIDS doesn't have yet a visualization library (or I couldn't find it): hence, what I ended up doing was converting the data to <code>cupy</code> array, then to <code>numpy</code> and using <code>.plot()</code> and <code>.value_counts()</code> to analyze</li>\n<li>there are still many functions to be yet implemented in all <code>cudf</code>, <code>cupy</code> and <code>cuml</code> libraries. Hence, you might find yourself not being able to use some of your favorite functions.</li>\n<li>k-fold or stratified k-fold are yet to be implemented; I had a lot of struggle here, although <a href=\"https://www.kaggle.com/c/landmark-recognition-2019/discussion/90691#1131854\" target=\"_blank\">Jiwei Liu already created a custom function for this.</a> However, when I tried to use it on my particular problem, I encountered error after error, with little to no reference on the internet on how to properly solve them.</li>\n<li>because it's rather new, RAPIDS doesn't have as much documentation/examples out there as we're used to with the og libraries.</li>\n</ul>\n<p>I will for sure continue to use more and more of RAPIDS. I will also update this post whenever I find something useful, and I strongly encourage you to do the same. :)</p>\n<p>Amazing competition and Happy Data Sciencin'!</p>",
  "messages": [
    {
      "id": "1146976",
      "postDate": "01/10/2021 08:13:00",
      "content": "<p>This competition was the perfect oppotunity to learn and practice for the very first time the RAPIDS suite of libraries. </p>\n<p>As the data was <strong>100mil+</strong> rows, usual libraries like Numpy, Pandas and Scikit-learn couldn't get the job done, or if they could, the process would take a very long time. Issues with insufficient memory would pop up as well, especially if you were using the Kaggle environment.</p>\n<p>This happens because these particular libraries are created to run on CPU, which becomes slower and slower as the volume of data increases.</p>\n<p>Hence, RAPIDS offers a suite of libraries very similar to the usual Numpy, Pandas, Scikit-learn etc., but combined with the power of GPU, which speeds up the process of analysis:</p>\n<p><img src=\"https://www.researchgate.net/ii/hosted.content.attachment/AS:682987469426688@1539848310063_xl\"></p>\n<p>I worked on a Z8 Workstation that has an NVIDIA Quadro RTX 8000 on it, but if you don't have local GPUs you can always go ahead and use the Kaggle environment, which is already equiped with a pretty powerful NVIDIA Tesla P-100 GPU.</p>\n<h2>What I've learned</h2>\n<ul>\n<li><code>cudf</code> is the replacer of <code>pandas</code>. It has implemented lots of helpful functions, such as <code>.read_csv()</code>, <code>.isna()</code>, <code>.fillna()</code>, <code>.groupby()</code> or <code>.agg()</code>.</li>\n<li><code>cudf</code> reads 100 mil rows in less than 10 seconds.</li>\n<li>saving the dataframe as <code>.parquet</code> instead of <code>.csv</code> is making the process 6x faster.</li>\n<li><code>cupy</code> is the replacer for <code>numpy</code>. It can be used to easily transfer data from <code>GPU</code> array to <code>CPU</code> array, so you can perform visualizations using the entire dataset.</li>\n<li><code>cuml</code> is the replacer for <code>sklearn</code>.</li>\n<li>the <code>roc_auc_score</code> from <code>cuml</code> is 16 times faster than the one from <code>sklearn</code>.</li>\n<li>model training was fast like the wind: 10 seconds for 10 million rows using <code>xgboost</code> on GPU</li>\n<li>helpful references are <a href=\"https://docs.rapids.ai/api/cudf/stable/10min.html#Dask-Performance-Tips\" target=\"_blank\">10 mins to cudf</a> and <a href=\"https://towardsdatascience.com/lightning-fast-xgboost-on-multiple-gpus-32710815c7c3\" target=\"_blank\">xgboost on GPU tutorial</a>.</li>\n</ul>\n<h2>Struggles I've encountered</h2>\n<ul>\n<li>RAPIDS doesn't have yet a visualization library (or I couldn't find it): hence, what I ended up doing was converting the data to <code>cupy</code> array, then to <code>numpy</code> and using <code>.plot()</code> and <code>.value_counts()</code> to analyze</li>\n<li>there are still many functions to be yet implemented in all <code>cudf</code>, <code>cupy</code> and <code>cuml</code> libraries. Hence, you might find yourself not being able to use some of your favorite functions.</li>\n<li>k-fold or stratified k-fold are yet to be implemented; I had a lot of struggle here, although <a href=\"https://www.kaggle.com/c/landmark-recognition-2019/discussion/90691#1131854\" target=\"_blank\">Jiwei Liu already created a custom function for this.</a> However, when I tried to use it on my particular problem, I encountered error after error, with little to no reference on the internet on how to properly solve them.</li>\n<li>because it's rather new, RAPIDS doesn't have as much documentation/examples out there as we're used to with the og libraries.</li>\n</ul>\n<p>I will for sure continue to use more and more of RAPIDS. I will also update this post whenever I find something useful, and I strongly encourage you to do the same. :)</p>\n<p>Amazing competition and Happy Data Sciencin'!</p>",
      "rawMarkdown": "This competition was the perfect oppotunity to learn and practice for the very first time the RAPIDS suite of libraries. \n\nAs the data was **100mil+** rows, usual libraries like Numpy, Pandas and Scikit-learn couldn't get the job done, or if they could, the process would take a very long time. Issues with insufficient memory would pop up as well, especially if you were using the Kaggle environment.\n\nThis happens because these particular libraries are created to run on CPU, which becomes slower and slower as the volume of data increases.\n\nHence, RAPIDS offers a suite of libraries very similar to the usual Numpy, Pandas, Scikit-learn etc., but combined with the power of GPU, which speeds up the process of analysis:\n\n<img src=\"https://www.researchgate.net/ii/hosted.content.attachment/AS:682987469426688@1539848310063_xl\">\n\nI worked on a Z8 Workstation that has an NVIDIA Quadro RTX 8000 on it, but if you don't have local GPUs you can always go ahead and use the Kaggle environment, which is already equiped with a pretty powerful NVIDIA Tesla P-100 GPU.\n\n## What I've learned\n* `cudf` is the replacer of `pandas`. It has implemented lots of helpful functions, such as `.read_csv()`, `.isna()`, `.fillna()`, `.groupby()` or `.agg()`.\n* `cudf` reads 100 mil rows in less than 10 seconds.\n* saving the dataframe as `.parquet` instead of `.csv` is making the process 6x faster.\n* `cupy` is the replacer for `numpy`. It can be used to easily transfer data from `GPU` array to `CPU` array, so you can perform visualizations using the entire dataset.\n* `cuml` is the replacer for `sklearn`.\n* the `roc_auc_score` from `cuml` is 16 times faster than the one from `sklearn`.\n* model training was fast like the wind: 10 seconds for 10 million rows using `xgboost` on GPU\n* helpful references are [10 mins to cudf](https://docs.rapids.ai/api/cudf/stable/10min.html#Dask-Performance-Tips) and [xgboost on GPU tutorial](https://towardsdatascience.com/lightning-fast-xgboost-on-multiple-gpus-32710815c7c3).\n\n## Struggles I've encountered\n* RAPIDS doesn't have yet a visualization library (or I couldn't find it): hence, what I ended up doing was converting the data to `cupy` array, then to `numpy` and using `.plot()` and `.value_counts()` to analyze\n* there are still many functions to be yet implemented in all `cudf`, `cupy` and `cuml` libraries. Hence, you might find yourself not being able to use some of your favorite functions.\n* k-fold or stratified k-fold are yet to be implemented; I had a lot of struggle here, although [Jiwei Liu already created a custom function for this.](https://www.kaggle.com/c/landmark-recognition-2019/discussion/90691#1131854) However, when I tried to use it on my particular problem, I encountered error after error, with little to no reference on the internet on how to properly solve them.\n* because it's rather new, RAPIDS doesn't have as much documentation/examples out there as we're used to with the og libraries.\n\nI will for sure continue to use more and more of RAPIDS. I will also update this post whenever I find something useful, and I strongly encourage you to do the same. :)\n\nAmazing competition and Happy Data Sciencin'!",
      "votes": null
    },
    {
      "id": "1147871",
      "postDate": "01/10/2021 19:01:53",
      "content": "<p>thank you for sharing this. I’m always thinking to try RAPIDS, this convince me even more</p>",
      "rawMarkdown": "thank you for sharing this. I’m always thinking to try RAPIDS, this convince me even more",
      "votes": null
    },
    {
      "id": "1151716",
      "postDate": "01/13/2021 14:11:00",
      "content": "<p>RAPIDS cudf doesn't have .plot() yet, in this meantime, to plot something you can convert to a pandas dataframe like this:<br>\n<code>df['feature'].to_pandas().plot()</code></p>",
      "rawMarkdown": "RAPIDS cudf doesn't have .plot() yet, in this meantime, to plot something you can convert to a pandas dataframe like this:\n`df['feature'].to_pandas().plot()  `",
      "votes": null
    },
    {
      "id": "1153690",
      "postDate": "01/15/2021 05:03:06",
      "content": "<p>quite intuitive , <br>\nthanks for sharing ✌ </p>",
      "rawMarkdown": "quite intuitive , \nthanks for sharing ✌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1147871,
      "author_name": "dwin183287",
      "author_url": "",
      "post_date": "01/10/2021 19:01:53",
      "content": "<p>thank you for sharing this. I’m always thinking to try RAPIDS, this convince me even more</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1151716,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "01/13/2021 14:11:00",
      "content": "<p>RAPIDS cudf doesn't have .plot() yet, in this meantime, to plot something you can convert to a pandas dataframe like this:<br>\n<code>df['feature'].to_pandas().plot()</code></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1153690,
      "author_name": "shyamgupta196",
      "author_url": "",
      "post_date": "01/15/2021 05:03:06",
      "content": "<p>quite intuitive , <br>\nthanks for sharing ✌ </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1146976": "This competition was the perfect oppotunity to learn and practice for the very first time the RAPIDS suite of libraries. \n\nAs the data was **100mil+** rows, usual libraries like Numpy, Pandas and Scikit-learn couldn't get the job done, or if they could, the process would take a very long time. Issues with insufficient memory would pop up as well, especially if you were using the Kaggle environment.\n\nThis happens because these particular libraries are created to run on CPU, which becomes slower and slower as the volume of data increases.\n\nHence, RAPIDS offers a suite of libraries very similar to the usual Numpy, Pandas, Scikit-learn etc., but combined with the power of GPU, which speeds up the process of analysis:\n\n<img src=\"https://www.researchgate.net/ii/hosted.content.attachment/AS:682987469426688@1539848310063_xl\">\n\nI worked on a Z8 Workstation that has an NVIDIA Quadro RTX 8000 on it, but if you don't have local GPUs you can always go ahead and use the Kaggle environment, which is already equiped with a pretty powerful NVIDIA Tesla P-100 GPU.\n\n## What I've learned\n* `cudf` is the replacer of `pandas`. It has implemented lots of helpful functions, such as `.read_csv()`, `.isna()`, `.fillna()`, `.groupby()` or `.agg()`.\n* `cudf` reads 100 mil rows in less than 10 seconds.\n* saving the dataframe as `.parquet` instead of `.csv` is making the process 6x faster.\n* `cupy` is the replacer for `numpy`. It can be used to easily transfer data from `GPU` array to `CPU` array, so you can perform visualizations using the entire dataset.\n* `cuml` is the replacer for `sklearn`.\n* the `roc_auc_score` from `cuml` is 16 times faster than the one from `sklearn`.\n* model training was fast like the wind: 10 seconds for 10 million rows using `xgboost` on GPU\n* helpful references are [10 mins to cudf](https://docs.rapids.ai/api/cudf/stable/10min.html#Dask-Performance-Tips) and [xgboost on GPU tutorial](https://towardsdatascience.com/lightning-fast-xgboost-on-multiple-gpus-32710815c7c3).\n\n## Struggles I've encountered\n* RAPIDS doesn't have yet a visualization library (or I couldn't find it): hence, what I ended up doing was converting the data to `cupy` array, then to `numpy` and using `.plot()` and `.value_counts()` to analyze\n* there are still many functions to be yet implemented in all `cudf`, `cupy` and `cuml` libraries. Hence, you might find yourself not being able to use some of your favorite functions.\n* k-fold or stratified k-fold are yet to be implemented; I had a lot of struggle here, although [Jiwei Liu already created a custom function for this.](https://www.kaggle.com/c/landmark-recognition-2019/discussion/90691#1131854) However, when I tried to use it on my particular problem, I encountered error after error, with little to no reference on the internet on how to properly solve them.\n* because it's rather new, RAPIDS doesn't have as much documentation/examples out there as we're used to with the og libraries.\n\nI will for sure continue to use more and more of RAPIDS. I will also update this post whenever I find something useful, and I strongly encourage you to do the same. :)\n\nAmazing competition and Happy Data Sciencin'!",
    "1147871": "thank you for sharing this. I’m always thinking to try RAPIDS, this convince me even more",
    "1151716": "RAPIDS cudf doesn't have .plot() yet, in this meantime, to plot something you can convert to a pandas dataframe like this:\n`df['feature'].to_pandas().plot()  `",
    "1153690": "quite intuitive , \nthanks for sharing ✌"
  },
  "source": "meta"
}