{
  "id": 71753,
  "title": "Faster Dataframe Processing",
  "url": "/competitions/PLAsTiCC-2018/discussion/71753",
  "author_name": "CPMP",
  "post_date": "2018-11-16T07:33:05.486000",
  "votes": 27,
  "comment_count": 9,
  "views": 0,
  "content": "<p>We discussed how to improve speed of pandas computation in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398\">this topic</a>.  Read all comments if you have not, they contain lots of very interesting insights.  Another way to speed up computation is to use more efficient implementations than pandas.  In particular Olivier mentioned the <a href=\"https://github.com/modin-project/modin\">modin</a> package.</p>\n\n<p>I'd like to mention two other packages that may be relevant:\n- the <a href=\"https://github.com/rapidsai/cudf\">RAPIDS</a> package that aims at providing a gpu accelerated pandas and numpy.\n- <a href=\"https://github.com/h2oai/datatable\">python datatable</a>, that mimics the well known data.table R package.</p>\n\n<p>Both have impressive benchmarks against pandas.  I have not used them yet but I will certainly use them at some point.</p>",
  "messages": [
    {
      "id": 422419,
      "postDate": "2018-11-16T07:33:05.487Z",
      "content": "<p>We discussed how to improve speed of pandas computation in <a href=\"https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398\">this topic</a>.  Read all comments if you have not, they contain lots of very interesting insights.  Another way to speed up computation is to use more efficient implementations than pandas.  In particular Olivier mentioned the <a href=\"https://github.com/modin-project/modin\">modin</a> package.</p>\n\n<p>I'd like to mention two other packages that may be relevant:\n- the <a href=\"https://github.com/rapidsai/cudf\">RAPIDS</a> package that aims at providing a gpu accelerated pandas and numpy.\n- <a href=\"https://github.com/h2oai/datatable\">python datatable</a>, that mimics the well known data.table R package.</p>\n\n<p>Both have impressive benchmarks against pandas.  I have not used them yet but I will certainly use them at some point.</p>",
      "rawMarkdown": "We discussed how to improve speed of pandas computation in [this topic][1].  Read all comments if you have not, they contain lots of very interesting insights.  Another way to speed up computation is to use more efficient implementations than pandas.  In particular Olivier mentioned the [modin][2] package.\n\nI'd like to mention two other packages that may be relevant:\n- the [RAPIDS][3] package that aims at providing a gpu accelerated pandas and numpy.\n- [python datatable][4], that mimics the well known data.table R package.\n\nBoth have impressive benchmarks against pandas.  I have not used them yet but I will certainly use them at some point.\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398\n  [2]: https://github.com/modin-project/modin\n  [3]: https://github.com/rapidsai/cudf\n  [4]: https://github.com/h2oai/datatable",
      "votes": 27
    },
    {
      "id": 465283,
      "postDate": "2019-02-02T18:17:07.740Z",
      "content": "<p>Thanks a lot, this exactly what I was looking for</p>",
      "rawMarkdown": "Thanks a lot, this exactly what I was looking for\n",
      "votes": 1
    },
    {
      "id": 422428,
      "postDate": "2018-11-16T07:53:50.483Z",
      "content": "<p>I guess the link to RAPIDS is <a href=\"https://rapids.ai/\">https://rapids.ai/</a> ?</p>",
      "rawMarkdown": "I guess the link to RAPIDS is https://rapids.ai/ ?",
      "votes": 1,
      "replies": [
        {
          "id": 422498,
          "postDate": "2018-11-16T10:03:22.863Z",
          "content": "<p>Sorry, yes.  Actually cuDF fromm Rapids: <a href=\"https://github.com/rapidsai/cudf\">https://github.com/rapidsai/cudf</a></p>\n\n<p>I fixed the link in the post.  Thanks for pointing this out.</p>",
          "rawMarkdown": "Sorry, yes.  Actually cuDF fromm Rapids: https://github.com/rapidsai/cudf\n\nI fixed the link in the post.  Thanks for pointing this out.\n"
        },
        {
          "id": 422616,
          "postDate": "2018-11-16T14:11:52.957Z",
          "content": "<p>@CPMP</p>\n\n<p>Thanks for mentioning the excellent R datatable package whis is extremely fast...</p>",
          "rawMarkdown": "@CPMP\n\nThanks for mentioning the excellent R datatable package whis is extremely fast..."
        },
        {
          "id": 424054,
          "postDate": "2018-11-19T13:31:32.010Z",
          "content": "<p>@mezoganet, data.table author, Matt Dowle, is now working at H2O and his new baby is python datatable...</p>",
          "rawMarkdown": "@mezoganet, data.table author, Matt Dowle, is now working at H2O and his new baby is python datatable...",
          "votes": 2
        }
      ]
    },
    {
      "id": 523690,
      "postDate": "2019-04-26T18:50:45.463Z",
      "content": "<p>Great info! I've tried <a href=\"https://github.com/modin-project/modin\">modin</a> and it's really impressive too. Unfortunately they haven't implemented all the functionalities yet, so a lot still defaults to pandas. I had some minor issues with the underlying distribution framework  <a href=\"https://github.com/ray-project/ray\">ray</a>, so I had to convert back sometimes to pandas DataFrame.  <strong>modin</strong> is still work in progress, but very promising </p>",
      "rawMarkdown": "Great info! I've tried [modin](https://github.com/modin-project/modin) and it's really impressive too. Unfortunately they haven't implemented all the functionalities yet, so a lot still defaults to pandas. I had some minor issues with the underlying distribution framework  [ray](https://github.com/ray-project/ray ), so I had to convert back sometimes to pandas DataFrame.  **modin** is still work in progress, but very promising "
    },
    {
      "id": 424995,
      "postDate": "2018-11-21T01:40:51.910Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 423219,
      "postDate": "2018-11-17T18:11:15.790Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 422717,
      "postDate": "2018-11-16T17:01:22.403Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 465283,
      "author_name": "Iurii Cojocari",
      "author_url": "",
      "post_date": "2019-02-02T18:17:07.740000",
      "content": "<p>Thanks a lot, this exactly what I was looking for</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 422428,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "2018-11-16T07:53:50.483000",
      "content": "<p>I guess the link to RAPIDS is <a href=\"https://rapids.ai/\">https://rapids.ai/</a> ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 422498,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-16T10:03:22.863000",
          "content": "<p>Sorry, yes.  Actually cuDF fromm Rapids: <a href=\"https://github.com/rapidsai/cudf\">https://github.com/rapidsai/cudf</a></p>\n\n<p>I fixed the link in the post.  Thanks for pointing this out.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 422616,
          "author_name": "mezoganet",
          "author_url": "",
          "post_date": "2018-11-16T14:11:52.957000",
          "content": "<p>@CPMP</p>\n\n<p>Thanks for mentioning the excellent R datatable package whis is extremely fast...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 424054,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2018-11-19T13:31:32.010000",
          "content": "<p>@mezoganet, data.table author, Matt Dowle, is now working at H2O and his new baby is python datatable...</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 523690,
      "author_name": "Markus",
      "author_url": "",
      "post_date": "2019-04-26T18:50:45.463000",
      "content": "<p>Great info! I've tried <a href=\"https://github.com/modin-project/modin\">modin</a> and it's really impressive too. Unfortunately they haven't implemented all the functionalities yet, so a lot still defaults to pandas. I had some minor issues with the underlying distribution framework  <a href=\"https://github.com/ray-project/ray\">ray</a>, so I had to convert back sometimes to pandas DataFrame.  <strong>modin</strong> is still work in progress, but very promising </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 424995,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-21T01:40:51.910000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 423219,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-17T18:11:15.790000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 422717,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-16T17:01:22.403000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422419": "We discussed how to improve speed of pandas computation in [this topic][1].  Read all comments if you have not, they contain lots of very interesting insights.  Another way to speed up computation is to use more efficient implementations than pandas.  In particular Olivier mentioned the [modin][2] package.\n\nI'd like to mention two other packages that may be relevant:\n- the [RAPIDS][3] package that aims at providing a gpu accelerated pandas and numpy.\n- [python datatable][4], that mimics the well known data.table R package.\n\nBoth have impressive benchmarks against pandas.  I have not used them yet but I will certainly use them at some point.\n\n  [1]: https://www.kaggle.com/c/PLAsTiCC-2018/discussion/71398\n  [2]: https://github.com/modin-project/modin\n  [3]: https://github.com/rapidsai/cudf\n  [4]: https://github.com/h2oai/datatable",
    "465283": "Thanks a lot, this exactly what I was looking for\n",
    "422428": "I guess the link to RAPIDS is https://rapids.ai/ ?",
    "523690": "Great info! I've tried [modin](https://github.com/modin-project/modin) and it's really impressive too. Unfortunately they haven't implemented all the functionalities yet, so a lot still defaults to pandas. I had some minor issues with the underlying distribution framework  [ray](https://github.com/ray-project/ray ), so I had to convert back sometimes to pandas DataFrame.  **modin** is still work in progress, but very promising ",
    "424995": "",
    "423219": "",
    "422717": ""
  }
}