{
  "id": 201842,
  "title": "Success with Dask or Datatable? ",
  "url": "/competitions/riiid-test-answer-prediction/discussion/201842",
  "author_name": "",
  "post_date": "2020-12-07T03:25:21.307304500Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I want to prepare the whole train data to create features (most of them are rolling features) but as others have face, I ran into RAM overload issues. I tried to use some optimization techniques (I have <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444\" target=\"_blank\">documented some here</a>, will update the thread with my latest findings), they helped till one point, but wasn't able to complete till the end. So I went through all the discussions and publicly available notebooks. Through <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s <a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">great tutorial on reading large datasets</a>, I came across new packages (Dask and Datatable). I tried to use Dask for few days but I am not successful with it. So wanted to know if I am missing anything.</p>\n<ol>\n<li>Has anyone tried Dask? If you have faced ram overload of workers in Dask, how did you solve it?</li>\n<li>I tried to go through stackoverflow to see if Datatable is worth trying for. Its <a href=\"https://datatable.readthedocs.io/en/latest/manual/index-manual.html\" target=\"_blank\">user guide</a> has almost what I need but wasn't sure if it also sure too much RAM similar to Pandas. Anyone has tried Datatable before or for this competition? I would be grateful if you could give your perspective here. </li>\n</ol>\n<p>Most probably I am gonna shift to Numpy if these don't work out. This make my code messy but if it works for this competition, I can't complain.</p>\n<p>Links: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">Tutorial on reading large datasets</a></li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444\" target=\"_blank\">Discussion on debugging pandas RAM issues</a></li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/200917\" target=\"_blank\">pandas .loc and .query vs numexpr speed comparison\n</a></li>\n</ul>\n<p>Thank you.</p>",
  "messages": [
    {
      "id": "1104532",
      "postDate": "12/07/2020 03:25:21",
      "content": "<p>Hi everyone,</p>\n<p>I want to prepare the whole train data to create features (most of them are rolling features) but as others have face, I ran into RAM overload issues. I tried to use some optimization techniques (I have <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444\" target=\"_blank\">documented some here</a>, will update the thread with my latest findings), they helped till one point, but wasn't able to complete till the end. So I went through all the discussions and publicly available notebooks. Through <a href=\"https://www.kaggle.com/rohanrao\" target=\"_blank\">@rohanrao</a>'s <a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">great tutorial on reading large datasets</a>, I came across new packages (Dask and Datatable). I tried to use Dask for few days but I am not successful with it. So wanted to know if I am missing anything.</p>\n<ol>\n<li>Has anyone tried Dask? If you have faced ram overload of workers in Dask, how did you solve it?</li>\n<li>I tried to go through stackoverflow to see if Datatable is worth trying for. Its <a href=\"https://datatable.readthedocs.io/en/latest/manual/index-manual.html\" target=\"_blank\">user guide</a> has almost what I need but wasn't sure if it also sure too much RAM similar to Pandas. Anyone has tried Datatable before or for this competition? I would be grateful if you could give your perspective here. </li>\n</ol>\n<p>Most probably I am gonna shift to Numpy if these don't work out. This make my code messy but if it works for this competition, I can't complain.</p>\n<p>Links: </p>\n<ul>\n<li><a href=\"https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets\" target=\"_blank\">Tutorial on reading large datasets</a></li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444\" target=\"_blank\">Discussion on debugging pandas RAM issues</a></li>\n<li><a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/200917\" target=\"_blank\">pandas .loc and .query vs numexpr speed comparison\n</a></li>\n</ul>\n<p>Thank you.</p>",
      "rawMarkdown": "Hi everyone,\n\nI want to prepare the whole train data to create features (most of them are rolling features) but as others have face, I ran into RAM overload issues. I tried to use some optimization techniques (I have [documented some here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444), will update the thread with my latest findings), they helped till one point, but wasn't able to complete till the end. So I went through all the discussions and publicly available notebooks. Through @rohanrao's [great tutorial on reading large datasets](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets), I came across new packages (Dask and Datatable). I tried to use Dask for few days but I am not successful with it. So wanted to know if I am missing anything.\n1. Has anyone tried Dask? If you have faced ram overload of workers in Dask, how did you solve it?\n2. I tried to go through stackoverflow to see if Datatable is worth trying for. Its [user guide](https://datatable.readthedocs.io/en/latest/manual/index-manual.html) has almost what I need but wasn't sure if it also sure too much RAM similar to Pandas. Anyone has tried Datatable before or for this competition? I would be grateful if you could give your perspective here. \n\nMost probably I am gonna shift to Numpy if these don't work out. This make my code messy but if it works for this competition, I can't complain.\n\nLinks: \n- [Tutorial on reading large datasets](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets)\n- [Discussion on debugging pandas RAM issues](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444)\n- [pandas .loc and .query vs numexpr speed comparison\n](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/200917)\n\nThank you.",
      "votes": null
    },
    {
      "id": "1104871",
      "postDate": "12/07/2020 09:55:46",
      "content": "<p>I tried the both and I ended up with cuDF.</p>\n<p>As concern as memory efficiency, both of them are fine.<br>\nDask evaluates the value when we call <code>compute</code>. Until then, we can do anything we want forgetting the memory limitation. <br>\nDatatable is also good as it doesn't consume that much memory.</p>\n<p>The reason why I gave up these two frameworks are:<br>\nDask, I don't know why but it keeps giving me <code>ValueError</code> time to time. Sometimes it happens, sometimes it doesn't with exactly the same code. I couldn't solve the problem so I gave up.<br>\nDatatable, basically it's really great but because of some lack of implementation, it doesn't allow me to do some tasks like cumsum, so I needed to convert datatable into pandas after all.</p>\n<p>cuDF can do basically similar thing with pandas and it's super fast. <br>\nAs it consumes GPU memory, we need different kind of care but can be dealt with. <br>\nAnd some APIs are different from pandas, like <code>apply</code> or <code>cumsum</code>. But there's a way to do that. We can find how to perform this in <a href=\"https://docs.rapids.ai/api/cudf/stable/api.html\" target=\"_blank\">its document</a>. </p>",
      "rawMarkdown": "I tried the both and I ended up with cuDF.\n\nAs concern as memory efficiency, both of them are fine.\nDask evaluates the value when we call `compute`. Until then, we can do anything we want forgetting the memory limitation. \nDatatable is also good as it doesn't consume that much memory.\n\nThe reason why I gave up these two frameworks are:\nDask, I don't know why but it keeps giving me `ValueError` time to time. Sometimes it happens, sometimes it doesn't with exactly the same code. I couldn't solve the problem so I gave up.\nDatatable, basically it's really great but because of some lack of implementation, it doesn't allow me to do some tasks like cumsum, so I needed to convert datatable into pandas after all.\n\ncuDF can do basically similar thing with pandas and it's super fast. \nAs it consumes GPU memory, we need different kind of care but can be dealt with. \nAnd some APIs are different from pandas, like `apply` or `cumsum`. But there's a way to do that. We can find how to perform this in [its document](https://docs.rapids.ai/api/cudf/stable/api.html).",
      "votes": null
    },
    {
      "id": "1104924",
      "postDate": "12/07/2020 11:24:31",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a> for sharing your experience. I tried to use cuDF but the notebook was continuously restarting without initializing cuDF, so I gave up in between. But will try it again. </p>",
      "rawMarkdown": "Thanks @kokitanisaka for sharing your experience. I tried to use cuDF but the notebook was continuously restarting without initializing cuDF, so I gave up in between. But will try it again.",
      "votes": null
    },
    {
      "id": "1104926",
      "postDate": "12/07/2020 11:26:16",
      "content": "<p>I am successful (kind of) with using pandas by changing data types and reseting memory (ram) multiple times after saving data to disk. My office systems have more ram (460 GB machines), so I stopped caring about ram in my last year. This competition is making more aware the real world doesn't work like that. With clever engineering of pandas we could at least use 50M rows for training process. </p>",
      "rawMarkdown": "I am successful (kind of) with using pandas by changing data types and reseting memory (ram) multiple times after saving data to disk. My office systems have more ram (460 GB machines), so I stopped caring about ram in my last year. This competition is making more aware the real world doesn't work like that. With clever engineering of pandas we could at least use 50M rows for training process.",
      "votes": null
    },
    {
      "id": "1104931",
      "postDate": "12/07/2020 11:30:55",
      "content": "<p>Great! Actually it seems some people are doing well with pandas. <br>\nAs concern as initializing cuDF, you need to set accelerator GPU. Or the import declaration raises error.</p>",
      "rawMarkdown": "Great! Actually it seems some people are doing well with pandas. \nAs concern as initializing cuDF, you need to set accelerator GPU. Or the import declaration raises error.",
      "votes": null
    },
    {
      "id": "1104933",
      "postDate": "12/07/2020 11:33:21",
      "content": "<p>Yes. I was using GPU when trying cuDF. </p>",
      "rawMarkdown": "Yes. I was using GPU when trying cuDF.",
      "votes": null
    },
    {
      "id": "1108212",
      "postDate": "12/10/2020 12:06:27",
      "content": "<p>i tried both but both did not wok</p>",
      "rawMarkdown": "i tried both but both did not wok",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1104871,
      "author_name": "kokitanisaka",
      "author_url": "",
      "post_date": "12/07/2020 09:55:46",
      "content": "<p>I tried the both and I ended up with cuDF.</p>\n<p>As concern as memory efficiency, both of them are fine.<br>\nDask evaluates the value when we call <code>compute</code>. Until then, we can do anything we want forgetting the memory limitation. <br>\nDatatable is also good as it doesn't consume that much memory.</p>\n<p>The reason why I gave up these two frameworks are:<br>\nDask, I don't know why but it keeps giving me <code>ValueError</code> time to time. Sometimes it happens, sometimes it doesn't with exactly the same code. I couldn't solve the problem so I gave up.<br>\nDatatable, basically it's really great but because of some lack of implementation, it doesn't allow me to do some tasks like cumsum, so I needed to convert datatable into pandas after all.</p>\n<p>cuDF can do basically similar thing with pandas and it's super fast. <br>\nAs it consumes GPU memory, we need different kind of care but can be dealt with. <br>\nAnd some APIs are different from pandas, like <code>apply</code> or <code>cumsum</code>. But there's a way to do that. We can find how to perform this in <a href=\"https://docs.rapids.ai/api/cudf/stable/api.html\" target=\"_blank\">its document</a>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1104924,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/07/2020 11:24:31",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/kokitanisaka\" target=\"_blank\">@kokitanisaka</a> for sharing your experience. I tried to use cuDF but the notebook was continuously restarting without initializing cuDF, so I gave up in between. But will try it again. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104926,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/07/2020 11:26:16",
          "content": "<p>I am successful (kind of) with using pandas by changing data types and reseting memory (ram) multiple times after saving data to disk. My office systems have more ram (460 GB machines), so I stopped caring about ram in my last year. This competition is making more aware the real world doesn't work like that. With clever engineering of pandas we could at least use 50M rows for training process. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104931,
          "author_name": "kokitanisaka",
          "author_url": "",
          "post_date": "12/07/2020 11:30:55",
          "content": "<p>Great! Actually it seems some people are doing well with pandas. <br>\nAs concern as initializing cuDF, you need to set accelerator GPU. Or the import declaration raises error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1104933,
          "author_name": "manikanthr5",
          "author_url": "",
          "post_date": "12/07/2020 11:33:21",
          "content": "<p>Yes. I was using GPU when trying cuDF. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1108212,
      "author_name": "andleebhayath",
      "author_url": "",
      "post_date": "12/10/2020 12:06:27",
      "content": "<p>i tried both but both did not wok</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1104532": "Hi everyone,\n\nI want to prepare the whole train data to create features (most of them are rolling features) but as others have face, I ran into RAM overload issues. I tried to use some optimization techniques (I have [documented some here](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444), will update the thread with my latest findings), they helped till one point, but wasn't able to complete till the end. So I went through all the discussions and publicly available notebooks. Through @rohanrao's [great tutorial on reading large datasets](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets), I came across new packages (Dask and Datatable). I tried to use Dask for few days but I am not successful with it. So wanted to know if I am missing anything.\n1. Has anyone tried Dask? If you have faced ram overload of workers in Dask, how did you solve it?\n2. I tried to go through stackoverflow to see if Datatable is worth trying for. Its [user guide](https://datatable.readthedocs.io/en/latest/manual/index-manual.html) has almost what I need but wasn't sure if it also sure too much RAM similar to Pandas. Anyone has tried Datatable before or for this competition? I would be grateful if you could give your perspective here. \n\nMost probably I am gonna shift to Numpy if these don't work out. This make my code messy but if it works for this competition, I can't complain.\n\nLinks: \n- [Tutorial on reading large datasets](https://www.kaggle.com/rohanrao/tutorial-on-reading-large-datasets)\n- [Discussion on debugging pandas RAM issues](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/201444)\n- [pandas .loc and .query vs numexpr speed comparison\n](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/200917)\n\nThank you.",
    "1104871": "I tried the both and I ended up with cuDF.\n\nAs concern as memory efficiency, both of them are fine.\nDask evaluates the value when we call `compute`. Until then, we can do anything we want forgetting the memory limitation. \nDatatable is also good as it doesn't consume that much memory.\n\nThe reason why I gave up these two frameworks are:\nDask, I don't know why but it keeps giving me `ValueError` time to time. Sometimes it happens, sometimes it doesn't with exactly the same code. I couldn't solve the problem so I gave up.\nDatatable, basically it's really great but because of some lack of implementation, it doesn't allow me to do some tasks like cumsum, so I needed to convert datatable into pandas after all.\n\ncuDF can do basically similar thing with pandas and it's super fast. \nAs it consumes GPU memory, we need different kind of care but can be dealt with. \nAnd some APIs are different from pandas, like `apply` or `cumsum`. But there's a way to do that. We can find how to perform this in [its document](https://docs.rapids.ai/api/cudf/stable/api.html).",
    "1104924": "Thanks @kokitanisaka for sharing your experience. I tried to use cuDF but the notebook was continuously restarting without initializing cuDF, so I gave up in between. But will try it again.",
    "1104926": "I am successful (kind of) with using pandas by changing data types and reseting memory (ram) multiple times after saving data to disk. My office systems have more ram (460 GB machines), so I stopped caring about ram in my last year. This competition is making more aware the real world doesn't work like that. With clever engineering of pandas we could at least use 50M rows for training process.",
    "1104931": "Great! Actually it seems some people are doing well with pandas. \nAs concern as initializing cuDF, you need to set accelerator GPU. Or the import declaration raises error.",
    "1104933": "Yes. I was using GPU when trying cuDF.",
    "1108212": "i tried both but both did not wok"
  },
  "source": "meta"
}