{
  "id": 74969,
  "title": "Did anyone successfully use CUDF from https://rapids.ai/ to speedup preprocessing?",
  "url": "/competitions/PLAsTiCC-2018/discussion/74969",
  "author_name": "",
  "post_date": "2018-12-17T16:59:20.260851400Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>I tried using <a href=\"https://rapids.ai/\">CUDF</a> to speed up my groupby object_id for LSTM autoencoder as suggested by @CPMP. my pandas code works great but slow to compute on test data. CUDF does not seem to have the same api as pandas, I could not figure out how to get it to work. Did any anyone get reasonable speed up from CUDF? Did test data fit gpu memory?</p>",
  "messages": [
    {
      "id": "440522",
      "postDate": "12/17/2018 16:59:20",
      "content": "<p>I tried using <a href=\"https://rapids.ai/\">CUDF</a> to speed up my groupby object_id for LSTM autoencoder as suggested by @CPMP. my pandas code works great but slow to compute on test data. CUDF does not seem to have the same api as pandas, I could not figure out how to get it to work. Did any anyone get reasonable speed up from CUDF? Did test data fit gpu memory?</p>",
      "rawMarkdown": "I tried using [CUDF][1] to speed up my groupby object_id for LSTM autoencoder as suggested by @CPMP. my pandas code works great but slow to compute on test data. CUDF does not seem to have the same api as pandas, I could not figure out how to get it to work. Did any anyone get reasonable speed up from CUDF? Did test data fit gpu memory?\n\n\n  [1]: https://rapids.ai/",
      "votes": null
    },
    {
      "id": "440561",
      "postDate": "12/17/2018 17:46:04",
      "content": "<p>I split the test set in 32 files and used threading to load the files and process using pandas. In my case, it decreased my test set processing time to less than 20 minutes.</p>",
      "rawMarkdown": "I split the test set in 32 files and used threading to load the files and process using pandas. In my case, it decreased my test set processing time to less than 20 minutes.",
      "votes": null
    },
    {
      "id": "440564",
      "postDate": "12/17/2018 17:54:23",
      "content": "<p>Wow &lt; 20 min, did you prevent object_ids being split across the files?</p>",
      "rawMarkdown": "Wow &lt; 20 min, did you prevent object_ids being split across the files?",
      "votes": null
    },
    {
      "id": "440605",
      "postDate": "12/17/2018 19:09:42",
      "content": "<p>yes, I stratified by object_id and using 32 cores :-)</p>",
      "rawMarkdown": "yes, I stratified by object_id and using 32 cores :-)",
      "votes": null
    },
    {
      "id": "440617",
      "postDate": "12/17/2018 19:35:31",
      "content": "<p>Same here, I generate a sub in 20 mins using 20 cores.  Parallelism is key.</p>",
      "rawMarkdown": "Same here, I generate a sub in 20 mins using 20 cores.  Parallelism is key.",
      "votes": null
    },
    {
      "id": "440817",
      "postDate": "12/18/2018 01:54:20",
      "content": "<p>Apparently yes! Actually 8GB memory is good enough most of the time. I'll publish a demo later and will keep you updated. Thank you!</p>",
      "rawMarkdown": "Apparently yes! Actually 8GB memory is good enough most of the time. I'll publish a demo later and will keep you updated. Thank you!",
      "votes": null
    },
    {
      "id": "440866",
      "postDate": "12/18/2018 03:21:40",
      "content": "<p>I find Dask to be sufficient most of the time, but more speed never hurts.\nAS for neural network training, I think we cannot really expect everything to fit into memory, especially if you are training an autoenocder rather than working with features. A dynamic loading process like PyTorch DataLoader is really necessary. The trick we used is to expand the raw data into the required format, save them as memmap on disk (75GB for the test set...) then read them as needed. This way you never have to redo groupby's and training starts almost instantly the moment you hit the run button.</p>",
      "rawMarkdown": "I find Dask to be sufficient most of the time, but more speed never hurts.\nAS for neural network training, I think we cannot really expect everything to fit into memory, especially if you are training an autoenocder rather than working with features. A dynamic loading process like PyTorch DataLoader is really necessary. The trick we used is to expand the raw data into the required format, save them as memmap on disk (75GB for the test set...) then read them as needed. This way you never have to redo groupby's and training starts almost instantly the moment you hit the run button.",
      "votes": null
    },
    {
      "id": "440883",
      "postDate": "12/18/2018 04:04:41",
      "content": "<p>Thanks Jiwei, need the demo, I found the documentation a little sparse, I see you are on the Rapids team at nvidia, I think it is a potential game changer once all the features are implemented.</p>",
      "rawMarkdown": "Thanks Jiwei, need the demo, I found the documentation a little sparse, I see you are on the Rapids team at nvidia, I think it is a potential game changer once all the features are implemented.",
      "votes": null
    },
    {
      "id": "441707",
      "postDate": "12/19/2018 00:10:21",
      "content": "<p>Here is a demo <a href=\"https://github.com/daxiongshu/Rapids_PLAsTiCC_2018\">https://github.com/daxiongshu/Rapids_PLAsTiCC_2018</a>  My experiment shows 10x speedup than pandas.\nSorry for not getting the fact straight yesterday. Loading the test time series along will cost 11 GB. But you can skip rows to fit the data on GPU. I highlight this part at the beginning of the notebook. We are working on dask-cudf to solve this OOM problem for both single and multi gpu setting.  Let me know if you have any questions!</p>",
      "rawMarkdown": "Here is a demo https://github.com/daxiongshu/Rapids_PLAsTiCC_2018  My experiment shows 10x speedup than pandas.\nSorry for not getting the fact straight yesterday. Loading the test time series along will cost 11 GB. But you can skip rows to fit the data on GPU. I highlight this part at the beginning of the notebook. We are working on dask-cudf to solve this OOM problem for both single and multi gpu setting.  Let me know if you have any questions!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 440561,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "12/17/2018 17:46:04",
      "content": "<p>I split the test set in 32 files and used threading to load the files and process using pandas. In my case, it decreased my test set processing time to less than 20 minutes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 440564,
          "author_name": "godaibo",
          "author_url": "",
          "post_date": "12/17/2018 17:54:23",
          "content": "<p>Wow &lt; 20 min, did you prevent object_ids being split across the files?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440605,
          "author_name": "titericz",
          "author_url": "",
          "post_date": "12/17/2018 19:09:42",
          "content": "<p>yes, I stratified by object_id and using 32 cores :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440617,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/17/2018 19:35:31",
          "content": "<p>Same here, I generate a sub in 20 mins using 20 cores.  Parallelism is key.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440817,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "12/18/2018 01:54:20",
      "content": "<p>Apparently yes! Actually 8GB memory is good enough most of the time. I'll publish a demo later and will keep you updated. Thank you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 440883,
          "author_name": "godaibo",
          "author_url": "",
          "post_date": "12/18/2018 04:04:41",
          "content": "<p>Thanks Jiwei, need the demo, I found the documentation a little sparse, I see you are on the Rapids team at nvidia, I think it is a potential game changer once all the features are implemented.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 440866,
      "author_name": "mithrillion",
      "author_url": "",
      "post_date": "12/18/2018 03:21:40",
      "content": "<p>I find Dask to be sufficient most of the time, but more speed never hurts.\nAS for neural network training, I think we cannot really expect everything to fit into memory, especially if you are training an autoenocder rather than working with features. A dynamic loading process like PyTorch DataLoader is really necessary. The trick we used is to expand the raw data into the required format, save them as memmap on disk (75GB for the test set...) then read them as needed. This way you never have to redo groupby's and training starts almost instantly the moment you hit the run button.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 441707,
      "author_name": "jiweiliu",
      "author_url": "",
      "post_date": "12/19/2018 00:10:21",
      "content": "<p>Here is a demo <a href=\"https://github.com/daxiongshu/Rapids_PLAsTiCC_2018\">https://github.com/daxiongshu/Rapids_PLAsTiCC_2018</a>  My experiment shows 10x speedup than pandas.\nSorry for not getting the fact straight yesterday. Loading the test time series along will cost 11 GB. But you can skip rows to fit the data on GPU. I highlight this part at the beginning of the notebook. We are working on dask-cudf to solve this OOM problem for both single and multi gpu setting.  Let me know if you have any questions!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "440522": "I tried using [CUDF][1] to speed up my groupby object_id for LSTM autoencoder as suggested by @CPMP. my pandas code works great but slow to compute on test data. CUDF does not seem to have the same api as pandas, I could not figure out how to get it to work. Did any anyone get reasonable speed up from CUDF? Did test data fit gpu memory?\n\n\n  [1]: https://rapids.ai/",
    "440561": "I split the test set in 32 files and used threading to load the files and process using pandas. In my case, it decreased my test set processing time to less than 20 minutes.",
    "440564": "Wow &lt; 20 min, did you prevent object_ids being split across the files?",
    "440605": "yes, I stratified by object_id and using 32 cores :-)",
    "440617": "Same here, I generate a sub in 20 mins using 20 cores.  Parallelism is key.",
    "440817": "Apparently yes! Actually 8GB memory is good enough most of the time. I'll publish a demo later and will keep you updated. Thank you!",
    "440866": "I find Dask to be sufficient most of the time, but more speed never hurts.\nAS for neural network training, I think we cannot really expect everything to fit into memory, especially if you are training an autoenocder rather than working with features. A dynamic loading process like PyTorch DataLoader is really necessary. The trick we used is to expand the raw data into the required format, save them as memmap on disk (75GB for the test set...) then read them as needed. This way you never have to redo groupby's and training starts almost instantly the moment you hit the run button.",
    "440883": "Thanks Jiwei, need the demo, I found the documentation a little sparse, I see you are on the Rapids team at nvidia, I think it is a potential game changer once all the features are implemented.",
    "441707": "Here is a demo https://github.com/daxiongshu/Rapids_PLAsTiCC_2018  My experiment shows 10x speedup than pandas.\nSorry for not getting the fact straight yesterday. Loading the test time series along will cost 11 GB. But you can skip rows to fit the data on GPU. I highlight this part at the beginning of the notebook. We are working on dask-cudf to solve this OOM problem for both single and multi gpu setting.  Let me know if you have any questions!"
  },
  "source": "meta"
}