{
  "id": 172373,
  "title": "rapids question",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/172373",
  "author_name": "",
  "post_date": "2020-08-04T20:11:33.538244300Z",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hei!</p>\n\n<p>I've very recently started to learn how to implement RAPIDS in my workflows, however I don't think I am doing it right, as my GPU is not used at all.</p>\n\n<p>The idea is to use <code>cudf</code> library to extract the metadata from the .dcm files (there are 289,826 dicom images in the train data only).</p>\n\n<p>After I've extracted all raw paths in <code>paths</code> variable and created a <code>get_observation_data()</code> function that returns a dictionary with all the information, I wrote this bit of code:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3564129%2Ff811eac167783a50d1cef263607ae9d4%2FAnnotation%202020-08-04%20211532.png?generation=1596572171529569&amp;alt=media\" alt=\"\"></p>\n\n<p>However, the GPU is not used at all and the process is super slow. I can't seem to find anything on Google the Big Guy and I am super new in this topic. Can anybody help me in explaining why is this happening?</p>\n\n<p>Thanks a bunch!</p>",
  "messages": [
    {
      "id": "958169",
      "postDate": "08/04/2020 20:11:33",
      "content": "<p>Hei!</p>\n\n<p>I've very recently started to learn how to implement RAPIDS in my workflows, however I don't think I am doing it right, as my GPU is not used at all.</p>\n\n<p>The idea is to use <code>cudf</code> library to extract the metadata from the .dcm files (there are 289,826 dicom images in the train data only).</p>\n\n<p>After I've extracted all raw paths in <code>paths</code> variable and created a <code>get_observation_data()</code> function that returns a dictionary with all the information, I wrote this bit of code:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3564129%2Ff811eac167783a50d1cef263607ae9d4%2FAnnotation%202020-08-04%20211532.png?generation=1596572171529569&amp;alt=media\" alt=\"\"></p>\n\n<p>However, the GPU is not used at all and the process is super slow. I can't seem to find anything on Google the Big Guy and I am super new in this topic. Can anybody help me in explaining why is this happening?</p>\n\n<p>Thanks a bunch!</p>",
      "rawMarkdown": "Hei!\n\nI've very recently started to learn how to implement RAPIDS in my workflows, however I don't think I am doing it right, as my GPU is not used at all.\n\nThe idea is to use `cudf` library to extract the metadata from the .dcm files (there are 289,826 dicom images in the train data only).\n\nAfter I've extracted all raw paths in `paths` variable and created a `get_observation_data()` function that returns a dictionary with all the information, I wrote this bit of code:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3564129%2Ff811eac167783a50d1cef263607ae9d4%2FAnnotation%202020-08-04%20211532.png?generation=1596572171529569&amp;alt=media)\n\nHowever, the GPU is not used at all and the process is super slow. I can't seem to find anything on Google the Big Guy and I am super new in this topic. Can anybody help me in explaining why is this happening?\n\nThanks a bunch!",
      "votes": null
    },
    {
      "id": "958194",
      "postDate": "08/04/2020 20:41:26",
      "content": "<p>Not sure what's in your get_observation_data() but that's worth checking since it's loading the .dcm files. Make sure it's not loading the pixels thats the slow bit. I also need to learn the rapids stuff at some point! </p>",
      "rawMarkdown": "Not sure what's in your get_observation_data() but that's worth checking since it's loading the .dcm files. Make sure it's not loading the pixels thats the slow bit. I also need to learn the rapids stuff at some point!",
      "votes": null
    },
    {
      "id": "958952",
      "postDate": "08/05/2020 08:32:06",
      "content": "<p>Hi, are you running in Kaggle notebook or local?  maybe we need to see your configuration as well.. and what <code>get_observation_data</code> does (I assume is where you extract the pixel data from the images - btw I'm not participating in the competition) </p>\n\n<p>I'd suggest to start by creating a random <code>cudf</code> from the examples page and check with <code>!nvidia-smi</code> if that works in gpu </p>",
      "rawMarkdown": "Hi, are you running in Kaggle notebook or local?  maybe we need to see your configuration as well.. and what `get_observation_data` does (I assume is where you extract the pixel data from the images - btw I'm not participating in the competition) \n\nI'd suggest to start by creating a random `cudf` from the examples page and check with `!nvidia-smi` if that works in gpu",
      "votes": null
    },
    {
      "id": "959042",
      "postDate": "08/05/2020 09:49:44",
      "content": "<p>Hei, thanks for your answer! I think the issue is in the function as well, I need to change it for various reasons (<a href=\"https://www.kaggle.com/andradaolteanu/pulmonary-fibrosis-competition-eda-dicom-prep\">here is the notebook, chapter 4 is where I use rapids and create the function</a>). Problem is I don't know yet how to use rapids to efficiently extract the data from the files. But I'm getting there!</p>",
      "rawMarkdown": "Hei, thanks for your answer! I think the issue is in the function as well, I need to change it for various reasons ([here is the notebook, chapter 4 is where I use rapids and create the function](https://www.kaggle.com/andradaolteanu/pulmonary-fibrosis-competition-eda-dicom-prep)). Problem is I don't know yet how to use rapids to efficiently extract the data from the files. But I'm getting there!",
      "votes": null
    },
    {
      "id": "959125",
      "postDate": "08/05/2020 11:20:51",
      "content": "<p>You wrote a code that has quadratic runtime complexity.  It would be slow whatever dataframe framework you use.  Indeed, at each concat you copy all the data from the concatenated dataframe.  This data grows linearly with the number of concats, hence you are summing a linearly increasing series.  Result is quadratic.</p>\n\n<p>I fixed it by first collecting all your dictionaries in a list, then constructing a data frame from it.  Once the dictionaries are built dataframe creation is almost instantaneuous.</p>\n\n<p><a href=\"https://www.kaggle.com/cpmpml/pulmonary-fibrosis-competition-eda-dicom-prep\">https://www.kaggle.com/cpmpml/pulmonary-fibrosis-competition-eda-dicom-prep</a></p>",
      "rawMarkdown": "You wrote a code that has quadratic runtime complexity.  It would be slow whatever dataframe framework you use.  Indeed, at each concat you copy all the data from the concatenated dataframe.  This data grows linearly with the number of concats, hence you are summing a linearly increasing series.  Result is quadratic.\n\nI fixed it by first collecting all your dictionaries in a list, then constructing a data frame from it.  Once the dictionaries are built dataframe creation is almost instantaneuous.\n\nhttps://www.kaggle.com/cpmpml/pulmonary-fibrosis-competition-eda-dicom-prep",
      "votes": null
    },
    {
      "id": "959337",
      "postDate": "08/05/2020 13:57:27",
      "content": "<p>Amazing! Thank you sooo much 🙏 I thought by using <code>cudf</code> within the loop I can make things go faster; I was clearly wrong. Thank you again for your help!</p>",
      "rawMarkdown": "Amazing! Thank you sooo much 🙏 I thought by using `cudf` within the loop I can make things go faster; I was clearly wrong. Thank you again for your help!",
      "votes": null
    },
    {
      "id": "959467",
      "postDate": "08/05/2020 15:51:58",
      "content": "<p>cuDf is not the issue.  Your code would have been slow with pandas inside the loop as well.  Anyway, happy to have helped, looking forward to seeing future versions of your notebook.</p>",
      "rawMarkdown": "cuDf is not the issue.  Your code would have been slow with pandas inside the loop as well.  Anyway, happy to have helped, looking forward to seeing future versions of your notebook.",
      "votes": null
    },
    {
      "id": "964577",
      "postDate": "08/10/2020 02:58:28",
      "content": "<p>what does metadata actually constitute? Because I can't find it on Data page of Competition Website</p>",
      "rawMarkdown": "what does metadata actually constitute? Because I can't find it on Data page of Competition Website",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 958194,
      "author_name": "jameschapman19",
      "author_url": "",
      "post_date": "08/04/2020 20:41:26",
      "content": "<p>Not sure what's in your get_observation_data() but that's worth checking since it's loading the .dcm files. Make sure it's not loading the pixels thats the slow bit. I also need to learn the rapids stuff at some point! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 958952,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "08/05/2020 08:32:06",
      "content": "<p>Hi, are you running in Kaggle notebook or local?  maybe we need to see your configuration as well.. and what <code>get_observation_data</code> does (I assume is where you extract the pixel data from the images - btw I'm not participating in the competition) </p>\n\n<p>I'd suggest to start by creating a random <code>cudf</code> from the examples page and check with <code>!nvidia-smi</code> if that works in gpu </p>",
      "votes": null,
      "replies": [
        {
          "id": 959042,
          "author_name": "andradaolteanu",
          "author_url": "",
          "post_date": "08/05/2020 09:49:44",
          "content": "<p>Hei, thanks for your answer! I think the issue is in the function as well, I need to change it for various reasons (<a href=\"https://www.kaggle.com/andradaolteanu/pulmonary-fibrosis-competition-eda-dicom-prep\">here is the notebook, chapter 4 is where I use rapids and create the function</a>). Problem is I don't know yet how to use rapids to efficiently extract the data from the files. But I'm getting there!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 959125,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "08/05/2020 11:20:51",
      "content": "<p>You wrote a code that has quadratic runtime complexity.  It would be slow whatever dataframe framework you use.  Indeed, at each concat you copy all the data from the concatenated dataframe.  This data grows linearly with the number of concats, hence you are summing a linearly increasing series.  Result is quadratic.</p>\n\n<p>I fixed it by first collecting all your dictionaries in a list, then constructing a data frame from it.  Once the dictionaries are built dataframe creation is almost instantaneuous.</p>\n\n<p><a href=\"https://www.kaggle.com/cpmpml/pulmonary-fibrosis-competition-eda-dicom-prep\">https://www.kaggle.com/cpmpml/pulmonary-fibrosis-competition-eda-dicom-prep</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 959337,
          "author_name": "andradaolteanu",
          "author_url": "",
          "post_date": "08/05/2020 13:57:27",
          "content": "<p>Amazing! Thank you sooo much 🙏 I thought by using <code>cudf</code> within the loop I can make things go faster; I was clearly wrong. Thank you again for your help!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 959467,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/05/2020 15:51:58",
          "content": "<p>cuDf is not the issue.  Your code would have been slow with pandas inside the loop as well.  Anyway, happy to have helped, looking forward to seeing future versions of your notebook.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 964577,
      "author_name": "digvijayyadav",
      "author_url": "",
      "post_date": "08/10/2020 02:58:28",
      "content": "<p>what does metadata actually constitute? Because I can't find it on Data page of Competition Website</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "958169": "Hei!\n\nI've very recently started to learn how to implement RAPIDS in my workflows, however I don't think I am doing it right, as my GPU is not used at all.\n\nThe idea is to use `cudf` library to extract the metadata from the .dcm files (there are 289,826 dicom images in the train data only).\n\nAfter I've extracted all raw paths in `paths` variable and created a `get_observation_data()` function that returns a dictionary with all the information, I wrote this bit of code:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3564129%2Ff811eac167783a50d1cef263607ae9d4%2FAnnotation%202020-08-04%20211532.png?generation=1596572171529569&amp;alt=media)\n\nHowever, the GPU is not used at all and the process is super slow. I can't seem to find anything on Google the Big Guy and I am super new in this topic. Can anybody help me in explaining why is this happening?\n\nThanks a bunch!",
    "958194": "Not sure what's in your get_observation_data() but that's worth checking since it's loading the .dcm files. Make sure it's not loading the pixels thats the slow bit. I also need to learn the rapids stuff at some point!",
    "958952": "Hi, are you running in Kaggle notebook or local?  maybe we need to see your configuration as well.. and what `get_observation_data` does (I assume is where you extract the pixel data from the images - btw I'm not participating in the competition) \n\nI'd suggest to start by creating a random `cudf` from the examples page and check with `!nvidia-smi` if that works in gpu",
    "959042": "Hei, thanks for your answer! I think the issue is in the function as well, I need to change it for various reasons ([here is the notebook, chapter 4 is where I use rapids and create the function](https://www.kaggle.com/andradaolteanu/pulmonary-fibrosis-competition-eda-dicom-prep)). Problem is I don't know yet how to use rapids to efficiently extract the data from the files. But I'm getting there!",
    "959125": "You wrote a code that has quadratic runtime complexity.  It would be slow whatever dataframe framework you use.  Indeed, at each concat you copy all the data from the concatenated dataframe.  This data grows linearly with the number of concats, hence you are summing a linearly increasing series.  Result is quadratic.\n\nI fixed it by first collecting all your dictionaries in a list, then constructing a data frame from it.  Once the dictionaries are built dataframe creation is almost instantaneuous.\n\nhttps://www.kaggle.com/cpmpml/pulmonary-fibrosis-competition-eda-dicom-prep",
    "959337": "Amazing! Thank you sooo much 🙏 I thought by using `cudf` within the loop I can make things go faster; I was clearly wrong. Thank you again for your help!",
    "959467": "cuDf is not the issue.  Your code would have been slow with pandas inside the loop as well.  Anyway, happy to have helped, looking forward to seeing future versions of your notebook.",
    "964577": "what does metadata actually constitute? Because I can't find it on Data page of Competition Website"
  },
  "source": "meta"
}