{
  "id": 170692,
  "title": "Utilizing GPU",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170692",
  "author_name": "Steven C",
  "post_date": "2020-07-28T15:41:37.629000",
  "votes": 5,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Currently I am very simply trying to consolidate my data using the Kaggle kernel. I have enabled the use of GPU, and it says that the GPU is on. However when I run my code, it doesn't look like the GPU is being utilized, and my notebook crashes with the error: Your notebook tried to allocate more memory than is available. It has restarted.</p>\n\n<p>Do I have to initialize anything else in order to use GPU? How can I get more memory? I've only been trying to preprocess 10 images and the notebook is crashing so I don't know how I could possibly even get to larger numbers.</p>",
  "messages": [
    {
      "id": 949374,
      "postDate": "2020-07-28T15:41:37.630Z",
      "content": "<p>Currently I am very simply trying to consolidate my data using the Kaggle kernel. I have enabled the use of GPU, and it says that the GPU is on. However when I run my code, it doesn't look like the GPU is being utilized, and my notebook crashes with the error: Your notebook tried to allocate more memory than is available. It has restarted.</p>\n\n<p>Do I have to initialize anything else in order to use GPU? How can I get more memory? I've only been trying to preprocess 10 images and the notebook is crashing so I don't know how I could possibly even get to larger numbers.</p>",
      "rawMarkdown": "Currently I am very simply trying to consolidate my data using the Kaggle kernel. I have enabled the use of GPU, and it says that the GPU is on. However when I run my code, it doesn't look like the GPU is being utilized, and my notebook crashes with the error: Your notebook tried to allocate more memory than is available. It has restarted.\n\nDo I have to initialize anything else in order to use GPU? How can I get more memory? I've only been trying to preprocess 10 images and the notebook is crashing so I don't know how I could possibly even get to larger numbers.",
      "votes": 5
    },
    {
      "id": 949457,
      "postDate": "2020-07-28T16:39:04.987Z",
      "content": "<p>What image sizes are you using and what CNN are you using? Neither TPU nor GPU can use the full images of size <code>4000x6000x3</code>.  You most likely need to use smaller images and smaller CNN. (Also lower batch size to avoid errors).</p>\n\n<p>Start by making a notebook that uses my <code>128x128</code> resized <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\">JPEGs</a> or <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">TFRecords</a>. Here is an example GPU notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a>. It uses 33,000 train images of size 128x128 with 15,000 addition external images 128x128. And it completes one fold of 20 epochs with 11 TTA and 11 OOF-TTA in 30 minutes with EfficientNetB0. Start from there. It currently scores LB 0.915+ and GPU can easily achieve much more.</p>",
      "rawMarkdown": "What image sizes are you using and what CNN are you using? Neither TPU nor GPU can use the full images of size `4000x6000x3`.  You most likely need to use smaller images and smaller CNN. (Also lower batch size to avoid errors).\n\nStart by making a notebook that uses my `128x128` resized [JPEGs][2] or [TFRecords][3]. Here is an example GPU notebook [here][1]. It uses 33,000 train images of size 128x128 with 15,000 addition external images 128x128. And it completes one fold of 20 epochs with 11 TTA and 11 OOF-TTA in 30 minutes with EfficientNetB0. Start from there. It currently scores LB 0.915+ and GPU can easily achieve much more.\n\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\n[2]: https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\n[3]: https://www.kaggle.com/cdeotte/melanoma-128x128",
      "votes": 3
    },
    {
      "id": 953486,
      "postDate": "2020-07-31T20:27:54.243Z",
      "content": "<p>I was facing the same problem earlier until I found the solution. The SIIM-ISIC dataset is really bad with I/O because the image sizes are huge (on an average 2000 x 3000px). As a result, resizing and other augmentations, as well as performing convolutions takes a lot of time and resources.</p>\n\n<p>The solution is performing all the augmentation/resizing on the images in a separate kernel to create a new dataset and then using this pre-processed data for training. </p>\n\n<p>Here's a link to my pre-processed dataset- <a href=\"https://www.kaggle.com/bravehart101/siimisic-melanoma-512x512-resized\">512 x 512px resized images</a>.</p>\n\n<p>My training time for ResNet50 on raw data was around 1.5 hours per EPOCH(!!!) while on the pre-processed data, the time was around 8.45 minutes per epoch for the same model.</p>",
      "rawMarkdown": "I was facing the same problem earlier until I found the solution. The SIIM-ISIC dataset is really bad with I/O because the image sizes are huge (on an average 2000 x 3000px). As a result, resizing and other augmentations, as well as performing convolutions takes a lot of time and resources.\n\nThe solution is performing all the augmentation/resizing on the images in a separate kernel to create a new dataset and then using this pre-processed data for training. \n\nHere's a link to my pre-processed dataset- [512 x 512px resized images](https://www.kaggle.com/bravehart101/siimisic-melanoma-512x512-resized).\n\nMy training time for ResNet50 on raw data was around 1.5 hours per EPOCH(!!!) while on the pre-processed data, the time was around 8.45 minutes per epoch for the same model.\n\n  ",
      "votes": 1,
      "replies": [
        {
          "id": 953490,
          "postDate": "2020-07-31T20:31:56.480Z",
          "content": "<p>Using TPU and TFRecords will get you down to around 2 minutes or less per epoch</p>",
          "rawMarkdown": "Using TPU and TFRecords will get you down to around 2 minutes or less per epoch",
          "votes": 1
        },
        {
          "id": 954007,
          "postDate": "2020-08-01T09:57:01.913Z",
          "content": "<p>Will definitely check it out for my next submission. Thanks!</p>",
          "rawMarkdown": "Will definitely check it out for my next submission. Thanks!"
        }
      ]
    },
    {
      "id": 949613,
      "postDate": "2020-07-28T18:41:26.220Z",
      "content": "<p>If you use Pytorch  then mixed precision and Gradient Accumulation are your friends. You can train faster with a relatively large batch size. \nIf you use TF, you can follow the examples by <a href=\"/cdeotte\">@cdeotte</a> in public kernels.</p>\n\n<p>However if you wanna  use huge models on higher resolutions images (512, 768) I would suggest TPU for quick experiments. Kaggle TPU config is tailored for Tensorflow while Colab is well suited for Pytorch. </p>",
      "rawMarkdown": "If you use Pytorch  then mixed precision and Gradient Accumulation are your friends. You can train faster with a relatively large batch size. \nIf you use TF, you can follow the examples by @cdeotte in public kernels.\n\nHowever if you wanna  use huge models on higher resolutions images (512, 768) I would suggest TPU for quick experiments. Kaggle TPU config is tailored for Tensorflow while Colab is well suited for Pytorch. \n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 949618,
          "postDate": "2020-07-28T18:44:31.820Z",
          "content": "<p><a href=\"/serigne\">@serigne</a> can you point me to Kaggle kernel with mixed precision training on pytorch?</p>",
          "rawMarkdown": "@serigne can you point me to Kaggle kernel with mixed precision training on pytorch?"
        },
        {
          "id": 949628,
          "postDate": "2020-07-28T18:53:31.783Z",
          "content": "<p>I would use Nvidia's Apex but it seems Pytorch has now native mixed precision .  You can take a look at this <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\">kernel</a></p>",
          "rawMarkdown": "I would use Nvidia's Apex but it seems Pytorch has now native mixed precision .  You can take a look at this [kernel](https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp)",
          "votes": 2
        },
        {
          "id": 949637,
          "postDate": "2020-07-28T19:03:23.247Z",
          "content": "<p>I am aware it exists for long time, but I was never able to learn it and I think it's quite important if I use my own GPU. Thanks, I was skipping Pytorch lightning kernels but maybe it's worth to give it a try</p>",
          "rawMarkdown": "I am aware it exists for long time, but I was never able to learn it and I think it's quite important if I use my own GPU. Thanks, I was skipping Pytorch lightning kernels but maybe it's worth to give it a try"
        },
        {
          "id": 950250,
          "postDate": "2020-07-29T09:44:33.517Z",
          "content": "<p>I just give a try to mixed float precision with 256x256 resolution and eff-b3 (not for this comp), i managed to multiply batchsize by 8 on a V100 card.</p>\n\n<p>Setup is even simpler with pytorch 1.6, just follow the first example : <a href=\"https://pytorch.org/docs/stable/notes/amp_examples.html\">pytorch doc</a></p>",
          "rawMarkdown": "I just give a try to mixed float precision with 256x256 resolution and eff-b3 (not for this comp), i managed to multiply batchsize by 8 on a V100 card.\n\nSetup is even simpler with pytorch 1.6, just follow the first example : [pytorch doc](https://pytorch.org/docs/stable/notes/amp_examples.html)"
        }
      ]
    },
    {
      "id": 949415,
      "postDate": "2020-07-28T16:14:51.427Z",
      "content": "<p>lower your batch size , thats the main problem\nand dont even think about to resize image in the preprocessing part..\nJust take the resized public data and preprocess</p>",
      "rawMarkdown": "lower your batch size , thats the main problem\nand dont even think about to resize image in the preprocessing part..\nJust take the resized public data and preprocess\n",
      "votes": 2
    },
    {
      "id": 951194,
      "postDate": "2020-07-30T01:25:16.393Z",
      "content": "<p>Apart from the crashing bit, GPU usually just shows it's not doing anything but it's actually is being used. You can confirm this if you run your model with lots of convolutions with and without the GPU on. At least that had been my experience with it</p>",
      "rawMarkdown": "Apart from the crashing bit, GPU usually just shows it's not doing anything but it's actually is being used. You can confirm this if you run your model with lots of convolutions with and without the GPU on. At least that had been my experience with it"
    },
    {
      "id": 949529,
      "postDate": "2020-07-28T17:29:59.910Z",
      "content": "<p>Please make your example kernel public so people can see what's wrong.</p>",
      "rawMarkdown": "Please make your example kernel public so people can see what's wrong."
    },
    {
      "id": 953488,
      "postDate": "2020-07-31T20:31:30.057Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 949457,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-28T16:39:04.987000",
      "content": "<p>What image sizes are you using and what CNN are you using? Neither TPU nor GPU can use the full images of size <code>4000x6000x3</code>.  You most likely need to use smaller images and smaller CNN. (Also lower batch size to avoid errors).</p>\n\n<p>Start by making a notebook that uses my <code>128x128</code> resized <a href=\"https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\">JPEGs</a> or <a href=\"https://www.kaggle.com/cdeotte/melanoma-128x128\">TFRecords</a>. Here is an example GPU notebook <a href=\"https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\">here</a>. It uses 33,000 train images of size 128x128 with 15,000 addition external images 128x128. And it completes one fold of 20 epochs with 11 TTA and 11 OOF-TTA in 30 minutes with EfficientNetB0. Start from there. It currently scores LB 0.915+ and GPU can easily achieve much more.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 953486,
      "author_name": "Aman Sharma",
      "author_url": "",
      "post_date": "2020-07-31T20:27:54.243000",
      "content": "<p>I was facing the same problem earlier until I found the solution. The SIIM-ISIC dataset is really bad with I/O because the image sizes are huge (on an average 2000 x 3000px). As a result, resizing and other augmentations, as well as performing convolutions takes a lot of time and resources.</p>\n\n<p>The solution is performing all the augmentation/resizing on the images in a separate kernel to create a new dataset and then using this pre-processed data for training. </p>\n\n<p>Here's a link to my pre-processed dataset- <a href=\"https://www.kaggle.com/bravehart101/siimisic-melanoma-512x512-resized\">512 x 512px resized images</a>.</p>\n\n<p>My training time for ResNet50 on raw data was around 1.5 hours per EPOCH(!!!) while on the pre-processed data, the time was around 8.45 minutes per epoch for the same model.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 953490,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-31T20:31:56.480000",
          "content": "<p>Using TPU and TFRecords will get you down to around 2 minutes or less per epoch</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 954007,
          "author_name": "Aman Sharma",
          "author_url": "",
          "post_date": "2020-08-01T09:57:01.913000",
          "content": "<p>Will definitely check it out for my next submission. Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 949613,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-07-28T18:41:26.220000",
      "content": "<p>If you use Pytorch  then mixed precision and Gradient Accumulation are your friends. You can train faster with a relatively large batch size. \nIf you use TF, you can follow the examples by <a href=\"/cdeotte\">@cdeotte</a> in public kernels.</p>\n\n<p>However if you wanna  use huge models on higher resolutions images (512, 768) I would suggest TPU for quick experiments. Kaggle TPU config is tailored for Tensorflow while Colab is well suited for Pytorch. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 949618,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-28T18:44:31.820000",
          "content": "<p><a href=\"/serigne\">@serigne</a> can you point me to Kaggle kernel with mixed precision training on pytorch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 949628,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-07-28T18:53:31.783000",
          "content": "<p>I would use Nvidia's Apex but it seems Pytorch has now native mixed precision .  You can take a look at this <a href=\"https://www.kaggle.com/hmendonca/melanoma-neat-pytorch-lightning-native-amp\">kernel</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 949637,
          "author_name": "Jacek Poplawski",
          "author_url": "",
          "post_date": "2020-07-28T19:03:23.247000",
          "content": "<p>I am aware it exists for long time, but I was never able to learn it and I think it's quite important if I use my own GPU. Thanks, I was skipping Pytorch lightning kernels but maybe it's worth to give it a try</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 950250,
          "author_name": "AnhMeow",
          "author_url": "",
          "post_date": "2020-07-29T09:44:33.517000",
          "content": "<p>I just give a try to mixed float precision with 256x256 resolution and eff-b3 (not for this comp), i managed to multiply batchsize by 8 on a V100 card.</p>\n\n<p>Setup is even simpler with pytorch 1.6, just follow the first example : <a href=\"https://pytorch.org/docs/stable/notes/amp_examples.html\">pytorch doc</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 949415,
      "author_name": "MhdSharuk",
      "author_url": "",
      "post_date": "2020-07-28T16:14:51.427000",
      "content": "<p>lower your batch size , thats the main problem\nand dont even think about to resize image in the preprocessing part..\nJust take the resized public data and preprocess</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 951194,
      "author_name": "Anh Le",
      "author_url": "",
      "post_date": "2020-07-30T01:25:16.393000",
      "content": "<p>Apart from the crashing bit, GPU usually just shows it's not doing anything but it's actually is being used. You can confirm this if you run your model with lots of convolutions with and without the GPU on. At least that had been my experience with it</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 949529,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2020-07-28T17:29:59.910000",
      "content": "<p>Please make your example kernel public so people can see what's wrong.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 953488,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-31T20:31:30.057000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "949374": "Currently I am very simply trying to consolidate my data using the Kaggle kernel. I have enabled the use of GPU, and it says that the GPU is on. However when I run my code, it doesn't look like the GPU is being utilized, and my notebook crashes with the error: Your notebook tried to allocate more memory than is available. It has restarted.\n\nDo I have to initialize anything else in order to use GPU? How can I get more memory? I've only been trying to preprocess 10 images and the notebook is crashing so I don't know how I could possibly even get to larger numbers.",
    "949457": "What image sizes are you using and what CNN are you using? Neither TPU nor GPU can use the full images of size `4000x6000x3`.  You most likely need to use smaller images and smaller CNN. (Also lower batch size to avoid errors).\n\nStart by making a notebook that uses my `128x128` resized [JPEGs][2] or [TFRecords][3]. Here is an example GPU notebook [here][1]. It uses 33,000 train images of size 128x128 with 15,000 addition external images 128x128. And it completes one fold of 20 epochs with 11 TTA and 11 OOF-TTA in 30 minutes with EfficientNetB0. Start from there. It currently scores LB 0.915+ and GPU can easily achieve much more.\n\n[1]: https://www.kaggle.com/cdeotte/tfrecord-experiments-upsample-and-coarse-dropout\n[2]: https://www.kaggle.com/cdeotte/jpeg-melanoma-128x128\n[3]: https://www.kaggle.com/cdeotte/melanoma-128x128",
    "953486": "I was facing the same problem earlier until I found the solution. The SIIM-ISIC dataset is really bad with I/O because the image sizes are huge (on an average 2000 x 3000px). As a result, resizing and other augmentations, as well as performing convolutions takes a lot of time and resources.\n\nThe solution is performing all the augmentation/resizing on the images in a separate kernel to create a new dataset and then using this pre-processed data for training. \n\nHere's a link to my pre-processed dataset- [512 x 512px resized images](https://www.kaggle.com/bravehart101/siimisic-melanoma-512x512-resized).\n\nMy training time for ResNet50 on raw data was around 1.5 hours per EPOCH(!!!) while on the pre-processed data, the time was around 8.45 minutes per epoch for the same model.\n\n  ",
    "949613": "If you use Pytorch  then mixed precision and Gradient Accumulation are your friends. You can train faster with a relatively large batch size. \nIf you use TF, you can follow the examples by @cdeotte in public kernels.\n\nHowever if you wanna  use huge models on higher resolutions images (512, 768) I would suggest TPU for quick experiments. Kaggle TPU config is tailored for Tensorflow while Colab is well suited for Pytorch. \n\n\n",
    "949415": "lower your batch size , thats the main problem\nand dont even think about to resize image in the preprocessing part..\nJust take the resized public data and preprocess\n",
    "951194": "Apart from the crashing bit, GPU usually just shows it's not doing anything but it's actually is being used. You can confirm this if you run your model with lots of convolutions with and without the GPU on. At least that had been my experience with it",
    "949529": "Please make your example kernel public so people can see what's wrong.",
    "953488": ""
  }
}