{
  "id": 168489,
  "title": "Tips for training big datasets",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/168489",
  "author_name": "",
  "post_date": "2020-07-20T21:03:37.406523500Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello all, this is my first competition on Kaggle. I was struggling with the sheer training time of my data: I am using PyTorch/fastai and initially reduced the samples I am training on to 100, but it still shockingly took ~10 mins per epoch (on a GPU session).</p>\n\n<p>Are there certain tricks to make training on such a huge dataset feasible? I've tried mixed-precision training, and that helped, but training a model on most of the dataset still seems unfeasible. Thanks in advance :)</p>",
  "messages": [
    {
      "id": "937244",
      "postDate": "07/20/2020 21:03:37",
      "content": "<p>Hello all, this is my first competition on Kaggle. I was struggling with the sheer training time of my data: I am using PyTorch/fastai and initially reduced the samples I am training on to 100, but it still shockingly took ~10 mins per epoch (on a GPU session).</p>\n\n<p>Are there certain tricks to make training on such a huge dataset feasible? I've tried mixed-precision training, and that helped, but training a model on most of the dataset still seems unfeasible. Thanks in advance :)</p>",
      "rawMarkdown": "Hello all, this is my first competition on Kaggle. I was struggling with the sheer training time of my data: I am using PyTorch/fastai and initially reduced the samples I am training on to 100, but it still shockingly took ~10 mins per epoch (on a GPU session).\n\nAre there certain tricks to make training on such a huge dataset feasible? I've tried mixed-precision training, and that helped, but training a model on most of the dataset still seems unfeasible. Thanks in advance :)",
      "votes": null
    },
    {
      "id": "937277",
      "postDate": "07/20/2020 22:25:38",
      "content": "<p>Hi! The original dataset for this competition has really big images so to accelerate training and experimentation you can use one <a href=\"/cdeotte\">@cdeotte</a> 's resized datasets that are in this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">discussion</a></p>\n\n<p>And another great way to run faster iterations is using TPUs! Here is a helper <a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\">notebook</a> from <a href=\"/abhishek\">@abhishek</a> where he uses PyTorch with TPUs.</p>\n\n<p>Hope this helps!</p>",
      "rawMarkdown": "Hi! The original dataset for this competition has really big images so to accelerate training and experimentation you can use one @cdeotte 's resized datasets that are in this [discussion](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579)\n\nAnd another great way to run faster iterations is using TPUs! Here is a helper [notebook](https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu) from @abhishek where he uses PyTorch with TPUs.\n\nHope this helps!",
      "votes": null
    },
    {
      "id": "937343",
      "postDate": "07/21/2020 00:20:43",
      "content": "<p>I like to debug using reduced image size rather than number of samples.  I feel like I learn even with a 32x32 image but never learn when only running 100 of the 30,000 images.  </p>\n\n<p>I do most of my development on local machines so lot less pain handling the quota, memory size restrictions.  </p>\n\n<p>Once the image size is too big for local PC than I can move to Kaggle TPU.</p>",
      "rawMarkdown": "I like to debug using reduced image size rather than number of samples.  I feel like I learn even with a 32x32 image but never learn when only running 100 of the 30,000 images.  \n\nI do most of my development on local machines so lot less pain handling the quota, memory size restrictions.  \n\nOnce the image size is too big for local PC than I can move to Kaggle TPU.",
      "votes": null
    },
    {
      "id": "937349",
      "postDate": "07/21/2020 00:30:20",
      "content": "<p>I ran into this problem as well using Keras image generator and figured out that resizing each batch of images on the fly was taking really long (20s per batch). I ended up pre-processing and resizing all the images in one go into the working folder, train the model then delete them at the end to keep everything clean</p>",
      "rawMarkdown": "I ran into this problem as well using Keras image generator and figured out that resizing each batch of images on the fly was taking really long (20s per batch). I ended up pre-processing and resizing all the images in one go into the working folder, train the model then delete them at the end to keep everything clean",
      "votes": null
    },
    {
      "id": "937487",
      "postDate": "07/21/2020 03:47:18",
      "content": "<p>I was remined from another post I just answered that my resize statement needed to be changed before I got any speed.  It's a tensorflow statement but there should be something similiar for you.  The nearest_neighbor callout was the change that got things zipping at a decent speed.</p>\n\n<p><code>image = tf.image.resize(image, [b, b],\n                                method=tf.image.ResizeMethod.NEAREST_NEIGHBOR)</code></p>",
      "rawMarkdown": "I was remined from another post I just answered that my resize statement needed to be changed before I got any speed.  It's a tensorflow statement but there should be something similiar for you.  The nearest_neighbor callout was the change that got things zipping at a decent speed.\n\n`    image = tf.image.resize(image, [b, b],\n                                method=tf.image.ResizeMethod.NEAREST_NEIGHBOR)`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 937277,
      "author_name": "santiviquez",
      "author_url": "",
      "post_date": "07/20/2020 22:25:38",
      "content": "<p>Hi! The original dataset for this competition has really big images so to accelerate training and experimentation you can use one <a href=\"/cdeotte\">@cdeotte</a> 's resized datasets that are in this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579\">discussion</a></p>\n\n<p>And another great way to run faster iterations is using TPUs! Here is a helper <a href=\"https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu\">notebook</a> from <a href=\"/abhishek\">@abhishek</a> where he uses PyTorch with TPUs.</p>\n\n<p>Hope this helps!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937343,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/21/2020 00:20:43",
      "content": "<p>I like to debug using reduced image size rather than number of samples.  I feel like I learn even with a 32x32 image but never learn when only running 100 of the 30,000 images.  </p>\n\n<p>I do most of my development on local machines so lot less pain handling the quota, memory size restrictions.  </p>\n\n<p>Once the image size is too big for local PC than I can move to Kaggle TPU.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937349,
      "author_name": "anhvle",
      "author_url": "",
      "post_date": "07/21/2020 00:30:20",
      "content": "<p>I ran into this problem as well using Keras image generator and figured out that resizing each batch of images on the fly was taking really long (20s per batch). I ended up pre-processing and resizing all the images in one go into the working folder, train the model then delete them at the end to keep everything clean</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 937487,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "07/21/2020 03:47:18",
      "content": "<p>I was remined from another post I just answered that my resize statement needed to be changed before I got any speed.  It's a tensorflow statement but there should be something similiar for you.  The nearest_neighbor callout was the change that got things zipping at a decent speed.</p>\n\n<p><code>image = tf.image.resize(image, [b, b],\n                                method=tf.image.ResizeMethod.NEAREST_NEIGHBOR)</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "937244": "Hello all, this is my first competition on Kaggle. I was struggling with the sheer training time of my data: I am using PyTorch/fastai and initially reduced the samples I am training on to 100, but it still shockingly took ~10 mins per epoch (on a GPU session).\n\nAre there certain tricks to make training on such a huge dataset feasible? I've tried mixed-precision training, and that helped, but training a model on most of the dataset still seems unfeasible. Thanks in advance :)",
    "937277": "Hi! The original dataset for this competition has really big images so to accelerate training and experimentation you can use one @cdeotte 's resized datasets that are in this [discussion](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155579)\n\nAnd another great way to run faster iterations is using TPUs! Here is a helper [notebook](https://www.kaggle.com/abhishek/accelerator-power-hour-pytorch-tpu) from @abhishek where he uses PyTorch with TPUs.\n\nHope this helps!",
    "937343": "I like to debug using reduced image size rather than number of samples.  I feel like I learn even with a 32x32 image but never learn when only running 100 of the 30,000 images.  \n\nI do most of my development on local machines so lot less pain handling the quota, memory size restrictions.  \n\nOnce the image size is too big for local PC than I can move to Kaggle TPU.",
    "937349": "I ran into this problem as well using Keras image generator and figured out that resizing each batch of images on the fly was taking really long (20s per batch). I ended up pre-processing and resizing all the images in one go into the working folder, train the model then delete them at the end to keep everything clean",
    "937487": "I was remined from another post I just answered that my resize statement needed to be changed before I got any speed.  It's a tensorflow statement but there should be something similiar for you.  The nearest_neighbor callout was the change that got things zipping at a decent speed.\n\n`    image = tf.image.resize(image, [b, b],\n                                method=tf.image.ResizeMethod.NEAREST_NEIGHBOR)`"
  },
  "source": "meta"
}