{
  "id": 160276,
  "title": "Help maximizing GPU. One epoch takes 5 hours with GPU.",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/160276",
  "author_name": "Edward Kahara",
  "post_date": "2020-06-20T14:56:08.904000",
  "votes": 3,
  "comment_count": 27,
  "views": 0,
  "content": "<p>I'm working on this competition with fast.ai. I'm training my model with the images in the jpeg folder.  Even with kaggle's GPU, it takes 5 hours to run one complete epoch. I can't run more than one epoch because it will go past the 9 hour session limit. Is there a way around this? If you're using pytorch or another library, how long does one epoch take? </p>",
  "messages": [
    {
      "id": 896241,
      "postDate": "2020-06-22T02:27:28.990Z",
      "content": "<p>The first thing i would do is time how long your data loader takes. For example (note there is no training, just loading):</p>\n\n<pre><code>import time\nstart = time.time()\nfor i,batch in enumerate(dataloader):\n    print(i, ', ',end='')\n    if i&gt;=STEPS_PER_EPOCH: break\nprint('Elapsed time minutes = ' (time.time()-start)/60.)\n</code></pre>\n\n<p>This should complete very fast. This will determine if your dataloader or training is the bottle neck. If your dataloader is the bottle neck, then speed it up. If your training is the bottle neck then (1) try smaller image size (2) try smaller model (3) freeze the bottom layers</p>\n\n<p>When you freeze bottom layers, then your GPU requires less memory and trains faster.</p>",
      "rawMarkdown": "The first thing i would do is time how long your data loader takes. For example (note there is no training, just loading):\n\n    import time\n    start = time.time()\n    for i,batch in enumerate(dataloader):\n        print(i, ', ',end='')\n        if i&gt;=STEPS_PER_EPOCH: break\n    print('Elapsed time minutes = ' (time.time()-start)/60.)\n\nThis should complete very fast. This will determine if your dataloader or training is the bottle neck. If your dataloader is the bottle neck, then speed it up. If your training is the bottle neck then (1) try smaller image size (2) try smaller model (3) freeze the bottom layers\n\nWhen you freeze bottom layers, then your GPU requires less memory and trains faster.\n",
      "votes": 8,
      "replies": [
        {
          "id": 896275,
          "postDate": "2020-06-22T03:37:49.817Z",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for such wonderful advice. Could I please ask if you could please quantize 'very fast'? For me this takes 50 seconds? I assume this is very slow and this process should take around 1-2secs if <code>STEPS_PER_EPOCH==64</code>.</p>",
          "rawMarkdown": "Thanks @cdeotte  for such wonderful advice. Could I please ask if you could please quantize 'very fast'? For me this takes 50 seconds? I assume this is very slow and this process should take around 1-2secs if `STEPS_PER_EPOCH==64`.",
          "votes": 2
        },
        {
          "id": 896285,
          "postDate": "2020-06-22T03:51:34.167Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 896291,
          "postDate": "2020-06-22T03:57:41.473Z",
          "content": "<p>50 seconds is very fast. Now train your model and see how long each epoch takes. If your data loader by itself is faster than your model's training (which uses both data loader and GPU train), then you have no problem with your data loader.</p>\n\n<p>The data loader runs on CPU and training runs on GPU/TPU. So each (data loader and GPU/TPU training) run at the same time. When you train your model (and use both data loader and training), the time for each epoch is the greater of data loader or train, i.e. <code>max(data loader, GPU/TPU train time)</code>. Therefore when your data loader by itself is faster than train (both dataloader and train running at same time), then your data loader is not a problem.</p>",
          "rawMarkdown": "50 seconds is very fast. Now train your model and see how long each epoch takes. If your data loader by itself is faster than your model's training (which uses both data loader and GPU train), then you have no problem with your data loader.\n\nThe data loader runs on CPU and training runs on GPU/TPU. So each (data loader and GPU/TPU training) run at the same time. When you train your model (and use both data loader and training), the time for each epoch is the greater of data loader or train, i.e. `max(data loader, GPU/TPU train time)`. Therefore when your data loader by itself is faster than train (both dataloader and train running at same time), then your data loader is not a problem.",
          "votes": 3
        },
        {
          "id": 896436,
          "postDate": "2020-06-22T07:12:40.330Z",
          "content": "<p>I received some advice <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159580\">earlier</a> to reduce the image sizes beforehand so that fast.ai doesn't take time resizing images in each batch during training. I'm trying that right now. I'll also try timing how long the dataloader takes. Thank you!</p>",
          "rawMarkdown": "I received some advice [earlier](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159580) to reduce the image sizes beforehand so that fast.ai doesn't take time resizing images in each batch during training. I'm trying that right now. I'll also try timing how long the dataloader takes. Thank you!"
        },
        {
          "id": 911320,
          "postDate": "2020-07-01T16:53:16.187Z",
          "content": "<p>```\nclass Melanoma_Dataset(Dataset):\n    def <strong>init</strong>(self, df, transform=None,train=True):\n        self.df = df\n        self.transform = transform\n        self.train = train</p>\n\n<pre><code>def __len__(self):\n    return (self.df).shape[0]\n\ndef __getitem__(self, index):\n    image = Image.open(INPUT_DIR + 'jpeg/train/'+self.df['image_name'][index]+'.jpg')\n    sex = SexToEnc[self.df['sex'][index]]\n    age = self.df['age_approx'][index]/self.df['age_approx'].max()\n    site = SiteToEnc[self.df['anatom_site_general_challenge'][index]]\n    features = torch.tensor(list(itertools.chain(sex, [age], site)))\n    if self.transform:\n        image = self.transform(image)\n    if self.train:\n        target = self.df['target'][index]\n        return image.float(),features.float(),target \n    else:\n        return image.float(),features.float()\n</code></pre>\n\n<p>```</p>\n\n<p>where SexToEnc, SiteToEnc are maps and\n<code>\ntransform = transforms.Compose([transforms.Resize((256,256)),\n                                transforms.ToTensor(),\n                                transforms.Normalize(mean=[0.485, 0.456, 0.406],\n                                 std=[0.229, 0.224, 0.225])])\n</code>\nI am using this to load my data. But it is taking too much time, can you suggest what changes should I make to reduce the time?</p>",
          "rawMarkdown": "```\nclass Melanoma_Dataset(Dataset):\n    def __init__(self, df, transform=None,train=True):\n        self.df = df\n        self.transform = transform\n        self.train = train\n        \n    def __len__(self):\n        return (self.df).shape[0]\n    \n    def __getitem__(self, index):\n        image = Image.open(INPUT_DIR + 'jpeg/train/'+self.df['image_name'][index]+'.jpg')\n        sex = SexToEnc[self.df['sex'][index]]\n        age = self.df['age_approx'][index]/self.df['age_approx'].max()\n        site = SiteToEnc[self.df['anatom_site_general_challenge'][index]]\n        features = torch.tensor(list(itertools.chain(sex, [age], site)))\n        if self.transform:\n            image = self.transform(image)\n        if self.train:\n            target = self.df['target'][index]\n            return image.float(),features.float(),target \n        else:\n            return image.float(),features.float()\n```\n\n\nwhere SexToEnc, SiteToEnc are maps and\n```\ntransform = transforms.Compose([transforms.Resize((256,256)),\n                                transforms.ToTensor(),\n                                transforms.Normalize(mean=[0.485, 0.456, 0.406],\n                                 std=[0.229, 0.224, 0.225])])\n```\nI am using this to load my data. But it is taking too much time, can you suggest what changes should I make to reduce the time?"
        },
        {
          "id": 911695,
          "postDate": "2020-07-02T00:18:50.200Z",
          "content": "<p>you should NOT be doing ANY pre-processing in training.  Preprocessing should be a step you do ONE TIME.  (Resize, transform, etc).  Then you store those images...............again do not be doing any preprocessing during training!</p>",
          "rawMarkdown": "you should NOT be doing ANY pre-processing in training.  Preprocessing should be a step you do ONE TIME.  (Resize, transform, etc).  Then you store those images...............again do not be doing any preprocessing during training!",
          "votes": -1
        },
        {
          "id": 913624,
          "postDate": "2020-07-03T10:10:54.043Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 914619,
          "postDate": "2020-07-04T05:12:08.670Z",
          "content": "<p>You can try to find the ideal num_workers.  This will GREATLY speed up your Dataloader.  Note that a higher number is not always better, and even using intuition like setting it to the number of physical/virtual cores isn't always the ideal.  So use. your dataset and batch size and do a basic walk of num_workers.  I explain a simple way to do this <a href=\"http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/\">here</a></p>",
          "rawMarkdown": "You can try to find the ideal num_workers.  This will GREATLY speed up your Dataloader.  Note that a higher number is not always better, and even using intuition like setting it to the number of physical/virtual cores isn't always the ideal.  So use. your dataset and batch size and do a basic walk of num_workers.  I explain a simple way to do this [here](http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/)",
          "votes": 1
        },
        {
          "id": 914942,
          "postDate": "2020-07-04T11:22:21.560Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 894591,
      "postDate": "2020-06-20T14:56:08.903Z",
      "content": "<p>I'm working on this competition with fast.ai. I'm training my model with the images in the jpeg folder.  Even with kaggle's GPU, it takes 5 hours to run one complete epoch. I can't run more than one epoch because it will go past the 9 hour session limit. Is there a way around this? If you're using pytorch or another library, how long does one epoch take? </p>",
      "rawMarkdown": "I'm working on this competition with fast.ai. I'm training my model with the images in the jpeg folder.  Even with kaggle's GPU, it takes 5 hours to run one complete epoch. I can't run more than one epoch because it will go past the 9 hour session limit. Is there a way around this? If you're using pytorch or another library, how long does one epoch take? ",
      "votes": 3
    },
    {
      "id": 896950,
      "postDate": "2020-06-22T14:35:14.867Z",
      "content": "<p>I used to use fast.ai for competitions, but sometimes it is difficult to dive into abstraction layers to speed up the code. If you are familiar with PyTorch, I would recommend switching to pure PyTorch code. If you want a high-level API for PyTorch there exist Catalyst and PyTorch Lightning. Using a TPU and PyTorch, one epoch takes about 5 minutes. </p>\n\n<p>As pointed out by <a href=\"/cdeotte\">@cdeotte</a> , Tensorflow with TFRecords are highly competitive as well.</p>",
      "rawMarkdown": "I used to use fast.ai for competitions, but sometimes it is difficult to dive into abstraction layers to speed up the code. If you are familiar with PyTorch, I would recommend switching to pure PyTorch code. If you want a high-level API for PyTorch there exist Catalyst and PyTorch Lightning. Using a TPU and PyTorch, one epoch takes about 5 minutes. \n\nAs pointed out by @cdeotte , Tensorflow with TFRecords are highly competitive as well.",
      "votes": 1,
      "replies": [
        {
          "id": 897003,
          "postDate": "2020-06-22T15:02:54.120Z",
          "content": "<p>I just started deep learning and the fast.ai course seemed like a good place to start. I expect to learn PyTorch and Tensorflow soon. Thank you!</p>",
          "rawMarkdown": "I just started deep learning and the fast.ai course seemed like a good place to start. I expect to learn PyTorch and Tensorflow soon. Thank you!"
        }
      ]
    },
    {
      "id": 896092,
      "postDate": "2020-06-21T20:18:45.317Z",
      "content": "<p>It seems like a long time, can you pass the code and see it?</p>",
      "rawMarkdown": "It seems like a long time, can you pass the code and see it?",
      "votes": 1,
      "replies": [
        {
          "id": 896438,
          "postDate": "2020-06-22T07:13:34.757Z",
          "content": "<p>What do you mean by pass the code?</p>",
          "rawMarkdown": "What do you mean by pass the code?\n"
        },
        {
          "id": 897252,
          "postDate": "2020-06-22T18:25:13.663Z",
          "content": "<p>I mean that you share what you have programmed because it is a long time.  I can only think of two options, that you've done something wrong, that's why I'm asking you for your code, or that you're not really using a GPU</p>",
          "rawMarkdown": "I mean that you share what you have programmed because it is a long time.  I can only think of two options, that you've done something wrong, that's why I'm asking you for your code, or that you're not really using a GPU",
          "votes": 1
        },
        {
          "id": 897288,
          "postDate": "2020-06-22T18:54:11.513Z",
          "content": "<p>I was using a GPU and there was nothing wrong with the code. I just hadn't resized the jpeg images beforehand. I'm working with resized jpeg images that are much smaller with the same code and everything is okay now. Thank you.</p>",
          "rawMarkdown": "I was using a GPU and there was nothing wrong with the code. I just hadn't resized the jpeg images beforehand. I'm working with resized jpeg images that are much smaller with the same code and everything is okay now. Thank you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 895041,
      "postDate": "2020-06-21T04:11:47.183Z",
      "content": "<p>Try to use smaller images sizes.</p>",
      "rawMarkdown": "Try to use smaller images sizes.",
      "votes": 1,
      "replies": [
        {
          "id": 895155,
          "postDate": "2020-06-21T06:29:45.727Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        }
      ]
    },
    {
      "id": 894645,
      "postDate": "2020-06-20T16:12:36.413Z",
      "content": "<p>Hello im using tensorflow tf record format (is faster than loading jpg images because you have the information in binary form). You could use a preprocessed data were you already have your images in numpy arrays and load them to memory (it will wolk with small images (256 should work and not allocate too much memory). This sould definitely make you training faster. Alos keep in mind that a big network with a lot of paramters make the training slow because it have to calculate a lot of gradient witch takes time. Try a simple network and a optimal learning rate for speed.</p>\n\n<p>On the other hand a lot of folks just save their models, and then load them in a new kernel and start straining over that, that's another solution.</p>",
      "rawMarkdown": "Hello im using tensorflow tf record format (is faster than loading jpg images because you have the information in binary form). You could use a preprocessed data were you already have your images in numpy arrays and load them to memory (it will wolk with small images (256 should work and not allocate too much memory). This sould definitely make you training faster. Alos keep in mind that a big network with a lot of paramters make the training slow because it have to calculate a lot of gradient witch takes time. Try a simple network and a optimal learning rate for speed.\n\nOn the other hand a lot of folks just save their models, and then load them in a new kernel and start straining over that, that's another solution.",
      "votes": 1,
      "replies": [
        {
          "id": 894737,
          "postDate": "2020-06-20T17:54:55.443Z",
          "content": "<p>I don't know Tensorflow, so I'll have to learn it and see. \nI like the idea of saving the model and then loading and training over it. Thank you.</p>",
          "rawMarkdown": "I don't know Tensorflow, so I'll have to learn it and see. \nI like the idea of saving the model and then loading and training over it. Thank you.",
          "votes": 1
        }
      ]
    },
    {
      "id": 894621,
      "postDate": "2020-06-20T15:41:29.420Z",
      "content": "<p>I used pytorch and made a very basic model and observed that most of the time is taken in data loading. In my case I observed that loading time is heavily increased if I use pytorch dataloader. I searched the problem and found that many people have found this issue that while loading with data loader time taken is increased by very much but couldnt find a solution. So if someone knows a faster way please tell.</p>",
      "rawMarkdown": "I used pytorch and made a very basic model and observed that most of the time is taken in data loading. In my case I observed that loading time is heavily increased if I use pytorch dataloader. I searched the problem and found that many people have found this issue that while loading with data loader time taken is increased by very much but couldnt find a solution. So if someone knows a faster way please tell.",
      "votes": 1,
      "replies": [
        {
          "id": 894732,
          "postDate": "2020-06-20T17:51:53.147Z",
          "content": "<p>Thank you. I will look into using pytorch.</p>",
          "rawMarkdown": "Thank you. I will look into using pytorch."
        },
        {
          "id": 911603,
          "postDate": "2020-07-01T21:28:30.270Z",
          "content": "<p>how one can split data into melanoma and benign from JPEG folder?</p>",
          "rawMarkdown": "how one can split data into melanoma and benign from JPEG folder?\n "
        },
        {
          "id": 911696,
          "postDate": "2020-07-02T00:20:50.637Z",
          "content": "<p>Please see my blog post on setting the ideal number of num_workers in your Pytorch dataloader:</p>\n\n<p><a href=\"http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/\">http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/</a></p>",
          "rawMarkdown": "Please see my blog post on setting the ideal number of num_workers in your Pytorch dataloader:\n\nhttp://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/\n\n"
        },
        {
          "id": 914314,
          "postDate": "2020-07-03T18:45:59.173Z",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> posted the Tfrecords resized dataset. Loading the jpegs is a bottleneck for the Dataloader. Instead I loaded the Tfrecords into memory and passed them to the Dataloader. This while not elegant, reduced my epoch time from 25 min on Jpegs to 2 min.  </p>",
          "rawMarkdown": "@cdeotte posted the Tfrecords resized dataset. Loading the jpegs is a bottleneck for the Dataloader. Instead I loaded the Tfrecords into memory and passed them to the Dataloader. This while not elegant, reduced my epoch time from 25 min on Jpegs to 2 min.  "
        }
      ]
    },
    {
      "id": 915355,
      "postDate": "2020-07-04T17:02:53.033Z",
      "content": "<p>What size images are you using? I posted 256x256 JPEGS and 256x256 TFRecords <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. Using these smaller images should speed up your GPU</p>",
      "rawMarkdown": "What size images are you using? I posted 256x256 JPEGS and 256x256 TFRecords [here][1]. Using these smaller images should speed up your GPU\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092"
    },
    {
      "id": 896233,
      "postDate": "2020-06-22T02:16:53.617Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 896241,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-06-22T02:27:28.990000",
      "content": "<p>The first thing i would do is time how long your data loader takes. For example (note there is no training, just loading):</p>\n\n<pre><code>import time\nstart = time.time()\nfor i,batch in enumerate(dataloader):\n    print(i, ', ',end='')\n    if i&gt;=STEPS_PER_EPOCH: break\nprint('Elapsed time minutes = ' (time.time()-start)/60.)\n</code></pre>\n\n<p>This should complete very fast. This will determine if your dataloader or training is the bottle neck. If your dataloader is the bottle neck, then speed it up. If your training is the bottle neck then (1) try smaller image size (2) try smaller model (3) freeze the bottom layers</p>\n\n<p>When you freeze bottom layers, then your GPU requires less memory and trains faster.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 896275,
          "author_name": "Aman Arora",
          "author_url": "",
          "post_date": "2020-06-22T03:37:49.817000",
          "content": "<p>Thanks <a href=\"/cdeotte\">@cdeotte</a>  for such wonderful advice. Could I please ask if you could please quantize 'very fast'? For me this takes 50 seconds? I assume this is very slow and this process should take around 1-2secs if <code>STEPS_PER_EPOCH==64</code>.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 896285,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-06-22T03:51:34.167000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 896291,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-06-22T03:57:41.473000",
          "content": "<p>50 seconds is very fast. Now train your model and see how long each epoch takes. If your data loader by itself is faster than your model's training (which uses both data loader and GPU train), then you have no problem with your data loader.</p>\n\n<p>The data loader runs on CPU and training runs on GPU/TPU. So each (data loader and GPU/TPU training) run at the same time. When you train your model (and use both data loader and training), the time for each epoch is the greater of data loader or train, i.e. <code>max(data loader, GPU/TPU train time)</code>. Therefore when your data loader by itself is faster than train (both dataloader and train running at same time), then your data loader is not a problem.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 896436,
          "author_name": "Edward Kahara",
          "author_url": "",
          "post_date": "2020-06-22T07:12:40.330000",
          "content": "<p>I received some advice <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/159580\">earlier</a> to reduce the image sizes beforehand so that fast.ai doesn't take time resizing images in each batch during training. I'm trying that right now. I'll also try timing how long the dataloader takes. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 911320,
          "author_name": "Aman Saxena",
          "author_url": "",
          "post_date": "2020-07-01T16:53:16.187000",
          "content": "<p>```\nclass Melanoma_Dataset(Dataset):\n    def <strong>init</strong>(self, df, transform=None,train=True):\n        self.df = df\n        self.transform = transform\n        self.train = train</p>\n\n<pre><code>def __len__(self):\n    return (self.df).shape[0]\n\ndef __getitem__(self, index):\n    image = Image.open(INPUT_DIR + 'jpeg/train/'+self.df['image_name'][index]+'.jpg')\n    sex = SexToEnc[self.df['sex'][index]]\n    age = self.df['age_approx'][index]/self.df['age_approx'].max()\n    site = SiteToEnc[self.df['anatom_site_general_challenge'][index]]\n    features = torch.tensor(list(itertools.chain(sex, [age], site)))\n    if self.transform:\n        image = self.transform(image)\n    if self.train:\n        target = self.df['target'][index]\n        return image.float(),features.float(),target \n    else:\n        return image.float(),features.float()\n</code></pre>\n\n<p>```</p>\n\n<p>where SexToEnc, SiteToEnc are maps and\n<code>\ntransform = transforms.Compose([transforms.Resize((256,256)),\n                                transforms.ToTensor(),\n                                transforms.Normalize(mean=[0.485, 0.456, 0.406],\n                                 std=[0.229, 0.224, 0.225])])\n</code>\nI am using this to load my data. But it is taking too much time, can you suggest what changes should I make to reduce the time?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 911695,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-02T00:18:50.200000",
          "content": "<p>you should NOT be doing ANY pre-processing in training.  Preprocessing should be a step you do ONE TIME.  (Resize, transform, etc).  Then you store those images...............again do not be doing any preprocessing during training!</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 913624,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-03T10:10:54.043000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 914619,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-04T05:12:08.670000",
          "content": "<p>You can try to find the ideal num_workers.  This will GREATLY speed up your Dataloader.  Note that a higher number is not always better, and even using intuition like setting it to the number of physical/virtual cores isn't always the ideal.  So use. your dataset and batch size and do a basic walk of num_workers.  I explain a simple way to do this <a href=\"http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/\">here</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 914942,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-04T11:22:21.560000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 896950,
      "author_name": "PAB97",
      "author_url": "",
      "post_date": "2020-06-22T14:35:14.867000",
      "content": "<p>I used to use fast.ai for competitions, but sometimes it is difficult to dive into abstraction layers to speed up the code. If you are familiar with PyTorch, I would recommend switching to pure PyTorch code. If you want a high-level API for PyTorch there exist Catalyst and PyTorch Lightning. Using a TPU and PyTorch, one epoch takes about 5 minutes. </p>\n\n<p>As pointed out by <a href=\"/cdeotte\">@cdeotte</a> , Tensorflow with TFRecords are highly competitive as well.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 897003,
          "author_name": "Edward Kahara",
          "author_url": "",
          "post_date": "2020-06-22T15:02:54.120000",
          "content": "<p>I just started deep learning and the fast.ai course seemed like a good place to start. I expect to learn PyTorch and Tensorflow soon. Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 896092,
      "author_name": "Maximofn",
      "author_url": "",
      "post_date": "2020-06-21T20:18:45.317000",
      "content": "<p>It seems like a long time, can you pass the code and see it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 896438,
          "author_name": "Edward Kahara",
          "author_url": "",
          "post_date": "2020-06-22T07:13:34.757000",
          "content": "<p>What do you mean by pass the code?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 897252,
          "author_name": "Maximofn",
          "author_url": "",
          "post_date": "2020-06-22T18:25:13.663000",
          "content": "<p>I mean that you share what you have programmed because it is a long time.  I can only think of two options, that you've done something wrong, that's why I'm asking you for your code, or that you're not really using a GPU</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 897288,
          "author_name": "Edward Kahara",
          "author_url": "",
          "post_date": "2020-06-22T18:54:11.513000",
          "content": "<p>I was using a GPU and there was nothing wrong with the code. I just hadn't resized the jpeg images beforehand. I'm working with resized jpeg images that are much smaller with the same code and everything is okay now. Thank you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 895041,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-06-21T04:11:47.183000",
      "content": "<p>Try to use smaller images sizes.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 895155,
          "author_name": "Edward Kahara",
          "author_url": "",
          "post_date": "2020-06-21T06:29:45.727000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 894645,
      "author_name": "Martin Kovacevic Buvinic",
      "author_url": "",
      "post_date": "2020-06-20T16:12:36.413000",
      "content": "<p>Hello im using tensorflow tf record format (is faster than loading jpg images because you have the information in binary form). You could use a preprocessed data were you already have your images in numpy arrays and load them to memory (it will wolk with small images (256 should work and not allocate too much memory). This sould definitely make you training faster. Alos keep in mind that a big network with a lot of paramters make the training slow because it have to calculate a lot of gradient witch takes time. Try a simple network and a optimal learning rate for speed.</p>\n\n<p>On the other hand a lot of folks just save their models, and then load them in a new kernel and start straining over that, that's another solution.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 894737,
          "author_name": "Edward Kahara",
          "author_url": "",
          "post_date": "2020-06-20T17:54:55.443000",
          "content": "<p>I don't know Tensorflow, so I'll have to learn it and see. \nI like the idea of saving the model and then loading and training over it. Thank you.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 894621,
      "author_name": "Praphul Jain",
      "author_url": "",
      "post_date": "2020-06-20T15:41:29.420000",
      "content": "<p>I used pytorch and made a very basic model and observed that most of the time is taken in data loading. In my case I observed that loading time is heavily increased if I use pytorch dataloader. I searched the problem and found that many people have found this issue that while loading with data loader time taken is increased by very much but couldnt find a solution. So if someone knows a faster way please tell.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 894732,
          "author_name": "Edward Kahara",
          "author_url": "",
          "post_date": "2020-06-20T17:51:53.147000",
          "content": "<p>Thank you. I will look into using pytorch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 911603,
          "author_name": "M.Talha Arshad",
          "author_url": "",
          "post_date": "2020-07-01T21:28:30.270000",
          "content": "<p>how one can split data into melanoma and benign from JPEG folder?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 911696,
          "author_name": "Signal",
          "author_url": "",
          "post_date": "2020-07-02T00:20:50.637000",
          "content": "<p>Please see my blog post on setting the ideal number of num_workers in your Pytorch dataloader:</p>\n\n<p><a href=\"http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/\">http://www.feeny.org/finding-the-ideal-num_workers-for-pytorch-dataloaders/</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 914314,
          "author_name": "Asad Ali",
          "author_url": "",
          "post_date": "2020-07-03T18:45:59.173000",
          "content": "<p><a href=\"/cdeotte\">@cdeotte</a> posted the Tfrecords resized dataset. Loading the jpegs is a bottleneck for the Dataloader. Instead I loaded the Tfrecords into memory and passed them to the Dataloader. This while not elegant, reduced my epoch time from 25 min on Jpegs to 2 min.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 915355,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-07-04T17:02:53.033000",
      "content": "<p>What size images are you using? I posted 256x256 JPEGS and 256x256 TFRecords <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092\">here</a>. Using these smaller images should speed up your GPU</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 896233,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-22T02:16:53.617000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "896241": "The first thing i would do is time how long your data loader takes. For example (note there is no training, just loading):\n\n    import time\n    start = time.time()\n    for i,batch in enumerate(dataloader):\n        print(i, ', ',end='')\n        if i&gt;=STEPS_PER_EPOCH: break\n    print('Elapsed time minutes = ' (time.time()-start)/60.)\n\nThis should complete very fast. This will determine if your dataloader or training is the bottle neck. If your dataloader is the bottle neck, then speed it up. If your training is the bottle neck then (1) try smaller image size (2) try smaller model (3) freeze the bottom layers\n\nWhen you freeze bottom layers, then your GPU requires less memory and trains faster.\n",
    "894591": "I'm working on this competition with fast.ai. I'm training my model with the images in the jpeg folder.  Even with kaggle's GPU, it takes 5 hours to run one complete epoch. I can't run more than one epoch because it will go past the 9 hour session limit. Is there a way around this? If you're using pytorch or another library, how long does one epoch take? ",
    "896950": "I used to use fast.ai for competitions, but sometimes it is difficult to dive into abstraction layers to speed up the code. If you are familiar with PyTorch, I would recommend switching to pure PyTorch code. If you want a high-level API for PyTorch there exist Catalyst and PyTorch Lightning. Using a TPU and PyTorch, one epoch takes about 5 minutes. \n\nAs pointed out by @cdeotte , Tensorflow with TFRecords are highly competitive as well.",
    "896092": "It seems like a long time, can you pass the code and see it?",
    "895041": "Try to use smaller images sizes.",
    "894645": "Hello im using tensorflow tf record format (is faster than loading jpg images because you have the information in binary form). You could use a preprocessed data were you already have your images in numpy arrays and load them to memory (it will wolk with small images (256 should work and not allocate too much memory). This sould definitely make you training faster. Alos keep in mind that a big network with a lot of paramters make the training slow because it have to calculate a lot of gradient witch takes time. Try a simple network and a optimal learning rate for speed.\n\nOn the other hand a lot of folks just save their models, and then load them in a new kernel and start straining over that, that's another solution.",
    "894621": "I used pytorch and made a very basic model and observed that most of the time is taken in data loading. In my case I observed that loading time is heavily increased if I use pytorch dataloader. I searched the problem and found that many people have found this issue that while loading with data loader time taken is increased by very much but couldnt find a solution. So if someone knows a faster way please tell.",
    "915355": "What size images are you using? I posted 256x256 JPEGS and 256x256 TFRecords [here][1]. Using these smaller images should speed up your GPU\n\n[1]: https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/164092",
    "896233": ""
  }
}