{
  "id": 205318,
  "title": "Any suggestions to speed up training in Pytorch?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/205318",
  "author_name": "",
  "post_date": "2020-12-19T14:18:49.023398500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am using EfficientNet_B4_ns with an image size of 512x512 for only 10 epochs(batch size is 16). I am also using automatic mixed-precision, but it takes 2 hours 12 minutes to complete a single fold training on GPU.<br>\nIs this expected? Or can I make some improvements to enhance the training time?</p>",
  "messages": [
    {
      "id": "1118918",
      "postDate": "12/19/2020 14:18:49",
      "content": "<p>I am using EfficientNet_B4_ns with an image size of 512x512 for only 10 epochs(batch size is 16). I am also using automatic mixed-precision, but it takes 2 hours 12 minutes to complete a single fold training on GPU.<br>\nIs this expected? Or can I make some improvements to enhance the training time?</p>",
      "rawMarkdown": "I am using EfficientNet_B4_ns with an image size of 512x512 for only 10 epochs(batch size is 16). I am also using automatic mixed-precision, but it takes 2 hours 12 minutes to complete a single fold training on GPU.\nIs this expected? Or can I make some improvements to enhance the training time?",
      "votes": null
    },
    {
      "id": "1119115",
      "postDate": "12/19/2020 18:12:30",
      "content": "<p>To some extent, there's not that much you can do about the actual deep learning part, unless there's some particular inefficiency. </p>\n<p>Quite often, the bottle neck is keeping the GPU fed. One possibility is to do all preprocessing including augmentation for your images up-front and to save the results in a fast to read format so that you can read it straight into the neural network (e.g. numpy arrays in hdf5). You end up with epochs*images records, but it would cut out image reading and pre-processing.</p>",
      "rawMarkdown": "To some extent, there's not that much you can do about the actual deep learning part, unless there's some particular inefficiency. \n\nQuite often, the bottle neck is keeping the GPU fed. One possibility is to do all preprocessing including augmentation for your images up-front and to save the results in a fast to read format so that you can read it straight into the neural network (e.g. numpy arrays in hdf5). You end up with epochs*images records, but it would cut out image reading and pre-processing.",
      "votes": null
    },
    {
      "id": "1119133",
      "postDate": "12/19/2020 18:31:56",
      "content": "<p>Thank you!<br>\nIs your model also taking similar training time?</p>",
      "rawMarkdown": "Thank you!\nIs your model also taking similar training time?",
      "votes": null
    },
    {
      "id": "1119162",
      "postDate": "12/19/2020 19:00:39",
      "content": "<p>Yes, with computing augmentations on the fly, it is.</p>",
      "rawMarkdown": "Yes, with computing augmentations on the fly, it is.",
      "votes": null
    },
    {
      "id": "1119618",
      "postDate": "12/20/2020 08:36:54",
      "content": "<p>The ResNet/ResNext models are generally faster with GPU and PyTorch from my personal experience. If you are not spending too much time on augmentation, then try a similar-sized resnet model.</p>",
      "rawMarkdown": "The ResNet/ResNext models are generally faster with GPU and PyTorch from my personal experience. If you are not spending too much time on augmentation, then try a similar-sized resnet model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1119115,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/19/2020 18:12:30",
      "content": "<p>To some extent, there's not that much you can do about the actual deep learning part, unless there's some particular inefficiency. </p>\n<p>Quite often, the bottle neck is keeping the GPU fed. One possibility is to do all preprocessing including augmentation for your images up-front and to save the results in a fast to read format so that you can read it straight into the neural network (e.g. numpy arrays in hdf5). You end up with epochs*images records, but it would cut out image reading and pre-processing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1119133,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "12/19/2020 18:31:56",
          "content": "<p>Thank you!<br>\nIs your model also taking similar training time?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1119162,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/19/2020 19:00:39",
          "content": "<p>Yes, with computing augmentations on the fly, it is.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1119618,
      "author_name": "vishnus",
      "author_url": "",
      "post_date": "12/20/2020 08:36:54",
      "content": "<p>The ResNet/ResNext models are generally faster with GPU and PyTorch from my personal experience. If you are not spending too much time on augmentation, then try a similar-sized resnet model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1118918": "I am using EfficientNet_B4_ns with an image size of 512x512 for only 10 epochs(batch size is 16). I am also using automatic mixed-precision, but it takes 2 hours 12 minutes to complete a single fold training on GPU.\nIs this expected? Or can I make some improvements to enhance the training time?",
    "1119115": "To some extent, there's not that much you can do about the actual deep learning part, unless there's some particular inefficiency. \n\nQuite often, the bottle neck is keeping the GPU fed. One possibility is to do all preprocessing including augmentation for your images up-front and to save the results in a fast to read format so that you can read it straight into the neural network (e.g. numpy arrays in hdf5). You end up with epochs*images records, but it would cut out image reading and pre-processing.",
    "1119133": "Thank you!\nIs your model also taking similar training time?",
    "1119162": "Yes, with computing augmentations on the fly, it is.",
    "1119618": "The ResNet/ResNext models are generally faster with GPU and PyTorch from my personal experience. If you are not spending too much time on augmentation, then try a similar-sized resnet model."
  },
  "source": "meta"
}