{
  "id": 201515,
  "title": "What machine are you using and how long per epoch?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/201515",
  "author_name": "",
  "post_date": "2020-12-05T12:04:38.189540800Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi guys, we all know that image models are trained faster with GPUs. At the same time not everyone can afford GPUs if not for the latest ones.<br>\nJust want to understand what machine everyone else is using to train their models, even if it's Google Collab or Kernel K80 GPUs. </p>\n<p>I'm using a 64gb RAM machine with no GPU to build my pipeline and apparently move it to Kaggle kernel. </p>\n<p>I got some credits left with the Azure machine containing 12 vcpus, 112 GiB memory 2X K80s. </p>\n<p>With no data augmentation and training on 80% of the images resized to 256*256 with train batch_size of 8 efficientNet_b7 is taking me 16mins per epoch</p>\n<p>Curious to know what machine you guys are training on</p>",
  "messages": [
    {
      "id": "1102860",
      "postDate": "12/05/2020 12:04:38",
      "content": "<p>Hi guys, we all know that image models are trained faster with GPUs. At the same time not everyone can afford GPUs if not for the latest ones.<br>\nJust want to understand what machine everyone else is using to train their models, even if it's Google Collab or Kernel K80 GPUs. </p>\n<p>I'm using a 64gb RAM machine with no GPU to build my pipeline and apparently move it to Kaggle kernel. </p>\n<p>I got some credits left with the Azure machine containing 12 vcpus, 112 GiB memory 2X K80s. </p>\n<p>With no data augmentation and training on 80% of the images resized to 256*256 with train batch_size of 8 efficientNet_b7 is taking me 16mins per epoch</p>\n<p>Curious to know what machine you guys are training on</p>",
      "rawMarkdown": "Hi guys, we all know that image models are trained faster with GPUs. At the same time not everyone can afford GPUs if not for the latest ones.\nJust want to understand what machine everyone else is using to train their models, even if it's Google Collab or Kernel K80 GPUs. \n\nI'm using a 64gb RAM machine with no GPU to build my pipeline and apparently move it to Kaggle kernel. \n\nI got some credits left with the Azure machine containing 12 vcpus, 112 GiB memory 2X K80s. \n\nWith no data augmentation and training on 80% of the images resized to 256*256 with train batch_size of 8 efficientNet_b7 is taking me 16mins per epoch\n\nCurious to know what machine you guys are training on",
      "votes": null
    },
    {
      "id": "1103229",
      "postDate": "12/05/2020 18:35:53",
      "content": "<p>Yeah, I am also facing the same problem in colab. It is taking around 16min per epoch🤕</p>",
      "rawMarkdown": "Yeah, I am also facing the same problem in colab. It is taking around 16min per epoch🤕",
      "votes": null
    },
    {
      "id": "1104563",
      "postDate": "12/07/2020 04:18:14",
      "content": "<p>I'm training on Kaggle Kernels (256*256, hell lot of Augmentations) and on EffNet-b7 with a batch size of 16 (because 32 gives an OOM Error). It takes about 14 minutes for a single epoch. <br>\nI think you should give Kaggle and Colab TPUs a try. It's hard to configure if you are using Torch but it's worth the time. </p>",
      "rawMarkdown": "I'm training on Kaggle Kernels (256*256, hell lot of Augmentations) and on EffNet-b7 with a batch size of 16 (because 32 gives an OOM Error). It takes about 14 minutes for a single epoch. \nI think you should give Kaggle and Colab TPUs a try. It's hard to configure if you are using Torch but it's worth the time.",
      "votes": null
    },
    {
      "id": "1104792",
      "postDate": "12/07/2020 08:18:06",
      "content": "<p>That's good to hear and thanks for the suggestion. I will definitely give it a atry</p>",
      "rawMarkdown": "That's good to hear and thanks for the suggestion. I will definitely give it a atry",
      "votes": null
    },
    {
      "id": "1105853",
      "postDate": "12/08/2020 09:00:55",
      "content": "<p>I am using Colab pro.</p>",
      "rawMarkdown": "I am using Colab pro.",
      "votes": null
    },
    {
      "id": "1105983",
      "postDate": "12/08/2020 11:43:23",
      "content": "<p>That's good. What GPUs you get consistently? P100 or T4? and how long does it taker per epoch?</p>",
      "rawMarkdown": "That's good. What GPUs you get consistently? P100 or T4? and how long does it taker per epoch?",
      "votes": null
    },
    {
      "id": "1106193",
      "postDate": "12/08/2020 15:26:34",
      "content": "<p>it's almost always V100. ~13 minutes per epoch. I don't use big NN architectures.</p>",
      "rawMarkdown": "it's almost always V100. ~13 minutes per epoch. I don't use big NN architectures.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1103229,
      "author_name": "shanmukh05",
      "author_url": "",
      "post_date": "12/05/2020 18:35:53",
      "content": "<p>Yeah, I am also facing the same problem in colab. It is taking around 16min per epoch🤕</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1104563,
      "author_name": "heyytanay",
      "author_url": "",
      "post_date": "12/07/2020 04:18:14",
      "content": "<p>I'm training on Kaggle Kernels (256*256, hell lot of Augmentations) and on EffNet-b7 with a batch size of 16 (because 32 gives an OOM Error). It takes about 14 minutes for a single epoch. <br>\nI think you should give Kaggle and Colab TPUs a try. It's hard to configure if you are using Torch but it's worth the time. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1104792,
          "author_name": "thanish",
          "author_url": "",
          "post_date": "12/07/2020 08:18:06",
          "content": "<p>That's good to hear and thanks for the suggestion. I will definitely give it a atry</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1105853,
      "author_name": "nroman",
      "author_url": "",
      "post_date": "12/08/2020 09:00:55",
      "content": "<p>I am using Colab pro.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1105983,
          "author_name": "thanish",
          "author_url": "",
          "post_date": "12/08/2020 11:43:23",
          "content": "<p>That's good. What GPUs you get consistently? P100 or T4? and how long does it taker per epoch?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1106193,
          "author_name": "nroman",
          "author_url": "",
          "post_date": "12/08/2020 15:26:34",
          "content": "<p>it's almost always V100. ~13 minutes per epoch. I don't use big NN architectures.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1102860": "Hi guys, we all know that image models are trained faster with GPUs. At the same time not everyone can afford GPUs if not for the latest ones.\nJust want to understand what machine everyone else is using to train their models, even if it's Google Collab or Kernel K80 GPUs. \n\nI'm using a 64gb RAM machine with no GPU to build my pipeline and apparently move it to Kaggle kernel. \n\nI got some credits left with the Azure machine containing 12 vcpus, 112 GiB memory 2X K80s. \n\nWith no data augmentation and training on 80% of the images resized to 256*256 with train batch_size of 8 efficientNet_b7 is taking me 16mins per epoch\n\nCurious to know what machine you guys are training on",
    "1103229": "Yeah, I am also facing the same problem in colab. It is taking around 16min per epoch🤕",
    "1104563": "I'm training on Kaggle Kernels (256*256, hell lot of Augmentations) and on EffNet-b7 with a batch size of 16 (because 32 gives an OOM Error). It takes about 14 minutes for a single epoch. \nI think you should give Kaggle and Colab TPUs a try. It's hard to configure if you are using Torch but it's worth the time.",
    "1104792": "That's good to hear and thanks for the suggestion. I will definitely give it a atry",
    "1105853": "I am using Colab pro.",
    "1105983": "That's good. What GPUs you get consistently? P100 or T4? and how long does it taker per epoch?",
    "1106193": "it's almost always V100. ~13 minutes per epoch. I don't use big NN architectures."
  },
  "source": "meta"
}