{
  "id": 161528,
  "title": "Why training takes too much time with GPU ?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/161528",
  "author_name": "",
  "post_date": "2020-06-25T07:00:50.718029Z",
  "votes": null,
  "comment_count": 9,
  "views": 0,
  "content": "<p>For this dataset why model training with GPU is taking too much time to train ?\nAny specific reason ?</p>",
  "messages": [
    {
      "id": "900955",
      "postDate": "06/25/2020 07:00:50",
      "content": "<p>For this dataset why model training with GPU is taking too much time to train ?\nAny specific reason ?</p>",
      "rawMarkdown": "For this dataset why model training with GPU is taking too much time to train ?\nAny specific reason ?",
      "votes": null
    },
    {
      "id": "901408",
      "postDate": "06/25/2020 13:02:37",
      "content": "<p>It is too abstract of a question tbh, training time can depend of a number of factors:\n1) # of model parameters\n2) Size of the image\n3) Dynamic pre-processing which might be taking place for each batch of image</p>\n\n<p>The remedy to 1) and 2) can potentially lead to drop in results, so it is a factor that you have to make a concious decision on.\nHowever, you can do a few things for 3) for example in my case, I wanted to pre-process the images in a way, but the step was slow and would have taken a lot of time to complete, so I downloaded the data, performed the pre-processing on the images, and then uploaded it on Kaggle.</p>\n\n<p>On Efficient Net (B5) Time to train with external data ~15 minutes (with validation) / epoch.\nI think for this much amount of data, it is fair enough. However, you can even reduce it using some techniques like mix precision training etc. </p>",
      "rawMarkdown": "It is too abstract of a question tbh, training time can depend of a number of factors:\n1) # of model parameters\n2) Size of the image\n3) Dynamic pre-processing which might be taking place for each batch of image\n\nThe remedy to 1) and 2) can potentially lead to drop in results, so it is a factor that you have to make a concious decision on.\nHowever, you can do a few things for 3) for example in my case, I wanted to pre-process the images in a way, but the step was slow and would have taken a lot of time to complete, so I downloaded the data, performed the pre-processing on the images, and then uploaded it on Kaggle.\n\nOn Efficient Net (B5) Time to train with external data ~15 minutes (with validation) / epoch.\nI think for this much amount of data, it is fair enough. However, you can even reduce it using some techniques like mix precision training etc.",
      "votes": null
    },
    {
      "id": "902812",
      "postDate": "06/26/2020 11:29:27",
      "content": "<p>third option is looking nice. Thanks for the idea</p>",
      "rawMarkdown": "third option is looking nice. Thanks for the idea",
      "votes": null
    },
    {
      "id": "904451",
      "postDate": "06/27/2020 16:27:10",
      "content": "<p>I have the same problem, im using EF_B1 with 240x240 image-size and an epoch takes about 1:30 hour eventhough im training on GPU and I'm sending all of my necessary tensors to the GPU. </p>\n\n<p>Do you have any idea how to fix that? </p>",
      "rawMarkdown": "I have the same problem, im using EF_B1 with 240x240 image-size and an epoch takes about 1:30 hour eventhough im training on GPU and I'm sending all of my necessary tensors to the GPU. \n\nDo you have any idea how to fix that?",
      "votes": null
    },
    {
      "id": "904470",
      "postDate": "06/27/2020 16:42:31",
      "content": "<p>Think you for sure have a problem - 10 minutes per epoch would make sense.  Doing kernels on my home PC so I can use nvidia to see that my GPU's are loaded and using lots of watts.  I don't use Kaggle often enough but my first quess would be that the GPU not functional and your time represents CPU.  Should have one of my machines available soon - will turn off the GPU and see what epoch time I get.</p>",
      "rawMarkdown": "Think you for sure have a problem - 10 minutes per epoch would make sense.  Doing kernels on my home PC so I can use nvidia to see that my GPU's are loaded and using lots of watts.  I don't use Kaggle often enough but my first quess would be that the GPU not functional and your time represents CPU.  Should have one of my machines available soon - will turn off the GPU and see what epoch time I get.",
      "votes": null
    },
    {
      "id": "904504",
      "postDate": "06/27/2020 17:12:44",
      "content": "<p>Yes, i think so too. I'm using the Kaggle GPU, is it possible that i share my notebook with you? Maybe you can see it instantly?</p>\n\n<p>*<em>EDIT:</em>*I shared my notebook with you, I have that problem on other competitions too but I just cant find the mistake. I also checked other notebooks but I dont find the difference.</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Yes, i think so too. I'm using the Kaggle GPU, is it possible that i share my notebook with you? Maybe you can see it instantly?\n\n**EDIT:**I shared my notebook with you, I have that problem on other competitions too but I just cant find the mistake. I also checked other notebooks but I dont find the difference.\n\nThank you",
      "votes": null
    },
    {
      "id": "904571",
      "postDate": "06/27/2020 18:13:49",
      "content": "<p>OK - got it.  Started it running just as you shared but need to go make lunch.  So first look will occur in one hour.</p>",
      "rawMarkdown": "OK - got it.  Started it running just as you shared but need to go make lunch.  So first look will occur in one hour.",
      "votes": null
    },
    {
      "id": "904574",
      "postDate": "06/27/2020 18:15:42",
      "content": "<p>Alright, thank you very much!</p>",
      "rawMarkdown": "Alright, thank you very much!",
      "votes": null
    },
    {
      "id": "904649",
      "postDate": "06/27/2020 19:43:45",
      "content": "<p>I am not a torch user - so some of what your doing not making obvious sense.  The for loop for the images is one of those:)</p>\n\n<p>The GPU is visible to torch.  The GPU memory is getting loaded - but very slowly.  After 1 hours seeing no GPU % use.  I got the batch size up to 512 which loads GPU memory to 13GB but that did not increase speed.</p>\n\n<p>You need a torch user to help - sorry!</p>",
      "rawMarkdown": "I am not a torch user - so some of what your doing not making obvious sense.  The for loop for the images is one of those:)\n\nThe GPU is visible to torch.  The GPU memory is getting loaded - but very slowly.  After 1 hours seeing no GPU % use.  I got the batch size up to 512 which loads GPU memory to 13GB but that did not increase speed.\n\nYou need a torch user to help - sorry!",
      "votes": null
    },
    {
      "id": "904690",
      "postDate": "06/27/2020 20:15:00",
      "content": "<p>Thank you very much though! I will create a topic on that</p>",
      "rawMarkdown": "Thank you very much though! I will create a topic on that",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 901408,
      "author_name": "saran95",
      "author_url": "",
      "post_date": "06/25/2020 13:02:37",
      "content": "<p>It is too abstract of a question tbh, training time can depend of a number of factors:\n1) # of model parameters\n2) Size of the image\n3) Dynamic pre-processing which might be taking place for each batch of image</p>\n\n<p>The remedy to 1) and 2) can potentially lead to drop in results, so it is a factor that you have to make a concious decision on.\nHowever, you can do a few things for 3) for example in my case, I wanted to pre-process the images in a way, but the step was slow and would have taken a lot of time to complete, so I downloaded the data, performed the pre-processing on the images, and then uploaded it on Kaggle.</p>\n\n<p>On Efficient Net (B5) Time to train with external data ~15 minutes (with validation) / epoch.\nI think for this much amount of data, it is fair enough. However, you can even reduce it using some techniques like mix precision training etc. </p>",
      "votes": null,
      "replies": [
        {
          "id": 902812,
          "author_name": "sid321axn",
          "author_url": "",
          "post_date": "06/26/2020 11:29:27",
          "content": "<p>third option is looking nice. Thanks for the idea</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904451,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "06/27/2020 16:27:10",
          "content": "<p>I have the same problem, im using EF_B1 with 240x240 image-size and an epoch takes about 1:30 hour eventhough im training on GPU and I'm sending all of my necessary tensors to the GPU. </p>\n\n<p>Do you have any idea how to fix that? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904470,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/27/2020 16:42:31",
          "content": "<p>Think you for sure have a problem - 10 minutes per epoch would make sense.  Doing kernels on my home PC so I can use nvidia to see that my GPU's are loaded and using lots of watts.  I don't use Kaggle often enough but my first quess would be that the GPU not functional and your time represents CPU.  Should have one of my machines available soon - will turn off the GPU and see what epoch time I get.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904504,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "06/27/2020 17:12:44",
          "content": "<p>Yes, i think so too. I'm using the Kaggle GPU, is it possible that i share my notebook with you? Maybe you can see it instantly?</p>\n\n<p>*<em>EDIT:</em>*I shared my notebook with you, I have that problem on other competitions too but I just cant find the mistake. I also checked other notebooks but I dont find the difference.</p>\n\n<p>Thank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904571,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/27/2020 18:13:49",
          "content": "<p>OK - got it.  Started it running just as you shared but need to go make lunch.  So first look will occur in one hour.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904574,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "06/27/2020 18:15:42",
          "content": "<p>Alright, thank you very much!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904649,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "06/27/2020 19:43:45",
          "content": "<p>I am not a torch user - so some of what your doing not making obvious sense.  The for loop for the images is one of those:)</p>\n\n<p>The GPU is visible to torch.  The GPU memory is getting loaded - but very slowly.  After 1 hours seeing no GPU % use.  I got the batch size up to 512 which loads GPU memory to 13GB but that did not increase speed.</p>\n\n<p>You need a torch user to help - sorry!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 904690,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "06/27/2020 20:15:00",
          "content": "<p>Thank you very much though! I will create a topic on that</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "900955": "For this dataset why model training with GPU is taking too much time to train ?\nAny specific reason ?",
    "901408": "It is too abstract of a question tbh, training time can depend of a number of factors:\n1) # of model parameters\n2) Size of the image\n3) Dynamic pre-processing which might be taking place for each batch of image\n\nThe remedy to 1) and 2) can potentially lead to drop in results, so it is a factor that you have to make a concious decision on.\nHowever, you can do a few things for 3) for example in my case, I wanted to pre-process the images in a way, but the step was slow and would have taken a lot of time to complete, so I downloaded the data, performed the pre-processing on the images, and then uploaded it on Kaggle.\n\nOn Efficient Net (B5) Time to train with external data ~15 minutes (with validation) / epoch.\nI think for this much amount of data, it is fair enough. However, you can even reduce it using some techniques like mix precision training etc.",
    "902812": "third option is looking nice. Thanks for the idea",
    "904451": "I have the same problem, im using EF_B1 with 240x240 image-size and an epoch takes about 1:30 hour eventhough im training on GPU and I'm sending all of my necessary tensors to the GPU. \n\nDo you have any idea how to fix that?",
    "904470": "Think you for sure have a problem - 10 minutes per epoch would make sense.  Doing kernels on my home PC so I can use nvidia to see that my GPU's are loaded and using lots of watts.  I don't use Kaggle often enough but my first quess would be that the GPU not functional and your time represents CPU.  Should have one of my machines available soon - will turn off the GPU and see what epoch time I get.",
    "904504": "Yes, i think so too. I'm using the Kaggle GPU, is it possible that i share my notebook with you? Maybe you can see it instantly?\n\n**EDIT:**I shared my notebook with you, I have that problem on other competitions too but I just cant find the mistake. I also checked other notebooks but I dont find the difference.\n\nThank you",
    "904571": "OK - got it.  Started it running just as you shared but need to go make lunch.  So first look will occur in one hour.",
    "904574": "Alright, thank you very much!",
    "904649": "I am not a torch user - so some of what your doing not making obvious sense.  The for loop for the images is one of those:)\n\nThe GPU is visible to torch.  The GPU memory is getting loaded - but very slowly.  After 1 hours seeing no GPU % use.  I got the batch size up to 512 which loads GPU memory to 13GB but that did not increase speed.\n\nYou need a torch user to help - sorry!",
    "904690": "Thank you very much though! I will create a topic on that"
  },
  "source": "meta"
}