{
  "id": 233152,
  "title": "Training time for one model?",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/233152",
  "author_name": "",
  "post_date": "2021-04-17T15:53:24.562728500Z",
  "votes": 10,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I am curious how long it takes to train one single model for other persons? </p>\n<p>One single model trained on a single 3090 takes <strong>around 20 minutes per epoch and per fold</strong>. </p>\n<p>It is a <strong>very simple model and no augmentation</strong> with <strong>1024 x 1024 image tiles</strong> and a **batch of 4 **(I guess I can go a little higher).</p>\n<p>Any feedback is more than welcome!  </p>",
  "messages": [
    {
      "id": "1276485",
      "postDate": "04/17/2021 15:53:24",
      "content": "<p>I am curious how long it takes to train one single model for other persons? </p>\n<p>One single model trained on a single 3090 takes <strong>around 20 minutes per epoch and per fold</strong>. </p>\n<p>It is a <strong>very simple model and no augmentation</strong> with <strong>1024 x 1024 image tiles</strong> and a **batch of 4 **(I guess I can go a little higher).</p>\n<p>Any feedback is more than welcome!  </p>",
      "rawMarkdown": "I am curious how long it takes to train one single model for other persons? \n\nOne single model trained on a single 3090 takes **around 20 minutes per epoch and per fold**. \n\nIt is a **very simple model and no augmentation** with **1024 x 1024 image tiles** and a **batch of 4 **(I guess I can go a little higher).\n\nAny feedback is more than welcome!",
      "votes": null
    },
    {
      "id": "1276525",
      "postDate": "04/17/2021 16:52:41",
      "content": "<p>So far what i am doing is training with 256*256 with augs in model like efficient net it usually takes 2-5 mins per epoch , model like resnest and resnet50 takes 3 mins with the batch size of 32-45 , gpu used :- v100 and even p100 takes 7 to 10 mins </p>",
      "rawMarkdown": "So far what i am doing is training with 256*256 with augs in model like efficient net it usually takes 2-5 mins per epoch , model like resnest and resnet50 takes 3 mins with the batch size of 32-45 , gpu used :- v100 and even p100 takes 7 to 10 mins",
      "votes": null
    },
    {
      "id": "1276529",
      "postDate": "04/17/2021 16:57:42",
      "content": "<p>Thanks for sharing these details!</p>\n<p>Have you tried other models and/or bigger image tiles?</p>",
      "rawMarkdown": "Thanks for sharing these details!\n\nHave you tried other models and/or bigger image tiles?",
      "votes": null
    },
    {
      "id": "1276574",
      "postDate": "04/17/2021 17:48:47",
      "content": "<p>yeah i tried SOTA models from <a href=\"https://paperswithcode.com/task/medical-image-segmentation\" target=\"_blank\">https://paperswithcode.com/task/medical-image-segmentation</a> , But smaller models are performing better then larger models from my experimentations , didnt tried tiles bigger then 256</p>",
      "rawMarkdown": "yeah i tried SOTA models from https://paperswithcode.com/task/medical-image-segmentation , But smaller models are performing better then larger models from my experimentations , didnt tried tiles bigger then 256",
      "votes": null
    },
    {
      "id": "1276633",
      "postDate": "04/17/2021 18:52:40",
      "content": "<p>That makes sense: smaller models work on smaller tiles. If you try bigger models, you should try bigger tiles. </p>",
      "rawMarkdown": "That makes sense: smaller models work on smaller tiles. If you try bigger models, you should try bigger tiles.",
      "votes": null
    },
    {
      "id": "1276638",
      "postDate": "04/17/2021 19:11:40",
      "content": "<p>yeah next step is to try  512 and 1024 most probably</p>",
      "rawMarkdown": "yeah next step is to try  512 and 1024 most probably",
      "votes": null
    },
    {
      "id": "1276642",
      "postDate": "04/17/2021 19:16:08",
      "content": "<p>Awesome, let us know how it goes. :)</p>",
      "rawMarkdown": "Awesome, let us know how it goes. :)",
      "votes": null
    },
    {
      "id": "1276767",
      "postDate": "04/17/2021 23:45:02",
      "content": "<p>4 min per epoch with rtx 3090 and 512x512 resolution</p>",
      "rawMarkdown": "4 min per epoch with rtx 3090 and 512x512 resolution",
      "votes": null
    },
    {
      "id": "1277190",
      "postDate": "04/18/2021 13:58:28",
      "content": "<p>Mine is taking around 4-5 minutes per epoch for efficientnet b4 on 512x512 with augs on kaggle gpus</p>",
      "rawMarkdown": "Mine is taking around 4-5 minutes per epoch for efficientnet b4 on 512x512 with augs on kaggle gpus",
      "votes": null
    },
    {
      "id": "1277299",
      "postDate": "04/18/2021 15:45:26",
      "content": "<p>Roughly 6 minutes per epoch for EffNet B5 with 512 pix tiles and medium Augmentation. Usually do max 20 epochs per fold on Kaggle TPU's </p>\n<p>With the 9 hours of max session time you can't do a full cross validation run….so implemented fold stopping. Usually I only run the first 2 or 3 folds…if a modification turns out to be a succesfull one then I perform another run with 3 folds and a different seed. Works for me sofar.</p>",
      "rawMarkdown": "Roughly 6 minutes per epoch for EffNet B5 with 512 pix tiles and medium Augmentation. Usually do max 20 epochs per fold on Kaggle TPU's \n\nWith the 9 hours of max session time you can't do a full cross validation run....so implemented fold stopping. Usually I only run the first 2 or 3 folds...if a modification turns out to be a succesfull one then I perform another run with 3 folds and a different seed. Works for me sofar.",
      "votes": null
    },
    {
      "id": "1277455",
      "postDate": "04/18/2021 18:34:46",
      "content": "<p>Mine is around 5 minutes per epoch for an EffNetb0. 256x256 with heavy augmentations.</p>",
      "rawMarkdown": "Mine is around 5 minutes per epoch for an EffNetb0. 256x256 with heavy augmentations.",
      "votes": null
    },
    {
      "id": "1278763",
      "postDate": "04/20/2021 08:57:32",
      "content": "<p>For me, resnext101, 512x512. batch_size 24, around 556 sec/epoch.<br>\nRTX6000x2.</p>",
      "rawMarkdown": "For me, resnext101, 512x512. batch_size 24, around 556 sec/epoch.\nRTX6000x2.",
      "votes": null
    },
    {
      "id": "1279268",
      "postDate": "04/20/2021 18:44:49",
      "content": "<p>I'm not even able to train a batch of 1 with  1024, what kind of hardware are you running?</p>",
      "rawMarkdown": "I'm not even able to train a batch of 1 with  1024, what kind of hardware are you running?",
      "votes": null
    },
    {
      "id": "1279276",
      "postDate": "04/20/2021 18:55:47",
      "content": "<p>Btw i forgot to answer your question, i'm using 256px with some augumentations, my epoch time is 130s, depending on my folds / augs / external data / learning rate it would go around 90s to 220s. It will also take 20-60 epochs to convege with patence=20 (i never changed the patience, i wonder what it would do to the score)</p>",
      "rawMarkdown": "Btw i forgot to answer your question, i'm using 256px with some augumentations, my epoch time is 130s, depending on my folds / augs / external data / learning rate it would go around 90s to 220s. It will also take 20-60 epochs to convege with patence=20 (i never changed the patience, i wonder what it would do to the score)",
      "votes": null
    },
    {
      "id": "1279285",
      "postDate": "04/20/2021 19:05:52",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a> Are you using TPU's? If not … you might want to try that out :-) An EffNet B5 with a Unet/FPN/Linknet can train with batch of 2 to 3 per node ( so times 8)</p>\n<p>You could try to switch the size to 512px.</p>\n<p>If you have implemented some Early Stopping then changing the patience won't effect the score…it will however save you GPU/TPU time.</p>",
      "rawMarkdown": "Hi @victorasso Are you using TPU's? If not ... you might want to try that out :-) An EffNet B5 with a Unet/FPN/Linknet can train with batch of 2 to 3 per node ( so times 8)\n\nYou could try to switch the size to 512px.\n\nIf you have implemented some Early Stopping then changing the patience won't effect the score...it will however save you GPU/TPU time.",
      "votes": null
    },
    {
      "id": "1279301",
      "postDate": "04/20/2021 19:13:17",
      "content": "<p>Is anyone able to use mixed-precision training using tf.keras ? I tried but the loss becomes NaN. It should speed up training on the latest hardwares if it is working properly.</p>\n<p>According to the <a href=\"https://www.tensorflow.org/guide/mixed_precisionl\" target=\"_blank\">TensorFlow blog</a>, I added only these three lines of code in the beginning.</p>\n<pre><code>from tensorflow.keras import mixed_precision\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_global_policy(policy)\n</code></pre>",
      "rawMarkdown": "Is anyone able to use mixed-precision training using tf.keras ? I tried but the loss becomes NaN. It should speed up training on the latest hardwares if it is working properly.\n\nAccording to the [TensorFlow blog](https://www.tensorflow.org/guide/mixed_precisionl), I added only these three lines of code in the beginning.\n```\nfrom tensorflow.keras import mixed_precision\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_global_policy(policy)\n```",
      "votes": null
    },
    {
      "id": "1279310",
      "postDate": "04/20/2021 19:17:40",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hetulvpatel\" target=\"_blank\">@hetulvpatel</a> Are you running on GPU? The nan loss is very likely caused by the last Dense layer not being a dtype of float32. This happens a lot. I believe that is somewhere mentioned in the TF documentation.</p>\n<p>Could also be the loss function.</p>",
      "rawMarkdown": "Hi @hetulvpatel Are you running on GPU? The nan loss is very likely caused by the last Dense layer not being a dtype of float32. This happens a lot. I believe that is somewhere mentioned in the TF documentation.\n\nCould also be the loss function.",
      "votes": null
    },
    {
      "id": "1279323",
      "postDate": "04/20/2021 19:31:50",
      "content": "<p>Not really, i'm having a hard time trying to change my code to TPUs for kaggle envs since i used my own machine to train my initial models, i even tried to rent an azure VM to do the job but after all that work all i got was a crippled env and a 50 dollar bill for a day of cringe.</p>\n<p>Do you have your own hardware or do you only use the kaggle env? If so, any tips for saving those precious hardware accelerated hours?</p>",
      "rawMarkdown": "Not really, i'm having a hard time trying to change my code to TPUs for kaggle envs since i used my own machine to train my initial models, i even tried to rent an azure VM to do the job but after all that work all i got was a crippled env and a 50 dollar bill for a day of cringe.\n\nDo you have your own hardware or do you only use the kaggle env? If so, any tips for saving those precious hardware accelerated hours?",
      "votes": null
    },
    {
      "id": "1279338",
      "postDate": "04/20/2021 19:39:38",
      "content": "<ul>\n<li>Yes, I train it on Ampere GPU. </li>\n<li>When I remove these three lines, the training goes perfectly without loss being NaN. </li>\n<li>Also, I checked my Dense layer it is float32. </li>\n<li>I tried all different losses e.g BCE, Dice, Combo.. But in every case loss becomes nan.</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Are you able to use Mixed Precision using tf.keras in your training?</p>",
      "rawMarkdown": "Yes, I train it on Ampere GPU. \n- When I remove these three lines, the training goes perfectly without loss being NaN. \n- Also, I checked my Dense layer it is float32. \n- I tried all different losses e.g BCE, Dice, Combo.. But in every case loss becomes nan.\n\n@rsmits Are you able to use Mixed Precision using tf.keras in your training?",
      "votes": null
    },
    {
      "id": "1279348",
      "postDate": "04/20/2021 19:45:27",
      "content": "<p><a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a> I do have a Data Science PC … but with an old 1070 Ti in it I haven't used it in the last year for any serious competition work…just for my hobby projects.</p>\n<p>For serious modelling and Kaggle I use Google Colab Pro to try out stuff and when I think I have an improvement then I move it over to Kaggle TPU's. </p>\n<p>For modifying your own code to get it up and running from your local GPU up to the Kaggle TPU's I would suggest the following 'research' path: Investigate and Create TFrecords from images, Switch your local code to use TF Datasets (if your not already doing so). Then take a look at some baseline and simple TPU code. In the Kaggle TPU Flowers competition there are a lot of notebooks explaning the concepts and showing how to do classification with TPU. That should get you started.</p>",
      "rawMarkdown": "victorasso I do have a Data Science PC ... but with an old 1070 Ti in it I haven't used it in the last year for any serious competition work...just for my hobby projects.\n\nFor serious modelling and Kaggle I use Google Colab Pro to try out stuff and when I think I have an improvement then I move it over to Kaggle TPU's. \n\nFor modifying your own code to get it up and running from your local GPU up to the Kaggle TPU's I would suggest the following 'research' path: Investigate and Create TFrecords from images, Switch your local code to use TF Datasets (if your not already doing so). Then take a look at some baseline and simple TPU code. In the Kaggle TPU Flowers competition there are a lot of notebooks explaning the concepts and showing how to do classification with TPU. That should get you started.",
      "votes": null
    },
    {
      "id": "1279387",
      "postDate": "04/20/2021 20:37:42",
      "content": "<p>I was stalling to really put some hours on learning that, it was already overwhelming the amount of time i had to allocate to get to know the tf basics, i guess i can't avoid that any longer.</p>\n<p>Thank you for your comments, so after all if i want to increase the image resolution on my training pipelines i will have to learn how to work with TPUs after all.</p>",
      "rawMarkdown": "I was stalling to really put some hours on learning that, it was already overwhelming the amount of time i had to allocate to get to know the tf basics, i guess i can't avoid that any longer.\n\nThank you for your comments, so after all if i want to increase the image resolution on my training pipelines i will have to learn how to work with TPUs after all.",
      "votes": null
    },
    {
      "id": "1279392",
      "postDate": "04/20/2021 20:54:47",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hetulvpatel\" target=\"_blank\">@hetulvpatel</a> Hmm in that case I'am not sure what could be a further reason you get nan loss.</p>\n<p>For training I only use TPU's. I don't specify the use of bloat16 since the TPU's use them internally anyway. </p>\n<p>Last year during the Kaggle Google Landmark competition I read some interresting discussion on whether or not it would benefit a TPU enabled code to specify bfloat16 explicitly. I tried it for a few models and didn't notice any performance difference.</p>",
      "rawMarkdown": "Hi @hetulvpatel Hmm in that case I'am not sure what could be a further reason you get nan loss.\n\nFor training I only use TPU's. I don't specify the use of bloat16 since the TPU's use them internally anyway. \n\nLast year during the Kaggle Google Landmark competition I read some interresting discussion on whether or not it would benefit a TPU enabled code to specify bfloat16 explicitly. I tried it for a few models and didn't notice any performance difference.",
      "votes": null
    },
    {
      "id": "1281785",
      "postDate": "04/23/2021 09:54:33",
      "content": "<p>18 mins on V100. batch size: 16</p>",
      "rawMarkdown": "18 mins on V100. batch size: 16",
      "votes": null
    },
    {
      "id": "1281902",
      "postDate": "04/23/2021 12:18:59",
      "content": "<p>Sweet! I guess on colab or GCP maybe?</p>",
      "rawMarkdown": "Sweet! I guess on colab or GCP maybe?",
      "votes": null
    },
    {
      "id": "1281904",
      "postDate": "04/23/2021 12:20:09",
      "content": "<p>You need GPU RAM big enough to hold the model at this size. Have you tried smaller tiles (like 512)?</p>",
      "rawMarkdown": "You need GPU RAM big enough to hold the model at this size. Have you tried smaller tiles (like 512)?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1276525,
      "author_name": "trooperog",
      "author_url": "",
      "post_date": "04/17/2021 16:52:41",
      "content": "<p>So far what i am doing is training with 256*256 with augs in model like efficient net it usually takes 2-5 mins per epoch , model like resnest and resnet50 takes 3 mins with the batch size of 32-45 , gpu used :- v100 and even p100 takes 7 to 10 mins </p>",
      "votes": null,
      "replies": [
        {
          "id": 1276529,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/17/2021 16:57:42",
          "content": "<p>Thanks for sharing these details!</p>\n<p>Have you tried other models and/or bigger image tiles?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1276574,
          "author_name": "trooperog",
          "author_url": "",
          "post_date": "04/17/2021 17:48:47",
          "content": "<p>yeah i tried SOTA models from <a href=\"https://paperswithcode.com/task/medical-image-segmentation\" target=\"_blank\">https://paperswithcode.com/task/medical-image-segmentation</a> , But smaller models are performing better then larger models from my experimentations , didnt tried tiles bigger then 256</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1276633,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/17/2021 18:52:40",
          "content": "<p>That makes sense: smaller models work on smaller tiles. If you try bigger models, you should try bigger tiles. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1276638,
          "author_name": "trooperog",
          "author_url": "",
          "post_date": "04/17/2021 19:11:40",
          "content": "<p>yeah next step is to try  512 and 1024 most probably</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1276642,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/17/2021 19:16:08",
          "content": "<p>Awesome, let us know how it goes. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1276767,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "04/17/2021 23:45:02",
      "content": "<p>4 min per epoch with rtx 3090 and 512x512 resolution</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1277190,
      "author_name": "deveshdarshan",
      "author_url": "",
      "post_date": "04/18/2021 13:58:28",
      "content": "<p>Mine is taking around 4-5 minutes per epoch for efficientnet b4 on 512x512 with augs on kaggle gpus</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1277299,
      "author_name": "rsmits",
      "author_url": "",
      "post_date": "04/18/2021 15:45:26",
      "content": "<p>Roughly 6 minutes per epoch for EffNet B5 with 512 pix tiles and medium Augmentation. Usually do max 20 epochs per fold on Kaggle TPU's </p>\n<p>With the 9 hours of max session time you can't do a full cross validation run….so implemented fold stopping. Usually I only run the first 2 or 3 folds…if a modification turns out to be a succesfull one then I perform another run with 3 folds and a different seed. Works for me sofar.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1277455,
      "author_name": "andrewshao05",
      "author_url": "",
      "post_date": "04/18/2021 18:34:46",
      "content": "<p>Mine is around 5 minutes per epoch for an EffNetb0. 256x256 with heavy augmentations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1278763,
      "author_name": "hihunjin",
      "author_url": "",
      "post_date": "04/20/2021 08:57:32",
      "content": "<p>For me, resnext101, 512x512. batch_size 24, around 556 sec/epoch.<br>\nRTX6000x2.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1279268,
      "author_name": "victorasso",
      "author_url": "",
      "post_date": "04/20/2021 18:44:49",
      "content": "<p>I'm not even able to train a batch of 1 with  1024, what kind of hardware are you running?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1279276,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "04/20/2021 18:55:47",
          "content": "<p>Btw i forgot to answer your question, i'm using 256px with some augumentations, my epoch time is 130s, depending on my folds / augs / external data / learning rate it would go around 90s to 220s. It will also take 20-60 epochs to convege with patence=20 (i never changed the patience, i wonder what it would do to the score)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279285,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/20/2021 19:05:52",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a> Are you using TPU's? If not … you might want to try that out :-) An EffNet B5 with a Unet/FPN/Linknet can train with batch of 2 to 3 per node ( so times 8)</p>\n<p>You could try to switch the size to 512px.</p>\n<p>If you have implemented some Early Stopping then changing the patience won't effect the score…it will however save you GPU/TPU time.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279323,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "04/20/2021 19:31:50",
          "content": "<p>Not really, i'm having a hard time trying to change my code to TPUs for kaggle envs since i used my own machine to train my initial models, i even tried to rent an azure VM to do the job but after all that work all i got was a crippled env and a 50 dollar bill for a day of cringe.</p>\n<p>Do you have your own hardware or do you only use the kaggle env? If so, any tips for saving those precious hardware accelerated hours?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279348,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/20/2021 19:45:27",
          "content": "<p><a href=\"https://www.kaggle.com/victorasso\" target=\"_blank\">@victorasso</a> I do have a Data Science PC … but with an old 1070 Ti in it I haven't used it in the last year for any serious competition work…just for my hobby projects.</p>\n<p>For serious modelling and Kaggle I use Google Colab Pro to try out stuff and when I think I have an improvement then I move it over to Kaggle TPU's. </p>\n<p>For modifying your own code to get it up and running from your local GPU up to the Kaggle TPU's I would suggest the following 'research' path: Investigate and Create TFrecords from images, Switch your local code to use TF Datasets (if your not already doing so). Then take a look at some baseline and simple TPU code. In the Kaggle TPU Flowers competition there are a lot of notebooks explaning the concepts and showing how to do classification with TPU. That should get you started.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279387,
          "author_name": "victorasso",
          "author_url": "",
          "post_date": "04/20/2021 20:37:42",
          "content": "<p>I was stalling to really put some hours on learning that, it was already overwhelming the amount of time i had to allocate to get to know the tf basics, i guess i can't avoid that any longer.</p>\n<p>Thank you for your comments, so after all if i want to increase the image resolution on my training pipelines i will have to learn how to work with TPUs after all.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1281904,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/23/2021 12:20:09",
          "content": "<p>You need GPU RAM big enough to hold the model at this size. Have you tried smaller tiles (like 512)?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1279301,
      "author_name": "hetulvpatel",
      "author_url": "",
      "post_date": "04/20/2021 19:13:17",
      "content": "<p>Is anyone able to use mixed-precision training using tf.keras ? I tried but the loss becomes NaN. It should speed up training on the latest hardwares if it is working properly.</p>\n<p>According to the <a href=\"https://www.tensorflow.org/guide/mixed_precisionl\" target=\"_blank\">TensorFlow blog</a>, I added only these three lines of code in the beginning.</p>\n<pre><code>from tensorflow.keras import mixed_precision\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_global_policy(policy)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1279310,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/20/2021 19:17:40",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hetulvpatel\" target=\"_blank\">@hetulvpatel</a> Are you running on GPU? The nan loss is very likely caused by the last Dense layer not being a dtype of float32. This happens a lot. I believe that is somewhere mentioned in the TF documentation.</p>\n<p>Could also be the loss function.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279338,
          "author_name": "hetulvpatel",
          "author_url": "",
          "post_date": "04/20/2021 19:39:38",
          "content": "<ul>\n<li>Yes, I train it on Ampere GPU. </li>\n<li>When I remove these three lines, the training goes perfectly without loss being NaN. </li>\n<li>Also, I checked my Dense layer it is float32. </li>\n<li>I tried all different losses e.g BCE, Dice, Combo.. But in every case loss becomes nan.</li>\n</ul>\n<p><a href=\"https://www.kaggle.com/rsmits\" target=\"_blank\">@rsmits</a> Are you able to use Mixed Precision using tf.keras in your training?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1279392,
          "author_name": "rsmits",
          "author_url": "",
          "post_date": "04/20/2021 20:54:47",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hetulvpatel\" target=\"_blank\">@hetulvpatel</a> Hmm in that case I'am not sure what could be a further reason you get nan loss.</p>\n<p>For training I only use TPU's. I don't specify the use of bloat16 since the TPU's use them internally anyway. </p>\n<p>Last year during the Kaggle Google Landmark competition I read some interresting discussion on whether or not it would benefit a TPU enabled code to specify bfloat16 explicitly. I tried it for a few models and didn't notice any performance difference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1281785,
      "author_name": "southsakura",
      "author_url": "",
      "post_date": "04/23/2021 09:54:33",
      "content": "<p>18 mins on V100. batch size: 16</p>",
      "votes": null,
      "replies": [
        {
          "id": 1281902,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "04/23/2021 12:18:59",
          "content": "<p>Sweet! I guess on colab or GCP maybe?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1276485": "I am curious how long it takes to train one single model for other persons? \n\nOne single model trained on a single 3090 takes **around 20 minutes per epoch and per fold**. \n\nIt is a **very simple model and no augmentation** with **1024 x 1024 image tiles** and a **batch of 4 **(I guess I can go a little higher).\n\nAny feedback is more than welcome!",
    "1276525": "So far what i am doing is training with 256*256 with augs in model like efficient net it usually takes 2-5 mins per epoch , model like resnest and resnet50 takes 3 mins with the batch size of 32-45 , gpu used :- v100 and even p100 takes 7 to 10 mins",
    "1276529": "Thanks for sharing these details!\n\nHave you tried other models and/or bigger image tiles?",
    "1276574": "yeah i tried SOTA models from https://paperswithcode.com/task/medical-image-segmentation , But smaller models are performing better then larger models from my experimentations , didnt tried tiles bigger then 256",
    "1276633": "That makes sense: smaller models work on smaller tiles. If you try bigger models, you should try bigger tiles.",
    "1276638": "yeah next step is to try  512 and 1024 most probably",
    "1276642": "Awesome, let us know how it goes. :)",
    "1276767": "4 min per epoch with rtx 3090 and 512x512 resolution",
    "1277190": "Mine is taking around 4-5 minutes per epoch for efficientnet b4 on 512x512 with augs on kaggle gpus",
    "1277299": "Roughly 6 minutes per epoch for EffNet B5 with 512 pix tiles and medium Augmentation. Usually do max 20 epochs per fold on Kaggle TPU's \n\nWith the 9 hours of max session time you can't do a full cross validation run....so implemented fold stopping. Usually I only run the first 2 or 3 folds...if a modification turns out to be a succesfull one then I perform another run with 3 folds and a different seed. Works for me sofar.",
    "1277455": "Mine is around 5 minutes per epoch for an EffNetb0. 256x256 with heavy augmentations.",
    "1278763": "For me, resnext101, 512x512. batch_size 24, around 556 sec/epoch.\nRTX6000x2.",
    "1279268": "I'm not even able to train a batch of 1 with  1024, what kind of hardware are you running?",
    "1279276": "Btw i forgot to answer your question, i'm using 256px with some augumentations, my epoch time is 130s, depending on my folds / augs / external data / learning rate it would go around 90s to 220s. It will also take 20-60 epochs to convege with patence=20 (i never changed the patience, i wonder what it would do to the score)",
    "1279285": "Hi @victorasso Are you using TPU's? If not ... you might want to try that out :-) An EffNet B5 with a Unet/FPN/Linknet can train with batch of 2 to 3 per node ( so times 8)\n\nYou could try to switch the size to 512px.\n\nIf you have implemented some Early Stopping then changing the patience won't effect the score...it will however save you GPU/TPU time.",
    "1279301": "Is anyone able to use mixed-precision training using tf.keras ? I tried but the loss becomes NaN. It should speed up training on the latest hardwares if it is working properly.\n\nAccording to the [TensorFlow blog](https://www.tensorflow.org/guide/mixed_precisionl), I added only these three lines of code in the beginning.\n```\nfrom tensorflow.keras import mixed_precision\npolicy = mixed_precision.Policy('mixed_float16')\nmixed_precision.set_global_policy(policy)\n```",
    "1279310": "Hi @hetulvpatel Are you running on GPU? The nan loss is very likely caused by the last Dense layer not being a dtype of float32. This happens a lot. I believe that is somewhere mentioned in the TF documentation.\n\nCould also be the loss function.",
    "1279323": "Not really, i'm having a hard time trying to change my code to TPUs for kaggle envs since i used my own machine to train my initial models, i even tried to rent an azure VM to do the job but after all that work all i got was a crippled env and a 50 dollar bill for a day of cringe.\n\nDo you have your own hardware or do you only use the kaggle env? If so, any tips for saving those precious hardware accelerated hours?",
    "1279338": "Yes, I train it on Ampere GPU. \n- When I remove these three lines, the training goes perfectly without loss being NaN. \n- Also, I checked my Dense layer it is float32. \n- I tried all different losses e.g BCE, Dice, Combo.. But in every case loss becomes nan.\n\n@rsmits Are you able to use Mixed Precision using tf.keras in your training?",
    "1279348": "victorasso I do have a Data Science PC ... but with an old 1070 Ti in it I haven't used it in the last year for any serious competition work...just for my hobby projects.\n\nFor serious modelling and Kaggle I use Google Colab Pro to try out stuff and when I think I have an improvement then I move it over to Kaggle TPU's. \n\nFor modifying your own code to get it up and running from your local GPU up to the Kaggle TPU's I would suggest the following 'research' path: Investigate and Create TFrecords from images, Switch your local code to use TF Datasets (if your not already doing so). Then take a look at some baseline and simple TPU code. In the Kaggle TPU Flowers competition there are a lot of notebooks explaning the concepts and showing how to do classification with TPU. That should get you started.",
    "1279387": "I was stalling to really put some hours on learning that, it was already overwhelming the amount of time i had to allocate to get to know the tf basics, i guess i can't avoid that any longer.\n\nThank you for your comments, so after all if i want to increase the image resolution on my training pipelines i will have to learn how to work with TPUs after all.",
    "1279392": "Hi @hetulvpatel Hmm in that case I'am not sure what could be a further reason you get nan loss.\n\nFor training I only use TPU's. I don't specify the use of bloat16 since the TPU's use them internally anyway. \n\nLast year during the Kaggle Google Landmark competition I read some interresting discussion on whether or not it would benefit a TPU enabled code to specify bfloat16 explicitly. I tried it for a few models and didn't notice any performance difference.",
    "1281785": "18 mins on V100. batch size: 16",
    "1281902": "Sweet! I guess on colab or GCP maybe?",
    "1281904": "You need GPU RAM big enough to hold the model at this size. Have you tried smaller tiles (like 512)?"
  },
  "source": "meta"
}