{
  "id": 131394,
  "title": "What's your training time?",
  "url": "/competitions/bengaliai-cv19/discussion/131394",
  "author_name": "Jie Wu",
  "post_date": "2020-02-19T16:39:03.417000",
  "votes": 13,
  "comment_count": 56,
  "views": 0,
  "content": "<p>Can someone share their time of training a model? Now I am using se-resnext50_32x4d, size 128*128, which takes 11mins without data aug for each epoch on a single RTX 2080Ti. With data aug, it takes 25-26min for each epoch. Train one model (single fold) will take 16-17 hours for just 40 epochs, this is very long. Just wondering those who trained their models on a 4-5 folds, with 80 epochs or 150 epochs, how long did they take?</p>",
  "messages": [
    {
      "id": 753029,
      "postDate": "2020-02-21T16:25:33.607Z",
      "content": "<p>I got <strong>~100s</strong> for ResNeXt50SE on 128x128 images using 2x 2080Ti + NVlink (though I'm not sure that the last one makes a difference for DL applications) with a simple augmentation like MixUp or CutMix. If one is using  2080Ti, <strong>half precision</strong> is the thing one nearly always must to do: it gives 1.5-2x boost. Though, APEX doesn't work well with multiple GPU, and I needed to implement my own half precision and dynamic rescaling procedures or alternatively use fast.ai. Also, make sure that you almost fully occupy the GPU RAM, though sometimes u need to train with giant batches, like 768 for the above considered case. With tiny batches the speed can degrade 1.5-2+ times. Finally, water/hybrid  cooled GPUs are the way to go: with OC ~2000MHz and 325W energy consumption it heats up only to ~65C under ~90% load. Also good CPU and SSD can be really helpful in some particular cases. And just a small final remark, do not use Windows for DL: I got ~6-8min per epoch for the same setup as I use on Linux.</p>",
      "rawMarkdown": "I got **~100s** for ResNeXt50SE on 128x128 images using 2x 2080Ti + NVlink (though I'm not sure that the last one makes a difference for DL applications) with a simple augmentation like MixUp or CutMix. If one is using  2080Ti, **half precision** is the thing one nearly always must to do: it gives 1.5-2x boost. Though, APEX doesn't work well with multiple GPU, and I needed to implement my own half precision and dynamic rescaling procedures or alternatively use fast.ai. Also, make sure that you almost fully occupy the GPU RAM, though sometimes u need to train with giant batches, like 768 for the above considered case. With tiny batches the speed can degrade 1.5-2+ times. Finally, water/hybrid  cooled GPUs are the way to go: with OC ~2000MHz and 325W energy consumption it heats up only to ~65C under ~90% load. Also good CPU and SSD can be really helpful in some particular cases. And just a small final remark, do not use Windows for DL: I got ~6-8min per epoch for the same setup as I use on Linux.",
      "votes": 13,
      "replies": [
        {
          "id": 753219,
          "postDate": "2020-02-21T21:43:44.217Z",
          "content": "<p>Wait... Is it faster to train on Linux than Windows? You said you got 6-8 min vs 100 s per epoch? That's quite a difference if I'm understanding correctly</p>",
          "rawMarkdown": "Wait... Is it faster to train on Linux than Windows? You said you got 6-8 min vs 100 s per epoch? That's quite a difference if I'm understanding correctly"
        },
        {
          "id": 753258,
          "postDate": "2020-02-22T00:02:24.620Z",
          "content": "<p>Yes it is correct. I spent some time trying to figure it out, but couldn't find the origin of this issue. So I found that it is simpler to install Ubuntu (or dual system with wubi - the simplest way to get Linux working with Windows).</p>",
          "rawMarkdown": "Yes it is correct. I spent some time trying to figure it out, but couldn't find the origin of this issue. So I found that it is simpler to install Ubuntu (or dual system with wubi - the simplest way to get Linux working with Windows).",
          "votes": 1
        },
        {
          "id": 753261,
          "postDate": "2020-02-22T00:13:14.730Z",
          "content": "<p>I see. Now I gotta try ubuntu. That would be huge if it is the same for me. Right now my best model (0.9814 LB) takes 700 secs for one epoch on a 2080 so I need to see how long it takes to do one epoch on linux. Thanks for the help!</p>",
          "rawMarkdown": "I see. Now I gotta try ubuntu. That would be huge if it is the same for me. Right now my best model (0.9814 LB) takes 700 secs for one epoch on a 2080 so I need to see how long it takes to do one epoch on linux. Thanks for the help!"
        },
        {
          "id": 753819,
          "postDate": "2020-02-22T17:14:25.717Z",
          "content": "<p>Maybe because we've to set num_workers=0 in windows, atleast if we're using notebooks?</p>",
          "rawMarkdown": "Maybe because we've to set num_workers=0 in windows, atleast if we're using notebooks?"
        },
        {
          "id": 755266,
          "postDate": "2020-02-24T15:59:27.310Z",
          "content": "<p>half precision means apex？</p>",
          "rawMarkdown": "half precision means apex？"
        },
        {
          "id": 755385,
          "postDate": "2020-02-24T18:18:19.767Z",
          "content": "<p>Half precision (mixed precision) is implemented in APEX library, if you use single GPU. And u just need to add only a few lines to make it work, but the thing is that u need to have 20x series GPU or V100 to get a speed boost.</p>",
          "rawMarkdown": "Half precision (mixed precision) is implemented in APEX library, if you use single GPU. And u just need to add only a few lines to make it work, but the thing is that u need to have 20x series GPU or V100 to get a speed boost."
        },
        {
          "id": 755520,
          "postDate": "2020-02-24T21:17:27.680Z",
          "content": "<p>Do you have any tips on building efficient data pipeline with augmentation in Pytorch? I recently started using it but I don't think my data pipeline is good since the GPU usage drops every time the code loads data onto the GPU, meaning my training process is not really concurrent.</p>",
          "rawMarkdown": "Do you have any tips on building efficient data pipeline with augmentation in Pytorch? I recently started using it but I don't think my data pipeline is good since the GPU usage drops every time the code loads data onto the GPU, meaning my training process is not really concurrent."
        },
        {
          "id": 755526,
          "postDate": "2020-02-24T21:23:20.910Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> <a href=\"https://docs.nvidia.com/deeplearning/sdk/dali-developer-guide/docs/index.html\">Nvidia DALI</a> can preprocess images on GPU</p>",
          "rawMarkdown": "@shujun717 [Nvidia DALI][1] can preprocess images on GPU\n\n[1]: https://docs.nvidia.com/deeplearning/sdk/dali-developer-guide/docs/index.html"
        },
        {
          "id": 755535,
          "postDate": "2020-02-24T21:32:23.080Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> Did you get a speed increase switching to Ubuntu?</p>",
          "rawMarkdown": "@shujun717 Did you get a speed increase switching to Ubuntu?"
        },
        {
          "id": 755547,
          "postDate": "2020-02-24T21:44:48.957Z",
          "content": "<p>Thanks for the advice and yeah, I think I got a speed increase of around 20% (as opposed to a few times by <a href=\"/iafoss\">@iafoss</a> ). But it may just be ubuntu is using more CPUs (windows only uses 1 for some reason but ubuntu uses everything it seems).</p>",
          "rawMarkdown": "Thanks for the advice and yeah, I think I got a speed increase of around 20% (as opposed to a few times by @iafoss ). But it may just be ubuntu is using more CPUs (windows only uses 1 for some reason but ubuntu uses everything it seems).",
          "votes": 1
        },
        {
          "id": 755597,
          "postDate": "2020-02-24T23:37:42.340Z",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> I'm just curious, did u use pure Pytorch or fast.ai based code for your test? The comparison Ubuntu vs Windows, I made, was for fast.ai one. Could it be some issue related not to Pytorch but to fast.ai instead?\nOn Ubuntu though pure Pytorch code I wrote for this competition has about the same speed as fast.ai (~5s slower because, probably, I did something in not the most optimal way).</p>",
          "rawMarkdown": "@shujun717 I'm just curious, did u use pure Pytorch or fast.ai based code for your test? The comparison Ubuntu vs Windows, I made, was for fast.ai one. Could it be some issue related not to Pytorch but to fast.ai instead?\nOn Ubuntu though pure Pytorch code I wrote for this competition has about the same speed as fast.ai (~5s slower because, probably, I did something in not the most optimal way)."
        },
        {
          "id": 755602,
          "postDate": "2020-02-24T23:56:29.517Z",
          "content": "<p>I used pure pytorch to code a resnet variant. My training loop is just a for loop with index shuffling and cutout. Do you use some type of dataloader to run asynchronously? </p>\n\n<p>I'm gonna do some more experiments with this code: <a href=\"https://github.com/yunjey/pytorch-tutorial/blob/master/tutorials/02-intermediate/deep_residual_network/main.py#L76-L113\">https://github.com/yunjey/pytorch-tutorial/blob/master/tutorials/02-intermediate/deep_residual_network/main.py#L76-L113</a>\nand let you know how it goes. </p>",
          "rawMarkdown": "I used pure pytorch to code a resnet variant. My training loop is just a for loop with index shuffling and cutout. Do you use some type of dataloader to run asynchronously? \n\nI'm gonna do some more experiments with this code: https://github.com/yunjey/pytorch-tutorial/blob/master/tutorials/02-intermediate/deep_residual_network/main.py#L76-L113\nand let you know how it goes. "
        },
        {
          "id": 755640,
          "postDate": "2020-02-25T01:22:00.620Z",
          "content": "<p>Thanks for clarification. In both cases I use Pytorch DataLoader class. The the main issue with it, if I understood correctly, is that in Windows a new processes, if multithread dataloader is used, is created by <code>spawn</code>system call, which is quite heavy and requites putting everything into <code>if __name__ == \"__main__\"</code> statement, while on Unix the threads can be created by <code>fork</code>system call, which is much cheaper. If Pytorch was somehow using a pool of processes, the issues could be mitigated to some extend, but it seams not many ppl are interested in Deep Learning on Windows,  and therefore not many efforts are put into this direction.\nBut it my case there may be something else, I tried to run fast.ai with num_workers=0  as well, but it didn't help(  Though, when I ran GPU  Deep Learning benchmarks on Windows, which do not have a dataloader but use something similar as your setup, I got comparable results to ones that should be expected for my GPUs.</p>",
          "rawMarkdown": "Thanks for clarification. In both cases I use Pytorch DataLoader class. The the main issue with it, if I understood correctly, is that in Windows a new processes, if multithread dataloader is used, is created by `spawn `system call, which is quite heavy and requites putting everything into `if __name__ == \"__main__\"` statement, while on Unix the threads can be created by `fork `system call, which is much cheaper. If Pytorch was somehow using a pool of processes, the issues could be mitigated to some extend, but it seams not many ppl are interested in Deep Learning on Windows,  and therefore not many efforts are put into this direction.\nBut it my case there may be something else, I tried to run fast.ai with num_workers=0  as well, but it didn't help(  Though, when I ran GPU  Deep Learning benchmarks on Windows, which do not have a dataloader but use something similar as your setup, I got comparable results to ones that should be expected for my GPUs.",
          "votes": 1
        },
        {
          "id": 755649,
          "postDate": "2020-02-25T01:32:14.980Z",
          "content": "<p>thanks</p>",
          "rawMarkdown": "thanks"
        },
        {
          "id": 755780,
          "postDate": "2020-02-25T05:45:52.103Z",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> Interesting. It seems like my code is poorly optimized then. A resnet variant with around 30 million params on 88x88 takes 470 s per epoch on 2x2080 for me. </p>\n\n<p>BTW you wanna team up? I think we can make run for .99</p>",
          "rawMarkdown": "@iafoss Interesting. It seems like my code is poorly optimized then. A resnet variant with around 30 million params on 88x88 takes 470 s per epoch on 2x2080 for me. \n\nBTW you wanna team up? I think we can make run for .99"
        }
      ]
    },
    {
      "id": 750730,
      "postDate": "2020-02-19T16:39:03.417Z",
      "content": "<p>Can someone share their time of training a model? Now I am using se-resnext50_32x4d, size 128*128, which takes 11mins without data aug for each epoch on a single RTX 2080Ti. With data aug, it takes 25-26min for each epoch. Train one model (single fold) will take 16-17 hours for just 40 epochs, this is very long. Just wondering those who trained their models on a 4-5 folds, with 80 epochs or 150 epochs, how long did they take?</p>",
      "rawMarkdown": "Can someone share their time of training a model? Now I am using se-resnext50_32x4d, size 128*128, which takes 11mins without data aug for each epoch on a single RTX 2080Ti. With data aug, it takes 25-26min for each epoch. Train one model (single fold) will take 16-17 hours for just 40 epochs, this is very long. Just wondering those who trained their models on a 4-5 folds, with 80 epochs or 150 epochs, how long did they take?",
      "votes": 13
    },
    {
      "id": 750805,
      "postDate": "2020-02-19T17:48:51.570Z",
      "content": "<p>se-resnext50_32x4d, 128x128 (loaded from png), basic augmentation (no mixup, no cutmix); I use apex (o2), SGD, 8% validation (184772 train images), batch size: 128, no mish activation; 25892010 params</p>\n\n<p>1 epoch: 7 minutes,\n100 epochs: ~ 12 hours.</p>\n\n<p>I have a single RTX-2080ti (avg. card usage above 90%)</p>",
      "rawMarkdown": "se-resnext50_32x4d, 128x128 (loaded from png), basic augmentation (no mixup, no cutmix); I use apex (o2), SGD, 8% validation (184772 train images), batch size: 128, no mish activation; 25892010 params\n\n1 epoch: 7 minutes,\n100 epochs: ~ 12 hours.\n\nI have a single RTX-2080ti (avg. card usage above 90%)",
      "votes": 5,
      "replies": [
        {
          "id": 750823,
          "postDate": "2020-02-19T18:03:51.483Z",
          "content": "<p>Hi <a href=\"/pestipeti\">@pestipeti</a> <br>\nDid you get better results with SGD than Adam ?</p>",
          "rawMarkdown": "Hi @pestipeti  \nDid you get better results with SGD than Adam ?"
        },
        {
          "id": 750850,
          "postDate": "2020-02-19T18:41:03.207Z",
          "content": "<p>SGD + Cosine is best for me; Adam + ReduceOnPlateau is faster,  more or less the same score, but it depends on the config.</p>",
          "rawMarkdown": "SGD + Cosine is best for me; Adam + ReduceOnPlateau is faster,  more or less the same score, but it depends on the config.",
          "votes": 2
        }
      ]
    },
    {
      "id": 753007,
      "postDate": "2020-02-21T16:01:00.433Z",
      "content": "<p>Wow</p>\n\n<p>Here I am trying to pass this huge chunk of people at 0.9663 (and then the one at 0.9696) only using Kaggle notebooks with my epoch time being less than 5 minutes.</p>\n\n<p>If anybody has a spare 2080Ti, feel free to reach out :P</p>",
      "rawMarkdown": "Wow\n\nHere I am trying to pass this huge chunk of people at 0.9663 (and then the one at 0.9696) only using Kaggle notebooks with my epoch time being less than 5 minutes.\n\nIf anybody has a spare 2080Ti, feel free to reach out :P",
      "votes": 3
    },
    {
      "id": 752439,
      "postDate": "2020-02-21T02:46:12.420Z",
      "content": "<p>Dataloader uses the num_works set for 16 is  very helpful(cpu i9 9900k)\nwhich save almost half time for us.\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 2x2080ti\napex: not used\nsplit: 9/10 train, 1/10 valid\ntrain(1 epoch): 220s</p>",
      "rawMarkdown": "Dataloader uses the num_works set for 16 is  very helpful(cpu i9 9900k)\nwhich save almost half time for us.\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 2x2080ti\napex: not used\nsplit: 9/10 train, 1/10 valid\ntrain(1 epoch): 220s",
      "votes": 3,
      "replies": [
        {
          "id": 755662,
          "postDate": "2020-02-25T01:52:03.593Z",
          "content": "<p>how much memory do you hava？</p>",
          "rawMarkdown": "how much memory do you hava？"
        }
      ]
    },
    {
      "id": 750966,
      "postDate": "2020-02-19T21:58:09.913Z",
      "content": "<p>our current setup with 64x64 images - 2x1080ti takes about ~2 minutes per epoch.</p>",
      "rawMarkdown": "our current setup with 64x64 images - 2x1080ti takes about ~2 minutes per epoch.",
      "votes": 1,
      "replies": [
        {
          "id": 750975,
          "postDate": "2020-02-19T22:07:24.557Z",
          "content": "<p>If I used 64x64, it went down to 2.5min per epoch on single RTX 2080Ti.</p>",
          "rawMarkdown": "If I used 64x64, it went down to 2.5min per epoch on single RTX 2080Ti.",
          "votes": 1
        }
      ]
    },
    {
      "id": 750776,
      "postDate": "2020-02-19T17:20:42.517Z",
      "content": "<p>I am using a se-resnext101_32x4d arhitecture with a 128*128 size. With augmentation is taking about 30 min per epoch on a single RTX 2080Ti . In order to converge for me I have to wait much more than 40 epochs so i spend about 24 hours per fold, so 5 entire days for all 5 folds...pretty god damn long time</p>",
      "rawMarkdown": "I am using a se-resnext101_32x4d arhitecture with a 128*128 size. With augmentation is taking about 30 min per epoch on a single RTX 2080Ti . In order to converge for me I have to wait much more than 40 epochs so i spend about 24 hours per fold, so 5 entire days for all 5 folds...pretty god damn long time",
      "votes": 1,
      "replies": [
        {
          "id": 750781,
          "postDate": "2020-02-19T17:23:37.543Z",
          "content": "<p>Seems we are the same. I also tried se-resnext101_32x4d, it's over 30min per epoch, but still running. I have not think to move to 5 folds. Because it takes so long.</p>",
          "rawMarkdown": "Seems we are the same. I also tried se-resnext101_32x4d, it's over 30min per epoch, but still running. I have not think to move to 5 folds. Because it takes so long.",
          "votes": 2
        },
        {
          "id": 750784,
          "postDate": "2020-02-19T17:27:58.007Z",
          "content": "<p>Yes, I usually make a series of modification, train for 50 epochs on a single fold and use that point as a benchmark to compare with another architecture and different settings that I made in the past. And when I have something superior on that point I let it train the entire fold and if also the entire fold result is best that the best model i had than i let it train on the entire 5 folds</p>",
          "rawMarkdown": "Yes, I usually make a series of modification, train for 50 epochs on a single fold and use that point as a benchmark to compare with another architecture and different settings that I made in the past. And when I have something superior on that point I let it train the entire fold and if also the entire fold result is best that the best model i had than i let it train on the entire 5 folds",
          "votes": 1
        },
        {
          "id": 750956,
          "postDate": "2020-02-19T21:43:25.027Z",
          "content": "<p>wow what augmentation are you using? I am using cutmix + mixup + some affine on 128x128, same model as yours, it's around 15 mins.  i have titan rtx, but its really not faster than 2080ti really. and not using half precision at the moment</p>",
          "rawMarkdown": "wow what augmentation are you using? I am using cutmix + mixup + some affine on 128x128, same model as yours, it's around 15 mins.  i have titan rtx, but its really not faster than 2080ti really. and not using half precision at the moment"
        }
      ]
    },
    {
      "id": 750739,
      "postDate": "2020-02-19T16:44:41.227Z",
      "content": "<p>I have similar times for my 2080 and on google colab. Haven't trained for more than 100 epochs yet.</p>",
      "rawMarkdown": "I have similar times for my 2080 and on google colab. Haven't trained for more than 100 epochs yet.",
      "votes": 1
    },
    {
      "id": 750733,
      "postDate": "2020-02-19T16:41:34.987Z",
      "content": "<p>I spend around 7 hours for training model for 40 epochs. </p>",
      "rawMarkdown": "I spend around 7 hours for training model for 40 epochs. ",
      "votes": 1,
      "replies": [
        {
          "id": 750743,
          "postDate": "2020-02-19T16:47:48.240Z",
          "content": "<p>Me too, without data aug about 7hrs</p>",
          "rawMarkdown": "Me too, without data aug about 7hrs"
        },
        {
          "id": 750774,
          "postDate": "2020-02-19T17:18:25.013Z",
          "content": "<p>I am using Data argumentation. I have different and smaller model I guess. </p>",
          "rawMarkdown": "I am using Data argumentation. I have different and smaller model I guess. "
        }
      ]
    },
    {
      "id": 751003,
      "postDate": "2020-02-19T23:04:43.483Z",
      "content": "<p><code>\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 1x1080ti\napex: not used\nsplit: 5/6 train, 1/6 valid\ntrain(1 epoch): 700s\ntrain(40 epochs): 8 hours\n</code></p>",
      "rawMarkdown": "```\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 1x1080ti\napex: not used\nsplit: 5/6 train, 1/6 valid\ntrain(1 epoch): 700s\ntrain(40 epochs): 8 hours\n```",
      "votes": 2,
      "replies": [
        {
          "id": 751141,
          "postDate": "2020-02-20T02:50:09.073Z",
          "content": "<p>I use 2x2080ti,  and use more training time about 715s one epoch.\nDo you remove the maxpooling layer?</p>",
          "rawMarkdown": "I use 2x2080ti,  and use more training time about 715s one epoch.\nDo you remove the maxpooling layer?"
        },
        {
          "id": 751482,
          "postDate": "2020-02-20T08:16:46.040Z",
          "content": "<p>No, I modified the model a bit but it should not affect the training speed too much. </p>",
          "rawMarkdown": "No, I modified the model a bit but it should not affect the training speed too much. "
        }
      ]
    },
    {
      "id": 750927,
      "postDate": "2020-02-19T20:42:08.363Z",
      "content": "<p>Update: After some changes, on single RTX 2080Ti GPU, 90% of data for training: se-resnext50_32x4d , size 128x128 with data aug, 7min per epoch; se-resnext101_32x4d , size 128x128 with data aug, 12.5min per epoch.</p>",
      "rawMarkdown": "Update: After some changes, on single RTX 2080Ti GPU, 90% of data for training: se-resnext50_32x4d , size 128x128 with data aug, 7min per epoch; se-resnext101_32x4d , size 128x128 with data aug, 12.5min per epoch.",
      "votes": 2,
      "replies": [
        {
          "id": 750992,
          "postDate": "2020-02-19T22:25:12.403Z",
          "content": "<p>Can you share what some changes you've made? I have a similar situation with you..</p>",
          "rawMarkdown": "Can you share what some changes you've made? I have a similar situation with you.."
        }
      ]
    },
    {
      "id": 756602,
      "postDate": "2020-02-25T22:10:37.267Z",
      "content": "<p>Setting num workers to a large amount (like 8) is important if your data processing is limiting your speed!</p>",
      "rawMarkdown": "Setting num workers to a large amount (like 8) is important if your data processing is limiting your speed!"
    },
    {
      "id": 756493,
      "postDate": "2020-02-25T19:07:49.830Z",
      "content": "<p>Image size: 137x236\nAug: simple scale, rotate, shift\nArch: resnet34\nGPU: rtx 2070 maxp\nBatch Size: 256\nnum_workers: 4\nTime/epoch: ~7.5 mins</p>",
      "rawMarkdown": "Image size: 137x236\nAug: simple scale, rotate, shift\nArch: resnet34\nGPU: rtx 2070 maxp\nBatch Size: 256\nnum_workers: 4\nTime/epoch: ~7.5 mins"
    },
    {
      "id": 755224,
      "postDate": "2020-02-24T15:12:57.830Z",
      "content": "<p>So I see that not many have 2080s and these super fast training speeds. But I find plentiful comments in CV score thread, stating nice CV and LB scores. So I am wonder if everyone is really doing cross validation and averaging the scores over the folds? Or are they just doing single fold and calling it CV score generally?   </p>\n\n<p>And I am sure many are running the experiments day and night probably everyday. I hope I am not sounding negative. I just started kaggle couple of weeks before and just want to know what everyone is doing. I hope someone will enlighten me here. Thanks a lot!</p>",
      "rawMarkdown": "So I see that not many have 2080s and these super fast training speeds. But I find plentiful comments in CV score thread, stating nice CV and LB scores. So I am wonder if everyone is really doing cross validation and averaging the scores over the folds? Or are they just doing single fold and calling it CV score generally?   \n\nAnd I am sure many are running the experiments day and night probably everyday. I hope I am not sounding negative. I just started kaggle couple of weeks before and just want to know what everyone is doing. I hope someone will enlighten me here. Thanks a lot!",
      "replies": [
        {
          "id": 755249,
          "postDate": "2020-02-24T15:44:28.260Z",
          "content": "<p>Mix of both I'd guess. 'Best single model' technically only means one model and not an ensemble of 5 fold models. You can calculate CV of just one model too.</p>",
          "rawMarkdown": "Mix of both I'd guess. 'Best single model' technically only means one model and not an ensemble of 5 fold models. You can calculate CV of just one model too.",
          "votes": 1
        },
        {
          "id": 755637,
          "postDate": "2020-02-25T01:10:03.610Z",
          "content": "<p>True, even to calculate CV for one model, you have to train it multiple times with different splits as the validation right?</p>",
          "rawMarkdown": "True, even to calculate CV for one model, you have to train it multiple times with different splits as the validation right?"
        },
        {
          "id": 755648,
          "postDate": "2020-02-25T01:29:29.600Z",
          "content": "<p>You can just train one fold and validate</p>",
          "rawMarkdown": "You can just train one fold and validate"
        },
        {
          "id": 755653,
          "postDate": "2020-02-25T01:35:08.710Z",
          "content": "<p>Yup then that's not \"cross-validation\" anymore. I understand what you mean tho. Thanks for the quick reply. </p>",
          "rawMarkdown": "Yup then that's not \"cross-validation\" anymore. I understand what you mean tho. Thanks for the quick reply. "
        }
      ]
    },
    {
      "id": 754726,
      "postDate": "2020-02-24T01:24:06.900Z",
      "content": "<p>model: efficient_net_b1  </p>\n\n<p>batch_size:100  </p>\n\n<p>imgsize: 128x128  </p>\n\n<p>gpu: 1x2080ti </p>\n\n<p>apex: not used  </p>\n\n<p>mix-precision: not used  </p>\n\n<p>system: ubuntu16  </p>\n\n<p>split: 0.01 valid  </p>\n\n<p>num_workers: 0</p>\n\n<p>train(1 epoch): about 2hours  </p>",
      "rawMarkdown": "\nmodel: efficient_net_b1  \n\nbatch_size:100  \n\nimgsize: 128x128  \n\ngpu: 1x2080ti \n \napex: not used  \n\nmix-precision: not used  \n\nsystem: ubuntu16  \n\nsplit: 0.01 valid  \n\nnum_workers: 0\n\ntrain(1 epoch): about 2hours  \n\n\n",
      "replies": [
        {
          "id": 754962,
          "postDate": "2020-02-24T09:19:39.403Z",
          "content": "<p>Oh, it's very slow. Try to increase num_workers and batch_size</p>",
          "rawMarkdown": "Oh, it's very slow. Try to increase num_workers and batch_size"
        }
      ]
    },
    {
      "id": 753727,
      "postDate": "2020-02-22T15:49:42.767Z",
      "content": "<p>maybe try Apex? and fine-tuning paras without apex</p>",
      "rawMarkdown": "maybe try Apex? and fine-tuning paras without apex"
    },
    {
      "id": 752627,
      "postDate": "2020-02-21T08:23:00.323Z",
      "content": "<p>Similar time. Now I'm using img 64x64 and 50 epochs training for fast experiments. This can cost around 5 hours and I can test several modifications in a single day.</p>",
      "rawMarkdown": "Similar time. Now I'm using img 64x64 and 50 epochs training for fast experiments. This can cost around 5 hours and I can test several modifications in a single day."
    },
    {
      "id": 752623,
      "postDate": "2020-02-21T08:18:20.070Z",
      "content": "<p>se-resnext50_32x4d, size 128x128x3 with augmentation on a single RTX 2080Ti takes 6 min on my side</p>",
      "rawMarkdown": "se-resnext50_32x4d, size 128x128x3 with augmentation on a single RTX 2080Ti takes 6 min on my side"
    },
    {
      "id": 752330,
      "postDate": "2020-02-20T22:07:31.593Z",
      "content": "<p>Why not try mixed precision, faster with less precision reduction. </p>",
      "rawMarkdown": "Why not try mixed precision, faster with less precision reduction. "
    },
    {
      "id": 752293,
      "postDate": "2020-02-20T20:54:24.603Z",
      "content": "<p>seresnext101\nsize 256\ngpu 2080ti\n1 epoch ~40min...</p>",
      "rawMarkdown": "seresnext101\nsize 256\ngpu 2080ti\n1 epoch ~40min..."
    },
    {
      "id": 751471,
      "postDate": "2020-02-20T08:03:49.340Z",
      "content": "<p>model: se_resnext50_32x4d\nimgsize: 137x236\ngpu: 2080Ti\napex: O1\nsplit: 80/20\ntrain(1 epoch): ~10min</p>",
      "rawMarkdown": "model: se_resnext50_32x4d\nimgsize: 137x236\ngpu: 2080Ti\napex: O1\nsplit: 80/20\ntrain(1 epoch): ~10min"
    },
    {
      "id": 751423,
      "postDate": "2020-02-20T07:29:33.167Z",
      "content": "<p>Anyone trains with 224 x 224? </p>\n\n<p>For my 128 x 128:\n<code>\ntraining time: ~15 mins/epoch (Depends on what you augment)\ngpu: RTX2080super\n</code></p>",
      "rawMarkdown": "Anyone trains with 224 x 224? \n\nFor my 128 x 128:\n```\ntraining time: ~15 mins/epoch (Depends on what you augment)\ngpu: RTX2080super\n```",
      "replies": [
        {
          "id": 753474,
          "postDate": "2020-02-22T08:31:54.833Z",
          "content": "<p>My 224x224 model on google colab takes ~1 hr per epoch using densenet161.</p>",
          "rawMarkdown": "My 224x224 model on google colab takes ~1 hr per epoch using densenet161."
        }
      ]
    },
    {
      "id": 750993,
      "postDate": "2020-02-19T22:37:20.763Z",
      "content": "<p>RTX2070 Super, ~resnext50_32x4d, size 128*128, Half Precision, 64 bs, 20% validation =&gt; 7 min per epoch</p>",
      "rawMarkdown": "RTX2070 Super, ~resnext50_32x4d, size 128*128, Half Precision, 64 bs, 20% validation =&gt; 7 min per epoch"
    },
    {
      "id": 751620,
      "postDate": "2020-02-20T10:44:29.167Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 753029,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-02-21T16:25:33.607000",
      "content": "<p>I got <strong>~100s</strong> for ResNeXt50SE on 128x128 images using 2x 2080Ti + NVlink (though I'm not sure that the last one makes a difference for DL applications) with a simple augmentation like MixUp or CutMix. If one is using  2080Ti, <strong>half precision</strong> is the thing one nearly always must to do: it gives 1.5-2x boost. Though, APEX doesn't work well with multiple GPU, and I needed to implement my own half precision and dynamic rescaling procedures or alternatively use fast.ai. Also, make sure that you almost fully occupy the GPU RAM, though sometimes u need to train with giant batches, like 768 for the above considered case. With tiny batches the speed can degrade 1.5-2+ times. Finally, water/hybrid  cooled GPUs are the way to go: with OC ~2000MHz and 325W energy consumption it heats up only to ~65C under ~90% load. Also good CPU and SSD can be really helpful in some particular cases. And just a small final remark, do not use Windows for DL: I got ~6-8min per epoch for the same setup as I use on Linux.</p>",
      "votes": 13,
      "replies": [
        {
          "id": 753219,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-02-21T21:43:44.217000",
          "content": "<p>Wait... Is it faster to train on Linux than Windows? You said you got 6-8 min vs 100 s per epoch? That's quite a difference if I'm understanding correctly</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753258,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-22T00:02:24.620000",
          "content": "<p>Yes it is correct. I spent some time trying to figure it out, but couldn't find the origin of this issue. So I found that it is simpler to install Ubuntu (or dual system with wubi - the simplest way to get Linux working with Windows).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 753261,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-02-22T00:13:14.730000",
          "content": "<p>I see. Now I gotta try ubuntu. That would be huge if it is the same for me. Right now my best model (0.9814 LB) takes 700 secs for one epoch on a 2080 so I need to see how long it takes to do one epoch on linux. Thanks for the help!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753819,
          "author_name": "Mighty Rains",
          "author_url": "",
          "post_date": "2020-02-22T17:14:25.717000",
          "content": "<p>Maybe because we've to set num_workers=0 in windows, atleast if we're using notebooks?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755266,
          "author_name": "dongliang",
          "author_url": "",
          "post_date": "2020-02-24T15:59:27.310000",
          "content": "<p>half precision means apex？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755385,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-24T18:18:19.767000",
          "content": "<p>Half precision (mixed precision) is implemented in APEX library, if you use single GPU. And u just need to add only a few lines to make it work, but the thing is that u need to have 20x series GPU or V100 to get a speed boost.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755520,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-02-24T21:17:27.680000",
          "content": "<p>Do you have any tips on building efficient data pipeline with augmentation in Pytorch? I recently started using it but I don't think my data pipeline is good since the GPU usage drops every time the code loads data onto the GPU, meaning my training process is not really concurrent.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755526,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-02-24T21:23:20.910000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> <a href=\"https://docs.nvidia.com/deeplearning/sdk/dali-developer-guide/docs/index.html\">Nvidia DALI</a> can preprocess images on GPU</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755535,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2020-02-24T21:32:23.080000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> Did you get a speed increase switching to Ubuntu?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755547,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-02-24T21:44:48.957000",
          "content": "<p>Thanks for the advice and yeah, I think I got a speed increase of around 20% (as opposed to a few times by <a href=\"/iafoss\">@iafoss</a> ). But it may just be ubuntu is using more CPUs (windows only uses 1 for some reason but ubuntu uses everything it seems).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 755597,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-24T23:37:42.340000",
          "content": "<p><a href=\"/shujun717\">@shujun717</a> I'm just curious, did u use pure Pytorch or fast.ai based code for your test? The comparison Ubuntu vs Windows, I made, was for fast.ai one. Could it be some issue related not to Pytorch but to fast.ai instead?\nOn Ubuntu though pure Pytorch code I wrote for this competition has about the same speed as fast.ai (~5s slower because, probably, I did something in not the most optimal way).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755602,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-02-24T23:56:29.517000",
          "content": "<p>I used pure pytorch to code a resnet variant. My training loop is just a for loop with index shuffling and cutout. Do you use some type of dataloader to run asynchronously? </p>\n\n<p>I'm gonna do some more experiments with this code: <a href=\"https://github.com/yunjey/pytorch-tutorial/blob/master/tutorials/02-intermediate/deep_residual_network/main.py#L76-L113\">https://github.com/yunjey/pytorch-tutorial/blob/master/tutorials/02-intermediate/deep_residual_network/main.py#L76-L113</a>\nand let you know how it goes. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755640,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2020-02-25T01:22:00.620000",
          "content": "<p>Thanks for clarification. In both cases I use Pytorch DataLoader class. The the main issue with it, if I understood correctly, is that in Windows a new processes, if multithread dataloader is used, is created by <code>spawn</code>system call, which is quite heavy and requites putting everything into <code>if __name__ == \"__main__\"</code> statement, while on Unix the threads can be created by <code>fork</code>system call, which is much cheaper. If Pytorch was somehow using a pool of processes, the issues could be mitigated to some extend, but it seams not many ppl are interested in Deep Learning on Windows,  and therefore not many efforts are put into this direction.\nBut it my case there may be something else, I tried to run fast.ai with num_workers=0  as well, but it didn't help(  Though, when I ran GPU  Deep Learning benchmarks on Windows, which do not have a dataloader but use something similar as your setup, I got comparable results to ones that should be expected for my GPUs.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 755649,
          "author_name": "dongliang",
          "author_url": "",
          "post_date": "2020-02-25T01:32:14.980000",
          "content": "<p>thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755780,
          "author_name": "Shujun",
          "author_url": "",
          "post_date": "2020-02-25T05:45:52.103000",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> Interesting. It seems like my code is poorly optimized then. A resnet variant with around 30 million params on 88x88 takes 470 s per epoch on 2x2080 for me. </p>\n\n<p>BTW you wanna team up? I think we can make run for .99</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 750805,
      "author_name": "Peter",
      "author_url": "",
      "post_date": "2020-02-19T17:48:51.570000",
      "content": "<p>se-resnext50_32x4d, 128x128 (loaded from png), basic augmentation (no mixup, no cutmix); I use apex (o2), SGD, 8% validation (184772 train images), batch size: 128, no mish activation; 25892010 params</p>\n\n<p>1 epoch: 7 minutes,\n100 epochs: ~ 12 hours.</p>\n\n<p>I have a single RTX-2080ti (avg. card usage above 90%)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 750823,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-19T18:03:51.483000",
          "content": "<p>Hi <a href=\"/pestipeti\">@pestipeti</a> <br>\nDid you get better results with SGD than Adam ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750850,
          "author_name": "Peter",
          "author_url": "",
          "post_date": "2020-02-19T18:41:03.207000",
          "content": "<p>SGD + Cosine is best for me; Adam + ReduceOnPlateau is faster,  more or less the same score, but it depends on the config.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 753007,
      "author_name": "Maxime Lenormand",
      "author_url": "",
      "post_date": "2020-02-21T16:01:00.433000",
      "content": "<p>Wow</p>\n\n<p>Here I am trying to pass this huge chunk of people at 0.9663 (and then the one at 0.9696) only using Kaggle notebooks with my epoch time being less than 5 minutes.</p>\n\n<p>If anybody has a spare 2080Ti, feel free to reach out :P</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 752439,
      "author_name": "Benlei Cui",
      "author_url": "",
      "post_date": "2020-02-21T02:46:12.420000",
      "content": "<p>Dataloader uses the num_works set for 16 is  very helpful(cpu i9 9900k)\nwhich save almost half time for us.\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 2x2080ti\napex: not used\nsplit: 9/10 train, 1/10 valid\ntrain(1 epoch): 220s</p>",
      "votes": 3,
      "replies": [
        {
          "id": 755662,
          "author_name": "dongliang",
          "author_url": "",
          "post_date": "2020-02-25T01:52:03.593000",
          "content": "<p>how much memory do you hava？</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 750966,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2020-02-19T21:58:09.913000",
      "content": "<p>our current setup with 64x64 images - 2x1080ti takes about ~2 minutes per epoch.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 750975,
          "author_name": "Jie Wu",
          "author_url": "",
          "post_date": "2020-02-19T22:07:24.557000",
          "content": "<p>If I used 64x64, it went down to 2.5min per epoch on single RTX 2080Ti.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 750776,
      "author_name": "Vlad Vaduva",
      "author_url": "",
      "post_date": "2020-02-19T17:20:42.517000",
      "content": "<p>I am using a se-resnext101_32x4d arhitecture with a 128*128 size. With augmentation is taking about 30 min per epoch on a single RTX 2080Ti . In order to converge for me I have to wait much more than 40 epochs so i spend about 24 hours per fold, so 5 entire days for all 5 folds...pretty god damn long time</p>",
      "votes": 1,
      "replies": [
        {
          "id": 750781,
          "author_name": "Jie Wu",
          "author_url": "",
          "post_date": "2020-02-19T17:23:37.543000",
          "content": "<p>Seems we are the same. I also tried se-resnext101_32x4d, it's over 30min per epoch, but still running. I have not think to move to 5 folds. Because it takes so long.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 750784,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-19T17:27:58.007000",
          "content": "<p>Yes, I usually make a series of modification, train for 50 epochs on a single fold and use that point as a benchmark to compare with another architecture and different settings that I made in the past. And when I have something superior on that point I let it train the entire fold and if also the entire fold result is best that the best model i had than i let it train on the entire 5 folds</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 750956,
          "author_name": "YL",
          "author_url": "",
          "post_date": "2020-02-19T21:43:25.027000",
          "content": "<p>wow what augmentation are you using? I am using cutmix + mixup + some affine on 128x128, same model as yours, it's around 15 mins.  i have titan rtx, but its really not faster than 2080ti really. and not using half precision at the moment</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 750739,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-02-19T16:44:41.227000",
      "content": "<p>I have similar times for my 2080 and on google colab. Haven't trained for more than 100 epochs yet.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 750733,
      "author_name": "Urvish",
      "author_url": "",
      "post_date": "2020-02-19T16:41:34.987000",
      "content": "<p>I spend around 7 hours for training model for 40 epochs. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 750743,
          "author_name": "Jie Wu",
          "author_url": "",
          "post_date": "2020-02-19T16:47:48.240000",
          "content": "<p>Me too, without data aug about 7hrs</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750774,
          "author_name": "Urvish",
          "author_url": "",
          "post_date": "2020-02-19T17:18:25.013000",
          "content": "<p>I am using Data argumentation. I have different and smaller model I guess. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 751003,
      "author_name": "Appian",
      "author_url": "",
      "post_date": "2020-02-19T23:04:43.483000",
      "content": "<p><code>\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 1x1080ti\napex: not used\nsplit: 5/6 train, 1/6 valid\ntrain(1 epoch): 700s\ntrain(40 epochs): 8 hours\n</code></p>",
      "votes": 2,
      "replies": [
        {
          "id": 751141,
          "author_name": "Bet4Honor",
          "author_url": "",
          "post_date": "2020-02-20T02:50:09.073000",
          "content": "<p>I use 2x2080ti,  and use more training time about 715s one epoch.\nDo you remove the maxpooling layer?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751482,
          "author_name": "Appian",
          "author_url": "",
          "post_date": "2020-02-20T08:16:46.040000",
          "content": "<p>No, I modified the model a bit but it should not affect the training speed too much. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 750927,
      "author_name": "Jie Wu",
      "author_url": "",
      "post_date": "2020-02-19T20:42:08.363000",
      "content": "<p>Update: After some changes, on single RTX 2080Ti GPU, 90% of data for training: se-resnext50_32x4d , size 128x128 with data aug, 7min per epoch; se-resnext101_32x4d , size 128x128 with data aug, 12.5min per epoch.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 750992,
          "author_name": "Taemyung Heo",
          "author_url": "",
          "post_date": "2020-02-19T22:25:12.403000",
          "content": "<p>Can you share what some changes you've made? I have a similar situation with you..</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 756602,
      "author_name": "arutema47",
      "author_url": "",
      "post_date": "2020-02-25T22:10:37.267000",
      "content": "<p>Setting num workers to a large amount (like 8) is important if your data processing is limiting your speed!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 756493,
      "author_name": "Prajwal Prashanth",
      "author_url": "",
      "post_date": "2020-02-25T19:07:49.830000",
      "content": "<p>Image size: 137x236\nAug: simple scale, rotate, shift\nArch: resnet34\nGPU: rtx 2070 maxp\nBatch Size: 256\nnum_workers: 4\nTime/epoch: ~7.5 mins</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 755224,
      "author_name": "datta",
      "author_url": "",
      "post_date": "2020-02-24T15:12:57.830000",
      "content": "<p>So I see that not many have 2080s and these super fast training speeds. But I find plentiful comments in CV score thread, stating nice CV and LB scores. So I am wonder if everyone is really doing cross validation and averaging the scores over the folds? Or are they just doing single fold and calling it CV score generally?   </p>\n\n<p>And I am sure many are running the experiments day and night probably everyday. I hope I am not sounding negative. I just started kaggle couple of weeks before and just want to know what everyone is doing. I hope someone will enlighten me here. Thanks a lot!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 755249,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-24T15:44:28.260000",
          "content": "<p>Mix of both I'd guess. 'Best single model' technically only means one model and not an ensemble of 5 fold models. You can calculate CV of just one model too.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 755637,
          "author_name": "datta",
          "author_url": "",
          "post_date": "2020-02-25T01:10:03.610000",
          "content": "<p>True, even to calculate CV for one model, you have to train it multiple times with different splits as the validation right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755648,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-25T01:29:29.600000",
          "content": "<p>You can just train one fold and validate</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 755653,
          "author_name": "datta",
          "author_url": "",
          "post_date": "2020-02-25T01:35:08.710000",
          "content": "<p>Yup then that's not \"cross-validation\" anymore. I understand what you mean tho. Thanks for the quick reply. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 754726,
      "author_name": "Bryce1010",
      "author_url": "",
      "post_date": "2020-02-24T01:24:06.900000",
      "content": "<p>model: efficient_net_b1  </p>\n\n<p>batch_size:100  </p>\n\n<p>imgsize: 128x128  </p>\n\n<p>gpu: 1x2080ti </p>\n\n<p>apex: not used  </p>\n\n<p>mix-precision: not used  </p>\n\n<p>system: ubuntu16  </p>\n\n<p>split: 0.01 valid  </p>\n\n<p>num_workers: 0</p>\n\n<p>train(1 epoch): about 2hours  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 754962,
          "author_name": "Mishunyayev Nikita",
          "author_url": "",
          "post_date": "2020-02-24T09:19:39.403000",
          "content": "<p>Oh, it's very slow. Try to increase num_workers and batch_size</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 753727,
      "author_name": "shiba",
      "author_url": "",
      "post_date": "2020-02-22T15:49:42.767000",
      "content": "<p>maybe try Apex? and fine-tuning paras without apex</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 752627,
      "author_name": "syoya",
      "author_url": "",
      "post_date": "2020-02-21T08:23:00.323000",
      "content": "<p>Similar time. Now I'm using img 64x64 and 50 epochs training for fast experiments. This can cost around 5 hours and I can test several modifications in a single day.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 752623,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2020-02-21T08:18:20.070000",
      "content": "<p>se-resnext50_32x4d, size 128x128x3 with augmentation on a single RTX 2080Ti takes 6 min on my side</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 752330,
      "author_name": "Anthony Chan",
      "author_url": "",
      "post_date": "2020-02-20T22:07:31.593000",
      "content": "<p>Why not try mixed precision, faster with less precision reduction. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 752293,
      "author_name": "Gold Retriever",
      "author_url": "",
      "post_date": "2020-02-20T20:54:24.603000",
      "content": "<p>seresnext101\nsize 256\ngpu 2080ti\n1 epoch ~40min...</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 751471,
      "author_name": "cab",
      "author_url": "",
      "post_date": "2020-02-20T08:03:49.340000",
      "content": "<p>model: se_resnext50_32x4d\nimgsize: 137x236\ngpu: 2080Ti\napex: O1\nsplit: 80/20\ntrain(1 epoch): ~10min</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 751423,
      "author_name": "Tanyapohn",
      "author_url": "",
      "post_date": "2020-02-20T07:29:33.167000",
      "content": "<p>Anyone trains with 224 x 224? </p>\n\n<p>For my 128 x 128:\n<code>\ntraining time: ~15 mins/epoch (Depends on what you augment)\ngpu: RTX2080super\n</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 753474,
          "author_name": "Rohit Agarwal",
          "author_url": "",
          "post_date": "2020-02-22T08:31:54.833000",
          "content": "<p>My 224x224 model on google colab takes ~1 hr per epoch using densenet161.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 750993,
      "author_name": "Jo Tom",
      "author_url": "",
      "post_date": "2020-02-19T22:37:20.763000",
      "content": "<p>RTX2070 Super, ~resnext50_32x4d, size 128*128, Half Precision, 64 bs, 20% validation =&gt; 7 min per epoch</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 751620,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T10:44:29.167000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "753029": "I got **~100s** for ResNeXt50SE on 128x128 images using 2x 2080Ti + NVlink (though I'm not sure that the last one makes a difference for DL applications) with a simple augmentation like MixUp or CutMix. If one is using  2080Ti, **half precision** is the thing one nearly always must to do: it gives 1.5-2x boost. Though, APEX doesn't work well with multiple GPU, and I needed to implement my own half precision and dynamic rescaling procedures or alternatively use fast.ai. Also, make sure that you almost fully occupy the GPU RAM, though sometimes u need to train with giant batches, like 768 for the above considered case. With tiny batches the speed can degrade 1.5-2+ times. Finally, water/hybrid  cooled GPUs are the way to go: with OC ~2000MHz and 325W energy consumption it heats up only to ~65C under ~90% load. Also good CPU and SSD can be really helpful in some particular cases. And just a small final remark, do not use Windows for DL: I got ~6-8min per epoch for the same setup as I use on Linux.",
    "750730": "Can someone share their time of training a model? Now I am using se-resnext50_32x4d, size 128*128, which takes 11mins without data aug for each epoch on a single RTX 2080Ti. With data aug, it takes 25-26min for each epoch. Train one model (single fold) will take 16-17 hours for just 40 epochs, this is very long. Just wondering those who trained their models on a 4-5 folds, with 80 epochs or 150 epochs, how long did they take?",
    "750805": "se-resnext50_32x4d, 128x128 (loaded from png), basic augmentation (no mixup, no cutmix); I use apex (o2), SGD, 8% validation (184772 train images), batch size: 128, no mish activation; 25892010 params\n\n1 epoch: 7 minutes,\n100 epochs: ~ 12 hours.\n\nI have a single RTX-2080ti (avg. card usage above 90%)",
    "753007": "Wow\n\nHere I am trying to pass this huge chunk of people at 0.9663 (and then the one at 0.9696) only using Kaggle notebooks with my epoch time being less than 5 minutes.\n\nIf anybody has a spare 2080Ti, feel free to reach out :P",
    "752439": "Dataloader uses the num_works set for 16 is  very helpful(cpu i9 9900k)\nwhich save almost half time for us.\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 2x2080ti\napex: not used\nsplit: 9/10 train, 1/10 valid\ntrain(1 epoch): 220s",
    "750966": "our current setup with 64x64 images - 2x1080ti takes about ~2 minutes per epoch.",
    "750776": "I am using a se-resnext101_32x4d arhitecture with a 128*128 size. With augmentation is taking about 30 min per epoch on a single RTX 2080Ti . In order to converge for me I have to wait much more than 40 epochs so i spend about 24 hours per fold, so 5 entire days for all 5 folds...pretty god damn long time",
    "750739": "I have similar times for my 2080 and on google colab. Haven't trained for more than 100 epochs yet.",
    "750733": "I spend around 7 hours for training model for 40 epochs. ",
    "751003": "```\nmodel: se_resnext50_32x4d\nimgsize: 128x128\ngpu: 1x1080ti\napex: not used\nsplit: 5/6 train, 1/6 valid\ntrain(1 epoch): 700s\ntrain(40 epochs): 8 hours\n```",
    "750927": "Update: After some changes, on single RTX 2080Ti GPU, 90% of data for training: se-resnext50_32x4d , size 128x128 with data aug, 7min per epoch; se-resnext101_32x4d , size 128x128 with data aug, 12.5min per epoch.",
    "756602": "Setting num workers to a large amount (like 8) is important if your data processing is limiting your speed!",
    "756493": "Image size: 137x236\nAug: simple scale, rotate, shift\nArch: resnet34\nGPU: rtx 2070 maxp\nBatch Size: 256\nnum_workers: 4\nTime/epoch: ~7.5 mins",
    "755224": "So I see that not many have 2080s and these super fast training speeds. But I find plentiful comments in CV score thread, stating nice CV and LB scores. So I am wonder if everyone is really doing cross validation and averaging the scores over the folds? Or are they just doing single fold and calling it CV score generally?   \n\nAnd I am sure many are running the experiments day and night probably everyday. I hope I am not sounding negative. I just started kaggle couple of weeks before and just want to know what everyone is doing. I hope someone will enlighten me here. Thanks a lot!",
    "754726": "\nmodel: efficient_net_b1  \n\nbatch_size:100  \n\nimgsize: 128x128  \n\ngpu: 1x2080ti \n \napex: not used  \n\nmix-precision: not used  \n\nsystem: ubuntu16  \n\nsplit: 0.01 valid  \n\nnum_workers: 0\n\ntrain(1 epoch): about 2hours  \n\n\n",
    "753727": "maybe try Apex? and fine-tuning paras without apex",
    "752627": "Similar time. Now I'm using img 64x64 and 50 epochs training for fast experiments. This can cost around 5 hours and I can test several modifications in a single day.",
    "752623": "se-resnext50_32x4d, size 128x128x3 with augmentation on a single RTX 2080Ti takes 6 min on my side",
    "752330": "Why not try mixed precision, faster with less precision reduction. ",
    "752293": "seresnext101\nsize 256\ngpu 2080ti\n1 epoch ~40min...",
    "751471": "model: se_resnext50_32x4d\nimgsize: 137x236\ngpu: 2080Ti\napex: O1\nsplit: 80/20\ntrain(1 epoch): ~10min",
    "751423": "Anyone trains with 224 x 224? \n\nFor my 128 x 128:\n```\ntraining time: ~15 mins/epoch (Depends on what you augment)\ngpu: RTX2080super\n```",
    "750993": "RTX2070 Super, ~resnext50_32x4d, size 128*128, Half Precision, 64 bs, 20% validation =&gt; 7 min per epoch",
    "751620": ""
  }
}