{
  "id": 173272,
  "title": "Pytorch user, try resnest instead of efficientnet",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173272",
  "author_name": "Arnaud Roussel",
  "post_date": "2020-08-08T14:54:01.544000",
  "votes": 25,
  "comment_count": 35,
  "views": 0,
  "content": "<p>It's a known issue with efficient nets that it is usually slower on Pytorch. Afaik depthwise convolution are really poorly optimized in Pytorch as can be seen in some tickets both on the efficientnet repo: <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/issues/159\" target=\"_blank\">https://github.com/lukemelas/EfficientNet-PyTorch/issues/159</a><br>\nand pytorch: <a href=\"https://github.com/pytorch/pytorch/issues/18631\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/18631</a></p>\n<p>Thankfully there is a recent paper proposing ResNest <a href=\"https://arxiv.org/pdf/2004.08955.pdf\" target=\"_blank\">https://arxiv.org/pdf/2004.08955.pdf</a> and a pytorch implementation here <a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">https://github.com/zhanghang1989/ResNeSt</a>. For example for me a B2 and a Resnest50 train in approx the same time but the resnest50 showed a good CV/LB increase since it's a bigger model at 25M instead of the 7M for the B2.</p>\n<p>Some numbers on my gtx1080ti:<br>\nB2 7M 7:30 per epoch<br>\nResnest50 25M 6:45 per epoch<br>\nB4 19M 20:00 per epoch<br>\nResnest100 46M 14:00 per epoch</p>\n<p>All this is of course slower than TF + TPU :)</p>",
  "messages": [
    {
      "id": 962929,
      "postDate": "2020-08-08T14:54:01.543Z",
      "content": "<p>It's a known issue with efficient nets that it is usually slower on Pytorch. Afaik depthwise convolution are really poorly optimized in Pytorch as can be seen in some tickets both on the efficientnet repo: <a href=\"https://github.com/lukemelas/EfficientNet-PyTorch/issues/159\" target=\"_blank\">https://github.com/lukemelas/EfficientNet-PyTorch/issues/159</a><br>\nand pytorch: <a href=\"https://github.com/pytorch/pytorch/issues/18631\" target=\"_blank\">https://github.com/pytorch/pytorch/issues/18631</a></p>\n<p>Thankfully there is a recent paper proposing ResNest <a href=\"https://arxiv.org/pdf/2004.08955.pdf\" target=\"_blank\">https://arxiv.org/pdf/2004.08955.pdf</a> and a pytorch implementation here <a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">https://github.com/zhanghang1989/ResNeSt</a>. For example for me a B2 and a Resnest50 train in approx the same time but the resnest50 showed a good CV/LB increase since it's a bigger model at 25M instead of the 7M for the B2.</p>\n<p>Some numbers on my gtx1080ti:<br>\nB2 7M 7:30 per epoch<br>\nResnest50 25M 6:45 per epoch<br>\nB4 19M 20:00 per epoch<br>\nResnest100 46M 14:00 per epoch</p>\n<p>All this is of course slower than TF + TPU :)</p>",
      "rawMarkdown": "It's a known issue with efficient nets that it is usually slower on Pytorch. Afaik depthwise convolution are really poorly optimized in Pytorch as can be seen in some tickets both on the efficientnet repo: https://github.com/lukemelas/EfficientNet-PyTorch/issues/159\nand pytorch: https://github.com/pytorch/pytorch/issues/18631\n\nThankfully there is a recent paper proposing ResNest https://arxiv.org/pdf/2004.08955.pdf and a pytorch implementation here https://github.com/zhanghang1989/ResNeSt. For example for me a B2 and a Resnest50 train in approx the same time but the resnest50 showed a good CV/LB increase since it's a bigger model at 25M instead of the 7M for the B2.\n\nSome numbers on my gtx1080ti:\nB2 7M 7:30 per epoch\nResnest50 25M 6:45 per epoch\nB4 19M 20:00 per epoch\nResnest100 46M 14:00 per epoch\n\nAll this is of course slower than TF + TPU :)",
      "votes": 25
    },
    {
      "id": 964937,
      "postDate": "2020-08-10T09:25:41.800Z",
      "content": "<p>First try isn't good, I guess I have to retune everything.</p>",
      "rawMarkdown": "First try isn't good, I guess I have to retune everything.",
      "votes": 3
    },
    {
      "id": 963322,
      "postDate": "2020-08-08T22:32:50.793Z",
      "content": "<blockquote>\n  <p>All this is of course slower than TF + TPU :)</p>\n</blockquote>\n<p>It's a little bit sad that Pytorch doesn't have a proper way right now to load data directly on TPUs( need to use CPU -&gt; Image data preprocessing is extremely slow at kaggle. So XLA could be really used at kaggle only for NLP, and even their 16 GB RAM limitation makes it difficult to deal with. While TF is not flexible, and it is difficult to add even simple things.</p>",
      "rawMarkdown": "&gt; All this is of course slower than TF + TPU :)\n\nIt's a little bit sad that Pytorch doesn't have a proper way right now to load data directly on TPUs( need to use CPU -&gt; Image data preprocessing is extremely slow at kaggle. So XLA could be really used at kaggle only for NLP, and even their 16 GB RAM limitation makes it difficult to deal with. While TF is not flexible, and it is difficult to add even simple things.\n",
      "votes": 3,
      "replies": [
        {
          "id": 963334,
          "postDate": "2020-08-08T23:03:00.687Z",
          "content": "<p>Yes, exactly my thoughts. If at least there was an equivalent to fast I/O + processing I'd gladly invest modifying my code to use XLA (or use pytorch lightning as I do locally but usually lightning for Kaggle TPU often fails for me).</p>",
          "rawMarkdown": "Yes, exactly my thoughts. If at least there was an equivalent to fast I/O + processing I'd gladly invest modifying my code to use XLA (or use pytorch lightning as I do locally but usually lightning for Kaggle TPU often fails for me).",
          "votes": 1
        },
        {
          "id": 963342,
          "postDate": "2020-08-08T23:33:27.493Z",
          "content": "<p>I/O + processing is clearly an issue, however Colab Pro has much better config for Pytorch TPU than Kaggle. More CPU cores and +35 GB of RAM. <br>\nAnd you can run it h24</p>",
          "rawMarkdown": " I/O + processing is clearly an issue, however Colab Pro has much better config for Pytorch TPU than Kaggle. More CPU cores and +35 GB of RAM. \nAnd you can run it h24",
          "votes": 2
        },
        {
          "id": 963999,
          "postDate": "2020-08-09T14:00:45.160Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> is it available in France? So far only available in the US, right?</p>",
          "rawMarkdown": "@serigne is it available in France? So far only available in the US, right?"
        },
        {
          "id": 964031,
          "postDate": "2020-08-09T14:28:48.907Z",
          "content": "<p>Yeah it's available.  Don't worry about the US zip code, you can put yours. </p>",
          "rawMarkdown": "Yeah it's available.  Don't worry about the US zip code, you can put yours. ",
          "votes": 2
        },
        {
          "id": 964273,
          "postDate": "2020-08-09T17:55:04.747Z",
          "content": "<p>Alright thanks!</p>",
          "rawMarkdown": "Alright thanks!"
        },
        {
          "id": 965539,
          "postDate": "2020-08-10T17:47:22.077Z",
          "content": "<p><a href=\"/serigne\">@serigne</a> Thanks for the colab pro tip. Didn't know that service existed. Price is pretty good too. Using it know for bigger model.</p>",
          "rawMarkdown": "@serigne Thanks for the colab pro tip. Didn't know that service existed. Price is pretty good too. Using it know for bigger model."
        },
        {
          "id": 966057,
          "postDate": "2020-08-11T05:23:10.470Z",
          "content": "<p>De rien <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> </p>\n<p>Note that on TPU , you can execute a notebook, close the browser and il will be still running. You can open it hours later elsewhere  and find all the training log.  </p>\n<p>This doesn't work for GPU though. It would disconnect after about 1 hour of no interactive session.</p>",
          "rawMarkdown": "De rien @arroqc \n\nNote that on TPU , you can execute a notebook, close the browser and il will be still running. You can open it hours later elsewhere  and find all the training log.  \n\nThis doesn't work for GPU though. It would disconnect after about 1 hour of no interactive session."
        },
        {
          "id": 972018,
          "postDate": "2020-08-16T07:02:18.087Z",
          "content": "<p>Have you tried NVIDIA-DALI which uses GPU for image augmentations and loads images directly on to GPU which makes data loading much faster as compared to the traditional approach.</p>",
          "rawMarkdown": "Have you tried NVIDIA-DALI which uses GPU for image augmentations and loads images directly on to GPU which makes data loading much faster as compared to the traditional approach.",
          "votes": 5,
          "replies": [
            {
              "id": 972062,
              "postDate": "2020-08-16T07:43:10.140Z",
              "content": "<p><a href=\"https://www.kaggle.com/chandanverma\" target=\"_blank\">@chandanverma</a> I personally have not tried DALI, but I am going to look into it because that tech sounds interesting.  However, my instinct is, I am not sure its a good thing for my current pipeline.  GPU memory is a premium for me.  I don’t have any GPU memory available to load the next batch and be doing pre-processing while the GPU is doing training.  With CPU based augmentations, they are actually super quick,and they use CPU memory.  I use pin_memory and a async copies into GPU memory, so my GPU is at 100%.  If your GPU is at 100% then its not waiting on you and I can’t imagine GPU augmentations being better.  In short, I think CPU + GPU is a powerful combination.  For just straight up doing pre-processing, but not online pre-processing, then I think DALI sounds great.  Either way, I am going to look into it.</p>",
              "rawMarkdown": "@chandanverma I personally have not tried DALI, but I am going to look into it because that tech sounds interesting.  However, my instinct is, I am not sure its a good thing for my current pipeline.  GPU memory is a premium for me.  I don’t have any GPU memory available to load the next batch and be doing pre-processing while the GPU is doing training.  With CPU based augmentations, they are actually super quick,and they use CPU memory.  I use pin_memory and a async copies into GPU memory, so my GPU is at 100%.  If your GPU is at 100% then its not waiting on you and I can’t imagine GPU augmentations being better.  In short, I think CPU + GPU is a powerful combination.  For just straight up doing pre-processing, but not online pre-processing, then I think DALI sounds great.  Either way, I am going to look into it.",
              "votes": 2
            },
            {
              "id": 972363,
              "postDate": "2020-08-16T13:43:54.437Z",
              "content": "<p>i tried it in deepfake competition but no success because of memory leak.  </p>",
              "rawMarkdown": "i tried it in deepfake competition but no success because of memory leak.  ",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 962954,
      "postDate": "2020-08-08T15:10:17.300Z",
      "content": "<p>Resnest are fast indeed and give good results :)</p>\n<p>You have even faster versions of Resnest50 from the hub</p>\n<pre><code>'resnest50_fast_1s1x64d',   'resnest50_fast_1s2x40d',\n'resnest50_fast_1s4x24d',  'resnest50_fast_2s1x64d','\nresnest50_fast_2s2x40d',  'resnest50_fast_4s1x64d',\n'resnest50_fast_4s2x40d'\n</code></pre>\n<p>For efficientnet , I would recommend using <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> by <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a>  instead of  this repo</p>",
      "rawMarkdown": "Resnest are fast indeed and give good results :)\n\nYou have even faster versions of Resnest50 from the hub\n```\n'resnest50_fast_1s1x64d',   'resnest50_fast_1s2x40d',\n'resnest50_fast_1s4x24d',  'resnest50_fast_2s1x64d','\nresnest50_fast_2s2x40d',  'resnest50_fast_4s1x64d',\n'resnest50_fast_4s2x40d'\n``` \n\n\nFor efficientnet , I would recommend using [timm](https://github.com/rwightman/pytorch-image-models) by @rwightman  instead of  this repo",
      "votes": 4,
      "replies": [
        {
          "id": 962979,
          "postDate": "2020-08-08T15:23:31.310Z",
          "content": "<p>Thanks for the tip. Do the fast version deliver the same performance or is there a trade-off in your experience ? The paper suggest lower performance on theirbenchmark but I was wondering what you think based on experience.</p>\n<p>Will give the efficientnet link a shot :)</p>",
          "rawMarkdown": "Thanks for the tip. Do the fast version deliver the same performance or is there a trade-off in your experience ? The paper suggest lower performance on theirbenchmark but I was wondering what you think based on experience.\n\nWill give the efficientnet link a shot :)",
          "votes": 2
        },
        {
          "id": 962996,
          "postDate": "2020-08-08T15:32:03.433Z",
          "content": "<p>I tried just the first one,  it run slightly faster  with similar result. </p>\n<p>For <code>timm</code>, you have much more models and weignts than Pytorch-Efficientnet repo. For instance I could get LB 0.952 with B5 noisy Student on <code>timm</code></p>",
          "rawMarkdown": "I tried just the first one,  it run slightly faster  with similar result. \n\nFor ` timm`, you have much more models and weignts than Pytorch-Efficientnet repo. For instance I could get LB 0.952 with B5 noisy Student on `timm`",
          "votes": 4
        },
        {
          "id": 963008,
          "postDate": "2020-08-08T15:36:38.310Z",
          "content": "<p>Does it run faster or it's just additional weights ? My main issue with B5 is that it takes ages to train locally for me.</p>",
          "rawMarkdown": "Does it run faster or it's just additional weights ? My main issue with B5 is that it takes ages to train locally for me.",
          "votes": 1,
          "replies": [
            {
              "id": 963016,
              "postDate": "2020-08-08T15:44:17.907Z",
              "content": "<blockquote>\n  <p>Does it run faster or it's just additional weights ? My main issue with B5 is that it takes ages to train locally for me.</p>\n</blockquote>\n<p>For timm ?  my recommendation was more based on options and flexibility but I didn't really compare speeds. </p>\n<p>For these big models I use either TPU or Multi-GPU. </p>",
              "rawMarkdown": "&gt; Does it run faster or it's just additional weights ? My main issue with B5 is that it takes ages to train locally for me.\n\nFor timm ?  my recommendation was more based on options and flexibility but I didn't really compare speeds. \n\nFor these big models I use either TPU or Multi-GPU. ",
              "votes": 1
            },
            {
              "id": 963454,
              "postDate": "2020-08-09T03:44:29.520Z",
              "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Hi Serigne may I know how is your TPU usage % when running the pytorch model? <br>\nDid you apply distributed parallel loader and multiprocessing.spawn as well?</p>\n<p>I tried but feel frustrated and at the end switch to tensorflow</p>\n<p>thanks</p>",
              "rawMarkdown": "@serigne Hi Serigne may I know how is your TPU usage % when running the pytorch model? \nDid you apply distributed parallel loader and multiprocessing.spawn as well?\n\nI tried but feel frustrated and at the end switch to tensorflow\n\nthanks"
            },
            {
              "id": 963627,
              "postDate": "2020-08-09T07:20:47.333Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 970243,
      "postDate": "2020-08-14T09:40:57.917Z",
      "content": "<p><code>Resnest50</code> converges much faster than Efficientnet and is less prone to noise. But weirdly, I'm getting large CV and LB gap using this model. I'm using <code>resnest50d_1s4x24d</code>. One of my experiments yielded <code>0.9254</code> on CV and <code>0.903</code> on LB. 😨</p>",
      "rawMarkdown": "`Resnest50` converges much faster than Efficientnet and is less prone to noise. But weirdly, I'm getting large CV and LB gap using this model. I'm using `resnest50d_1s4x24d`. One of my experiments yielded `0.9254` on CV and `0.903` on LB. 😨",
      "votes": 1
    },
    {
      "id": 965368,
      "postDate": "2020-08-10T15:26:17.877Z",
      "content": "<p>It's true my intial kernels in the beginning with resnet50 with 5 folds 128 image size gave a socre of 0.91.<br>\nEfficientNet in pytorch does run slow. Given Lower score.</p>",
      "rawMarkdown": "It's true my intial kernels in the beginning with resnet50 with 5 folds 128 image size gave a socre of 0.91.\nEfficientNet in pytorch does run slow. Given Lower score.",
      "votes": 1
    },
    {
      "id": 964305,
      "postDate": "2020-08-09T18:22:37.213Z",
      "content": "<p>Thanks, I think it explains some of my problems with pytorch in this competition.</p>",
      "rawMarkdown": "Thanks, I think it explains some of my problems with pytorch in this competition.",
      "votes": 1
    },
    {
      "id": 963867,
      "postDate": "2020-08-09T11:24:24.457Z",
      "content": "<p>For <code>TF/Keras</code> user, : <a href=\"https://github.com/QiaoranC/tf_ResNeSt_RegNet_model\" target=\"_blank\">ResNeSt</a> , TF version.</p>\n<p>But note, no pre-trained weights for <code>tf</code> version yet, one needs to <a href=\"https://github.com/nerox8664/pytorch2keras\" target=\"_blank\">convert</a> it to try. </p>",
      "rawMarkdown": "For `TF/Keras` user, : [ResNeSt](https://github.com/QiaoranC/tf_ResNeSt_RegNet_model) , TF version.\n\nBut note, no pre-trained weights for `tf` version yet, one needs to [convert](https://github.com/nerox8664/pytorch2keras) it to try. ",
      "votes": 1,
      "replies": [
        {
          "id": 965951,
          "postDate": "2020-08-11T02:16:37.053Z",
          "content": "<p>I have a converted version <a href=\"https://github.com/RichardXiao13/TensorFlow-ResNets\">here</a> but the CV LB gap was large.</p>",
          "rawMarkdown": "I have a converted version [here](https://github.com/RichardXiao13/TensorFlow-ResNets) but the CV LB gap was large.",
          "votes": 2
        },
        {
          "id": 972058,
          "postDate": "2020-08-16T07:41:15.700Z",
          "content": "<p><a href=\"https://www.kaggle.com/richardxiao03\" target=\"_blank\">@richardxiao03</a> <br>\nAmazing. Thanks for sharing brother. Starred. -)</p>",
          "rawMarkdown": "@richardxiao03 \nAmazing. Thanks for sharing brother. Starred. -)"
        }
      ]
    },
    {
      "id": 963618,
      "postDate": "2020-08-09T07:14:31.447Z",
      "content": "<p>I came across this issue a few months ago with pytorch, it's crazy they still didn't optimize it.</p>",
      "rawMarkdown": "I came across this issue a few months ago with pytorch, it's crazy they still didn't optimize it.",
      "votes": 1
    },
    {
      "id": 963509,
      "postDate": "2020-08-09T04:55:16.817Z",
      "content": "<p>I am trying resnext50 from <a href=\"https://pytorch.org/hub/facebookresearch_semi-supervised-ImageNet1K-models_resnext/\">here</a>. It can easily give .93LB on images of size 256. They are faster than resnest too. </p>",
      "rawMarkdown": "I am trying resnext50 from [here](https://pytorch.org/hub/facebookresearch_semi-supervised-ImageNet1K-models_resnext/). It can easily give .93LB on images of size 256. They are faster than resnest too. ",
      "votes": 1
    },
    {
      "id": 965773,
      "postDate": "2020-08-10T20:58:59.317Z",
      "content": "<p>I've just submitted my <code>resnest50_fast_1s1x64d</code> model trained on 384*384</p>\n<p>LB 0.946 : 5 folds  + 15 TTA for each fold  inference </p>\n<p>It run really Fast on TPU (with big batch size) </p>",
      "rawMarkdown": "I've just submitted my `resnest50_fast_1s1x64d` model trained on 384*384\n\nLB 0.946 : 5 folds  + 15 TTA for each fold  inference \n\nIt run really Fast on TPU (with big batch size) ",
      "votes": 2,
      "replies": [
        {
          "id": 965838,
          "postDate": "2020-08-10T22:32:17.110Z",
          "content": "<p>I'm getting similar resullts. Just started with google colab pro. Currently running one training through a GPU. Will try TPU after and hope pytorch lightning is able to seamlessly make the transition.</p>",
          "rawMarkdown": "I'm getting similar resullts. Just started with google colab pro. Currently running one training through a GPU. Will try TPU after and hope pytorch lightning is able to seamlessly make the transition."
        },
        {
          "id": 965839,
          "postDate": "2020-08-10T22:36:36.847Z",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> what is your CV for this 0.946 sub?</p>",
          "rawMarkdown": "@serigne what is your CV for this 0.946 sub?"
        },
        {
          "id": 966050,
          "postDate": "2020-08-11T05:11:39.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  OOF (with 5 TTA per fold) = 0.930</p>",
          "rawMarkdown": "@cpmpml  OOF (with 5 TTA per fold) = 0.930",
          "votes": 1
        },
        {
          "id": 966147,
          "postDate": "2020-08-11T07:24:10.347Z",
          "content": "<p>Can you share information about the package information from where you are using 'resnest50_fast_1s1x64d'. </p>",
          "rawMarkdown": "Can you share information about the package information from where you are using 'resnest50_fast_1s1x64d'. "
        },
        {
          "id": 966497,
          "postDate": "2020-08-11T13:13:58.023Z",
          "content": "<p>I use the <a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">official repo</a></p>",
          "rawMarkdown": "I use the [official repo](https://github.com/zhanghang1989/ResNeSt)\n\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 966496,
      "postDate": "2020-08-11T13:13:39.147Z",
      "content": "<p>All you all using the same head/classifier with these different models, that you were using with EFnets?</p>\n\n<p>Right now my head is pretty basic for EFnets, any advice for improving it?</p>\n\n<pre><code>    self.meta_fc1 = nn.Linear(n_meta_features, 4)\n    self.meta_bn0 = nn.BatchNorm1d(4)\n    self.combined_fc0 = nn.Linear(num_ftrs + 4, 1)\n    self.avgpool0 = nn.AdaptiveAvgPool2d((1,1))\n    self.dropout0 = nn.Dropout(p=0.2)\n\n    self.dense = nn.Linear(num_ftrs, 1)\n\ndef forward(self, inputs):\n\n    # Image + Meta\n    image, meta = inputs\n    cnn_features = self.arch.extract_features(image)         \n    cnn_features = self.avgpool0(cnn_features) \n    cnn_features = cnn_features.flatten(start_dim=1)\n    meta_features = self.meta_fc1(meta)\n    meta_features = F.selu(meta_features)\n    meta_features = self.meta_bn0(meta_features)\n    features = torch.cat((cnn_features, meta_features), dim=1)\n    output = self.combined_fc0(features)\n    output = torch.sigmoid(output)\n    return output  \n</code></pre>",
      "rawMarkdown": "All you all using the same head/classifier with these different models, that you were using with EFnets?\n\nRight now my head is pretty basic for EFnets, any advice for improving it?\n\n        self.meta_fc1 = nn.Linear(n_meta_features, 4)\n        self.meta_bn0 = nn.BatchNorm1d(4)\n        self.combined_fc0 = nn.Linear(num_ftrs + 4, 1)\n        self.avgpool0 = nn.AdaptiveAvgPool2d((1,1))\n        self.dropout0 = nn.Dropout(p=0.2)\n        \n        self.dense = nn.Linear(num_ftrs, 1)\n                                        \n    def forward(self, inputs):\n\n        # Image + Meta\n        image, meta = inputs\n        cnn_features = self.arch.extract_features(image)         \n        cnn_features = self.avgpool0(cnn_features) \n        cnn_features = cnn_features.flatten(start_dim=1)\n        meta_features = self.meta_fc1(meta)\n        meta_features = F.selu(meta_features)\n        meta_features = self.meta_bn0(meta_features)\n        features = torch.cat((cnn_features, meta_features), dim=1)\n        output = self.combined_fc0(features)\n        output = torch.sigmoid(output)\n        return output  \n     ",
      "replies": [
        {
          "id": 972219,
          "postDate": "2020-08-16T10:50:46.880Z",
          "content": "<p>Hi, you are new to Kaggle hence I feel I should give you a little friendly advice.  It is considered bad practice to share specific information during the last week of competition.  Why?  Because in some cases, people shared very valuable info on the last day and it blew up the leaderboard.  Sharing at least a week in advance makes sure every participant has a fair chance to leverage what has been shared.</p>\n<p>I am not saying you are giving away very valuable info, but you might.</p>\n<p>Disclaimer: I don't work at Kaggle nor am I tied to Kaggle in anyway, you can freely ignore my friendly advice ;)</p>",
          "rawMarkdown": "Hi, you are new to Kaggle hence I feel I should give you a little friendly advice.  It is considered bad practice to share specific information during the last week of competition.  Why?  Because in some cases, people shared very valuable info on the last day and it blew up the leaderboard.  Sharing at least a week in advance makes sure every participant has a fair chance to leverage what has been shared.\n\nI am not saying you are giving away very valuable info, but you might.\n\nDisclaimer: I don't work at Kaggle nor am I tied to Kaggle in anyway, you can freely ignore my friendly advice ;)\n",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 964937,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-10T09:25:41.800000",
      "content": "<p>First try isn't good, I guess I have to retune everything.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 963322,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2020-08-08T22:32:50.793000",
      "content": "<blockquote>\n  <p>All this is of course slower than TF + TPU :)</p>\n</blockquote>\n<p>It's a little bit sad that Pytorch doesn't have a proper way right now to load data directly on TPUs( need to use CPU -&gt; Image data preprocessing is extremely slow at kaggle. So XLA could be really used at kaggle only for NLP, and even their 16 GB RAM limitation makes it difficult to deal with. While TF is not flexible, and it is difficult to add even simple things.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 963334,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-08T23:03:00.687000",
          "content": "<p>Yes, exactly my thoughts. If at least there was an equivalent to fast I/O + processing I'd gladly invest modifying my code to use XLA (or use pytorch lightning as I do locally but usually lightning for Kaggle TPU often fails for me).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 963342,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-08T23:33:27.493000",
          "content": "<p>I/O + processing is clearly an issue, however Colab Pro has much better config for Pytorch TPU than Kaggle. More CPU cores and +35 GB of RAM. <br>\nAnd you can run it h24</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 963999,
          "author_name": "PAB97",
          "author_url": "",
          "post_date": "2020-08-09T14:00:45.160000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> is it available in France? So far only available in the US, right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 964031,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-09T14:28:48.907000",
          "content": "<p>Yeah it's available.  Don't worry about the US zip code, you can put yours. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 964273,
          "author_name": "PAB97",
          "author_url": "",
          "post_date": "2020-08-09T17:55:04.747000",
          "content": "<p>Alright thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 965539,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-10T17:47:22.077000",
          "content": "<p><a href=\"/serigne\">@serigne</a> Thanks for the colab pro tip. Didn't know that service existed. Price is pretty good too. Using it know for bigger model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966057,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-11T05:23:10.470000",
          "content": "<p>De rien <a href=\"https://www.kaggle.com/arroqc\" target=\"_blank\">@arroqc</a> </p>\n<p>Note that on TPU , you can execute a notebook, close the browser and il will be still running. You can open it hours later elsewhere  and find all the training log.  </p>\n<p>This doesn't work for GPU though. It would disconnect after about 1 hour of no interactive session.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 972018,
          "author_name": "Chandan Verma",
          "author_url": "",
          "post_date": "2020-08-16T07:02:18.087000",
          "content": "<p>Have you tried NVIDIA-DALI which uses GPU for image augmentations and loads images directly on to GPU which makes data loading much faster as compared to the traditional approach.</p>",
          "votes": 5,
          "replies": [
            {
              "id": 972062,
              "author_name": "Signal",
              "author_url": "",
              "post_date": "2020-08-16T07:43:10.140000",
              "content": "<p><a href=\"https://www.kaggle.com/chandanverma\" target=\"_blank\">@chandanverma</a> I personally have not tried DALI, but I am going to look into it because that tech sounds interesting.  However, my instinct is, I am not sure its a good thing for my current pipeline.  GPU memory is a premium for me.  I don’t have any GPU memory available to load the next batch and be doing pre-processing while the GPU is doing training.  With CPU based augmentations, they are actually super quick,and they use CPU memory.  I use pin_memory and a async copies into GPU memory, so my GPU is at 100%.  If your GPU is at 100% then its not waiting on you and I can’t imagine GPU augmentations being better.  In short, I think CPU + GPU is a powerful combination.  For just straight up doing pre-processing, but not online pre-processing, then I think DALI sounds great.  Either way, I am going to look into it.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 972363,
              "author_name": "yimacs",
              "author_url": "",
              "post_date": "2020-08-16T13:43:54.437000",
              "content": "<p>i tried it in deepfake competition but no success because of memory leak.  </p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 962954,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-08-08T15:10:17.300000",
      "content": "<p>Resnest are fast indeed and give good results :)</p>\n<p>You have even faster versions of Resnest50 from the hub</p>\n<pre><code>'resnest50_fast_1s1x64d',   'resnest50_fast_1s2x40d',\n'resnest50_fast_1s4x24d',  'resnest50_fast_2s1x64d','\nresnest50_fast_2s2x40d',  'resnest50_fast_4s1x64d',\n'resnest50_fast_4s2x40d'\n</code></pre>\n<p>For efficientnet , I would recommend using <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> by <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a>  instead of  this repo</p>",
      "votes": 4,
      "replies": [
        {
          "id": 962979,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-08T15:23:31.310000",
          "content": "<p>Thanks for the tip. Do the fast version deliver the same performance or is there a trade-off in your experience ? The paper suggest lower performance on theirbenchmark but I was wondering what you think based on experience.</p>\n<p>Will give the efficientnet link a shot :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 962996,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-08T15:32:03.433000",
          "content": "<p>I tried just the first one,  it run slightly faster  with similar result. </p>\n<p>For <code>timm</code>, you have much more models and weignts than Pytorch-Efficientnet repo. For instance I could get LB 0.952 with B5 noisy Student on <code>timm</code></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 963008,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-08T15:36:38.310000",
          "content": "<p>Does it run faster or it's just additional weights ? My main issue with B5 is that it takes ages to train locally for me.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 963016,
              "author_name": "Serigne ",
              "author_url": "",
              "post_date": "2020-08-08T15:44:17.907000",
              "content": "<blockquote>\n  <p>Does it run faster or it's just additional weights ? My main issue with B5 is that it takes ages to train locally for me.</p>\n</blockquote>\n<p>For timm ?  my recommendation was more based on options and flexibility but I didn't really compare speeds. </p>\n<p>For these big models I use either TPU or Multi-GPU. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 963454,
              "author_name": "Salaryman",
              "author_url": "",
              "post_date": "2020-08-09T03:44:29.520000",
              "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> Hi Serigne may I know how is your TPU usage % when running the pytorch model? <br>\nDid you apply distributed parallel loader and multiprocessing.spawn as well?</p>\n<p>I tried but feel frustrated and at the end switch to tensorflow</p>\n<p>thanks</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 963627,
              "author_name": "",
              "author_url": "",
              "post_date": "2020-08-09T07:20:47.333000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 970243,
      "author_name": "Tahsin Mostafiz",
      "author_url": "",
      "post_date": "2020-08-14T09:40:57.917000",
      "content": "<p><code>Resnest50</code> converges much faster than Efficientnet and is less prone to noise. But weirdly, I'm getting large CV and LB gap using this model. I'm using <code>resnest50d_1s4x24d</code>. One of my experiments yielded <code>0.9254</code> on CV and <code>0.903</code> on LB. 😨</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 965368,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-10T15:26:17.877000",
      "content": "<p>It's true my intial kernels in the beginning with resnet50 with 5 folds 128 image size gave a socre of 0.91.<br>\nEfficientNet in pytorch does run slow. Given Lower score.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 964305,
      "author_name": "Jacek Poplawski",
      "author_url": "",
      "post_date": "2020-08-09T18:22:37.213000",
      "content": "<p>Thanks, I think it explains some of my problems with pytorch in this competition.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 963867,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-08-09T11:24:24.457000",
      "content": "<p>For <code>TF/Keras</code> user, : <a href=\"https://github.com/QiaoranC/tf_ResNeSt_RegNet_model\" target=\"_blank\">ResNeSt</a> , TF version.</p>\n<p>But note, no pre-trained weights for <code>tf</code> version yet, one needs to <a href=\"https://github.com/nerox8664/pytorch2keras\" target=\"_blank\">convert</a> it to try. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 965951,
          "author_name": "Richard Xiao",
          "author_url": "",
          "post_date": "2020-08-11T02:16:37.053000",
          "content": "<p>I have a converted version <a href=\"https://github.com/RichardXiao13/TensorFlow-ResNets\">here</a> but the CV LB gap was large.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 972058,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-08-16T07:41:15.700000",
          "content": "<p><a href=\"https://www.kaggle.com/richardxiao03\" target=\"_blank\">@richardxiao03</a> <br>\nAmazing. Thanks for sharing brother. Starred. -)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 963618,
      "author_name": "Krisztián Fekete",
      "author_url": "",
      "post_date": "2020-08-09T07:14:31.447000",
      "content": "<p>I came across this issue a few months ago with pytorch, it's crazy they still didn't optimize it.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 963509,
      "author_name": "Vishnu Subramanian",
      "author_url": "",
      "post_date": "2020-08-09T04:55:16.817000",
      "content": "<p>I am trying resnext50 from <a href=\"https://pytorch.org/hub/facebookresearch_semi-supervised-ImageNet1K-models_resnext/\">here</a>. It can easily give .93LB on images of size 256. They are faster than resnest too. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 965773,
      "author_name": "Serigne ",
      "author_url": "",
      "post_date": "2020-08-10T20:58:59.317000",
      "content": "<p>I've just submitted my <code>resnest50_fast_1s1x64d</code> model trained on 384*384</p>\n<p>LB 0.946 : 5 folds  + 15 TTA for each fold  inference </p>\n<p>It run really Fast on TPU (with big batch size) </p>",
      "votes": 2,
      "replies": [
        {
          "id": 965838,
          "author_name": "Arnaud Roussel",
          "author_url": "",
          "post_date": "2020-08-10T22:32:17.110000",
          "content": "<p>I'm getting similar resullts. Just started with google colab pro. Currently running one training through a GPU. Will try TPU after and hope pytorch lightning is able to seamlessly make the transition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 965839,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-10T22:36:36.847000",
          "content": "<p><a href=\"https://www.kaggle.com/serigne\" target=\"_blank\">@serigne</a> what is your CV for this 0.946 sub?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966050,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-11T05:11:39.460000",
          "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>  OOF (with 5 TTA per fold) = 0.930</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 966147,
          "author_name": "Vishnu Subramanian",
          "author_url": "",
          "post_date": "2020-08-11T07:24:10.347000",
          "content": "<p>Can you share information about the package information from where you are using 'resnest50_fast_1s1x64d'. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 966497,
          "author_name": "Serigne ",
          "author_url": "",
          "post_date": "2020-08-11T13:13:58.023000",
          "content": "<p>I use the <a href=\"https://github.com/zhanghang1989/ResNeSt\" target=\"_blank\">official repo</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 966496,
      "author_name": "Signal",
      "author_url": "",
      "post_date": "2020-08-11T13:13:39.147000",
      "content": "<p>All you all using the same head/classifier with these different models, that you were using with EFnets?</p>\n\n<p>Right now my head is pretty basic for EFnets, any advice for improving it?</p>\n\n<pre><code>    self.meta_fc1 = nn.Linear(n_meta_features, 4)\n    self.meta_bn0 = nn.BatchNorm1d(4)\n    self.combined_fc0 = nn.Linear(num_ftrs + 4, 1)\n    self.avgpool0 = nn.AdaptiveAvgPool2d((1,1))\n    self.dropout0 = nn.Dropout(p=0.2)\n\n    self.dense = nn.Linear(num_ftrs, 1)\n\ndef forward(self, inputs):\n\n    # Image + Meta\n    image, meta = inputs\n    cnn_features = self.arch.extract_features(image)         \n    cnn_features = self.avgpool0(cnn_features) \n    cnn_features = cnn_features.flatten(start_dim=1)\n    meta_features = self.meta_fc1(meta)\n    meta_features = F.selu(meta_features)\n    meta_features = self.meta_bn0(meta_features)\n    features = torch.cat((cnn_features, meta_features), dim=1)\n    output = self.combined_fc0(features)\n    output = torch.sigmoid(output)\n    return output  \n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 972219,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-16T10:50:46.880000",
          "content": "<p>Hi, you are new to Kaggle hence I feel I should give you a little friendly advice.  It is considered bad practice to share specific information during the last week of competition.  Why?  Because in some cases, people shared very valuable info on the last day and it blew up the leaderboard.  Sharing at least a week in advance makes sure every participant has a fair chance to leverage what has been shared.</p>\n<p>I am not saying you are giving away very valuable info, but you might.</p>\n<p>Disclaimer: I don't work at Kaggle nor am I tied to Kaggle in anyway, you can freely ignore my friendly advice ;)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "962929": "It's a known issue with efficient nets that it is usually slower on Pytorch. Afaik depthwise convolution are really poorly optimized in Pytorch as can be seen in some tickets both on the efficientnet repo: https://github.com/lukemelas/EfficientNet-PyTorch/issues/159\nand pytorch: https://github.com/pytorch/pytorch/issues/18631\n\nThankfully there is a recent paper proposing ResNest https://arxiv.org/pdf/2004.08955.pdf and a pytorch implementation here https://github.com/zhanghang1989/ResNeSt. For example for me a B2 and a Resnest50 train in approx the same time but the resnest50 showed a good CV/LB increase since it's a bigger model at 25M instead of the 7M for the B2.\n\nSome numbers on my gtx1080ti:\nB2 7M 7:30 per epoch\nResnest50 25M 6:45 per epoch\nB4 19M 20:00 per epoch\nResnest100 46M 14:00 per epoch\n\nAll this is of course slower than TF + TPU :)",
    "964937": "First try isn't good, I guess I have to retune everything.",
    "963322": "&gt; All this is of course slower than TF + TPU :)\n\nIt's a little bit sad that Pytorch doesn't have a proper way right now to load data directly on TPUs( need to use CPU -&gt; Image data preprocessing is extremely slow at kaggle. So XLA could be really used at kaggle only for NLP, and even their 16 GB RAM limitation makes it difficult to deal with. While TF is not flexible, and it is difficult to add even simple things.\n",
    "962954": "Resnest are fast indeed and give good results :)\n\nYou have even faster versions of Resnest50 from the hub\n```\n'resnest50_fast_1s1x64d',   'resnest50_fast_1s2x40d',\n'resnest50_fast_1s4x24d',  'resnest50_fast_2s1x64d','\nresnest50_fast_2s2x40d',  'resnest50_fast_4s1x64d',\n'resnest50_fast_4s2x40d'\n``` \n\n\nFor efficientnet , I would recommend using [timm](https://github.com/rwightman/pytorch-image-models) by @rwightman  instead of  this repo",
    "970243": "`Resnest50` converges much faster than Efficientnet and is less prone to noise. But weirdly, I'm getting large CV and LB gap using this model. I'm using `resnest50d_1s4x24d`. One of my experiments yielded `0.9254` on CV and `0.903` on LB. 😨",
    "965368": "It's true my intial kernels in the beginning with resnet50 with 5 folds 128 image size gave a socre of 0.91.\nEfficientNet in pytorch does run slow. Given Lower score.",
    "964305": "Thanks, I think it explains some of my problems with pytorch in this competition.",
    "963867": "For `TF/Keras` user, : [ResNeSt](https://github.com/QiaoranC/tf_ResNeSt_RegNet_model) , TF version.\n\nBut note, no pre-trained weights for `tf` version yet, one needs to [convert](https://github.com/nerox8664/pytorch2keras) it to try. ",
    "963618": "I came across this issue a few months ago with pytorch, it's crazy they still didn't optimize it.",
    "963509": "I am trying resnext50 from [here](https://pytorch.org/hub/facebookresearch_semi-supervised-ImageNet1K-models_resnext/). It can easily give .93LB on images of size 256. They are faster than resnest too. ",
    "965773": "I've just submitted my `resnest50_fast_1s1x64d` model trained on 384*384\n\nLB 0.946 : 5 folds  + 15 TTA for each fold  inference \n\nIt run really Fast on TPU (with big batch size) ",
    "966496": "All you all using the same head/classifier with these different models, that you were using with EFnets?\n\nRight now my head is pretty basic for EFnets, any advice for improving it?\n\n        self.meta_fc1 = nn.Linear(n_meta_features, 4)\n        self.meta_bn0 = nn.BatchNorm1d(4)\n        self.combined_fc0 = nn.Linear(num_ftrs + 4, 1)\n        self.avgpool0 = nn.AdaptiveAvgPool2d((1,1))\n        self.dropout0 = nn.Dropout(p=0.2)\n        \n        self.dense = nn.Linear(num_ftrs, 1)\n                                        \n    def forward(self, inputs):\n\n        # Image + Meta\n        image, meta = inputs\n        cnn_features = self.arch.extract_features(image)         \n        cnn_features = self.avgpool0(cnn_features) \n        cnn_features = cnn_features.flatten(start_dim=1)\n        meta_features = self.meta_fc1(meta)\n        meta_features = F.selu(meta_features)\n        meta_features = self.meta_bn0(meta_features)\n        features = torch.cat((cnn_features, meta_features), dim=1)\n        output = self.combined_fc0(features)\n        output = torch.sigmoid(output)\n        return output  \n     "
  }
}