{
  "id": 218907,
  "title": "Introducing new SOTA model NFNets",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/218907",
  "author_name": "Mobassir",
  "post_date": "2021-02-12T15:41:53.474000",
  "votes": 74,
  "comment_count": 26,
  "views": 0,
  "content": "<p>Introducing NFNets, a family of image classification models that are:</p>\n<p>*SOTA on ImageNet (86.5% top-1 w/o extra data)<br>\n*Up to 8.7x faster to train than EfficientNets to a given accuracy<br>\n*Normalizer-free (no BatchNorm!)</p>\n<p>Paper: <a href=\"http://dpmd.ai/06171\" target=\"_blank\">http://dpmd.ai/06171</a><br>\nCode: <a href=\"http://dpmd.ai/nfnets\" target=\"_blank\">http://dpmd.ai/nfnets</a><br>\n<img src=\"https://pbs.twimg.com/media/EuB11mwXEAAaG_T?format=png&amp;name=small\" alt=\"\"></p>",
  "messages": [
    {
      "id": 1197989,
      "postDate": "2021-02-12T15:41:53.473Z",
      "content": "<p>Introducing NFNets, a family of image classification models that are:</p>\n<p>*SOTA on ImageNet (86.5% top-1 w/o extra data)<br>\n*Up to 8.7x faster to train than EfficientNets to a given accuracy<br>\n*Normalizer-free (no BatchNorm!)</p>\n<p>Paper: <a href=\"http://dpmd.ai/06171\" target=\"_blank\">http://dpmd.ai/06171</a><br>\nCode: <a href=\"http://dpmd.ai/nfnets\" target=\"_blank\">http://dpmd.ai/nfnets</a><br>\n<img src=\"https://pbs.twimg.com/media/EuB11mwXEAAaG_T?format=png&amp;name=small\" alt=\"\"></p>",
      "rawMarkdown": "Introducing NFNets, a family of image classification models that are:\n\n*SOTA on ImageNet (86.5% top-1 w/o extra data)\n*Up to 8.7x faster to train than EfficientNets to a given accuracy\n*Normalizer-free (no BatchNorm!)\n\nPaper: http://dpmd.ai/06171\nCode: http://dpmd.ai/nfnets\n![](https://pbs.twimg.com/media/EuB11mwXEAAaG_T?format=png&name=small)",
      "votes": 73
    },
    {
      "id": 1198008,
      "postDate": "2021-02-12T16:09:23.143Z",
      "content": "<p>Maybe i should have study <strong>jax</strong> instead of <strong>TensorFlow</strong>…</p>\n<p>Google deepmind team: We are opensourcing a new work<br>\nTensorFlow team: have you used <strong>TensorFlow</strong> right?<br>\nGoogle deepmind team: TensorWhat?</p>",
      "rawMarkdown": "Maybe i should have study **jax** instead of **TensorFlow**...\n\nGoogle deepmind team: We are opensourcing a new work\nTensorFlow team: have you used **TensorFlow** right?\nGoogle deepmind team: TensorWhat?",
      "votes": 16
    },
    {
      "id": 1198722,
      "postDate": "2021-02-13T08:21:24.150Z",
      "content": "<h1>Update</h1>\n<p>if you want to try NF RegNet models in pytorch as <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> mentioned below,then you can follow these steps : <br>\ninclude this updated dataset : <a href=\"https://www.kaggle.com/mobassir/timm2021\" target=\"_blank\">https://www.kaggle.com/mobassir/timm2021</a> in your notebook,for example say in this notebook : <a href=\"https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9\" target=\"_blank\">https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9</a> </p>\n<p>and then replace these imports : </p>\n<pre><code>package_paths = [\n    '../input/pytorch-image-models/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>with this : </p>\n<pre><code>package_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>then use for example,</p>\n<p>kernel_type = 'nf_resnet50' </p>\n<p>net_type = 'nf_resnet50'</p>\n<p>and model class should be like this : </p>\n<pre><code>class cassavamodel(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n</code></pre>\n<p>if you want to use NFNet-F* models,then use the latest version of the dataset and the code below should work : </p>\n<p>replace this : </p>\n<pre><code>package_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>with this : </p>\n<pre><code>package_paths = [\n    '../input/timm2021/timm2021', \n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>then use for example,</p>\n<p>kernel_type = 'nfnet_f1' </p>\n<p>net_type = 'nfnet_f1'</p>\n<p>and model class should be like this : </p>\n<pre><code>class CassvaImgClassifier(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n</code></pre>\n<p>you can use these steps in your inference kernel for making offline inference in this competition<br>\nhope it helps,thank you</p>",
      "rawMarkdown": "# Update\n\nif you want to try NF RegNet models in pytorch as @stanleyjzheng mentioned below,then you can follow these steps : \ninclude this updated dataset : https://www.kaggle.com/mobassir/timm2021 in your notebook,for example say in this notebook : https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9 \n\nand then replace these imports : \n```\npackage_paths = [\n    '../input/pytorch-image-models/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n```\nwith this : \n\n```\npackage_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n\n```\n\nthen use for example,\n\nkernel_type = 'nf_resnet50' \n\nnet_type = 'nf_resnet50'\n\nand model class should be like this : \n\n```\nclass cassavamodel(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n```\nif you want to use NFNet-F* models,then use the latest version of the dataset and the code below should work : \n\nreplace this : \n```\npackage_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n\n```\n\nwith this : \n\n```\npackage_paths = [\n    '../input/timm2021/timm2021', \n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n\n```\n\nthen use for example,\n\nkernel_type = 'nfnet_f1' \n\nnet_type = 'nfnet_f1'\n\nand model class should be like this : \n\n```\nclass CassvaImgClassifier(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n    \n```\n\nyou can use these steps in your inference kernel for making offline inference in this competition\nhope it helps,thank you",
      "votes": 9
    },
    {
      "id": 1199573,
      "postDate": "2021-02-13T23:49:39.987Z",
      "content": "<p><img src=\"https://i.imgflip.com/4xy4fu.jpg\" alt=\"\"></p>",
      "rawMarkdown": "![](https://i.imgflip.com/4xy4fu.jpg)",
      "votes": 5
    },
    {
      "id": 1198146,
      "postDate": "2021-02-12T18:56:26.913Z",
      "content": "<p>Absolutely amazing stuff, no batch norm means much less memory usage and much faster training, not to mention it doesn't use additional training data. No PyTorch/Tensorflow implementation kills it unfortunately; doubt there will be pretrained models out before the end of this competition. </p>\n<p>NF RegNet models and some pretrained are available at this repo: <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py</a><br>\nUnfortunately, doesn't cover the new NFNet-F* models in this new paper and the AGC implementation.</p>",
      "rawMarkdown": "Absolutely amazing stuff, no batch norm means much less memory usage and much faster training, not to mention it doesn't use additional training data. No PyTorch/Tensorflow implementation kills it unfortunately; doubt there will be pretrained models out before the end of this competition. \n\nNF RegNet models and some pretrained are available at this repo: https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\nUnfortunately, doesn't cover the new NFNet-F* models in this new paper and the AGC implementation.",
      "votes": 3,
      "replies": [
        {
          "id": 1198148,
          "postDate": "2021-02-12T18:58:07.610Z",
          "content": "<p>More details from the creator of pytorch-image-models on Twitter here: <a href=\"https://twitter.com/ak92501/status/1360077120253415425\" target=\"_blank\">https://twitter.com/ak92501/status/1360077120253415425</a></p>",
          "rawMarkdown": "More details from the creator of pytorch-image-models on Twitter here: https://twitter.com/ak92501/status/1360077120253415425",
          "votes": 5
        },
        {
          "id": 1198265,
          "postDate": "2021-02-12T20:58:15.790Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1198539,
          "postDate": "2021-02-13T06:02:51.103Z",
          "content": "<p>Anybody know the source of NFNet-F* models in keras frameworks!!</p>",
          "rawMarkdown": "Anybody know the source of NFNet-F* models in keras frameworks!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1201274,
      "postDate": "2021-02-15T09:15:29.220Z",
      "content": "<p>Although you can find NFNets in <code>timm</code>, the trick in this paper is the Adaptive Gradient Clipping. AFAIK, the example code for AGC hasn't been released yet</p>\n<p>Edit: maybe it's there now :) <a href=\"https://github.com/deepmind/deepmind-research/blob/68a1754d293e04d278dfb649b2d59512f32351a7/nfnets/optim.py#L476-L491\" target=\"_blank\">https://github.com/deepmind/deepmind-research/blob/68a1754d293e04d278dfb649b2d59512f32351a7/nfnets/optim.py#L476-L491</a> </p>",
      "rawMarkdown": "Although you can find NFNets in `timm`, the trick in this paper is the Adaptive Gradient Clipping. AFAIK, the example code for AGC hasn't been released yet\n\nEdit: maybe it's there now :) https://github.com/deepmind/deepmind-research/blob/68a1754d293e04d278dfb649b2d59512f32351a7/nfnets/optim.py#L476-L491 ",
      "votes": 4,
      "replies": [
        {
          "id": 1201287,
          "postDate": "2021-02-15T09:31:24.177Z",
          "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> i do agree with you</p>",
          "rawMarkdown": "@anjum48 i do agree with you",
          "votes": 1
        },
        {
          "id": 1204599,
          "postDate": "2021-02-16T08:46:28Z",
          "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> The researchers state in their paper, that AGC is only relevant with large batchsize and its \"usefulness\" drops with lowering the batchsize.</p>\n<p>Quote:</p>\n<blockquote>\n  <p>As anticipated, the benefits of using AGC are smaller when the batch size is small.</p>\n</blockquote>",
          "rawMarkdown": "@anjum48 The researchers state in their paper, that AGC is only relevant with large batchsize and its \"usefulness\" drops with lowering the batchsize.\n\nQuote:\n> As anticipated, the benefits of using AGC are smaller when the batch size is small.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1199647,
      "postDate": "2021-02-14T03:16:39.730Z",
      "content": "<p>Tried nf_resnet50 and nf_regnet and it doesnt match my efficentnet models. Wondering if others are seeing something similar. I was surprised by the number of parameters to train - regnet_b4 has 28M and regnet_b5 has 40M vs 17M for efficientnet_b4. </p>",
      "rawMarkdown": "Tried nf_resnet50 and nf_regnet and it doesnt match my efficentnet models. Wondering if others are seeing something similar. I was surprised by the number of parameters to train - regnet_b4 has 28M and regnet_b5 has 40M vs 17M for efficientnet_b4. ",
      "votes": 4
    },
    {
      "id": 1199530,
      "postDate": "2021-02-13T21:38:13.443Z",
      "content": "<p>I have tried to use it instead of simple resnet50. The new one is worse by around 0.001.</p>",
      "rawMarkdown": "I have tried to use it instead of simple resnet50. The new one is worse by around 0.001.",
      "votes": 4
    },
    {
      "id": 1209977,
      "postDate": "2021-02-19T06:17:26.697Z",
      "content": "<h1>Paper_Summary</h1>\n<p>It's a very interesting paper. Talks a lot about Batch normalization.<br>\n<a href=\"https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\" target=\"_blank\">https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/</a><br>\nWould love to hear your opinions.</p>",
      "rawMarkdown": "#Paper_Summary\nIt's a very interesting paper. Talks a lot about Batch normalization.\nhttps://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\nWould love to hear your opinions.",
      "votes": 1
    },
    {
      "id": 1206645,
      "postDate": "2021-02-17T13:41:57.280Z",
      "content": "<p>every once in a while something new shows up to surprise us!</p>",
      "rawMarkdown": "every once in a while something new shows up to surprise us!",
      "votes": 1
    },
    {
      "id": 1206618,
      "postDate": "2021-02-17T13:25:58.737Z",
      "content": "<p>DeepMind is awesome :) </p>",
      "rawMarkdown": "DeepMind is awesome :) ",
      "votes": 1
    },
    {
      "id": 1202087,
      "postDate": "2021-02-15T20:35:00.633Z",
      "content": "<p>Wow… This is revolutionary stuff! Thanks for sharing!</p>",
      "rawMarkdown": "Wow... This is revolutionary stuff! Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1199953,
      "postDate": "2021-02-14T09:21:55.437Z",
      "content": "<p>Great right! I'm amazed by the 8,7x faster training. Can't wait to give it a try!</p>",
      "rawMarkdown": "Great right! I'm amazed by the 8,7x faster training. Can't wait to give it a try!",
      "votes": 1
    },
    {
      "id": 1199595,
      "postDate": "2021-02-14T01:32:57.487Z",
      "content": "<p>Hi,is it have a good performance?I found regnet is not good in this game.</p>",
      "rawMarkdown": "Hi,is it have a good performance?I found regnet is not good in this game.",
      "votes": 1,
      "replies": [
        {
          "id": 1199706,
          "postDate": "2021-02-14T03:36:29.487Z",
          "content": "<p>seeing similar results for resnet and regnets vs effnets for this dataset. The noisy student approach for effnets might be particularly important for this dataset. </p>",
          "rawMarkdown": "seeing similar results for resnet and regnets vs effnets for this dataset. The noisy student approach for effnets might be particularly important for this dataset. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 1199464,
      "postDate": "2021-02-13T20:29:28.367Z",
      "content": "<p>The paper also just featured in Sebastian Raschka's <a href=\"https://youtu.be/2I8SqKLt-Nk\" target=\"_blank\">Deep Learning News #3</a> - a lot of other interesting stuff, too. However, this thread is great and gave me a lot more details/perspective on NFNets.</p>",
      "rawMarkdown": "The paper also just featured in Sebastian Raschka's [Deep Learning News #3](https://youtu.be/2I8SqKLt-Nk) - a lot of other interesting stuff, too. However, this thread is great and gave me a lot more details/perspective on NFNets.",
      "votes": 1
    },
    {
      "id": 1198002,
      "postDate": "2021-02-12T16:02:57.223Z",
      "content": "<p>Have you tried implementing it?<br>\nDid it give any significant boost compared to Efficientnets?</p>",
      "rawMarkdown": "Have you tried implementing it?\nDid it give any significant boost compared to Efficientnets?",
      "votes": 1,
      "replies": [
        {
          "id": 1198058,
          "postDate": "2021-02-12T17:06:43.637Z",
          "content": "<p>It released officially few hours ago, so yet to give it a try!!</p>",
          "rawMarkdown": "It released officially few hours ago, so yet to give it a try!!"
        }
      ]
    },
    {
      "id": 1198333,
      "postDate": "2021-02-12T23:43:19.023Z",
      "content": "<p>It looks like these have huge numbers of parameters (multiple of an EfficientNet with the same training latency), but are somehow really fast to train on TPU.  Do these huge numbers of parameters not result in the same amount of memory usage as with an EfficientNet due to something (like absence of BN layers)? Or do these just work this well with TPUs, but I should not get my hope up about using these with my own GPU? </p>\n<p>I guess on Kaggle we do have the TPUv3 and I've seen a few notebooks that use JAX, so it can be done in Kaggle notebooks. And sooner or later someone will do a TF and PyTorch port.</p>",
      "rawMarkdown": "It looks like these have huge numbers of parameters (multiple of an EfficientNet with the same training latency), but are somehow really fast to train on TPU.  Do these huge numbers of parameters not result in the same amount of memory usage as with an EfficientNet due to something (like absence of BN layers)? Or do these just work this well with TPUs, but I should not get my hope up about using these with my own GPU? \n\nI guess on Kaggle we do have the TPUv3 and I've seen a few notebooks that use JAX, so it can be done in Kaggle notebooks. And sooner or later someone will do a TF and PyTorch port.",
      "votes": 2,
      "replies": [
        {
          "id": 1198344,
          "postDate": "2021-02-13T00:11:42.743Z",
          "content": "<p>Did some research on this - it is a combination of a ton of different innovations. Of course, haven't had time to test, but this is my understanding. </p>\n<p>Why NFNet trains as fast as an Efficientnet that has 50x less FLOPs and 15x less parameters is purely different optimization. Efficientnet is optimized on number of FLOPs, but current computing hardware cannot take advantage. NFNets were optimized with current computing hardware in mind. I'm sure the innovations of AGC and removing batch norm help, but I couldn't find how much they helped. </p>\n<blockquote>\n  <p>The choice of which metric to optimize– theoretical FLOPS,inference latency on a target device, or training latency on an accelerator–is a matter of preference, and the nature of each metric will yield different design requirements. In this work we choose to focus on manually designing models whichTable are optimized for training latency on existing accelerators,as in Radosavovic et al. (2020).  It is possible that future accelerators may be able to take full advantage of the potential training speed that largely goes unrealized with models like EfficientNets, so we believe this direction should not be ignored (Hooker, 2020), however we anticipate that developing models with improved training speed on curren thardware will be beneficial for accelerating research.  We note that accelerators like GPU and TPU tend to favor dense computation, and while there are differences between these two platforms, they have enough in common that models designed for one device are likely to train fast on the other</p>\n</blockquote>",
          "rawMarkdown": "Did some research on this - it is a combination of a ton of different innovations. Of course, haven't had time to test, but this is my understanding. \n\nWhy NFNet trains as fast as an Efficientnet that has 50x less FLOPs and 15x less parameters is purely different optimization. Efficientnet is optimized on number of FLOPs, but current computing hardware cannot take advantage. NFNets were optimized with current computing hardware in mind. I'm sure the innovations of AGC and removing batch norm help, but I couldn't find how much they helped. \n\n> The choice of which metric to optimize– theoretical FLOPS,inference latency on a target device, or training latency on an accelerator–is a matter of preference, and the nature of each metric will yield different design requirements. In this work we choose to focus on manually designing models whichTable are optimized for training latency on existing accelerators,as in Radosavovic et al. (2020).  It is possible that future accelerators may be able to take full advantage of the potential training speed that largely goes unrealized with models like EfficientNets, so we believe this direction should not be ignored (Hooker, 2020), however we anticipate that developing models with improved training speed on curren thardware will be beneficial for accelerating research.  We note that accelerators like GPU and TPU tend to favor dense computation, and while there are differences between these two platforms, they have enough in common that models designed for one device are likely to train fast on the other\n\n",
          "votes": 2
        },
        {
          "id": 1200732,
          "postDate": "2021-02-14T23:43:58.690Z",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Nevermind, I stand corrected! There was a table in the appendix. It seems removing BN improves latency by 50-100%<br>\n<img src=\"https://i.imgur.com/KwkG5aR.png\" alt=\"\"></p>",
          "rawMarkdown": "@bjoernholzhauer Nevermind, I stand corrected! There was a table in the appendix. It seems removing BN improves latency by 50-100%\n![](https://i.imgur.com/KwkG5aR.png)",
          "votes": 2
        }
      ]
    },
    {
      "id": 1199606,
      "postDate": "2021-02-14T01:59:09.960Z",
      "content": "<p>thanks for this great news!</p>",
      "rawMarkdown": "thanks for this great news!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1198008,
      "author_name": "Hiram Coria 🧬",
      "author_url": "",
      "post_date": "2021-02-12T16:09:23.143000",
      "content": "<p>Maybe i should have study <strong>jax</strong> instead of <strong>TensorFlow</strong>…</p>\n<p>Google deepmind team: We are opensourcing a new work<br>\nTensorFlow team: have you used <strong>TensorFlow</strong> right?<br>\nGoogle deepmind team: TensorWhat?</p>",
      "votes": 16,
      "replies": []
    },
    {
      "id": 1198722,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2021-02-13T08:21:24.150000",
      "content": "<h1>Update</h1>\n<p>if you want to try NF RegNet models in pytorch as <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> mentioned below,then you can follow these steps : <br>\ninclude this updated dataset : <a href=\"https://www.kaggle.com/mobassir/timm2021\" target=\"_blank\">https://www.kaggle.com/mobassir/timm2021</a> in your notebook,for example say in this notebook : <a href=\"https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9\" target=\"_blank\">https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9</a> </p>\n<p>and then replace these imports : </p>\n<pre><code>package_paths = [\n    '../input/pytorch-image-models/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>with this : </p>\n<pre><code>package_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>then use for example,</p>\n<p>kernel_type = 'nf_resnet50' </p>\n<p>net_type = 'nf_resnet50'</p>\n<p>and model class should be like this : </p>\n<pre><code>class cassavamodel(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n</code></pre>\n<p>if you want to use NFNet-F* models,then use the latest version of the dataset and the code below should work : </p>\n<p>replace this : </p>\n<pre><code>package_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>with this : </p>\n<pre><code>package_paths = [\n    '../input/timm2021/timm2021', \n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n</code></pre>\n<p>then use for example,</p>\n<p>kernel_type = 'nfnet_f1' </p>\n<p>net_type = 'nfnet_f1'</p>\n<p>and model class should be like this : </p>\n<pre><code>class CassvaImgClassifier(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n</code></pre>\n<p>you can use these steps in your inference kernel for making offline inference in this competition<br>\nhope it helps,thank you</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 1199573,
      "author_name": "AmorfEvo",
      "author_url": "",
      "post_date": "2021-02-13T23:49:39.987000",
      "content": "<p><img src=\"https://i.imgflip.com/4xy4fu.jpg\" alt=\"\"></p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1198146,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2021-02-12T18:56:26.913000",
      "content": "<p>Absolutely amazing stuff, no batch norm means much less memory usage and much faster training, not to mention it doesn't use additional training data. No PyTorch/Tensorflow implementation kills it unfortunately; doubt there will be pretrained models out before the end of this competition. </p>\n<p>NF RegNet models and some pretrained are available at this repo: <a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py</a><br>\nUnfortunately, doesn't cover the new NFNet-F* models in this new paper and the AGC implementation.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1198148,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-12T18:58:07.610000",
          "content": "<p>More details from the creator of pytorch-image-models on Twitter here: <a href=\"https://twitter.com/ak92501/status/1360077120253415425\" target=\"_blank\">https://twitter.com/ak92501/status/1360077120253415425</a></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1198265,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-12T20:58:15.790000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1198539,
          "author_name": "Akhilesh D. Kapse",
          "author_url": "",
          "post_date": "2021-02-13T06:02:51.103000",
          "content": "<p>Anybody know the source of NFNet-F* models in keras frameworks!!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1201274,
      "author_name": "datasaurus",
      "author_url": "",
      "post_date": "2021-02-15T09:15:29.220000",
      "content": "<p>Although you can find NFNets in <code>timm</code>, the trick in this paper is the Adaptive Gradient Clipping. AFAIK, the example code for AGC hasn't been released yet</p>\n<p>Edit: maybe it's there now :) <a href=\"https://github.com/deepmind/deepmind-research/blob/68a1754d293e04d278dfb649b2d59512f32351a7/nfnets/optim.py#L476-L491\" target=\"_blank\">https://github.com/deepmind/deepmind-research/blob/68a1754d293e04d278dfb649b2d59512f32351a7/nfnets/optim.py#L476-L491</a> </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1201287,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2021-02-15T09:31:24.177000",
          "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> i do agree with you</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1204599,
          "author_name": "Ali Abdin",
          "author_url": "",
          "post_date": "2021-02-16T08:46:28",
          "content": "<p><a href=\"https://www.kaggle.com/anjum48\" target=\"_blank\">@anjum48</a> The researchers state in their paper, that AGC is only relevant with large batchsize and its \"usefulness\" drops with lowering the batchsize.</p>\n<p>Quote:</p>\n<blockquote>\n  <p>As anticipated, the benefits of using AGC are smaller when the batch size is small.</p>\n</blockquote>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1199647,
      "author_name": "Trushant Kalyanpur",
      "author_url": "",
      "post_date": "2021-02-14T03:16:39.730000",
      "content": "<p>Tried nf_resnet50 and nf_regnet and it doesnt match my efficentnet models. Wondering if others are seeing something similar. I was surprised by the number of parameters to train - regnet_b4 has 28M and regnet_b5 has 40M vs 17M for efficientnet_b4. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1199530,
      "author_name": "Roman",
      "author_url": "",
      "post_date": "2021-02-13T21:38:13.443000",
      "content": "<p>I have tried to use it instead of simple resnet50. The new one is worse by around 0.001.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1209977,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-02-19T06:17:26.697000",
      "content": "<h1>Paper_Summary</h1>\n<p>It's a very interesting paper. Talks a lot about Batch normalization.<br>\n<a href=\"https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\" target=\"_blank\">https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/</a><br>\nWould love to hear your opinions.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1206645,
      "author_name": "Mau Rua",
      "author_url": "",
      "post_date": "2021-02-17T13:41:57.280000",
      "content": "<p>every once in a while something new shows up to surprise us!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1206618,
      "author_name": "Mostafa Tanasan",
      "author_url": "",
      "post_date": "2021-02-17T13:25:58.737000",
      "content": "<p>DeepMind is awesome :) </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1202087,
      "author_name": "Andy Jian Zhou",
      "author_url": "",
      "post_date": "2021-02-15T20:35:00.633000",
      "content": "<p>Wow… This is revolutionary stuff! Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1199953,
      "author_name": "Maarten",
      "author_url": "",
      "post_date": "2021-02-14T09:21:55.437000",
      "content": "<p>Great right! I'm amazed by the 8,7x faster training. Can't wait to give it a try!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1199595,
      "author_name": "Zekun",
      "author_url": "",
      "post_date": "2021-02-14T01:32:57.487000",
      "content": "<p>Hi,is it have a good performance?I found regnet is not good in this game.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1199706,
          "author_name": "Trushant Kalyanpur",
          "author_url": "",
          "post_date": "2021-02-14T03:36:29.487000",
          "content": "<p>seeing similar results for resnet and regnets vs effnets for this dataset. The noisy student approach for effnets might be particularly important for this dataset. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1199464,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-13T20:29:28.367000",
      "content": "<p>The paper also just featured in Sebastian Raschka's <a href=\"https://youtu.be/2I8SqKLt-Nk\" target=\"_blank\">Deep Learning News #3</a> - a lot of other interesting stuff, too. However, this thread is great and gave me a lot more details/perspective on NFNets.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1198002,
      "author_name": "Mohneesh_Sreegirisetty",
      "author_url": "",
      "post_date": "2021-02-12T16:02:57.223000",
      "content": "<p>Have you tried implementing it?<br>\nDid it give any significant boost compared to Efficientnets?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1198058,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2021-02-12T17:06:43.637000",
          "content": "<p>It released officially few hours ago, so yet to give it a try!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1198333,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-12T23:43:19.023000",
      "content": "<p>It looks like these have huge numbers of parameters (multiple of an EfficientNet with the same training latency), but are somehow really fast to train on TPU.  Do these huge numbers of parameters not result in the same amount of memory usage as with an EfficientNet due to something (like absence of BN layers)? Or do these just work this well with TPUs, but I should not get my hope up about using these with my own GPU? </p>\n<p>I guess on Kaggle we do have the TPUv3 and I've seen a few notebooks that use JAX, so it can be done in Kaggle notebooks. And sooner or later someone will do a TF and PyTorch port.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1198344,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-13T00:11:42.743000",
          "content": "<p>Did some research on this - it is a combination of a ton of different innovations. Of course, haven't had time to test, but this is my understanding. </p>\n<p>Why NFNet trains as fast as an Efficientnet that has 50x less FLOPs and 15x less parameters is purely different optimization. Efficientnet is optimized on number of FLOPs, but current computing hardware cannot take advantage. NFNets were optimized with current computing hardware in mind. I'm sure the innovations of AGC and removing batch norm help, but I couldn't find how much they helped. </p>\n<blockquote>\n  <p>The choice of which metric to optimize– theoretical FLOPS,inference latency on a target device, or training latency on an accelerator–is a matter of preference, and the nature of each metric will yield different design requirements. In this work we choose to focus on manually designing models whichTable are optimized for training latency on existing accelerators,as in Radosavovic et al. (2020).  It is possible that future accelerators may be able to take full advantage of the potential training speed that largely goes unrealized with models like EfficientNets, so we believe this direction should not be ignored (Hooker, 2020), however we anticipate that developing models with improved training speed on curren thardware will be beneficial for accelerating research.  We note that accelerators like GPU and TPU tend to favor dense computation, and while there are differences between these two platforms, they have enough in common that models designed for one device are likely to train fast on the other</p>\n</blockquote>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1200732,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-14T23:43:58.690000",
          "content": "<p><a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> Nevermind, I stand corrected! There was a table in the appendix. It seems removing BN improves latency by 50-100%<br>\n<img src=\"https://i.imgur.com/KwkG5aR.png\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1199606,
      "author_name": "Memento",
      "author_url": "",
      "post_date": "2021-02-14T01:59:09.960000",
      "content": "<p>thanks for this great news!</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1197989": "Introducing NFNets, a family of image classification models that are:\n\n*SOTA on ImageNet (86.5% top-1 w/o extra data)\n*Up to 8.7x faster to train than EfficientNets to a given accuracy\n*Normalizer-free (no BatchNorm!)\n\nPaper: http://dpmd.ai/06171\nCode: http://dpmd.ai/nfnets\n![](https://pbs.twimg.com/media/EuB11mwXEAAaG_T?format=png&name=small)",
    "1198008": "Maybe i should have study **jax** instead of **TensorFlow**...\n\nGoogle deepmind team: We are opensourcing a new work\nTensorFlow team: have you used **TensorFlow** right?\nGoogle deepmind team: TensorWhat?",
    "1198722": "# Update\n\nif you want to try NF RegNet models in pytorch as @stanleyjzheng mentioned below,then you can follow these steps : \ninclude this updated dataset : https://www.kaggle.com/mobassir/timm2021 in your notebook,for example say in this notebook : https://www.kaggle.com/mobassir/faster-pytorch-tpu-baseline-for-cld-cv-0-9 \n\nand then replace these imports : \n```\npackage_paths = [\n    '../input/pytorch-image-models/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n```\nwith this : \n\n```\npackage_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n\n```\n\nthen use for example,\n\nkernel_type = 'nf_resnet50' \n\nnet_type = 'nf_resnet50'\n\nand model class should be like this : \n\n```\nclass cassavamodel(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n```\nif you want to use NFNet-F* models,then use the latest version of the dataset and the code below should work : \n\nreplace this : \n```\npackage_paths = [\n    '../input/timm2021/pytorch-image-models-master', #'../input/efficientnet-pytorch-07/efficientnet_pytorch-0.7.0'\n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n\n```\n\nwith this : \n\n```\npackage_paths = [\n    '../input/timm2021/timm2021', \n    '../input/image-fmix/FMix-master'\n]\nfor pth in package_paths:\n    sys.path.append(pth)\n\n```\n\nthen use for example,\n\nkernel_type = 'nfnet_f1' \n\nnet_type = 'nfnet_f1'\n\nand model class should be like this : \n\n```\nclass CassvaImgClassifier(nn.Module):\n    def __init__(self, model_arch, n_class, pretrained=True):\n        super().__init__()\n        self.model = timm.create_model(model_arch, pretrained=pretrained)\n        n_features = self.model.head.fc.in_features\n        self.model.head.fc = nn.Linear(n_features, n_class)\n\n    def forward(self, x):\n        x = self.model(x)\n        return x\n    \n```\n\nyou can use these steps in your inference kernel for making offline inference in this competition\nhope it helps,thank you",
    "1199573": "![](https://i.imgflip.com/4xy4fu.jpg)",
    "1198146": "Absolutely amazing stuff, no batch norm means much less memory usage and much faster training, not to mention it doesn't use additional training data. No PyTorch/Tensorflow implementation kills it unfortunately; doubt there will be pretrained models out before the end of this competition. \n\nNF RegNet models and some pretrained are available at this repo: https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\nUnfortunately, doesn't cover the new NFNet-F* models in this new paper and the AGC implementation.",
    "1201274": "Although you can find NFNets in `timm`, the trick in this paper is the Adaptive Gradient Clipping. AFAIK, the example code for AGC hasn't been released yet\n\nEdit: maybe it's there now :) https://github.com/deepmind/deepmind-research/blob/68a1754d293e04d278dfb649b2d59512f32351a7/nfnets/optim.py#L476-L491 ",
    "1199647": "Tried nf_resnet50 and nf_regnet and it doesnt match my efficentnet models. Wondering if others are seeing something similar. I was surprised by the number of parameters to train - regnet_b4 has 28M and regnet_b5 has 40M vs 17M for efficientnet_b4. ",
    "1199530": "I have tried to use it instead of simple resnet50. The new one is worse by around 0.001.",
    "1209977": "#Paper_Summary\nIt's a very interesting paper. Talks a lot about Batch normalization.\nhttps://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\nWould love to hear your opinions.",
    "1206645": "every once in a while something new shows up to surprise us!",
    "1206618": "DeepMind is awesome :) ",
    "1202087": "Wow... This is revolutionary stuff! Thanks for sharing!",
    "1199953": "Great right! I'm amazed by the 8,7x faster training. Can't wait to give it a try!",
    "1199595": "Hi,is it have a good performance?I found regnet is not good in this game.",
    "1199464": "The paper also just featured in Sebastian Raschka's [Deep Learning News #3](https://youtu.be/2I8SqKLt-Nk) - a lot of other interesting stuff, too. However, this thread is great and gave me a lot more details/perspective on NFNets.",
    "1198002": "Have you tried implementing it?\nDid it give any significant boost compared to Efficientnets?",
    "1198333": "It looks like these have huge numbers of parameters (multiple of an EfficientNet with the same training latency), but are somehow really fast to train on TPU.  Do these huge numbers of parameters not result in the same amount of memory usage as with an EfficientNet due to something (like absence of BN layers)? Or do these just work this well with TPUs, but I should not get my hope up about using these with my own GPU? \n\nI guess on Kaggle we do have the TPUv3 and I've seen a few notebooks that use JAX, so it can be done in Kaggle notebooks. And sooner or later someone will do a TF and PyTorch port.",
    "1199606": "thanks for this great news!"
  }
}