{
  "id": 219371,
  "title": "NFNet-F* models released in pytorch-image-models",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/219371",
  "author_name": "Stanley Zheng",
  "post_date": "2021-02-14T16:43:27.191000",
  "votes": 34,
  "comment_count": 30,
  "views": 0,
  "content": "<p><a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py</a></p>\n<p>NFNet models were recently released in pytorch-image-models. These are only the model definitions and no imagenet pretrained weights are available at the moment, but potentially something to try in the final days of this competition. </p>\n<p>How to use timm on Kaggle: <br>\nAdd this dataset: <a href=\"https://www.kaggle.com/stanleyjzheng/timm-nfnet\" target=\"_blank\">https://www.kaggle.com/stanleyjzheng/timm-nfnet</a></p>\n<pre><code>import sys; sys.path.insert(0,'../input/timm-nfnet')\nimport timm\n\nmodel = timm.create_model('nfnet_f1', pretrained=False)\n</code></pre>\n<p>I'm currently comparing NFNet-F0 to Effnet-B0 on VBD chest xray - I will update this thread with my results in a few hours. NFNet is not in the PyPi package yet, so I had to install from source. Will have results out on speed/vRAM usage soon.</p>\n<p>Edit: added dataset for kaggle installation</p>",
  "messages": [
    {
      "id": 1200425,
      "postDate": "2021-02-14T16:43:27.193Z",
      "content": "<p><a href=\"https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py</a></p>\n<p>NFNet models were recently released in pytorch-image-models. These are only the model definitions and no imagenet pretrained weights are available at the moment, but potentially something to try in the final days of this competition. </p>\n<p>How to use timm on Kaggle: <br>\nAdd this dataset: <a href=\"https://www.kaggle.com/stanleyjzheng/timm-nfnet\" target=\"_blank\">https://www.kaggle.com/stanleyjzheng/timm-nfnet</a></p>\n<pre><code>import sys; sys.path.insert(0,'../input/timm-nfnet')\nimport timm\n\nmodel = timm.create_model('nfnet_f1', pretrained=False)\n</code></pre>\n<p>I'm currently comparing NFNet-F0 to Effnet-B0 on VBD chest xray - I will update this thread with my results in a few hours. NFNet is not in the PyPi package yet, so I had to install from source. Will have results out on speed/vRAM usage soon.</p>\n<p>Edit: added dataset for kaggle installation</p>",
      "rawMarkdown": "https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\n\nNFNet models were recently released in pytorch-image-models. These are only the model definitions and no imagenet pretrained weights are available at the moment, but potentially something to try in the final days of this competition. \n\nHow to use timm on Kaggle: \nAdd this dataset: https://www.kaggle.com/stanleyjzheng/timm-nfnet\n\n```python\nimport sys; sys.path.insert(0,'../input/timm-nfnet')\nimport timm\n\nmodel = timm.create_model('nfnet_f1', pretrained=False)\n```\n\nI'm currently comparing NFNet-F0 to Effnet-B0 on VBD chest xray - I will update this thread with my results in a few hours. NFNet is not in the PyPi package yet, so I had to install from source. Will have results out on speed/vRAM usage soon.\n\nEdit: added dataset for kaggle installation",
      "votes": 34
    },
    {
      "id": 1200709,
      "postDate": "2021-02-14T22:49:14.847Z",
      "content": "<p>Thanks for sharing Stanley!</p>\n<p>if anyone is curious about this method, Yannic Kilcher just released a great video going over the paper:<br>\n<a href=\"https://www.youtube.com/watch?v=rNkHjZtH0RQ\" target=\"_blank\">https://www.youtube.com/watch?v=rNkHjZtH0RQ</a></p>",
      "rawMarkdown": "Thanks for sharing Stanley!\n\nif anyone is curious about this method, Yannic Kilcher just released a great video going over the paper:\nhttps://www.youtube.com/watch?v=rNkHjZtH0RQ",
      "votes": 7
    },
    {
      "id": 1200448,
      "postDate": "2021-02-14T17:17:11.800Z",
      "content": "<p>Speed results on RTX3090 using AMP.<br>\nEffnet b0 512x512 (5.3m params), batch size 16 (7.1gb vRAM used): avg 0.2004s/step<br>\nEfficientnet b3 512x512 (12m params), batch size 16 (11.2gb vRAM used): avg 0.3372s/step<br>\nNFNet F0 512x512 (71.5m params), batch size 16 (12.3gb vRAM used): avg 0.6365s/step<br>\nNFNet F1 512x512, (132.6m params), batch size 16 (22.4gb vRAM used): avg 1.0725s/step</p>\n<p>Update:<br>\nAll NFNet above use GeLU.<br>\nNFNet F0 SiLU 512x512 (71.5m params), batch size 16 (10.7gb vRAM used): avg 0.5501s/step</p>\n<p>So SiLU has lower vRAM usage and training latency.</p>",
      "rawMarkdown": "Speed results on RTX3090 using AMP.\nEffnet b0 512x512 (5.3m params), batch size 16 (7.1gb vRAM used): avg 0.2004s/step\nEfficientnet b3 512x512 (12m params), batch size 16 (11.2gb vRAM used): avg 0.3372s/step\nNFNet F0 512x512 (71.5m params), batch size 16 (12.3gb vRAM used): avg 0.6365s/step\nNFNet F1 512x512, (132.6m params), batch size 16 (22.4gb vRAM used): avg 1.0725s/step\n\nUpdate:\nAll NFNet above use GeLU.\nNFNet F0 SiLU 512x512 (71.5m params), batch size 16 (10.7gb vRAM used): avg 0.5501s/step\n\nSo SiLU has lower vRAM usage and training latency.\n",
      "votes": 3,
      "replies": [
        {
          "id": 1200497,
          "postDate": "2021-02-14T18:07:34.017Z",
          "content": "<p>What data did you use to test these, As you said there wasn't imagenet weights in them?</p>",
          "rawMarkdown": "What data did you use to test these, As you said there wasn't imagenet weights in them?"
        },
        {
          "id": 1200501,
          "postDate": "2021-02-14T18:11:34.943Z",
          "content": "<p>This is for the vbd chest x-ray competition, which is my main focus at the moment. Dataset shouldn't matter for speed though. <br>\nThere are no imagenet weights yet so this is a random weights initialization</p>",
          "rawMarkdown": "This is for the vbd chest x-ray competition, which is my main focus at the moment. Dataset shouldn't matter for speed though. \nThere are no imagenet weights yet so this is a random weights initialization"
        },
        {
          "id": 1200587,
          "postDate": "2021-02-14T19:01:16.930Z",
          "content": "<p>It's weird that in paper they claimed <code>NFNet</code> is faster than <code>EffNet</code>. 😳</p>",
          "rawMarkdown": "It's weird that in paper they claimed `NFNet` is faster than `EffNet`. 😳"
        },
        {
          "id": 1200592,
          "postDate": "2021-02-14T19:05:34.017Z",
          "content": "<p>I think it's down to JAX. Gelu GPU on pytorch is apparently slower, and I'm training at 512x512 on both models, while the comparison in the paper is 192x192 vs 224x224 for effnetb0 vs nfnetb0 if I remember correctly.</p>\n<p>I also just found out there is no AGC implementation in the pytorch version, which causes loss to explode at high LR. This may also affect speed.</p>",
          "rawMarkdown": "I think it's down to JAX. Gelu GPU on pytorch is apparently slower, and I'm training at 512x512 on both models, while the comparison in the paper is 192x192 vs 224x224 for effnetb0 vs nfnetb0 if I remember correctly.\n\nI also just found out there is no AGC implementation in the pytorch version, which causes loss to explode at high LR. This may also affect speed."
        },
        {
          "id": 1200597,
          "postDate": "2021-02-14T19:13:56.217Z",
          "content": "<p>btw did you use the official weight for <code>NfNet</code>?</p>",
          "rawMarkdown": "btw did you use the official weight for `NfNet`?"
        },
        {
          "id": 1200599,
          "postDate": "2021-02-14T19:17:55.923Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1200601,
          "postDate": "2021-02-14T19:22:59.837Z",
          "content": "<blockquote>\n  <p>Gelu GPU on pytorch is apparently slower</p>\n</blockquote>\n<p>I lose the reference but <strong>Mish</strong> it's better than <strong>Gelu</strong> for computer vision tasks. Maybe mish could improve the speed.</p>",
          "rawMarkdown": "> Gelu GPU on pytorch is apparently slower\n \nI lose the reference but **Mish** it's better than **Gelu** for computer vision tasks. Maybe mish could improve the speed."
        },
        {
          "id": 1200605,
          "postDate": "2021-02-14T19:28:16.183Z",
          "content": "<p>The F0 model also has more parameters than b0 and b3 combined.</p>\n<p>AGC will probably be implemented at the backward stage for PyTorch? So you will most likely have to implement it yourself.</p>",
          "rawMarkdown": "The F0 model also has more parameters than b0 and b3 combined.\n\nAGC will probably be implemented at the backward stage for PyTorch? So you will most likely have to implement it yourself.",
          "votes": 1
        },
        {
          "id": 1200634,
          "postDate": "2021-02-14T20:02:01.830Z",
          "content": "<blockquote>\n  <p>AGC will probably be implemented at the backward stage for PyTorch? So you will most likely have to implement it yourself.</p>\n</blockquote>\n<p>That's a great point, I somehow completely overlooked that, thanks. Will give it a shot.</p>",
          "rawMarkdown": "> AGC will probably be implemented at the backward stage for PyTorch? So you will most likely have to implement it yourself.\n\nThat's a great point, I somehow completely overlooked that, thanks. Will give it a shot."
        },
        {
          "id": 1200654,
          "postDate": "2021-02-14T20:25:34.970Z",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> are you also using <code>effnetb0</code> without pretrained weight?</p>",
          "rawMarkdown": "@stanleyjzheng are you also using `effnetb0` without pretrained weight?"
        },
        {
          "id": 1200656,
          "postDate": "2021-02-14T20:29:26.327Z",
          "content": "<p>Yes, correct, though it shouldn't matter speed-wise</p>",
          "rawMarkdown": "Yes, correct, though it shouldn't matter speed-wise",
          "votes": 1
        },
        {
          "id": 1201774,
          "postDate": "2021-02-15T16:40:45.813Z",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> My experiments seem to line up with yours. I cant seem to use high lr without a Nan loss returning after some time. Looks like AGC is really important to train with out BN</p>",
          "rawMarkdown": "@stanleyjzheng My experiments seem to line up with yours. I cant seem to use high lr without a Nan loss returning after some time. Looks like AGC is really important to train with out BN",
          "votes": 1
        },
        {
          "id": 1208102,
          "postDate": "2021-02-18T06:50:55.870Z",
          "content": "<p>goodness, I need 2x3090 for NFNet man, it is way too slow.</p>",
          "rawMarkdown": "goodness, I need 2x3090 for NFNet man, it is way too slow."
        }
      ]
    },
    {
      "id": 1202011,
      "postDate": "2021-02-15T19:49:35.603Z",
      "content": "<p>Tried F3s yesterday. Trained on 448 images. It's about 12 minutes/epoch on RTX3090*2 with DDP, AMP. Slow, indeed.</p>",
      "rawMarkdown": "Tried F3s yesterday. Trained on 448 images. It's about 12 minutes/epoch on RTX3090*2 with DDP, AMP. Slow, indeed.",
      "votes": 4,
      "replies": [
        {
          "id": 1202109,
          "postDate": "2021-02-15T21:05:42.123Z",
          "content": "<p>Im thinking that this model doesnt scale well with higher resolutions in comparison to effnets. Or maybe the implementation isnt right/optimized.</p>",
          "rawMarkdown": "Im thinking that this model doesnt scale well with higher resolutions in comparison to effnets. Or maybe the implementation isnt right/optimized."
        }
      ]
    },
    {
      "id": 1207721,
      "postDate": "2021-02-18T01:50:48.510Z",
      "content": "<p>You can use pretrained weights now from this great dataset Stanley made! <br>\n<a href=\"https://www.kaggle.com/stanleyjzheng/nfnet-pretrained\" target=\"_blank\">https://www.kaggle.com/stanleyjzheng/nfnet-pretrained</a><br>\nEasy implementation! Give it a try. Thanks <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "rawMarkdown": "You can use pretrained weights now from this great dataset Stanley made! \nhttps://www.kaggle.com/stanleyjzheng/nfnet-pretrained\nEasy implementation! Give it a try. Thanks @stanleyjzheng ",
      "votes": 1
    },
    {
      "id": 1207223,
      "postDate": "2021-02-17T18:48:10.667Z",
      "content": "<p>The pre-trained weights are now <a href=\"https://github.com/deepmind/deepmind-research/tree/master/nfnets\" target=\"_blank\">available</a> from the DeepMind github, as well.</p>",
      "rawMarkdown": "The pre-trained weights are now [available](https://github.com/deepmind/deepmind-research/tree/master/nfnets) from the DeepMind github, as well.",
      "votes": 1,
      "replies": [
        {
          "id": 1207299,
          "postDate": "2021-02-17T19:37:32.740Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1205626,
      "postDate": "2021-02-16T20:30:07.127Z",
      "content": "<p>AGC just added in pytorch-image-models (updated <a href=\"https://www.kaggle.com/stanleyjzheng/timm-nfnet\" target=\"_blank\">my dataset</a> too)</p>\n<p>It's in <code>timm/utils/agc.py</code>.</p>\n<p>Example:</p>\n<pre><code>from utils import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n</code></pre>",
      "rawMarkdown": "AGC just added in pytorch-image-models (updated [my dataset](https://www.kaggle.com/stanleyjzheng/timm-nfnet) too)\n\nIt's in `timm/utils/agc.py`.\n\nExample:\n```python\nfrom utils import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n```\n",
      "votes": 1
    },
    {
      "id": 1200428,
      "postDate": "2021-02-14T16:45:55.937Z",
      "content": "<p>Is something like this available for keras(TF2)?<br>\nOr is this another one of those \"learn pytorch\" scenarios where keras won't get anything till its too late.</p>",
      "rawMarkdown": "Is something like this available for keras(TF2)?\nOr is this another one of those \"learn pytorch\" scenarios where keras won't get anything till its too late.",
      "votes": 2,
      "replies": [
        {
          "id": 1200430,
          "postDate": "2021-02-14T16:47:32.700Z",
          "content": "<p>This is a brand new model released 3 days ago, so unfortuantely no TF implementation yet, a bit of unfortunate timing. </p>",
          "rawMarkdown": "This is a brand new model released 3 days ago, so unfortuantely no TF implementation yet, a bit of unfortunate timing. ",
          "votes": 1
        },
        {
          "id": 1200436,
          "postDate": "2021-02-14T16:51:58.840Z",
          "content": "<p>yeah I know it released like 3 days ago, the team which released this was deepmind, if I am not wrong.<br>\nEven google guys are not using TF2 for research, I think they are using JAX.</p>",
          "rawMarkdown": "yeah I know it released like 3 days ago, the team which released this was deepmind, if I am not wrong.\nEven google guys are not using TF2 for research, I think they are using JAX.",
          "votes": 2
        },
        {
          "id": 1201278,
          "postDate": "2021-02-15T09:17:33.720Z",
          "content": "<p>Hi, I was also looking for keras implementation and I found this could be interesting I still didn't put it into test. <a href=\"https://libraries.io/pypi/nfnets-keras\" target=\"_blank\">https://libraries.io/pypi/nfnets-keras</a>. I will update my reply once I do. </p>",
          "rawMarkdown": "Hi, I was also looking for keras implementation and I found this could be interesting I still didn't put it into test. https://libraries.io/pypi/nfnets-keras. I will update my reply once I do. ",
          "votes": 1
        },
        {
          "id": 1201494,
          "postDate": "2021-02-15T12:38:26.037Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1201838,
          "postDate": "2021-02-15T17:24:22.337Z",
          "content": "<p>Woah, looks legit, will try to implement this and tell the results if it really works. Gonna check the code as well, I have got the gist of the paper, let's see.</p>",
          "rawMarkdown": "Woah, looks legit, will try to implement this and tell the results if it really works. Gonna check the code as well, I have got the gist of the paper, let's see."
        },
        {
          "id": 1204624,
          "postDate": "2021-02-16T09:19:05.627Z",
          "content": "<p>I tried it. But, it is not working, there are some issues. I was able to resolve the first issue which was a simple change. <a href=\"https://github.com/prateekkrjain/nfnets-keras\" target=\"_blank\">Variant error fixed</a>.</p>\n<p>But, after this, there are more issues. The one I am stuck at is:<br>\n<code>TypeError: Failed to convert object of type &lt;class 'tuple'&gt; to Tensor. Contents: (None, 1, 1, 1). Consider casting elements to a supported type.</code></p>\n<p>The error is in line:</p>\n<pre><code>nfnets_keras/nfnet_layers.py:58 call  *\n        r = tf.random.uniform(shape = [batch_size, 1, 1, 1], dtype = x.dtype)\n</code></pre>",
          "rawMarkdown": "I tried it. But, it is not working, there are some issues. I was able to resolve the first issue which was a simple change. [Variant error fixed](https://github.com/prateekkrjain/nfnets-keras).\n\nBut, after this, there are more issues. The one I am stuck at is:\n`TypeError: Failed to convert object of type <class 'tuple'> to Tensor. Contents: (None, 1, 1, 1). Consider casting elements to a supported type.`\n\nThe error is in line:\n```\nnfnets_keras/nfnet_layers.py:58 call  *\n        r = tf.random.uniform(shape = [batch_size, 1, 1, 1], dtype = x.dtype)\n```"
        }
      ]
    },
    {
      "id": 1209979,
      "postDate": "2021-02-19T06:18:06.293Z",
      "content": "<h1>Paper_Summary</h1>\n<p>It's a very interesting paper. Talks a lot about Batch normalization.<br>\n<a href=\"https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\" target=\"_blank\">https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/</a><br>\nWould love to hear your opinions.</p>",
      "rawMarkdown": "#Paper_Summary\nIt's a very interesting paper. Talks a lot about Batch normalization.\nhttps://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\nWould love to hear your opinions."
    },
    {
      "id": 1207225,
      "postDate": "2021-02-17T18:48:10.667Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1200709,
      "author_name": "Matt",
      "author_url": "",
      "post_date": "2021-02-14T22:49:14.847000",
      "content": "<p>Thanks for sharing Stanley!</p>\n<p>if anyone is curious about this method, Yannic Kilcher just released a great video going over the paper:<br>\n<a href=\"https://www.youtube.com/watch?v=rNkHjZtH0RQ\" target=\"_blank\">https://www.youtube.com/watch?v=rNkHjZtH0RQ</a></p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1200448,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2021-02-14T17:17:11.800000",
      "content": "<p>Speed results on RTX3090 using AMP.<br>\nEffnet b0 512x512 (5.3m params), batch size 16 (7.1gb vRAM used): avg 0.2004s/step<br>\nEfficientnet b3 512x512 (12m params), batch size 16 (11.2gb vRAM used): avg 0.3372s/step<br>\nNFNet F0 512x512 (71.5m params), batch size 16 (12.3gb vRAM used): avg 0.6365s/step<br>\nNFNet F1 512x512, (132.6m params), batch size 16 (22.4gb vRAM used): avg 1.0725s/step</p>\n<p>Update:<br>\nAll NFNet above use GeLU.<br>\nNFNet F0 SiLU 512x512 (71.5m params), batch size 16 (10.7gb vRAM used): avg 0.5501s/step</p>\n<p>So SiLU has lower vRAM usage and training latency.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1200497,
          "author_name": "Mohneesh_Sreegirisetty",
          "author_url": "",
          "post_date": "2021-02-14T18:07:34.017000",
          "content": "<p>What data did you use to test these, As you said there wasn't imagenet weights in them?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200501,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-14T18:11:34.943000",
          "content": "<p>This is for the vbd chest x-ray competition, which is my main focus at the moment. Dataset shouldn't matter for speed though. <br>\nThere are no imagenet weights yet so this is a random weights initialization</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200587,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-02-14T19:01:16.930000",
          "content": "<p>It's weird that in paper they claimed <code>NFNet</code> is faster than <code>EffNet</code>. 😳</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200592,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-14T19:05:34.017000",
          "content": "<p>I think it's down to JAX. Gelu GPU on pytorch is apparently slower, and I'm training at 512x512 on both models, while the comparison in the paper is 192x192 vs 224x224 for effnetb0 vs nfnetb0 if I remember correctly.</p>\n<p>I also just found out there is no AGC implementation in the pytorch version, which causes loss to explode at high LR. This may also affect speed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200597,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-02-14T19:13:56.217000",
          "content": "<p>btw did you use the official weight for <code>NfNet</code>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200599,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-14T19:17:55.923000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200601,
          "author_name": "Hiram Coria 🧬",
          "author_url": "",
          "post_date": "2021-02-14T19:22:59.837000",
          "content": "<blockquote>\n  <p>Gelu GPU on pytorch is apparently slower</p>\n</blockquote>\n<p>I lose the reference but <strong>Mish</strong> it's better than <strong>Gelu</strong> for computer vision tasks. Maybe mish could improve the speed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200605,
          "author_name": "JunYong Tong",
          "author_url": "",
          "post_date": "2021-02-14T19:28:16.183000",
          "content": "<p>The F0 model also has more parameters than b0 and b3 combined.</p>\n<p>AGC will probably be implemented at the backward stage for PyTorch? So you will most likely have to implement it yourself.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1200634,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-14T20:02:01.830000",
          "content": "<blockquote>\n  <p>AGC will probably be implemented at the backward stage for PyTorch? So you will most likely have to implement it yourself.</p>\n</blockquote>\n<p>That's a great point, I somehow completely overlooked that, thanks. Will give it a shot.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200654,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2021-02-14T20:25:34.970000",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> are you also using <code>effnetb0</code> without pretrained weight?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1200656,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-14T20:29:26.327000",
          "content": "<p>Yes, correct, though it shouldn't matter speed-wise</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1201774,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-02-15T16:40:45.813000",
          "content": "<p><a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> My experiments seem to line up with yours. I cant seem to use high lr without a Nan loss returning after some time. Looks like AGC is really important to train with out BN</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1208102,
          "author_name": "gao-hongnan",
          "author_url": "",
          "post_date": "2021-02-18T06:50:55.870000",
          "content": "<p>goodness, I need 2x3090 for NFNet man, it is way too slow.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1202011,
      "author_name": "sin",
      "author_url": "",
      "post_date": "2021-02-15T19:49:35.603000",
      "content": "<p>Tried F3s yesterday. Trained on 448 images. It's about 12 minutes/epoch on RTX3090*2 with DDP, AMP. Slow, indeed.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 1202109,
          "author_name": "Yann Majewski",
          "author_url": "",
          "post_date": "2021-02-15T21:05:42.123000",
          "content": "<p>Im thinking that this model doesnt scale well with higher resolutions in comparison to effnets. Or maybe the implementation isnt right/optimized.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1207721,
      "author_name": "Andy Jian Zhou",
      "author_url": "",
      "post_date": "2021-02-18T01:50:48.510000",
      "content": "<p>You can use pretrained weights now from this great dataset Stanley made! <br>\n<a href=\"https://www.kaggle.com/stanleyjzheng/nfnet-pretrained\" target=\"_blank\">https://www.kaggle.com/stanleyjzheng/nfnet-pretrained</a><br>\nEasy implementation! Give it a try. Thanks <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1207223,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-02-17T18:48:10.667000",
      "content": "<p>The pre-trained weights are now <a href=\"https://github.com/deepmind/deepmind-research/tree/master/nfnets\" target=\"_blank\">available</a> from the DeepMind github, as well.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1207299,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-17T19:37:32.740000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1205626,
      "author_name": "Stanley Zheng",
      "author_url": "",
      "post_date": "2021-02-16T20:30:07.127000",
      "content": "<p>AGC just added in pytorch-image-models (updated <a href=\"https://www.kaggle.com/stanleyjzheng/timm-nfnet\" target=\"_blank\">my dataset</a> too)</p>\n<p>It's in <code>timm/utils/agc.py</code>.</p>\n<p>Example:</p>\n<pre><code>from utils import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n</code></pre>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1200428,
      "author_name": "Mohneesh_Sreegirisetty",
      "author_url": "",
      "post_date": "2021-02-14T16:45:55.937000",
      "content": "<p>Is something like this available for keras(TF2)?<br>\nOr is this another one of those \"learn pytorch\" scenarios where keras won't get anything till its too late.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1200430,
          "author_name": "Stanley Zheng",
          "author_url": "",
          "post_date": "2021-02-14T16:47:32.700000",
          "content": "<p>This is a brand new model released 3 days ago, so unfortuantely no TF implementation yet, a bit of unfortunate timing. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1200436,
          "author_name": "Mohneesh_Sreegirisetty",
          "author_url": "",
          "post_date": "2021-02-14T16:51:58.840000",
          "content": "<p>yeah I know it released like 3 days ago, the team which released this was deepmind, if I am not wrong.<br>\nEven google guys are not using TF2 for research, I think they are using JAX.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1201278,
          "author_name": "Omar Bougacha",
          "author_url": "",
          "post_date": "2021-02-15T09:17:33.720000",
          "content": "<p>Hi, I was also looking for keras implementation and I found this could be interesting I still didn't put it into test. <a href=\"https://libraries.io/pypi/nfnets-keras\" target=\"_blank\">https://libraries.io/pypi/nfnets-keras</a>. I will update my reply once I do. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1201494,
          "author_name": "",
          "author_url": "",
          "post_date": "2021-02-15T12:38:26.037000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1201838,
          "author_name": "Mohneesh_Sreegirisetty",
          "author_url": "",
          "post_date": "2021-02-15T17:24:22.337000",
          "content": "<p>Woah, looks legit, will try to implement this and tell the results if it really works. Gonna check the code as well, I have got the gist of the paper, let's see.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1204624,
          "author_name": "Prateek Jain",
          "author_url": "",
          "post_date": "2021-02-16T09:19:05.627000",
          "content": "<p>I tried it. But, it is not working, there are some issues. I was able to resolve the first issue which was a simple change. <a href=\"https://github.com/prateekkrjain/nfnets-keras\" target=\"_blank\">Variant error fixed</a>.</p>\n<p>But, after this, there are more issues. The one I am stuck at is:<br>\n<code>TypeError: Failed to convert object of type &lt;class 'tuple'&gt; to Tensor. Contents: (None, 1, 1, 1). Consider casting elements to a supported type.</code></p>\n<p>The error is in line:</p>\n<pre><code>nfnets_keras/nfnet_layers.py:58 call  *\n        r = tf.random.uniform(shape = [batch_size, 1, 1, 1], dtype = x.dtype)\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1209979,
      "author_name": "Dr. Amritpal Singh",
      "author_url": "",
      "post_date": "2021-02-19T06:18:06.293000",
      "content": "<h1>Paper_Summary</h1>\n<p>It's a very interesting paper. Talks a lot about Batch normalization.<br>\n<a href=\"https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\" target=\"_blank\">https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/</a><br>\nWould love to hear your opinions.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1207225,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-17T18:48:10.667000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1200425": "https://github.com/rwightman/pytorch-image-models/blob/master/timm/models/nfnet.py\n\nNFNet models were recently released in pytorch-image-models. These are only the model definitions and no imagenet pretrained weights are available at the moment, but potentially something to try in the final days of this competition. \n\nHow to use timm on Kaggle: \nAdd this dataset: https://www.kaggle.com/stanleyjzheng/timm-nfnet\n\n```python\nimport sys; sys.path.insert(0,'../input/timm-nfnet')\nimport timm\n\nmodel = timm.create_model('nfnet_f1', pretrained=False)\n```\n\nI'm currently comparing NFNet-F0 to Effnet-B0 on VBD chest xray - I will update this thread with my results in a few hours. NFNet is not in the PyPi package yet, so I had to install from source. Will have results out on speed/vRAM usage soon.\n\nEdit: added dataset for kaggle installation",
    "1200709": "Thanks for sharing Stanley!\n\nif anyone is curious about this method, Yannic Kilcher just released a great video going over the paper:\nhttps://www.youtube.com/watch?v=rNkHjZtH0RQ",
    "1200448": "Speed results on RTX3090 using AMP.\nEffnet b0 512x512 (5.3m params), batch size 16 (7.1gb vRAM used): avg 0.2004s/step\nEfficientnet b3 512x512 (12m params), batch size 16 (11.2gb vRAM used): avg 0.3372s/step\nNFNet F0 512x512 (71.5m params), batch size 16 (12.3gb vRAM used): avg 0.6365s/step\nNFNet F1 512x512, (132.6m params), batch size 16 (22.4gb vRAM used): avg 1.0725s/step\n\nUpdate:\nAll NFNet above use GeLU.\nNFNet F0 SiLU 512x512 (71.5m params), batch size 16 (10.7gb vRAM used): avg 0.5501s/step\n\nSo SiLU has lower vRAM usage and training latency.\n",
    "1202011": "Tried F3s yesterday. Trained on 448 images. It's about 12 minutes/epoch on RTX3090*2 with DDP, AMP. Slow, indeed.",
    "1207721": "You can use pretrained weights now from this great dataset Stanley made! \nhttps://www.kaggle.com/stanleyjzheng/nfnet-pretrained\nEasy implementation! Give it a try. Thanks @stanleyjzheng ",
    "1207223": "The pre-trained weights are now [available](https://github.com/deepmind/deepmind-research/tree/master/nfnets) from the DeepMind github, as well.",
    "1205626": "AGC just added in pytorch-image-models (updated [my dataset](https://www.kaggle.com/stanleyjzheng/timm-nfnet) too)\n\nIt's in `timm/utils/agc.py`.\n\nExample:\n```python\nfrom utils import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n```\n",
    "1200428": "Is something like this available for keras(TF2)?\nOr is this another one of those \"learn pytorch\" scenarios where keras won't get anything till its too late.",
    "1209979": "#Paper_Summary\nIt's a very interesting paper. Talks a lot about Batch normalization.\nhttps://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\nWould love to hear your opinions.",
    "1207225": ""
  }
}