{
  "id": 220268,
  "title": "PRETRAINED - PyTorch/Tensorflow NFNet F* dataset",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/220268",
  "author_name": "",
  "post_date": "2021-02-17T20:19:29.082362700Z",
  "votes": 39,
  "comment_count": 26,
  "views": 0,
  "content": "<p>The pretrained weights for the new State of the Art NFNet F* models were released on DeepMind's <a href=\"https://github.com/deepmind/deepmind-research/tree/master/nfnets\" target=\"_blank\">GitHub repo</a> today in Haiku (Jax) format. I've converted them to PyTorch for ease of use.</p>\n<p>Dataset link: <a href=\"http://www.kaggle.com/stanleyjzheng/nfnet-pretrained\" target=\"_blank\">www.kaggle.com/stanleyjzheng/nfnet-pretrained</a></p>\n<p>You will also need to compile TIMM (pytorch-image-models) from scratch. Here is a snippet:</p>\n<p>Add this <a href=\"https://www.kaggle.com/stanleyjzheng/timm-nfnet\" target=\"_blank\">TIMM dataset</a></p>\n<pre><code>import sys; sys.path.insert(0,'../input/timm-nfnet')\nimport torch\nimport timm\n\nmodel = timm.create_model('nfnet_f0', pretrained=False)\nmodel.load_state_dict(torch.load('../input/nfnet-pretrained/NFNet-f0.pt'))\nmodel.head.fc = nn.Linear(3072, num_classes)\n</code></pre>\n<p>Converting them to Tensorflow is relatively trivial as well: <a href=\"https://dm-haiku.readthedocs.io/en/latest/notebooks/jax2tf.html\" target=\"_blank\">https://dm-haiku.readthedocs.io/en/latest/notebooks/jax2tf.html</a> </p>\n<p>TIMM also has adaptive gradient clipping (AGC) to be used with NFNets:</p>\n<pre><code>from timm.utils.agc import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n</code></pre>\n<p>Edit: Thank you all so much for datasets expert! 4x expert now!!!</p>",
  "messages": [
    {
      "id": "1207383",
      "postDate": "02/17/2021 20:19:29",
      "content": "<p>The pretrained weights for the new State of the Art NFNet F* models were released on DeepMind's <a href=\"https://github.com/deepmind/deepmind-research/tree/master/nfnets\" target=\"_blank\">GitHub repo</a> today in Haiku (Jax) format. I've converted them to PyTorch for ease of use.</p>\n<p>Dataset link: <a href=\"http://www.kaggle.com/stanleyjzheng/nfnet-pretrained\" target=\"_blank\">www.kaggle.com/stanleyjzheng/nfnet-pretrained</a></p>\n<p>You will also need to compile TIMM (pytorch-image-models) from scratch. Here is a snippet:</p>\n<p>Add this <a href=\"https://www.kaggle.com/stanleyjzheng/timm-nfnet\" target=\"_blank\">TIMM dataset</a></p>\n<pre><code>import sys; sys.path.insert(0,'../input/timm-nfnet')\nimport torch\nimport timm\n\nmodel = timm.create_model('nfnet_f0', pretrained=False)\nmodel.load_state_dict(torch.load('../input/nfnet-pretrained/NFNet-f0.pt'))\nmodel.head.fc = nn.Linear(3072, num_classes)\n</code></pre>\n<p>Converting them to Tensorflow is relatively trivial as well: <a href=\"https://dm-haiku.readthedocs.io/en/latest/notebooks/jax2tf.html\" target=\"_blank\">https://dm-haiku.readthedocs.io/en/latest/notebooks/jax2tf.html</a> </p>\n<p>TIMM also has adaptive gradient clipping (AGC) to be used with NFNets:</p>\n<pre><code>from timm.utils.agc import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n</code></pre>\n<p>Edit: Thank you all so much for datasets expert! 4x expert now!!!</p>",
      "rawMarkdown": "The pretrained weights for the new State of the Art NFNet F* models were released on DeepMind's [GitHub repo](https://github.com/deepmind/deepmind-research/tree/master/nfnets) today in Haiku (Jax) format. I've converted them to PyTorch for ease of use.\n\nDataset link: www.kaggle.com/stanleyjzheng/nfnet-pretrained\n\nYou will also need to compile TIMM (pytorch-image-models) from scratch. Here is a snippet:\n\nAdd this [TIMM dataset](https://www.kaggle.com/stanleyjzheng/timm-nfnet)\n\n```python\nimport sys; sys.path.insert(0,'../input/timm-nfnet')\nimport torch\nimport timm\n\nmodel = timm.create_model('nfnet_f0', pretrained=False)\nmodel.load_state_dict(torch.load('../input/nfnet-pretrained/NFNet-f0.pt'))\nmodel.head.fc = nn.Linear(3072, num_classes)\n```\n\nConverting them to Tensorflow is relatively trivial as well: https://dm-haiku.readthedocs.io/en/latest/notebooks/jax2tf.html \n\nTIMM also has adaptive gradient clipping (AGC) to be used with NFNets:\n```python\nfrom timm.utils.agc import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n```\n\nEdit: Thank you all so much for datasets expert! 4x expert now!!!",
      "votes": null
    },
    {
      "id": "1207386",
      "postDate": "02/17/2021 20:21:52",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1207403",
      "postDate": "02/17/2021 20:33:51",
      "content": "<p>How did you convert weights to PyTorch?<br>\nCan you add the other models too?</p>",
      "rawMarkdown": "How did you convert weights to PyTorch?\nCan you add the other models too?",
      "votes": null
    },
    {
      "id": "1207405",
      "postDate": "02/17/2021 20:35:10",
      "content": "<p>I used their utils.py, all the other models are uploading, but 5gb takes a while to upload.</p>",
      "rawMarkdown": "I used their utils.py, all the other models are uploading, but 5gb takes a while to upload.",
      "votes": null
    },
    {
      "id": "1207407",
      "postDate": "02/17/2021 20:36:16",
      "content": "<p>That's great!<br>\nThank you</p>",
      "rawMarkdown": "That's great!\nThank you",
      "votes": null
    },
    {
      "id": "1207411",
      "postDate": "02/17/2021 20:39:35",
      "content": "<p>My pleasure :)</p>",
      "rawMarkdown": "My pleasure :)",
      "votes": null
    },
    {
      "id": "1207521",
      "postDate": "02/17/2021 22:24:46",
      "content": "<p>Thanks Stanley. I didn't read this too carefully, but is AGC required for NFNets?</p>",
      "rawMarkdown": "Thanks Stanley. I didn't read this too carefully, but is AGC required for NFNets?",
      "votes": null
    },
    {
      "id": "1207526",
      "postDate": "02/17/2021 22:30:57",
      "content": "<p>No problem. Yes - it should fix the gradient explosion problems we had earlier. </p>",
      "rawMarkdown": "No problem. Yes - it should fix the gradient explosion problems we had earlier.",
      "votes": null
    },
    {
      "id": "1207685",
      "postDate": "02/18/2021 01:06:32",
      "content": "<p>Thanks Stanley. Will try it out. I am curious how these compare to the other models on this dataset. Looks like F0 has 68M params! </p>",
      "rawMarkdown": "Thanks Stanley. Will try it out. I am curious how these compare to the other models on this dataset. Looks like F0 has 68M params!",
      "votes": null
    },
    {
      "id": "1207988",
      "postDate": "02/18/2021 05:28:12",
      "content": "<p>Really nice work <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a>, thanks a lot man!</p>",
      "rawMarkdown": "Really nice work @stanleyjzheng, thanks a lot man!",
      "votes": null
    },
    {
      "id": "1207992",
      "postDate": "02/18/2021 05:31:44",
      "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> Thanks!</p>\n<p>Your upvote pushed my dataset to bronze, thanks for 4x expert haha :)</p>",
      "rawMarkdown": "reighns Thanks!\n\nYour upvote pushed my dataset to bronze, thanks for 4x expert haha :)",
      "votes": null
    },
    {
      "id": "1207995",
      "postDate": "02/18/2021 05:34:01",
      "content": "<p>LOL hahahahha, nice nice, anyways I am unsure what adaptive gradient clipping does. Would you pointing me to the right resource on it. So in the training pipeline, I just need to add the snippet you showed above to make it work right? Can I use PyTorch's AMP as well?</p>",
      "rawMarkdown": "LOL hahahahha, nice nice, anyways I am unsure what adaptive gradient clipping does. Would you pointing me to the right resource on it. So in the training pipeline, I just need to add the snippet you showed above to make it work right? Can I use PyTorch's AMP as well?",
      "votes": null
    },
    {
      "id": "1208004",
      "postDate": "02/18/2021 05:37:59",
      "content": "<p>Yep, definitely! Just between your optimizer and your loss, you need to add that adaptive clip grad line. For example, mine with AMP is</p>\n<pre><code>scaler.scale(loss).backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\nscaler.step(optimizer)\nscaler.update()\n</code></pre>\n<p>If you're interested in the technical details, the paper is great.<br>\nDiscussion here (it's the same as regular pytorch gradient clipping): <a href=\"https://discuss.pytorch.org/t/proper-way-to-do-gradient-clipping/191\" target=\"_blank\">https://discuss.pytorch.org/t/proper-way-to-do-gradient-clipping/191</a></p>",
      "rawMarkdown": "Yep, definitely! Just between your optimizer and your loss, you need to add that adaptive clip grad line. For example, mine with AMP is\n```python\nscaler.scale(loss).backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\nscaler.step(optimizer)\nscaler.update()\n```\n\nIf you're interested in the technical details, the paper is great.\nDiscussion here (it's the same as regular pytorch gradient clipping): https://discuss.pytorch.org/t/proper-way-to-do-gradient-clipping/191",
      "votes": null
    },
    {
      "id": "1208005",
      "postDate": "02/18/2021 05:40:58",
      "content": "<p>I tried exactly this and my train loss goes to Nan in second epoch and I didn't have a chance to debug this yet. </p>",
      "rawMarkdown": "I tried exactly this and my train loss goes to Nan in second epoch and I didn't have a chance to debug this yet.",
      "votes": null
    },
    {
      "id": "1208008",
      "postDate": "02/18/2021 05:43:18",
      "content": "<p>Interesting - try lower LR maybe? I'm not in this competition but I'm in VBD and it works great. LR  &gt; 5e-4 gives me NaN loss. Sorry, I have no clue</p>",
      "rawMarkdown": "Interesting - try lower LR maybe? I'm not in this competition but I'm in VBD and it works great. LR  > 5e-4 gives me NaN loss. Sorry, I have no clue",
      "votes": null
    },
    {
      "id": "1208024",
      "postDate": "02/18/2021 05:54:56",
      "content": "<p>no worries. Most likely an implementation error on my part. I dont have enough time to use it for this competition but I will definitely use it for others. Thanks!</p>",
      "rawMarkdown": "no worries. Most likely an implementation error on my part. I dont have enough time to use it for this competition but I will definitely use it for others. Thanks!",
      "votes": null
    },
    {
      "id": "1208112",
      "postDate": "02/18/2021 07:00:49",
      "content": "<p><a href=\"https://www.kaggle.com/trushk\" target=\"_blank\">@trushk</a> Even I am getting nans<br>\nI was using lr 1e-4 to 1e-6 with cosine scheduling</p>",
      "rawMarkdown": "trushk Even I am getting nans\nI was using lr 1e-4 to 1e-6 with cosine scheduling",
      "votes": null
    },
    {
      "id": "1208124",
      "postDate": "02/18/2021 07:06:26",
      "content": "<p>Thanks a bunch! I am training with it now. I noticed that Roff (Timm) will prompt a line: No pretrained weights exist for this model. Using random initialization. But I think it is ok since we overrode it with our pretrained weights here.</p>",
      "rawMarkdown": "Thanks a bunch! I am training with it now. I noticed that Roff (Timm) will prompt a line: No pretrained weights exist for this model. Using random initialization. But I think it is ok since we overrode it with our pretrained weights here.",
      "votes": null
    },
    {
      "id": "1209512",
      "postDate": "02/18/2021 23:55:02",
      "content": "<p>Thank you Stanley for sharing this.</p>",
      "rawMarkdown": "Thank you Stanley for sharing this.",
      "votes": null
    },
    {
      "id": "1209974",
      "postDate": "02/19/2021 06:16:44",
      "content": "<h1>Paper_Summary</h1>\n<p>It's a very interesting paper. Talks a lot about Batch normalization.<br>\n<a href=\"https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\" target=\"_blank\">https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/</a><br>\nWould love to hear your opinions.</p>",
      "rawMarkdown": "#Paper_Summary\nIt's a very interesting paper. Talks a lot about Batch normalization.\nhttps://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\nWould love to hear your opinions.",
      "votes": null
    },
    {
      "id": "1210453",
      "postDate": "02/19/2021 12:55:04",
      "content": "<p>Thank you for sharing.</p>",
      "rawMarkdown": "Thank you for sharing.",
      "votes": null
    },
    {
      "id": "1211531",
      "postDate": "02/20/2021 10:03:38",
      "content": "<p>Thank you for sharing such insightful content <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "rawMarkdown": "Thank you for sharing such insightful content @stanleyjzheng",
      "votes": null
    },
    {
      "id": "1216171",
      "postDate": "02/24/2021 06:17:44",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "1219208",
      "postDate": "02/26/2021 14:51:13",
      "content": "<p>nice work!. Thank you.</p>",
      "rawMarkdown": "nice work!. Thank you.",
      "votes": null
    },
    {
      "id": "1225166",
      "postDate": "03/03/2021 11:32:07",
      "content": "<p>I'm not in this competition either, using it elsewhere, but what optimizer do You guys use? As to my first experiments, Adam is likely to produce the instability and NaNs. The original paper suggests Nesterov' momentum SGD with momentum=0.9, weight_decay=2e-5. It seems to be much more stable when in use with NFNets and AGC. Also, try to include a \"warm up\" phase (increasing lr gradually during first, say, 5 epochs). This setup works fine for me, i'm training nf_net_f3 with lr=0.1*BATCH_SIZE/256 (also as paper suggests). Will be glad if You also share Your experiences in this</p>",
      "rawMarkdown": "I'm not in this competition either, using it elsewhere, but what optimizer do You guys use? As to my first experiments, Adam is likely to produce the instability and NaNs. The original paper suggests Nesterov' momentum SGD with momentum=0.9, weight_decay=2e-5. It seems to be much more stable when in use with NFNets and AGC. Also, try to include a \"warm up\" phase (increasing lr gradually during first, say, 5 epochs). This setup works fine for me, i'm training nf_net_f3 with lr=0.1*BATCH_SIZE/256 (also as paper suggests). Will be glad if You also share Your experiences in this",
      "votes": null
    },
    {
      "id": "1239102",
      "postDate": "03/15/2021 13:10:30",
      "content": "<p>The paper states, that AGC is only necessary for large batchsizes and heavy data augmentations</p>\n<blockquote>\n  <p>Using AGC, we can train NF-ResNets stably with larger<br>\n  batch sizes (up to 4096), as well as with very strong data<br>\n  augmentations like RandAugment (Cubuk et al., 2020) for<br>\n  which NF-ResNets without AGC fail to train […] As anticipated, the benefits of using AGC are smaller when the batch size is small. </p>\n</blockquote>",
      "rawMarkdown": "The paper states, that AGC is only necessary for large batchsizes and heavy data augmentations\n\n> Using AGC, we can train NF-ResNets stably with larger\nbatch sizes (up to 4096), as well as with very strong data\naugmentations like RandAugment (Cubuk et al., 2020) for\nwhich NF-ResNets without AGC fail to train [...] As anticipated, the benefits of using AGC are smaller when the batch size is small.",
      "votes": null
    },
    {
      "id": "1621224",
      "postDate": "12/17/2021 14:43:47",
      "content": "<p>great. Nice work~!! <br>\ngod mind ✨✨✨</p>",
      "rawMarkdown": "great. Nice work~!! \ngod mind ✨✨✨",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1207386,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "02/17/2021 20:21:52",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207403,
      "author_name": "debarshichanda",
      "author_url": "",
      "post_date": "02/17/2021 20:33:51",
      "content": "<p>How did you convert weights to PyTorch?<br>\nCan you add the other models too?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207405,
          "author_name": "stanleyjzheng",
          "author_url": "",
          "post_date": "02/17/2021 20:35:10",
          "content": "<p>I used their utils.py, all the other models are uploading, but 5gb takes a while to upload.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207407,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/17/2021 20:36:16",
          "content": "<p>That's great!<br>\nThank you</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207411,
          "author_name": "stanleyjzheng",
          "author_url": "",
          "post_date": "02/17/2021 20:39:35",
          "content": "<p>My pleasure :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207521,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "02/17/2021 22:24:46",
      "content": "<p>Thanks Stanley. I didn't read this too carefully, but is AGC required for NFNets?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207526,
          "author_name": "stanleyjzheng",
          "author_url": "",
          "post_date": "02/17/2021 22:30:57",
          "content": "<p>No problem. Yes - it should fix the gradient explosion problems we had earlier. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1239102,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "03/15/2021 13:10:30",
          "content": "<p>The paper states, that AGC is only necessary for large batchsizes and heavy data augmentations</p>\n<blockquote>\n  <p>Using AGC, we can train NF-ResNets stably with larger<br>\n  batch sizes (up to 4096), as well as with very strong data<br>\n  augmentations like RandAugment (Cubuk et al., 2020) for<br>\n  which NF-ResNets without AGC fail to train […] As anticipated, the benefits of using AGC are smaller when the batch size is small. </p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1207685,
      "author_name": "trushk",
      "author_url": "",
      "post_date": "02/18/2021 01:06:32",
      "content": "<p>Thanks Stanley. Will try it out. I am curious how these compare to the other models on this dataset. Looks like F0 has 68M params! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1207988,
      "author_name": "reighns",
      "author_url": "",
      "post_date": "02/18/2021 05:28:12",
      "content": "<p>Really nice work <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a>, thanks a lot man!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1207992,
          "author_name": "stanleyjzheng",
          "author_url": "",
          "post_date": "02/18/2021 05:31:44",
          "content": "<p><a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> Thanks!</p>\n<p>Your upvote pushed my dataset to bronze, thanks for 4x expert haha :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1207995,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "02/18/2021 05:34:01",
          "content": "<p>LOL hahahahha, nice nice, anyways I am unsure what adaptive gradient clipping does. Would you pointing me to the right resource on it. So in the training pipeline, I just need to add the snippet you showed above to make it work right? Can I use PyTorch's AMP as well?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208004,
          "author_name": "stanleyjzheng",
          "author_url": "",
          "post_date": "02/18/2021 05:37:59",
          "content": "<p>Yep, definitely! Just between your optimizer and your loss, you need to add that adaptive clip grad line. For example, mine with AMP is</p>\n<pre><code>scaler.scale(loss).backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\nscaler.step(optimizer)\nscaler.update()\n</code></pre>\n<p>If you're interested in the technical details, the paper is great.<br>\nDiscussion here (it's the same as regular pytorch gradient clipping): <a href=\"https://discuss.pytorch.org/t/proper-way-to-do-gradient-clipping/191\" target=\"_blank\">https://discuss.pytorch.org/t/proper-way-to-do-gradient-clipping/191</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208005,
          "author_name": "trushk",
          "author_url": "",
          "post_date": "02/18/2021 05:40:58",
          "content": "<p>I tried exactly this and my train loss goes to Nan in second epoch and I didn't have a chance to debug this yet. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208008,
          "author_name": "stanleyjzheng",
          "author_url": "",
          "post_date": "02/18/2021 05:43:18",
          "content": "<p>Interesting - try lower LR maybe? I'm not in this competition but I'm in VBD and it works great. LR  &gt; 5e-4 gives me NaN loss. Sorry, I have no clue</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208024,
          "author_name": "trushk",
          "author_url": "",
          "post_date": "02/18/2021 05:54:56",
          "content": "<p>no worries. Most likely an implementation error on my part. I dont have enough time to use it for this competition but I will definitely use it for others. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208112,
          "author_name": "debarshichanda",
          "author_url": "",
          "post_date": "02/18/2021 07:00:49",
          "content": "<p><a href=\"https://www.kaggle.com/trushk\" target=\"_blank\">@trushk</a> Even I am getting nans<br>\nI was using lr 1e-4 to 1e-6 with cosine scheduling</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1208124,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "02/18/2021 07:06:26",
          "content": "<p>Thanks a bunch! I am training with it now. I noticed that Roff (Timm) will prompt a line: No pretrained weights exist for this model. Using random initialization. But I think it is ok since we overrode it with our pretrained weights here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1225166,
          "author_name": "snufkin77",
          "author_url": "",
          "post_date": "03/03/2021 11:32:07",
          "content": "<p>I'm not in this competition either, using it elsewhere, but what optimizer do You guys use? As to my first experiments, Adam is likely to produce the instability and NaNs. The original paper suggests Nesterov' momentum SGD with momentum=0.9, weight_decay=2e-5. It seems to be much more stable when in use with NFNets and AGC. Also, try to include a \"warm up\" phase (increasing lr gradually during first, say, 5 epochs). This setup works fine for me, i'm training nf_net_f3 with lr=0.1*BATCH_SIZE/256 (also as paper suggests). Will be glad if You also share Your experiences in this</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1209512,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/18/2021 23:55:02",
      "content": "<p>Thank you Stanley for sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1209974,
      "author_name": "amritpal333",
      "author_url": "",
      "post_date": "02/19/2021 06:16:44",
      "content": "<h1>Paper_Summary</h1>\n<p>It's a very interesting paper. Talks a lot about Batch normalization.<br>\n<a href=\"https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\" target=\"_blank\">https://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/</a><br>\nWould love to hear your opinions.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1210453,
      "author_name": "piyalitt",
      "author_url": "",
      "post_date": "02/19/2021 12:55:04",
      "content": "<p>Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211531,
      "author_name": "alifrahman",
      "author_url": "",
      "post_date": "02/20/2021 10:03:38",
      "content": "<p>Thank you for sharing such insightful content <a href=\"https://www.kaggle.com/stanleyjzheng\" target=\"_blank\">@stanleyjzheng</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1216171,
      "author_name": "himanshu1999",
      "author_url": "",
      "post_date": "02/24/2021 06:17:44",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1219208,
      "author_name": "mssjss",
      "author_url": "",
      "post_date": "02/26/2021 14:51:13",
      "content": "<p>nice work!. Thank you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1621224,
      "author_name": "kimalpha",
      "author_url": "",
      "post_date": "12/17/2021 14:43:47",
      "content": "<p>great. Nice work~!! <br>\ngod mind ✨✨✨</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1207383": "The pretrained weights for the new State of the Art NFNet F* models were released on DeepMind's [GitHub repo](https://github.com/deepmind/deepmind-research/tree/master/nfnets) today in Haiku (Jax) format. I've converted them to PyTorch for ease of use.\n\nDataset link: www.kaggle.com/stanleyjzheng/nfnet-pretrained\n\nYou will also need to compile TIMM (pytorch-image-models) from scratch. Here is a snippet:\n\nAdd this [TIMM dataset](https://www.kaggle.com/stanleyjzheng/timm-nfnet)\n\n```python\nimport sys; sys.path.insert(0,'../input/timm-nfnet')\nimport torch\nimport timm\n\nmodel = timm.create_model('nfnet_f0', pretrained=False)\nmodel.load_state_dict(torch.load('../input/nfnet-pretrained/NFNet-f0.pt'))\nmodel.head.fc = nn.Linear(3072, num_classes)\n```\n\nConverting them to Tensorflow is relatively trivial as well: https://dm-haiku.readthedocs.io/en/latest/notebooks/jax2tf.html \n\nTIMM also has adaptive gradient clipping (AGC) to be used with NFNets:\n```python\nfrom timm.utils.agc import adaptive_clip_grad\n\nloss.backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\noptimizer.step()\n```\n\nEdit: Thank you all so much for datasets expert! 4x expert now!!!",
    "1207386": "Thanks for sharing!",
    "1207403": "How did you convert weights to PyTorch?\nCan you add the other models too?",
    "1207405": "I used their utils.py, all the other models are uploading, but 5gb takes a while to upload.",
    "1207407": "That's great!\nThank you",
    "1207411": "My pleasure :)",
    "1207521": "Thanks Stanley. I didn't read this too carefully, but is AGC required for NFNets?",
    "1207526": "No problem. Yes - it should fix the gradient explosion problems we had earlier.",
    "1207685": "Thanks Stanley. Will try it out. I am curious how these compare to the other models on this dataset. Looks like F0 has 68M params!",
    "1207988": "Really nice work @stanleyjzheng, thanks a lot man!",
    "1207992": "reighns Thanks!\n\nYour upvote pushed my dataset to bronze, thanks for 4x expert haha :)",
    "1207995": "LOL hahahahha, nice nice, anyways I am unsure what adaptive gradient clipping does. Would you pointing me to the right resource on it. So in the training pipeline, I just need to add the snippet you showed above to make it work right? Can I use PyTorch's AMP as well?",
    "1208004": "Yep, definitely! Just between your optimizer and your loss, you need to add that adaptive clip grad line. For example, mine with AMP is\n```python\nscaler.scale(loss).backward()\nadaptive_clip_grad(model.parameters(), clip_factor=0.01, eps=1e-3, norm_type=2.0)\nscaler.step(optimizer)\nscaler.update()\n```\n\nIf you're interested in the technical details, the paper is great.\nDiscussion here (it's the same as regular pytorch gradient clipping): https://discuss.pytorch.org/t/proper-way-to-do-gradient-clipping/191",
    "1208005": "I tried exactly this and my train loss goes to Nan in second epoch and I didn't have a chance to debug this yet.",
    "1208008": "Interesting - try lower LR maybe? I'm not in this competition but I'm in VBD and it works great. LR  > 5e-4 gives me NaN loss. Sorry, I have no clue",
    "1208024": "no worries. Most likely an implementation error on my part. I dont have enough time to use it for this competition but I will definitely use it for others. Thanks!",
    "1208112": "trushk Even I am getting nans\nI was using lr 1e-4 to 1e-6 with cosine scheduling",
    "1208124": "Thanks a bunch! I am training with it now. I noticed that Roff (Timm) will prompt a line: No pretrained weights exist for this model. Using random initialization. But I think it is ok since we overrode it with our pretrained weights here.",
    "1209512": "Thank you Stanley for sharing this.",
    "1209974": "#Paper_Summary\nIt's a very interesting paper. Talks a lot about Batch normalization.\nhttps://www.linkedin.com/posts/amritpal-singh-001_paperabrsummary-computervision-google-activity-6767776421348741120-SFos/\nWould love to hear your opinions.",
    "1210453": "Thank you for sharing.",
    "1211531": "Thank you for sharing such insightful content @stanleyjzheng",
    "1216171": "Thanks for sharing!",
    "1219208": "nice work!. Thank you.",
    "1225166": "I'm not in this competition either, using it elsewhere, but what optimizer do You guys use? As to my first experiments, Adam is likely to produce the instability and NaNs. The original paper suggests Nesterov' momentum SGD with momentum=0.9, weight_decay=2e-5. It seems to be much more stable when in use with NFNets and AGC. Also, try to include a \"warm up\" phase (increasing lr gradually during first, say, 5 epochs). This setup works fine for me, i'm training nf_net_f3 with lr=0.1*BATCH_SIZE/256 (also as paper suggests). Will be glad if You also share Your experiences in this",
    "1239102": "The paper states, that AGC is only necessary for large batchsizes and heavy data augmentations\n\n> Using AGC, we can train NF-ResNets stably with larger\nbatch sizes (up to 4096), as well as with very strong data\naugmentations like RandAugment (Cubuk et al., 2020) for\nwhich NF-ResNets without AGC fail to train [...] As anticipated, the benefits of using AGC are smaller when the batch size is small.",
    "1621224": "great. Nice work~!! \ngod mind ✨✨✨"
  },
  "source": "meta"
}