{
  "id": 222563,
  "title": "Your Experiences With NFNet Models",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/222563",
  "author_name": "",
  "post_date": "2021-02-27T21:51:45.932149700Z",
  "votes": 18,
  "comment_count": 39,
  "views": 0,
  "content": "<p>I have been trying to get some decent scores from recently released nfnets. I did get some promising results but in general it's still lacking behind some good old model architectures. I tried to stick on official paper generally but implementing some parts leading decreased cv scores. I have published one of my trials here as notebook if you wanna check:</p>\n<p><a href=\"https://www.kaggle.com/datafan07/ranzcr-nfnets-tutorial-single-fold-training\" target=\"_blank\">https://www.kaggle.com/datafan07/ranzcr-nfnets-tutorial-single-fold-training</a></p>\n<p>For example using Adaptive Gradient Clipping didn't make a difference on training part. Maybe changing lambda value to something else helps…</p>\n<p>What you guys think? Have any of you tried these architectures for this competition? What are your observations?</p>\n<p>Good luck all…</p>",
  "messages": [
    {
      "id": "1220342",
      "postDate": "02/27/2021 21:51:45",
      "content": "<p>I have been trying to get some decent scores from recently released nfnets. I did get some promising results but in general it's still lacking behind some good old model architectures. I tried to stick on official paper generally but implementing some parts leading decreased cv scores. I have published one of my trials here as notebook if you wanna check:</p>\n<p><a href=\"https://www.kaggle.com/datafan07/ranzcr-nfnets-tutorial-single-fold-training\" target=\"_blank\">https://www.kaggle.com/datafan07/ranzcr-nfnets-tutorial-single-fold-training</a></p>\n<p>For example using Adaptive Gradient Clipping didn't make a difference on training part. Maybe changing lambda value to something else helps…</p>\n<p>What you guys think? Have any of you tried these architectures for this competition? What are your observations?</p>\n<p>Good luck all…</p>",
      "rawMarkdown": "I have been trying to get some decent scores from recently released nfnets. I did get some promising results but in general it's still lacking behind some good old model architectures. I tried to stick on official paper generally but implementing some parts leading decreased cv scores. I have published one of my trials here as notebook if you wanna check:\n\nhttps://www.kaggle.com/datafan07/ranzcr-nfnets-tutorial-single-fold-training\n\nFor example using Adaptive Gradient Clipping didn't make a difference on training part. Maybe changing lambda value to something else helps...\n\nWhat you guys think? Have any of you tried these architectures for this competition? What are your observations?\n\nGood luck all...",
      "votes": null
    },
    {
      "id": "1220419",
      "postDate": "02/28/2021 00:59:11",
      "content": "<p>Just hard models to train in my experience.. Ill stick to my good old resnets</p>",
      "rawMarkdown": "Just hard models to train in my experience.. Ill stick to my good old resnets",
      "votes": null
    },
    {
      "id": "1220421",
      "postDate": "02/28/2021 01:06:38",
      "content": "<p>I've been seeing some people claiming promising results, but probably going to wait to see more conclusive evidence since they're huge models</p>",
      "rawMarkdown": "I've been seeing some people claiming promising results, but probably going to wait to see more conclusive evidence since they're huge models",
      "votes": null
    },
    {
      "id": "1220451",
      "postDate": "02/28/2021 02:20:12",
      "content": "<p>I think <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a> has also observed that AGC isn't making much difference over using regular gradient clipping…<br>\n<a href=\"https://twitter.com/wightmanr/status/1361746367769546753\" target=\"_blank\">https://twitter.com/wightmanr/status/1361746367769546753</a></p>",
      "rawMarkdown": "I think @rwightman has also observed that AGC isn't making much difference over using regular gradient clipping...\nhttps://twitter.com/wightmanr/status/1361746367769546753",
      "votes": null
    },
    {
      "id": "1220727",
      "postDate": "02/28/2021 09:52:00",
      "content": "<p>i have spend most of my times training nfnets f0 model and no matter what i try i can't get cv 0.95 on first fold, i would like to know what is the maximum roc auc that one can get in this competition? sadly the best single model thread is kind of frozen :(<br>\ndoes image size matter much here? i tried f0 model with different image sizes and didn't observe improvement  in roc auc</p>",
      "rawMarkdown": "i have spend most of my times training nfnets f0 model and no matter what i try i can't get cv 0.95 on first fold, i would like to know what is the maximum roc auc that one can get in this competition? sadly the best single model thread is kind of frozen :(\ndoes image size matter much here? i tried f0 model with different image sizes and didn't observe improvement  in roc auc",
      "votes": null
    },
    {
      "id": "1220728",
      "postDate": "02/28/2021 09:52:07",
      "content": "<p>That's good to know thanks for sharing</p>",
      "rawMarkdown": "That's good to know thanks for sharing",
      "votes": null
    },
    {
      "id": "1220889",
      "postDate": "02/28/2021 13:00:39",
      "content": "<p>Image size did matter in my trials, but idk what image size you have been using with f0's?</p>",
      "rawMarkdown": "Image size did matter in my trials, but idk what image size you have been using with f0's?",
      "votes": null
    },
    {
      "id": "1220893",
      "postDate": "02/28/2021 13:05:19",
      "content": "<p>i tried 512,576,384,456 but nothing helps me to get 0.95+ on first fold out of 5,,were you able to get more than that while using nfnets f0? you used much bigger image size?</p>",
      "rawMarkdown": "i tried 512,576,384,456 but nothing helps me to get 0.95+ on first fold out of 5,,were you able to get more than that while using nfnets f0? you used much bigger image size?",
      "votes": null
    },
    {
      "id": "1221226",
      "postDate": "02/28/2021 19:10:20",
      "content": "<p>I tried a few folds using F1, but the performance was worse than ResNet based models</p>",
      "rawMarkdown": "I tried a few folds using F1, but the performance was worse than ResNet based models",
      "votes": null
    },
    {
      "id": "1221233",
      "postDate": "02/28/2021 19:17:15",
      "content": "<p>I usually choose pretrained image test sizes related to f number, but using smaller images on same f's reduced my cv score so I assumed increasing them would get better scores. But didn't try above 512 I think…</p>",
      "rawMarkdown": "I usually choose pretrained image test sizes related to f number, but using smaller images on same f's reduced my cv score so I assumed increasing them would get better scores. But didn't try above 512 I think...",
      "votes": null
    },
    {
      "id": "1221244",
      "postDate": "02/28/2021 19:38:09",
      "content": "<p>How long does it take to converge?<br>\nI see it is very parameter sensitive, small change in lr hurt in cv badly,i was getting close to 0.95 cv but not MORE than that after 30-40 epochs iirc<br>\nYou ever got 0.95+ using nfnets?</p>",
      "rawMarkdown": "How long does it take to converge?\nI see it is very parameter sensitive, small change in lr hurt in cv badly,i was getting close to 0.95 cv but not MORE than that after 30-40 epochs iirc\nYou ever got 0.95+ using nfnets?",
      "votes": null
    },
    {
      "id": "1221245",
      "postDate": "02/28/2021 19:39:27",
      "content": "<p>Just curious <br>\n\"Is everyone's best model resnet200d or resnet50d like 'd' models?\"<br>\nOr other models working well too?<br>\nHow about tf models?<br>\nWe Don't have any 'd' models there<br>\nWhat's the best cv one can get with tensorflow in this competition? <br>\nJust curious </p>",
      "rawMarkdown": "Just curious \n\"Is everyone's best model resnet200d or resnet50d like 'd' models?\"\nOr other models working well too?\nHow about tf models?\nWe Don't have any 'd' models there\nWhat's the best cv one can get with tensorflow in this competition? \nJust curious",
      "votes": null
    },
    {
      "id": "1221249",
      "postDate": "02/28/2021 19:42:25",
      "content": "<blockquote>\n  <p>Have any of you tried these architectures for this competition? What are your observations?</p>\n</blockquote>\n<p>Slower to train and worse performance than EfficientNets. Perhaps torch implementation is not perfect yet, or those models dont transfer to competition domain well.</p>",
      "rawMarkdown": "> Have any of you tried these architectures for this competition? What are your observations?\n\nSlower to train and worse performance than EfficientNets. Perhaps torch implementation is not perfect yet, or those models dont transfer to competition domain well.",
      "votes": null
    },
    {
      "id": "1221304",
      "postDate": "02/28/2021 21:33:50",
      "content": "<p>My best models so far are using ResNet200d, but I think the main reason that it is so popular is that it trains well on GPUs, i.e. good performance with large image sizes (and therefore small batch sizes) and relatively quick training. For me, it outperforms EfficientNet-B4 for the same amount of compute time using 512x512 images.</p>\n<p>So far I haven't got anything else to work as well within the constraints of the VRAM in my machine (yet!) ¯_(ツ)_/¯ </p>",
      "rawMarkdown": "My best models so far are using ResNet200d, but I think the main reason that it is so popular is that it trains well on GPUs, i.e. good performance with large image sizes (and therefore small batch sizes) and relatively quick training. For me, it outperforms EfficientNet-B4 for the same amount of compute time using 512x512 images.\n\nSo far I haven't got anything else to work as well within the constraints of the VRAM in my machine (yet!) ¯\\_(ツ)_/¯",
      "votes": null
    },
    {
      "id": "1221307",
      "postDate": "02/28/2021 21:38:08",
      "content": "<p>That's difficult to say as it is related to your LR scheduler settings…..</p>",
      "rawMarkdown": "That's difficult to say as it is related to your LR scheduler settings.....",
      "votes": null
    },
    {
      "id": "1221318",
      "postDate": "02/28/2021 22:03:57",
      "content": "<p>Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(</p>",
      "rawMarkdown": "Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(",
      "votes": null
    },
    {
      "id": "1221561",
      "postDate": "03/01/2021 05:11:15",
      "content": "<p>Current NFNets in timm are full of GELU activation which make them insanely slow ☹️. I gave f2 and f3 a try few weeks ago and found that at 640x640, they were slower than b6/ b7 at 1024. Since then, I stopped exploring them. <br>\nWith similar training schedule and heavy augmentation (drop path + randaug), EffNets gave me better results than ResNet (the d version) in this challenge. </p>",
      "rawMarkdown": "Current NFNets in timm are full of GELU activation which make them insanely slow ☹️. I gave f2 and f3 a try few weeks ago and found that at 640x640, they were slower than b6/ b7 at 1024. Since then, I stopped exploring them. \nWith similar training schedule and heavy augmentation (drop path + randaug), EffNets gave me better results than ResNet (the d version) in this challenge.",
      "votes": null
    },
    {
      "id": "1222012",
      "postDate": "03/01/2021 13:21:19",
      "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> that's interesting, I feel like most of the people in this comp are still using ResNet's. It's good to see Effnet alternative… </p>",
      "rawMarkdown": "andy2709 that's interesting, I feel like most of the people in this comp are still using ResNet's. It's good to see Effnet alternative...",
      "votes": null
    },
    {
      "id": "1222379",
      "postDate": "03/01/2021 18:26:30",
      "content": "<blockquote>\n  <p>Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(</p>\n</blockquote>\n<p>Did you enable bfloat16 ?  It doesn't work for some models </p>",
      "rawMarkdown": "> Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(\n\nDid you enable bfloat16 ?  It doesn't work for some models",
      "votes": null
    },
    {
      "id": "1222594",
      "postDate": "03/01/2021 23:56:11",
      "content": "<p>Yeah bfloat16 was enabled, all model works on torch xla but couldn’t make  resnet200d working, also each epoch training takes a lot of time<br>\nI will try without bfloat16. Thank you for letting me know :)</p>",
      "rawMarkdown": "Yeah bfloat16 was enabled, all model works on torch xla but couldn’t make  resnet200d working, also each epoch training takes a lot of time\nI will try without bfloat16. Thank you for letting me know :)",
      "votes": null
    },
    {
      "id": "1222677",
      "postDate": "03/02/2021 03:13:28",
      "content": "<p>resnet101 give me cv near to 96, have no too much luck with efficientnet, and the nfnnet too expensive, so i forgive use it</p>",
      "rawMarkdown": "resnet101 give me cv near to 96, have no too much luck with efficientnet, and the nfnnet too expensive, so i forgive use it",
      "votes": null
    },
    {
      "id": "1222872",
      "postDate": "03/02/2021 08:30:58",
      "content": "<p>Some models didn't work for me the last time I used xla with bfloat16 .</p>\n<p>I had to disable it ( and reduce the batch size + longer training  unfortunately )</p>",
      "rawMarkdown": "Some models didn't work for me the last time I used xla with bfloat16 .\n\nI had to disable it ( and reduce the batch size + longer training  unfortunately )",
      "votes": null
    },
    {
      "id": "1222887",
      "postDate": "03/02/2021 08:58:09",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I have been exclusively using Tensorflow for classical models… Best I could do was 0.955 on EfficientNet B7 and 0.958 using ensemble. But hard to train it on GPU, and this drains TPU hours like anything too 😆</p>",
      "rawMarkdown": "mobassir I have been exclusively using Tensorflow for classical models... Best I could do was 0.955 on EfficientNet B7 and 0.958 using ensemble. But hard to train it on GPU, and this drains TPU hours like anything too 😆",
      "votes": null
    },
    {
      "id": "1222925",
      "postDate": "03/02/2021 09:57:02",
      "content": "<p><a href=\"https://www.kaggle.com/manabendrarout\" target=\"_blank\">@manabendrarout</a> interesting,you sure your validation strategy is good? i tried to improve ragnar's tf baseline,i couldn't get 0.95 cv without 'n' step TTA, is your 0.955 CV without TTA or with TTA? what lb you get with that? interesting</p>",
      "rawMarkdown": "manabendrarout interesting,you sure your validation strategy is good? i tried to improve ragnar's tf baseline,i couldn't get 0.95 cv without 'n' step TTA, is your 0.955 CV without TTA or with TTA? what lb you get with that? interesting",
      "votes": null
    },
    {
      "id": "1222941",
      "postDate": "03/02/2021 10:12:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> ,<br>\nThis is the config for my 0.955 CV (Single Model):-<br>\nModel:- EfficientNet B7 (with two extra dense layers and dropouts added after)<br>\nDrop Connect Rate:- 0.3<br>\nImage Size:- 600 😅<br>\nInitial Weights:- Noisy Student<br>\nLR Scheduler:- Decrease LR on plateau by 10x<br>\nOptimizer:- Adam<br>\nLoss:- Binary Cross entropy<br>\nAugmentation:- Basic (Random left-right, top-bottom flips and brightness change)<br>\nPlatform:- Tensorflow<br>\nTTA:- None</p>\n<p>This gets me a 0.946 LB.<br>\nStill working on improving it… But the gains have almost stagnated and processing time is so high that there are lot of diminishing returns there.</p>",
      "rawMarkdown": "Hi @mobassir ,\nThis is the config for my 0.955 CV (Single Model):-\nModel:- EfficientNet B7 (with two extra dense layers and dropouts added after)\nDrop Connect Rate:- 0.3\nImage Size:- 600 😅\nInitial Weights:- Noisy Student\nLR Scheduler:- Decrease LR on plateau by 10x\nOptimizer:- Adam\nLoss:- Binary Cross entropy\nAugmentation:- Basic (Random left-right, top-bottom flips and brightness change)\nPlatform:- Tensorflow\nTTA:- None\n\nThis gets me a 0.946 LB.\nStill working on improving it... But the gains have almost stagnated and processing time is so high that there are lot of diminishing returns there.",
      "votes": null
    },
    {
      "id": "1224655",
      "postDate": "03/03/2021 00:20:02",
      "content": "<p>I have been able to get LB.96+ using nfnets implemented by timm. It just seems to be more sensitive to hyperparameters than networks such as efficientnet and resnet.</p>",
      "rawMarkdown": "I have been able to get LB.96+ using nfnets implemented by timm. It just seems to be more sensitive to hyperparameters than networks such as efficientnet and resnet.",
      "votes": null
    },
    {
      "id": "1225184",
      "postDate": "03/03/2021 11:54:21",
      "content": "<p>Yeah they are, I also get over lb 96.5's but have to say they are pretty heavy models…</p>",
      "rawMarkdown": "Yeah they are, I also get over lb 96.5's but have to say they are pretty heavy models...",
      "votes": null
    },
    {
      "id": "1228732",
      "postDate": "03/06/2021 17:44:19",
      "content": "<p>Exactly for me, ResNet had better performance than efficient-net, Nfnets, however, end up using too much computing power so am not able to use them</p>",
      "rawMarkdown": "Exactly for me, ResNet had better performance than efficient-net, Nfnets, however, end up using too much computing power so am not able to use them",
      "votes": null
    },
    {
      "id": "1229891",
      "postDate": "03/07/2021 16:59:56",
      "content": "<p>Yeah they are consuming computing power a lot…</p>",
      "rawMarkdown": "Yeah they are consuming computing power a lot...",
      "votes": null
    },
    {
      "id": "1230147",
      "postDate": "03/07/2021 20:04:46",
      "content": "<p>what is the good way to start with CNN?</p>",
      "rawMarkdown": "what is the good way to start with CNN?",
      "votes": null
    },
    {
      "id": "1237729",
      "postDate": "03/14/2021 12:04:22",
      "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> Just to confirm it's possible to get CV close to 0.97 with nfnets, but still doesn't change the fact that they are pretty heavy</p>",
      "rawMarkdown": "mobassir Just to confirm it's possible to get CV close to 0.97 with nfnets, but still doesn't change the fact that they are pretty heavy",
      "votes": null
    },
    {
      "id": "1237739",
      "postDate": "03/14/2021 12:11:23",
      "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> i guess it's very hyperparameter sensitive,i tried onecycleLR,deepmemory optimizer,nfnetsf-1, and lr ranges between 0.001-0.0011 and some other lr as well,then trained for around 50 hours at a stretch but couldn't get such impressive CV,if your 0.97 is not very hyperparameter sensitive then it's a great model,thanks</p>",
      "rawMarkdown": "datafan07 i guess it's very hyperparameter sensitive,i tried onecycleLR,deepmemory optimizer,nfnetsf-1, and lr ranges between 0.001-0.0011 and some other lr as well,then trained for around 50 hours at a stretch but couldn't get such impressive CV,if your 0.97 is not very hyperparameter sensitive then it's a great model,thanks",
      "votes": null
    },
    {
      "id": "1239220",
      "postDate": "03/15/2021 14:44:27",
      "content": "<p>There are many ways. One that worked for me is to try to implement a simple CNN (using numpy). Once you know the basics, read some tutorials from popular libraries. Here is one example: <a href=\"https://www.tensorflow.org/tutorials/images/cnn\" target=\"_blank\">https://www.tensorflow.org/tutorials/images/cnn</a>.</p>\n<p>Happy learning!</p>",
      "rawMarkdown": "There are many ways. One that worked for me is to try to implement a simple CNN (using numpy). Once you know the basics, read some tutorials from popular libraries. Here is one example: https://www.tensorflow.org/tutorials/images/cnn.\n\nHappy learning!",
      "votes": null
    },
    {
      "id": "1239791",
      "postDate": "03/16/2021 02:51:10",
      "content": "<p>I got LB: 0.969 with nfnet f1 (5 fold average) , however I couldn't get similar score with nfnet f0, f2. (~LB: 0.967)</p>",
      "rawMarkdown": "I got LB: 0.969 with nfnet f1 (5 fold average) , however I couldn't get similar score with nfnet f0, f2. (~LB: 0.967)",
      "votes": null
    },
    {
      "id": "1239883",
      "postDate": "03/16/2021 05:34:20",
      "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> My Rainforest teammate <a href=\"https://www.kaggle.com/vineeth1999\" target=\"_blank\">@vineeth1999</a> implemented it for Rainforest . It didnt give us satisfactory results . I didnt try any new experiments for RANCZR . I only went through public kernels and discussion forums ..</p>",
      "rawMarkdown": "datafan07 My Rainforest teammate @vineeth1999 implemented it for Rainforest . It didnt give us satisfactory results . I didnt try any new experiments for RANCZR . I only went through public kernels and discussion forums ..",
      "votes": null
    },
    {
      "id": "1240417",
      "postDate": "03/16/2021 12:04:59",
      "content": "<p>That's nice, how was your cv and resolution if you don't mind sharing?</p>",
      "rawMarkdown": "That's nice, how was your cv and resolution if you don't mind sharing?",
      "votes": null
    },
    {
      "id": "1240419",
      "postDate": "03/16/2021 12:05:56",
      "content": "<p>My single model comes from nfnets for this competition interestingly. Followed by resnet200d…</p>",
      "rawMarkdown": "My single model comes from nfnets for this competition interestingly. Followed by resnet200d...",
      "votes": null
    },
    {
      "id": "1240491",
      "postDate": "03/16/2021 12:48:51",
      "content": "<p>Can you share what CV and LB score you achieved with nfnet as single?</p>",
      "rawMarkdown": "Can you share what CV and LB score you achieved with nfnet as single?",
      "votes": null
    },
    {
      "id": "1240547",
      "postDate": "03/16/2021 13:33:22",
      "content": "<p>img size: 736<br>\nCV: 0.962</p>",
      "rawMarkdown": "img size: 736\nCV: 0.962",
      "votes": null
    },
    {
      "id": "1241202",
      "postDate": "03/17/2021 00:14:28",
      "content": "<p>Just a note, nfnet saved my team from shakeup…</p>",
      "rawMarkdown": "Just a note, nfnet saved my team from shakeup...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1220419,
      "author_name": "yannmajewski",
      "author_url": "",
      "post_date": "02/28/2021 00:59:11",
      "content": "<p>Just hard models to train in my experience.. Ill stick to my good old resnets</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1220421,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/28/2021 01:06:38",
      "content": "<p>I've been seeing some people claiming promising results, but probably going to wait to see more conclusive evidence since they're huge models</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1220451,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "02/28/2021 02:20:12",
      "content": "<p>I think <a href=\"https://www.kaggle.com/rwightman\" target=\"_blank\">@rwightman</a> has also observed that AGC isn't making much difference over using regular gradient clipping…<br>\n<a href=\"https://twitter.com/wightmanr/status/1361746367769546753\" target=\"_blank\">https://twitter.com/wightmanr/status/1361746367769546753</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1220728,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "02/28/2021 09:52:07",
          "content": "<p>That's good to know thanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1220727,
      "author_name": "mobassir",
      "author_url": "",
      "post_date": "02/28/2021 09:52:00",
      "content": "<p>i have spend most of my times training nfnets f0 model and no matter what i try i can't get cv 0.95 on first fold, i would like to know what is the maximum roc auc that one can get in this competition? sadly the best single model thread is kind of frozen :(<br>\ndoes image size matter much here? i tried f0 model with different image sizes and didn't observe improvement  in roc auc</p>",
      "votes": null,
      "replies": [
        {
          "id": 1220889,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "02/28/2021 13:00:39",
          "content": "<p>Image size did matter in my trials, but idk what image size you have been using with f0's?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1220893,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/28/2021 13:05:19",
          "content": "<p>i tried 512,576,384,456 but nothing helps me to get 0.95+ on first fold out of 5,,were you able to get more than that while using nfnets f0? you used much bigger image size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221233,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "02/28/2021 19:17:15",
          "content": "<p>I usually choose pretrained image test sizes related to f number, but using smaller images on same f's reduced my cv score so I assumed increasing them would get better scores. But didn't try above 512 I think…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221244,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/28/2021 19:38:09",
          "content": "<p>How long does it take to converge?<br>\nI see it is very parameter sensitive, small change in lr hurt in cv badly,i was getting close to 0.95 cv but not MORE than that after 30-40 epochs iirc<br>\nYou ever got 0.95+ using nfnets?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221307,
          "author_name": "pukkinming",
          "author_url": "",
          "post_date": "02/28/2021 21:38:08",
          "content": "<p>That's difficult to say as it is related to your LR scheduler settings…..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1237729,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "03/14/2021 12:04:22",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> Just to confirm it's possible to get CV close to 0.97 with nfnets, but still doesn't change the fact that they are pretty heavy</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1237739,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "03/14/2021 12:11:23",
          "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> i guess it's very hyperparameter sensitive,i tried onecycleLR,deepmemory optimizer,nfnetsf-1, and lr ranges between 0.001-0.0011 and some other lr as well,then trained for around 50 hours at a stretch but couldn't get such impressive CV,if your 0.97 is not very hyperparameter sensitive then it's a great model,thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1221226,
      "author_name": "anjum48",
      "author_url": "",
      "post_date": "02/28/2021 19:10:20",
      "content": "<p>I tried a few folds using F1, but the performance was worse than ResNet based models</p>",
      "votes": null,
      "replies": [
        {
          "id": 1221245,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/28/2021 19:39:27",
          "content": "<p>Just curious <br>\n\"Is everyone's best model resnet200d or resnet50d like 'd' models?\"<br>\nOr other models working well too?<br>\nHow about tf models?<br>\nWe Don't have any 'd' models there<br>\nWhat's the best cv one can get with tensorflow in this competition? <br>\nJust curious </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221304,
          "author_name": "anjum48",
          "author_url": "",
          "post_date": "02/28/2021 21:33:50",
          "content": "<p>My best models so far are using ResNet200d, but I think the main reason that it is so popular is that it trains well on GPUs, i.e. good performance with large image sizes (and therefore small batch sizes) and relatively quick training. For me, it outperforms EfficientNet-B4 for the same amount of compute time using 512x512 images.</p>\n<p>So far I haven't got anything else to work as well within the constraints of the VRAM in my machine (yet!) ¯_(ツ)_/¯ </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221318,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "02/28/2021 22:03:57",
          "content": "<p>Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1221561,
          "author_name": "andy2709",
          "author_url": "",
          "post_date": "03/01/2021 05:11:15",
          "content": "<p>Current NFNets in timm are full of GELU activation which make them insanely slow ☹️. I gave f2 and f3 a try few weeks ago and found that at 640x640, they were slower than b6/ b7 at 1024. Since then, I stopped exploring them. <br>\nWith similar training schedule and heavy augmentation (drop path + randaug), EffNets gave me better results than ResNet (the d version) in this challenge. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222012,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "03/01/2021 13:21:19",
          "content": "<p><a href=\"https://www.kaggle.com/andy2709\" target=\"_blank\">@andy2709</a> that's interesting, I feel like most of the people in this comp are still using ResNet's. It's good to see Effnet alternative… </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222379,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "03/01/2021 18:26:30",
          "content": "<blockquote>\n  <p>Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(</p>\n</blockquote>\n<p>Did you enable bfloat16 ?  It doesn't work for some models </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222594,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "03/01/2021 23:56:11",
          "content": "<p>Yeah bfloat16 was enabled, all model works on torch xla but couldn’t make  resnet200d working, also each epoch training takes a lot of time<br>\nI will try without bfloat16. Thank you for letting me know :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222872,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "03/02/2021 08:30:58",
          "content": "<p>Some models didn't work for me the last time I used xla with bfloat16 .</p>\n<p>I had to disable it ( and reduce the batch size + longer training  unfortunately )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222887,
          "author_name": "manabendrarout",
          "author_url": "",
          "post_date": "03/02/2021 08:58:09",
          "content": "<p><a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> I have been exclusively using Tensorflow for classical models… Best I could do was 0.955 on EfficientNet B7 and 0.958 using ensemble. But hard to train it on GPU, and this drains TPU hours like anything too 😆</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222925,
          "author_name": "mobassir",
          "author_url": "",
          "post_date": "03/02/2021 09:57:02",
          "content": "<p><a href=\"https://www.kaggle.com/manabendrarout\" target=\"_blank\">@manabendrarout</a> interesting,you sure your validation strategy is good? i tried to improve ragnar's tf baseline,i couldn't get 0.95 cv without 'n' step TTA, is your 0.955 CV without TTA or with TTA? what lb you get with that? interesting</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222941,
          "author_name": "manabendrarout",
          "author_url": "",
          "post_date": "03/02/2021 10:12:27",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/mobassir\" target=\"_blank\">@mobassir</a> ,<br>\nThis is the config for my 0.955 CV (Single Model):-<br>\nModel:- EfficientNet B7 (with two extra dense layers and dropouts added after)<br>\nDrop Connect Rate:- 0.3<br>\nImage Size:- 600 😅<br>\nInitial Weights:- Noisy Student<br>\nLR Scheduler:- Decrease LR on plateau by 10x<br>\nOptimizer:- Adam<br>\nLoss:- Binary Cross entropy<br>\nAugmentation:- Basic (Random left-right, top-bottom flips and brightness change)<br>\nPlatform:- Tensorflow<br>\nTTA:- None</p>\n<p>This gets me a 0.946 LB.<br>\nStill working on improving it… But the gains have almost stagnated and processing time is so high that there are lot of diminishing returns there.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1221249,
      "author_name": "fffrrt",
      "author_url": "",
      "post_date": "02/28/2021 19:42:25",
      "content": "<blockquote>\n  <p>Have any of you tried these architectures for this competition? What are your observations?</p>\n</blockquote>\n<p>Slower to train and worse performance than EfficientNets. Perhaps torch implementation is not perfect yet, or those models dont transfer to competition domain well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1222677,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "03/02/2021 03:13:28",
      "content": "<p>resnet101 give me cv near to 96, have no too much luck with efficientnet, and the nfnnet too expensive, so i forgive use it</p>",
      "votes": null,
      "replies": [
        {
          "id": 1228732,
          "author_name": "digvijayyadav",
          "author_url": "",
          "post_date": "03/06/2021 17:44:19",
          "content": "<p>Exactly for me, ResNet had better performance than efficient-net, Nfnets, however, end up using too much computing power so am not able to use them</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1229891,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "03/07/2021 16:59:56",
          "content": "<p>Yeah they are consuming computing power a lot…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1224655,
      "author_name": "sinpcw",
      "author_url": "",
      "post_date": "03/03/2021 00:20:02",
      "content": "<p>I have been able to get LB.96+ using nfnets implemented by timm. It just seems to be more sensitive to hyperparameters than networks such as efficientnet and resnet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1225184,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "03/03/2021 11:54:21",
          "content": "<p>Yeah they are, I also get over lb 96.5's but have to say they are pretty heavy models…</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1230147,
      "author_name": "shivamverma3164",
      "author_url": "",
      "post_date": "03/07/2021 20:04:46",
      "content": "<p>what is the good way to start with CNN?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1239220,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/15/2021 14:44:27",
          "content": "<p>There are many ways. One that worked for me is to try to implement a simple CNN (using numpy). Once you know the basics, read some tutorials from popular libraries. Here is one example: <a href=\"https://www.tensorflow.org/tutorials/images/cnn\" target=\"_blank\">https://www.tensorflow.org/tutorials/images/cnn</a>.</p>\n<p>Happy learning!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239791,
      "author_name": "sugawarya",
      "author_url": "",
      "post_date": "03/16/2021 02:51:10",
      "content": "<p>I got LB: 0.969 with nfnet f1 (5 fold average) , however I couldn't get similar score with nfnet f0, f2. (~LB: 0.967)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1240417,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "03/16/2021 12:04:59",
          "content": "<p>That's nice, how was your cv and resolution if you don't mind sharing?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240547,
          "author_name": "sugawarya",
          "author_url": "",
          "post_date": "03/16/2021 13:33:22",
          "content": "<p>img size: 736<br>\nCV: 0.962</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1239883,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/16/2021 05:34:20",
      "content": "<p><a href=\"https://www.kaggle.com/datafan07\" target=\"_blank\">@datafan07</a> My Rainforest teammate <a href=\"https://www.kaggle.com/vineeth1999\" target=\"_blank\">@vineeth1999</a> implemented it for Rainforest . It didnt give us satisfactory results . I didnt try any new experiments for RANCZR . I only went through public kernels and discussion forums ..</p>",
      "votes": null,
      "replies": [
        {
          "id": 1240419,
          "author_name": "datafan07",
          "author_url": "",
          "post_date": "03/16/2021 12:05:56",
          "content": "<p>My single model comes from nfnets for this competition interestingly. Followed by resnet200d…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1240491,
          "author_name": "morizin",
          "author_url": "",
          "post_date": "03/16/2021 12:48:51",
          "content": "<p>Can you share what CV and LB score you achieved with nfnet as single?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241202,
      "author_name": "datafan07",
      "author_url": "",
      "post_date": "03/17/2021 00:14:28",
      "content": "<p>Just a note, nfnet saved my team from shakeup…</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1220342": "I have been trying to get some decent scores from recently released nfnets. I did get some promising results but in general it's still lacking behind some good old model architectures. I tried to stick on official paper generally but implementing some parts leading decreased cv scores. I have published one of my trials here as notebook if you wanna check:\n\nhttps://www.kaggle.com/datafan07/ranzcr-nfnets-tutorial-single-fold-training\n\nFor example using Adaptive Gradient Clipping didn't make a difference on training part. Maybe changing lambda value to something else helps...\n\nWhat you guys think? Have any of you tried these architectures for this competition? What are your observations?\n\nGood luck all...",
    "1220419": "Just hard models to train in my experience.. Ill stick to my good old resnets",
    "1220421": "I've been seeing some people claiming promising results, but probably going to wait to see more conclusive evidence since they're huge models",
    "1220451": "I think @rwightman has also observed that AGC isn't making much difference over using regular gradient clipping...\nhttps://twitter.com/wightmanr/status/1361746367769546753",
    "1220727": "i have spend most of my times training nfnets f0 model and no matter what i try i can't get cv 0.95 on first fold, i would like to know what is the maximum roc auc that one can get in this competition? sadly the best single model thread is kind of frozen :(\ndoes image size matter much here? i tried f0 model with different image sizes and didn't observe improvement  in roc auc",
    "1220728": "That's good to know thanks for sharing",
    "1220889": "Image size did matter in my trials, but idk what image size you have been using with f0's?",
    "1220893": "i tried 512,576,384,456 but nothing helps me to get 0.95+ on first fold out of 5,,were you able to get more than that while using nfnets f0? you used much bigger image size?",
    "1221226": "I tried a few folds using F1, but the performance was worse than ResNet based models",
    "1221233": "I usually choose pretrained image test sizes related to f number, but using smaller images on same f's reduced my cv score so I assumed increasing them would get better scores. But didn't try above 512 I think...",
    "1221244": "How long does it take to converge?\nI see it is very parameter sensitive, small change in lr hurt in cv badly,i was getting close to 0.95 cv but not MORE than that after 30-40 epochs iirc\nYou ever got 0.95+ using nfnets?",
    "1221245": "Just curious \n\"Is everyone's best model resnet200d or resnet50d like 'd' models?\"\nOr other models working well too?\nHow about tf models?\nWe Don't have any 'd' models there\nWhat's the best cv one can get with tensorflow in this competition? \nJust curious",
    "1221249": "> Have any of you tried these architectures for this competition? What are your observations?\n\nSlower to train and worse performance than EfficientNets. Perhaps torch implementation is not perfect yet, or those models dont transfer to competition domain well.",
    "1221304": "My best models so far are using ResNet200d, but I think the main reason that it is so popular is that it trains well on GPUs, i.e. good performance with large image sizes (and therefore small batch sizes) and relatively quick training. For me, it outperforms EfficientNet-B4 for the same amount of compute time using 512x512 images.\n\nSo far I haven't got anything else to work as well within the constraints of the VRAM in my machine (yet!) ¯\\_(ツ)_/¯",
    "1221307": "That's difficult to say as it is related to your LR scheduler settings.....",
    "1221318": "Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(",
    "1221561": "Current NFNets in timm are full of GELU activation which make them insanely slow ☹️. I gave f2 and f3 a try few weeks ago and found that at 640x640, they were slower than b6/ b7 at 1024. Since then, I stopped exploring them. \nWith similar training schedule and heavy augmentation (drop path + randaug), EffNets gave me better results than ResNet (the d version) in this challenge.",
    "1222012": "andy2709 that's interesting, I feel like most of the people in this comp are still using ResNet's. It's good to see Effnet alternative...",
    "1222379": "> Not a heavy gpu guy.i tried resnet200d on pytorch xla and got loss NaN very fast,this architecture Doesn't work for me on tpu but other architectures are working on torch xla :(\n\nDid you enable bfloat16 ?  It doesn't work for some models",
    "1222594": "Yeah bfloat16 was enabled, all model works on torch xla but couldn’t make  resnet200d working, also each epoch training takes a lot of time\nI will try without bfloat16. Thank you for letting me know :)",
    "1222677": "resnet101 give me cv near to 96, have no too much luck with efficientnet, and the nfnnet too expensive, so i forgive use it",
    "1222872": "Some models didn't work for me the last time I used xla with bfloat16 .\n\nI had to disable it ( and reduce the batch size + longer training  unfortunately )",
    "1222887": "mobassir I have been exclusively using Tensorflow for classical models... Best I could do was 0.955 on EfficientNet B7 and 0.958 using ensemble. But hard to train it on GPU, and this drains TPU hours like anything too 😆",
    "1222925": "manabendrarout interesting,you sure your validation strategy is good? i tried to improve ragnar's tf baseline,i couldn't get 0.95 cv without 'n' step TTA, is your 0.955 CV without TTA or with TTA? what lb you get with that? interesting",
    "1222941": "Hi @mobassir ,\nThis is the config for my 0.955 CV (Single Model):-\nModel:- EfficientNet B7 (with two extra dense layers and dropouts added after)\nDrop Connect Rate:- 0.3\nImage Size:- 600 😅\nInitial Weights:- Noisy Student\nLR Scheduler:- Decrease LR on plateau by 10x\nOptimizer:- Adam\nLoss:- Binary Cross entropy\nAugmentation:- Basic (Random left-right, top-bottom flips and brightness change)\nPlatform:- Tensorflow\nTTA:- None\n\nThis gets me a 0.946 LB.\nStill working on improving it... But the gains have almost stagnated and processing time is so high that there are lot of diminishing returns there.",
    "1224655": "I have been able to get LB.96+ using nfnets implemented by timm. It just seems to be more sensitive to hyperparameters than networks such as efficientnet and resnet.",
    "1225184": "Yeah they are, I also get over lb 96.5's but have to say they are pretty heavy models...",
    "1228732": "Exactly for me, ResNet had better performance than efficient-net, Nfnets, however, end up using too much computing power so am not able to use them",
    "1229891": "Yeah they are consuming computing power a lot...",
    "1230147": "what is the good way to start with CNN?",
    "1237729": "mobassir Just to confirm it's possible to get CV close to 0.97 with nfnets, but still doesn't change the fact that they are pretty heavy",
    "1237739": "datafan07 i guess it's very hyperparameter sensitive,i tried onecycleLR,deepmemory optimizer,nfnetsf-1, and lr ranges between 0.001-0.0011 and some other lr as well,then trained for around 50 hours at a stretch but couldn't get such impressive CV,if your 0.97 is not very hyperparameter sensitive then it's a great model,thanks",
    "1239220": "There are many ways. One that worked for me is to try to implement a simple CNN (using numpy). Once you know the basics, read some tutorials from popular libraries. Here is one example: https://www.tensorflow.org/tutorials/images/cnn.\n\nHappy learning!",
    "1239791": "I got LB: 0.969 with nfnet f1 (5 fold average) , however I couldn't get similar score with nfnet f0, f2. (~LB: 0.967)",
    "1239883": "datafan07 My Rainforest teammate @vineeth1999 implemented it for Rainforest . It didnt give us satisfactory results . I didnt try any new experiments for RANCZR . I only went through public kernels and discussion forums ..",
    "1240417": "That's nice, how was your cv and resolution if you don't mind sharing?",
    "1240419": "My single model comes from nfnets for this competition interestingly. Followed by resnet200d...",
    "1240491": "Can you share what CV and LB score you achieved with nfnet as single?",
    "1240547": "img size: 736\nCV: 0.962",
    "1241202": "Just a note, nfnet saved my team from shakeup..."
  },
  "source": "meta"
}