{
  "id": 239514,
  "title": "Some CV results sharing",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/239514",
  "author_name": "",
  "post_date": "2021-05-16T15:35:01.257032800Z",
  "votes": 8,
  "comment_count": 40,
  "views": 0,
  "content": "<p>Up to now, my best single model gets a cv score around 0.92. Since cv is more reliable than lb, I'm curious about the highest cv score in your experiments. Anyone gets a score &gt;= 0.93 ?</p>",
  "messages": [
    {
      "id": "1310312",
      "postDate": "05/16/2021 15:35:01",
      "content": "<p>Up to now, my best single model gets a cv score around 0.92. Since cv is more reliable than lb, I'm curious about the highest cv score in your experiments. Anyone gets a score &gt;= 0.93 ?</p>",
      "rawMarkdown": "Up to now, my best single model gets a cv score around 0.92. Since cv is more reliable than lb, I'm curious about the highest cv score in your experiments. Anyone gets a score >= 0.93 ?",
      "votes": null
    },
    {
      "id": "1310600",
      "postDate": "05/16/2021 18:46:16",
      "content": "<p>Your CV Score means Macro F1-Score?</p>",
      "rawMarkdown": "Your CV Score means Macro F1-Score?",
      "votes": null
    },
    {
      "id": "1310609",
      "postDate": "05/16/2021 18:55:32",
      "content": "<p>Single model - do you mean single fold, or averaged multi-fold?</p>",
      "rawMarkdown": "Single model - do you mean single fold, or averaged multi-fold?",
      "votes": null
    },
    {
      "id": "1310808",
      "postDate": "05/17/2021 01:55:37",
      "content": "<p>5 folds effecientnet_b5_ns</p>",
      "rawMarkdown": "5 folds effecientnet_b5_ns",
      "votes": null
    },
    {
      "id": "1310811",
      "postDate": "05/17/2021 01:56:41",
      "content": "<p>Mean f1-score, which is averaged over f1 score of all categories.</p>",
      "rawMarkdown": "Mean f1-score, which is averaged over f1 score of all categories.",
      "votes": null
    },
    {
      "id": "1310816",
      "postDate": "05/17/2021 02:01:18",
      "content": "<p>I got around 0.92 for EfficientNetB7 too.<br>\nYou can test the model's generalization on the last year's dataset.<br>\nIf posible, you can share the results here.</p>",
      "rawMarkdown": "I got around 0.92 for EfficientNetB7 too.\nYou can test the model's generalization on the last year's dataset.\nIf posible, you can share the results here.",
      "votes": null
    },
    {
      "id": "1310822",
      "postDate": "05/17/2021 02:08:33",
      "content": "<p>Seems that it's hard to get a score above 0.93.</p>",
      "rawMarkdown": "Seems that it's hard to get a score above 0.93.",
      "votes": null
    },
    {
      "id": "1310827",
      "postDate": "05/17/2021 02:14:15",
      "content": "<p>Maybe the noisy labels is an important reason.<br>\nAnd I found the model cannot find an accurate boundary between 'scab' and 'rust'.</p>",
      "rawMarkdown": "Maybe the noisy labels is an important reason.\nAnd I found the model cannot find an accurate boundary between 'scab' and 'rust'.",
      "votes": null
    },
    {
      "id": "1311253",
      "postDate": "05/17/2021 09:12:35",
      "content": "<p>My max CV is 0.9377 for EfficientNetB5 (Tensorflow, 5 folds averaged).</p>",
      "rawMarkdown": "My max CV is 0.9377 for EfficientNetB5 (Tensorflow, 5 folds averaged).",
      "votes": null
    },
    {
      "id": "1311267",
      "postDate": "05/17/2021 09:24:51",
      "content": "<p>Wow, that's really impressive.</p>",
      "rawMarkdown": "Wow, that's really impressive.",
      "votes": null
    },
    {
      "id": "1311302",
      "postDate": "05/17/2021 09:57:13",
      "content": "<p>I'm using <code>tfa.metrics.F1Score(len(CFG.CLASSES), average='weighted')</code> if you're interested.</p>",
      "rawMarkdown": "I'm using `tfa.metrics.F1Score(len(CFG.CLASSES), average='weighted')` if you're interested.",
      "votes": null
    },
    {
      "id": "1311305",
      "postDate": "05/17/2021 10:01:59",
      "content": "<p>Thanks for your advice.</p>",
      "rawMarkdown": "Thanks for your advice.",
      "votes": null
    },
    {
      "id": "1311306",
      "postDate": "05/17/2021 10:03:17",
      "content": "<p><code>complex</code> label detection is also a problem.</p>",
      "rawMarkdown": "`complex` label detection is also a problem.",
      "votes": null
    },
    {
      "id": "1311336",
      "postDate": "05/17/2021 10:33:54",
      "content": "<p>hi, the second place you achieve is so impressive.I also use the efficientnet but the score in cv is 0.826 and in lb is 0.819.Can i know what the resolution in your network is?</p>",
      "rawMarkdown": "hi, the second place you achieve is so impressive.I also use the efficientnet but the score in cv is 0.826 and in lb is 0.819.Can i know what the resolution in your network is?",
      "votes": null
    },
    {
      "id": "1311348",
      "postDate": "05/17/2021 10:44:15",
      "content": "<p>456 * 456 (Using pretrained model from timm: <a href=\"https://github.com/rwightman/pytorch-image-models)\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models)</a>,  you can try different image size and see which one performs best. BTW, I think we should focus more on cv instead of lb.</p>",
      "rawMarkdown": "456 * 456 (Using pretrained model from timm: https://github.com/rwightman/pytorch-image-models),  you can try different image size and see which one performs best. BTW, I think we should focus more on cv instead of lb.",
      "votes": null
    },
    {
      "id": "1311356",
      "postDate": "05/17/2021 10:50:37",
      "content": "<p>thx very much! yeah we should more focus on cv score.My question is that i have tried the 528x528 resolution in efficientnetb6(which is the recommended resolution from the tutorial), but the result is very bad. I dont know why ,because 224x224 is a really small resolution</p>",
      "rawMarkdown": "thx very much! yeah we should more focus on cv score.My question is that i have tried the 528x528 resolution in efficientnetb6(which is the recommended resolution from the tutorial), but the result is very bad. I dont know why ,because 224x224 is a really small resolution",
      "votes": null
    },
    {
      "id": "1311359",
      "postDate": "05/17/2021 10:59:46",
      "content": "<p>I'm using 768x768.</p>",
      "rawMarkdown": "I'm using 768x768.",
      "votes": null
    },
    {
      "id": "1311365",
      "postDate": "05/17/2021 11:04:14",
      "content": "<p>thanks for your reply!</p>",
      "rawMarkdown": "thanks for your reply!",
      "votes": null
    },
    {
      "id": "1311441",
      "postDate": "05/17/2021 12:01:04",
      "content": "<p>Sometimes, choosing model is very  tricky. Some models with more parameters just may not fit this problem. You need to try multiple different models to find the best one.</p>",
      "rawMarkdown": "Sometimes, choosing model is very  tricky. Some models with more parameters just may not fit this problem. You need to try multiple different models to find the best one.",
      "votes": null
    },
    {
      "id": "1311746",
      "postDate": "05/17/2021 15:27:00",
      "content": "<p>thx so much！！</p>",
      "rawMarkdown": "thx so much！！",
      "votes": null
    },
    {
      "id": "1311810",
      "postDate": "05/17/2021 16:17:02",
      "content": "<p>What's really strange is that my model ensemble with CV ~0.956 scores on LB as 0.835. Maybe it's because of low percentage of data evaluated on public LB (just 18%).</p>",
      "rawMarkdown": "What's really strange is that my model ensemble with CV ~0.956 scores on LB as 0.835. Maybe it's because of low percentage of data evaluated on public LB (just 18%).",
      "votes": null
    },
    {
      "id": "1311863",
      "postDate": "05/17/2021 16:55:42",
      "content": "<p><a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> is there any reason you are using <code>average='weighted'</code> instead of using <code>average='micro'</code> for the evaluation metric?</p>",
      "rawMarkdown": "atamazian is there any reason you are using `average='weighted'` instead of using `average='micro'` for the evaluation metric?",
      "votes": null
    },
    {
      "id": "1311936",
      "postDate": "05/17/2021 17:55:07",
      "content": "<p>Well, the training dataset is imbalanced, so I thought that <code>'weighted'</code> option will give more accurate score.</p>",
      "rawMarkdown": "Well, the training dataset is imbalanced, so I thought that `'weighted'` option will give more accurate score.",
      "votes": null
    },
    {
      "id": "1312359",
      "postDate": "05/18/2021 02:28:22",
      "content": "<p><a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> I still don't get you, which f1 score average method did you use? <br>\naverage = None, macro, micro or weighted?</p>\n<p>As per <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a>, 'weighted' would be a better metric. <br>\n<a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> is using 'micro'</p>\n<p>Weighted would take into consideration the class imbalance. <br>\nMacro would not bother about class imbalance. It's simple averaging.<br>\nMicro would ignore the rare samples. </p>",
      "rawMarkdown": "ytepzhi I still don't get you, which f1 score average method did you use? \naverage = None, macro, micro or weighted?\n\n\nAs per @atamazian, 'weighted' would be a better metric. \n@brendanartley is using 'micro'\n\nWeighted would take into consideration the class imbalance. \nMacro would not bother about class imbalance. It's simple averaging.\nMicro would ignore the rare samples.",
      "votes": null
    },
    {
      "id": "1312362",
      "postDate": "05/18/2021 02:32:25",
      "content": "<p>Simple averaging without weighting.</p>",
      "rawMarkdown": "Simple averaging without weighting.",
      "votes": null
    },
    {
      "id": "1312369",
      "postDate": "05/18/2021 02:41:35",
      "content": "<p>You should use sklearn.metrics.f1_score('samples') or another function like this.</p>",
      "rawMarkdown": "You should use sklearn.metrics.f1_score('samples') or another function like this.",
      "votes": null
    },
    {
      "id": "1312450",
      "postDate": "05/18/2021 04:26:02",
      "content": "<p>Most of my CV score are around 0.89~0.92. I also use the efficientNet model from timm. B.T.W why did you choose the resolution of the input image to be 456 * 456? <br>\nMany thanks!</p>",
      "rawMarkdown": "Most of my CV score are around 0.89~0.92. I also use the efficientNet model from timm. B.T.W why did you choose the resolution of the input image to be 456 * 456? \nMany thanks!",
      "votes": null
    },
    {
      "id": "1320218",
      "postDate": "05/23/2021 21:46:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> .<br>\nMy CV is around 0.928 to 0.93 but LB is 0.807.<br>\nI don't know why.<br>\nIs there any trick that everyone else is using like setting different thresholds for complex etc.</p>",
      "rawMarkdown": "Hi @ytepzhi .\nMy CV is around 0.928 to 0.93 but LB is 0.807.\nI don't know why.\nIs there any trick that everyone else is using like setting different thresholds for complex etc.",
      "votes": null
    },
    {
      "id": "1320222",
      "postDate": "05/23/2021 21:49:52",
      "content": "<p>I think that public LB score is not representative since it uses only 19% of data for evaluation.</p>",
      "rawMarkdown": "I think that public LB score is not representative since it uses only 19% of data for evaluation.",
      "votes": null
    },
    {
      "id": "1320225",
      "postDate": "05/23/2021 21:57:36",
      "content": "<p>Yes. But I was just kind of confused that if i am missing something or not.</p>",
      "rawMarkdown": "Yes. But I was just kind of confused that if i am missing something or not.",
      "votes": null
    },
    {
      "id": "1320423",
      "postDate": "05/24/2021 04:46:16",
      "content": "<p>Choosing threshold is important for this competition, I'm currently using threshold 0.4 for all categories. If probability(label A) &gt; 0.4, A is in the prediction.</p>",
      "rawMarkdown": "Choosing threshold is important for this competition, I'm currently using threshold 0.4 for all categories. If probability(label A) > 0.4, A is in the prediction.",
      "votes": null
    },
    {
      "id": "1320429",
      "postDate": "05/24/2021 04:49:03",
      "content": "<p>Exactly ! 19 % test data is so unreliable.</p>",
      "rawMarkdown": "Exactly ! 19 % test data is so unreliable.",
      "votes": null
    },
    {
      "id": "1320995",
      "postDate": "05/24/2021 13:18:58",
      "content": "<p>I have a question. How are you people cross-validating and experimenting at such a speed? For Efficientnetb4 using resized images, it is taking too much time. How much time is it taking you for 5 folds? </p>",
      "rawMarkdown": "I have a question. How are you people cross-validating and experimenting at such a speed? For Efficientnetb4 using resized images, it is taking too much time. How much time is it taking you for 5 folds?",
      "votes": null
    },
    {
      "id": "1321037",
      "postDate": "05/24/2021 13:44:29",
      "content": "<p><a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> You are using a multi-label approach, right? 0.40 was tuned across all the folds or it's test and trial? Are you using healthy class?</p>",
      "rawMarkdown": "ytepzhi You are using a multi-label approach, right? 0.40 was tuned across all the folds or it's test and trial? Are you using healthy class?",
      "votes": null
    },
    {
      "id": "1321038",
      "postDate": "05/24/2021 13:44:36",
      "content": "<p>TF, 5 folds, EfficientNetB5, preaugmented TFRecords takes ~ 5 hours.</p>",
      "rawMarkdown": "TF, 5 folds, EfficientNetB5, preaugmented TFRecords takes ~ 5 hours.",
      "votes": null
    },
    {
      "id": "1321069",
      "postDate": "05/24/2021 14:00:11",
      "content": "<p>Thanks. Is pre augmented the right approach? I mean for every epoch we should have unique augmentations. 5 hours is a good time. I think TPU is the key here to make things faster?</p>",
      "rawMarkdown": "Thanks. Is pre augmented the right approach? I mean for every epoch we should have unique augmentations. 5 hours is a good time. I think TPU is the key here to make things faster?",
      "votes": null
    },
    {
      "id": "1321150",
      "postDate": "05/24/2021 14:19:31",
      "content": "<p>My CV is around 0.95 but LB is stuck at 0.807<br>\nThreshold is 0.5</p>",
      "rawMarkdown": "My CV is around 0.95 but LB is stuck at 0.807\nThreshold is 0.5",
      "votes": null
    },
    {
      "id": "1321164",
      "postDate": "05/24/2021 14:28:03",
      "content": "<p>Check <a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-ultimate-preprocessing\" target=\"_blank\">this notebook</a> for more info. And yes, TPU greatly speeds up the training.</p>",
      "rawMarkdown": "Check [this notebook](https://www.kaggle.com/nickuzmenkov/pp2021-ultimate-preprocessing) for more info. And yes, TPU greatly speeds up the training.",
      "votes": null
    },
    {
      "id": "1323032",
      "postDate": "05/25/2021 22:20:17",
      "content": "<p>My 5-fold CV is slightly north of 0.93 with F1-macro and threshold=0.5. Corresponding LB which is based on an ensemble of these 5 models with optimized thresholds, such that each fold and class-label has its own threshold value, gave me 0.826 on public LB.</p>",
      "rawMarkdown": "My 5-fold CV is slightly north of 0.93 with F1-macro and threshold=0.5. Corresponding LB which is based on an ensemble of these 5 models with optimized thresholds, such that each fold and class-label has its own threshold value, gave me 0.826 on public LB.",
      "votes": null
    },
    {
      "id": "1327218",
      "postDate": "05/29/2021 04:44:27",
      "content": "<p>Did you guys used prescribed image size by keras for the respective models for efficientnet or was it different ??</p>",
      "rawMarkdown": "Did you guys used prescribed image size by keras for the respective models for efficientnet or was it different ??",
      "votes": null
    },
    {
      "id": "1327531",
      "postDate": "05/29/2021 11:50:47",
      "content": "<p>I used torch, the images are resized in different shapes for different models.</p>",
      "rawMarkdown": "I used torch, the images are resized in different shapes for different models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1310600,
      "author_name": "rainyq",
      "author_url": "",
      "post_date": "05/16/2021 18:46:16",
      "content": "<p>Your CV Score means Macro F1-Score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1310811,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/17/2021 01:56:41",
          "content": "<p>Mean f1-score, which is averaged over f1 score of all categories.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310816,
          "author_name": "rainyq",
          "author_url": "",
          "post_date": "05/17/2021 02:01:18",
          "content": "<p>I got around 0.92 for EfficientNetB7 too.<br>\nYou can test the model's generalization on the last year's dataset.<br>\nIf posible, you can share the results here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310822,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/17/2021 02:08:33",
          "content": "<p>Seems that it's hard to get a score above 0.93.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1310827,
          "author_name": "rainyq",
          "author_url": "",
          "post_date": "05/17/2021 02:14:15",
          "content": "<p>Maybe the noisy labels is an important reason.<br>\nAnd I found the model cannot find an accurate boundary between 'scab' and 'rust'.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311306,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/17/2021 10:03:17",
          "content": "<p><code>complex</code> label detection is also a problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1312359,
          "author_name": "antoreepjana",
          "author_url": "",
          "post_date": "05/18/2021 02:28:22",
          "content": "<p><a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> I still don't get you, which f1 score average method did you use? <br>\naverage = None, macro, micro or weighted?</p>\n<p>As per <a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a>, 'weighted' would be a better metric. <br>\n<a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> is using 'micro'</p>\n<p>Weighted would take into consideration the class imbalance. <br>\nMacro would not bother about class imbalance. It's simple averaging.<br>\nMicro would ignore the rare samples. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1312362,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/18/2021 02:32:25",
          "content": "<p>Simple averaging without weighting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1312369,
          "author_name": "rainyq",
          "author_url": "",
          "post_date": "05/18/2021 02:41:35",
          "content": "<p>You should use sklearn.metrics.f1_score('samples') or another function like this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1310609,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "05/16/2021 18:55:32",
      "content": "<p>Single model - do you mean single fold, or averaged multi-fold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1310808,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/17/2021 01:55:37",
          "content": "<p>5 folds effecientnet_b5_ns</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311253,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/17/2021 09:12:35",
          "content": "<p>My max CV is 0.9377 for EfficientNetB5 (Tensorflow, 5 folds averaged).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311267,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/17/2021 09:24:51",
          "content": "<p>Wow, that's really impressive.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311302,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/17/2021 09:57:13",
          "content": "<p>I'm using <code>tfa.metrics.F1Score(len(CFG.CLASSES), average='weighted')</code> if you're interested.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311305,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/17/2021 10:01:59",
          "content": "<p>Thanks for your advice.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311810,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/17/2021 16:17:02",
          "content": "<p>What's really strange is that my model ensemble with CV ~0.956 scores on LB as 0.835. Maybe it's because of low percentage of data evaluated on public LB (just 18%).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311863,
          "author_name": "brendanartley",
          "author_url": "",
          "post_date": "05/17/2021 16:55:42",
          "content": "<p><a href=\"https://www.kaggle.com/atamazian\" target=\"_blank\">@atamazian</a> is there any reason you are using <code>average='weighted'</code> instead of using <code>average='micro'</code> for the evaluation metric?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311936,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/17/2021 17:55:07",
          "content": "<p>Well, the training dataset is imbalanced, so I thought that <code>'weighted'</code> option will give more accurate score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1311336,
      "author_name": "kagglezzr",
      "author_url": "",
      "post_date": "05/17/2021 10:33:54",
      "content": "<p>hi, the second place you achieve is so impressive.I also use the efficientnet but the score in cv is 0.826 and in lb is 0.819.Can i know what the resolution in your network is?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1311348,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/17/2021 10:44:15",
          "content": "<p>456 * 456 (Using pretrained model from timm: <a href=\"https://github.com/rwightman/pytorch-image-models)\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models)</a>,  you can try different image size and see which one performs best. BTW, I think we should focus more on cv instead of lb.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311356,
          "author_name": "kagglezzr",
          "author_url": "",
          "post_date": "05/17/2021 10:50:37",
          "content": "<p>thx very much! yeah we should more focus on cv score.My question is that i have tried the 528x528 resolution in efficientnetb6(which is the recommended resolution from the tutorial), but the result is very bad. I dont know why ,because 224x224 is a really small resolution</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311359,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/17/2021 10:59:46",
          "content": "<p>I'm using 768x768.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311365,
          "author_name": "kagglezzr",
          "author_url": "",
          "post_date": "05/17/2021 11:04:14",
          "content": "<p>thanks for your reply!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311441,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/17/2021 12:01:04",
          "content": "<p>Sometimes, choosing model is very  tricky. Some models with more parameters just may not fit this problem. You need to try multiple different models to find the best one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1311746,
          "author_name": "kagglezzr",
          "author_url": "",
          "post_date": "05/17/2021 15:27:00",
          "content": "<p>thx so much！！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1312450,
      "author_name": "crissallan",
      "author_url": "",
      "post_date": "05/18/2021 04:26:02",
      "content": "<p>Most of my CV score are around 0.89~0.92. I also use the efficientNet model from timm. B.T.W why did you choose the resolution of the input image to be 456 * 456? <br>\nMany thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1320218,
      "author_name": "micheomaano",
      "author_url": "",
      "post_date": "05/23/2021 21:46:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> .<br>\nMy CV is around 0.928 to 0.93 but LB is 0.807.<br>\nI don't know why.<br>\nIs there any trick that everyone else is using like setting different thresholds for complex etc.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1320222,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/23/2021 21:49:52",
          "content": "<p>I think that public LB score is not representative since it uses only 19% of data for evaluation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1320225,
          "author_name": "micheomaano",
          "author_url": "",
          "post_date": "05/23/2021 21:57:36",
          "content": "<p>Yes. But I was just kind of confused that if i am missing something or not.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1320423,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/24/2021 04:46:16",
          "content": "<p>Choosing threshold is important for this competition, I'm currently using threshold 0.4 for all categories. If probability(label A) &gt; 0.4, A is in the prediction.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1320429,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/24/2021 04:49:03",
          "content": "<p>Exactly ! 19 % test data is so unreliable.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1321037,
          "author_name": "deepdreamx",
          "author_url": "",
          "post_date": "05/24/2021 13:44:29",
          "content": "<p><a href=\"https://www.kaggle.com/ytepzhi\" target=\"_blank\">@ytepzhi</a> You are using a multi-label approach, right? 0.40 was tuned across all the folds or it's test and trial? Are you using healthy class?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1320995,
      "author_name": "deepdreamx",
      "author_url": "",
      "post_date": "05/24/2021 13:18:58",
      "content": "<p>I have a question. How are you people cross-validating and experimenting at such a speed? For Efficientnetb4 using resized images, it is taking too much time. How much time is it taking you for 5 folds? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1321038,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/24/2021 13:44:36",
          "content": "<p>TF, 5 folds, EfficientNetB5, preaugmented TFRecords takes ~ 5 hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1321069,
          "author_name": "deepdreamx",
          "author_url": "",
          "post_date": "05/24/2021 14:00:11",
          "content": "<p>Thanks. Is pre augmented the right approach? I mean for every epoch we should have unique augmentations. 5 hours is a good time. I think TPU is the key here to make things faster?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1321164,
          "author_name": "atamazian",
          "author_url": "",
          "post_date": "05/24/2021 14:28:03",
          "content": "<p>Check <a href=\"https://www.kaggle.com/nickuzmenkov/pp2021-ultimate-preprocessing\" target=\"_blank\">this notebook</a> for more info. And yes, TPU greatly speeds up the training.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1321150,
      "author_name": "micheomaano",
      "author_url": "",
      "post_date": "05/24/2021 14:19:31",
      "content": "<p>My CV is around 0.95 but LB is stuck at 0.807<br>\nThreshold is 0.5</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1323032,
      "author_name": "datasciencegeek",
      "author_url": "",
      "post_date": "05/25/2021 22:20:17",
      "content": "<p>My 5-fold CV is slightly north of 0.93 with F1-macro and threshold=0.5. Corresponding LB which is based on an ensemble of these 5 models with optimized thresholds, such that each fold and class-label has its own threshold value, gave me 0.826 on public LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1327218,
      "author_name": "jl18pg052",
      "author_url": "",
      "post_date": "05/29/2021 04:44:27",
      "content": "<p>Did you guys used prescribed image size by keras for the respective models for efficientnet or was it different ??</p>",
      "votes": null,
      "replies": [
        {
          "id": 1327531,
          "author_name": "ytepzhi",
          "author_url": "",
          "post_date": "05/29/2021 11:50:47",
          "content": "<p>I used torch, the images are resized in different shapes for different models.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1310312": "Up to now, my best single model gets a cv score around 0.92. Since cv is more reliable than lb, I'm curious about the highest cv score in your experiments. Anyone gets a score >= 0.93 ?",
    "1310600": "Your CV Score means Macro F1-Score?",
    "1310609": "Single model - do you mean single fold, or averaged multi-fold?",
    "1310808": "5 folds effecientnet_b5_ns",
    "1310811": "Mean f1-score, which is averaged over f1 score of all categories.",
    "1310816": "I got around 0.92 for EfficientNetB7 too.\nYou can test the model's generalization on the last year's dataset.\nIf posible, you can share the results here.",
    "1310822": "Seems that it's hard to get a score above 0.93.",
    "1310827": "Maybe the noisy labels is an important reason.\nAnd I found the model cannot find an accurate boundary between 'scab' and 'rust'.",
    "1311253": "My max CV is 0.9377 for EfficientNetB5 (Tensorflow, 5 folds averaged).",
    "1311267": "Wow, that's really impressive.",
    "1311302": "I'm using `tfa.metrics.F1Score(len(CFG.CLASSES), average='weighted')` if you're interested.",
    "1311305": "Thanks for your advice.",
    "1311306": "`complex` label detection is also a problem.",
    "1311336": "hi, the second place you achieve is so impressive.I also use the efficientnet but the score in cv is 0.826 and in lb is 0.819.Can i know what the resolution in your network is?",
    "1311348": "456 * 456 (Using pretrained model from timm: https://github.com/rwightman/pytorch-image-models),  you can try different image size and see which one performs best. BTW, I think we should focus more on cv instead of lb.",
    "1311356": "thx very much! yeah we should more focus on cv score.My question is that i have tried the 528x528 resolution in efficientnetb6(which is the recommended resolution from the tutorial), but the result is very bad. I dont know why ,because 224x224 is a really small resolution",
    "1311359": "I'm using 768x768.",
    "1311365": "thanks for your reply!",
    "1311441": "Sometimes, choosing model is very  tricky. Some models with more parameters just may not fit this problem. You need to try multiple different models to find the best one.",
    "1311746": "thx so much！！",
    "1311810": "What's really strange is that my model ensemble with CV ~0.956 scores on LB as 0.835. Maybe it's because of low percentage of data evaluated on public LB (just 18%).",
    "1311863": "atamazian is there any reason you are using `average='weighted'` instead of using `average='micro'` for the evaluation metric?",
    "1311936": "Well, the training dataset is imbalanced, so I thought that `'weighted'` option will give more accurate score.",
    "1312359": "ytepzhi I still don't get you, which f1 score average method did you use? \naverage = None, macro, micro or weighted?\n\n\nAs per @atamazian, 'weighted' would be a better metric. \n@brendanartley is using 'micro'\n\nWeighted would take into consideration the class imbalance. \nMacro would not bother about class imbalance. It's simple averaging.\nMicro would ignore the rare samples.",
    "1312362": "Simple averaging without weighting.",
    "1312369": "You should use sklearn.metrics.f1_score('samples') or another function like this.",
    "1312450": "Most of my CV score are around 0.89~0.92. I also use the efficientNet model from timm. B.T.W why did you choose the resolution of the input image to be 456 * 456? \nMany thanks!",
    "1320218": "Hi @ytepzhi .\nMy CV is around 0.928 to 0.93 but LB is 0.807.\nI don't know why.\nIs there any trick that everyone else is using like setting different thresholds for complex etc.",
    "1320222": "I think that public LB score is not representative since it uses only 19% of data for evaluation.",
    "1320225": "Yes. But I was just kind of confused that if i am missing something or not.",
    "1320423": "Choosing threshold is important for this competition, I'm currently using threshold 0.4 for all categories. If probability(label A) > 0.4, A is in the prediction.",
    "1320429": "Exactly ! 19 % test data is so unreliable.",
    "1320995": "I have a question. How are you people cross-validating and experimenting at such a speed? For Efficientnetb4 using resized images, it is taking too much time. How much time is it taking you for 5 folds?",
    "1321037": "ytepzhi You are using a multi-label approach, right? 0.40 was tuned across all the folds or it's test and trial? Are you using healthy class?",
    "1321038": "TF, 5 folds, EfficientNetB5, preaugmented TFRecords takes ~ 5 hours.",
    "1321069": "Thanks. Is pre augmented the right approach? I mean for every epoch we should have unique augmentations. 5 hours is a good time. I think TPU is the key here to make things faster?",
    "1321150": "My CV is around 0.95 but LB is stuck at 0.807\nThreshold is 0.5",
    "1321164": "Check [this notebook](https://www.kaggle.com/nickuzmenkov/pp2021-ultimate-preprocessing) for more info. And yes, TPU greatly speeds up the training.",
    "1323032": "My 5-fold CV is slightly north of 0.93 with F1-macro and threshold=0.5. Corresponding LB which is based on an ensemble of these 5 models with optimized thresholds, such that each fold and class-label has its own threshold value, gave me 0.826 on public LB.",
    "1327218": "Did you guys used prescribed image size by keras for the respective models for efficientnet or was it different ??",
    "1327531": "I used torch, the images are resized in different shapes for different models."
  },
  "source": "meta"
}