{
  "id": 202333,
  "title": "Things I've tried with FastAI till now",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/202333",
  "author_name": "",
  "post_date": "2020-12-09T13:32:19.332082500Z",
  "votes": 3,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi guys, I'm using FastAI and Pytorch to create an efficientnet-b3_ns model for this competition. I've used CutMix as suggested in <a href=\"https://www.kaggle.com/ababino/cutmix-with-fastai-and-efficientnet/notebook\" target=\"_blank\">this</a> notebook. I've also used a few tweaks from <a href=\"https://www.kaggle.com/muellerzr\" target=\"_blank\">@muellerzr</a> <a href=\"https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai\" target=\"_blank\">notebook</a>. I've used a tta with n=15, achieving an accuracy of close to 0.893. However, now I'm running out of any major tweaks that I can implement to improve my accuracy. Any suggestions to cross the 0.9 mark are highly appreciated.</p>",
  "messages": [
    {
      "id": "1107198",
      "postDate": "12/09/2020 13:32:19",
      "content": "<p>Hi guys, I'm using FastAI and Pytorch to create an efficientnet-b3_ns model for this competition. I've used CutMix as suggested in <a href=\"https://www.kaggle.com/ababino/cutmix-with-fastai-and-efficientnet/notebook\" target=\"_blank\">this</a> notebook. I've also used a few tweaks from <a href=\"https://www.kaggle.com/muellerzr\" target=\"_blank\">@muellerzr</a> <a href=\"https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai\" target=\"_blank\">notebook</a>. I've used a tta with n=15, achieving an accuracy of close to 0.893. However, now I'm running out of any major tweaks that I can implement to improve my accuracy. Any suggestions to cross the 0.9 mark are highly appreciated.</p>",
      "rawMarkdown": "Hi guys, I'm using FastAI and Pytorch to create an efficientnet-b3_ns model for this competition. I've used CutMix as suggested in [this](https://www.kaggle.com/ababino/cutmix-with-fastai-and-efficientnet/notebook) notebook. I've also used a few tweaks from @muellerzr [notebook](https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai). I've used a tta with n=15, achieving an accuracy of close to 0.893. However, now I'm running out of any major tweaks that I can implement to improve my accuracy. Any suggestions to cross the 0.9 mark are highly appreciated.",
      "votes": null
    },
    {
      "id": "1107209",
      "postDate": "12/09/2020 13:39:56",
      "content": "<p>0.893 is a single or k-folds?</p>",
      "rawMarkdown": "0.893 is a single or k-folds?",
      "votes": null
    },
    {
      "id": "1107213",
      "postDate": "12/09/2020 13:45:20",
      "content": "<p>single fold. Has k-fold improved performance?</p>",
      "rawMarkdown": "single fold. Has k-fold improved performance?",
      "votes": null
    },
    {
      "id": "1107228",
      "postDate": "12/09/2020 13:57:22",
      "content": "<p>Try different image size</p>",
      "rawMarkdown": "Try different image size",
      "votes": null
    },
    {
      "id": "1107238",
      "postDate": "12/09/2020 14:03:59",
      "content": "<p>I've used sizes of 256 and 512. the 512 ones gave better accuracy. What size can push the accuracy even higher?</p>",
      "rawMarkdown": "I've used sizes of 256 and 512. the 512 ones gave better accuracy. What size can push the accuracy even higher?",
      "votes": null
    },
    {
      "id": "1107327",
      "postDate": "12/09/2020 15:42:35",
      "content": "<p>I am looking for the same but you can try bigger model, more augs that gave good results, train with SGD momentum they give good results for longer trainings, don't early stop, use k-fold gives 0.1-0.2~ improvement, use label smoothing gives 0.1~, try another model?.</p>",
      "rawMarkdown": "I am looking for the same but you can try bigger model, more augs that gave good results, train with SGD momentum they give good results for longer trainings, don't early stop, use k-fold gives 0.1-0.2~ improvement, use label smoothing gives 0.1~, try another model?.",
      "votes": null
    },
    {
      "id": "1107558",
      "postDate": "12/09/2020 19:18:02",
      "content": "<p>Have you tired Pseudo Labeling?</p>",
      "rawMarkdown": "Have you tired Pseudo Labeling?",
      "votes": null
    },
    {
      "id": "1107746",
      "postDate": "12/09/2020 23:39:26",
      "content": "<p>Trying a bigger efficientnet models does indeed improve the score and is the obvious thing to do. At some point there's a point of diminishing returns in terms of training &amp; inference time (when you may end up trading off [small?] improvements vs. not enough inference time to run other models to ensemble - but who knows exactly where that point is). I think a bunch of other people have already mentioned improvements with the larger models and I've also seen those.</p>",
      "rawMarkdown": "Trying a bigger efficientnet models does indeed improve the score and is the obvious thing to do. At some point there's a point of diminishing returns in terms of training & inference time (when you may end up trading off [small?] improvements vs. not enough inference time to run other models to ensemble - but who knows exactly where that point is). I think a bunch of other people have already mentioned improvements with the larger models and I've also seen those.",
      "votes": null
    },
    {
      "id": "1107897",
      "postDate": "12/10/2020 04:03:55",
      "content": "<p>Yes you're right bigger efficientnet models give some accuracy boost, I did some experiments, found<br>\nEFFNETB2 gave LB = 0.886, EFFNETB4 LB = 0.891, EFFNETB7 LB = 0.893 with nothing changed in pipeline. But as you said the improvements vs. not enough inference time, it do matter yes, when we want to ensemble other models. </p>",
      "rawMarkdown": "Yes you're right bigger efficientnet models give some accuracy boost, I did some experiments, found\nEFFNETB2 gave LB = 0.886, EFFNETB4 LB = 0.891, EFFNETB7 LB = 0.893 with nothing changed in pipeline. But as you said the improvements vs. not enough inference time, it do matter yes, when we want to ensemble other models.",
      "votes": null
    },
    {
      "id": "1107932",
      "postDate": "12/10/2020 05:09:50",
      "content": "<p>Bigger efficientnet models also increase learning time. Like I was trying b8_ap model, it took almost 18 minutes per epoch with just 4 batch size. The b3_ns on the other hand took 7 minutes per epoch for a batch size of 28. So I stopped training of b8_ap. Also, most public notebooks have used b3_ns. So I thought b3_ns was suiting this data.</p>",
      "rawMarkdown": "Bigger efficientnet models also increase learning time. Like I was trying b8_ap model, it took almost 18 minutes per epoch with just 4 batch size. The b3_ns on the other hand took 7 minutes per epoch for a batch size of 28. So I stopped training of b8_ap. Also, most public notebooks have used b3_ns. So I thought b3_ns was suiting this data.",
      "votes": null
    },
    {
      "id": "1107939",
      "postDate": "12/10/2020 05:29:25",
      "content": "<p>Yes they do but i think you're not using TPUs, they can drastically minimize the time, check some notebooks and give it a try. But always first experiment on smaller models, then switch to bigger models and TPU. With B7 batch size can be &gt;32(or 32) on TPU. B3,B4 are the go to though.</p>",
      "rawMarkdown": "Yes they do but i think you're not using TPUs, they can drastically minimize the time, check some notebooks and give it a try. But always first experiment on smaller models, then switch to bigger models and TPU. With B7 batch size can be >32(or 32) on TPU. B3,B4 are the go to though.",
      "votes": null
    },
    {
      "id": "1108123",
      "postDate": "12/10/2020 09:40:42",
      "content": "<p>Well, the image we are given are 800 by 600. So, if you use models that use rectangular input, a sensible upper bound may be 600 (or 800 if you somehow pad or something). Larger then just gets more complicated - I suppose you could start to do things like making the images larger and using super-resolution etc., but doing all that during inference sounds challenging and you'd have to convince me using cross-validation results that it truly adds value.</p>",
      "rawMarkdown": "Well, the image we are given are 800 by 600. So, if you use models that use rectangular input, a sensible upper bound may be 600 (or 800 if you somehow pad or something). Larger then just gets more complicated - I suppose you could start to do things like making the images larger and using super-resolution etc., but doing all that during inference sounds challenging and you'd have to convince me using cross-validation results that it truly adds value.",
      "votes": null
    },
    {
      "id": "1108134",
      "postDate": "12/10/2020 10:06:59",
      "content": "<p>Is FastAI compatible with TPUs? I'm not very sure about that….<br>\nOr is there anything separate that we should do to make FastAI learners run on TPUs?? If yes, then can you share some resources/links on how to do that?</p>",
      "rawMarkdown": "Is FastAI compatible with TPUs? I'm not very sure about that....\nOr is there anything separate that we should do to make FastAI learners run on TPUs?? If yes, then can you share some resources/links on how to do that?",
      "votes": null
    },
    {
      "id": "1108150",
      "postDate": "12/10/2020 10:40:39",
      "content": "<p>No i don't think it does yet, but it should soon as it is based on pytorch and pytorch has got support recently.</p>",
      "rawMarkdown": "No i don't think it does yet, but it should soon as it is based on pytorch and pytorch has got support recently.",
      "votes": null
    },
    {
      "id": "1109406",
      "postDate": "12/11/2020 16:29:09",
      "content": "<p>No not yet. How to do that in FastAI?</p>",
      "rawMarkdown": "No not yet. How to do that in FastAI?",
      "votes": null
    },
    {
      "id": "1109469",
      "postDate": "12/11/2020 18:10:23",
      "content": "<p>In your inference code you would predict for the test set, use the predictions as labels (either pick the most probable category, or only do so if the model is pretty sure, or use soft labels), fine tune the model with those extra images added to the training data and predict again.</p>",
      "rawMarkdown": "In your inference code you would predict for the test set, use the predictions as labels (either pick the most probable category, or only do so if the model is pretty sure, or use soft labels), fine tune the model with those extra images added to the training data and predict again.",
      "votes": null
    },
    {
      "id": "1109591",
      "postDate": "12/11/2020 21:26:27",
      "content": "<p>Hi.</p>\n<p>We got improvements by using 5-fold. In the old scoring leaderboard, it was 0.3-0.4. But it might give less improvement with the current test data. I would suggest you to give it a try.</p>",
      "rawMarkdown": "Hi.\n\nWe got improvements by using 5-fold. In the old scoring leaderboard, it was 0.3-0.4. But it might give less improvement with the current test data. I would suggest you to give it a try.",
      "votes": null
    },
    {
      "id": "1109595",
      "postDate": "12/11/2020 21:31:49",
      "content": "<p>Maybe one improvement can come from the model ensembling. You can try different backbone algorithms for the fine-tuning part other than efficientnet like resnet or resnext and ensemble the results. There are some discussions in the discussion section which were mentioning about the smaller models were giving better performances for the resnet family (e.g. 18 vs 34) </p>",
      "rawMarkdown": "Maybe one improvement can come from the model ensembling. You can try different backbone algorithms for the fine-tuning part other than efficientnet like resnet or resnext and ensemble the results. There are some discussions in the discussion section which were mentioning about the smaller models were giving better performances for the resnet family (e.g. 18 vs 34)",
      "votes": null
    },
    {
      "id": "1127470",
      "postDate": "12/26/2020 15:06:53",
      "content": "<p>Are you guys using k-fold the way Zach Mueller has instructed in his course?<br>\nCan you guide me a bit on how to add k-fold to my current learner?</p>",
      "rawMarkdown": "Are you guys using k-fold the way Zach Mueller has instructed in his course?\nCan you guide me a bit on how to add k-fold to my current learner?",
      "votes": null
    },
    {
      "id": "1127489",
      "postDate": "12/26/2020 15:21:51",
      "content": "<pre><code>def create_cv_split(df, n_splits, *args, **kwargs):\n\n    fold = StratifiedKFold(n_splits=n_splits, *args, **kwargs)\n    for idx, (train_idx, val_idx) in enumerate(fold.split(df, df[\"label\"])):\n        df.loc[val_idx, \"fold\"] = idx\n\n    df[\"fold\"] = df[\"fold\"].astype(int)\n    return df\n\n# df_train is the `train.csv`.\ndf_folds = create_cv_split(df_train.copy(),\n                           n_splits=5,\n                           shuffle=True,\n                           random_state=42)\n</code></pre>\n<p>Hi, I use this code which was taken/adapted from <a href=\"https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training/\" target=\"_blank\">this</a> notebook.</p>",
      "rawMarkdown": "```python\n\ndef create_cv_split(df, n_splits, *args, **kwargs):\n    \n    fold = StratifiedKFold(n_splits=n_splits, *args, **kwargs)\n    for idx, (train_idx, val_idx) in enumerate(fold.split(df, df[\"label\"])):\n        df.loc[val_idx, \"fold\"] = idx\n        \n    df[\"fold\"] = df[\"fold\"].astype(int)\n    return df\n\n# df_train is the `train.csv`.\ndf_folds = create_cv_split(df_train.copy(),\n                           n_splits=5,\n                           shuffle=True,\n                           random_state=42)\n\n```\n\nHi, I use this code which was taken/adapted from [this](https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training/) notebook.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1107209,
      "author_name": "dldmw579",
      "author_url": "",
      "post_date": "12/09/2020 13:39:56",
      "content": "<p>0.893 is a single or k-folds?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107213,
          "author_name": "adamdavis99",
          "author_url": "",
          "post_date": "12/09/2020 13:45:20",
          "content": "<p>single fold. Has k-fold improved performance?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109591,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "12/11/2020 21:26:27",
          "content": "<p>Hi.</p>\n<p>We got improvements by using 5-fold. In the old scoring leaderboard, it was 0.3-0.4. But it might give less improvement with the current test data. I would suggest you to give it a try.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127470,
          "author_name": "adamdavis99",
          "author_url": "",
          "post_date": "12/26/2020 15:06:53",
          "content": "<p>Are you guys using k-fold the way Zach Mueller has instructed in his course?<br>\nCan you guide me a bit on how to add k-fold to my current learner?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127489,
          "author_name": "snnclsr",
          "author_url": "",
          "post_date": "12/26/2020 15:21:51",
          "content": "<pre><code>def create_cv_split(df, n_splits, *args, **kwargs):\n\n    fold = StratifiedKFold(n_splits=n_splits, *args, **kwargs)\n    for idx, (train_idx, val_idx) in enumerate(fold.split(df, df[\"label\"])):\n        df.loc[val_idx, \"fold\"] = idx\n\n    df[\"fold\"] = df[\"fold\"].astype(int)\n    return df\n\n# df_train is the `train.csv`.\ndf_folds = create_cv_split(df_train.copy(),\n                           n_splits=5,\n                           shuffle=True,\n                           random_state=42)\n</code></pre>\n<p>Hi, I use this code which was taken/adapted from <a href=\"https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training/\" target=\"_blank\">this</a> notebook.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1107228,
      "author_name": "khursani8",
      "author_url": "",
      "post_date": "12/09/2020 13:57:22",
      "content": "<p>Try different image size</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107238,
          "author_name": "adamdavis99",
          "author_url": "",
          "post_date": "12/09/2020 14:03:59",
          "content": "<p>I've used sizes of 256 and 512. the 512 ones gave better accuracy. What size can push the accuracy even higher?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108123,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/10/2020 09:40:42",
          "content": "<p>Well, the image we are given are 800 by 600. So, if you use models that use rectangular input, a sensible upper bound may be 600 (or 800 if you somehow pad or something). Larger then just gets more complicated - I suppose you could start to do things like making the images larger and using super-resolution etc., but doing all that during inference sounds challenging and you'd have to convince me using cross-validation results that it truly adds value.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1107327,
      "author_name": "anku5hk",
      "author_url": "",
      "post_date": "12/09/2020 15:42:35",
      "content": "<p>I am looking for the same but you can try bigger model, more augs that gave good results, train with SGD momentum they give good results for longer trainings, don't early stop, use k-fold gives 0.1-0.2~ improvement, use label smoothing gives 0.1~, try another model?.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107746,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/09/2020 23:39:26",
          "content": "<p>Trying a bigger efficientnet models does indeed improve the score and is the obvious thing to do. At some point there's a point of diminishing returns in terms of training &amp; inference time (when you may end up trading off [small?] improvements vs. not enough inference time to run other models to ensemble - but who knows exactly where that point is). I think a bunch of other people have already mentioned improvements with the larger models and I've also seen those.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107897,
          "author_name": "anku5hk",
          "author_url": "",
          "post_date": "12/10/2020 04:03:55",
          "content": "<p>Yes you're right bigger efficientnet models give some accuracy boost, I did some experiments, found<br>\nEFFNETB2 gave LB = 0.886, EFFNETB4 LB = 0.891, EFFNETB7 LB = 0.893 with nothing changed in pipeline. But as you said the improvements vs. not enough inference time, it do matter yes, when we want to ensemble other models. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107932,
          "author_name": "adamdavis99",
          "author_url": "",
          "post_date": "12/10/2020 05:09:50",
          "content": "<p>Bigger efficientnet models also increase learning time. Like I was trying b8_ap model, it took almost 18 minutes per epoch with just 4 batch size. The b3_ns on the other hand took 7 minutes per epoch for a batch size of 28. So I stopped training of b8_ap. Also, most public notebooks have used b3_ns. So I thought b3_ns was suiting this data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107939,
          "author_name": "anku5hk",
          "author_url": "",
          "post_date": "12/10/2020 05:29:25",
          "content": "<p>Yes they do but i think you're not using TPUs, they can drastically minimize the time, check some notebooks and give it a try. But always first experiment on smaller models, then switch to bigger models and TPU. With B7 batch size can be &gt;32(or 32) on TPU. B3,B4 are the go to though.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108134,
          "author_name": "adamdavis99",
          "author_url": "",
          "post_date": "12/10/2020 10:06:59",
          "content": "<p>Is FastAI compatible with TPUs? I'm not very sure about that….<br>\nOr is there anything separate that we should do to make FastAI learners run on TPUs?? If yes, then can you share some resources/links on how to do that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1108150,
          "author_name": "anku5hk",
          "author_url": "",
          "post_date": "12/10/2020 10:40:39",
          "content": "<p>No i don't think it does yet, but it should soon as it is based on pytorch and pytorch has got support recently.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1107558,
      "author_name": "yurilla",
      "author_url": "",
      "post_date": "12/09/2020 19:18:02",
      "content": "<p>Have you tired Pseudo Labeling?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1109406,
          "author_name": "adamdavis99",
          "author_url": "",
          "post_date": "12/11/2020 16:29:09",
          "content": "<p>No not yet. How to do that in FastAI?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1109469,
          "author_name": "bjoernholzhauer",
          "author_url": "",
          "post_date": "12/11/2020 18:10:23",
          "content": "<p>In your inference code you would predict for the test set, use the predictions as labels (either pick the most probable category, or only do so if the model is pretty sure, or use soft labels), fine tune the model with those extra images added to the training data and predict again.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1109595,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "12/11/2020 21:31:49",
      "content": "<p>Maybe one improvement can come from the model ensembling. You can try different backbone algorithms for the fine-tuning part other than efficientnet like resnet or resnext and ensemble the results. There are some discussions in the discussion section which were mentioning about the smaller models were giving better performances for the resnet family (e.g. 18 vs 34) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1107198": "Hi guys, I'm using FastAI and Pytorch to create an efficientnet-b3_ns model for this competition. I've used CutMix as suggested in [this](https://www.kaggle.com/ababino/cutmix-with-fastai-and-efficientnet/notebook) notebook. I've also used a few tweaks from @muellerzr [notebook](https://www.kaggle.com/muellerzr/recreating-abhishek-s-tez-with-fastai). I've used a tta with n=15, achieving an accuracy of close to 0.893. However, now I'm running out of any major tweaks that I can implement to improve my accuracy. Any suggestions to cross the 0.9 mark are highly appreciated.",
    "1107209": "0.893 is a single or k-folds?",
    "1107213": "single fold. Has k-fold improved performance?",
    "1107228": "Try different image size",
    "1107238": "I've used sizes of 256 and 512. the 512 ones gave better accuracy. What size can push the accuracy even higher?",
    "1107327": "I am looking for the same but you can try bigger model, more augs that gave good results, train with SGD momentum they give good results for longer trainings, don't early stop, use k-fold gives 0.1-0.2~ improvement, use label smoothing gives 0.1~, try another model?.",
    "1107558": "Have you tired Pseudo Labeling?",
    "1107746": "Trying a bigger efficientnet models does indeed improve the score and is the obvious thing to do. At some point there's a point of diminishing returns in terms of training & inference time (when you may end up trading off [small?] improvements vs. not enough inference time to run other models to ensemble - but who knows exactly where that point is). I think a bunch of other people have already mentioned improvements with the larger models and I've also seen those.",
    "1107897": "Yes you're right bigger efficientnet models give some accuracy boost, I did some experiments, found\nEFFNETB2 gave LB = 0.886, EFFNETB4 LB = 0.891, EFFNETB7 LB = 0.893 with nothing changed in pipeline. But as you said the improvements vs. not enough inference time, it do matter yes, when we want to ensemble other models.",
    "1107932": "Bigger efficientnet models also increase learning time. Like I was trying b8_ap model, it took almost 18 minutes per epoch with just 4 batch size. The b3_ns on the other hand took 7 minutes per epoch for a batch size of 28. So I stopped training of b8_ap. Also, most public notebooks have used b3_ns. So I thought b3_ns was suiting this data.",
    "1107939": "Yes they do but i think you're not using TPUs, they can drastically minimize the time, check some notebooks and give it a try. But always first experiment on smaller models, then switch to bigger models and TPU. With B7 batch size can be >32(or 32) on TPU. B3,B4 are the go to though.",
    "1108123": "Well, the image we are given are 800 by 600. So, if you use models that use rectangular input, a sensible upper bound may be 600 (or 800 if you somehow pad or something). Larger then just gets more complicated - I suppose you could start to do things like making the images larger and using super-resolution etc., but doing all that during inference sounds challenging and you'd have to convince me using cross-validation results that it truly adds value.",
    "1108134": "Is FastAI compatible with TPUs? I'm not very sure about that....\nOr is there anything separate that we should do to make FastAI learners run on TPUs?? If yes, then can you share some resources/links on how to do that?",
    "1108150": "No i don't think it does yet, but it should soon as it is based on pytorch and pytorch has got support recently.",
    "1109406": "No not yet. How to do that in FastAI?",
    "1109469": "In your inference code you would predict for the test set, use the predictions as labels (either pick the most probable category, or only do so if the model is pretty sure, or use soft labels), fine tune the model with those extra images added to the training data and predict again.",
    "1109591": "Hi.\n\nWe got improvements by using 5-fold. In the old scoring leaderboard, it was 0.3-0.4. But it might give less improvement with the current test data. I would suggest you to give it a try.",
    "1109595": "Maybe one improvement can come from the model ensembling. You can try different backbone algorithms for the fine-tuning part other than efficientnet like resnet or resnext and ensemble the results. There are some discussions in the discussion section which were mentioning about the smaller models were giving better performances for the resnet family (e.g. 18 vs 34)",
    "1127470": "Are you guys using k-fold the way Zach Mueller has instructed in his course?\nCan you guide me a bit on how to add k-fold to my current learner?",
    "1127489": "```python\n\ndef create_cv_split(df, n_splits, *args, **kwargs):\n    \n    fold = StratifiedKFold(n_splits=n_splits, *args, **kwargs)\n    for idx, (train_idx, val_idx) in enumerate(fold.split(df, df[\"label\"])):\n        df.loc[val_idx, \"fold\"] = idx\n        \n    df[\"fold\"] = df[\"fold\"].astype(int)\n    return df\n\n# df_train is the `train.csv`.\ndf_folds = create_cv_split(df_train.copy(),\n                           n_splits=5,\n                           shuffle=True,\n                           random_state=42)\n\n```\n\nHi, I use this code which was taken/adapted from [this](https://www.kaggle.com/yasufuminakama/cassava-resnext50-32x4d-starter-training/) notebook."
  },
  "source": "meta"
}