{
  "id": 205630,
  "title": "How I got LB 0.901 in a simple way.",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/205630",
  "author_name": "",
  "post_date": "2020-12-21T03:09:41.654580800Z",
  "votes": 102,
  "comment_count": 46,
  "views": 0,
  "content": "<p>I got the LB score of 0.901 by simply using <strong>SnapMix</strong> augmentation and model ensemble. </p>\n<p>Detailed results can be found as below.</p>\n<p><strong>Single model result (single fold, no tta, no extra training data)</strong></p>\n<p><strong><em><em></em></em></strong>Resnet50  : LB 0.891  <a href=\"https://www.kaggle.com/shaolihuang/training-with-snapmix\" target=\"_blank\">Notebook for training </a></p>\n<p><strong>Model ensemble ( no tta, no extra training data):</strong></p>\n<ol>\n<li><p>Resnet50 (5 folds) :  LB 0.897</p></li>\n<li><p>Resnet50 (5 folds) + Resnet101(5 folds):  LB 0.899</p></li>\n<li><p>Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds):  LB 0.901</p></li>\n</ol>",
  "messages": [
    {
      "id": "1120662",
      "postDate": "12/21/2020 03:09:41",
      "content": "<p>I got the LB score of 0.901 by simply using <strong>SnapMix</strong> augmentation and model ensemble. </p>\n<p>Detailed results can be found as below.</p>\n<p><strong>Single model result (single fold, no tta, no extra training data)</strong></p>\n<p><strong><em><em></em></em></strong>Resnet50  : LB 0.891  <a href=\"https://www.kaggle.com/shaolihuang/training-with-snapmix\" target=\"_blank\">Notebook for training </a></p>\n<p><strong>Model ensemble ( no tta, no extra training data):</strong></p>\n<ol>\n<li><p>Resnet50 (5 folds) :  LB 0.897</p></li>\n<li><p>Resnet50 (5 folds) + Resnet101(5 folds):  LB 0.899</p></li>\n<li><p>Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds):  LB 0.901</p></li>\n</ol>",
      "rawMarkdown": "I got the LB score of 0.901 by simply using **SnapMix** augmentation and model ensemble. \n\nDetailed results can be found as below.\n\n**Single model result (single fold, no tta, no extra training data)**\n \n  ********Resnet50  : LB 0.891  [Notebook for training ](https://www.kaggle.com/shaolihuang/training-with-snapmix)\n\n**Model ensemble ( no tta, no extra training data):**\n\n1. Resnet50 (5 folds) :  LB 0.897\n\n2. Resnet50 (5 folds) + Resnet101(5 folds):  LB 0.899\n\n3. Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds):  LB 0.901",
      "votes": null
    },
    {
      "id": "1120737",
      "postDate": "12/21/2020 04:54:45",
      "content": "<p>Hi,the SanpMix augmentation is what?Thanks!</p>",
      "rawMarkdown": "Hi,the SanpMix augmentation is what?Thanks!",
      "votes": null
    },
    {
      "id": "1120749",
      "postDate": "12/21/2020 04:59:12",
      "content": "<p>SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained Data<br>\nCode: <a href=\"url\" target=\"_blank\">https://github.com/Shaoli-Huang/SnapMix</a></p>",
      "rawMarkdown": "SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained Data\nCode: [https://github.com/Shaoli-Huang/SnapMix](url)",
      "votes": null
    },
    {
      "id": "1120752",
      "postDate": "12/21/2020 05:03:43",
      "content": "<p>Thanks!I will test it</p>",
      "rawMarkdown": "Thanks!I will test it",
      "votes": null
    },
    {
      "id": "1120758",
      "postDate": "12/21/2020 05:07:34",
      "content": "<p>you can find some training detail that I tried from the <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/204631\" target=\"_blank\">link</a></p>",
      "rawMarkdown": "you can find some training detail that I tried from the [link](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/204631)",
      "votes": null
    },
    {
      "id": "1120786",
      "postDate": "12/21/2020 05:42:47",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/David\" target=\"_blank\">@David</a> . Thank you for sharing. I'm new in the kaggle contest. Can you tell me how marge Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds) models?  I not understand the tecniques of marge two or three models. Thanks in an advance. </p>",
      "rawMarkdown": "Hey @David . Thank you for sharing. I'm new in the kaggle contest. Can you tell me how marge Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds) models?  I not understand the tecniques of marge two or three models. Thanks in an advance.",
      "votes": null
    },
    {
      "id": "1120797",
      "postDate": "12/21/2020 05:51:43",
      "content": "<p>I used Voting Ensembles.<br>\nif you have n models and m images, you can put all the prediction results as an array with shape nxm.<br>\nThen I used the following code to pick the final prediction.</p>\n<p>pres = []<br>\nfor i in range(allresults.shape[1]):<br>\n    pre= Counter(allresults[:,i]).most_common(1)<br>\n    pres.append(pre[0][0])</p>",
      "rawMarkdown": "I used Voting Ensembles.\nif you have n models and m images, you can put all the prediction results as an array with shape nxm.\nThen I used the following code to pick the final prediction.\n\npres = []\nfor i in range(allresults.shape[1]):\n    pre= Counter(allresults[:,i]).most_common(1)\n    pres.append(pre[0][0])",
      "votes": null
    },
    {
      "id": "1120814",
      "postDate": "12/21/2020 06:09:52",
      "content": "<p>Thank you so much that's make sense. Can you please share any notebook which use this technique!  Then it help me a lot.  </p>",
      "rawMarkdown": "Thank you so much that's make sense. Can you please share any notebook which use this technique!  Then it help me a lot.",
      "votes": null
    },
    {
      "id": "1120931",
      "postDate": "12/21/2020 08:01:25",
      "content": "<p>I'm curious, when you use different sizes of the same architecture, does the smallest resent still add value in the final ensemble (e.g. did you try the Resnet101 alone?)? I would have been tempted to do different architectures. </p>",
      "rawMarkdown": "I'm curious, when you use different sizes of the same architecture, does the smallest resent still add value in the final ensemble (e.g. did you try the Resnet101 alone?)? I would have been tempted to do different architectures.",
      "votes": null
    },
    {
      "id": "1121318",
      "postDate": "12/21/2020 14:46:39",
      "content": "<p>I would consider keeping the raw probabilities of each label, adding those, and then taking the label with the max as your final answer. Seems intuitively that it is using more information than if you just take the most common label.</p>",
      "rawMarkdown": "I would consider keeping the raw probabilities of each label, adding those, and then taking the label with the max as your final answer. Seems intuitively that it is using more information than if you just take the most common label.",
      "votes": null
    },
    {
      "id": "1121320",
      "postDate": "12/21/2020 14:49:48",
      "content": "<p>Thank you for sharing. Just a quick question, how much time it took to train the whole thing and what was the submission time ?</p>",
      "rawMarkdown": "Thank you for sharing. Just a quick question, how much time it took to train the whole thing and what was the submission time ?",
      "votes": null
    },
    {
      "id": "1121882",
      "postDate": "12/22/2020 01:54:57",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> , nice results, you mean that the only data augmentation that you used was <code>SnapMix</code>?</p>",
      "rawMarkdown": "Hey @shaolihuang , nice results, you mean that the only data augmentation that you used was `SnapMix`?",
      "votes": null
    },
    {
      "id": "1121922",
      "postDate": "12/22/2020 02:52:57",
      "content": "<p>oh, I just realize you're the 1st author of the method, I was wondering what kind of awesome guys can get quickly updated about latest papers, now all makes sense.👍</p>",
      "rawMarkdown": "oh, I just realize you're the 1st author of the method, I was wondering what kind of awesome guys can get quickly updated about latest papers, now all makes sense.👍",
      "votes": null
    },
    {
      "id": "1121963",
      "postDate": "12/22/2020 04:08:43",
      "content": "<p>Are you using keras or pytorch? 5-10 fold cross validation with 10 epochs takes a long time on GPU (perhaps more than 9 hours limit). Are you using your own GPU or are doing something else to avoid the 9 hour limit?</p>",
      "rawMarkdown": "Are you using keras or pytorch? 5-10 fold cross validation with 10 epochs takes a long time on GPU (perhaps more than 9 hours limit). Are you using your own GPU or are doing something else to avoid the 9 hour limit?",
      "votes": null
    },
    {
      "id": "1122138",
      "postDate": "12/22/2020 08:13:48",
      "content": "<p><a href=\"https://www.kaggle.com/David\" target=\"_blank\">@David</a> thanks to u  for telling everyone  the SnapMix augmentation method. 😃</p>",
      "rawMarkdown": "David thanks to u  for telling everyone  the SnapMix augmentation method. 😃",
      "votes": null
    },
    {
      "id": "1122326",
      "postDate": "12/22/2020 11:01:48",
      "content": "<p>I trained models offline and upload them as a dataset.</p>",
      "rawMarkdown": "I trained models offline and upload them as a dataset.",
      "votes": null
    },
    {
      "id": "1122333",
      "postDate": "12/22/2020 11:06:19",
      "content": "<p>Thx, I am trying to test whether Snapmix is effective in other datasets. Then, I saw someone mention it in this competition,  so I just did some quick experiments to see if it works on this dataset.</p>",
      "rawMarkdown": "Thx, I am trying to test whether Snapmix is effective in other datasets. Then, I saw someone mention it in this competition,  so I just did some quick experiments to see if it works on this dataset.",
      "votes": null
    },
    {
      "id": "1122337",
      "postDate": "12/22/2020 11:10:25",
      "content": "<p>Concretely speaking, I used SnapMix coupled with some standard augmentation practice (including random crop and random flip).</p>",
      "rawMarkdown": "Concretely speaking, I used SnapMix coupled with some standard augmentation practice (including random crop and random flip).",
      "votes": null
    },
    {
      "id": "1122342",
      "postDate": "12/22/2020 11:13:44",
      "content": "<p>The default transforms used in the repository has only a few basic transforms as shown <a href=\"https://github.com/Shaoli-Huang/SnapMix/blob/8d4aa030731f835eeefeee404e0fab98a1a1d4de/datasets/tfs.py#L18\" target=\"_blank\">here</a>, i suppose only these were used ?</p>\n<p>EDIT : I've open sourced my pipeline <a href=\"https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline\" target=\"_blank\">here</a> for people to experiment</p>",
      "rawMarkdown": "The default transforms used in the repository has only a few basic transforms as shown [here](https://github.com/Shaoli-Huang/SnapMix/blob/8d4aa030731f835eeefeee404e0fab98a1a1d4de/datasets/tfs.py#L18), i suppose only these were used ?\n\nEDIT : I've open sourced my pipeline [here](https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline) for people to experiment",
      "votes": null
    },
    {
      "id": "1122346",
      "postDate": "12/22/2020 11:16:30",
      "content": "<p>your paper is fantastic and does make more sense, nice job!</p>",
      "rawMarkdown": "your paper is fantastic and does make more sense, nice job!",
      "votes": null
    },
    {
      "id": "1122351",
      "postDate": "12/22/2020 11:24:09",
      "content": "<p>I do not do statistics on the specific time for training all models. I found the following related information from the training log, I hope it helps.</p>\n<p>Training a single fold of resnet50 for 40 epochs took around 4 hours.</p>\n<p>Training a single fold of resnet101 for 40 epochs took around 6 hours.</p>\n<p>I upload all the trained models as a dataset.  The submission of ensembling all models took around two hours.</p>",
      "rawMarkdown": "I do not do statistics on the specific time for training all models. I found the following related information from the training log, I hope it helps.\n\nTraining a single fold of resnet50 for 40 epochs took around 4 hours.\n\nTraining a single fold of resnet101 for 40 epochs took around 6 hours.\n\n\n\nI upload all the trained models as a dataset.  The submission of ensembling all models took around two hours.",
      "votes": null
    },
    {
      "id": "1122485",
      "postDate": "12/22/2020 13:24:33",
      "content": "<p>Hello! How much time did you spend on inference?</p>",
      "rawMarkdown": "Hello! How much time did you spend on inference?",
      "votes": null
    },
    {
      "id": "1122521",
      "postDate": "12/22/2020 13:49:23",
      "content": "<p>Applying SnapMix did not increase any inference time. In other words, my method has exactly the same inference time as the standard Resnet-50.</p>\n<p>when I evaluated the validation dataset containing 4280 images,  it took 20 seconds (using my own GPU). </p>",
      "rawMarkdown": "Applying SnapMix did not increase any inference time. In other words, my method has exactly the same inference time as the standard Resnet-50.\n\nwhen I evaluated the validation dataset containing 4280 images,  it took 20 seconds (using my own GPU).",
      "votes": null
    },
    {
      "id": "1122533",
      "postDate": "12/22/2020 14:01:37",
      "content": "<p><a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> Thank you for sharing. Have you tested Resnet101(5 folds) and ResNext101(5 folds) separately? And what's the LB score of them?</p>",
      "rawMarkdown": "shaolihuang Thank you for sharing. Have you tested Resnet101(5 folds) and ResNext101(5 folds) separately? And what's the LB score of them?",
      "votes": null
    },
    {
      "id": "1122537",
      "postDate": "12/22/2020 14:04:56",
      "content": "<p>Oh,how long about inference on the whole test set when ranking ?</p>",
      "rawMarkdown": "Oh,how long about inference on the whole test set when ranking ?",
      "votes": null
    },
    {
      "id": "1123013",
      "postDate": "12/22/2020 21:10:50",
      "content": "<p>thank you,  <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a>,  for your feedback. Are you using a packaged method such as sci-kit learn to do CV? or implemented it yourself? I am using PyTorch and am wondering if there are CV packages that support it?</p>",
      "rawMarkdown": "thank you,  @shaolihuang,  for your feedback. Are you using a packaged method such as sci-kit learn to do CV? or implemented it yourself? I am using PyTorch and am wondering if there are CV packages that support it?",
      "votes": null
    },
    {
      "id": "1124919",
      "postDate": "12/24/2020 09:15:18",
      "content": "<p>what is fold?</p>",
      "rawMarkdown": "what is fold?",
      "votes": null
    },
    {
      "id": "1124920",
      "postDate": "12/24/2020 09:15:38",
      "content": "<p>I just know epochs, what is fold</p>",
      "rawMarkdown": "I just know epochs, what is fold",
      "votes": null
    },
    {
      "id": "1124967",
      "postDate": "12/24/2020 09:52:25",
      "content": "<p><a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> were your models pretrained on ImageNet?</p>",
      "rawMarkdown": "shaolihuang were your models pretrained on ImageNet?",
      "votes": null
    },
    {
      "id": "1125069",
      "postDate": "12/24/2020 11:22:57",
      "content": "<p>Yes,  pretrained model from torchvision </p>",
      "rawMarkdown": "Yes,  pretrained model from torchvision",
      "votes": null
    },
    {
      "id": "1126828",
      "postDate": "12/26/2020 02:30:07",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a>, thank you for sharing your insight. Just to confirm are you using a voting ensemble of 18 models (5 folds x 3 models) then? So at test time you are making 18 predictions per image? </p>",
      "rawMarkdown": "Hey @shaolihuang, thank you for sharing your insight. Just to confirm are you using a voting ensemble of 18 models (5 folds x 3 models) then? So at test time you are making 18 predictions per image?",
      "votes": null
    },
    {
      "id": "1126854",
      "postDate": "12/26/2020 03:31:55",
      "content": "<p>Actually it should be 15 models.<br>\nAnd yes,  make 15 predictions in total.</p>",
      "rawMarkdown": "Actually it should be 15 models.\nAnd yes,  make 15 predictions in total.",
      "votes": null
    },
    {
      "id": "1126866",
      "postDate": "12/26/2020 03:58:17",
      "content": "<p>comment deleted</p>",
      "rawMarkdown": "comment deleted",
      "votes": null
    },
    {
      "id": "1127020",
      "postDate": "12/26/2020 07:12:55",
      "content": "<p>PyTorch does not have any inbuilt package. You can use the sci-kit learn stratified k fold like <a href=\"https://github.com/svishnu88/Cassava/blob/3734fabe28d6f6dd8c6ca21b3daacfba9822816b/cassavadata.py#L67\" target=\"_blank\">here</a> . </p>",
      "rawMarkdown": "PyTorch does not have any inbuilt package. You can use the sci-kit learn stratified k fold like [here](https://github.com/svishnu88/Cassava/blob/3734fabe28d6f6dd8c6ca21b3daacfba9822816b/cassavadata.py#L67) .",
      "votes": null
    },
    {
      "id": "1127524",
      "postDate": "12/26/2020 15:58:24",
      "content": "<p><a href=\"https://towardsdatascience.com/complete-guide-to-pythons-cross-validation-with-examples-a9676b5cac12\" target=\"_blank\">https://towardsdatascience.com/complete-guide-to-pythons-cross-validation-with-examples-a9676b5cac12</a></p>",
      "rawMarkdown": "https://towardsdatascience.com/complete-guide-to-pythons-cross-validation-with-examples-a9676b5cac12",
      "votes": null
    },
    {
      "id": "1128367",
      "postDate": "12/27/2020 11:33:08",
      "content": "<p>The number of <strong><em>\"epochs\"</em></strong> is a hyperparameter that defines the number of times that the learning algorithm will work through the entire training dataset. One epoch means that each sample in the training dataset has had an opportunity to update the internal model parameters. Furthermore, an epoch is comprised of one or more batches (to help the training).</p>\n<p>Differently, with <strong><em>\"fold\"</em></strong>, we refer to a specific resampling of the validation data. This term is related to the process of cross-validation: a process through which you create a k (interchangeable) number of resampled validation sets to validate your model. In this way, you will have a k number of validation results instead of just one, making your results more robust. Nevertheless, this robustness comes with some contra, such as the increasing time in training. In this sense, <em>leave one out cross-validation</em> is the most demanding type of validation, where the number of folds equals the number of instances in the data set. </p>",
      "rawMarkdown": "The number of ***\"epochs\"*** is a hyperparameter that defines the number of times that the learning algorithm will work through the entire training dataset. One epoch means that each sample in the training dataset has had an opportunity to update the internal model parameters. Furthermore, an epoch is comprised of one or more batches (to help the training).\n\nDifferently, with ***\"fold\"***, we refer to a specific resampling of the validation data. This term is related to the process of cross-validation: a process through which you create a k (interchangeable) number of resampled validation sets to validate your model. In this way, you will have a k number of validation results instead of just one, making your results more robust. Nevertheless, this robustness comes with some contra, such as the increasing time in training. In this sense, *leave one out cross-validation* is the most demanding type of validation, where the number of folds equals the number of instances in the data set.",
      "votes": null
    },
    {
      "id": "1132262",
      "postDate": "12/30/2020 09:17:24",
      "content": "<p>How long will it take? For me only a single resnet50 model will cost about half an hour….. And what is your LB by single fold resnet50 without snapMix ?</p>",
      "rawMarkdown": "How long will it take? For me only a single resnet50 model will cost about half an hour..... And what is your LB by single fold resnet50 without snapMix ?",
      "votes": null
    },
    {
      "id": "1133566",
      "postDate": "12/31/2020 10:18:09",
      "content": "<p>Hi David,<br>\nCan you please refer some good sources to read about snapmix?</p>",
      "rawMarkdown": "Hi David,\nCan you please refer some good sources to read about snapmix?",
      "votes": null
    },
    {
      "id": "1133690",
      "postDate": "12/31/2020 12:44:51",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/Marto93\" target=\"_blank\">@Marto93</a> for replying there. I have a (probably dumb) question about how k-fold CV is used by kagglers. My understanding was always that people use CV to estimate their model performance, but then they may re-train using the full data available before they deploy it. The OP is posting his leaderboard performance, and not the CV performance. Do folks on here typically average the predictions from their k models? Or are they re-training on the full dataset? Thank you so much in advance if you choose to answer me! </p>",
      "rawMarkdown": "Thanks @Marto93 for replying there. I have a (probably dumb) question about how k-fold CV is used by kagglers. My understanding was always that people use CV to estimate their model performance, but then they may re-train using the full data available before they deploy it. The OP is posting his leaderboard performance, and not the CV performance. Do folks on here typically average the predictions from their k models? Or are they re-training on the full dataset? Thank you so much in advance if you choose to answer me!",
      "votes": null
    },
    {
      "id": "1135159",
      "postDate": "01/02/2021 00:09:01",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1135381",
      "postDate": "01/02/2021 07:56:18",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> <br>\nThanks for sharing insights.<br>\nCan you tell me what image size do you prefer for training all models and whats your PC config?</p>",
      "rawMarkdown": "Hi @shaolihuang \nThanks for sharing insights.\nCan you tell me what image size do you prefer for training all models and whats your PC config?",
      "votes": null
    },
    {
      "id": "1141237",
      "postDate": "01/06/2021 15:22:03",
      "content": "<p>Hi! Thank you for the introducing Snapmix method. I have one question. Your Resnet50 (5-folds) performs better then single fold Resnet50. My understanding of cross validation fold is that it is being used for a more robust estimate of the model itself, instead of improving performance.<br>\nSo my assumption on your 5-fold Resnet model is as follows:</p>\n<ol>\n<li>Train 5 models with Kfolds, each model trained on 80% of the data.</li>\n<li>During inference, use these 5 models to output the predicted probabilities separately</li>\n<li>Label the class with the maximum probabilities for each test example. (soft-voting)</li>\n</ol>\n<p>i hope you can enlighten me with this clarification that i have, would really appreciate it! Thank you:)</p>",
      "rawMarkdown": "Hi! Thank you for the introducing Snapmix method. I have one question. Your Resnet50 (5-folds) performs better then single fold Resnet50. My understanding of cross validation fold is that it is being used for a more robust estimate of the model itself, instead of improving performance.\nSo my assumption on your 5-fold Resnet model is as follows:\n1. Train 5 models with Kfolds, each model trained on 80% of the data.\n2. During inference, use these 5 models to output the predicted probabilities separately\n3. Label the class with the maximum probabilities for each test example. (soft-voting)\n\ni hope you can enlighten me with this clarification that i have, would really appreciate it! Thank you:)",
      "votes": null
    },
    {
      "id": "1144595",
      "postDate": "01/08/2021 14:57:00",
      "content": "<p>Paper: <a href=\"https://arxiv.org/abs/2012.04846\" target=\"_blank\">https://arxiv.org/abs/2012.04846</a><br>\nPytorch Github repository: <a href=\"https://github.com/Shaoli-Huang/SnapMix\" target=\"_blank\">https://github.com/Shaoli-Huang/SnapMix</a></p>",
      "rawMarkdown": "Paper: https://arxiv.org/abs/2012.04846\nPytorch Github repository: https://github.com/Shaoli-Huang/SnapMix",
      "votes": null
    },
    {
      "id": "1144600",
      "postDate": "01/08/2021 14:58:08",
      "content": "<p>Check out this post <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209136\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209136</a></p>",
      "rawMarkdown": "Check out this post https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209136",
      "votes": null
    },
    {
      "id": "1145156",
      "postDate": "01/09/2021 00:11:07",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/aqx\" target=\"_blank\">@aqx</a>. I'm far from sure, but I think the author is probably averaging the probabilities of the five folds at test time instead of selecting the model with the most confident prediction as you are suggesting. You can check this out in this <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>\n<p>Check out this line of code in the last cell:</p>\n<p><code>tst_preds = np.mean(tst_preds, axis=0)</code></p>",
      "rawMarkdown": "Hey @aqx. I'm far from sure, but I think the author is probably averaging the probabilities of the five folds at test time instead of selecting the model with the most confident prediction as you are suggesting. You can check this out in this [notebook](https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta) by @khyeh0719 \n\nCheck out this line of code in the last cell:\n\n`tst_preds = np.mean(tst_preds, axis=0) `",
      "votes": null
    },
    {
      "id": "1145419",
      "postDate": "01/09/2021 06:07:49",
      "content": "<p>Thank you for the explanation! </p>",
      "rawMarkdown": "Thank you for the explanation!",
      "votes": null
    },
    {
      "id": "1156179",
      "postDate": "01/17/2021 01:56:30",
      "content": "<p>Very helpful information. Thanks for sharing.</p>\n<p>Can I ask what machine are you using since I saw from the other comment that you're doing all the training offline?</p>",
      "rawMarkdown": "Very helpful information. Thanks for sharing.\n\nCan I ask what machine are you using since I saw from the other comment that you're doing all the training offline?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1120737,
      "author_name": "bcwang",
      "author_url": "",
      "post_date": "12/21/2020 04:54:45",
      "content": "<p>Hi,the SanpMix augmentation is what?Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1120749,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/21/2020 04:59:12",
          "content": "<p>SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained Data<br>\nCode: <a href=\"url\" target=\"_blank\">https://github.com/Shaoli-Huang/SnapMix</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120752,
          "author_name": "bcwang",
          "author_url": "",
          "post_date": "12/21/2020 05:03:43",
          "content": "<p>Thanks!I will test it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1120758,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/21/2020 05:07:34",
          "content": "<p>you can find some training detail that I tried from the <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/204631\" target=\"_blank\">link</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120786,
      "author_name": "aifahim",
      "author_url": "",
      "post_date": "12/21/2020 05:42:47",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/David\" target=\"_blank\">@David</a> . Thank you for sharing. I'm new in the kaggle contest. Can you tell me how marge Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds) models?  I not understand the tecniques of marge two or three models. Thanks in an advance. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1120797,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/21/2020 05:51:43",
          "content": "<p>I used Voting Ensembles.<br>\nif you have n models and m images, you can put all the prediction results as an array with shape nxm.<br>\nThen I used the following code to pick the final prediction.</p>\n<p>pres = []<br>\nfor i in range(allresults.shape[1]):<br>\n    pre= Counter(allresults[:,i]).most_common(1)<br>\n    pres.append(pre[0][0])</p>",
          "votes": null,
          "replies": [
            {
              "id": 1121318,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "12/21/2020 14:46:39",
              "content": "<p>I would consider keeping the raw probabilities of each label, adding those, and then taking the label with the max as your final answer. Seems intuitively that it is using more information than if you just take the most common label.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1120814,
          "author_name": "aifahim",
          "author_url": "",
          "post_date": "12/21/2020 06:09:52",
          "content": "<p>Thank you so much that's make sense. Can you please share any notebook which use this technique!  Then it help me a lot.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126828,
          "author_name": "kagglethomas88",
          "author_url": "",
          "post_date": "12/26/2020 02:30:07",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a>, thank you for sharing your insight. Just to confirm are you using a voting ensemble of 18 models (5 folds x 3 models) then? So at test time you are making 18 predictions per image? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1126854,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/26/2020 03:31:55",
          "content": "<p>Actually it should be 15 models.<br>\nAnd yes,  make 15 predictions in total.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1132262,
          "author_name": "clwwlc",
          "author_url": "",
          "post_date": "12/30/2020 09:17:24",
          "content": "<p>How long will it take? For me only a single resnet50 model will cost about half an hour….. And what is your LB by single fold resnet50 without snapMix ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1120931,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/21/2020 08:01:25",
      "content": "<p>I'm curious, when you use different sizes of the same architecture, does the smallest resent still add value in the final ensemble (e.g. did you try the Resnet101 alone?)? I would have been tempted to do different architectures. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1121320,
      "author_name": "atharvaingle",
      "author_url": "",
      "post_date": "12/21/2020 14:49:48",
      "content": "<p>Thank you for sharing. Just a quick question, how much time it took to train the whole thing and what was the submission time ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122351,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/22/2020 11:24:09",
          "content": "<p>I do not do statistics on the specific time for training all models. I found the following related information from the training log, I hope it helps.</p>\n<p>Training a single fold of resnet50 for 40 epochs took around 4 hours.</p>\n<p>Training a single fold of resnet101 for 40 epochs took around 6 hours.</p>\n<p>I upload all the trained models as a dataset.  The submission of ensembling all models took around two hours.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1156179,
          "author_name": "jasondolorso",
          "author_url": "",
          "post_date": "01/17/2021 01:56:30",
          "content": "<p>Very helpful information. Thanks for sharing.</p>\n<p>Can I ask what machine are you using since I saw from the other comment that you're doing all the training offline?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121882,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "12/22/2020 01:54:57",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> , nice results, you mean that the only data augmentation that you used was <code>SnapMix</code>?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122337,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/22/2020 11:10:25",
          "content": "<p>Concretely speaking, I used SnapMix coupled with some standard augmentation practice (including random crop and random flip).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122342,
          "author_name": "sachinprabhu",
          "author_url": "",
          "post_date": "12/22/2020 11:13:44",
          "content": "<p>The default transforms used in the repository has only a few basic transforms as shown <a href=\"https://github.com/Shaoli-Huang/SnapMix/blob/8d4aa030731f835eeefeee404e0fab98a1a1d4de/datasets/tfs.py#L18\" target=\"_blank\">here</a>, i suppose only these were used ?</p>\n<p>EDIT : I've open sourced my pipeline <a href=\"https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline\" target=\"_blank\">here</a> for people to experiment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121922,
      "author_name": "plugin1689",
      "author_url": "",
      "post_date": "12/22/2020 02:52:57",
      "content": "<p>oh, I just realize you're the 1st author of the method, I was wondering what kind of awesome guys can get quickly updated about latest papers, now all makes sense.👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122333,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/22/2020 11:06:19",
          "content": "<p>Thx, I am trying to test whether Snapmix is effective in other datasets. Then, I saw someone mention it in this competition,  so I just did some quick experiments to see if it works on this dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122346,
          "author_name": "plugin1689",
          "author_url": "",
          "post_date": "12/22/2020 11:16:30",
          "content": "<p>your paper is fantastic and does make more sense, nice job!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1121963,
      "author_name": "kamranmo",
      "author_url": "",
      "post_date": "12/22/2020 04:08:43",
      "content": "<p>Are you using keras or pytorch? 5-10 fold cross validation with 10 epochs takes a long time on GPU (perhaps more than 9 hours limit). Are you using your own GPU or are doing something else to avoid the 9 hour limit?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122326,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/22/2020 11:01:48",
          "content": "<p>I trained models offline and upload them as a dataset.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1123013,
          "author_name": "kamranmo",
          "author_url": "",
          "post_date": "12/22/2020 21:10:50",
          "content": "<p>thank you,  <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a>,  for your feedback. Are you using a packaged method such as sci-kit learn to do CV? or implemented it yourself? I am using PyTorch and am wondering if there are CV packages that support it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1127020,
          "author_name": "vishnus",
          "author_url": "",
          "post_date": "12/26/2020 07:12:55",
          "content": "<p>PyTorch does not have any inbuilt package. You can use the sci-kit learn stratified k fold like <a href=\"https://github.com/svishnu88/Cassava/blob/3734fabe28d6f6dd8c6ca21b3daacfba9822816b/cassavadata.py#L67\" target=\"_blank\">here</a> . </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122138,
      "author_name": "saurabhshahane",
      "author_url": "",
      "post_date": "12/22/2020 08:13:48",
      "content": "<p><a href=\"https://www.kaggle.com/David\" target=\"_blank\">@David</a> thanks to u  for telling everyone  the SnapMix augmentation method. 😃</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1122485,
      "author_name": "zekun98",
      "author_url": "",
      "post_date": "12/22/2020 13:24:33",
      "content": "<p>Hello! How much time did you spend on inference?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1122521,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/22/2020 13:49:23",
          "content": "<p>Applying SnapMix did not increase any inference time. In other words, my method has exactly the same inference time as the standard Resnet-50.</p>\n<p>when I evaluated the validation dataset containing 4280 images,  it took 20 seconds (using my own GPU). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1122537,
          "author_name": "zekun98",
          "author_url": "",
          "post_date": "12/22/2020 14:04:56",
          "content": "<p>Oh,how long about inference on the whole test set when ranking ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1122533,
      "author_name": "wjfearth",
      "author_url": "",
      "post_date": "12/22/2020 14:01:37",
      "content": "<p><a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> Thank you for sharing. Have you tested Resnet101(5 folds) and ResNext101(5 folds) separately? And what's the LB score of them?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1124919,
      "author_name": "mme20172021",
      "author_url": "",
      "post_date": "12/24/2020 09:15:18",
      "content": "<p>what is fold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1127524,
          "author_name": "mknzfr",
          "author_url": "",
          "post_date": "12/26/2020 15:58:24",
          "content": "<p><a href=\"https://towardsdatascience.com/complete-guide-to-pythons-cross-validation-with-examples-a9676b5cac12\" target=\"_blank\">https://towardsdatascience.com/complete-guide-to-pythons-cross-validation-with-examples-a9676b5cac12</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1128367,
          "author_name": "marto24",
          "author_url": "",
          "post_date": "12/27/2020 11:33:08",
          "content": "<p>The number of <strong><em>\"epochs\"</em></strong> is a hyperparameter that defines the number of times that the learning algorithm will work through the entire training dataset. One epoch means that each sample in the training dataset has had an opportunity to update the internal model parameters. Furthermore, an epoch is comprised of one or more batches (to help the training).</p>\n<p>Differently, with <strong><em>\"fold\"</em></strong>, we refer to a specific resampling of the validation data. This term is related to the process of cross-validation: a process through which you create a k (interchangeable) number of resampled validation sets to validate your model. In this way, you will have a k number of validation results instead of just one, making your results more robust. Nevertheless, this robustness comes with some contra, such as the increasing time in training. In this sense, <em>leave one out cross-validation</em> is the most demanding type of validation, where the number of folds equals the number of instances in the data set. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1133690,
          "author_name": "funky15",
          "author_url": "",
          "post_date": "12/31/2020 12:44:51",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/Marto93\" target=\"_blank\">@Marto93</a> for replying there. I have a (probably dumb) question about how k-fold CV is used by kagglers. My understanding was always that people use CV to estimate their model performance, but then they may re-train using the full data available before they deploy it. The OP is posting his leaderboard performance, and not the CV performance. Do folks on here typically average the predictions from their k models? Or are they re-training on the full dataset? Thank you so much in advance if you choose to answer me! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1124920,
      "author_name": "mme20172021",
      "author_url": "",
      "post_date": "12/24/2020 09:15:38",
      "content": "<p>I just know epochs, what is fold</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144600,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "01/08/2021 14:58:08",
          "content": "<p>Check out this post <a href=\"https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209136\" target=\"_blank\">https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209136</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1124967,
      "author_name": "dumbluck",
      "author_url": "",
      "post_date": "12/24/2020 09:52:25",
      "content": "<p><a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> were your models pretrained on ImageNet?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1125069,
          "author_name": "shaolihuang",
          "author_url": "",
          "post_date": "12/24/2020 11:22:57",
          "content": "<p>Yes,  pretrained model from torchvision </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1126866,
      "author_name": "plugin1689",
      "author_url": "",
      "post_date": "12/26/2020 03:58:17",
      "content": "<p>comment deleted</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1133566,
      "author_name": "prvnkmr",
      "author_url": "",
      "post_date": "12/31/2020 10:18:09",
      "content": "<p>Hi David,<br>\nCan you please refer some good sources to read about snapmix?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1144595,
          "author_name": "bessenyeiszilrd",
          "author_url": "",
          "post_date": "01/08/2021 14:57:00",
          "content": "<p>Paper: <a href=\"https://arxiv.org/abs/2012.04846\" target=\"_blank\">https://arxiv.org/abs/2012.04846</a><br>\nPytorch Github repository: <a href=\"https://github.com/Shaoli-Huang/SnapMix\" target=\"_blank\">https://github.com/Shaoli-Huang/SnapMix</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1135159,
      "author_name": "kingarthurie",
      "author_url": "",
      "post_date": "01/02/2021 00:09:01",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1135381,
      "author_name": "suchitnasre",
      "author_url": "",
      "post_date": "01/02/2021 07:56:18",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/shaolihuang\" target=\"_blank\">@shaolihuang</a> <br>\nThanks for sharing insights.<br>\nCan you tell me what image size do you prefer for training all models and whats your PC config?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1141237,
      "author_name": "angqx95",
      "author_url": "",
      "post_date": "01/06/2021 15:22:03",
      "content": "<p>Hi! Thank you for the introducing Snapmix method. I have one question. Your Resnet50 (5-folds) performs better then single fold Resnet50. My understanding of cross validation fold is that it is being used for a more robust estimate of the model itself, instead of improving performance.<br>\nSo my assumption on your 5-fold Resnet model is as follows:</p>\n<ol>\n<li>Train 5 models with Kfolds, each model trained on 80% of the data.</li>\n<li>During inference, use these 5 models to output the predicted probabilities separately</li>\n<li>Label the class with the maximum probabilities for each test example. (soft-voting)</li>\n</ol>\n<p>i hope you can enlighten me with this clarification that i have, would really appreciate it! Thank you:)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1145156,
          "author_name": "vedder",
          "author_url": "",
          "post_date": "01/09/2021 00:11:07",
          "content": "<p>Hey <a href=\"https://www.kaggle.com/aqx\" target=\"_blank\">@aqx</a>. I'm far from sure, but I think the author is probably averaging the probabilities of the five folds at test time instead of selecting the model with the most confident prediction as you are suggesting. You can check this out in this <a href=\"https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta\" target=\"_blank\">notebook</a> by <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> </p>\n<p>Check out this line of code in the last cell:</p>\n<p><code>tst_preds = np.mean(tst_preds, axis=0)</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1145419,
          "author_name": "angqx95",
          "author_url": "",
          "post_date": "01/09/2021 06:07:49",
          "content": "<p>Thank you for the explanation! </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1120662": "I got the LB score of 0.901 by simply using **SnapMix** augmentation and model ensemble. \n\nDetailed results can be found as below.\n\n**Single model result (single fold, no tta, no extra training data)**\n \n  ********Resnet50  : LB 0.891  [Notebook for training ](https://www.kaggle.com/shaolihuang/training-with-snapmix)\n\n**Model ensemble ( no tta, no extra training data):**\n\n1. Resnet50 (5 folds) :  LB 0.897\n\n2. Resnet50 (5 folds) + Resnet101(5 folds):  LB 0.899\n\n3. Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds):  LB 0.901",
    "1120737": "Hi,the SanpMix augmentation is what?Thanks!",
    "1120749": "SnapMix: Semantically Proportional Mixing for Augmenting Fine-grained Data\nCode: [https://github.com/Shaoli-Huang/SnapMix](url)",
    "1120752": "Thanks!I will test it",
    "1120758": "you can find some training detail that I tried from the [link](https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/204631)",
    "1120786": "Hey @David . Thank you for sharing. I'm new in the kaggle contest. Can you tell me how marge Resnet50 (5 folds) + Resnet101(5 folds) + ResNext101(5 folds) models?  I not understand the tecniques of marge two or three models. Thanks in an advance.",
    "1120797": "I used Voting Ensembles.\nif you have n models and m images, you can put all the prediction results as an array with shape nxm.\nThen I used the following code to pick the final prediction.\n\npres = []\nfor i in range(allresults.shape[1]):\n    pre= Counter(allresults[:,i]).most_common(1)\n    pres.append(pre[0][0])",
    "1120814": "Thank you so much that's make sense. Can you please share any notebook which use this technique!  Then it help me a lot.",
    "1120931": "I'm curious, when you use different sizes of the same architecture, does the smallest resent still add value in the final ensemble (e.g. did you try the Resnet101 alone?)? I would have been tempted to do different architectures.",
    "1121318": "I would consider keeping the raw probabilities of each label, adding those, and then taking the label with the max as your final answer. Seems intuitively that it is using more information than if you just take the most common label.",
    "1121320": "Thank you for sharing. Just a quick question, how much time it took to train the whole thing and what was the submission time ?",
    "1121882": "Hey @shaolihuang , nice results, you mean that the only data augmentation that you used was `SnapMix`?",
    "1121922": "oh, I just realize you're the 1st author of the method, I was wondering what kind of awesome guys can get quickly updated about latest papers, now all makes sense.👍",
    "1121963": "Are you using keras or pytorch? 5-10 fold cross validation with 10 epochs takes a long time on GPU (perhaps more than 9 hours limit). Are you using your own GPU or are doing something else to avoid the 9 hour limit?",
    "1122138": "David thanks to u  for telling everyone  the SnapMix augmentation method. 😃",
    "1122326": "I trained models offline and upload them as a dataset.",
    "1122333": "Thx, I am trying to test whether Snapmix is effective in other datasets. Then, I saw someone mention it in this competition,  so I just did some quick experiments to see if it works on this dataset.",
    "1122337": "Concretely speaking, I used SnapMix coupled with some standard augmentation practice (including random crop and random flip).",
    "1122342": "The default transforms used in the repository has only a few basic transforms as shown [here](https://github.com/Shaoli-Huang/SnapMix/blob/8d4aa030731f835eeefeee404e0fab98a1a1d4de/datasets/tfs.py#L18), i suppose only these were used ?\n\nEDIT : I've open sourced my pipeline [here](https://www.kaggle.com/sachinprabhu/pytorch-resnet50-snapmix-train-pipeline) for people to experiment",
    "1122346": "your paper is fantastic and does make more sense, nice job!",
    "1122351": "I do not do statistics on the specific time for training all models. I found the following related information from the training log, I hope it helps.\n\nTraining a single fold of resnet50 for 40 epochs took around 4 hours.\n\nTraining a single fold of resnet101 for 40 epochs took around 6 hours.\n\n\n\nI upload all the trained models as a dataset.  The submission of ensembling all models took around two hours.",
    "1122485": "Hello! How much time did you spend on inference?",
    "1122521": "Applying SnapMix did not increase any inference time. In other words, my method has exactly the same inference time as the standard Resnet-50.\n\nwhen I evaluated the validation dataset containing 4280 images,  it took 20 seconds (using my own GPU).",
    "1122533": "shaolihuang Thank you for sharing. Have you tested Resnet101(5 folds) and ResNext101(5 folds) separately? And what's the LB score of them?",
    "1122537": "Oh,how long about inference on the whole test set when ranking ?",
    "1123013": "thank you,  @shaolihuang,  for your feedback. Are you using a packaged method such as sci-kit learn to do CV? or implemented it yourself? I am using PyTorch and am wondering if there are CV packages that support it?",
    "1124919": "what is fold?",
    "1124920": "I just know epochs, what is fold",
    "1124967": "shaolihuang were your models pretrained on ImageNet?",
    "1125069": "Yes,  pretrained model from torchvision",
    "1126828": "Hey @shaolihuang, thank you for sharing your insight. Just to confirm are you using a voting ensemble of 18 models (5 folds x 3 models) then? So at test time you are making 18 predictions per image?",
    "1126854": "Actually it should be 15 models.\nAnd yes,  make 15 predictions in total.",
    "1126866": "comment deleted",
    "1127020": "PyTorch does not have any inbuilt package. You can use the sci-kit learn stratified k fold like [here](https://github.com/svishnu88/Cassava/blob/3734fabe28d6f6dd8c6ca21b3daacfba9822816b/cassavadata.py#L67) .",
    "1127524": "https://towardsdatascience.com/complete-guide-to-pythons-cross-validation-with-examples-a9676b5cac12",
    "1128367": "The number of ***\"epochs\"*** is a hyperparameter that defines the number of times that the learning algorithm will work through the entire training dataset. One epoch means that each sample in the training dataset has had an opportunity to update the internal model parameters. Furthermore, an epoch is comprised of one or more batches (to help the training).\n\nDifferently, with ***\"fold\"***, we refer to a specific resampling of the validation data. This term is related to the process of cross-validation: a process through which you create a k (interchangeable) number of resampled validation sets to validate your model. In this way, you will have a k number of validation results instead of just one, making your results more robust. Nevertheless, this robustness comes with some contra, such as the increasing time in training. In this sense, *leave one out cross-validation* is the most demanding type of validation, where the number of folds equals the number of instances in the data set.",
    "1132262": "How long will it take? For me only a single resnet50 model will cost about half an hour..... And what is your LB by single fold resnet50 without snapMix ?",
    "1133566": "Hi David,\nCan you please refer some good sources to read about snapmix?",
    "1133690": "Thanks @Marto93 for replying there. I have a (probably dumb) question about how k-fold CV is used by kagglers. My understanding was always that people use CV to estimate their model performance, but then they may re-train using the full data available before they deploy it. The OP is posting his leaderboard performance, and not the CV performance. Do folks on here typically average the predictions from their k models? Or are they re-training on the full dataset? Thank you so much in advance if you choose to answer me!",
    "1135159": "Thanks for sharing",
    "1135381": "Hi @shaolihuang \nThanks for sharing insights.\nCan you tell me what image size do you prefer for training all models and whats your PC config?",
    "1141237": "Hi! Thank you for the introducing Snapmix method. I have one question. Your Resnet50 (5-folds) performs better then single fold Resnet50. My understanding of cross validation fold is that it is being used for a more robust estimate of the model itself, instead of improving performance.\nSo my assumption on your 5-fold Resnet model is as follows:\n1. Train 5 models with Kfolds, each model trained on 80% of the data.\n2. During inference, use these 5 models to output the predicted probabilities separately\n3. Label the class with the maximum probabilities for each test example. (soft-voting)\n\ni hope you can enlighten me with this clarification that i have, would really appreciate it! Thank you:)",
    "1144595": "Paper: https://arxiv.org/abs/2012.04846\nPytorch Github repository: https://github.com/Shaoli-Huang/SnapMix",
    "1144600": "Check out this post https://www.kaggle.com/c/cassava-leaf-disease-classification/discussion/209136",
    "1145156": "Hey @aqx. I'm far from sure, but I think the author is probably averaging the probabilities of the five folds at test time instead of selecting the model with the most confident prediction as you are suggesting. You can check this out in this [notebook](https://www.kaggle.com/khyeh0719/pytorch-efficientnet-baseline-inference-tta) by @khyeh0719 \n\nCheck out this line of code in the last cell:\n\n`tst_preds = np.mean(tst_preds, axis=0) `",
    "1145419": "Thank you for the explanation!",
    "1156179": "Very helpful information. Thanks for sharing.\n\nCan I ask what machine are you using since I saw from the other comment that you're doing all the training offline?"
  },
  "source": "meta"
}