{
  "id": 220735,
  "title": "Private 7th |  Public 71th  Solution",
  "url": "/competitions/cassava-leaf-disease-classification/writeups/private-7th-public-71th-solution",
  "author_name": "",
  "post_date": "2021-02-19T10:35:59.861685700Z",
  "votes": 17,
  "comment_count": 11,
  "views": 0,
  "content": "<blockquote>\n  <p>The experience of this competition is extraordinary for me, due to the noise data of PB, I think I am really lucky to enter the gold medal zone, because most people's PB scores are very close. I really appreciate all the discussion topics and public notebooks, and I learn a lot from it. Thank you very much！Congratulations to all the people and teams who won the medals !</p>\n</blockquote>\n<h1>Solution</h1>\n<p>two schemes were selected to submit:</p>\n<h2>1. the first submission: no secondary processing for noise label</h2>\n<h4>CV 0.90398  |  LB 0.906  |  PB 0.901</h4>\n<p><strong>this scheme is the best CV in local</strong>. It uses a lot of augmentation, including randomcrop, H/V flip, a large number of RGB transform, cutout, grid distortion, etc.,  </p>\n<p><strong>6modelx4TTA:</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>focalcos</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext50</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>focalcos</td>\n<td>384</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>bi-tempered</td>\n<td>384</td>\n</tr>\n</tbody>\n</table>\n<p>the results show that the scores of all ensemble submitted PB without secondary processing for noise label were between 0.898 and 0.902. </p>\n<p>there are two results of PB 0.902:</p>\n<p><strong>3model and 5xrandomTTA:</strong><br>\n<strong>CV ---  |  LB 0.902  |  PB 0.902</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>bi-tempered</td>\n<td>384</td>\n</tr>\n</tbody>\n</table>\n<p>randomTTA is the same as the augmentation method in training, randomly 5 times</p>\n<p><strong>5model and 5xTTA:</strong><br>\n<strong>CV ---  |  LB 0.901  |  PB 0.902</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>focalcos</td>\n<td>384</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>focalcos</td>\n<td>384</td>\n</tr>\n<tr>\n<td>regnety120</td>\n<td>focalcos</td>\n<td>512</td>\n</tr>\n</tbody>\n</table>\n<h2>2. the second submission: based on the prediction results of the first scheme, knowledge distillation is carried out, and the student model results and the teacher model results are weighted by a ratio of 4:6</h2>\n<h4>CV 0.9063  |  LB 0.902  |  PB 0.900</h4>\n<p>the augmentation and the model structure is unchanged</p>\n<p><strong>6 student model and 4xTTA</strong><br>\n<strong>CV 0.9079  |  LB 0.901  |  PB 0.898</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>CE</td>\n<td>384</td>\n</tr>\n<tr>\n<td>resnext101</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n<tr>\n<td>resnest101</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n</tbody>\n</table>\n<p>I think the student model <strong>overfitting</strong> the training set, so I tried various weighting methods with the predicted probability of the first scheme. The range of pb is between 0.898 and 0.902</p>\n<h2>3. something else (all changes to labels or delete labels are only for the training set, not the label of the validation set, keep the TRUE label in validation set)</h2>\n<p>(as a comparison of the following,  no processing has been applied to the noisy label: <br>\n<strong>efficientnet-b4 5fold 5tta:  PB 0.894</strong>, the following attempts all use this benchmark)</p>\n<ul>\n<li>images with prediction probability greater than 0.9 and incorrect predictions are deleted. <strong>PB 0.895</strong></li>\n<li>the labels of images with prediction probability greater than 0.9 and incorrect predictions are changed. <strong>PB 0.896</strong></li>\n<li>simply aug + snapmix, and gradually increase the probability of snapmix by epoch. <strong>PB 0.896</strong></li>\n<li>2019+2020 Dataset, since the 2019 data set contains multiple sizes, and any resize/crop form will affect the final CV, I finally did not use the 2019 data ( The data for 2019 was deleted in the verification set) <strong>PB 0.895</strong></li>\n<li>focalcos loss will increase CV, but some models will reduce the prediction probability of true label. </li>\n<li>in my submit, when using random TTA, the score of PB and LB are close.</li>\n<li>RegNet PB is higher than LB, but RegNet LB is too low, RegNet is not selected for submit.</li>\n</ul>\n<p>I am very glad to get my first solo gold, and I really look forward to 1st solution</p>",
  "messages": [
    {
      "id": "1210304",
      "postDate": "02/19/2021 10:35:59",
      "content": "<blockquote>\n  <p>The experience of this competition is extraordinary for me, due to the noise data of PB, I think I am really lucky to enter the gold medal zone, because most people's PB scores are very close. I really appreciate all the discussion topics and public notebooks, and I learn a lot from it. Thank you very much！Congratulations to all the people and teams who won the medals !</p>\n</blockquote>\n<h1>Solution</h1>\n<p>two schemes were selected to submit:</p>\n<h2>1. the first submission: no secondary processing for noise label</h2>\n<h4>CV 0.90398  |  LB 0.906  |  PB 0.901</h4>\n<p><strong>this scheme is the best CV in local</strong>. It uses a lot of augmentation, including randomcrop, H/V flip, a large number of RGB transform, cutout, grid distortion, etc.,  </p>\n<p><strong>6modelx4TTA:</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>focalcos</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext50</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>focalcos</td>\n<td>384</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>bi-tempered</td>\n<td>384</td>\n</tr>\n</tbody>\n</table>\n<p>the results show that the scores of all ensemble submitted PB without secondary processing for noise label were between 0.898 and 0.902. </p>\n<p>there are two results of PB 0.902:</p>\n<p><strong>3model and 5xrandomTTA:</strong><br>\n<strong>CV ---  |  LB 0.902  |  PB 0.902</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>bi-tempered</td>\n<td>384</td>\n</tr>\n</tbody>\n</table>\n<p>randomTTA is the same as the augmentation method in training, randomly 5 times</p>\n<p><strong>5model and 5xTTA:</strong><br>\n<strong>CV ---  |  LB 0.901  |  PB 0.902</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>focalcos</td>\n<td>384</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>bi-tempered</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>focalcos</td>\n<td>384</td>\n</tr>\n<tr>\n<td>regnety120</td>\n<td>focalcos</td>\n<td>512</td>\n</tr>\n</tbody>\n</table>\n<h2>2. the second submission: based on the prediction results of the first scheme, knowledge distillation is carried out, and the student model results and the teacher model results are weighted by a ratio of 4:6</h2>\n<h4>CV 0.9063  |  LB 0.902  |  PB 0.900</h4>\n<p>the augmentation and the model structure is unchanged</p>\n<p><strong>6 student model and 4xTTA</strong><br>\n<strong>CV 0.9079  |  LB 0.901  |  PB 0.898</strong></p>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>loss</th>\n<th>size</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnet_b4_ns</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n<tr>\n<td>seresnext101</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n<tr>\n<td>VIT_base16</td>\n<td>CE</td>\n<td>384</td>\n</tr>\n<tr>\n<td>resnext101</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n<tr>\n<td>resnest101</td>\n<td>CE</td>\n<td>512</td>\n</tr>\n</tbody>\n</table>\n<p>I think the student model <strong>overfitting</strong> the training set, so I tried various weighting methods with the predicted probability of the first scheme. The range of pb is between 0.898 and 0.902</p>\n<h2>3. something else (all changes to labels or delete labels are only for the training set, not the label of the validation set, keep the TRUE label in validation set)</h2>\n<p>(as a comparison of the following,  no processing has been applied to the noisy label: <br>\n<strong>efficientnet-b4 5fold 5tta:  PB 0.894</strong>, the following attempts all use this benchmark)</p>\n<ul>\n<li>images with prediction probability greater than 0.9 and incorrect predictions are deleted. <strong>PB 0.895</strong></li>\n<li>the labels of images with prediction probability greater than 0.9 and incorrect predictions are changed. <strong>PB 0.896</strong></li>\n<li>simply aug + snapmix, and gradually increase the probability of snapmix by epoch. <strong>PB 0.896</strong></li>\n<li>2019+2020 Dataset, since the 2019 data set contains multiple sizes, and any resize/crop form will affect the final CV, I finally did not use the 2019 data ( The data for 2019 was deleted in the verification set) <strong>PB 0.895</strong></li>\n<li>focalcos loss will increase CV, but some models will reduce the prediction probability of true label. </li>\n<li>in my submit, when using random TTA, the score of PB and LB are close.</li>\n<li>RegNet PB is higher than LB, but RegNet LB is too low, RegNet is not selected for submit.</li>\n</ul>\n<p>I am very glad to get my first solo gold, and I really look forward to 1st solution</p>",
      "rawMarkdown": "> The experience of this competition is extraordinary for me, due to the noise data of PB, I think I am really lucky to enter the gold medal zone, because most people's PB scores are very close. I really appreciate all the discussion topics and public notebooks, and I learn a lot from it. Thank you very much！Congratulations to all the people and teams who won the medals !\n\n# Solution \ntwo schemes were selected to submit:\n## 1. the first submission: no secondary processing for noise label    \n#### CV 0.90398  |  LB 0.906  |  PB 0.901\n**this scheme is the best CV in local**. It uses a lot of augmentation, including randomcrop, H/V flip, a large number of RGB transform, cutout, grid distortion, etc.,  \n\n**6modelx4TTA:**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | bi-tempered | 512\n| seresnext101 | bi-tempered | 512\n| seresnext101 | focalcos | 512\n| seresnext50 | bi-tempered| 512\n| VIT_base16| focalcos | 384\n| VIT_base16| bi-tempered| 384\n\nthe results show that the scores of all ensemble submitted PB without secondary processing for noise label were between 0.898 and 0.902. \n\nthere are two results of PB 0.902:\n\n**3model and 5xrandomTTA:**\n**CV ---  |  LB 0.902  |  PB 0.902**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | bi-tempered | 512\n| seresnext101 | bi-tempered | 512\n| VIT_base16| bi-tempered| 384\n\nrandomTTA is the same as the augmentation method in training, randomly 5 times\n\n**5model and 5xTTA:**\n**CV ---  |  LB 0.901  |  PB 0.902**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | bi-tempered | 512\n| efficientnet_b4_ns | focalcos | 384\n| seresnext101 | bi-tempered | 512\n| VIT_base16| focalcos | 384\n| regnety120| focalcos | 512\n\n## 2. the second submission: based on the prediction results of the first scheme, knowledge distillation is carried out, and the student model results and the teacher model results are weighted by a ratio of 4:6\n#### CV 0.9063  |  LB 0.902  |  PB 0.900\nthe augmentation and the model structure is unchanged\n\n**6 student model and 4xTTA**\n**CV 0.9079  |  LB 0.901  |  PB 0.898**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | CE | 512\n| seresnext101  | CE | 512\n| VIT_base16 | CE  | 384\n| resnext101 | CE | 512\n| resnest101 | CE | 512\n\nI think the student model **overfitting** the training set, so I tried various weighting methods with the predicted probability of the first scheme. The range of pb is between 0.898 and 0.902\n\n## 3. something else (all changes to labels or delete labels are only for the training set, not the label of the validation set, keep the TRUE label in validation set)\n\n(as a comparison of the following,  no processing has been applied to the noisy label: \n**efficientnet-b4 5fold 5tta:  PB 0.894**, the following attempts all use this benchmark)\n- images with prediction probability greater than 0.9 and incorrect predictions are deleted. **PB 0.895**\n- the labels of images with prediction probability greater than 0.9 and incorrect predictions are changed. **PB 0.896**\n- simply aug + snapmix, and gradually increase the probability of snapmix by epoch. **PB 0.896**\n- 2019+2020 Dataset, since the 2019 data set contains multiple sizes, and any resize/crop form will affect the final CV, I finally did not use the 2019 data ( The data for 2019 was deleted in the verification set) **PB 0.895**\n- focalcos loss will increase CV, but some models will reduce the prediction probability of true label. \n- in my submit, when using random TTA, the score of PB and LB are close.\n- RegNet PB is higher than LB, but RegNet LB is too low, RegNet is not selected for submit.\n\nI am very glad to get my first solo gold, and I really look forward to 1st solution",
      "votes": null
    },
    {
      "id": "1210364",
      "postDate": "02/19/2021 11:31:15",
      "content": "<p>Congrats on your solo gold medal and great work!<br>\nI like the experimental approach and really good job.</p>\n<p>I have one question.<br>\nFor distillation, did you use hard OOF ensemble lables or soft type that load all teacher models?</p>\n<p><a href=\"https://www.kaggle.com/dbwlalagaga\" target=\"_blank\">@dbwlalagaga</a> </p>",
      "rawMarkdown": "Congrats on your solo gold medal and great work!\nI like the experimental approach and really good job.\n\nI have one question.\nFor distillation, did you use hard OOF ensemble lables or soft type that load all teacher models?\n\n@dbwlalagaga",
      "votes": null
    },
    {
      "id": "1210390",
      "postDate": "02/19/2021 11:51:44",
      "content": "<p>Thank you!  I use 6 models avg softmax predictions * 0.7 and don't change the original label, like this:</p>\n<pre><code>original label:  [0, 0, 0, 1]\n6 models avg softmax pred: [0.6, 0.3, 0, 0.1]\nnew label: [0.42, 0.21, 0, 1]\n</code></pre>",
      "rawMarkdown": "Thank you!  I use 6 models avg softmax predictions * 0.7 and don't change the original label, like this:\n```\noriginal label:  [0, 0, 0, 1]\n6 models avg softmax pred: [0.6, 0.3, 0, 0.1]\nnew label: [0.42, 0.21, 0, 1]\n```",
      "votes": null
    },
    {
      "id": "1210400",
      "postDate": "02/19/2021 12:04:09",
      "content": "<p>Good work!</p>\n<p>I have tried distillation for several types. <br>\nIn my case, there are models that have achieved 0.899 in private lb as a single model, but I didn't  ensemble my final submissions. </p>\n<p>Congratulations again. :)</p>",
      "rawMarkdown": "Good work!\n\nI have tried distillation for several types. \nIn my case, there are models that have achieved 0.899 in private lb as a single model, but I didn't  ensemble my final submissions. \n\nCongratulations again. :)",
      "votes": null
    },
    {
      "id": "1210527",
      "postDate": "02/19/2021 13:52:44",
      "content": "<p>Thank you again ! :)</p>",
      "rawMarkdown": "Thank you again ! :)",
      "votes": null
    },
    {
      "id": "1210607",
      "postDate": "02/19/2021 14:52:33",
      "content": "<p>Congratulations! </p>\n<p>May I ask each model for final submission is only using 1 fold or 5 fold (5 files for each model) for ensemble?</p>",
      "rawMarkdown": "Congratulations! \n\nMay I ask each model for final submission is only using 1 fold or 5 fold (5 files for each model) for ensemble?",
      "votes": null
    },
    {
      "id": "1211102",
      "postDate": "02/20/2021 00:23:18",
      "content": "<p>Hi Welkin, The highest 1-fold auc of the same model may reach 0.905, while the lowest 0.891. 0.905 can only show a better effect on the 1 fold validation set. In more test sets, a model with 1fold cv=0.891 may have better results, so it is recommended to use an average of 5fold for ensemble</p>",
      "rawMarkdown": "Hi Welkin, The highest 1-fold auc of the same model may reach 0.905, while the lowest 0.891. 0.905 can only show a better effect on the 1 fold validation set. In more test sets, a model with 1fold cv=0.891 may have better results, so it is recommended to use an average of 5fold for ensemble",
      "votes": null
    },
    {
      "id": "1211231",
      "postDate": "02/20/2021 04:44:05",
      "content": "<p>So, in <strong><code>5model and 5xTTA</code></strong> submission, in fact there are 5 fold x 5 models = 25 models, and each model uses 5 TTA, which means you have run 25 x (1+5) = 150 times test_set and ensemble 150 prediction results by np.mean(axis=-1), is that right?</p>",
      "rawMarkdown": "So, in **`5model and 5xTTA`** submission, in fact there are 5 fold x 5 models = 25 models, and each model uses 5 TTA, which means you have run 25 x (1+5) = 150 times test_set and ensemble 150 prediction results by np.mean(axis=-1), is that right?",
      "votes": null
    },
    {
      "id": "1211238",
      "postDate": "02/20/2021 04:56:26",
      "content": "<p>That's very good. It has a lot of inspiration for me!</p>",
      "rawMarkdown": "That's very good. It has a lot of inspiration for me!",
      "votes": null
    },
    {
      "id": "1211269",
      "postDate": "02/20/2021 05:31:11",
      "content": "<p>Yes~ <br>\nThe time limit for this competition is 9 hours. When the average size of single model is about 150M and bs = 64, 4model + 7tta takes about 8-9 hours</p>",
      "rawMarkdown": "Yes~ \nThe time limit for this competition is 9 hours. When the average size of single model is about 150M and bs = 64, 4model + 7tta takes about 8-9 hours",
      "votes": null
    },
    {
      "id": "1211272",
      "postDate": "02/20/2021 05:39:12",
      "content": "<p>Thanks for your reply, learned a lot : )</p>",
      "rawMarkdown": "Thanks for your reply, learned a lot : )",
      "votes": null
    },
    {
      "id": "1211572",
      "postDate": "02/20/2021 10:35:37",
      "content": "<p>Congratz ! Your solution looks really robust which explain your final ranking :)</p>",
      "rawMarkdown": "Congratz ! Your solution looks really robust which explain your final ranking :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1210364,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "02/19/2021 11:31:15",
      "content": "<p>Congrats on your solo gold medal and great work!<br>\nI like the experimental approach and really good job.</p>\n<p>I have one question.<br>\nFor distillation, did you use hard OOF ensemble lables or soft type that load all teacher models?</p>\n<p><a href=\"https://www.kaggle.com/dbwlalagaga\" target=\"_blank\">@dbwlalagaga</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 1210390,
          "author_name": "dbwlalagaga",
          "author_url": "",
          "post_date": "02/19/2021 11:51:44",
          "content": "<p>Thank you!  I use 6 models avg softmax predictions * 0.7 and don't change the original label, like this:</p>\n<pre><code>original label:  [0, 0, 0, 1]\n6 models avg softmax pred: [0.6, 0.3, 0, 0.1]\nnew label: [0.42, 0.21, 0, 1]\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210400,
          "author_name": "piantic",
          "author_url": "",
          "post_date": "02/19/2021 12:04:09",
          "content": "<p>Good work!</p>\n<p>I have tried distillation for several types. <br>\nIn my case, there are models that have achieved 0.899 in private lb as a single model, but I didn't  ensemble my final submissions. </p>\n<p>Congratulations again. :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1210527,
          "author_name": "dbwlalagaga",
          "author_url": "",
          "post_date": "02/19/2021 13:52:44",
          "content": "<p>Thank you again ! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1210607,
      "author_name": "welkinfeng",
      "author_url": "",
      "post_date": "02/19/2021 14:52:33",
      "content": "<p>Congratulations! </p>\n<p>May I ask each model for final submission is only using 1 fold or 5 fold (5 files for each model) for ensemble?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1211102,
          "author_name": "dbwlalagaga",
          "author_url": "",
          "post_date": "02/20/2021 00:23:18",
          "content": "<p>Hi Welkin, The highest 1-fold auc of the same model may reach 0.905, while the lowest 0.891. 0.905 can only show a better effect on the 1 fold validation set. In more test sets, a model with 1fold cv=0.891 may have better results, so it is recommended to use an average of 5fold for ensemble</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211231,
          "author_name": "welkinfeng",
          "author_url": "",
          "post_date": "02/20/2021 04:44:05",
          "content": "<p>So, in <strong><code>5model and 5xTTA</code></strong> submission, in fact there are 5 fold x 5 models = 25 models, and each model uses 5 TTA, which means you have run 25 x (1+5) = 150 times test_set and ensemble 150 prediction results by np.mean(axis=-1), is that right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211269,
          "author_name": "dbwlalagaga",
          "author_url": "",
          "post_date": "02/20/2021 05:31:11",
          "content": "<p>Yes~ <br>\nThe time limit for this competition is 9 hours. When the average size of single model is about 150M and bs = 64, 4model + 7tta takes about 8-9 hours</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1211272,
          "author_name": "welkinfeng",
          "author_url": "",
          "post_date": "02/20/2021 05:39:12",
          "content": "<p>Thanks for your reply, learned a lot : )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1211238,
      "author_name": "fengcoco",
      "author_url": "",
      "post_date": "02/20/2021 04:56:26",
      "content": "<p>That's very good. It has a lot of inspiration for me!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1211572,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "02/20/2021 10:35:37",
      "content": "<p>Congratz ! Your solution looks really robust which explain your final ranking :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1210304": "> The experience of this competition is extraordinary for me, due to the noise data of PB, I think I am really lucky to enter the gold medal zone, because most people's PB scores are very close. I really appreciate all the discussion topics and public notebooks, and I learn a lot from it. Thank you very much！Congratulations to all the people and teams who won the medals !\n\n# Solution \ntwo schemes were selected to submit:\n## 1. the first submission: no secondary processing for noise label    \n#### CV 0.90398  |  LB 0.906  |  PB 0.901\n**this scheme is the best CV in local**. It uses a lot of augmentation, including randomcrop, H/V flip, a large number of RGB transform, cutout, grid distortion, etc.,  \n\n**6modelx4TTA:**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | bi-tempered | 512\n| seresnext101 | bi-tempered | 512\n| seresnext101 | focalcos | 512\n| seresnext50 | bi-tempered| 512\n| VIT_base16| focalcos | 384\n| VIT_base16| bi-tempered| 384\n\nthe results show that the scores of all ensemble submitted PB without secondary processing for noise label were between 0.898 and 0.902. \n\nthere are two results of PB 0.902:\n\n**3model and 5xrandomTTA:**\n**CV ---  |  LB 0.902  |  PB 0.902**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | bi-tempered | 512\n| seresnext101 | bi-tempered | 512\n| VIT_base16| bi-tempered| 384\n\nrandomTTA is the same as the augmentation method in training, randomly 5 times\n\n**5model and 5xTTA:**\n**CV ---  |  LB 0.901  |  PB 0.902**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | bi-tempered | 512\n| efficientnet_b4_ns | focalcos | 384\n| seresnext101 | bi-tempered | 512\n| VIT_base16| focalcos | 384\n| regnety120| focalcos | 512\n\n## 2. the second submission: based on the prediction results of the first scheme, knowledge distillation is carried out, and the student model results and the teacher model results are weighted by a ratio of 4:6\n#### CV 0.9063  |  LB 0.902  |  PB 0.900\nthe augmentation and the model structure is unchanged\n\n**6 student model and 4xTTA**\n**CV 0.9079  |  LB 0.901  |  PB 0.898**\n| model | loss | size\n| --- | --- |\n| efficientnet_b4_ns | CE | 512\n| seresnext101  | CE | 512\n| VIT_base16 | CE  | 384\n| resnext101 | CE | 512\n| resnest101 | CE | 512\n\nI think the student model **overfitting** the training set, so I tried various weighting methods with the predicted probability of the first scheme. The range of pb is between 0.898 and 0.902\n\n## 3. something else (all changes to labels or delete labels are only for the training set, not the label of the validation set, keep the TRUE label in validation set)\n\n(as a comparison of the following,  no processing has been applied to the noisy label: \n**efficientnet-b4 5fold 5tta:  PB 0.894**, the following attempts all use this benchmark)\n- images with prediction probability greater than 0.9 and incorrect predictions are deleted. **PB 0.895**\n- the labels of images with prediction probability greater than 0.9 and incorrect predictions are changed. **PB 0.896**\n- simply aug + snapmix, and gradually increase the probability of snapmix by epoch. **PB 0.896**\n- 2019+2020 Dataset, since the 2019 data set contains multiple sizes, and any resize/crop form will affect the final CV, I finally did not use the 2019 data ( The data for 2019 was deleted in the verification set) **PB 0.895**\n- focalcos loss will increase CV, but some models will reduce the prediction probability of true label. \n- in my submit, when using random TTA, the score of PB and LB are close.\n- RegNet PB is higher than LB, but RegNet LB is too low, RegNet is not selected for submit.\n\nI am very glad to get my first solo gold, and I really look forward to 1st solution",
    "1210364": "Congrats on your solo gold medal and great work!\nI like the experimental approach and really good job.\n\nI have one question.\nFor distillation, did you use hard OOF ensemble lables or soft type that load all teacher models?\n\n@dbwlalagaga",
    "1210390": "Thank you!  I use 6 models avg softmax predictions * 0.7 and don't change the original label, like this:\n```\noriginal label:  [0, 0, 0, 1]\n6 models avg softmax pred: [0.6, 0.3, 0, 0.1]\nnew label: [0.42, 0.21, 0, 1]\n```",
    "1210400": "Good work!\n\nI have tried distillation for several types. \nIn my case, there are models that have achieved 0.899 in private lb as a single model, but I didn't  ensemble my final submissions. \n\nCongratulations again. :)",
    "1210527": "Thank you again ! :)",
    "1210607": "Congratulations! \n\nMay I ask each model for final submission is only using 1 fold or 5 fold (5 files for each model) for ensemble?",
    "1211102": "Hi Welkin, The highest 1-fold auc of the same model may reach 0.905, while the lowest 0.891. 0.905 can only show a better effect on the 1 fold validation set. In more test sets, a model with 1fold cv=0.891 may have better results, so it is recommended to use an average of 5fold for ensemble",
    "1211231": "So, in **`5model and 5xTTA`** submission, in fact there are 5 fold x 5 models = 25 models, and each model uses 5 TTA, which means you have run 25 x (1+5) = 150 times test_set and ensemble 150 prediction results by np.mean(axis=-1), is that right?",
    "1211238": "That's very good. It has a lot of inspiration for me!",
    "1211269": "Yes~ \nThe time limit for this competition is 9 hours. When the average size of single model is about 150M and bs = 64, 4model + 7tta takes about 8-9 hours",
    "1211272": "Thanks for your reply, learned a lot : )",
    "1211572": "Congratz ! Your solution looks really robust which explain your final ranking :)"
  },
  "source": "meta"
}