{
  "id": 264243,
  "title": "4th place solution",
  "url": "/competitions/siim-covid19-detection/writeups/rtx-4090-4th-place-solution",
  "author_name": "",
  "post_date": "2021-08-15T01:00:54.460Z",
  "votes": 50,
  "comment_count": 19,
  "views": 0,
  "content": "<p>First of all, I would like to thank Kaggle, SIIM, FISABIO &amp; RSNA for this amazing competition and also to this amazing community for constantly sharing. Congratulations to all the winners, medalists, and amazing people of this community.</p>\n<p>Thanks to all the members of our team <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/zaber666\" target=\"_blank\">@zaber666</a> <a href=\"https://www.kaggle.com/nexh98\" target=\"_blank\">@nexh98</a> <a href=\"https://www.kaggle.com/artemenon\" target=\"_blank\">@artemenon</a>. This would not be possible without you. Special thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for teaming up with us. We've learned a lot from you. No wonder you're one of the best <strong>kagglers</strong>.</p>\n<h2>Data Processing:</h2>\n<ul>\n<li><strong>Data-Cleaning</strong>: We removed duplicates manually for both classification and detection. Also for detection, instead of taking one image per patient we only removed unannotated images. For this reason, our <strong>CV</strong> was low comparing most people but we had a good <strong>lb</strong> correlation.</li>\n<li><strong>Cross-Validation</strong>: StratifiedGroupKFold.</li>\n</ul>\n<h2>Study-Level:</h2>\n<ul>\n<li><strong>Pretraining</strong>: We pretrained our models on chexpert. So, all our models use it except(Chris's aux-loss model).</li>\n<li><strong>Models</strong>: EfficientnetB6 &amp; EfficientnetB7. (Others models didn't do well for us not even <code>effnetv2</code>, <code>resnet200d</code>, <code>nfnet</code>, <code>vit</code>).</li>\n<li><strong>Loss</strong>: CategoricalCrossEntropy.</li>\n<li><strong>Label-Smoothing</strong>: 0.01</li>\n<li><strong>Augmentation</strong>: HorizontalFlip, VerticalFlip, RandomRotation, CoarseDropout, RandomShift, RandomZoom, Random(Brightness, Contrast), Cutmix-Mixup.</li>\n<li><strong>Scheduler</strong>: WarmpupExponentialDecay.</li>\n<li><strong>Image-Size</strong>: 512, 640, 768.</li>\n<li><strong>Pseudo-Labeling</strong>: BIMCV + RICORD + RSNA.</li>\n<li><strong>Knowledge-Distillation(KD)</strong>: We used around 5 of our best models to generate <strong>soft-labels</strong> for it and one of our submissions uses 3 KD models.</li>\n<li><strong>Aux-Loss</strong>: One of our submissions uses <strong>aux-loss</strong> model by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> which uses FPN and EfficientnetB4 as backbone. You can check out the post <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263676\" target=\"_blank\">here</a> for more details. We also used <strong>aux-loss</strong> from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's discussion and used it for generating <strong>soft-labels</strong> for KD model. </li>\n<li><strong>Post-Processing:</strong> We used <strong>geometric-mean</strong> to re-rank our confidence score(inspired from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s solution of VinBigData <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">here</a>). Our <strong>4cls</strong> models were very dominating that this had very little impact on the score. So finally we decided not to use it even though it did improve the <strong>CV</strong> a little bit.<br>\n<img src=\"https://i.ibb.co/WtPVw7x/eq1.png\" alt=\"eq1\"></li>\n</ul>\n<h2>2cls Model:</h2>\n<p>We used only one <strong>2cls</strong> model. Including more didn't have that much impact.</p>\n<ul>\n<li><strong>Model</strong>: EfficientNetb7</li>\n<li><strong>Image-Size</strong>: 640</li>\n<li><strong>External Data</strong>: We used <strong>RICORD</strong> dataset and their labels posted by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> from <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240187\" target=\"_blank\">here</a>. We simply took <strong>max-voting</strong> to get the labels from different annotators.</li>\n</ul>\n<h2>Image-Level:</h2>\n<ul>\n<li><strong>Pretraining</strong>: We pretrained all backbones of detection models.</li>\n<li><strong>External Data</strong>: RSNA - <code>opacity</code> only (6k+).</li>\n<li><strong>Models</strong>: yolov5x-transformer(thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for posting it), yolov5x6, yolov3-spp.</li>\n<li><strong>Augmentation</strong>: HorizontalFlip, VerticalFlip, Random(Brightness, Contrast), Mosaic-Mixup.</li>\n<li><strong>Scheduler</strong>: WarmpupCosineDecay.</li>\n<li><strong>Image-Size</strong>: 512. (Increasing image size worsen our result cuz our backbones were pretrained on <strong>512</strong> image size. We didn't have the time to pretrain on large image-size)</li>\n<li><strong>Ensemble</strong>: We used <strong>WBF</strong> to merge boxes from different models. (thanks for <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for making it work).</li>\n<li><strong>Post-Processing:</strong> We used <strong>geometric-mean</strong> to re-rank our confidence score(inspired from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s solution of VinBigData <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">here</a>). We used both <strong>4cls</strong> and <strong>2cls</strong> models' prediction here.<br>\n<img src=\"https://i.ibb.co/jR3RXBN/eq2.png\" alt=\"eq2\"></li>\n<li><strong>BBox-Filter</strong>: We filtered out abnormal boxes based on their features. This has little effect on the both <strong>CV</strong> and <strong>LB</strong> still we kept it.</li>\n</ul>\n<h2>Final Submissions:</h2>\n<p>We tried to keep our two final submissions as different as possible. So two of our submissions use different classification models. When we looked back we noticed our best <strong>CV</strong> submission doesn't use any pseudo so we tried to incorporate them in our best <strong>LB</strong> submission. But as we had very little time we went for full data training with <strong>BIMCV</strong> + <strong>RICORD</strong> + <strong>RSNA</strong> pseudo. But we didn't want to take too much risk as we had very few submission left so ensembled full data models with our <strong>LB 0.652</strong> submission. For both submissions, 2cls model &amp; detection models are kept the same.</p>\n<ul>\n<li>Best <strong>CV</strong>:<ul>\n<li>b6 | 512 | KD</li>\n<li>b6 | 640 | KD</li>\n<li>b7 | 512 | KD</li>\n<li>FPN | b4 | 512 | aux_loss</li></ul></li>\n<li>Best <strong>LB</strong>:<ul>\n<li>(b6, b7) x (512, 640, 768) - 6 models | BIMCV+RICORD+RSNA pseudo | Full Data</li>\n<li>b6 | 512 | KD</li>\n<li>b6 | 512 | BIMCV+RICORD pseudo</li>\n<li>b6 | 512</li></ul></li>\n</ul>\n<h2>Result:</h2>\n<ul>\n<li><strong>CV</strong>:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>none</th>\n<th>opacity</th>\n<th>study</th>\n<th>final</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.823</td>\n<td>0.562</td>\n<td>0.607</td>\n<td>0.635</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>LB</strong>:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>public/private</th>\n<th>none</th>\n<th>opacity</th>\n<th>study</th>\n<th>final</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>public</td>\n<td>0.810</td>\n<td>0.594</td>\n<td>0.628</td>\n<td>0.654</td>\n</tr>\n<tr>\n<td>private</td>\n<td>0.822</td>\n<td>0.594</td>\n<td>0.592</td>\n<td>0.631</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Inference Code:</strong> <a href=\"https://www.kaggle.com/awsaf49/siim-covid19-detection-final-infer\" target=\"_blank\">https://www.kaggle.com/awsaf49/siim-covid19-detection-final-infer</a></p>\n<h2>Thank you for reading such a long post :)</h2>",
  "messages": [
    {
      "id": "1466468",
      "postDate": "08/11/2021 13:32:29",
      "content": "<p>First of all, I would like to thank Kaggle, SIIM, FISABIO &amp; RSNA for this amazing competition and also to this amazing community for constantly sharing. Congratulations to all the winners, medalists, and amazing people of this community.</p>\n<p>Thanks to all the members of our team <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> <a href=\"https://www.kaggle.com/zaber666\" target=\"_blank\">@zaber666</a> <a href=\"https://www.kaggle.com/nexh98\" target=\"_blank\">@nexh98</a> <a href=\"https://www.kaggle.com/artemenon\" target=\"_blank\">@artemenon</a>. This would not be possible without you. Special thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for teaming up with us. We've learned a lot from you. No wonder you're one of the best <strong>kagglers</strong>.</p>\n<h2>Data Processing:</h2>\n<ul>\n<li><strong>Data-Cleaning</strong>: We removed duplicates manually for both classification and detection. Also for detection, instead of taking one image per patient we only removed unannotated images. For this reason, our <strong>CV</strong> was low comparing most people but we had a good <strong>lb</strong> correlation.</li>\n<li><strong>Cross-Validation</strong>: StratifiedGroupKFold.</li>\n</ul>\n<h2>Study-Level:</h2>\n<ul>\n<li><strong>Pretraining</strong>: We pretrained our models on chexpert. So, all our models use it except(Chris's aux-loss model).</li>\n<li><strong>Models</strong>: EfficientnetB6 &amp; EfficientnetB7. (Others models didn't do well for us not even <code>effnetv2</code>, <code>resnet200d</code>, <code>nfnet</code>, <code>vit</code>).</li>\n<li><strong>Loss</strong>: CategoricalCrossEntropy.</li>\n<li><strong>Label-Smoothing</strong>: 0.01</li>\n<li><strong>Augmentation</strong>: HorizontalFlip, VerticalFlip, RandomRotation, CoarseDropout, RandomShift, RandomZoom, Random(Brightness, Contrast), Cutmix-Mixup.</li>\n<li><strong>Scheduler</strong>: WarmpupExponentialDecay.</li>\n<li><strong>Image-Size</strong>: 512, 640, 768.</li>\n<li><strong>Pseudo-Labeling</strong>: BIMCV + RICORD + RSNA.</li>\n<li><strong>Knowledge-Distillation(KD)</strong>: We used around 5 of our best models to generate <strong>soft-labels</strong> for it and one of our submissions uses 3 KD models.</li>\n<li><strong>Aux-Loss</strong>: One of our submissions uses <strong>aux-loss</strong> model by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> which uses FPN and EfficientnetB4 as backbone. You can check out the post <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/263676\" target=\"_blank\">here</a> for more details. We also used <strong>aux-loss</strong> from <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> 's discussion and used it for generating <strong>soft-labels</strong> for KD model. </li>\n<li><strong>Post-Processing:</strong> We used <strong>geometric-mean</strong> to re-rank our confidence score(inspired from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s solution of VinBigData <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">here</a>). Our <strong>4cls</strong> models were very dominating that this had very little impact on the score. So finally we decided not to use it even though it did improve the <strong>CV</strong> a little bit.<br>\n<img src=\"https://i.ibb.co/WtPVw7x/eq1.png\" alt=\"eq1\"></li>\n</ul>\n<h2>2cls Model:</h2>\n<p>We used only one <strong>2cls</strong> model. Including more didn't have that much impact.</p>\n<ul>\n<li><strong>Model</strong>: EfficientNetb7</li>\n<li><strong>Image-Size</strong>: 640</li>\n<li><strong>External Data</strong>: We used <strong>RICORD</strong> dataset and their labels posted by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> from <a href=\"https://www.kaggle.com/c/siim-covid19-detection/discussion/240187\" target=\"_blank\">here</a>. We simply took <strong>max-voting</strong> to get the labels from different annotators.</li>\n</ul>\n<h2>Image-Level:</h2>\n<ul>\n<li><strong>Pretraining</strong>: We pretrained all backbones of detection models.</li>\n<li><strong>External Data</strong>: RSNA - <code>opacity</code> only (6k+).</li>\n<li><strong>Models</strong>: yolov5x-transformer(thanks to <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for posting it), yolov5x6, yolov3-spp.</li>\n<li><strong>Augmentation</strong>: HorizontalFlip, VerticalFlip, Random(Brightness, Contrast), Mosaic-Mixup.</li>\n<li><strong>Scheduler</strong>: WarmpupCosineDecay.</li>\n<li><strong>Image-Size</strong>: 512. (Increasing image size worsen our result cuz our backbones were pretrained on <strong>512</strong> image size. We didn't have the time to pretrain on large image-size)</li>\n<li><strong>Ensemble</strong>: We used <strong>WBF</strong> to merge boxes from different models. (thanks for <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for making it work).</li>\n<li><strong>Post-Processing:</strong> We used <strong>geometric-mean</strong> to re-rank our confidence score(inspired from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>'s solution of VinBigData <a href=\"https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\" target=\"_blank\">here</a>). We used both <strong>4cls</strong> and <strong>2cls</strong> models' prediction here.<br>\n<img src=\"https://i.ibb.co/jR3RXBN/eq2.png\" alt=\"eq2\"></li>\n<li><strong>BBox-Filter</strong>: We filtered out abnormal boxes based on their features. This has little effect on the both <strong>CV</strong> and <strong>LB</strong> still we kept it.</li>\n</ul>\n<h2>Final Submissions:</h2>\n<p>We tried to keep our two final submissions as different as possible. So two of our submissions use different classification models. When we looked back we noticed our best <strong>CV</strong> submission doesn't use any pseudo so we tried to incorporate them in our best <strong>LB</strong> submission. But as we had very little time we went for full data training with <strong>BIMCV</strong> + <strong>RICORD</strong> + <strong>RSNA</strong> pseudo. But we didn't want to take too much risk as we had very few submission left so ensembled full data models with our <strong>LB 0.652</strong> submission. For both submissions, 2cls model &amp; detection models are kept the same.</p>\n<ul>\n<li>Best <strong>CV</strong>:<ul>\n<li>b6 | 512 | KD</li>\n<li>b6 | 640 | KD</li>\n<li>b7 | 512 | KD</li>\n<li>FPN | b4 | 512 | aux_loss</li></ul></li>\n<li>Best <strong>LB</strong>:<ul>\n<li>(b6, b7) x (512, 640, 768) - 6 models | BIMCV+RICORD+RSNA pseudo | Full Data</li>\n<li>b6 | 512 | KD</li>\n<li>b6 | 512 | BIMCV+RICORD pseudo</li>\n<li>b6 | 512</li></ul></li>\n</ul>\n<h2>Result:</h2>\n<ul>\n<li><strong>CV</strong>:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>none</th>\n<th>opacity</th>\n<th>study</th>\n<th>final</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.823</td>\n<td>0.562</td>\n<td>0.607</td>\n<td>0.635</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li><strong>LB</strong>:</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>public/private</th>\n<th>none</th>\n<th>opacity</th>\n<th>study</th>\n<th>final</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>public</td>\n<td>0.810</td>\n<td>0.594</td>\n<td>0.628</td>\n<td>0.654</td>\n</tr>\n<tr>\n<td>private</td>\n<td>0.822</td>\n<td>0.594</td>\n<td>0.592</td>\n<td>0.631</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Inference Code:</strong> <a href=\"https://www.kaggle.com/awsaf49/siim-covid19-detection-final-infer\" target=\"_blank\">https://www.kaggle.com/awsaf49/siim-covid19-detection-final-infer</a></p>\n<h2>Thank you for reading such a long post :)</h2>",
      "rawMarkdown": "First of all, I would like to thank Kaggle, SIIM, FISABIO & RSNA for this amazing competition and also to this amazing community for constantly sharing. Congratulations to all the winners, medalists, and amazing people of this community.\n\nThanks to all the members of our team @cdeotte @zaber666 @nexh98 @artemenon. This would not be possible without you. Special thanks to @cdeotte for teaming up with us. We've learned a lot from you. No wonder you're one of the best **kagglers**.\n\n## Data Processing:\n* **Data-Cleaning**: We removed duplicates manually for both classification and detection. Also for detection, instead of taking one image per patient we only removed unannotated images. For this reason, our **CV** was low comparing most people but we had a good **lb** correlation.\n* **Cross-Validation**: StratifiedGroupKFold.\n\n## Study-Level:\n* **Pretraining**: We pretrained our models on chexpert. So, all our models use it except(Chris's aux-loss model).\n* **Models**: EfficientnetB6 & EfficientnetB7. (Others models didn't do well for us not even `effnetv2`, `resnet200d`, `nfnet`, `vit`).\n* **Loss**: CategoricalCrossEntropy.\n* **Label-Smoothing**: 0.01\n* **Augmentation**: HorizontalFlip, VerticalFlip, RandomRotation, CoarseDropout, RandomShift, RandomZoom, Random(Brightness, Contrast), Cutmix-Mixup.\n* **Scheduler**: WarmpupExponentialDecay.\n* **Image-Size**: 512, 640, 768.\n* **Pseudo-Labeling**: BIMCV + RICORD + RSNA.\n* **Knowledge-Distillation(KD)**: We used around 5 of our best models to generate **soft-labels** for it and one of our submissions uses 3 KD models.\n* **Aux-Loss**: One of our submissions uses **aux-loss** model by @cdeotte which uses FPN and EfficientnetB4 as backbone. You can check out the post [here][1] for more details. We also used **aux-loss** from @hengck23 's discussion and used it for generating **soft-labels** for KD model. \n* **Post-Processing:** We used **geometric-mean** to re-rank our confidence score(inspired from @cdeotte's solution of VinBigData [here][2]). Our **4cls** models were very dominating that this had very little impact on the score. So finally we decided not to use it even though it did improve the **CV** a little bit.\n<img src=\"https://i.ibb.co/WtPVw7x/eq1.png\" alt=\"eq1\" border=\"0\">\n    \n## 2cls Model:\nWe used only one **2cls** model. Including more didn't have that much impact.\n* **Model**: EfficientNetb7\n* **Image-Size**: 640\n* **External Data**: We used **RICORD** dataset and their labels posted by @raddar from [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/240187). We simply took **max-voting** to get the labels from different annotators.\n\n## Image-Level:\n* **Pretraining**: We pretrained all backbones of detection models.\n* **External Data**: RSNA - `opacity` only (6k+).\n* **Models**: yolov5x-transformer(thanks to @hengck23 for posting it), yolov5x6, yolov3-spp.\n* **Augmentation**: HorizontalFlip, VerticalFlip, Random(Brightness, Contrast), Mosaic-Mixup.\n* **Scheduler**: WarmpupCosineDecay.\n* **Image-Size**: 512. (Increasing image size worsen our result cuz our backbones were pretrained on **512** image size. We didn't have the time to pretrain on large image-size)\n* **Ensemble**: We used **WBF** to merge boxes from different models. (thanks for @cdeotte for making it work).\n* **Post-Processing:** We used **geometric-mean** to re-rank our confidence score(inspired from @cdeotte's solution of VinBigData [here][2]). We used both **4cls** and **2cls** models' prediction here.\n <img src=\"https://i.ibb.co/jR3RXBN/eq2.png\" alt=\"eq2\" border=\"0\">\n* **BBox-Filter**: We filtered out abnormal boxes based on their features. This has little effect on the both **CV** and **LB** still we kept it.\n\n## Final Submissions:\nWe tried to keep our two final submissions as different as possible. So two of our submissions use different classification models. When we looked back we noticed our best **CV** submission doesn't use any pseudo so we tried to incorporate them in our best **LB** submission. But as we had very little time we went for full data training with **BIMCV** + **RICORD** + **RSNA** pseudo. But we didn't want to take too much risk as we had very few submission left so ensembled full data models with our **LB 0.652** submission. For both submissions, 2cls model & detection models are kept the same.\n* Best **CV**:\n    * b6 | 512 | KD\n    * b6 | 640 | KD\n    * b7 | 512 | KD\n    * FPN | b4 | 512 | aux_loss\n* Best **LB**:\n    * (b6, b7) x (512, 640, 768) - 6 models | BIMCV+RICORD+RSNA pseudo | Full Data\n    * b6 | 512 | KD\n    * b6 | 512 | BIMCV+RICORD pseudo\n    * b6 | 512\n    \n## Result:\n\n* **CV**:\n\n| none  | opacity | study | final |\n| ----- | ------- | ----- | ----- |\n| 0.823 |  0.562  | 0.607 | 0.635 |\n\n\n* **LB**:\n\n| public/private | none  | opacity | study | final |\n| :--------------: | :-----: |  -----  | ----- | ----- |\n| public | 0.810 |  0.594  | 0.628 | 0.654 | \n| private  | 0.822 |  0.594  | 0.592 | 0.631 | \n\n\n[1]: https://www.kaggle.com/c/siim-covid19-detection/discussion/263676\n\n[2]: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\n\n**Inference Code:** https://www.kaggle.com/awsaf49/siim-covid19-detection-final-infer\n\n## Thank you for reading such a long post :)",
      "votes": null
    },
    {
      "id": "1467324",
      "postDate": "08/11/2021 23:41:50",
      "content": "<p>Thank you for the explanation, good luck for the next time 🙏</p>",
      "rawMarkdown": "Thank you for the explanation, good luck for the next time 🙏",
      "votes": null
    },
    {
      "id": "1467499",
      "postDate": "08/12/2021 03:00:41",
      "content": "<p>How do your team \"removed duplicates manually for both classification and detection\" ? Is this required domain knowledge ?</p>",
      "rawMarkdown": "How do your team \"removed duplicates manually for both classification and detection\" ? Is this required domain knowledge ?",
      "votes": null
    },
    {
      "id": "1467510",
      "postDate": "08/12/2021 03:07:17",
      "content": "<p>not exactly, finding duplicates wasn't that tough. As duplicates will be from the same patient we had to compare images from the same patients. </p>",
      "rawMarkdown": "not exactly, finding duplicates wasn't that tough. As duplicates will be from the same patient we had to compare images from the same patients.",
      "votes": null
    },
    {
      "id": "1467514",
      "postDate": "08/12/2021 03:13:57",
      "content": "<p>Thank you for anwsering and congrats for the gold medal</p>",
      "rawMarkdown": "Thank you for anwsering and congrats for the gold medal",
      "votes": null
    },
    {
      "id": "1468141",
      "postDate": "08/12/2021 09:20:41",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/zaber666\" target=\"_blank\">@zaber666</a>, <a href=\"https://www.kaggle.com/nexh98\" target=\"_blank\">@nexh98</a>, and <a href=\"https://www.kaggle.com/artemenon\" target=\"_blank\">@artemenon</a> for your winning🎉,<br>\nThank you so much for sharing your solution, <br>\nThis is the neatest and clean solution in this competition</p>",
      "rawMarkdown": "Congratulations @awsaf49, @cdeotte, @zaber666, @nexh98, and @artemenon for your winning🎉,\nThank you so much for sharing your solution, \nThis is the neatest and clean solution in this competition",
      "votes": null
    },
    {
      "id": "1468314",
      "postDate": "08/12/2021 10:54:53",
      "content": "<p>Hey congrats to all of you ! </p>\n<p>Everybody that started the competition before the reset of the metric remember your team above 0.410 in the first few weeks, I'm wondering what was the trick to those really fast high score ? Is it the pre-training on the Chexpert dataset ? </p>",
      "rawMarkdown": "Hey congrats to all of you ! \n\nEverybody that started the competition before the reset of the metric remember your team above 0.410 in the first few weeks, I'm wondering what was the trick to those really fast high score ? Is it the pre-training on the Chexpert dataset ?",
      "votes": null
    },
    {
      "id": "1468327",
      "postDate": "08/12/2021 11:02:54",
      "content": "<p>nice to see that my external data helped :) max voting for simplicity or did it just work better compared to soft labels - i.e. targets = 1/3, 1/3, 1/3 if different categories?</p>",
      "rawMarkdown": "nice to see that my external data helped :) max voting for simplicity or did it just work better compared to soft labels - i.e. targets = 1/3, 1/3, 1/3 if different categories?",
      "votes": null
    },
    {
      "id": "1468369",
      "postDate": "08/12/2021 11:34:22",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> and team for maintaning your high position throughout all the competition.</p>\n<ul>\n<li>Could you please elaborate more on this? and how much boost did you get from it?</li>\n</ul>\n<blockquote>\n  <p>Knowledge-Distillation(KD): We used around 5 of our best models to generate soft-labels for it and one of our submissions uses 3 KD models.</p>\n</blockquote>",
      "rawMarkdown": "Congratulations @awsaf49 and team for maintaning your high position throughout all the competition.\n* Could you please elaborate more on this? and how much boost did you get from it?\n> Knowledge-Distillation(KD): We used around 5 of our best models to generate soft-labels for it and one of our submissions uses 3 KD models.",
      "votes": null
    },
    {
      "id": "1468425",
      "postDate": "08/12/2021 12:17:14",
      "content": "<p>Well, we used it as <strong>hard labels</strong> but <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thought of using it as <strong>soft-labels</strong> but somehow we forgot to give it shot. I guess <strong>soft-labels</strong> would've been better.</p>",
      "rawMarkdown": "Well, we used it as **hard labels** but @cdeotte thought of using it as **soft-labels** but somehow we forgot to give it shot. I guess **soft-labels** would've been better.",
      "votes": null
    },
    {
      "id": "1468446",
      "postDate": "08/12/2021 12:25:22",
      "content": "<p>I don't think it was only due to  <code>chexpert</code>, our <strong>data-split</strong> always created lower cv but better <strong>lb</strong>. Even when I tried <strong>public-kernels</strong> in our <strong>split</strong> it scores lower <strong>cv</strong> but better <strong>lb</strong>.</p>",
      "rawMarkdown": "I don't think it was only due to  `chexpert`, our **data-split** always created lower cv but better **lb**. Even when I tried **public-kernels** in our **split** it scores lower **cv** but better **lb**.",
      "votes": null
    },
    {
      "id": "1468453",
      "postDate": "08/12/2021 12:27:54",
      "content": "<p>We had a bunch of models in the end so we used them for <strong>Knowledge Distillation</strong> . Here's our tentative improvement,</p>\n<ul>\n<li>baseline : <code>0.38+</code></li>\n<li>aux-loss : <code>0.39+</code></li>\n<li>kd: <code>0.40+</code></li>\n</ul>",
      "rawMarkdown": "We had a bunch of models in the end so we used them for **Knowledge Distillation** . Here's our tentative improvement,\n* baseline : `0.38+`\n* aux-loss : `0.39+`\n* kd: `0.40+`",
      "votes": null
    },
    {
      "id": "1468457",
      "postDate": "08/12/2021 12:29:27",
      "content": "<p>Initially, we tried to use it for <strong>4class</strong> it didn't work, but when we used it for <strong>2class</strong> it did improve the <strong>cv</strong> and <strong>lb</strong></p>",
      "rawMarkdown": "Initially, we tried to use it for **4class** it didn't work, but when we used it for **2class** it did improve the **cv** and **lb**",
      "votes": null
    },
    {
      "id": "1474532",
      "postDate": "08/16/2021 06:49:05",
      "content": "<p>Congrats and thanks a lot for this detailed walkthrough. Cheers!</p>",
      "rawMarkdown": "Congrats and thanks a lot for this detailed walkthrough. Cheers!",
      "votes": null
    },
    {
      "id": "1652293",
      "postDate": "01/16/2022 13:49:50",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>  !!</p>\n<p>I have one question.</p>\n<p>Did you use the same kfold division for all models when performing knowledge distillation? I would like to hear the wise man's opinion, because I'm worried about leaks when doing knowledge distillation.</p>",
      "rawMarkdown": "Congratulations @awsaf49  !!\n\nI have one question.\n\nDid you use the same kfold division for all models when performing knowledge distillation? I would like to hear the wise man's opinion, because I'm worried about leaks when doing knowledge distillation.",
      "votes": null
    },
    {
      "id": "1652306",
      "postDate": "01/16/2022 13:58:43",
      "content": "<p>Yes, if you <strong>train_labels</strong> then you can easily avoid leaks…</p>",
      "rawMarkdown": "Yes, if you **train_labels** then you can easily avoid leaks...",
      "votes": null
    },
    {
      "id": "1652318",
      "postDate": "01/16/2022 14:07:12",
      "content": "<p>thanks for replying!!</p>",
      "rawMarkdown": "thanks for replying!!",
      "votes": null
    },
    {
      "id": "1652371",
      "postDate": "01/16/2022 14:50:37",
      "content": "<p>Just sharing my experience with knowledge distillation, even if there is no leakage correlation between cv lb is somehow deteriorated (cv score boost &gt;&gt; unseen data score boost). I am unable to figure out the exact reason behind that. </p>",
      "rawMarkdown": "Just sharing my experience with knowledge distillation, even if there is no leakage correlation between cv lb is somehow deteriorated (cv score boost >> unseen data score boost). I am unable to figure out the exact reason behind that.",
      "votes": null
    },
    {
      "id": "1652489",
      "postDate": "01/16/2022 17:14:55",
      "content": "<p>that's weird. if we use <strong>valid_labels</strong>(leaky) then we get overestimated score hence <code>cv&gt;lb</code> but if we use <strong>train_labels</strong> then cv &amp; lb is comparable for us.<br>\nbtw if we use <strong>train_labels</strong> again we may face <strong>overfitting</strong> it could be the reason …</p>",
      "rawMarkdown": "that's weird. if we use **valid_labels**(leaky) then we get overestimated score hence `cv>lb` but if we use **train_labels** then cv & lb is comparable for us.\nbtw if we use **train_labels** again we may face **overfitting** it could be the reason ...",
      "votes": null
    },
    {
      "id": "1660325",
      "postDate": "01/22/2022 15:38:51",
      "content": "<p>When I tried with the data of this competition, LB increased by KD, but CV increased remarkably and the gap between CV and LB became large. Have you experienced the same thing?</p>",
      "rawMarkdown": "When I tried with the data of this competition, LB increased by KD, but CV increased remarkably and the gap between CV and LB became large. Have you experienced the same thing?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1467324,
      "author_name": "iniestamoh",
      "author_url": "",
      "post_date": "08/11/2021 23:41:50",
      "content": "<p>Thank you for the explanation, good luck for the next time 🙏</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1467499,
      "author_name": "researchbntz",
      "author_url": "",
      "post_date": "08/12/2021 03:00:41",
      "content": "<p>How do your team \"removed duplicates manually for both classification and detection\" ? Is this required domain knowledge ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1467510,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "08/12/2021 03:07:17",
          "content": "<p>not exactly, finding duplicates wasn't that tough. As duplicates will be from the same patient we had to compare images from the same patients. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1467514,
          "author_name": "researchbntz",
          "author_url": "",
          "post_date": "08/12/2021 03:13:57",
          "content": "<p>Thank you for anwsering and congrats for the gold medal</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1468141,
      "author_name": "iftiben10",
      "author_url": "",
      "post_date": "08/12/2021 09:20:41",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>, <a href=\"https://www.kaggle.com/zaber666\" target=\"_blank\">@zaber666</a>, <a href=\"https://www.kaggle.com/nexh98\" target=\"_blank\">@nexh98</a>, and <a href=\"https://www.kaggle.com/artemenon\" target=\"_blank\">@artemenon</a> for your winning🎉,<br>\nThank you so much for sharing your solution, <br>\nThis is the neatest and clean solution in this competition</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1468314,
      "author_name": "tomdarmon",
      "author_url": "",
      "post_date": "08/12/2021 10:54:53",
      "content": "<p>Hey congrats to all of you ! </p>\n<p>Everybody that started the competition before the reset of the metric remember your team above 0.410 in the first few weeks, I'm wondering what was the trick to those really fast high score ? Is it the pre-training on the Chexpert dataset ? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1468446,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "08/12/2021 12:25:22",
          "content": "<p>I don't think it was only due to  <code>chexpert</code>, our <strong>data-split</strong> always created lower cv but better <strong>lb</strong>. Even when I tried <strong>public-kernels</strong> in our <strong>split</strong> it scores lower <strong>cv</strong> but better <strong>lb</strong>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1468327,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "08/12/2021 11:02:54",
      "content": "<p>nice to see that my external data helped :) max voting for simplicity or did it just work better compared to soft labels - i.e. targets = 1/3, 1/3, 1/3 if different categories?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1468425,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "08/12/2021 12:17:14",
          "content": "<p>Well, we used it as <strong>hard labels</strong> but <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> thought of using it as <strong>soft-labels</strong> but somehow we forgot to give it shot. I guess <strong>soft-labels</strong> would've been better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1468457,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "08/12/2021 12:29:27",
          "content": "<p>Initially, we tried to use it for <strong>4class</strong> it didn't work, but when we used it for <strong>2class</strong> it did improve the <strong>cv</strong> and <strong>lb</strong></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1468369,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "08/12/2021 11:34:22",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> and team for maintaning your high position throughout all the competition.</p>\n<ul>\n<li>Could you please elaborate more on this? and how much boost did you get from it?</li>\n</ul>\n<blockquote>\n  <p>Knowledge-Distillation(KD): We used around 5 of our best models to generate soft-labels for it and one of our submissions uses 3 KD models.</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 1468453,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "08/12/2021 12:27:54",
          "content": "<p>We had a bunch of models in the end so we used them for <strong>Knowledge Distillation</strong> . Here's our tentative improvement,</p>\n<ul>\n<li>baseline : <code>0.38+</code></li>\n<li>aux-loss : <code>0.39+</code></li>\n<li>kd: <code>0.40+</code></li>\n</ul>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1474532,
      "author_name": "ashutoshsahay",
      "author_url": "",
      "post_date": "08/16/2021 06:49:05",
      "content": "<p>Congrats and thanks a lot for this detailed walkthrough. Cheers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1652293,
      "author_name": "abebe9849",
      "author_url": "",
      "post_date": "01/16/2022 13:49:50",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a>  !!</p>\n<p>I have one question.</p>\n<p>Did you use the same kfold division for all models when performing knowledge distillation? I would like to hear the wise man's opinion, because I'm worried about leaks when doing knowledge distillation.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1652306,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "01/16/2022 13:58:43",
          "content": "<p>Yes, if you <strong>train_labels</strong> then you can easily avoid leaks…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1652318,
          "author_name": "abebe9849",
          "author_url": "",
          "post_date": "01/16/2022 14:07:12",
          "content": "<p>thanks for replying!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1652371,
          "author_name": "nischaydnk",
          "author_url": "",
          "post_date": "01/16/2022 14:50:37",
          "content": "<p>Just sharing my experience with knowledge distillation, even if there is no leakage correlation between cv lb is somehow deteriorated (cv score boost &gt;&gt; unseen data score boost). I am unable to figure out the exact reason behind that. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1652489,
          "author_name": "awsaf49",
          "author_url": "",
          "post_date": "01/16/2022 17:14:55",
          "content": "<p>that's weird. if we use <strong>valid_labels</strong>(leaky) then we get overestimated score hence <code>cv&gt;lb</code> but if we use <strong>train_labels</strong> then cv &amp; lb is comparable for us.<br>\nbtw if we use <strong>train_labels</strong> again we may face <strong>overfitting</strong> it could be the reason …</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1660325,
          "author_name": "abebe9849",
          "author_url": "",
          "post_date": "01/22/2022 15:38:51",
          "content": "<p>When I tried with the data of this competition, LB increased by KD, but CV increased remarkably and the gap between CV and LB became large. Have you experienced the same thing?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1466468": "First of all, I would like to thank Kaggle, SIIM, FISABIO & RSNA for this amazing competition and also to this amazing community for constantly sharing. Congratulations to all the winners, medalists, and amazing people of this community.\n\nThanks to all the members of our team @cdeotte @zaber666 @nexh98 @artemenon. This would not be possible without you. Special thanks to @cdeotte for teaming up with us. We've learned a lot from you. No wonder you're one of the best **kagglers**.\n\n## Data Processing:\n* **Data-Cleaning**: We removed duplicates manually for both classification and detection. Also for detection, instead of taking one image per patient we only removed unannotated images. For this reason, our **CV** was low comparing most people but we had a good **lb** correlation.\n* **Cross-Validation**: StratifiedGroupKFold.\n\n## Study-Level:\n* **Pretraining**: We pretrained our models on chexpert. So, all our models use it except(Chris's aux-loss model).\n* **Models**: EfficientnetB6 & EfficientnetB7. (Others models didn't do well for us not even `effnetv2`, `resnet200d`, `nfnet`, `vit`).\n* **Loss**: CategoricalCrossEntropy.\n* **Label-Smoothing**: 0.01\n* **Augmentation**: HorizontalFlip, VerticalFlip, RandomRotation, CoarseDropout, RandomShift, RandomZoom, Random(Brightness, Contrast), Cutmix-Mixup.\n* **Scheduler**: WarmpupExponentialDecay.\n* **Image-Size**: 512, 640, 768.\n* **Pseudo-Labeling**: BIMCV + RICORD + RSNA.\n* **Knowledge-Distillation(KD)**: We used around 5 of our best models to generate **soft-labels** for it and one of our submissions uses 3 KD models.\n* **Aux-Loss**: One of our submissions uses **aux-loss** model by @cdeotte which uses FPN and EfficientnetB4 as backbone. You can check out the post [here][1] for more details. We also used **aux-loss** from @hengck23 's discussion and used it for generating **soft-labels** for KD model. \n* **Post-Processing:** We used **geometric-mean** to re-rank our confidence score(inspired from @cdeotte's solution of VinBigData [here][2]). Our **4cls** models were very dominating that this had very little impact on the score. So finally we decided not to use it even though it did improve the **CV** a little bit.\n<img src=\"https://i.ibb.co/WtPVw7x/eq1.png\" alt=\"eq1\" border=\"0\">\n    \n## 2cls Model:\nWe used only one **2cls** model. Including more didn't have that much impact.\n* **Model**: EfficientNetb7\n* **Image-Size**: 640\n* **External Data**: We used **RICORD** dataset and their labels posted by @raddar from [here](https://www.kaggle.com/c/siim-covid19-detection/discussion/240187). We simply took **max-voting** to get the labels from different annotators.\n\n## Image-Level:\n* **Pretraining**: We pretrained all backbones of detection models.\n* **External Data**: RSNA - `opacity` only (6k+).\n* **Models**: yolov5x-transformer(thanks to @hengck23 for posting it), yolov5x6, yolov3-spp.\n* **Augmentation**: HorizontalFlip, VerticalFlip, Random(Brightness, Contrast), Mosaic-Mixup.\n* **Scheduler**: WarmpupCosineDecay.\n* **Image-Size**: 512. (Increasing image size worsen our result cuz our backbones were pretrained on **512** image size. We didn't have the time to pretrain on large image-size)\n* **Ensemble**: We used **WBF** to merge boxes from different models. (thanks for @cdeotte for making it work).\n* **Post-Processing:** We used **geometric-mean** to re-rank our confidence score(inspired from @cdeotte's solution of VinBigData [here][2]). We used both **4cls** and **2cls** models' prediction here.\n <img src=\"https://i.ibb.co/jR3RXBN/eq2.png\" alt=\"eq2\" border=\"0\">\n* **BBox-Filter**: We filtered out abnormal boxes based on their features. This has little effect on the both **CV** and **LB** still we kept it.\n\n## Final Submissions:\nWe tried to keep our two final submissions as different as possible. So two of our submissions use different classification models. When we looked back we noticed our best **CV** submission doesn't use any pseudo so we tried to incorporate them in our best **LB** submission. But as we had very little time we went for full data training with **BIMCV** + **RICORD** + **RSNA** pseudo. But we didn't want to take too much risk as we had very few submission left so ensembled full data models with our **LB 0.652** submission. For both submissions, 2cls model & detection models are kept the same.\n* Best **CV**:\n    * b6 | 512 | KD\n    * b6 | 640 | KD\n    * b7 | 512 | KD\n    * FPN | b4 | 512 | aux_loss\n* Best **LB**:\n    * (b6, b7) x (512, 640, 768) - 6 models | BIMCV+RICORD+RSNA pseudo | Full Data\n    * b6 | 512 | KD\n    * b6 | 512 | BIMCV+RICORD pseudo\n    * b6 | 512\n    \n## Result:\n\n* **CV**:\n\n| none  | opacity | study | final |\n| ----- | ------- | ----- | ----- |\n| 0.823 |  0.562  | 0.607 | 0.635 |\n\n\n* **LB**:\n\n| public/private | none  | opacity | study | final |\n| :--------------: | :-----: |  -----  | ----- | ----- |\n| public | 0.810 |  0.594  | 0.628 | 0.654 | \n| private  | 0.822 |  0.594  | 0.592 | 0.631 | \n\n\n[1]: https://www.kaggle.com/c/siim-covid19-detection/discussion/263676\n\n[2]: https://www.kaggle.com/c/vinbigdata-chest-xray-abnormalities-detection/discussion/229637\n\n**Inference Code:** https://www.kaggle.com/awsaf49/siim-covid19-detection-final-infer\n\n## Thank you for reading such a long post :)",
    "1467324": "Thank you for the explanation, good luck for the next time 🙏",
    "1467499": "How do your team \"removed duplicates manually for both classification and detection\" ? Is this required domain knowledge ?",
    "1467510": "not exactly, finding duplicates wasn't that tough. As duplicates will be from the same patient we had to compare images from the same patients.",
    "1467514": "Thank you for anwsering and congrats for the gold medal",
    "1468141": "Congratulations @awsaf49, @cdeotte, @zaber666, @nexh98, and @artemenon for your winning🎉,\nThank you so much for sharing your solution, \nThis is the neatest and clean solution in this competition",
    "1468314": "Hey congrats to all of you ! \n\nEverybody that started the competition before the reset of the metric remember your team above 0.410 in the first few weeks, I'm wondering what was the trick to those really fast high score ? Is it the pre-training on the Chexpert dataset ?",
    "1468327": "nice to see that my external data helped :) max voting for simplicity or did it just work better compared to soft labels - i.e. targets = 1/3, 1/3, 1/3 if different categories?",
    "1468369": "Congratulations @awsaf49 and team for maintaning your high position throughout all the competition.\n* Could you please elaborate more on this? and how much boost did you get from it?\n> Knowledge-Distillation(KD): We used around 5 of our best models to generate soft-labels for it and one of our submissions uses 3 KD models.",
    "1468425": "Well, we used it as **hard labels** but @cdeotte thought of using it as **soft-labels** but somehow we forgot to give it shot. I guess **soft-labels** would've been better.",
    "1468446": "I don't think it was only due to  `chexpert`, our **data-split** always created lower cv but better **lb**. Even when I tried **public-kernels** in our **split** it scores lower **cv** but better **lb**.",
    "1468453": "We had a bunch of models in the end so we used them for **Knowledge Distillation** . Here's our tentative improvement,\n* baseline : `0.38+`\n* aux-loss : `0.39+`\n* kd: `0.40+`",
    "1468457": "Initially, we tried to use it for **4class** it didn't work, but when we used it for **2class** it did improve the **cv** and **lb**",
    "1474532": "Congrats and thanks a lot for this detailed walkthrough. Cheers!",
    "1652293": "Congratulations @awsaf49  !!\n\nI have one question.\n\nDid you use the same kfold division for all models when performing knowledge distillation? I would like to hear the wise man's opinion, because I'm worried about leaks when doing knowledge distillation.",
    "1652306": "Yes, if you **train_labels** then you can easily avoid leaks...",
    "1652318": "thanks for replying!!",
    "1652371": "Just sharing my experience with knowledge distillation, even if there is no leakage correlation between cv lb is somehow deteriorated (cv score boost >> unseen data score boost). I am unable to figure out the exact reason behind that.",
    "1652489": "that's weird. if we use **valid_labels**(leaky) then we get overestimated score hence `cv>lb` but if we use **train_labels** then cv & lb is comparable for us.\nbtw if we use **train_labels** again we may face **overfitting** it could be the reason ...",
    "1660325": "When I tried with the data of this competition, LB increased by KD, but CV increased remarkably and the gap between CV and LB became large. Have you experienced the same thing?"
  },
  "source": "meta"
}