{
  "id": 226561,
  "title": "Solo 60th Place Solution - Forward Ensembling & Pseudo Labeling External Data",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/writeups/tucker-arrants-solo-60th-place-solution-forward-en",
  "author_name": "",
  "post_date": "2021-03-20T03:18:29.553Z",
  "votes": 31,
  "comment_count": 13,
  "views": 0,
  "content": "<p><strong>60th Place Solution</strong></p>\n<p>Thank you Kaggle and RANZCR for hosting this competition. It was the first competition I really immersed myself in and I have learned a tremendous amount. Thank you to all the competitors, especially those that so openly shared their results and techniques. Kaggle is a wonderful medium through which to learn data science, and it would not be the same without everyone sharing and communicating. </p>\n<p>I am not pleased with my final standing, but I believe that all Kagglers should share their procedure, irrespective of their final position, because it is highly unlikely that anyone else did exactly what you did, so there is always something to learn from reading other's write ups. Here it goes.</p>\n<p><strong>Image resolutions</strong></p>\n<p>My approach was a brute force approach and not particularly elegant. I began with <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>’s kernels on 3 stage training with annotations and changed several parameters like learning rate, batch size, training augmentations, and optimizer (AdamP showed better CV results than Adam). I also generated stratified folds by coupling <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@sin</a>'s <a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">kernel</a> with some <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064\" target=\"_blank\">useful information</a> about incorrect labels / inverted images shared by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and others. Once I had established a CV strategy, I started to experiment with different image resolutions, all with a ResNet200D. I tried image sizes from <code>512</code> to <code>736</code> in increments of <code>16</code> (all with a batch size of <code>32</code>). I saw an increase in AUC + decrease in BCE loss each time I bumped up the image resolution. I validated all models against my first fold to save time (more on that later). </p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217296\" target=\"_blank\">There was some discussion</a> as to whether or not this multi stage approach with teacher-student training on annotated images actually gave any performance boost. I continued to use it anyways for one reason: you can use the stage 2 weights to lower training time. I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights. In other words, instead of repeating all 3 stages anytime I changed image resolutions / other parameters, I simply reused stage 2 weights and skipped to stage 3. This is not ideal, but it saved me a ton of training time and allowed me to run more experiments. I figured it would work seeing as we frequently fine tune models that have been trained on ImageNet with different image resolutions than the ones they were initially trained on. </p>\n<p><strong>Preprocessing</strong></p>\n<p>I did minimal preprocessing. <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146\" target=\"_blank\">It was mentioned</a> that some images have black borders around them. I removed them during inference and saw an increase in score. I hesitated to remove them from training as these black boxes could serve as a form of regularization, but decided to remove them for the last couple models I trained. I figured if they helped with inference, they would help with training, but I did not run any real experiments to verify this.</p>\n<p><strong>Forward Selection</strong></p>\n<p>I then created a fold prediction correlation heatmap for these different image sizes and saw a surprisingly low correlation between the ResNet200D model when trained on different image sizes (0.85 - 0.91). This showed me that one could get good model diversity by simply changing image resolutions, which got me thinking about ensembling. I remembered reading <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@chrisdeotte</a>’s <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">kernel on forward selection</a> from the Melanoma competition and decided to try it. I saw nice improvements when doing so and decided to experiment with different model architectures as well. I trained a SEResNet152D on <code>672</code> and <code>720</code> image resolutions and checked the fold prediction correlations between them with my previous ResNet200D models: the correlation was even lower, so I threw them into forward selection with my other ResNet200D models. </p>\n<p>The above procedure got me to CV <code>0.966</code> and public leaderboard <code>0.968</code>. I then experimented with RegNetYs, EfficientNets, and ResNeSts. RegNetYs took too long to train and EfficientNets were never able to reach the same CV score as the ResNet200D / SEResNet152D. Only the ResNeSt’s got similar scores, but simply took too long to train. I had reached a bit of an impasse and was not sure how to proceed.</p>\n<p><strong>Pseudo labels</strong></p>\n<p>I then saw <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">this post</a> on duplicate images between the NIH Chest XRays dataset and the training / public test images given to use in this competition. Seeing as I already had an ensemble of models with good diversity, I was confident I could generate fairly high quality pseudo labels from this external dataset. (Note that you must generate pseudo labels from models that have not seen the fold you are validating against otherwise you get subtle CV leakage). I first tried to remove duplicates using CNN embeddings and RAPIDS NearestNeighbors, but this approach still left around 500 duplicate images, so I combined this approach with ImageHash to remove the remaining duplicates. I then started generating pseudo labels. I decided to use soft pseudo labels where at least one catheter class is present (any column with prediction above <code>0.7</code>) and all other classes without any catheters (remaining columns with predictions below <code>0.3</code>). I added Gaussian noise to the pseudo labels in an attempt to reduce confirmation bias.</p>\n<p>Instead of re-training the model jointly with concatenated labeled + pseudo labels, I decided to introduce another training stage in which I train on only NIH pseudo labels. The final stage is then fine tuning on labeled training data. I also decided to train the last few epochs without any data augmentation, based on <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212532\" target=\"_blank\">another discussion post</a>, and saw further improvements. I carried out this procedure for my top-3 highest scoring models and threw them back into my forward selection algorithm to compare CV and public leaderboard score. It indeed increased. </p>\n<p>I am still not sure if my CV was entirely leak free. I think many teams that used external data had some form of leakage. For this reason, I only included 2 NIH pseudo label pretrained models in the folds I forward selected against. Had I invested more time in examining these duplicate images found by RAPIDS NN and ImageHash, I think I would have had a more robust CV and better public / private score. </p>\n<p><strong>Training / validation folds</strong></p>\n<p>Up to this point, I had only validated against the first fold, so in some sense I was ‘overfitting’ my own validation set. This is not ideal; it would be much better to validate each model against all folds and use the full OOF to ensemble, as opposed to single fold predictions. That being said, I joined this competition fairly late so I needed to save time somehow while simultaneously running many experiments. I also trusted my fold splits and was not overly concerned with validating against only one fold. However, I did want my final predictions to encompass all training data, so I repeated what I did on the first fold with the second fold, so that my models (in total) had been trained on all data folds. </p>\n<p>At the end, my approach was:</p>\n<ul>\n<li>Train teacher on annotated samples  (light augmentations for 4 epochs) </li>\n<li>Transfer this information to student  (light augmentations for 6-7 epochs)</li>\n<li>Use these student weights as a starting point for NIH pseudo labeling training (light augmentations for 7-8 epochs)</li>\n<li>Train on full labeled data (heavy augmentation for 7-8 epochs)</li>\n<li>Train on full labeled data (no augmentation for 1-2 epochs)</li>\n<li>Repeat on different image sizes and model architectures and forward select based on validation fold predictions</li>\n</ul>\n<p>I had around 10 models validated against fold 1 and 10 validated against fold 2. I forward selected each fold separately and averaged their predictions. Seeing as I never got around to full KFold training, I thought that including public models would add some beneficial diversity, so I added <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@sin</a>’s public 5Fold ResNet200D models to inference. My final submission was a convex combination of my own models weighted <code>0.85</code> and <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">sin</a>’s 5 models weighted at <code>0.15</code>.</p>\n<p>I chose my best public leaderboard score for my first submission (seeing as I trusted CV and public correlation) and chose my best CV score for my second submission. They were both <code>0.971</code> on the private leaderboard and not surprisingly, my best CV score was also my best private leaderboard score. </p>\n<p><strong>What didn’t really work</strong></p>\n<ul>\n<li>TTA beyond simple horizontal flipping only showed improvements when the number of steps was fairly high (&gt;5) and this took too long to infer with</li>\n<li>Models other than ResNet200D and SEResNet152D (ResNeSt200E worked as well but took too long to train)</li>\n<li>Image resolutions below <code>512</code></li>\n</ul>\n<p><strong>What I wanted to test, but didn’t</strong></p>\n<ul>\n<li>Experiment with changing image resolutions during training as a means of regularization</li>\n<li>Using shape priors and a segmentation model to create features (categorize catheter overlaps with HoG descriptors and intensity histograms)</li>\n<li>Use masks and segmentation model outputs to categorize distance between various anatomical features and the catheters to create more features</li>\n<li>Upsample images without any catheters present with NIH Chest XRay images and compare performance</li>\n<li>Experiment with label smoothing / special loss functions</li>\n<li>Use MixUp on pseudo labels to further reduce confirmation bias</li>\n</ul>\n<p><strong>Acknowledgements</strong></p>\n<p>I would first like to thank <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing their approach on how to use annotated images to increase CV scores. Without their generous contributions in this competition, I would not have made much progress during these past couple months. </p>\n<p>I also want to thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@chrisdeotte</a> for his wonderful notebooks on forward selection, <a href=\"https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\" target=\"_blank\">pseudo-labeling</a>, and <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">finding duplicate images with RAPIDS</a>. Forward selection was very useful for generating high quality pseudo labels and he inspired me to experiment with ensembling models trained on different image sizes. Thank you Chris.</p>\n<p>Additionally, I would like to thank <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for sharing domain knowledge so openly during this competition and <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@sin</a> for sharing his use of heavy augmentations and his cross-validation strategy. In coupling their ideas, I was able to generate a robust CV strategy that informed all my decisions during this competition.</p>\n<p>Lastly, I want to thank Z by HP &amp; NVIDIA for providing me with a Z8 workstation and ZBook Studio laptop. Without these GPUs, I would not have been able to experiment with these larger networks and different image resolutions so quickly. I ran smaller experiments / debugged my code on the laptop and ran the main experiments on the workstation. GPU memory was the bottleneck in this competition so having powerful local GPUs was integral to my solution. (RTX Quadro 8000 and RTX Quadro 5000 mobile, for reference). When it comes to deep learning and GPUs, bigger seems to be better. That being said, I wish I found more elegant approaches to circumvent this problem, but I am still learning. Next time. </p>",
  "messages": [
    {
      "id": "1241203",
      "postDate": "03/17/2021 00:15:48",
      "content": "<p><strong>60th Place Solution</strong></p>\n<p>Thank you Kaggle and RANZCR for hosting this competition. It was the first competition I really immersed myself in and I have learned a tremendous amount. Thank you to all the competitors, especially those that so openly shared their results and techniques. Kaggle is a wonderful medium through which to learn data science, and it would not be the same without everyone sharing and communicating. </p>\n<p>I am not pleased with my final standing, but I believe that all Kagglers should share their procedure, irrespective of their final position, because it is highly unlikely that anyone else did exactly what you did, so there is always something to learn from reading other's write ups. Here it goes.</p>\n<p><strong>Image resolutions</strong></p>\n<p>My approach was a brute force approach and not particularly elegant. I began with <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a>’s kernels on 3 stage training with annotations and changed several parameters like learning rate, batch size, training augmentations, and optimizer (AdamP showed better CV results than Adam). I also generated stratified folds by coupling <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@sin</a>'s <a href=\"https://www.kaggle.com/underwearfitting/how-to-properly-split-folds\" target=\"_blank\">kernel</a> with some <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064\" target=\"_blank\">useful information</a> about incorrect labels / inverted images shared by <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> and others. Once I had established a CV strategy, I started to experiment with different image resolutions, all with a ResNet200D. I tried image sizes from <code>512</code> to <code>736</code> in increments of <code>16</code> (all with a batch size of <code>32</code>). I saw an increase in AUC + decrease in BCE loss each time I bumped up the image resolution. I validated all models against my first fold to save time (more on that later). </p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217296\" target=\"_blank\">There was some discussion</a> as to whether or not this multi stage approach with teacher-student training on annotated images actually gave any performance boost. I continued to use it anyways for one reason: you can use the stage 2 weights to lower training time. I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights. In other words, instead of repeating all 3 stages anytime I changed image resolutions / other parameters, I simply reused stage 2 weights and skipped to stage 3. This is not ideal, but it saved me a ton of training time and allowed me to run more experiments. I figured it would work seeing as we frequently fine tune models that have been trained on ImageNet with different image resolutions than the ones they were initially trained on. </p>\n<p><strong>Preprocessing</strong></p>\n<p>I did minimal preprocessing. <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146\" target=\"_blank\">It was mentioned</a> that some images have black borders around them. I removed them during inference and saw an increase in score. I hesitated to remove them from training as these black boxes could serve as a form of regularization, but decided to remove them for the last couple models I trained. I figured if they helped with inference, they would help with training, but I did not run any real experiments to verify this.</p>\n<p><strong>Forward Selection</strong></p>\n<p>I then created a fold prediction correlation heatmap for these different image sizes and saw a surprisingly low correlation between the ResNet200D model when trained on different image sizes (0.85 - 0.91). This showed me that one could get good model diversity by simply changing image resolutions, which got me thinking about ensembling. I remembered reading <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@chrisdeotte</a>’s <a href=\"https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private\" target=\"_blank\">kernel on forward selection</a> from the Melanoma competition and decided to try it. I saw nice improvements when doing so and decided to experiment with different model architectures as well. I trained a SEResNet152D on <code>672</code> and <code>720</code> image resolutions and checked the fold prediction correlations between them with my previous ResNet200D models: the correlation was even lower, so I threw them into forward selection with my other ResNet200D models. </p>\n<p>The above procedure got me to CV <code>0.966</code> and public leaderboard <code>0.968</code>. I then experimented with RegNetYs, EfficientNets, and ResNeSts. RegNetYs took too long to train and EfficientNets were never able to reach the same CV score as the ResNet200D / SEResNet152D. Only the ResNeSt’s got similar scores, but simply took too long to train. I had reached a bit of an impasse and was not sure how to proceed.</p>\n<p><strong>Pseudo labels</strong></p>\n<p>I then saw <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808\" target=\"_blank\">this post</a> on duplicate images between the NIH Chest XRays dataset and the training / public test images given to use in this competition. Seeing as I already had an ensemble of models with good diversity, I was confident I could generate fairly high quality pseudo labels from this external dataset. (Note that you must generate pseudo labels from models that have not seen the fold you are validating against otherwise you get subtle CV leakage). I first tried to remove duplicates using CNN embeddings and RAPIDS NearestNeighbors, but this approach still left around 500 duplicate images, so I combined this approach with ImageHash to remove the remaining duplicates. I then started generating pseudo labels. I decided to use soft pseudo labels where at least one catheter class is present (any column with prediction above <code>0.7</code>) and all other classes without any catheters (remaining columns with predictions below <code>0.3</code>). I added Gaussian noise to the pseudo labels in an attempt to reduce confirmation bias.</p>\n<p>Instead of re-training the model jointly with concatenated labeled + pseudo labels, I decided to introduce another training stage in which I train on only NIH pseudo labels. The final stage is then fine tuning on labeled training data. I also decided to train the last few epochs without any data augmentation, based on <a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212532\" target=\"_blank\">another discussion post</a>, and saw further improvements. I carried out this procedure for my top-3 highest scoring models and threw them back into my forward selection algorithm to compare CV and public leaderboard score. It indeed increased. </p>\n<p>I am still not sure if my CV was entirely leak free. I think many teams that used external data had some form of leakage. For this reason, I only included 2 NIH pseudo label pretrained models in the folds I forward selected against. Had I invested more time in examining these duplicate images found by RAPIDS NN and ImageHash, I think I would have had a more robust CV and better public / private score. </p>\n<p><strong>Training / validation folds</strong></p>\n<p>Up to this point, I had only validated against the first fold, so in some sense I was ‘overfitting’ my own validation set. This is not ideal; it would be much better to validate each model against all folds and use the full OOF to ensemble, as opposed to single fold predictions. That being said, I joined this competition fairly late so I needed to save time somehow while simultaneously running many experiments. I also trusted my fold splits and was not overly concerned with validating against only one fold. However, I did want my final predictions to encompass all training data, so I repeated what I did on the first fold with the second fold, so that my models (in total) had been trained on all data folds. </p>\n<p>At the end, my approach was:</p>\n<ul>\n<li>Train teacher on annotated samples  (light augmentations for 4 epochs) </li>\n<li>Transfer this information to student  (light augmentations for 6-7 epochs)</li>\n<li>Use these student weights as a starting point for NIH pseudo labeling training (light augmentations for 7-8 epochs)</li>\n<li>Train on full labeled data (heavy augmentation for 7-8 epochs)</li>\n<li>Train on full labeled data (no augmentation for 1-2 epochs)</li>\n<li>Repeat on different image sizes and model architectures and forward select based on validation fold predictions</li>\n</ul>\n<p>I had around 10 models validated against fold 1 and 10 validated against fold 2. I forward selected each fold separately and averaged their predictions. Seeing as I never got around to full KFold training, I thought that including public models would add some beneficial diversity, so I added <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@sin</a>’s public 5Fold ResNet200D models to inference. My final submission was a convex combination of my own models weighted <code>0.85</code> and <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">sin</a>’s 5 models weighted at <code>0.15</code>.</p>\n<p>I chose my best public leaderboard score for my first submission (seeing as I trusted CV and public correlation) and chose my best CV score for my second submission. They were both <code>0.971</code> on the private leaderboard and not surprisingly, my best CV score was also my best private leaderboard score. </p>\n<p><strong>What didn’t really work</strong></p>\n<ul>\n<li>TTA beyond simple horizontal flipping only showed improvements when the number of steps was fairly high (&gt;5) and this took too long to infer with</li>\n<li>Models other than ResNet200D and SEResNet152D (ResNeSt200E worked as well but took too long to train)</li>\n<li>Image resolutions below <code>512</code></li>\n</ul>\n<p><strong>What I wanted to test, but didn’t</strong></p>\n<ul>\n<li>Experiment with changing image resolutions during training as a means of regularization</li>\n<li>Using shape priors and a segmentation model to create features (categorize catheter overlaps with HoG descriptors and intensity histograms)</li>\n<li>Use masks and segmentation model outputs to categorize distance between various anatomical features and the catheters to create more features</li>\n<li>Upsample images without any catheters present with NIH Chest XRay images and compare performance</li>\n<li>Experiment with label smoothing / special loss functions</li>\n<li>Use MixUp on pseudo labels to further reduce confirmation bias</li>\n</ul>\n<p><strong>Acknowledgements</strong></p>\n<p>I would first like to thank <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> and <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> for sharing their approach on how to use annotated images to increase CV scores. Without their generous contributions in this competition, I would not have made much progress during these past couple months. </p>\n<p>I also want to thank <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@chrisdeotte</a> for his wonderful notebooks on forward selection, <a href=\"https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969\" target=\"_blank\">pseudo-labeling</a>, and <a href=\"https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates\" target=\"_blank\">finding duplicate images with RAPIDS</a>. Forward selection was very useful for generating high quality pseudo labels and he inspired me to experiment with ensembling models trained on different image sizes. Thank you Chris.</p>\n<p>Additionally, I would like to thank <a href=\"https://www.kaggle.com/raddar\" target=\"_blank\">@raddar</a> for sharing domain knowledge so openly during this competition and <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@sin</a> for sharing his use of heavy augmentations and his cross-validation strategy. In coupling their ideas, I was able to generate a robust CV strategy that informed all my decisions during this competition.</p>\n<p>Lastly, I want to thank Z by HP &amp; NVIDIA for providing me with a Z8 workstation and ZBook Studio laptop. Without these GPUs, I would not have been able to experiment with these larger networks and different image resolutions so quickly. I ran smaller experiments / debugged my code on the laptop and ran the main experiments on the workstation. GPU memory was the bottleneck in this competition so having powerful local GPUs was integral to my solution. (RTX Quadro 8000 and RTX Quadro 5000 mobile, for reference). When it comes to deep learning and GPUs, bigger seems to be better. That being said, I wish I found more elegant approaches to circumvent this problem, but I am still learning. Next time. </p>",
      "rawMarkdown": "**60th Place Solution**\n\nThank you Kaggle and RANZCR for hosting this competition. It was the first competition I really immersed myself in and I have learned a tremendous amount. Thank you to all the competitors, especially those that so openly shared their results and techniques. Kaggle is a wonderful medium through which to learn data science, and it would not be the same without everyone sharing and communicating. \n\nI am not pleased with my final standing, but I believe that all Kagglers should share their procedure, irrespective of their final position, because it is highly unlikely that anyone else did exactly what you did, so there is always something to learn from reading other's write ups. Here it goes.\n\n**Image resolutions**\n\nMy approach was a brute force approach and not particularly elegant. I began with [@yasufuminakama](https://www.kaggle.com/yasufuminakama)’s kernels on 3 stage training with annotations and changed several parameters like learning rate, batch size, training augmentations, and optimizer (AdamP showed better CV results than Adam). I also generated stratified folds by coupling [@sin](https://www.kaggle.com/underwearfitting)'s [kernel](https://www.kaggle.com/underwearfitting/how-to-properly-split-folds) with some [useful information](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064) about incorrect labels / inverted images shared by [@raddar](https://www.kaggle.com/raddar) and others. Once I had established a CV strategy, I started to experiment with different image resolutions, all with a ResNet200D. I tried image sizes from `512` to `736` in increments of `16` (all with a batch size of `32`). I saw an increase in AUC + decrease in BCE loss each time I bumped up the image resolution. I validated all models against my first fold to save time (more on that later). \n\n[There was some discussion](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217296) as to whether or not this multi stage approach with teacher-student training on annotated images actually gave any performance boost. I continued to use it anyways for one reason: you can use the stage 2 weights to lower training time. I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights. In other words, instead of repeating all 3 stages anytime I changed image resolutions / other parameters, I simply reused stage 2 weights and skipped to stage 3. This is not ideal, but it saved me a ton of training time and allowed me to run more experiments. I figured it would work seeing as we frequently fine tune models that have been trained on ImageNet with different image resolutions than the ones they were initially trained on. \n\n**Preprocessing**\n\nI did minimal preprocessing. [It was mentioned](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146) that some images have black borders around them. I removed them during inference and saw an increase in score. I hesitated to remove them from training as these black boxes could serve as a form of regularization, but decided to remove them for the last couple models I trained. I figured if they helped with inference, they would help with training, but I did not run any real experiments to verify this.\n\n**Forward Selection**\n\nI then created a fold prediction correlation heatmap for these different image sizes and saw a surprisingly low correlation between the ResNet200D model when trained on different image sizes (0.85 - 0.91). This showed me that one could get good model diversity by simply changing image resolutions, which got me thinking about ensembling. I remembered reading [@chrisdeotte](https://www.kaggle.com/cdeotte)’s [kernel on forward selection](https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private) from the Melanoma competition and decided to try it. I saw nice improvements when doing so and decided to experiment with different model architectures as well. I trained a SEResNet152D on `672` and `720` image resolutions and checked the fold prediction correlations between them with my previous ResNet200D models: the correlation was even lower, so I threw them into forward selection with my other ResNet200D models. \n\nThe above procedure got me to CV `0.966` and public leaderboard `0.968`. I then experimented with RegNetYs, EfficientNets, and ResNeSts. RegNetYs took too long to train and EfficientNets were never able to reach the same CV score as the ResNet200D / SEResNet152D. Only the ResNeSt’s got similar scores, but simply took too long to train. I had reached a bit of an impasse and was not sure how to proceed.\n\n**Pseudo labels**\n\nI then saw [this post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808) on duplicate images between the NIH Chest XRays dataset and the training / public test images given to use in this competition. Seeing as I already had an ensemble of models with good diversity, I was confident I could generate fairly high quality pseudo labels from this external dataset. (Note that you must generate pseudo labels from models that have not seen the fold you are validating against otherwise you get subtle CV leakage). I first tried to remove duplicates using CNN embeddings and RAPIDS NearestNeighbors, but this approach still left around 500 duplicate images, so I combined this approach with ImageHash to remove the remaining duplicates. I then started generating pseudo labels. I decided to use soft pseudo labels where at least one catheter class is present (any column with prediction above `0.7`) and all other classes without any catheters (remaining columns with predictions below `0.3`). I added Gaussian noise to the pseudo labels in an attempt to reduce confirmation bias.\n\nInstead of re-training the model jointly with concatenated labeled + pseudo labels, I decided to introduce another training stage in which I train on only NIH pseudo labels. The final stage is then fine tuning on labeled training data. I also decided to train the last few epochs without any data augmentation, based on [another discussion post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212532), and saw further improvements. I carried out this procedure for my top-3 highest scoring models and threw them back into my forward selection algorithm to compare CV and public leaderboard score. It indeed increased. \n\nI am still not sure if my CV was entirely leak free. I think many teams that used external data had some form of leakage. For this reason, I only included 2 NIH pseudo label pretrained models in the folds I forward selected against. Had I invested more time in examining these duplicate images found by RAPIDS NN and ImageHash, I think I would have had a more robust CV and better public / private score. \n\n**Training / validation folds**\n\nUp to this point, I had only validated against the first fold, so in some sense I was ‘overfitting’ my own validation set. This is not ideal; it would be much better to validate each model against all folds and use the full OOF to ensemble, as opposed to single fold predictions. That being said, I joined this competition fairly late so I needed to save time somehow while simultaneously running many experiments. I also trusted my fold splits and was not overly concerned with validating against only one fold. However, I did want my final predictions to encompass all training data, so I repeated what I did on the first fold with the second fold, so that my models (in total) had been trained on all data folds. \n\nAt the end, my approach was:\n\n* Train teacher on annotated samples  (light augmentations for 4 epochs) \n* Transfer this information to student  (light augmentations for 6-7 epochs)\n* Use these student weights as a starting point for NIH pseudo labeling training (light augmentations for 7-8 epochs)\n* Train on full labeled data (heavy augmentation for 7-8 epochs)\n* Train on full labeled data (no augmentation for 1-2 epochs)\n* Repeat on different image sizes and model architectures and forward select based on validation fold predictions\n\nI had around 10 models validated against fold 1 and 10 validated against fold 2. I forward selected each fold separately and averaged their predictions. Seeing as I never got around to full KFold training, I thought that including public models would add some beneficial diversity, so I added [@sin](https://www.kaggle.com/underwearfitting)’s public 5Fold ResNet200D models to inference. My final submission was a convex combination of my own models weighted `0.85` and [sin](https://www.kaggle.com/underwearfitting)’s 5 models weighted at `0.15`.\n\nI chose my best public leaderboard score for my first submission (seeing as I trusted CV and public correlation) and chose my best CV score for my second submission. They were both `0.971` on the private leaderboard and not surprisingly, my best CV score was also my best private leaderboard score. \n\n**What didn’t really work**\n\n* TTA beyond simple horizontal flipping only showed improvements when the number of steps was fairly high (>5) and this took too long to infer with\n* Models other than ResNet200D and SEResNet152D (ResNeSt200E worked as well but took too long to train)\n* Image resolutions below `512`\n\n**What I wanted to test, but didn’t**\n\n* Experiment with changing image resolutions during training as a means of regularization\n* Using shape priors and a segmentation model to create features (categorize catheter overlaps with HoG descriptors and intensity histograms)\n* Use masks and segmentation model outputs to categorize distance between various anatomical features and the catheters to create more features\n* Upsample images without any catheters present with NIH Chest XRay images and compare performance\n* Experiment with label smoothing / special loss functions\n* Use MixUp on pseudo labels to further reduce confirmation bias\n\n**Acknowledgements**\n\nI would first like to thank [@yasufuminakama](https://www.kaggle.com/yasufuminakama) and [@hengck23](https://www.kaggle.com/hengck23) for sharing their approach on how to use annotated images to increase CV scores. Without their generous contributions in this competition, I would not have made much progress during these past couple months. \n\nI also want to thank [@chrisdeotte](https://www.kaggle.com/cdeotte) for his wonderful notebooks on forward selection, [pseudo-labeling](https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969), and [finding duplicate images with RAPIDS](https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates). Forward selection was very useful for generating high quality pseudo labels and he inspired me to experiment with ensembling models trained on different image sizes. Thank you Chris.\n\nAdditionally, I would like to thank [@raddar](https://www.kaggle.com/raddar) for sharing domain knowledge so openly during this competition and [@sin](https://www.kaggle.com/underwearfitting) for sharing his use of heavy augmentations and his cross-validation strategy. In coupling their ideas, I was able to generate a robust CV strategy that informed all my decisions during this competition.\n\nLastly, I want to thank Z by HP & NVIDIA for providing me with a Z8 workstation and ZBook Studio laptop. Without these GPUs, I would not have been able to experiment with these larger networks and different image resolutions so quickly. I ran smaller experiments / debugged my code on the laptop and ran the main experiments on the workstation. GPU memory was the bottleneck in this competition so having powerful local GPUs was integral to my solution. (RTX Quadro 8000 and RTX Quadro 5000 mobile, for reference). When it comes to deep learning and GPUs, bigger seems to be better. That being said, I wish I found more elegant approaches to circumvent this problem, but I am still learning. Next time.",
      "votes": null
    },
    {
      "id": "1241214",
      "postDate": "03/17/2021 00:30:35",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> and thanks for sharing this detailed writeup :)</p>\n<blockquote>\n  <p>I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights.</p>\n</blockquote>\n<p>How low was the image resolution you used in stage 1 and 2? Because with my models, whenever I increase the image size in step 3, both CV and LB decrease.<br>\nDid you tweak some parameters other than batch size when you increase the image size?</p>",
      "rawMarkdown": "Congratulations @tuckerarrants and thanks for sharing this detailed writeup :)\n\n>I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights.\n\nHow low was the image resolution you used in stage 1 and 2? Because with my models, whenever I increase the image size in step 3, both CV and LB decrease.\nDid you tweak some parameters other than batch size when you increase the image size?",
      "votes": null
    },
    {
      "id": "1241219",
      "postDate": "03/17/2021 00:34:05",
      "content": "<p>Thank you Amin. The lowest resolution I used in stages 1 and 2 was <code>576</code>.  Sometimes I went as high as <code>768</code> for these stages, but only before I started reusing lower resolution weights to save time.</p>\n<p>I used slightly heavier augmentations each time I increased the image size.</p>",
      "rawMarkdown": "Thank you Amin. The lowest resolution I used in stages 1 and 2 was `576`.  Sometimes I went as high as `768` for these stages, but only before I started reusing lower resolution weights to save time.\n\nI used slightly heavier augmentations each time I increased the image size.",
      "votes": null
    },
    {
      "id": "1241225",
      "postDate": "03/17/2021 00:39:27",
      "content": "<p>Great job Tucker. Thanks for sharing. Strong solo finish.</p>",
      "rawMarkdown": "Great job Tucker. Thanks for sharing. Strong solo finish.",
      "votes": null
    },
    {
      "id": "1241227",
      "postDate": "03/17/2021 00:41:03",
      "content": "<p>No thank you, I could not have done it without all your contributions. Congratulations on 2nd place, very impressive indeed. </p>",
      "rawMarkdown": "No thank you, I could not have done it without all your contributions. Congratulations on 2nd place, very impressive indeed.",
      "votes": null
    },
    {
      "id": "1241232",
      "postDate": "03/17/2021 00:44:40",
      "content": "<p>Great write-up. When looking at your ensembling efforts did you assign weights at the macro level across all columns or did you ever try to optimize at the column level? I found that in my experiment when assigning weights my ensembling technique would basically just hard select whichever model had the best performance for a given column and any combination of models didn't really seem to yield very much. </p>\n<p>Never got around to applying that knowledge to the leaderboard, but I was surprised to see that behavior. </p>\n<p><img src=\"https://i.imgur.com/zhEn4Ow.png\" alt=\"\"></p>",
      "rawMarkdown": "Great write-up. When looking at your ensembling efforts did you assign weights at the macro level across all columns or did you ever try to optimize at the column level? I found that in my experiment when assigning weights my ensembling technique would basically just hard select whichever model had the best performance for a given column and any combination of models didn't really seem to yield very much. \n\nNever got around to applying that knowledge to the leaderboard, but I was surprised to see that behavior. \n\n![](https://i.imgur.com/zhEn4Ow.png)",
      "votes": null
    },
    {
      "id": "1241247",
      "postDate": "03/17/2021 00:55:26",
      "content": "<p>Thank you. Sadly, I only got around to assigning them at the macro level. Only at the last minute did I think to do it for each column. I may play around with the column-wise approach now that the competition is over to compare performance. </p>",
      "rawMarkdown": "Thank you. Sadly, I only got around to assigning them at the macro level. Only at the last minute did I think to do it for each column. I may play around with the column-wise approach now that the competition is over to compare performance.",
      "votes": null
    },
    {
      "id": "1241336",
      "postDate": "03/17/2021 02:10:35",
      "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> Congratulations Tucker Arrants for becoming Competition expert and also for Solo Silver . Great Writeup </p>",
      "rawMarkdown": "tuckerarrants Congratulations Tucker Arrants for becoming Competition expert and also for Solo Silver . Great Writeup",
      "votes": null
    },
    {
      "id": "1241355",
      "postDate": "03/17/2021 02:30:03",
      "content": "<p>At the column level I tried. Will be posting a discussion soon. Somehow at the column level the optimization doesn’t lead to good results. </p>",
      "rawMarkdown": "At the column level I tried. Will be posting a discussion soon. Somehow at the column level the optimization doesn’t lead to good results.",
      "votes": null
    },
    {
      "id": "1241361",
      "postDate": "03/17/2021 02:33:44",
      "content": "<p>Interesting…I look forward to reading your discussion post.</p>",
      "rawMarkdown": "Interesting...I look forward to reading your discussion post.",
      "votes": null
    },
    {
      "id": "1241363",
      "postDate": "03/17/2021 02:34:42",
      "content": "<p>Thank you Usha, I really appreciate it. </p>",
      "rawMarkdown": "Thank you Usha, I really appreciate it.",
      "votes": null
    },
    {
      "id": "1241422",
      "postDate": "03/17/2021 03:18:09",
      "content": "<p>Great write up, thank you. Can I ask how you were able to get the Z8 and Zbook from HP / NVIDIA 😛 <br>\nTrying to figure out how I can overcome this bottleneck!</p>",
      "rawMarkdown": "Great write up, thank you. Can I ask how you were able to get the Z8 and Zbook from HP / NVIDIA 😛 \nTrying to figure out how I can overcome this bottleneck!",
      "votes": null
    },
    {
      "id": "1241444",
      "postDate": "03/17/2021 03:33:17",
      "content": "<p>It was quite the surprise. They launched an ambassadorship program last year and I was asked to join, presumably based on some of my public Kaggle notebooks. (It was not something I applied for). </p>\n<p>My best advice is to continue your data science journey and to share what you learn: hopefully a similar opportunity will present itself to you. I think that such programs will become more prevalent as more companies recognize the importance of hardware in data science / machine learning / deep learning. </p>\n<p>I feel your pain, I have been trying to secure NVIDIA GPUs for my own personal use and they are very illusive these days. Hopefully this shortage is temporary and we can all get our hands on some solid GPUs in the near future.</p>",
      "rawMarkdown": "It was quite the surprise. They launched an ambassadorship program last year and I was asked to join, presumably based on some of my public Kaggle notebooks. (It was not something I applied for). \n\nMy best advice is to continue your data science journey and to share what you learn: hopefully a similar opportunity will present itself to you. I think that such programs will become more prevalent as more companies recognize the importance of hardware in data science / machine learning / deep learning. \n\nI feel your pain, I have been trying to secure NVIDIA GPUs for my own personal use and they are very illusive these days. Hopefully this shortage is temporary and we can all get our hands on some solid GPUs in the near future.",
      "votes": null
    },
    {
      "id": "1241518",
      "postDate": "03/17/2021 04:41:52",
      "content": "<p>Congratulations. Well explained  Great job )) </p>",
      "rawMarkdown": "Congratulations. Well explained  Great job ))",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1241214,
      "author_name": "amiiiney",
      "author_url": "",
      "post_date": "03/17/2021 00:30:35",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> and thanks for sharing this detailed writeup :)</p>\n<blockquote>\n  <p>I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights.</p>\n</blockquote>\n<p>How low was the image resolution you used in stage 1 and 2? Because with my models, whenever I increase the image size in step 3, both CV and LB decrease.<br>\nDid you tweak some parameters other than batch size when you increase the image size?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241219,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/17/2021 00:34:05",
          "content": "<p>Thank you Amin. The lowest resolution I used in stages 1 and 2 was <code>576</code>.  Sometimes I went as high as <code>768</code> for these stages, but only before I started reusing lower resolution weights to save time.</p>\n<p>I used slightly heavier augmentations each time I increased the image size.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241225,
      "author_name": "underwearfitting",
      "author_url": "",
      "post_date": "03/17/2021 00:39:27",
      "content": "<p>Great job Tucker. Thanks for sharing. Strong solo finish.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241227,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/17/2021 00:41:03",
          "content": "<p>No thank you, I could not have done it without all your contributions. Congratulations on 2nd place, very impressive indeed. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241232,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "03/17/2021 00:44:40",
      "content": "<p>Great write-up. When looking at your ensembling efforts did you assign weights at the macro level across all columns or did you ever try to optimize at the column level? I found that in my experiment when assigning weights my ensembling technique would basically just hard select whichever model had the best performance for a given column and any combination of models didn't really seem to yield very much. </p>\n<p>Never got around to applying that knowledge to the leaderboard, but I was surprised to see that behavior. </p>\n<p><img src=\"https://i.imgur.com/zhEn4Ow.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1241247,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/17/2021 00:55:26",
          "content": "<p>Thank you. Sadly, I only got around to assigning them at the macro level. Only at the last minute did I think to do it for each column. I may play around with the column-wise approach now that the competition is over to compare performance. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1241355,
          "author_name": "reighns",
          "author_url": "",
          "post_date": "03/17/2021 02:30:03",
          "content": "<p>At the column level I tried. Will be posting a discussion soon. Somehow at the column level the optimization doesn’t lead to good results. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1241361,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/17/2021 02:33:44",
          "content": "<p>Interesting…I look forward to reading your discussion post.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241336,
      "author_name": "usharengaraju",
      "author_url": "",
      "post_date": "03/17/2021 02:10:35",
      "content": "<p><a href=\"https://www.kaggle.com/tuckerarrants\" target=\"_blank\">@tuckerarrants</a> Congratulations Tucker Arrants for becoming Competition expert and also for Solo Silver . Great Writeup </p>",
      "votes": null,
      "replies": [
        {
          "id": 1241363,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/17/2021 02:34:42",
          "content": "<p>Thank you Usha, I really appreciate it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241422,
      "author_name": "reubenschmidt",
      "author_url": "",
      "post_date": "03/17/2021 03:18:09",
      "content": "<p>Great write up, thank you. Can I ask how you were able to get the Z8 and Zbook from HP / NVIDIA 😛 <br>\nTrying to figure out how I can overcome this bottleneck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1241444,
          "author_name": "tuckerarrants",
          "author_url": "",
          "post_date": "03/17/2021 03:33:17",
          "content": "<p>It was quite the surprise. They launched an ambassadorship program last year and I was asked to join, presumably based on some of my public Kaggle notebooks. (It was not something I applied for). </p>\n<p>My best advice is to continue your data science journey and to share what you learn: hopefully a similar opportunity will present itself to you. I think that such programs will become more prevalent as more companies recognize the importance of hardware in data science / machine learning / deep learning. </p>\n<p>I feel your pain, I have been trying to secure NVIDIA GPUs for my own personal use and they are very illusive these days. Hopefully this shortage is temporary and we can all get our hands on some solid GPUs in the near future.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1241518,
      "author_name": "ammarali32",
      "author_url": "",
      "post_date": "03/17/2021 04:41:52",
      "content": "<p>Congratulations. Well explained  Great job )) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1241203": "**60th Place Solution**\n\nThank you Kaggle and RANZCR for hosting this competition. It was the first competition I really immersed myself in and I have learned a tremendous amount. Thank you to all the competitors, especially those that so openly shared their results and techniques. Kaggle is a wonderful medium through which to learn data science, and it would not be the same without everyone sharing and communicating. \n\nI am not pleased with my final standing, but I believe that all Kagglers should share their procedure, irrespective of their final position, because it is highly unlikely that anyone else did exactly what you did, so there is always something to learn from reading other's write ups. Here it goes.\n\n**Image resolutions**\n\nMy approach was a brute force approach and not particularly elegant. I began with [@yasufuminakama](https://www.kaggle.com/yasufuminakama)’s kernels on 3 stage training with annotations and changed several parameters like learning rate, batch size, training augmentations, and optimizer (AdamP showed better CV results than Adam). I also generated stratified folds by coupling [@sin](https://www.kaggle.com/underwearfitting)'s [kernel](https://www.kaggle.com/underwearfitting/how-to-properly-split-folds) with some [useful information](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/210064) about incorrect labels / inverted images shared by [@raddar](https://www.kaggle.com/raddar) and others. Once I had established a CV strategy, I started to experiment with different image resolutions, all with a ResNet200D. I tried image sizes from `512` to `736` in increments of `16` (all with a batch size of `32`). I saw an increase in AUC + decrease in BCE loss each time I bumped up the image resolution. I validated all models against my first fold to save time (more on that later). \n\n[There was some discussion](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217296) as to whether or not this multi stage approach with teacher-student training on annotated images actually gave any performance boost. I continued to use it anyways for one reason: you can use the stage 2 weights to lower training time. I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights. In other words, instead of repeating all 3 stages anytime I changed image resolutions / other parameters, I simply reused stage 2 weights and skipped to stage 3. This is not ideal, but it saved me a ton of training time and allowed me to run more experiments. I figured it would work seeing as we frequently fine tune models that have been trained on ImageNet with different image resolutions than the ones they were initially trained on. \n\n**Preprocessing**\n\nI did minimal preprocessing. [It was mentioned](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/224146) that some images have black borders around them. I removed them during inference and saw an increase in score. I hesitated to remove them from training as these black boxes could serve as a form of regularization, but decided to remove them for the last couple models I trained. I figured if they helped with inference, they would help with training, but I did not run any real experiments to verify this.\n\n**Forward Selection**\n\nI then created a fold prediction correlation heatmap for these different image sizes and saw a surprisingly low correlation between the ResNet200D model when trained on different image sizes (0.85 - 0.91). This showed me that one could get good model diversity by simply changing image resolutions, which got me thinking about ensembling. I remembered reading [@chrisdeotte](https://www.kaggle.com/cdeotte)’s [kernel on forward selection](https://www.kaggle.com/cdeotte/forward-selection-oof-ensemble-0-942-private) from the Melanoma competition and decided to try it. I saw nice improvements when doing so and decided to experiment with different model architectures as well. I trained a SEResNet152D on `672` and `720` image resolutions and checked the fold prediction correlations between them with my previous ResNet200D models: the correlation was even lower, so I threw them into forward selection with my other ResNet200D models. \n\nThe above procedure got me to CV `0.966` and public leaderboard `0.968`. I then experimented with RegNetYs, EfficientNets, and ResNeSts. RegNetYs took too long to train and EfficientNets were never able to reach the same CV score as the ResNet200D / SEResNet152D. Only the ResNeSt’s got similar scores, but simply took too long to train. I had reached a bit of an impasse and was not sure how to proceed.\n\n**Pseudo labels**\n\nI then saw [this post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/221808) on duplicate images between the NIH Chest XRays dataset and the training / public test images given to use in this competition. Seeing as I already had an ensemble of models with good diversity, I was confident I could generate fairly high quality pseudo labels from this external dataset. (Note that you must generate pseudo labels from models that have not seen the fold you are validating against otherwise you get subtle CV leakage). I first tried to remove duplicates using CNN embeddings and RAPIDS NearestNeighbors, but this approach still left around 500 duplicate images, so I combined this approach with ImageHash to remove the remaining duplicates. I then started generating pseudo labels. I decided to use soft pseudo labels where at least one catheter class is present (any column with prediction above `0.7`) and all other classes without any catheters (remaining columns with predictions below `0.3`). I added Gaussian noise to the pseudo labels in an attempt to reduce confirmation bias.\n\nInstead of re-training the model jointly with concatenated labeled + pseudo labels, I decided to introduce another training stage in which I train on only NIH pseudo labels. The final stage is then fine tuning on labeled training data. I also decided to train the last few epochs without any data augmentation, based on [another discussion post](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/212532), and saw further improvements. I carried out this procedure for my top-3 highest scoring models and threw them back into my forward selection algorithm to compare CV and public leaderboard score. It indeed increased. \n\nI am still not sure if my CV was entirely leak free. I think many teams that used external data had some form of leakage. For this reason, I only included 2 NIH pseudo label pretrained models in the folds I forward selected against. Had I invested more time in examining these duplicate images found by RAPIDS NN and ImageHash, I think I would have had a more robust CV and better public / private score. \n\n**Training / validation folds**\n\nUp to this point, I had only validated against the first fold, so in some sense I was ‘overfitting’ my own validation set. This is not ideal; it would be much better to validate each model against all folds and use the full OOF to ensemble, as opposed to single fold predictions. That being said, I joined this competition fairly late so I needed to save time somehow while simultaneously running many experiments. I also trusted my fold splits and was not overly concerned with validating against only one fold. However, I did want my final predictions to encompass all training data, so I repeated what I did on the first fold with the second fold, so that my models (in total) had been trained on all data folds. \n\nAt the end, my approach was:\n\n* Train teacher on annotated samples  (light augmentations for 4 epochs) \n* Transfer this information to student  (light augmentations for 6-7 epochs)\n* Use these student weights as a starting point for NIH pseudo labeling training (light augmentations for 7-8 epochs)\n* Train on full labeled data (heavy augmentation for 7-8 epochs)\n* Train on full labeled data (no augmentation for 1-2 epochs)\n* Repeat on different image sizes and model architectures and forward select based on validation fold predictions\n\nI had around 10 models validated against fold 1 and 10 validated against fold 2. I forward selected each fold separately and averaged their predictions. Seeing as I never got around to full KFold training, I thought that including public models would add some beneficial diversity, so I added [@sin](https://www.kaggle.com/underwearfitting)’s public 5Fold ResNet200D models to inference. My final submission was a convex combination of my own models weighted `0.85` and [sin](https://www.kaggle.com/underwearfitting)’s 5 models weighted at `0.15`.\n\nI chose my best public leaderboard score for my first submission (seeing as I trusted CV and public correlation) and chose my best CV score for my second submission. They were both `0.971` on the private leaderboard and not surprisingly, my best CV score was also my best private leaderboard score. \n\n**What didn’t really work**\n\n* TTA beyond simple horizontal flipping only showed improvements when the number of steps was fairly high (>5) and this took too long to infer with\n* Models other than ResNet200D and SEResNet152D (ResNeSt200E worked as well but took too long to train)\n* Image resolutions below `512`\n\n**What I wanted to test, but didn’t**\n\n* Experiment with changing image resolutions during training as a means of regularization\n* Using shape priors and a segmentation model to create features (categorize catheter overlaps with HoG descriptors and intensity histograms)\n* Use masks and segmentation model outputs to categorize distance between various anatomical features and the catheters to create more features\n* Upsample images without any catheters present with NIH Chest XRay images and compare performance\n* Experiment with label smoothing / special loss functions\n* Use MixUp on pseudo labels to further reduce confirmation bias\n\n**Acknowledgements**\n\nI would first like to thank [@yasufuminakama](https://www.kaggle.com/yasufuminakama) and [@hengck23](https://www.kaggle.com/hengck23) for sharing their approach on how to use annotated images to increase CV scores. Without their generous contributions in this competition, I would not have made much progress during these past couple months. \n\nI also want to thank [@chrisdeotte](https://www.kaggle.com/cdeotte) for his wonderful notebooks on forward selection, [pseudo-labeling](https://www.kaggle.com/cdeotte/pseudo-labeling-qda-0-969), and [finding duplicate images with RAPIDS](https://www.kaggle.com/cdeotte/rapids-cuml-knn-find-duplicates). Forward selection was very useful for generating high quality pseudo labels and he inspired me to experiment with ensembling models trained on different image sizes. Thank you Chris.\n\nAdditionally, I would like to thank [@raddar](https://www.kaggle.com/raddar) for sharing domain knowledge so openly during this competition and [@sin](https://www.kaggle.com/underwearfitting) for sharing his use of heavy augmentations and his cross-validation strategy. In coupling their ideas, I was able to generate a robust CV strategy that informed all my decisions during this competition.\n\nLastly, I want to thank Z by HP & NVIDIA for providing me with a Z8 workstation and ZBook Studio laptop. Without these GPUs, I would not have been able to experiment with these larger networks and different image resolutions so quickly. I ran smaller experiments / debugged my code on the laptop and ran the main experiments on the workstation. GPU memory was the bottleneck in this competition so having powerful local GPUs was integral to my solution. (RTX Quadro 8000 and RTX Quadro 5000 mobile, for reference). When it comes to deep learning and GPUs, bigger seems to be better. That being said, I wish I found more elegant approaches to circumvent this problem, but I am still learning. Next time.",
    "1241214": "Congratulations @tuckerarrants and thanks for sharing this detailed writeup :)\n\n>I trained stages 1 and 2 on a lower image resolution and then trained many different stage 3 models of various image resolutions from these weights.\n\nHow low was the image resolution you used in stage 1 and 2? Because with my models, whenever I increase the image size in step 3, both CV and LB decrease.\nDid you tweak some parameters other than batch size when you increase the image size?",
    "1241219": "Thank you Amin. The lowest resolution I used in stages 1 and 2 was `576`.  Sometimes I went as high as `768` for these stages, but only before I started reusing lower resolution weights to save time.\n\nI used slightly heavier augmentations each time I increased the image size.",
    "1241225": "Great job Tucker. Thanks for sharing. Strong solo finish.",
    "1241227": "No thank you, I could not have done it without all your contributions. Congratulations on 2nd place, very impressive indeed.",
    "1241232": "Great write-up. When looking at your ensembling efforts did you assign weights at the macro level across all columns or did you ever try to optimize at the column level? I found that in my experiment when assigning weights my ensembling technique would basically just hard select whichever model had the best performance for a given column and any combination of models didn't really seem to yield very much. \n\nNever got around to applying that knowledge to the leaderboard, but I was surprised to see that behavior. \n\n![](https://i.imgur.com/zhEn4Ow.png)",
    "1241247": "Thank you. Sadly, I only got around to assigning them at the macro level. Only at the last minute did I think to do it for each column. I may play around with the column-wise approach now that the competition is over to compare performance.",
    "1241336": "tuckerarrants Congratulations Tucker Arrants for becoming Competition expert and also for Solo Silver . Great Writeup",
    "1241355": "At the column level I tried. Will be posting a discussion soon. Somehow at the column level the optimization doesn’t lead to good results.",
    "1241361": "Interesting...I look forward to reading your discussion post.",
    "1241363": "Thank you Usha, I really appreciate it.",
    "1241422": "Great write up, thank you. Can I ask how you were able to get the Z8 and Zbook from HP / NVIDIA 😛 \nTrying to figure out how I can overcome this bottleneck!",
    "1241444": "It was quite the surprise. They launched an ambassadorship program last year and I was asked to join, presumably based on some of my public Kaggle notebooks. (It was not something I applied for). \n\nMy best advice is to continue your data science journey and to share what you learn: hopefully a similar opportunity will present itself to you. I think that such programs will become more prevalent as more companies recognize the importance of hardware in data science / machine learning / deep learning. \n\nI feel your pain, I have been trying to secure NVIDIA GPUs for my own personal use and they are very illusive these days. Hopefully this shortage is temporary and we can all get our hands on some solid GPUs in the near future.",
    "1241518": "Congratulations. Well explained  Great job ))"
  },
  "source": "meta"
}