{
  "id": 175324,
  "title": "[2nd place] Solution Overview",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/175324",
  "author_name": "Ian Pan",
  "post_date": "2020-08-18T00:19:51.980000",
  "votes": 225,
  "comment_count": 79,
  "views": 0,
  "content": "<p>Code available here: <a href=\"https://github.com/i-pan/kaggle-melanoma\" target=\"_blank\">https://github.com/i-pan/kaggle-melanoma</a></p>\n<p>Wow. Did not expect this. But super excited to finally get my solo gold medal and complete the journey to Grandmaster. I'm grateful to benefit from the shakeup this time after suffering in PANDA. Thank you to all the organizers, and congratulations to all the other winners and participants. </p>\n<p>This year, I accomplished 2 milestones: 1) I graduated medical school and became a doctor; and 2) I became a Kaggle Competitions Grandmaster. I actually just started my intern year as a doctor working 70-80 hours/week so did not have too much time to dedicate to this competition. The key was having a pipeline that allowed me to quickly iterate on experiments so I could start them in the morning, go to work, then analyze the results when I came back. Long story short, to all those who are starting out, keep going. Put in the time and effort. Compete, read winners' solutions, and learn, over and over again. </p>\n<p>I have to thank my teammates in previous competitions <a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a>, <a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a>, and <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a>. And I also have to thank the Kaggle community as I have learned so much from participating.</p>\n<p>Also huge thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for all the insights he shared. I used his triple stratified dataset to split my data for CV. </p>\n<p><strong>Environment</strong><br>\nPyTorch 1.6 with automatic mixed precision<br>\n4x NVIDIA GTX 1080 Ti 11GB provided by HOSTKEY (<a href=\"http://www.hostkey.com\" target=\"_blank\">www.hostkey.com</a>) as part of their Kaggle competitions grant<br>\n4x Quadro RTX 6000 24GB </p>\n<p><strong>Batch Size</strong><br>\nIn the beginning, I was training models on GTX 1080 Ti GPUs provided by HOSTKEY. I wasn’t able to get multi-GPU training to work, so I was just experimenting with batch sizes of 8-16 on a single 11GB 1080 Ti. My CV scores were not that high, usually a 5-fold average of around 0.92, with similar results on LB. The TPU kernels were doing so well (on public LB, at least), that I thought it was in part due to increased batch size. At that point, I switched over to my 4x Quadro RTX6000 24GB setup to leverage more GPU memory and multi-GPU training. I aimed for a batch size of 64 while maximizing image resolution for a particular backbone (and if that would not fit, I settled for BS 32 with gradient accumulation 2).</p>\n<p><strong>Backbones &amp; Image Resolution</strong><br>\nI tried several backbones in the EfficientNet, SE-ResNeXt, and ResNeSt families. I also tried BiT-ResNet (recently released by Google). EfficientNet performed better on CV so I decided to stick with EfficientNets for the remainder of the competition. I used the 1080 Ti GPUs to experiment with different backbones, all other hyperparameters held constant. Some of you asked whether this necessarily transfers when I switch over to a larger GPU and increase the batch size/change the image resolution. I did briefly compare backbones on the larger GPU and the results seemed to be consistent (EfficientNet &gt; all others). </p>\n<p>I experimented with backbones of different sizes from pruned EfficientNet-B3 to EfficientNet-B8, using the implementation from <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a>. For smaller backbones, I tried larger resolutions (up to 1024 x 1024) and for EfficientNet-B8 I went down to 384 x 384. The image size for each backbone was selected so that I could use batch size of 64 (16/GPU) during training. I found that smaller EfficientNets at higher resolutions were not as good. The best models for me were EfficientNet-B6 (initialized with noisy student weights) and EfficientNet-B7, at image resolutions of 512 x 512 and 640 x 640, respectively, so I only used these moving forward.</p>\n<p><strong>Base Model</strong><br>\nMy base model was your standard convolutional neural network backbone with linear classification head. I used generalized mean pooling with trainable parameter p (not sure if this was any better than average or max, as I just stuck with GeM from the beginning). I also used multisample dropout following <a href=\"https://www.kaggle.com/qishenha\" target=\"_blank\">@qishenha</a>’s implementation in one of this winning NLP solutions. </p>\n<p><strong>From 2 Class to 3 Class</strong><br>\nI felt that more granular classes would result in better feature representations that could help improve performance. The majority of melanomas are dark (exception being rare amelanotic melanomas), so differentiating them from benign nevi is probably the most challenging task. The 2019 data all had auxiliary diagnoses, including nevi, whereas a large fraction of the 2020 data was unknown. </p>\n<p>I trained a model on 2019 data only using the diagnosis as the target. Then, I applied this model to the 2020 data. My main focus was on labeling the unknowns as nevus or not nevus, as I know they are not melanoma. To find the threshold at which I would label an image as nevus, I used the 5th percentile of the 2019 model’s predictions on the 2020 data which had a known label of nevus.</p>\n<p>Now, all of the 2019 and 2020 data has a label of other, benign nevus, or melanoma, and I trained my model on these 3 classes using vanilla cross-entropy loss. I did not try label smoothing.</p>\n<p><strong>Upsampling</strong><br>\nThere was a lot of discussion over whether or not to upsample malignant images or not. I did upsample malignant images for 2020 data. Because I used 2019 data and the percentage of melanomas in that dataset was much higher, I wanted to make sure that the 2019 melanomas did not overwhelm the 2020 melanomas. To that end, I upsampled the 2020 melanomas 7 times so that there was about an equal number of melanomas from both datasets.</p>\n<p><strong>Training</strong><br>\nI used AdamW optimizer and cosine annealing with warm restarts scheduler, initial learning rate 3.0E-4. I did 3 snapshots, 2 epochs each for EfficientNet-B6 and 3 epochs each for EfficientNet-B7. I noticed that it did not take long for the models to start overfitting. I found that this gave better results than using one cycle, so I stuck with it. Out of the 3 snapshots, I just took the one that did the best on that validation fold. Every experiment I ran, I did 5-fold CV using Chris Deotte’s triple stratified data splits. Single validation folds were not stable for me, so in order to really understand my model performance and the effects of my adjustments I had to look at 5-fold CV average. I only validated on 2020 data. </p>\n<p>In the beginning, I was using metadata by using embeddings for age, sex, and location. Each embedding was mapped to a 32-D vector and concatenated to the final feature vector before input into the linear classification layer. I did not want to spend too much time tuning this because I was afraid I would overfit to the distribution of the training set. I just used mean/mode imputation for missing values. </p>\n<p>It wasn’t until the last several days of the competition that I decided to train models without metadata, so I could be eligible for the without context special prize. It turns out that these models were actually my highest scoring private LB solutions!</p>\n<p><strong>Augmentations</strong><br>\nI knew that augmentation would be important given the small percentage of melanomas in the 2020 data. I used the RandAugment strategy, implemented here: <a href=\"https://github.com/ildoonet/pytorch-randaugment\" target=\"_blank\">https://github.com/ildoonet/pytorch-randaugment</a>. I used N=3 augmentations with magnitude M/30 where M was sampled from a Poisson distribution with mean 12 for extra stochasticity. For those unfamiliar with RandAugment, M is essentially the “hardness” of the augmentation (angle for rotation, % zoom, gamma for contrast adjustment, etc.). For augmentations like flips, M is not relevant. I tried other augmentations such as mixup, cutmix, and grid mask, but those did not help.</p>\n<p>I also used square cropping during training and inference. During training, a square was randomly cropped from the image if it was rectangular (otherwise, the entire image was used), where the length of the square image was the size of the shortest side (i.e., 768x512 would be cropped to 512x512). During inference, I spaced out 10 square crops as TTA and took the average as the final prediction (again, unless the image was already square - then no TTA was applied). I found that this gave me better results than rectangular crops or using the whole image. </p>\n<p><strong>Pseudolabeling</strong><br>\nPseudolabeling was key to my solution. Given the limited number of 2020 melanomas, I felt that pseudolabeling would help increase performance. 2019 melanomas were helpful but still different from 2020 melanomas. I took my 5-fold EfficientNet-B6 model, trained without metadata, and obtained soft pseudolabels (3 classes) for the test set. When combining the test data with the training data (2019+2020), I upsampled images with melanoma prediction &gt; 0.5 7 times (same factor as I did for 2020 training data). I used <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>'s implementation (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733</a>) of cross entropy in PyTorch (without label smoothing) so I could use soft pseudolabels. </p>\n<p><strong>CV vs LB</strong><br>\nI knew early on that it would be easy to fit public LB, given the small number of melanomas that would be in the public test. At the same time, the CV for my different experiments was much tighter than the LB, so I was nervous to fully trust CV as there may have been differences between training and test data. With that in mind, I favored solutions that had reasonably high CV and LB. </p>\n<p>There is a fair amount of luck that goes into picking the right solution, but you should be able to justify to yourself why you are picking a certain solution over another (going by CV score, LB score, some combination of CV/LB, or some hypotheses about the private test set that would favor one solution over another). </p>\n<p>My 2nd place solution was an ensemble of 3 5-fold models:</p>\n<ul>\n<li>EfficientNet-B6, 512x512, BS64, no metadata (CV 0.9336 / public 0.9534)</li>\n<li>EfficientNet-B7, 640x640, BS32, gradient accumulation 2, no metadata (CV 0.9389 / public 0.9525)</li>\n<li>Model 1, trained on combined training and pseudolabeled test data (CV 0.9438 / public 0.9493)</li>\n</ul>\n<p>Note that the CV score does not account for the 5-fold blend effect. </p>\n<p>My highest scoring private LB solution was actually model 3 alone, which I did not select.</p>\n<p>I trained other models with metadata, but best private LB score was 0.945 (public LB 0.959), with similar CV. </p>",
  "messages": [
    {
      "id": 974426,
      "postDate": "2020-08-18T00:19:51.980Z",
      "content": "<p>Code available here: <a href=\"https://github.com/i-pan/kaggle-melanoma\" target=\"_blank\">https://github.com/i-pan/kaggle-melanoma</a></p>\n<p>Wow. Did not expect this. But super excited to finally get my solo gold medal and complete the journey to Grandmaster. I'm grateful to benefit from the shakeup this time after suffering in PANDA. Thank you to all the organizers, and congratulations to all the other winners and participants. </p>\n<p>This year, I accomplished 2 milestones: 1) I graduated medical school and became a doctor; and 2) I became a Kaggle Competitions Grandmaster. I actually just started my intern year as a doctor working 70-80 hours/week so did not have too much time to dedicate to this competition. The key was having a pipeline that allowed me to quickly iterate on experiments so I could start them in the morning, go to work, then analyze the results when I came back. Long story short, to all those who are starting out, keep going. Put in the time and effort. Compete, read winners' solutions, and learn, over and over again. </p>\n<p>I have to thank my teammates in previous competitions <a href=\"https://www.kaggle.com/felipekitamura\" target=\"_blank\">@felipekitamura</a>, <a href=\"https://www.kaggle.com/alexandrecc\" target=\"_blank\">@alexandrecc</a>, and <a href=\"https://www.kaggle.com/jamesphoward\" target=\"_blank\">@jamesphoward</a>. And I also have to thank the Kaggle community as I have learned so much from participating.</p>\n<p>Also huge thanks to <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> for all the insights he shared. I used his triple stratified dataset to split my data for CV. </p>\n<p><strong>Environment</strong><br>\nPyTorch 1.6 with automatic mixed precision<br>\n4x NVIDIA GTX 1080 Ti 11GB provided by HOSTKEY (<a href=\"http://www.hostkey.com\" target=\"_blank\">www.hostkey.com</a>) as part of their Kaggle competitions grant<br>\n4x Quadro RTX 6000 24GB </p>\n<p><strong>Batch Size</strong><br>\nIn the beginning, I was training models on GTX 1080 Ti GPUs provided by HOSTKEY. I wasn’t able to get multi-GPU training to work, so I was just experimenting with batch sizes of 8-16 on a single 11GB 1080 Ti. My CV scores were not that high, usually a 5-fold average of around 0.92, with similar results on LB. The TPU kernels were doing so well (on public LB, at least), that I thought it was in part due to increased batch size. At that point, I switched over to my 4x Quadro RTX6000 24GB setup to leverage more GPU memory and multi-GPU training. I aimed for a batch size of 64 while maximizing image resolution for a particular backbone (and if that would not fit, I settled for BS 32 with gradient accumulation 2).</p>\n<p><strong>Backbones &amp; Image Resolution</strong><br>\nI tried several backbones in the EfficientNet, SE-ResNeXt, and ResNeSt families. I also tried BiT-ResNet (recently released by Google). EfficientNet performed better on CV so I decided to stick with EfficientNets for the remainder of the competition. I used the 1080 Ti GPUs to experiment with different backbones, all other hyperparameters held constant. Some of you asked whether this necessarily transfers when I switch over to a larger GPU and increase the batch size/change the image resolution. I did briefly compare backbones on the larger GPU and the results seemed to be consistent (EfficientNet &gt; all others). </p>\n<p>I experimented with backbones of different sizes from pruned EfficientNet-B3 to EfficientNet-B8, using the implementation from <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models</a>. For smaller backbones, I tried larger resolutions (up to 1024 x 1024) and for EfficientNet-B8 I went down to 384 x 384. The image size for each backbone was selected so that I could use batch size of 64 (16/GPU) during training. I found that smaller EfficientNets at higher resolutions were not as good. The best models for me were EfficientNet-B6 (initialized with noisy student weights) and EfficientNet-B7, at image resolutions of 512 x 512 and 640 x 640, respectively, so I only used these moving forward.</p>\n<p><strong>Base Model</strong><br>\nMy base model was your standard convolutional neural network backbone with linear classification head. I used generalized mean pooling with trainable parameter p (not sure if this was any better than average or max, as I just stuck with GeM from the beginning). I also used multisample dropout following <a href=\"https://www.kaggle.com/qishenha\" target=\"_blank\">@qishenha</a>’s implementation in one of this winning NLP solutions. </p>\n<p><strong>From 2 Class to 3 Class</strong><br>\nI felt that more granular classes would result in better feature representations that could help improve performance. The majority of melanomas are dark (exception being rare amelanotic melanomas), so differentiating them from benign nevi is probably the most challenging task. The 2019 data all had auxiliary diagnoses, including nevi, whereas a large fraction of the 2020 data was unknown. </p>\n<p>I trained a model on 2019 data only using the diagnosis as the target. Then, I applied this model to the 2020 data. My main focus was on labeling the unknowns as nevus or not nevus, as I know they are not melanoma. To find the threshold at which I would label an image as nevus, I used the 5th percentile of the 2019 model’s predictions on the 2020 data which had a known label of nevus.</p>\n<p>Now, all of the 2019 and 2020 data has a label of other, benign nevus, or melanoma, and I trained my model on these 3 classes using vanilla cross-entropy loss. I did not try label smoothing.</p>\n<p><strong>Upsampling</strong><br>\nThere was a lot of discussion over whether or not to upsample malignant images or not. I did upsample malignant images for 2020 data. Because I used 2019 data and the percentage of melanomas in that dataset was much higher, I wanted to make sure that the 2019 melanomas did not overwhelm the 2020 melanomas. To that end, I upsampled the 2020 melanomas 7 times so that there was about an equal number of melanomas from both datasets.</p>\n<p><strong>Training</strong><br>\nI used AdamW optimizer and cosine annealing with warm restarts scheduler, initial learning rate 3.0E-4. I did 3 snapshots, 2 epochs each for EfficientNet-B6 and 3 epochs each for EfficientNet-B7. I noticed that it did not take long for the models to start overfitting. I found that this gave better results than using one cycle, so I stuck with it. Out of the 3 snapshots, I just took the one that did the best on that validation fold. Every experiment I ran, I did 5-fold CV using Chris Deotte’s triple stratified data splits. Single validation folds were not stable for me, so in order to really understand my model performance and the effects of my adjustments I had to look at 5-fold CV average. I only validated on 2020 data. </p>\n<p>In the beginning, I was using metadata by using embeddings for age, sex, and location. Each embedding was mapped to a 32-D vector and concatenated to the final feature vector before input into the linear classification layer. I did not want to spend too much time tuning this because I was afraid I would overfit to the distribution of the training set. I just used mean/mode imputation for missing values. </p>\n<p>It wasn’t until the last several days of the competition that I decided to train models without metadata, so I could be eligible for the without context special prize. It turns out that these models were actually my highest scoring private LB solutions!</p>\n<p><strong>Augmentations</strong><br>\nI knew that augmentation would be important given the small percentage of melanomas in the 2020 data. I used the RandAugment strategy, implemented here: <a href=\"https://github.com/ildoonet/pytorch-randaugment\" target=\"_blank\">https://github.com/ildoonet/pytorch-randaugment</a>. I used N=3 augmentations with magnitude M/30 where M was sampled from a Poisson distribution with mean 12 for extra stochasticity. For those unfamiliar with RandAugment, M is essentially the “hardness” of the augmentation (angle for rotation, % zoom, gamma for contrast adjustment, etc.). For augmentations like flips, M is not relevant. I tried other augmentations such as mixup, cutmix, and grid mask, but those did not help.</p>\n<p>I also used square cropping during training and inference. During training, a square was randomly cropped from the image if it was rectangular (otherwise, the entire image was used), where the length of the square image was the size of the shortest side (i.e., 768x512 would be cropped to 512x512). During inference, I spaced out 10 square crops as TTA and took the average as the final prediction (again, unless the image was already square - then no TTA was applied). I found that this gave me better results than rectangular crops or using the whole image. </p>\n<p><strong>Pseudolabeling</strong><br>\nPseudolabeling was key to my solution. Given the limited number of 2020 melanomas, I felt that pseudolabeling would help increase performance. 2019 melanomas were helpful but still different from 2020 melanomas. I took my 5-fold EfficientNet-B6 model, trained without metadata, and obtained soft pseudolabels (3 classes) for the test set. When combining the test data with the training data (2019+2020), I upsampled images with melanoma prediction &gt; 0.5 7 times (same factor as I did for 2020 training data). I used <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a>'s implementation (<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733</a>) of cross entropy in PyTorch (without label smoothing) so I could use soft pseudolabels. </p>\n<p><strong>CV vs LB</strong><br>\nI knew early on that it would be easy to fit public LB, given the small number of melanomas that would be in the public test. At the same time, the CV for my different experiments was much tighter than the LB, so I was nervous to fully trust CV as there may have been differences between training and test data. With that in mind, I favored solutions that had reasonably high CV and LB. </p>\n<p>There is a fair amount of luck that goes into picking the right solution, but you should be able to justify to yourself why you are picking a certain solution over another (going by CV score, LB score, some combination of CV/LB, or some hypotheses about the private test set that would favor one solution over another). </p>\n<p>My 2nd place solution was an ensemble of 3 5-fold models:</p>\n<ul>\n<li>EfficientNet-B6, 512x512, BS64, no metadata (CV 0.9336 / public 0.9534)</li>\n<li>EfficientNet-B7, 640x640, BS32, gradient accumulation 2, no metadata (CV 0.9389 / public 0.9525)</li>\n<li>Model 1, trained on combined training and pseudolabeled test data (CV 0.9438 / public 0.9493)</li>\n</ul>\n<p>Note that the CV score does not account for the 5-fold blend effect. </p>\n<p>My highest scoring private LB solution was actually model 3 alone, which I did not select.</p>\n<p>I trained other models with metadata, but best private LB score was 0.945 (public LB 0.959), with similar CV. </p>",
      "rawMarkdown": "Code available here: https://github.com/i-pan/kaggle-melanoma\n\nWow. Did not expect this. But super excited to finally get my solo gold medal and complete the journey to Grandmaster. I'm grateful to benefit from the shakeup this time after suffering in PANDA. Thank you to all the organizers, and congratulations to all the other winners and participants. \n\nThis year, I accomplished 2 milestones: 1) I graduated medical school and became a doctor; and 2) I became a Kaggle Competitions Grandmaster. I actually just started my intern year as a doctor working 70-80 hours/week so did not have too much time to dedicate to this competition. The key was having a pipeline that allowed me to quickly iterate on experiments so I could start them in the morning, go to work, then analyze the results when I came back. Long story short, to all those who are starting out, keep going. Put in the time and effort. Compete, read winners' solutions, and learn, over and over again. \n\nI have to thank my teammates in previous competitions @felipekitamura, @alexandrecc, and @jamesphoward. And I also have to thank the Kaggle community as I have learned so much from participating.\n\nAlso huge thanks to @cdeotte for all the insights he shared. I used his triple stratified dataset to split my data for CV. \n\n**Environment**\nPyTorch 1.6 with automatic mixed precision\n4x NVIDIA GTX 1080 Ti 11GB provided by HOSTKEY (www.hostkey.com) as part of their Kaggle competitions grant\n4x Quadro RTX 6000 24GB \n\n**Batch Size**\nIn the beginning, I was training models on GTX 1080 Ti GPUs provided by HOSTKEY. I wasn’t able to get multi-GPU training to work, so I was just experimenting with batch sizes of 8-16 on a single 11GB 1080 Ti. My CV scores were not that high, usually a 5-fold average of around 0.92, with similar results on LB. The TPU kernels were doing so well (on public LB, at least), that I thought it was in part due to increased batch size. At that point, I switched over to my 4x Quadro RTX6000 24GB setup to leverage more GPU memory and multi-GPU training. I aimed for a batch size of 64 while maximizing image resolution for a particular backbone (and if that would not fit, I settled for BS 32 with gradient accumulation 2).\n\n**Backbones & Image Resolution**\nI tried several backbones in the EfficientNet, SE-ResNeXt, and ResNeSt families. I also tried BiT-ResNet (recently released by Google). EfficientNet performed better on CV so I decided to stick with EfficientNets for the remainder of the competition. I used the 1080 Ti GPUs to experiment with different backbones, all other hyperparameters held constant. Some of you asked whether this necessarily transfers when I switch over to a larger GPU and increase the batch size/change the image resolution. I did briefly compare backbones on the larger GPU and the results seemed to be consistent (EfficientNet > all others). \n\nI experimented with backbones of different sizes from pruned EfficientNet-B3 to EfficientNet-B8, using the implementation from https://github.com/rwightman/pytorch-image-models. For smaller backbones, I tried larger resolutions (up to 1024 x 1024) and for EfficientNet-B8 I went down to 384 x 384. The image size for each backbone was selected so that I could use batch size of 64 (16/GPU) during training. I found that smaller EfficientNets at higher resolutions were not as good. The best models for me were EfficientNet-B6 (initialized with noisy student weights) and EfficientNet-B7, at image resolutions of 512 x 512 and 640 x 640, respectively, so I only used these moving forward.\n\n**Base Model**\nMy base model was your standard convolutional neural network backbone with linear classification head. I used generalized mean pooling with trainable parameter p (not sure if this was any better than average or max, as I just stuck with GeM from the beginning). I also used multisample dropout following @qishenha’s implementation in one of this winning NLP solutions. \n\n**From 2 Class to 3 Class**\nI felt that more granular classes would result in better feature representations that could help improve performance. The majority of melanomas are dark (exception being rare amelanotic melanomas), so differentiating them from benign nevi is probably the most challenging task. The 2019 data all had auxiliary diagnoses, including nevi, whereas a large fraction of the 2020 data was unknown. \n\nI trained a model on 2019 data only using the diagnosis as the target. Then, I applied this model to the 2020 data. My main focus was on labeling the unknowns as nevus or not nevus, as I know they are not melanoma. To find the threshold at which I would label an image as nevus, I used the 5th percentile of the 2019 model’s predictions on the 2020 data which had a known label of nevus.\n\nNow, all of the 2019 and 2020 data has a label of other, benign nevus, or melanoma, and I trained my model on these 3 classes using vanilla cross-entropy loss. I did not try label smoothing.\n\n**Upsampling**\nThere was a lot of discussion over whether or not to upsample malignant images or not. I did upsample malignant images for 2020 data. Because I used 2019 data and the percentage of melanomas in that dataset was much higher, I wanted to make sure that the 2019 melanomas did not overwhelm the 2020 melanomas. To that end, I upsampled the 2020 melanomas 7 times so that there was about an equal number of melanomas from both datasets.\n\n**Training**\nI used AdamW optimizer and cosine annealing with warm restarts scheduler, initial learning rate 3.0E-4. I did 3 snapshots, 2 epochs each for EfficientNet-B6 and 3 epochs each for EfficientNet-B7. I noticed that it did not take long for the models to start overfitting. I found that this gave better results than using one cycle, so I stuck with it. Out of the 3 snapshots, I just took the one that did the best on that validation fold. Every experiment I ran, I did 5-fold CV using Chris Deotte’s triple stratified data splits. Single validation folds were not stable for me, so in order to really understand my model performance and the effects of my adjustments I had to look at 5-fold CV average. I only validated on 2020 data. \n\nIn the beginning, I was using metadata by using embeddings for age, sex, and location. Each embedding was mapped to a 32-D vector and concatenated to the final feature vector before input into the linear classification layer. I did not want to spend too much time tuning this because I was afraid I would overfit to the distribution of the training set. I just used mean/mode imputation for missing values. \n\nIt wasn’t until the last several days of the competition that I decided to train models without metadata, so I could be eligible for the without context special prize. It turns out that these models were actually my highest scoring private LB solutions!\n\n**Augmentations**\nI knew that augmentation would be important given the small percentage of melanomas in the 2020 data. I used the RandAugment strategy, implemented here: https://github.com/ildoonet/pytorch-randaugment. I used N=3 augmentations with magnitude M/30 where M was sampled from a Poisson distribution with mean 12 for extra stochasticity. For those unfamiliar with RandAugment, M is essentially the “hardness” of the augmentation (angle for rotation, % zoom, gamma for contrast adjustment, etc.). For augmentations like flips, M is not relevant. I tried other augmentations such as mixup, cutmix, and grid mask, but those did not help.\n\nI also used square cropping during training and inference. During training, a square was randomly cropped from the image if it was rectangular (otherwise, the entire image was used), where the length of the square image was the size of the shortest side (i.e., 768x512 would be cropped to 512x512). During inference, I spaced out 10 square crops as TTA and took the average as the final prediction (again, unless the image was already square - then no TTA was applied). I found that this gave me better results than rectangular crops or using the whole image. \n\n**Pseudolabeling**\nPseudolabeling was key to my solution. Given the limited number of 2020 melanomas, I felt that pseudolabeling would help increase performance. 2019 melanomas were helpful but still different from 2020 melanomas. I took my 5-fold EfficientNet-B6 model, trained without metadata, and obtained soft pseudolabels (3 classes) for the test set. When combining the test data with the training data (2019+2020), I upsampled images with melanoma prediction > 0.5 7 times (same factor as I did for 2020 training data). I used @cpmpml's implementation (https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733) of cross entropy in PyTorch (without label smoothing) so I could use soft pseudolabels. \n\n**CV vs LB**\nI knew early on that it would be easy to fit public LB, given the small number of melanomas that would be in the public test. At the same time, the CV for my different experiments was much tighter than the LB, so I was nervous to fully trust CV as there may have been differences between training and test data. With that in mind, I favored solutions that had reasonably high CV and LB. \n\nThere is a fair amount of luck that goes into picking the right solution, but you should be able to justify to yourself why you are picking a certain solution over another (going by CV score, LB score, some combination of CV/LB, or some hypotheses about the private test set that would favor one solution over another). \n\nMy 2nd place solution was an ensemble of 3 5-fold models:\n- EfficientNet-B6, 512x512, BS64, no metadata (CV 0.9336 / public 0.9534)\n- EfficientNet-B7, 640x640, BS32, gradient accumulation 2, no metadata (CV 0.9389 / public 0.9525)\n- Model 1, trained on combined training and pseudolabeled test data (CV 0.9438 / public 0.9493)\n\nNote that the CV score does not account for the 5-fold blend effect. \n\nMy highest scoring private LB solution was actually model 3 alone, which I did not select.\n\nI trained other models with metadata, but best private LB score was 0.945 (public LB 0.959), with similar CV. \n\n",
      "votes": 224
    },
    {
      "id": 980705,
      "postDate": "2020-08-21T19:36:47.260Z",
      "content": "<p>You can edit your title now ;)  Welcome to competitions GM club, well deserved!</p>",
      "rawMarkdown": "You can edit your title now ;)  Welcome to competitions GM club, well deserved!",
      "votes": 7,
      "replies": [
        {
          "id": 981141,
          "postDate": "2020-08-22T07:45:25.253Z",
          "content": "<p>speaking from his profile picture (and we all know profile pics show the true age), he might be the youngest GM ever</p>",
          "rawMarkdown": "speaking from his profile picture (and we all know profile pics show the true age), he might be the youngest GM ever",
          "votes": 9
        },
        {
          "id": 981250,
          "postDate": "2020-08-22T10:13:53.060Z",
          "content": "<p>Anti Dieter indeed.  You two should team at some point.</p>",
          "rawMarkdown": "Anti Dieter indeed.  You two should team at some point.",
          "votes": 3
        },
        {
          "id": 982341,
          "postDate": "2020-08-23T09:28:27.393Z",
          "content": "<p>He needs a cool hat though. ;)</p>",
          "rawMarkdown": "He needs a cool hat though. ;)",
          "votes": 2
        }
      ]
    },
    {
      "id": 974762,
      "postDate": "2020-08-18T03:33:37.440Z",
      "content": "<p>Congrats Ian on becoming Kaggle Grandmaster and achieving 2nd place money finish. </p>\n<p>It looks like Pseudo labeling worked well for you. How did you assign targets 0 and 1 based on model 1's predictions?</p>",
      "rawMarkdown": "Congrats Ian on becoming Kaggle Grandmaster and achieving 2nd place money finish. \n\nIt looks like Pseudo labeling worked well for you. How did you assign targets 0 and 1 based on model 1's predictions?",
      "votes": 5,
      "replies": [
        {
          "id": 975628,
          "postDate": "2020-08-18T11:43:38.587Z",
          "content": "<p>Same question here. </p>",
          "rawMarkdown": "Same question here. "
        },
        {
          "id": 976522,
          "postDate": "2020-08-18T23:28:21.767Z",
          "content": "<p>Thank you, Chris! I really appreciate all your discussion posts and notebooks, in this competition and others. </p>\n<p>I updated my post to detail the pseudolabeling more. I used soft pseudolabels (predicted probabilities) and a modified cross entropy loss. </p>",
          "rawMarkdown": "Thank you, Chris! I really appreciate all your discussion posts and notebooks, in this competition and others. \n\nI updated my post to detail the pseudolabeling more. I used soft pseudolabels (predicted probabilities) and a modified cross entropy loss. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 974508,
      "postDate": "2020-08-18T01:00:55.630Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> , very good results, strong models, and for becoming GM!<br>\nI have a question, but it is more about curiosity, I wonder how you usually do your experiments at the beginning? like the ones you said done with the 1080TI, I always have a hard time in the beginning, experimenting with simpler and weaker models before moving to the strong ones.</p>",
      "rawMarkdown": "Congratulations @vaillant , very good results, strong models, and for becoming GM!\nI have a question, but it is more about curiosity, I wonder how you usually do your experiments at the beginning? like the ones you said done with the 1080TI, I always have a hard time in the beginning, experimenting with simpler and weaker models before moving to the strong ones.",
      "votes": 6,
      "replies": [
        {
          "id": 974694,
          "postDate": "2020-08-18T02:39:07.503Z",
          "content": "<p>I'm having the same question here. Will the experiments setup on a simple model generalize on the complex model?</p>",
          "rawMarkdown": "I'm having the same question here. Will the experiments setup on a simple model generalize on the complex model?",
          "votes": 1
        },
        {
          "id": 975600,
          "postDate": "2020-08-18T11:24:28.607Z",
          "content": "<p>Yeah <a href=\"https://www.kaggle.com/waylongo\" target=\"_blank\">@waylongo</a> , that is my point, in my experience spending too much time on weaker models hyper-parameters can be a waste of time as most don't generalize to the other, what may be a time saving using weaker models is doing experiments with datasets or training/evaluation pipelines.</p>",
          "rawMarkdown": "Yeah @waylongo , that is my point, in my experience spending too much time on weaker models hyper-parameters can be a waste of time as most don't generalize to the other, what may be a time saving using weaker models is doing experiments with datasets or training/evaluation pipelines.",
          "votes": 1
        },
        {
          "id": 975627,
          "postDate": "2020-08-18T11:42:18.907Z",
          "content": "<p>I think he might use Chris's kernel, which is highly modulized. Totally respect!</p>",
          "rawMarkdown": "I think he might use Chris's kernel, which is highly modulized. Totally respect!"
        },
        {
          "id": 976541,
          "postDate": "2020-08-18T23:55:42.313Z",
          "content": "<p>I had gotten a grant from HOSTKEY for 4x 1080 Ti, so to maximize usage of all my GPUs, I also ran additional experiments on those GPUs. </p>\n<p>I mainly used it to test different backbones. I used small image size 256x256 and just experimented with different backbones to see which one performed best. This doesn't guarantee that I will pick the best backbone for the final configuration, but it led me to conclude that EfficientNet was the best family of backbones, without much benefit going to B7/B8. </p>\n<p>Usually I don't have the luxury of having additional GPUs for extra experimentation. I didn't use the 1080s for the final models because I wanted to leverage maximum batch size. </p>",
          "rawMarkdown": "I had gotten a grant from HOSTKEY for 4x 1080 Ti, so to maximize usage of all my GPUs, I also ran additional experiments on those GPUs. \n\nI mainly used it to test different backbones. I used small image size 256x256 and just experimented with different backbones to see which one performed best. This doesn't guarantee that I will pick the best backbone for the final configuration, but it led me to conclude that EfficientNet was the best family of backbones, without much benefit going to B7/B8. \n\nUsually I don't have the luxury of having additional GPUs for extra experimentation. I didn't use the 1080s for the final models because I wanted to leverage maximum batch size. ",
          "votes": 2
        },
        {
          "id": 976577,
          "postDate": "2020-08-19T00:46:04.193Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> this makes a lot of sense.</p>",
          "rawMarkdown": "Thanks @vaillant this makes a lot of sense."
        }
      ]
    },
    {
      "id": 975923,
      "postDate": "2020-08-18T14:25:19.007Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congratulations!!! That's incredible. 😃</p>",
      "rawMarkdown": "@vaillant Congratulations!!! That's incredible. 😃",
      "votes": 4
    },
    {
      "id": 983480,
      "postDate": "2020-08-24T10:42:32.930Z",
      "content": "<p>Congratulations , Ian! It is very important to achieve this success by spending much less time than many people. Enjoy the \"GM club\" because you deserve it!</p>",
      "rawMarkdown": "Congratulations , Ian! It is very important to achieve this success by spending much less time than many people. Enjoy the \"GM club\" because you deserve it!\n",
      "votes": 1
    },
    {
      "id": 976591,
      "postDate": "2020-08-19T01:09:21.223Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> , <strong>Congratulations</strong> for becoming the <strong>GM</strong>. More surprisingly you are a medical students too. What a fantastic combination you have !  Really appreciable. </p>",
      "rawMarkdown": "@vaillant , **Congratulations** for becoming the **GM**. More surprisingly you are a medical students too. What a fantastic combination you have !  Really appreciable. ",
      "votes": 1
    },
    {
      "id": 975613,
      "postDate": "2020-08-18T11:35:12.133Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congrats on becoming a Grandmaster! You're my superstar because I also started kaggle while I was a medical student and I'm a medical resident now. It's unbelievable for me to compete for kaggle while doing a residency.</p>",
      "rawMarkdown": "@vaillant Congrats on becoming a Grandmaster! You're my superstar because I also started kaggle while I was a medical student and I'm a medical resident now. It's unbelievable for me to compete for kaggle while doing a residency.",
      "votes": 1,
      "replies": [
        {
          "id": 976537,
          "postDate": "2020-08-18T23:50:58.543Z",
          "content": "<p>Thank you! Now that I've started residency, it's definitely tough to keep up with Kaggle. I got lucky with this competition…</p>",
          "rawMarkdown": "Thank you! Now that I've started residency, it's definitely tough to keep up with Kaggle. I got lucky with this competition...",
          "votes": 1
        }
      ]
    },
    {
      "id": 975167,
      "postDate": "2020-08-18T07:21:13.837Z",
      "content": "<p>Congrats on the 2nd place finish and the well deserved GM title, Doctor Pan!</p>",
      "rawMarkdown": "Congrats on the 2nd place finish and the well deserved GM title, Doctor Pan!",
      "votes": 1
    },
    {
      "id": 974745,
      "postDate": "2020-08-18T03:21:51.360Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> on becoming GM!</p>",
      "rawMarkdown": "Congrats @vaillant on becoming GM!",
      "votes": 1
    },
    {
      "id": 974488,
      "postDate": "2020-08-18T00:51:27.653Z",
      "content": "<p>Glad to see straightforward data science at work.</p>\n<blockquote>\n  <p>4x Quadro 24GB</p>\n</blockquote>\n<p>Sweet setup =]</p>",
      "rawMarkdown": "Glad to see straightforward data science at work.\n\n> 4x Quadro 24GB\n\nSweet setup =]",
      "votes": 1
    },
    {
      "id": 974480,
      "postDate": "2020-08-18T00:46:24.263Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Double congratulation. Wish you all the best for your upcoming journey. It's a great skill set, being a doctor in the profession and same time practitioner in a deep learning sphere. Good luck brother. </p>",
      "rawMarkdown": "@vaillant Double congratulation. Wish you all the best for your upcoming journey. It's a great skill set, being a doctor in the profession and same time practitioner in a deep learning sphere. Good luck brother. ",
      "votes": 1
    },
    {
      "id": 974455,
      "postDate": "2020-08-18T00:28:25.503Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> seeing pytorch solution that high is great!</p>",
      "rawMarkdown": "Congrats @vaillant seeing pytorch solution that high is great!",
      "votes": 1
    },
    {
      "id": 974446,
      "postDate": "2020-08-18T00:26:32.540Z",
      "content": "<p>Awesome work and congrats on earning grandmaster… oh and becoming a medical doctor! Superhuman.</p>",
      "rawMarkdown": "Awesome work and congrats on earning grandmaster... oh and becoming a medical doctor! Superhuman.",
      "votes": 1
    },
    {
      "id": 974444,
      "postDate": "2020-08-18T00:26:28.533Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> ! Nice and clean solution.</p>",
      "rawMarkdown": "Congrats @vaillant ! Nice and clean solution.",
      "votes": 1
    },
    {
      "id": 1005160,
      "postDate": "2020-09-10T09:44:05.377Z",
      "content": "<p>Very excellent work. I will try the codes in GitHub</p>",
      "rawMarkdown": "Very excellent work. I will try the codes in GitHub\n"
    },
    {
      "id": 1000324,
      "postDate": "2020-09-06T13:15:03.690Z",
      "content": "<p>Congratulations i-pan! I having following your progress on Kaggle since RSNA comp. I am really inspired by your journey and would like to learn more from you. I am trying to build a smooth experimentations pipeline that will allow quick iterations. Thank you for being part of this community! Well deserved.</p>",
      "rawMarkdown": "Congratulations i-pan! I having following your progress on Kaggle since RSNA comp. I am really inspired by your journey and would like to learn more from you. I am trying to build a smooth experimentations pipeline that will allow quick iterations. Thank you for being part of this community! Well deserved."
    },
    {
      "id": 987863,
      "postDate": "2020-08-27T15:23:47.073Z",
      "content": "<p>Great work!!</p>",
      "rawMarkdown": "Great work!!\n"
    },
    {
      "id": 986422,
      "postDate": "2020-08-26T13:28:19.033Z",
      "content": "<p>Thank you for taking the time to do the write up and congrats!</p>",
      "rawMarkdown": "Thank you for taking the time to do the write up and congrats!"
    },
    {
      "id": 984026,
      "postDate": "2020-08-24T19:19:59.303Z",
      "content": "<p>Very motivating to hear of someone that has gone through so much in addition to this competition, and persevered.  Congrats on both this and becoming a MD! </p>",
      "rawMarkdown": "Very motivating to hear of someone that has gone through so much in addition to this competition, and persevered.  Congrats on both this and becoming a MD! "
    },
    {
      "id": 983230,
      "postDate": "2020-08-24T05:54:15.633Z",
      "content": "<p>Congratuataions <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> on twin win 👍👍.<br>\nHave one query on use of cross entropy loss as it was highly imbalance data so how you fine tune AUC on cross entropy.<br>\nSimple Cross entropy was not improving score and then I stuck to focal loss only which was doing better with Radam.</p>\n<p>Thanks</p>",
      "rawMarkdown": "Congratuataions @vaillant on twin win 👍👍.\nHave one query on use of cross entropy loss as it was highly imbalance data so how you fine tune AUC on cross entropy.\nSimple Cross entropy was not improving score and then I stuck to focal loss only which was doing better with Radam.\n\nThanks"
    },
    {
      "id": 983147,
      "postDate": "2020-08-24T04:48:33.820Z",
      "content": "<p>Clean Explanation!<br>\nCongrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  🎊</p>",
      "rawMarkdown": "Clean Explanation!\nCongrats @vaillant  🎊"
    },
    {
      "id": 982896,
      "postDate": "2020-08-23T19:04:42.463Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>, I've seen at least 3 very differing implementations of gradient accumulation in pytorch. Considering your final LB position, whatever you were doing was clearly working. Can you share the details of your gradaccum, e.g. model.zero_grad() vs optimizer.zero_grad(), if you divided by the number of steps or not, when you propagated, etc? Even pseudo code would suffice.</p>\n<p>Thanks!</p>",
      "rawMarkdown": "@vaillant, I've seen at least 3 very differing implementations of gradient accumulation in pytorch. Considering your final LB position, whatever you were doing was clearly working. Can you share the details of your gradaccum, e.g. model.zero_grad() vs optimizer.zero_grad(), if you divided by the number of steps or not, when you propagated, etc? Even pseudo code would suffice.\n\nThanks!",
      "replies": [
        {
          "id": 984104,
          "postDate": "2020-08-24T20:40:24.387Z",
          "content": "<p>I have a <code>Trainer</code> class that uses this for gradient accumulation:</p>\n<pre><code>self.optimizer.zero_grad()\nfor i in range(int(self.gradient_accumulation)):\n    output = self.model(accum_batch[i])\n    loss = self.criterion(output, accum_labels[i])\n    (loss / self.gradient_accumulation).backward()\nself.optimizer.step()\n</code></pre>",
          "rawMarkdown": "I have a `Trainer` class that uses this for gradient accumulation:\n\n```\nself.optimizer.zero_grad()\nfor i in range(int(self.gradient_accumulation)):\n    output = self.model(accum_batch[i])\n    loss = self.criterion(output, accum_labels[i])\n    (loss / self.gradient_accumulation).backward()\nself.optimizer.step()\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 982643,
      "postDate": "2020-08-23T14:44:31.503Z",
      "content": "<p>Congratulations, great job. An interesting note about the difference between melanomas and nevus.</p>",
      "rawMarkdown": "Congratulations, great job. An interesting note about the difference between melanomas and nevus."
    },
    {
      "id": 982583,
      "postDate": "2020-08-23T14:00:38.963Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congratulations</p>",
      "rawMarkdown": "@vaillant Congratulations"
    },
    {
      "id": 982446,
      "postDate": "2020-08-23T11:31:50.770Z",
      "content": "<p>Incredible work Ian, people like you motivate me to push myself harder. Thanks for sharing</p>",
      "rawMarkdown": "Incredible work Ian, people like you motivate me to push myself harder. Thanks for sharing"
    },
    {
      "id": 981886,
      "postDate": "2020-08-22T19:42:03.670Z",
      "content": "<p>Congratulations, I'm so happy for you.\nYou just gave me a push to go forward.</p>",
      "rawMarkdown": "Congratulations, I'm so happy for you.\nYou just gave me a push to go forward."
    },
    {
      "id": 980674,
      "postDate": "2020-08-21T19:14:42.587Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats"
    },
    {
      "id": 980040,
      "postDate": "2020-08-21T09:41:37.130Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "rawMarkdown": "Congrats @vaillant "
    },
    {
      "id": 979816,
      "postDate": "2020-08-21T06:08:48.580Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Thank you for this detailed explanation.👍</p>",
      "rawMarkdown": "@vaillant Thank you for this detailed explanation.👍"
    },
    {
      "id": 977758,
      "postDate": "2020-08-19T17:34:23.407Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> so you used 2019 new data as well or only 2019 old data? </p>",
      "rawMarkdown": "@vaillant so you used 2019 new data as well or only 2019 old data? ",
      "replies": [
        {
          "id": 977773,
          "postDate": "2020-08-19T17:40:52.597Z",
          "content": "<p>I used all the training data for ISIC 2019 in addition to the data provided for this competition. </p>",
          "rawMarkdown": "I used all the training data for ISIC 2019 in addition to the data provided for this competition. "
        }
      ]
    },
    {
      "id": 977618,
      "postDate": "2020-08-19T15:44:41.500Z",
      "content": "<p>Congrat's <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> on both of your achievements. My brother is also a medical resident back in France so I can understand that time is not on your side 😷 Also thanks for sharing your insights. It looks like EfficientNet is becoming the norm lately on many competitions. Looking forward to your cleaned notebook to get into the details of your solution 🤓 Again great job!! 👋👋</p>",
      "rawMarkdown": "Congrat's @vaillant on both of your achievements. My brother is also a medical resident back in France so I can understand that time is not on your side 😷 Also thanks for sharing your insights. It looks like EfficientNet is becoming the norm lately on many competitions. Looking forward to your cleaned notebook to get into the details of your solution 🤓 Again great job!! 👋👋"
    },
    {
      "id": 977494,
      "postDate": "2020-08-19T14:13:18.100Z",
      "content": "<p>THX for sharing！</p>\n\n<p>It was inspiring as well！</p>",
      "rawMarkdown": "THX for sharing！\n\nIt was inspiring as well！"
    },
    {
      "id": 976804,
      "postDate": "2020-08-19T05:30:14.547Z",
      "content": "<p>Congrats!!</p>",
      "rawMarkdown": "Congrats!!"
    },
    {
      "id": 976718,
      "postDate": "2020-08-19T04:05:04.733Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> <br>\nI am working with Dr Trivedi (HITI lab) on the spine hardware project for which you had created the initial object detection models. I was really impressed at the way your code was structured but when I came to know that you are a medical student, I am in awe since then. Congratulations on this win!!<br>\nI must say you have set the bar really high for me in the project as well given your expertise in data science. Again, congratulations! </p>",
      "rawMarkdown": "Hi @vaillant \nI am working with Dr Trivedi (HITI lab) on the spine hardware project for which you had created the initial object detection models. I was really impressed at the way your code was structured but when I came to know that you are a medical student, I am in awe since then. Congratulations on this win!!\nI must say you have set the bar really high for me in the project as well given your expertise in data science. Again, congratulations! "
    },
    {
      "id": 976074,
      "postDate": "2020-08-18T16:08:48.203Z",
      "content": "<p>Awesome.  Why provisional?  You expect top team to be removed?  ;)</p>\n<p>Just kidding, congrats on the result.  I tried to use model trained on 2019 predictions on 2020 as feature, but did not think of pseudo label.  Great idea!</p>",
      "rawMarkdown": "Awesome.  Why provisional?  You expect top team to be removed?  ;)\n\nJust kidding, congrats on the result.  I tried to use model trained on 2019 predictions on 2020 as feature, but did not think of pseudo label.  Great idea!",
      "replies": [
        {
          "id": 976521,
          "postDate": "2020-08-18T23:27:42.473Z",
          "content": "<p>Haha, provisional since I might get DQ'd! I don't think I broke any rules, but after the deepfake fiasco, you can never be sure…</p>",
          "rawMarkdown": "Haha, provisional since I might get DQ'd! I don't think I broke any rules, but after the deepfake fiasco, you can never be sure...",
          "votes": 4
        }
      ]
    },
    {
      "id": 976059,
      "postDate": "2020-08-18T15:58:51.940Z",
      "content": "<p>great job!!</p>",
      "rawMarkdown": "great job!!"
    },
    {
      "id": 975820,
      "postDate": "2020-08-18T13:37:10.700Z",
      "content": "<p>Congrats!!</p>",
      "rawMarkdown": "Congrats!!"
    },
    {
      "id": 975618,
      "postDate": "2020-08-18T11:37:56.247Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Big congrats for you to become GM! I'm very curious why you didn't adopt the metadata at first? And in you later models with metadata, how did you use the metadata? Something like 2 branches network, something like one for CNN features &amp; the other for meta featrues? Thank you!</p>",
      "rawMarkdown": "Hi @vaillant Big congrats for you to become GM! I'm very curious why you didn't adopt the metadata at first? And in you later models with metadata, how did you use the metadata? Something like 2 branches network, something like one for CNN features & the other for meta featrues? Thank you!",
      "replies": [
        {
          "id": 976538,
          "postDate": "2020-08-18T23:52:07.150Z",
          "content": "<p>Thank you! I was using metadata in the beginning, and during the last week I decided to train some models without metadata so I could be eligible for without context prizes. Turns out these were my best models in the end. </p>\n<p>I used embedding layers to map the metadata to 32-dimensional vectors, which I then concatenated to the image feature vector before inputting the concatenated vector into the final linear classification head. </p>",
          "rawMarkdown": "Thank you! I was using metadata in the beginning, and during the last week I decided to train some models without metadata so I could be eligible for without context prizes. Turns out these were my best models in the end. \n\nI used embedding layers to map the metadata to 32-dimensional vectors, which I then concatenated to the image feature vector before inputting the concatenated vector into the final linear classification head. "
        },
        {
          "id": 976851,
          "postDate": "2020-08-19T06:17:24.797Z",
          "content": "<p>Hi Ian, just saw you updated post about this part. I'm clear. Thanks for you explanation &amp; once again, big congrats!</p>",
          "rawMarkdown": "Hi Ian, just saw you updated post about this part. I'm clear. Thanks for you explanation & once again, big congrats!"
        }
      ]
    },
    {
      "id": 975511,
      "postDate": "2020-08-18T10:26:55.873Z",
      "content": "<p>congrats on making it to GM!</p>",
      "rawMarkdown": "congrats on making it to GM!"
    },
    {
      "id": 975184,
      "postDate": "2020-08-18T07:28:45.653Z",
      "content": "<p>Great Job. Congrats to GM. A huge milestone in ones kaggle carrer.</p>",
      "rawMarkdown": "Great Job. Congrats to GM. A huge milestone in ones kaggle carrer."
    },
    {
      "id": 975105,
      "postDate": "2020-08-18T06:39:45.527Z",
      "content": "<p>Congratulations :D </p>",
      "rawMarkdown": "Congratulations :D "
    },
    {
      "id": 974835,
      "postDate": "2020-08-18T04:19:55.827Z",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for becoming grand master!!!</p>",
      "rawMarkdown": "congrats @vaillant for becoming grand master!!!"
    },
    {
      "id": 974826,
      "postDate": "2020-08-18T04:14:37.907Z",
      "content": "<p>Wow，Congratulations. Nice GMMMMMMMMMMMM!</p>",
      "rawMarkdown": "Wow，Congratulations. Nice GMMMMMMMMMMMM!"
    },
    {
      "id": 974772,
      "postDate": "2020-08-18T03:42:41.607Z",
      "content": "<p>Congrats on your milestones, Ian! Very nice job in the achieving this result and becoming competitions GM while working 70-80 hours/week. That is very impressive!</p>",
      "rawMarkdown": "Congrats on your milestones, Ian! Very nice job in the achieving this result and becoming competitions GM while working 70-80 hours/week. That is very impressive!"
    },
    {
      "id": 974710,
      "postDate": "2020-08-18T02:52:53.080Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  Congrats on your GM!</p>",
      "rawMarkdown": "@vaillant  Congrats on your GM!"
    },
    {
      "id": 974697,
      "postDate": "2020-08-18T02:39:23.233Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> for the win and making Grandmaster. You deserve it.</p>\n<p>Thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @vaillant for the win and making Grandmaster. You deserve it.\n\nThanks for sharing your solution overview."
    },
    {
      "id": 974684,
      "postDate": "2020-08-18T02:34:21.797Z",
      "content": "<p>Great job. Since you also used pytorch I'm really curious to see the little tricks I may have been missing :) </p>\n<p>You had really strong hardware though ! How long did it take for these large model + large image to train ?</p>",
      "rawMarkdown": "Great job. Since you also used pytorch I'm really curious to see the little tricks I may have been missing :) \n\nYou had really strong hardware though ! How long did it take for these large model + large image to train ?",
      "replies": [
        {
          "id": 976535,
          "postDate": "2020-08-18T23:49:26.720Z",
          "content": "<p>For my EfficientNet-B6, each fold took about 2 hours. For B7, about 4 hours. </p>",
          "rawMarkdown": "For my EfficientNet-B6, each fold took about 2 hours. For B7, about 4 hours. "
        }
      ]
    },
    {
      "id": 974640,
      "postDate": "2020-08-18T02:09:32.970Z",
      "content": "<p>Thanks for insights and inspiration I will going try my best and maybe next time took medal too😄 Congratulations🎉</p>\n<p>Can you explain how from 3 classes (mel/nev/other) you jumped to 2 classes, thanks.</p>",
      "rawMarkdown": "Thanks for insights and inspiration I will going try my best and maybe next time took medal too😄 Congratulations🎉\n\nCan you explain how from 3 classes (mel/nev/other) you jumped to 2 classes, thanks.",
      "replies": [
        {
          "id": 976536,
          "postDate": "2020-08-18T23:49:53.103Z",
          "content": "<p>Thanks! I just used the melanoma probability as the predicted score for submission.</p>",
          "rawMarkdown": "Thanks! I just used the melanoma probability as the predicted score for submission."
        }
      ]
    },
    {
      "id": 974607,
      "postDate": "2020-08-18T01:59:50.380Z",
      "content": "<p>congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "rawMarkdown": "congrats @vaillant "
    },
    {
      "id": 974572,
      "postDate": "2020-08-18T01:36:53.110Z",
      "content": "<p>Congrats for GM!</p>",
      "rawMarkdown": "Congrats for GM!"
    },
    {
      "id": 974541,
      "postDate": "2020-08-18T01:22:09.340Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> - excellent result!</p>\n<p>Did you use label smoothing with bce?</p>",
      "rawMarkdown": "Congratulations @vaillant - excellent result!\n\nDid you use label smoothing with bce?",
      "replies": [
        {
          "id": 976539,
          "postDate": "2020-08-18T23:52:41.980Z",
          "content": "<p>Thank you! I did not use label smoothing. </p>",
          "rawMarkdown": "Thank you! I did not use label smoothing. "
        }
      ]
    },
    {
      "id": 974521,
      "postDate": "2020-08-18T01:09:23.033Z",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>, congratulations with becoming a grandmaster and with the first solo gold, great job!</p>",
      "rawMarkdown": "@vaillant, congratulations with becoming a grandmaster and with the first solo gold, great job!"
    },
    {
      "id": 974471,
      "postDate": "2020-08-18T00:38:30.300Z",
      "content": "<p>Many congratulations to you!! It is rare to see such a fantastic blend of Kaggle grand master and a medical doctor! </p>",
      "rawMarkdown": "Many congratulations to you!! It is rare to see such a fantastic blend of Kaggle grand master and a medical doctor! "
    },
    {
      "id": 974465,
      "postDate": "2020-08-18T00:35:41.510Z",
      "content": "<p>Many congratulations!</p>",
      "rawMarkdown": "Many congratulations!"
    },
    {
      "id": 974439,
      "postDate": "2020-08-18T00:23:58.017Z",
      "content": "<p>Congrats for GrandMaster!!</p>",
      "rawMarkdown": "Congrats for GrandMaster!!"
    },
    {
      "id": 974437,
      "postDate": "2020-08-18T00:23:33.520Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "rawMarkdown": "Congrats @vaillant "
    },
    {
      "id": 974434,
      "postDate": "2020-08-18T00:22:39.527Z",
      "content": "<p>Congrats. very good job!</p>",
      "rawMarkdown": "Congrats. very good job!"
    },
    {
      "id": 974429,
      "postDate": "2020-08-18T00:20:46.213Z",
      "content": "<p>Congrats 🎉</p>",
      "rawMarkdown": "Congrats 🎉"
    },
    {
      "id": 988886,
      "postDate": "2020-08-28T11:22:53.293Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1712792,
      "postDate": "2022-03-05T11:01:42.323Z",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!"
    }
  ],
  "comments": [
    {
      "id": 980705,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-08-21T19:36:47.260000",
      "content": "<p>You can edit your title now ;)  Welcome to competitions GM club, well deserved!</p>",
      "votes": 7,
      "replies": [
        {
          "id": 981141,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2020-08-22T07:45:25.253000",
          "content": "<p>speaking from his profile picture (and we all know profile pics show the true age), he might be the youngest GM ever</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 981250,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2020-08-22T10:13:53.060000",
          "content": "<p>Anti Dieter indeed.  You two should team at some point.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 982341,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-08-23T09:28:27.393000",
          "content": "<p>He needs a cool hat though. ;)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 974762,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2020-08-18T03:33:37.440000",
      "content": "<p>Congrats Ian on becoming Kaggle Grandmaster and achieving 2nd place money finish. </p>\n<p>It looks like Pseudo labeling worked well for you. How did you assign targets 0 and 1 based on model 1's predictions?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 975628,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T11:43:38.587000",
          "content": "<p>Same question here. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 976522,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-08-18T23:28:21.767000",
          "content": "<p>Thank you, Chris! I really appreciate all your discussion posts and notebooks, in this competition and others. </p>\n<p>I updated my post to detail the pseudolabeling more. I used soft pseudolabels (predicted probabilities) and a modified cross entropy loss. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 974508,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-08-18T01:00:55.630000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> , very good results, strong models, and for becoming GM!<br>\nI have a question, but it is more about curiosity, I wonder how you usually do your experiments at the beginning? like the ones you said done with the 1080TI, I always have a hard time in the beginning, experimenting with simpler and weaker models before moving to the strong ones.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 974694,
          "author_name": "Waylon Wu",
          "author_url": "",
          "post_date": "2020-08-18T02:39:07.503000",
          "content": "<p>I'm having the same question here. Will the experiments setup on a simple model generalize on the complex model?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975600,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-18T11:24:28.607000",
          "content": "<p>Yeah <a href=\"https://www.kaggle.com/waylongo\" target=\"_blank\">@waylongo</a> , that is my point, in my experience spending too much time on weaker models hyper-parameters can be a waste of time as most don't generalize to the other, what may be a time saving using weaker models is doing experiments with datasets or training/evaluation pipelines.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 975627,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T11:42:18.907000",
          "content": "<p>I think he might use Chris's kernel, which is highly modulized. Totally respect!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 976541,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-08-18T23:55:42.313000",
          "content": "<p>I had gotten a grant from HOSTKEY for 4x 1080 Ti, so to maximize usage of all my GPUs, I also ran additional experiments on those GPUs. </p>\n<p>I mainly used it to test different backbones. I used small image size 256x256 and just experimented with different backbones to see which one performed best. This doesn't guarantee that I will pick the best backbone for the final configuration, but it led me to conclude that EfficientNet was the best family of backbones, without much benefit going to B7/B8. </p>\n<p>Usually I don't have the luxury of having additional GPUs for extra experimentation. I didn't use the 1080s for the final models because I wanted to leverage maximum batch size. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 976577,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2020-08-19T00:46:04.193000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> this makes a lot of sense.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 975923,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2020-08-18T14:25:19.007000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congratulations!!! That's incredible. 😃</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 983480,
      "author_name": "acanacar",
      "author_url": "",
      "post_date": "2020-08-24T10:42:32.930000",
      "content": "<p>Congratulations , Ian! It is very important to achieve this success by spending much less time than many people. Enjoy the \"GM club\" because you deserve it!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 976591,
      "author_name": "Md Fahim",
      "author_url": "",
      "post_date": "2020-08-19T01:09:21.223000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> , <strong>Congratulations</strong> for becoming the <strong>GM</strong>. More surprisingly you are a medical students too. What a fantastic combination you have !  Really appreciable. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 975613,
      "author_name": "OsciiArt",
      "author_url": "",
      "post_date": "2020-08-18T11:35:12.133000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congrats on becoming a Grandmaster! You're my superstar because I also started kaggle while I was a medical student and I'm a medical resident now. It's unbelievable for me to compete for kaggle while doing a residency.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 976537,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-08-18T23:50:58.543000",
          "content": "<p>Thank you! Now that I've started residency, it's definitely tough to keep up with Kaggle. I got lucky with this competition…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 975167,
      "author_name": "Bo",
      "author_url": "",
      "post_date": "2020-08-18T07:21:13.837000",
      "content": "<p>Congrats on the 2nd place finish and the well deserved GM title, Doctor Pan!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974745,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2020-08-18T03:21:51.360000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> on becoming GM!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974488,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-08-18T00:51:27.653000",
      "content": "<p>Glad to see straightforward data science at work.</p>\n<blockquote>\n  <p>4x Quadro 24GB</p>\n</blockquote>\n<p>Sweet setup =]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974480,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-08-18T00:46:24.263000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Double congratulation. Wish you all the best for your upcoming journey. It's a great skill set, being a doctor in the profession and same time practitioner in a deep learning sphere. Good luck brother. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974455,
      "author_name": "Ertuğrul Demir",
      "author_url": "",
      "post_date": "2020-08-18T00:28:25.503000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> seeing pytorch solution that high is great!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974446,
      "author_name": "Rob Mulla",
      "author_url": "",
      "post_date": "2020-08-18T00:26:32.540000",
      "content": "<p>Awesome work and congrats on earning grandmaster… oh and becoming a medical doctor! Superhuman.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 974444,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-08-18T00:26:28.533000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> ! Nice and clean solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1005160,
      "author_name": "amirahalsulami",
      "author_url": "",
      "post_date": "2020-09-10T09:44:05.377000",
      "content": "<p>Very excellent work. I will try the codes in GitHub</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1000324,
      "author_name": "sm_erlo",
      "author_url": "",
      "post_date": "2020-09-06T13:15:03.690000",
      "content": "<p>Congratulations i-pan! I having following your progress on Kaggle since RSNA comp. I am really inspired by your journey and would like to learn more from you. I am trying to build a smooth experimentations pipeline that will allow quick iterations. Thank you for being part of this community! Well deserved.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 987863,
      "author_name": "Bivek Subedi",
      "author_url": "",
      "post_date": "2020-08-27T15:23:47.073000",
      "content": "<p>Great work!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 986422,
      "author_name": "5starkarma",
      "author_url": "",
      "post_date": "2020-08-26T13:28:19.033000",
      "content": "<p>Thank you for taking the time to do the write up and congrats!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 984026,
      "author_name": "Baron Smith",
      "author_url": "",
      "post_date": "2020-08-24T19:19:59.303000",
      "content": "<p>Very motivating to hear of someone that has gone through so much in addition to this competition, and persevered.  Congrats on both this and becoming a MD! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 983230,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2020-08-24T05:54:15.633000",
      "content": "<p>Congratuataions <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> on twin win 👍👍.<br>\nHave one query on use of cross entropy loss as it was highly imbalance data so how you fine tune AUC on cross entropy.<br>\nSimple Cross entropy was not improving score and then I stuck to focal loss only which was doing better with Radam.</p>\n<p>Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 983147,
      "author_name": "Naman Jaswani",
      "author_url": "",
      "post_date": "2020-08-24T04:48:33.820000",
      "content": "<p>Clean Explanation!<br>\nCongrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>  🎊</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 982896,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2020-08-23T19:04:42.463000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a>, I've seen at least 3 very differing implementations of gradient accumulation in pytorch. Considering your final LB position, whatever you were doing was clearly working. Can you share the details of your gradaccum, e.g. model.zero_grad() vs optimizer.zero_grad(), if you divided by the number of steps or not, when you propagated, etc? Even pseudo code would suffice.</p>\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 984104,
          "author_name": "Ian Pan",
          "author_url": "",
          "post_date": "2020-08-24T20:40:24.387000",
          "content": "<p>I have a <code>Trainer</code> class that uses this for gradient accumulation:</p>\n<pre><code>self.optimizer.zero_grad()\nfor i in range(int(self.gradient_accumulation)):\n    output = self.model(accum_batch[i])\n    loss = self.criterion(output, accum_labels[i])\n    (loss / self.gradient_accumulation).backward()\nself.optimizer.step()\n</code></pre>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 982643,
      "author_name": "Ladanova Sveta",
      "author_url": "",
      "post_date": "2020-08-23T14:44:31.503000",
      "content": "<p>Congratulations, great job. An interesting note about the difference between melanomas and nevus.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 982583,
      "author_name": "Nandanam",
      "author_url": "",
      "post_date": "2020-08-23T14:00:38.963000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Congratulations</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 982446,
      "author_name": "Alan Choon",
      "author_url": "",
      "post_date": "2020-08-23T11:31:50.770000",
      "content": "<p>Incredible work Ian, people like you motivate me to push myself harder. Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 981886,
      "author_name": "YASSYN IDAR",
      "author_url": "",
      "post_date": "2020-08-22T19:42:03.670000",
      "content": "<p>Congratulations, I'm so happy for you.\nYou just gave me a push to go forward.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 980674,
      "author_name": "Bijeesha Vs",
      "author_url": "",
      "post_date": "2020-08-21T19:14:42.587000",
      "content": "<p>Congrats</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 980040,
      "author_name": "Utku_Kubilay",
      "author_url": "",
      "post_date": "2020-08-21T09:41:37.130000",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 979816,
      "author_name": "@pp1e",
      "author_url": "",
      "post_date": "2020-08-21T06:08:48.580000",
      "content": "<p><a href=\"https://www.kaggle.com/vaillant\" target=\"_blank\">@vaillant</a> Thank you for this detailed explanation.👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 977758,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-19T17:34:23.407000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 977773,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-19T17:40:52.597000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 977618,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-19T15:44:41.500000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 977494,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-19T14:13:18.100000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 976804,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-19T05:30:14.547000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 976718,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-19T04:05:04.733000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 976074,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T16:08:48.203000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 976521,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T23:27:42.473000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 976059,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T15:58:51.940000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975820,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T13:37:10.700000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975618,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T11:37:56.247000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 976538,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T23:52:07.150000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 976851,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-19T06:17:24.797000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 975511,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T10:26:55.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975184,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T07:28:45.653000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 975105,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T06:39:45.527000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974835,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T04:19:55.827000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974826,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T04:14:37.907000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974772,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T03:42:41.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974710,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:52:53.080000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974697,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:39:23.233000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974684,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:34:21.797000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 976535,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T23:49:26.720000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974640,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T02:09:32.970000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 976536,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T23:49:53.103000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974607,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:59:50.380000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974572,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:36:53.110000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974541,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:22:09.340000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 976539,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-08-18T23:52:41.980000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 974521,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T01:09:23.033000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974471,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:38:30.300000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974465,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:35:41.510000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974439,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:23:58.017000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974437,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:23:33.520000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974434,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:22:39.527000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 974429,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-18T00:20:46.213000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 988886,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-28T11:22:53.293000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1712792,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-05T11:01:42.323000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "974426": "Code available here: https://github.com/i-pan/kaggle-melanoma\n\nWow. Did not expect this. But super excited to finally get my solo gold medal and complete the journey to Grandmaster. I'm grateful to benefit from the shakeup this time after suffering in PANDA. Thank you to all the organizers, and congratulations to all the other winners and participants. \n\nThis year, I accomplished 2 milestones: 1) I graduated medical school and became a doctor; and 2) I became a Kaggle Competitions Grandmaster. I actually just started my intern year as a doctor working 70-80 hours/week so did not have too much time to dedicate to this competition. The key was having a pipeline that allowed me to quickly iterate on experiments so I could start them in the morning, go to work, then analyze the results when I came back. Long story short, to all those who are starting out, keep going. Put in the time and effort. Compete, read winners' solutions, and learn, over and over again. \n\nI have to thank my teammates in previous competitions @felipekitamura, @alexandrecc, and @jamesphoward. And I also have to thank the Kaggle community as I have learned so much from participating.\n\nAlso huge thanks to @cdeotte for all the insights he shared. I used his triple stratified dataset to split my data for CV. \n\n**Environment**\nPyTorch 1.6 with automatic mixed precision\n4x NVIDIA GTX 1080 Ti 11GB provided by HOSTKEY (www.hostkey.com) as part of their Kaggle competitions grant\n4x Quadro RTX 6000 24GB \n\n**Batch Size**\nIn the beginning, I was training models on GTX 1080 Ti GPUs provided by HOSTKEY. I wasn’t able to get multi-GPU training to work, so I was just experimenting with batch sizes of 8-16 on a single 11GB 1080 Ti. My CV scores were not that high, usually a 5-fold average of around 0.92, with similar results on LB. The TPU kernels were doing so well (on public LB, at least), that I thought it was in part due to increased batch size. At that point, I switched over to my 4x Quadro RTX6000 24GB setup to leverage more GPU memory and multi-GPU training. I aimed for a batch size of 64 while maximizing image resolution for a particular backbone (and if that would not fit, I settled for BS 32 with gradient accumulation 2).\n\n**Backbones & Image Resolution**\nI tried several backbones in the EfficientNet, SE-ResNeXt, and ResNeSt families. I also tried BiT-ResNet (recently released by Google). EfficientNet performed better on CV so I decided to stick with EfficientNets for the remainder of the competition. I used the 1080 Ti GPUs to experiment with different backbones, all other hyperparameters held constant. Some of you asked whether this necessarily transfers when I switch over to a larger GPU and increase the batch size/change the image resolution. I did briefly compare backbones on the larger GPU and the results seemed to be consistent (EfficientNet > all others). \n\nI experimented with backbones of different sizes from pruned EfficientNet-B3 to EfficientNet-B8, using the implementation from https://github.com/rwightman/pytorch-image-models. For smaller backbones, I tried larger resolutions (up to 1024 x 1024) and for EfficientNet-B8 I went down to 384 x 384. The image size for each backbone was selected so that I could use batch size of 64 (16/GPU) during training. I found that smaller EfficientNets at higher resolutions were not as good. The best models for me were EfficientNet-B6 (initialized with noisy student weights) and EfficientNet-B7, at image resolutions of 512 x 512 and 640 x 640, respectively, so I only used these moving forward.\n\n**Base Model**\nMy base model was your standard convolutional neural network backbone with linear classification head. I used generalized mean pooling with trainable parameter p (not sure if this was any better than average or max, as I just stuck with GeM from the beginning). I also used multisample dropout following @qishenha’s implementation in one of this winning NLP solutions. \n\n**From 2 Class to 3 Class**\nI felt that more granular classes would result in better feature representations that could help improve performance. The majority of melanomas are dark (exception being rare amelanotic melanomas), so differentiating them from benign nevi is probably the most challenging task. The 2019 data all had auxiliary diagnoses, including nevi, whereas a large fraction of the 2020 data was unknown. \n\nI trained a model on 2019 data only using the diagnosis as the target. Then, I applied this model to the 2020 data. My main focus was on labeling the unknowns as nevus or not nevus, as I know they are not melanoma. To find the threshold at which I would label an image as nevus, I used the 5th percentile of the 2019 model’s predictions on the 2020 data which had a known label of nevus.\n\nNow, all of the 2019 and 2020 data has a label of other, benign nevus, or melanoma, and I trained my model on these 3 classes using vanilla cross-entropy loss. I did not try label smoothing.\n\n**Upsampling**\nThere was a lot of discussion over whether or not to upsample malignant images or not. I did upsample malignant images for 2020 data. Because I used 2019 data and the percentage of melanomas in that dataset was much higher, I wanted to make sure that the 2019 melanomas did not overwhelm the 2020 melanomas. To that end, I upsampled the 2020 melanomas 7 times so that there was about an equal number of melanomas from both datasets.\n\n**Training**\nI used AdamW optimizer and cosine annealing with warm restarts scheduler, initial learning rate 3.0E-4. I did 3 snapshots, 2 epochs each for EfficientNet-B6 and 3 epochs each for EfficientNet-B7. I noticed that it did not take long for the models to start overfitting. I found that this gave better results than using one cycle, so I stuck with it. Out of the 3 snapshots, I just took the one that did the best on that validation fold. Every experiment I ran, I did 5-fold CV using Chris Deotte’s triple stratified data splits. Single validation folds were not stable for me, so in order to really understand my model performance and the effects of my adjustments I had to look at 5-fold CV average. I only validated on 2020 data. \n\nIn the beginning, I was using metadata by using embeddings for age, sex, and location. Each embedding was mapped to a 32-D vector and concatenated to the final feature vector before input into the linear classification layer. I did not want to spend too much time tuning this because I was afraid I would overfit to the distribution of the training set. I just used mean/mode imputation for missing values. \n\nIt wasn’t until the last several days of the competition that I decided to train models without metadata, so I could be eligible for the without context special prize. It turns out that these models were actually my highest scoring private LB solutions!\n\n**Augmentations**\nI knew that augmentation would be important given the small percentage of melanomas in the 2020 data. I used the RandAugment strategy, implemented here: https://github.com/ildoonet/pytorch-randaugment. I used N=3 augmentations with magnitude M/30 where M was sampled from a Poisson distribution with mean 12 for extra stochasticity. For those unfamiliar with RandAugment, M is essentially the “hardness” of the augmentation (angle for rotation, % zoom, gamma for contrast adjustment, etc.). For augmentations like flips, M is not relevant. I tried other augmentations such as mixup, cutmix, and grid mask, but those did not help.\n\nI also used square cropping during training and inference. During training, a square was randomly cropped from the image if it was rectangular (otherwise, the entire image was used), where the length of the square image was the size of the shortest side (i.e., 768x512 would be cropped to 512x512). During inference, I spaced out 10 square crops as TTA and took the average as the final prediction (again, unless the image was already square - then no TTA was applied). I found that this gave me better results than rectangular crops or using the whole image. \n\n**Pseudolabeling**\nPseudolabeling was key to my solution. Given the limited number of 2020 melanomas, I felt that pseudolabeling would help increase performance. 2019 melanomas were helpful but still different from 2020 melanomas. I took my 5-fold EfficientNet-B6 model, trained without metadata, and obtained soft pseudolabels (3 classes) for the test set. When combining the test data with the training data (2019+2020), I upsampled images with melanoma prediction > 0.5 7 times (same factor as I did for 2020 training data). I used @cpmpml's implementation (https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/173733) of cross entropy in PyTorch (without label smoothing) so I could use soft pseudolabels. \n\n**CV vs LB**\nI knew early on that it would be easy to fit public LB, given the small number of melanomas that would be in the public test. At the same time, the CV for my different experiments was much tighter than the LB, so I was nervous to fully trust CV as there may have been differences between training and test data. With that in mind, I favored solutions that had reasonably high CV and LB. \n\nThere is a fair amount of luck that goes into picking the right solution, but you should be able to justify to yourself why you are picking a certain solution over another (going by CV score, LB score, some combination of CV/LB, or some hypotheses about the private test set that would favor one solution over another). \n\nMy 2nd place solution was an ensemble of 3 5-fold models:\n- EfficientNet-B6, 512x512, BS64, no metadata (CV 0.9336 / public 0.9534)\n- EfficientNet-B7, 640x640, BS32, gradient accumulation 2, no metadata (CV 0.9389 / public 0.9525)\n- Model 1, trained on combined training and pseudolabeled test data (CV 0.9438 / public 0.9493)\n\nNote that the CV score does not account for the 5-fold blend effect. \n\nMy highest scoring private LB solution was actually model 3 alone, which I did not select.\n\nI trained other models with metadata, but best private LB score was 0.945 (public LB 0.959), with similar CV. \n\n",
    "980705": "You can edit your title now ;)  Welcome to competitions GM club, well deserved!",
    "974762": "Congrats Ian on becoming Kaggle Grandmaster and achieving 2nd place money finish. \n\nIt looks like Pseudo labeling worked well for you. How did you assign targets 0 and 1 based on model 1's predictions?",
    "974508": "Congratulations @vaillant , very good results, strong models, and for becoming GM!\nI have a question, but it is more about curiosity, I wonder how you usually do your experiments at the beginning? like the ones you said done with the 1080TI, I always have a hard time in the beginning, experimenting with simpler and weaker models before moving to the strong ones.",
    "975923": "@vaillant Congratulations!!! That's incredible. 😃",
    "983480": "Congratulations , Ian! It is very important to achieve this success by spending much less time than many people. Enjoy the \"GM club\" because you deserve it!\n",
    "976591": "@vaillant , **Congratulations** for becoming the **GM**. More surprisingly you are a medical students too. What a fantastic combination you have !  Really appreciable. ",
    "975613": "@vaillant Congrats on becoming a Grandmaster! You're my superstar because I also started kaggle while I was a medical student and I'm a medical resident now. It's unbelievable for me to compete for kaggle while doing a residency.",
    "975167": "Congrats on the 2nd place finish and the well deserved GM title, Doctor Pan!",
    "974745": "Congrats @vaillant on becoming GM!",
    "974488": "Glad to see straightforward data science at work.\n\n> 4x Quadro 24GB\n\nSweet setup =]",
    "974480": "@vaillant Double congratulation. Wish you all the best for your upcoming journey. It's a great skill set, being a doctor in the profession and same time practitioner in a deep learning sphere. Good luck brother. ",
    "974455": "Congrats @vaillant seeing pytorch solution that high is great!",
    "974446": "Awesome work and congrats on earning grandmaster... oh and becoming a medical doctor! Superhuman.",
    "974444": "Congrats @vaillant ! Nice and clean solution.",
    "1005160": "Very excellent work. I will try the codes in GitHub\n",
    "1000324": "Congratulations i-pan! I having following your progress on Kaggle since RSNA comp. I am really inspired by your journey and would like to learn more from you. I am trying to build a smooth experimentations pipeline that will allow quick iterations. Thank you for being part of this community! Well deserved.",
    "987863": "Great work!!\n",
    "986422": "Thank you for taking the time to do the write up and congrats!",
    "984026": "Very motivating to hear of someone that has gone through so much in addition to this competition, and persevered.  Congrats on both this and becoming a MD! ",
    "983230": "Congratuataions @vaillant on twin win 👍👍.\nHave one query on use of cross entropy loss as it was highly imbalance data so how you fine tune AUC on cross entropy.\nSimple Cross entropy was not improving score and then I stuck to focal loss only which was doing better with Radam.\n\nThanks",
    "983147": "Clean Explanation!\nCongrats @vaillant  🎊",
    "982896": "@vaillant, I've seen at least 3 very differing implementations of gradient accumulation in pytorch. Considering your final LB position, whatever you were doing was clearly working. Can you share the details of your gradaccum, e.g. model.zero_grad() vs optimizer.zero_grad(), if you divided by the number of steps or not, when you propagated, etc? Even pseudo code would suffice.\n\nThanks!",
    "982643": "Congratulations, great job. An interesting note about the difference between melanomas and nevus.",
    "982583": "@vaillant Congratulations",
    "982446": "Incredible work Ian, people like you motivate me to push myself harder. Thanks for sharing",
    "981886": "Congratulations, I'm so happy for you.\nYou just gave me a push to go forward.",
    "980674": "Congrats",
    "980040": "Congrats @vaillant ",
    "979816": "@vaillant Thank you for this detailed explanation.👍",
    "977758": "@vaillant so you used 2019 new data as well or only 2019 old data? ",
    "977618": "Congrat's @vaillant on both of your achievements. My brother is also a medical resident back in France so I can understand that time is not on your side 😷 Also thanks for sharing your insights. It looks like EfficientNet is becoming the norm lately on many competitions. Looking forward to your cleaned notebook to get into the details of your solution 🤓 Again great job!! 👋👋",
    "977494": "THX for sharing！\n\nIt was inspiring as well！",
    "976804": "Congrats!!",
    "976718": "Hi @vaillant \nI am working with Dr Trivedi (HITI lab) on the spine hardware project for which you had created the initial object detection models. I was really impressed at the way your code was structured but when I came to know that you are a medical student, I am in awe since then. Congratulations on this win!!\nI must say you have set the bar really high for me in the project as well given your expertise in data science. Again, congratulations! ",
    "976074": "Awesome.  Why provisional?  You expect top team to be removed?  ;)\n\nJust kidding, congrats on the result.  I tried to use model trained on 2019 predictions on 2020 as feature, but did not think of pseudo label.  Great idea!",
    "976059": "great job!!",
    "975820": "Congrats!!",
    "975618": "Hi @vaillant Big congrats for you to become GM! I'm very curious why you didn't adopt the metadata at first? And in you later models with metadata, how did you use the metadata? Something like 2 branches network, something like one for CNN features & the other for meta featrues? Thank you!",
    "975511": "congrats on making it to GM!",
    "975184": "Great Job. Congrats to GM. A huge milestone in ones kaggle carrer.",
    "975105": "Congratulations :D ",
    "974835": "congrats @vaillant for becoming grand master!!!",
    "974826": "Wow，Congratulations. Nice GMMMMMMMMMMMM!",
    "974772": "Congrats on your milestones, Ian! Very nice job in the achieving this result and becoming competitions GM while working 70-80 hours/week. That is very impressive!",
    "974710": "@vaillant  Congrats on your GM!",
    "974697": "Congrats @vaillant for the win and making Grandmaster. You deserve it.\n\nThanks for sharing your solution overview.",
    "974684": "Great job. Since you also used pytorch I'm really curious to see the little tricks I may have been missing :) \n\nYou had really strong hardware though ! How long did it take for these large model + large image to train ?",
    "974640": "Thanks for insights and inspiration I will going try my best and maybe next time took medal too😄 Congratulations🎉\n\nCan you explain how from 3 classes (mel/nev/other) you jumped to 2 classes, thanks.",
    "974607": "congrats @vaillant ",
    "974572": "Congrats for GM!",
    "974541": "Congratulations @vaillant - excellent result!\n\nDid you use label smoothing with bce?",
    "974521": "@vaillant, congratulations with becoming a grandmaster and with the first solo gold, great job!",
    "974471": "Many congratulations to you!! It is rare to see such a fantastic blend of Kaggle grand master and a medical doctor! ",
    "974465": "Many congratulations!",
    "974439": "Congrats for GrandMaster!!",
    "974437": "Congrats @vaillant ",
    "974434": "Congrats. very good job!",
    "974429": "Congrats 🎉",
    "988886": "",
    "1712792": "Thank you!"
  }
}