{
  "id": 320192,
  "title": "1st Place Solution",
  "url": "/competitions/happy-whale-and-dolphin/discussion/320192",
  "author_name": "knshnb",
  "post_date": "2022-04-20T12:43:31.707000",
  "votes": 198,
  "comment_count": 83,
  "views": 0,
  "content": "<p>This was my first participation in a Kaggle competition and I was so fortunate to win 1st place!!!!<br>\nI appreciate my team member <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>!</p>\n<p>We really enjoyed the competition and worked hard literally until the last minutes. We would like to deeply appreciate the Kaggle staff for organizing this great competition and other teams for competing with us.</p>\n<h2>Overview</h2>\n<p>Our solution is based on sub-center ArcFace with Dynamic margins, which was shown to be effective in <a href=\"https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\" target=\"_blank\">Google Landmark Recognition 2020 the 3rd place solution</a>  by <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/discussion/277098\" target=\"_blank\">the first place solution of Google Landmark Recognition 2021</a> by <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>\n<p>Basically, our solution was an ensemble of two pipelines implemented by <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and me. Since we shared knowledge with each other during the competition, our pipelines share much in common. Below, I will mainly explain my pipeline, which was slightly better in CV score. </p>\n<h2>Key Points</h2>\n<ul>\n<li>Tuning of dynamic margin hyperparameters with <a href=\"https://optuna.org/\" target=\"_blank\">Optuna</a></li>\n<li>Larger learning rate for ArcFace head</li>\n<li>Bounding box mixing augmentation</li>\n<li>Ensemble of knn and logit</li>\n<li>Two-round pseudo labeling</li>\n<li>Ensemble of many models</li>\n</ul>\n<h2>Dataset</h2>\n<p>We used several types of bounding boxes to crop images. We thank <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> and <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> a lot for providing valuable datasets. We also trained our own yolov5 model using fullbody annotations, which we refer to as fullbody_charm.</p>\n<p>For train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins. Adding non-cropped images by a small ratio also worked as a regularization. For test data, we took the mean of predictions between fullbody and fullbody_charm.</p>\n<p>We used the cropped images resized to a fixed size. We mainly used the image size of (1024, 1024). Some models were trained with the image size of (1200, 1200) and (1440, 1440) for ensembling.</p>\n<h2>Backbone &amp; Neck</h2>\n<p>We trained several different imagenet-pretrained backbones for ensembling (efficientnet_b5, efficientnet_b6, efficientnet_b7, efficientnetv2_m, efficientnetv2_l, etc). The best performance in a single model was achieved by efficientnet_b7.</p>\n<p>Using GeM pooling (p=3) instead of GAP enhanced the performance.</p>\n<p>The normalization layer before the ArcFace head was important. Batchnorm was slightly better than Layernorm in our experiments.</p>\n<p>In addition to the final feature map of the backbone, we used the second final feature map to capture more local information. We simply concatenated those two GeM-pooled feature maps and passed them to head.</p>\n<h2>Head</h2>\n<p>For handling imbalanced classes, we adopted ArcFace with dynamic margins. Since it seemed sensitive to hyperparameters, we tuned them on images of (256, 256) and efficientnet_b0 using Optuna. It seemed the acquired hyperparameters also worked well on large images and architectures.</p>\n<p>In the last competition, it was reported that handling flipped images as different classes significantly enhanced the performance. In this competition, we did not think that this technique works well because some images are taken from different angles. To handle this issue, we adapted the sub-center ArcFace of k=2 with the usual flip data augmentation.</p>\n<p>We also added a second head for classifying species. Sub-center ArcFace with dynamic margins worked better than simple Linear head.</p>\n<h2>Training</h2>\n<p>While we mainly checked a single-fold validation score locally, we trained our models by using whole train data for submission.<br>\nSetting the learning rate of the head 10 times bigger than the learning rate of the backbone significantly improved the performance.<br>\nOptimal training settings of us differed possibly due to slight differences in our pipelines. While I trained the models for 30 epochs by AdamW optimizer of lr_backbone=1.6e-3 with warmup cosine annealing scheduler, charmq trained the models for 20 epochs by Adam of lr_backbone=1e-4 with cosine annealing scheduler. Most of the models were trained with the batch size of 16-32 on 2-8x NVIDIA Tesla V100 (32GB).</p>\n<p>In addition to the bounding box mix augmentation described above, we adopted many data augmentations because the models were expressive enough to reach almost 100% of training accuracy. Below are the list of data augmentations we used by Albumentations implementation:</p>\n<pre><code>A.Affine(rotate=(-15, 15), translate_percent=(0.0, 0.25), shear=(-3, 3), p=0.5),\nA.RandomResizedCrop(image_size[0], image_size[1], scale=(0.9, 1.0), ratio=(0.75, 1.3333333333)),\nA.ToGray(p=0.1),\nA.GaussianBlur(blur_limit=(3, 7), p=0.05),\nA.GaussNoise(p=0.05),\nA.RandomGridShuffle(grid=(2, 2), p=0.3),\nA.Posterize(p=0.2),\nA.RandomBrightnessContrast(p=0.5),\nA.Cutout(p=0.05),\nA.RandomSnow(p=0.1),\nA.RandomRain(p=0.05),\nA.HorizontalFlip(p=0.5),\n</code></pre>\n<h2>Postprocess</h2>\n<p>We combined the following two metrics.<br>\nknn: Using feature vectors, we calculated the largest cosine similarity for each class in training data. This can be done for each model and is easier to ensemble many models than feature concatenation. In implementation, we looked at only the top 500 training data for each test individual by <code>sklearn.neighbors.NearestNeighbors</code>. As a test-time augmentation, we took the mean of original test images and flipped test images searched over both original train images and flipped train images.<br>\nlogit: We used the simple output of the model without margins.<br>\n(We calculated the mean of predictions for two bounding boxes as explained above.) While knn worked much better in CV but only slightly better in the public leaderboard than logit. This is probably caused by highly imbalanced data and the distribution differences between train and test (knn is more likely to output classes with more train data). To mitigate this, we mixed the prediction of knn and logit with knn_ratio=0.5. After pseudo labeling, we increased the knn_ratio to 0.8.</p>\n<p>We labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.</p>\n<h2>Other</h2>\n<p>Pseudo labeling enhanced the public LB score a lot in this competition probably because of extremely imbalanced data. We stuck to improving CV scores until the very final stage of the competition. On the day before the deadline, we got a big boost in the leaderboard score (0.88589/0.85959 -&gt; 0.89343/0.87062) by a pseudo-label submission. The second round of pseudo labeling on the final day also improved the score (0.89680/0.87579). Additional rounds of pseudo labeling might have further improved performance, but unfortunately, we did not have time to do that.</p>\n<p>As a final submission, we ensembled around 50 models including the ones trained without pseudo labels and with the first-round pseudo labels. After the competition, we confirmed that the ensemble of only 2 models (the best one from each of us) scored 0.89385/0.87336, which could still win first place!</p>\n<h2>What did not work</h2>\n<ul>\n<li>input 4-channel images with segmentation mask (1st place solution of the last competition)</li>\n<li>input 6-channel images combining 2 types of images cropped by fullbody and backfin bboxes</li>\n<li>input rectangle images such as (512, 1024)</li>\n<li>maintain aspect ratio of images by bounding box expansion instead of resizing</li>\n<li>AdaCos, triplet loss</li>\n<li>Focal loss</li>\n<li>training p of GeM pooling</li>\n<li>ConvNeXt</li>\n<li>Swin Transformer (384 was too small)</li>\n<li>dolg</li>\n<li>using pseudo labels for knn</li>\n<li>cutmix</li>\n</ul>\n<h2>Acknowledgement</h2>\n<p>We deeply acknowledge great OSS such as PyTorch, PyTorch Lightning, PyTorch Image Models, Albumentations, etc. We would also like to appreciate Preferred Networks, Inc for allowing us to use computational resources.</p>\n<p>[Update 2022/05/24]<br>\nI added explanations on some minor points.<br>\ncode (knshnb): <a href=\"https://github.com/knshnb/kaggle-happywhale-1st-place\" target=\"_blank\">https://github.com/knshnb/kaggle-happywhale-1st-place</a><br>\ncode (charmq): <a href=\"https://github.com/tyamaguchi17/kaggle-happywhale-1st-place-solution-charmq\" target=\"_blank\">https://github.com/tyamaguchi17/kaggle-happywhale-1st-place-solution-charmq</a></p>",
  "messages": [
    {
      "id": 1762107,
      "postDate": "2022-04-20T12:43:31.707Z",
      "content": "<p>This was my first participation in a Kaggle competition and I was so fortunate to win 1st place!!!!<br>\nI appreciate my team member <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a>!</p>\n<p>We really enjoyed the competition and worked hard literally until the last minutes. We would like to deeply appreciate the Kaggle staff for organizing this great competition and other teams for competing with us.</p>\n<h2>Overview</h2>\n<p>Our solution is based on sub-center ArcFace with Dynamic margins, which was shown to be effective in <a href=\"https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\" target=\"_blank\">Google Landmark Recognition 2020 the 3rd place solution</a>  by <a href=\"https://www.kaggle.com/boliu0\" target=\"_blank\">@boliu0</a> and <a href=\"https://www.kaggle.com/c/landmark-recognition-2021/discussion/277098\" target=\"_blank\">the first place solution of Google Landmark Recognition 2021</a> by <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a>.</p>\n<p>Basically, our solution was an ensemble of two pipelines implemented by <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> and me. Since we shared knowledge with each other during the competition, our pipelines share much in common. Below, I will mainly explain my pipeline, which was slightly better in CV score. </p>\n<h2>Key Points</h2>\n<ul>\n<li>Tuning of dynamic margin hyperparameters with <a href=\"https://optuna.org/\" target=\"_blank\">Optuna</a></li>\n<li>Larger learning rate for ArcFace head</li>\n<li>Bounding box mixing augmentation</li>\n<li>Ensemble of knn and logit</li>\n<li>Two-round pseudo labeling</li>\n<li>Ensemble of many models</li>\n</ul>\n<h2>Dataset</h2>\n<p>We used several types of bounding boxes to crop images. We thank <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> and <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> a lot for providing valuable datasets. We also trained our own yolov5 model using fullbody annotations, which we refer to as fullbody_charm.</p>\n<p>For train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins. Adding non-cropped images by a small ratio also worked as a regularization. For test data, we took the mean of predictions between fullbody and fullbody_charm.</p>\n<p>We used the cropped images resized to a fixed size. We mainly used the image size of (1024, 1024). Some models were trained with the image size of (1200, 1200) and (1440, 1440) for ensembling.</p>\n<h2>Backbone &amp; Neck</h2>\n<p>We trained several different imagenet-pretrained backbones for ensembling (efficientnet_b5, efficientnet_b6, efficientnet_b7, efficientnetv2_m, efficientnetv2_l, etc). The best performance in a single model was achieved by efficientnet_b7.</p>\n<p>Using GeM pooling (p=3) instead of GAP enhanced the performance.</p>\n<p>The normalization layer before the ArcFace head was important. Batchnorm was slightly better than Layernorm in our experiments.</p>\n<p>In addition to the final feature map of the backbone, we used the second final feature map to capture more local information. We simply concatenated those two GeM-pooled feature maps and passed them to head.</p>\n<h2>Head</h2>\n<p>For handling imbalanced classes, we adopted ArcFace with dynamic margins. Since it seemed sensitive to hyperparameters, we tuned them on images of (256, 256) and efficientnet_b0 using Optuna. It seemed the acquired hyperparameters also worked well on large images and architectures.</p>\n<p>In the last competition, it was reported that handling flipped images as different classes significantly enhanced the performance. In this competition, we did not think that this technique works well because some images are taken from different angles. To handle this issue, we adapted the sub-center ArcFace of k=2 with the usual flip data augmentation.</p>\n<p>We also added a second head for classifying species. Sub-center ArcFace with dynamic margins worked better than simple Linear head.</p>\n<h2>Training</h2>\n<p>While we mainly checked a single-fold validation score locally, we trained our models by using whole train data for submission.<br>\nSetting the learning rate of the head 10 times bigger than the learning rate of the backbone significantly improved the performance.<br>\nOptimal training settings of us differed possibly due to slight differences in our pipelines. While I trained the models for 30 epochs by AdamW optimizer of lr_backbone=1.6e-3 with warmup cosine annealing scheduler, charmq trained the models for 20 epochs by Adam of lr_backbone=1e-4 with cosine annealing scheduler. Most of the models were trained with the batch size of 16-32 on 2-8x NVIDIA Tesla V100 (32GB).</p>\n<p>In addition to the bounding box mix augmentation described above, we adopted many data augmentations because the models were expressive enough to reach almost 100% of training accuracy. Below are the list of data augmentations we used by Albumentations implementation:</p>\n<pre><code>A.Affine(rotate=(-15, 15), translate_percent=(0.0, 0.25), shear=(-3, 3), p=0.5),\nA.RandomResizedCrop(image_size[0], image_size[1], scale=(0.9, 1.0), ratio=(0.75, 1.3333333333)),\nA.ToGray(p=0.1),\nA.GaussianBlur(blur_limit=(3, 7), p=0.05),\nA.GaussNoise(p=0.05),\nA.RandomGridShuffle(grid=(2, 2), p=0.3),\nA.Posterize(p=0.2),\nA.RandomBrightnessContrast(p=0.5),\nA.Cutout(p=0.05),\nA.RandomSnow(p=0.1),\nA.RandomRain(p=0.05),\nA.HorizontalFlip(p=0.5),\n</code></pre>\n<h2>Postprocess</h2>\n<p>We combined the following two metrics.<br>\nknn: Using feature vectors, we calculated the largest cosine similarity for each class in training data. This can be done for each model and is easier to ensemble many models than feature concatenation. In implementation, we looked at only the top 500 training data for each test individual by <code>sklearn.neighbors.NearestNeighbors</code>. As a test-time augmentation, we took the mean of original test images and flipped test images searched over both original train images and flipped train images.<br>\nlogit: We used the simple output of the model without margins.<br>\n(We calculated the mean of predictions for two bounding boxes as explained above.) While knn worked much better in CV but only slightly better in the public leaderboard than logit. This is probably caused by highly imbalanced data and the distribution differences between train and test (knn is more likely to output classes with more train data). To mitigate this, we mixed the prediction of knn and logit with knn_ratio=0.5. After pseudo labeling, we increased the knn_ratio to 0.8.</p>\n<p>We labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.</p>\n<h2>Other</h2>\n<p>Pseudo labeling enhanced the public LB score a lot in this competition probably because of extremely imbalanced data. We stuck to improving CV scores until the very final stage of the competition. On the day before the deadline, we got a big boost in the leaderboard score (0.88589/0.85959 -&gt; 0.89343/0.87062) by a pseudo-label submission. The second round of pseudo labeling on the final day also improved the score (0.89680/0.87579). Additional rounds of pseudo labeling might have further improved performance, but unfortunately, we did not have time to do that.</p>\n<p>As a final submission, we ensembled around 50 models including the ones trained without pseudo labels and with the first-round pseudo labels. After the competition, we confirmed that the ensemble of only 2 models (the best one from each of us) scored 0.89385/0.87336, which could still win first place!</p>\n<h2>What did not work</h2>\n<ul>\n<li>input 4-channel images with segmentation mask (1st place solution of the last competition)</li>\n<li>input 6-channel images combining 2 types of images cropped by fullbody and backfin bboxes</li>\n<li>input rectangle images such as (512, 1024)</li>\n<li>maintain aspect ratio of images by bounding box expansion instead of resizing</li>\n<li>AdaCos, triplet loss</li>\n<li>Focal loss</li>\n<li>training p of GeM pooling</li>\n<li>ConvNeXt</li>\n<li>Swin Transformer (384 was too small)</li>\n<li>dolg</li>\n<li>using pseudo labels for knn</li>\n<li>cutmix</li>\n</ul>\n<h2>Acknowledgement</h2>\n<p>We deeply acknowledge great OSS such as PyTorch, PyTorch Lightning, PyTorch Image Models, Albumentations, etc. We would also like to appreciate Preferred Networks, Inc for allowing us to use computational resources.</p>\n<p>[Update 2022/05/24]<br>\nI added explanations on some minor points.<br>\ncode (knshnb): <a href=\"https://github.com/knshnb/kaggle-happywhale-1st-place\" target=\"_blank\">https://github.com/knshnb/kaggle-happywhale-1st-place</a><br>\ncode (charmq): <a href=\"https://github.com/tyamaguchi17/kaggle-happywhale-1st-place-solution-charmq\" target=\"_blank\">https://github.com/tyamaguchi17/kaggle-happywhale-1st-place-solution-charmq</a></p>",
      "rawMarkdown": "This was my first participation in a Kaggle competition and I was so fortunate to win 1st place!!!!\nI appreciate my team member @charmq!\n\nWe really enjoyed the competition and worked hard literally until the last minutes. We would like to deeply appreciate the Kaggle staff for organizing this great competition and other teams for competing with us.\n\n## Overview\nOur solution is based on sub-center ArcFace with Dynamic margins, which was shown to be effective in [Google Landmark Recognition 2020 the 3rd place solution](https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757)  by @boliu0 and [the first place solution of Google Landmark Recognition 2021](https://www.kaggle.com/c/landmark-recognition-2021/discussion/277098) by @christofhenkel.\n\nBasically, our solution was an ensemble of two pipelines implemented by @charmq and me. Since we shared knowledge with each other during the competition, our pipelines share much in common. Below, I will mainly explain my pipeline, which was slightly better in CV score. \n\n## Key Points\n- Tuning of dynamic margin hyperparameters with [Optuna](https://optuna.org/)\n- Larger learning rate for ArcFace head\n- Bounding box mixing augmentation\n- Ensemble of knn and logit\n- Two-round pseudo labeling\n- Ensemble of many models\n\n## Dataset\nWe used several types of bounding boxes to crop images. We thank @jpbremer and @phalanx a lot for providing valuable datasets. We also trained our own yolov5 model using fullbody annotations, which we refer to as fullbody_charm.\n\nFor train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins. Adding non-cropped images by a small ratio also worked as a regularization. For test data, we took the mean of predictions between fullbody and fullbody_charm.\n\nWe used the cropped images resized to a fixed size. We mainly used the image size of (1024, 1024). Some models were trained with the image size of (1200, 1200) and (1440, 1440) for ensembling.\n\n## Backbone & Neck\nWe trained several different imagenet-pretrained backbones for ensembling (efficientnet_b5, efficientnet_b6, efficientnet_b7, efficientnetv2_m, efficientnetv2_l, etc). The best performance in a single model was achieved by efficientnet_b7.\n\nUsing GeM pooling (p=3) instead of GAP enhanced the performance.\n\nThe normalization layer before the ArcFace head was important. Batchnorm was slightly better than Layernorm in our experiments.\n\nIn addition to the final feature map of the backbone, we used the second final feature map to capture more local information. We simply concatenated those two GeM-pooled feature maps and passed them to head.\n\n## Head\nFor handling imbalanced classes, we adopted ArcFace with dynamic margins. Since it seemed sensitive to hyperparameters, we tuned them on images of (256, 256) and efficientnet_b0 using Optuna. It seemed the acquired hyperparameters also worked well on large images and architectures.\n\nIn the last competition, it was reported that handling flipped images as different classes significantly enhanced the performance. In this competition, we did not think that this technique works well because some images are taken from different angles. To handle this issue, we adapted the sub-center ArcFace of k=2 with the usual flip data augmentation.\n\nWe also added a second head for classifying species. Sub-center ArcFace with dynamic margins worked better than simple Linear head.\n\n## Training\nWhile we mainly checked a single-fold validation score locally, we trained our models by using whole train data for submission.\nSetting the learning rate of the head 10 times bigger than the learning rate of the backbone significantly improved the performance.\nOptimal training settings of us differed possibly due to slight differences in our pipelines. While I trained the models for 30 epochs by AdamW optimizer of lr_backbone=1.6e-3 with warmup cosine annealing scheduler, charmq trained the models for 20 epochs by Adam of lr_backbone=1e-4 with cosine annealing scheduler. Most of the models were trained with the batch size of 16-32 on 2-8x NVIDIA Tesla V100 (32GB).\n\nIn addition to the bounding box mix augmentation described above, we adopted many data augmentations because the models were expressive enough to reach almost 100% of training accuracy. Below are the list of data augmentations we used by Albumentations implementation:\n```\nA.Affine(rotate=(-15, 15), translate_percent=(0.0, 0.25), shear=(-3, 3), p=0.5),\nA.RandomResizedCrop(image_size[0], image_size[1], scale=(0.9, 1.0), ratio=(0.75, 1.3333333333)),\nA.ToGray(p=0.1),\nA.GaussianBlur(blur_limit=(3, 7), p=0.05),\nA.GaussNoise(p=0.05),\nA.RandomGridShuffle(grid=(2, 2), p=0.3),\nA.Posterize(p=0.2),\nA.RandomBrightnessContrast(p=0.5),\nA.Cutout(p=0.05),\nA.RandomSnow(p=0.1),\nA.RandomRain(p=0.05),\nA.HorizontalFlip(p=0.5),\n```\n\n## Postprocess\nWe combined the following two metrics.\nknn: Using feature vectors, we calculated the largest cosine similarity for each class in training data. This can be done for each model and is easier to ensemble many models than feature concatenation. In implementation, we looked at only the top 500 training data for each test individual by `sklearn.neighbors.NearestNeighbors`. As a test-time augmentation, we took the mean of original test images and flipped test images searched over both original train images and flipped train images.\nlogit: We used the simple output of the model without margins.\n(We calculated the mean of predictions for two bounding boxes as explained above.) While knn worked much better in CV but only slightly better in the public leaderboard than logit. This is probably caused by highly imbalanced data and the distribution differences between train and test (knn is more likely to output classes with more train data). To mitigate this, we mixed the prediction of knn and logit with knn_ratio=0.5. After pseudo labeling, we increased the knn_ratio to 0.8.\n\nWe labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.\n\n## Other\nPseudo labeling enhanced the public LB score a lot in this competition probably because of extremely imbalanced data. We stuck to improving CV scores until the very final stage of the competition. On the day before the deadline, we got a big boost in the leaderboard score (0.88589/0.85959 -> 0.89343/0.87062) by a pseudo-label submission. The second round of pseudo labeling on the final day also improved the score (0.89680/0.87579). Additional rounds of pseudo labeling might have further improved performance, but unfortunately, we did not have time to do that.\n\nAs a final submission, we ensembled around 50 models including the ones trained without pseudo labels and with the first-round pseudo labels. After the competition, we confirmed that the ensemble of only 2 models (the best one from each of us) scored 0.89385/0.87336, which could still win first place!\n\n## What did not work\n- input 4-channel images with segmentation mask (1st place solution of the last competition)\n- input 6-channel images combining 2 types of images cropped by fullbody and backfin bboxes\n- input rectangle images such as (512, 1024)\n- maintain aspect ratio of images by bounding box expansion instead of resizing\n- AdaCos, triplet loss\n- Focal loss\n- training p of GeM pooling\n- ConvNeXt\n- Swin Transformer (384 was too small)\n- dolg\n- using pseudo labels for knn\n- cutmix\n\n## Acknowledgement\nWe deeply acknowledge great OSS such as PyTorch, PyTorch Lightning, PyTorch Image Models, Albumentations, etc. We would also like to appreciate Preferred Networks, Inc for allowing us to use computational resources.\n\n[Update 2022/05/24]\nI added explanations on some minor points.\ncode (knshnb): https://github.com/knshnb/kaggle-happywhale-1st-place\ncode (charmq): https://github.com/tyamaguchi17/kaggle-happywhale-1st-place-solution-charmq",
      "votes": 198
    },
    {
      "id": 1762120,
      "postDate": "2022-04-20T12:52:52.323Z",
      "content": "<p>Thanks for teaming up with me! <br>\nI can't believe you were novice. Amazing performance!</p>",
      "rawMarkdown": "Thanks for teaming up with me! \nI can't believe you were novice. Amazing performance!",
      "votes": 15,
      "replies": [
        {
          "id": 1765430,
          "postDate": "2022-04-23T13:43:43.187Z",
          "content": "<p>Thank you for merging with me too. It was a great experience working with you in my first competition!</p>",
          "rawMarkdown": "Thank you for merging with me too. It was a great experience working with you in my first competition!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1770239,
      "postDate": "2022-04-28T05:41:38.980Z",
      "content": "<p>nice, good job</p>",
      "rawMarkdown": "nice, good job",
      "votes": 5
    },
    {
      "id": 1762190,
      "postDate": "2022-04-20T13:55:08.980Z",
      "content": "<p>Big congrats! Amazing job!</p>\n<p>At least it seems we were the only team to reach 0.890 without any pseudo labeling :D<br>\nMight be the reason we dropped a bit more than others on private. It's crazy to see that everyone used it with success.</p>",
      "rawMarkdown": "Big congrats! Amazing job!\n\nAt least it seems we were the only team to reach 0.890 without any pseudo labeling :D\nMight be the reason we dropped a bit more than others on private. It's crazy to see that everyone used it with success.",
      "votes": 5,
      "replies": [
        {
          "id": 1762214,
          "postDate": "2022-04-20T14:13:14.247Z",
          "content": "<p>I agree with you. I think pseudo labeling is a risky method in some sense because we can't see the improvement without submission. Actually, we tried pseudo labeling in the last 2 days. We saw the improvement in the lb score and were convinced that everyone else was using pseudo labeling.</p>",
          "rawMarkdown": "I agree with you. I think pseudo labeling is a risky method in some sense because we can't see the improvement without submission. Actually, we tried pseudo labeling in the last 2 days. We saw the improvement in the lb score and were convinced that everyone else was using pseudo labeling.",
          "votes": 3
        },
        {
          "id": 1762343,
          "postDate": "2022-04-20T15:59:45.347Z",
          "content": "<p>Yeah, I was even expecting you to use pseudo as you improved step by step. But for us it did not work when trying it on CV, so we dropped it. And as we had perfect correlation between CV and LB saw no reason to try it on LB. But we should have :)</p>",
          "rawMarkdown": "Yeah, I was even expecting you to use pseudo as you improved step by step. But for us it did not work when trying it on CV, so we dropped it. And as we had perfect correlation between CV and LB saw no reason to try it on LB. But we should have :)\n",
          "votes": 2
        },
        {
          "id": 1762401,
          "postDate": "2022-04-20T16:47:32.580Z",
          "content": "<p>Wow <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Amazing model without pseudo labels! Our team also did not use pseudo labels either. And after reading all the top solutions, i realize this is what held us back. </p>\n<p>I read solutions from last year and it seemed that it did not help last year, so we didn't try it. Upon reflection, this year has 30k test images whereas last year had only 5k. And this year half (8k) of the individual id had only 1 image where last year was more dense. So it makes sense why it would be important this year and not last year.</p>\n<p>I'm so curious how much it would improve our pipeline. I'm thinking training all our models using pseudo today. But i'm also tired of training models, so i may not.</p>",
          "rawMarkdown": "Wow @philippsinger Amazing model without pseudo labels! Our team also did not use pseudo labels either. And after reading all the top solutions, i realize this is what held us back. \n\n I read solutions from last year and it seemed that it did not help last year, so we didn't try it. Upon reflection, this year has 30k test images whereas last year had only 5k. And this year half (8k) of the individual id had only 1 image where last year was more dense. So it makes sense why it would be important this year and not last year.\n\nI'm so curious how much it would improve our pipeline. I'm thinking training all our models using pseudo today. But i'm also tired of training models, so i may not.",
          "votes": 5
        },
        {
          "id": 1762491,
          "postDate": "2022-04-20T18:12:03.027Z",
          "content": "<p>Congratulations <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> ! Amazing push in the last week, too.</p>\n<p>Haha, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , yeah the motivation to retrain everything again just to check the score is not very high. We also thought about it after the competition. <br>\nWe did try pseudos at one point, but on a clean validation setup and only using pseudos for that validation set it wasn't doing much of a difference, thus we didn't look it at much more. I guess it only makes sense to use it with the test set included, at which point a validation score boost is not necessarily meaningful for test after a retraining. </p>",
          "rawMarkdown": "Congratulations @charmq @knshnb ! Amazing push in the last week, too.\n\nHaha, @cdeotte , yeah the motivation to retrain everything again just to check the score is not very high. We also thought about it after the competition. \nWe did try pseudos at one point, but on a clean validation setup and only using pseudos for that validation set it wasn't doing much of a difference, thus we didn't look it at much more. I guess it only makes sense to use it with the test set included, at which point a validation score boost is not necessarily meaningful for test after a retraining. ",
          "votes": 3
        },
        {
          "id": 1762511,
          "postDate": "2022-04-20T18:39:14.793Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> you guys dropped nearly the same in absolute score on private as we did, it might indeed be due to us both not using pseudos. I still do not understand why it would have more impact on private vs public though. Maybe QE alone would have worked on test, this also did not help on validation, but never tried it on a sub on test, and pseudo might have a similar effect.</p>",
          "rawMarkdown": "@cdeotte you guys dropped nearly the same in absolute score on private as we did, it might indeed be due to us both not using pseudos. I still do not understand why it would have more impact on private vs public though. Maybe QE alone would have worked on test, this also did not help on validation, but never tried it on a sub on test, and pseudo might have a similar effect.",
          "votes": 1
        },
        {
          "id": 1762520,
          "postDate": "2022-04-20T18:50:42.970Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> can you share your solution? I'm really want to know how was you able to achieve such a high score without it, thanks</p>",
          "rawMarkdown": "@philippsinger can you share your solution? I'm really want to know how was you able to achieve such a high score without it, thanks",
          "votes": 1
        },
        {
          "id": 1762526,
          "postDate": "2022-04-20T18:57:11.180Z",
          "content": "<p>Best performing were some sort-of siamese models on fullbody and backfin crops with adaptive margin increase and ArcFace. Also some more tricks to stabilize training, such as different learning rates for the cosine head and backbone to counter the extreme overfitting of the head. For backbones eca_nfnet_l2 worked best. We only trained on 512x512 for 10 epochs. We only joined late and did not have much compute power, so could not explore larger resolutions and models much.</p>",
          "rawMarkdown": "Best performing were some sort-of siamese models on fullbody and backfin crops with adaptive margin increase and ArcFace. Also some more tricks to stabilize training, such as different learning rates for the cosine head and backbone to counter the extreme overfitting of the head. For backbones eca_nfnet_l2 worked best. We only trained on 512x512 for 10 epochs. We only joined late and did not have much compute power, so could not explore larger resolutions and models much.",
          "votes": 5
        },
        {
          "id": 1762535,
          "postDate": "2022-04-20T19:11:46.230Z",
          "content": "<p>Thanks for sharing. So for different learning rate for cosine head and backbone, which should be higher? Can you share your params</p>",
          "rawMarkdown": "Thanks for sharing. So for different learning rate for cosine head and backbone, which should be higher? Can you share your params"
        },
        {
          "id": 1762541,
          "postDate": "2022-04-20T19:17:04.497Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> I think the distribution of <strong>test</strong> dataset individual id was much different than <strong>train</strong>. Specifically LB probes showed the most popular train individual ids (with ct = 100) did <strong>not</strong> appear in test. I also assume many individual id with count=1 in train may have had count=50+ in test.</p>\n<p>So I think that private test had more individual ids with count=1 in train. From 1st place solution, we see that pseudo boosted public LB +0.011 and private LB +0.016 which supports this idea.</p>\n<p>This is another reason that pseudo didn't help on CV. Because rare counts in CV train folds also had rare counts in CV valid folds. Our CV scheme did not mimic the train test relationship. Specially it is never the case that a valid fold had ct=50 for an individual id with ct=1 in train folds.</p>",
          "rawMarkdown": "@philippsinger I think the distribution of **test** dataset individual id was much different than **train**. Specifically LB probes showed the most popular train individual ids (with ct = 100) did **not** appear in test. I also assume many individual id with count=1 in train may have had count=50+ in test.\n\nSo I think that private test had more individual ids with count=1 in train. From 1st place solution, we see that pseudo boosted public LB +0.011 and private LB +0.016 which supports this idea.\n\nThis is another reason that pseudo didn't help on CV. Because rare counts in CV train folds also had rare counts in CV valid folds. Our CV scheme did not mimic the train test relationship. Specially it is never the case that a valid fold had ct=50 for an individual id with ct=1 in train folds.",
          "votes": 5
        },
        {
          "id": 1762557,
          "postDate": "2022-04-20T19:38:49.097Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> this is not exactly what we found. It is correct, that the distribution is different, but not too much. These two very frequent IDs do indeed appear on test. Just with a lower count (probably below 27), so the score is zero, when probing for them. We speculate, that they put a maximum frequency on the test IDs to discourage probing. You can prove the existance by replacing all these predictions with an impossible ID from another species. Your score will drop about 0.001 (when replacing both IDs).<br>\nOther than that, we had almost perfect agreement with CV, pub LB and private LB score, which does support equal or similar distribution.</p>",
          "rawMarkdown": "@cdeotte this is not exactly what we found. It is correct, that the distribution is different, but not too much. These two very frequent IDs do indeed appear on test. Just with a lower count (probably below 27), so the score is zero, when probing for them. We speculate, that they put a maximum frequency on the test IDs to discourage probing. You can prove the existance by replacing all these predictions with an impossible ID from another species. Your score will drop about 0.001 (when replacing both IDs).\nOther than that, we had almost perfect agreement with CV, pub LB and private LB score, which does support equal or similar distribution.",
          "votes": 1
        },
        {
          "id": 1762561,
          "postDate": "2022-04-20T19:41:24.760Z",
          "content": "<blockquote>\n  <p>I also assume many individual id with count=1 in train may have had count=50+ in test.</p>\n</blockquote>\n<p>We did not consider this though, this might definitely explain why pseudo works well on test. As said above, this should also mean though that QE on test should work well, I would be curious if anyone tried that.</p>",
          "rawMarkdown": "> I also assume many individual id with count=1 in train may have had count=50+ in test.\n\nWe did not consider this though, this might definitely explain why pseudo works well on test. As said above, this should also mean though that QE on test should work well, I would be curious if anyone tried that.",
          "votes": 2
        },
        {
          "id": 1765459,
          "postDate": "2022-04-23T14:23:45.930Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Thanks for your comments! In our experiment, pseudo labeling by an ensembled prediction for validation data improved a CV score of a single model. We did not test the effect of the ensemble of models trained by pseudo labeling in CV, though.<br>\nAnyway, I was amazed that you achieved such a high score without pseudo labeling with the image size 512! I would be very interested if you could elaborate more on your solution in discussion or somewhere.</p>",
          "rawMarkdown": "@philippsinger Thanks for your comments! In our experiment, pseudo labeling by an ensembled prediction for validation data improved a CV score of a single model. We did not test the effect of the ensemble of models trained by pseudo labeling in CV, though.\nAnyway, I was amazed that you achieved such a high score without pseudo labeling with the image size 512! I would be very interested if you could elaborate more on your solution in discussion or somewhere."
        },
        {
          "id": 1768157,
          "postDate": "2022-04-26T03:04:08.343Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> I took two models from our best (19th place) submission ensemble of 12 models and retrained them with pseudo labels. The ensemble's public LB boost <strong>+0.009</strong> and the private LB boost <strong>+0.016</strong>. So the private LB benefit more from pseudo. It looks like pseudo is very powerful in this competition!</p>",
          "rawMarkdown": "@philippsinger @ilu000 I took two models from our best (19th place) submission ensemble of 12 models and retrained them with pseudo labels. The ensemble's public LB boost **+0.009** and the private LB boost **+0.016**. So the private LB benefit more from pseudo. It looks like pseudo is very powerful in this competition!",
          "votes": 2
        },
        {
          "id": 1768586,
          "postDate": "2022-04-26T12:53:21.350Z",
          "content": "<p>ooof, thanks for testing</p>",
          "rawMarkdown": "ooof, thanks for testing",
          "votes": 1
        }
      ]
    },
    {
      "id": 1762334,
      "postDate": "2022-04-20T15:48:16.033Z",
      "content": "<p>Congratulations on finishing first!</p>\n<blockquote>\n  <p>For train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. </p>\n</blockquote>\n<p>Do you mean for a given image id, there's a 60% chance you're training with the fullbody crop, 15% chance with the fullbody_charm crop, and so on?</p>\n<blockquote>\n  <p>Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins.</p>\n</blockquote>\n<p>If the backfin bbox helps a lot, why isn't it assigned a ratio much higher than the 0.05 above?</p>\n<blockquote>\n  <p>We labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.</p>\n</blockquote>\n<p>Is \"new_individual\" only ever placed in the top 2 in the final list of predicted individuals?</p>",
      "rawMarkdown": "Congratulations on finishing first!\n\n> For train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. \n\nDo you mean for a given image id, there's a 60% chance you're training with the fullbody crop, 15% chance with the fullbody_charm crop, and so on?\n\n> Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins.\n\nIf the backfin bbox helps a lot, why isn't it assigned a ratio much higher than the 0.05 above?\n\n> We labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.\n\nIs \"new_individual\" only ever placed in the top 2 in the final list of predicted individuals?",
      "votes": 3,
      "replies": [
        {
          "id": 1762550,
          "postDate": "2022-04-20T19:24:16.197Z",
          "content": "<p>fullbody 0.60 :fullbody_charm 0.15 :backfin 0.15 :detic 0.05 :none 0.05 <br>\nbackfin is assigned 0.15</p>",
          "rawMarkdown": "fullbody 0.60 :fullbody_charm 0.15 :backfin 0.15 :detic 0.05 :none 0.05 \nbackfin is assigned 0.15"
        },
        {
          "id": 1764038,
          "postDate": "2022-04-22T05:18:12.853Z",
          "content": "<p>I would like to know these things, too!</p>",
          "rawMarkdown": "I would like to know these things, too!"
        },
        {
          "id": 1765238,
          "postDate": "2022-04-23T09:28:28.410Z",
          "content": "<p>You are right on the bbox mix. As <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> pointed out, backfin is assigned 15%.<br>\nFor \"new_individual\", we just inserted the prediction of \"new_individual\" as a certain threshold, so it appears as the 3rd, 4th, and 5th predictions as well.</p>",
          "rawMarkdown": "You are right on the bbox mix. As @cdeotte pointed out, backfin is assigned 15%.\nFor \"new_individual\", we just inserted the prediction of \"new_individual\" as a certain threshold, so it appears as the 3rd, 4th, and 5th predictions as well.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1762125,
      "postDate": "2022-04-20T12:59:32.107Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a>, and thanks for sharing amazing solution.<br>\nI'm proud to have you both as my colleagues, and looking forward to fighting together at another competitions !! </p>",
      "rawMarkdown": "Congratulations @charmq @knshnb, and thanks for sharing amazing solution.\nI'm proud to have you both as my colleagues, and looking forward to fighting together at another competitions !! ",
      "votes": 4,
      "replies": [
        {
          "id": 1765460,
          "postDate": "2022-04-23T14:24:53.973Z",
          "content": "<p>Thanks for your comments. Congratulations on your solo gold, too!!!</p>",
          "rawMarkdown": "Thanks for your comments. Congratulations on your solo gold, too!!!"
        }
      ]
    },
    {
      "id": 1793564,
      "postDate": "2022-05-18T02:43:19.490Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!",
      "votes": 1
    },
    {
      "id": 1793065,
      "postDate": "2022-05-17T15:10:46.847Z",
      "content": "<p>Congratulations~</p>",
      "rawMarkdown": "Congratulations~",
      "votes": 1
    },
    {
      "id": 1775816,
      "postDate": "2022-05-03T12:31:42.843Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!",
      "votes": 1
    },
    {
      "id": 1773859,
      "postDate": "2022-05-01T14:14:12.623Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": 1
    },
    {
      "id": 1773837,
      "postDate": "2022-05-01T13:30:15.937Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": 1
    },
    {
      "id": 1773149,
      "postDate": "2022-04-30T23:21:34.440Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": 1
    },
    {
      "id": 1773090,
      "postDate": "2022-04-30T21:47:36.487Z",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": 1
    },
    {
      "id": 1772796,
      "postDate": "2022-04-30T15:25:26.530Z",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a>  thanks for sharing👌👌👌</p>",
      "rawMarkdown": "Congratulations @knshnb  thanks for sharing👌👌👌",
      "votes": 1
    },
    {
      "id": 1772609,
      "postDate": "2022-04-30T11:46:29.997Z",
      "content": "<p>Congratulations! </p>",
      "rawMarkdown": "Congratulations! ",
      "votes": 1
    },
    {
      "id": 1772441,
      "postDate": "2022-04-30T08:46:49.517Z",
      "content": "<p>Congratulations, so best models</p>",
      "rawMarkdown": "Congratulations, so best models",
      "votes": 1
    },
    {
      "id": 1772186,
      "postDate": "2022-04-30T00:22:26.033Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": 1
    },
    {
      "id": 1771932,
      "postDate": "2022-04-29T16:50:34.310Z",
      "content": "<p>congratulations! Thank you for sharing your methodologies!! </p>",
      "rawMarkdown": "congratulations! Thank you for sharing your methodologies!! ",
      "votes": 1
    },
    {
      "id": 1771457,
      "postDate": "2022-04-29T08:22:36.017Z",
      "content": "<p>Congratulations</p>",
      "rawMarkdown": "Congratulations",
      "votes": 1
    },
    {
      "id": 1771428,
      "postDate": "2022-04-29T07:43:23.513Z",
      "content": "<p>Thanks a lot, that was very clear and helpful, got many fresh ideas</p>",
      "rawMarkdown": "Thanks a lot, that was very clear and helpful, got many fresh ideas",
      "votes": 1
    },
    {
      "id": 1770976,
      "postDate": "2022-04-28T19:36:06.317Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!",
      "votes": 1
    },
    {
      "id": 1770932,
      "postDate": "2022-04-28T18:36:53.690Z",
      "content": "<p>Nice work. Congratulations for the win!</p>",
      "rawMarkdown": "Nice work. Congratulations for the win!\n",
      "votes": 1
    },
    {
      "id": 1770316,
      "postDate": "2022-04-28T07:03:44.683Z",
      "content": "<p>good job!!!</p>",
      "rawMarkdown": "good job!!!",
      "votes": 1
    },
    {
      "id": 1768687,
      "postDate": "2022-04-26T14:58:15.210Z",
      "content": "<p>Congratulations for the first place! Thanks for sharing the details as well. </p>",
      "rawMarkdown": "Congratulations for the first place! Thanks for sharing the details as well. ",
      "votes": 1
    },
    {
      "id": 1768038,
      "postDate": "2022-04-25T22:55:05.997Z",
      "content": "<p>Congratulations for the first place!!!</p>",
      "rawMarkdown": "Congratulations for the first place!!!",
      "votes": 1
    },
    {
      "id": 1767789,
      "postDate": "2022-04-25T16:30:03.253Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": 1
    },
    {
      "id": 1766110,
      "postDate": "2022-04-24T07:50:49.860Z",
      "content": "<p>Congratulations for the first place!</p>",
      "rawMarkdown": "Congratulations for the first place!",
      "votes": 1
    },
    {
      "id": 1766001,
      "postDate": "2022-04-24T05:27:42.967Z",
      "content": "<p>congratulations! your share was really helpfull!!</p>",
      "rawMarkdown": "congratulations! your share was really helpfull!!",
      "votes": 1
    },
    {
      "id": 1765869,
      "postDate": "2022-04-24T01:56:36.313Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": 1
    },
    {
      "id": 1765860,
      "postDate": "2022-04-24T01:41:31.467Z",
      "content": "<p>congratulations!</p>",
      "rawMarkdown": "congratulations!",
      "votes": 1
    },
    {
      "id": 1765548,
      "postDate": "2022-04-23T16:23:41.590Z",
      "content": "<p>amazing solution!</p>",
      "rawMarkdown": "amazing solution!",
      "votes": 1
    },
    {
      "id": 1765494,
      "postDate": "2022-04-23T15:04:47.603Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": 1
    },
    {
      "id": 1765428,
      "postDate": "2022-04-23T13:42:41.113Z",
      "content": "<p>Your share is very helpful for me!!</p>",
      "rawMarkdown": "Your share is very helpful for me!!",
      "votes": 1
    },
    {
      "id": 1765243,
      "postDate": "2022-04-23T09:31:42.770Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": 1
    },
    {
      "id": 1765146,
      "postDate": "2022-04-23T08:18:54.267Z",
      "content": "<p>Thanks for this nice dataset.I have successfully moved from novice to Contributor</p>",
      "rawMarkdown": "Thanks for this nice dataset.I have successfully moved from novice to Contributor",
      "votes": 1
    },
    {
      "id": 1764970,
      "postDate": "2022-04-23T04:05:34.967Z",
      "content": "<p>Great Job!</p>",
      "rawMarkdown": "Great Job!",
      "votes": 1
    },
    {
      "id": 1764866,
      "postDate": "2022-04-22T23:51:19.367Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations",
      "votes": 1
    },
    {
      "id": 1764277,
      "postDate": "2022-04-22T10:34:05.787Z",
      "content": "<p>congratulations</p>",
      "rawMarkdown": "congratulations\n",
      "votes": 1
    },
    {
      "id": 1764068,
      "postDate": "2022-04-22T05:47:54.980Z",
      "content": "<p>Congratulations for the first place!</p>",
      "rawMarkdown": "Congratulations for the first place!",
      "votes": 1
    },
    {
      "id": 1763946,
      "postDate": "2022-04-22T02:04:54.403Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a>!!!<br>\nThanks for sharing.</p>",
      "rawMarkdown": "Congrats @charmq @knshnb!!!\nThanks for sharing.",
      "votes": 1
    },
    {
      "id": 1763891,
      "postDate": "2022-04-22T00:15:43.647Z",
      "content": "<p>Congratulation! <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> </p>",
      "rawMarkdown": "Congratulation! @knshnb ",
      "votes": 1
    },
    {
      "id": 1763524,
      "postDate": "2022-04-21T15:32:24.793Z",
      "content": "<p><a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> Can you share your experience with different lr at backbone and head? I tried it but it shows superior performance on the early epoch but far from the best single lr at the last epoch.</p>",
      "rawMarkdown": "@knshnb Can you share your experience with different lr at backbone and head? I tried it but it shows superior performance on the early epoch but far from the best single lr at the last epoch.",
      "votes": 1,
      "replies": [
        {
          "id": 1763788,
          "postDate": "2022-04-21T20:10:39.460Z",
          "content": "<p><code>we adapted the sub-center ArcFace of k=2 with the usual flip data augmentation.</code> and how about this part. I didn't see you use flip in your augmentation?</p>",
          "rawMarkdown": "`we adapted the sub-center ArcFace of k=2 with the usual flip data augmentation.` and how about this part. I didn't see you use flip in your augmentation?"
        },
        {
          "id": 1765248,
          "postDate": "2022-04-23T09:33:43.230Z",
          "content": "<p>I'm sorry that we can not remember the detail, but different lrs enhanced our CV scores significantly in both of our pipelines.<br>\nThank you for pointing out the flip augmentation! We just forgot to add that to the list. Updated.</p>",
          "rawMarkdown": "I'm sorry that we can not remember the detail, but different lrs enhanced our CV scores significantly in both of our pipelines.\nThank you for pointing out the flip augmentation! We just forgot to add that to the list. Updated."
        },
        {
          "id": 1765389,
          "postDate": "2022-04-23T12:41:56.413Z",
          "content": "<p><a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> thanks for your sharing. Can you explain more about your second species head? what margin did you choose and what is the implementation? I tried to add it as follow: backbone -&gt; features  (head_1) -&gt; bn_2 -&gt; arcface (id), features -&gt; (head_2) -&gt; bn_2 -&gt; arcface (species).</p>",
          "rawMarkdown": "@knshnb thanks for your sharing. Can you explain more about your second species head? what margin did you choose and what is the implementation? I tried to add it as follow: backbone -> features  (head_1) -> bn_2 -> arcface (id), features -> (head_2) -> bn_2 -> arcface (species)."
        },
        {
          "id": 1765547,
          "postDate": "2022-04-23T16:22:11.237Z",
          "content": "<p>backbone -&gt; bn -&gt; arcface (id), arcface (species)<br>\nPlease also refer to the source code, which we are planning to release soon.</p>",
          "rawMarkdown": "backbone -> bn -> arcface (id), arcface (species)\nPlease also refer to the source code, which we are planning to release soon.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1763431,
      "postDate": "2022-04-21T14:19:09.210Z",
      "content": "<p>congratulation!! <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> </p>",
      "rawMarkdown": "congratulation!! @charmq @knshnb ",
      "votes": 1
    },
    {
      "id": 1762514,
      "postDate": "2022-04-20T18:40:59.017Z",
      "content": "<p>Congrats! efficientnet_b7 on bs=16,  res=1024x1024, gpu=v100(32) ?  Should this fit? How long did it take?</p>",
      "rawMarkdown": "Congrats! efficientnet_b7 on bs=16,  res=1024x1024, gpu=v100(32) ?  Should this fit? How long did it take?",
      "votes": 1,
      "replies": [
        {
          "id": 1765255,
          "postDate": "2022-04-23T09:35:21.773Z",
          "content": "<p>Thanks! Using fp16 training, that case was fitted in 4x V100 (32GB).</p>",
          "rawMarkdown": "Thanks! Using fp16 training, that case was fitted in 4x V100 (32GB)."
        }
      ]
    },
    {
      "id": 1762216,
      "postDate": "2022-04-20T14:13:48.217Z",
      "content": "<p>Congrats! Pseudo really boost a lot!</p>",
      "rawMarkdown": "Congrats! Pseudo really boost a lot!",
      "votes": 1
    },
    {
      "id": 1763189,
      "postDate": "2022-04-21T10:29:38.087Z",
      "content": "<p>congratulations <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> </p>",
      "rawMarkdown": "congratulations @charmq ",
      "votes": 2
    },
    {
      "id": 1762605,
      "postDate": "2022-04-20T20:21:03.627Z",
      "content": "<p>Congrats <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> for the first place!!!</p>",
      "rawMarkdown": "Congrats @charmq @knshnb for the first place!!!",
      "votes": 2,
      "replies": [
        {
          "id": 1765462,
          "postDate": "2022-04-23T14:26:29.857Z",
          "content": "<p>Thank you!!!</p>",
          "rawMarkdown": "Thank you!!!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1762117,
      "postDate": "2022-04-20T12:51:51.260Z",
      "content": "<p>Congratulations!! </p>",
      "rawMarkdown": "Congratulations!! ",
      "votes": 2
    },
    {
      "id": 2404751,
      "postDate": "2023-08-23T13:45:24.660Z",
      "content": "<p>I tried to reproduce your score on leaderboard with the repo instructions you share on github but unfortunately was not able to do so.  I trained your architecture for 30 epochs, then chose 'pred_idx' as predictions taken from test_fullbody_results.npz. For inference I chose logit threshold for 'new_individual' to be 0.32. The number of 1st predictions of 'new_individual' ratio was 0.17 (close to 0.165 as you specified). The score I get with the resulting submissionon on kaggle is only 0.35. Am I missing something?</p>",
      "rawMarkdown": "I tried to reproduce your score on leaderboard with the repo instructions you share on github but unfortunately was not able to do so.  I trained your architecture for 30 epochs, then chose 'pred_idx' as predictions taken from test_fullbody_results.npz. For inference I chose logit threshold for 'new_individual' to be 0.32. The number of 1st predictions of 'new_individual' ratio was 0.17 (close to 0.165 as you specified). The score I get with the resulting submissionon on kaggle is only 0.35. Am I missing something?"
    },
    {
      "id": 2164997,
      "postDate": "2023-03-01T23:06:02.390Z",
      "content": "<p>Excellent!!!, this new tool will be of crucial help for all developers whether they are experienced, teams or newbies. See option  in the Kaggle menu, just below .</p>",
      "rawMarkdown": "Excellent!!!, this new tool will be of crucial help for all developers whether they are experienced, teams or newbies. See option <Models> in the Kaggle menu, just below <DataSets>."
    },
    {
      "id": 1782046,
      "postDate": "2022-05-09T07:21:19.303Z",
      "content": "<p>I want to know some questions about optuna search. Do you use all data for training or some data for sampling.</p>",
      "rawMarkdown": "I want to know some questions about optuna search. Do you use all data for training or some data for sampling.",
      "replies": [
        {
          "id": 1790169,
          "postDate": "2022-05-14T15:44:08.837Z",
          "content": "<p>We optimized a single-fold validation score on all train data by Optuna search. As I explained above, we used the small model and images so that training completes quickly on a single V100 (16GB). I might not remember accurately, but I think it took around one day to run 100+ trials by 8 processes in parallel.</p>",
          "rawMarkdown": "We optimized a single-fold validation score on all train data by Optuna search. As I explained above, we used the small model and images so that training completes quickly on a single V100 (16GB). I might not remember accurately, but I think it took around one day to run 100+ trials by 8 processes in parallel."
        }
      ]
    },
    {
      "id": 1766112,
      "postDate": "2022-04-24T07:51:38.733Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1766422,
      "postDate": "2022-04-24T14:48:44.060Z",
      "content": "<p>Thanks for sharing, congratulations!</p>",
      "rawMarkdown": "Thanks for sharing, congratulations!",
      "votes": 1
    },
    {
      "id": 1766367,
      "postDate": "2022-04-24T13:31:19.160Z",
      "content": "<p>Thanks for sharing. <br>\nCongratulation!</p>",
      "rawMarkdown": "Thanks for sharing. \nCongratulation!",
      "votes": 1
    },
    {
      "id": 1766253,
      "postDate": "2022-04-24T11:02:36.013Z",
      "content": "<p>Thanks for this dataset</p>",
      "rawMarkdown": "Thanks for this dataset",
      "votes": 1
    },
    {
      "id": 1766102,
      "postDate": "2022-04-24T07:47:10.537Z",
      "content": "<p>This was very helpful. Thanks!</p>",
      "rawMarkdown": "This was very helpful. Thanks!",
      "votes": 1
    },
    {
      "id": 1763020,
      "postDate": "2022-04-21T08:01:19.867Z",
      "content": "<p>Thanks for sharing, I learned a lot!</p>",
      "rawMarkdown": "Thanks for sharing, I learned a lot!",
      "votes": 1
    },
    {
      "id": 1762181,
      "postDate": "2022-04-20T13:41:28.973Z",
      "content": "<p>Thanks for  Sharing solution!!</p>",
      "rawMarkdown": "Thanks for  Sharing solution!!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1762120,
      "author_name": "charmq",
      "author_url": "",
      "post_date": "2022-04-20T12:52:52.323000",
      "content": "<p>Thanks for teaming up with me! <br>\nI can't believe you were novice. Amazing performance!</p>",
      "votes": 15,
      "replies": [
        {
          "id": 1765430,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2022-04-23T13:43:43.187000",
          "content": "<p>Thank you for merging with me too. It was a great experience working with you in my first competition!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1770239,
      "author_name": "SaiNithish",
      "author_url": "",
      "post_date": "2022-04-28T05:41:38.980000",
      "content": "<p>nice, good job</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 1762190,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2022-04-20T13:55:08.980000",
      "content": "<p>Big congrats! Amazing job!</p>\n<p>At least it seems we were the only team to reach 0.890 without any pseudo labeling :D<br>\nMight be the reason we dropped a bit more than others on private. It's crazy to see that everyone used it with success.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 1762214,
          "author_name": "charmq",
          "author_url": "",
          "post_date": "2022-04-20T14:13:14.247000",
          "content": "<p>I agree with you. I think pseudo labeling is a risky method in some sense because we can't see the improvement without submission. Actually, we tried pseudo labeling in the last 2 days. We saw the improvement in the lb score and were convinced that everyone else was using pseudo labeling.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1762343,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-04-20T15:59:45.347000",
          "content": "<p>Yeah, I was even expecting you to use pseudo as you improved step by step. But for us it did not work when trying it on CV, so we dropped it. And as we had perfect correlation between CV and LB saw no reason to try it on LB. But we should have :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1762401,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-20T16:47:32.580000",
          "content": "<p>Wow <a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Amazing model without pseudo labels! Our team also did not use pseudo labels either. And after reading all the top solutions, i realize this is what held us back. </p>\n<p>I read solutions from last year and it seemed that it did not help last year, so we didn't try it. Upon reflection, this year has 30k test images whereas last year had only 5k. And this year half (8k) of the individual id had only 1 image where last year was more dense. So it makes sense why it would be important this year and not last year.</p>\n<p>I'm so curious how much it would improve our pipeline. I'm thinking training all our models using pseudo today. But i'm also tired of training models, so i may not.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1762491,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2022-04-20T18:12:03.027000",
          "content": "<p>Congratulations <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a> ! Amazing push in the last week, too.</p>\n<p>Haha, <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> , yeah the motivation to retrain everything again just to check the score is not very high. We also thought about it after the competition. <br>\nWe did try pseudos at one point, but on a clean validation setup and only using pseudos for that validation set it wasn't doing much of a difference, thus we didn't look it at much more. I guess it only makes sense to use it with the test set included, at which point a validation score boost is not necessarily meaningful for test after a retraining. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1762511,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-04-20T18:39:14.793000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> you guys dropped nearly the same in absolute score on private as we did, it might indeed be due to us both not using pseudos. I still do not understand why it would have more impact on private vs public though. Maybe QE alone would have worked on test, this also did not help on validation, but never tried it on a sub on test, and pseudo might have a similar effect.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1762520,
          "author_name": "cuongnn",
          "author_url": "",
          "post_date": "2022-04-20T18:50:42.970000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> can you share your solution? I'm really want to know how was you able to achieve such a high score without it, thanks</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1762526,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-04-20T18:57:11.180000",
          "content": "<p>Best performing were some sort-of siamese models on fullbody and backfin crops with adaptive margin increase and ArcFace. Also some more tricks to stabilize training, such as different learning rates for the cosine head and backbone to counter the extreme overfitting of the head. For backbones eca_nfnet_l2 worked best. We only trained on 512x512 for 10 epochs. We only joined late and did not have much compute power, so could not explore larger resolutions and models much.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1762535,
          "author_name": "cuongnn",
          "author_url": "",
          "post_date": "2022-04-20T19:11:46.230000",
          "content": "<p>Thanks for sharing. So for different learning rate for cosine head and backbone, which should be higher? Can you share your params</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1762541,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-20T19:17:04.497000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> I think the distribution of <strong>test</strong> dataset individual id was much different than <strong>train</strong>. Specifically LB probes showed the most popular train individual ids (with ct = 100) did <strong>not</strong> appear in test. I also assume many individual id with count=1 in train may have had count=50+ in test.</p>\n<p>So I think that private test had more individual ids with count=1 in train. From 1st place solution, we see that pseudo boosted public LB +0.011 and private LB +0.016 which supports this idea.</p>\n<p>This is another reason that pseudo didn't help on CV. Because rare counts in CV train folds also had rare counts in CV valid folds. Our CV scheme did not mimic the train test relationship. Specially it is never the case that a valid fold had ct=50 for an individual id with ct=1 in train folds.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1762557,
          "author_name": "Pascal Pfeiffer",
          "author_url": "",
          "post_date": "2022-04-20T19:38:49.097000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> this is not exactly what we found. It is correct, that the distribution is different, but not too much. These two very frequent IDs do indeed appear on test. Just with a lower count (probably below 27), so the score is zero, when probing for them. We speculate, that they put a maximum frequency on the test IDs to discourage probing. You can prove the existance by replacing all these predictions with an impossible ID from another species. Your score will drop about 0.001 (when replacing both IDs).<br>\nOther than that, we had almost perfect agreement with CV, pub LB and private LB score, which does support equal or similar distribution.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1762561,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-04-20T19:41:24.760000",
          "content": "<blockquote>\n  <p>I also assume many individual id with count=1 in train may have had count=50+ in test.</p>\n</blockquote>\n<p>We did not consider this though, this might definitely explain why pseudo works well on test. As said above, this should also mean though that QE on test should work well, I would be curious if anyone tried that.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1765459,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2022-04-23T14:23:45.930000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> Thanks for your comments! In our experiment, pseudo labeling by an ensembled prediction for validation data improved a CV score of a single model. We did not test the effect of the ensemble of models trained by pseudo labeling in CV, though.<br>\nAnyway, I was amazed that you achieved such a high score without pseudo labeling with the image size 512! I would be very interested if you could elaborate more on your solution in discussion or somewhere.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1768157,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-26T03:04:08.343000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> <a href=\"https://www.kaggle.com/ilu000\" target=\"_blank\">@ilu000</a> I took two models from our best (19th place) submission ensemble of 12 models and retrained them with pseudo labels. The ensemble's public LB boost <strong>+0.009</strong> and the private LB boost <strong>+0.016</strong>. So the private LB benefit more from pseudo. It looks like pseudo is very powerful in this competition!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1768586,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2022-04-26T12:53:21.350000",
          "content": "<p>ooof, thanks for testing</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1762334,
      "author_name": "wafflebufflo",
      "author_url": "",
      "post_date": "2022-04-20T15:48:16.033000",
      "content": "<p>Congratulations on finishing first!</p>\n<blockquote>\n  <p>For train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. </p>\n</blockquote>\n<p>Do you mean for a given image id, there's a 60% chance you're training with the fullbody crop, 15% chance with the fullbody_charm crop, and so on?</p>\n<blockquote>\n  <p>Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins.</p>\n</blockquote>\n<p>If the backfin bbox helps a lot, why isn't it assigned a ratio much higher than the 0.05 above?</p>\n<blockquote>\n  <p>We labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.</p>\n</blockquote>\n<p>Is \"new_individual\" only ever placed in the top 2 in the final list of predicted individuals?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1762550,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-20T19:24:16.197000",
          "content": "<p>fullbody 0.60 :fullbody_charm 0.15 :backfin 0.15 :detic 0.05 :none 0.05 <br>\nbackfin is assigned 0.15</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1764038,
          "author_name": "Brian Salamone",
          "author_url": "",
          "post_date": "2022-04-22T05:18:12.853000",
          "content": "<p>I would like to know these things, too!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1765238,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2022-04-23T09:28:28.410000",
          "content": "<p>You are right on the bbox mix. As <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> pointed out, backfin is assigned 15%.<br>\nFor \"new_individual\", we just inserted the prediction of \"new_individual\" as a certain threshold, so it appears as the 3rd, 4th, and 5th predictions as well.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1762125,
      "author_name": "Yiemon773",
      "author_url": "",
      "post_date": "2022-04-20T12:59:32.107000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/charmq\" target=\"_blank\">@charmq</a> <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a>, and thanks for sharing amazing solution.<br>\nI'm proud to have you both as my colleagues, and looking forward to fighting together at another competitions !! </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1765460,
          "author_name": "knshnb",
          "author_url": "",
          "post_date": "2022-04-23T14:24:53.973000",
          "content": "<p>Thanks for your comments. Congratulations on your solo gold, too!!!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1793564,
      "author_name": "HPEVERYDAY",
      "author_url": "",
      "post_date": "2022-05-18T02:43:19.490000",
      "content": "<p>congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1793065,
      "author_name": "Aubrey12138",
      "author_url": "",
      "post_date": "2022-05-17T15:10:46.847000",
      "content": "<p>Congratulations~</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1775816,
      "author_name": "Tarek Mohamed Amin",
      "author_url": "",
      "post_date": "2022-05-03T12:31:42.843000",
      "content": "<p>congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1773859,
      "author_name": "KUNAL KASHYAP",
      "author_url": "",
      "post_date": "2022-05-01T14:14:12.623000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1773837,
      "author_name": "Ervin Macic",
      "author_url": "",
      "post_date": "2022-05-01T13:30:15.937000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1773149,
      "author_name": "Prosper Alikizang",
      "author_url": "",
      "post_date": "2022-04-30T23:21:34.440000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1773090,
      "author_name": "Ed Al. Wang",
      "author_url": "",
      "post_date": "2022-04-30T21:47:36.487000",
      "content": "<p>Congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1772796,
      "author_name": "Alex Parkhomenko",
      "author_url": "",
      "post_date": "2022-04-30T15:25:26.530000",
      "content": "<p>Congratulations <a href=\"https://www.kaggle.com/knshnb\" target=\"_blank\">@knshnb</a>  thanks for sharing👌👌👌</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1772609,
      "author_name": "Lakshya Bansal",
      "author_url": "",
      "post_date": "2022-04-30T11:46:29.997000",
      "content": "<p>Congratulations! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1772441,
      "author_name": "TeguhPermana",
      "author_url": "",
      "post_date": "2022-04-30T08:46:49.517000",
      "content": "<p>Congratulations, so best models</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1772186,
      "author_name": "Edson Apari Bacilio",
      "author_url": "",
      "post_date": "2022-04-30T00:22:26.033000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1771932,
      "author_name": "msmccann10",
      "author_url": "",
      "post_date": "2022-04-29T16:50:34.310000",
      "content": "<p>congratulations! Thank you for sharing your methodologies!! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1771457,
      "author_name": "ElMartian",
      "author_url": "",
      "post_date": "2022-04-29T08:22:36.017000",
      "content": "<p>Congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1771428,
      "author_name": "Alexander Shein",
      "author_url": "",
      "post_date": "2022-04-29T07:43:23.513000",
      "content": "<p>Thanks a lot, that was very clear and helpful, got many fresh ideas</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1770976,
      "author_name": "DanilovEgor",
      "author_url": "",
      "post_date": "2022-04-28T19:36:06.317000",
      "content": "<p>congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1770932,
      "author_name": "Mghoi Mwadime",
      "author_url": "",
      "post_date": "2022-04-28T18:36:53.690000",
      "content": "<p>Nice work. Congratulations for the win!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1770316,
      "author_name": "Feng Qilong",
      "author_url": "",
      "post_date": "2022-04-28T07:03:44.683000",
      "content": "<p>good job!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1768687,
      "author_name": "Teja Surya",
      "author_url": "",
      "post_date": "2022-04-26T14:58:15.210000",
      "content": "<p>Congratulations for the first place! Thanks for sharing the details as well. </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1768038,
      "author_name": "k_ohmori",
      "author_url": "",
      "post_date": "2022-04-25T22:55:05.997000",
      "content": "<p>Congratulations for the first place!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1767789,
      "author_name": "cp11122",
      "author_url": "",
      "post_date": "2022-04-25T16:30:03.253000",
      "content": "<p>congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1766110,
      "author_name": "Periklis Drakousis",
      "author_url": "",
      "post_date": "2022-04-24T07:50:49.860000",
      "content": "<p>Congratulations for the first place!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1766001,
      "author_name": "Hetvi Bhora",
      "author_url": "",
      "post_date": "2022-04-24T05:27:42.967000",
      "content": "<p>congratulations! your share was really helpfull!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765869,
      "author_name": "wang.dakun",
      "author_url": "",
      "post_date": "2022-04-24T01:56:36.313000",
      "content": "<p>congratulations</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765860,
      "author_name": "primerbi",
      "author_url": "",
      "post_date": "2022-04-24T01:41:31.467000",
      "content": "<p>congratulations!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765548,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T16:23:41.590000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765494,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T15:04:47.603000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765428,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T13:42:41.113000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765243,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T09:31:42.770000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765146,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T08:18:54.267000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1764970,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-23T04:05:34.967000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1764866,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-22T23:51:19.367000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1764277,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-22T10:34:05.787000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1764068,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-22T05:47:54.980000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1763946,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-22T02:04:54.403000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1763891,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-22T00:15:43.647000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1763524,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-21T15:32:24.793000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1763788,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-21T20:10:39.460000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1765248,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-23T09:33:43.230000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1765389,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-23T12:41:56.413000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1765547,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-23T16:22:11.237000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1763431,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-21T14:19:09.210000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1762514,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-20T18:40:59.017000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 1765255,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-23T09:35:21.773000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1762216,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-20T14:13:48.217000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1763189,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-21T10:29:38.087000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1762605,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-20T20:21:03.627000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1765462,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-04-23T14:26:29.857000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1762117,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-20T12:51:51.260000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2404751,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-08-23T13:45:24.660000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2164997,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-01T23:06:02.390000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1782046,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-09T07:21:19.303000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 1790169,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-14T15:44:08.837000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1766112,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-24T07:51:38.733000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1766422,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-24T14:48:44.060000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1766367,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-24T13:31:19.160000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1766253,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-24T11:02:36.013000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1766102,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-24T07:47:10.537000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1763020,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-21T08:01:19.867000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1762181,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-20T13:41:28.973000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1762107": "This was my first participation in a Kaggle competition and I was so fortunate to win 1st place!!!!\nI appreciate my team member @charmq!\n\nWe really enjoyed the competition and worked hard literally until the last minutes. We would like to deeply appreciate the Kaggle staff for organizing this great competition and other teams for competing with us.\n\n## Overview\nOur solution is based on sub-center ArcFace with Dynamic margins, which was shown to be effective in [Google Landmark Recognition 2020 the 3rd place solution](https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757)  by @boliu0 and [the first place solution of Google Landmark Recognition 2021](https://www.kaggle.com/c/landmark-recognition-2021/discussion/277098) by @christofhenkel.\n\nBasically, our solution was an ensemble of two pipelines implemented by @charmq and me. Since we shared knowledge with each other during the competition, our pipelines share much in common. Below, I will mainly explain my pipeline, which was slightly better in CV score. \n\n## Key Points\n- Tuning of dynamic margin hyperparameters with [Optuna](https://optuna.org/)\n- Larger learning rate for ArcFace head\n- Bounding box mixing augmentation\n- Ensemble of knn and logit\n- Two-round pseudo labeling\n- Ensemble of many models\n\n## Dataset\nWe used several types of bounding boxes to crop images. We thank @jpbremer and @phalanx a lot for providing valuable datasets. We also trained our own yolov5 model using fullbody annotations, which we refer to as fullbody_charm.\n\nFor train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins. Adding non-cropped images by a small ratio also worked as a regularization. For test data, we took the mean of predictions between fullbody and fullbody_charm.\n\nWe used the cropped images resized to a fixed size. We mainly used the image size of (1024, 1024). Some models were trained with the image size of (1200, 1200) and (1440, 1440) for ensembling.\n\n## Backbone & Neck\nWe trained several different imagenet-pretrained backbones for ensembling (efficientnet_b5, efficientnet_b6, efficientnet_b7, efficientnetv2_m, efficientnetv2_l, etc). The best performance in a single model was achieved by efficientnet_b7.\n\nUsing GeM pooling (p=3) instead of GAP enhanced the performance.\n\nThe normalization layer before the ArcFace head was important. Batchnorm was slightly better than Layernorm in our experiments.\n\nIn addition to the final feature map of the backbone, we used the second final feature map to capture more local information. We simply concatenated those two GeM-pooled feature maps and passed them to head.\n\n## Head\nFor handling imbalanced classes, we adopted ArcFace with dynamic margins. Since it seemed sensitive to hyperparameters, we tuned them on images of (256, 256) and efficientnet_b0 using Optuna. It seemed the acquired hyperparameters also worked well on large images and architectures.\n\nIn the last competition, it was reported that handling flipped images as different classes significantly enhanced the performance. In this competition, we did not think that this technique works well because some images are taken from different angles. To handle this issue, we adapted the sub-center ArcFace of k=2 with the usual flip data augmentation.\n\nWe also added a second head for classifying species. Sub-center ArcFace with dynamic margins worked better than simple Linear head.\n\n## Training\nWhile we mainly checked a single-fold validation score locally, we trained our models by using whole train data for submission.\nSetting the learning rate of the head 10 times bigger than the learning rate of the backbone significantly improved the performance.\nOptimal training settings of us differed possibly due to slight differences in our pipelines. While I trained the models for 30 epochs by AdamW optimizer of lr_backbone=1.6e-3 with warmup cosine annealing scheduler, charmq trained the models for 20 epochs by Adam of lr_backbone=1e-4 with cosine annealing scheduler. Most of the models were trained with the batch size of 16-32 on 2-8x NVIDIA Tesla V100 (32GB).\n\nIn addition to the bounding box mix augmentation described above, we adopted many data augmentations because the models were expressive enough to reach almost 100% of training accuracy. Below are the list of data augmentations we used by Albumentations implementation:\n```\nA.Affine(rotate=(-15, 15), translate_percent=(0.0, 0.25), shear=(-3, 3), p=0.5),\nA.RandomResizedCrop(image_size[0], image_size[1], scale=(0.9, 1.0), ratio=(0.75, 1.3333333333)),\nA.ToGray(p=0.1),\nA.GaussianBlur(blur_limit=(3, 7), p=0.05),\nA.GaussNoise(p=0.05),\nA.RandomGridShuffle(grid=(2, 2), p=0.3),\nA.Posterize(p=0.2),\nA.RandomBrightnessContrast(p=0.5),\nA.Cutout(p=0.05),\nA.RandomSnow(p=0.1),\nA.RandomRain(p=0.05),\nA.HorizontalFlip(p=0.5),\n```\n\n## Postprocess\nWe combined the following two metrics.\nknn: Using feature vectors, we calculated the largest cosine similarity for each class in training data. This can be done for each model and is easier to ensemble many models than feature concatenation. In implementation, we looked at only the top 500 training data for each test individual by `sklearn.neighbors.NearestNeighbors`. As a test-time augmentation, we took the mean of original test images and flipped test images searched over both original train images and flipped train images.\nlogit: We used the simple output of the model without margins.\n(We calculated the mean of predictions for two bounding boxes as explained above.) While knn worked much better in CV but only slightly better in the public leaderboard than logit. This is probably caused by highly imbalanced data and the distribution differences between train and test (knn is more likely to output classes with more train data). To mitigate this, we mixed the prediction of knn and logit with knn_ratio=0.5. After pseudo labeling, we increased the knn_ratio to 0.8.\n\nWe labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.\n\n## Other\nPseudo labeling enhanced the public LB score a lot in this competition probably because of extremely imbalanced data. We stuck to improving CV scores until the very final stage of the competition. On the day before the deadline, we got a big boost in the leaderboard score (0.88589/0.85959 -> 0.89343/0.87062) by a pseudo-label submission. The second round of pseudo labeling on the final day also improved the score (0.89680/0.87579). Additional rounds of pseudo labeling might have further improved performance, but unfortunately, we did not have time to do that.\n\nAs a final submission, we ensembled around 50 models including the ones trained without pseudo labels and with the first-round pseudo labels. After the competition, we confirmed that the ensemble of only 2 models (the best one from each of us) scored 0.89385/0.87336, which could still win first place!\n\n## What did not work\n- input 4-channel images with segmentation mask (1st place solution of the last competition)\n- input 6-channel images combining 2 types of images cropped by fullbody and backfin bboxes\n- input rectangle images such as (512, 1024)\n- maintain aspect ratio of images by bounding box expansion instead of resizing\n- AdaCos, triplet loss\n- Focal loss\n- training p of GeM pooling\n- ConvNeXt\n- Swin Transformer (384 was too small)\n- dolg\n- using pseudo labels for knn\n- cutmix\n\n## Acknowledgement\nWe deeply acknowledge great OSS such as PyTorch, PyTorch Lightning, PyTorch Image Models, Albumentations, etc. We would also like to appreciate Preferred Networks, Inc for allowing us to use computational resources.\n\n[Update 2022/05/24]\nI added explanations on some minor points.\ncode (knshnb): https://github.com/knshnb/kaggle-happywhale-1st-place\ncode (charmq): https://github.com/tyamaguchi17/kaggle-happywhale-1st-place-solution-charmq",
    "1762120": "Thanks for teaming up with me! \nI can't believe you were novice. Amazing performance!",
    "1770239": "nice, good job",
    "1762190": "Big congrats! Amazing job!\n\nAt least it seems we were the only team to reach 0.890 without any pseudo labeling :D\nMight be the reason we dropped a bit more than others on private. It's crazy to see that everyone used it with success.",
    "1762334": "Congratulations on finishing first!\n\n> For train data, we randomly mixed several bboxes with the ratio of fullbody:fullbody_charm:backfin:detic:none=0.60:0.15:0.15:0.05:0.05. \n\nDo you mean for a given image id, there's a 60% chance you're training with the fullbody crop, 15% chance with the fullbody_charm crop, and so on?\n\n> Especially, combining backfin bbox to train data significantly improved the performance possibly because it enhances the robustness to images that only contain backfins.\n\nIf the backfin bbox helps a lot, why isn't it assigned a ratio much higher than the 0.05 above?\n\n> We labeled individuals whose ensembled predictions are lower than a certain threshold as “new_individual”. Through several submission trials, we decided to set the threshold so that the ratio of “new_individual” as the first prediction is 0.165.\n\nIs \"new_individual\" only ever placed in the top 2 in the final list of predicted individuals?",
    "1762125": "Congratulations @charmq @knshnb, and thanks for sharing amazing solution.\nI'm proud to have you both as my colleagues, and looking forward to fighting together at another competitions !! ",
    "1793564": "congratulations!",
    "1793065": "Congratulations~",
    "1775816": "congratulations!",
    "1773859": "Congratulations",
    "1773837": "Congratulations",
    "1773149": "Congratulations",
    "1773090": "Congratulations!",
    "1772796": "Congratulations @knshnb  thanks for sharing👌👌👌",
    "1772609": "Congratulations! ",
    "1772441": "Congratulations, so best models",
    "1772186": "Congratulations",
    "1771932": "congratulations! Thank you for sharing your methodologies!! ",
    "1771457": "Congratulations",
    "1771428": "Thanks a lot, that was very clear and helpful, got many fresh ideas",
    "1770976": "congratulations!",
    "1770932": "Nice work. Congratulations for the win!\n",
    "1770316": "good job!!!",
    "1768687": "Congratulations for the first place! Thanks for sharing the details as well. ",
    "1768038": "Congratulations for the first place!!!",
    "1767789": "congratulations",
    "1766110": "Congratulations for the first place!",
    "1766001": "congratulations! your share was really helpfull!!",
    "1765869": "congratulations",
    "1765860": "congratulations!",
    "1765548": "amazing solution!",
    "1765494": "congratulations",
    "1765428": "Your share is very helpful for me!!",
    "1765243": "congratulations",
    "1765146": "Thanks for this nice dataset.I have successfully moved from novice to Contributor",
    "1764970": "Great Job!",
    "1764866": "congratulations",
    "1764277": "congratulations\n",
    "1764068": "Congratulations for the first place!",
    "1763946": "Congrats @charmq @knshnb!!!\nThanks for sharing.",
    "1763891": "Congratulation! @knshnb ",
    "1763524": "@knshnb Can you share your experience with different lr at backbone and head? I tried it but it shows superior performance on the early epoch but far from the best single lr at the last epoch.",
    "1763431": "congratulation!! @charmq @knshnb ",
    "1762514": "Congrats! efficientnet_b7 on bs=16,  res=1024x1024, gpu=v100(32) ?  Should this fit? How long did it take?",
    "1762216": "Congrats! Pseudo really boost a lot!",
    "1763189": "congratulations @charmq ",
    "1762605": "Congrats @charmq @knshnb for the first place!!!",
    "1762117": "Congratulations!! ",
    "2404751": "I tried to reproduce your score on leaderboard with the repo instructions you share on github but unfortunately was not able to do so.  I trained your architecture for 30 epochs, then chose 'pred_idx' as predictions taken from test_fullbody_results.npz. For inference I chose logit threshold for 'new_individual' to be 0.32. The number of 1st predictions of 'new_individual' ratio was 0.17 (close to 0.165 as you specified). The score I get with the resulting submissionon on kaggle is only 0.35. Am I missing something?",
    "2164997": "Excellent!!!, this new tool will be of crucial help for all developers whether they are experienced, teams or newbies. See option <Models> in the Kaggle menu, just below <DataSets>.",
    "1782046": "I want to know some questions about optuna search. Do you use all data for training or some data for sampling.",
    "1766112": "",
    "1766422": "Thanks for sharing, congratulations!",
    "1766367": "Thanks for sharing. \nCongratulation!",
    "1766253": "Thanks for this dataset",
    "1766102": "This was very helpful. Thanks!",
    "1763020": "Thanks for sharing, I learned a lot!",
    "1762181": "Thanks for  Sharing solution!!"
  }
}