{
  "id": 319896,
  "title": "3rd solution",
  "url": "/competitions/happy-whale-and-dolphin/discussion/319896",
  "author_name": "",
  "post_date": "2022-04-19T10:10:07.387926600Z",
  "votes": 64,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.</p>\n<h1>Dataset:</h1>\n<p><a href=\"https://postimg.cc/t11JNKjg\" target=\"_blank\"><img src=\"https://i.postimg.cc/t11JNKjg/image.png\" alt=\"image\"></a><br>\nFigure 1:  Training set generation pipeline</p>\n<ul>\n<li><p>As shown in Figure 1, we first randomly labeled around 5000 whale bodies. Next, we used labeled bodies to train a salient whale body detector using YOLOv5, with mosaic=0, degrees=0, and mixup=0. Next, we used the trained detector to predict on the entire training set and labeled images that either their box score is less than 0.4 or their number of boxes is not equal to one. Finally, the well-labeled training set was used to train the whale body detector again. </p></li>\n<li><p>Embedding extractor training set generation pipeline:<br>\nWe used the whale body detector mentioned above to predict the test set. If the number of boxes is not equal to one, we adjusted the detector's score threshold and the image's resolution (we used different resolutions for inference). For images that contain multiple bboxes,  we first cropped those boxes and get their embeddings, and then computed the cosine similarity with training set. If the cosine distance is greater than 0.5, we chose the closest box as the salient target, else, we chose the target that has the highest detection score. So far, we completed the salient detection by making each image in the training set got only one corresponding whale body bbox. Finally, we can use those bboxes to get our cropped whale body training set.</p></li>\n</ul>\n<h1>Models</h1>\n<h3>EfficientNet</h3>\n<ul>\n<li>framework: pytorch</li>\n<li>backbone:  tf_efficientnet_b7_ns/tf_efficientnet_b6_ns imagenet pretrained          </li>\n<li>image size:   768</li>\n<li>embedding size: 512</li>\n<li>hyperparameter: 20 epoch, Adam, lr=5e-4, cosine, warmup 4 epoch<br>\nDuring training, we found that train_loss is very low and the val_loss is much higher. In order to alleviate the over-fitting, optimization is carried out from the perspective of data enhancement and features space.</li>\n<li>feature space constraint <br>\nbackbone -&gt; gempooling -&gt; bnn-neck -&gt; arcface(s=30, m=0.3) or adaface(m=0.3, h=0.333,s=30, t_alpha=0.01)<br>\n<a href=\"https://postimg.cc/gLFcYj19\" target=\"_blank\"><img src=\"https://i.postimg.cc/gLFcYj19/image.png\" alt=\"image\"></a><br>\nref: <a href=\"https://arxiv.org/abs/1903.07071\" target=\"_blank\">https://arxiv.org/abs/1903.07071</a><br>\nBecause Arcface only measures cosine Angle, the feature space does not carry on the distance constraint. Therefore, we use <a href=\"https://github.com/michuanhaohao/reid-strong-baseline/blob/3da7e6f03164a92e696cb6da059b1cd771b0346d/modeling/baseline.py\" target=\"_blank\">BNNeck</a> to shape the feature space and increase the difficulty of feature distinction, as result alleviating the over-fitting.</li>\n</ul>\n<h3>NFNet</h3>\n<ul>\n<li>backbone: eca_nfnet_l2, imagenet pretrained</li>\n<li>image size: 1024 </li>\n<li>embedding size: 512</li>\n<li>hyperparameter: 25epoch, 4epoch warmup,  1e-3lr, adam<br>\nOthers are the same as EfficientNet<br>\nWe found that the training speed of nfnet_l2 1024 is as fast as effb6 768</li>\n</ul>\n<h1>Augmentation</h1>\n<p>mixup: 0.5prob, 0.5 alpha</p>\n<pre><code>aug8p3 = A.OneOf([\n            A.Sharpen(p=0.3),\n            A.ToGray(p=0.3),\n            A.CLAHE(p=0.3),\n        ], p=0.5)\n\nargs['transform'] = {\n    'train': A.Compose([\n        A.ShiftScaleRotate(rotate_limit=15, scale_limit=0.1, border_mode=cv2.BORDER_REFLECT, p=0.5),\n        A.Resize(size, size),\n        aug8p3,\n        A.HorizontalFlip(p=0.5),\n        A.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1),\n        A.Normalize()\n    ]),\n\n    'val': A.Compose([\n        A.Resize(size, size),\n        A.Normalize()\n    ])\n}\n</code></pre>\n<p>Through the dataset analysis, we found that many individual distinctions highly rely on texture differences. Therefore, we believe that sharpening and grayscaling can make the model increases the impact of texture and reduce the dependence on color.</p>\n<h1>Tricks</h1>\n<p>Due to the excessive number of pictures, tricks were trained by a small model with a low resolution, and then the large model was trained with a high resolution.</p>\n<h3>Base small model</h3>\n<p>tf_efficientnet_b5_ns, 512, backbone -&gt; AdaptiveAvgPool2d -&gt; linear -&gt; arcface(s=30, m=0.3)</p>\n<h3>Different trick effects on new_id, not_new_id, and cv</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>FALSE</th>\n<th>TRUE</th>\n<th>cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>base</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>aug8p3</td>\n<td>0.6</td>\n<td>-0.92</td>\n<td>0.45</td>\n</tr>\n<tr>\n<td>bnn</td>\n<td>0.38</td>\n<td>2.9</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>bnn+aug8p3</td>\n<td>0.97</td>\n<td>2.46</td>\n<td>1.12</td>\n</tr>\n<tr>\n<td>neckD+aug8p3</td>\n<td>0.74</td>\n<td>2.65</td>\n<td>0.93</td>\n</tr>\n<tr>\n<td>aug8p3_neckbnn_freebn</td>\n<td>1.5</td>\n<td>1.64</td>\n<td>1.51</td>\n</tr>\n<tr>\n<td>aug8p3_neckbnn_gem</td>\n<td>2.19</td>\n<td>2.94</td>\n<td>2.26</td>\n</tr>\n<tr>\n<td>mixup</td>\n<td>1.18</td>\n<td>0.39</td>\n<td>1.1</td>\n</tr>\n</tbody>\n</table>\n<p>In the table above, FALSE represents not_new_id, and TURE represents new_id. By balancing the effects on False and True, the final tricks combination is mixup + aug8p3 + neckbnn +gem + freeze_backbone_bn. freeze_backbone_bn is not used in NFNetL2.</p>\n<h3>Improvements on LB</h3>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>img size</th>\n<th>fold0  lb</th>\n<th>3fold  lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efb6</td>\n<td>768</td>\n<td>0.784</td>\n<td></td>\n</tr>\n<tr>\n<td>efb6 + tricks</td>\n<td>768</td>\n<td>0.817</td>\n<td>0.834</td>\n</tr>\n<tr>\n<td>nfnl2 + tricks</td>\n<td>768</td>\n<td>0.811</td>\n<td>0.824</td>\n</tr>\n</tbody>\n</table>\n<h1>Pseudo-label</h1>\n<p>Since a high threshold will reduce the precision of True, we use a gradual iteration method to gradually add pseudo labels as shown in Figure 3.</p>\n<p><a href=\"https://postimg.cc/H8HTFxr6\" target=\"_blank\"><img src=\"https://i.postimg.cc/H8HTFxr6/image.png\" alt=\"image\"></a><br>\nFigure 3:   Pseudo-label pipelines. Above is pipeline 1, below is pipeline 2.</p>\n<p>The method to generate pseudo labels is to use different similarity thresholds to generate submission results and remove the category whose top1 is new_id from the submission results. As there is a certain intersection between pseudo labels and validation sets, we compute the similarity of pseudo labels and validation sets, and the images of validation sets with high similarity scores are removed to ensure the offline scores are aligned with the online score.</p>\n<ul>\n<li>Step 1: <br>\n'body' data is used to train the model. After multi-fold model stacking, pseudo labels are obtained on the test set with the high threshold value. Train the 'body' again and iterated twice.</li>\n<li>Step 2: <br>\nWe further use the pseudo-labels from step 1 to train part and body models, then do the stacking ensemble. The ensembled model is then used to get new pseudo labels by setting a relatively low threshold.<br>\npart reference: <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319789\" target=\"_blank\">Part</a></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>img size</th>\n<th>pseudo</th>\n<th>fold0  lb</th>\n<th>3fold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efb6</td>\n<td>768</td>\n<td></td>\n<td>0.784</td>\n<td></td>\n</tr>\n<tr>\n<td>efb6 + tricks</td>\n<td>768</td>\n<td></td>\n<td>0.817</td>\n<td>0.834</td>\n</tr>\n<tr>\n<td>nfnl2 + tricks</td>\n<td>768</td>\n<td></td>\n<td>0.811</td>\n<td>0.824</td>\n</tr>\n<tr>\n<td>nfnl2 + tricks</td>\n<td>768</td>\n<td>824pse thre0.5</td>\n<td></td>\n<td>0.845</td>\n</tr>\n<tr>\n<td>step1_1 :merge  b6, nfn</td>\n<td></td>\n<td>824pse thre0.5</td>\n<td></td>\n<td>0.858</td>\n</tr>\n<tr>\n<td>step1_2 :merge  b6, nfn</td>\n<td></td>\n<td>858pse thre0.46</td>\n<td></td>\n<td>0.864</td>\n</tr>\n<tr>\n<td>step1_2: multi merge</td>\n<td></td>\n<td>858pse thre0.46</td>\n<td></td>\n<td>0.875</td>\n</tr>\n<tr>\n<td>step2_1: multi body part merge</td>\n<td></td>\n<td>875pse thre0.42</td>\n<td></td>\n<td>0.883</td>\n</tr>\n<tr>\n<td>step2_1: body merge</td>\n<td></td>\n<td>0.883 pseu thre 0.43</td>\n<td></td>\n<td>0.886</td>\n</tr>\n</tbody>\n</table>\n<h1>Ensemble</h1>\n<p>We use ckpt merge for the same fold ensembling, and use submit merge for diff fold ensembling. The weight of submit merge is determined by LB.</p>\n<ul>\n<li>ckpt merge<br>\nFor same fold, we get embeddings by backbone -&gt; gempooling -&gt; bnn-neck -&gt; norm(feature), then different model' embs are concated channelwise ([batchsize, 512] -&gt; [batchsize, n*512]). After that, we search the threshold and get single fold submits</li>\n<li>submit merge<br>\nBy following <a href=\"https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels\" target=\"_blank\">simple-ensemble</a>, we ensemble and rerank different folds. </li>\n</ul>\n<p>[b6_aug8p3_mixup8prob_neckbnn_gem_freezebn_878pseu0-42_768,<br>\nb6_aug8p3_mixup8prob_neckbnn_gem_freezebn_881pseu0-44_768,</p>\n<p>b7_aug8p3__mixup_neckbnn_gem_freezebn_858pseu0-5_768,<br>\nb7_aug8p3_mix8p_pool2_neckbnn_gem_freezebn_884pseu_768,</p>\n<p>nfnet_l2_aug8p3_mixup_neckbnn_25e_883pseu0-43_1024,<br>\nnfnet_l2_aug8p3_mixup_neckbnn_25e_885pseu_adaface_1024]</p>\n<h1>Note:</h1>\n<ol>\n<li>Pseudo labels make submit merge have little improvement (0.001-0.002).</li>\n<li>We found that the new_id in the training set is 11%. By analyzing submit results, we believe that new_id in the test set is 16%, thus we use 0.15<em>True +0.85</em>False as our threshold search setup.</li>\n</ol>",
  "messages": [
    {
      "id": "1760465",
      "postDate": "04/19/2022 10:10:07",
      "content": "<p>Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.</p>\n<h1>Dataset:</h1>\n<p><a href=\"https://postimg.cc/t11JNKjg\" target=\"_blank\"><img src=\"https://i.postimg.cc/t11JNKjg/image.png\" alt=\"image\"></a><br>\nFigure 1:  Training set generation pipeline</p>\n<ul>\n<li><p>As shown in Figure 1, we first randomly labeled around 5000 whale bodies. Next, we used labeled bodies to train a salient whale body detector using YOLOv5, with mosaic=0, degrees=0, and mixup=0. Next, we used the trained detector to predict on the entire training set and labeled images that either their box score is less than 0.4 or their number of boxes is not equal to one. Finally, the well-labeled training set was used to train the whale body detector again. </p></li>\n<li><p>Embedding extractor training set generation pipeline:<br>\nWe used the whale body detector mentioned above to predict the test set. If the number of boxes is not equal to one, we adjusted the detector's score threshold and the image's resolution (we used different resolutions for inference). For images that contain multiple bboxes,  we first cropped those boxes and get their embeddings, and then computed the cosine similarity with training set. If the cosine distance is greater than 0.5, we chose the closest box as the salient target, else, we chose the target that has the highest detection score. So far, we completed the salient detection by making each image in the training set got only one corresponding whale body bbox. Finally, we can use those bboxes to get our cropped whale body training set.</p></li>\n</ul>\n<h1>Models</h1>\n<h3>EfficientNet</h3>\n<ul>\n<li>framework: pytorch</li>\n<li>backbone:  tf_efficientnet_b7_ns/tf_efficientnet_b6_ns imagenet pretrained          </li>\n<li>image size:   768</li>\n<li>embedding size: 512</li>\n<li>hyperparameter: 20 epoch, Adam, lr=5e-4, cosine, warmup 4 epoch<br>\nDuring training, we found that train_loss is very low and the val_loss is much higher. In order to alleviate the over-fitting, optimization is carried out from the perspective of data enhancement and features space.</li>\n<li>feature space constraint <br>\nbackbone -&gt; gempooling -&gt; bnn-neck -&gt; arcface(s=30, m=0.3) or adaface(m=0.3, h=0.333,s=30, t_alpha=0.01)<br>\n<a href=\"https://postimg.cc/gLFcYj19\" target=\"_blank\"><img src=\"https://i.postimg.cc/gLFcYj19/image.png\" alt=\"image\"></a><br>\nref: <a href=\"https://arxiv.org/abs/1903.07071\" target=\"_blank\">https://arxiv.org/abs/1903.07071</a><br>\nBecause Arcface only measures cosine Angle, the feature space does not carry on the distance constraint. Therefore, we use <a href=\"https://github.com/michuanhaohao/reid-strong-baseline/blob/3da7e6f03164a92e696cb6da059b1cd771b0346d/modeling/baseline.py\" target=\"_blank\">BNNeck</a> to shape the feature space and increase the difficulty of feature distinction, as result alleviating the over-fitting.</li>\n</ul>\n<h3>NFNet</h3>\n<ul>\n<li>backbone: eca_nfnet_l2, imagenet pretrained</li>\n<li>image size: 1024 </li>\n<li>embedding size: 512</li>\n<li>hyperparameter: 25epoch, 4epoch warmup,  1e-3lr, adam<br>\nOthers are the same as EfficientNet<br>\nWe found that the training speed of nfnet_l2 1024 is as fast as effb6 768</li>\n</ul>\n<h1>Augmentation</h1>\n<p>mixup: 0.5prob, 0.5 alpha</p>\n<pre><code>aug8p3 = A.OneOf([\n            A.Sharpen(p=0.3),\n            A.ToGray(p=0.3),\n            A.CLAHE(p=0.3),\n        ], p=0.5)\n\nargs['transform'] = {\n    'train': A.Compose([\n        A.ShiftScaleRotate(rotate_limit=15, scale_limit=0.1, border_mode=cv2.BORDER_REFLECT, p=0.5),\n        A.Resize(size, size),\n        aug8p3,\n        A.HorizontalFlip(p=0.5),\n        A.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1),\n        A.Normalize()\n    ]),\n\n    'val': A.Compose([\n        A.Resize(size, size),\n        A.Normalize()\n    ])\n}\n</code></pre>\n<p>Through the dataset analysis, we found that many individual distinctions highly rely on texture differences. Therefore, we believe that sharpening and grayscaling can make the model increases the impact of texture and reduce the dependence on color.</p>\n<h1>Tricks</h1>\n<p>Due to the excessive number of pictures, tricks were trained by a small model with a low resolution, and then the large model was trained with a high resolution.</p>\n<h3>Base small model</h3>\n<p>tf_efficientnet_b5_ns, 512, backbone -&gt; AdaptiveAvgPool2d -&gt; linear -&gt; arcface(s=30, m=0.3)</p>\n<h3>Different trick effects on new_id, not_new_id, and cv</h3>\n<table>\n<thead>\n<tr>\n<th></th>\n<th>FALSE</th>\n<th>TRUE</th>\n<th>cv</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>base</td>\n<td></td>\n<td></td>\n<td></td>\n</tr>\n<tr>\n<td>aug8p3</td>\n<td>0.6</td>\n<td>-0.92</td>\n<td>0.45</td>\n</tr>\n<tr>\n<td>bnn</td>\n<td>0.38</td>\n<td>2.9</td>\n<td>0.63</td>\n</tr>\n<tr>\n<td>bnn+aug8p3</td>\n<td>0.97</td>\n<td>2.46</td>\n<td>1.12</td>\n</tr>\n<tr>\n<td>neckD+aug8p3</td>\n<td>0.74</td>\n<td>2.65</td>\n<td>0.93</td>\n</tr>\n<tr>\n<td>aug8p3_neckbnn_freebn</td>\n<td>1.5</td>\n<td>1.64</td>\n<td>1.51</td>\n</tr>\n<tr>\n<td>aug8p3_neckbnn_gem</td>\n<td>2.19</td>\n<td>2.94</td>\n<td>2.26</td>\n</tr>\n<tr>\n<td>mixup</td>\n<td>1.18</td>\n<td>0.39</td>\n<td>1.1</td>\n</tr>\n</tbody>\n</table>\n<p>In the table above, FALSE represents not_new_id, and TURE represents new_id. By balancing the effects on False and True, the final tricks combination is mixup + aug8p3 + neckbnn +gem + freeze_backbone_bn. freeze_backbone_bn is not used in NFNetL2.</p>\n<h3>Improvements on LB</h3>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>img size</th>\n<th>fold0  lb</th>\n<th>3fold  lb</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efb6</td>\n<td>768</td>\n<td>0.784</td>\n<td></td>\n</tr>\n<tr>\n<td>efb6 + tricks</td>\n<td>768</td>\n<td>0.817</td>\n<td>0.834</td>\n</tr>\n<tr>\n<td>nfnl2 + tricks</td>\n<td>768</td>\n<td>0.811</td>\n<td>0.824</td>\n</tr>\n</tbody>\n</table>\n<h1>Pseudo-label</h1>\n<p>Since a high threshold will reduce the precision of True, we use a gradual iteration method to gradually add pseudo labels as shown in Figure 3.</p>\n<p><a href=\"https://postimg.cc/H8HTFxr6\" target=\"_blank\"><img src=\"https://i.postimg.cc/H8HTFxr6/image.png\" alt=\"image\"></a><br>\nFigure 3:   Pseudo-label pipelines. Above is pipeline 1, below is pipeline 2.</p>\n<p>The method to generate pseudo labels is to use different similarity thresholds to generate submission results and remove the category whose top1 is new_id from the submission results. As there is a certain intersection between pseudo labels and validation sets, we compute the similarity of pseudo labels and validation sets, and the images of validation sets with high similarity scores are removed to ensure the offline scores are aligned with the online score.</p>\n<ul>\n<li>Step 1: <br>\n'body' data is used to train the model. After multi-fold model stacking, pseudo labels are obtained on the test set with the high threshold value. Train the 'body' again and iterated twice.</li>\n<li>Step 2: <br>\nWe further use the pseudo-labels from step 1 to train part and body models, then do the stacking ensemble. The ensembled model is then used to get new pseudo labels by setting a relatively low threshold.<br>\npart reference: <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319789\" target=\"_blank\">Part</a></li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>name</th>\n<th>img size</th>\n<th>pseudo</th>\n<th>fold0  lb</th>\n<th>3fold</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efb6</td>\n<td>768</td>\n<td></td>\n<td>0.784</td>\n<td></td>\n</tr>\n<tr>\n<td>efb6 + tricks</td>\n<td>768</td>\n<td></td>\n<td>0.817</td>\n<td>0.834</td>\n</tr>\n<tr>\n<td>nfnl2 + tricks</td>\n<td>768</td>\n<td></td>\n<td>0.811</td>\n<td>0.824</td>\n</tr>\n<tr>\n<td>nfnl2 + tricks</td>\n<td>768</td>\n<td>824pse thre0.5</td>\n<td></td>\n<td>0.845</td>\n</tr>\n<tr>\n<td>step1_1 :merge  b6, nfn</td>\n<td></td>\n<td>824pse thre0.5</td>\n<td></td>\n<td>0.858</td>\n</tr>\n<tr>\n<td>step1_2 :merge  b6, nfn</td>\n<td></td>\n<td>858pse thre0.46</td>\n<td></td>\n<td>0.864</td>\n</tr>\n<tr>\n<td>step1_2: multi merge</td>\n<td></td>\n<td>858pse thre0.46</td>\n<td></td>\n<td>0.875</td>\n</tr>\n<tr>\n<td>step2_1: multi body part merge</td>\n<td></td>\n<td>875pse thre0.42</td>\n<td></td>\n<td>0.883</td>\n</tr>\n<tr>\n<td>step2_1: body merge</td>\n<td></td>\n<td>0.883 pseu thre 0.43</td>\n<td></td>\n<td>0.886</td>\n</tr>\n</tbody>\n</table>\n<h1>Ensemble</h1>\n<p>We use ckpt merge for the same fold ensembling, and use submit merge for diff fold ensembling. The weight of submit merge is determined by LB.</p>\n<ul>\n<li>ckpt merge<br>\nFor same fold, we get embeddings by backbone -&gt; gempooling -&gt; bnn-neck -&gt; norm(feature), then different model' embs are concated channelwise ([batchsize, 512] -&gt; [batchsize, n*512]). After that, we search the threshold and get single fold submits</li>\n<li>submit merge<br>\nBy following <a href=\"https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels\" target=\"_blank\">simple-ensemble</a>, we ensemble and rerank different folds. </li>\n</ul>\n<p>[b6_aug8p3_mixup8prob_neckbnn_gem_freezebn_878pseu0-42_768,<br>\nb6_aug8p3_mixup8prob_neckbnn_gem_freezebn_881pseu0-44_768,</p>\n<p>b7_aug8p3__mixup_neckbnn_gem_freezebn_858pseu0-5_768,<br>\nb7_aug8p3_mix8p_pool2_neckbnn_gem_freezebn_884pseu_768,</p>\n<p>nfnet_l2_aug8p3_mixup_neckbnn_25e_883pseu0-43_1024,<br>\nnfnet_l2_aug8p3_mixup_neckbnn_25e_885pseu_adaface_1024]</p>\n<h1>Note:</h1>\n<ol>\n<li>Pseudo labels make submit merge have little improvement (0.001-0.002).</li>\n<li>We found that the new_id in the training set is 11%. By analyzing submit results, we believe that new_id in the test set is 16%, thus we use 0.15<em>True +0.85</em>False as our threshold search setup.</li>\n</ol>",
      "rawMarkdown": "Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.\n\n# Dataset: \n<a href='https://postimg.cc/t11JNKjg' target='_blank'><img src='https://i.postimg.cc/t11JNKjg/image.png' border='0' alt='image'/></a>\nFigure 1:  Training set generation pipeline\n- As shown in Figure 1, we first randomly labeled around 5000 whale bodies. Next, we used labeled bodies to train a salient whale body detector using YOLOv5, with mosaic=0, degrees=0, and mixup=0. Next, we used the trained detector to predict on the entire training set and labeled images that either their box score is less than 0.4 or their number of boxes is not equal to one. Finally, the well-labeled training set was used to train the whale body detector again. \n\n- Embedding extractor training set generation pipeline:\nWe used the whale body detector mentioned above to predict the test set. If the number of boxes is not equal to one, we adjusted the detector's score threshold and the image's resolution (we used different resolutions for inference). For images that contain multiple bboxes,  we first cropped those boxes and get their embeddings, and then computed the cosine similarity with training set. If the cosine distance is greater than 0.5, we chose the closest box as the salient target, else, we chose the target that has the highest detection score. So far, we completed the salient detection by making each image in the training set got only one corresponding whale body bbox. Finally, we can use those bboxes to get our cropped whale body training set.\n\n# Models\n### EfficientNet\n- framework: pytorch\n- backbone:  tf_efficientnet_b7_ns/tf_efficientnet_b6_ns imagenet pretrained          \n- image size:   768\n- embedding size: 512\n- hyperparameter: 20 epoch, Adam, lr=5e-4, cosine, warmup 4 epoch\nDuring training, we found that train_loss is very low and the val_loss is much higher. In order to alleviate the over-fitting, optimization is carried out from the perspective of data enhancement and features space.\n- feature space constraint \nbackbone -> gempooling -> bnn-neck -> arcface(s=30, m=0.3) or adaface(m=0.3, h=0.333,s=30, t_alpha=0.01)\n<a href='https://postimg.cc/gLFcYj19' target='_blank'><img src='https://i.postimg.cc/gLFcYj19/image.png' border='0' alt='image'/></a>\nref: https://arxiv.org/abs/1903.07071\nBecause Arcface only measures cosine Angle, the feature space does not carry on the distance constraint. Therefore, we use [BNNeck](https://github.com/michuanhaohao/reid-strong-baseline/blob/3da7e6f03164a92e696cb6da059b1cd771b0346d/modeling/baseline.py) to shape the feature space and increase the difficulty of feature distinction, as result alleviating the over-fitting.\n\n### NFNet\n- backbone: eca_nfnet_l2, imagenet pretrained\n- image size: 1024 \n- embedding size: 512\n- hyperparameter: 25epoch, 4epoch warmup,  1e-3lr, adam\nOthers are the same as EfficientNet\nWe found that the training speed of nfnet_l2 1024 is as fast as effb6 768\n\n# Augmentation\nmixup: 0.5prob, 0.5 alpha\n\n```\naug8p3 = A.OneOf([\n            A.Sharpen(p=0.3),\n            A.ToGray(p=0.3),\n            A.CLAHE(p=0.3),\n        ], p=0.5)\n        \nargs['transform'] = {\n    'train': A.Compose([\n        A.ShiftScaleRotate(rotate_limit=15, scale_limit=0.1, border_mode=cv2.BORDER_REFLECT, p=0.5),\n        A.Resize(size, size),\n        aug8p3,\n        A.HorizontalFlip(p=0.5),\n        A.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1),\n        A.Normalize()\n    ]),\n\n    'val': A.Compose([\n        A.Resize(size, size),\n        A.Normalize()\n    ])\n}\n```\nThrough the dataset analysis, we found that many individual distinctions highly rely on texture differences. Therefore, we believe that sharpening and grayscaling can make the model increases the impact of texture and reduce the dependence on color.\n\n# Tricks\nDue to the excessive number of pictures, tricks were trained by a small model with a low resolution, and then the large model was trained with a high resolution.\n\n### Base small model \ntf_efficientnet_b5_ns, 512, backbone -> AdaptiveAvgPool2d -> linear -> arcface(s=30, m=0.3)\n\n### Different trick effects on new_id, not_new_id, and cv\n\n|                       | FALSE | TRUE  | cv   |\n| --------------------- | ----- | ----- | ---- |\n| base                  |       |       |      |\n| aug8p3                | 0.6   | -0.92 | 0.45 |\n| bnn                   | 0.38  | 2.9   | 0.63 |\n| bnn+aug8p3            | 0.97  | 2.46  | 1.12 |\n| neckD+aug8p3          | 0.74  | 2.65  | 0.93 |\n| aug8p3_neckbnn_freebn | 1.5   | 1.64  | 1.51 |\n| aug8p3_neckbnn_gem    | 2.19  | 2.94  | 2.26 |\n| mixup                   | 1.18  | 0.39  | 1.1  | \n   \nIn the table above, FALSE represents not_new_id, and TURE represents new_id. By balancing the effects on False and True, the final tricks combination is mixup + aug8p3 + neckbnn +gem + freeze_backbone_bn. freeze_backbone_bn is not used in NFNetL2.\n\n### Improvements on LB\n| name            | img size  | fold0  lb  | 3fold  lb |\n| --------------- | --------- | ---------- | --------- |\n| efb6            | 768       | 0.784      |           |\n| efb6 + tricks   | 768       | 0.817      | 0.834     |\n| nfnl2 + tricks  | 768       | 0.811      | 0.824     |\n\n# Pseudo-label\nSince a high threshold will reduce the precision of True, we use a gradual iteration method to gradually add pseudo labels as shown in Figure 3.\n\n<a href='https://postimg.cc/H8HTFxr6' target='_blank'><img src='https://i.postimg.cc/H8HTFxr6/image.png' border='0' alt='image'/></a>\nFigure 3:   Pseudo-label pipelines. Above is pipeline 1, below is pipeline 2.\n\nThe method to generate pseudo labels is to use different similarity thresholds to generate submission results and remove the category whose top1 is new_id from the submission results. As there is a certain intersection between pseudo labels and validation sets, we compute the similarity of pseudo labels and validation sets, and the images of validation sets with high similarity scores are removed to ensure the offline scores are aligned with the online score.\n- Step 1: \n'body' data is used to train the model. After multi-fold model stacking, pseudo labels are obtained on the test set with the high threshold value. Train the 'body' again and iterated twice.\n- Step 2: \nWe further use the pseudo-labels from step 1 to train part and body models, then do the stacking ensemble. The ensembled model is then used to get new pseudo labels by setting a relatively low threshold.\npart reference: [Part](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319789)\n\n| name                           | img size  | pseudo               | fold0  lb  | 3fold  |\n| ------------------------------ | --------- | -------------------- | ---------- | ------ |\n| efb6                           | 768       |                      | 0.784      |        |\n| efb6 + tricks                  | 768       |                      | 0.817      | 0.834  |\n| nfnl2 + tricks                 | 768       |                      | 0.811      | 0.824  |\n| nfnl2 + tricks                 | 768       | 824pse thre0.5       |            | 0.845  |\n| step1_1 :merge  b6, nfn        |           | 824pse thre0.5       |            | 0.858  |\n| step1_2 :merge  b6, nfn        |           | 858pse thre0.46      |            | 0.864  |\n| step1_2: multi merge           |           | 858pse thre0.46      |            | 0.875  |\n| step2_1: multi body part merge |           | 875pse thre0.42      |            | 0.883  |\n| step2_1: body merge            |           | 0.883 pseu thre 0.43 |            | 0.886  |\n\n# Ensemble\n\nWe use ckpt merge for the same fold ensembling, and use submit merge for diff fold ensembling. The weight of submit merge is determined by LB.\n- ckpt merge\nFor same fold, we get embeddings by backbone -> gempooling -> bnn-neck -> norm(feature), then different model' embs are concated channelwise ([batchsize, 512] -> [batchsize, n*512]). After that, we search the threshold and get single fold submits\n- submit merge\nBy following [simple-ensemble](https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels), we ensemble and rerank different folds. \n\n[b6_aug8p3_mixup8prob_neckbnn_gem_freezebn_878pseu0-42_768,\nb6_aug8p3_mixup8prob_neckbnn_gem_freezebn_881pseu0-44_768,\n\nb7_aug8p3__mixup_neckbnn_gem_freezebn_858pseu0-5_768,\nb7_aug8p3_mix8p_pool2_neckbnn_gem_freezebn_884pseu_768,\n\nnfnet_l2_aug8p3_mixup_neckbnn_25e_883pseu0-43_1024,\nnfnet_l2_aug8p3_mixup_neckbnn_25e_885pseu_adaface_1024]\n\n# Note: \n1. Pseudo labels make submit merge have little improvement (0.001-0.002).\n2. We found that the new_id in the training set is 11%. By analyzing submit results, we believe that new_id in the test set is 16%, thus we use 0.15*True +0.85*False as our threshold search setup.",
      "votes": null
    },
    {
      "id": "1761427",
      "postDate": "04/19/2022 22:47:27",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1764617",
      "postDate": "04/22/2022 16:28:28",
      "content": "<p>Congratulation and thanks for your work! </p>",
      "rawMarkdown": "Congratulation and thanks for your work!",
      "votes": null
    },
    {
      "id": "1767268",
      "postDate": "04/25/2022 07:51:51",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1761427,
      "author_name": "hkayesh",
      "author_url": "",
      "post_date": "04/19/2022 22:47:27",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1764617,
      "author_name": "fisherh",
      "author_url": "",
      "post_date": "04/22/2022 16:28:28",
      "content": "<p>Congratulation and thanks for your work! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1767268,
      "author_name": "mingxuding",
      "author_url": "",
      "post_date": "04/25/2022 07:51:51",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1760465": "Congratulations to all the winners. Thanks Happywhale and Kaggle for hosting this interesting competition.\n\n# Dataset: \n<a href='https://postimg.cc/t11JNKjg' target='_blank'><img src='https://i.postimg.cc/t11JNKjg/image.png' border='0' alt='image'/></a>\nFigure 1:  Training set generation pipeline\n- As shown in Figure 1, we first randomly labeled around 5000 whale bodies. Next, we used labeled bodies to train a salient whale body detector using YOLOv5, with mosaic=0, degrees=0, and mixup=0. Next, we used the trained detector to predict on the entire training set and labeled images that either their box score is less than 0.4 or their number of boxes is not equal to one. Finally, the well-labeled training set was used to train the whale body detector again. \n\n- Embedding extractor training set generation pipeline:\nWe used the whale body detector mentioned above to predict the test set. If the number of boxes is not equal to one, we adjusted the detector's score threshold and the image's resolution (we used different resolutions for inference). For images that contain multiple bboxes,  we first cropped those boxes and get their embeddings, and then computed the cosine similarity with training set. If the cosine distance is greater than 0.5, we chose the closest box as the salient target, else, we chose the target that has the highest detection score. So far, we completed the salient detection by making each image in the training set got only one corresponding whale body bbox. Finally, we can use those bboxes to get our cropped whale body training set.\n\n# Models\n### EfficientNet\n- framework: pytorch\n- backbone:  tf_efficientnet_b7_ns/tf_efficientnet_b6_ns imagenet pretrained          \n- image size:   768\n- embedding size: 512\n- hyperparameter: 20 epoch, Adam, lr=5e-4, cosine, warmup 4 epoch\nDuring training, we found that train_loss is very low and the val_loss is much higher. In order to alleviate the over-fitting, optimization is carried out from the perspective of data enhancement and features space.\n- feature space constraint \nbackbone -> gempooling -> bnn-neck -> arcface(s=30, m=0.3) or adaface(m=0.3, h=0.333,s=30, t_alpha=0.01)\n<a href='https://postimg.cc/gLFcYj19' target='_blank'><img src='https://i.postimg.cc/gLFcYj19/image.png' border='0' alt='image'/></a>\nref: https://arxiv.org/abs/1903.07071\nBecause Arcface only measures cosine Angle, the feature space does not carry on the distance constraint. Therefore, we use [BNNeck](https://github.com/michuanhaohao/reid-strong-baseline/blob/3da7e6f03164a92e696cb6da059b1cd771b0346d/modeling/baseline.py) to shape the feature space and increase the difficulty of feature distinction, as result alleviating the over-fitting.\n\n### NFNet\n- backbone: eca_nfnet_l2, imagenet pretrained\n- image size: 1024 \n- embedding size: 512\n- hyperparameter: 25epoch, 4epoch warmup,  1e-3lr, adam\nOthers are the same as EfficientNet\nWe found that the training speed of nfnet_l2 1024 is as fast as effb6 768\n\n# Augmentation\nmixup: 0.5prob, 0.5 alpha\n\n```\naug8p3 = A.OneOf([\n            A.Sharpen(p=0.3),\n            A.ToGray(p=0.3),\n            A.CLAHE(p=0.3),\n        ], p=0.5)\n        \nargs['transform'] = {\n    'train': A.Compose([\n        A.ShiftScaleRotate(rotate_limit=15, scale_limit=0.1, border_mode=cv2.BORDER_REFLECT, p=0.5),\n        A.Resize(size, size),\n        aug8p3,\n        A.HorizontalFlip(p=0.5),\n        A.ColorJitter(brightness=0.1, contrast=0.1, saturation=0.1),\n        A.Normalize()\n    ]),\n\n    'val': A.Compose([\n        A.Resize(size, size),\n        A.Normalize()\n    ])\n}\n```\nThrough the dataset analysis, we found that many individual distinctions highly rely on texture differences. Therefore, we believe that sharpening and grayscaling can make the model increases the impact of texture and reduce the dependence on color.\n\n# Tricks\nDue to the excessive number of pictures, tricks were trained by a small model with a low resolution, and then the large model was trained with a high resolution.\n\n### Base small model \ntf_efficientnet_b5_ns, 512, backbone -> AdaptiveAvgPool2d -> linear -> arcface(s=30, m=0.3)\n\n### Different trick effects on new_id, not_new_id, and cv\n\n|                       | FALSE | TRUE  | cv   |\n| --------------------- | ----- | ----- | ---- |\n| base                  |       |       |      |\n| aug8p3                | 0.6   | -0.92 | 0.45 |\n| bnn                   | 0.38  | 2.9   | 0.63 |\n| bnn+aug8p3            | 0.97  | 2.46  | 1.12 |\n| neckD+aug8p3          | 0.74  | 2.65  | 0.93 |\n| aug8p3_neckbnn_freebn | 1.5   | 1.64  | 1.51 |\n| aug8p3_neckbnn_gem    | 2.19  | 2.94  | 2.26 |\n| mixup                   | 1.18  | 0.39  | 1.1  | \n   \nIn the table above, FALSE represents not_new_id, and TURE represents new_id. By balancing the effects on False and True, the final tricks combination is mixup + aug8p3 + neckbnn +gem + freeze_backbone_bn. freeze_backbone_bn is not used in NFNetL2.\n\n### Improvements on LB\n| name            | img size  | fold0  lb  | 3fold  lb |\n| --------------- | --------- | ---------- | --------- |\n| efb6            | 768       | 0.784      |           |\n| efb6 + tricks   | 768       | 0.817      | 0.834     |\n| nfnl2 + tricks  | 768       | 0.811      | 0.824     |\n\n# Pseudo-label\nSince a high threshold will reduce the precision of True, we use a gradual iteration method to gradually add pseudo labels as shown in Figure 3.\n\n<a href='https://postimg.cc/H8HTFxr6' target='_blank'><img src='https://i.postimg.cc/H8HTFxr6/image.png' border='0' alt='image'/></a>\nFigure 3:   Pseudo-label pipelines. Above is pipeline 1, below is pipeline 2.\n\nThe method to generate pseudo labels is to use different similarity thresholds to generate submission results and remove the category whose top1 is new_id from the submission results. As there is a certain intersection between pseudo labels and validation sets, we compute the similarity of pseudo labels and validation sets, and the images of validation sets with high similarity scores are removed to ensure the offline scores are aligned with the online score.\n- Step 1: \n'body' data is used to train the model. After multi-fold model stacking, pseudo labels are obtained on the test set with the high threshold value. Train the 'body' again and iterated twice.\n- Step 2: \nWe further use the pseudo-labels from step 1 to train part and body models, then do the stacking ensemble. The ensembled model is then used to get new pseudo labels by setting a relatively low threshold.\npart reference: [Part](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319789)\n\n| name                           | img size  | pseudo               | fold0  lb  | 3fold  |\n| ------------------------------ | --------- | -------------------- | ---------- | ------ |\n| efb6                           | 768       |                      | 0.784      |        |\n| efb6 + tricks                  | 768       |                      | 0.817      | 0.834  |\n| nfnl2 + tricks                 | 768       |                      | 0.811      | 0.824  |\n| nfnl2 + tricks                 | 768       | 824pse thre0.5       |            | 0.845  |\n| step1_1 :merge  b6, nfn        |           | 824pse thre0.5       |            | 0.858  |\n| step1_2 :merge  b6, nfn        |           | 858pse thre0.46      |            | 0.864  |\n| step1_2: multi merge           |           | 858pse thre0.46      |            | 0.875  |\n| step2_1: multi body part merge |           | 875pse thre0.42      |            | 0.883  |\n| step2_1: body merge            |           | 0.883 pseu thre 0.43 |            | 0.886  |\n\n# Ensemble\n\nWe use ckpt merge for the same fold ensembling, and use submit merge for diff fold ensembling. The weight of submit merge is determined by LB.\n- ckpt merge\nFor same fold, we get embeddings by backbone -> gempooling -> bnn-neck -> norm(feature), then different model' embs are concated channelwise ([batchsize, 512] -> [batchsize, n*512]). After that, we search the threshold and get single fold submits\n- submit merge\nBy following [simple-ensemble](https://www.kaggle.com/code/yamsam/simple-ensemble-of-public-best-kernels), we ensemble and rerank different folds. \n\n[b6_aug8p3_mixup8prob_neckbnn_gem_freezebn_878pseu0-42_768,\nb6_aug8p3_mixup8prob_neckbnn_gem_freezebn_881pseu0-44_768,\n\nb7_aug8p3__mixup_neckbnn_gem_freezebn_858pseu0-5_768,\nb7_aug8p3_mix8p_pool2_neckbnn_gem_freezebn_884pseu_768,\n\nnfnet_l2_aug8p3_mixup_neckbnn_25e_883pseu0-43_1024,\nnfnet_l2_aug8p3_mixup_neckbnn_25e_885pseu_adaface_1024]\n\n# Note: \n1. Pseudo labels make submit merge have little improvement (0.001-0.002).\n2. We found that the new_id in the training set is 11%. By analyzing submit results, we believe that new_id in the test set is 16%, thus we use 0.15*True +0.85*False as our threshold search setup.",
    "1761427": "Thanks for sharing",
    "1764617": "Congratulation and thanks for your work!",
    "1767268": "Thanks for sharing"
  },
  "source": "meta"
}