{
  "id": 320040,
  "title": "4th place solution",
  "url": "/competitions/happy-whale-and-dolphin/writeups/all-data-are-ext-4th-place-solution",
  "author_name": "",
  "post_date": "2022-04-19T19:44:46.326318700Z",
  "votes": 59,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Congrats to all the winners. It has been a tough competition. The top 2 teams had a very strong last push. Thanks for the great collaboration once again, my long time teammate <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>\n<h1>TL;DR</h1>\n<ul>\n<li>Yolov5 Detection</li>\n<li>Dynamic Margin ArcFace with DOLG CNN backbone<ul>\n<li>Best backbone is ConvNext</li></ul></li>\n<li>Two heads: predicting ids and species</li>\n<li>3 rounds of pseudo labeling</li>\n<li>Tune new_individual threshold separately for each species</li>\n</ul>\n<h1>Detection</h1>\n<p>We trained yolov5m with image size 512 for 40 epochs. We started by using this <a href=\"https://www.kaggle.com/datasets/awsaf49/happywhale-boundingbox-yolov5-dataset\" target=\"_blank\">public notebook's</a> predictions as labels. Then visualize examples with low OOF confidence. If the predicted bbox are wrong, remove this example from training set, or fix the ground truth bbox. </p>\n<p>We iterate this for 9 rounds. In the end, most OOF predictions look correct. The OOF iou = 0.93863. But some test set prediction bboxs are still off. In hindsight, we probably should start from scratch and label thousands of images, like other top teams did.</p>\n<p>After getting the predicted bbox, we extend it by 20% and feed the crop to arcface model.</p>\n<h1>Model</h1>\n<p>The modeling part is heavily influenced by the recent landmarks competitions. The architecture is Dynamic Margin ArcFace with DOLG CNN backbone. The dynamic margin arcface was introduced by us in last year's Landmark, see detail <a href=\"https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\" target=\"_blank\">here</a>. The DOLG was introduced to the Kaggle community by <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> in this year's Landmark, see detail <a href=\"https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099\" target=\"_blank\">here</a>.</p>\n<p>Other components of the top landmark solutions didn't work here though, including sub center arcface, and vision transformers. All the vision transformers underperform CNNs. The best CNN in our solution is ConvNext.</p>\n<p>The only twist in the model is that we output two heads, predicting both individual_id and species.</p>\n<p>For augmentations we used the following plus mixup</p>\n<pre><code>    A.HorizontalFlip(p=0.5),\n    A.RandomContrast(limit=0.2, p=0.75),\n    A.ShiftScaleRotate(shift_limit=0.0, scale_limit=0.3, rotate_limit=10, border_mode=0, p=0.7),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n</code></pre>\n<h1>Training steps</h1>\n<ol>\n<li>Tune everything on 5 fold. i.e. train on 80% of training data per fold, \"sacrificing\" some individual_ids in order to do cross validation.</li>\n<li>Ensemble 9 fold models to make pseudo labels.</li>\n<li>Train 6 best models on 100% training data + pseudo labeled test data</li>\n<li>Make new pseudo labels from step 3 ensemble</li>\n<li>Repeat step 3-4 for two more rounds</li>\n<li>Ensemble the 12 models from last two rounds.</li>\n</ol>\n<p>Average single model's public LB in 3 pseudo label rounds are 0.855, 0.862, 0.865</p>\n<h1>Ensemble</h1>\n<p>The best performing backbone is ConvNext. We also trained 3 non-ConvNext for diversity. The final six models are ConvNext Base, Large, XLarge, EfficientNet B7, V2L, NFNet L2. We used diversified image sizes for different models: 640, 672, 704, 736, 768, 800, 832, 864, 896, 960, 1024</p>\n<p>The ensemble is done by concatenating single models feature, before computing cosine similarity.</p>\n<h1>new_individual threshold</h1>\n<p>Since final models are trained on 100% of training data, there's no validation scores. New_individual's thresholds therefore need to be tuned on the LB. </p>\n<p>We noticed that different species have different levels of difficulty to predict, hence the optimal new_individual thresholds are different for different species. Tuning each species on the LB is riskly, so we: (1) tuned each species' <em>relative</em> threshold separately on CV, (2) tuned overall threshold level on LB, (3) adjust the LB-optimal overall threshold level by CV-optimal species relative thresholds.</p>\n<p>This adjustment boosted our public LB from 0.882 or 0.883 to 0.888</p>\n<p>We used an ensemble of 15 models’ species head to predict test set's species.</p>\n<h1>Things that didn't work</h1>\n<ul>\n<li>Vision Transformers</li>\n<li>Reranking Post processing</li>\n<li>Using the previous humpback whale competition's data to pretrain</li>\n<li>Use HFlip to double the number of individual_ids, a trick that worked well in previous humpback whale competition</li>\n</ul>",
  "messages": [
    {
      "id": "1761302",
      "postDate": "04/19/2022 19:44:46",
      "content": "<p>Congrats to all the winners. It has been a tough competition. The top 2 teams had a very strong last push. Thanks for the great collaboration once again, my long time teammate <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> </p>\n<h1>TL;DR</h1>\n<ul>\n<li>Yolov5 Detection</li>\n<li>Dynamic Margin ArcFace with DOLG CNN backbone<ul>\n<li>Best backbone is ConvNext</li></ul></li>\n<li>Two heads: predicting ids and species</li>\n<li>3 rounds of pseudo labeling</li>\n<li>Tune new_individual threshold separately for each species</li>\n</ul>\n<h1>Detection</h1>\n<p>We trained yolov5m with image size 512 for 40 epochs. We started by using this <a href=\"https://www.kaggle.com/datasets/awsaf49/happywhale-boundingbox-yolov5-dataset\" target=\"_blank\">public notebook's</a> predictions as labels. Then visualize examples with low OOF confidence. If the predicted bbox are wrong, remove this example from training set, or fix the ground truth bbox. </p>\n<p>We iterate this for 9 rounds. In the end, most OOF predictions look correct. The OOF iou = 0.93863. But some test set prediction bboxs are still off. In hindsight, we probably should start from scratch and label thousands of images, like other top teams did.</p>\n<p>After getting the predicted bbox, we extend it by 20% and feed the crop to arcface model.</p>\n<h1>Model</h1>\n<p>The modeling part is heavily influenced by the recent landmarks competitions. The architecture is Dynamic Margin ArcFace with DOLG CNN backbone. The dynamic margin arcface was introduced by us in last year's Landmark, see detail <a href=\"https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757\" target=\"_blank\">here</a>. The DOLG was introduced to the Kaggle community by <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> in this year's Landmark, see detail <a href=\"https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099\" target=\"_blank\">here</a>.</p>\n<p>Other components of the top landmark solutions didn't work here though, including sub center arcface, and vision transformers. All the vision transformers underperform CNNs. The best CNN in our solution is ConvNext.</p>\n<p>The only twist in the model is that we output two heads, predicting both individual_id and species.</p>\n<p>For augmentations we used the following plus mixup</p>\n<pre><code>    A.HorizontalFlip(p=0.5),\n    A.RandomContrast(limit=0.2, p=0.75),\n    A.ShiftScaleRotate(shift_limit=0.0, scale_limit=0.3, rotate_limit=10, border_mode=0, p=0.7),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n</code></pre>\n<h1>Training steps</h1>\n<ol>\n<li>Tune everything on 5 fold. i.e. train on 80% of training data per fold, \"sacrificing\" some individual_ids in order to do cross validation.</li>\n<li>Ensemble 9 fold models to make pseudo labels.</li>\n<li>Train 6 best models on 100% training data + pseudo labeled test data</li>\n<li>Make new pseudo labels from step 3 ensemble</li>\n<li>Repeat step 3-4 for two more rounds</li>\n<li>Ensemble the 12 models from last two rounds.</li>\n</ol>\n<p>Average single model's public LB in 3 pseudo label rounds are 0.855, 0.862, 0.865</p>\n<h1>Ensemble</h1>\n<p>The best performing backbone is ConvNext. We also trained 3 non-ConvNext for diversity. The final six models are ConvNext Base, Large, XLarge, EfficientNet B7, V2L, NFNet L2. We used diversified image sizes for different models: 640, 672, 704, 736, 768, 800, 832, 864, 896, 960, 1024</p>\n<p>The ensemble is done by concatenating single models feature, before computing cosine similarity.</p>\n<h1>new_individual threshold</h1>\n<p>Since final models are trained on 100% of training data, there's no validation scores. New_individual's thresholds therefore need to be tuned on the LB. </p>\n<p>We noticed that different species have different levels of difficulty to predict, hence the optimal new_individual thresholds are different for different species. Tuning each species on the LB is riskly, so we: (1) tuned each species' <em>relative</em> threshold separately on CV, (2) tuned overall threshold level on LB, (3) adjust the LB-optimal overall threshold level by CV-optimal species relative thresholds.</p>\n<p>This adjustment boosted our public LB from 0.882 or 0.883 to 0.888</p>\n<p>We used an ensemble of 15 models’ species head to predict test set's species.</p>\n<h1>Things that didn't work</h1>\n<ul>\n<li>Vision Transformers</li>\n<li>Reranking Post processing</li>\n<li>Using the previous humpback whale competition's data to pretrain</li>\n<li>Use HFlip to double the number of individual_ids, a trick that worked well in previous humpback whale competition</li>\n</ul>",
      "rawMarkdown": "Congrats to all the winners. It has been a tough competition. The top 2 teams had a very strong last push. Thanks for the great collaboration once again, my long time teammate @haqishen \n\n# TL;DR\n- Yolov5 Detection\n- Dynamic Margin ArcFace with DOLG CNN backbone\n - Best backbone is ConvNext\n- Two heads: predicting ids and species\n- 3 rounds of pseudo labeling\n- Tune new_individual threshold separately for each species\n\n# Detection\nWe trained yolov5m with image size 512 for 40 epochs. We started by using this [public notebook's](https://www.kaggle.com/datasets/awsaf49/happywhale-boundingbox-yolov5-dataset) predictions as labels. Then visualize examples with low OOF confidence. If the predicted bbox are wrong, remove this example from training set, or fix the ground truth bbox. \n\nWe iterate this for 9 rounds. In the end, most OOF predictions look correct. The OOF iou = 0.93863. But some test set prediction bboxs are still off. In hindsight, we probably should start from scratch and label thousands of images, like other top teams did.\n\nAfter getting the predicted bbox, we extend it by 20% and feed the crop to arcface model.\n\n# Model\nThe modeling part is heavily influenced by the recent landmarks competitions. The architecture is Dynamic Margin ArcFace with DOLG CNN backbone. The dynamic margin arcface was introduced by us in last year's Landmark, see detail [here](https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757). The DOLG was introduced to the Kaggle community by @christofhenkel in this year's Landmark, see detail [here](https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099).\n\nOther components of the top landmark solutions didn't work here though, including sub center arcface, and vision transformers. All the vision transformers underperform CNNs. The best CNN in our solution is ConvNext.\n\nThe only twist in the model is that we output two heads, predicting both individual_id and species.\n\nFor augmentations we used the following plus mixup\n```\n    A.HorizontalFlip(p=0.5),\n    A.RandomContrast(limit=0.2, p=0.75),\n    A.ShiftScaleRotate(shift_limit=0.0, scale_limit=0.3, rotate_limit=10, border_mode=0, p=0.7),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n```\n\n# Training steps\n1. Tune everything on 5 fold. i.e. train on 80% of training data per fold, \"sacrificing\" some individual_ids in order to do cross validation.\n2. Ensemble 9 fold models to make pseudo labels.\n3. Train 6 best models on 100% training data + pseudo labeled test data\n4. Make new pseudo labels from step 3 ensemble\n5. Repeat step 3-4 for two more rounds\n6. Ensemble the 12 models from last two rounds.\n\nAverage single model's public LB in 3 pseudo label rounds are 0.855, 0.862, 0.865\n\n# Ensemble\nThe best performing backbone is ConvNext. We also trained 3 non-ConvNext for diversity. The final six models are ConvNext Base, Large, XLarge, EfficientNet B7, V2L, NFNet L2. We used diversified image sizes for different models: 640, 672, 704, 736, 768, 800, 832, 864, 896, 960, 1024\n\nThe ensemble is done by concatenating single models feature, before computing cosine similarity.\n\n# new_individual threshold\nSince final models are trained on 100% of training data, there's no validation scores. New_individual's thresholds therefore need to be tuned on the LB. \n\nWe noticed that different species have different levels of difficulty to predict, hence the optimal new_individual thresholds are different for different species. Tuning each species on the LB is riskly, so we: (1) tuned each species' *relative* threshold separately on CV, (2) tuned overall threshold level on LB, (3) adjust the LB-optimal overall threshold level by CV-optimal species relative thresholds.\n\nThis adjustment boosted our public LB from 0.882 or 0.883 to 0.888\n\nWe used an ensemble of 15 models’ species head to predict test set's species.\n\n# Things that didn't work\n- Vision Transformers\n- Reranking Post processing\n- Using the previous humpback whale competition's data to pretrain\n- Use HFlip to double the number of individual_ids, a trick that worked well in previous humpback whale competition",
      "votes": null
    },
    {
      "id": "1761335",
      "postDate": "04/19/2022 20:04:10",
      "content": "<p>Congratulations! Great result and solution description. I am reading and … probably have more question soon. First - are you implemented training in Pytorch or Tensorflow? Have you used TPU or just GPU?</p>",
      "rawMarkdown": "Congratulations! Great result and solution description. I am reading and ... probably have more question soon. First - are you implemented training in Pytorch or Tensorflow? Have you used TPU or just GPU?",
      "votes": null
    },
    {
      "id": "1761344",
      "postDate": "04/19/2022 20:11:47",
      "content": "<p>Thanks. We used pytorch with <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library. We used GPUs only. The largest models were trained on 8x V100s with 32G mem. Some smaller models including yolov5 were trained on 4x A100s with 40G mem.</p>",
      "rawMarkdown": "Thanks. We used pytorch with [timm](https://github.com/rwightman/pytorch-image-models) library. We used GPUs only. The largest models were trained on 8x V100s with 32G mem. Some smaller models including yolov5 were trained on 4x A100s with 40G mem.",
      "votes": null
    },
    {
      "id": "1761374",
      "postDate": "04/19/2022 20:49:03",
      "content": "<p>Thank you very much for answer. We were talking internally with <a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> about your possible solution implementation - Pytorch / TF :) (since we are Pytorch lovers but…. this competition was for TF lovers - we used TPU) I was 100% sure that you use Pytorch. I know this is not the most important thing in this competition (maybe I am wrong) but I had feeling that you use Pytorch :)</p>\n<p>Anyway … I am impressed by your effectiveness. Tomorrow I will post more question. Thank you again and congratulations. You are great team! <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> 👍👍👍</p>",
      "rawMarkdown": "Thank you very much for answer. We were talking internally with @lukaszborecki about your possible solution implementation - Pytorch / TF :) (since we are Pytorch lovers but.... this competition was for TF lovers - we used TPU) I was 100% sure that you use Pytorch. I know this is not the most important thing in this competition (maybe I am wrong) but I had feeling that you use Pytorch :)\n\nAnyway ... I am impressed by your effectiveness. Tomorrow I will post more question. Thank you again and congratulations. You are great team! @haqishen 👍👍👍",
      "votes": null
    },
    {
      "id": "1761399",
      "postDate": "04/19/2022 21:45:39",
      "content": "<p>Thanks. Yeah I much prefer Pytorch. The first thing I do when I start working on a TF project is to convert everything into pytorch :) </p>",
      "rawMarkdown": "Thanks. Yeah I much prefer Pytorch. The first thing I do when I start working on a TF project is to convert everything into pytorch :)",
      "votes": null
    },
    {
      "id": "1761414",
      "postDate": "04/19/2022 22:20:12",
      "content": "<p>Congrats and thanks for your description! May I ask if you monitored top1 or top5 Acc and could share your results? <br>\nAnd if it is possible to share an inference pipeline? Actually, I followed nearly the same process and modified my code from yours and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 's repo, but the results were much lower than yours. I wonder if I did something wrong in inference and appreciate it if you can help. </p>",
      "rawMarkdown": "Congrats and thanks for your description! May I ask if you monitored top1 or top5 Acc and could share your results? \nAnd if it is possible to share an inference pipeline? Actually, I followed nearly the same process and modified my code from yours and @christofhenkel 's repo, but the results were much lower than yours. I wonder if I did something wrong in inference and appreciate it if you can help.",
      "votes": null
    },
    {
      "id": "1761471",
      "postDate": "04/20/2022 00:50:37",
      "content": "<p>Congrats, Bo and Qishen.<br>\nI learned a lot from you, All Data Are Ext.<br>\nI deeply thank you for sharing your creative solutions.  </p>",
      "rawMarkdown": "Congrats, Bo and Qishen.\nI learned a lot from you, All Data Are Ext.\nI deeply thank you for sharing your creative solutions.",
      "votes": null
    },
    {
      "id": "1761525",
      "postDate": "04/20/2022 01:37:59",
      "content": "<p>Thank you. We haven't tracked top1 or top5 accuracy, only the map5 score.</p>\n<p>We don't plan to open source the code this time. We didn't do anything special during inference. It's similar to landmark. I'm afraid it's hard to say what went wrong with your pipeline. It could be that the hyperparameters are not well tuned, the quality of the crop dataset is not high, …, so many places where things may go wrong.</p>",
      "rawMarkdown": "Thank you. We haven't tracked top1 or top5 accuracy, only the map5 score.\n\nWe don't plan to open source the code this time. We didn't do anything special during inference. It's similar to landmark. I'm afraid it's hard to say what went wrong with your pipeline. It could be that the hyperparameters are not well tuned, the quality of the crop dataset is not high, ..., so many places where things may go wrong.",
      "votes": null
    },
    {
      "id": "1761543",
      "postDate": "04/20/2022 01:55:31",
      "content": "<p>Ok got it. Thanks.</p>",
      "rawMarkdown": "Ok got it. Thanks.",
      "votes": null
    },
    {
      "id": "1762088",
      "postDate": "04/20/2022 12:29:07",
      "content": "<p>Good job. Would you please tell me what the skills or training ConvNext ?</p>",
      "rawMarkdown": "Good job. Would you please tell me what the skills or training ConvNext ?",
      "votes": null
    },
    {
      "id": "1762094",
      "postDate": "04/20/2022 12:34:12",
      "content": "<p>Congrats!<br>\nWhile results of tf and pytorch should be very similar (maybe with the exception of some default initialization), imho, pytorch is (or at least was) much more flexible which helps during prototyping.</p>",
      "rawMarkdown": "Congrats!\nWhile results of tf and pytorch should be very similar (maybe with the exception of some default initialization), imho, pytorch is (or at least was) much more flexible which helps during prototyping.",
      "votes": null
    },
    {
      "id": "1762137",
      "postDate": "04/20/2022 13:08:54",
      "content": "<p>Thanks. Nothing special to train ConvNext. We used the <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library, which makes it super easy to exchange backbones. The backbones we used are <code>convnext_xlarge_384_in22ft1k</code>, <code>convnext_large_384_in22ft1k</code>, <code>convnext_base_384_in22ft1k</code></p>\n<p>People reported good ConvNext performance in a previous competition, I think PetFinder. I also got equal or better results with it on a non-competition project than EffNets.</p>",
      "rawMarkdown": "Thanks. Nothing special to train ConvNext. We used the [timm](https://github.com/rwightman/pytorch-image-models) library, which makes it super easy to exchange backbones. The backbones we used are `convnext_xlarge_384_in22ft1k`, `convnext_large_384_in22ft1k`, `convnext_base_384_in22ft1k`\n\nPeople reported good ConvNext performance in a previous competition, I think PetFinder. I also got equal or better results with it on a non-competition project than EffNets.",
      "votes": null
    },
    {
      "id": "1762138",
      "postDate": "04/20/2022 13:10:29",
      "content": "<p>Thanks. Yeah, couldn't agree more. 💯</p>",
      "rawMarkdown": "Thanks. Yeah, couldn't agree more. 💯",
      "votes": null
    },
    {
      "id": "1762556",
      "postDate": "04/20/2022 19:36:52",
      "content": "<p>Great work comes from a great team!</p>\n<p>Thank you so much for sharing this information with us!</p>",
      "rawMarkdown": "Great work comes from a great team!\n\nThank you so much for sharing this information with us!",
      "votes": null
    },
    {
      "id": "1766261",
      "postDate": "04/24/2022 11:10:51",
      "content": "<p>Great Work. <br>\nThank you so much for sharing this information with us!</p>",
      "rawMarkdown": "Great Work. \nThank you so much for sharing this information with us!",
      "votes": null
    },
    {
      "id": "1852852",
      "postDate": "07/12/2022 12:29:30",
      "content": "<p><code>Reranking Post processing</code><br>\ncan you explain it a bit, or give any resource to study, thanks!</p>",
      "rawMarkdown": "`Reranking Post processing`\ncan you explain it a bit, or give any resource to study, thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1761335,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/19/2022 20:04:10",
      "content": "<p>Congratulations! Great result and solution description. I am reading and … probably have more question soon. First - are you implemented training in Pytorch or Tensorflow? Have you used TPU or just GPU?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1761344,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "04/19/2022 20:11:47",
          "content": "<p>Thanks. We used pytorch with <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library. We used GPUs only. The largest models were trained on 8x V100s with 32G mem. Some smaller models including yolov5 were trained on 4x A100s with 40G mem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761374,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "04/19/2022 20:49:03",
          "content": "<p>Thank you very much for answer. We were talking internally with <a href=\"https://www.kaggle.com/lukaszborecki\" target=\"_blank\">@lukaszborecki</a> about your possible solution implementation - Pytorch / TF :) (since we are Pytorch lovers but…. this competition was for TF lovers - we used TPU) I was 100% sure that you use Pytorch. I know this is not the most important thing in this competition (maybe I am wrong) but I had feeling that you use Pytorch :)</p>\n<p>Anyway … I am impressed by your effectiveness. Tomorrow I will post more question. Thank you again and congratulations. You are great team! <a href=\"https://www.kaggle.com/haqishen\" target=\"_blank\">@haqishen</a> 👍👍👍</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761399,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "04/19/2022 21:45:39",
          "content": "<p>Thanks. Yeah I much prefer Pytorch. The first thing I do when I start working on a TF project is to convert everything into pytorch :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1762094,
          "author_name": "ilu000",
          "author_url": "",
          "post_date": "04/20/2022 12:34:12",
          "content": "<p>Congrats!<br>\nWhile results of tf and pytorch should be very similar (maybe with the exception of some default initialization), imho, pytorch is (or at least was) much more flexible which helps during prototyping.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1762138,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "04/20/2022 13:10:29",
          "content": "<p>Thanks. Yeah, couldn't agree more. 💯</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761414,
      "author_name": "leonshangguan",
      "author_url": "",
      "post_date": "04/19/2022 22:20:12",
      "content": "<p>Congrats and thanks for your description! May I ask if you monitored top1 or top5 Acc and could share your results? <br>\nAnd if it is possible to share an inference pipeline? Actually, I followed nearly the same process and modified my code from yours and <a href=\"https://www.kaggle.com/christofhenkel\" target=\"_blank\">@christofhenkel</a> 's repo, but the results were much lower than yours. I wonder if I did something wrong in inference and appreciate it if you can help. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1761525,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "04/20/2022 01:37:59",
          "content": "<p>Thank you. We haven't tracked top1 or top5 accuracy, only the map5 score.</p>\n<p>We don't plan to open source the code this time. We didn't do anything special during inference. It's similar to landmark. I'm afraid it's hard to say what went wrong with your pipeline. It could be that the hyperparameters are not well tuned, the quality of the crop dataset is not high, …, so many places where things may go wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761543,
          "author_name": "leonshangguan",
          "author_url": "",
          "post_date": "04/20/2022 01:55:31",
          "content": "<p>Ok got it. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761471,
      "author_name": "yoichi7yamakawa",
      "author_url": "",
      "post_date": "04/20/2022 00:50:37",
      "content": "<p>Congrats, Bo and Qishen.<br>\nI learned a lot from you, All Data Are Ext.<br>\nI deeply thank you for sharing your creative solutions.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1762088,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "04/20/2022 12:29:07",
      "content": "<p>Good job. Would you please tell me what the skills or training ConvNext ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1762137,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "04/20/2022 13:08:54",
          "content": "<p>Thanks. Nothing special to train ConvNext. We used the <a href=\"https://github.com/rwightman/pytorch-image-models\" target=\"_blank\">timm</a> library, which makes it super easy to exchange backbones. The backbones we used are <code>convnext_xlarge_384_in22ft1k</code>, <code>convnext_large_384_in22ft1k</code>, <code>convnext_base_384_in22ft1k</code></p>\n<p>People reported good ConvNext performance in a previous competition, I think PetFinder. I also got equal or better results with it on a non-competition project than EffNets.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1762556,
      "author_name": "dsxavier",
      "author_url": "",
      "post_date": "04/20/2022 19:36:52",
      "content": "<p>Great work comes from a great team!</p>\n<p>Thank you so much for sharing this information with us!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1766261,
      "author_name": "amirhosainmahdiani",
      "author_url": "",
      "post_date": "04/24/2022 11:10:51",
      "content": "<p>Great Work. <br>\nThank you so much for sharing this information with us!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1852852,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "07/12/2022 12:29:30",
      "content": "<p><code>Reranking Post processing</code><br>\ncan you explain it a bit, or give any resource to study, thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1761302": "Congrats to all the winners. It has been a tough competition. The top 2 teams had a very strong last push. Thanks for the great collaboration once again, my long time teammate @haqishen \n\n# TL;DR\n- Yolov5 Detection\n- Dynamic Margin ArcFace with DOLG CNN backbone\n - Best backbone is ConvNext\n- Two heads: predicting ids and species\n- 3 rounds of pseudo labeling\n- Tune new_individual threshold separately for each species\n\n# Detection\nWe trained yolov5m with image size 512 for 40 epochs. We started by using this [public notebook's](https://www.kaggle.com/datasets/awsaf49/happywhale-boundingbox-yolov5-dataset) predictions as labels. Then visualize examples with low OOF confidence. If the predicted bbox are wrong, remove this example from training set, or fix the ground truth bbox. \n\nWe iterate this for 9 rounds. In the end, most OOF predictions look correct. The OOF iou = 0.93863. But some test set prediction bboxs are still off. In hindsight, we probably should start from scratch and label thousands of images, like other top teams did.\n\nAfter getting the predicted bbox, we extend it by 20% and feed the crop to arcface model.\n\n# Model\nThe modeling part is heavily influenced by the recent landmarks competitions. The architecture is Dynamic Margin ArcFace with DOLG CNN backbone. The dynamic margin arcface was introduced by us in last year's Landmark, see detail [here](https://www.kaggle.com/competitions/landmark-recognition-2020/discussion/187757). The DOLG was introduced to the Kaggle community by @christofhenkel in this year's Landmark, see detail [here](https://www.kaggle.com/competitions/landmark-retrieval-2021/discussion/277099).\n\nOther components of the top landmark solutions didn't work here though, including sub center arcface, and vision transformers. All the vision transformers underperform CNNs. The best CNN in our solution is ConvNext.\n\nThe only twist in the model is that we output two heads, predicting both individual_id and species.\n\nFor augmentations we used the following plus mixup\n```\n    A.HorizontalFlip(p=0.5),\n    A.RandomContrast(limit=0.2, p=0.75),\n    A.ShiftScaleRotate(shift_limit=0.0, scale_limit=0.3, rotate_limit=10, border_mode=0, p=0.7),\n    A.HueSaturationValue(hue_shift_limit=20, sat_shift_limit=30, val_shift_limit=20, p=0.5),\n```\n\n# Training steps\n1. Tune everything on 5 fold. i.e. train on 80% of training data per fold, \"sacrificing\" some individual_ids in order to do cross validation.\n2. Ensemble 9 fold models to make pseudo labels.\n3. Train 6 best models on 100% training data + pseudo labeled test data\n4. Make new pseudo labels from step 3 ensemble\n5. Repeat step 3-4 for two more rounds\n6. Ensemble the 12 models from last two rounds.\n\nAverage single model's public LB in 3 pseudo label rounds are 0.855, 0.862, 0.865\n\n# Ensemble\nThe best performing backbone is ConvNext. We also trained 3 non-ConvNext for diversity. The final six models are ConvNext Base, Large, XLarge, EfficientNet B7, V2L, NFNet L2. We used diversified image sizes for different models: 640, 672, 704, 736, 768, 800, 832, 864, 896, 960, 1024\n\nThe ensemble is done by concatenating single models feature, before computing cosine similarity.\n\n# new_individual threshold\nSince final models are trained on 100% of training data, there's no validation scores. New_individual's thresholds therefore need to be tuned on the LB. \n\nWe noticed that different species have different levels of difficulty to predict, hence the optimal new_individual thresholds are different for different species. Tuning each species on the LB is riskly, so we: (1) tuned each species' *relative* threshold separately on CV, (2) tuned overall threshold level on LB, (3) adjust the LB-optimal overall threshold level by CV-optimal species relative thresholds.\n\nThis adjustment boosted our public LB from 0.882 or 0.883 to 0.888\n\nWe used an ensemble of 15 models’ species head to predict test set's species.\n\n# Things that didn't work\n- Vision Transformers\n- Reranking Post processing\n- Using the previous humpback whale competition's data to pretrain\n- Use HFlip to double the number of individual_ids, a trick that worked well in previous humpback whale competition",
    "1761335": "Congratulations! Great result and solution description. I am reading and ... probably have more question soon. First - are you implemented training in Pytorch or Tensorflow? Have you used TPU or just GPU?",
    "1761344": "Thanks. We used pytorch with [timm](https://github.com/rwightman/pytorch-image-models) library. We used GPUs only. The largest models were trained on 8x V100s with 32G mem. Some smaller models including yolov5 were trained on 4x A100s with 40G mem.",
    "1761374": "Thank you very much for answer. We were talking internally with @lukaszborecki about your possible solution implementation - Pytorch / TF :) (since we are Pytorch lovers but.... this competition was for TF lovers - we used TPU) I was 100% sure that you use Pytorch. I know this is not the most important thing in this competition (maybe I am wrong) but I had feeling that you use Pytorch :)\n\nAnyway ... I am impressed by your effectiveness. Tomorrow I will post more question. Thank you again and congratulations. You are great team! @haqishen 👍👍👍",
    "1761399": "Thanks. Yeah I much prefer Pytorch. The first thing I do when I start working on a TF project is to convert everything into pytorch :)",
    "1761414": "Congrats and thanks for your description! May I ask if you monitored top1 or top5 Acc and could share your results? \nAnd if it is possible to share an inference pipeline? Actually, I followed nearly the same process and modified my code from yours and @christofhenkel 's repo, but the results were much lower than yours. I wonder if I did something wrong in inference and appreciate it if you can help.",
    "1761471": "Congrats, Bo and Qishen.\nI learned a lot from you, All Data Are Ext.\nI deeply thank you for sharing your creative solutions.",
    "1761525": "Thank you. We haven't tracked top1 or top5 accuracy, only the map5 score.\n\nWe don't plan to open source the code this time. We didn't do anything special during inference. It's similar to landmark. I'm afraid it's hard to say what went wrong with your pipeline. It could be that the hyperparameters are not well tuned, the quality of the crop dataset is not high, ..., so many places where things may go wrong.",
    "1761543": "Ok got it. Thanks.",
    "1762088": "Good job. Would you please tell me what the skills or training ConvNext ?",
    "1762094": "Congrats!\nWhile results of tf and pytorch should be very similar (maybe with the exception of some default initialization), imho, pytorch is (or at least was) much more flexible which helps during prototyping.",
    "1762137": "Thanks. Nothing special to train ConvNext. We used the [timm](https://github.com/rwightman/pytorch-image-models) library, which makes it super easy to exchange backbones. The backbones we used are `convnext_xlarge_384_in22ft1k`, `convnext_large_384_in22ft1k`, `convnext_base_384_in22ft1k`\n\nPeople reported good ConvNext performance in a previous competition, I think PetFinder. I also got equal or better results with it on a non-competition project than EffNets.",
    "1762138": "Thanks. Yeah, couldn't agree more. 💯",
    "1762556": "Great work comes from a great team!\n\nThank you so much for sharing this information with us!",
    "1766261": "Great Work. \nThank you so much for sharing this information with us!",
    "1852852": "`Reranking Post processing`\ncan you explain it a bit, or give any resource to study, thanks!"
  },
  "source": "meta"
}