{
  "id": 320407,
  "title": "14th place solution",
  "url": "/competitions/happy-whale-and-dolphin/writeups/nk35jk-14th-place-solution",
  "author_name": "",
  "post_date": "2022-04-21T13:39:51.849046600Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to the organizers for hosting the interesting competition and congratulations to all the winners. It was tough, but was a great experience for me.</p>\n<h2>Datasets</h2>\n<p>Three Yolov5 models are trained to detect fullbody and backfin, respectively.<br>\nScores improved greatly with changing datasets as: original image -&gt; detic crop -&gt; fullbody crop -&gt; fullbody/backfin crop.</p>\n<h2>Models</h2>\n<ul>\n<li>ArcFace models are trained using fullbody crop (640×640) and backfin crop (448×448), respectively. The parameters of ArcFace are set as scale = 25 and margin = 0.5 (The large margin worked with improved datasets in my case). Embedding size is 2048.</li>\n<li>Focal loss for species classification (loss weight = 0.1) is also used.</li>\n<li>Backbone: Efficientnet-b7 and ConvNeXt-L</li>\n<li>In order to make embeddings more discriminative, multiply the feature maps by the attention weights computed with GAP features as a query. This worked only with Efficientnet.</li>\n<li>Pseudo labeling: Take about 30% top predictions (would be better to use more data and multiple iterations).</li>\n<li>Distillation: Use the features of the teacher model as soft target and compute MSE. This greatly improves the performance in the early stages of training, but the contribution to the final score does not seem to be that large, so adopt cosine schedule to set the final loss weight to 0.</li>\n</ul>\n<h3>Augmentation</h3>\n<p>Average blur, motion blur, gaussian noise, saturation, brightness, contrast, grayscale and affine transform (flip, rotation, shear, scaling and translation) are used.</p>\n<h3>Training details</h3>\n<ul>\n<li>Hyper parameters: AdamW with weight decay = 0.05, base lr = 2e-4, cosine scheduling with linear warmup, batch size = 24 along with gradient accumulation</li>\n<li>PyTorch is used and training 30 epochs with a single A100 GPU take about 30 hours for a fullbody model and 15 hours for a backfin model.</li>\n<li>Private LB of the best single fullbody model is 0.781.</li>\n</ul>\n<h2>Post-process</h2>\n<ol>\n<li>Cosine similarity matrix for the individuals is computed using the features of the test images and ArcFace weights using each model.</li>\n<li>Then average the matrices to ensemble eight fullbody/backfiin models.</li>\n<li>Insert new_individual with a fixed threshold (should have changed threshold for each species).</li>\n<li>Finally, switch adjacent top predictions based on the difference of cosine similarity and the degree of assignment to the top for class balancing.</li>\n</ol>\n<h2>What did not work</h2>\n<p>The following did not work in my case.</p>\n<ul>\n<li>Elasticface</li>\n<li>Magface</li>\n<li>Setting sample-wise margin based on the magnitude of losses</li>\n<li>Triplet loss</li>\n<li>Semi-supervised learning with self-distillation loss of DINO (somewhat worked, but employed pseudo-labeling)</li>\n<li>Adding a SOD mask channel to input</li>\n<li>Adding an attention weight channel to input</li>\n</ul>\n<p>Thank you for reading.</p>",
  "messages": [
    {
      "id": "1763399",
      "postDate": "04/21/2022 13:39:51",
      "content": "<p>Thanks to the organizers for hosting the interesting competition and congratulations to all the winners. It was tough, but was a great experience for me.</p>\n<h2>Datasets</h2>\n<p>Three Yolov5 models are trained to detect fullbody and backfin, respectively.<br>\nScores improved greatly with changing datasets as: original image -&gt; detic crop -&gt; fullbody crop -&gt; fullbody/backfin crop.</p>\n<h2>Models</h2>\n<ul>\n<li>ArcFace models are trained using fullbody crop (640×640) and backfin crop (448×448), respectively. The parameters of ArcFace are set as scale = 25 and margin = 0.5 (The large margin worked with improved datasets in my case). Embedding size is 2048.</li>\n<li>Focal loss for species classification (loss weight = 0.1) is also used.</li>\n<li>Backbone: Efficientnet-b7 and ConvNeXt-L</li>\n<li>In order to make embeddings more discriminative, multiply the feature maps by the attention weights computed with GAP features as a query. This worked only with Efficientnet.</li>\n<li>Pseudo labeling: Take about 30% top predictions (would be better to use more data and multiple iterations).</li>\n<li>Distillation: Use the features of the teacher model as soft target and compute MSE. This greatly improves the performance in the early stages of training, but the contribution to the final score does not seem to be that large, so adopt cosine schedule to set the final loss weight to 0.</li>\n</ul>\n<h3>Augmentation</h3>\n<p>Average blur, motion blur, gaussian noise, saturation, brightness, contrast, grayscale and affine transform (flip, rotation, shear, scaling and translation) are used.</p>\n<h3>Training details</h3>\n<ul>\n<li>Hyper parameters: AdamW with weight decay = 0.05, base lr = 2e-4, cosine scheduling with linear warmup, batch size = 24 along with gradient accumulation</li>\n<li>PyTorch is used and training 30 epochs with a single A100 GPU take about 30 hours for a fullbody model and 15 hours for a backfin model.</li>\n<li>Private LB of the best single fullbody model is 0.781.</li>\n</ul>\n<h2>Post-process</h2>\n<ol>\n<li>Cosine similarity matrix for the individuals is computed using the features of the test images and ArcFace weights using each model.</li>\n<li>Then average the matrices to ensemble eight fullbody/backfiin models.</li>\n<li>Insert new_individual with a fixed threshold (should have changed threshold for each species).</li>\n<li>Finally, switch adjacent top predictions based on the difference of cosine similarity and the degree of assignment to the top for class balancing.</li>\n</ol>\n<h2>What did not work</h2>\n<p>The following did not work in my case.</p>\n<ul>\n<li>Elasticface</li>\n<li>Magface</li>\n<li>Setting sample-wise margin based on the magnitude of losses</li>\n<li>Triplet loss</li>\n<li>Semi-supervised learning with self-distillation loss of DINO (somewhat worked, but employed pseudo-labeling)</li>\n<li>Adding a SOD mask channel to input</li>\n<li>Adding an attention weight channel to input</li>\n</ul>\n<p>Thank you for reading.</p>",
      "rawMarkdown": "Thanks to the organizers for hosting the interesting competition and congratulations to all the winners. It was tough, but was a great experience for me.\n\n## Datasets\n\nThree Yolov5 models are trained to detect fullbody and backfin, respectively.\nScores improved greatly with changing datasets as: original image -> detic crop -> fullbody crop -> fullbody/backfin crop.\n\n## Models\n\n- ArcFace models are trained using fullbody crop (640×640) and backfin crop (448×448), respectively. The parameters of ArcFace are set as scale = 25 and margin = 0.5 (The large margin worked with improved datasets in my case). Embedding size is 2048.\n- Focal loss for species classification (loss weight = 0.1) is also used.\n- Backbone: Efficientnet-b7 and ConvNeXt-L\n- In order to make embeddings more discriminative, multiply the feature maps by the attention weights computed with GAP features as a query. This worked only with Efficientnet.\n- Pseudo labeling: Take about 30% top predictions (would be better to use more data and multiple iterations).\n- Distillation: Use the features of the teacher model as soft target and compute MSE. This greatly improves the performance in the early stages of training, but the contribution to the final score does not seem to be that large, so adopt cosine schedule to set the final loss weight to 0.\n\n### Augmentation\n\nAverage blur, motion blur, gaussian noise, saturation, brightness, contrast, grayscale and affine transform (flip, rotation, shear, scaling and translation) are used.\n\n### Training details\n\n- Hyper parameters: AdamW with weight decay = 0.05, base lr = 2e-4, cosine scheduling with linear warmup, batch size = 24 along with gradient accumulation\n- PyTorch is used and training 30 epochs with a single A100 GPU take about 30 hours for a fullbody model and 15 hours for a backfin model.\n- Private LB of the best single fullbody model is 0.781.\n\n## Post-process\n\n1. Cosine similarity matrix for the individuals is computed using the features of the test images and ArcFace weights using each model.\n1. Then average the matrices to ensemble eight fullbody/backfiin models.\n1. Insert new_individual with a fixed threshold (should have changed threshold for each species).\n1. Finally, switch adjacent top predictions based on the difference of cosine similarity and the degree of assignment to the top for class balancing.\n\n## What did not work\n\nThe following did not work in my case.\n\n- Elasticface\n- Magface\n- Setting sample-wise margin based on the magnitude of losses\n- Triplet loss\n- Semi-supervised learning with self-distillation loss of DINO (somewhat worked, but employed pseudo-labeling)\n- Adding a SOD mask channel to input\n- Adding an attention weight channel to input\n\nThank you for reading.",
      "votes": null
    },
    {
      "id": "1763439",
      "postDate": "04/21/2022 14:26:30",
      "content": "<p>You did the very good job at first competition. Great!<br>\nI learned a lot from you!</p>",
      "rawMarkdown": "You did the very good job at first competition. Great!\nI learned a lot from you!",
      "votes": null
    },
    {
      "id": "1763450",
      "postDate": "04/21/2022 14:39:58",
      "content": "<p>Congratulations!! Thanks for sharing</p>",
      "rawMarkdown": "Congratulations!! Thanks for sharing",
      "votes": null
    },
    {
      "id": "1763458",
      "postDate": "04/21/2022 14:46:52",
      "content": "<p>I'm glad to hear that! Your posts really helped me.</p>",
      "rawMarkdown": "I'm glad to hear that! Your posts really helped me.",
      "votes": null
    },
    {
      "id": "1763460",
      "postDate": "04/21/2022 14:47:19",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1763439,
      "author_name": "deepkim",
      "author_url": "",
      "post_date": "04/21/2022 14:26:30",
      "content": "<p>You did the very good job at first competition. Great!<br>\nI learned a lot from you!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1763458,
          "author_name": "nk35jk",
          "author_url": "",
          "post_date": "04/21/2022 14:46:52",
          "content": "<p>I'm glad to hear that! Your posts really helped me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1763450,
      "author_name": "ayessa",
      "author_url": "",
      "post_date": "04/21/2022 14:39:58",
      "content": "<p>Congratulations!! Thanks for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 1763460,
          "author_name": "nk35jk",
          "author_url": "",
          "post_date": "04/21/2022 14:47:19",
          "content": "<p>Thanks a lot!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1763399": "Thanks to the organizers for hosting the interesting competition and congratulations to all the winners. It was tough, but was a great experience for me.\n\n## Datasets\n\nThree Yolov5 models are trained to detect fullbody and backfin, respectively.\nScores improved greatly with changing datasets as: original image -> detic crop -> fullbody crop -> fullbody/backfin crop.\n\n## Models\n\n- ArcFace models are trained using fullbody crop (640×640) and backfin crop (448×448), respectively. The parameters of ArcFace are set as scale = 25 and margin = 0.5 (The large margin worked with improved datasets in my case). Embedding size is 2048.\n- Focal loss for species classification (loss weight = 0.1) is also used.\n- Backbone: Efficientnet-b7 and ConvNeXt-L\n- In order to make embeddings more discriminative, multiply the feature maps by the attention weights computed with GAP features as a query. This worked only with Efficientnet.\n- Pseudo labeling: Take about 30% top predictions (would be better to use more data and multiple iterations).\n- Distillation: Use the features of the teacher model as soft target and compute MSE. This greatly improves the performance in the early stages of training, but the contribution to the final score does not seem to be that large, so adopt cosine schedule to set the final loss weight to 0.\n\n### Augmentation\n\nAverage blur, motion blur, gaussian noise, saturation, brightness, contrast, grayscale and affine transform (flip, rotation, shear, scaling and translation) are used.\n\n### Training details\n\n- Hyper parameters: AdamW with weight decay = 0.05, base lr = 2e-4, cosine scheduling with linear warmup, batch size = 24 along with gradient accumulation\n- PyTorch is used and training 30 epochs with a single A100 GPU take about 30 hours for a fullbody model and 15 hours for a backfin model.\n- Private LB of the best single fullbody model is 0.781.\n\n## Post-process\n\n1. Cosine similarity matrix for the individuals is computed using the features of the test images and ArcFace weights using each model.\n1. Then average the matrices to ensemble eight fullbody/backfiin models.\n1. Insert new_individual with a fixed threshold (should have changed threshold for each species).\n1. Finally, switch adjacent top predictions based on the difference of cosine similarity and the degree of assignment to the top for class balancing.\n\n## What did not work\n\nThe following did not work in my case.\n\n- Elasticface\n- Magface\n- Setting sample-wise margin based on the magnitude of losses\n- Triplet loss\n- Semi-supervised learning with self-distillation loss of DINO (somewhat worked, but employed pseudo-labeling)\n- Adding a SOD mask channel to input\n- Adding an attention weight channel to input\n\nThank you for reading.",
    "1763439": "You did the very good job at first competition. Great!\nI learned a lot from you!",
    "1763450": "Congratulations!! Thanks for sharing",
    "1763458": "I'm glad to hear that! Your posts really helped me.",
    "1763460": "Thanks a lot!"
  },
  "source": "meta"
}