{
  "id": 321322,
  "title": "55th place solution",
  "url": "/competitions/happy-whale-and-dolphin/writeups/kakiage-onigiri-55th-place-solution",
  "author_name": "",
  "post_date": "2022-04-26T09:22:06.194716900Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congratulations to all the winners. Thanks to Kaggle for hosting such a great competition.<br>\nThis is our first time to get medal at Kaggle.</p>\n<p>Here is our solution.</p>\n<h1>Dataset</h1>\n<p>We used fullbody and backfin dataset created by <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>. We thank <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> a lot, because these dataset contributed to boost our score.</p>\n<h1>Models</h1>\n<p>We trained 5 types of model with Arcface.</p>\n<ul>\n<li>Efficient-netV1B5</li>\n<li>Efficient-netV1B6</li>\n<li>Efficient-netV1B7</li>\n<li>Efficient-netV2L</li>\n<li>ConvNext</li>\n</ul>\n<h1>Training</h1>\n<h2>Augmentation</h2>\n<ul>\n<li>Random Rotate(-10 to 10 degree)</li>\n<li>Random hue</li>\n<li>Random contrast</li>\n<li>Random brightness</li>\n</ul>\n<h2>Training settings</h2>\n<ul>\n<li>batch_size = 32</li>\n<li>image_size = 700</li>\n<li>Optimizer = Adam</li>\n<li>loss = SparseCategoricalFocalLoss (γ=1)</li>\n<li>embedding size = 2048</li>\n<li>Batch Normalization before Dropout</li>\n</ul>\n<h2>How to train models</h2>\n<p>We trained each models by 2 steps in each dataset. We used all train data (no validation data) for training models.</p>\n<ul>\n<li>1st step  <br>\nWe trained models with all augmentation for 30 epochs.</li>\n<li>2nd step  <br>\nWe trained models without Random Rotate augmentation for 25 epochs (In the case of Efficient-netV2L and ConvNext, 20 epochs).</li>\n</ul>\n<p>Finally, we got 10 models (5 models for 2 dataset).</p>\n<h1>Postprocess</h1>\n<p>We used trained models to get embeddings of training dataset and test dataset.</p>\n<ul>\n<li><p>TTA(Test Time Augmentation)  <br>\nWe used TTA to get embeddings from each trained models. We input 4 types of image into a trained model, and calculate the average of embeddings.</p>\n<ul>\n<li>normal image</li>\n<li>flip left right image</li>\n<li>-10 degree rotate</li>\n<li>10 degree rotate</li></ul></li>\n<li><p>L2 normalization  <br>\nWe calculated L2 normalization.</p></li>\n<li><p>Concat embeddings  <br>\nWe concatenated all embeddings in each dataset.</p></li>\n<li><p>K-NN  <br>\nIn fullbody dataset and backfin dataset, we calculated distance between training embeddings and test embeddings, and got top 100 (k=100) individual ids for each test ids.</p></li>\n<li><p>Ensamble   </p></li>\n</ul>\n<ol>\n<li>We calculated top 100 individuals confidence (confidence = 1 - distance) in each test ids</li>\n<li>We added backfin's confidences to fullbody's confidences.</li>\n<li>We inserted new_individual with a fixed threshold (threshold=0.5).</li>\n<li>We sorted confidences and extract top 5 individuals.</li>\n</ol>\n<h1>What did not work</h1>\n<h2>Models</h2>\n<ul>\n<li>ViT</li>\n<li>BEiT</li>\n<li>DeiT</li>\n<li>Swin Transformer(swin_large_384)</li>\n<li>Efficient-netV2XL</li>\n</ul>\n<h2>Training settings</h2>\n<ul>\n<li>GeMpooling</li>\n<li>CosFace</li>\n<li>AdaCos</li>\n<li>SparseCategoricalCrossentropy</li>\n<li>Flatten Layer</li>\n<li>RMSprop</li>\n</ul>\n<h2>Augmentaion</h2>\n<ul>\n<li>-90 to 90 degree randomrotate augmentation</li>\n<li>Cutout augmentation</li>\n<li>Mix stochastically YOLOv5 and Detic in happywhale-tfrecords-bb</li>\n<li>Random Crop</li>\n</ul>\n<h2>Postprocess</h2>\n<ul>\n<li>ZCA</li>\n<li>query expansion(αQE)</li>\n</ul>\n<h1>Computational resouces</h1>\n<ul>\n<li>Google Colab Pro +</li>\n</ul>\n<h1>Acknowkedgement</h1>\n<p>Thank you <a href=\"https://www.kaggle.com/youtaro\" target=\"_blank\">@youtaro</a> <a href=\"https://www.kaggle.com/shoheishimatani\" target=\"_blank\">@shoheishimatani</a> <a href=\"https://www.kaggle.com/room208\" target=\"_blank\">@room208</a> <a href=\"https://www.kaggle.com/aprilsixteen\" target=\"_blank\">@aprilsixteen</a> for this competition!</p>",
  "messages": [
    {
      "id": "1768422",
      "postDate": "04/26/2022 09:22:06",
      "content": "<p>Congratulations to all the winners. Thanks to Kaggle for hosting such a great competition.<br>\nThis is our first time to get medal at Kaggle.</p>\n<p>Here is our solution.</p>\n<h1>Dataset</h1>\n<p>We used fullbody and backfin dataset created by <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>. We thank <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a> a lot, because these dataset contributed to boost our score.</p>\n<h1>Models</h1>\n<p>We trained 5 types of model with Arcface.</p>\n<ul>\n<li>Efficient-netV1B5</li>\n<li>Efficient-netV1B6</li>\n<li>Efficient-netV1B7</li>\n<li>Efficient-netV2L</li>\n<li>ConvNext</li>\n</ul>\n<h1>Training</h1>\n<h2>Augmentation</h2>\n<ul>\n<li>Random Rotate(-10 to 10 degree)</li>\n<li>Random hue</li>\n<li>Random contrast</li>\n<li>Random brightness</li>\n</ul>\n<h2>Training settings</h2>\n<ul>\n<li>batch_size = 32</li>\n<li>image_size = 700</li>\n<li>Optimizer = Adam</li>\n<li>loss = SparseCategoricalFocalLoss (γ=1)</li>\n<li>embedding size = 2048</li>\n<li>Batch Normalization before Dropout</li>\n</ul>\n<h2>How to train models</h2>\n<p>We trained each models by 2 steps in each dataset. We used all train data (no validation data) for training models.</p>\n<ul>\n<li>1st step  <br>\nWe trained models with all augmentation for 30 epochs.</li>\n<li>2nd step  <br>\nWe trained models without Random Rotate augmentation for 25 epochs (In the case of Efficient-netV2L and ConvNext, 20 epochs).</li>\n</ul>\n<p>Finally, we got 10 models (5 models for 2 dataset).</p>\n<h1>Postprocess</h1>\n<p>We used trained models to get embeddings of training dataset and test dataset.</p>\n<ul>\n<li><p>TTA(Test Time Augmentation)  <br>\nWe used TTA to get embeddings from each trained models. We input 4 types of image into a trained model, and calculate the average of embeddings.</p>\n<ul>\n<li>normal image</li>\n<li>flip left right image</li>\n<li>-10 degree rotate</li>\n<li>10 degree rotate</li></ul></li>\n<li><p>L2 normalization  <br>\nWe calculated L2 normalization.</p></li>\n<li><p>Concat embeddings  <br>\nWe concatenated all embeddings in each dataset.</p></li>\n<li><p>K-NN  <br>\nIn fullbody dataset and backfin dataset, we calculated distance between training embeddings and test embeddings, and got top 100 (k=100) individual ids for each test ids.</p></li>\n<li><p>Ensamble   </p></li>\n</ul>\n<ol>\n<li>We calculated top 100 individuals confidence (confidence = 1 - distance) in each test ids</li>\n<li>We added backfin's confidences to fullbody's confidences.</li>\n<li>We inserted new_individual with a fixed threshold (threshold=0.5).</li>\n<li>We sorted confidences and extract top 5 individuals.</li>\n</ol>\n<h1>What did not work</h1>\n<h2>Models</h2>\n<ul>\n<li>ViT</li>\n<li>BEiT</li>\n<li>DeiT</li>\n<li>Swin Transformer(swin_large_384)</li>\n<li>Efficient-netV2XL</li>\n</ul>\n<h2>Training settings</h2>\n<ul>\n<li>GeMpooling</li>\n<li>CosFace</li>\n<li>AdaCos</li>\n<li>SparseCategoricalCrossentropy</li>\n<li>Flatten Layer</li>\n<li>RMSprop</li>\n</ul>\n<h2>Augmentaion</h2>\n<ul>\n<li>-90 to 90 degree randomrotate augmentation</li>\n<li>Cutout augmentation</li>\n<li>Mix stochastically YOLOv5 and Detic in happywhale-tfrecords-bb</li>\n<li>Random Crop</li>\n</ul>\n<h2>Postprocess</h2>\n<ul>\n<li>ZCA</li>\n<li>query expansion(αQE)</li>\n</ul>\n<h1>Computational resouces</h1>\n<ul>\n<li>Google Colab Pro +</li>\n</ul>\n<h1>Acknowkedgement</h1>\n<p>Thank you <a href=\"https://www.kaggle.com/youtaro\" target=\"_blank\">@youtaro</a> <a href=\"https://www.kaggle.com/shoheishimatani\" target=\"_blank\">@shoheishimatani</a> <a href=\"https://www.kaggle.com/room208\" target=\"_blank\">@room208</a> <a href=\"https://www.kaggle.com/aprilsixteen\" target=\"_blank\">@aprilsixteen</a> for this competition!</p>",
      "rawMarkdown": "Congratulations to all the winners. Thanks to Kaggle for hosting such a great competition.\nThis is our first time to get medal at Kaggle.\n\nHere is our solution.\n\n# Dataset\nWe used fullbody and backfin dataset created by @jpbremer. We thank @jpbremer a lot, because these dataset contributed to boost our score.\n\n# Models\nWe trained 5 types of model with Arcface.\n\n- Efficient-netV1B5\n- Efficient-netV1B6\n- Efficient-netV1B7\n- Efficient-netV2L\n- ConvNext\n\n# Training\n\n## Augmentation\n- Random Rotate(-10 to 10 degree)\n- Random hue\n- Random contrast\n- Random brightness\n\n## Training settings\n- batch_size = 32\n- image_size = 700\n- Optimizer = Adam\n- loss = SparseCategoricalFocalLoss (&gamma;=1)\n- embedding size = 2048\n- Batch Normalization before Dropout\n\n## How to train models\nWe trained each models by 2 steps in each dataset. We used all train data (no validation data) for training models.\n\n- 1st step  \n  We trained models with all augmentation for 30 epochs.\n- 2nd step  \n  We trained models without Random Rotate augmentation for 25 epochs (In the case of Efficient-netV2L and ConvNext, 20 epochs).\n\nFinally, we got 10 models (5 models for 2 dataset).\n\n# Postprocess\nWe used trained models to get embeddings of training dataset and test dataset.\n\n- TTA(Test Time Augmentation)  \n  We used TTA to get embeddings from each trained models. We input 4 types of image into a trained model, and calculate the average of embeddings.\n  - normal image\n  - flip left right image\n  - -10 degree rotate\n  - 10 degree rotate\n\n- L2 normalization  \n  We calculated L2 normalization.\n\n- Concat embeddings  \n  We concatenated all embeddings in each dataset.\n\n- K-NN  \n  In fullbody dataset and backfin dataset, we calculated distance between training embeddings and test embeddings, and got top 100 (k=100) individual ids for each test ids.\n\n- Ensamble   \n 1.  We calculated top 100 individuals confidence (confidence = 1 - distance) in each test ids\n 2. We added backfin's confidences to fullbody's confidences.\n 3. We inserted new_individual with a fixed threshold (threshold=0.5).\n 4. We sorted confidences and extract top 5 individuals.\n\n# What did not work\n## Models\n- ViT\n- BEiT\n- DeiT\n- Swin Transformer(swin_large_384)\n- Efficient-netV2XL\n\n## Training settings\n- GeMpooling\n- CosFace\n- AdaCos\n- SparseCategoricalCrossentropy\n- Flatten Layer\n- RMSprop\n\n## Augmentaion\n- -90 to 90 degree randomrotate augmentation\n- Cutout augmentation\n- Mix stochastically YOLOv5 and Detic in happywhale-tfrecords-bb\n- Random Crop\n\n## Postprocess \n- ZCA\n- query expansion(&alpha;QE)\n\n# Computational resouces\n- Google Colab Pro +\n\n# Acknowkedgement\nThank you @youtaro @shoheishimatani @room208 @aprilsixteen for this competition!",
      "votes": null
    },
    {
      "id": "1768552",
      "postDate": "04/26/2022 12:01:49",
      "content": "<p>A lot of work and experiments! Thank you for one more great solution desctiption.</p>\n<p>Could you show usage of SparseCategoricalFocalLoss? Is this one from focal_loss?</p>\n<p>TTA during the embedding retrieval stage cool! 👍 As far as I understand you get 4 embeddings for each image.</p>",
      "rawMarkdown": "A lot of work and experiments! Thank you for one more great solution desctiption.\n\nCould you show usage of SparseCategoricalFocalLoss? Is this one from focal_loss?\n\nTTA during the embedding retrieval stage cool! 👍 As far as I understand you get 4 embeddings for each image.",
      "votes": null
    },
    {
      "id": "1769727",
      "postDate": "04/27/2022 14:23:37",
      "content": "<p>Thank you for your questions.</p>\n<p>Yes. SparseCategoricalFocalLoss is one from focal_loss.<br>\nWe used <a href=\"https://focal-loss.readthedocs.io/en/latest/generated/focal_loss.SparseCategoricalFocalLoss.html\" target=\"_blank\">focal_loss</a> library. </p>\n<p>I'm sorry to cause your misunderstanding, because we got one embedding for each image by taking an average of 4 types of embeddings by TTA. We also tried getting 4 embeddings by TTA, but it did not improve our score.</p>",
      "rawMarkdown": "Thank you for your questions.\n\nYes. SparseCategoricalFocalLoss is one from focal_loss.\nWe used [focal_loss](https://focal-loss.readthedocs.io/en/latest/generated/focal_loss.SparseCategoricalFocalLoss.html) library. \n\nI'm sorry to cause your misunderstanding, because we got one embedding for each image by taking an average of 4 types of embeddings by TTA. We also tried getting 4 embeddings by TTA, but it did not improve our score.",
      "votes": null
    },
    {
      "id": "1770035",
      "postDate": "04/27/2022 22:19:01",
      "content": "<p>Congratulations. Great solution!</p>",
      "rawMarkdown": "Congratulations. Great solution!",
      "votes": null
    },
    {
      "id": "1785994",
      "postDate": "05/12/2022 14:41:46",
      "content": "<p>Very informative solution post. Thank you for sharing!</p>",
      "rawMarkdown": "Very informative solution post. Thank you for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1768552,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/26/2022 12:01:49",
      "content": "<p>A lot of work and experiments! Thank you for one more great solution desctiption.</p>\n<p>Could you show usage of SparseCategoricalFocalLoss? Is this one from focal_loss?</p>\n<p>TTA during the embedding retrieval stage cool! 👍 As far as I understand you get 4 embeddings for each image.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1769727,
          "author_name": "koheiogino",
          "author_url": "",
          "post_date": "04/27/2022 14:23:37",
          "content": "<p>Thank you for your questions.</p>\n<p>Yes. SparseCategoricalFocalLoss is one from focal_loss.<br>\nWe used <a href=\"https://focal-loss.readthedocs.io/en/latest/generated/focal_loss.SparseCategoricalFocalLoss.html\" target=\"_blank\">focal_loss</a> library. </p>\n<p>I'm sorry to cause your misunderstanding, because we got one embedding for each image by taking an average of 4 types of embeddings by TTA. We also tried getting 4 embeddings by TTA, but it did not improve our score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1770035,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "04/27/2022 22:19:01",
      "content": "<p>Congratulations. Great solution!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1785994,
      "author_name": "",
      "author_url": "",
      "post_date": "05/12/2022 14:41:46",
      "content": "<p>Very informative solution post. Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1768422": "Congratulations to all the winners. Thanks to Kaggle for hosting such a great competition.\nThis is our first time to get medal at Kaggle.\n\nHere is our solution.\n\n# Dataset\nWe used fullbody and backfin dataset created by @jpbremer. We thank @jpbremer a lot, because these dataset contributed to boost our score.\n\n# Models\nWe trained 5 types of model with Arcface.\n\n- Efficient-netV1B5\n- Efficient-netV1B6\n- Efficient-netV1B7\n- Efficient-netV2L\n- ConvNext\n\n# Training\n\n## Augmentation\n- Random Rotate(-10 to 10 degree)\n- Random hue\n- Random contrast\n- Random brightness\n\n## Training settings\n- batch_size = 32\n- image_size = 700\n- Optimizer = Adam\n- loss = SparseCategoricalFocalLoss (&gamma;=1)\n- embedding size = 2048\n- Batch Normalization before Dropout\n\n## How to train models\nWe trained each models by 2 steps in each dataset. We used all train data (no validation data) for training models.\n\n- 1st step  \n  We trained models with all augmentation for 30 epochs.\n- 2nd step  \n  We trained models without Random Rotate augmentation for 25 epochs (In the case of Efficient-netV2L and ConvNext, 20 epochs).\n\nFinally, we got 10 models (5 models for 2 dataset).\n\n# Postprocess\nWe used trained models to get embeddings of training dataset and test dataset.\n\n- TTA(Test Time Augmentation)  \n  We used TTA to get embeddings from each trained models. We input 4 types of image into a trained model, and calculate the average of embeddings.\n  - normal image\n  - flip left right image\n  - -10 degree rotate\n  - 10 degree rotate\n\n- L2 normalization  \n  We calculated L2 normalization.\n\n- Concat embeddings  \n  We concatenated all embeddings in each dataset.\n\n- K-NN  \n  In fullbody dataset and backfin dataset, we calculated distance between training embeddings and test embeddings, and got top 100 (k=100) individual ids for each test ids.\n\n- Ensamble   \n 1.  We calculated top 100 individuals confidence (confidence = 1 - distance) in each test ids\n 2. We added backfin's confidences to fullbody's confidences.\n 3. We inserted new_individual with a fixed threshold (threshold=0.5).\n 4. We sorted confidences and extract top 5 individuals.\n\n# What did not work\n## Models\n- ViT\n- BEiT\n- DeiT\n- Swin Transformer(swin_large_384)\n- Efficient-netV2XL\n\n## Training settings\n- GeMpooling\n- CosFace\n- AdaCos\n- SparseCategoricalCrossentropy\n- Flatten Layer\n- RMSprop\n\n## Augmentaion\n- -90 to 90 degree randomrotate augmentation\n- Cutout augmentation\n- Mix stochastically YOLOv5 and Detic in happywhale-tfrecords-bb\n- Random Crop\n\n## Postprocess \n- ZCA\n- query expansion(&alpha;QE)\n\n# Computational resouces\n- Google Colab Pro +\n\n# Acknowkedgement\nThank you @youtaro @shoheishimatani @room208 @aprilsixteen for this competition!",
    "1768552": "A lot of work and experiments! Thank you for one more great solution desctiption.\n\nCould you show usage of SparseCategoricalFocalLoss? Is this one from focal_loss?\n\nTTA during the embedding retrieval stage cool! 👍 As far as I understand you get 4 embeddings for each image.",
    "1769727": "Thank you for your questions.\n\nYes. SparseCategoricalFocalLoss is one from focal_loss.\nWe used [focal_loss](https://focal-loss.readthedocs.io/en/latest/generated/focal_loss.SparseCategoricalFocalLoss.html) library. \n\nI'm sorry to cause your misunderstanding, because we got one embedding for each image by taking an average of 4 types of embeddings by TTA. We also tried getting 4 embeddings by TTA, but it did not improve our score.",
    "1770035": "Congratulations. Great solution!",
    "1785994": "Very informative solution post. Thank you for sharing!"
  },
  "source": "meta"
}