{
  "id": 320051,
  "title": "Hierarchical learning solution",
  "url": "/competitions/happy-whale-and-dolphin/discussion/320051",
  "author_name": "Vsst",
  "post_date": "2022-04-19T20:25:40.628000",
  "votes": 10,
  "comment_count": 0,
  "views": 0,
  "content": "<p>First of all, I would like to say congratulations to all the winners!🎉🔥 You did a good job with such a complex dataset!</p>\n<p>Although I didn’t do well in this competition (4 places to bronze😩), I would like to share my solution. I think this information may be useful for further competitions.</p>\n<p><strong>Architecture</strong><br>\nAt the beginning of the competition I was interested in finding neural net architectures, which outperform the baseline effnet. I was sure that there are better architectures (Well, in the end I found that the differences are not so drastic). I stopped at the effnetv2_m/effnet_b6/convneXt_m in my final ensemble.<br>\nWhen I saw the data the first idea came to my mind was, that it is a hierarchical learning competition. After some search I found this interesting article: <strong><a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\">Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not</a></strong>. This article is applied to a classification problem, but I suggested it will work well in re-identification either (I modified an article approach a little making the individual embedding part larger than species). The final architecture of my solution you can see in the picture below.<br>\n<a href=\"https://postimg.cc/4Y5FhwRD\" target=\"_blank\"><img src=\"https://i.postimg.cc/bvcXWFCw/arch1.png\" alt=\"arch1.png\"></a></p>\n<p><strong>Augmentations</strong><br>\n     2.1 <strong>Shear scale rotate</strong><br>\nFrom the dataset you can see that there are a lot of different viewpoints for one whale. Using this augmentation I tried to simulate them.<br>\n    2.2 <strong>CLAHE</strong> at test and train time before all augmentation<br>\nThe purpose of using this augmentation is that some back fins are under water, some are darkened. <br>\n   2.3 <strong>Local grayscale transformation</strong><br>\nThis transformation is based on this <a href=\"https://arxiv.org/pdf/2101.08533.pdf\" target=\"_blank\">article</a>, which gave a little boost to LB<br>\n   2.4 <strong>Color/brightness/hue/horizontal flip</strong></p>\n<p><strong>Dataset</strong><br>\nFor all novice kagglers I would like to say that <strong>dataset is crucially important</strong>. I didn’t pay a lot of attention to it, but when I switched from detic to backfins dataset, the score was boosted a lot. </p>\n<p><strong>Loss</strong><br>\nA lot of people used ArcFace over CosFace, but in my experiments CosFace performed better. I used <a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\" target=\"_blank\">subcenter CosFace</a>, because of the noise in the dataset. Also I followed dynamic margins as in the top solutions of Google Landmark competition</p>\n<p><strong>Thing that didn’t work 🙁</strong><br>\nExperiments with triplet loss. I tried several approaches:</p>\n<ol>\n<li>Simple hard triplet loss</li>\n<li>Triplet mining the following way. In the batch there was the same proportion of whales and dolphins. </li>\n<li><a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\"> Centroid triplet loss </a><br>\nAll this experiments performed worse than cosface</li>\n</ol>\n<p>👀<strong>Summary of mentioned articles:</strong><br>\n[1] Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not <a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\">https://arxiv.org/pdf/2104.13643.pdf</a><br>\n[2] Sub-center ArcFace: Boosting Face Recognition by Large-scale Noisy Web Faces <a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf</a><br>\n[3] A general multi-modal data learning method for Person Re-identification<br>\n<a href=\"https://arxiv.org/pdf/2101.08533.pdf\" target=\"_blank\">https://arxiv.org/pdf/2101.08533.pdf</a><br>\n[4] On the Unreasonable Effectiveness of Centroids in Image Retrieval<br>\n<a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\">https://arxiv.org/pdf/2104.13643.pdf</a></p>",
  "messages": [
    {
      "id": 1761359,
      "postDate": "2022-04-19T20:25:40.630Z",
      "content": "<p>First of all, I would like to say congratulations to all the winners!🎉🔥 You did a good job with such a complex dataset!</p>\n<p>Although I didn’t do well in this competition (4 places to bronze😩), I would like to share my solution. I think this information may be useful for further competitions.</p>\n<p><strong>Architecture</strong><br>\nAt the beginning of the competition I was interested in finding neural net architectures, which outperform the baseline effnet. I was sure that there are better architectures (Well, in the end I found that the differences are not so drastic). I stopped at the effnetv2_m/effnet_b6/convneXt_m in my final ensemble.<br>\nWhen I saw the data the first idea came to my mind was, that it is a hierarchical learning competition. After some search I found this interesting article: <strong><a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\">Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not</a></strong>. This article is applied to a classification problem, but I suggested it will work well in re-identification either (I modified an article approach a little making the individual embedding part larger than species). The final architecture of my solution you can see in the picture below.<br>\n<a href=\"https://postimg.cc/4Y5FhwRD\" target=\"_blank\"><img src=\"https://i.postimg.cc/bvcXWFCw/arch1.png\" alt=\"arch1.png\"></a></p>\n<p><strong>Augmentations</strong><br>\n     2.1 <strong>Shear scale rotate</strong><br>\nFrom the dataset you can see that there are a lot of different viewpoints for one whale. Using this augmentation I tried to simulate them.<br>\n    2.2 <strong>CLAHE</strong> at test and train time before all augmentation<br>\nThe purpose of using this augmentation is that some back fins are under water, some are darkened. <br>\n   2.3 <strong>Local grayscale transformation</strong><br>\nThis transformation is based on this <a href=\"https://arxiv.org/pdf/2101.08533.pdf\" target=\"_blank\">article</a>, which gave a little boost to LB<br>\n   2.4 <strong>Color/brightness/hue/horizontal flip</strong></p>\n<p><strong>Dataset</strong><br>\nFor all novice kagglers I would like to say that <strong>dataset is crucially important</strong>. I didn’t pay a lot of attention to it, but when I switched from detic to backfins dataset, the score was boosted a lot. </p>\n<p><strong>Loss</strong><br>\nA lot of people used ArcFace over CosFace, but in my experiments CosFace performed better. I used <a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\" target=\"_blank\">subcenter CosFace</a>, because of the noise in the dataset. Also I followed dynamic margins as in the top solutions of Google Landmark competition</p>\n<p><strong>Thing that didn’t work 🙁</strong><br>\nExperiments with triplet loss. I tried several approaches:</p>\n<ol>\n<li>Simple hard triplet loss</li>\n<li>Triplet mining the following way. In the batch there was the same proportion of whales and dolphins. </li>\n<li><a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\"> Centroid triplet loss </a><br>\nAll this experiments performed worse than cosface</li>\n</ol>\n<p>👀<strong>Summary of mentioned articles:</strong><br>\n[1] Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not <a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\">https://arxiv.org/pdf/2104.13643.pdf</a><br>\n[2] Sub-center ArcFace: Boosting Face Recognition by Large-scale Noisy Web Faces <a href=\"https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\" target=\"_blank\">https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf</a><br>\n[3] A general multi-modal data learning method for Person Re-identification<br>\n<a href=\"https://arxiv.org/pdf/2101.08533.pdf\" target=\"_blank\">https://arxiv.org/pdf/2101.08533.pdf</a><br>\n[4] On the Unreasonable Effectiveness of Centroids in Image Retrieval<br>\n<a href=\"https://arxiv.org/pdf/2104.13643.pdf\" target=\"_blank\">https://arxiv.org/pdf/2104.13643.pdf</a></p>",
      "rawMarkdown": "First of all, I would like to say congratulations to all the winners!🎉🔥 You did a good job with such a complex dataset!\n\nAlthough I didn’t do well in this competition (4 places to bronze😩), I would like to share my solution. I think this information may be useful for further competitions.\n\n**Architecture**\nAt the beginning of the competition I was interested in finding neural net architectures, which outperform the baseline effnet. I was sure that there are better architectures (Well, in the end I found that the differences are not so drastic). I stopped at the effnetv2_m/effnet_b6/convneXt_m in my final ensemble.\nWhen I saw the data the first idea came to my mind was, that it is a hierarchical learning competition. After some search I found this interesting article: **[Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not]( https://arxiv.org/pdf/2104.13643.pdf)**. This article is applied to a classification problem, but I suggested it will work well in re-identification either (I modified an article approach a little making the individual embedding part larger than species). The final architecture of my solution you can see in the picture below.\n[![arch1.png](https://i.postimg.cc/bvcXWFCw/arch1.png)](https://postimg.cc/4Y5FhwRD)\n\n\n**Augmentations**\n     2.1 **Shear scale rotate**\nFrom the dataset you can see that there are a lot of different viewpoints for one whale. Using this augmentation I tried to simulate them.\n    2.2 **CLAHE** at test and train time before all augmentation\nThe purpose of using this augmentation is that some back fins are under water, some are darkened. \n   2.3 **Local grayscale transformation**\nThis transformation is based on this [article](https://arxiv.org/pdf/2101.08533.pdf), which gave a little boost to LB\n   2.4 **Color/brightness/hue/horizontal flip**\n\n**Dataset**\nFor all novice kagglers I would like to say that **dataset is crucially important**. I didn’t pay a lot of attention to it, but when I switched from detic to backfins dataset, the score was boosted a lot. \n\n**Loss**\nA lot of people used ArcFace over CosFace, but in my experiments CosFace performed better. I used [subcenter CosFace](https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf), because of the noise in the dataset. Also I followed dynamic margins as in the top solutions of Google Landmark competition\n\n**Thing that didn’t work 🙁**\nExperiments with triplet loss. I tried several approaches:\n1. Simple hard triplet loss\n2. Triplet mining the following way. In the batch there was the same proportion of whales and dolphins. \n3. [ Centroid triplet loss ](https://arxiv.org/pdf/2104.13643.pdf)\nAll this experiments performed worse than cosface\n\n👀**Summary of mentioned articles:**\n[1] Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not https://arxiv.org/pdf/2104.13643.pdf\n[2] Sub-center ArcFace: Boosting Face Recognition by Large-scale Noisy Web Faces https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\n[3] A general multi-modal data learning method for Person Re-identification\nhttps://arxiv.org/pdf/2101.08533.pdf\n[4] On the Unreasonable Effectiveness of Centroids in Image Retrieval\nhttps://arxiv.org/pdf/2104.13643.pdf\n",
      "votes": 10
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1761359": "First of all, I would like to say congratulations to all the winners!🎉🔥 You did a good job with such a complex dataset!\n\nAlthough I didn’t do well in this competition (4 places to bronze😩), I would like to share my solution. I think this information may be useful for further competitions.\n\n**Architecture**\nAt the beginning of the competition I was interested in finding neural net architectures, which outperform the baseline effnet. I was sure that there are better architectures (Well, in the end I found that the differences are not so drastic). I stopped at the effnetv2_m/effnet_b6/convneXt_m in my final ensemble.\nWhen I saw the data the first idea came to my mind was, that it is a hierarchical learning competition. After some search I found this interesting article: **[Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not]( https://arxiv.org/pdf/2104.13643.pdf)**. This article is applied to a classification problem, but I suggested it will work well in re-identification either (I modified an article approach a little making the individual embedding part larger than species). The final architecture of my solution you can see in the picture below.\n[![arch1.png](https://i.postimg.cc/bvcXWFCw/arch1.png)](https://postimg.cc/4Y5FhwRD)\n\n\n**Augmentations**\n     2.1 **Shear scale rotate**\nFrom the dataset you can see that there are a lot of different viewpoints for one whale. Using this augmentation I tried to simulate them.\n    2.2 **CLAHE** at test and train time before all augmentation\nThe purpose of using this augmentation is that some back fins are under water, some are darkened. \n   2.3 **Local grayscale transformation**\nThis transformation is based on this [article](https://arxiv.org/pdf/2101.08533.pdf), which gave a little boost to LB\n   2.4 **Color/brightness/hue/horizontal flip**\n\n**Dataset**\nFor all novice kagglers I would like to say that **dataset is crucially important**. I didn’t pay a lot of attention to it, but when I switched from detic to backfins dataset, the score was boosted a lot. \n\n**Loss**\nA lot of people used ArcFace over CosFace, but in my experiments CosFace performed better. I used [subcenter CosFace](https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf), because of the noise in the dataset. Also I followed dynamic margins as in the top solutions of Google Landmark competition\n\n**Thing that didn’t work 🙁**\nExperiments with triplet loss. I tried several approaches:\n1. Simple hard triplet loss\n2. Triplet mining the following way. In the batch there was the same proportion of whales and dolphins. \n3. [ Centroid triplet loss ](https://arxiv.org/pdf/2104.13643.pdf)\nAll this experiments performed worse than cosface\n\n👀**Summary of mentioned articles:**\n[1] Your \"Flamingo\" is My \"Bird\": Fine-Grained, or Not https://arxiv.org/pdf/2104.13643.pdf\n[2] Sub-center ArcFace: Boosting Face Recognition by Large-scale Noisy Web Faces https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/123560715.pdf\n[3] A general multi-modal data learning method for Person Re-identification\nhttps://arxiv.org/pdf/2101.08533.pdf\n[4] On the Unreasonable Effectiveness of Centroids in Image Retrieval\nhttps://arxiv.org/pdf/2104.13643.pdf\n"
  }
}