{
  "id": 319868,
  "title": "8th place solution [My part]",
  "url": "/competitions/happy-whale-and-dolphin/writeups/ntechlab-8th-place-solution-my-part",
  "author_name": "",
  "post_date": "2022-04-20T13:03:48.137Z",
  "votes": 21,
  "comment_count": 13,
  "views": 0,
  "content": "<p><strong>My key features</strong> in importance order:</p>\n<ol>\n<li>A lot of GPUs (or TPUs)</li>\n<li>Custom detector</li>\n<li>Pseudo labeling</li>\n<li>Big input image size</li>\n<li>Invariant loss part from <a href=\"https://arxiv.org/abs/2101.05419\" target=\"_blank\">DAIL -- Dataset-Aware and Invariant Learning for\nFace Recognition</a></li>\n</ol>\n<p><strong>Custom detector</strong><br>\nI've labeled by hand 1k train images, train Yolo, verify by hand 3k images and train final result with 4k labeled images. There are two classes: dorsal fin and full body. Detector was trained very well with perfect quality, I have few hundreds images without boxes, usually it is images under water or tails.</p>\n<p><strong>DAIL Invariant loss part</strong><br>\nThe idea -- we have two datasets: dorsal fins and bodies, it give us more data than train (80k+ vs 50k+ images). Let's train it together with kind of different heads: one head for fins and one for bodies. Also, It increases train time 1.5x.</p>\n<p><strong>Scores</strong><br>\nBest solo model - 799<br>\nBest solo model with pseudo data - 852<br>\nEnsemble score (concat) - 859<br>\nEnsemble with teammates - 872<br>\nDifferent threshold for new individual based on species - best solution (884)</p>\n<p><strong>Pseudo</strong><br>\n I have two iterations, from submit ~840 I took 60% top predictions, got around 830 solo model score. The second iteration after team merge, from submut ~860 I took 70% top predictions (around 15k image).</p>\n<p><strong>Train details</strong><br>\nBest backbones: dm_nfnet_f6, image size 576 and tf_efficientnet_l2_ns, image size 800 (from timm)<br>\nEmbedding size: 4096<br>\nLoss: AMSoftmax aka CosFace (no different in score with ArcFace), m=0.35 and s=25-30<br>\nLR Scheduler: CosineAnnealingLR with SGD<br>\nEpochs: 20-30<br>\nAugmentation: Horizontal flip, blur; increasing amount of augmentation decreased my metrics</p>\n<p>Ensemble - 4 models, two different backbones and two scales (25 and 30)<br>\nI've train on whole data without folds</p>\n<p><strong>Species classification</strong><br>\nOur last big improve -- thresholds based on species, I've trained it at start of competition, <a href=\"https://www.kaggle.com/code/kwentar/species-classification\" target=\"_blank\">train notebook</a>, <a href=\"https://www.kaggle.com/datasets/kwentar/happywhale-test-species\" target=\"_blank\">dataset</a></p>\n<p><strong>What didn't work</strong></p>\n<ol>\n<li>GeM and other poolings, average best for me</li>\n<li>Model on individual species</li>\n</ol>\n<p>Thank you all and especially my teammates <a href=\"https://www.kaggle.com/olegshapovalov\" target=\"_blank\">@olegshapovalov</a>, <a href=\"https://www.kaggle.com/lenny27\" target=\"_blank\">@lenny27</a>, <a href=\"https://www.kaggle.com/ilyadobrynin\" target=\"_blank\">@ilyadobrynin</a> for this competition, that was a great three months!</p>",
  "messages": [
    {
      "id": "1760298",
      "postDate": "04/19/2022 07:39:01",
      "content": "<p><strong>My key features</strong> in importance order:</p>\n<ol>\n<li>A lot of GPUs (or TPUs)</li>\n<li>Custom detector</li>\n<li>Pseudo labeling</li>\n<li>Big input image size</li>\n<li>Invariant loss part from <a href=\"https://arxiv.org/abs/2101.05419\" target=\"_blank\">DAIL -- Dataset-Aware and Invariant Learning for\nFace Recognition</a></li>\n</ol>\n<p><strong>Custom detector</strong><br>\nI've labeled by hand 1k train images, train Yolo, verify by hand 3k images and train final result with 4k labeled images. There are two classes: dorsal fin and full body. Detector was trained very well with perfect quality, I have few hundreds images without boxes, usually it is images under water or tails.</p>\n<p><strong>DAIL Invariant loss part</strong><br>\nThe idea -- we have two datasets: dorsal fins and bodies, it give us more data than train (80k+ vs 50k+ images). Let's train it together with kind of different heads: one head for fins and one for bodies. Also, It increases train time 1.5x.</p>\n<p><strong>Scores</strong><br>\nBest solo model - 799<br>\nBest solo model with pseudo data - 852<br>\nEnsemble score (concat) - 859<br>\nEnsemble with teammates - 872<br>\nDifferent threshold for new individual based on species - best solution (884)</p>\n<p><strong>Pseudo</strong><br>\n I have two iterations, from submit ~840 I took 60% top predictions, got around 830 solo model score. The second iteration after team merge, from submut ~860 I took 70% top predictions (around 15k image).</p>\n<p><strong>Train details</strong><br>\nBest backbones: dm_nfnet_f6, image size 576 and tf_efficientnet_l2_ns, image size 800 (from timm)<br>\nEmbedding size: 4096<br>\nLoss: AMSoftmax aka CosFace (no different in score with ArcFace), m=0.35 and s=25-30<br>\nLR Scheduler: CosineAnnealingLR with SGD<br>\nEpochs: 20-30<br>\nAugmentation: Horizontal flip, blur; increasing amount of augmentation decreased my metrics</p>\n<p>Ensemble - 4 models, two different backbones and two scales (25 and 30)<br>\nI've train on whole data without folds</p>\n<p><strong>Species classification</strong><br>\nOur last big improve -- thresholds based on species, I've trained it at start of competition, <a href=\"https://www.kaggle.com/code/kwentar/species-classification\" target=\"_blank\">train notebook</a>, <a href=\"https://www.kaggle.com/datasets/kwentar/happywhale-test-species\" target=\"_blank\">dataset</a></p>\n<p><strong>What didn't work</strong></p>\n<ol>\n<li>GeM and other poolings, average best for me</li>\n<li>Model on individual species</li>\n</ol>\n<p>Thank you all and especially my teammates <a href=\"https://www.kaggle.com/olegshapovalov\" target=\"_blank\">@olegshapovalov</a>, <a href=\"https://www.kaggle.com/lenny27\" target=\"_blank\">@lenny27</a>, <a href=\"https://www.kaggle.com/ilyadobrynin\" target=\"_blank\">@ilyadobrynin</a> for this competition, that was a great three months!</p>",
      "rawMarkdown": "**My key features** in importance order:\n0. A lot of GPUs (or TPUs)\n1. Custom detector\n2. Pseudo labeling\n3. Big input image size\n4. Invariant loss part from [DAIL -- Dataset-Aware and Invariant Learning for\nFace Recognition](https://arxiv.org/abs/2101.05419)\n\n**Custom detector**\nI've labeled by hand 1k train images, train Yolo, verify by hand 3k images and train final result with 4k labeled images. There are two classes: dorsal fin and full body. Detector was trained very well with perfect quality, I have few hundreds images without boxes, usually it is images under water or tails.\n\n**DAIL Invariant loss part**\nThe idea -- we have two datasets: dorsal fins and bodies, it give us more data than train (80k+ vs 50k+ images). Let's train it together with kind of different heads: one head for fins and one for bodies. Also, It increases train time 1.5x.\n\n**Scores**\nBest solo model - 799\nBest solo model with pseudo data - 852\nEnsemble score (concat) - 859\nEnsemble with teammates - 872\nDifferent threshold for new individual based on species - best solution (884)\n\n**Pseudo**\n I have two iterations, from submit ~840 I took 60% top predictions, got around 830 solo model score. The second iteration after team merge, from submut ~860 I took 70% top predictions (around 15k image).\n\n**Train details**\nBest backbones: dm_nfnet_f6, image size 576 and tf_efficientnet_l2_ns, image size 800 (from timm)\nEmbedding size: 4096\nLoss: AMSoftmax aka CosFace (no different in score with ArcFace), m=0.35 and s=25-30\nLR Scheduler: CosineAnnealingLR with SGD\nEpochs: 20-30\nAugmentation: Horizontal flip, blur; increasing amount of augmentation decreased my metrics\n\nEnsemble - 4 models, two different backbones and two scales (25 and 30)\nI've train on whole data without folds\n\n**Species classification**\nOur last big improve -- thresholds based on species, I've trained it at start of competition, [train notebook](https://www.kaggle.com/code/kwentar/species-classification), [dataset](https://www.kaggle.com/datasets/kwentar/happywhale-test-species)\n\n**What didn't work**\n1. GeM and other poolings, average best for me\n2. Model on individual species\n\n\n\nThank you all and especially my teammates @olegshapovalov, @lenny27, @ilyadobrynin for this competition, that was a great three months!",
      "votes": null
    },
    {
      "id": "1760340",
      "postDate": "04/19/2022 08:14:56",
      "content": "<pre><code>Best solo model - 799\nBest solo model with pseudo data - 852\n</code></pre>\n<p>Seems like a big boost, how many images you pick from test data as pseudo?</p>",
      "rawMarkdown": "```\nBest solo model - 799\nBest solo model with pseudo data - 852\n```\n\nSeems like a big boost, how many images you pick from test data as pseudo?",
      "votes": null
    },
    {
      "id": "1760343",
      "postDate": "04/19/2022 08:20:19",
      "content": "<p>Yeap, pseudo is important here. I have two iterations, from submit like 840 I took 60% top predictions, got around 830. The second iteration after team merge, from submut 860+- I took 70% top predictions (around 15k image).</p>",
      "rawMarkdown": "Yeap, pseudo is important here. I have two iterations, from submit like 840 I took 60% top predictions, got around 830. The second iteration after team merge, from submut 860+- I took 70% top predictions (around 15k image).",
      "votes": null
    },
    {
      "id": "1760687",
      "postDate": "04/19/2022 13:19:53",
      "content": "<p>Have you compared the impact of different embedding sizes on the effect? We use 512.</p>",
      "rawMarkdown": "Have you compared the impact of different embedding sizes on the effect? We use 512.",
      "votes": null
    },
    {
      "id": "1760733",
      "postDate": "04/19/2022 13:45:12",
      "content": "<p>Can you explain more about species classification part, thanks?</p>",
      "rawMarkdown": "Can you explain more about species classification part, thanks?",
      "votes": null
    },
    {
      "id": "1760786",
      "postDate": "04/19/2022 14:12:48",
      "content": "<p>What exactly you interested in? I've created classification model for species prediction and predict test data species, we used it for species threshold. Train notebook you can see here <a href=\"https://www.kaggle.com/code/kwentar/species-classification\" target=\"_blank\">https://www.kaggle.com/code/kwentar/species-classification</a>, it is straightforward image classification, nothing special</p>",
      "rawMarkdown": "What exactly you interested in? I've created classification model for species prediction and predict test data species, we used it for species threshold. Train notebook you can see here https://www.kaggle.com/code/kwentar/species-classification, it is straightforward image classification, nothing special",
      "votes": null
    },
    {
      "id": "1760787",
      "postDate": "04/19/2022 14:13:41",
      "content": "<p>Yes, I've tried from 512 to 8192, 4096 has better result than others, the lower the worse</p>",
      "rawMarkdown": "Yes, I've tried from 512 to 8192, 4096 has better result than others, the lower the worse",
      "votes": null
    },
    {
      "id": "1760788",
      "postDate": "04/19/2022 14:15:17",
      "content": "<p>How much CV difference is there?</p>",
      "rawMarkdown": "How much CV difference is there?",
      "votes": null
    },
    {
      "id": "1760857",
      "postDate": "04/19/2022 14:50:48",
      "content": "<p>I've lost the cv results for this part, but it is quite significant, btw my teammates have different embedding size, I guess here is no one good solution</p>",
      "rawMarkdown": "I've lost the cv results for this part, but it is quite significant, btw my teammates have different embedding size, I guess here is no one good solution",
      "votes": null
    },
    {
      "id": "1760869",
      "postDate": "04/19/2022 14:59:42",
      "content": "<p>Sorry for my vague question. I really want to know how did you use species score for decide new_individual? As I understand, you have 26 classes for species and have 26 threshold for them. Feel free to correct me if I'm wrong.</p>",
      "rawMarkdown": "Sorry for my vague question. I really want to know how did you use species score for decide new_individual? As I understand, you have 26 classes for species and have 26 threshold for them. Feel free to correct me if I'm wrong.",
      "votes": null
    },
    {
      "id": "1760876",
      "postDate": "04/19/2022 15:05:27",
      "content": "<p>Kind of, yes, but no 26, we started from 2 thresholds: one for dolphins and one for whales, it improves score (872-&gt;879), then try to tune threshold for biggest species (beluga, blue whale, etc). It kind of LB overfit, but here it works.</p>",
      "rawMarkdown": "Kind of, yes, but no 26, we started from 2 thresholds: one for dolphins and one for whales, it improves score (872->879), then try to tune threshold for biggest species (beluga, blue whale, etc). It kind of LB overfit, but here it works.",
      "votes": null
    },
    {
      "id": "1760915",
      "postDate": "04/19/2022 15:33:02",
      "content": "<p>So your threshold is based on similarity score, right?. For each image, you gonna predict top 1 species and apply different threshold corresponding top 1 species on cosine similarity score.</p>",
      "rawMarkdown": "So your threshold is based on similarity score, right?. For each image, you gonna predict top 1 species and apply different threshold corresponding top 1 species on cosine similarity score.",
      "votes": null
    },
    {
      "id": "1760927",
      "postDate": "04/19/2022 15:39:53",
      "content": "<p>Oh, my mistake, these thresholds used for decision new individual or not, if similarity &gt; threshold -&gt; predict top 1 id from train else new_individual on top 1 place</p>",
      "rawMarkdown": "Oh, my mistake, these thresholds used for decision new individual or not, if similarity > threshold -> predict top 1 id from train else new_individual on top 1 place",
      "votes": null
    },
    {
      "id": "1761059",
      "postDate": "04/19/2022 16:51:29",
      "content": "<p>can you share your code for this part to clarify my concern, thank? </p>",
      "rawMarkdown": "can you share your code for this part to clarify my concern, thank?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1760340,
      "author_name": "ptran1203",
      "author_url": "",
      "post_date": "04/19/2022 08:14:56",
      "content": "<pre><code>Best solo model - 799\nBest solo model with pseudo data - 852\n</code></pre>\n<p>Seems like a big boost, how many images you pick from test data as pseudo?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760343,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "04/19/2022 08:20:19",
          "content": "<p>Yeap, pseudo is important here. I have two iterations, from submit like 840 I took 60% top predictions, got around 830. The second iteration after team merge, from submut 860+- I took 70% top predictions (around 15k image).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760687,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "04/19/2022 13:19:53",
      "content": "<p>Have you compared the impact of different embedding sizes on the effect? We use 512.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760787,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "04/19/2022 14:13:41",
          "content": "<p>Yes, I've tried from 512 to 8192, 4096 has better result than others, the lower the worse</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760788,
          "author_name": "biglafe",
          "author_url": "",
          "post_date": "04/19/2022 14:15:17",
          "content": "<p>How much CV difference is there?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760857,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "04/19/2022 14:50:48",
          "content": "<p>I've lost the cv results for this part, but it is quite significant, btw my teammates have different embedding size, I guess here is no one good solution</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760733,
      "author_name": "cuongnn218",
      "author_url": "",
      "post_date": "04/19/2022 13:45:12",
      "content": "<p>Can you explain more about species classification part, thanks?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760786,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "04/19/2022 14:12:48",
          "content": "<p>What exactly you interested in? I've created classification model for species prediction and predict test data species, we used it for species threshold. Train notebook you can see here <a href=\"https://www.kaggle.com/code/kwentar/species-classification\" target=\"_blank\">https://www.kaggle.com/code/kwentar/species-classification</a>, it is straightforward image classification, nothing special</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760869,
          "author_name": "cuongnn218",
          "author_url": "",
          "post_date": "04/19/2022 14:59:42",
          "content": "<p>Sorry for my vague question. I really want to know how did you use species score for decide new_individual? As I understand, you have 26 classes for species and have 26 threshold for them. Feel free to correct me if I'm wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760876,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "04/19/2022 15:05:27",
          "content": "<p>Kind of, yes, but no 26, we started from 2 thresholds: one for dolphins and one for whales, it improves score (872-&gt;879), then try to tune threshold for biggest species (beluga, blue whale, etc). It kind of LB overfit, but here it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760915,
          "author_name": "cuongnn218",
          "author_url": "",
          "post_date": "04/19/2022 15:33:02",
          "content": "<p>So your threshold is based on similarity score, right?. For each image, you gonna predict top 1 species and apply different threshold corresponding top 1 species on cosine similarity score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760927,
          "author_name": "kwentar",
          "author_url": "",
          "post_date": "04/19/2022 15:39:53",
          "content": "<p>Oh, my mistake, these thresholds used for decision new individual or not, if similarity &gt; threshold -&gt; predict top 1 id from train else new_individual on top 1 place</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761059,
          "author_name": "cuongnn218",
          "author_url": "",
          "post_date": "04/19/2022 16:51:29",
          "content": "<p>can you share your code for this part to clarify my concern, thank? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1760298": "**My key features** in importance order:\n0. A lot of GPUs (or TPUs)\n1. Custom detector\n2. Pseudo labeling\n3. Big input image size\n4. Invariant loss part from [DAIL -- Dataset-Aware and Invariant Learning for\nFace Recognition](https://arxiv.org/abs/2101.05419)\n\n**Custom detector**\nI've labeled by hand 1k train images, train Yolo, verify by hand 3k images and train final result with 4k labeled images. There are two classes: dorsal fin and full body. Detector was trained very well with perfect quality, I have few hundreds images without boxes, usually it is images under water or tails.\n\n**DAIL Invariant loss part**\nThe idea -- we have two datasets: dorsal fins and bodies, it give us more data than train (80k+ vs 50k+ images). Let's train it together with kind of different heads: one head for fins and one for bodies. Also, It increases train time 1.5x.\n\n**Scores**\nBest solo model - 799\nBest solo model with pseudo data - 852\nEnsemble score (concat) - 859\nEnsemble with teammates - 872\nDifferent threshold for new individual based on species - best solution (884)\n\n**Pseudo**\n I have two iterations, from submit ~840 I took 60% top predictions, got around 830 solo model score. The second iteration after team merge, from submut ~860 I took 70% top predictions (around 15k image).\n\n**Train details**\nBest backbones: dm_nfnet_f6, image size 576 and tf_efficientnet_l2_ns, image size 800 (from timm)\nEmbedding size: 4096\nLoss: AMSoftmax aka CosFace (no different in score with ArcFace), m=0.35 and s=25-30\nLR Scheduler: CosineAnnealingLR with SGD\nEpochs: 20-30\nAugmentation: Horizontal flip, blur; increasing amount of augmentation decreased my metrics\n\nEnsemble - 4 models, two different backbones and two scales (25 and 30)\nI've train on whole data without folds\n\n**Species classification**\nOur last big improve -- thresholds based on species, I've trained it at start of competition, [train notebook](https://www.kaggle.com/code/kwentar/species-classification), [dataset](https://www.kaggle.com/datasets/kwentar/happywhale-test-species)\n\n**What didn't work**\n1. GeM and other poolings, average best for me\n2. Model on individual species\n\n\n\nThank you all and especially my teammates @olegshapovalov, @lenny27, @ilyadobrynin for this competition, that was a great three months!",
    "1760340": "```\nBest solo model - 799\nBest solo model with pseudo data - 852\n```\n\nSeems like a big boost, how many images you pick from test data as pseudo?",
    "1760343": "Yeap, pseudo is important here. I have two iterations, from submit like 840 I took 60% top predictions, got around 830. The second iteration after team merge, from submut 860+- I took 70% top predictions (around 15k image).",
    "1760687": "Have you compared the impact of different embedding sizes on the effect? We use 512.",
    "1760733": "Can you explain more about species classification part, thanks?",
    "1760786": "What exactly you interested in? I've created classification model for species prediction and predict test data species, we used it for species threshold. Train notebook you can see here https://www.kaggle.com/code/kwentar/species-classification, it is straightforward image classification, nothing special",
    "1760787": "Yes, I've tried from 512 to 8192, 4096 has better result than others, the lower the worse",
    "1760788": "How much CV difference is there?",
    "1760857": "I've lost the cv results for this part, but it is quite significant, btw my teammates have different embedding size, I guess here is no one good solution",
    "1760869": "Sorry for my vague question. I really want to know how did you use species score for decide new_individual? As I understand, you have 26 classes for species and have 26 threshold for them. Feel free to correct me if I'm wrong.",
    "1760876": "Kind of, yes, but no 26, we started from 2 thresholds: one for dolphins and one for whales, it improves score (872->879), then try to tune threshold for biggest species (beluga, blue whale, etc). It kind of LB overfit, but here it works.",
    "1760915": "So your threshold is based on similarity score, right?. For each image, you gonna predict top 1 species and apply different threshold corresponding top 1 species on cosine similarity score.",
    "1760927": "Oh, my mistake, these thresholds used for decision new individual or not, if similarity > threshold -> predict top 1 id from train else new_individual on top 1 place",
    "1761059": "can you share your code for this part to clarify my concern, thank?"
  },
  "source": "meta"
}