{
  "id": 319941,
  "title": "10th place solution",
  "url": "/competitions/happy-whale-and-dolphin/writeups/yiemon773-10th-place-solution",
  "author_name": "",
  "post_date": "2022-04-20T11:57:41.477Z",
  "votes": 32,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Congratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition. <br>\nThis is My First Solo Gold Medal. <br>\nI'm very impressed to win the gold medal among so many kaggle GMs, masters.  </p>\n<p>Here is my 10th place solution.</p>\n<h1>Dataset</h1>\n<p>I used 2-type datasets and especially thank <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>, because of his sharing valuable approaches and datasets.  </p>\n<ul>\n<li>fullbody dataset  <ul>\n<li>public dataset by <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a></li>\n<li>private dataset  <ul>\n<li>trained yolov5 and WBF with bboxes by Detic  </li></ul></li></ul></li>\n<li>backfin dataset  <ul>\n<li>public dataset by <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>　　</li></ul></li>\n</ul>\n<h1>Model</h1>\n<p>I trained 2-type(fullbody/backfin) models with image sizes 512 or 784.<br>\nfullbody/backfin model: trained with fullbody/backfin dataset</p>\n<p>When we search nearest neighbors,  </p>\n<ul>\n<li><p>for the specices with backfin -&gt; use concatenated embeddings by fullbody models and backfin models   </p></li>\n<li><p>for the specices without backfin (like Beluga, …) -&gt; use embeddings by fullbody models</p>\n<ul>\n<li>backbone  </li>\n<li>efficientnetv2_m</li>\n<li>efficientnetv2_l</li>\n<li>convnext_base</li>\n<li>convnext_large</li></ul></li>\n</ul>\n<p>In my experiments, EfficientnetV2 &gt; EfficientnetV1 ≧ ConvNext, but ensembling them boosted my CV/LB scores. <br>\nAnd I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model. <br>\nThis also works well. </p>\n<h1>Augmentation</h1>\n<p>The main augmentation is below.  </p>\n<ul>\n<li>HorizontalFlip</li>\n<li>ImageCompression</li>\n<li>ShiftScaleRotate</li>\n<li>RandomBrightnessContrast</li>\n<li>HueSaturationValue</li>\n<li>MotionBlur  </li>\n</ul>\n<h3>Mixup</h3>\n<ul>\n<li>mixup the embeddings (not images) and Arcface with soft label worked (CV:+0.003-0.005)</li>\n</ul>\n<h1>Loss</h1>\n<p>In my case, some aux-loss helped.<br>\nLoss = ArcfaceLoss + FocalLoss + SpeciesLoss</p>\n<ul>\n<li>SpeciesLoss: classification of 26-species</li>\n</ul>\n<h3>progressive dynamic margins</h3>\n<p>At some early epochs, training didn't go well if the margins are somewhat large. <br>\nInspired by <a href=\"https://arxiv.org/pdf/2010.05350.pdf\" target=\"_blank\">https://arxiv.org/pdf/2010.05350.pdf</a>, I re-designed the function to increase margins gradually. </p>\n<ul>\n<li>1~5 epoch: increase coefficient of margins linearly from 0.2 to 1</li>\n<li>6~20 epoch: coefficient of margins = 1 (That is, this function is equal to original-dynamic margins)  </li>\n</ul>\n<h1>Pseudo Label</h1>\n<p>This is also the key to raise the scores.  <br>\nAt first, I trained models on pseudo-label, <br>\nand then trained on original training dataset using this pretrained weights.</p>\n<h1>Ensemble</h1>\n<p>I concatenate weighted embeddings, this operation works a little (CV:+0.001-0.002)</p>\n<h1>Thresholding</h1>\n<p>To predict <code>new_individual_id</code>, search the best threshold by the predicted top1-species on validation datasets.  </p>\n<p>Thank you for your reading.<br>\nAny questions are welcome !</p>",
  "messages": [
    {
      "id": "1760738",
      "postDate": "04/19/2022 13:48:02",
      "content": "<p>Congratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition. <br>\nThis is My First Solo Gold Medal. <br>\nI'm very impressed to win the gold medal among so many kaggle GMs, masters.  </p>\n<p>Here is my 10th place solution.</p>\n<h1>Dataset</h1>\n<p>I used 2-type datasets and especially thank <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>, because of his sharing valuable approaches and datasets.  </p>\n<ul>\n<li>fullbody dataset  <ul>\n<li>public dataset by <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a></li>\n<li>private dataset  <ul>\n<li>trained yolov5 and WBF with bboxes by Detic  </li></ul></li></ul></li>\n<li>backfin dataset  <ul>\n<li>public dataset by <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>　　</li></ul></li>\n</ul>\n<h1>Model</h1>\n<p>I trained 2-type(fullbody/backfin) models with image sizes 512 or 784.<br>\nfullbody/backfin model: trained with fullbody/backfin dataset</p>\n<p>When we search nearest neighbors,  </p>\n<ul>\n<li><p>for the specices with backfin -&gt; use concatenated embeddings by fullbody models and backfin models   </p></li>\n<li><p>for the specices without backfin (like Beluga, …) -&gt; use embeddings by fullbody models</p>\n<ul>\n<li>backbone  </li>\n<li>efficientnetv2_m</li>\n<li>efficientnetv2_l</li>\n<li>convnext_base</li>\n<li>convnext_large</li></ul></li>\n</ul>\n<p>In my experiments, EfficientnetV2 &gt; EfficientnetV1 ≧ ConvNext, but ensembling them boosted my CV/LB scores. <br>\nAnd I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model. <br>\nThis also works well. </p>\n<h1>Augmentation</h1>\n<p>The main augmentation is below.  </p>\n<ul>\n<li>HorizontalFlip</li>\n<li>ImageCompression</li>\n<li>ShiftScaleRotate</li>\n<li>RandomBrightnessContrast</li>\n<li>HueSaturationValue</li>\n<li>MotionBlur  </li>\n</ul>\n<h3>Mixup</h3>\n<ul>\n<li>mixup the embeddings (not images) and Arcface with soft label worked (CV:+0.003-0.005)</li>\n</ul>\n<h1>Loss</h1>\n<p>In my case, some aux-loss helped.<br>\nLoss = ArcfaceLoss + FocalLoss + SpeciesLoss</p>\n<ul>\n<li>SpeciesLoss: classification of 26-species</li>\n</ul>\n<h3>progressive dynamic margins</h3>\n<p>At some early epochs, training didn't go well if the margins are somewhat large. <br>\nInspired by <a href=\"https://arxiv.org/pdf/2010.05350.pdf\" target=\"_blank\">https://arxiv.org/pdf/2010.05350.pdf</a>, I re-designed the function to increase margins gradually. </p>\n<ul>\n<li>1~5 epoch: increase coefficient of margins linearly from 0.2 to 1</li>\n<li>6~20 epoch: coefficient of margins = 1 (That is, this function is equal to original-dynamic margins)  </li>\n</ul>\n<h1>Pseudo Label</h1>\n<p>This is also the key to raise the scores.  <br>\nAt first, I trained models on pseudo-label, <br>\nand then trained on original training dataset using this pretrained weights.</p>\n<h1>Ensemble</h1>\n<p>I concatenate weighted embeddings, this operation works a little (CV:+0.001-0.002)</p>\n<h1>Thresholding</h1>\n<p>To predict <code>new_individual_id</code>, search the best threshold by the predicted top1-species on validation datasets.  </p>\n<p>Thank you for your reading.<br>\nAny questions are welcome !</p>",
      "rawMarkdown": "Congratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition. \nThis is My First Solo Gold Medal. \nI'm very impressed to win the gold medal among so many kaggle GMs, masters.  \n\nHere is my 10th place solution.\n\n# Dataset  \nI used 2-type datasets and especially thank @jpbremer, because of his sharing valuable approaches and datasets.  \n - fullbody dataset  \n   - public dataset by @jpbremer\n   - private dataset  \n     - trained yolov5 and WBF with bboxes by Detic  \n - backfin dataset  \n      - public dataset by @jpbremer　　\n\n# Model\nI trained 2-type(fullbody/backfin) models with image sizes 512 or 784.\nfullbody/backfin model: trained with fullbody/backfin dataset\n\nWhen we search nearest neighbors,  \n - for the specices with backfin -> use concatenated embeddings by fullbody models and backfin models   \n - for the specices without backfin (like Beluga, ...) -> use embeddings by fullbody models\n\n- backbone  \n  - efficientnetv2_m\n  - efficientnetv2_l\n  - convnext_base\n  - convnext_large\n\nIn my experiments, EfficientnetV2 > EfficientnetV1 ≧ ConvNext, but ensembling them boosted my CV/LB scores. \nAnd I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model. \nThis also works well. \n\n# Augmentation\nThe main augmentation is below.  \n - HorizontalFlip\n - ImageCompression\n - ShiftScaleRotate\n - RandomBrightnessContrast\n - HueSaturationValue\n - MotionBlur  \n### Mixup  \n - mixup the embeddings (not images) and Arcface with soft label worked (CV:+0.003-0.005)\n\n# Loss \nIn my case, some aux-loss helped.\nLoss = ArcfaceLoss + FocalLoss + SpeciesLoss\n - SpeciesLoss: classification of 26-species\n\n### progressive dynamic margins  \nAt some early epochs, training didn't go well if the margins are somewhat large. \nInspired by https://arxiv.org/pdf/2010.05350.pdf, I re-designed the function to increase margins gradually. \n- 1~5 epoch: increase coefficient of margins linearly from 0.2 to 1\n- 6~20 epoch: coefficient of margins = 1 (That is, this function is equal to original-dynamic margins)  \n\n# Pseudo Label\nThis is also the key to raise the scores.  \nAt first, I trained models on pseudo-label, \nand then trained on original training dataset using this pretrained weights.\n\n# Ensemble\nI concatenate weighted embeddings, this operation works a little (CV:+0.001-0.002)\n\n# Thresholding \nTo predict `new_individual_id`, search the best threshold by the predicted top1-species on validation datasets.  \n\n\nThank you for your reading.\nAny questions are welcome !",
      "votes": null
    },
    {
      "id": "1760763",
      "postDate": "04/19/2022 13:59:22",
      "content": "<p>Congratulations!<br>\nHow many added pseudo-labels, and with what confidence did you use them for training?</p>",
      "rawMarkdown": "Congratulations!\nHow many added pseudo-labels, and with what confidence did you use them for training?",
      "votes": null
    },
    {
      "id": "1760770",
      "postDate": "04/19/2022 14:03:33",
      "content": "<p>Can you introduce the parameters of training effV2? I feel that the training skills of V2 are a little different from that of V1.</p>",
      "rawMarkdown": "Can you introduce the parameters of training effV2? I feel that the training skills of V2 are a little different from that of V1.",
      "votes": null
    },
    {
      "id": "1760781",
      "postDate": "04/19/2022 14:10:59",
      "content": "<p>For me in V2 loss after few epochs transform into nan((</p>",
      "rawMarkdown": "For me in V2 loss after few epochs transform into nan((",
      "votes": null
    },
    {
      "id": "1760784",
      "postDate": "04/19/2022 14:11:52",
      "content": "<p>Thanks for sharing I learned a lot. But can you describe this part in details or share you inference code? <code>To predict new_individual_id, search the best threshold by the predicted top1-species on validation datasets.</code></p>",
      "rawMarkdown": "Thanks for sharing I learned a lot. But can you describe this part in details or share you inference code? `To predict new_individual_id, search the best threshold by the predicted top1-species on validation datasets.`",
      "votes": null
    },
    {
      "id": "1760785",
      "postDate": "04/19/2022 14:12:36",
      "content": "<p>I use all pseudo labeled data (i.e. all the test data) to increase data for pretraining. </p>",
      "rawMarkdown": "I use all pseudo labeled data (i.e. all the test data) to increase data for pretraining.",
      "votes": null
    },
    {
      "id": "1760793",
      "postDate": "04/19/2022 14:18:10",
      "content": "<p>An example of my effv2 training config is below.  <br>\nI don't know if the params are useful for another task.</p>\n<pre><code># margin setting\nmargin_config:\n  params:\n    dynamic_margin_max: 0.8\n    dynamic_margin_min: 0.2\n    margin_power: 0.25\n    prog_epoch_n: 3\n\n# lr scheduler\nlr_scheduler: cosine_schedule_with_warmup \n  lr_scheduler_params: \n    num_training_steps: 20\n    num_warmup_steps: 5\n\n# optimizer\noptimizer_config: AdamW\n  optimizer_params: \n    lr: 7.0e-4\n</code></pre>",
      "rawMarkdown": "An example of my effv2 training config is below.  \nI don't know if the params are useful for another task.\n```\n# margin setting\nmargin_config:\n  params:\n    dynamic_margin_max: 0.8\n    dynamic_margin_min: 0.2\n    margin_power: 0.25\n    prog_epoch_n: 3\n\n# lr scheduler\nlr_scheduler: cosine_schedule_with_warmup \n  lr_scheduler_params: \n    num_training_steps: 20\n    num_warmup_steps: 5\n\n# optimizer\noptimizer_config: AdamW\n  optimizer_params: \n    lr: 7.0e-4\n```",
      "votes": null
    },
    {
      "id": "1760827",
      "postDate": "04/19/2022 14:32:30",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> for this amazing summary.<br>\nCould you please tell more about embeddings mixup technique? I know about simple mixup, when you mix images, and how the features mixup works? Maybe there is come code with loss realisation?<br>\nThis is simply awesome!</p>",
      "rawMarkdown": "Thank you, @yoichi7yamakawa for this amazing summary.\nCould you please tell more about embeddings mixup technique? I know about simple mixup, when you mix images, and how the features mixup works? Maybe there is come code with loss realisation?\nThis is simply awesome!",
      "votes": null
    },
    {
      "id": "1761072",
      "postDate": "04/19/2022 16:58:05",
      "content": "<p>Such a clean and elegant solution. Great work and congrats for solo gold 🤜🤛</p>",
      "rawMarkdown": "Such a clean and elegant solution. Great work and congrats for solo gold 🤜🤛",
      "votes": null
    },
    {
      "id": "1761257",
      "postDate": "04/19/2022 18:31:16",
      "content": "<p>Congrats for solo gold!!</p>",
      "rawMarkdown": "Congrats for solo gold!!",
      "votes": null
    },
    {
      "id": "1761434",
      "postDate": "04/19/2022 23:09:59",
      "content": "<p>Thanks for sharing. So what is the coefficient of the margin? As I understood, let say you used the formula f(n) = a * n ^ -x + b with a, b is lower and upper bound of margin, x is the power, so the coefficient c make the function become f'(x) = cf(n) and c is change by epoch right? Do you still use n as the number of class like in the paper ?  <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> </p>",
      "rawMarkdown": "Thanks for sharing. So what is the coefficient of the margin? As I understood, let say you used the formula f(n) = a * n ^ -x + b with a, b is lower and upper bound of margin, x is the power, so the coefficient c make the function become f'(x) = cf(n) and c is change by epoch right? Do you still use n as the number of class like in the paper ?  @yoichi7yamakawa",
      "votes": null
    },
    {
      "id": "1761481",
      "postDate": "04/20/2022 01:07:49",
      "content": "<p>Here it is. <br>\nYou can mixup embeddings in the same way as the basic image base mixup, and forward mixed-embeddings using the loss below. </p>\n<pre><code>class ArcFaceLossAdaptiveMarginMixup(nn.Module):\n    def __init__(self, margins, s=30.0):\n        super().__init__()\n        self.crit = CrossEntropyLossWithSoftlabel()  \n        self.s = s\n        self.margins = margins\n\n    def forward(self, logits, labels, perm, coeffs): \n        \"\"\"\n        perm: permutated index in batch by using mixup  \n        coeffs: soft-labels by using mixup \n        \"\"\"\n        ms = []\n        ms = self.margins[labels.cpu().numpy()]\n        cos_m = torch.from_numpy(np.cos(ms)).float().type_as(logits)\n        sin_m = torch.from_numpy(np.sin(ms)).float().type_as(logits)\n        th = torch.from_numpy(np.cos(math.pi - ms)).float().type_as(logits)\n        mm = torch.from_numpy(np.sin(math.pi - ms) * ms).float().type_as(logits)\n\n        perm_labels = labels[perm]\n        perm_ms = self.margins[perm_labels.cpu().numpy()]\n        perm_cos_m = torch.from_numpy(np.cos(perm_ms)).float().type_as(logits)\n        perm_sin_m = torch.from_numpy(np.sin(perm_ms)).float().type_as(logits)\n        perm_th = torch.from_numpy(np.cos(math.pi - perm_ms)).float().type_as(logits)\n        perm_mm = torch.from_numpy(np.sin(math.pi - perm_ms) * perm_ms).float().type_as(logits)\n\n        logits = logits.float()\n        cosine = logits\n        sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n\n        # original label\n        labels2 = torch.zeros_like(logits)\n        labels2.scatter_(1, labels.view(-1, 1).long(), 1)\n        phi = cosine * cos_m.view(-1, 1) - sine * sin_m.view(-1, 1)\n        phi = torch.where(cosine &gt; th.view(-1, 1), phi, cosine - mm.view(-1, 1))\n\n        # perm label\n        perm_labels2 = torch.zeros_like(logits)\n        perm_labels2.scatter_(1, perm_labels.view(-1, 1).long(), 1)\n\n        # fix perm labels for not double-count the same labels \n        perm_labels2 = perm_labels2 - torch.logical_and(perm_labels2, labels2).int()\n\n        perm_phi = cosine * perm_cos_m.view(-1, 1) - sine * perm_sin_m.view(-1, 1)\n        perm_phi = torch.where(cosine &gt; perm_th.view(-1, 1), perm_phi, cosine - perm_mm.view(-1, 1))\n\n        # get index with no label\n        with_no_label = 1 - (labels2 + perm_labels2 &gt; 0).type_as(logits)\n\n        output = (labels2 * phi) + (perm_labels2 * perm_phi) + (with_no_label * cosine)\n        output *= self.s\n\n        loss = self.crit(output, labels, perm_labels, coeffs)\n\n        return loss\n</code></pre>",
      "rawMarkdown": "Here it is. \nYou can mixup embeddings in the same way as the basic image base mixup, and forward mixed-embeddings using the loss below. \n\n```\nclass ArcFaceLossAdaptiveMarginMixup(nn.Module):\n    def __init__(self, margins, s=30.0):\n        super().__init__()\n        self.crit = CrossEntropyLossWithSoftlabel()  \n        self.s = s\n        self.margins = margins\n\n    def forward(self, logits, labels, perm, coeffs): \n        \"\"\"\n        perm: permutated index in batch by using mixup  \n        coeffs: soft-labels by using mixup \n        \"\"\"\n        ms = []\n        ms = self.margins[labels.cpu().numpy()]\n        cos_m = torch.from_numpy(np.cos(ms)).float().type_as(logits)\n        sin_m = torch.from_numpy(np.sin(ms)).float().type_as(logits)\n        th = torch.from_numpy(np.cos(math.pi - ms)).float().type_as(logits)\n        mm = torch.from_numpy(np.sin(math.pi - ms) * ms).float().type_as(logits)\n\n        perm_labels = labels[perm]\n        perm_ms = self.margins[perm_labels.cpu().numpy()]\n        perm_cos_m = torch.from_numpy(np.cos(perm_ms)).float().type_as(logits)\n        perm_sin_m = torch.from_numpy(np.sin(perm_ms)).float().type_as(logits)\n        perm_th = torch.from_numpy(np.cos(math.pi - perm_ms)).float().type_as(logits)\n        perm_mm = torch.from_numpy(np.sin(math.pi - perm_ms) * perm_ms).float().type_as(logits)\n\n        logits = logits.float()\n        cosine = logits\n        sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n\n        # original label\n        labels2 = torch.zeros_like(logits)\n        labels2.scatter_(1, labels.view(-1, 1).long(), 1)\n        phi = cosine * cos_m.view(-1, 1) - sine * sin_m.view(-1, 1)\n        phi = torch.where(cosine > th.view(-1, 1), phi, cosine - mm.view(-1, 1))\n\n        # perm label\n        perm_labels2 = torch.zeros_like(logits)\n        perm_labels2.scatter_(1, perm_labels.view(-1, 1).long(), 1)\n\n        # fix perm labels for not double-count the same labels \n        perm_labels2 = perm_labels2 - torch.logical_and(perm_labels2, labels2).int()\n\n        perm_phi = cosine * perm_cos_m.view(-1, 1) - sine * perm_sin_m.view(-1, 1)\n        perm_phi = torch.where(cosine > perm_th.view(-1, 1), perm_phi, cosine - perm_mm.view(-1, 1))\n\n        # get index with no label\n        with_no_label = 1 - (labels2 + perm_labels2 > 0).type_as(logits)\n\n        output = (labels2 * phi) + (perm_labels2 * perm_phi) + (with_no_label * cosine)\n        output *= self.s\n\n        loss = self.crit(output, labels, perm_labels, coeffs)\n\n        return loss\n```",
      "votes": null
    },
    {
      "id": "1761483",
      "postDate": "04/20/2022 01:08:31",
      "content": "<p>Thanks !! <br>\nI'm looking forward to fight together someday again !</p>",
      "rawMarkdown": "Thanks !! \nI'm looking forward to fight together someday again !",
      "votes": null
    },
    {
      "id": "1761484",
      "postDate": "04/20/2022 01:09:16",
      "content": "<p>Thank you ! </p>",
      "rawMarkdown": "Thank you !",
      "votes": null
    },
    {
      "id": "1761490",
      "postDate": "04/20/2022 01:13:02",
      "content": "<p>Yes, I used <code>n</code> (number of classes), and I implemented the margin function like below.</p>\n<pre><code>f(epoch_coeff, n_class) = epoch_coeff * (a * n_class ^ -x + b)    \n\nepoch_coeff: linearly increased in first 5 epochs\n</code></pre>",
      "rawMarkdown": "Yes, I used `n` (number of classes), and I implemented the margin function like below.\n```\nf(epoch_coeff, n_class) = epoch_coeff * (a * n_class ^ -x + b)    \n\nepoch_coeff: linearly increased in first 5 epochs\n```",
      "votes": null
    },
    {
      "id": "1761559",
      "postDate": "04/20/2022 02:16:32",
      "content": "<p><a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> How about this part. can you share the code or pseudo code for this?</p>",
      "rawMarkdown": "yoichi7yamakawa How about this part. can you share the code or pseudo code for this?",
      "votes": null
    },
    {
      "id": "1761614",
      "postDate": "04/20/2022 03:26:38",
      "content": "<p>Congrats to our new GM candidate! 🙌</p>",
      "rawMarkdown": "Congrats to our new GM candidate! 🙌",
      "votes": null
    },
    {
      "id": "1761644",
      "postDate": "04/20/2022 04:01:11",
      "content": "<p>Sorry for late reply. I'm going to prepare pseudo code. </p>",
      "rawMarkdown": "Sorry for late reply. I'm going to prepare pseudo code.",
      "votes": null
    },
    {
      "id": "1761652",
      "postDate": "04/20/2022 04:07:47",
      "content": "<p>Thanks for your reply, I really appreciate your sharing, very informative </p>",
      "rawMarkdown": "Thanks for your reply, I really appreciate your sharing, very informative",
      "votes": null
    },
    {
      "id": "1761798",
      "postDate": "04/20/2022 07:22:27",
      "content": "<p><a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> first of all congratulations - solo gold is … like highway to becoming Competition GM. I keep fingers crossed for your success soon - will follow you. As we can see … even in such competitve competition it is possible. Great!</p>\n<p>I collected question for you to understand it. I really appreciate for answering:</p>\n<ul>\n<li>\"mixup the embeddings \" - how does it exactly mean? <ul>\n<li>\"And I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model.\" - could you show part of code explaining this idea? Why did you decide to implement such architecture?</li>\n<li>What lenght of concatenated embeddings (for both path: a. dorsal fin b. beluga) did you use?</li>\n<li>What NN heads did you implement? As I can see (loss function) there are at least two heads (simmilarity and species classification).</li></ul></li>\n</ul>",
      "rawMarkdown": "yoichi7yamakawa first of all congratulations - solo gold is ... like highway to becoming Competition GM. I keep fingers crossed for your success soon - will follow you. As we can see ... even in such competitve competition it is possible. Great!\n\nI collected question for you to understand it. I really appreciate for answering:\n - \"mixup the embeddings \" - how does it exactly mean? \n- \"And I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model.\" - could you show part of code explaining this idea? Why did you decide to implement such architecture?\n- What lenght of concatenated embeddings (for both path: a. dorsal fin b. beluga) did you use?\n- What NN heads did you implement? As I can see (loss function) there are at least two heads (simmilarity and species classification).",
      "votes": null
    },
    {
      "id": "1762159",
      "postDate": "04/20/2022 13:27:45",
      "content": "<p>Thank you for warmful message, Qishen ! <br>\nI read your past GLR solutions again and again. </p>",
      "rawMarkdown": "Thank you for warmful message, Qishen ! \nI read your past GLR solutions again and again.",
      "votes": null
    },
    {
      "id": "1765022",
      "postDate": "04/23/2022 06:10:23",
      "content": "<p><a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> can you share your code for species loss head? </p>",
      "rawMarkdown": "yoichi7yamakawa can you share your code for species loss head?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1760763,
      "author_name": "maxmar",
      "author_url": "",
      "post_date": "04/19/2022 13:59:22",
      "content": "<p>Congratulations!<br>\nHow many added pseudo-labels, and with what confidence did you use them for training?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760785,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/19/2022 14:12:36",
          "content": "<p>I use all pseudo labeled data (i.e. all the test data) to increase data for pretraining. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760770,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "04/19/2022 14:03:33",
      "content": "<p>Can you introduce the parameters of training effV2? I feel that the training skills of V2 are a little different from that of V1.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760781,
          "author_name": "maxmar",
          "author_url": "",
          "post_date": "04/19/2022 14:10:59",
          "content": "<p>For me in V2 loss after few epochs transform into nan((</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760793,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/19/2022 14:18:10",
          "content": "<p>An example of my effv2 training config is below.  <br>\nI don't know if the params are useful for another task.</p>\n<pre><code># margin setting\nmargin_config:\n  params:\n    dynamic_margin_max: 0.8\n    dynamic_margin_min: 0.2\n    margin_power: 0.25\n    prog_epoch_n: 3\n\n# lr scheduler\nlr_scheduler: cosine_schedule_with_warmup \n  lr_scheduler_params: \n    num_training_steps: 20\n    num_warmup_steps: 5\n\n# optimizer\noptimizer_config: AdamW\n  optimizer_params: \n    lr: 7.0e-4\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761434,
          "author_name": "cuongnn218",
          "author_url": "",
          "post_date": "04/19/2022 23:09:59",
          "content": "<p>Thanks for sharing. So what is the coefficient of the margin? As I understood, let say you used the formula f(n) = a * n ^ -x + b with a, b is lower and upper bound of margin, x is the power, so the coefficient c make the function become f'(x) = cf(n) and c is change by epoch right? Do you still use n as the number of class like in the paper ?  <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761490,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/20/2022 01:13:02",
          "content": "<p>Yes, I used <code>n</code> (number of classes), and I implemented the margin function like below.</p>\n<pre><code>f(epoch_coeff, n_class) = epoch_coeff * (a * n_class ^ -x + b)    \n\nepoch_coeff: linearly increased in first 5 epochs\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760784,
      "author_name": "cuongnn218",
      "author_url": "",
      "post_date": "04/19/2022 14:11:52",
      "content": "<p>Thanks for sharing I learned a lot. But can you describe this part in details or share you inference code? <code>To predict new_individual_id, search the best threshold by the predicted top1-species on validation datasets.</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 1761559,
          "author_name": "cuongnn218",
          "author_url": "",
          "post_date": "04/20/2022 02:16:32",
          "content": "<p><a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> How about this part. can you share the code or pseudo code for this?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761644,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/20/2022 04:01:11",
          "content": "<p>Sorry for late reply. I'm going to prepare pseudo code. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1761652,
          "author_name": "cuongnn218",
          "author_url": "",
          "post_date": "04/20/2022 04:07:47",
          "content": "<p>Thanks for your reply, I really appreciate your sharing, very informative </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760827,
      "author_name": "ilyadobrynin",
      "author_url": "",
      "post_date": "04/19/2022 14:32:30",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> for this amazing summary.<br>\nCould you please tell more about embeddings mixup technique? I know about simple mixup, when you mix images, and how the features mixup works? Maybe there is come code with loss realisation?<br>\nThis is simply awesome!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1761481,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/20/2022 01:07:49",
          "content": "<p>Here it is. <br>\nYou can mixup embeddings in the same way as the basic image base mixup, and forward mixed-embeddings using the loss below. </p>\n<pre><code>class ArcFaceLossAdaptiveMarginMixup(nn.Module):\n    def __init__(self, margins, s=30.0):\n        super().__init__()\n        self.crit = CrossEntropyLossWithSoftlabel()  \n        self.s = s\n        self.margins = margins\n\n    def forward(self, logits, labels, perm, coeffs): \n        \"\"\"\n        perm: permutated index in batch by using mixup  \n        coeffs: soft-labels by using mixup \n        \"\"\"\n        ms = []\n        ms = self.margins[labels.cpu().numpy()]\n        cos_m = torch.from_numpy(np.cos(ms)).float().type_as(logits)\n        sin_m = torch.from_numpy(np.sin(ms)).float().type_as(logits)\n        th = torch.from_numpy(np.cos(math.pi - ms)).float().type_as(logits)\n        mm = torch.from_numpy(np.sin(math.pi - ms) * ms).float().type_as(logits)\n\n        perm_labels = labels[perm]\n        perm_ms = self.margins[perm_labels.cpu().numpy()]\n        perm_cos_m = torch.from_numpy(np.cos(perm_ms)).float().type_as(logits)\n        perm_sin_m = torch.from_numpy(np.sin(perm_ms)).float().type_as(logits)\n        perm_th = torch.from_numpy(np.cos(math.pi - perm_ms)).float().type_as(logits)\n        perm_mm = torch.from_numpy(np.sin(math.pi - perm_ms) * perm_ms).float().type_as(logits)\n\n        logits = logits.float()\n        cosine = logits\n        sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n\n        # original label\n        labels2 = torch.zeros_like(logits)\n        labels2.scatter_(1, labels.view(-1, 1).long(), 1)\n        phi = cosine * cos_m.view(-1, 1) - sine * sin_m.view(-1, 1)\n        phi = torch.where(cosine &gt; th.view(-1, 1), phi, cosine - mm.view(-1, 1))\n\n        # perm label\n        perm_labels2 = torch.zeros_like(logits)\n        perm_labels2.scatter_(1, perm_labels.view(-1, 1).long(), 1)\n\n        # fix perm labels for not double-count the same labels \n        perm_labels2 = perm_labels2 - torch.logical_and(perm_labels2, labels2).int()\n\n        perm_phi = cosine * perm_cos_m.view(-1, 1) - sine * perm_sin_m.view(-1, 1)\n        perm_phi = torch.where(cosine &gt; perm_th.view(-1, 1), perm_phi, cosine - perm_mm.view(-1, 1))\n\n        # get index with no label\n        with_no_label = 1 - (labels2 + perm_labels2 &gt; 0).type_as(logits)\n\n        output = (labels2 * phi) + (perm_labels2 * perm_phi) + (with_no_label * cosine)\n        output *= self.s\n\n        loss = self.crit(output, labels, perm_labels, coeffs)\n\n        return loss\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761072,
      "author_name": "benihime91",
      "author_url": "",
      "post_date": "04/19/2022 16:58:05",
      "content": "<p>Such a clean and elegant solution. Great work and congrats for solo gold 🤜🤛</p>",
      "votes": null,
      "replies": [
        {
          "id": 1761484,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/20/2022 01:09:16",
          "content": "<p>Thank you ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761257,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "04/19/2022 18:31:16",
      "content": "<p>Congrats for solo gold!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1761483,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/20/2022 01:08:31",
          "content": "<p>Thanks !! <br>\nI'm looking forward to fight together someday again !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761614,
      "author_name": "haqishen",
      "author_url": "",
      "post_date": "04/20/2022 03:26:38",
      "content": "<p>Congrats to our new GM candidate! 🙌</p>",
      "votes": null,
      "replies": [
        {
          "id": 1762159,
          "author_name": "yoichi7yamakawa",
          "author_url": "",
          "post_date": "04/20/2022 13:27:45",
          "content": "<p>Thank you for warmful message, Qishen ! <br>\nI read your past GLR solutions again and again. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761798,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/20/2022 07:22:27",
      "content": "<p><a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> first of all congratulations - solo gold is … like highway to becoming Competition GM. I keep fingers crossed for your success soon - will follow you. As we can see … even in such competitve competition it is possible. Great!</p>\n<p>I collected question for you to understand it. I really appreciate for answering:</p>\n<ul>\n<li>\"mixup the embeddings \" - how does it exactly mean? <ul>\n<li>\"And I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model.\" - could you show part of code explaining this idea? Why did you decide to implement such architecture?</li>\n<li>What lenght of concatenated embeddings (for both path: a. dorsal fin b. beluga) did you use?</li>\n<li>What NN heads did you implement? As I can see (loss function) there are at least two heads (simmilarity and species classification).</li></ul></li>\n</ul>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1765022,
      "author_name": "cuongnn218",
      "author_url": "",
      "post_date": "04/23/2022 06:10:23",
      "content": "<p><a href=\"https://www.kaggle.com/yoichi7yamakawa\" target=\"_blank\">@yoichi7yamakawa</a> can you share your code for species loss head? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1760738": "Congratulations to all the winners. And thank you to kaggle host and all participants for this exciting competition. \nThis is My First Solo Gold Medal. \nI'm very impressed to win the gold medal among so many kaggle GMs, masters.  \n\nHere is my 10th place solution.\n\n# Dataset  \nI used 2-type datasets and especially thank @jpbremer, because of his sharing valuable approaches and datasets.  \n - fullbody dataset  \n   - public dataset by @jpbremer\n   - private dataset  \n     - trained yolov5 and WBF with bboxes by Detic  \n - backfin dataset  \n      - public dataset by @jpbremer　　\n\n# Model\nI trained 2-type(fullbody/backfin) models with image sizes 512 or 784.\nfullbody/backfin model: trained with fullbody/backfin dataset\n\nWhen we search nearest neighbors,  \n - for the specices with backfin -> use concatenated embeddings by fullbody models and backfin models   \n - for the specices without backfin (like Beluga, ...) -> use embeddings by fullbody models\n\n- backbone  \n  - efficientnetv2_m\n  - efficientnetv2_l\n  - convnext_base\n  - convnext_large\n\nIn my experiments, EfficientnetV2 > EfficientnetV1 ≧ ConvNext, but ensembling them boosted my CV/LB scores. \nAnd I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model. \nThis also works well. \n\n# Augmentation\nThe main augmentation is below.  \n - HorizontalFlip\n - ImageCompression\n - ShiftScaleRotate\n - RandomBrightnessContrast\n - HueSaturationValue\n - MotionBlur  \n### Mixup  \n - mixup the embeddings (not images) and Arcface with soft label worked (CV:+0.003-0.005)\n\n# Loss \nIn my case, some aux-loss helped.\nLoss = ArcfaceLoss + FocalLoss + SpeciesLoss\n - SpeciesLoss: classification of 26-species\n\n### progressive dynamic margins  \nAt some early epochs, training didn't go well if the margins are somewhat large. \nInspired by https://arxiv.org/pdf/2010.05350.pdf, I re-designed the function to increase margins gradually. \n- 1~5 epoch: increase coefficient of margins linearly from 0.2 to 1\n- 6~20 epoch: coefficient of margins = 1 (That is, this function is equal to original-dynamic margins)  \n\n# Pseudo Label\nThis is also the key to raise the scores.  \nAt first, I trained models on pseudo-label, \nand then trained on original training dataset using this pretrained weights.\n\n# Ensemble\nI concatenate weighted embeddings, this operation works a little (CV:+0.001-0.002)\n\n# Thresholding \nTo predict `new_individual_id`, search the best threshold by the predicted top1-species on validation datasets.  \n\n\nThank you for your reading.\nAny questions are welcome !",
    "1760763": "Congratulations!\nHow many added pseudo-labels, and with what confidence did you use them for training?",
    "1760770": "Can you introduce the parameters of training effV2? I feel that the training skills of V2 are a little different from that of V1.",
    "1760781": "For me in V2 loss after few epochs transform into nan((",
    "1760784": "Thanks for sharing I learned a lot. But can you describe this part in details or share you inference code? `To predict new_individual_id, search the best threshold by the predicted top1-species on validation datasets.`",
    "1760785": "I use all pseudo labeled data (i.e. all the test data) to increase data for pretraining.",
    "1760793": "An example of my effv2 training config is below.  \nI don't know if the params are useful for another task.\n```\n# margin setting\nmargin_config:\n  params:\n    dynamic_margin_max: 0.8\n    dynamic_margin_min: 0.2\n    margin_power: 0.25\n    prog_epoch_n: 3\n\n# lr scheduler\nlr_scheduler: cosine_schedule_with_warmup \n  lr_scheduler_params: \n    num_training_steps: 20\n    num_warmup_steps: 5\n\n# optimizer\noptimizer_config: AdamW\n  optimizer_params: \n    lr: 7.0e-4\n```",
    "1760827": "Thank you, @yoichi7yamakawa for this amazing summary.\nCould you please tell more about embeddings mixup technique? I know about simple mixup, when you mix images, and how the features mixup works? Maybe there is come code with loss realisation?\nThis is simply awesome!",
    "1761072": "Such a clean and elegant solution. Great work and congrats for solo gold 🤜🤛",
    "1761257": "Congrats for solo gold!!",
    "1761434": "Thanks for sharing. So what is the coefficient of the margin? As I understood, let say you used the formula f(n) = a * n ^ -x + b with a, b is lower and upper bound of margin, x is the power, so the coefficient c make the function become f'(x) = cf(n) and c is change by epoch right? Do you still use n as the number of class like in the paper ?  @yoichi7yamakawa",
    "1761481": "Here it is. \nYou can mixup embeddings in the same way as the basic image base mixup, and forward mixed-embeddings using the loss below. \n\n```\nclass ArcFaceLossAdaptiveMarginMixup(nn.Module):\n    def __init__(self, margins, s=30.0):\n        super().__init__()\n        self.crit = CrossEntropyLossWithSoftlabel()  \n        self.s = s\n        self.margins = margins\n\n    def forward(self, logits, labels, perm, coeffs): \n        \"\"\"\n        perm: permutated index in batch by using mixup  \n        coeffs: soft-labels by using mixup \n        \"\"\"\n        ms = []\n        ms = self.margins[labels.cpu().numpy()]\n        cos_m = torch.from_numpy(np.cos(ms)).float().type_as(logits)\n        sin_m = torch.from_numpy(np.sin(ms)).float().type_as(logits)\n        th = torch.from_numpy(np.cos(math.pi - ms)).float().type_as(logits)\n        mm = torch.from_numpy(np.sin(math.pi - ms) * ms).float().type_as(logits)\n\n        perm_labels = labels[perm]\n        perm_ms = self.margins[perm_labels.cpu().numpy()]\n        perm_cos_m = torch.from_numpy(np.cos(perm_ms)).float().type_as(logits)\n        perm_sin_m = torch.from_numpy(np.sin(perm_ms)).float().type_as(logits)\n        perm_th = torch.from_numpy(np.cos(math.pi - perm_ms)).float().type_as(logits)\n        perm_mm = torch.from_numpy(np.sin(math.pi - perm_ms) * perm_ms).float().type_as(logits)\n\n        logits = logits.float()\n        cosine = logits\n        sine = torch.sqrt(1.0 - torch.pow(cosine, 2))\n\n        # original label\n        labels2 = torch.zeros_like(logits)\n        labels2.scatter_(1, labels.view(-1, 1).long(), 1)\n        phi = cosine * cos_m.view(-1, 1) - sine * sin_m.view(-1, 1)\n        phi = torch.where(cosine > th.view(-1, 1), phi, cosine - mm.view(-1, 1))\n\n        # perm label\n        perm_labels2 = torch.zeros_like(logits)\n        perm_labels2.scatter_(1, perm_labels.view(-1, 1).long(), 1)\n\n        # fix perm labels for not double-count the same labels \n        perm_labels2 = perm_labels2 - torch.logical_and(perm_labels2, labels2).int()\n\n        perm_phi = cosine * perm_cos_m.view(-1, 1) - sine * perm_sin_m.view(-1, 1)\n        perm_phi = torch.where(cosine > perm_th.view(-1, 1), perm_phi, cosine - perm_mm.view(-1, 1))\n\n        # get index with no label\n        with_no_label = 1 - (labels2 + perm_labels2 > 0).type_as(logits)\n\n        output = (labels2 * phi) + (perm_labels2 * perm_phi) + (with_no_label * cosine)\n        output *= self.s\n\n        loss = self.crit(output, labels, perm_labels, coeffs)\n\n        return loss\n```",
    "1761483": "Thanks !! \nI'm looking forward to fight together someday again !",
    "1761484": "Thank you !",
    "1761490": "Yes, I used `n` (number of classes), and I implemented the margin function like below.\n```\nf(epoch_coeff, n_class) = epoch_coeff * (a * n_class ^ -x + b)    \n\nepoch_coeff: linearly increased in first 5 epochs\n```",
    "1761559": "yoichi7yamakawa How about this part. can you share the code or pseudo code for this?",
    "1761614": "Congrats to our new GM candidate! 🙌",
    "1761644": "Sorry for late reply. I'm going to prepare pseudo code.",
    "1761652": "Thanks for your reply, I really appreciate your sharing, very informative",
    "1761798": "yoichi7yamakawa first of all congratulations - solo gold is ... like highway to becoming Competition GM. I keep fingers crossed for your success soon - will follow you. As we can see ... even in such competitve competition it is possible. Great!\n\nI collected question for you to understand it. I really appreciate for answering:\n - \"mixup the embeddings \" - how does it exactly mean? \n- \"And I concatenated outputs of conv-layers before pooling layer, and then forward this to the neck of the model.\" - could you show part of code explaining this idea? Why did you decide to implement such architecture?\n- What lenght of concatenated embeddings (for both path: a. dorsal fin b. beluga) did you use?\n- What NN heads did you implement? As I can see (loss function) there are at least two heads (simmilarity and species classification).",
    "1762159": "Thank you for warmful message, Qishen ! \nI read your past GLR solutions again and again.",
    "1765022": "yoichi7yamakawa can you share your code for species loss head?"
  },
  "source": "meta"
}