{
  "id": 82409,
  "title": "25-th place solution: CosFace + ProtoNets",
  "url": "/competitions/humpback-whale-identification/discussion/82409",
  "author_name": "Bartek",
  "post_date": "2019-03-01T08:02:31.189000",
  "votes": 32,
  "comment_count": 37,
  "views": 0,
  "content": "<p>Here I would describe my part of the solution: CosFace approch.\nThe ProtoNets and Ensebling would be described by <a href=\"/daisukelab\">@daisukelab</a></p>\n\n<p><strong>1. Preprocessing</strong>\nFist of all, I use BB from @radek (thank you!). I just did train model on his annotated data and that's all, I did not invest time for this part of competiton. In later stage I also used updated BB from @radek (I call them v2), but the difference in final results was very small.</p>\n\n<p><strong>2. Model and Data-Loading</strong>\nMy final model was se-resnext101 (I also try se154 but it did not work nice). In fact, my model and augmentation was exactly the same like in this kernel: <a href=\"https://www.kaggle.com/stalkermustang/pytorch-pretraiedmodels-se-resnext101-baseline\">https://www.kaggle.com/stalkermustang/pytorch-pretraiedmodels-se-resnext101-baseline</a>\n Other stuff which I tried:\n -  CutOut -&gt; fail\n -  MixUp -&gt; fail\n - OverSample -&gt; fail\n - Cluster-Based Sampling: here I try to sample batch so that there were similar classes (based on cosine similarity) -&gt; fail, overfitting</p>\n\n<p><strong>3. Loss Function</strong>\nIn the past, I was developing the Face-Recognition system and the main purpose of approaching the whale problem was testing such technology on whales:) So in general I was testing ArcFace and CosFace. Both of them works pretty good but CosFace was slightly better. For CosFace I use very high margin 0.6 (in original paper it was 0.35). \nIn fact, this step take me the longest time (3 weeks), where I was trying:</p>\n\n<ul>\n<li>BatchNorm vs LayerNorm before L2 normalization: LayerNorm better</li>\n<li>AlphaDropout vs DropOut: Alpha better but with no influence in debugging model (resnet50) so I did not use both of them, what now I think was one of the biggest mistake I made</li>\n<li>CosFace vs ArcFace vs SphereFace: CosFace with m=0.6 was clear winner (I also the the idea of the CosFace the most)</li>\n</ul>\n\n<p><strong>4. Optimalization</strong>\nHere I use AdamW (with fixed weight-decay). I also try OneCycle but it was not working (I think that   code was wrong). In general I train the model by 30 epochs, so it was pretty quick.</p>\n\n<ul>\n<li>New-whale\nI did not use 'new-whale' for training. In sumbission I just want to have ~27% of 'new_whale'\nI have two approches for this problem which did not work:</li>\n<li>Each new-whale as different class: In general it was ok, the accuracy on validation set was just 0.2% less. But it does not work well on LB.</li>\n<li>Use second loss-function (KL-Divergence) which would force the outout distribution afer SoftMax of 'new-whale' to be uniform (so exactly the same probability for each class). It also work fine for validation set, but not in LB. </li>\n</ul>\n\n<p>Look like I would need more time for this approach (especially second one), because I really like it:)</p>\n\n<p>The final models was se101resnext trained on:\n- gray and 448x448\n- gray and 256x748\n- rgb and 448x448\n- rgb and 256x748</p>\n\n<p>In general, the approch was pretty simple and work moderate. But look like @pudae had the similar idea, so I'm now looking into his approach :)</p>\n\n<p>The ProtoNet and Ensemble part would be explained by <a href=\"/daisukelab\">@daisukelab</a>.</p>\n\n<p>Code: <a href=\"https://github.com/melgor/kaggle-whale-tail\">https://github.com/melgor/kaggle-whale-tail</a>\nThis is minimal code for training single model and create sumbission without 'new-whale'</p>",
  "messages": [
    {
      "id": 481275,
      "postDate": "2019-03-01T08:02:31.190Z",
      "content": "<p>Here I would describe my part of the solution: CosFace approch.\nThe ProtoNets and Ensebling would be described by <a href=\"/daisukelab\">@daisukelab</a></p>\n\n<p><strong>1. Preprocessing</strong>\nFist of all, I use BB from @radek (thank you!). I just did train model on his annotated data and that's all, I did not invest time for this part of competiton. In later stage I also used updated BB from @radek (I call them v2), but the difference in final results was very small.</p>\n\n<p><strong>2. Model and Data-Loading</strong>\nMy final model was se-resnext101 (I also try se154 but it did not work nice). In fact, my model and augmentation was exactly the same like in this kernel: <a href=\"https://www.kaggle.com/stalkermustang/pytorch-pretraiedmodels-se-resnext101-baseline\">https://www.kaggle.com/stalkermustang/pytorch-pretraiedmodels-se-resnext101-baseline</a>\n Other stuff which I tried:\n -  CutOut -&gt; fail\n -  MixUp -&gt; fail\n - OverSample -&gt; fail\n - Cluster-Based Sampling: here I try to sample batch so that there were similar classes (based on cosine similarity) -&gt; fail, overfitting</p>\n\n<p><strong>3. Loss Function</strong>\nIn the past, I was developing the Face-Recognition system and the main purpose of approaching the whale problem was testing such technology on whales:) So in general I was testing ArcFace and CosFace. Both of them works pretty good but CosFace was slightly better. For CosFace I use very high margin 0.6 (in original paper it was 0.35). \nIn fact, this step take me the longest time (3 weeks), where I was trying:</p>\n\n<ul>\n<li>BatchNorm vs LayerNorm before L2 normalization: LayerNorm better</li>\n<li>AlphaDropout vs DropOut: Alpha better but with no influence in debugging model (resnet50) so I did not use both of them, what now I think was one of the biggest mistake I made</li>\n<li>CosFace vs ArcFace vs SphereFace: CosFace with m=0.6 was clear winner (I also the the idea of the CosFace the most)</li>\n</ul>\n\n<p><strong>4. Optimalization</strong>\nHere I use AdamW (with fixed weight-decay). I also try OneCycle but it was not working (I think that   code was wrong). In general I train the model by 30 epochs, so it was pretty quick.</p>\n\n<ul>\n<li>New-whale\nI did not use 'new-whale' for training. In sumbission I just want to have ~27% of 'new_whale'\nI have two approches for this problem which did not work:</li>\n<li>Each new-whale as different class: In general it was ok, the accuracy on validation set was just 0.2% less. But it does not work well on LB.</li>\n<li>Use second loss-function (KL-Divergence) which would force the outout distribution afer SoftMax of 'new-whale' to be uniform (so exactly the same probability for each class). It also work fine for validation set, but not in LB. </li>\n</ul>\n\n<p>Look like I would need more time for this approach (especially second one), because I really like it:)</p>\n\n<p>The final models was se101resnext trained on:\n- gray and 448x448\n- gray and 256x748\n- rgb and 448x448\n- rgb and 256x748</p>\n\n<p>In general, the approch was pretty simple and work moderate. But look like @pudae had the similar idea, so I'm now looking into his approach :)</p>\n\n<p>The ProtoNet and Ensemble part would be explained by <a href=\"/daisukelab\">@daisukelab</a>.</p>\n\n<p>Code: <a href=\"https://github.com/melgor/kaggle-whale-tail\">https://github.com/melgor/kaggle-whale-tail</a>\nThis is minimal code for training single model and create sumbission without 'new-whale'</p>",
      "rawMarkdown": "Here I would describe my part of the solution: CosFace approch.\nThe ProtoNets and Ensebling would be described by @daisukelab\n\n**1. Preprocessing**\nFist of all, I use BB from @radek (thank you!). I just did train model on his annotated data and that's all, I did not invest time for this part of competiton. In later stage I also used updated BB from @radek (I call them v2), but the difference in final results was very small.\n\n**2. Model and Data-Loading**\nMy final model was se-resnext101 (I also try se154 but it did not work nice). In fact, my model and augmentation was exactly the same like in this kernel: https://www.kaggle.com/stalkermustang/pytorch-pretraiedmodels-se-resnext101-baseline\n Other stuff which I tried:\n -  CutOut -&gt; fail\n -  MixUp -&gt; fail\n - OverSample -&gt; fail\n - Cluster-Based Sampling: here I try to sample batch so that there were similar classes (based on cosine similarity) -&gt; fail, overfitting\n\n**3. Loss Function**\nIn the past, I was developing the Face-Recognition system and the main purpose of approaching the whale problem was testing such technology on whales:) So in general I was testing ArcFace and CosFace. Both of them works pretty good but CosFace was slightly better. For CosFace I use very high margin 0.6 (in original paper it was 0.35). \nIn fact, this step take me the longest time (3 weeks), where I was trying:\n\n - BatchNorm vs LayerNorm before L2 normalization: LayerNorm better\n - AlphaDropout vs DropOut: Alpha better but with no influence in debugging model (resnet50) so I did not use both of them, what now I think was one of the biggest mistake I made\n - CosFace vs ArcFace vs SphereFace: CosFace with m=0.6 was clear winner (I also the the idea of the CosFace the most)\n\n**4. Optimalization**\nHere I use AdamW (with fixed weight-decay). I also try OneCycle but it was not working (I think that   code was wrong). In general I train the model by 30 epochs, so it was pretty quick.\n\n - New-whale\nI did not use 'new-whale' for training. In sumbission I just want to have ~27% of 'new_whale'\nI have two approches for this problem which did not work:\n - Each new-whale as different class: In general it was ok, the accuracy on validation set was just 0.2% less. But it does not work well on LB.\n - Use second loss-function (KL-Divergence) which would force the outout distribution afer SoftMax of 'new-whale' to be uniform (so exactly the same probability for each class). It also work fine for validation set, but not in LB. \n\nLook like I would need more time for this approach (especially second one), because I really like it:)\n\nThe final models was se101resnext trained on:\n- gray and 448x448\n- gray and 256x748\n- rgb and 448x448\n- rgb and 256x748\n\nIn general, the approch was pretty simple and work moderate. But look like @pudae had the similar idea, so I'm now looking into his approach :)\n\nThe ProtoNet and Ensemble part would be explained by @daisukelab.\n\nCode: https://github.com/melgor/kaggle-whale-tail\nThis is minimal code for training single model and create sumbission without 'new-whale'",
      "votes": 32
    },
    {
      "id": 481640,
      "postDate": "2019-03-01T16:44:15.433Z",
      "content": "<p>Here's about team ensemble.</p>\n\n<p>Basically we found that ensemble of different approach works really well, combination of CosFace approach and ProtoNets.</p>\n\n<ul>\n<li>First team ensemble, 2 CosFace models + 8 ProtoNets models was LB private:public = 0.94168:0.93314</li>\n<li>Second team ensemble, 3 CosFace models + 3 ProtoNets models was LB private:public = 0.94627:0.94045</li>\n<li>Final team ensemble, 11 CosFace models + 3 ProtoNets models was LB private:public = 0.94909:0.94810</li>\n</ul>\n\n<p>All results are once processed by softmax and geometric mean was calculated.\nThen 'new_whale' threshold was set to 30% for almost all results.</p>\n\n<p>Lastly, this two weeks was special to me.\nI have been working with ProtoNets from the beginning of this year,\njoined this competition just basically for checking performance of that.\nBut once I shared my ProtoNets repo in the <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/81085\">thread</a>,\nmany people seems to have enjoyed and it became very nice discussion. It was wonderful, many thanks to Heng as <a href=\"/hengck23\">@hengck23</a>.</p>\n\n<p>And then @Bartek interested in my solution and kindly invited me to make a team. And now it resulted in effective ensemble.\nAs I was spending most of time for ProtoNets only, I couldn't even try CNN+effective loss approach in time by myself, so it was very good that we could try something different.\nThank you teammate @Bartek!</p>\n\n<p>So... I spend almost all time, less sleep, think and try and tried again, really (healthily) tired.</p>\n\n<p>Thank you all, congratulations to winners! Now I can go to bed, zzzz...</p>",
      "rawMarkdown": "Here's about team ensemble.\n\nBasically we found that ensemble of different approach works really well, combination of CosFace approach and ProtoNets.\n\n- First team ensemble, 2 CosFace models + 8 ProtoNets models was LB private:public = 0.94168:0.93314\n- Second team ensemble, 3 CosFace models + 3 ProtoNets models was LB private:public = 0.94627:0.94045\n- Final team ensemble, 11 CosFace models + 3 ProtoNets models was LB private:public = 0.94909:0.94810\n\nAll results are once processed by softmax and geometric mean was calculated.\nThen 'new_whale' threshold was set to 30% for almost all results.\n\nLastly, this two weeks was special to me.\nI have been working with ProtoNets from the beginning of this year,\njoined this competition just basically for checking performance of that.\nBut once I shared my ProtoNets repo in the [thread](https://www.kaggle.com/c/humpback-whale-identification/discussion/81085),\nmany people seems to have enjoyed and it became very nice discussion. It was wonderful, many thanks to Heng as @hengck23.\n\nAnd then @Bartek interested in my solution and kindly invited me to make a team. And now it resulted in effective ensemble.\nAs I was spending most of time for ProtoNets only, I couldn't even try CNN+effective loss approach in time by myself, so it was very good that we could try something different.\nThank you teammate @Bartek!\n\nSo... I spend almost all time, less sleep, think and try and tried again, really (healthily) tired.\n\nThank you all, congratulations to winners! Now I can go to bed, zzzz...",
      "votes": 4
    },
    {
      "id": 485617,
      "postDate": "2019-03-07T16:51:23.813Z",
      "content": "<p><a href=\"/melgor\">@melgor</a> so far I still haven't been able to reach the 0.68LB without TTA and new_whale. I can only reach 0.667LB but I won't give up haha.</p>\n\n<p>Are you using data from the playground competition?</p>",
      "rawMarkdown": "@melgor so far I still haven't been able to reach the 0.68LB without TTA and new_whale. I can only reach 0.667LB but I won't give up haha.\n\nAre you using data from the playground competition?",
      "votes": 1,
      "replies": [
        {
          "id": 485661,
          "postDate": "2019-03-07T18:38:33.940Z",
          "content": "<p>In fact, I'm continuting experiments with Whales, that is why I did not release the code. I will try to find tomorrow some time to release the code.</p>",
          "rawMarkdown": "In fact, I'm continuting experiments with Whales, that is why I did not release the code. I will try to find tomorrow some time to release the code.",
          "votes": 1
        },
        {
          "id": 485745,
          "postDate": "2019-03-07T21:13:51.260Z",
          "content": "<p>Implementing top solutions is very important in the learning process. But don't worry, take your time! :)</p>\n\n<p>I've noticed that the top1 solution trained on both this competition and playground's data. I thought that the playground's images were a subset of this compeition's training set, I haven't used any playground data so far, I'll take a look into that</p>",
          "rawMarkdown": "Implementing top solutions is very important in the learning process. But don't worry, take your time! :)\n\nI've noticed that the top1 solution trained on both this competition and playground's data. I thought that the playground's images were a subset of this compeition's training set, I haven't used any playground data so far, I'll take a look into that"
        },
        {
          "id": 487420,
          "postDate": "2019-03-10T20:53:30.023Z",
          "content": "<p>Hi, I added code: <a href=\"https://github.com/melgor/kaggle-whale-tail\">https://github.com/melgor/kaggle-whale-tail</a>\nBut this is minimal example, without creating submission with new-whale, learn single model. My code is rather mess as I spend not so much time for this copetition (mostly weekends).</p>\n\n<p>In general, I also did not used playgroud-data. Also the 3rd solution learn for a very long time (500 epochs vs 40 mine). I did some experiments and look like training longer res50 can match se101 using same input-size (so it is interesting). Then I decided to try again with removing avg-pool but training longer.</p>",
          "rawMarkdown": "Hi, I added code: https://github.com/melgor/kaggle-whale-tail\nBut this is minimal example, without creating submission with new-whale, learn single model. My code is rather mess as I spend not so much time for this copetition (mostly weekends).\n\nIn general, I also did not used playgroud-data. Also the 3rd solution learn for a very long time (500 epochs vs 40 mine). I did some experiments and look like training longer res50 can match se101 using same input-size (so it is interesting). Then I decided to try again with removing avg-pool but training longer.",
          "votes": 1
        },
        {
          "id": 487466,
          "postDate": "2019-03-10T23:52:52.837Z",
          "content": "<p>Thanks Bartek, I'll take a look later.</p>\n\n<p>So far I've managed to get 0.67502/0.67966 on public/private respectively (no new_whale and no tta, approx. 0.94 with it) using CosFace with your margin and a DenseNet121 (320x320) with a 512FC layer. It took me 31 epochs using AdamW with 5e-4 weight decay.  I also used pudae's augmentation (added flip since I'm not doubling up the identities yet) and radek's bbox as augmentation (0.5 prob of using it or not during training). For inference I computed the center of every class.</p>\n\n<p>Now, about using flattening and training longer I've tried both and I got worse results. What helped was changing LayerNorm for BatchNorm without shift (bias=0 and non-trainable).</p>\n\n<p>May I ask how much are you getting with a single model so far? I think part of pudae's good results are due to averaging predictions on 5 different bbox during inference. I got considerably better results just by changing martins bbox annotations for radek's, so I guess there are still room for improvements following this path.</p>",
          "rawMarkdown": "Thanks Bartek, I'll take a look later.\n\nSo far I've managed to get 0.67502/0.67966 on public/private respectively (no new_whale and no tta, approx. 0.94 with it) using CosFace with your margin and a DenseNet121 (320x320) with a 512FC layer. It took me 31 epochs using AdamW with 5e-4 weight decay.  I also used pudae's augmentation (added flip since I'm not doubling up the identities yet) and radek's bbox as augmentation (0.5 prob of using it or not during training). For inference I computed the center of every class.\n\nNow, about using flattening and training longer I've tried both and I got worse results. What helped was changing LayerNorm for BatchNorm without shift (bias=0 and non-trainable).\n\nMay I ask how much are you getting with a single model so far? I think part of pudae's good results are due to averaging predictions on 5 different bbox during inference. I got considerably better results just by changing martins bbox annotations for radek's, so I guess there are still room for improvements following this path.",
          "votes": 1
        },
        {
          "id": 487581,
          "postDate": "2019-03-11T06:39:30.377Z",
          "content": "<p>Currently I'm maily focus on traning longer using Res50,  my best model on 224x224, my standard augmentation, pure classification (no center for every class) I get 0.656. \nThe results is not best-one, but based on my all history, I'm able to get 0.02 better score using same model just for longer training (I use 100 epochs) but not training on all data (1k images for validation).  And Res50 because it learn fast:)</p>\n\n<p>About LayerNorm vs BatchNorm, I will test it again to confirm if there is any difference.\nThen I will also try pudae's augmentation.\nAnd finaly center of every class.</p>\n\n<p>About BB, I I'm not sure if training beter detector would help a lot. I trained ones (using radek code) and compare to the second release of BB. In both the results was pretty much same.</p>\n\n<p>Also, best model within competition could get 0.682 (se101, 448x448)-&gt; 0.93. But I did not use center features, so maybe you would get 0.94:)</p>\n\n<p>If you would get score &gt; 0.95 with single model, let me know:)</p>",
          "rawMarkdown": "Currently I'm maily focus on traning longer using Res50,  my best model on 224x224, my standard augmentation, pure classification (no center for every class) I get 0.656. \nThe results is not best-one, but based on my all history, I'm able to get 0.02 better score using same model just for longer training (I use 100 epochs) but not training on all data (1k images for validation).  And Res50 because it learn fast:)\n\nAbout LayerNorm vs BatchNorm, I will test it again to confirm if there is any difference.\nThen I will also try pudae's augmentation.\nAnd finaly center of every class.\n\nAbout BB, I I'm not sure if training beter detector would help a lot. I trained ones (using radek code) and compare to the second release of BB. In both the results was pretty much same.\n\nAlso, best model within competition could get 0.682 (se101, 448x448)-&gt; 0.93. But I did not use center features, so maybe you would get 0.94:)\n\nIf you would get score &gt; 0.95 with single model, let me know:)\n",
          "votes": 1
        },
        {
          "id": 487731,
          "postDate": "2019-03-11T11:46:23.403Z",
          "content": "<p>Sure! For my 0.94 model I still need to try higher resolution and retrain on all data, I'm also using around 1k validation images.</p>\n\n<p>Comparing with the center of each class instead of individual images gives aprox. 0.01LB boost, if I'm not mistaken.</p>\n\n<p>About bbox, I guess it could improve not because it is a better model but because it is an ensemble/TTA of some sorts, since he is predicting on 5 different crops (bboxes) for each image! But so far it is just guess work, I'll take some time latter to train some detector models and validate this assumption.</p>",
          "rawMarkdown": "Sure! For my 0.94 model I still need to try higher resolution and retrain on all data, I'm also using around 1k validation images.\n\nComparing with the center of each class instead of individual images gives aprox. 0.01LB boost, if I'm not mistaken.\n\nAbout bbox, I guess it could improve not because it is a better model but because it is an ensemble/TTA of some sorts, since he is predicting on 5 different crops (bboxes) for each image! But so far it is just guess work, I'll take some time latter to train some detector models and validate this assumption."
        },
        {
          "id": 487843,
          "postDate": "2019-03-11T14:32:12.243Z",
          "content": "<p>How do you use 512FC? Are you adding this layer after Pooling? (Also with DropOut)?</p>",
          "rawMarkdown": "How do you use 512FC? Are you adding this layer after Pooling? (Also with DropOut)?"
        },
        {
          "id": 488027,
          "postDate": "2019-03-11T20:20:53.040Z",
          "content": "<p>```\nclass DenseNet121BN(nn.Module):\n    embeddings_dim = 512</p>\n\n<pre><code>def __init__(self, pretrained=True, **kwargs):\n    super(DenseNet121BN, self).__init__(**kwargs)\n    self.encoder = vision.densenet121(pretrained=pretrained).features\n\n    self.pool = nn.AdaptiveAvgPool2d(1)\n    self.bn1 = nn.BatchNorm1d(1024)\n    self.bn2 = nn.BatchNorm1d(self.embeddings_dim)\n    self.norm = L2Norm()\n    self.dropout = nn.Dropout(0.5)\n    self.embedding = nn.Linear(1024, self.embeddings_dim)\n    self.loss = CosineSoftmaxLoss(self.embeddings_dim, 5004, s=30, m=0.6)\n\n    weights_init_kaiming(self.bn1)\n    weights_init_kaiming(self.bn2)\n    self.bn1.bias.requires_grad_(False)\n    self.bn2.bias.requires_grad_(False)\n\ndef forward(self, x):\n    # Feature extraction\n    mean = [0.485, 0.456, 0.406]\n    std = [0.229, 0.224, 0.225]\n    x[:, 0, :, :] = (x[:, 0, :, :] - mean[0]) / std[0]\n    x[:, 1, :, :] = (x[:, 1, :, :] - mean[1]) / std[1]\n    x[:, 2, :, :] = (x[:, 2, :, :] - mean[2]) / std[2]\n\n    x = self.encoder(x)\n\n    x = self.pool(x)\n    x = x.view(x.size(0), -1)\n    x = self.bn1(x)\n    x = self.dropout(x)\n    x = self.embedding(x)\n    x = self.bn2(x)\n    x = self.norm(x)\n    return x\n\ndef get_loss(self, x, y):\n    x = self.forward(x)\n    loss = self.loss(x, y)\n    return loss\n</code></pre>\n\n<p>```</p>",
          "rawMarkdown": "```\nclass DenseNet121BN(nn.Module):\n    embeddings_dim = 512\n\n    def __init__(self, pretrained=True, **kwargs):\n        super(DenseNet121BN, self).__init__(**kwargs)\n        self.encoder = vision.densenet121(pretrained=pretrained).features\n\n        self.pool = nn.AdaptiveAvgPool2d(1)\n        self.bn1 = nn.BatchNorm1d(1024)\n        self.bn2 = nn.BatchNorm1d(self.embeddings_dim)\n        self.norm = L2Norm()\n        self.dropout = nn.Dropout(0.5)\n        self.embedding = nn.Linear(1024, self.embeddings_dim)\n        self.loss = CosineSoftmaxLoss(self.embeddings_dim, 5004, s=30, m=0.6)\n\n        weights_init_kaiming(self.bn1)\n        weights_init_kaiming(self.bn2)\n        self.bn1.bias.requires_grad_(False)\n        self.bn2.bias.requires_grad_(False)\n\n    def forward(self, x):\n        # Feature extraction\n        mean = [0.485, 0.456, 0.406]\n        std = [0.229, 0.224, 0.225]\n        x[:, 0, :, :] = (x[:, 0, :, :] - mean[0]) / std[0]\n        x[:, 1, :, :] = (x[:, 1, :, :] - mean[1]) / std[1]\n        x[:, 2, :, :] = (x[:, 2, :, :] - mean[2]) / std[2]\n\n        x = self.encoder(x)\n\n        x = self.pool(x)\n        x = x.view(x.size(0), -1)\n        x = self.bn1(x)\n        x = self.dropout(x)\n        x = self.embedding(x)\n        x = self.bn2(x)\n        x = self.norm(x)\n        return x\n\n    def get_loss(self, x, y):\n        x = self.forward(x)\n        loss = self.loss(x, y)\n        return loss\n```",
          "votes": 1
        }
      ]
    },
    {
      "id": 481449,
      "postDate": "2019-03-01T12:18:45.350Z",
      "content": "<p>Here's my part, detail of my ensemble of ProtoNets models.</p>\n\n<p>1.My local ProtoNets models</p>\n\n<p>Basically my models are same as the code on the github, here are two major differences.</p>\n\n<p>1.1 TTA for both prototypes and test samples</p>\n\n<p>I call it as PTTA (Prototype Test Time Augmentation) below. This pushed score about 0.01.\n- When making prototypes, do once training non new whale class images as is.\n- Then update prototypes with augmented images for 4 times.\n- For test samples, it's normal TTA; one with image as is and 4 times more with augmented images.</p>\n\n<p>1.2 Augmentation from Martin's solution</p>\n\n<p>Martin's solution has scratch built augmentation which made different result from my augmentation.\nModels trained with this augmentation helped increasing about 0.02 when ensembled.</p>\n\n<p>2.Ensemble of ProtoNets models</p>\n\n<p>Ensemble of 10 models was best LB score, private:public = 0.91001:0.90719.\nHere are major model's recipes.</p>\n\n<p>2.1 Model #1 LB private:public = 0.88834:0.87837\n- Training: k=60, n=1, data sampling=more than 2 samples, resized to 256 and random cropped 224, model=resnet50\n- Test: data cropping with Radek's bbox resized to 224, applied PTTA</p>\n\n<p>2.2 Model #2 LB score unknown, used for ensemble only\n- Training: k=60, n=1, data sampling=more than 2 samples, Martin's augmentation and cropping, model=resnet50\n- Test: Martin's cropping, applied PTTA</p>\n\n<p>2.3 Model #3 LB score unknown, used for ensemble only\n- Training: k=20, n=1, data sampling=more than 2 samples, resized to 448 and random cropped 384, model=resnet50\n- Test: data cropping with Radek's bbox resized to 384, applied PTTA</p>\n\n<p>Other models are combination of augmentation (Martin's or mine) and image size (384 or 224).</p>\n\n<p>Continues to explain final team ensemble...</p>",
      "rawMarkdown": "Here's my part, detail of my ensemble of ProtoNets models.\n\n1.My local ProtoNets models\n\nBasically my models are same as the code on the github, here are two major differences.\n\n1.1 TTA for both prototypes and test samples\n\nI call it as PTTA (Prototype Test Time Augmentation) below. This pushed score about 0.01.\n- When making prototypes, do once training non new whale class images as is.\n- Then update prototypes with augmented images for 4 times.\n- For test samples, it's normal TTA; one with image as is and 4 times more with augmented images.\n\n1.2 Augmentation from Martin's solution\n\nMartin's solution has scratch built augmentation which made different result from my augmentation.\nModels trained with this augmentation helped increasing about 0.02 when ensembled.\n\n2.Ensemble of ProtoNets models\n\nEnsemble of 10 models was best LB score, private:public = 0.91001:0.90719.\nHere are major model's recipes.\n\n2.1 Model #1 LB private:public = 0.88834:0.87837\n- Training: k=60, n=1, data sampling=more than 2 samples, resized to 256 and random cropped 224, model=resnet50\n- Test: data cropping with Radek's bbox resized to 224, applied PTTA\n\n2.2 Model #2 LB score unknown, used for ensemble only\n- Training: k=60, n=1, data sampling=more than 2 samples, Martin's augmentation and cropping, model=resnet50\n- Test: Martin's cropping, applied PTTA\n\n2.3 Model #3 LB score unknown, used for ensemble only\n- Training: k=20, n=1, data sampling=more than 2 samples, resized to 448 and random cropped 384, model=resnet50\n- Test: data cropping with Radek's bbox resized to 384, applied PTTA\n\nOther models are combination of augmentation (Martin's or mine) and image size (384 or 224).\n\nContinues to explain final team ensemble...",
      "votes": 2,
      "replies": [
        {
          "id": 482716,
          "postDate": "2019-03-03T14:49:45.573Z",
          "content": "<p>I am really happy for both of you Bartek and  <a href=\"/daisukelab\">@daisukelab</a> .. You deserve it.. I didn't know about protonets. Learned a lot from your great code and discussions....  </p>",
          "rawMarkdown": "I am really happy for both of you Bartek and  @daisukelab .. You deserve it.. I didn't know about protonets. Learned a lot from your great code and discussions....  ",
          "votes": 1
        },
        {
          "id": 482906,
          "postDate": "2019-03-03T20:28:23.657Z",
          "content": "<p>Thank you Hainder as <a href=\"/hwasiti\">@hwasiti</a>, I also appreciate that you joined the discussion and made it interesting.</p>",
          "rawMarkdown": "Thank you Hainder as @hwasiti, I also appreciate that you joined the discussion and made it interesting.",
          "votes": 1
        },
        {
          "id": 482955,
          "postDate": "2019-03-03T23:14:11.237Z",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a> for sharing your training procedure. The best I was able to get with Protonet was 0.576 LB score before I run out of time. But the model really helped improve my ensemble score. I tried only resNet18 and ResNet35. Would you be sharing your code with the PTTA (Prototype Test Time Augmentation) ? I want to continue experimenting with this data in a couple of weeks.</p>",
          "rawMarkdown": "Thanks @daisukelab for sharing your training procedure. The best I was able to get with Protonet was 0.576 LB score before I run out of time. But the model really helped improve my ensemble score. I tried only resNet18 and ResNet35. Would you be sharing your code with the PTTA (Prototype Test Time Augmentation) ? I want to continue experimenting with this data in a couple of weeks."
        },
        {
          "id": 483272,
          "postDate": "2019-03-04T12:00:07.640Z",
          "content": "<p>Hi, how much epochs did you try? It would require 300-500 epochs to get your model trained enough.\nSo I was training for one or two days for one model with a single GTX1080Ti.\nAnd regarding PTTA, I will clean up and update repository. Please hold on.</p>",
          "rawMarkdown": "Hi, how much epochs did you try? It would require 300-500 epochs to get your model trained enough.\nSo I was training for one or two days for one model with a single GTX1080Ti.\nAnd regarding PTTA, I will clean up and update repository. Please hold on."
        },
        {
          "id": 483280,
          "postDate": "2019-03-04T12:15:48.203Z",
          "content": "<p>The longest one was 160 epochs. Because your baseline model notebook can reach 0.748 in 100 epochs, I thought that would be enough.</p>",
          "rawMarkdown": "The longest one was 160 epochs. Because your baseline model notebook can reach 0.748 in 100 epochs, I thought that would be enough."
        },
        {
          "id": 483372,
          "postDate": "2019-03-04T14:29:17.030Z",
          "content": "<p>Then you got 0.576 LB, umm... If I could take a look at your code, I might be able to say something...\nHave you try my code as it is? If the result of that scores low, there could be something different from mine...</p>",
          "rawMarkdown": "Then you got 0.576 LB, umm... If I could take a look at your code, I might be able to say something...\nHave you try my code as it is? If the result of that scores low, there could be something different from mine...",
          "votes": 1
        },
        {
          "id": 483692,
          "postDate": "2019-03-05T01:50:21.453Z",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a>. I went back and checked my logs to answer you properly.</p>\n\n<ul>\n<li>Model-1 trained with the only change LR = 5e-3 for 40 epochs, LB=0.470</li>\n<li>Model-2 trained with the only changes LR = 5e-3 , k_train = 100 for 50 epochs, LB=0.647 - For some reason this one brings down my ensemble score unlike model-1 which improved my scores.</li>\n<li>Model-3 trained with the only changes LR = 5e-3 , k_train = 120, ResNet50, trained for 55 epochs, LB=0.572</li>\n<li>Model-4 trained with the default parameters, trained for 120 epochs, LB=0.576 - This one is the continued training by initializing with the weights from model-1 hence trained for a total of 160 epochs.</li>\n</ul>",
          "rawMarkdown": "Thanks @daisukelab. I went back and checked my logs to answer you properly.\n\n- Model-1 trained with the only change LR = 5e-3 for 40 epochs, LB=0.470\n- Model-2 trained with the only changes LR = 5e-3 , k_train = 100 for 50 epochs, LB=0.647 - For some reason this one brings down my ensemble score unlike model-1 which improved my scores.\n- Model-3 trained with the only changes LR = 5e-3 , k_train = 120, ResNet50, trained for 55 epochs, LB=0.572\n- Model-4 trained with the default parameters, trained for 120 epochs, LB=0.576 - This one is the continued training by initializing with the weights from model-1 hence trained for a total of 160 epochs.",
          "votes": 1
        },
        {
          "id": 483741,
          "postDate": "2019-03-05T04:15:39.080Z",
          "content": "<p>Hi,</p>\n\n<p>Regarding model-1&amp;4, it should perform better. One thing could affect is LR, 5e-3 is too high??\nSo I tried LR=5e-3 locally.</p>\n\n<pre><code>Log with LR=5e-3@18 ... categorical_accuracy=0.566, val_1-shot_10-way_acc=0.934\n  vs\nLog with LR=3e-3@18 ... categorical_accuracy=0.69, val_1-shot_10-way_acc=0.967\n</code></pre>\n\n<p>Regarding model-2&amp;3, I cannot try so big k_train, I hope you could see better results, I guess it is with smaller LR. This is one reason I hope to port to fast.ai...</p>",
          "rawMarkdown": "Hi,\n\nRegarding model-1&amp;4, it should perform better. One thing could affect is LR, 5e-3 is too high??\nSo I tried LR=5e-3 locally.\n\n    Log with LR=5e-3@18 ... categorical_accuracy=0.566, val_1-shot_10-way_acc=0.934\n      vs\n    Log with LR=3e-3@18 ... categorical_accuracy=0.69, val_1-shot_10-way_acc=0.967\n\nRegarding model-2&amp;3, I cannot try so big k_train, I hope you could see better results, I guess it is with smaller LR. This is one reason I hope to port to fast.ai...",
          "votes": 1
        },
        {
          "id": 483751,
          "postDate": "2019-03-05T04:34:13.673Z",
          "content": "<p>Thanks again <a href=\"/daisukelab\">@daisukelab</a>. I actually reduced the learning rate back to 3e-3 for the 120 training of model-4. The log actually looked good for model-4 see below:-</p>\n\n<p>&gt;loss=0.117, categorical_accuracy=0.964, val_1-shot_10-way_acc=1]</p>\n\n<p>I will try it again when I start my post competition test on this data. I am in the MSFT Malware competition right now which is a bit challenging :-)</p>",
          "rawMarkdown": "Thanks again @daisukelab. I actually reduced the learning rate back to 3e-3 for the 120 training of model-4. The log actually looked good for model-4 see below:-\n\n&gt;loss=0.117, categorical_accuracy=0.964, val_1-shot_10-way_acc=1]\n\nI will try it again when I start my post competition test on this data. I am in the MSFT Malware competition right now which is a bit challenging :-)"
        }
      ]
    },
    {
      "id": 481374,
      "postDate": "2019-03-01T10:16:53.083Z",
      "content": "<p>@Bartek</p>\n\n<p>Nice solution and great results!</p>\n\n<blockquote>\n  <blockquote>\n    <p>In the past, I was developing the Face-Recognition system  ...</p>\n  </blockquote>\n</blockquote>\n\n<p>I did face recognition before for a short period of time. At first I though it is not going to work for whales because face data has low dimension. However i was surprised by the good results of metric based softmax like center-loss, etc in my later experiments. I think it is because the fluke are planar, making the the discriminative pattern quite stable.</p>\n\n<p>by the way, i work out that \"78% of the data are for private LB\" = 0.78*7960 = 6209 images.1/6209=0.00016. We have the same score of 0.94909 for rank 24 and 25, i.e. practically same number of average correct images. </p>",
      "rawMarkdown": "@Bartek\n\nNice solution and great results!\n\n&gt;&gt;In the past, I was developing the Face-Recognition system  ...\n\nI did face recognition before for a short period of time. At first I though it is not going to work for whales because face data has low dimension. However i was surprised by the good results of metric based softmax like center-loss, etc in my later experiments. I think it is because the fluke are planar, making the the discriminative pattern quite stable.\n\nby the way, i work out that \"78% of the data are for private LB\" = 0.78*7960 = 6209 images.1/6209=0.00016. We have the same score of 0.94909 for rank 24 and 25, i.e. practically same number of average correct images. ",
      "votes": 2,
      "replies": [
        {
          "id": 481383,
          "postDate": "2019-03-01T10:36:28.347Z",
          "content": "<p>Look like you predict one or two images better than us:) I with <a href=\"/daisukelab\">@daisukelab</a> was observing you for the last week, worrying that you would jump over us. And the story came true.... :P </p>",
          "rawMarkdown": "Look like you predict one or two images better than us:) I with @daisukelab was observing you for the last week, worrying that you would jump over us. And the story came true.... :P ",
          "votes": 2
        },
        {
          "id": 481386,
          "postDate": "2019-03-01T10:37:10.247Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 481416,
          "postDate": "2019-03-01T11:34:21.027Z",
          "content": "<p>It was interesting to see that your score improves day by day, it encouraged me to think more and more. :)\nThank you Heng!</p>",
          "rawMarkdown": "It was interesting to see that your score improves day by day, it encouraged me to think more and more. :)\nThank you Heng!"
        }
      ]
    },
    {
      "id": 484122,
      "postDate": "2019-03-05T15:43:14.827Z",
      "content": "<p>Hi Bartek, when you use the bbox coordinates you first crop the images and then resizes it to 448x448. But do you keep the original aspect ratio, i.e. first pad with zeros to a square and than resizes or do you simple squash it into a square changing the original aspect ratio? Thanks</p>",
      "rawMarkdown": "Hi Bartek, when you use the bbox coordinates you first crop the images and then resizes it to 448x448. But do you keep the original aspect ratio, i.e. first pad with zeros to a square and than resizes or do you simple squash it into a square changing the original aspect ratio? Thanks",
      "replies": [
        {
          "id": 484167,
          "postDate": "2019-03-05T16:32:36.770Z",
          "content": "<p>I have just quash it into a square changing the original aspect ratio.  This is why I also trained second model with image size 256x748 to keep the ratio better. However, the accuracy of both model are comparable. </p>",
          "rawMarkdown": "I have just quash it into a square changing the original aspect ratio.  This is why I also trained second model with image size 256x748 to keep the ratio better. However, the accuracy of both model are comparable. ",
          "votes": 1
        },
        {
          "id": 484172,
          "postDate": "2019-03-05T16:38:47.190Z",
          "content": "<p>I'm tracking down what are the things that impaired my model performance in comparison with yours and Pudae's. </p>\n\n<p>It seems there are two main things: first Radek's bbox annotations are superior to the ones from Martin's kernel and the second is that I did not use any normalization after the last pooling layer. Using Radek's annotations and layer normalization improved performance considerably! </p>\n\n<p>Thanks for your help!</p>",
          "rawMarkdown": "I'm tracking down what are the things that impaired my model performance in comparison with yours and Pudae's. \n\nIt seems there are two main things: first Radek's bbox annotations are superior to the ones from Martin's kernel and the second is that I did not use any normalization after the last pooling layer. Using Radek's annotations and layer normalization improved performance considerably! \n\nThanks for your help!",
          "votes": 1
        },
        {
          "id": 484244,
          "postDate": "2019-03-05T18:20:52.720Z",
          "content": "<p>Did you also CosFace/ArcFace?\nWhat model? What Augmentation? \nI'm aslo compering my solution to Pudae's and look like mean-feature per class give him nice boost. I need to check it also.\nAlso, I was using just AvgPool (where Pudae flatten-&gt;Dropout&gt;BN-&gt;FC-&gt;BN-&gt;FC), what did you use?</p>",
          "rawMarkdown": "Did you also CosFace/ArcFace?\nWhat model? What Augmentation? \nI'm aslo compering my solution to Pudae's and look like mean-feature per class give him nice boost. I need to check it also.\nAlso, I was using just AvgPool (where Pudae flatten-&gt;Dropout&gt;BN-&gt;FC-&gt;BN-&gt;FC), what did you use?",
          "votes": 1
        },
        {
          "id": 484362,
          "postDate": "2019-03-05T22:01:39.880Z",
          "content": "<p>I've tried both. CosFace with 0.6 margin is slightly better than ArcFace from what I've been testing. I guess it is the same for you.</p>\n\n<p>So far I'm using DenseNet121 with gray scale images with: </p>\n\n<p><code>Aug1 = iaa.Sequential(\n        [\n            iaa.Fliplr(0.5), \n            iaa.Sometimes(0.5, iaa.Affine(\n                scale={\"x\": (0.95, 1.05), \"y\": (0.95, 1.05)},\n                shear=(-10, 10),\n                rotate=(-10, 10),\n                order=1, \n                cval=0,\n                mode='constant'\n            )),\n            iaa.OneOf([\n                iaa.GammaContrast((0.5, 1.5)),\n                iaa.LinearContrast((0.5, 1.5)),\n                iaa.ContrastNormalization((0.70, 1.30)),\n                ]),\n        ])</code></p>\n\n<p>Next step I'll use colored images with grayscale augmentation</p>",
          "rawMarkdown": "I've tried both. CosFace with 0.6 margin is slightly better than ArcFace from what I've been testing. I guess it is the same for you.\n\nSo far I'm using DenseNet121 with gray scale images with: \n\n```Aug1 = iaa.Sequential(\n        [\n            iaa.Fliplr(0.5), \n            iaa.Sometimes(0.5, iaa.Affine(\n                scale={\"x\": (0.95, 1.05), \"y\": (0.95, 1.05)},\n                shear=(-10, 10),\n                rotate=(-10, 10),\n                order=1, \n                cval=0,\n                mode='constant'\n            )),\n            iaa.OneOf([\n                iaa.GammaContrast((0.5, 1.5)),\n                iaa.LinearContrast((0.5, 1.5)),\n                iaa.ContrastNormalization((0.70, 1.30)),\n                ]),\n        ])```\n\nNext step I'll use colored images with grayscale augmentation\n"
        },
        {
          "id": 484363,
          "postDate": "2019-03-05T22:02:44.533Z",
          "content": "<p>I've also compared AvgPool with flatten and AvgPool seems slightly better. But the best that I have used seems to be GeM pool.</p>",
          "rawMarkdown": "I've also compared AvgPool with flatten and AvgPool seems slightly better. But the best that I have used seems to be GeM pool."
        }
      ]
    },
    {
      "id": 481823,
      "postDate": "2019-03-01T22:16:14.933Z",
      "content": "<p>Interesting solution, congrats</p>",
      "rawMarkdown": "Interesting solution, congrats"
    },
    {
      "id": 481426,
      "postDate": "2019-03-01T11:48:53.587Z",
      "content": "<p>Congratulations on your work Bartek and Daisuke! Let me ask you something, did you use any fully connected layer or just plugged the CosFace layer on top of the global pooling? And what kind of pooling did you use? Thanks!</p>",
      "rawMarkdown": "Congratulations on your work Bartek and Daisuke! Let me ask you something, did you use any fully connected layer or just plugged the CosFace layer on top of the global pooling? And what kind of pooling did you use? Thanks!",
      "replies": [
        {
          "id": 481437,
          "postDate": "2019-03-01T12:02:17.393Z",
          "content": "<p>I was using normal Fully connected layer. Exactyl it look like that in PyTorch:</p>\n\n<p>class NormLinear(nn.Module):</p>\n\n<pre><code>def __init__(self, in_features, out_features, temperature = 0.05, temperature_trainable = False):\n    super(NormLinear, self).__init__()\n    self.weight = nn.Parameter(torch.Tensor(out_features, in_features))\n    nn.init.kaiming_uniform_(self.weight, a=math.sqrt(5))\n\n    self.scale = 1 / temperature\n    if temperature_trainable:\n        self.scale = nn.Parameter(torch.Tensor(1))\n        nn.init.constant_(self.scale, 1 / temperature)\n\ndef forward(self, x):\n    x_norm = F.normalize(x)\n    w_norm = F.normalize(self.weight)\n    cosine = F.linear(x_norm, w_norm, None)\n    out = cosine #* self.scale\n    return out\n\n\n# model preparation\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel_name = 'se_resnext101_32x4d'\nmodel = pretrainedmodels.__dict__[model_name](num_classes=1000, pretrained='imagenet')\nmodel.avg_pool = nn.AdaptiveAvgPool2d((1, 1))\nmodel.last_linear = nn.Sequential(*\n          [nn.LayerNorm(model.last_linear.in_features, elementwise_affine = False),\n          NormLinear(model.last_linear.in_features, 5004)])\n</code></pre>\n\n<p>'</p>",
          "rawMarkdown": "I was using normal Fully connected layer. Exactyl it look like that in PyTorch:\n\n\n\nclass NormLinear(nn.Module):\n\n    def __init__(self, in_features, out_features, temperature = 0.05, temperature_trainable = False):\n        super(NormLinear, self).__init__()\n        self.weight = nn.Parameter(torch.Tensor(out_features, in_features))\n        nn.init.kaiming_uniform_(self.weight, a=math.sqrt(5))\n\n        self.scale = 1 / temperature\n        if temperature_trainable:\n            self.scale = nn.Parameter(torch.Tensor(1))\n            nn.init.constant_(self.scale, 1 / temperature)\n\n    def forward(self, x):\n        x_norm = F.normalize(x)\n        w_norm = F.normalize(self.weight)\n        cosine = F.linear(x_norm, w_norm, None)\n        out = cosine #* self.scale\n        return out\n\n\n    # model preparation\n    device = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\n    model_name = 'se_resnext101_32x4d'\n    model = pretrainedmodels.__dict__[model_name](num_classes=1000, pretrained='imagenet')\n    model.avg_pool = nn.AdaptiveAvgPool2d((1, 1))\n    model.last_linear = nn.Sequential(*\n              [nn.LayerNorm(model.last_linear.in_features, elementwise_affine = False),\n              NormLinear(model.last_linear.in_features, 5004)])\n'\n",
          "votes": 3
        },
        {
          "id": 481459,
          "postDate": "2019-03-01T12:31:51.017Z",
          "content": "<p>Thanks! But what about the cosine margin? </p>",
          "rawMarkdown": "Thanks! But what about the cosine margin? "
        },
        {
          "id": 481471,
          "postDate": "2019-03-01T13:05:18.573Z",
          "content": "<p>I will relese full code later, but this is CosFace implementation:</p>\n\n<pre><code>class CosineMarginCrossEntropy(nn.Module):\ndef __init__(self, m=0.60, s=30.0):\n    super(CosineMarginCrossEntropy, self).__init__()\n    self.m = m\n    self.s = s\n    self.ce = torch.nn.CrossEntropyLoss()\n\ndef forward(self, input, target):\n    one_hot = torch.zeros_like(input)\n    one_hot.scatter_(1, target.view(-1, 1), 1.0)\n    # -------------torch.where(out_i = {x_i if condition_i else y_i) -------------\n    output = self.s * (input - one_hot * self.m)\n    loss = self.ce(output, target)\n    return loss\n\ncriterion = CosineMarginCrossEntropy().cuda()\n</code></pre>\n\n<p>The code which I post there is the main idea of all my score.</p>",
          "rawMarkdown": "I will relese full code later, but this is CosFace implementation:\n\n\n\n    class CosineMarginCrossEntropy(nn.Module):\n    def __init__(self, m=0.60, s=30.0):\n        super(CosineMarginCrossEntropy, self).__init__()\n        self.m = m\n        self.s = s\n        self.ce = torch.nn.CrossEntropyLoss()\n\n    def forward(self, input, target):\n        one_hot = torch.zeros_like(input)\n        one_hot.scatter_(1, target.view(-1, 1), 1.0)\n        # -------------torch.where(out_i = {x_i if condition_i else y_i) -------------\n        output = self.s * (input - one_hot * self.m)\n        loss = self.ce(output, target)\n        return loss\n\n    criterion = CosineMarginCrossEntropy().cuda()\n\nThe code which I post there is the main idea of all my score.",
          "votes": 2
        }
      ]
    },
    {
      "id": 481392,
      "postDate": "2019-03-01T10:46:08.277Z",
      "content": "<p>Congrats <a href=\"/melgor\">@melgor</a> and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @melgor and thanks for sharing your solution overview."
    },
    {
      "id": 481300,
      "postDate": "2019-03-01T08:19:26.767Z",
      "content": "<p>Congratulations <a href=\"/melgor\">@melgor</a> thanks for sharing your summary.</p>",
      "rawMarkdown": "Congratulations @melgor thanks for sharing your summary.\n"
    }
  ],
  "comments": [
    {
      "id": 481640,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "2019-03-01T16:44:15.433000",
      "content": "<p>Here's about team ensemble.</p>\n\n<p>Basically we found that ensemble of different approach works really well, combination of CosFace approach and ProtoNets.</p>\n\n<ul>\n<li>First team ensemble, 2 CosFace models + 8 ProtoNets models was LB private:public = 0.94168:0.93314</li>\n<li>Second team ensemble, 3 CosFace models + 3 ProtoNets models was LB private:public = 0.94627:0.94045</li>\n<li>Final team ensemble, 11 CosFace models + 3 ProtoNets models was LB private:public = 0.94909:0.94810</li>\n</ul>\n\n<p>All results are once processed by softmax and geometric mean was calculated.\nThen 'new_whale' threshold was set to 30% for almost all results.</p>\n\n<p>Lastly, this two weeks was special to me.\nI have been working with ProtoNets from the beginning of this year,\njoined this competition just basically for checking performance of that.\nBut once I shared my ProtoNets repo in the <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/81085\">thread</a>,\nmany people seems to have enjoyed and it became very nice discussion. It was wonderful, many thanks to Heng as <a href=\"/hengck23\">@hengck23</a>.</p>\n\n<p>And then @Bartek interested in my solution and kindly invited me to make a team. And now it resulted in effective ensemble.\nAs I was spending most of time for ProtoNets only, I couldn't even try CNN+effective loss approach in time by myself, so it was very good that we could try something different.\nThank you teammate @Bartek!</p>\n\n<p>So... I spend almost all time, less sleep, think and try and tried again, really (healthily) tired.</p>\n\n<p>Thank you all, congratulations to winners! Now I can go to bed, zzzz...</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 485617,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2019-03-07T16:51:23.813000",
      "content": "<p><a href=\"/melgor\">@melgor</a> so far I still haven't been able to reach the 0.68LB without TTA and new_whale. I can only reach 0.667LB but I won't give up haha.</p>\n\n<p>Are you using data from the playground competition?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 485661,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-07T18:38:33.940000",
          "content": "<p>In fact, I'm continuting experiments with Whales, that is why I did not release the code. I will try to find tomorrow some time to release the code.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 485745,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-07T21:13:51.260000",
          "content": "<p>Implementing top solutions is very important in the learning process. But don't worry, take your time! :)</p>\n\n<p>I've noticed that the top1 solution trained on both this competition and playground's data. I thought that the playground's images were a subset of this compeition's training set, I haven't used any playground data so far, I'll take a look into that</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 487420,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-10T20:53:30.023000",
          "content": "<p>Hi, I added code: <a href=\"https://github.com/melgor/kaggle-whale-tail\">https://github.com/melgor/kaggle-whale-tail</a>\nBut this is minimal example, without creating submission with new-whale, learn single model. My code is rather mess as I spend not so much time for this copetition (mostly weekends).</p>\n\n<p>In general, I also did not used playgroud-data. Also the 3rd solution learn for a very long time (500 epochs vs 40 mine). I did some experiments and look like training longer res50 can match se101 using same input-size (so it is interesting). Then I decided to try again with removing avg-pool but training longer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 487466,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-10T23:52:52.837000",
          "content": "<p>Thanks Bartek, I'll take a look later.</p>\n\n<p>So far I've managed to get 0.67502/0.67966 on public/private respectively (no new_whale and no tta, approx. 0.94 with it) using CosFace with your margin and a DenseNet121 (320x320) with a 512FC layer. It took me 31 epochs using AdamW with 5e-4 weight decay.  I also used pudae's augmentation (added flip since I'm not doubling up the identities yet) and radek's bbox as augmentation (0.5 prob of using it or not during training). For inference I computed the center of every class.</p>\n\n<p>Now, about using flattening and training longer I've tried both and I got worse results. What helped was changing LayerNorm for BatchNorm without shift (bias=0 and non-trainable).</p>\n\n<p>May I ask how much are you getting with a single model so far? I think part of pudae's good results are due to averaging predictions on 5 different bbox during inference. I got considerably better results just by changing martins bbox annotations for radek's, so I guess there are still room for improvements following this path.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 487581,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-11T06:39:30.377000",
          "content": "<p>Currently I'm maily focus on traning longer using Res50,  my best model on 224x224, my standard augmentation, pure classification (no center for every class) I get 0.656. \nThe results is not best-one, but based on my all history, I'm able to get 0.02 better score using same model just for longer training (I use 100 epochs) but not training on all data (1k images for validation).  And Res50 because it learn fast:)</p>\n\n<p>About LayerNorm vs BatchNorm, I will test it again to confirm if there is any difference.\nThen I will also try pudae's augmentation.\nAnd finaly center of every class.</p>\n\n<p>About BB, I I'm not sure if training beter detector would help a lot. I trained ones (using radek code) and compare to the second release of BB. In both the results was pretty much same.</p>\n\n<p>Also, best model within competition could get 0.682 (se101, 448x448)-&gt; 0.93. But I did not use center features, so maybe you would get 0.94:)</p>\n\n<p>If you would get score &gt; 0.95 with single model, let me know:)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 487731,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-11T11:46:23.403000",
          "content": "<p>Sure! For my 0.94 model I still need to try higher resolution and retrain on all data, I'm also using around 1k validation images.</p>\n\n<p>Comparing with the center of each class instead of individual images gives aprox. 0.01LB boost, if I'm not mistaken.</p>\n\n<p>About bbox, I guess it could improve not because it is a better model but because it is an ensemble/TTA of some sorts, since he is predicting on 5 different crops (bboxes) for each image! But so far it is just guess work, I'll take some time latter to train some detector models and validate this assumption.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 487843,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-11T14:32:12.243000",
          "content": "<p>How do you use 512FC? Are you adding this layer after Pooling? (Also with DropOut)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488027,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-11T20:20:53.040000",
          "content": "<p>```\nclass DenseNet121BN(nn.Module):\n    embeddings_dim = 512</p>\n\n<pre><code>def __init__(self, pretrained=True, **kwargs):\n    super(DenseNet121BN, self).__init__(**kwargs)\n    self.encoder = vision.densenet121(pretrained=pretrained).features\n\n    self.pool = nn.AdaptiveAvgPool2d(1)\n    self.bn1 = nn.BatchNorm1d(1024)\n    self.bn2 = nn.BatchNorm1d(self.embeddings_dim)\n    self.norm = L2Norm()\n    self.dropout = nn.Dropout(0.5)\n    self.embedding = nn.Linear(1024, self.embeddings_dim)\n    self.loss = CosineSoftmaxLoss(self.embeddings_dim, 5004, s=30, m=0.6)\n\n    weights_init_kaiming(self.bn1)\n    weights_init_kaiming(self.bn2)\n    self.bn1.bias.requires_grad_(False)\n    self.bn2.bias.requires_grad_(False)\n\ndef forward(self, x):\n    # Feature extraction\n    mean = [0.485, 0.456, 0.406]\n    std = [0.229, 0.224, 0.225]\n    x[:, 0, :, :] = (x[:, 0, :, :] - mean[0]) / std[0]\n    x[:, 1, :, :] = (x[:, 1, :, :] - mean[1]) / std[1]\n    x[:, 2, :, :] = (x[:, 2, :, :] - mean[2]) / std[2]\n\n    x = self.encoder(x)\n\n    x = self.pool(x)\n    x = x.view(x.size(0), -1)\n    x = self.bn1(x)\n    x = self.dropout(x)\n    x = self.embedding(x)\n    x = self.bn2(x)\n    x = self.norm(x)\n    return x\n\ndef get_loss(self, x, y):\n    x = self.forward(x)\n    loss = self.loss(x, y)\n    return loss\n</code></pre>\n\n<p>```</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 481449,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "2019-03-01T12:18:45.350000",
      "content": "<p>Here's my part, detail of my ensemble of ProtoNets models.</p>\n\n<p>1.My local ProtoNets models</p>\n\n<p>Basically my models are same as the code on the github, here are two major differences.</p>\n\n<p>1.1 TTA for both prototypes and test samples</p>\n\n<p>I call it as PTTA (Prototype Test Time Augmentation) below. This pushed score about 0.01.\n- When making prototypes, do once training non new whale class images as is.\n- Then update prototypes with augmented images for 4 times.\n- For test samples, it's normal TTA; one with image as is and 4 times more with augmented images.</p>\n\n<p>1.2 Augmentation from Martin's solution</p>\n\n<p>Martin's solution has scratch built augmentation which made different result from my augmentation.\nModels trained with this augmentation helped increasing about 0.02 when ensembled.</p>\n\n<p>2.Ensemble of ProtoNets models</p>\n\n<p>Ensemble of 10 models was best LB score, private:public = 0.91001:0.90719.\nHere are major model's recipes.</p>\n\n<p>2.1 Model #1 LB private:public = 0.88834:0.87837\n- Training: k=60, n=1, data sampling=more than 2 samples, resized to 256 and random cropped 224, model=resnet50\n- Test: data cropping with Radek's bbox resized to 224, applied PTTA</p>\n\n<p>2.2 Model #2 LB score unknown, used for ensemble only\n- Training: k=60, n=1, data sampling=more than 2 samples, Martin's augmentation and cropping, model=resnet50\n- Test: Martin's cropping, applied PTTA</p>\n\n<p>2.3 Model #3 LB score unknown, used for ensemble only\n- Training: k=20, n=1, data sampling=more than 2 samples, resized to 448 and random cropped 384, model=resnet50\n- Test: data cropping with Radek's bbox resized to 384, applied PTTA</p>\n\n<p>Other models are combination of augmentation (Martin's or mine) and image size (384 or 224).</p>\n\n<p>Continues to explain final team ensemble...</p>",
      "votes": 2,
      "replies": [
        {
          "id": 482716,
          "author_name": "Haider Alwasiti",
          "author_url": "",
          "post_date": "2019-03-03T14:49:45.573000",
          "content": "<p>I am really happy for both of you Bartek and  <a href=\"/daisukelab\">@daisukelab</a> .. You deserve it.. I didn't know about protonets. Learned a lot from your great code and discussions....  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 482906,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-03-03T20:28:23.657000",
          "content": "<p>Thank you Hainder as <a href=\"/hwasiti\">@hwasiti</a>, I also appreciate that you joined the discussion and made it interesting.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 482955,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-03-03T23:14:11.237000",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a> for sharing your training procedure. The best I was able to get with Protonet was 0.576 LB score before I run out of time. But the model really helped improve my ensemble score. I tried only resNet18 and ResNet35. Would you be sharing your code with the PTTA (Prototype Test Time Augmentation) ? I want to continue experimenting with this data in a couple of weeks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 483272,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-03-04T12:00:07.640000",
          "content": "<p>Hi, how much epochs did you try? It would require 300-500 epochs to get your model trained enough.\nSo I was training for one or two days for one model with a single GTX1080Ti.\nAnd regarding PTTA, I will clean up and update repository. Please hold on.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 483280,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-03-04T12:15:48.203000",
          "content": "<p>The longest one was 160 epochs. Because your baseline model notebook can reach 0.748 in 100 epochs, I thought that would be enough.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 483372,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-03-04T14:29:17.030000",
          "content": "<p>Then you got 0.576 LB, umm... If I could take a look at your code, I might be able to say something...\nHave you try my code as it is? If the result of that scores low, there could be something different from mine...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 483692,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-03-05T01:50:21.453000",
          "content": "<p>Thanks <a href=\"/daisukelab\">@daisukelab</a>. I went back and checked my logs to answer you properly.</p>\n\n<ul>\n<li>Model-1 trained with the only change LR = 5e-3 for 40 epochs, LB=0.470</li>\n<li>Model-2 trained with the only changes LR = 5e-3 , k_train = 100 for 50 epochs, LB=0.647 - For some reason this one brings down my ensemble score unlike model-1 which improved my scores.</li>\n<li>Model-3 trained with the only changes LR = 5e-3 , k_train = 120, ResNet50, trained for 55 epochs, LB=0.572</li>\n<li>Model-4 trained with the default parameters, trained for 120 epochs, LB=0.576 - This one is the continued training by initializing with the weights from model-1 hence trained for a total of 160 epochs.</li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 483741,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-03-05T04:15:39.080000",
          "content": "<p>Hi,</p>\n\n<p>Regarding model-1&amp;4, it should perform better. One thing could affect is LR, 5e-3 is too high??\nSo I tried LR=5e-3 locally.</p>\n\n<pre><code>Log with LR=5e-3@18 ... categorical_accuracy=0.566, val_1-shot_10-way_acc=0.934\n  vs\nLog with LR=3e-3@18 ... categorical_accuracy=0.69, val_1-shot_10-way_acc=0.967\n</code></pre>\n\n<p>Regarding model-2&amp;3, I cannot try so big k_train, I hope you could see better results, I guess it is with smaller LR. This is one reason I hope to port to fast.ai...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 483751,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-03-05T04:34:13.673000",
          "content": "<p>Thanks again <a href=\"/daisukelab\">@daisukelab</a>. I actually reduced the learning rate back to 3e-3 for the 120 training of model-4. The log actually looked good for model-4 see below:-</p>\n\n<p>&gt;loss=0.117, categorical_accuracy=0.964, val_1-shot_10-way_acc=1]</p>\n\n<p>I will try it again when I start my post competition test on this data. I am in the MSFT Malware competition right now which is a bit challenging :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 481374,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-03-01T10:16:53.083000",
      "content": "<p>@Bartek</p>\n\n<p>Nice solution and great results!</p>\n\n<blockquote>\n  <blockquote>\n    <p>In the past, I was developing the Face-Recognition system  ...</p>\n  </blockquote>\n</blockquote>\n\n<p>I did face recognition before for a short period of time. At first I though it is not going to work for whales because face data has low dimension. However i was surprised by the good results of metric based softmax like center-loss, etc in my later experiments. I think it is because the fluke are planar, making the the discriminative pattern quite stable.</p>\n\n<p>by the way, i work out that \"78% of the data are for private LB\" = 0.78*7960 = 6209 images.1/6209=0.00016. We have the same score of 0.94909 for rank 24 and 25, i.e. practically same number of average correct images. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 481383,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-01T10:36:28.347000",
          "content": "<p>Look like you predict one or two images better than us:) I with <a href=\"/daisukelab\">@daisukelab</a> was observing you for the last week, worrying that you would jump over us. And the story came true.... :P </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 481386,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-03-01T10:37:10.247000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 481416,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "2019-03-01T11:34:21.027000",
          "content": "<p>It was interesting to see that your score improves day by day, it encouraged me to think more and more. :)\nThank you Heng!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 484122,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2019-03-05T15:43:14.827000",
      "content": "<p>Hi Bartek, when you use the bbox coordinates you first crop the images and then resizes it to 448x448. But do you keep the original aspect ratio, i.e. first pad with zeros to a square and than resizes or do you simple squash it into a square changing the original aspect ratio? Thanks</p>",
      "votes": 0,
      "replies": [
        {
          "id": 484167,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-05T16:32:36.770000",
          "content": "<p>I have just quash it into a square changing the original aspect ratio.  This is why I also trained second model with image size 256x748 to keep the ratio better. However, the accuracy of both model are comparable. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 484172,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-05T16:38:47.190000",
          "content": "<p>I'm tracking down what are the things that impaired my model performance in comparison with yours and Pudae's. </p>\n\n<p>It seems there are two main things: first Radek's bbox annotations are superior to the ones from Martin's kernel and the second is that I did not use any normalization after the last pooling layer. Using Radek's annotations and layer normalization improved performance considerably! </p>\n\n<p>Thanks for your help!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 484244,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-05T18:20:52.720000",
          "content": "<p>Did you also CosFace/ArcFace?\nWhat model? What Augmentation? \nI'm aslo compering my solution to Pudae's and look like mean-feature per class give him nice boost. I need to check it also.\nAlso, I was using just AvgPool (where Pudae flatten-&gt;Dropout&gt;BN-&gt;FC-&gt;BN-&gt;FC), what did you use?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 484362,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-05T22:01:39.880000",
          "content": "<p>I've tried both. CosFace with 0.6 margin is slightly better than ArcFace from what I've been testing. I guess it is the same for you.</p>\n\n<p>So far I'm using DenseNet121 with gray scale images with: </p>\n\n<p><code>Aug1 = iaa.Sequential(\n        [\n            iaa.Fliplr(0.5), \n            iaa.Sometimes(0.5, iaa.Affine(\n                scale={\"x\": (0.95, 1.05), \"y\": (0.95, 1.05)},\n                shear=(-10, 10),\n                rotate=(-10, 10),\n                order=1, \n                cval=0,\n                mode='constant'\n            )),\n            iaa.OneOf([\n                iaa.GammaContrast((0.5, 1.5)),\n                iaa.LinearContrast((0.5, 1.5)),\n                iaa.ContrastNormalization((0.70, 1.30)),\n                ]),\n        ])</code></p>\n\n<p>Next step I'll use colored images with grayscale augmentation</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 484363,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-05T22:02:44.533000",
          "content": "<p>I've also compared AvgPool with flatten and AvgPool seems slightly better. But the best that I have used seems to be GeM pool.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 481823,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-03-01T22:16:14.933000",
      "content": "<p>Interesting solution, congrats</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 481426,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2019-03-01T11:48:53.587000",
      "content": "<p>Congratulations on your work Bartek and Daisuke! Let me ask you something, did you use any fully connected layer or just plugged the CosFace layer on top of the global pooling? And what kind of pooling did you use? Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 481437,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-01T12:02:17.393000",
          "content": "<p>I was using normal Fully connected layer. Exactyl it look like that in PyTorch:</p>\n\n<p>class NormLinear(nn.Module):</p>\n\n<pre><code>def __init__(self, in_features, out_features, temperature = 0.05, temperature_trainable = False):\n    super(NormLinear, self).__init__()\n    self.weight = nn.Parameter(torch.Tensor(out_features, in_features))\n    nn.init.kaiming_uniform_(self.weight, a=math.sqrt(5))\n\n    self.scale = 1 / temperature\n    if temperature_trainable:\n        self.scale = nn.Parameter(torch.Tensor(1))\n        nn.init.constant_(self.scale, 1 / temperature)\n\ndef forward(self, x):\n    x_norm = F.normalize(x)\n    w_norm = F.normalize(self.weight)\n    cosine = F.linear(x_norm, w_norm, None)\n    out = cosine #* self.scale\n    return out\n\n\n# model preparation\ndevice = torch.device(\"cuda\" if torch.cuda.is_available() else \"cpu\")\nmodel_name = 'se_resnext101_32x4d'\nmodel = pretrainedmodels.__dict__[model_name](num_classes=1000, pretrained='imagenet')\nmodel.avg_pool = nn.AdaptiveAvgPool2d((1, 1))\nmodel.last_linear = nn.Sequential(*\n          [nn.LayerNorm(model.last_linear.in_features, elementwise_affine = False),\n          NormLinear(model.last_linear.in_features, 5004)])\n</code></pre>\n\n<p>'</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 481459,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-01T12:31:51.017000",
          "content": "<p>Thanks! But what about the cosine margin? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 481471,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-01T13:05:18.573000",
          "content": "<p>I will relese full code later, but this is CosFace implementation:</p>\n\n<pre><code>class CosineMarginCrossEntropy(nn.Module):\ndef __init__(self, m=0.60, s=30.0):\n    super(CosineMarginCrossEntropy, self).__init__()\n    self.m = m\n    self.s = s\n    self.ce = torch.nn.CrossEntropyLoss()\n\ndef forward(self, input, target):\n    one_hot = torch.zeros_like(input)\n    one_hot.scatter_(1, target.view(-1, 1), 1.0)\n    # -------------torch.where(out_i = {x_i if condition_i else y_i) -------------\n    output = self.s * (input - one_hot * self.m)\n    loss = self.ce(output, target)\n    return loss\n\ncriterion = CosineMarginCrossEntropy().cuda()\n</code></pre>\n\n<p>The code which I post there is the main idea of all my score.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 481392,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-03-01T10:46:08.277000",
      "content": "<p>Congrats <a href=\"/melgor\">@melgor</a> and thanks for sharing your solution overview.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 481300,
      "author_name": "Karthik Chowdary Tsaliki",
      "author_url": "",
      "post_date": "2019-03-01T08:19:26.767000",
      "content": "<p>Congratulations <a href=\"/melgor\">@melgor</a> thanks for sharing your summary.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "481275": "Here I would describe my part of the solution: CosFace approch.\nThe ProtoNets and Ensebling would be described by @daisukelab\n\n**1. Preprocessing**\nFist of all, I use BB from @radek (thank you!). I just did train model on his annotated data and that's all, I did not invest time for this part of competiton. In later stage I also used updated BB from @radek (I call them v2), but the difference in final results was very small.\n\n**2. Model and Data-Loading**\nMy final model was se-resnext101 (I also try se154 but it did not work nice). In fact, my model and augmentation was exactly the same like in this kernel: https://www.kaggle.com/stalkermustang/pytorch-pretraiedmodels-se-resnext101-baseline\n Other stuff which I tried:\n -  CutOut -&gt; fail\n -  MixUp -&gt; fail\n - OverSample -&gt; fail\n - Cluster-Based Sampling: here I try to sample batch so that there were similar classes (based on cosine similarity) -&gt; fail, overfitting\n\n**3. Loss Function**\nIn the past, I was developing the Face-Recognition system and the main purpose of approaching the whale problem was testing such technology on whales:) So in general I was testing ArcFace and CosFace. Both of them works pretty good but CosFace was slightly better. For CosFace I use very high margin 0.6 (in original paper it was 0.35). \nIn fact, this step take me the longest time (3 weeks), where I was trying:\n\n - BatchNorm vs LayerNorm before L2 normalization: LayerNorm better\n - AlphaDropout vs DropOut: Alpha better but with no influence in debugging model (resnet50) so I did not use both of them, what now I think was one of the biggest mistake I made\n - CosFace vs ArcFace vs SphereFace: CosFace with m=0.6 was clear winner (I also the the idea of the CosFace the most)\n\n**4. Optimalization**\nHere I use AdamW (with fixed weight-decay). I also try OneCycle but it was not working (I think that   code was wrong). In general I train the model by 30 epochs, so it was pretty quick.\n\n - New-whale\nI did not use 'new-whale' for training. In sumbission I just want to have ~27% of 'new_whale'\nI have two approches for this problem which did not work:\n - Each new-whale as different class: In general it was ok, the accuracy on validation set was just 0.2% less. But it does not work well on LB.\n - Use second loss-function (KL-Divergence) which would force the outout distribution afer SoftMax of 'new-whale' to be uniform (so exactly the same probability for each class). It also work fine for validation set, but not in LB. \n\nLook like I would need more time for this approach (especially second one), because I really like it:)\n\nThe final models was se101resnext trained on:\n- gray and 448x448\n- gray and 256x748\n- rgb and 448x448\n- rgb and 256x748\n\nIn general, the approch was pretty simple and work moderate. But look like @pudae had the similar idea, so I'm now looking into his approach :)\n\nThe ProtoNet and Ensemble part would be explained by @daisukelab.\n\nCode: https://github.com/melgor/kaggle-whale-tail\nThis is minimal code for training single model and create sumbission without 'new-whale'",
    "481640": "Here's about team ensemble.\n\nBasically we found that ensemble of different approach works really well, combination of CosFace approach and ProtoNets.\n\n- First team ensemble, 2 CosFace models + 8 ProtoNets models was LB private:public = 0.94168:0.93314\n- Second team ensemble, 3 CosFace models + 3 ProtoNets models was LB private:public = 0.94627:0.94045\n- Final team ensemble, 11 CosFace models + 3 ProtoNets models was LB private:public = 0.94909:0.94810\n\nAll results are once processed by softmax and geometric mean was calculated.\nThen 'new_whale' threshold was set to 30% for almost all results.\n\nLastly, this two weeks was special to me.\nI have been working with ProtoNets from the beginning of this year,\njoined this competition just basically for checking performance of that.\nBut once I shared my ProtoNets repo in the [thread](https://www.kaggle.com/c/humpback-whale-identification/discussion/81085),\nmany people seems to have enjoyed and it became very nice discussion. It was wonderful, many thanks to Heng as @hengck23.\n\nAnd then @Bartek interested in my solution and kindly invited me to make a team. And now it resulted in effective ensemble.\nAs I was spending most of time for ProtoNets only, I couldn't even try CNN+effective loss approach in time by myself, so it was very good that we could try something different.\nThank you teammate @Bartek!\n\nSo... I spend almost all time, less sleep, think and try and tried again, really (healthily) tired.\n\nThank you all, congratulations to winners! Now I can go to bed, zzzz...",
    "485617": "@melgor so far I still haven't been able to reach the 0.68LB without TTA and new_whale. I can only reach 0.667LB but I won't give up haha.\n\nAre you using data from the playground competition?",
    "481449": "Here's my part, detail of my ensemble of ProtoNets models.\n\n1.My local ProtoNets models\n\nBasically my models are same as the code on the github, here are two major differences.\n\n1.1 TTA for both prototypes and test samples\n\nI call it as PTTA (Prototype Test Time Augmentation) below. This pushed score about 0.01.\n- When making prototypes, do once training non new whale class images as is.\n- Then update prototypes with augmented images for 4 times.\n- For test samples, it's normal TTA; one with image as is and 4 times more with augmented images.\n\n1.2 Augmentation from Martin's solution\n\nMartin's solution has scratch built augmentation which made different result from my augmentation.\nModels trained with this augmentation helped increasing about 0.02 when ensembled.\n\n2.Ensemble of ProtoNets models\n\nEnsemble of 10 models was best LB score, private:public = 0.91001:0.90719.\nHere are major model's recipes.\n\n2.1 Model #1 LB private:public = 0.88834:0.87837\n- Training: k=60, n=1, data sampling=more than 2 samples, resized to 256 and random cropped 224, model=resnet50\n- Test: data cropping with Radek's bbox resized to 224, applied PTTA\n\n2.2 Model #2 LB score unknown, used for ensemble only\n- Training: k=60, n=1, data sampling=more than 2 samples, Martin's augmentation and cropping, model=resnet50\n- Test: Martin's cropping, applied PTTA\n\n2.3 Model #3 LB score unknown, used for ensemble only\n- Training: k=20, n=1, data sampling=more than 2 samples, resized to 448 and random cropped 384, model=resnet50\n- Test: data cropping with Radek's bbox resized to 384, applied PTTA\n\nOther models are combination of augmentation (Martin's or mine) and image size (384 or 224).\n\nContinues to explain final team ensemble...",
    "481374": "@Bartek\n\nNice solution and great results!\n\n&gt;&gt;In the past, I was developing the Face-Recognition system  ...\n\nI did face recognition before for a short period of time. At first I though it is not going to work for whales because face data has low dimension. However i was surprised by the good results of metric based softmax like center-loss, etc in my later experiments. I think it is because the fluke are planar, making the the discriminative pattern quite stable.\n\nby the way, i work out that \"78% of the data are for private LB\" = 0.78*7960 = 6209 images.1/6209=0.00016. We have the same score of 0.94909 for rank 24 and 25, i.e. practically same number of average correct images. ",
    "484122": "Hi Bartek, when you use the bbox coordinates you first crop the images and then resizes it to 448x448. But do you keep the original aspect ratio, i.e. first pad with zeros to a square and than resizes or do you simple squash it into a square changing the original aspect ratio? Thanks",
    "481823": "Interesting solution, congrats",
    "481426": "Congratulations on your work Bartek and Daisuke! Let me ask you something, did you use any fully connected layer or just plugged the CosFace layer on top of the global pooling? And what kind of pooling did you use? Thanks!",
    "481392": "Congrats @melgor and thanks for sharing your solution overview.",
    "481300": "Congratulations @melgor thanks for sharing your summary.\n"
  }
}