{
  "id": 82484,
  "title": "3rd place solution with code: ArcFace",
  "url": "/competitions/humpback-whale-identification/discussion/82484",
  "author_name": "pudae",
  "post_date": "2019-03-01T17:48:26.279000",
  "votes": 95,
  "comment_count": 35,
  "views": 0,
  "content": "<h2>UPDATE: code available on github</h2>\n\n<p><a href=\"https://github.com/pudae/kaggle-humpback\">https://github.com/pudae/kaggle-humpback</a></p>\n\n<hr>\n\n<p>Congrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.</p>\n\n<p>Here is my solution summary.</p>\n\n<h1>Solution Summary</h1>\n\n<h2>Dataset</h2>\n\n<ul>\n<li><strong>Validation set</strong>: randomly sampled 400 identities that has 2 images + 110 new whales (= 400 * 0.276).</li>\n<li><strong>training set</strong>: all images except new whales.</li>\n<li>I doubled up the identities by horizontal flip.</li>\n</ul>\n\n<h2>Model</h2>\n\n<p><strong>bounding box &amp; landmark</strong></p>\n\n<ul>\n<li>I used annotations by <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/78699\">Paul Johnson</a> and <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/76281\">Radek Osmulski</a>. (Thanks to Paul and Radek. Without your contribution, I couldn't achieve such high score.)</li>\n<li>I made 5 fold CV and trained 5 models using them.</li>\n<li>IOU: 0.93</li>\n</ul>\n\n<p><strong>whale identifier</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/pdf/1801.07698.pdf\">ArcFace</a> approach is used.</li>\n<li>Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.</li>\n<li>densenet121</li>\n<li>m 0.5 (the default value of the paper)</li>\n<li>weight decay 0.0005, droupout 0.5</li>\n</ul>\n\n<h2>Augmentation</h2>\n\n<ul>\n<li>average blur, motion blur</li>\n<li>add, multiply, grayscale</li>\n<li>scale, translate, shear, rotate</li>\n<li>align or no-align</li>\n</ul>\n\n<h2>Training</h2>\n\n<ul>\n<li>adam optimizer</li>\n<li>learning rate of 0.00025 -&gt; 0.000125 -&gt; 0.0000625</li>\n</ul>\n\n<h2>Inference</h2>\n\n<p><strong>getting embedding feature for identity</strong></p>\n\n<ul>\n<li>For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.</li>\n<li>For each identities, the center of all feature vectors was used as final embedding feature.</li>\n</ul>\n\n<p><strong>getting embedding feature for test image</strong></p>\n\n<ul>\n<li>For each images, multiple feature vectors were generate and the center of the feature vectors was used.</li>\n</ul>\n\n<p><strong>computing similarity</strong></p>\n\n<ul>\n<li>The cosine similarity of above two feature vectors was used as the measure of similarity.</li>\n</ul>\n\n<p><strong>selecting threshold</strong></p>\n\n<ul>\n<li>The threshold for new whale was selected so that the proportion of new whale is about 0.276.</li>\n</ul>\n\n<h1>The process to the final method</h1>\n\n<p>Followings are the process to the final method.</p>\n\n<p><strong>without landmark</strong></p>\n\n<p>At first, I excluded the identities having only one image and new whales from the training set. For inference, the identity of the most similar image of the training set was used as the predicted identity.</p>\n\n<p>&gt; Public LB: 0.90, Private LB: 0.90 </p>\n\n<p>After using the center of all feature vectors in the same identity, I got</p>\n\n<p>&gt; Public LB: 0.942 / Private LB: 0.939</p>\n\n<p>After using weight decay 0.0005</p>\n\n<p>&gt; Public LB: 0.946 / Private LB: 0.946</p>\n\n<p>After including the identities having one image to training set</p>\n\n<p>&gt; Public LB: 0.963 / Private LB: 0.961</p>\n\n<p><strong>with landmark</strong></p>\n\n<p>When I used aligned image, network was trained faster but the score was not improved.</p>\n\n<p>&gt; Public LB: 0.962 / Private LB: 0.959</p>\n\n<p>The bounding boxes and landmarks of some images are very poor and it seems to prevent improving scores. So I also used non-aligned images.</p>\n\n<p>&gt; Public LB: 0.965 / Private LB: 0.961</p>\n\n<p>Finally, I doubled up identities by horizontal flip. Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.</p>\n\n<p>&gt; Public LB: 0.968 ~ 0.971 / Private LB: 0.965 ~ 0.968</p>\n\n<p>Congrats to winners again.\nThanks.</p>",
  "messages": [
    {
      "id": 481676,
      "postDate": "2019-03-01T17:48:26.280Z",
      "content": "<h2>UPDATE: code available on github</h2>\n\n<p><a href=\"https://github.com/pudae/kaggle-humpback\">https://github.com/pudae/kaggle-humpback</a></p>\n\n<hr>\n\n<p>Congrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.</p>\n\n<p>Here is my solution summary.</p>\n\n<h1>Solution Summary</h1>\n\n<h2>Dataset</h2>\n\n<ul>\n<li><strong>Validation set</strong>: randomly sampled 400 identities that has 2 images + 110 new whales (= 400 * 0.276).</li>\n<li><strong>training set</strong>: all images except new whales.</li>\n<li>I doubled up the identities by horizontal flip.</li>\n</ul>\n\n<h2>Model</h2>\n\n<p><strong>bounding box &amp; landmark</strong></p>\n\n<ul>\n<li>I used annotations by <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/78699\">Paul Johnson</a> and <a href=\"https://www.kaggle.com/c/humpback-whale-identification/discussion/76281\">Radek Osmulski</a>. (Thanks to Paul and Radek. Without your contribution, I couldn't achieve such high score.)</li>\n<li>I made 5 fold CV and trained 5 models using them.</li>\n<li>IOU: 0.93</li>\n</ul>\n\n<p><strong>whale identifier</strong></p>\n\n<ul>\n<li><a href=\"https://arxiv.org/pdf/1801.07698.pdf\">ArcFace</a> approach is used.</li>\n<li>Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.</li>\n<li>densenet121</li>\n<li>m 0.5 (the default value of the paper)</li>\n<li>weight decay 0.0005, droupout 0.5</li>\n</ul>\n\n<h2>Augmentation</h2>\n\n<ul>\n<li>average blur, motion blur</li>\n<li>add, multiply, grayscale</li>\n<li>scale, translate, shear, rotate</li>\n<li>align or no-align</li>\n</ul>\n\n<h2>Training</h2>\n\n<ul>\n<li>adam optimizer</li>\n<li>learning rate of 0.00025 -&gt; 0.000125 -&gt; 0.0000625</li>\n</ul>\n\n<h2>Inference</h2>\n\n<p><strong>getting embedding feature for identity</strong></p>\n\n<ul>\n<li>For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.</li>\n<li>For each identities, the center of all feature vectors was used as final embedding feature.</li>\n</ul>\n\n<p><strong>getting embedding feature for test image</strong></p>\n\n<ul>\n<li>For each images, multiple feature vectors were generate and the center of the feature vectors was used.</li>\n</ul>\n\n<p><strong>computing similarity</strong></p>\n\n<ul>\n<li>The cosine similarity of above two feature vectors was used as the measure of similarity.</li>\n</ul>\n\n<p><strong>selecting threshold</strong></p>\n\n<ul>\n<li>The threshold for new whale was selected so that the proportion of new whale is about 0.276.</li>\n</ul>\n\n<h1>The process to the final method</h1>\n\n<p>Followings are the process to the final method.</p>\n\n<p><strong>without landmark</strong></p>\n\n<p>At first, I excluded the identities having only one image and new whales from the training set. For inference, the identity of the most similar image of the training set was used as the predicted identity.</p>\n\n<p>&gt; Public LB: 0.90, Private LB: 0.90 </p>\n\n<p>After using the center of all feature vectors in the same identity, I got</p>\n\n<p>&gt; Public LB: 0.942 / Private LB: 0.939</p>\n\n<p>After using weight decay 0.0005</p>\n\n<p>&gt; Public LB: 0.946 / Private LB: 0.946</p>\n\n<p>After including the identities having one image to training set</p>\n\n<p>&gt; Public LB: 0.963 / Private LB: 0.961</p>\n\n<p><strong>with landmark</strong></p>\n\n<p>When I used aligned image, network was trained faster but the score was not improved.</p>\n\n<p>&gt; Public LB: 0.962 / Private LB: 0.959</p>\n\n<p>The bounding boxes and landmarks of some images are very poor and it seems to prevent improving scores. So I also used non-aligned images.</p>\n\n<p>&gt; Public LB: 0.965 / Private LB: 0.961</p>\n\n<p>Finally, I doubled up identities by horizontal flip. Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.</p>\n\n<p>&gt; Public LB: 0.968 ~ 0.971 / Private LB: 0.965 ~ 0.968</p>\n\n<p>Congrats to winners again.\nThanks.</p>",
      "rawMarkdown": "## UPDATE: code available on github\n[https://github.com/pudae/kaggle-humpback](https://github.com/pudae/kaggle-humpback)\n\n---\n\nCongrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.\n\nHere is my solution summary.\nSolution Summary\n==============\n\nDataset\n---------\n- **Validation set**: randomly sampled 400 identities that has 2 images + 110 new whales (= 400 * 0.276).\n- **training set**: all images except new whales.\n- I doubled up the identities by horizontal flip.\n\nModel\n-------\n**bounding box &amp; landmark**\n\n- I used annotations by [Paul Johnson](https://www.kaggle.com/c/humpback-whale-identification/discussion/78699) and [Radek Osmulski](https://www.kaggle.com/c/humpback-whale-identification/discussion/76281). (Thanks to Paul and Radek. Without your contribution, I couldn't achieve such high score.)\n- I made 5 fold CV and trained 5 models using them.\n- IOU: 0.93\n\n**whale identifier**\n\n- [ArcFace](https://arxiv.org/pdf/1801.07698.pdf) approach is used.\n- Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.\n- densenet121\n- m 0.5 (the default value of the paper)\n- weight decay 0.0005, droupout 0.5\n\nAugmentation\n-----------------\n- average blur, motion blur\n- add, multiply, grayscale\n- scale, translate, shear, rotate\n- align or no-align\n\nTraining\n---------\n- adam optimizer\n- learning rate of 0.00025 -&gt; 0.000125 -&gt; 0.0000625\n\nInference\n-----------\n**getting embedding feature for identity**\n\n- For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.\n- For each identities, the center of all feature vectors was used as final embedding feature.\n\n**getting embedding feature for test image**\n\n- For each images, multiple feature vectors were generate and the center of the feature vectors was used.\n\n**computing similarity**\n\n- The cosine similarity of above two feature vectors was used as the measure of similarity.\n\n**selecting threshold**\n\n- The threshold for new whale was selected so that the proportion of new whale is about 0.276.\n\nThe process to the final method\n========================\nFollowings are the process to the final method.\n\n**without landmark**\n\nAt first, I excluded the identities having only one image and new whales from the training set. For inference, the identity of the most similar image of the training set was used as the predicted identity.\n\n &gt; Public LB: 0.90, Private LB: 0.90 \n\nAfter using the center of all feature vectors in the same identity, I got\n\n&gt; Public LB: 0.942 / Private LB: 0.939\n\nAfter using weight decay 0.0005\n\n&gt; Public LB: 0.946 / Private LB: 0.946\n\nAfter including the identities having one image to training set\n\n&gt; Public LB: 0.963 / Private LB: 0.961\n\n**with landmark**\n\nWhen I used aligned image, network was trained faster but the score was not improved.\n\n&gt; Public LB: 0.962 / Private LB: 0.959\n\nThe bounding boxes and landmarks of some images are very poor and it seems to prevent improving scores. So I also used non-aligned images.\n\n&gt; Public LB: 0.965 / Private LB: 0.961\n\nFinally, I doubled up identities by horizontal flip. Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.\n\n&gt; Public LB: 0.968 ~ 0.971 / Private LB: 0.965 ~ 0.968\n\nCongrats to winners again.\nThanks.",
      "votes": 95
    },
    {
      "id": 481706,
      "postDate": "2019-03-01T18:26:09.657Z",
      "content": "<p>Thanks for the write-up Pudae. Congrats on the 3rd place finish!</p>\n\n<p>Glad to hear that my keypoints played a part in a top winning solution!</p>",
      "rawMarkdown": "Thanks for the write-up Pudae. Congrats on the 3rd place finish!\n\nGlad to hear that my keypoints played a part in a top winning solution!",
      "votes": 6,
      "replies": [
        {
          "id": 481912,
          "postDate": "2019-03-02T03:03:17.753Z",
          "content": "<p>Thank you again very much!!</p>",
          "rawMarkdown": "Thank you again very much!!",
          "votes": 2
        }
      ]
    },
    {
      "id": 485577,
      "postDate": "2019-03-07T15:24:01.270Z",
      "content": "<p>Congrats <a href=\"/pudae81\">@pudae81</a>!</p>\n\n<blockquote>\n  <p>I made 5 fold CV and trained 5 models using them.\n  For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.</p>\n</blockquote>\n\n<p>Nice approach!\nHow did you use five bounding boxes in training? Randomly use one as augmentation?</p>",
      "rawMarkdown": "Congrats @pudae81!\n\n&gt; I made 5 fold CV and trained 5 models using them.\n&gt; For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.\n\nNice approach!\nHow did you use five bounding boxes in training? Randomly use one as augmentation?",
      "votes": 1,
      "replies": [
        {
          "id": 485865,
          "postDate": "2019-03-08T02:35:32.907Z",
          "content": "<p>The method using 5 bounding boxes came to mind too late to use it on training.</p>",
          "rawMarkdown": "The method using 5 bounding boxes came to mind too late to use it on training.",
          "votes": 2
        },
        {
          "id": 485898,
          "postDate": "2019-03-08T04:09:42.870Z",
          "content": "<p>I see! Thx for your reply!</p>",
          "rawMarkdown": "I see! Thx for your reply!"
        },
        {
          "id": 486510,
          "postDate": "2019-03-08T22:25:17.143Z",
          "content": "<p><a href=\"/pudae81\">@pudae81</a> all the LB values that you reported use flip-TTA and 5 bbox-ensembling/average? Even for the lower scores like 0.9LB? </p>",
          "rawMarkdown": "@pudae81 all the LB values that you reported use flip-TTA and 5 bbox-ensembling/average? Even for the lower scores like 0.9LB? "
        },
        {
          "id": 488444,
          "postDate": "2019-03-12T13:25:43.363Z",
          "content": "<p>From the submissions over LB 0.968, that method was used.</p>",
          "rawMarkdown": "From the submissions over LB 0.968, that method was used.",
          "votes": 1
        }
      ]
    },
    {
      "id": 482077,
      "postDate": "2019-03-02T09:30:57.323Z",
      "content": "<p>Hi, <a href=\"/pudae81\">@pudae81</a>!  Thanks for sharing your solution!\nI have several questions:\n1)    so your solution is a pure classifier, right? I wonder why ArcFace can tackle such class skew  problem. did you used any oversampling or heavy augmentation for the imbalance problem?\n2)   I noticed that using the mean vector of an identity for its representation makes a dramatic boost on your prediction. would  you please share the motivation of this idea? \n3)  what's the hyper param of <em>s</em> in your ArcFace metric and how you tune it?</p>",
      "rawMarkdown": "Hi, @pudae81!  Thanks for sharing your solution!\nI have several questions:\n1)    so your solution is a pure classifier, right? I wonder why ArcFace can tackle such class skew  problem. did you used any oversampling or heavy augmentation for the imbalance problem?\n2)   I noticed that using the mean vector of an identity for its representation makes a dramatic boost on your prediction. would  you please share the motivation of this idea? \n3)  what's the hyper param of *s* in your ArcFace metric and how you tune it?",
      "votes": 1,
      "replies": [
        {
          "id": 483398,
          "postDate": "2019-03-04T15:04:55.323Z",
          "content": "<blockquote>\n  <p>1) so your solution is a pure classifier, right? I wonder why ArcFace can tackle such class skew problem. did you used any oversampling or heavy augmentation for the imbalance problem?</p>\n</blockquote>\n\n<ul>\n<li>I have tried nothing for the imbalance problem.</li>\n</ul>\n\n<blockquote>\n  <p>2) I noticed that using the mean vector of an identity for its representation makes a dramatic boost on your prediction. would you please share the motivation of this idea? </p>\n</blockquote>\n\n<ul>\n<li><p>From the <a href=\"https://arxiv.org/pdf/1801.07698.pdf\">paper</a>: </p>\n\n<blockquote>\n  <p>To get the embedding features for templates (e.g. IJB-B and IJB-C) or videos (e.g. YTF and iQIYI-VID), we simply calculate the feature centre of all images from the template or all frames from the video.</p>\n</blockquote></li>\n</ul>\n\n<blockquote>\n  <p>3) what's the hyper param of s in your ArcFace metric and how you tune it?</p>\n</blockquote>\n\n<ul>\n<li>I have used s(=65) and m(=0.5) which are the default value of the paper. I have tried several values, but it didn't improve the scores.</li>\n</ul>",
          "rawMarkdown": "&gt; 1) so your solution is a pure classifier, right? I wonder why ArcFace can tackle such class skew problem. did you used any oversampling or heavy augmentation for the imbalance problem?\n\n- I have tried nothing for the imbalance problem.\n\n&gt; 2) I noticed that using the mean vector of an identity for its representation makes a dramatic boost on your prediction. would you please share the motivation of this idea? \n\n- From the [paper](https://arxiv.org/pdf/1801.07698.pdf): \n\n    &gt; To get the embedding features for templates (e.g. IJB-B and IJB-C) or videos (e.g. YTF and iQIYI-VID), we simply calculate the feature centre of all images from the template or all frames from the video.\n\n&gt; 3) what's the hyper param of s in your ArcFace metric and how you tune it?\n\n- I have used s(=65) and m(=0.5) which are the default value of the paper. I have tried several values, but it didn't improve the scores.\n",
          "votes": 2
        },
        {
          "id": 484538,
          "postDate": "2019-03-06T06:46:35.350Z",
          "content": "<p>What Image-size did you use for training?</p>",
          "rawMarkdown": "What Image-size did you use for training?"
        },
        {
          "id": 484846,
          "postDate": "2019-03-06T15:10:59.593Z",
          "content": "<p>The image size was 320x320.</p>",
          "rawMarkdown": "The image size was 320x320."
        },
        {
          "id": 485041,
          "postDate": "2019-03-06T21:04:13.177Z",
          "content": "<p>&gt; For each identities, the center of all feature vectors was used as final embedding feature.</p>\n\n<p>When you say center, you mean that for each individual you run the net for all images of that individual, L2 normalize it and take the mean? Then during test you compare (dot product) the test image's embedding with each class/individual center?</p>",
          "rawMarkdown": "&gt; For each identities, the center of all feature vectors was used as final embedding feature.\n\nWhen you say center, you mean that for each individual you run the net for all images of that individual, L2 normalize it and take the mean? Then during test you compare (dot product) the test image's embedding with each class/individual center?"
        },
        {
          "id": 485092,
          "postDate": "2019-03-06T22:54:35.667Z",
          "content": "<blockquote>\n  <p>When you say center, you mean that for each individual you run the net for all images of that individual, L2 normalize it and take the mean? </p>\n</blockquote>\n\n<ul>\n<li>After taking the mean, I normalized it again to locate it on the hypersphere.</li>\n</ul>\n\n<blockquote>\n  <p>Then during test you compare (dot product) the test image's embedding with each class/individual center?</p>\n</blockquote>\n\n<ul>\n<li>You're right</li>\n</ul>",
          "rawMarkdown": "&gt; When you say center, you mean that for each individual you run the net for all images of that individual, L2 normalize it and take the mean? \n\n- After taking the mean, I normalized it again to locate it on the hypersphere.\n\n&gt; Then during test you compare (dot product) the test image's embedding with each class/individual center?\n\n- You're right",
          "votes": 1
        }
      ]
    },
    {
      "id": 481814,
      "postDate": "2019-03-01T22:08:42.293Z",
      "content": "<p>Congrats @pudae , great solution!</p>",
      "rawMarkdown": "Congrats @pudae , great solution!",
      "votes": 1,
      "replies": [
        {
          "id": 481909,
          "postDate": "2019-03-02T03:00:49.250Z",
          "content": "<p>Thanks :)</p>",
          "rawMarkdown": "Thanks :)",
          "votes": 2
        },
        {
          "id": 486636,
          "postDate": "2019-03-09T06:27:13.503Z",
          "content": "<p>hi pudae\nGreat solution..\nCould you point which file in repository implementing the Arc loss ? I couldnt find one there</p>",
          "rawMarkdown": "hi pudae\nGreat solution..\nCould you point which file in repository implementing the Arc loss ? I couldnt find one there"
        },
        {
          "id": 502552,
          "postDate": "2019-03-28T18:08:14.410Z",
          "content": "<p>hi pudae..\nTo calculate the feature center ,used as class reference for Test  . I used below code. But I got almost same score without. \nDid i undrestand correctly what you meant feature centre ?</p>\n\n<p>```\nfor c in list(set(data2.train_ds.y.items)): # unique classes found in train set\n     l.append(torch.mean(preds1t[torch.nonzero(y1t == c)],dim=0))</p>\n\n<p>trn_centre=torch.cat(l)\ntrn_centre.size() # 5004 * 5004 matrix</p>\n\n<p>for feat in preds1: # test preds of size no of images * no of classes\n        dists = F.cosine_similarity(trn_centre, feat.unsqueeze(0).repeat(5004, 1))\n        predicted_similarity = dists.cuda()#learn.model.head(dists.cuda())\n        sims.append(predicted_similarity.squeeze().detach().cpu())</p>\n\n<h1>generating the preds</h1>\n\n<p>classes = df.Id.unique()\nnew_whale_idx = np.where(classes == 'new_whale')[0][0]\ntop_5s = []\nfor sim in sims:\n    idxs = sim.argsort(descending=True)\n    probs = sim[idxs]\n    top_5 = []\n    for i, p in zip(idxs, probs):\n        if 'new_whale' not in top_5 and p &lt;0.4 and len(top_5) &lt; 5: #615#575 res,.58 63 for dense\n          top_5.append('new_whale')\n        if len(top_5) == 5: break\n        if i == new_whale_idx: continue\n        predicted_class = labels_list[i]\n        if predicted_class not in top_5: top_5.append(predicted_class)\n    top_5s.append(top_5)\n```</p>",
          "rawMarkdown": "hi pudae..\nTo calculate the feature center ,used as class reference for Test  . I used below code. But I got almost same score without. \nDid i undrestand correctly what you meant feature centre ?\n\n```\nfor c in list(set(data2.train_ds.y.items)): # unique classes found in train set\n     l.append(torch.mean(preds1t[torch.nonzero(y1t == c)],dim=0))\n\ntrn_centre=torch.cat(l)\ntrn_centre.size() # 5004 * 5004 matrix\n\nfor feat in preds1: # test preds of size no of images * no of classes\n        dists = F.cosine_similarity(trn_centre, feat.unsqueeze(0).repeat(5004, 1))\n        predicted_similarity = dists.cuda()#learn.model.head(dists.cuda())\n        sims.append(predicted_similarity.squeeze().detach().cpu())\n\n# generating the preds\nclasses = df.Id.unique()\nnew_whale_idx = np.where(classes == 'new_whale')[0][0]\ntop_5s = []\nfor sim in sims:\n    idxs = sim.argsort(descending=True)\n    probs = sim[idxs]\n    top_5 = []\n    for i, p in zip(idxs, probs):\n        if 'new_whale' not in top_5 and p &lt;0.4 and len(top_5) &lt; 5: #615#575 res,.58 63 for dense\n          top_5.append('new_whale')\n        if len(top_5) == 5: break\n        if i == new_whale_idx: continue\n        predicted_class = labels_list[i]\n        if predicted_class not in top_5: top_5.append(predicted_class)\n    top_5s.append(top_5)\n```"
        }
      ]
    },
    {
      "id": 481805,
      "postDate": "2019-03-01T21:55:24.507Z",
      "content": "<p>Congratulations on your work! </p>\n\n<blockquote>\n  <p>Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.</p>\n</blockquote>\n\n<ul>\n<li>The FC you mentioned is before the 5004 dimensions normalized FC of the paper, right? If yes how many dimensions did you use?</li>\n</ul>\n\n<blockquote>\n  <p>Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.</p>\n</blockquote>\n\n<ul>\n<li>So now you have ~10008 identities, right? But I didn't understand what you did with the logit. Can you elaborate?</li>\n</ul>\n\n<p>Thank you!</p>",
      "rawMarkdown": "Congratulations on your work! \n&gt; Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.\n\n- The FC you mentioned is before the 5004 dimensions normalized FC of the paper, right? If yes how many dimensions did you use?\n\n&gt; Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.\n\n- So now you have ~10008 identities, right? But I didn't understand what you did with the logit. Can you elaborate?\n\nThank you!",
      "votes": 1,
      "replies": [
        {
          "id": 481908,
          "postDate": "2019-03-02T03:00:35.483Z",
          "content": "<blockquote>\n  <p>The FC you mentioned is before the 5004 dimensions normalized FC of the paper, right?</p>\n</blockquote>\n\n<p>Yes, you're right.</p>\n\n<blockquote>\n  <p>If yes how many dimensions did you use?</p>\n</blockquote>\n\n<p>the dimension of the feature vector was 512.</p>\n\n<blockquote>\n  <p>So now you have ~10008 identities, right? </p>\n</blockquote>\n\n<p>Yes~</p>\n\n<blockquote>\n  <p>But I didn't understand what you did with the logit. Can you elaborate?</p>\n</blockquote>\n\n<p>Without zero setting, network would be trained such that the distance between flip and non-flip are apart from each others. It seemed to make training difficult in my  case. Setting the logit value of flipped label to zero was helpful in this case.</p>\n\n<p>Thank you!!</p>",
          "rawMarkdown": "&gt; The FC you mentioned is before the 5004 dimensions normalized FC of the paper, right?\n\nYes, you're right.\n\n&gt; If yes how many dimensions did you use?\n\nthe dimension of the feature vector was 512.\n\n&gt; So now you have ~10008 identities, right? \n\nYes~\n\n&gt; But I didn't understand what you did with the logit. Can you elaborate?\n\nWithout zero setting, network would be trained such that the distance between flip and non-flip are apart from each others. It seemed to make training difficult in my  case. Setting the logit value of flipped label to zero was helpful in this case.\n\nThank you!!",
          "votes": 3
        }
      ]
    },
    {
      "id": 481680,
      "postDate": "2019-03-01T17:54:19.727Z",
      "content": "<p>Great solution! I have once tried flip-doubling, but first attempt was unsuccessfull, so I gave up :(\nHow much blur-augmentation added?</p>",
      "rawMarkdown": "Great solution! I have once tried flip-doubling, but first attempt was unsuccessfull, so I gave up :(\nHow much blur-augmentation added?",
      "votes": 1,
      "replies": [
        {
          "id": 481913,
          "postDate": "2019-03-02T03:05:02.910Z",
          "content": "<p>The kernel size for blurring was between 3 to 5.\nThank you~ :)</p>",
          "rawMarkdown": "The kernel size for blurring was between 3 to 5.\nThank you~ :)",
          "votes": 4
        }
      ]
    },
    {
      "id": 1676889,
      "postDate": "2022-02-05T11:21:32.423Z",
      "content": "<p>Thanks for sharing its amazing how Object Detection approach is used here</p>",
      "rawMarkdown": "Thanks for sharing its amazing how Object Detection approach is used here"
    },
    {
      "id": 1675197,
      "postDate": "2022-02-04T05:12:02.850Z",
      "content": "<p><code>For each identities, the center of all feature vectors was used as final embedding feature.</code><br>\nwhat do you mean by center here? </p>",
      "rawMarkdown": "`For each identities, the center of all feature vectors was used as final embedding feature.`\nwhat do you mean by center here? "
    },
    {
      "id": 482830,
      "postDate": "2019-03-03T17:54:16.703Z",
      "content": "<p>&gt; After including the identities having one image to training set</p>\n\n<p>Did you fin tune the previous model which had smaller number of classes, or trained from scratch with all clases?</p>",
      "rawMarkdown": "&gt; After including the identities having one image to training set\n\nDid you fin tune the previous model which had smaller number of classes, or trained from scratch with all clases?",
      "replies": [
        {
          "id": 483389,
          "postDate": "2019-03-04T14:54:20.997Z",
          "content": "<p>I trained from scratch.</p>",
          "rawMarkdown": "I trained from scratch."
        },
        {
          "id": 488516,
          "postDate": "2019-03-12T15:39:18.740Z",
          "content": "<p>hi pudae...\n1) How fast does Arcface loss converges... I see it converges very slowly . \n2) why do we need to renomalize the features ,we are using bn which i think does normalizes the activations of layer.</p>\n\n<p><code>\nfeatures = self.bn2(features)\nfeatures = F.normalize(features\n</code></p>",
          "rawMarkdown": "hi pudae...\n1) How fast does Arcface loss converges... I see it converges very slowly . \n2) why do we need to renomalize the features ,we are using bn which i think does normalizes the activations of layer.\n\n```\nfeatures = self.bn2(features)\nfeatures = F.normalize(features\n```\n"
        },
        {
          "id": 488610,
          "postDate": "2019-03-12T18:38:21.150Z",
          "content": "<p>Please read original paper which should explain you the idea of normalization: <a href=\"https://arxiv.org/abs/1801.07698\">https://arxiv.org/abs/1801.07698</a></p>\n\n<p>Also, loss converge very slow, it's true because the task defined by ArcFace is much harder than classfication. But classification accuracy converge much faster.</p>",
          "rawMarkdown": "Please read original paper which should explain you the idea of normalization: https://arxiv.org/abs/1801.07698\n\nAlso, loss converge very slow, it's true because the task defined by ArcFace is much harder than classfication. But classification accuracy converge much faster.",
          "votes": 2
        }
      ]
    },
    {
      "id": 482827,
      "postDate": "2019-03-03T17:52:49.127Z",
      "content": "<p>Nice and concise! Do you plan to upload the source code?</p>",
      "rawMarkdown": "Nice and concise! Do you plan to upload the source code?",
      "replies": [
        {
          "id": 483390,
          "postDate": "2019-03-04T14:54:39.533Z",
          "content": "<p>After clean up the codes, I will share it. :)</p>",
          "rawMarkdown": "After clean up the codes, I will share it. :)",
          "votes": 2
        }
      ]
    },
    {
      "id": 482540,
      "postDate": "2019-03-03T07:18:16.263Z",
      "content": "<p>Congratulations !\nIt is really amazing to see your center works. I tried to use center to calc similarity but fail to work :(.</p>",
      "rawMarkdown": "Congratulations !\nIt is really amazing to see your center works. I tried to use center to calc similarity but fail to work :(.",
      "replies": [
        {
          "id": 483391,
          "postDate": "2019-03-04T14:55:09.950Z",
          "content": "<p>Thank you!! </p>",
          "rawMarkdown": "Thank you!! "
        }
      ]
    },
    {
      "id": 481972,
      "postDate": "2019-03-02T05:53:48.957Z",
      "content": "<p>Congrats @pudai81 for a very strong solo finish and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @pudai81 for a very strong solo finish and thanks for sharing your solution overview."
    },
    {
      "id": 481944,
      "postDate": "2019-03-02T04:50:01.443Z",
      "content": "<blockquote>\n  <blockquote>\n    <p>When I used aligned image, network was trained faster but the score was not improved.</p>\n  </blockquote>\n</blockquote>\n\n<p>have you tried ensembled of aligned and non-aligned images? that is a common trick in face recognition.</p>",
      "rawMarkdown": "&gt;&gt;When I used aligned image, network was trained faster but the score was not improved.\n\nhave you tried ensembled of aligned and non-aligned images? that is a common trick in face recognition.",
      "replies": [
        {
          "id": 481961,
          "postDate": "2019-03-02T05:29:24.417Z",
          "content": "<p>Yes. LB &gt; 0.965 models used aligned and non-aligned image on training and inference. </p>",
          "rawMarkdown": "Yes. LB &gt; 0.965 models used aligned and non-aligned image on training and inference. "
        }
      ]
    },
    {
      "id": 956861,
      "postDate": "2020-08-03T21:00:06.377Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 481706,
      "author_name": "Paul Johnson",
      "author_url": "",
      "post_date": "2019-03-01T18:26:09.657000",
      "content": "<p>Thanks for the write-up Pudae. Congrats on the 3rd place finish!</p>\n\n<p>Glad to hear that my keypoints played a part in a top winning solution!</p>",
      "votes": 6,
      "replies": [
        {
          "id": 481912,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-02T03:03:17.753000",
          "content": "<p>Thank you again very much!!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 485577,
      "author_name": "yu4u",
      "author_url": "",
      "post_date": "2019-03-07T15:24:01.270000",
      "content": "<p>Congrats <a href=\"/pudae81\">@pudae81</a>!</p>\n\n<blockquote>\n  <p>I made 5 fold CV and trained 5 models using them.\n  For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.</p>\n</blockquote>\n\n<p>Nice approach!\nHow did you use five bounding boxes in training? Randomly use one as augmentation?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 485865,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-08T02:35:32.907000",
          "content": "<p>The method using 5 bounding boxes came to mind too late to use it on training.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 485898,
          "author_name": "yu4u",
          "author_url": "",
          "post_date": "2019-03-08T04:09:42.870000",
          "content": "<p>I see! Thx for your reply!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 486510,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-08T22:25:17.143000",
          "content": "<p><a href=\"/pudae81\">@pudae81</a> all the LB values that you reported use flip-TTA and 5 bbox-ensembling/average? Even for the lower scores like 0.9LB? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488444,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-12T13:25:43.363000",
          "content": "<p>From the submissions over LB 0.968, that method was used.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 482077,
      "author_name": "gengshi",
      "author_url": "",
      "post_date": "2019-03-02T09:30:57.323000",
      "content": "<p>Hi, <a href=\"/pudae81\">@pudae81</a>!  Thanks for sharing your solution!\nI have several questions:\n1)    so your solution is a pure classifier, right? I wonder why ArcFace can tackle such class skew  problem. did you used any oversampling or heavy augmentation for the imbalance problem?\n2)   I noticed that using the mean vector of an identity for its representation makes a dramatic boost on your prediction. would  you please share the motivation of this idea? \n3)  what's the hyper param of <em>s</em> in your ArcFace metric and how you tune it?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 483398,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-04T15:04:55.323000",
          "content": "<blockquote>\n  <p>1) so your solution is a pure classifier, right? I wonder why ArcFace can tackle such class skew problem. did you used any oversampling or heavy augmentation for the imbalance problem?</p>\n</blockquote>\n\n<ul>\n<li>I have tried nothing for the imbalance problem.</li>\n</ul>\n\n<blockquote>\n  <p>2) I noticed that using the mean vector of an identity for its representation makes a dramatic boost on your prediction. would you please share the motivation of this idea? </p>\n</blockquote>\n\n<ul>\n<li><p>From the <a href=\"https://arxiv.org/pdf/1801.07698.pdf\">paper</a>: </p>\n\n<blockquote>\n  <p>To get the embedding features for templates (e.g. IJB-B and IJB-C) or videos (e.g. YTF and iQIYI-VID), we simply calculate the feature centre of all images from the template or all frames from the video.</p>\n</blockquote></li>\n</ul>\n\n<blockquote>\n  <p>3) what's the hyper param of s in your ArcFace metric and how you tune it?</p>\n</blockquote>\n\n<ul>\n<li>I have used s(=65) and m(=0.5) which are the default value of the paper. I have tried several values, but it didn't improve the scores.</li>\n</ul>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 484538,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-06T06:46:35.350000",
          "content": "<p>What Image-size did you use for training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 484846,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-06T15:10:59.593000",
          "content": "<p>The image size was 320x320.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485041,
          "author_name": "Eduardo Rocha de Andrade",
          "author_url": "",
          "post_date": "2019-03-06T21:04:13.177000",
          "content": "<p>&gt; For each identities, the center of all feature vectors was used as final embedding feature.</p>\n\n<p>When you say center, you mean that for each individual you run the net for all images of that individual, L2 normalize it and take the mean? Then during test you compare (dot product) the test image's embedding with each class/individual center?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 485092,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-06T22:54:35.667000",
          "content": "<blockquote>\n  <p>When you say center, you mean that for each individual you run the net for all images of that individual, L2 normalize it and take the mean? </p>\n</blockquote>\n\n<ul>\n<li>After taking the mean, I normalized it again to locate it on the hypersphere.</li>\n</ul>\n\n<blockquote>\n  <p>Then during test you compare (dot product) the test image's embedding with each class/individual center?</p>\n</blockquote>\n\n<ul>\n<li>You're right</li>\n</ul>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 481814,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2019-03-01T22:08:42.293000",
      "content": "<p>Congrats @pudae , great solution!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 481909,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-02T03:00:49.250000",
          "content": "<p>Thanks :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 486636,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-03-09T06:27:13.503000",
          "content": "<p>hi pudae\nGreat solution..\nCould you point which file in repository implementing the Arc loss ? I couldnt find one there</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 502552,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-03-28T18:08:14.410000",
          "content": "<p>hi pudae..\nTo calculate the feature center ,used as class reference for Test  . I used below code. But I got almost same score without. \nDid i undrestand correctly what you meant feature centre ?</p>\n\n<p>```\nfor c in list(set(data2.train_ds.y.items)): # unique classes found in train set\n     l.append(torch.mean(preds1t[torch.nonzero(y1t == c)],dim=0))</p>\n\n<p>trn_centre=torch.cat(l)\ntrn_centre.size() # 5004 * 5004 matrix</p>\n\n<p>for feat in preds1: # test preds of size no of images * no of classes\n        dists = F.cosine_similarity(trn_centre, feat.unsqueeze(0).repeat(5004, 1))\n        predicted_similarity = dists.cuda()#learn.model.head(dists.cuda())\n        sims.append(predicted_similarity.squeeze().detach().cpu())</p>\n\n<h1>generating the preds</h1>\n\n<p>classes = df.Id.unique()\nnew_whale_idx = np.where(classes == 'new_whale')[0][0]\ntop_5s = []\nfor sim in sims:\n    idxs = sim.argsort(descending=True)\n    probs = sim[idxs]\n    top_5 = []\n    for i, p in zip(idxs, probs):\n        if 'new_whale' not in top_5 and p &lt;0.4 and len(top_5) &lt; 5: #615#575 res,.58 63 for dense\n          top_5.append('new_whale')\n        if len(top_5) == 5: break\n        if i == new_whale_idx: continue\n        predicted_class = labels_list[i]\n        if predicted_class not in top_5: top_5.append(predicted_class)\n    top_5s.append(top_5)\n```</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 481805,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2019-03-01T21:55:24.507000",
      "content": "<p>Congratulations on your work! </p>\n\n<blockquote>\n  <p>Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.</p>\n</blockquote>\n\n<ul>\n<li>The FC you mentioned is before the 5004 dimensions normalized FC of the paper, right? If yes how many dimensions did you use?</li>\n</ul>\n\n<blockquote>\n  <p>Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.</p>\n</blockquote>\n\n<ul>\n<li>So now you have ~10008 identities, right? But I didn't understand what you did with the logit. Can you elaborate?</li>\n</ul>\n\n<p>Thank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 481908,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-02T03:00:35.483000",
          "content": "<blockquote>\n  <p>The FC you mentioned is before the 5004 dimensions normalized FC of the paper, right?</p>\n</blockquote>\n\n<p>Yes, you're right.</p>\n\n<blockquote>\n  <p>If yes how many dimensions did you use?</p>\n</blockquote>\n\n<p>the dimension of the feature vector was 512.</p>\n\n<blockquote>\n  <p>So now you have ~10008 identities, right? </p>\n</blockquote>\n\n<p>Yes~</p>\n\n<blockquote>\n  <p>But I didn't understand what you did with the logit. Can you elaborate?</p>\n</blockquote>\n\n<p>Without zero setting, network would be trained such that the distance between flip and non-flip are apart from each others. It seemed to make training difficult in my  case. Setting the logit value of flipped label to zero was helpful in this case.</p>\n\n<p>Thank you!!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 481680,
      "author_name": "old-ufo",
      "author_url": "",
      "post_date": "2019-03-01T17:54:19.727000",
      "content": "<p>Great solution! I have once tried flip-doubling, but first attempt was unsuccessfull, so I gave up :(\nHow much blur-augmentation added?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 481913,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-02T03:05:02.910000",
          "content": "<p>The kernel size for blurring was between 3 to 5.\nThank you~ :)</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 1676889,
      "author_name": "Athar Sayed",
      "author_url": "",
      "post_date": "2022-02-05T11:21:32.423000",
      "content": "<p>Thanks for sharing its amazing how Object Detection approach is used here</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1675197,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-02-04T05:12:02.850000",
      "content": "<p><code>For each identities, the center of all feature vectors was used as final embedding feature.</code><br>\nwhat do you mean by center here? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 482830,
      "author_name": "Artem.Sanakoev",
      "author_url": "",
      "post_date": "2019-03-03T17:54:16.703000",
      "content": "<p>&gt; After including the identities having one image to training set</p>\n\n<p>Did you fin tune the previous model which had smaller number of classes, or trained from scratch with all clases?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 483389,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-04T14:54:20.997000",
          "content": "<p>I trained from scratch.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488516,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2019-03-12T15:39:18.740000",
          "content": "<p>hi pudae...\n1) How fast does Arcface loss converges... I see it converges very slowly . \n2) why do we need to renomalize the features ,we are using bn which i think does normalizes the activations of layer.</p>\n\n<p><code>\nfeatures = self.bn2(features)\nfeatures = F.normalize(features\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 488610,
          "author_name": "Bartek",
          "author_url": "",
          "post_date": "2019-03-12T18:38:21.150000",
          "content": "<p>Please read original paper which should explain you the idea of normalization: <a href=\"https://arxiv.org/abs/1801.07698\">https://arxiv.org/abs/1801.07698</a></p>\n\n<p>Also, loss converge very slow, it's true because the task defined by ArcFace is much harder than classfication. But classification accuracy converge much faster.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 482827,
      "author_name": "Artem.Sanakoev",
      "author_url": "",
      "post_date": "2019-03-03T17:52:49.127000",
      "content": "<p>Nice and concise! Do you plan to upload the source code?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 483390,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-04T14:54:39.533000",
          "content": "<p>After clean up the codes, I will share it. :)</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 482540,
      "author_name": "xftts",
      "author_url": "",
      "post_date": "2019-03-03T07:18:16.263000",
      "content": "<p>Congratulations !\nIt is really amazing to see your center works. I tried to use center to calc similarity but fail to work :(.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 483391,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-04T14:55:09.950000",
          "content": "<p>Thank you!! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 481972,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-03-02T05:53:48.957000",
      "content": "<p>Congrats @pudai81 for a very strong solo finish and thanks for sharing your solution overview.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 481944,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2019-03-02T04:50:01.443000",
      "content": "<blockquote>\n  <blockquote>\n    <p>When I used aligned image, network was trained faster but the score was not improved.</p>\n  </blockquote>\n</blockquote>\n\n<p>have you tried ensembled of aligned and non-aligned images? that is a common trick in face recognition.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 481961,
          "author_name": "pudae",
          "author_url": "",
          "post_date": "2019-03-02T05:29:24.417000",
          "content": "<p>Yes. LB &gt; 0.965 models used aligned and non-aligned image on training and inference. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 956861,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-03T21:00:06.377000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "481676": "## UPDATE: code available on github\n[https://github.com/pudae/kaggle-humpback](https://github.com/pudae/kaggle-humpback)\n\n---\n\nCongrats to all the winners.\nThanks to Kaggle and hosting team for an interesting competition.\n\nHere is my solution summary.\nSolution Summary\n==============\n\nDataset\n---------\n- **Validation set**: randomly sampled 400 identities that has 2 images + 110 new whales (= 400 * 0.276).\n- **training set**: all images except new whales.\n- I doubled up the identities by horizontal flip.\n\nModel\n-------\n**bounding box &amp; landmark**\n\n- I used annotations by [Paul Johnson](https://www.kaggle.com/c/humpback-whale-identification/discussion/78699) and [Radek Osmulski](https://www.kaggle.com/c/humpback-whale-identification/discussion/76281). (Thanks to Paul and Radek. Without your contribution, I couldn't achieve such high score.)\n- I made 5 fold CV and trained 5 models using them.\n- IOU: 0.93\n\n**whale identifier**\n\n- [ArcFace](https://arxiv.org/pdf/1801.07698.pdf) approach is used.\n- Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.\n- densenet121\n- m 0.5 (the default value of the paper)\n- weight decay 0.0005, droupout 0.5\n\nAugmentation\n-----------------\n- average blur, motion blur\n- add, multiply, grayscale\n- scale, translate, shear, rotate\n- align or no-align\n\nTraining\n---------\n- adam optimizer\n- learning rate of 0.00025 -&gt; 0.000125 -&gt; 0.0000625\n\nInference\n-----------\n**getting embedding feature for identity**\n\n- For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.\n- For each identities, the center of all feature vectors was used as final embedding feature.\n\n**getting embedding feature for test image**\n\n- For each images, multiple feature vectors were generate and the center of the feature vectors was used.\n\n**computing similarity**\n\n- The cosine similarity of above two feature vectors was used as the measure of similarity.\n\n**selecting threshold**\n\n- The threshold for new whale was selected so that the proportion of new whale is about 0.276.\n\nThe process to the final method\n========================\nFollowings are the process to the final method.\n\n**without landmark**\n\nAt first, I excluded the identities having only one image and new whales from the training set. For inference, the identity of the most similar image of the training set was used as the predicted identity.\n\n &gt; Public LB: 0.90, Private LB: 0.90 \n\nAfter using the center of all feature vectors in the same identity, I got\n\n&gt; Public LB: 0.942 / Private LB: 0.939\n\nAfter using weight decay 0.0005\n\n&gt; Public LB: 0.946 / Private LB: 0.946\n\nAfter including the identities having one image to training set\n\n&gt; Public LB: 0.963 / Private LB: 0.961\n\n**with landmark**\n\nWhen I used aligned image, network was trained faster but the score was not improved.\n\n&gt; Public LB: 0.962 / Private LB: 0.959\n\nThe bounding boxes and landmarks of some images are very poor and it seems to prevent improving scores. So I also used non-aligned images.\n\n&gt; Public LB: 0.965 / Private LB: 0.961\n\nFinally, I doubled up identities by horizontal flip. Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.\n\n&gt; Public LB: 0.968 ~ 0.971 / Private LB: 0.965 ~ 0.968\n\nCongrats to winners again.\nThanks.",
    "481706": "Thanks for the write-up Pudae. Congrats on the 3rd place finish!\n\nGlad to hear that my keypoints played a part in a top winning solution!",
    "485577": "Congrats @pudae81!\n\n&gt; I made 5 fold CV and trained 5 models using them.\n&gt; For each images, I got multiple feature vector by using 5 bounding boxes and landmarks.\n\nNice approach!\nHow did you use five bounding boxes in training? Randomly use one as augmentation?",
    "482077": "Hi, @pudae81!  Thanks for sharing your solution!\nI have several questions:\n1)    so your solution is a pure classifier, right? I wonder why ArcFace can tackle such class skew  problem. did you used any oversampling or heavy augmentation for the imbalance problem?\n2)   I noticed that using the mean vector of an identity for its representation makes a dramatic boost on your prediction. would  you please share the motivation of this idea? \n3)  what's the hyper param of *s* in your ArcFace metric and how you tune it?",
    "481814": "Congrats @pudae , great solution!",
    "481805": "Congratulations on your work! \n&gt; Following the paper, the layers after last convolution were replaced to flattening -&gt; BN -&gt; dropout -&gt; FC -&gt; BN.\n\n- The FC you mentioned is before the 5004 dimensions normalized FC of the paper, right? If yes how many dimensions did you use?\n\n&gt; Flipped images have different identities but visually very similar. So I set the logit value of flipped to zero to prevent flowing gradient.\n\n- So now you have ~10008 identities, right? But I didn't understand what you did with the logit. Can you elaborate?\n\nThank you!",
    "481680": "Great solution! I have once tried flip-doubling, but first attempt was unsuccessfull, so I gave up :(\nHow much blur-augmentation added?",
    "1676889": "Thanks for sharing its amazing how Object Detection approach is used here",
    "1675197": "`For each identities, the center of all feature vectors was used as final embedding feature.`\nwhat do you mean by center here? ",
    "482830": "&gt; After including the identities having one image to training set\n\nDid you fin tune the previous model which had smaller number of classes, or trained from scratch with all clases?",
    "482827": "Nice and concise! Do you plan to upload the source code?",
    "482540": "Congratulations !\nIt is really amazing to see your center works. I tried to use center to calc similarity but fail to work :(.",
    "481972": "Congrats @pudai81 for a very strong solo finish and thanks for sharing your solution overview.",
    "481944": "&gt;&gt;When I used aligned image, network was trained faster but the score was not improved.\n\nhave you tried ensembled of aligned and non-aligned images? that is a common trick in face recognition.",
    "956861": ""
  }
}