{
  "id": 319916,
  "title": "11th Place Solution",
  "url": "/competitions/happy-whale-and-dolphin/writeups/tereka-ahmet-yu4u-11th-place-solution",
  "author_name": "",
  "post_date": "2022-04-20T12:00:56.327Z",
  "votes": 39,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Thank you for hosts, also team members(  <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>   <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>)<br>\nCompetition is very tough, so we try to work harder every day, but it's very fun for us.<br>\nI write a summary of our team solution</p>\n<h1>Summary</h1>\n<ul>\n<li>Detection: YOLOV5 + WBF</li>\n<li>EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL + Pseudo Labeling</li>\n<li>Rescore Siamese Network</li>\n</ul>\n<h1>Detection Part</h1>\n<h2>Dataset</h2>\n<p>Firstly, We found small whale in images<br>\nso we need to focus on whale for accurate identification.</p>\n<p>I use labelimg for annotating whales. this annotation tool can export yolo format.<br>\n<a href=\"https://github.com/tzutalin/labelImg\" target=\"_blank\">https://github.com/tzutalin/labelImg</a></p>\n<p>Finally, we annotated 5800 images.</p>\n<h2>Training</h2>\n<ul>\n<li>img size: 1280</li>\n<li>YOLOV5x6</li>\n<li>BS8</li>\n<li>SyncBN</li>\n<li>Epoch 20</li>\n<li>6 Fold</li>\n</ul>\n<h2>Inference</h2>\n<p>6fold models + WBF -&gt; Filter top 1box</p>\n<h1>Whale Identity Part</h1>\n<h2>Dataset</h2>\n<ul>\n<li>Detection result</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL</li>\n<li>ArcFace</li>\n<li>Adam</li>\n<li>Pseudo labeling(Multi-step, Threshold)</li>\n<li>CosineAnnealingWarmRestarts</li>\n<li>GPU/TPU</li>\n</ul>\n<p>ensemble using concat(32000dim). a single model is about 0.805.<br>\nI split 100folds for training.</p>\n<h2>Prediction</h2>\n<p>We use cosine similarity for the identification of whales with some post-process.</p>\n<h2>Rescore</h2>\n<p>Before prediction, We use Siamese Network for top20<br>\nWe combine the similarity matrix and Siamese Network Score. it's a huge improvement in our score.<br>\nIt's achieved Public 0.881/Private 0.853</p>\n<h1>we didn't work</h1>\n<ul>\n<li>DoLG</li>\n<li>twice identity using Flip</li>\n<li>Swin/ConvNeXt</li>\n</ul>",
  "messages": [
    {
      "id": "1760569",
      "postDate": "04/19/2022 11:44:39",
      "content": "<p>Thank you for hosts, also team members(  <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a>   <a href=\"https://www.kaggle.com/ren4yu\" target=\"_blank\">@ren4yu</a>)<br>\nCompetition is very tough, so we try to work harder every day, but it's very fun for us.<br>\nI write a summary of our team solution</p>\n<h1>Summary</h1>\n<ul>\n<li>Detection: YOLOV5 + WBF</li>\n<li>EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL + Pseudo Labeling</li>\n<li>Rescore Siamese Network</li>\n</ul>\n<h1>Detection Part</h1>\n<h2>Dataset</h2>\n<p>Firstly, We found small whale in images<br>\nso we need to focus on whale for accurate identification.</p>\n<p>I use labelimg for annotating whales. this annotation tool can export yolo format.<br>\n<a href=\"https://github.com/tzutalin/labelImg\" target=\"_blank\">https://github.com/tzutalin/labelImg</a></p>\n<p>Finally, we annotated 5800 images.</p>\n<h2>Training</h2>\n<ul>\n<li>img size: 1280</li>\n<li>YOLOV5x6</li>\n<li>BS8</li>\n<li>SyncBN</li>\n<li>Epoch 20</li>\n<li>6 Fold</li>\n</ul>\n<h2>Inference</h2>\n<p>6fold models + WBF -&gt; Filter top 1box</p>\n<h1>Whale Identity Part</h1>\n<h2>Dataset</h2>\n<ul>\n<li>Detection result</li>\n</ul>\n<h2>Training</h2>\n<ul>\n<li>EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL</li>\n<li>ArcFace</li>\n<li>Adam</li>\n<li>Pseudo labeling(Multi-step, Threshold)</li>\n<li>CosineAnnealingWarmRestarts</li>\n<li>GPU/TPU</li>\n</ul>\n<p>ensemble using concat(32000dim). a single model is about 0.805.<br>\nI split 100folds for training.</p>\n<h2>Prediction</h2>\n<p>We use cosine similarity for the identification of whales with some post-process.</p>\n<h2>Rescore</h2>\n<p>Before prediction, We use Siamese Network for top20<br>\nWe combine the similarity matrix and Siamese Network Score. it's a huge improvement in our score.<br>\nIt's achieved Public 0.881/Private 0.853</p>\n<h1>we didn't work</h1>\n<ul>\n<li>DoLG</li>\n<li>twice identity using Flip</li>\n<li>Swin/ConvNeXt</li>\n</ul>",
      "rawMarkdown": "Thank you for hosts, also team members(  @aerdem4   @ren4yu)\nCompetition is very tough, so we try to work harder every day, but it's very fun for us.\nI write a summary of our team solution\n\n# Summary\n- Detection: YOLOV5 + WBF\n- EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL + Pseudo Labeling\n- Rescore Siamese Network\n\n# Detection Part\n## Dataset\nFirstly, We found small whale in images\nso we need to focus on whale for accurate identification.\n\nI use labelimg for annotating whales. this annotation tool can export yolo format.\nhttps://github.com/tzutalin/labelImg\n\nFinally, we annotated 5800 images.\n\n## Training\n- img size: 1280\n- YOLOV5x6\n- BS8\n- SyncBN\n- Epoch 20\n- 6 Fold\n\n## Inference\n6fold models + WBF -> Filter top 1box\n\n# Whale Identity Part\n## Dataset\n- Detection result\n\n## Training\n- EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL\n- ArcFace\n- Adam\n- Pseudo labeling(Multi-step, Threshold)\n- CosineAnnealingWarmRestarts\n- GPU/TPU\n \nensemble using concat(32000dim). a single model is about 0.805.\nI split 100folds for training.\n\n## Prediction\nWe use cosine similarity for the identification of whales with some post-process.\n\n## Rescore\nBefore prediction, We use Siamese Network for top20\nWe combine the similarity matrix and Siamese Network Score. it's a huge improvement in our score.\nIt's achieved Public 0.881/Private 0.853\n\n# we didn't work\n- DoLG\n- twice identity using Flip\n- Swin/ConvNeXt",
      "votes": null
    },
    {
      "id": "1760579",
      "postDate": "04/19/2022 11:50:57",
      "content": "<p><code>ensemble using concat(32000dim)</code></p>\n<p>Your computer must be really strong, thank for sharing your solution</p>",
      "rawMarkdown": "```ensemble using concat(32000dim)```\n\nYour computer must be really strong, thank for sharing your solution",
      "votes": null
    },
    {
      "id": "1760583",
      "postDate": "04/19/2022 11:52:58",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> for this summary.<br>\nCan you give a hint about Rescore Siamese Network?<br>\nWhat is it? I have never hear about this awesome technique before (maybe useful links?)</p>",
      "rawMarkdown": "Thank you @tereka for this summary.\nCan you give a hint about Rescore Siamese Network?\nWhat is it? I have never hear about this awesome technique before (maybe useful links?)",
      "votes": null
    },
    {
      "id": "1760644",
      "postDate": "04/19/2022 12:44:14",
      "content": "<p>We have trained a Siamese comparison model on the individual pairs. It basically applies the same backbone on both images and then compares generated local features.</p>",
      "rawMarkdown": "We have trained a Siamese comparison model on the individual pairs. It basically applies the same backbone on both images and then compares generated local features.",
      "votes": null
    },
    {
      "id": "1760653",
      "postDate": "04/19/2022 12:47:24",
      "content": "<p>Thanks for your sharing, and I have some questions.</p>\n<blockquote>\n  <p>I split 100folds for training.</p>\n</blockquote>\n<p>Q1：100 folds, is that writing error?</p>\n<blockquote>\n  <p>ensemble using concat(32000dim)</p>\n</blockquote>\n<p>Q2: You are using different fold embedding concat, or I miss something?</p>\n<blockquote>\n  <p>Before prediction, We use Siamese Network for top20k</p>\n</blockquote>\n<p>Q3: It is interesting， can you explain more on that， thanks.</p>",
      "rawMarkdown": "Thanks for your sharing, and I have some questions.\n\n> I split 100folds for training.\n\nQ1：100 folds, is that writing error?\n\n> ensemble using concat(32000dim)\n\nQ2: You are using different fold embedding concat, or I miss something?\n\n> Before prediction, We use Siamese Network for top20k\n\nQ3: It is interesting， can you explain more on that， thanks.",
      "votes": null
    },
    {
      "id": "1760673",
      "postDate": "04/19/2022 13:05:11",
      "content": "<p>Is the input of  Siamese images or features?</p>",
      "rawMarkdown": "Is the input of  Siamese images or features?",
      "votes": null
    },
    {
      "id": "1760675",
      "postDate": "04/19/2022 13:06:17",
      "content": "<p>Good solution. Have you compared the impact of different embedding sizes on the effect? We use 512.</p>",
      "rawMarkdown": "Good solution. Have you compared the impact of different embedding sizes on the effect? We use 512.",
      "votes": null
    },
    {
      "id": "1760676",
      "postDate": "04/19/2022 13:06:51",
      "content": "<blockquote>\n  <p>100 folds, is that writing error?</p>\n</blockquote>\n<p>No. It's correct<br>\none model use only one fold.</p>\n<p>For Example, </p>\n<p>I trained EfficientNetB5 896(imsize)<br>\nI use Training: 99 number of index, Validation: 1 index fold .</p>\n<p>Next, I trained EfficientNetB6 768(imsize)<br>\nI use Training: 99 number of index, Validation: 2 index fold.</p>\n<blockquote>\n  <p>You are using different fold embedding concat, or I miss something?</p>\n</blockquote>\n<p>All models predict all training/testing</p>\n<blockquote>\n  <p>Q3: It is interesting， can you explain more on that， thanks.</p>\n</blockquote>\n<p>Siamese Network trained pair is the same identity or not.<br>\nWe extracted Top20k similarity using embeddings.  </p>\n<p>Siamise Network ourput is here</p>\n<p>image1.jpg, image2.jpg, 0.99(same confidence)<br>\nimage1.jpg, image3.jpg, 0.94<br>\nimage1.jpg, image4.jpg, 0.12</p>\n<p>we convert it to a similarity matrix</p>\n<p>we sum similarity matrix(embedding) + siamese network matrix</p>",
      "rawMarkdown": "> 100 folds, is that writing error?\n\nNo. It's correct\none model use only one fold.\n\nFor Example, \n\nI trained EfficientNetB5 896(imsize)\nI use Training: 99 number of index, Validation: 1 index fold .\n\nNext, I trained EfficientNetB6 768(imsize)\nI use Training: 99 number of index, Validation: 2 index fold.\n\n> You are using different fold embedding concat, or I miss something?\n\nAll models predict all training/testing\n\n> Q3: It is interesting， can you explain more on that， thanks.\n\nSiamese Network trained pair is the same identity or not.\nWe extracted Top20k similarity using embeddings.  \n\nSiamise Network ourput is here\n\nimage1.jpg, image2.jpg, 0.99(same confidence)\nimage1.jpg, image3.jpg, 0.94\nimage1.jpg, image4.jpg, 0.12\n\nwe convert it to a similarity matrix\n\nwe sum similarity matrix(embedding) + siamese network matrix",
      "votes": null
    },
    {
      "id": "1760678",
      "postDate": "04/19/2022 13:08:13",
      "content": "<p>We mainly use 512. bigger embeddings are better in CV but I cannot independently check LB in competition.</p>",
      "rawMarkdown": "We mainly use 512. bigger embeddings are better in CV but I cannot independently check LB in competition.",
      "votes": null
    },
    {
      "id": "1760709",
      "postDate": "04/19/2022 13:30:46",
      "content": "<p>Siamese model takes the images as input and creates the local feature vectors itself. But we pretrain its weights via ArcFace model. And we finetune the weights on the image pairs suggested by the ArcFace model.</p>",
      "rawMarkdown": "Siamese model takes the images as input and creates the local feature vectors itself. But we pretrain its weights via ArcFace model. And we finetune the weights on the image pairs suggested by the ArcFace model.",
      "votes": null
    },
    {
      "id": "1760714",
      "postDate": "04/19/2022 13:33:21",
      "content": "<p>Thanks for your detailed explanation！</p>",
      "rawMarkdown": "Thanks for your detailed explanation！",
      "votes": null
    },
    {
      "id": "1760823",
      "postDate": "04/19/2022 14:28:24",
      "content": "<p>This is definitely one of my best findings from this competition!<br>\nThank you <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> and <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> for this awesome approach!</p>",
      "rawMarkdown": "This is definitely one of my best findings from this competition!\nThank you @aerdem4 and @tereka for this awesome approach!",
      "votes": null
    },
    {
      "id": "1760855",
      "postDate": "04/19/2022 14:48:34",
      "content": "<p>Dealing with such a large embedding array was problematic for me as well especially with limited hardware resource, but it is not impossible. Personally I saved the embeddings to disk then loaded them from disk using numpy memmap to reduce memory usage. You can also cast the embeddings to float16 to further reduce ram usage. I was able to concat up to about 20000dim with these little tricks using only kaggle notebook (12GB RAM iirc).</p>",
      "rawMarkdown": "Dealing with such a large embedding array was problematic for me as well especially with limited hardware resource, but it is not impossible. Personally I saved the embeddings to disk then loaded them from disk using numpy memmap to reduce memory usage. You can also cast the embeddings to float16 to further reduce ram usage. I was able to concat up to about 20000dim with these little tricks using only kaggle notebook (12GB RAM iirc).",
      "votes": null
    },
    {
      "id": "1760860",
      "postDate": "04/19/2022 14:51:17",
      "content": "<p>I have 128GB of memory, it can calculate once a time.<br>\nbut if you don't have a memory, you can split the data then you can calculate similarity.<br>\nit indicates that you don't need to calculate at the same time.</p>",
      "rawMarkdown": "I have 128GB of memory, it can calculate once a time.\nbut if you don't have a memory, you can split the data then you can calculate similarity.\nit indicates that you don't need to calculate at the same time.",
      "votes": null
    },
    {
      "id": "1761778",
      "postDate": "04/20/2022 06:56:48",
      "content": "<p>Great result! Congratulations. Thank you very much for solution description. </p>\n<p>I will be really happy if you can answer my question:</p>\n<ol>\n<li>SyncBN - (suppose you use multi-GPU training) what was your HW configuration in this competition?</li>\n<li>Why did you use both GPU / TPU? </li>\n<li>What does it mean - ensemble using concat(32000dim)? What do you ensembled?</li>\n<li>\"We use Siamese Network for top20\" - what does it mean? For each id (in test) did you predict 20 candidates and then  use Siamese for pre-final recognition (final recognition as far as I understand was perfomred on cos similarity and siamese)?</li>\n<li>How did you perform Cosine Similarity?</li>\n</ol>\n<p>Is any chance you publish notebook or part of notebook with two parts:<br>\na.  cosine similarity <br>\nb.  Siamese Network</p>",
      "rawMarkdown": "Great result! Congratulations. Thank you very much for solution description. \n\nI will be really happy if you can answer my question:\n\n1. SyncBN - (suppose you use multi-GPU training) what was your HW configuration in this competition?\n2. Why did you use both GPU / TPU? \n3. What does it mean - ensemble using concat(32000dim)? What do you ensembled?\n4. \"We use Siamese Network for top20\" - what does it mean? For each id (in test) did you predict 20 candidates and then  use Siamese for pre-final recognition (final recognition as far as I understand was perfomred on cos similarity and siamese)?\n5. How did you perform Cosine Similarity?\n\nIs any chance you publish notebook or part of notebook with two parts:\na.  cosine similarity \nb.  Siamese Network",
      "votes": null
    },
    {
      "id": "1761956",
      "postDate": "04/20/2022 10:38:13",
      "content": "<p>Thank you! </p>\n<ol>\n<li>RTX3090 x 2</li>\n<li>my team member yu4u used TPU for training/ I used GPU for training.</li>\n<li>we create many models(EfficientNetB5 + ArcFace etc…), about 20. we extracted vectors from each model, and finally, I concatenated all</li>\n<li>I wrote in detail. I extracted top20 using cosine similarity(32000dims).<br>\nbecause we cannot calculate all pairs using Siamese Network<br>\nThe Final is cosine similarity matrix + Siamese Network prediction(convert pair confidence to matrix)</li>\n<li>I have 128GB memory machine.  I can calculate once<br>\n(it's already applied l2)</li>\n</ol>\n<p>S_all = V_tests_mat@V_trains_mat.T<br>\nnote: numpy float32</p>",
      "rawMarkdown": "Thank you! \n1. RTX3090 x 2\n2. my team member yu4u used TPU for training/ I used GPU for training.\n3. we create many models(EfficientNetB5 + ArcFace etc...), about 20. we extracted vectors from each model, and finally, I concatenated all\n4. I wrote in detail. I extracted top20 using cosine similarity(32000dims).\nbecause we cannot calculate all pairs using Siamese Network\nThe Final is cosine similarity matrix + Siamese Network prediction(convert pair confidence to matrix)\n5.  I have 128GB memory machine.  I can calculate once\n(it's already applied l2)\n\nS_all = V_tests_mat@V_trains_mat.T\nnote: numpy float32",
      "votes": null
    },
    {
      "id": "1762927",
      "postDate": "04/21/2022 05:12:45",
      "content": "<p>nice notebook</p>",
      "rawMarkdown": "nice notebook",
      "votes": null
    },
    {
      "id": "1764164",
      "postDate": "04/22/2022 08:09:24",
      "content": "<p>I tried the same solution as you with 20folds split ,</p>\n<ul>\n<li>EfficientNet B5/B6/B7/V2S/V2M</li>\n<li>ArcFace</li>\n<li>Adam</li>\n<li>Kaggle TPU<br>\nbut my result does not exceed 0.8<br>\nI'm not sure what I missed, I didn't use the label pseudo ?</li>\n</ul>",
      "rawMarkdown": "I tried the same solution as you with 20folds split ,\n- EfficientNet B5/B6/B7/V2S/V2M\n- ArcFace\n- Adam\n- Kaggle TPU\nbut my result does not exceed 0.8\nI'm not sure what I missed, I didn't use the label pseudo ?",
      "votes": null
    },
    {
      "id": "1764214",
      "postDate": "04/22/2022 09:05:06",
      "content": "<p>nice notebook</p>",
      "rawMarkdown": "nice notebook",
      "votes": null
    },
    {
      "id": "1764293",
      "postDate": "04/22/2022 11:08:27",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1764314",
      "postDate": "04/22/2022 11:38:02",
      "content": "<p>To what extent has SiameseNetworkScore improved your score?</p>",
      "rawMarkdown": "To what extent has SiameseNetworkScore improved your score?",
      "votes": null
    },
    {
      "id": "1764407",
      "postDate": "04/22/2022 13:05:57",
      "content": "<p>without pseudo labeling, I reach 0.805.<br>\none of our advantage, yolov5 is better cropping</p>",
      "rawMarkdown": "without pseudo labeling, I reach 0.805.\none of our advantage, yolov5 is better cropping",
      "votes": null
    },
    {
      "id": "1764412",
      "postDate": "04/22/2022 13:10:04",
      "content": "<p>Without(Public/Private)<br>\n0.868<br>\n0.837</p>\n<p>I just add one Siamese Network<br>\n0.874<br>\n0.846</p>\n<p>add some Siamese Network(4models)<br>\n0.879<br>\n0.851<br>\n(Final is + hyperparameter tune)</p>",
      "rawMarkdown": "Without(Public/Private)\n0.868\n0.837\n\nI just add one Siamese Network\n0.874\n0.846\n\nadd some Siamese Network(4models)\n0.879\n0.851\n(Final is + hyperparameter tune)",
      "votes": null
    },
    {
      "id": "1764552",
      "postDate": "04/22/2022 15:24:45",
      "content": "<p>Thanks. That's up over 1%!</p>",
      "rawMarkdown": "Thanks. That's up over 1%!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1760579,
      "author_name": "ptran1203",
      "author_url": "",
      "post_date": "04/19/2022 11:50:57",
      "content": "<p><code>ensemble using concat(32000dim)</code></p>\n<p>Your computer must be really strong, thank for sharing your solution</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760855,
          "author_name": "nanguyen",
          "author_url": "",
          "post_date": "04/19/2022 14:48:34",
          "content": "<p>Dealing with such a large embedding array was problematic for me as well especially with limited hardware resource, but it is not impossible. Personally I saved the embeddings to disk then loaded them from disk using numpy memmap to reduce memory usage. You can also cast the embeddings to float16 to further reduce ram usage. I was able to concat up to about 20000dim with these little tricks using only kaggle notebook (12GB RAM iirc).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760860,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/19/2022 14:51:17",
          "content": "<p>I have 128GB of memory, it can calculate once a time.<br>\nbut if you don't have a memory, you can split the data then you can calculate similarity.<br>\nit indicates that you don't need to calculate at the same time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760583,
      "author_name": "ilyadobrynin",
      "author_url": "",
      "post_date": "04/19/2022 11:52:58",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> for this summary.<br>\nCan you give a hint about Rescore Siamese Network?<br>\nWhat is it? I have never hear about this awesome technique before (maybe useful links?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760644,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "04/19/2022 12:44:14",
          "content": "<p>We have trained a Siamese comparison model on the individual pairs. It basically applies the same backbone on both images and then compares generated local features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760673,
          "author_name": "biglafe",
          "author_url": "",
          "post_date": "04/19/2022 13:05:11",
          "content": "<p>Is the input of  Siamese images or features?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760709,
          "author_name": "aerdem4",
          "author_url": "",
          "post_date": "04/19/2022 13:30:46",
          "content": "<p>Siamese model takes the images as input and creates the local feature vectors itself. But we pretrain its weights via ArcFace model. And we finetune the weights on the image pairs suggested by the ArcFace model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760823,
          "author_name": "ilyadobrynin",
          "author_url": "",
          "post_date": "04/19/2022 14:28:24",
          "content": "<p>This is definitely one of my best findings from this competition!<br>\nThank you <a href=\"https://www.kaggle.com/aerdem4\" target=\"_blank\">@aerdem4</a> and <a href=\"https://www.kaggle.com/tereka\" target=\"_blank\">@tereka</a> for this awesome approach!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760653,
      "author_name": "librauee",
      "author_url": "",
      "post_date": "04/19/2022 12:47:24",
      "content": "<p>Thanks for your sharing, and I have some questions.</p>\n<blockquote>\n  <p>I split 100folds for training.</p>\n</blockquote>\n<p>Q1：100 folds, is that writing error?</p>\n<blockquote>\n  <p>ensemble using concat(32000dim)</p>\n</blockquote>\n<p>Q2: You are using different fold embedding concat, or I miss something?</p>\n<blockquote>\n  <p>Before prediction, We use Siamese Network for top20k</p>\n</blockquote>\n<p>Q3: It is interesting， can you explain more on that， thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760676,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/19/2022 13:06:51",
          "content": "<blockquote>\n  <p>100 folds, is that writing error?</p>\n</blockquote>\n<p>No. It's correct<br>\none model use only one fold.</p>\n<p>For Example, </p>\n<p>I trained EfficientNetB5 896(imsize)<br>\nI use Training: 99 number of index, Validation: 1 index fold .</p>\n<p>Next, I trained EfficientNetB6 768(imsize)<br>\nI use Training: 99 number of index, Validation: 2 index fold.</p>\n<blockquote>\n  <p>You are using different fold embedding concat, or I miss something?</p>\n</blockquote>\n<p>All models predict all training/testing</p>\n<blockquote>\n  <p>Q3: It is interesting， can you explain more on that， thanks.</p>\n</blockquote>\n<p>Siamese Network trained pair is the same identity or not.<br>\nWe extracted Top20k similarity using embeddings.  </p>\n<p>Siamise Network ourput is here</p>\n<p>image1.jpg, image2.jpg, 0.99(same confidence)<br>\nimage1.jpg, image3.jpg, 0.94<br>\nimage1.jpg, image4.jpg, 0.12</p>\n<p>we convert it to a similarity matrix</p>\n<p>we sum similarity matrix(embedding) + siamese network matrix</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760714,
          "author_name": "librauee",
          "author_url": "",
          "post_date": "04/19/2022 13:33:21",
          "content": "<p>Thanks for your detailed explanation！</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1760675,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "04/19/2022 13:06:17",
      "content": "<p>Good solution. Have you compared the impact of different embedding sizes on the effect? We use 512.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760678,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/19/2022 13:08:13",
          "content": "<p>We mainly use 512. bigger embeddings are better in CV but I cannot independently check LB in competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1761778,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/20/2022 06:56:48",
      "content": "<p>Great result! Congratulations. Thank you very much for solution description. </p>\n<p>I will be really happy if you can answer my question:</p>\n<ol>\n<li>SyncBN - (suppose you use multi-GPU training) what was your HW configuration in this competition?</li>\n<li>Why did you use both GPU / TPU? </li>\n<li>What does it mean - ensemble using concat(32000dim)? What do you ensembled?</li>\n<li>\"We use Siamese Network for top20\" - what does it mean? For each id (in test) did you predict 20 candidates and then  use Siamese for pre-final recognition (final recognition as far as I understand was perfomred on cos similarity and siamese)?</li>\n<li>How did you perform Cosine Similarity?</li>\n</ol>\n<p>Is any chance you publish notebook or part of notebook with two parts:<br>\na.  cosine similarity <br>\nb.  Siamese Network</p>",
      "votes": null,
      "replies": [
        {
          "id": 1761956,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/20/2022 10:38:13",
          "content": "<p>Thank you! </p>\n<ol>\n<li>RTX3090 x 2</li>\n<li>my team member yu4u used TPU for training/ I used GPU for training.</li>\n<li>we create many models(EfficientNetB5 + ArcFace etc…), about 20. we extracted vectors from each model, and finally, I concatenated all</li>\n<li>I wrote in detail. I extracted top20 using cosine similarity(32000dims).<br>\nbecause we cannot calculate all pairs using Siamese Network<br>\nThe Final is cosine similarity matrix + Siamese Network prediction(convert pair confidence to matrix)</li>\n<li>I have 128GB memory machine.  I can calculate once<br>\n(it's already applied l2)</li>\n</ol>\n<p>S_all = V_tests_mat@V_trains_mat.T<br>\nnote: numpy float32</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1762927,
      "author_name": "imhungrynow",
      "author_url": "",
      "post_date": "04/21/2022 05:12:45",
      "content": "<p>nice notebook</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1764164,
      "author_name": "nghiahoangtrung",
      "author_url": "",
      "post_date": "04/22/2022 08:09:24",
      "content": "<p>I tried the same solution as you with 20folds split ,</p>\n<ul>\n<li>EfficientNet B5/B6/B7/V2S/V2M</li>\n<li>ArcFace</li>\n<li>Adam</li>\n<li>Kaggle TPU<br>\nbut my result does not exceed 0.8<br>\nI'm not sure what I missed, I didn't use the label pseudo ?</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1764407,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/22/2022 13:05:57",
          "content": "<p>without pseudo labeling, I reach 0.805.<br>\none of our advantage, yolov5 is better cropping</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1764214,
      "author_name": "djaberomarkahlouche",
      "author_url": "",
      "post_date": "04/22/2022 09:05:06",
      "content": "<p>nice notebook</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1764293,
      "author_name": "arujmahajan",
      "author_url": "",
      "post_date": "04/22/2022 11:08:27",
      "content": "<p>Thanks for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1764314,
      "author_name": "yujiariyasu",
      "author_url": "",
      "post_date": "04/22/2022 11:38:02",
      "content": "<p>To what extent has SiameseNetworkScore improved your score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1764412,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "04/22/2022 13:10:04",
          "content": "<p>Without(Public/Private)<br>\n0.868<br>\n0.837</p>\n<p>I just add one Siamese Network<br>\n0.874<br>\n0.846</p>\n<p>add some Siamese Network(4models)<br>\n0.879<br>\n0.851<br>\n(Final is + hyperparameter tune)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764552,
          "author_name": "yujiariyasu",
          "author_url": "",
          "post_date": "04/22/2022 15:24:45",
          "content": "<p>Thanks. That's up over 1%!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1760569": "Thank you for hosts, also team members(  @aerdem4   @ren4yu)\nCompetition is very tough, so we try to work harder every day, but it's very fun for us.\nI write a summary of our team solution\n\n# Summary\n- Detection: YOLOV5 + WBF\n- EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL + Pseudo Labeling\n- Rescore Siamese Network\n\n# Detection Part\n## Dataset\nFirstly, We found small whale in images\nso we need to focus on whale for accurate identification.\n\nI use labelimg for annotating whales. this annotation tool can export yolo format.\nhttps://github.com/tzutalin/labelImg\n\nFinally, we annotated 5800 images.\n\n## Training\n- img size: 1280\n- YOLOV5x6\n- BS8\n- SyncBN\n- Epoch 20\n- 6 Fold\n\n## Inference\n6fold models + WBF -> Filter top 1box\n\n# Whale Identity Part\n## Dataset\n- Detection result\n\n## Training\n- EfficientNet B5/B6/B7/V2S/V2M/V2L/V2XL\n- ArcFace\n- Adam\n- Pseudo labeling(Multi-step, Threshold)\n- CosineAnnealingWarmRestarts\n- GPU/TPU\n \nensemble using concat(32000dim). a single model is about 0.805.\nI split 100folds for training.\n\n## Prediction\nWe use cosine similarity for the identification of whales with some post-process.\n\n## Rescore\nBefore prediction, We use Siamese Network for top20\nWe combine the similarity matrix and Siamese Network Score. it's a huge improvement in our score.\nIt's achieved Public 0.881/Private 0.853\n\n# we didn't work\n- DoLG\n- twice identity using Flip\n- Swin/ConvNeXt",
    "1760579": "```ensemble using concat(32000dim)```\n\nYour computer must be really strong, thank for sharing your solution",
    "1760583": "Thank you @tereka for this summary.\nCan you give a hint about Rescore Siamese Network?\nWhat is it? I have never hear about this awesome technique before (maybe useful links?)",
    "1760644": "We have trained a Siamese comparison model on the individual pairs. It basically applies the same backbone on both images and then compares generated local features.",
    "1760653": "Thanks for your sharing, and I have some questions.\n\n> I split 100folds for training.\n\nQ1：100 folds, is that writing error?\n\n> ensemble using concat(32000dim)\n\nQ2: You are using different fold embedding concat, or I miss something?\n\n> Before prediction, We use Siamese Network for top20k\n\nQ3: It is interesting， can you explain more on that， thanks.",
    "1760673": "Is the input of  Siamese images or features?",
    "1760675": "Good solution. Have you compared the impact of different embedding sizes on the effect? We use 512.",
    "1760676": "> 100 folds, is that writing error?\n\nNo. It's correct\none model use only one fold.\n\nFor Example, \n\nI trained EfficientNetB5 896(imsize)\nI use Training: 99 number of index, Validation: 1 index fold .\n\nNext, I trained EfficientNetB6 768(imsize)\nI use Training: 99 number of index, Validation: 2 index fold.\n\n> You are using different fold embedding concat, or I miss something?\n\nAll models predict all training/testing\n\n> Q3: It is interesting， can you explain more on that， thanks.\n\nSiamese Network trained pair is the same identity or not.\nWe extracted Top20k similarity using embeddings.  \n\nSiamise Network ourput is here\n\nimage1.jpg, image2.jpg, 0.99(same confidence)\nimage1.jpg, image3.jpg, 0.94\nimage1.jpg, image4.jpg, 0.12\n\nwe convert it to a similarity matrix\n\nwe sum similarity matrix(embedding) + siamese network matrix",
    "1760678": "We mainly use 512. bigger embeddings are better in CV but I cannot independently check LB in competition.",
    "1760709": "Siamese model takes the images as input and creates the local feature vectors itself. But we pretrain its weights via ArcFace model. And we finetune the weights on the image pairs suggested by the ArcFace model.",
    "1760714": "Thanks for your detailed explanation！",
    "1760823": "This is definitely one of my best findings from this competition!\nThank you @aerdem4 and @tereka for this awesome approach!",
    "1760855": "Dealing with such a large embedding array was problematic for me as well especially with limited hardware resource, but it is not impossible. Personally I saved the embeddings to disk then loaded them from disk using numpy memmap to reduce memory usage. You can also cast the embeddings to float16 to further reduce ram usage. I was able to concat up to about 20000dim with these little tricks using only kaggle notebook (12GB RAM iirc).",
    "1760860": "I have 128GB of memory, it can calculate once a time.\nbut if you don't have a memory, you can split the data then you can calculate similarity.\nit indicates that you don't need to calculate at the same time.",
    "1761778": "Great result! Congratulations. Thank you very much for solution description. \n\nI will be really happy if you can answer my question:\n\n1. SyncBN - (suppose you use multi-GPU training) what was your HW configuration in this competition?\n2. Why did you use both GPU / TPU? \n3. What does it mean - ensemble using concat(32000dim)? What do you ensembled?\n4. \"We use Siamese Network for top20\" - what does it mean? For each id (in test) did you predict 20 candidates and then  use Siamese for pre-final recognition (final recognition as far as I understand was perfomred on cos similarity and siamese)?\n5. How did you perform Cosine Similarity?\n\nIs any chance you publish notebook or part of notebook with two parts:\na.  cosine similarity \nb.  Siamese Network",
    "1761956": "Thank you! \n1. RTX3090 x 2\n2. my team member yu4u used TPU for training/ I used GPU for training.\n3. we create many models(EfficientNetB5 + ArcFace etc...), about 20. we extracted vectors from each model, and finally, I concatenated all\n4. I wrote in detail. I extracted top20 using cosine similarity(32000dims).\nbecause we cannot calculate all pairs using Siamese Network\nThe Final is cosine similarity matrix + Siamese Network prediction(convert pair confidence to matrix)\n5.  I have 128GB memory machine.  I can calculate once\n(it's already applied l2)\n\nS_all = V_tests_mat@V_trains_mat.T\nnote: numpy float32",
    "1762927": "nice notebook",
    "1764164": "I tried the same solution as you with 20folds split ,\n- EfficientNet B5/B6/B7/V2S/V2M\n- ArcFace\n- Adam\n- Kaggle TPU\nbut my result does not exceed 0.8\nI'm not sure what I missed, I didn't use the label pseudo ?",
    "1764214": "nice notebook",
    "1764293": "Thanks for sharing",
    "1764314": "To what extent has SiameseNetworkScore improved your score?",
    "1764407": "without pseudo labeling, I reach 0.805.\none of our advantage, yolov5 is better cropping",
    "1764412": "Without(Public/Private)\n0.868\n0.837\n\nI just add one Siamese Network\n0.874\n0.846\n\nadd some Siamese Network(4models)\n0.879\n0.851\n(Final is + hyperparameter tune)",
    "1764552": "Thanks. That's up over 1%!"
  },
  "source": "meta"
}