{
  "id": 320298,
  "title": "19th Place - Single Model LB 860 Without Pseudo",
  "url": "/competitions/happy-whale-and-dolphin/discussion/320298",
  "author_name": "Chris Deotte",
  "post_date": "2022-04-21T01:20:16.395000",
  "votes": 71,
  "comment_count": 44,
  "views": 0,
  "content": "<h1>19th Place HappyWhale Solution</h1>\n<p>We're very happy to present our 19th place HappyWhale solution. Our team consists of <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/bolkonsky\" target=\"_blank\">@bolkonsky</a> It was a pleasure to work with and learn from all these Kaggle Grandmasters. Thank you Kaggle, Ted Cheeseman, collaborators, and everyone at Happywhale for a great competition!</p>\n<h1>Summary</h1>\n<ul>\n<li>Large Backbones like <strong>EffNetB7</strong>, Large Image sizes like <strong>768x768</strong></li>\n<li>Two <strong>ArcFace</strong> modules. One for species and one for individual_id both <strong>m=0.19, s=19</strong></li>\n<li>Six datasets; Mpware <strong>fullbody</strong>, Mpware <strong>fin</strong>, Jan fullbody, Jan fin, Phalanx detic, Awsaf yolo</li>\n<li>Infer all six datasets and <strong>Bayesian optimize</strong> weighted average of six embeddings</li>\n<li><strong>Eight fold CV</strong>. Tune <code>new_individual</code> threshold on CV. Best single model <strong>CV 0.865 LB 0.859</strong></li>\n</ul>\n<h1>Datasets</h1>\n<p>We would love to train our models with only the original images and have the model learn to place attention on the dorsal fin and/or fullbody. However the images are very large like 2500x3500. So instead we used six datasets where important parts of the image were cropped and then resized to either 512x512, 640x640, or 768x768.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/box2.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/crop2.png\" alt=\"\"></p>\n<h1>Models</h1>\n<ul>\n<li>Image Sizes 512x512, 640x640, 768x768</li>\n<li>EffNetV1-B5, EffNetV1-B6, EffNetV1-B7</li>\n<li>EffNetV2-L, EffNetV2-XL</li>\n<li>ConvNext-L</li>\n<li>Batchsize = 64</li>\n<li>TensorFlow trained with <strong>Nvidia 8xV100 GPU</strong> and CoLab TPU</li>\n</ul>\n<h1>Train - One Model All Datasets</h1>\n<p>We train one model on all 6 datasets for 20 epochs exponential learning rate decay with 5 epochs warmup. So the model sees every train image 120 times. We also use data augmentation and two ArcFace heads.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/train2.png\" alt=\"\"></p>\n<h1>Infer - Bayesian Optimized Average of Six Embeddings</h1>\n<p>We infer the model on each dataset separately and get six embeddings for each train and test image. Next we use Bayesian optimization on CV score to find the optimal weighted average of these 6 embeddings to use for KNN matching.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/infer2.png\" alt=\"\"></p>\n<h1>Ensemble</h1>\n<p>Our best single model had 8-Fold CV 0.866 and LB 0.859. We trained about a dozen different models and then ensembled the <code>96 = 12 x 8</code> fold models by using a voting ensemble of each fold model's 5 predictions. Our ensemble result was LB 0.868.</p>\n<h1>How To Improve</h1>\n<p>After reading other top teams' winning solutions, we believe that the next step to boost our model's LB would be to pseudo label the test images and then retrain our models using train and pseudo labeled test. I believe this can boost LB as much as <code>+0.010 to +0.020</code> !</p>\n<h1>UPDATE</h1>\n<p>After the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost <strong>+0.009</strong> and the private LB boost <strong>+0.016</strong>. This confirms that pseudo labeling is very powerful in this competition. </p>\n<p>Just adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!</p>",
  "messages": [
    {
      "id": 1762819,
      "postDate": "2022-04-21T01:20:16.397Z",
      "content": "<h1>19th Place HappyWhale Solution</h1>\n<p>We're very happy to present our 19th place HappyWhale solution. Our team consists of <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/mpware\" target=\"_blank\">@mpware</a> <a href=\"https://www.kaggle.com/bolkonsky\" target=\"_blank\">@bolkonsky</a> It was a pleasure to work with and learn from all these Kaggle Grandmasters. Thank you Kaggle, Ted Cheeseman, collaborators, and everyone at Happywhale for a great competition!</p>\n<h1>Summary</h1>\n<ul>\n<li>Large Backbones like <strong>EffNetB7</strong>, Large Image sizes like <strong>768x768</strong></li>\n<li>Two <strong>ArcFace</strong> modules. One for species and one for individual_id both <strong>m=0.19, s=19</strong></li>\n<li>Six datasets; Mpware <strong>fullbody</strong>, Mpware <strong>fin</strong>, Jan fullbody, Jan fin, Phalanx detic, Awsaf yolo</li>\n<li>Infer all six datasets and <strong>Bayesian optimize</strong> weighted average of six embeddings</li>\n<li><strong>Eight fold CV</strong>. Tune <code>new_individual</code> threshold on CV. Best single model <strong>CV 0.865 LB 0.859</strong></li>\n</ul>\n<h1>Datasets</h1>\n<p>We would love to train our models with only the original images and have the model learn to place attention on the dorsal fin and/or fullbody. However the images are very large like 2500x3500. So instead we used six datasets where important parts of the image were cropped and then resized to either 512x512, 640x640, or 768x768.</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/box2.png\" alt=\"\"><br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/crop2.png\" alt=\"\"></p>\n<h1>Models</h1>\n<ul>\n<li>Image Sizes 512x512, 640x640, 768x768</li>\n<li>EffNetV1-B5, EffNetV1-B6, EffNetV1-B7</li>\n<li>EffNetV2-L, EffNetV2-XL</li>\n<li>ConvNext-L</li>\n<li>Batchsize = 64</li>\n<li>TensorFlow trained with <strong>Nvidia 8xV100 GPU</strong> and CoLab TPU</li>\n</ul>\n<h1>Train - One Model All Datasets</h1>\n<p>We train one model on all 6 datasets for 20 epochs exponential learning rate decay with 5 epochs warmup. So the model sees every train image 120 times. We also use data augmentation and two ArcFace heads.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/train2.png\" alt=\"\"></p>\n<h1>Infer - Bayesian Optimized Average of Six Embeddings</h1>\n<p>We infer the model on each dataset separately and get six embeddings for each train and test image. Next we use Bayesian optimization on CV score to find the optimal weighted average of these 6 embeddings to use for KNN matching.<br>\n<img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/infer2.png\" alt=\"\"></p>\n<h1>Ensemble</h1>\n<p>Our best single model had 8-Fold CV 0.866 and LB 0.859. We trained about a dozen different models and then ensembled the <code>96 = 12 x 8</code> fold models by using a voting ensemble of each fold model's 5 predictions. Our ensemble result was LB 0.868.</p>\n<h1>How To Improve</h1>\n<p>After reading other top teams' winning solutions, we believe that the next step to boost our model's LB would be to pseudo label the test images and then retrain our models using train and pseudo labeled test. I believe this can boost LB as much as <code>+0.010 to +0.020</code> !</p>\n<h1>UPDATE</h1>\n<p>After the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost <strong>+0.009</strong> and the private LB boost <strong>+0.016</strong>. This confirms that pseudo labeling is very powerful in this competition. </p>\n<p>Just adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!</p>",
      "rawMarkdown": "# 19th Place HappyWhale Solution\nWe're very happy to present our 19th place HappyWhale solution. Our team consists of @ragnar123 @mpware @bolkonsky It was a pleasure to work with and learn from all these Kaggle Grandmasters. Thank you Kaggle, Ted Cheeseman, collaborators, and everyone at Happywhale for a great competition!\n\n# Summary\n* Large Backbones like **EffNetB7**, Large Image sizes like **768x768**\n* Two **ArcFace** modules. One for species and one for individual_id both **m=0.19, s=19**\n* Six datasets; Mpware **fullbody**, Mpware **fin**, Jan fullbody, Jan fin, Phalanx detic, Awsaf yolo\n* Infer all six datasets and **Bayesian optimize** weighted average of six embeddings\n* **Eight fold CV**. Tune `new_individual` threshold on CV. Best single model **CV 0.865 LB 0.859**\n\n# Datasets\nWe would love to train our models with only the original images and have the model learn to place attention on the dorsal fin and/or fullbody. However the images are very large like 2500x3500. So instead we used six datasets where important parts of the image were cropped and then resized to either 512x512, 640x640, or 768x768.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/box2.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/crop2.png)\n\n# Models\n* Image Sizes 512x512, 640x640, 768x768\n* EffNetV1-B5, EffNetV1-B6, EffNetV1-B7\n* EffNetV2-L, EffNetV2-XL\n* ConvNext-L\n* Batchsize = 64\n* TensorFlow trained with **Nvidia 8xV100 GPU** and CoLab TPU\n\n# Train - One Model All Datasets\nWe train one model on all 6 datasets for 20 epochs exponential learning rate decay with 5 epochs warmup. So the model sees every train image 120 times. We also use data augmentation and two ArcFace heads.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/train2.png)\n\n# Infer - Bayesian Optimized Average of Six Embeddings\nWe infer the model on each dataset separately and get six embeddings for each train and test image. Next we use Bayesian optimization on CV score to find the optimal weighted average of these 6 embeddings to use for KNN matching.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/infer2.png)\n\n# Ensemble\nOur best single model had 8-Fold CV 0.866 and LB 0.859. We trained about a dozen different models and then ensembled the `96 = 12 x 8` fold models by using a voting ensemble of each fold model's 5 predictions. Our ensemble result was LB 0.868.\n\n# How To Improve\nAfter reading other top teams' winning solutions, we believe that the next step to boost our model's LB would be to pseudo label the test images and then retrain our models using train and pseudo labeled test. I believe this can boost LB as much as `+0.010 to +0.020` !\n\n# UPDATE\nAfter the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost **+0.009** and the private LB boost **+0.016**. This confirms that pseudo labeling is very powerful in this competition. \n\nJust adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!",
      "votes": 70
    },
    {
      "id": 1762863,
      "postDate": "2022-04-21T03:22:35.737Z",
      "content": "<p>Can I ask how do you config the 6 datasets training phase? As a beginner, your solution was the most easy to understand and to try after this competition, but I still don't know how to do the dataset part. Thank you for sharing solution.</p>",
      "rawMarkdown": "Can I ask how do you config the 6 datasets training phase? As a beginner, your solution was the most easy to understand and to try after this competition, but I still don't know how to do the dataset part. Thank you for sharing solution.",
      "votes": 3,
      "replies": [
        {
          "id": 1762868,
          "postDate": "2022-04-21T03:46:31.107Z",
          "content": "<p>In popular public notebooks they use 1 dataset like this</p>\n<pre><code>train_files = tf.io.gfile.glob(GCS_PATH + '/happywhale-2022-train*.tfrec')\ntrain_dataset = get_training_dataset(train_files)\n</code></pre>\n<p>To use 2 datasets, just do this</p>\n<pre><code>GCS_PATH1 = '../input/DATASET1'\nGCS_PATH2 = '../input/DATASET2'\ntrain_files1 = list( tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-train*.tfrec') )\ntrain_files2 = list( tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-train*.tfrec') )\n</code></pre>\n<p>Then during training, just do </p>\n<pre><code>train_dataset = get_training_dataset(train_files1+train_files2)\nmodel.fit(train_dataset)\n</code></pre>",
          "rawMarkdown": "In popular public notebooks they use 1 dataset like this\n\n    train_files = tf.io.gfile.glob(GCS_PATH + '/happywhale-2022-train*.tfrec')\n    train_dataset = get_training_dataset(train_files)\n\nTo use 2 datasets, just do this\n\n    GCS_PATH1 = '../input/DATASET1'\n    GCS_PATH2 = '../input/DATASET2'\n    train_files1 = list( tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-train*.tfrec') )\n    train_files2 = list( tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-train*.tfrec') )\n\nThen during training, just do \n\n    train_dataset = get_training_dataset(train_files1+train_files2)\n    model.fit(train_dataset)\n",
          "votes": 4
        },
        {
          "id": 1762871,
          "postDate": "2022-04-21T03:52:23.767Z",
          "content": "<p>Then during inference just do this</p>\n<pre><code>test_files1 = tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-test*.tfrec')\ntest_dataset1 = get_test_dataset(test_files1)\ntest_embeddings1 = embed_model.predict(test_dataset1)\n\ntest_files2 = tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-test*.tfrec')\ntest_dataset2 = get_test_dataset(test_files2)\ntest_embeddings2 = embed_model.predict(test_dataset2)\n</code></pre>\n<p>And then average the embeddings</p>\n<pre><code>w1 = 0.5\nw2 = 0.5\ntest_embeddings = w1*test_embeddings1 + w2*test_embeddings2\n</code></pre>\n<p>And infer the train embeddings the same way.</p>",
          "rawMarkdown": "Then during inference just do this\n\n    test_files1 = tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-test*.tfrec')\n    test_dataset1 = get_test_dataset(test_files1)\n    test_embeddings1 = embed_model.predict(test_dataset1)\n\n    test_files2 = tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-test*.tfrec')\n    test_dataset2 = get_test_dataset(test_files2)\n    test_embeddings2 = embed_model.predict(test_dataset2)\n\nAnd then average the embeddings\n\n    w1 = 0.5\n    w2 = 0.5\n    test_embeddings = w1*test_embeddings1 + w2*test_embeddings2\n\nAnd infer the train embeddings the same way.\n    ",
          "votes": 3
        },
        {
          "id": 1762877,
          "postDate": "2022-04-21T03:57:42.583Z",
          "content": "<p>You can try this with the 4 datasets that Martin posted <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315524\" target=\"_blank\">here</a> and get a great CV LB score!</p>",
          "rawMarkdown": "You can try this with the 4 datasets that Martin posted [here][1] and get a great CV LB score!\n\n[1]: https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315524",
          "votes": 2
        },
        {
          "id": 1762878,
          "postDate": "2022-04-21T03:58:08.617Z",
          "content": "<p>Thanks for your explanation. ❤️</p>",
          "rawMarkdown": "Thanks for your explanation. ❤️",
          "votes": 1
        }
      ]
    },
    {
      "id": 1763094,
      "postDate": "2022-04-21T09:07:48.340Z",
      "content": "<p>This is nice score with single model, data diversity is work well. Congrats on results and thanks for sharing solution <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "rawMarkdown": "This is nice score with single model, data diversity is work well. Congrats on results and thanks for sharing solution @cdeotte ",
      "votes": 4,
      "replies": [
        {
          "id": 1763100,
          "postDate": "2022-04-21T09:14:04.877Z",
          "content": "<p><a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> Thanks KhanhVD. Congratulations to you and your team's strong finish. You were very close to gold. You'll get gold next time!</p>\n<p>Did your team use pseudo labels? I'm curious because it seems that teams that used pseudo labels shook upward and teams that did not use pseudo labels shook down. Do you know why your team shook down from public to private?</p>",
          "rawMarkdown": "@duykhanh99 Thanks KhanhVD. Congratulations to you and your team's strong finish. You were very close to gold. You'll get gold next time!\n\nDid your team use pseudo labels? I'm curious because it seems that teams that used pseudo labels shook upward and teams that did not use pseudo labels shook down. Do you know why your team shook down from public to private?",
          "votes": 1
        }
      ]
    },
    {
      "id": 1768164,
      "postDate": "2022-04-26T03:08:18.093Z",
      "content": "<p>UPDATE: After the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost <strong>+0.009</strong> and the private LB boost <strong>+0.016</strong>. This confirms that pseudo labeling is very powerful in this competition. </p>\n<p>Just adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!</p>",
      "rawMarkdown": "UPDATE: After the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost **+0.009** and the private LB boost **+0.016**. This confirms that pseudo labeling is very powerful in this competition. \n\nJust adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!",
      "votes": 1,
      "replies": [
        {
          "id": 1768556,
          "postDate": "2022-04-26T12:04:25.277Z",
          "content": "<p>Great result <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Thank you for sharing.<br>\nHave you tried different tricks we can described by participants?</p>",
          "rawMarkdown": "Great result @cdeotte! Thank you for sharing.\nHave you tried different tricks we can described by participants?",
          "votes": 1
        },
        {
          "id": 1768637,
          "postDate": "2022-04-26T14:10:56.637Z",
          "content": "<p>No. I just did the simplest technique which is take the 66% most confident test predictions and make datasets with those test images and labels. Then i added all of those test images to every fold of a single model during training. Next I ensemble that new single model with our old ensemble.</p>\n<p>The result is improved LB score. Next I take the top 66% from this new ensemble and repeat the procedure. I did 3 iterations with two of our single models. I think there are tricks to make pseudo more effective but i haven't tried them.</p>",
          "rawMarkdown": "No. I just did the simplest technique which is take the 66% most confident test predictions and make datasets with those test images and labels. Then i added all of those test images to every fold of a single model during training. Next I ensemble that new single model with our old ensemble.\n\nThe result is improved LB score. Next I take the top 66% from this new ensemble and repeat the procedure. I did 3 iterations with two of our single models. I think there are tricks to make pseudo more effective but i haven't tried them.",
          "votes": 2
        }
      ]
    },
    {
      "id": 1766553,
      "postDate": "2022-04-24T16:14:29.243Z",
      "content": "<p>Very nice job! I like that you guys used a lot of data augmentation. Your final ensemble model had a really good score!</p>",
      "rawMarkdown": "Very nice job! I like that you guys used a lot of data augmentation. Your final ensemble model had a really good score!",
      "votes": 1,
      "replies": [
        {
          "id": 1766877,
          "postDate": "2022-04-24T22:35:59.033Z",
          "content": "<p>Thanks Satoshi</p>",
          "rawMarkdown": "Thanks Satoshi",
          "votes": 1
        }
      ]
    },
    {
      "id": 1766195,
      "postDate": "2022-04-24T09:33:14.810Z",
      "content": "<p>Congratulations ! and Thank for sharing!</p>",
      "rawMarkdown": "Congratulations ! and Thank for sharing!",
      "votes": 1
    },
    {
      "id": 1765652,
      "postDate": "2022-04-23T18:21:28.707Z",
      "content": "<p>Congratulations Chris! Glad to see my name in your solution 😁</p>",
      "rawMarkdown": "Congratulations Chris! Glad to see my name in your solution 😁",
      "votes": 1,
      "replies": [
        {
          "id": 1766417,
          "postDate": "2022-04-24T14:39:42.457Z",
          "content": "<p>Thanks Awsaf! Thanks for your YOLO bbox dataset!</p>",
          "rawMarkdown": "Thanks Awsaf! Thanks for your YOLO bbox dataset!",
          "votes": 2
        },
        {
          "id": 1766437,
          "postDate": "2022-04-24T15:03:13.190Z",
          "content": "<p>Glad To Help :D</p>",
          "rawMarkdown": "Glad To Help :D",
          "votes": 1
        }
      ]
    },
    {
      "id": 1763920,
      "postDate": "2022-04-22T00:42:06.833Z",
      "content": "<p>Only Nvidia 8xV100 GPUs 😄<br>\nWell done <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
      "rawMarkdown": "Only Nvidia 8xV100 GPUs 😄\nWell done @cdeotte !",
      "votes": 1
    },
    {
      "id": 1763219,
      "postDate": "2022-04-21T10:58:02.683Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chri0077\" target=\"_blank\">@chri0077</a> congratulations for your great score! Thank you for solution description. As always easy to understand after … but for me hard to invent during competition :)</p>\n<p>I am trying to understand:</p>\n<blockquote>\n  <p>We train one model on all 6 datasets for 20 epochs.</p>\n</blockquote>\n<p>What does exactly mean? </p>",
      "rawMarkdown": "Hi @chri0077 congratulations for your great score! Thank you for solution description. As always easy to understand after ... but for me hard to invent during competition :)\n\nI am trying to understand:\n\n> We train one model on all 6 datasets for 20 epochs.\n\nWhat does exactly mean? \n",
      "votes": 1,
      "replies": [
        {
          "id": 1763477,
          "postDate": "2022-04-21T14:57:34.643Z",
          "content": "<p>Thanks Remek. Congratulations to you and your team.</p>\n<p>Training on 6 datasets just means concatenating the six datasets and then training on all <code>6x50k = 300k</code> images. So each epoch, the model will see 6 copies of every train image (with each one being a different crop bbox resize):</p>\n<pre><code>train_dataset = get_training_dataset(\n    train_files1 + train_files2 + train_files3 \n    + train_files4 + train_files5 + train_files6)\nmodel.fit(train_dataset, epochs=20)\n</code></pre>",
          "rawMarkdown": "Thanks Remek. Congratulations to you and your team.\n\nTraining on 6 datasets just means concatenating the six datasets and then training on all `6x50k = 300k` images. So each epoch, the model will see 6 copies of every train image (with each one being a different crop bbox resize):\n\n    train_dataset = get_training_dataset(\n        train_files1 + train_files2 + train_files3 \n        + train_files4 + train_files5 + train_files6)\n    model.fit(train_dataset, epochs=20)",
          "votes": 1
        },
        {
          "id": 1763526,
          "postDate": "2022-04-21T15:32:41.407Z",
          "content": "<p>Thank you! I put a lot of effort in this competition and counted on higher position but …. unfortunately had really difficulties to find better solution. We have learned a lot but … whatever I did … score was worse :) :) :) This is time I pay for unexperience in competition. </p>\n<p>Thank you for explnataions! :) </p>\n<p>BTW:<br>\nreading code for me is easier then … description :) - thank you posting </p>",
          "rawMarkdown": "Thank you! I put a lot of effort in this competition and counted on higher position but .... unfortunately had really difficulties to find better solution. We have learned a lot but ... whatever I did ... score was worse :) :) :) This is time I pay for unexperience in competition. \n\nThank you for explnataions! :) \n\nBTW:\nreading code for me is easier then ... description :) - thank you posting ",
          "votes": 1
        },
        {
          "id": 1764499,
          "postDate": "2022-04-22T14:33:29.487Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> one more question - sorry for asking and asking but … I am trying to understand all solutiosn to get gold some day :) </p>\n<p>I can not see explanations why did you use two heads - where did you use classification head?</p>",
          "rawMarkdown": "@cdeotte one more question - sorry for asking and asking but ... I am trying to understand all solutiosn to get gold some day :) \n\nI can not see explanations why did you use two heads - where did you use classification head?"
        },
        {
          "id": 1764548,
          "postDate": "2022-04-22T15:22:48.510Z",
          "content": "<p>In general adding additional heads (auxiliary learning) challenges the model to do multiple learning tasks at once and usually makes the first head (the one we care about) smarter. This is a classic trick. I used it in Pet Comp <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301015\" target=\"_blank\">here</a>. </p>\n<p>Furthermore ArcFace loss clusters similar <code>individual_id</code> embeddings together in N-dimensional space (image <a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/226279\" target=\"_blank\">here</a>). So adding a second ArcFace loss on species will move similar species together which can help the model get the <code>individual_ids</code> located better (and improve the quality of the embeddings).</p>\n<p>After adding a second head, our CV LB boosted about <strong>+0.003</strong> if I remember correctly. The TFRecords have both the <code>individual_id</code> and the <code>species</code> for each train image. Here is the model. (This code is taken from one of <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> models):</p>\n<pre><code>def get_model():\n    with strategy.scope():\n        margin1 = ArcMarginProduct(\n            n_classes = N_CLASSES, \n            s = 19, \n            m = 0.19, \n            name='head/arc_margin1', \n            dtype='float32'\n            )\n        margin2 = ArcMarginProduct(\n            n_classes = 26, \n            s = 19, \n            m = 0.19, \n            name='head/arc_margin2', \n            dtype='float32'\n          )\n        inp = tf.keras.layers.Input(shape = (*IMAGE_SIZE, 3), name = 'inp1')\n        label1 = tf.keras.layers.Input(shape = (), name = 'inp2')\n        label2 = tf.keras.layers.Input(shape = (), name = 'inp3')\n        x = efn.EfficientNetB7(weights = 'noisy-student', \n                               include_top = False)(inp)\n        EMBEDDING = tf.keras.layers.GlobalAveragePooling2D()(x)\n        x = tf.keras.layers.BatchNormalization()(EMBEDDING)\n        x1 = margin1([x, label1])\n        x2 = margin2([x, label2])\n        output1 = tf.keras.layers.Softmax(dtype = 'float32')(x1)\n        output2 = tf.keras.layers.Softmax(dtype = 'float32')(x2)\n        model = tf.keras.models.Model(inputs = [inp, label1, label2], \n                                               outputs = [output1, output2])\n        opt = tf.keras.optimizers.Adam(learning_rate = LR)\n        model.compile(\n           optimizer = opt,\n           loss = [tf.keras.losses.SparseCategoricalCrossentropy(), \n                        tf.keras.losses.SparseCategoricalCrossentropy()],\n           metrics=[tf.keras.metrics.SparseTopKCategoricalAccuracy(k = 5)]\n        ) \n        return model\n</code></pre>",
          "rawMarkdown": "In general adding additional heads (auxiliary learning) challenges the model to do multiple learning tasks at once and usually makes the first head (the one we care about) smarter. This is a classic trick. I used it in Pet Comp [here][1]. \n\nFurthermore ArcFace loss clusters similar `individual_id` embeddings together in N-dimensional space (image [here][2]). So adding a second ArcFace loss on species will move similar species together which can help the model get the `individual_ids` located better (and improve the quality of the embeddings).\n\nAfter adding a second head, our CV LB boosted about **+0.003** if I remember correctly. The TFRecords have both the `individual_id` and the `species` for each train image. Here is the model. (This code is taken from one of @ragnar123 models):\n\n    def get_model():\n        with strategy.scope():\n            margin1 = ArcMarginProduct(\n                n_classes = N_CLASSES, \n                s = 19, \n                m = 0.19, \n                name='head/arc_margin1', \n                dtype='float32'\n                )\n            margin2 = ArcMarginProduct(\n                n_classes = 26, \n                s = 19, \n                m = 0.19, \n                name='head/arc_margin2', \n                dtype='float32'\n              )\n            inp = tf.keras.layers.Input(shape = (*IMAGE_SIZE, 3), name = 'inp1')\n            label1 = tf.keras.layers.Input(shape = (), name = 'inp2')\n            label2 = tf.keras.layers.Input(shape = (), name = 'inp3')\n            x = efn.EfficientNetB7(weights = 'noisy-student', \n                                   include_top = False)(inp)\n            EMBEDDING = tf.keras.layers.GlobalAveragePooling2D()(x)\n            x = tf.keras.layers.BatchNormalization()(EMBEDDING)\n            x1 = margin1([x, label1])\n            x2 = margin2([x, label2])\n            output1 = tf.keras.layers.Softmax(dtype = 'float32')(x1)\n            output2 = tf.keras.layers.Softmax(dtype = 'float32')(x2)\n            model = tf.keras.models.Model(inputs = [inp, label1, label2], \n                                                   outputs = [output1, output2])\n            opt = tf.keras.optimizers.Adam(learning_rate = LR)\n            model.compile(\n               optimizer = opt,\n               loss = [tf.keras.losses.SparseCategoricalCrossentropy(), \n                            tf.keras.losses.SparseCategoricalCrossentropy()],\n               metrics=[tf.keras.metrics.SparseTopKCategoricalAccuracy(k = 5)]\n            ) \n            return model\n\n[1]: https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301015\n[2]: https://www.kaggle.com/competitions/shopee-product-matching/discussion/226279",
          "votes": 1
        },
        {
          "id": 1764558,
          "postDate": "2022-04-22T15:26:38.667Z",
          "content": "<p>Big thanks!!! 🙏🙏🙏🙏</p>",
          "rawMarkdown": "Big thanks!!! 🙏🙏🙏🙏",
          "votes": 1
        },
        {
          "id": 1765440,
          "postDate": "2022-04-23T13:53:15.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is any explanation you use for a classifier (for species classification) AfcFace instead of SoftMax? </p>\n<p>Sorry for digging 😂 but since you are sharing knowledge I am going learn as much as possible. </p>",
          "rawMarkdown": "@cdeotte is any explanation you use for a classifier (for species classification) AfcFace instead of SoftMax? \n\nSorry for digging 😂 but since you are sharing knowledge I am going learn as much as possible. ",
          "votes": 1
        },
        {
          "id": 1765546,
          "postDate": "2022-04-23T16:17:31.843Z",
          "content": "<p>Species classification uses both ArcFace and Softmax. To do classification, you need Softmax. Using ArcFace is optional. So the question is, why did we additionally use ArcFace?</p>\n<p>The reason comes from understanding ArcFace explained <a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/226279\" target=\"_blank\">here</a>. We want high quality embeddings. In dolphin/whale matching, we don't use the output of either the <code>individual_id</code> classifier nor the <code>species</code> classifier (we only use embeddings from NN hidden layer). </p>\n<p>Using ArcFace improves the quality of embeddings. It moves the N-dimensional vectors around in N-dimensional space and clusters them together. Using two ArcFace helps move the embeddings into better clusters to help us better find matches.</p>",
          "rawMarkdown": "Species classification uses both ArcFace and Softmax. To do classification, you need Softmax. Using ArcFace is optional. So the question is, why did we additionally use ArcFace?\n\nThe reason comes from understanding ArcFace explained [here][1]. We want high quality embeddings. In dolphin/whale matching, we don't use the output of either the `individual_id` classifier nor the `species` classifier (we only use embeddings from NN hidden layer). \n\nUsing ArcFace improves the quality of embeddings. It moves the N-dimensional vectors around in N-dimensional space and clusters them together. Using two ArcFace helps move the embeddings into better clusters to help us better find matches.\n\n[1]: https://www.kaggle.com/competitions/shopee-product-matching/discussion/226279",
          "votes": 2
        },
        {
          "id": 1765764,
          "postDate": "2022-04-23T20:23:32.087Z",
          "content": "<p>👍👍👍👍 Thank you 💪💪💪</p>",
          "rawMarkdown": "👍👍👍👍 Thank you 💪💪💪",
          "votes": 1
        }
      ]
    },
    {
      "id": 1763073,
      "postDate": "2022-04-21T08:51:47.910Z",
      "content": "<p>Really elegant solution and nice writeup! <br>\nI wonder why did not you tried pseudo labels.</p>\n<p>For us pseudo helped 0.01-0.02 indeed but we had weaker single models.</p>",
      "rawMarkdown": "Really elegant solution and nice writeup! \nI wonder why did not you tried pseudo labels.\n\nFor us pseudo helped 0.01-0.02 indeed but we had weaker single models.",
      "votes": 1,
      "replies": [
        {
          "id": 1763102,
          "postDate": "2022-04-21T09:19:03.570Z",
          "content": "<p>Unfortunately, I thought it would not help and didn't want to waste 24 hours training a new model. </p>\n<p>I read solutions from last year's comp where pseudo labels did not make a big difference and we did some experiments with KNN test embeddings. For example <code>KNN.fit(train + confident test)</code> then <code>KNN.predict(valid)</code> and we did not see a boost in CV score.</p>\n<p>But now i know that to get results from ArcFace, you cannot just use test embeddings (from old ArcFace). You need to train new ArcFace on the actual test images with pseudo labels.</p>",
          "rawMarkdown": "Unfortunately, I thought it would not help and didn't want to waste 24 hours training a new model. \n\nI read solutions from last year's comp where pseudo labels did not make a big difference and we did some experiments with KNN test embeddings. For example `KNN.fit(train + confident test)` then `KNN.predict(valid)` and we did not see a boost in CV score.\n\nBut now i know that to get results from ArcFace, you cannot just use test embeddings (from old ArcFace). You need to train new ArcFace on the actual test images with pseudo labels."
        },
        {
          "id": 1763112,
          "postDate": "2022-04-21T09:25:34.940Z",
          "content": "<p><a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> Can you describe the procedure you used to pseudo. I assume you took the X% most confident test predictions. </p>\n<p>What i'm curious about is did you then train a model with these X% test and use that same model to predict those X% test? </p>\n<p>Or did you split the X% test into 5 folds. Then include 4 folds in the training of a new model and use the model to predict the hold out test fold. And repeat this 5 times. Then concatenate the test predictions. (This way all test images are predicted using a model trained on different test pseudo than itself and gives the model a chance to correct its mistakes more).</p>",
          "rawMarkdown": "@gaborfodor Can you describe the procedure you used to pseudo. I assume you took the X% most confident test predictions. \n\nWhat i'm curious about is did you then train a model with these X% test and use that same model to predict those X% test? \n\nOr did you split the X% test into 5 folds. Then include 4 folds in the training of a new model and use the model to predict the hold out test fold. And repeat this 5 times. Then concatenate the test predictions. (This way all test images are predicted using a model trained on different test pseudo than itself and gives the model a chance to correct its mistakes more)."
        },
        {
          "id": 1763157,
          "postDate": "2022-04-21T09:56:30.300Z",
          "content": "<p>We used small fixed validation 1 image for each individual id. Single training images we could not split so we just used them for training. For local validation we trained on 44704 images and validated on the fix holdout set with 6329 images. Before submission we retrained on full training set with same config.</p>\n<p>We noticed that Top1 accuracy is quite close to Top5 MAP. Once we were able to beat the best public notebooks with single models we started ensembling to mine pseudo labels from the test set</p>\n<p>First we added the most confident 10K predictions then image_size++ model_size++ retrain ensemble …<br>\nWe had two more rounds with 15K than 20K additional test labels.<br>\nThey showed boost both on local validations and on LB.</p>\n<p>Having proper cross validation like you did would have been better but with 5+ hours training times I am too impatient for 5 fold CV :)</p>",
          "rawMarkdown": "We used small fixed validation 1 image for each individual id. Single training images we could not split so we just used them for training. For local validation we trained on 44704 images and validated on the fix holdout set with 6329 images. Before submission we retrained on full training set with same config.\n\nWe noticed that Top1 accuracy is quite close to Top5 MAP. Once we were able to beat the best public notebooks with single models we started ensembling to mine pseudo labels from the test set\n\nFirst we added the most confident 10K predictions then image_size++ model_size++ retrain ensemble ...\nWe had two more rounds with 15K than 20K additional test labels.\nThey showed boost both on local validations and on LB.\n\nHaving proper cross validation like you did would have been better but with 5+ hours training times I am too impatient for 5 fold CV :)\n\n\n\n\n",
          "votes": 1
        },
        {
          "id": 1763165,
          "postDate": "2022-04-21T09:59:54.187Z",
          "content": "<p>We used the simplest hackiest solution so once we added pseudo label to the train set it pretty much remained top1 for sure. The rest of the top 5 predictions could be corrected…</p>",
          "rawMarkdown": "We used the simplest hackiest solution so once we added pseudo label to the train set it pretty much remained top1 for sure. The rest of the top 5 predictions could be corrected...",
          "votes": 1
        }
      ]
    },
    {
      "id": 1762946,
      "postDate": "2022-04-21T05:34:08.393Z",
      "content": "<p>did you try concatenation of the 6 embeddings and whats the cv/lb score compared to the weighted average</p>",
      "rawMarkdown": "did you try concatenation of the 6 embeddings and whats the cv/lb score compared to the weighted average",
      "votes": 1,
      "replies": [
        {
          "id": 1762957,
          "postDate": "2022-04-21T05:53:02.500Z",
          "content": "<p>Great idea. We did try concatenation (of the 6 dataset embeddings) and it was worse CV. Perhaps the weighted average has less resultant columns (features) and generalizes better.</p>\n<p>Our final ensemble was 12 models with 8 folds each. And each fold has 6 embeddings. This produced about 576 ( = 12 x 8 x 6) sets of embeddings. Using weighted average we reduced each of the 96 fold models to 1 embedding. After we had a total of 96 embeddings.</p>\n<p>Using these embeddings we computed 5 predictions from each model. After we had 480 ( = 96 x 5) predictions. We used a voting ensemble to reduce these 480 predictions to 5 predictions (for each test image).</p>\n<p>I do believe we could have used more embeddings concatenation and/or embeddings averaging instead of a voting ensemble to achieve a higher ensemble LB score.</p>",
          "rawMarkdown": "Great idea. We did try concatenation (of the 6 dataset embeddings) and it was worse CV. Perhaps the weighted average has less resultant columns (features) and generalizes better.\n\nOur final ensemble was 12 models with 8 folds each. And each fold has 6 embeddings. This produced about 576 ( = 12 x 8 x 6) sets of embeddings. Using weighted average we reduced each of the 96 fold models to 1 embedding. After we had a total of 96 embeddings.\n\nUsing these embeddings we computed 5 predictions from each model. After we had 480 ( = 96 x 5) predictions. We used a voting ensemble to reduce these 480 predictions to 5 predictions (for each test image).\n\nI do believe we could have used more embeddings concatenation and/or embeddings averaging instead of a voting ensemble to achieve a higher ensemble LB score.",
          "votes": 1
        },
        {
          "id": 1762976,
          "postDate": "2022-04-21T06:50:11.160Z",
          "content": "<p>thanks for the reply, any specific reason you used 8 folds instead of 5 or 10, I never heard someone using 8 folds so curious</p>",
          "rawMarkdown": "thanks for the reply, any specific reason you used 8 folds instead of 5 or 10, I never heard someone using 8 folds so curious"
        }
      ]
    },
    {
      "id": 1762918,
      "postDate": "2022-04-21T05:01:06.437Z",
      "content": "<p>Thanks for sharing! Could you elaborate a bit on the Bayesian optimization for getting the weights? I get the sense of the idea but want to understand what was the objective function you were using since the end goal is to KNN.</p>",
      "rawMarkdown": "Thanks for sharing! Could you elaborate a bit on the Bayesian optimization for getting the weights? I get the sense of the idea but want to understand what was the objective function you were using since the end goal is to KNN.",
      "votes": 1,
      "replies": [
        {
          "id": 1762956,
          "postDate": "2022-04-21T05:51:40.087Z",
          "content": "<p>The objective function is the CV score. Inside the Bayesian function, 6 weights are explored, both train and valid average embeddings are created, KNN is performed and CV score is calculated (over a range of thresholds). The result for each set of proposed set of 6 weights is the CV score.</p>\n<p>The Bayesian optimization idea and code are from <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> . Below is an excerpt from his code (modified for 2 datasets for easy reading). Bayesian optimization boosted CV LB about <strong>+0.003</strong> (compared with an equal average of embeddings)</p>\n<pre><code>from bayes_opt import BayesianOptimization\n\npbounds = {'w1': (0.0, 1.0), 'w2': (0.0, 1.0)}\ndef get_optimization(w1, w2):\n    training_embeddings = w1 * tr_emb1 + w2 * tr_emb2\n    validation_embeddings = w1 * val_emb1 + w2 * val_emb2 \n    # Nearest neighbors model\n    model = NearestNeighbors(n_neighbors = 100, \n         metric = 'cosine').fit(training_embeddings)\n    distances, indices = model.kneighbors(validation_embeddings)\n    best_map = 0\n    for thres in np.arange(0.20, 0.40, 0.01):\n        map5 = get_map(n_validation_images, distances, indices,\n             training_individual_id, validation_individual_id, thres)\n        if map5 &gt; best_map:  \n             best_map = map5\n    return best_map\noptimizer = BayesianOptimization(\n    f = get_optimization, pbounds = pbounds)\noptimizer.maximize(\n    init_points = 100, n_iter = 50)\nw1 = optimizer.max['params']['w1']\nw2 = optimizer.max['params']['w2']\n</code></pre>",
          "rawMarkdown": "The objective function is the CV score. Inside the Bayesian function, 6 weights are explored, both train and valid average embeddings are created, KNN is performed and CV score is calculated (over a range of thresholds). The result for each set of proposed set of 6 weights is the CV score.\n\nThe Bayesian optimization idea and code are from @ragnar123 . Below is an excerpt from his code (modified for 2 datasets for easy reading). Bayesian optimization boosted CV LB about **+0.003** (compared with an equal average of embeddings)\n\n\n    from bayes_opt import BayesianOptimization\n\n    pbounds = {'w1': (0.0, 1.0), 'w2': (0.0, 1.0)}\n    def get_optimization(w1, w2):\n        training_embeddings = w1 * tr_emb1 + w2 * tr_emb2\n        validation_embeddings = w1 * val_emb1 + w2 * val_emb2 \n        # Nearest neighbors model\n        model = NearestNeighbors(n_neighbors = 100, \n             metric = 'cosine').fit(training_embeddings)\n        distances, indices = model.kneighbors(validation_embeddings)\n        best_map = 0\n        for thres in np.arange(0.20, 0.40, 0.01):\n            map5 = get_map(n_validation_images, distances, indices,\n                 training_individual_id, validation_individual_id, thres)\n            if map5 > best_map:  \n                 best_map = map5\n        return best_map\n    optimizer = BayesianOptimization(\n        f = get_optimization, pbounds = pbounds)\n    optimizer.maximize(\n        init_points = 100, n_iter = 50)\n    w1 = optimizer.max['params']['w1']\n    w2 = optimizer.max['params']['w2']\n\n",
          "votes": 3
        },
        {
          "id": 1762958,
          "postDate": "2022-04-21T05:56:50.040Z",
          "content": "<p>Helps. Thank you!</p>",
          "rawMarkdown": "Helps. Thank you!",
          "votes": 1
        }
      ]
    },
    {
      "id": 1762893,
      "postDate": "2022-04-21T04:11:15.533Z",
      "content": "<p>Always thanks for sharing after competitions finish.<br>\nPseudo labelling incrased our score around 0.02 with 60% test set.<br>\nThe bayesian optimization seems nice.  I learned a lot from you.</p>",
      "rawMarkdown": "Always thanks for sharing after competitions finish.\nPseudo labelling incrased our score around 0.02 with 60% test set.\nThe bayesian optimization seems nice.  I learned a lot from you.",
      "votes": 1
    },
    {
      "id": 1762881,
      "postDate": "2022-04-21T04:01:13.213Z",
      "content": "<p>Thanks for sharing, great performance without pseudo</p>",
      "rawMarkdown": "Thanks for sharing, great performance without pseudo",
      "votes": 1,
      "replies": [
        {
          "id": 1762882,
          "postDate": "2022-04-21T04:02:14.450Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/liuzhangzhen\" target=\"_blank\">@liuzhangzhen</a> . Great job achieving solo 36th place Silver!</p>",
          "rawMarkdown": "Thanks @liuzhangzhen . Great job achieving solo 36th place Silver!"
        }
      ]
    },
    {
      "id": 1766881,
      "postDate": "2022-04-24T22:39:58.520Z",
      "content": "<p>Thanks for sharing. My solution was practically the same as yours, except that I only used 4 different datasets (an experiment for a fold with TPU in colab was taking 2 and a half days). At least it was good to know my solution was competitive</p>",
      "rawMarkdown": "Thanks for sharing. My solution was practically the same as yours, except that I only used 4 different datasets (an experiment for a fold with TPU in colab was taking 2 and a half days). At least it was good to know my solution was competitive",
      "votes": 2,
      "replies": [
        {
          "id": 1766882,
          "postDate": "2022-04-24T22:46:27.850Z",
          "content": "<p>Great job Robson! Another trick that gave us a big boost was to replace the bbox in the dorsal fin dataset of the 3 species that don't have dorsal fin. For those 3 species we used the full body dataset bbox instead. This kept all 6 datasets clean with high quality meaningful bbox.</p>",
          "rawMarkdown": "Great job Robson! Another trick that gave us a big boost was to replace the bbox in the dorsal fin dataset of the 3 species that don't have dorsal fin. For those 3 species we used the full body dataset bbox instead. This kept all 6 datasets clean with high quality meaningful bbox.",
          "votes": 2
        },
        {
          "id": 1766887,
          "postDate": "2022-04-24T22:58:06.677Z",
          "content": "<p>Its a nice trick. Each new competition I learn to think out the box, what makes things more fun. Thanks for sharing with us</p>",
          "rawMarkdown": "Its a nice trick. Each new competition I learn to think out the box, what makes things more fun. Thanks for sharing with us",
          "votes": 1
        }
      ]
    },
    {
      "id": 1762852,
      "postDate": "2022-04-21T02:56:56.683Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1762863,
      "author_name": "Bao Loc Pham",
      "author_url": "",
      "post_date": "2022-04-21T03:22:35.737000",
      "content": "<p>Can I ask how do you config the 6 datasets training phase? As a beginner, your solution was the most easy to understand and to try after this competition, but I still don't know how to do the dataset part. Thank you for sharing solution.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1762868,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T03:46:31.107000",
          "content": "<p>In popular public notebooks they use 1 dataset like this</p>\n<pre><code>train_files = tf.io.gfile.glob(GCS_PATH + '/happywhale-2022-train*.tfrec')\ntrain_dataset = get_training_dataset(train_files)\n</code></pre>\n<p>To use 2 datasets, just do this</p>\n<pre><code>GCS_PATH1 = '../input/DATASET1'\nGCS_PATH2 = '../input/DATASET2'\ntrain_files1 = list( tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-train*.tfrec') )\ntrain_files2 = list( tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-train*.tfrec') )\n</code></pre>\n<p>Then during training, just do </p>\n<pre><code>train_dataset = get_training_dataset(train_files1+train_files2)\nmodel.fit(train_dataset)\n</code></pre>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1762871,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T03:52:23.767000",
          "content": "<p>Then during inference just do this</p>\n<pre><code>test_files1 = tf.io.gfile.glob(GCS_PATH1 + '/happywhale-2022-test*.tfrec')\ntest_dataset1 = get_test_dataset(test_files1)\ntest_embeddings1 = embed_model.predict(test_dataset1)\n\ntest_files2 = tf.io.gfile.glob(GCS_PATH2 + '/happywhale-2022-test*.tfrec')\ntest_dataset2 = get_test_dataset(test_files2)\ntest_embeddings2 = embed_model.predict(test_dataset2)\n</code></pre>\n<p>And then average the embeddings</p>\n<pre><code>w1 = 0.5\nw2 = 0.5\ntest_embeddings = w1*test_embeddings1 + w2*test_embeddings2\n</code></pre>\n<p>And infer the train embeddings the same way.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1762877,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T03:57:42.583000",
          "content": "<p>You can try this with the 4 datasets that Martin posted <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/315524\" target=\"_blank\">here</a> and get a great CV LB score!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1762878,
          "author_name": "Bao Loc Pham",
          "author_url": "",
          "post_date": "2022-04-21T03:58:08.617000",
          "content": "<p>Thanks for your explanation. ❤️</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1763094,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2022-04-21T09:07:48.340000",
      "content": "<p>This is nice score with single model, data diversity is work well. Congrats on results and thanks for sharing solution <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> </p>",
      "votes": 4,
      "replies": [
        {
          "id": 1763100,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T09:14:04.877000",
          "content": "<p><a href=\"https://www.kaggle.com/duykhanh99\" target=\"_blank\">@duykhanh99</a> Thanks KhanhVD. Congratulations to you and your team's strong finish. You were very close to gold. You'll get gold next time!</p>\n<p>Did your team use pseudo labels? I'm curious because it seems that teams that used pseudo labels shook upward and teams that did not use pseudo labels shook down. Do you know why your team shook down from public to private?</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1768164,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2022-04-26T03:08:18.093000",
      "content": "<p>UPDATE: After the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost <strong>+0.009</strong> and the private LB boost <strong>+0.016</strong>. This confirms that pseudo labeling is very powerful in this competition. </p>\n<p>Just adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1768556,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-04-26T12:04:25.277000",
          "content": "<p>Great result <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>! Thank you for sharing.<br>\nHave you tried different tricks we can described by participants?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1768637,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-26T14:10:56.637000",
          "content": "<p>No. I just did the simplest technique which is take the 66% most confident test predictions and make datasets with those test images and labels. Then i added all of those test images to every fold of a single model during training. Next I ensemble that new single model with our old ensemble.</p>\n<p>The result is improved LB score. Next I take the top 66% from this new ensemble and repeat the procedure. I did 3 iterations with two of our single models. I think there are tricks to make pseudo more effective but i haven't tried them.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1766553,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-24T16:14:29.243000",
      "content": "<p>Very nice job! I like that you guys used a lot of data augmentation. Your final ensemble model had a really good score!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1766877,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-24T22:35:59.033000",
          "content": "<p>Thanks Satoshi</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1766195,
      "author_name": "Making TARS",
      "author_url": "",
      "post_date": "2022-04-24T09:33:14.810000",
      "content": "<p>Congratulations ! and Thank for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1765652,
      "author_name": "Awsaf",
      "author_url": "",
      "post_date": "2022-04-23T18:21:28.707000",
      "content": "<p>Congratulations Chris! Glad to see my name in your solution 😁</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1766417,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-24T14:39:42.457000",
          "content": "<p>Thanks Awsaf! Thanks for your YOLO bbox dataset!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1766437,
          "author_name": "Awsaf",
          "author_url": "",
          "post_date": "2022-04-24T15:03:13.190000",
          "content": "<p>Glad To Help :D</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1763920,
      "author_name": "Alin Cijov",
      "author_url": "",
      "post_date": "2022-04-22T00:42:06.833000",
      "content": "<p>Only Nvidia 8xV100 GPUs 😄<br>\nWell done <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1763219,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-04-21T10:58:02.683000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/chri0077\" target=\"_blank\">@chri0077</a> congratulations for your great score! Thank you for solution description. As always easy to understand after … but for me hard to invent during competition :)</p>\n<p>I am trying to understand:</p>\n<blockquote>\n  <p>We train one model on all 6 datasets for 20 epochs.</p>\n</blockquote>\n<p>What does exactly mean? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1763477,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T14:57:34.643000",
          "content": "<p>Thanks Remek. Congratulations to you and your team.</p>\n<p>Training on 6 datasets just means concatenating the six datasets and then training on all <code>6x50k = 300k</code> images. So each epoch, the model will see 6 copies of every train image (with each one being a different crop bbox resize):</p>\n<pre><code>train_dataset = get_training_dataset(\n    train_files1 + train_files2 + train_files3 \n    + train_files4 + train_files5 + train_files6)\nmodel.fit(train_dataset, epochs=20)\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1763526,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-04-21T15:32:41.407000",
          "content": "<p>Thank you! I put a lot of effort in this competition and counted on higher position but …. unfortunately had really difficulties to find better solution. We have learned a lot but … whatever I did … score was worse :) :) :) This is time I pay for unexperience in competition. </p>\n<p>Thank you for explnataions! :) </p>\n<p>BTW:<br>\nreading code for me is easier then … description :) - thank you posting </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1764499,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-04-22T14:33:29.487000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> one more question - sorry for asking and asking but … I am trying to understand all solutiosn to get gold some day :) </p>\n<p>I can not see explanations why did you use two heads - where did you use classification head?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1764548,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-22T15:22:48.510000",
          "content": "<p>In general adding additional heads (auxiliary learning) challenges the model to do multiple learning tasks at once and usually makes the first head (the one we care about) smarter. This is a classic trick. I used it in Pet Comp <a href=\"https://www.kaggle.com/competitions/petfinder-pawpularity-score/discussion/301015\" target=\"_blank\">here</a>. </p>\n<p>Furthermore ArcFace loss clusters similar <code>individual_id</code> embeddings together in N-dimensional space (image <a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/226279\" target=\"_blank\">here</a>). So adding a second ArcFace loss on species will move similar species together which can help the model get the <code>individual_ids</code> located better (and improve the quality of the embeddings).</p>\n<p>After adding a second head, our CV LB boosted about <strong>+0.003</strong> if I remember correctly. The TFRecords have both the <code>individual_id</code> and the <code>species</code> for each train image. Here is the model. (This code is taken from one of <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> models):</p>\n<pre><code>def get_model():\n    with strategy.scope():\n        margin1 = ArcMarginProduct(\n            n_classes = N_CLASSES, \n            s = 19, \n            m = 0.19, \n            name='head/arc_margin1', \n            dtype='float32'\n            )\n        margin2 = ArcMarginProduct(\n            n_classes = 26, \n            s = 19, \n            m = 0.19, \n            name='head/arc_margin2', \n            dtype='float32'\n          )\n        inp = tf.keras.layers.Input(shape = (*IMAGE_SIZE, 3), name = 'inp1')\n        label1 = tf.keras.layers.Input(shape = (), name = 'inp2')\n        label2 = tf.keras.layers.Input(shape = (), name = 'inp3')\n        x = efn.EfficientNetB7(weights = 'noisy-student', \n                               include_top = False)(inp)\n        EMBEDDING = tf.keras.layers.GlobalAveragePooling2D()(x)\n        x = tf.keras.layers.BatchNormalization()(EMBEDDING)\n        x1 = margin1([x, label1])\n        x2 = margin2([x, label2])\n        output1 = tf.keras.layers.Softmax(dtype = 'float32')(x1)\n        output2 = tf.keras.layers.Softmax(dtype = 'float32')(x2)\n        model = tf.keras.models.Model(inputs = [inp, label1, label2], \n                                               outputs = [output1, output2])\n        opt = tf.keras.optimizers.Adam(learning_rate = LR)\n        model.compile(\n           optimizer = opt,\n           loss = [tf.keras.losses.SparseCategoricalCrossentropy(), \n                        tf.keras.losses.SparseCategoricalCrossentropy()],\n           metrics=[tf.keras.metrics.SparseTopKCategoricalAccuracy(k = 5)]\n        ) \n        return model\n</code></pre>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1764558,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-04-22T15:26:38.667000",
          "content": "<p>Big thanks!!! 🙏🙏🙏🙏</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1765440,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-04-23T13:53:15.460000",
          "content": "<p><a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> is any explanation you use for a classifier (for species classification) AfcFace instead of SoftMax? </p>\n<p>Sorry for digging 😂 but since you are sharing knowledge I am going learn as much as possible. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1765546,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-23T16:17:31.843000",
          "content": "<p>Species classification uses both ArcFace and Softmax. To do classification, you need Softmax. Using ArcFace is optional. So the question is, why did we additionally use ArcFace?</p>\n<p>The reason comes from understanding ArcFace explained <a href=\"https://www.kaggle.com/competitions/shopee-product-matching/discussion/226279\" target=\"_blank\">here</a>. We want high quality embeddings. In dolphin/whale matching, we don't use the output of either the <code>individual_id</code> classifier nor the <code>species</code> classifier (we only use embeddings from NN hidden layer). </p>\n<p>Using ArcFace improves the quality of embeddings. It moves the N-dimensional vectors around in N-dimensional space and clusters them together. Using two ArcFace helps move the embeddings into better clusters to help us better find matches.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1765764,
          "author_name": "Remek Kinas",
          "author_url": "",
          "post_date": "2022-04-23T20:23:32.087000",
          "content": "<p>👍👍👍👍 Thank you 💪💪💪</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1763073,
      "author_name": "beluga",
      "author_url": "",
      "post_date": "2022-04-21T08:51:47.910000",
      "content": "<p>Really elegant solution and nice writeup! <br>\nI wonder why did not you tried pseudo labels.</p>\n<p>For us pseudo helped 0.01-0.02 indeed but we had weaker single models.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1763102,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T09:19:03.570000",
          "content": "<p>Unfortunately, I thought it would not help and didn't want to waste 24 hours training a new model. </p>\n<p>I read solutions from last year's comp where pseudo labels did not make a big difference and we did some experiments with KNN test embeddings. For example <code>KNN.fit(train + confident test)</code> then <code>KNN.predict(valid)</code> and we did not see a boost in CV score.</p>\n<p>But now i know that to get results from ArcFace, you cannot just use test embeddings (from old ArcFace). You need to train new ArcFace on the actual test images with pseudo labels.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1763112,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T09:25:34.940000",
          "content": "<p><a href=\"https://www.kaggle.com/gaborfodor\" target=\"_blank\">@gaborfodor</a> Can you describe the procedure you used to pseudo. I assume you took the X% most confident test predictions. </p>\n<p>What i'm curious about is did you then train a model with these X% test and use that same model to predict those X% test? </p>\n<p>Or did you split the X% test into 5 folds. Then include 4 folds in the training of a new model and use the model to predict the hold out test fold. And repeat this 5 times. Then concatenate the test predictions. (This way all test images are predicted using a model trained on different test pseudo than itself and gives the model a chance to correct its mistakes more).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1763157,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2022-04-21T09:56:30.300000",
          "content": "<p>We used small fixed validation 1 image for each individual id. Single training images we could not split so we just used them for training. For local validation we trained on 44704 images and validated on the fix holdout set with 6329 images. Before submission we retrained on full training set with same config.</p>\n<p>We noticed that Top1 accuracy is quite close to Top5 MAP. Once we were able to beat the best public notebooks with single models we started ensembling to mine pseudo labels from the test set</p>\n<p>First we added the most confident 10K predictions then image_size++ model_size++ retrain ensemble …<br>\nWe had two more rounds with 15K than 20K additional test labels.<br>\nThey showed boost both on local validations and on LB.</p>\n<p>Having proper cross validation like you did would have been better but with 5+ hours training times I am too impatient for 5 fold CV :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1763165,
          "author_name": "beluga",
          "author_url": "",
          "post_date": "2022-04-21T09:59:54.187000",
          "content": "<p>We used the simplest hackiest solution so once we added pseudo label to the train set it pretty much remained top1 for sure. The rest of the top 5 predictions could be corrected…</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1762946,
      "author_name": "DeepUnderstanding",
      "author_url": "",
      "post_date": "2022-04-21T05:34:08.393000",
      "content": "<p>did you try concatenation of the 6 embeddings and whats the cv/lb score compared to the weighted average</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1762957,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T05:53:02.500000",
          "content": "<p>Great idea. We did try concatenation (of the 6 dataset embeddings) and it was worse CV. Perhaps the weighted average has less resultant columns (features) and generalizes better.</p>\n<p>Our final ensemble was 12 models with 8 folds each. And each fold has 6 embeddings. This produced about 576 ( = 12 x 8 x 6) sets of embeddings. Using weighted average we reduced each of the 96 fold models to 1 embedding. After we had a total of 96 embeddings.</p>\n<p>Using these embeddings we computed 5 predictions from each model. After we had 480 ( = 96 x 5) predictions. We used a voting ensemble to reduce these 480 predictions to 5 predictions (for each test image).</p>\n<p>I do believe we could have used more embeddings concatenation and/or embeddings averaging instead of a voting ensemble to achieve a higher ensemble LB score.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1762976,
          "author_name": "DeepUnderstanding",
          "author_url": "",
          "post_date": "2022-04-21T06:50:11.160000",
          "content": "<p>thanks for the reply, any specific reason you used 8 folds instead of 5 or 10, I never heard someone using 8 folds so curious</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1762918,
      "author_name": "Shreyansh Singh",
      "author_url": "",
      "post_date": "2022-04-21T05:01:06.437000",
      "content": "<p>Thanks for sharing! Could you elaborate a bit on the Bayesian optimization for getting the weights? I get the sense of the idea but want to understand what was the objective function you were using since the end goal is to KNN.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1762956,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T05:51:40.087000",
          "content": "<p>The objective function is the CV score. Inside the Bayesian function, 6 weights are explored, both train and valid average embeddings are created, KNN is performed and CV score is calculated (over a range of thresholds). The result for each set of proposed set of 6 weights is the CV score.</p>\n<p>The Bayesian optimization idea and code are from <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> . Below is an excerpt from his code (modified for 2 datasets for easy reading). Bayesian optimization boosted CV LB about <strong>+0.003</strong> (compared with an equal average of embeddings)</p>\n<pre><code>from bayes_opt import BayesianOptimization\n\npbounds = {'w1': (0.0, 1.0), 'w2': (0.0, 1.0)}\ndef get_optimization(w1, w2):\n    training_embeddings = w1 * tr_emb1 + w2 * tr_emb2\n    validation_embeddings = w1 * val_emb1 + w2 * val_emb2 \n    # Nearest neighbors model\n    model = NearestNeighbors(n_neighbors = 100, \n         metric = 'cosine').fit(training_embeddings)\n    distances, indices = model.kneighbors(validation_embeddings)\n    best_map = 0\n    for thres in np.arange(0.20, 0.40, 0.01):\n        map5 = get_map(n_validation_images, distances, indices,\n             training_individual_id, validation_individual_id, thres)\n        if map5 &gt; best_map:  \n             best_map = map5\n    return best_map\noptimizer = BayesianOptimization(\n    f = get_optimization, pbounds = pbounds)\noptimizer.maximize(\n    init_points = 100, n_iter = 50)\nw1 = optimizer.max['params']['w1']\nw2 = optimizer.max['params']['w2']\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1762958,
          "author_name": "Shreyansh Singh",
          "author_url": "",
          "post_date": "2022-04-21T05:56:50.040000",
          "content": "<p>Helps. Thank you!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1762893,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2022-04-21T04:11:15.533000",
      "content": "<p>Always thanks for sharing after competitions finish.<br>\nPseudo labelling incrased our score around 0.02 with 60% test set.<br>\nThe bayesian optimization seems nice.  I learned a lot from you.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1762881,
      "author_name": "liuzhangzhen",
      "author_url": "",
      "post_date": "2022-04-21T04:01:13.213000",
      "content": "<p>Thanks for sharing, great performance without pseudo</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1762882,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-21T04:02:14.450000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/liuzhangzhen\" target=\"_blank\">@liuzhangzhen</a> . Great job achieving solo 36th place Silver!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1766881,
      "author_name": "Robson",
      "author_url": "",
      "post_date": "2022-04-24T22:39:58.520000",
      "content": "<p>Thanks for sharing. My solution was practically the same as yours, except that I only used 4 different datasets (an experiment for a fold with TPU in colab was taking 2 and a half days). At least it was good to know my solution was competitive</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1766882,
          "author_name": "Chris Deotte",
          "author_url": "",
          "post_date": "2022-04-24T22:46:27.850000",
          "content": "<p>Great job Robson! Another trick that gave us a big boost was to replace the bbox in the dorsal fin dataset of the 3 species that don't have dorsal fin. For those 3 species we used the full body dataset bbox instead. This kept all 6 datasets clean with high quality meaningful bbox.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1766887,
          "author_name": "Robson",
          "author_url": "",
          "post_date": "2022-04-24T22:58:06.677000",
          "content": "<p>Its a nice trick. Each new competition I learn to think out the box, what makes things more fun. Thanks for sharing with us</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1762852,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-04-21T02:56:56.683000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1762819": "# 19th Place HappyWhale Solution\nWe're very happy to present our 19th place HappyWhale solution. Our team consists of @ragnar123 @mpware @bolkonsky It was a pleasure to work with and learn from all these Kaggle Grandmasters. Thank you Kaggle, Ted Cheeseman, collaborators, and everyone at Happywhale for a great competition!\n\n# Summary\n* Large Backbones like **EffNetB7**, Large Image sizes like **768x768**\n* Two **ArcFace** modules. One for species and one for individual_id both **m=0.19, s=19**\n* Six datasets; Mpware **fullbody**, Mpware **fin**, Jan fullbody, Jan fin, Phalanx detic, Awsaf yolo\n* Infer all six datasets and **Bayesian optimize** weighted average of six embeddings\n* **Eight fold CV**. Tune `new_individual` threshold on CV. Best single model **CV 0.865 LB 0.859**\n\n# Datasets\nWe would love to train our models with only the original images and have the model learn to place attention on the dorsal fin and/or fullbody. However the images are very large like 2500x3500. So instead we used six datasets where important parts of the image were cropped and then resized to either 512x512, 640x640, or 768x768.\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/box2.png)\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/crop2.png)\n\n# Models\n* Image Sizes 512x512, 640x640, 768x768\n* EffNetV1-B5, EffNetV1-B6, EffNetV1-B7\n* EffNetV2-L, EffNetV2-XL\n* ConvNext-L\n* Batchsize = 64\n* TensorFlow trained with **Nvidia 8xV100 GPU** and CoLab TPU\n\n# Train - One Model All Datasets\nWe train one model on all 6 datasets for 20 epochs exponential learning rate decay with 5 epochs warmup. So the model sees every train image 120 times. We also use data augmentation and two ArcFace heads.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/train2.png)\n\n# Infer - Bayesian Optimized Average of Six Embeddings\nWe infer the model on each dataset separately and get six embeddings for each train and test image. Next we use Bayesian optimization on CV score to find the optimal weighted average of these 6 embeddings to use for KNN matching.\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Apr-2022/infer2.png)\n\n# Ensemble\nOur best single model had 8-Fold CV 0.866 and LB 0.859. We trained about a dozen different models and then ensembled the `96 = 12 x 8` fold models by using a voting ensemble of each fold model's 5 predictions. Our ensemble result was LB 0.868.\n\n# How To Improve\nAfter reading other top teams' winning solutions, we believe that the next step to boost our model's LB would be to pseudo label the test images and then retrain our models using train and pseudo labeled test. I believe this can boost LB as much as `+0.010 to +0.020` !\n\n# UPDATE\nAfter the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost **+0.009** and the private LB boost **+0.016**. This confirms that pseudo labeling is very powerful in this competition. \n\nJust adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!",
    "1762863": "Can I ask how do you config the 6 datasets training phase? As a beginner, your solution was the most easy to understand and to try after this competition, but I still don't know how to do the dataset part. Thank you for sharing solution.",
    "1763094": "This is nice score with single model, data diversity is work well. Congrats on results and thanks for sharing solution @cdeotte ",
    "1768164": "UPDATE: After the competition ended, I retrained two of our 12 ensemble models using pseudo labels. The public LB boost **+0.009** and the private LB boost **+0.016**. This confirms that pseudo labeling is very powerful in this competition. \n\nJust adding pseudo to two of our 12 ensemble models boosts our final placement to Gold Medal finish! I assume if we add multiple rounds of pseudo to all our ensemble models, we can climb further into Gold Medal zone!",
    "1766553": "Very nice job! I like that you guys used a lot of data augmentation. Your final ensemble model had a really good score!",
    "1766195": "Congratulations ! and Thank for sharing!",
    "1765652": "Congratulations Chris! Glad to see my name in your solution 😁",
    "1763920": "Only Nvidia 8xV100 GPUs 😄\nWell done @cdeotte !",
    "1763219": "Hi @chri0077 congratulations for your great score! Thank you for solution description. As always easy to understand after ... but for me hard to invent during competition :)\n\nI am trying to understand:\n\n> We train one model on all 6 datasets for 20 epochs.\n\nWhat does exactly mean? \n",
    "1763073": "Really elegant solution and nice writeup! \nI wonder why did not you tried pseudo labels.\n\nFor us pseudo helped 0.01-0.02 indeed but we had weaker single models.",
    "1762946": "did you try concatenation of the 6 embeddings and whats the cv/lb score compared to the weighted average",
    "1762918": "Thanks for sharing! Could you elaborate a bit on the Bayesian optimization for getting the weights? I get the sense of the idea but want to understand what was the objective function you were using since the end goal is to KNN.",
    "1762893": "Always thanks for sharing after competitions finish.\nPseudo labelling incrased our score around 0.02 with 60% test set.\nThe bayesian optimization seems nice.  I learned a lot from you.",
    "1762881": "Thanks for sharing, great performance without pseudo",
    "1766881": "Thanks for sharing. My solution was practically the same as yours, except that I only used 4 different datasets (an experiment for a fold with TPU in colab was taking 2 and a half days). At least it was good to know my solution was competitive",
    "1762852": ""
  }
}