{
  "id": 319828,
  "title": "18th place solution",
  "url": "/competitions/happy-whale-and-dolphin/writeups/18th-place-solution",
  "author_name": "",
  "post_date": "2022-04-21T07:02:33.743Z",
  "votes": 26,
  "comment_count": 5,
  "views": 0,
  "content": "<h2>Intro</h2>\n<p>During this competition, we have seen several approaches to preprocess the competition's dataset. It includes full body cropping, back fin cropping and its variant with the removed background.<br>\nFull body and back fin cropping approaches are very valuable independently of each other. We have made that conclusion when we have found that the model trained on the back fin dataset was better when trained on the bigger image sizes with optimal image size around 768x768. So there is no way to train models on full body images while keeping backfin size close to 768x768, and it is also unreasonable to get the final predictions using only back fin images because there are body pigmentations and other body features that can be helpful for identification. That’s why you 100% need to train at least two separate models on full body crops and back fin crops with consequent ensemble mechanisms.</p>\n<h2>MLP</h2>\n<p>So, we have trained several classifiers on different datasets and ensembled their embeddings using Multi Layer Perceptron (MLP). MLP was chosen because it has a better potential to get insights from embeddings of the models trained on different dataset sources. Specifically, we concatenated features from different models before feeding them to the MLP. It was trained using ArcFace. The model is very shallow and has dropouts at the beginning (0.15) and at the end (0.50). You can find more details in the attached GitHub repository. In general, MLP gave us roughly +5% to the best performing CNN classifier.</p>\n<h2>CNNs</h2>\n<p>In total, we had 8 CNN classifiers that were used in MLP training.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train set images</th>\n<th>Test set images</th>\n<th>LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv1-b7</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>81.6%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b6</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>79.8%</td>\n</tr>\n<tr>\n<td>eca_nfnet_l2</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>79.3%</td>\n</tr>\n<tr>\n<td>eca_nfnet_l2</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>81.2%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b6</td>\n<td>back fin</td>\n<td>backfin</td>\n<td>77.8%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b6</td>\n<td>mix-dataset</td>\n<td>backfin</td>\n<td>80.0%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b7</td>\n<td>original images</td>\n<td>original images</td>\n<td>66.0%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b7</td>\n<td>full body</td>\n<td>full body</td>\n<td>76.8%</td>\n</tr>\n</tbody>\n</table>\n<h2>Mix-dataset</h2>\n<p>A big improvement came from the thing we called mix-dataset. During the mix-dataset training procedure our data loader was sampling images from different data sources (back fin and full body images), but testing images were always from a single dataset source. Efficientnetv1-b6 trained and tested on only backfin images had 77.8% LB score, but the same model with the same hyperparameters had 80.0% LB score when trained on mix-dataset and tested on backfin images. <br>\nIn my understanding, such an approach introduced more training diversity and prevented models from overfitting. An overfitting apparently took place in our training procedure because most of the time we saw &gt;99% training accuracy with a tiny training loss. </p>\n<h2>Class balancing</h2>\n<p>We had two tricks to improve predictions generated by the MLP model. In the prediction .csv file we had 5 potential match candidates sorted by confidence, from the most confident to the least confident.</p>\n<ol>\n<li>In each row, If we see the first element which was also the first element in more than 10 other rows, and the second element which was never the first in other rows, then we swap them.<br>\n[ind_1, ind_2, ind_3, …] -&gt; [ind_2, ind_1, ind_3, …]</li>\n<li>If the first element in the row is new_individual and the second element in the row is the element which was never the first in other rows, then we swap them.<br>\n[new_ind, ind_1, ind_2, …] -&gt; [ind_1, new_ind, ind_2, …]<br>\nMLP without class balancing had 0.865 LB score and MLP with class balancing had 0.868 LB score.</li>\n</ol>\n<p>github repo: <a href=\"https://github.com/achilleess/happywhale-2022\" target=\"_blank\">https://github.com/achilleess/happywhale-2022</a><br>\nThis code does not include all of the above-mentioned models. You can find here training procedures for MLP, back fin models, and some of the full body models. The rest of the code may be released later.</p>\n<p>Teammates:<br>\n<a href=\"https://www.kaggle.com/neomaoro\" target=\"_blank\">@neomaoro</a> <br>\n<a href=\"https://www.kaggle.com/chihantsai\" target=\"_blank\">@chihantsai</a><br>\n<a href=\"https://www.kaggle.com/atom1231\" target=\"_blank\">@atom1231</a><br>\n<a href=\"https://www.kaggle.com/kzvdar42\" target=\"_blank\">@kzvdar42</a> </p>",
  "messages": [
    {
      "id": "1760097",
      "postDate": "04/19/2022 04:08:47",
      "content": "<h2>Intro</h2>\n<p>During this competition, we have seen several approaches to preprocess the competition's dataset. It includes full body cropping, back fin cropping and its variant with the removed background.<br>\nFull body and back fin cropping approaches are very valuable independently of each other. We have made that conclusion when we have found that the model trained on the back fin dataset was better when trained on the bigger image sizes with optimal image size around 768x768. So there is no way to train models on full body images while keeping backfin size close to 768x768, and it is also unreasonable to get the final predictions using only back fin images because there are body pigmentations and other body features that can be helpful for identification. That’s why you 100% need to train at least two separate models on full body crops and back fin crops with consequent ensemble mechanisms.</p>\n<h2>MLP</h2>\n<p>So, we have trained several classifiers on different datasets and ensembled their embeddings using Multi Layer Perceptron (MLP). MLP was chosen because it has a better potential to get insights from embeddings of the models trained on different dataset sources. Specifically, we concatenated features from different models before feeding them to the MLP. It was trained using ArcFace. The model is very shallow and has dropouts at the beginning (0.15) and at the end (0.50). You can find more details in the attached GitHub repository. In general, MLP gave us roughly +5% to the best performing CNN classifier.</p>\n<h2>CNNs</h2>\n<p>In total, we had 8 CNN classifiers that were used in MLP training.</p>\n<table>\n<thead>\n<tr>\n<th>Model</th>\n<th>Train set images</th>\n<th>Test set images</th>\n<th>LB score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>efficientnetv1-b7</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>81.6%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b6</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>79.8%</td>\n</tr>\n<tr>\n<td>eca_nfnet_l2</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>79.3%</td>\n</tr>\n<tr>\n<td>eca_nfnet_l2</td>\n<td>mix-dataset</td>\n<td>full body</td>\n<td>81.2%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b6</td>\n<td>back fin</td>\n<td>backfin</td>\n<td>77.8%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b6</td>\n<td>mix-dataset</td>\n<td>backfin</td>\n<td>80.0%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b7</td>\n<td>original images</td>\n<td>original images</td>\n<td>66.0%</td>\n</tr>\n<tr>\n<td>efficientnetv1-b7</td>\n<td>full body</td>\n<td>full body</td>\n<td>76.8%</td>\n</tr>\n</tbody>\n</table>\n<h2>Mix-dataset</h2>\n<p>A big improvement came from the thing we called mix-dataset. During the mix-dataset training procedure our data loader was sampling images from different data sources (back fin and full body images), but testing images were always from a single dataset source. Efficientnetv1-b6 trained and tested on only backfin images had 77.8% LB score, but the same model with the same hyperparameters had 80.0% LB score when trained on mix-dataset and tested on backfin images. <br>\nIn my understanding, such an approach introduced more training diversity and prevented models from overfitting. An overfitting apparently took place in our training procedure because most of the time we saw &gt;99% training accuracy with a tiny training loss. </p>\n<h2>Class balancing</h2>\n<p>We had two tricks to improve predictions generated by the MLP model. In the prediction .csv file we had 5 potential match candidates sorted by confidence, from the most confident to the least confident.</p>\n<ol>\n<li>In each row, If we see the first element which was also the first element in more than 10 other rows, and the second element which was never the first in other rows, then we swap them.<br>\n[ind_1, ind_2, ind_3, …] -&gt; [ind_2, ind_1, ind_3, …]</li>\n<li>If the first element in the row is new_individual and the second element in the row is the element which was never the first in other rows, then we swap them.<br>\n[new_ind, ind_1, ind_2, …] -&gt; [ind_1, new_ind, ind_2, …]<br>\nMLP without class balancing had 0.865 LB score and MLP with class balancing had 0.868 LB score.</li>\n</ol>\n<p>github repo: <a href=\"https://github.com/achilleess/happywhale-2022\" target=\"_blank\">https://github.com/achilleess/happywhale-2022</a><br>\nThis code does not include all of the above-mentioned models. You can find here training procedures for MLP, back fin models, and some of the full body models. The rest of the code may be released later.</p>\n<p>Teammates:<br>\n<a href=\"https://www.kaggle.com/neomaoro\" target=\"_blank\">@neomaoro</a> <br>\n<a href=\"https://www.kaggle.com/chihantsai\" target=\"_blank\">@chihantsai</a><br>\n<a href=\"https://www.kaggle.com/atom1231\" target=\"_blank\">@atom1231</a><br>\n<a href=\"https://www.kaggle.com/kzvdar42\" target=\"_blank\">@kzvdar42</a> </p>",
      "rawMarkdown": "## Intro\nDuring this competition, we have seen several approaches to preprocess the competition's dataset. It includes full body cropping, back fin cropping and its variant with the removed background.\nFull body and back fin cropping approaches are very valuable independently of each other. We have made that conclusion when we have found that the model trained on the back fin dataset was better when trained on the bigger image sizes with optimal image size around 768x768. So there is no way to train models on full body images while keeping backfin size close to 768x768, and it is also unreasonable to get the final predictions using only back fin images because there are body pigmentations and other body features that can be helpful for identification. That’s why you 100% need to train at least two separate models on full body crops and back fin crops with consequent ensemble mechanisms.\n\n## MLP\nSo, we have trained several classifiers on different datasets and ensembled their embeddings using Multi Layer Perceptron (MLP). MLP was chosen because it has a better potential to get insights from embeddings of the models trained on different dataset sources. Specifically, we concatenated features from different models before feeding them to the MLP. It was trained using ArcFace. The model is very shallow and has dropouts at the beginning (0.15) and at the end (0.50). You can find more details in the attached GitHub repository. In general, MLP gave us roughly +5% to the best performing CNN classifier.\n\n## CNNs\nIn total, we had 8 CNN classifiers that were used in MLP training.\n\n| Model | Train set images | Test set images | LB score |\n| --- | --- | --- | --- |\n| efficientnetv1-b7 | mix-dataset | full body | 81.6% |\n| efficientnetv1-b6 | mix-dataset | full body | 79.8% |\n| eca_nfnet_l2 | mix-dataset | full body | 79.3% |\n| eca_nfnet_l2 |  mix-dataset| full body | 81.2% |\n| efficientnetv1-b6 | back fin | backfin | 77.8% |\n| efficientnetv1-b6 | mix-dataset | backfin | 80.0% |\n| efficientnetv1-b7 | original images | original images | 66.0% |\n| efficientnetv1-b7 | full body | full body | 76.8% |\n\n## Mix-dataset\nA big improvement came from the thing we called mix-dataset. During the mix-dataset training procedure our data loader was sampling images from different data sources (back fin and full body images), but testing images were always from a single dataset source. Efficientnetv1-b6 trained and tested on only backfin images had 77.8% LB score, but the same model with the same hyperparameters had 80.0% LB score when trained on mix-dataset and tested on backfin images. \nIn my understanding, such an approach introduced more training diversity and prevented models from overfitting. An overfitting apparently took place in our training procedure because most of the time we saw >99% training accuracy with a tiny training loss. \n\n## Class balancing\nWe had two tricks to improve predictions generated by the MLP model. In the prediction .csv file we had 5 potential match candidates sorted by confidence, from the most confident to the least confident.\n1. In each row, If we see the first element which was also the first element in more than 10 other rows, and the second element which was never the first in other rows, then we swap them.\n [ind_1, ind_2, ind_3, …] -> [ind_2, ind_1, ind_3, …]\n2. If the first element in the row is new_individual and the second element in the row is the element which was never the first in other rows, then we swap them.\n[new_ind, ind_1, ind_2, …] -> [ind_1, new_ind, ind_2, …]\nMLP without class balancing had 0.865 LB score and MLP with class balancing had 0.868 LB score.\n\ngithub repo: [https://github.com/achilleess/happywhale-2022](https://github.com/achilleess/happywhale-2022)\nThis code does not include all of the above-mentioned models. You can find here training procedures for MLP, back fin models, and some of the full body models. The rest of the code may be released later.\n\nTeammates:\n@neomaoro \n@chihantsai\n@atom1231\n@kzvdar42",
      "votes": null
    },
    {
      "id": "1760113",
      "postDate": "04/19/2022 04:31:59",
      "content": "<p>Congrats on your first kaggle (silver ?) medal and thanks for your wonderful work !!!</p>",
      "rawMarkdown": "Congrats on your first kaggle (silver ?) medal and thanks for your wonderful work !!!",
      "votes": null
    },
    {
      "id": "1760146",
      "postDate": "04/19/2022 05:05:59",
      "content": "<p>Congrats! Appreciate great effort you have made to this comp!!</p>",
      "rawMarkdown": "Congrats! Appreciate great effort you have made to this comp!!",
      "votes": null
    },
    {
      "id": "1760919",
      "postDate": "04/19/2022 15:35:11",
      "content": "<p>Great job! I will learn your sharing, especially MLP</p>",
      "rawMarkdown": "Great job! I will learn your sharing, especially MLP",
      "votes": null
    },
    {
      "id": "1761675",
      "postDate": "04/20/2022 04:45:59",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/achilles38\" target=\"_blank\">@achilles38</a> </p>",
      "rawMarkdown": "Great job @achilles38",
      "votes": null
    },
    {
      "id": "1761800",
      "postDate": "04/20/2022 07:24:59",
      "content": "<p>Congrats! I train efficientnetv1-b5 with image size 512 on backfin dataset, also get nearly 80.0% score on public LB. It's amazing that MLP ensembling has such a potential !</p>",
      "rawMarkdown": "Congrats! I train efficientnetv1-b5 with image size 512 on backfin dataset, also get nearly 80.0% score on public LB. It's amazing that MLP ensembling has such a potential !",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1760113,
      "author_name": "atom1231",
      "author_url": "",
      "post_date": "04/19/2022 04:31:59",
      "content": "<p>Congrats on your first kaggle (silver ?) medal and thanks for your wonderful work !!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1760146,
      "author_name": "neomaoro",
      "author_url": "",
      "post_date": "04/19/2022 05:05:59",
      "content": "<p>Congrats! Appreciate great effort you have made to this comp!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1760919,
      "author_name": "autoatom",
      "author_url": "",
      "post_date": "04/19/2022 15:35:11",
      "content": "<p>Great job! I will learn your sharing, especially MLP</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1761675,
      "author_name": "ishanmehta115",
      "author_url": "",
      "post_date": "04/20/2022 04:45:59",
      "content": "<p>Great job <a href=\"https://www.kaggle.com/achilles38\" target=\"_blank\">@achilles38</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1761800,
      "author_name": "jisongxie",
      "author_url": "",
      "post_date": "04/20/2022 07:24:59",
      "content": "<p>Congrats! I train efficientnetv1-b5 with image size 512 on backfin dataset, also get nearly 80.0% score on public LB. It's amazing that MLP ensembling has such a potential !</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1760097": "## Intro\nDuring this competition, we have seen several approaches to preprocess the competition's dataset. It includes full body cropping, back fin cropping and its variant with the removed background.\nFull body and back fin cropping approaches are very valuable independently of each other. We have made that conclusion when we have found that the model trained on the back fin dataset was better when trained on the bigger image sizes with optimal image size around 768x768. So there is no way to train models on full body images while keeping backfin size close to 768x768, and it is also unreasonable to get the final predictions using only back fin images because there are body pigmentations and other body features that can be helpful for identification. That’s why you 100% need to train at least two separate models on full body crops and back fin crops with consequent ensemble mechanisms.\n\n## MLP\nSo, we have trained several classifiers on different datasets and ensembled their embeddings using Multi Layer Perceptron (MLP). MLP was chosen because it has a better potential to get insights from embeddings of the models trained on different dataset sources. Specifically, we concatenated features from different models before feeding them to the MLP. It was trained using ArcFace. The model is very shallow and has dropouts at the beginning (0.15) and at the end (0.50). You can find more details in the attached GitHub repository. In general, MLP gave us roughly +5% to the best performing CNN classifier.\n\n## CNNs\nIn total, we had 8 CNN classifiers that were used in MLP training.\n\n| Model | Train set images | Test set images | LB score |\n| --- | --- | --- | --- |\n| efficientnetv1-b7 | mix-dataset | full body | 81.6% |\n| efficientnetv1-b6 | mix-dataset | full body | 79.8% |\n| eca_nfnet_l2 | mix-dataset | full body | 79.3% |\n| eca_nfnet_l2 |  mix-dataset| full body | 81.2% |\n| efficientnetv1-b6 | back fin | backfin | 77.8% |\n| efficientnetv1-b6 | mix-dataset | backfin | 80.0% |\n| efficientnetv1-b7 | original images | original images | 66.0% |\n| efficientnetv1-b7 | full body | full body | 76.8% |\n\n## Mix-dataset\nA big improvement came from the thing we called mix-dataset. During the mix-dataset training procedure our data loader was sampling images from different data sources (back fin and full body images), but testing images were always from a single dataset source. Efficientnetv1-b6 trained and tested on only backfin images had 77.8% LB score, but the same model with the same hyperparameters had 80.0% LB score when trained on mix-dataset and tested on backfin images. \nIn my understanding, such an approach introduced more training diversity and prevented models from overfitting. An overfitting apparently took place in our training procedure because most of the time we saw >99% training accuracy with a tiny training loss. \n\n## Class balancing\nWe had two tricks to improve predictions generated by the MLP model. In the prediction .csv file we had 5 potential match candidates sorted by confidence, from the most confident to the least confident.\n1. In each row, If we see the first element which was also the first element in more than 10 other rows, and the second element which was never the first in other rows, then we swap them.\n [ind_1, ind_2, ind_3, …] -> [ind_2, ind_1, ind_3, …]\n2. If the first element in the row is new_individual and the second element in the row is the element which was never the first in other rows, then we swap them.\n[new_ind, ind_1, ind_2, …] -> [ind_1, new_ind, ind_2, …]\nMLP without class balancing had 0.865 LB score and MLP with class balancing had 0.868 LB score.\n\ngithub repo: [https://github.com/achilleess/happywhale-2022](https://github.com/achilleess/happywhale-2022)\nThis code does not include all of the above-mentioned models. You can find here training procedures for MLP, back fin models, and some of the full body models. The rest of the code may be released later.\n\nTeammates:\n@neomaoro \n@chihantsai\n@atom1231\n@kzvdar42",
    "1760113": "Congrats on your first kaggle (silver ?) medal and thanks for your wonderful work !!!",
    "1760146": "Congrats! Appreciate great effort you have made to this comp!!",
    "1760919": "Great job! I will learn your sharing, especially MLP",
    "1761675": "Great job @achilles38",
    "1761800": "Congrats! I train efficientnetv1-b5 with image size 512 on backfin dataset, also get nearly 80.0% score on public LB. It's amazing that MLP ensembling has such a potential !"
  },
  "source": "meta"
}