{
  "id": 320502,
  "title": "2nd place solution",
  "url": "/competitions/happy-whale-and-dolphin/writeups/rist-whales-2nd-place-solution",
  "author_name": "",
  "post_date": "2022-04-22T03:51:51.987Z",
  "votes": 48,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Thank you to the competition host, Kaggle and all participants. And congrats to the winners.<br>\nMany thanks to <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>. We used the public fullbody dataset and annotation as our main dataset.  Without it we could not achieve a high score in a short period of time.</p>\n<h1>Single model</h1>\n<h2>Dataset</h2>\n<p>Two different datasets were used. </p>\n<ul>\n<li>fullbody dataset (yolov5): <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184\" target=\"_blank\">this dataset</a> cropped by yolov5. </li>\n<li>fullbody dataset (yolox): the above annotations were used for GT and trained yolox.</li>\n</ul>\n<p>All models were trained using the fullbody dataset (yolov5) before the last day of the contest. Finally, a few models were fine-tuned a few epochs using the fullbody dataset (yolox).</p>\n<h2>Model</h2>\n<p>efficientnet_l2 worked the best in validation</p>\n<h3>Training Recipe</h3>\n<ul>\n<li>backbone = tf_efficientnet_l2_ns</li>\n<li>img_size = 768</li>\n<li>loss = Arcface with adaptive margin</li>\n<li>augmentation = Horizontal flip, RandAugment</li>\n<li>optimizer = SGD</li>\n<li>scheduler = CosineDecay with warmup</li>\n<li>batch_size = 16 per GPU</li>\n<li>n_epoch = 20</li>\n</ul>\n<p>RandAugment improved the validation score quite a bit.</p>\n<h2>Pseudo labeling</h2>\n<p>This is the key to getting a high score on the leaderboard.</p>\n<p>We use FC layer prediction (<code>(logits * scale).softmax(-1)</code>) of trained models to generate pseudo labels. The confidence threshold was set to 0.8.<br>\nEvery pseudo labeling round was trained from imagenet pretrained weights.</p>\n<p>Following are the leaderboard scores of each round. We used flip testing starting round3.</p>\n<p>Pseudo label rounds: backbone = efficientnet_l2</p>\n<table>\n<thead>\n<tr>\n<th>Pseudo label round</th>\n<th>Dataset</th>\n<th>Public score</th>\n<th>Private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>initial model (round1)</td>\n<td>fullbody (yolov5)</td>\n<td>0.846</td>\n<td>0.812</td>\n</tr>\n<tr>\n<td>round2</td>\n<td>fullbody (yolov5)</td>\n<td>0.875</td>\n<td>0.849</td>\n</tr>\n<tr>\n<td>round3</td>\n<td>fullbody (yolov5)</td>\n<td>0.885</td>\n<td>0.860</td>\n</tr>\n<tr>\n<td>round4</td>\n<td>fullbody (yolov5)</td>\n<td>0.889</td>\n<td>0.862</td>\n</tr>\n<tr>\n<td>round5</td>\n<td>fullbody (yolov5)</td>\n<td>0.887</td>\n<td>0.863</td>\n</tr>\n<tr>\n<td>round5 (fine-tuning)</td>\n<td>fullbody (yolox)</td>\n<td>0.891</td>\n<td>0.870</td>\n</tr>\n</tbody>\n</table>\n<h2>Make submission</h2>\n<p>The normal image retrieval method. Calculate cos similarity with train dataset and get top5 ids.</p>\n<p>A fixed cosine similarity of 0.5 was used to insert “new_individual”. There is nothing special post-processing, such as species-specific thresholds.</p>\n<h1>Last ensemble</h1>\n<p>The final submission is a 4-model ensemble of efficientnet_l2 belows. (public=0.897 / private=0.872)</p>\n<p>Ensemble models: all backbone = efficientnet_l2</p>\n<table>\n<thead>\n<tr>\n<th>Pseudo label round</th>\n<th>Image size</th>\n<th>Dataset</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>round3</td>\n<td>(1024, 1024)</td>\n<td>fullbody (yolov5)</td>\n</tr>\n<tr>\n<td>round4</td>\n<td>(768, 768)</td>\n<td>fullbody (yolov5)</td>\n</tr>\n<tr>\n<td>round5</td>\n<td>(768, 768)</td>\n<td>fullbody (yolox) fine-tuned</td>\n</tr>\n<tr>\n<td>round5</td>\n<td>(768x2, 768)</td>\n<td>vertical concatenated original dataset and fullbody (yolox)</td>\n</tr>\n</tbody>\n</table>\n<h1>Performance tips</h1>\n<h2>Gradient checkpointing</h2>\n<p>To train efficientnet_l2 on RTX3090, gradient checkpoint is a must. With gradient checkpointing and mixed precision, we could train the network with batch_size 16 on a single RTX3090. Without it, even batch size 2 gives OOM.</p>\n<h2>PyTorch build</h2>\n<p>We found that different PyTorch builds can greatly affect training throughput. We compared the official release and the build by NVIDIA.</p>\n<table>\n<thead>\n<tr>\n<th>docker image</th>\n<th>cudnn_benchmark</th>\n<th>training throughput (imgs/sec)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel</td>\n<td>False</td>\n<td>2.5</td>\n</tr>\n<tr>\n<td>pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel</td>\n<td>True</td>\n<td>3.4</td>\n</tr>\n<tr>\n<td>nvcr.io/nvidia/pytorch:22.02-py3</td>\n<td>False</td>\n<td>4</td>\n</tr>\n<tr>\n<td>nvcr.io/nvidia/pytorch:22.02-py3</td>\n<td>True</td>\n<td>4.1</td>\n</tr>\n</tbody>\n</table>\n<p>The build from NVIDIA is 20% faster. Probably because of the updated CUDA/CuDNN version.</p>\n<h1>Acknowledge</h1>\n<p>takuoko is a member of Z by HP Data Science Global Ambassadors. Special Thanks to Z by HP for sponsoring me a Z8G4 Workstation with dual A6000 GPU and a ZBook with RTX5000 GPU.</p>",
  "messages": [
    {
      "id": "1763934",
      "postDate": "04/22/2022 01:36:07",
      "content": "<p>Thank you to the competition host, Kaggle and all participants. And congrats to the winners.<br>\nMany thanks to <a href=\"https://www.kaggle.com/jpbremer\" target=\"_blank\">@jpbremer</a>. We used the public fullbody dataset and annotation as our main dataset.  Without it we could not achieve a high score in a short period of time.</p>\n<h1>Single model</h1>\n<h2>Dataset</h2>\n<p>Two different datasets were used. </p>\n<ul>\n<li>fullbody dataset (yolov5): <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184\" target=\"_blank\">this dataset</a> cropped by yolov5. </li>\n<li>fullbody dataset (yolox): the above annotations were used for GT and trained yolox.</li>\n</ul>\n<p>All models were trained using the fullbody dataset (yolov5) before the last day of the contest. Finally, a few models were fine-tuned a few epochs using the fullbody dataset (yolox).</p>\n<h2>Model</h2>\n<p>efficientnet_l2 worked the best in validation</p>\n<h3>Training Recipe</h3>\n<ul>\n<li>backbone = tf_efficientnet_l2_ns</li>\n<li>img_size = 768</li>\n<li>loss = Arcface with adaptive margin</li>\n<li>augmentation = Horizontal flip, RandAugment</li>\n<li>optimizer = SGD</li>\n<li>scheduler = CosineDecay with warmup</li>\n<li>batch_size = 16 per GPU</li>\n<li>n_epoch = 20</li>\n</ul>\n<p>RandAugment improved the validation score quite a bit.</p>\n<h2>Pseudo labeling</h2>\n<p>This is the key to getting a high score on the leaderboard.</p>\n<p>We use FC layer prediction (<code>(logits * scale).softmax(-1)</code>) of trained models to generate pseudo labels. The confidence threshold was set to 0.8.<br>\nEvery pseudo labeling round was trained from imagenet pretrained weights.</p>\n<p>Following are the leaderboard scores of each round. We used flip testing starting round3.</p>\n<p>Pseudo label rounds: backbone = efficientnet_l2</p>\n<table>\n<thead>\n<tr>\n<th>Pseudo label round</th>\n<th>Dataset</th>\n<th>Public score</th>\n<th>Private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>initial model (round1)</td>\n<td>fullbody (yolov5)</td>\n<td>0.846</td>\n<td>0.812</td>\n</tr>\n<tr>\n<td>round2</td>\n<td>fullbody (yolov5)</td>\n<td>0.875</td>\n<td>0.849</td>\n</tr>\n<tr>\n<td>round3</td>\n<td>fullbody (yolov5)</td>\n<td>0.885</td>\n<td>0.860</td>\n</tr>\n<tr>\n<td>round4</td>\n<td>fullbody (yolov5)</td>\n<td>0.889</td>\n<td>0.862</td>\n</tr>\n<tr>\n<td>round5</td>\n<td>fullbody (yolov5)</td>\n<td>0.887</td>\n<td>0.863</td>\n</tr>\n<tr>\n<td>round5 (fine-tuning)</td>\n<td>fullbody (yolox)</td>\n<td>0.891</td>\n<td>0.870</td>\n</tr>\n</tbody>\n</table>\n<h2>Make submission</h2>\n<p>The normal image retrieval method. Calculate cos similarity with train dataset and get top5 ids.</p>\n<p>A fixed cosine similarity of 0.5 was used to insert “new_individual”. There is nothing special post-processing, such as species-specific thresholds.</p>\n<h1>Last ensemble</h1>\n<p>The final submission is a 4-model ensemble of efficientnet_l2 belows. (public=0.897 / private=0.872)</p>\n<p>Ensemble models: all backbone = efficientnet_l2</p>\n<table>\n<thead>\n<tr>\n<th>Pseudo label round</th>\n<th>Image size</th>\n<th>Dataset</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>round3</td>\n<td>(1024, 1024)</td>\n<td>fullbody (yolov5)</td>\n</tr>\n<tr>\n<td>round4</td>\n<td>(768, 768)</td>\n<td>fullbody (yolov5)</td>\n</tr>\n<tr>\n<td>round5</td>\n<td>(768, 768)</td>\n<td>fullbody (yolox) fine-tuned</td>\n</tr>\n<tr>\n<td>round5</td>\n<td>(768x2, 768)</td>\n<td>vertical concatenated original dataset and fullbody (yolox)</td>\n</tr>\n</tbody>\n</table>\n<h1>Performance tips</h1>\n<h2>Gradient checkpointing</h2>\n<p>To train efficientnet_l2 on RTX3090, gradient checkpoint is a must. With gradient checkpointing and mixed precision, we could train the network with batch_size 16 on a single RTX3090. Without it, even batch size 2 gives OOM.</p>\n<h2>PyTorch build</h2>\n<p>We found that different PyTorch builds can greatly affect training throughput. We compared the official release and the build by NVIDIA.</p>\n<table>\n<thead>\n<tr>\n<th>docker image</th>\n<th>cudnn_benchmark</th>\n<th>training throughput (imgs/sec)</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel</td>\n<td>False</td>\n<td>2.5</td>\n</tr>\n<tr>\n<td>pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel</td>\n<td>True</td>\n<td>3.4</td>\n</tr>\n<tr>\n<td>nvcr.io/nvidia/pytorch:22.02-py3</td>\n<td>False</td>\n<td>4</td>\n</tr>\n<tr>\n<td>nvcr.io/nvidia/pytorch:22.02-py3</td>\n<td>True</td>\n<td>4.1</td>\n</tr>\n</tbody>\n</table>\n<p>The build from NVIDIA is 20% faster. Probably because of the updated CUDA/CuDNN version.</p>\n<h1>Acknowledge</h1>\n<p>takuoko is a member of Z by HP Data Science Global Ambassadors. Special Thanks to Z by HP for sponsoring me a Z8G4 Workstation with dual A6000 GPU and a ZBook with RTX5000 GPU.</p>",
      "rawMarkdown": "Thank you to the competition host, Kaggle and all participants. And congrats to the winners.\nMany thanks to @jpbremer. We used the public fullbody dataset and annotation as our main dataset.  Without it we could not achieve a high score in a short period of time.\n\n# Single model\n## Dataset\nTwo different datasets were used. \n- fullbody dataset (yolov5): [this dataset](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184) cropped by yolov5. \n- fullbody dataset (yolox): the above annotations were used for GT and trained yolox.\n\nAll models were trained using the fullbody dataset (yolov5) before the last day of the contest. Finally, a few models were fine-tuned a few epochs using the fullbody dataset (yolox).\n\n## Model\nefficientnet_l2 worked the best in validation\n\n### Training Recipe\n- backbone = tf_efficientnet_l2_ns\n- img_size = 768\n- loss = Arcface with adaptive margin\n- augmentation = Horizontal flip, RandAugment\n- optimizer = SGD\n- scheduler = CosineDecay with warmup\n- batch_size = 16 per GPU\n- n_epoch = 20\n\nRandAugment improved the validation score quite a bit.\n\n## Pseudo labeling\n\nThis is the key to getting a high score on the leaderboard.\n\nWe use FC layer prediction (`(logits * scale).softmax(-1)`) of trained models to generate pseudo labels. The confidence threshold was set to 0.8.\nEvery pseudo labeling round was trained from imagenet pretrained weights.\n\nFollowing are the leaderboard scores of each round. We used flip testing starting round3.\n\nPseudo label rounds: backbone = efficientnet_l2\n| Pseudo label round     | Dataset           | Public score | Private score |\n|------------------------|-------------------|--------------|---------------|\n| initial model (round1) | fullbody (yolov5) | 0.846        | 0.812         |\n| round2                 | fullbody (yolov5) | 0.875        | 0.849         |\n| round3                 | fullbody (yolov5) | 0.885        | 0.860         |\n| round4                 | fullbody (yolov5) | 0.889        | 0.862         |\n| round5                 | fullbody (yolov5) | 0.887        | 0.863         |\n| round5 (fine-tuning)   | fullbody (yolox)  | 0.891        | 0.870         |\n\n## Make submission\nThe normal image retrieval method. Calculate cos similarity with train dataset and get top5 ids.\n\nA fixed cosine similarity of 0.5 was used to insert “new_individual”. There is nothing special post-processing, such as species-specific thresholds.\n\n# Last ensemble\nThe final submission is a 4-model ensemble of efficientnet_l2 belows. (public=0.897 / private=0.872)\n\nEnsemble models: all backbone = efficientnet_l2\n| Pseudo label round | Image size   | Dataset                                                     |\n|--------------------|--------------|-------------------------------------------------------------|\n| round3             | (1024, 1024) | fullbody (yolov5)                                           |\n| round4             | (768, 768)   | fullbody (yolov5)                                           |\n| round5             | (768, 768)   | fullbody (yolox) fine-tuned                                 |\n| round5             | (768x2, 768) | vertical concatenated original dataset and fullbody (yolox) |\n\n\n# Performance tips\n## Gradient checkpointing\n\nTo train efficientnet_l2 on RTX3090, gradient checkpoint is a must. With gradient checkpointing and mixed precision, we could train the network with batch_size 16 on a single RTX3090. Without it, even batch size 2 gives OOM.\n\n## PyTorch build\n\nWe found that different PyTorch builds can greatly affect training throughput. We compared the official release and the build by NVIDIA.\n\n| docker image                                 | cudnn_benchmark   |   training throughput (imgs/sec) |\n|:---------------------------------------------|:------------------|----------------------:|\n| pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel | False             |                   2.5 |\n| pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel | True              |                   3.4 |\n| nvcr.io/nvidia/pytorch:22.02-py3             | False             |                   4   |\n| nvcr.io/nvidia/pytorch:22.02-py3             | True              |                   4.1 |\n\nThe build from NVIDIA is 20% faster. Probably because of the updated CUDA/CuDNN version.\n\n# Acknowledge\ntakuoko is a member of Z by HP Data Science Global Ambassadors. Special Thanks to Z by HP for sponsoring me a Z8G4 Workstation with dual A6000 GPU and a ZBook with RTX5000 GPU.",
      "votes": null
    },
    {
      "id": "1763956",
      "postDate": "04/22/2022 02:26:37",
      "content": "<p>GOOD JOB!!! How much is the difference between your loss in the training set and the loss in the validation?<br>\nMy feeling is to use effb6 to produce the over-fitting, so I use hard enhancement(mixup …) and feature space constraints to alleviate.  </p>",
      "rawMarkdown": "GOOD JOB!!! How much is the difference between your loss in the training set and the loss in the validation?\nMy feeling is to use effb6 to produce the over-fitting, so I use hard enhancement(mixup ...) and feature space constraints to alleviate.",
      "votes": null
    },
    {
      "id": "1764023",
      "postDate": "04/22/2022 04:42:31",
      "content": "<p>Thanks for sharing your approach, very insightful! Congrats on your 2nd place 🎉</p>",
      "rawMarkdown": "Thanks for sharing your approach, very insightful! Congrats on your 2nd place 🎉",
      "votes": null
    },
    {
      "id": "1764237",
      "postDate": "04/22/2022 09:26:55",
      "content": "<p><code>Gradient checkpointing</code><br>\nCan you recommend some information or code for pytorch gradient checkpoint? It sounds amazing!</p>",
      "rawMarkdown": "`Gradient checkpointing`\nCan you recommend some information or code for pytorch gradient checkpoint? It sounds amazing!",
      "votes": null
    },
    {
      "id": "1764328",
      "postDate": "04/22/2022 11:55:35",
      "content": "<p>First of all congratulations! Outstanding result - initial model LB score 0.846 👍👍👍</p>\n<p>May I ask you about explaining:</p>\n<blockquote>\n  <p>RandAugment improved the validation score quite a bit.</p>\n</blockquote>\n<p>How did you implement this? Could you provide part of the code which shows us RandAugment? </p>\n<blockquote>\n  <p>((logits * scale).softmax(-1)) </p>\n</blockquote>\n<p>What is a scale?</p>\n<blockquote>\n  <p>gradient checkpointing</p>\n</blockquote>\n<p>Thank you for providing this tip. Looks interesting.</p>\n<p>As I can see you made your pipeline using Pytorch. There is information about RTX3090. Did you use only one RTX3090 to train all efnet_l2 models?</p>",
      "rawMarkdown": "First of all congratulations! Outstanding result - initial model LB score 0.846 👍👍👍\n\nMay I ask you about explaining:\n\n> RandAugment improved the validation score quite a bit.\n\nHow did you implement this? Could you provide part of the code which shows us RandAugment? \n\n> ((logits * scale).softmax(-1)) \n\nWhat is a scale?\n\n> gradient checkpointing\n\nThank you for providing this tip. Looks interesting.\n\nAs I can see you made your pipeline using Pytorch. There is information about RTX3090. Did you use only one RTX3090 to train all efnet_l2 models?",
      "votes": null
    },
    {
      "id": "1764330",
      "postDate": "04/22/2022 11:58:09",
      "content": "<p>In the latest master branch of timm, gradient checkpointing is available.</p>\n<p><a href=\"https://github.com/rwightman/pytorch-image-models/blob/01a0e25a67305b94ea767083f4113ff002e4435c/timm/models/efficientnet.py#L527-L528\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/01a0e25a67305b94ea767083f4113ff002e4435c/timm/models/efficientnet.py#L527-L528</a></p>",
      "rawMarkdown": "In the latest master branch of timm, gradient checkpointing is available.\n\nhttps://github.com/rwightman/pytorch-image-models/blob/01a0e25a67305b94ea767083f4113ff002e4435c/timm/models/efficientnet.py#L527-L528",
      "votes": null
    },
    {
      "id": "1764350",
      "postDate": "04/22/2022 12:15:17",
      "content": "<blockquote>\n  <p>How did you implement this? Could you provide part of the code which shows us RandAugment?</p>\n</blockquote>\n<p>We used the RandAugment implementation in <a href=\"https://github.com/open-mmlab/mmclassification\" target=\"_blank\">mmclassification</a>. The policies and hyperparamters are copied from imagenet configs in mmclassification.</p>\n<blockquote>\n  <p>What is a scale?</p>\n</blockquote>\n<p>ArcFace scale parameter.</p>\n<blockquote>\n  <p>Did you use only one RTX3090 to train all efnet_l2 models?</p>\n</blockquote>\n<p>I trained the first two rounds of effcientnet_l2 using a single RTX3090. It took 7~8 days. <a href=\"https://www.kaggle.com/takuok\" target=\"_blank\">@takuok</a> used multiple A100 GPUs to train other models since the deadline is approaching.</p>",
      "rawMarkdown": "> How did you implement this? Could you provide part of the code which shows us RandAugment?\n\nWe used the RandAugment implementation in [mmclassification](https://github.com/open-mmlab/mmclassification). The policies and hyperparamters are copied from imagenet configs in mmclassification.\n\n> What is a scale?\n\nArcFace scale parameter.\n\n> Did you use only one RTX3090 to train all efnet_l2 models?\n\nI trained the first two rounds of effcientnet_l2 using a single RTX3090. It took 7~8 days. @takuok used multiple A100 GPUs to train other models since the deadline is approaching.",
      "votes": null
    },
    {
      "id": "1764353",
      "postDate": "04/22/2022 12:19:59",
      "content": "<p>Thank you very much!!!!</p>\n<p>7-8 days 😳😳😳</p>",
      "rawMarkdown": "Thank you very much!!!!\n\n7-8 days 😳😳😳",
      "votes": null
    },
    {
      "id": "1764358",
      "postDate": "04/22/2022 12:29:49",
      "content": "<p>It's in pytorch for years. You can check the DenseNet implementation in torchvision.</p>",
      "rawMarkdown": "It's in pytorch for years. You can check the DenseNet implementation in torchvision.",
      "votes": null
    },
    {
      "id": "1764436",
      "postDate": "04/22/2022 13:25:19",
      "content": "<p>I found really nice post here: <a href=\"https://spell.ml/blog/gradient-checkpointing-pytorch-YGypLBAAACEAefHs\" target=\"_blank\">Training larger-than-memory PyTorch models using gradient checkpointing</a></p>",
      "rawMarkdown": "I found really nice post here: [Training larger-than-memory PyTorch models using gradient checkpointing](https://spell.ml/blog/gradient-checkpointing-pytorch-YGypLBAAACEAefHs)",
      "votes": null
    },
    {
      "id": "1765768",
      "postDate": "04/23/2022 20:26:15",
      "content": "<p>Could you share implementation of  Arcface with adaptive margin?  It would be helpful for me to run my notebook and understand how it works (what are the benefits of using  Arcface with adaptive margin). Thank you!</p>",
      "rawMarkdown": "Could you share implementation of  Arcface with adaptive margin?  It would be helpful for me to run my notebook and understand how it works (what are the benefits of using  Arcface with adaptive margin). Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1763956,
      "author_name": "biglafe",
      "author_url": "",
      "post_date": "04/22/2022 02:26:37",
      "content": "<p>GOOD JOB!!! How much is the difference between your loss in the training set and the loss in the validation?<br>\nMy feeling is to use effb6 to produce the over-fitting, so I use hard enhancement(mixup …) and feature space constraints to alleviate.  </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1764023,
      "author_name": "niekvanderzwaag",
      "author_url": "",
      "post_date": "04/22/2022 04:42:31",
      "content": "<p>Thanks for sharing your approach, very insightful! Congrats on your 2nd place 🎉</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1764237,
      "author_name": "syxuming",
      "author_url": "",
      "post_date": "04/22/2022 09:26:55",
      "content": "<p><code>Gradient checkpointing</code><br>\nCan you recommend some information or code for pytorch gradient checkpoint? It sounds amazing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1764330,
          "author_name": "amanatsu",
          "author_url": "",
          "post_date": "04/22/2022 11:58:09",
          "content": "<p>In the latest master branch of timm, gradient checkpointing is available.</p>\n<p><a href=\"https://github.com/rwightman/pytorch-image-models/blob/01a0e25a67305b94ea767083f4113ff002e4435c/timm/models/efficientnet.py#L527-L528\" target=\"_blank\">https://github.com/rwightman/pytorch-image-models/blob/01a0e25a67305b94ea767083f4113ff002e4435c/timm/models/efficientnet.py#L527-L528</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764358,
          "author_name": "tascj0",
          "author_url": "",
          "post_date": "04/22/2022 12:29:49",
          "content": "<p>It's in pytorch for years. You can check the DenseNet implementation in torchvision.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764436,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "04/22/2022 13:25:19",
          "content": "<p>I found really nice post here: <a href=\"https://spell.ml/blog/gradient-checkpointing-pytorch-YGypLBAAACEAefHs\" target=\"_blank\">Training larger-than-memory PyTorch models using gradient checkpointing</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1764328,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/22/2022 11:55:35",
      "content": "<p>First of all congratulations! Outstanding result - initial model LB score 0.846 👍👍👍</p>\n<p>May I ask you about explaining:</p>\n<blockquote>\n  <p>RandAugment improved the validation score quite a bit.</p>\n</blockquote>\n<p>How did you implement this? Could you provide part of the code which shows us RandAugment? </p>\n<blockquote>\n  <p>((logits * scale).softmax(-1)) </p>\n</blockquote>\n<p>What is a scale?</p>\n<blockquote>\n  <p>gradient checkpointing</p>\n</blockquote>\n<p>Thank you for providing this tip. Looks interesting.</p>\n<p>As I can see you made your pipeline using Pytorch. There is information about RTX3090. Did you use only one RTX3090 to train all efnet_l2 models?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1764350,
          "author_name": "tascj0",
          "author_url": "",
          "post_date": "04/22/2022 12:15:17",
          "content": "<blockquote>\n  <p>How did you implement this? Could you provide part of the code which shows us RandAugment?</p>\n</blockquote>\n<p>We used the RandAugment implementation in <a href=\"https://github.com/open-mmlab/mmclassification\" target=\"_blank\">mmclassification</a>. The policies and hyperparamters are copied from imagenet configs in mmclassification.</p>\n<blockquote>\n  <p>What is a scale?</p>\n</blockquote>\n<p>ArcFace scale parameter.</p>\n<blockquote>\n  <p>Did you use only one RTX3090 to train all efnet_l2 models?</p>\n</blockquote>\n<p>I trained the first two rounds of effcientnet_l2 using a single RTX3090. It took 7~8 days. <a href=\"https://www.kaggle.com/takuok\" target=\"_blank\">@takuok</a> used multiple A100 GPUs to train other models since the deadline is approaching.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1764353,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "04/22/2022 12:19:59",
          "content": "<p>Thank you very much!!!!</p>\n<p>7-8 days 😳😳😳</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1765768,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "04/23/2022 20:26:15",
      "content": "<p>Could you share implementation of  Arcface with adaptive margin?  It would be helpful for me to run my notebook and understand how it works (what are the benefits of using  Arcface with adaptive margin). Thank you!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1763934": "Thank you to the competition host, Kaggle and all participants. And congrats to the winners.\nMany thanks to @jpbremer. We used the public fullbody dataset and annotation as our main dataset.  Without it we could not achieve a high score in a short period of time.\n\n# Single model\n## Dataset\nTwo different datasets were used. \n- fullbody dataset (yolov5): [this dataset](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/311184) cropped by yolov5. \n- fullbody dataset (yolox): the above annotations were used for GT and trained yolox.\n\nAll models were trained using the fullbody dataset (yolov5) before the last day of the contest. Finally, a few models were fine-tuned a few epochs using the fullbody dataset (yolox).\n\n## Model\nefficientnet_l2 worked the best in validation\n\n### Training Recipe\n- backbone = tf_efficientnet_l2_ns\n- img_size = 768\n- loss = Arcface with adaptive margin\n- augmentation = Horizontal flip, RandAugment\n- optimizer = SGD\n- scheduler = CosineDecay with warmup\n- batch_size = 16 per GPU\n- n_epoch = 20\n\nRandAugment improved the validation score quite a bit.\n\n## Pseudo labeling\n\nThis is the key to getting a high score on the leaderboard.\n\nWe use FC layer prediction (`(logits * scale).softmax(-1)`) of trained models to generate pseudo labels. The confidence threshold was set to 0.8.\nEvery pseudo labeling round was trained from imagenet pretrained weights.\n\nFollowing are the leaderboard scores of each round. We used flip testing starting round3.\n\nPseudo label rounds: backbone = efficientnet_l2\n| Pseudo label round     | Dataset           | Public score | Private score |\n|------------------------|-------------------|--------------|---------------|\n| initial model (round1) | fullbody (yolov5) | 0.846        | 0.812         |\n| round2                 | fullbody (yolov5) | 0.875        | 0.849         |\n| round3                 | fullbody (yolov5) | 0.885        | 0.860         |\n| round4                 | fullbody (yolov5) | 0.889        | 0.862         |\n| round5                 | fullbody (yolov5) | 0.887        | 0.863         |\n| round5 (fine-tuning)   | fullbody (yolox)  | 0.891        | 0.870         |\n\n## Make submission\nThe normal image retrieval method. Calculate cos similarity with train dataset and get top5 ids.\n\nA fixed cosine similarity of 0.5 was used to insert “new_individual”. There is nothing special post-processing, such as species-specific thresholds.\n\n# Last ensemble\nThe final submission is a 4-model ensemble of efficientnet_l2 belows. (public=0.897 / private=0.872)\n\nEnsemble models: all backbone = efficientnet_l2\n| Pseudo label round | Image size   | Dataset                                                     |\n|--------------------|--------------|-------------------------------------------------------------|\n| round3             | (1024, 1024) | fullbody (yolov5)                                           |\n| round4             | (768, 768)   | fullbody (yolov5)                                           |\n| round5             | (768, 768)   | fullbody (yolox) fine-tuned                                 |\n| round5             | (768x2, 768) | vertical concatenated original dataset and fullbody (yolox) |\n\n\n# Performance tips\n## Gradient checkpointing\n\nTo train efficientnet_l2 on RTX3090, gradient checkpoint is a must. With gradient checkpointing and mixed precision, we could train the network with batch_size 16 on a single RTX3090. Without it, even batch size 2 gives OOM.\n\n## PyTorch build\n\nWe found that different PyTorch builds can greatly affect training throughput. We compared the official release and the build by NVIDIA.\n\n| docker image                                 | cudnn_benchmark   |   training throughput (imgs/sec) |\n|:---------------------------------------------|:------------------|----------------------:|\n| pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel | False             |                   2.5 |\n| pytorch/pytorch:1.11.0-cuda11.3-cudnn8-devel | True              |                   3.4 |\n| nvcr.io/nvidia/pytorch:22.02-py3             | False             |                   4   |\n| nvcr.io/nvidia/pytorch:22.02-py3             | True              |                   4.1 |\n\nThe build from NVIDIA is 20% faster. Probably because of the updated CUDA/CuDNN version.\n\n# Acknowledge\ntakuoko is a member of Z by HP Data Science Global Ambassadors. Special Thanks to Z by HP for sponsoring me a Z8G4 Workstation with dual A6000 GPU and a ZBook with RTX5000 GPU.",
    "1763956": "GOOD JOB!!! How much is the difference between your loss in the training set and the loss in the validation?\nMy feeling is to use effb6 to produce the over-fitting, so I use hard enhancement(mixup ...) and feature space constraints to alleviate.",
    "1764023": "Thanks for sharing your approach, very insightful! Congrats on your 2nd place 🎉",
    "1764237": "`Gradient checkpointing`\nCan you recommend some information or code for pytorch gradient checkpoint? It sounds amazing!",
    "1764328": "First of all congratulations! Outstanding result - initial model LB score 0.846 👍👍👍\n\nMay I ask you about explaining:\n\n> RandAugment improved the validation score quite a bit.\n\nHow did you implement this? Could you provide part of the code which shows us RandAugment? \n\n> ((logits * scale).softmax(-1)) \n\nWhat is a scale?\n\n> gradient checkpointing\n\nThank you for providing this tip. Looks interesting.\n\nAs I can see you made your pipeline using Pytorch. There is information about RTX3090. Did you use only one RTX3090 to train all efnet_l2 models?",
    "1764330": "In the latest master branch of timm, gradient checkpointing is available.\n\nhttps://github.com/rwightman/pytorch-image-models/blob/01a0e25a67305b94ea767083f4113ff002e4435c/timm/models/efficientnet.py#L527-L528",
    "1764350": "> How did you implement this? Could you provide part of the code which shows us RandAugment?\n\nWe used the RandAugment implementation in [mmclassification](https://github.com/open-mmlab/mmclassification). The policies and hyperparamters are copied from imagenet configs in mmclassification.\n\n> What is a scale?\n\nArcFace scale parameter.\n\n> Did you use only one RTX3090 to train all efnet_l2 models?\n\nI trained the first two rounds of effcientnet_l2 using a single RTX3090. It took 7~8 days. @takuok used multiple A100 GPUs to train other models since the deadline is approaching.",
    "1764353": "Thank you very much!!!!\n\n7-8 days 😳😳😳",
    "1764358": "It's in pytorch for years. You can check the DenseNet implementation in torchvision.",
    "1764436": "I found really nice post here: [Training larger-than-memory PyTorch models using gradient checkpointing](https://spell.ml/blog/gradient-checkpointing-pytorch-YGypLBAAACEAefHs)",
    "1765768": "Could you share implementation of  Arcface with adaptive margin?  It would be helpful for me to run my notebook and understand how it works (what are the benefits of using  Arcface with adaptive margin). Thank you!"
  },
  "source": "meta"
}