{
  "id": 319787,
  "title": "How to take the ArcFace Baseline from private lb: 0.470 (public: 0.522) to 0.804 (public: 0.834)",
  "url": "/competitions/happy-whale-and-dolphin/discussion/319787",
  "author_name": "",
  "post_date": "2022-04-19T00:54:43.137241Z",
  "votes": 29,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I thought others might be interested to learn the steps I took to take the <a href=\"https://www.kaggle.com/code/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">public solution by ks</a> from 0.470 on the private leaderboard to 0.804. I'm sure others were able to push it much further!</p>\n<p><strong>Link to kernel <a href=\"https://www.kaggle.com/code/lextoumbourou/happywhale-tpu-baseline-to-0-804-elasticface\" target=\"_blank\">here</a></strong></p>\n<p>I had to train it in GCP due to the TPU wait time, but it should train just fine in Kaggle.</p>\n<h2>Summary of changes</h2>\n<ul>\n<li><p>Image size: 864x846.</p></li>\n<li><p>Pretrained model: EfficientNet b5 with noisy-student.</p></li>\n<li><p>Pooling: concat pooling (from fastai)</p>\n<pre><code>avg_pool = tf.keras.layers.GlobalAveragePooling2D()(x)\nmax_pool = tf.keras.layers.GlobalMaxPooling2D()(x)\npretrained_out = tf.keras.layers.Concatenate()([avg_pool, max_pool])\n</code></pre></li>\n<li><p>Head:</p>\n<ul>\n<li>Dual-head model: Species classification output and metric.</li>\n<li>Add <a href=\"https://arxiv.org/abs/1905.09788\" target=\"_blank\">Multi-Sample Dropout</a> (idea thanks to <a href=\"https://www.kaggle.com/dhakshiin\" target=\"_blank\">@dhakshiin</a>) before metric.</li></ul></li>\n<li><p>Metric function: <a href=\"https://arxiv.org/abs/2109.09416\" target=\"_blank\">Elastic Margin Loss</a> (w=0.3, std=0.025, s=30) (minor improvement over ArcFace).</p></li>\n<li><p>Embed size: 1024.</p></li>\n<li><p>Data:</p>\n<ul>\n<li>Clean species label thanks to <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305574\" target=\"_blank\">this post</a> by <a href=\"https://www.kaggle.com/kwentar\" target=\"_blank\">@kwentar</a></li>\n<li>Train on entire dataset (no validation) with 12157 extra pseudo labels.</li>\n<li>Use pseudo labelling-algorithm by <a href=\"https://www.kaggle.com/dhakshiin\" target=\"_blank\">@dhakshiin</a>: Use multiple models, find top-1 examples that the majority agree on, ensuring 0.2 separations between the confidence of 1st class and 2nd class.</li>\n<li>Select between 8x crops randomly selected at training time: full-body crops, detic crops, original yolov5 crops, 3x TokenCut crops and full-sized images (<a href=\"https://www.kaggle.com/lextoumbourou/happywhale-generate-tfrecords-with-pseudo\" target=\"_blank\">See my TFRecords notebook here</a>)</li></ul></li>\n<li><p>Augmentations: <code>random_flip_left_right</code>, <code>random_hue</code>, <code>random_saturation</code>, <code>random_contrast</code>, <code>random_brightness</code> and <code>random_rgb_to_gray</code> (parameters tuned slightly from original kernel)</p></li>\n<li><p>Optimiser: Adam (unchanged from original)</p></li>\n<li><p>Learning Rate: exponential decay with 4 epochs warm-up (unchanged from original)</p></li>\n<li><p>Epochs: 30</p></li>\n<li><p>Inference:</p>\n<ul>\n<li>Generated embeddings for each crop with a horizontal flip.</li>\n<li>Take a weighted mean of all crop embeddings to create final embeddings.</li>\n<li>Inference: standard KNN inference (use KNN=150 when added to the final ensemble)</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "1759852",
      "postDate": "04/19/2022 00:54:43",
      "content": "<p>I thought others might be interested to learn the steps I took to take the <a href=\"https://www.kaggle.com/code/ks2019/happywhale-arcface-baseline-tpu\" target=\"_blank\">public solution by ks</a> from 0.470 on the private leaderboard to 0.804. I'm sure others were able to push it much further!</p>\n<p><strong>Link to kernel <a href=\"https://www.kaggle.com/code/lextoumbourou/happywhale-tpu-baseline-to-0-804-elasticface\" target=\"_blank\">here</a></strong></p>\n<p>I had to train it in GCP due to the TPU wait time, but it should train just fine in Kaggle.</p>\n<h2>Summary of changes</h2>\n<ul>\n<li><p>Image size: 864x846.</p></li>\n<li><p>Pretrained model: EfficientNet b5 with noisy-student.</p></li>\n<li><p>Pooling: concat pooling (from fastai)</p>\n<pre><code>avg_pool = tf.keras.layers.GlobalAveragePooling2D()(x)\nmax_pool = tf.keras.layers.GlobalMaxPooling2D()(x)\npretrained_out = tf.keras.layers.Concatenate()([avg_pool, max_pool])\n</code></pre></li>\n<li><p>Head:</p>\n<ul>\n<li>Dual-head model: Species classification output and metric.</li>\n<li>Add <a href=\"https://arxiv.org/abs/1905.09788\" target=\"_blank\">Multi-Sample Dropout</a> (idea thanks to <a href=\"https://www.kaggle.com/dhakshiin\" target=\"_blank\">@dhakshiin</a>) before metric.</li></ul></li>\n<li><p>Metric function: <a href=\"https://arxiv.org/abs/2109.09416\" target=\"_blank\">Elastic Margin Loss</a> (w=0.3, std=0.025, s=30) (minor improvement over ArcFace).</p></li>\n<li><p>Embed size: 1024.</p></li>\n<li><p>Data:</p>\n<ul>\n<li>Clean species label thanks to <a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305574\" target=\"_blank\">this post</a> by <a href=\"https://www.kaggle.com/kwentar\" target=\"_blank\">@kwentar</a></li>\n<li>Train on entire dataset (no validation) with 12157 extra pseudo labels.</li>\n<li>Use pseudo labelling-algorithm by <a href=\"https://www.kaggle.com/dhakshiin\" target=\"_blank\">@dhakshiin</a>: Use multiple models, find top-1 examples that the majority agree on, ensuring 0.2 separations between the confidence of 1st class and 2nd class.</li>\n<li>Select between 8x crops randomly selected at training time: full-body crops, detic crops, original yolov5 crops, 3x TokenCut crops and full-sized images (<a href=\"https://www.kaggle.com/lextoumbourou/happywhale-generate-tfrecords-with-pseudo\" target=\"_blank\">See my TFRecords notebook here</a>)</li></ul></li>\n<li><p>Augmentations: <code>random_flip_left_right</code>, <code>random_hue</code>, <code>random_saturation</code>, <code>random_contrast</code>, <code>random_brightness</code> and <code>random_rgb_to_gray</code> (parameters tuned slightly from original kernel)</p></li>\n<li><p>Optimiser: Adam (unchanged from original)</p></li>\n<li><p>Learning Rate: exponential decay with 4 epochs warm-up (unchanged from original)</p></li>\n<li><p>Epochs: 30</p></li>\n<li><p>Inference:</p>\n<ul>\n<li>Generated embeddings for each crop with a horizontal flip.</li>\n<li>Take a weighted mean of all crop embeddings to create final embeddings.</li>\n<li>Inference: standard KNN inference (use KNN=150 when added to the final ensemble)</li></ul></li>\n</ul>",
      "rawMarkdown": "I thought others might be interested to learn the steps I took to take the [public solution by ks](https://www.kaggle.com/code/ks2019/happywhale-arcface-baseline-tpu) from 0.470 on the private leaderboard to 0.804. I'm sure others were able to push it much further!\n\n**Link to kernel [here](https://www.kaggle.com/code/lextoumbourou/happywhale-tpu-baseline-to-0-804-elasticface)**\n\nI had to train it in GCP due to the TPU wait time, but it should train just fine in Kaggle.\n\n## Summary of changes\n\n- Image size: 864x846.\n- Pretrained model: EfficientNet b5 with noisy-student.\n- Pooling: concat pooling (from fastai)\n\n    ```\n    avg_pool = tf.keras.layers.GlobalAveragePooling2D()(x)\n    max_pool = tf.keras.layers.GlobalMaxPooling2D()(x)\n    pretrained_out = tf.keras.layers.Concatenate()([avg_pool, max_pool])\n    ```\n\n- Head:\n  - Dual-head model: Species classification output and metric.\n  - Add [Multi-Sample Dropout](https://arxiv.org/abs/1905.09788) (idea thanks to @dhakshiin) before metric.\n\n- Metric function: [Elastic Margin Loss](https://arxiv.org/abs/2109.09416) (w=0.3, std=0.025, s=30) (minor improvement over ArcFace).\n- Embed size: 1024.\n- Data:\n    - Clean species label thanks to [this post](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305574) by @kwentar\n    - Train on entire dataset (no validation) with 12157 extra pseudo labels.\n    - Use pseudo labelling-algorithm by @dhakshiin: Use multiple models, find top-1 examples that the majority agree on, ensuring 0.2 separations between the confidence of 1st class and 2nd class.\n    - Select between 8x crops randomly selected at training time: full-body crops, detic crops, original yolov5 crops, 3x TokenCut crops and full-sized images ([See my TFRecords notebook here](https://www.kaggle.com/lextoumbourou/happywhale-generate-tfrecords-with-pseudo))\n- Augmentations: `random_flip_left_right`, `random_hue`, `random_saturation`, `random_contrast`, `random_brightness` and `random_rgb_to_gray` (parameters tuned slightly from original kernel)\n- Optimiser: Adam (unchanged from original)\n- Learning Rate: exponential decay with 4 epochs warm-up (unchanged from original)\n- Epochs: 30\n\n- Inference:\n  - Generated embeddings for each crop with a horizontal flip.\n  - Take a weighted mean of all crop embeddings to create final embeddings.\n  - Inference: standard KNN inference (use KNN=150 when added to the final ensemble)",
      "votes": null
    },
    {
      "id": "1759877",
      "postDate": "04/19/2022 01:09:10",
      "content": "<p>Nice work! May I ask you explain GAP, GMP part more detail? Isn't the output of pretrained efficientnet backbone is 1D feature vector? 😯</p>",
      "rawMarkdown": "Nice work! May I ask you explain GAP, GMP part more detail? Isn't the output of pretrained efficientnet backbone is 1D feature vector? 😯",
      "votes": null
    },
    {
      "id": "1759892",
      "postDate": "04/19/2022 01:19:12",
      "content": "<p>For more pretrained models, there's an option to throw away the original head. For <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet/EfficientNetB7\" target=\"_blank\">Tensorflow's EfficientNet</a> it's <code>include_top=False</code>, which by default also throws away the pooling layers.</p>\n<p>So the output of the pretrained model becomes a 4D tensor output of the last convolutional layer, which can be passed to multiple pooling layers.</p>\n<p>Does that make sense?</p>",
      "rawMarkdown": "For more pretrained models, there's an option to throw away the original head. For [Tensorflow's EfficientNet](https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet/EfficientNetB7) it's `include_top=False`, which by default also throws away the pooling layers.\n\nSo the output of the pretrained model becomes a 4D tensor output of the last convolutional layer, which can be passed to multiple pooling layers.\n\nDoes that make sense?",
      "votes": null
    },
    {
      "id": "1759936",
      "postDate": "04/19/2022 02:09:54",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!",
      "votes": null
    },
    {
      "id": "1759937",
      "postDate": "04/19/2022 02:11:10",
      "content": "<p>Congratulate,can you share the GCS_DS_PATH and Data?thank you</p>",
      "rawMarkdown": "Congratulate,can you share the GCS_DS_PATH and Data?thank you",
      "votes": null
    },
    {
      "id": "1759981",
      "postDate": "04/19/2022 02:37:57",
      "content": "<p>Thanks for sharing knowledge XD</p>",
      "rawMarkdown": "Thanks for sharing knowledge XD",
      "votes": null
    },
    {
      "id": "1760037",
      "postDate": "04/19/2022 03:36:29",
      "content": "<p>Done. All the attached data should be public now including the embeddings. Let me know if you have any issues.</p>",
      "rawMarkdown": "Done. All the attached data should be public now including the embeddings. Let me know if you have any issues.",
      "votes": null
    },
    {
      "id": "1760324",
      "postDate": "04/19/2022 08:04:23",
      "content": "<p>idk why your dataset is public but still cannot be accessed. Maybe it's take time to fully public on Kaggle.</p>",
      "rawMarkdown": "idk why your dataset is public but still cannot be accessed. Maybe it's take time to fully public on Kaggle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1759877,
      "author_name": "yeongjingwakcs",
      "author_url": "",
      "post_date": "04/19/2022 01:09:10",
      "content": "<p>Nice work! May I ask you explain GAP, GMP part more detail? Isn't the output of pretrained efficientnet backbone is 1D feature vector? 😯</p>",
      "votes": null,
      "replies": [
        {
          "id": 1759892,
          "author_name": "lextoumbourou",
          "author_url": "",
          "post_date": "04/19/2022 01:19:12",
          "content": "<p>For more pretrained models, there's an option to throw away the original head. For <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet/EfficientNetB7\" target=\"_blank\">Tensorflow's EfficientNet</a> it's <code>include_top=False</code>, which by default also throws away the pooling layers.</p>\n<p>So the output of the pretrained model becomes a 4D tensor output of the last convolutional layer, which can be passed to multiple pooling layers.</p>\n<p>Does that make sense?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1759981,
          "author_name": "yeongjingwakcs",
          "author_url": "",
          "post_date": "04/19/2022 02:37:57",
          "content": "<p>Thanks for sharing knowledge XD</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1759936,
      "author_name": "yangranran",
      "author_url": "",
      "post_date": "04/19/2022 02:09:54",
      "content": "<p>Great work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1759937,
      "author_name": "qi0239",
      "author_url": "",
      "post_date": "04/19/2022 02:11:10",
      "content": "<p>Congratulate,can you share the GCS_DS_PATH and Data?thank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760037,
          "author_name": "lextoumbourou",
          "author_url": "",
          "post_date": "04/19/2022 03:36:29",
          "content": "<p>Done. All the attached data should be public now including the embeddings. Let me know if you have any issues.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1760324,
          "author_name": "locbaop",
          "author_url": "",
          "post_date": "04/19/2022 08:04:23",
          "content": "<p>idk why your dataset is public but still cannot be accessed. Maybe it's take time to fully public on Kaggle.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1759852": "I thought others might be interested to learn the steps I took to take the [public solution by ks](https://www.kaggle.com/code/ks2019/happywhale-arcface-baseline-tpu) from 0.470 on the private leaderboard to 0.804. I'm sure others were able to push it much further!\n\n**Link to kernel [here](https://www.kaggle.com/code/lextoumbourou/happywhale-tpu-baseline-to-0-804-elasticface)**\n\nI had to train it in GCP due to the TPU wait time, but it should train just fine in Kaggle.\n\n## Summary of changes\n\n- Image size: 864x846.\n- Pretrained model: EfficientNet b5 with noisy-student.\n- Pooling: concat pooling (from fastai)\n\n    ```\n    avg_pool = tf.keras.layers.GlobalAveragePooling2D()(x)\n    max_pool = tf.keras.layers.GlobalMaxPooling2D()(x)\n    pretrained_out = tf.keras.layers.Concatenate()([avg_pool, max_pool])\n    ```\n\n- Head:\n  - Dual-head model: Species classification output and metric.\n  - Add [Multi-Sample Dropout](https://arxiv.org/abs/1905.09788) (idea thanks to @dhakshiin) before metric.\n\n- Metric function: [Elastic Margin Loss](https://arxiv.org/abs/2109.09416) (w=0.3, std=0.025, s=30) (minor improvement over ArcFace).\n- Embed size: 1024.\n- Data:\n    - Clean species label thanks to [this post](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/305574) by @kwentar\n    - Train on entire dataset (no validation) with 12157 extra pseudo labels.\n    - Use pseudo labelling-algorithm by @dhakshiin: Use multiple models, find top-1 examples that the majority agree on, ensuring 0.2 separations between the confidence of 1st class and 2nd class.\n    - Select between 8x crops randomly selected at training time: full-body crops, detic crops, original yolov5 crops, 3x TokenCut crops and full-sized images ([See my TFRecords notebook here](https://www.kaggle.com/lextoumbourou/happywhale-generate-tfrecords-with-pseudo))\n- Augmentations: `random_flip_left_right`, `random_hue`, `random_saturation`, `random_contrast`, `random_brightness` and `random_rgb_to_gray` (parameters tuned slightly from original kernel)\n- Optimiser: Adam (unchanged from original)\n- Learning Rate: exponential decay with 4 epochs warm-up (unchanged from original)\n- Epochs: 30\n\n- Inference:\n  - Generated embeddings for each crop with a horizontal flip.\n  - Take a weighted mean of all crop embeddings to create final embeddings.\n  - Inference: standard KNN inference (use KNN=150 when added to the final ensemble)",
    "1759877": "Nice work! May I ask you explain GAP, GMP part more detail? Isn't the output of pretrained efficientnet backbone is 1D feature vector? 😯",
    "1759892": "For more pretrained models, there's an option to throw away the original head. For [Tensorflow's EfficientNet](https://www.tensorflow.org/api_docs/python/tf/keras/applications/efficientnet/EfficientNetB7) it's `include_top=False`, which by default also throws away the pooling layers.\n\nSo the output of the pretrained model becomes a 4D tensor output of the last convolutional layer, which can be passed to multiple pooling layers.\n\nDoes that make sense?",
    "1759936": "Great work!",
    "1759937": "Congratulate,can you share the GCS_DS_PATH and Data?thank you",
    "1759981": "Thanks for sharing knowledge XD",
    "1760037": "Done. All the attached data should be public now including the embeddings. Let me know if you have any issues.",
    "1760324": "idk why your dataset is public but still cannot be accessed. Maybe it's take time to fully public on Kaggle."
  },
  "source": "meta"
}