{
  "id": 154186,
  "title": "2nd Place Solution",
  "url": "/competitions/herbarium-2020-fgvc7/writeups/sesamind-2nd-place-solution",
  "author_name": "",
  "post_date": "2020-05-28T20:22:10.497Z",
  "votes": 17,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First of all special thanks to my teammate Tal :)</p>\n\n<h2>Backbone:</h2>\n\n<p>As a backbone, we used TResNet architecture (High Performance GPU-Dedicated Architecture) which allowed us to do fast training with large batches, while maintaining top scores along all resolutions  (more details in our <a href=\"https://arxiv.org/pdf/2003.13630.pdf\">paper</a> and code <a href=\"https://github.com/mrT23/TResNet\">here</a>).\nModels used in this competition: TResNet-M, TResNet-L.</p>\n\n<h2>BottleNeck Head:</h2>\n\n<p>We avoided having a massive classification layer: 2048x32094 (65.72M params), by using a bottleneck head that reduces the embedding feature vector from 2048 dims to 512 dims. total parameters= 2048x512 + 512x32094 (17.48M params).\nThis modification not only decreased the size of the network, but allowed us to use samples in larger batch sizes (Faster Training and helps the Soft-Triplet loss).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2Fd861d950b9f015672631c7b59815e5a5%2FBasicBlcok_Tresnet%20(1\" alt=\"\">%20copy.png?generation=1590586933112476&amp;alt=media)</p>\n\n<h2>Loss function</h2>\n\n<p>We used Cross-Entropy loss with label smoothing 0.2, and in addition we added Soft-Triplet Loss with Soft Margin which emphasizes hard examples by focusing of the most difficult positives and negatives samples in the batch, for more details check out our ReID <a href=\"https://arxiv.org/pdf/1910.07038.pdf\">paper</a> Section 3.1.\nOnline-Hard Negative mining works best when training with large batch size (more negative examples); hence, we leveraged the high memory utilization of our  Network and the bottleneck head. for example: with TResNet-M @ 448x448, we could reach maximum batch size of 128 on a single V100 16GB, significantly higher than other common networks like EfficientNet and ResNext.</p>\n\n<h2>Class Balancing</h2>\n\n<p>Herbarium 2020 dataset suffers from prominent class im-balance, therefore we used a simple loss weighting, multiplying by the inverse of class frequency.\nNote: We used class balancing only in the last 15~ epochs of the training, to enable the network to learn the basic features first.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2F332b1711fa4eefa0b2de9984d491742d%2FScreen%20Shot%202020-05-27%20at%2016.46.38.png?generation=1590587199078597&amp;alt=media\" alt=\"\"></p>\n\n<h2>Augmentations</h2>\n\n<p>• Squish instead of Crop\n• FlipLR\n• Cutout: 0.5</p>\n\nSpecial augmentation scheme (+~0.5% on public LB):\n\n<ul>\n<li>Random Zoom-in (p=0.4, zoom_factor= random between 1-1.15)</li>\n<li>Reduced affinity transformations: p_shear: 0.1, p_rotate=0.2, rotate_degrees=10</li>\n<li>Decolorize: p_decolorize: 0.4</li>\n<li>Rectangular resizing: small gains\n<h3>TTA</h3></li>\n</ul>\n\n<p>Averaging horizontal flip predictions</p>\n\n<h2>Our best single model:</h2>\n\n<p>Model: TResNet-L</p>\n\n<p><strong>Stage 1:</strong>\n100 epochs at 448x448, CE+SoftTriplet (class-balancing starting from epoch 85)\nPublic LB score (without  post-processing):       0.84089\nPublic LB score (with  post-processing P1+P2): 0.85093</p>\n\n<p><strong>Stage 2 (high resolution fine-tuning):</strong>\n10 epoch fine-tune + class balancing on 544x416\nPublic LB score (without  post-processing):       0.84727\nPublic LB score (with  post-processing P1+P2): 0.85604</p>\n\n<h2>Post-Process (+~0.9% on the pubic LB)</h2>\n\n<p>The origanizers wrote in the description that:\n\"Each category has at least 1 instance in both the training and test datasets.\"</p>\n\n<p>Based on this prior knowledge we used two post-processing algorithms that tries to fill some of the empty classes.</p>\n\n<p><strong>P1: Switching classes based on low top-1 confidence (~0.9% on the public LB)</strong>\nthe idea is basically whenever the network is not confident about its predictions, and if one of the empty classes is close to top1 prediction, then switch predictions:\n1) Go through all the top 2 predictions in the test set and if all these criteria are met switch between top1 and top2:\n• top2 &gt; top1 * m (m=0.3)\n• top2 class never appeared in the top1 predictions\n• top1 class appeared at least once as top2 prediction for another sample image in the test set.\n2) Repeat the same switching algorithm above for top 3, 4 and 5 (performed sequentially, first perform all of top2 switches and then top3 switches and so on).</p>\n\n<p><strong>P2: Filling empty classes (~+0.1-0.2% on the public LB)</strong>\n1) Go through the predictions after the first post-process, and for all classes that don't appear in the top1 predictions and find the image sample with the highest prediction from all of the top5 predictions of the test set.</p>\n\n<h2>Pre-trained model</h2>\n\n<p>We used pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slightly better scores on the public LB. </p>\n\n<h2>Ensemble (+0.8% on public LB)</h2>\n\n<p>We used an ensemble of 9 models trained from different pretrains (imagenet and iNaturalist2018) and different resolutions, final resolutions: 544x416 mostly and some 620x476. \nPublic LB score (with post-processing P1+P2):  0.86310</p>\n\n<h2>What didn't work for us (relevant to this competition only):</h2>\n\n<ul>\n<li>Fixed cropping by 10-20%</li>\n<li>Pseudo-labeling</li>\n<li>Mixup</li>\n<li>Focal loss</li>\n</ul>",
  "messages": [
    {
      "id": "863710",
      "postDate": "05/27/2020 14:07:43",
      "content": "<p>First of all special thanks to my teammate Tal :)</p>\n\n<h2>Backbone:</h2>\n\n<p>As a backbone, we used TResNet architecture (High Performance GPU-Dedicated Architecture) which allowed us to do fast training with large batches, while maintaining top scores along all resolutions  (more details in our <a href=\"https://arxiv.org/pdf/2003.13630.pdf\">paper</a> and code <a href=\"https://github.com/mrT23/TResNet\">here</a>).\nModels used in this competition: TResNet-M, TResNet-L.</p>\n\n<h2>BottleNeck Head:</h2>\n\n<p>We avoided having a massive classification layer: 2048x32094 (65.72M params), by using a bottleneck head that reduces the embedding feature vector from 2048 dims to 512 dims. total parameters= 2048x512 + 512x32094 (17.48M params).\nThis modification not only decreased the size of the network, but allowed us to use samples in larger batch sizes (Faster Training and helps the Soft-Triplet loss).\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2Fd861d950b9f015672631c7b59815e5a5%2FBasicBlcok_Tresnet%20(1\" alt=\"\">%20copy.png?generation=1590586933112476&amp;alt=media)</p>\n\n<h2>Loss function</h2>\n\n<p>We used Cross-Entropy loss with label smoothing 0.2, and in addition we added Soft-Triplet Loss with Soft Margin which emphasizes hard examples by focusing of the most difficult positives and negatives samples in the batch, for more details check out our ReID <a href=\"https://arxiv.org/pdf/1910.07038.pdf\">paper</a> Section 3.1.\nOnline-Hard Negative mining works best when training with large batch size (more negative examples); hence, we leveraged the high memory utilization of our  Network and the bottleneck head. for example: with TResNet-M @ 448x448, we could reach maximum batch size of 128 on a single V100 16GB, significantly higher than other common networks like EfficientNet and ResNext.</p>\n\n<h2>Class Balancing</h2>\n\n<p>Herbarium 2020 dataset suffers from prominent class im-balance, therefore we used a simple loss weighting, multiplying by the inverse of class frequency.\nNote: We used class balancing only in the last 15~ epochs of the training, to enable the network to learn the basic features first.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2F332b1711fa4eefa0b2de9984d491742d%2FScreen%20Shot%202020-05-27%20at%2016.46.38.png?generation=1590587199078597&amp;alt=media\" alt=\"\"></p>\n\n<h2>Augmentations</h2>\n\n<p>• Squish instead of Crop\n• FlipLR\n• Cutout: 0.5</p>\n\nSpecial augmentation scheme (+~0.5% on public LB):\n\n<ul>\n<li>Random Zoom-in (p=0.4, zoom_factor= random between 1-1.15)</li>\n<li>Reduced affinity transformations: p_shear: 0.1, p_rotate=0.2, rotate_degrees=10</li>\n<li>Decolorize: p_decolorize: 0.4</li>\n<li>Rectangular resizing: small gains\n<h3>TTA</h3></li>\n</ul>\n\n<p>Averaging horizontal flip predictions</p>\n\n<h2>Our best single model:</h2>\n\n<p>Model: TResNet-L</p>\n\n<p><strong>Stage 1:</strong>\n100 epochs at 448x448, CE+SoftTriplet (class-balancing starting from epoch 85)\nPublic LB score (without  post-processing):       0.84089\nPublic LB score (with  post-processing P1+P2): 0.85093</p>\n\n<p><strong>Stage 2 (high resolution fine-tuning):</strong>\n10 epoch fine-tune + class balancing on 544x416\nPublic LB score (without  post-processing):       0.84727\nPublic LB score (with  post-processing P1+P2): 0.85604</p>\n\n<h2>Post-Process (+~0.9% on the pubic LB)</h2>\n\n<p>The origanizers wrote in the description that:\n\"Each category has at least 1 instance in both the training and test datasets.\"</p>\n\n<p>Based on this prior knowledge we used two post-processing algorithms that tries to fill some of the empty classes.</p>\n\n<p><strong>P1: Switching classes based on low top-1 confidence (~0.9% on the public LB)</strong>\nthe idea is basically whenever the network is not confident about its predictions, and if one of the empty classes is close to top1 prediction, then switch predictions:\n1) Go through all the top 2 predictions in the test set and if all these criteria are met switch between top1 and top2:\n• top2 &gt; top1 * m (m=0.3)\n• top2 class never appeared in the top1 predictions\n• top1 class appeared at least once as top2 prediction for another sample image in the test set.\n2) Repeat the same switching algorithm above for top 3, 4 and 5 (performed sequentially, first perform all of top2 switches and then top3 switches and so on).</p>\n\n<p><strong>P2: Filling empty classes (~+0.1-0.2% on the public LB)</strong>\n1) Go through the predictions after the first post-process, and for all classes that don't appear in the top1 predictions and find the image sample with the highest prediction from all of the top5 predictions of the test set.</p>\n\n<h2>Pre-trained model</h2>\n\n<p>We used pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slightly better scores on the public LB. </p>\n\n<h2>Ensemble (+0.8% on public LB)</h2>\n\n<p>We used an ensemble of 9 models trained from different pretrains (imagenet and iNaturalist2018) and different resolutions, final resolutions: 544x416 mostly and some 620x476. \nPublic LB score (with post-processing P1+P2):  0.86310</p>\n\n<h2>What didn't work for us (relevant to this competition only):</h2>\n\n<ul>\n<li>Fixed cropping by 10-20%</li>\n<li>Pseudo-labeling</li>\n<li>Mixup</li>\n<li>Focal loss</li>\n</ul>",
      "rawMarkdown": "First of all special thanks to my teammate Tal :)\n## Backbone:\nAs a backbone, we used TResNet architecture (High Performance GPU-Dedicated Architecture) which allowed us to do fast training with large batches, while maintaining top scores along all resolutions  (more details in our [paper](https://arxiv.org/pdf/2003.13630.pdf) and code [here](https://github.com/mrT23/TResNet)).\nModels used in this competition: TResNet-M, TResNet-L.\n## BottleNeck Head:\nWe avoided having a massive classification layer: 2048x32094 (65.72M params), by using a bottleneck head that reduces the embedding feature vector from 2048 dims to 512 dims. total parameters= 2048x512 + 512x32094 (17.48M params).\nThis modification not only decreased the size of the network, but allowed us to use samples in larger batch sizes (Faster Training and helps the Soft-Triplet loss).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2Fd861d950b9f015672631c7b59815e5a5%2FBasicBlcok_Tresnet%20(1)%20copy.png?generation=1590586933112476&amp;alt=media)\n\n\n## Loss function\nWe used Cross-Entropy loss with label smoothing 0.2, and in addition we added Soft-Triplet Loss with Soft Margin which emphasizes hard examples by focusing of the most difficult positives and negatives samples in the batch, for more details check out our ReID [paper](https://arxiv.org/pdf/1910.07038.pdf) Section 3.1.\nOnline-Hard Negative mining works best when training with large batch size (more negative examples); hence, we leveraged the high memory utilization of our  Network and the bottleneck head. for example: with TResNet-M @ 448x448, we could reach maximum batch size of 128 on a single V100 16GB, significantly higher than other common networks like EfficientNet and ResNext.\n## Class Balancing\nHerbarium 2020 dataset suffers from prominent class im-balance, therefore we used a simple loss weighting, multiplying by the inverse of class frequency.\nNote: We used class balancing only in the last 15~ epochs of the training, to enable the network to learn the basic features first.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2F332b1711fa4eefa0b2de9984d491742d%2FScreen%20Shot%202020-05-27%20at%2016.46.38.png?generation=1590587199078597&amp;alt=media)\n\n## Augmentations\n• Squish instead of Crop\n• FlipLR\n• Cutout: 0.5\n#### Special augmentation scheme (+~0.5% on public LB):\n- Random Zoom-in (p=0.4, zoom_factor= random between 1-1.15)\n- Reduced affinity transformations: p_shear: 0.1, p_rotate=0.2, rotate_degrees=10\n- Decolorize: p_decolorize: 0.4\n- Rectangular resizing: small gains\n### TTA\nAveraging horizontal flip predictions\n## Our best single model:\nModel: TResNet-L\n\n**Stage 1:**\n100 epochs at 448x448, CE+SoftTriplet (class-balancing starting from epoch 85)\nPublic LB score (without  post-processing):       0.84089\nPublic LB score (with  post-processing P1+P2): 0.85093\n\n**Stage 2 (high resolution fine-tuning):**\n10 epoch fine-tune + class balancing on 544x416\nPublic LB score (without  post-processing):       0.84727\nPublic LB score (with  post-processing P1+P2): 0.85604\n\n## Post-Process (+~0.9% on the pubic LB)\nThe origanizers wrote in the description that:\n\"Each category has at least 1 instance in both the training and test datasets.\"\n\nBased on this prior knowledge we used two post-processing algorithms that tries to fill some of the empty classes.\n\n**P1: Switching classes based on low top-1 confidence (~0.9% on the public LB)**\nthe idea is basically whenever the network is not confident about its predictions, and if one of the empty classes is close to top1 prediction, then switch predictions:\n1) Go through all the top 2 predictions in the test set and if all these criteria are met switch between top1 and top2:\n• top2 &gt; top1 * m (m=0.3)\n• top2 class never appeared in the top1 predictions\n• top1 class appeared at least once as top2 prediction for another sample image in the test set.\n2) Repeat the same switching algorithm above for top 3, 4 and 5 (performed sequentially, first perform all of top2 switches and then top3 switches and so on).\n\n**P2: Filling empty classes (~+0.1-0.2% on the public LB)**\n1) Go through the predictions after the first post-process, and for all classes that don't appear in the top1 predictions and find the image sample with the highest prediction from all of the top5 predictions of the test set.\n## Pre-trained model\nWe used pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slightly better scores on the public LB. \n\n## Ensemble (+0.8% on public LB)\nWe used an ensemble of 9 models trained from different pretrains (imagenet and iNaturalist2018) and different resolutions, final resolutions: 544x416 mostly and some 620x476. \nPublic LB score (with post-processing P1+P2):  0.86310\n## What didn't work for us (relevant to this competition only):\n- Fixed cropping by 10-20%\n- Pseudo-labeling\n- Mixup\n- Focal loss",
      "votes": null
    },
    {
      "id": "863868",
      "postDate": "05/27/2020 16:03:22",
      "content": "<p>Great job! :)</p>",
      "rawMarkdown": "Great job! :)",
      "votes": null
    },
    {
      "id": "864064",
      "postDate": "05/27/2020 19:02:16",
      "content": "<p>Great post, thanks!</p>",
      "rawMarkdown": "Great post, thanks!",
      "votes": null
    },
    {
      "id": "864450",
      "postDate": "05/28/2020 02:20:09",
      "content": "<p>good post, thanks</p>",
      "rawMarkdown": "good post, thanks",
      "votes": null
    },
    {
      "id": "867971",
      "postDate": "05/30/2020 19:32:23",
      "content": "<p>Hi, <a href=\"/hussam789\">@hussam789</a> </p>\n\n<p>Would you mind to share how do you validate your model?</p>",
      "rawMarkdown": "Hi, @hussam789 \n\nWould you mind to share how do you validate your model?",
      "votes": null
    },
    {
      "id": "868446",
      "postDate": "05/31/2020 08:38:32",
      "content": "<p>Good question, At first we were using a stratified split of 80/20. but then when we started using class balanced loss weights, we used all of the training data since we want to include the few-shot classes (more than 10k classes have only 2-3 samples per class).\nwe didn't do k-fold splits since the dataset is relatively large and we didn't feel that it's necessary for this task.</p>",
      "rawMarkdown": "Good question, At first we were using a stratified split of 80/20. but then when we started using class balanced loss weights, we used all of the training data since we want to include the few-shot classes (more than 10k classes have only 2-3 samples per class).\nwe didn't do k-fold splits since the dataset is relatively large and we didn't feel that it's necessary for this task.",
      "votes": null
    },
    {
      "id": "868514",
      "postDate": "05/31/2020 09:54:27",
      "content": "<p>Oh, I see. But how do you track overfit then? I mean when you use whole train to validate</p>",
      "rawMarkdown": "Oh, I see. But how do you track overfit then? I mean when you use whole train to validate",
      "votes": null
    },
    {
      "id": "868657",
      "postDate": "05/31/2020 12:00:02",
      "content": "<p><a href=\"/discoholic\">@discoholic</a> before we ask our selves whether we \"overfit\" or not, first we should understand the metric we are trying to optimize. \n\"In macro F1 a separate F1 score is <strong>calculated for each species</strong> value and then <strong>averaged</strong>.\"</p>\n\n<p>what do we know about the test set?\n1. We know that each class is represented in the test set, so miss-classifying images of the few-shot classes can gravely reduce our \"macro F1\" score in the test set.\n2. We know that the test set is probably also im-balanced since we already know that the num of samples per class is capped at 10 and there is at least one image per class: 138296 / 32094 &lt; 10, so probably the train set and test set came from the same long-tail distribution (~probably)\n3. The test set is ~8 times smaller than the training set</p>\n\n<p>Given this information about the test set, then if we \"over-fit\" then it's probably because of the frequent classes in the training set are more \"dominant\" in the test set predictions than the few-shot classes.</p>\n\n<p>If the score was top-1 metric instead of macro f1 score then I would agree that there is a higher risk of \"over-fitting\" if we train on all of the training set.</p>\n\n<p>To answer your question, we used soft-target predictions of the test set from our best model (ensemble + TTA) and validated our score of individual training against it. and whoever we had a better Top model (with ensemble), we updated the soft-target predictions of the test set and now every run is validated against that (maybe its not ideal but it worked for us).\nour internal score and the public LB score had a good correlation, and we saw also a good correlation with the private LB score, our final submission was indeed our best private LB score.</p>",
      "rawMarkdown": "discoholic before we ask our selves whether we \"overfit\" or not, first we should understand the metric we are trying to optimize. \n\"In macro F1 a separate F1 score is **calculated for each species** value and then **averaged**.\"\n\nwhat do we know about the test set?\n1. We know that each class is represented in the test set, so miss-classifying images of the few-shot classes can gravely reduce our \"macro F1\" score in the test set.\n2. We know that the test set is probably also im-balanced since we already know that the num of samples per class is capped at 10 and there is at least one image per class: 138296 / 32094 &lt; 10, so probably the train set and test set came from the same long-tail distribution (~probably)\n3. The test set is ~8 times smaller than the training set\n\nGiven this information about the test set, then if we \"over-fit\" then it's probably because of the frequent classes in the training set are more \"dominant\" in the test set predictions than the few-shot classes.\n\nIf the score was top-1 metric instead of macro f1 score then I would agree that there is a higher risk of \"over-fitting\" if we train on all of the training set.\n\nTo answer your question, we used soft-target predictions of the test set from our best model (ensemble + TTA) and validated our score of individual training against it. and whoever we had a better Top model (with ensemble), we updated the soft-target predictions of the test set and now every run is validated against that (maybe its not ideal but it worked for us).\nour internal score and the public LB score had a good correlation, and we saw also a good correlation with the private LB score, our final submission was indeed our best private LB score.",
      "votes": null
    },
    {
      "id": "871414",
      "postDate": "06/02/2020 10:45:31",
      "content": "<p>Do I understand it correctly: you've trained some model, than made a prediction over all test set, called it soft-target, than validate next model against first?</p>",
      "rawMarkdown": "Do I understand it correctly: you've trained some model, than made a prediction over all test set, called it soft-target, than validate next model against first?",
      "votes": null
    },
    {
      "id": "876276",
      "postDate": "06/06/2020 15:39:24",
      "content": "<p>Yes, only the model used to produce the soft-targets is our current best model (Largest model) or ensemble of multiple models.\nThen validated the next single models and sometimes smaller models against the soft-targets</p>",
      "rawMarkdown": "Yes, only the model used to produce the soft-targets is our current best model (Largest model) or ensemble of multiple models.\nThen validated the next single models and sometimes smaller models against the soft-targets",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 863868,
      "author_name": "piantic",
      "author_url": "",
      "post_date": "05/27/2020 16:03:22",
      "content": "<p>Great job! :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 864064,
      "author_name": "asafnoy",
      "author_url": "",
      "post_date": "05/27/2020 19:02:16",
      "content": "<p>Great post, thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 864450,
      "author_name": "buaashijie",
      "author_url": "",
      "post_date": "05/28/2020 02:20:09",
      "content": "<p>good post, thanks</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 867971,
      "author_name": "discoholic",
      "author_url": "",
      "post_date": "05/30/2020 19:32:23",
      "content": "<p>Hi, <a href=\"/hussam789\">@hussam789</a> </p>\n\n<p>Would you mind to share how do you validate your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 868446,
          "author_name": "hussam789",
          "author_url": "",
          "post_date": "05/31/2020 08:38:32",
          "content": "<p>Good question, At first we were using a stratified split of 80/20. but then when we started using class balanced loss weights, we used all of the training data since we want to include the few-shot classes (more than 10k classes have only 2-3 samples per class).\nwe didn't do k-fold splits since the dataset is relatively large and we didn't feel that it's necessary for this task.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868514,
          "author_name": "discoholic",
          "author_url": "",
          "post_date": "05/31/2020 09:54:27",
          "content": "<p>Oh, I see. But how do you track overfit then? I mean when you use whole train to validate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 868657,
          "author_name": "hussam789",
          "author_url": "",
          "post_date": "05/31/2020 12:00:02",
          "content": "<p><a href=\"/discoholic\">@discoholic</a> before we ask our selves whether we \"overfit\" or not, first we should understand the metric we are trying to optimize. \n\"In macro F1 a separate F1 score is <strong>calculated for each species</strong> value and then <strong>averaged</strong>.\"</p>\n\n<p>what do we know about the test set?\n1. We know that each class is represented in the test set, so miss-classifying images of the few-shot classes can gravely reduce our \"macro F1\" score in the test set.\n2. We know that the test set is probably also im-balanced since we already know that the num of samples per class is capped at 10 and there is at least one image per class: 138296 / 32094 &lt; 10, so probably the train set and test set came from the same long-tail distribution (~probably)\n3. The test set is ~8 times smaller than the training set</p>\n\n<p>Given this information about the test set, then if we \"over-fit\" then it's probably because of the frequent classes in the training set are more \"dominant\" in the test set predictions than the few-shot classes.</p>\n\n<p>If the score was top-1 metric instead of macro f1 score then I would agree that there is a higher risk of \"over-fitting\" if we train on all of the training set.</p>\n\n<p>To answer your question, we used soft-target predictions of the test set from our best model (ensemble + TTA) and validated our score of individual training against it. and whoever we had a better Top model (with ensemble), we updated the soft-target predictions of the test set and now every run is validated against that (maybe its not ideal but it worked for us).\nour internal score and the public LB score had a good correlation, and we saw also a good correlation with the private LB score, our final submission was indeed our best private LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 871414,
          "author_name": "discoholic",
          "author_url": "",
          "post_date": "06/02/2020 10:45:31",
          "content": "<p>Do I understand it correctly: you've trained some model, than made a prediction over all test set, called it soft-target, than validate next model against first?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 876276,
          "author_name": "hussam789",
          "author_url": "",
          "post_date": "06/06/2020 15:39:24",
          "content": "<p>Yes, only the model used to produce the soft-targets is our current best model (Largest model) or ensemble of multiple models.\nThen validated the next single models and sometimes smaller models against the soft-targets</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "863710": "First of all special thanks to my teammate Tal :)\n## Backbone:\nAs a backbone, we used TResNet architecture (High Performance GPU-Dedicated Architecture) which allowed us to do fast training with large batches, while maintaining top scores along all resolutions  (more details in our [paper](https://arxiv.org/pdf/2003.13630.pdf) and code [here](https://github.com/mrT23/TResNet)).\nModels used in this competition: TResNet-M, TResNet-L.\n## BottleNeck Head:\nWe avoided having a massive classification layer: 2048x32094 (65.72M params), by using a bottleneck head that reduces the embedding feature vector from 2048 dims to 512 dims. total parameters= 2048x512 + 512x32094 (17.48M params).\nThis modification not only decreased the size of the network, but allowed us to use samples in larger batch sizes (Faster Training and helps the Soft-Triplet loss).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2Fd861d950b9f015672631c7b59815e5a5%2FBasicBlcok_Tresnet%20(1)%20copy.png?generation=1590586933112476&amp;alt=media)\n\n\n## Loss function\nWe used Cross-Entropy loss with label smoothing 0.2, and in addition we added Soft-Triplet Loss with Soft Margin which emphasizes hard examples by focusing of the most difficult positives and negatives samples in the batch, for more details check out our ReID [paper](https://arxiv.org/pdf/1910.07038.pdf) Section 3.1.\nOnline-Hard Negative mining works best when training with large batch size (more negative examples); hence, we leveraged the high memory utilization of our  Network and the bottleneck head. for example: with TResNet-M @ 448x448, we could reach maximum batch size of 128 on a single V100 16GB, significantly higher than other common networks like EfficientNet and ResNext.\n## Class Balancing\nHerbarium 2020 dataset suffers from prominent class im-balance, therefore we used a simple loss weighting, multiplying by the inverse of class frequency.\nNote: We used class balancing only in the last 15~ epochs of the training, to enable the network to learn the basic features first.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1578028%2F332b1711fa4eefa0b2de9984d491742d%2FScreen%20Shot%202020-05-27%20at%2016.46.38.png?generation=1590587199078597&amp;alt=media)\n\n## Augmentations\n• Squish instead of Crop\n• FlipLR\n• Cutout: 0.5\n#### Special augmentation scheme (+~0.5% on public LB):\n- Random Zoom-in (p=0.4, zoom_factor= random between 1-1.15)\n- Reduced affinity transformations: p_shear: 0.1, p_rotate=0.2, rotate_degrees=10\n- Decolorize: p_decolorize: 0.4\n- Rectangular resizing: small gains\n### TTA\nAveraging horizontal flip predictions\n## Our best single model:\nModel: TResNet-L\n\n**Stage 1:**\n100 epochs at 448x448, CE+SoftTriplet (class-balancing starting from epoch 85)\nPublic LB score (without  post-processing):       0.84089\nPublic LB score (with  post-processing P1+P2): 0.85093\n\n**Stage 2 (high resolution fine-tuning):**\n10 epoch fine-tune + class balancing on 544x416\nPublic LB score (without  post-processing):       0.84727\nPublic LB score (with  post-processing P1+P2): 0.85604\n\n## Post-Process (+~0.9% on the pubic LB)\nThe origanizers wrote in the description that:\n\"Each category has at least 1 instance in both the training and test datasets.\"\n\nBased on this prior knowledge we used two post-processing algorithms that tries to fill some of the empty classes.\n\n**P1: Switching classes based on low top-1 confidence (~0.9% on the public LB)**\nthe idea is basically whenever the network is not confident about its predictions, and if one of the empty classes is close to top1 prediction, then switch predictions:\n1) Go through all the top 2 predictions in the test set and if all these criteria are met switch between top1 and top2:\n• top2 &gt; top1 * m (m=0.3)\n• top2 class never appeared in the top1 predictions\n• top1 class appeared at least once as top2 prediction for another sample image in the test set.\n2) Repeat the same switching algorithm above for top 3, 4 and 5 (performed sequentially, first perform all of top2 switches and then top3 switches and so on).\n\n**P2: Filling empty classes (~+0.1-0.2% on the public LB)**\n1) Go through the predictions after the first post-process, and for all classes that don't appear in the top1 predictions and find the image sample with the highest prediction from all of the top5 predictions of the test set.\n## Pre-trained model\nWe used pre-trains from ImageNet and iNaturalist 2018. we got almost similar scores, with iNat18 having slightly better scores on the public LB. \n\n## Ensemble (+0.8% on public LB)\nWe used an ensemble of 9 models trained from different pretrains (imagenet and iNaturalist2018) and different resolutions, final resolutions: 544x416 mostly and some 620x476. \nPublic LB score (with post-processing P1+P2):  0.86310\n## What didn't work for us (relevant to this competition only):\n- Fixed cropping by 10-20%\n- Pseudo-labeling\n- Mixup\n- Focal loss",
    "863868": "Great job! :)",
    "864064": "Great post, thanks!",
    "864450": "good post, thanks",
    "867971": "Hi, @hussam789 \n\nWould you mind to share how do you validate your model?",
    "868446": "Good question, At first we were using a stratified split of 80/20. but then when we started using class balanced loss weights, we used all of the training data since we want to include the few-shot classes (more than 10k classes have only 2-3 samples per class).\nwe didn't do k-fold splits since the dataset is relatively large and we didn't feel that it's necessary for this task.",
    "868514": "Oh, I see. But how do you track overfit then? I mean when you use whole train to validate",
    "868657": "discoholic before we ask our selves whether we \"overfit\" or not, first we should understand the metric we are trying to optimize. \n\"In macro F1 a separate F1 score is **calculated for each species** value and then **averaged**.\"\n\nwhat do we know about the test set?\n1. We know that each class is represented in the test set, so miss-classifying images of the few-shot classes can gravely reduce our \"macro F1\" score in the test set.\n2. We know that the test set is probably also im-balanced since we already know that the num of samples per class is capped at 10 and there is at least one image per class: 138296 / 32094 &lt; 10, so probably the train set and test set came from the same long-tail distribution (~probably)\n3. The test set is ~8 times smaller than the training set\n\nGiven this information about the test set, then if we \"over-fit\" then it's probably because of the frequent classes in the training set are more \"dominant\" in the test set predictions than the few-shot classes.\n\nIf the score was top-1 metric instead of macro f1 score then I would agree that there is a higher risk of \"over-fitting\" if we train on all of the training set.\n\nTo answer your question, we used soft-target predictions of the test set from our best model (ensemble + TTA) and validated our score of individual training against it. and whoever we had a better Top model (with ensemble), we updated the soft-target predictions of the test set and now every run is validated against that (maybe its not ideal but it worked for us).\nour internal score and the public LB score had a good correlation, and we saw also a good correlation with the private LB score, our final submission was indeed our best private LB score.",
    "871414": "Do I understand it correctly: you've trained some model, than made a prediction over all test set, called it soft-target, than validate next model against first?",
    "876276": "Yes, only the model used to produce the soft-targets is our current best model (Largest model) or ensemble of multiple models.\nThen validated the next single models and sometimes smaller models against the soft-targets"
  },
  "source": "meta"
}