{
  "id": 319894,
  "title": "8th place solution [part]",
  "url": "/competitions/happy-whale-and-dolphin/discussion/319894",
  "author_name": "",
  "post_date": "2022-04-19T09:45:33.093980900Z",
  "votes": 17,
  "comment_count": 2,
  "views": 0,
  "content": "<p>First of all let me thank organizers and participants. That was interesting and quite tough three months.</p>\n<p><strong>Solutions from my teammates:</strong></p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319868\" target=\"_blank\">Aleksey</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319900\" target=\"_blank\">Oleg</a></li>\n</ol>\n<p><strong>TLDR</strong>: the more is better here. </p>\n<p><strong>Data</strong><br>\nFirst experiments were made on full images and it gives quite poor performance on local validation and on LB. The decision was to make some crops from the original image.<br>\nCrops were made for bodies and for fins using Yolo_v5.<br>\nAfter getting boxes for train and test data, I decided to train models on body crops where there were bodies and on fins otherwise.<br>\nDuring the experiments I decided to change this strategy. I make two datasets with shared individual_ids: in first dataset there were only body crops, and in second dataset only fin crops. So, for one individual there could be body and fin in separate images. If there were no detected objects on frame, I simply took full frame to the batch.<br>\nThis approach give me the boost in data amount and diversity and significantly raise val and LB scores.<br>\nAnother key feature from my perspective was training final models on all train data without splitting in on folds.<br>\nThis scheme is harder than one with folds, because I need to train two models, one for the validation and another for the submission. But since here we need do match test data with ALL train data, the performance of the models that see all the train individuals were much much better here.</p>\n<p><strong>Models</strong><br>\nI started with experiments on classic small person re-identification model <a href=\"https://github.com/KaiyangZhou/deep-person-reid/blob/master/torchreid/models/osnet_ain.py\" target=\"_blank\">OSNet</a>.  I add a small modification to this model - channel attention from <a href=\"https://arxiv.org/pdf/1905.12830.pdf\" target=\"_blank\">this awesome paper</a>. Despite that this architecture is very small, it can reach a competitive performance compare with the even bigger efficientnets_(b4,b5,b6).<br>\nOn that experiments I ended up with 600-800px image size and 2046-4096 feature size. Looks like the big images were critical here.<br>\nModels in my final ensemble: osnet_ain_attn_x1_0, timm_tf_efficientnet_b4_ns, timm_tf_efficientnet_b5_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b7_ns, timm_tf_efficientnetv2_m_in21ft1k, timm_convnext_base_384_in22ft1k, timm_tf_efficientnetv2_m_in21.</p>\n<p><strong>Losses</strong><br>\nIn this competition I try many losses, such as <a href=\"https://github.com/cvqluu/Angular-Penalty-Softmax-Losses-Pytorch\" target=\"_blank\">am-softmax</a>, <a href=\"https://github.com/ronghuaiyang/arcface-pytorch\" target=\"_blank\">arcface</a>, <a href=\"https://github.com/4uiiurz1/pytorch-adacos\" target=\"_blank\">adacos</a> end other CE-based losses. <br>\nThe best choice for me was AM-Softmax with m=0.35 and S=30.<br>\nThere was a simple trick which give me boost both on local validation and LB. As I wrote earlier, there were two datasets - bodies and fins. I make two separate losses - one for body samples and second for fin samples. The final loss was a mean of this two losses.<br>\nDuring the inference for each test sample I predict features for fin and body and make mean feature for them. This approach works better than single body or fin feature.</p>\n<p><strong>Train details</strong></p>\n<ol>\n<li>Epochs 25-40</li>\n<li>Optim Adam, LR 0.001</li>\n<li>Warmup for 5 ep + CosineAnnealingLR</li>\n<li>Augs: RandomBrightnessContrast, ColorJitter, IAAAdditiveGaussianNoise, GaussNoise, Blur, MotionBlur, ShiftScaleRotate, HorizontalFlip</li>\n</ol>\n<p><strong>Pseudolabels</strong><br>\nAnother huge boost came from pseudolabels. The strategy was simple: train good ensemble and take some percentage of the most confident predictions. My percentage of test for pseudo was 60% (~16k samples). This is very simple yet <strong>very</strong> effective.</p>\n<p><strong>Scores</strong><br>\nMy best ensemble score was 0.863 </p>\n<p><strong>Postprocessing</strong><br>\nIn this competition it was critical to choose best threshold for the new_individual class. But unfortunately it was very hard to find a way to do that.<br>\nAt the start of the competition (and some time after team merging) I use simple threshold for all the species. <br>\nNearly the competition deadline we start to find the threshold optimum. For us the best decision was to make species classifier, choose whales and dolphins (<a href=\"https://www.kaggle.com/code/kwentar/species-classification\" target=\"_blank\">based on Alexey's nb</a>) and define separate thresholds for whales and for dolphines, so that there were balance between new_individuals in this two groups.</p>\n<p><strong>Tricks that didn't help</strong></p>\n<ol>\n<li>TTA - no boost, yet not worse</li>\n<li>SWA - no boost at all</li>\n<li><a href=\"https://arxiv.org/abs/1701.08398\" target=\"_blank\">Re-Ranking</a>. There is one very effective postprocessing trick to improve mAP metric - Re-ranking. Unfortunately, in this competition it simply didn't work.</li>\n<li>Poolings different from GAP</li>\n<li>Tricky postprocessing based on the test images similarity.</li>\n<li>Things that I forgot to mention :)</li>\n</ol>",
  "messages": [
    {
      "id": "1760440",
      "postDate": "04/19/2022 09:45:33",
      "content": "<p>First of all let me thank organizers and participants. That was interesting and quite tough three months.</p>\n<p><strong>Solutions from my teammates:</strong></p>\n<ol>\n<li><a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319868\" target=\"_blank\">Aleksey</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319900\" target=\"_blank\">Oleg</a></li>\n</ol>\n<p><strong>TLDR</strong>: the more is better here. </p>\n<p><strong>Data</strong><br>\nFirst experiments were made on full images and it gives quite poor performance on local validation and on LB. The decision was to make some crops from the original image.<br>\nCrops were made for bodies and for fins using Yolo_v5.<br>\nAfter getting boxes for train and test data, I decided to train models on body crops where there were bodies and on fins otherwise.<br>\nDuring the experiments I decided to change this strategy. I make two datasets with shared individual_ids: in first dataset there were only body crops, and in second dataset only fin crops. So, for one individual there could be body and fin in separate images. If there were no detected objects on frame, I simply took full frame to the batch.<br>\nThis approach give me the boost in data amount and diversity and significantly raise val and LB scores.<br>\nAnother key feature from my perspective was training final models on all train data without splitting in on folds.<br>\nThis scheme is harder than one with folds, because I need to train two models, one for the validation and another for the submission. But since here we need do match test data with ALL train data, the performance of the models that see all the train individuals were much much better here.</p>\n<p><strong>Models</strong><br>\nI started with experiments on classic small person re-identification model <a href=\"https://github.com/KaiyangZhou/deep-person-reid/blob/master/torchreid/models/osnet_ain.py\" target=\"_blank\">OSNet</a>.  I add a small modification to this model - channel attention from <a href=\"https://arxiv.org/pdf/1905.12830.pdf\" target=\"_blank\">this awesome paper</a>. Despite that this architecture is very small, it can reach a competitive performance compare with the even bigger efficientnets_(b4,b5,b6).<br>\nOn that experiments I ended up with 600-800px image size and 2046-4096 feature size. Looks like the big images were critical here.<br>\nModels in my final ensemble: osnet_ain_attn_x1_0, timm_tf_efficientnet_b4_ns, timm_tf_efficientnet_b5_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b7_ns, timm_tf_efficientnetv2_m_in21ft1k, timm_convnext_base_384_in22ft1k, timm_tf_efficientnetv2_m_in21.</p>\n<p><strong>Losses</strong><br>\nIn this competition I try many losses, such as <a href=\"https://github.com/cvqluu/Angular-Penalty-Softmax-Losses-Pytorch\" target=\"_blank\">am-softmax</a>, <a href=\"https://github.com/ronghuaiyang/arcface-pytorch\" target=\"_blank\">arcface</a>, <a href=\"https://github.com/4uiiurz1/pytorch-adacos\" target=\"_blank\">adacos</a> end other CE-based losses. <br>\nThe best choice for me was AM-Softmax with m=0.35 and S=30.<br>\nThere was a simple trick which give me boost both on local validation and LB. As I wrote earlier, there were two datasets - bodies and fins. I make two separate losses - one for body samples and second for fin samples. The final loss was a mean of this two losses.<br>\nDuring the inference for each test sample I predict features for fin and body and make mean feature for them. This approach works better than single body or fin feature.</p>\n<p><strong>Train details</strong></p>\n<ol>\n<li>Epochs 25-40</li>\n<li>Optim Adam, LR 0.001</li>\n<li>Warmup for 5 ep + CosineAnnealingLR</li>\n<li>Augs: RandomBrightnessContrast, ColorJitter, IAAAdditiveGaussianNoise, GaussNoise, Blur, MotionBlur, ShiftScaleRotate, HorizontalFlip</li>\n</ol>\n<p><strong>Pseudolabels</strong><br>\nAnother huge boost came from pseudolabels. The strategy was simple: train good ensemble and take some percentage of the most confident predictions. My percentage of test for pseudo was 60% (~16k samples). This is very simple yet <strong>very</strong> effective.</p>\n<p><strong>Scores</strong><br>\nMy best ensemble score was 0.863 </p>\n<p><strong>Postprocessing</strong><br>\nIn this competition it was critical to choose best threshold for the new_individual class. But unfortunately it was very hard to find a way to do that.<br>\nAt the start of the competition (and some time after team merging) I use simple threshold for all the species. <br>\nNearly the competition deadline we start to find the threshold optimum. For us the best decision was to make species classifier, choose whales and dolphins (<a href=\"https://www.kaggle.com/code/kwentar/species-classification\" target=\"_blank\">based on Alexey's nb</a>) and define separate thresholds for whales and for dolphines, so that there were balance between new_individuals in this two groups.</p>\n<p><strong>Tricks that didn't help</strong></p>\n<ol>\n<li>TTA - no boost, yet not worse</li>\n<li>SWA - no boost at all</li>\n<li><a href=\"https://arxiv.org/abs/1701.08398\" target=\"_blank\">Re-Ranking</a>. There is one very effective postprocessing trick to improve mAP metric - Re-ranking. Unfortunately, in this competition it simply didn't work.</li>\n<li>Poolings different from GAP</li>\n<li>Tricky postprocessing based on the test images similarity.</li>\n<li>Things that I forgot to mention :)</li>\n</ol>",
      "rawMarkdown": "First of all let me thank organizers and participants. That was interesting and quite tough three months.\n\n**Solutions from my teammates:**\n1. [Aleksey](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319868)\n2. [Oleg](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319900)\n\n**TLDR**: the more is better here. \n\n**Data**\nFirst experiments were made on full images and it gives quite poor performance on local validation and on LB. The decision was to make some crops from the original image.\nCrops were made for bodies and for fins using Yolo_v5.\nAfter getting boxes for train and test data, I decided to train models on body crops where there were bodies and on fins otherwise.\nDuring the experiments I decided to change this strategy. I make two datasets with shared individual_ids: in first dataset there were only body crops, and in second dataset only fin crops. So, for one individual there could be body and fin in separate images. If there were no detected objects on frame, I simply took full frame to the batch.\nThis approach give me the boost in data amount and diversity and significantly raise val and LB scores.\nAnother key feature from my perspective was training final models on all train data without splitting in on folds.\nThis scheme is harder than one with folds, because I need to train two models, one for the validation and another for the submission. But since here we need do match test data with ALL train data, the performance of the models that see all the train individuals were much much better here.\n\n**Models**\nI started with experiments on classic small person re-identification model [OSNet](https://github.com/KaiyangZhou/deep-person-reid/blob/master/torchreid/models/osnet_ain.py).  I add a small modification to this model - channel attention from [this awesome paper](https://arxiv.org/pdf/1905.12830.pdf). Despite that this architecture is very small, it can reach a competitive performance compare with the even bigger efficientnets_(b4,b5,b6).\nOn that experiments I ended up with 600-800px image size and 2046-4096 feature size. Looks like the big images were critical here.\nModels in my final ensemble: osnet_ain_attn_x1_0, timm_tf_efficientnet_b4_ns, timm_tf_efficientnet_b5_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b7_ns, timm_tf_efficientnetv2_m_in21ft1k, timm_convnext_base_384_in22ft1k, timm_tf_efficientnetv2_m_in21.\n\n**Losses**\nIn this competition I try many losses, such as [am-softmax](https://github.com/cvqluu/Angular-Penalty-Softmax-Losses-Pytorch), [arcface](https://github.com/ronghuaiyang/arcface-pytorch), [adacos](https://github.com/4uiiurz1/pytorch-adacos) end other CE-based losses. \nThe best choice for me was AM-Softmax with m=0.35 and S=30.\nThere was a simple trick which give me boost both on local validation and LB. As I wrote earlier, there were two datasets - bodies and fins. I make two separate losses - one for body samples and second for fin samples. The final loss was a mean of this two losses.\nDuring the inference for each test sample I predict features for fin and body and make mean feature for them. This approach works better than single body or fin feature.\n\n**Train details**\n1. Epochs 25-40\n2. Optim Adam, LR 0.001\n3. Warmup for 5 ep + CosineAnnealingLR\n4. Augs: RandomBrightnessContrast, ColorJitter, IAAAdditiveGaussianNoise, GaussNoise, Blur, MotionBlur, ShiftScaleRotate, HorizontalFlip\n\n**Pseudolabels**\nAnother huge boost came from pseudolabels. The strategy was simple: train good ensemble and take some percentage of the most confident predictions. My percentage of test for pseudo was 60% (~16k samples). This is very simple yet **very** effective.\n\n**Scores**\nMy best ensemble score was 0.863 \n\n**Postprocessing**\nIn this competition it was critical to choose best threshold for the new_individual class. But unfortunately it was very hard to find a way to do that.\nAt the start of the competition (and some time after team merging) I use simple threshold for all the species. \nNearly the competition deadline we start to find the threshold optimum. For us the best decision was to make species classifier, choose whales and dolphins ([based on Alexey's nb](https://www.kaggle.com/code/kwentar/species-classification)) and define separate thresholds for whales and for dolphines, so that there were balance between new_individuals in this two groups.\n\n**Tricks that didn't help**\n1. TTA - no boost, yet not worse\n2. SWA - no boost at all\n3. [Re-Ranking](https://arxiv.org/abs/1701.08398). There is one very effective postprocessing trick to improve mAP metric - Re-ranking. Unfortunately, in this competition it simply didn't work.\n4. Poolings different from GAP\n5. Tricky postprocessing based on the test images similarity.\n6. Things that I forgot to mention :)",
      "votes": null
    },
    {
      "id": "1760455",
      "postDate": "04/19/2022 10:02:54",
      "content": "<p>what kind of image data set you used …good read section Tricks that didn't help  is a better tool to save time …</p>",
      "rawMarkdown": "what kind of image data set you used ...good read section Tricks that didn't help  is a better tool to save time ...",
      "votes": null
    },
    {
      "id": "1760460",
      "postDate": "04/19/2022 10:07:30",
      "content": "<p>I use default dataset from this competition, without external data.<br>\nAt first try to figure out how to use data from previous whales competition, but for me it was useless.</p>",
      "rawMarkdown": "I use default dataset from this competition, without external data.\nAt first try to figure out how to use data from previous whales competition, but for me it was useless.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1760455,
      "author_name": "vidiqlogy",
      "author_url": "",
      "post_date": "04/19/2022 10:02:54",
      "content": "<p>what kind of image data set you used …good read section Tricks that didn't help  is a better tool to save time …</p>",
      "votes": null,
      "replies": [
        {
          "id": 1760460,
          "author_name": "ilyadobrynin",
          "author_url": "",
          "post_date": "04/19/2022 10:07:30",
          "content": "<p>I use default dataset from this competition, without external data.<br>\nAt first try to figure out how to use data from previous whales competition, but for me it was useless.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1760440": "First of all let me thank organizers and participants. That was interesting and quite tough three months.\n\n**Solutions from my teammates:**\n1. [Aleksey](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319868)\n2. [Oleg](https://www.kaggle.com/competitions/happy-whale-and-dolphin/discussion/319900)\n\n**TLDR**: the more is better here. \n\n**Data**\nFirst experiments were made on full images and it gives quite poor performance on local validation and on LB. The decision was to make some crops from the original image.\nCrops were made for bodies and for fins using Yolo_v5.\nAfter getting boxes for train and test data, I decided to train models on body crops where there were bodies and on fins otherwise.\nDuring the experiments I decided to change this strategy. I make two datasets with shared individual_ids: in first dataset there were only body crops, and in second dataset only fin crops. So, for one individual there could be body and fin in separate images. If there were no detected objects on frame, I simply took full frame to the batch.\nThis approach give me the boost in data amount and diversity and significantly raise val and LB scores.\nAnother key feature from my perspective was training final models on all train data without splitting in on folds.\nThis scheme is harder than one with folds, because I need to train two models, one for the validation and another for the submission. But since here we need do match test data with ALL train data, the performance of the models that see all the train individuals were much much better here.\n\n**Models**\nI started with experiments on classic small person re-identification model [OSNet](https://github.com/KaiyangZhou/deep-person-reid/blob/master/torchreid/models/osnet_ain.py).  I add a small modification to this model - channel attention from [this awesome paper](https://arxiv.org/pdf/1905.12830.pdf). Despite that this architecture is very small, it can reach a competitive performance compare with the even bigger efficientnets_(b4,b5,b6).\nOn that experiments I ended up with 600-800px image size and 2046-4096 feature size. Looks like the big images were critical here.\nModels in my final ensemble: osnet_ain_attn_x1_0, timm_tf_efficientnet_b4_ns, timm_tf_efficientnet_b5_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b6_ns, timm_tf_efficientnet_b7_ns, timm_tf_efficientnetv2_m_in21ft1k, timm_convnext_base_384_in22ft1k, timm_tf_efficientnetv2_m_in21.\n\n**Losses**\nIn this competition I try many losses, such as [am-softmax](https://github.com/cvqluu/Angular-Penalty-Softmax-Losses-Pytorch), [arcface](https://github.com/ronghuaiyang/arcface-pytorch), [adacos](https://github.com/4uiiurz1/pytorch-adacos) end other CE-based losses. \nThe best choice for me was AM-Softmax with m=0.35 and S=30.\nThere was a simple trick which give me boost both on local validation and LB. As I wrote earlier, there were two datasets - bodies and fins. I make two separate losses - one for body samples and second for fin samples. The final loss was a mean of this two losses.\nDuring the inference for each test sample I predict features for fin and body and make mean feature for them. This approach works better than single body or fin feature.\n\n**Train details**\n1. Epochs 25-40\n2. Optim Adam, LR 0.001\n3. Warmup for 5 ep + CosineAnnealingLR\n4. Augs: RandomBrightnessContrast, ColorJitter, IAAAdditiveGaussianNoise, GaussNoise, Blur, MotionBlur, ShiftScaleRotate, HorizontalFlip\n\n**Pseudolabels**\nAnother huge boost came from pseudolabels. The strategy was simple: train good ensemble and take some percentage of the most confident predictions. My percentage of test for pseudo was 60% (~16k samples). This is very simple yet **very** effective.\n\n**Scores**\nMy best ensemble score was 0.863 \n\n**Postprocessing**\nIn this competition it was critical to choose best threshold for the new_individual class. But unfortunately it was very hard to find a way to do that.\nAt the start of the competition (and some time after team merging) I use simple threshold for all the species. \nNearly the competition deadline we start to find the threshold optimum. For us the best decision was to make species classifier, choose whales and dolphins ([based on Alexey's nb](https://www.kaggle.com/code/kwentar/species-classification)) and define separate thresholds for whales and for dolphines, so that there were balance between new_individuals in this two groups.\n\n**Tricks that didn't help**\n1. TTA - no boost, yet not worse\n2. SWA - no boost at all\n3. [Re-Ranking](https://arxiv.org/abs/1701.08398). There is one very effective postprocessing trick to improve mAP metric - Re-ranking. Unfortunately, in this competition it simply didn't work.\n4. Poolings different from GAP\n5. Tricky postprocessing based on the test images similarity.\n6. Things that I forgot to mention :)",
    "1760455": "what kind of image data set you used ...good read section Tricks that didn't help  is a better tool to save time ...",
    "1760460": "I use default dataset from this competition, without external data.\nAt first try to figure out how to use data from previous whales competition, but for me it was useless."
  },
  "source": "meta"
}