{
  "id": 82427,
  "title": "9th place solution or how we spent last one and a half month",
  "url": "/competitions/humpback-whale-identification/writeups/ods-ai-bratannet-9th-place-solution-or-how-we-spen",
  "author_name": "",
  "post_date": "2019-03-01T09:37:35.985571100Z",
  "votes": 38,
  "comment_count": 8,
  "views": 0,
  "content": "<p><strong>TL;DR</strong> Adam, Cosine with restarts, CosFace, ArcFace, High-resolution images, Weighted sampling, new_whale distillation, Pseudo labeled test, Resnet34, BNInception, Densenet121, AutoAugment, CoordConv, GAPNet </p>\n\n<p>We’d like to share our solution as a story how we gradually improve our models. </p>\n\n<p>But first of all I’d like to thank my teammate <a href=\"https://www.kaggle.com/vlad0922\">Vladislav</a> for fruitful collaboration, kaggle community for motivation and ods.ai for kind support:)</p>\n\n<p>To start with, it was obvious idea to consider whale’s flukes as human faces. Fortunately there are tons of papers for face identification, re-identification and verification. </p>\n\n<p>So in the beginning of this competition this paper <a href=\"https://arxiv.org/abs/1804.06655\">https://arxiv.org/abs/1804.06655</a> helped us a lot. It features comprehensive survey of state-of-the-art face identification techniques. \nAccording to it softmax-based losses look really promising. Due to their classification nature and the fact that we already have classification pipeline from Protein Atlas and Draw challenges we decided to focus on them. </p>\n\n<p>Among others <strong>Cosface</strong> and <strong>Arcface</strong> stand out as newly discovered SOTA for face recognition task. The main idea is to bring examples of the same class close to each other in cosine similarity space and to pull apart distinct classes. Training with cosface or arcface generally is classification, so the final loss was CrossEntropy. One can read more details in their papers: <a href=\"https://arxiv.org/abs/1801.09414\">https://arxiv.org/abs/1801.09414</a>, <a href=\"https://arxiv.org/abs/1801.07698\">https://arxiv.org/abs/1801.07698</a>. After we train cosface or arcface net we took embeddings and calculate cosine similarity between train and test images. Then average similarities for each class in train and took 5 most similar.</p>\n\n<p>In the beginning of every competition you should always devise robust validation procedure. We did it poorly. But nevertheless, to do so we select about 1000 sample from classes with number of instances greater than 3, one example for each class. Also we chose about the same number of new_whale images. This setup show good correlation between local score and public LB score. Unfortunately threshold for new_whale derived from local validation was slightly biased. That was really bad, because threshold was unreliable. Another way for threshold determination was to adjust it so top-1 new_whale percentage should be around 30%. </p>\n\n<p>To keep up with kaggle community in conquering LB we decided to construct our pipeline in the following way:</p>\n\n<ol>\n<li>Decrease training time of model as much as possible</li>\n<li>Test a lot of hypothesis as much as possible</li>\n</ol>\n\n<p>To do that we restrict image size to 256x256 and number of epochs up to 64. This kind of restrictions gave us a model that can be trained in 2 hours on 1080ti or even faster on 2080ti. This setting let us iterate quickly in testing new hypothesis or optimizing hyperparameters. After we have established our general training we start endless array of experiments for low-res images. </p>\n\n<p>Let’s divide our experiments into two broad groups:</p>\n\n<ol>\n<li>Model engineering: what to train</li>\n<li>Training engineering: how to train</li>\n</ol>\n\n<p>Here comes Model engineering. We started with some heavy encoders such as inceptionv4, seresnext50 etc. But it appeared that for us in classification task they seem to overfit a lot. Then we decided to main some light networks such as resnet34, bninception and densenet121. After several competitions I begin to realise that sometimes when you don’t have much data light encoders may really boost your score. Like they don’t tend to overfit much to rare classes and label noise. This is just a hypothesis that need to be carefully verified.</p>\n\n<p>To get final models after initial 64 epochs on 256x256 images we increase image size up to 1024 for resnet34, up to 512 for bninception and up to 640 for densenet121 and train for 64 epochs more. </p>\n\n<p>To boost model performant we tried a lot of modification. According to our findings <strong>CoordConv</strong> <a href=\"https://arxiv.org/abs/1807.03247\">https://arxiv.org/abs/1807.03247</a> and <strong>GapNet</strong> architecture <a href=\"https://openreview.net/forum?id=ryl5khRcKm\">https://openreview.net/forum?id=ryl5khRcKm</a> helped to improve resnet34 score. Unfortunately we didn’t have time to test this mods on bninception and densenet121. Also adding some sophisticated convolution blocks to our nets didn’t help. So Squeeze-and-Excitation, Convolutional Block Attention Module didn’t help. That was a sad story because a lot of time was spent on trying to optimize model architecture instead of optimizing training itself. </p>\n\n<p>When it comes to training one of the first things that comes to mind is how not to overfit to training data. Especially when one working with zero and few-shot learning. Inspired by AutoAugment paper <a href=\"https://arxiv.org/abs/1805.09501\">https://arxiv.org/abs/1805.09501</a> we search augmentation space by random sampling and came up with the following augmentations:</p>\n\n<ol>\n<li>HorizontalFlip</li>\n<li>Rotate with 16 degree limit</li>\n<li>ShiftScaleRotate with 16 degree limit</li>\n<li>RandomBrightnessContrast</li>\n<li>RandomGamma</li>\n<li>Blur</li>\n<li>Perspective transform: tile left, right and corner</li>\n<li>Shear</li>\n<li>MotionBlur</li>\n<li>GridDistortion</li>\n<li>ElasticTransform</li>\n<li>Cutout</li>\n</ol>\n\n<p>Those augmentations we took from albumentations <a href=\"https://github.com/albu/albumentations\">https://github.com/albu/albumentations</a> and Augmentor <a href=\"https://github.com/mdbloice/Augmentor\">https://github.com/mdbloice/Augmentor</a> modules. Firstly we thought that this is to much for our networks. But actually for our models it was essential part not to overfit to training data.</p>\n\n<p>Cosface and Arcface parameters was optimised as well. Cosface: S = 32.0, M=0.35. Arcface: M1 = 1.0, M2 = 0.4, M3 = 0.15.</p>\n\n<p>We experimented a lot with optimizers and their hyperparameters: Adam, AdamW, SGD, SGDW. But the best optimizer for us appeared to be good old <strong>Adam with Cosine annealing</strong>.</p>\n\n<p>In the end we tried different kinds of TTA, but it didn’t help to improve the score. \nMixed precision learning didn’t show good results either. </p>\n\n<p>Starting from the beginning we realised that it is essential to do something with new_whales in order to incorporate them into training process. Simple solution was to assign each new_whale probability of each class as 1 / 5004. With help of weighted sampling technique it gave us some boost. But then we realised why don’t we use softmax predictions for new_whales derived from trained ensemble. So we came up with <strong>distillation</strong>. We choose distillation instead of pseudo labels, because new_whale is considered to have different labels from train labels. Though it might not to be really true. </p>\n\n<p>To further boost model capabilities we add test images with <strong>pseudo labels</strong> into train. Eventually our single model can hit 0.958 with snapshot ensembling. Unfortunately ensembling of models trained in this way didn’t give score improvement. Maybe it was due to less variety because of pseudo labels and distillation. </p>\n\n<p>In the end I should mention that this competition was really interesting and give us an opportunity to develop face/fluke recognition skills. Thank your Kaggle!</p>",
  "messages": [
    {
      "id": "481346",
      "postDate": "03/01/2019 09:37:35",
      "content": "<p><strong>TL;DR</strong> Adam, Cosine with restarts, CosFace, ArcFace, High-resolution images, Weighted sampling, new_whale distillation, Pseudo labeled test, Resnet34, BNInception, Densenet121, AutoAugment, CoordConv, GAPNet </p>\n\n<p>We’d like to share our solution as a story how we gradually improve our models. </p>\n\n<p>But first of all I’d like to thank my teammate <a href=\"https://www.kaggle.com/vlad0922\">Vladislav</a> for fruitful collaboration, kaggle community for motivation and ods.ai for kind support:)</p>\n\n<p>To start with, it was obvious idea to consider whale’s flukes as human faces. Fortunately there are tons of papers for face identification, re-identification and verification. </p>\n\n<p>So in the beginning of this competition this paper <a href=\"https://arxiv.org/abs/1804.06655\">https://arxiv.org/abs/1804.06655</a> helped us a lot. It features comprehensive survey of state-of-the-art face identification techniques. \nAccording to it softmax-based losses look really promising. Due to their classification nature and the fact that we already have classification pipeline from Protein Atlas and Draw challenges we decided to focus on them. </p>\n\n<p>Among others <strong>Cosface</strong> and <strong>Arcface</strong> stand out as newly discovered SOTA for face recognition task. The main idea is to bring examples of the same class close to each other in cosine similarity space and to pull apart distinct classes. Training with cosface or arcface generally is classification, so the final loss was CrossEntropy. One can read more details in their papers: <a href=\"https://arxiv.org/abs/1801.09414\">https://arxiv.org/abs/1801.09414</a>, <a href=\"https://arxiv.org/abs/1801.07698\">https://arxiv.org/abs/1801.07698</a>. After we train cosface or arcface net we took embeddings and calculate cosine similarity between train and test images. Then average similarities for each class in train and took 5 most similar.</p>\n\n<p>In the beginning of every competition you should always devise robust validation procedure. We did it poorly. But nevertheless, to do so we select about 1000 sample from classes with number of instances greater than 3, one example for each class. Also we chose about the same number of new_whale images. This setup show good correlation between local score and public LB score. Unfortunately threshold for new_whale derived from local validation was slightly biased. That was really bad, because threshold was unreliable. Another way for threshold determination was to adjust it so top-1 new_whale percentage should be around 30%. </p>\n\n<p>To keep up with kaggle community in conquering LB we decided to construct our pipeline in the following way:</p>\n\n<ol>\n<li>Decrease training time of model as much as possible</li>\n<li>Test a lot of hypothesis as much as possible</li>\n</ol>\n\n<p>To do that we restrict image size to 256x256 and number of epochs up to 64. This kind of restrictions gave us a model that can be trained in 2 hours on 1080ti or even faster on 2080ti. This setting let us iterate quickly in testing new hypothesis or optimizing hyperparameters. After we have established our general training we start endless array of experiments for low-res images. </p>\n\n<p>Let’s divide our experiments into two broad groups:</p>\n\n<ol>\n<li>Model engineering: what to train</li>\n<li>Training engineering: how to train</li>\n</ol>\n\n<p>Here comes Model engineering. We started with some heavy encoders such as inceptionv4, seresnext50 etc. But it appeared that for us in classification task they seem to overfit a lot. Then we decided to main some light networks such as resnet34, bninception and densenet121. After several competitions I begin to realise that sometimes when you don’t have much data light encoders may really boost your score. Like they don’t tend to overfit much to rare classes and label noise. This is just a hypothesis that need to be carefully verified.</p>\n\n<p>To get final models after initial 64 epochs on 256x256 images we increase image size up to 1024 for resnet34, up to 512 for bninception and up to 640 for densenet121 and train for 64 epochs more. </p>\n\n<p>To boost model performant we tried a lot of modification. According to our findings <strong>CoordConv</strong> <a href=\"https://arxiv.org/abs/1807.03247\">https://arxiv.org/abs/1807.03247</a> and <strong>GapNet</strong> architecture <a href=\"https://openreview.net/forum?id=ryl5khRcKm\">https://openreview.net/forum?id=ryl5khRcKm</a> helped to improve resnet34 score. Unfortunately we didn’t have time to test this mods on bninception and densenet121. Also adding some sophisticated convolution blocks to our nets didn’t help. So Squeeze-and-Excitation, Convolutional Block Attention Module didn’t help. That was a sad story because a lot of time was spent on trying to optimize model architecture instead of optimizing training itself. </p>\n\n<p>When it comes to training one of the first things that comes to mind is how not to overfit to training data. Especially when one working with zero and few-shot learning. Inspired by AutoAugment paper <a href=\"https://arxiv.org/abs/1805.09501\">https://arxiv.org/abs/1805.09501</a> we search augmentation space by random sampling and came up with the following augmentations:</p>\n\n<ol>\n<li>HorizontalFlip</li>\n<li>Rotate with 16 degree limit</li>\n<li>ShiftScaleRotate with 16 degree limit</li>\n<li>RandomBrightnessContrast</li>\n<li>RandomGamma</li>\n<li>Blur</li>\n<li>Perspective transform: tile left, right and corner</li>\n<li>Shear</li>\n<li>MotionBlur</li>\n<li>GridDistortion</li>\n<li>ElasticTransform</li>\n<li>Cutout</li>\n</ol>\n\n<p>Those augmentations we took from albumentations <a href=\"https://github.com/albu/albumentations\">https://github.com/albu/albumentations</a> and Augmentor <a href=\"https://github.com/mdbloice/Augmentor\">https://github.com/mdbloice/Augmentor</a> modules. Firstly we thought that this is to much for our networks. But actually for our models it was essential part not to overfit to training data.</p>\n\n<p>Cosface and Arcface parameters was optimised as well. Cosface: S = 32.0, M=0.35. Arcface: M1 = 1.0, M2 = 0.4, M3 = 0.15.</p>\n\n<p>We experimented a lot with optimizers and their hyperparameters: Adam, AdamW, SGD, SGDW. But the best optimizer for us appeared to be good old <strong>Adam with Cosine annealing</strong>.</p>\n\n<p>In the end we tried different kinds of TTA, but it didn’t help to improve the score. \nMixed precision learning didn’t show good results either. </p>\n\n<p>Starting from the beginning we realised that it is essential to do something with new_whales in order to incorporate them into training process. Simple solution was to assign each new_whale probability of each class as 1 / 5004. With help of weighted sampling technique it gave us some boost. But then we realised why don’t we use softmax predictions for new_whales derived from trained ensemble. So we came up with <strong>distillation</strong>. We choose distillation instead of pseudo labels, because new_whale is considered to have different labels from train labels. Though it might not to be really true. </p>\n\n<p>To further boost model capabilities we add test images with <strong>pseudo labels</strong> into train. Eventually our single model can hit 0.958 with snapshot ensembling. Unfortunately ensembling of models trained in this way didn’t give score improvement. Maybe it was due to less variety because of pseudo labels and distillation. </p>\n\n<p>In the end I should mention that this competition was really interesting and give us an opportunity to develop face/fluke recognition skills. Thank your Kaggle!</p>",
      "rawMarkdown": "**TL;DR** Adam, Cosine with restarts, CosFace, ArcFace, High-resolution images, Weighted sampling, new_whale distillation, Pseudo labeled test, Resnet34, BNInception, Densenet121, AutoAugment, CoordConv, GAPNet \n\nWe’d like to share our solution as a story how we gradually improve our models. \n\nBut first of all I’d like to thank my teammate [Vladislav][1] for fruitful collaboration, kaggle community for motivation and ods.ai for kind support:)\n\nTo start with, it was obvious idea to consider whale’s flukes as human faces. Fortunately there are tons of papers for face identification, re-identification and verification. \n\nSo in the beginning of this competition this paper https://arxiv.org/abs/1804.06655 helped us a lot. It features comprehensive survey of state-of-the-art face identification techniques. \nAccording to it softmax-based losses look really promising. Due to their classification nature and the fact that we already have classification pipeline from Protein Atlas and Draw challenges we decided to focus on them. \n\nAmong others **Cosface** and **Arcface** stand out as newly discovered SOTA for face recognition task. The main idea is to bring examples of the same class close to each other in cosine similarity space and to pull apart distinct classes. Training with cosface or arcface generally is classification, so the final loss was CrossEntropy. One can read more details in their papers: https://arxiv.org/abs/1801.09414, https://arxiv.org/abs/1801.07698. After we train cosface or arcface net we took embeddings and calculate cosine similarity between train and test images. Then average similarities for each class in train and took 5 most similar.\n\nIn the beginning of every competition you should always devise robust validation procedure. We did it poorly. But nevertheless, to do so we select about 1000 sample from classes with number of instances greater than 3, one example for each class. Also we chose about the same number of new_whale images. This setup show good correlation between local score and public LB score. Unfortunately threshold for new_whale derived from local validation was slightly biased. That was really bad, because threshold was unreliable. Another way for threshold determination was to adjust it so top-1 new_whale percentage should be around 30%. \n\nTo keep up with kaggle community in conquering LB we decided to construct our pipeline in the following way:\n\n 1. Decrease training time of model as much as possible\n 2. Test a lot of hypothesis as much as possible\n\nTo do that we restrict image size to 256x256 and number of epochs up to 64. This kind of restrictions gave us a model that can be trained in 2 hours on 1080ti or even faster on 2080ti. This setting let us iterate quickly in testing new hypothesis or optimizing hyperparameters. After we have established our general training we start endless array of experiments for low-res images. \n\nLet’s divide our experiments into two broad groups:\n\n 1. Model engineering: what to train\n 2. Training engineering: how to train\n\nHere comes Model engineering. We started with some heavy encoders such as inceptionv4, seresnext50 etc. But it appeared that for us in classification task they seem to overfit a lot. Then we decided to main some light networks such as resnet34, bninception and densenet121. After several competitions I begin to realise that sometimes when you don’t have much data light encoders may really boost your score. Like they don’t tend to overfit much to rare classes and label noise. This is just a hypothesis that need to be carefully verified.\n\nTo get final models after initial 64 epochs on 256x256 images we increase image size up to 1024 for resnet34, up to 512 for bninception and up to 640 for densenet121 and train for 64 epochs more. \n\nTo boost model performant we tried a lot of modification. According to our findings **CoordConv** https://arxiv.org/abs/1807.03247 and **GapNet** architecture https://openreview.net/forum?id=ryl5khRcKm helped to improve resnet34 score. Unfortunately we didn’t have time to test this mods on bninception and densenet121. Also adding some sophisticated convolution blocks to our nets didn’t help. So Squeeze-and-Excitation, Convolutional Block Attention Module didn’t help. That was a sad story because a lot of time was spent on trying to optimize model architecture instead of optimizing training itself. \n\nWhen it comes to training one of the first things that comes to mind is how not to overfit to training data. Especially when one working with zero and few-shot learning. Inspired by AutoAugment paper https://arxiv.org/abs/1805.09501 we search augmentation space by random sampling and came up with the following augmentations:\n\n 1. HorizontalFlip\n 2. Rotate with 16 degree limit\n 3. ShiftScaleRotate with 16 degree limit\n 4. RandomBrightnessContrast\n 5. RandomGamma\n 6. Blur\n 7. Perspective transform: tile left, right and corner\n 8. Shear\n 9. MotionBlur\n 10. GridDistortion\n 11. ElasticTransform\n 12. Cutout\n\nThose augmentations we took from albumentations https://github.com/albu/albumentations and Augmentor https://github.com/mdbloice/Augmentor modules. Firstly we thought that this is to much for our networks. But actually for our models it was essential part not to overfit to training data.\n\nCosface and Arcface parameters was optimised as well. Cosface: S = 32.0, M=0.35. Arcface: M1 = 1.0, M2 = 0.4, M3 = 0.15.\n\nWe experimented a lot with optimizers and their hyperparameters: Adam, AdamW, SGD, SGDW. But the best optimizer for us appeared to be good old **Adam with Cosine annealing**.\n\nIn the end we tried different kinds of TTA, but it didn’t help to improve the score. \nMixed precision learning didn’t show good results either. \n\nStarting from the beginning we realised that it is essential to do something with new_whales in order to incorporate them into training process. Simple solution was to assign each new_whale probability of each class as 1 / 5004. With help of weighted sampling technique it gave us some boost. But then we realised why don’t we use softmax predictions for new_whales derived from trained ensemble. So we came up with **distillation**. We choose distillation instead of pseudo labels, because new_whale is considered to have different labels from train labels. Though it might not to be really true. \n\nTo further boost model capabilities we add test images with **pseudo labels** into train. Eventually our single model can hit 0.958 with snapshot ensembling. Unfortunately ensembling of models trained in this way didn’t give score improvement. Maybe it was due to less variety because of pseudo labels and distillation. \n\nIn the end I should mention that this competition was really interesting and give us an opportunity to develop face/fluke recognition skills. Thank your Kaggle!\n\t\n\t\n\n\n  [1]: https://www.kaggle.com/vlad0922",
      "votes": null
    },
    {
      "id": "481348",
      "postDate": "03/01/2019 09:40:04",
      "content": "<p>Congratulations <a href=\"/sawseen\">@sawseen</a>. Hard work paid finally.</p>",
      "rawMarkdown": "Congratulations @sawseen. Hard work paid finally.",
      "votes": null
    },
    {
      "id": "481352",
      "postDate": "03/01/2019 09:49:59",
      "content": "<p>I like, how you call 4 year old algorithm \"good old Adam\"))))</p>",
      "rawMarkdown": "I like, how you call 4 year old algorithm \"good old Adam\"))))",
      "votes": null
    },
    {
      "id": "481358",
      "postDate": "03/01/2019 09:55:04",
      "content": "<p>Congrats on the great result! Great write up :) Really like the set up you had with quick iterations to test hypotheses.</p>",
      "rawMarkdown": "Congrats on the great result! Great write up :) Really like the set up you had with quick iterations to test hypotheses.",
      "votes": null
    },
    {
      "id": "481394",
      "postDate": "03/01/2019 10:47:12",
      "content": "<p>Congrats <a href=\"/sawseen\">@sawseen</a> and thanks for sharing your solution overview.</p>",
      "rawMarkdown": "Congrats @sawseen and thanks for sharing your solution overview.",
      "votes": null
    },
    {
      "id": "481461",
      "postDate": "03/01/2019 12:32:40",
      "content": "<p>Thank you! As what kaggle have taught me is to iterate as fast as possible and fully utilize your hardware. </p>",
      "rawMarkdown": "Thank you! As what kaggle have taught me is to iterate as fast as possible and fully utilize your hardware.",
      "votes": null
    },
    {
      "id": "481463",
      "postDate": "03/01/2019 12:33:44",
      "content": "<p>4 years in deep learning seems like an eternity:)</p>",
      "rawMarkdown": "4 years in deep learning seems like an eternity:)",
      "votes": null
    },
    {
      "id": "481833",
      "postDate": "03/01/2019 22:37:34",
      "content": "<p>Great solution, congrats and thanks for sharing.</p>",
      "rawMarkdown": "Great solution, congrats and thanks for sharing.",
      "votes": null
    },
    {
      "id": "505626",
      "postDate": "04/02/2019 10:09:22",
      "content": "<p>Hi Ivan, \n\" So we came up with distillation\" -- &gt; What exactly distillation mean here, are you talking about knowledge distillation? Can you please eloberate more on this. </p>",
      "rawMarkdown": "Hi Ivan, \n\" So we came up with distillation\" -- &gt; What exactly distillation mean here, are you talking about knowledge distillation? Can you please eloberate more on this.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 481348,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "03/01/2019 09:40:04",
      "content": "<p>Congratulations <a href=\"/sawseen\">@sawseen</a>. Hard work paid finally.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481352,
      "author_name": "oldufo",
      "author_url": "",
      "post_date": "03/01/2019 09:49:59",
      "content": "<p>I like, how you call 4 year old algorithm \"good old Adam\"))))</p>",
      "votes": null,
      "replies": [
        {
          "id": 481463,
          "author_name": "sawseen",
          "author_url": "",
          "post_date": "03/01/2019 12:33:44",
          "content": "<p>4 years in deep learning seems like an eternity:)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 481358,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "03/01/2019 09:55:04",
      "content": "<p>Congrats on the great result! Great write up :) Really like the set up you had with quick iterations to test hypotheses.</p>",
      "votes": null,
      "replies": [
        {
          "id": 481461,
          "author_name": "sawseen",
          "author_url": "",
          "post_date": "03/01/2019 12:32:40",
          "content": "<p>Thank you! As what kaggle have taught me is to iterate as fast as possible and fully utilize your hardware. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 481394,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/01/2019 10:47:12",
      "content": "<p>Congrats <a href=\"/sawseen\">@sawseen</a> and thanks for sharing your solution overview.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481833,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/01/2019 22:37:34",
      "content": "<p>Great solution, congrats and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 505626,
      "author_name": "gautam840",
      "author_url": "",
      "post_date": "04/02/2019 10:09:22",
      "content": "<p>Hi Ivan, \n\" So we came up with distillation\" -- &gt; What exactly distillation mean here, are you talking about knowledge distillation? Can you please eloberate more on this. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "481346": "**TL;DR** Adam, Cosine with restarts, CosFace, ArcFace, High-resolution images, Weighted sampling, new_whale distillation, Pseudo labeled test, Resnet34, BNInception, Densenet121, AutoAugment, CoordConv, GAPNet \n\nWe’d like to share our solution as a story how we gradually improve our models. \n\nBut first of all I’d like to thank my teammate [Vladislav][1] for fruitful collaboration, kaggle community for motivation and ods.ai for kind support:)\n\nTo start with, it was obvious idea to consider whale’s flukes as human faces. Fortunately there are tons of papers for face identification, re-identification and verification. \n\nSo in the beginning of this competition this paper https://arxiv.org/abs/1804.06655 helped us a lot. It features comprehensive survey of state-of-the-art face identification techniques. \nAccording to it softmax-based losses look really promising. Due to their classification nature and the fact that we already have classification pipeline from Protein Atlas and Draw challenges we decided to focus on them. \n\nAmong others **Cosface** and **Arcface** stand out as newly discovered SOTA for face recognition task. The main idea is to bring examples of the same class close to each other in cosine similarity space and to pull apart distinct classes. Training with cosface or arcface generally is classification, so the final loss was CrossEntropy. One can read more details in their papers: https://arxiv.org/abs/1801.09414, https://arxiv.org/abs/1801.07698. After we train cosface or arcface net we took embeddings and calculate cosine similarity between train and test images. Then average similarities for each class in train and took 5 most similar.\n\nIn the beginning of every competition you should always devise robust validation procedure. We did it poorly. But nevertheless, to do so we select about 1000 sample from classes with number of instances greater than 3, one example for each class. Also we chose about the same number of new_whale images. This setup show good correlation between local score and public LB score. Unfortunately threshold for new_whale derived from local validation was slightly biased. That was really bad, because threshold was unreliable. Another way for threshold determination was to adjust it so top-1 new_whale percentage should be around 30%. \n\nTo keep up with kaggle community in conquering LB we decided to construct our pipeline in the following way:\n\n 1. Decrease training time of model as much as possible\n 2. Test a lot of hypothesis as much as possible\n\nTo do that we restrict image size to 256x256 and number of epochs up to 64. This kind of restrictions gave us a model that can be trained in 2 hours on 1080ti or even faster on 2080ti. This setting let us iterate quickly in testing new hypothesis or optimizing hyperparameters. After we have established our general training we start endless array of experiments for low-res images. \n\nLet’s divide our experiments into two broad groups:\n\n 1. Model engineering: what to train\n 2. Training engineering: how to train\n\nHere comes Model engineering. We started with some heavy encoders such as inceptionv4, seresnext50 etc. But it appeared that for us in classification task they seem to overfit a lot. Then we decided to main some light networks such as resnet34, bninception and densenet121. After several competitions I begin to realise that sometimes when you don’t have much data light encoders may really boost your score. Like they don’t tend to overfit much to rare classes and label noise. This is just a hypothesis that need to be carefully verified.\n\nTo get final models after initial 64 epochs on 256x256 images we increase image size up to 1024 for resnet34, up to 512 for bninception and up to 640 for densenet121 and train for 64 epochs more. \n\nTo boost model performant we tried a lot of modification. According to our findings **CoordConv** https://arxiv.org/abs/1807.03247 and **GapNet** architecture https://openreview.net/forum?id=ryl5khRcKm helped to improve resnet34 score. Unfortunately we didn’t have time to test this mods on bninception and densenet121. Also adding some sophisticated convolution blocks to our nets didn’t help. So Squeeze-and-Excitation, Convolutional Block Attention Module didn’t help. That was a sad story because a lot of time was spent on trying to optimize model architecture instead of optimizing training itself. \n\nWhen it comes to training one of the first things that comes to mind is how not to overfit to training data. Especially when one working with zero and few-shot learning. Inspired by AutoAugment paper https://arxiv.org/abs/1805.09501 we search augmentation space by random sampling and came up with the following augmentations:\n\n 1. HorizontalFlip\n 2. Rotate with 16 degree limit\n 3. ShiftScaleRotate with 16 degree limit\n 4. RandomBrightnessContrast\n 5. RandomGamma\n 6. Blur\n 7. Perspective transform: tile left, right and corner\n 8. Shear\n 9. MotionBlur\n 10. GridDistortion\n 11. ElasticTransform\n 12. Cutout\n\nThose augmentations we took from albumentations https://github.com/albu/albumentations and Augmentor https://github.com/mdbloice/Augmentor modules. Firstly we thought that this is to much for our networks. But actually for our models it was essential part not to overfit to training data.\n\nCosface and Arcface parameters was optimised as well. Cosface: S = 32.0, M=0.35. Arcface: M1 = 1.0, M2 = 0.4, M3 = 0.15.\n\nWe experimented a lot with optimizers and their hyperparameters: Adam, AdamW, SGD, SGDW. But the best optimizer for us appeared to be good old **Adam with Cosine annealing**.\n\nIn the end we tried different kinds of TTA, but it didn’t help to improve the score. \nMixed precision learning didn’t show good results either. \n\nStarting from the beginning we realised that it is essential to do something with new_whales in order to incorporate them into training process. Simple solution was to assign each new_whale probability of each class as 1 / 5004. With help of weighted sampling technique it gave us some boost. But then we realised why don’t we use softmax predictions for new_whales derived from trained ensemble. So we came up with **distillation**. We choose distillation instead of pseudo labels, because new_whale is considered to have different labels from train labels. Though it might not to be really true. \n\nTo further boost model capabilities we add test images with **pseudo labels** into train. Eventually our single model can hit 0.958 with snapshot ensembling. Unfortunately ensembling of models trained in this way didn’t give score improvement. Maybe it was due to less variety because of pseudo labels and distillation. \n\nIn the end I should mention that this competition was really interesting and give us an opportunity to develop face/fluke recognition skills. Thank your Kaggle!\n\t\n\t\n\n\n  [1]: https://www.kaggle.com/vlad0922",
    "481348": "Congratulations @sawseen. Hard work paid finally.",
    "481352": "I like, how you call 4 year old algorithm \"good old Adam\"))))",
    "481358": "Congrats on the great result! Great write up :) Really like the set up you had with quick iterations to test hypotheses.",
    "481394": "Congrats @sawseen and thanks for sharing your solution overview.",
    "481461": "Thank you! As what kaggle have taught me is to iterate as fast as possible and fully utilize your hardware.",
    "481463": "4 years in deep learning seems like an eternity:)",
    "481833": "Great solution, congrats and thanks for sharing.",
    "505626": "Hi Ivan, \n\" So we came up with distillation\" -- &gt; What exactly distillation mean here, are you talking about knowledge distillation? Can you please eloberate more on this."
  },
  "source": "meta"
}