{
  "id": 96424,
  "title": "3rd place solution description",
  "url": "/competitions/imet-2019-fgvc6/discussion/96424",
  "author_name": "Ilya Kibardin",
  "post_date": "2019-06-20T12:31:48.697000",
  "votes": -32,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Thanks to the organizers for an interesting competiton!</p>\n\n<h2>Validation</h2>\n\n<p>I used 5 fold CV split by <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a></p>\n\n<h2>Models Used</h2>\n\n<p>All the experiments were carried out with the se_resnext50 model, then se_resnext101 and senet154 were trained with optimal parameters. Nasnet did not work for me good enough in this competition. I also tried Efficientnet and densenet161, but they were significantly worse than se_resnext and senet. I used models pretrained on imagenet from Cadene repo <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>. I also added dropout to all the models as it improved the validation score. I also tried adding an additional output for predicting the resulting f2 score (to use as a metric for overall confidence), but that did not work, obviously, I guess.</p>\n\n<h2>Preprocessing</h2>\n\n<p>I used a custom version of RandomSizedCrop from albumentations <a href=\"https://github.com/albu/albumentations\">https://github.com/albu/albumentations</a> . Crop sizes were chosen randomly in interval [200, 800] with height/width aspect ratio in [0.4, 1.6] and then resized into 320x320 input size. I also used HorizontalFlip, RandomBrightnessContrast and RandomGamma augmentations. Also, in the middle of my experiments, I discovered that Cutout (random erase) and light \nrotations of +- 5 degrees slightly improved validation score, so I added them to the final models.\nI used several label smoothing techniques:\n1. 0 -&gt; eps; 1 -&gt; 1 - eps for eps = 0.05 for every label\n2. Smoothing with empirical distribution as described in the Inception paper <a href=\"https://arxiv.org/abs/1512.00567\">https://arxiv.org/abs/1512.00567</a>\n3. Smoothing with using OOF predictions of my ensemble as an empirical distribution for the previous approach.</p>\n\n<p>All of the approaches significantly boosted the validation score approximately the same. So I decided to keep all of them and to blend them together. The tricky part here is that you have to optimize the threshold for every smoothing technique separately.\nI experimented a lot with Mixup in this competition, but it did not work for me at all. </p>\n\n<h2>Loss Function</h2>\n\n<p>I modified the FocalLoss with replacing (1 - p)<em>2 coefficient by (1 + (1 - p)</em>*2), which worked way better than the vanilla one. Also, from my first experiment, I noticed that the tags seemed to be way noisier than the cultures and my model had 2x higher error rate on tags. So I started weighting tags 2x more in my loss function. \nDue to the fact that the competition metric was F2 score which weights recall higher, I tried to lower the number of false negative predictions while caring less for false positive ones. So I added another 2x multiplier for coefficient where (labels == 1).</p>\n\n<p>At last, then I explored the dataset I found that:\n1. Annotations are quite noisy\n2. There are a lot of related tags/cultures such as France/Paris or Textile/Textile Fragments</p>\n\n<p>So I applied an idea for loss funtion proposed in <a href=\"https://arxiv.org/pdf/1809.00778.pdf\">https://arxiv.org/pdf/1809.00778.pdf</a> by last year open images runner-ups. I did not punish the model for co-occurrent classes. This hack improved my score significantly.</p>\n\n<h2>Training</h2>\n\n<p>I trained model in 3 stages:\n1. fc layer only with Adam. I stopped the stage when the model reached ~0.4 F2 score on validation, otherwise, too much training with frozen encoder led to very overfitted and degenerate models. \n2. Full model with Adam and ReduceLROnPlateau, until the lr drops to 1e-7 for 1e-4.\n3. Cosine annealing with SGD. This stage provided me with several different local minimums like described in <a href=\"https://arxiv.org/abs/1704.00109\">https://arxiv.org/abs/1704.00109</a> . I then ensembled the checkpoints in order to obtain a heavy ensemble for data cleaning and pseudo labeling</p>\n\n<h2>General workflow</h2>\n\n<ol>\n<li>Using all the stuff I described above, I trained a heavy ensemble of se_resnext50, se_resnext101, senet154, nasnet, densenet161 with different smoothings. I also used snapshot ensembling with 6 best checkpoints and TTA10 (horizontal flip * 5 crops). </li>\n<li>I obtained OOF predictions for train and public subset. I whitened the train subset by dropping samples with high errors. In total, I dropped around 5% of the samples. I used pseudo labels for ~5k images from the public dataset. While thresholding for pseudo labeling I did not use the optimal validation threshold, but I thresholded to obtain similar distributions of labels in train / public. This approach worked better for me.  </li>\n<li>I retrained se_resnext101 and senet154 on the cleared data with pseudo labels using all the described techniques. Here I did not use snapshot ensembling. Also, I used only TTA2 (horizontal flip).</li>\n</ol>\n\n<p>I also tried to train a second layer lgdm model, but I did not manage to do it because by the end of the competition I had to pass the exams at my university so I did not have enough free time to implement it.</p>",
  "messages": [
    {
      "id": 556628,
      "postDate": "2019-06-20T13:23:56.463Z",
      "content": "<blockquote>\n  <p>To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.</p>\n</blockquote>",
      "rawMarkdown": "&gt; To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.",
      "votes": 4
    },
    {
      "id": 556627,
      "postDate": "2019-06-20T13:22:31.203Z",
      "content": "<p>Congratulations! Will you release your code? </p>",
      "rawMarkdown": "Congratulations! Will you release your code? ",
      "votes": 2,
      "replies": [
        {
          "id": 556638,
          "postDate": "2019-06-20T13:35:20.597Z",
          "content": "<p>How to overfit public LB with 2-3 submissions? :)</p>",
          "rawMarkdown": "How to overfit public LB with 2-3 submissions? :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 556623,
      "postDate": "2019-06-20T13:17:38.377Z",
      "content": "<p>Five people became top in a few days at the same time, and all five people had overfitted the local cv. Four of them are also from X5, it’s a coincidence！</p>",
      "rawMarkdown": "Five people became top in a few days at the same time, and all five people had overfitted the local cv. Four of them are also from X5, it’s a coincidence！",
      "votes": 2,
      "replies": [
        {
          "id": 556642,
          "postDate": "2019-06-20T13:37:52.587Z",
          "content": "<p>another guy is their friend! They should buy a lottery ticket!</p>",
          "rawMarkdown": "another guy is their friend! They should buy a lottery ticket!",
          "votes": 3
        }
      ]
    },
    {
      "id": 556606,
      "postDate": "2019-06-20T13:00:02.857Z",
      "content": "<p>telling story is all you need!</p>",
      "rawMarkdown": "telling story is all you need!",
      "votes": 3
    },
    {
      "id": 556693,
      "postDate": "2019-06-20T14:31:30.143Z",
      "content": "<p>So how to overfit public LB to 0.70+?</p>",
      "rawMarkdown": "So how to overfit public LB to 0.70+?",
      "votes": 2
    },
    {
      "id": 556588,
      "postDate": "2019-06-20T12:31:48.697Z",
      "content": "<p>Thanks to the organizers for an interesting competiton!</p>\n\n<h2>Validation</h2>\n\n<p>I used 5 fold CV split by <a href=\"https://github.com/trent-b/iterative-stratification\">https://github.com/trent-b/iterative-stratification</a></p>\n\n<h2>Models Used</h2>\n\n<p>All the experiments were carried out with the se_resnext50 model, then se_resnext101 and senet154 were trained with optimal parameters. Nasnet did not work for me good enough in this competition. I also tried Efficientnet and densenet161, but they were significantly worse than se_resnext and senet. I used models pretrained on imagenet from Cadene repo <a href=\"https://github.com/Cadene/pretrained-models.pytorch\">https://github.com/Cadene/pretrained-models.pytorch</a>. I also added dropout to all the models as it improved the validation score. I also tried adding an additional output for predicting the resulting f2 score (to use as a metric for overall confidence), but that did not work, obviously, I guess.</p>\n\n<h2>Preprocessing</h2>\n\n<p>I used a custom version of RandomSizedCrop from albumentations <a href=\"https://github.com/albu/albumentations\">https://github.com/albu/albumentations</a> . Crop sizes were chosen randomly in interval [200, 800] with height/width aspect ratio in [0.4, 1.6] and then resized into 320x320 input size. I also used HorizontalFlip, RandomBrightnessContrast and RandomGamma augmentations. Also, in the middle of my experiments, I discovered that Cutout (random erase) and light \nrotations of +- 5 degrees slightly improved validation score, so I added them to the final models.\nI used several label smoothing techniques:\n1. 0 -&gt; eps; 1 -&gt; 1 - eps for eps = 0.05 for every label\n2. Smoothing with empirical distribution as described in the Inception paper <a href=\"https://arxiv.org/abs/1512.00567\">https://arxiv.org/abs/1512.00567</a>\n3. Smoothing with using OOF predictions of my ensemble as an empirical distribution for the previous approach.</p>\n\n<p>All of the approaches significantly boosted the validation score approximately the same. So I decided to keep all of them and to blend them together. The tricky part here is that you have to optimize the threshold for every smoothing technique separately.\nI experimented a lot with Mixup in this competition, but it did not work for me at all. </p>\n\n<h2>Loss Function</h2>\n\n<p>I modified the FocalLoss with replacing (1 - p)<em>2 coefficient by (1 + (1 - p)</em>*2), which worked way better than the vanilla one. Also, from my first experiment, I noticed that the tags seemed to be way noisier than the cultures and my model had 2x higher error rate on tags. So I started weighting tags 2x more in my loss function. \nDue to the fact that the competition metric was F2 score which weights recall higher, I tried to lower the number of false negative predictions while caring less for false positive ones. So I added another 2x multiplier for coefficient where (labels == 1).</p>\n\n<p>At last, then I explored the dataset I found that:\n1. Annotations are quite noisy\n2. There are a lot of related tags/cultures such as France/Paris or Textile/Textile Fragments</p>\n\n<p>So I applied an idea for loss funtion proposed in <a href=\"https://arxiv.org/pdf/1809.00778.pdf\">https://arxiv.org/pdf/1809.00778.pdf</a> by last year open images runner-ups. I did not punish the model for co-occurrent classes. This hack improved my score significantly.</p>\n\n<h2>Training</h2>\n\n<p>I trained model in 3 stages:\n1. fc layer only with Adam. I stopped the stage when the model reached ~0.4 F2 score on validation, otherwise, too much training with frozen encoder led to very overfitted and degenerate models. \n2. Full model with Adam and ReduceLROnPlateau, until the lr drops to 1e-7 for 1e-4.\n3. Cosine annealing with SGD. This stage provided me with several different local minimums like described in <a href=\"https://arxiv.org/abs/1704.00109\">https://arxiv.org/abs/1704.00109</a> . I then ensembled the checkpoints in order to obtain a heavy ensemble for data cleaning and pseudo labeling</p>\n\n<h2>General workflow</h2>\n\n<ol>\n<li>Using all the stuff I described above, I trained a heavy ensemble of se_resnext50, se_resnext101, senet154, nasnet, densenet161 with different smoothings. I also used snapshot ensembling with 6 best checkpoints and TTA10 (horizontal flip * 5 crops). </li>\n<li>I obtained OOF predictions for train and public subset. I whitened the train subset by dropping samples with high errors. In total, I dropped around 5% of the samples. I used pseudo labels for ~5k images from the public dataset. While thresholding for pseudo labeling I did not use the optimal validation threshold, but I thresholded to obtain similar distributions of labels in train / public. This approach worked better for me.  </li>\n<li>I retrained se_resnext101 and senet154 on the cleared data with pseudo labels using all the described techniques. Here I did not use snapshot ensembling. Also, I used only TTA2 (horizontal flip).</li>\n</ol>\n\n<p>I also tried to train a second layer lgdm model, but I did not manage to do it because by the end of the competition I had to pass the exams at my university so I did not have enough free time to implement it.</p>",
      "rawMarkdown": "Thanks to the organizers for an interesting competiton!\n\n## Validation\nI used 5 fold CV split by https://github.com/trent-b/iterative-stratification\n\n## Models Used\nAll the experiments were carried out with the se_resnext50 model, then se_resnext101 and senet154 were trained with optimal parameters. Nasnet did not work for me good enough in this competition. I also tried Efficientnet and densenet161, but they were significantly worse than se_resnext and senet. I used models pretrained on imagenet from Cadene repo https://github.com/Cadene/pretrained-models.pytorch. I also added dropout to all the models as it improved the validation score. I also tried adding an additional output for predicting the resulting f2 score (to use as a metric for overall confidence), but that did not work, obviously, I guess.\n\n## Preprocessing\nI used a custom version of RandomSizedCrop from albumentations https://github.com/albu/albumentations . Crop sizes were chosen randomly in interval [200, 800] with height/width aspect ratio in [0.4, 1.6] and then resized into 320x320 input size. I also used HorizontalFlip, RandomBrightnessContrast and RandomGamma augmentations. Also, in the middle of my experiments, I discovered that Cutout (random erase) and light \nrotations of +- 5 degrees slightly improved validation score, so I added them to the final models.\nI used several label smoothing techniques:\n1. 0 -&gt; eps; 1 -&gt; 1 - eps for eps = 0.05 for every label\n2. Smoothing with empirical distribution as described in the Inception paper https://arxiv.org/abs/1512.00567\n3. Smoothing with using OOF predictions of my ensemble as an empirical distribution for the previous approach.\n\nAll of the approaches significantly boosted the validation score approximately the same. So I decided to keep all of them and to blend them together. The tricky part here is that you have to optimize the threshold for every smoothing technique separately.\nI experimented a lot with Mixup in this competition, but it did not work for me at all. \n\n## Loss Function\nI modified the FocalLoss with replacing (1 - p)*2 coefficient by (1 + (1 - p)**2), which worked way better than the vanilla one. Also, from my first experiment, I noticed that the tags seemed to be way noisier than the cultures and my model had 2x higher error rate on tags. So I started weighting tags 2x more in my loss function. \nDue to the fact that the competition metric was F2 score which weights recall higher, I tried to lower the number of false negative predictions while caring less for false positive ones. So I added another 2x multiplier for coefficient where (labels == 1).\n\nAt last, then I explored the dataset I found that:\n1. Annotations are quite noisy\n2. There are a lot of related tags/cultures such as France/Paris or Textile/Textile Fragments\n\nSo I applied an idea for loss funtion proposed in https://arxiv.org/pdf/1809.00778.pdf by last year open images runner-ups. I did not punish the model for co-occurrent classes. This hack improved my score significantly.\n## Training\nI trained model in 3 stages:\n1. fc layer only with Adam. I stopped the stage when the model reached ~0.4 F2 score on validation, otherwise, too much training with frozen encoder led to very overfitted and degenerate models. \n2. Full model with Adam and ReduceLROnPlateau, until the lr drops to 1e-7 for 1e-4.\n3. Cosine annealing with SGD. This stage provided me with several different local minimums like described in https://arxiv.org/abs/1704.00109 . I then ensembled the checkpoints in order to obtain a heavy ensemble for data cleaning and pseudo labeling\n\n## General workflow\n1. Using all the stuff I described above, I trained a heavy ensemble of se_resnext50, se_resnext101, senet154, nasnet, densenet161 with different smoothings. I also used snapshot ensembling with 6 best checkpoints and TTA10 (horizontal flip * 5 crops). \n2. I obtained OOF predictions for train and public subset. I whitened the train subset by dropping samples with high errors. In total, I dropped around 5% of the samples. I used pseudo labels for ~5k images from the public dataset. While thresholding for pseudo labeling I did not use the optimal validation threshold, but I thresholded to obtain similar distributions of labels in train / public. This approach worked better for me.  \n3. I retrained se_resnext101 and senet154 on the cleared data with pseudo labels using all the described techniques. Here I did not use snapshot ensembling. Also, I used only TTA2 (horizontal flip).\n \nI also tried to train a second layer lgdm model, but I did not manage to do it because by the end of the competition I had to pass the exams at my university so I did not have enough free time to implement it.\n",
      "votes": -32
    },
    {
      "id": 556605,
      "postDate": "2019-06-20T12:58:10.913Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 556628,
      "author_name": "Kulbear",
      "author_url": "",
      "post_date": "2019-06-20T13:23:56.463000",
      "content": "<blockquote>\n  <p>To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.</p>\n</blockquote>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 556627,
      "author_name": "earhian",
      "author_url": "",
      "post_date": "2019-06-20T13:22:31.203000",
      "content": "<p>Congratulations! Will you release your code? </p>",
      "votes": 2,
      "replies": [
        {
          "id": 556638,
          "author_name": "earhian",
          "author_url": "",
          "post_date": "2019-06-20T13:35:20.597000",
          "content": "<p>How to overfit public LB with 2-3 submissions? :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 556623,
      "author_name": "qrfaction",
      "author_url": "",
      "post_date": "2019-06-20T13:17:38.377000",
      "content": "<p>Five people became top in a few days at the same time, and all five people had overfitted the local cv. Four of them are also from X5, it’s a coincidence！</p>",
      "votes": 2,
      "replies": [
        {
          "id": 556642,
          "author_name": "earhian",
          "author_url": "",
          "post_date": "2019-06-20T13:37:52.587000",
          "content": "<p>another guy is their friend! They should buy a lottery ticket!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 556606,
      "author_name": "qrfaction",
      "author_url": "",
      "post_date": "2019-06-20T13:00:02.857000",
      "content": "<p>telling story is all you need!</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 556693,
      "author_name": "Strideradu",
      "author_url": "",
      "post_date": "2019-06-20T14:31:30.143000",
      "content": "<p>So how to overfit public LB to 0.70+?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 556605,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-20T12:58:10.913000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "556628": "&gt; To be honest, I don't see anything special in this solution. Really smells fishy to me. Anyway, congrats for the second place.",
    "556627": "Congratulations! Will you release your code? ",
    "556623": "Five people became top in a few days at the same time, and all five people had overfitted the local cv. Four of them are also from X5, it’s a coincidence！",
    "556606": "telling story is all you need!",
    "556693": "So how to overfit public LB to 0.70+?",
    "556588": "Thanks to the organizers for an interesting competiton!\n\n## Validation\nI used 5 fold CV split by https://github.com/trent-b/iterative-stratification\n\n## Models Used\nAll the experiments were carried out with the se_resnext50 model, then se_resnext101 and senet154 were trained with optimal parameters. Nasnet did not work for me good enough in this competition. I also tried Efficientnet and densenet161, but they were significantly worse than se_resnext and senet. I used models pretrained on imagenet from Cadene repo https://github.com/Cadene/pretrained-models.pytorch. I also added dropout to all the models as it improved the validation score. I also tried adding an additional output for predicting the resulting f2 score (to use as a metric for overall confidence), but that did not work, obviously, I guess.\n\n## Preprocessing\nI used a custom version of RandomSizedCrop from albumentations https://github.com/albu/albumentations . Crop sizes were chosen randomly in interval [200, 800] with height/width aspect ratio in [0.4, 1.6] and then resized into 320x320 input size. I also used HorizontalFlip, RandomBrightnessContrast and RandomGamma augmentations. Also, in the middle of my experiments, I discovered that Cutout (random erase) and light \nrotations of +- 5 degrees slightly improved validation score, so I added them to the final models.\nI used several label smoothing techniques:\n1. 0 -&gt; eps; 1 -&gt; 1 - eps for eps = 0.05 for every label\n2. Smoothing with empirical distribution as described in the Inception paper https://arxiv.org/abs/1512.00567\n3. Smoothing with using OOF predictions of my ensemble as an empirical distribution for the previous approach.\n\nAll of the approaches significantly boosted the validation score approximately the same. So I decided to keep all of them and to blend them together. The tricky part here is that you have to optimize the threshold for every smoothing technique separately.\nI experimented a lot with Mixup in this competition, but it did not work for me at all. \n\n## Loss Function\nI modified the FocalLoss with replacing (1 - p)*2 coefficient by (1 + (1 - p)**2), which worked way better than the vanilla one. Also, from my first experiment, I noticed that the tags seemed to be way noisier than the cultures and my model had 2x higher error rate on tags. So I started weighting tags 2x more in my loss function. \nDue to the fact that the competition metric was F2 score which weights recall higher, I tried to lower the number of false negative predictions while caring less for false positive ones. So I added another 2x multiplier for coefficient where (labels == 1).\n\nAt last, then I explored the dataset I found that:\n1. Annotations are quite noisy\n2. There are a lot of related tags/cultures such as France/Paris or Textile/Textile Fragments\n\nSo I applied an idea for loss funtion proposed in https://arxiv.org/pdf/1809.00778.pdf by last year open images runner-ups. I did not punish the model for co-occurrent classes. This hack improved my score significantly.\n## Training\nI trained model in 3 stages:\n1. fc layer only with Adam. I stopped the stage when the model reached ~0.4 F2 score on validation, otherwise, too much training with frozen encoder led to very overfitted and degenerate models. \n2. Full model with Adam and ReduceLROnPlateau, until the lr drops to 1e-7 for 1e-4.\n3. Cosine annealing with SGD. This stage provided me with several different local minimums like described in https://arxiv.org/abs/1704.00109 . I then ensembled the checkpoints in order to obtain a heavy ensemble for data cleaning and pseudo labeling\n\n## General workflow\n1. Using all the stuff I described above, I trained a heavy ensemble of se_resnext50, se_resnext101, senet154, nasnet, densenet161 with different smoothings. I also used snapshot ensembling with 6 best checkpoints and TTA10 (horizontal flip * 5 crops). \n2. I obtained OOF predictions for train and public subset. I whitened the train subset by dropping samples with high errors. In total, I dropped around 5% of the samples. I used pseudo labels for ~5k images from the public dataset. While thresholding for pseudo labeling I did not use the optimal validation threshold, but I thresholded to obtain similar distributions of labels in train / public. This approach worked better for me.  \n3. I retrained se_resnext101 and senet154 on the cleared data with pseudo labels using all the described techniques. Here I did not use snapshot ensembling. Also, I used only TTA2 (horizontal flip).\n \nI also tried to train a second layer lgdm model, but I did not manage to do it because by the end of the competition I had to pass the exams at my university so I did not have enough free time to implement it.\n",
    "556605": ""
  }
}