{
  "id": 95311,
  "title": "10th Solution",
  "url": "/competitions/imet-2019-fgvc6/writeups/alchemists-creed-obey-the-rules-10th-solution",
  "author_name": "",
  "post_date": "2019-06-11T14:32:06.967Z",
  "votes": 30,
  "comment_count": 7,
  "views": 0,
  "content": "<p>First of all, I would like to thank FGVC6,CVPR2019 and Kaggle for hosting such an interesting competition. I would like to thank my teammates Jiwei Liu and Yuanhao Wu for their hard work and brilliant ideas. Special congrats to Yuanhao for earning GrandMaster title! </p>\n\n<p>We feel really lucky to get the last gold medal. We only scored ~0.615 on the team merger deadline, lots of efforts were done in the last week. </p>\n\n<p><strong>Dataset</strong>: 5 fold CV with Multilabel Iterative Stratification\n<strong>Models</strong>: seresnext50, seresnext101, pnasnet5large, inceptionv4 and senet154. \n<strong>Augmentation</strong>: \n<strong>Train</strong>: Resize, RandomSizedCrop, HorizontalFlip, Blur, GaussNoise, HueSaturationValue and CLAHE\n<strong>Test</strong>: Only Resize\n<strong>Loss</strong>: BCE loss for the 1st stage, F2 loss for the 2nd stage\n<strong>Optimizor</strong>: Adam, LR=2e-4. \n<strong>Learning rate schedule</strong>: ReduceLROnPlateau(patience=0, ratio=0.1)\n<strong>Input size</strong>: 331x331 for pnasnet5large, 384x384 for other models. \n<strong>Batch size</strong>: 128. </p>\n\n<p>We started the journey with seresnext50 using BCE loss. Small modifications were made to the network, our logit takes input from the last two layers rather than the last one layer. When we finetuned the entire network, we found that a significant boost in f2 score is achieved right after the LR reduces, so we decide to set patience to 0. Later, we found that if we only train the last two layers, we can achieve similar score compared to training the entire network. Hence, we decided to only train the last two layers and classifier, this saves us lots of time and allows us to apply big batch size without batch accumulation. Meanwhile, in another experiment, we found that training the best model achieved by BCE loss using F2 loss for 3-5 epoches could give another big boost (~0.01). This finalizes our solution: train the model with BCE loss until converge (~0.603 on LB for single fold of seresnext50) -&gt; switch loss to F2 for another 3-5 epoches (~0.612 on LB for single fold of seresnext50). </p>\n\n<p>Then it comes to the final week. We don’t have time for more experiments, so we just applied the above strategy to other models and keep generating 5-fold results. A summary of our model performance (most models only have CV scores, since we don’t have enough submissions for all the models in the last 2 days):</p>\n\n<p><strong>Seresnext50</strong>: 0.6078\n<strong>Seresnext101</strong>: 0.6207 (0.646 on LB)\n<strong>Pnasnet5large</strong>: 0.6066\n<strong>Inceptionv4</strong>: 0.6030 \n<strong>Senet154</strong>: 0.6212 (0.642 on LB)</p>\n\n<p>Our ensemble method is a simple weighted average over the above models, it scores 0.658 on public LB. </p>\n\n<ol>\n<li><p>What we tried but not work:</p></li>\n<li><p>All crop related augmentation methods failed. </p></li>\n<li><p>TTA does not work for us. </p></li>\n<li><p>Efficientnet only has a CV of 0.5956, maybe we have not use its potential fully. </p></li>\n<li><p>We tried batch accumulation, our test shows there around 0.001 loss in local CV score. We didn’t accumulate too many batches, perhaps this is the reason. Personally, I think there is a trade off between the big batch size and BN synchronization problem. </p></li>\n</ol>\n\n<p>Our solution is a little bit brute force compared to other teams. Congrats to all winners. Thanks for reading! </p>",
  "messages": [
    {
      "id": "550200",
      "postDate": "06/11/2019 11:58:07",
      "content": "<p>First of all, I would like to thank FGVC6,CVPR2019 and Kaggle for hosting such an interesting competition. I would like to thank my teammates Jiwei Liu and Yuanhao Wu for their hard work and brilliant ideas. Special congrats to Yuanhao for earning GrandMaster title! </p>\n\n<p>We feel really lucky to get the last gold medal. We only scored ~0.615 on the team merger deadline, lots of efforts were done in the last week. </p>\n\n<p><strong>Dataset</strong>: 5 fold CV with Multilabel Iterative Stratification\n<strong>Models</strong>: seresnext50, seresnext101, pnasnet5large, inceptionv4 and senet154. \n<strong>Augmentation</strong>: \n<strong>Train</strong>: Resize, RandomSizedCrop, HorizontalFlip, Blur, GaussNoise, HueSaturationValue and CLAHE\n<strong>Test</strong>: Only Resize\n<strong>Loss</strong>: BCE loss for the 1st stage, F2 loss for the 2nd stage\n<strong>Optimizor</strong>: Adam, LR=2e-4. \n<strong>Learning rate schedule</strong>: ReduceLROnPlateau(patience=0, ratio=0.1)\n<strong>Input size</strong>: 331x331 for pnasnet5large, 384x384 for other models. \n<strong>Batch size</strong>: 128. </p>\n\n<p>We started the journey with seresnext50 using BCE loss. Small modifications were made to the network, our logit takes input from the last two layers rather than the last one layer. When we finetuned the entire network, we found that a significant boost in f2 score is achieved right after the LR reduces, so we decide to set patience to 0. Later, we found that if we only train the last two layers, we can achieve similar score compared to training the entire network. Hence, we decided to only train the last two layers and classifier, this saves us lots of time and allows us to apply big batch size without batch accumulation. Meanwhile, in another experiment, we found that training the best model achieved by BCE loss using F2 loss for 3-5 epoches could give another big boost (~0.01). This finalizes our solution: train the model with BCE loss until converge (~0.603 on LB for single fold of seresnext50) -&gt; switch loss to F2 for another 3-5 epoches (~0.612 on LB for single fold of seresnext50). </p>\n\n<p>Then it comes to the final week. We don’t have time for more experiments, so we just applied the above strategy to other models and keep generating 5-fold results. A summary of our model performance (most models only have CV scores, since we don’t have enough submissions for all the models in the last 2 days):</p>\n\n<p><strong>Seresnext50</strong>: 0.6078\n<strong>Seresnext101</strong>: 0.6207 (0.646 on LB)\n<strong>Pnasnet5large</strong>: 0.6066\n<strong>Inceptionv4</strong>: 0.6030 \n<strong>Senet154</strong>: 0.6212 (0.642 on LB)</p>\n\n<p>Our ensemble method is a simple weighted average over the above models, it scores 0.658 on public LB. </p>\n\n<ol>\n<li><p>What we tried but not work:</p></li>\n<li><p>All crop related augmentation methods failed. </p></li>\n<li><p>TTA does not work for us. </p></li>\n<li><p>Efficientnet only has a CV of 0.5956, maybe we have not use its potential fully. </p></li>\n<li><p>We tried batch accumulation, our test shows there around 0.001 loss in local CV score. We didn’t accumulate too many batches, perhaps this is the reason. Personally, I think there is a trade off between the big batch size and BN synchronization problem. </p></li>\n</ol>\n\n<p>Our solution is a little bit brute force compared to other teams. Congrats to all winners. Thanks for reading! </p>",
      "rawMarkdown": "First of all, I would like to thank FGVC6,CVPR2019 and Kaggle for hosting such an interesting competition. I would like to thank my teammates Jiwei Liu and Yuanhao Wu for their hard work and brilliant ideas. Special congrats to Yuanhao for earning GrandMaster title! \n\nWe feel really lucky to get the last gold medal. We only scored ~0.615 on the team merger deadline, lots of efforts were done in the last week. \n\n**Dataset**: 5 fold CV with Multilabel Iterative Stratification\n**Models**: seresnext50, seresnext101, pnasnet5large, inceptionv4 and senet154. \n**Augmentation**: \n**Train**: Resize, RandomSizedCrop, HorizontalFlip, Blur, GaussNoise, HueSaturationValue and CLAHE\n**Test**: Only Resize\n**Loss**: BCE loss for the 1st stage, F2 loss for the 2nd stage\n**Optimizor**: Adam, LR=2e-4. \n**Learning rate schedule**: ReduceLROnPlateau(patience=0, ratio=0.1)\n**Input size**: 331x331 for pnasnet5large, 384x384 for other models. \n**Batch size**: 128. \n\nWe started the journey with seresnext50 using BCE loss. Small modifications were made to the network, our logit takes input from the last two layers rather than the last one layer. When we finetuned the entire network, we found that a significant boost in f2 score is achieved right after the LR reduces, so we decide to set patience to 0. Later, we found that if we only train the last two layers, we can achieve similar score compared to training the entire network. Hence, we decided to only train the last two layers and classifier, this saves us lots of time and allows us to apply big batch size without batch accumulation. Meanwhile, in another experiment, we found that training the best model achieved by BCE loss using F2 loss for 3-5 epoches could give another big boost (~0.01). This finalizes our solution: train the model with BCE loss until converge (~0.603 on LB for single fold of seresnext50) -&gt; switch loss to F2 for another 3-5 epoches (~0.612 on LB for single fold of seresnext50). \n\nThen it comes to the final week. We don’t have time for more experiments, so we just applied the above strategy to other models and keep generating 5-fold results. A summary of our model performance (most models only have CV scores, since we don’t have enough submissions for all the models in the last 2 days):\n\n**Seresnext50**: 0.6078\n**Seresnext101**: 0.6207 (0.646 on LB)\n**Pnasnet5large**: 0.6066\n**Inceptionv4**: 0.6030 \n**Senet154**: 0.6212 (0.642 on LB)\n\nOur ensemble method is a simple weighted average over the above models, it scores 0.658 on public LB. \n\n1. What we tried but not work:\n\n2. All crop related augmentation methods failed. \n\n3. TTA does not work for us. \n\n4. Efficientnet only has a CV of 0.5956, maybe we have not use its potential fully. \n\n5. We tried batch accumulation, our test shows there around 0.001 loss in local CV score. We didn’t accumulate too many batches, perhaps this is the reason. Personally, I think there is a trade off between the big batch size and BN synchronization problem. \n\nOur solution is a little bit brute force compared to other teams. Congrats to all winners. Thanks for reading!",
      "votes": null
    },
    {
      "id": "550205",
      "postDate": "06/11/2019 12:01:32",
      "content": "<p>congrats!</p>",
      "rawMarkdown": "congrats!",
      "votes": null
    },
    {
      "id": "550213",
      "postDate": "06/11/2019 12:11:05",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "550307",
      "postDate": "06/11/2019 13:33:55",
      "content": "<p>Congrats !\nWill you opensource the code?</p>",
      "rawMarkdown": "Congrats !\nWill you opensource the code?",
      "votes": null
    },
    {
      "id": "550694",
      "postDate": "06/11/2019 23:58:06",
      "content": "<p>We are in the process of cleaning the code. May take a while. </p>",
      "rawMarkdown": "We are in the process of cleaning the code. May take a while.",
      "votes": null
    },
    {
      "id": "550695",
      "postDate": "06/11/2019 23:58:19",
      "content": "<p>Thanks! Congrats on your solo gold as well! </p>",
      "rawMarkdown": "Thanks! Congrats on your solo gold as well!",
      "votes": null
    },
    {
      "id": "550696",
      "postDate": "06/11/2019 23:58:25",
      "content": "<p>Thanks! </p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "568748",
      "postDate": "07/05/2019 11:31:27",
      "content": "<p><a href=\"/naivelamb\">@naivelamb</a> is it open now? </p>",
      "rawMarkdown": "naivelamb is it open now?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 550205,
      "author_name": "lucaskg",
      "author_url": "",
      "post_date": "06/11/2019 12:01:32",
      "content": "<p>congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 550696,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "06/11/2019 23:58:25",
          "content": "<p>Thanks! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550213,
      "author_name": "angelecarre",
      "author_url": "",
      "post_date": "06/11/2019 12:11:05",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": [
        {
          "id": 550695,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "06/11/2019 23:58:19",
          "content": "<p>Thanks! Congrats on your solo gold as well! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 550307,
      "author_name": "rinnqd",
      "author_url": "",
      "post_date": "06/11/2019 13:33:55",
      "content": "<p>Congrats !\nWill you opensource the code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 550694,
          "author_name": "naivelamb",
          "author_url": "",
          "post_date": "06/11/2019 23:58:06",
          "content": "<p>We are in the process of cleaning the code. May take a while. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 568748,
          "author_name": "rinnqd",
          "author_url": "",
          "post_date": "07/05/2019 11:31:27",
          "content": "<p><a href=\"/naivelamb\">@naivelamb</a> is it open now? </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "550200": "First of all, I would like to thank FGVC6,CVPR2019 and Kaggle for hosting such an interesting competition. I would like to thank my teammates Jiwei Liu and Yuanhao Wu for their hard work and brilliant ideas. Special congrats to Yuanhao for earning GrandMaster title! \n\nWe feel really lucky to get the last gold medal. We only scored ~0.615 on the team merger deadline, lots of efforts were done in the last week. \n\n**Dataset**: 5 fold CV with Multilabel Iterative Stratification\n**Models**: seresnext50, seresnext101, pnasnet5large, inceptionv4 and senet154. \n**Augmentation**: \n**Train**: Resize, RandomSizedCrop, HorizontalFlip, Blur, GaussNoise, HueSaturationValue and CLAHE\n**Test**: Only Resize\n**Loss**: BCE loss for the 1st stage, F2 loss for the 2nd stage\n**Optimizor**: Adam, LR=2e-4. \n**Learning rate schedule**: ReduceLROnPlateau(patience=0, ratio=0.1)\n**Input size**: 331x331 for pnasnet5large, 384x384 for other models. \n**Batch size**: 128. \n\nWe started the journey with seresnext50 using BCE loss. Small modifications were made to the network, our logit takes input from the last two layers rather than the last one layer. When we finetuned the entire network, we found that a significant boost in f2 score is achieved right after the LR reduces, so we decide to set patience to 0. Later, we found that if we only train the last two layers, we can achieve similar score compared to training the entire network. Hence, we decided to only train the last two layers and classifier, this saves us lots of time and allows us to apply big batch size without batch accumulation. Meanwhile, in another experiment, we found that training the best model achieved by BCE loss using F2 loss for 3-5 epoches could give another big boost (~0.01). This finalizes our solution: train the model with BCE loss until converge (~0.603 on LB for single fold of seresnext50) -&gt; switch loss to F2 for another 3-5 epoches (~0.612 on LB for single fold of seresnext50). \n\nThen it comes to the final week. We don’t have time for more experiments, so we just applied the above strategy to other models and keep generating 5-fold results. A summary of our model performance (most models only have CV scores, since we don’t have enough submissions for all the models in the last 2 days):\n\n**Seresnext50**: 0.6078\n**Seresnext101**: 0.6207 (0.646 on LB)\n**Pnasnet5large**: 0.6066\n**Inceptionv4**: 0.6030 \n**Senet154**: 0.6212 (0.642 on LB)\n\nOur ensemble method is a simple weighted average over the above models, it scores 0.658 on public LB. \n\n1. What we tried but not work:\n\n2. All crop related augmentation methods failed. \n\n3. TTA does not work for us. \n\n4. Efficientnet only has a CV of 0.5956, maybe we have not use its potential fully. \n\n5. We tried batch accumulation, our test shows there around 0.001 loss in local CV score. We didn’t accumulate too many batches, perhaps this is the reason. Personally, I think there is a trade off between the big batch size and BN synchronization problem. \n\nOur solution is a little bit brute force compared to other teams. Congrats to all winners. Thanks for reading!",
    "550205": "congrats!",
    "550213": "Congrats!",
    "550307": "Congrats !\nWill you opensource the code?",
    "550694": "We are in the process of cleaning the code. May take a while.",
    "550695": "Thanks! Congrats on your solo gold as well!",
    "550696": "Thanks!",
    "568748": "naivelamb is it open now?"
  },
  "source": "meta"
}