{
  "id": 158370,
  "title": "LB #1 Solution",
  "url": "/competitions/iwildcam-2020-fgvc7/writeups/megvii-research-nanjing-lb-1-solution",
  "author_name": "",
  "post_date": "2020-06-14T05:08:46.904873700Z",
  "votes": 14,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Sorry for the delay in posting our solution.\nThanks to other teams for posting their solutions and the host for hosting this competition.</p>\n\n<p>We would like to share our solution with the competition participants and also the fine-grained recognition community.</p>\n\n<p>In general, we utilize Efficientnet-b6 and SeResNeXt101 as our backbone networks, several training methods and tricks are used to help the model to learn generalizable feature representations across camera locations. Note that, our method achieves 0.953 public and 0.906 private score, while the highest submission which we did not pick was 0.953 and 0.917 on public and private leaderboards, respectively. In details,</p>\n\n<ol>\n<li>A mixup strategy for the long-tailed data distribution which utilizes two data samplers, one is the uniformed sampler which samples each image with a uniformed probability, another is a reversed sampler which samples each image with a probability proportional to the reciprocal of corresponding class sample size. The images from those two data samplers are then mixed-up for better performance.</li>\n<li>An auxiliary classifier for different locations is added to our network, in the front of classification for location, the feature will go through a gradient reversal layer to ensure the model can learn a generalizable feature across locations.</li>\n<li>Adversarial training (AdvProp) was utilized to learn a noise robust representation.</li>\n<li>We ran MegaDetector V4 on all images.</li>\n<li>The bbox with the max confidence in each image is cropped to train a bbox model.</li>\n<li>We train two versions of each model, one with original full image and one with the cropped max confidence bbox.</li>\n<li>During test, for the image which has at least one bbox, the prediction is a weighted average of bbox model and full image model(0.3 full image + 0.7  bbox), for the images which have no bbox, we will use the full image model.</li>\n<li>We first obtain the predictions of Efficientnet-b6 and SeResNeXt101 using the aforementioned procedure, then the final prediction is a simple average of the two network architectures.</li>\n<li>We average the predictions which are clustered by location and datetime, as the sequence annotation appears to be noisy.</li>\n</ol>",
  "messages": [
    {
      "id": "885312",
      "postDate": "06/14/2020 05:08:46",
      "content": "<p>Sorry for the delay in posting our solution.\nThanks to other teams for posting their solutions and the host for hosting this competition.</p>\n\n<p>We would like to share our solution with the competition participants and also the fine-grained recognition community.</p>\n\n<p>In general, we utilize Efficientnet-b6 and SeResNeXt101 as our backbone networks, several training methods and tricks are used to help the model to learn generalizable feature representations across camera locations. Note that, our method achieves 0.953 public and 0.906 private score, while the highest submission which we did not pick was 0.953 and 0.917 on public and private leaderboards, respectively. In details,</p>\n\n<ol>\n<li>A mixup strategy for the long-tailed data distribution which utilizes two data samplers, one is the uniformed sampler which samples each image with a uniformed probability, another is a reversed sampler which samples each image with a probability proportional to the reciprocal of corresponding class sample size. The images from those two data samplers are then mixed-up for better performance.</li>\n<li>An auxiliary classifier for different locations is added to our network, in the front of classification for location, the feature will go through a gradient reversal layer to ensure the model can learn a generalizable feature across locations.</li>\n<li>Adversarial training (AdvProp) was utilized to learn a noise robust representation.</li>\n<li>We ran MegaDetector V4 on all images.</li>\n<li>The bbox with the max confidence in each image is cropped to train a bbox model.</li>\n<li>We train two versions of each model, one with original full image and one with the cropped max confidence bbox.</li>\n<li>During test, for the image which has at least one bbox, the prediction is a weighted average of bbox model and full image model(0.3 full image + 0.7  bbox), for the images which have no bbox, we will use the full image model.</li>\n<li>We first obtain the predictions of Efficientnet-b6 and SeResNeXt101 using the aforementioned procedure, then the final prediction is a simple average of the two network architectures.</li>\n<li>We average the predictions which are clustered by location and datetime, as the sequence annotation appears to be noisy.</li>\n</ol>",
      "rawMarkdown": "Sorry for the delay in posting our solution.\nThanks to other teams for posting their solutions and the host for hosting this competition.\n\nWe would like to share our solution with the competition participants and also the fine-grained recognition community.\n\nIn general, we utilize Efficientnet-b6 and SeResNeXt101 as our backbone networks, several training methods and tricks are used to help the model to learn generalizable feature representations across camera locations. Note that, our method achieves 0.953 public and 0.906 private score, while the highest submission which we did not pick was 0.953 and 0.917 on public and private leaderboards, respectively. In details,\n\n1. A mixup strategy for the long-tailed data distribution which utilizes two data samplers, one is the uniformed sampler which samples each image with a uniformed probability, another is a reversed sampler which samples each image with a probability proportional to the reciprocal of corresponding class sample size. The images from those two data samplers are then mixed-up for better performance.\n2. An auxiliary classifier for different locations is added to our network, in the front of classification for location, the feature will go through a gradient reversal layer to ensure the model can learn a generalizable feature across locations.\n3. Adversarial training (AdvProp) was utilized to learn a noise robust representation.\n4. We ran MegaDetector V4 on all images.\n5. The bbox with the max confidence in each image is cropped to train a bbox model.\n6. We train two versions of each model, one with original full image and one with the cropped max confidence bbox.\n7. During test, for the image which has at least one bbox, the prediction is a weighted average of bbox model and full image model(0.3 full image + 0.7  bbox), for the images which have no bbox, we will use the full image model.\n8. We first obtain the predictions of Efficientnet-b6 and SeResNeXt101 using the aforementioned procedure, then the final prediction is a simple average of the two network architectures.\n9. We average the predictions which are clustered by location and datetime, as the sequence annotation appears to be noisy.",
      "votes": null
    },
    {
      "id": "885316",
      "postDate": "06/14/2020 05:11:44",
      "content": "<p><a href=\"/stevenyin\">@stevenyin</a> Hi steven, Here's our solution, cheers.</p>",
      "rawMarkdown": "stevenyin Hi steven, Here's our solution, cheers.",
      "votes": null
    },
    {
      "id": "886709",
      "postDate": "06/15/2020 08:15:12",
      "content": "<p>Great thanks for your sharing !</p>",
      "rawMarkdown": "Great thanks for your sharing !",
      "votes": null
    },
    {
      "id": "887497",
      "postDate": "06/15/2020 17:42:08",
      "content": "<p>Thanks for sharing - can you explain number 2 in a little more detail? </p>",
      "rawMarkdown": "Thanks for sharing - can you explain number 2 in a little more detail?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 885316,
      "author_name": "tiantianzhang4",
      "author_url": "",
      "post_date": "06/14/2020 05:11:44",
      "content": "<p><a href=\"/stevenyin\">@stevenyin</a> Hi steven, Here's our solution, cheers.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 886709,
      "author_name": "",
      "author_url": "",
      "post_date": "06/15/2020 08:15:12",
      "content": "<p>Great thanks for your sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 887497,
      "author_name": "justinkay92",
      "author_url": "",
      "post_date": "06/15/2020 17:42:08",
      "content": "<p>Thanks for sharing - can you explain number 2 in a little more detail? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "885312": "Sorry for the delay in posting our solution.\nThanks to other teams for posting their solutions and the host for hosting this competition.\n\nWe would like to share our solution with the competition participants and also the fine-grained recognition community.\n\nIn general, we utilize Efficientnet-b6 and SeResNeXt101 as our backbone networks, several training methods and tricks are used to help the model to learn generalizable feature representations across camera locations. Note that, our method achieves 0.953 public and 0.906 private score, while the highest submission which we did not pick was 0.953 and 0.917 on public and private leaderboards, respectively. In details,\n\n1. A mixup strategy for the long-tailed data distribution which utilizes two data samplers, one is the uniformed sampler which samples each image with a uniformed probability, another is a reversed sampler which samples each image with a probability proportional to the reciprocal of corresponding class sample size. The images from those two data samplers are then mixed-up for better performance.\n2. An auxiliary classifier for different locations is added to our network, in the front of classification for location, the feature will go through a gradient reversal layer to ensure the model can learn a generalizable feature across locations.\n3. Adversarial training (AdvProp) was utilized to learn a noise robust representation.\n4. We ran MegaDetector V4 on all images.\n5. The bbox with the max confidence in each image is cropped to train a bbox model.\n6. We train two versions of each model, one with original full image and one with the cropped max confidence bbox.\n7. During test, for the image which has at least one bbox, the prediction is a weighted average of bbox model and full image model(0.3 full image + 0.7  bbox), for the images which have no bbox, we will use the full image model.\n8. We first obtain the predictions of Efficientnet-b6 and SeResNeXt101 using the aforementioned procedure, then the final prediction is a simple average of the two network architectures.\n9. We average the predictions which are clustered by location and datetime, as the sequence annotation appears to be noisy.",
    "885316": "stevenyin Hi steven, Here's our solution, cheers.",
    "886709": "Great thanks for your sharing !",
    "887497": "Thanks for sharing - can you explain number 2 in a little more detail?"
  },
  "source": "meta"
}