{
  "id": 157932,
  "title": "LB #3 Solution",
  "url": "/competitions/iwildcam-2020-fgvc7/discussion/157932",
  "author_name": "",
  "post_date": "2020-06-12T15:14:23.684000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi Kagglers, \nSorry for the late post about our solution. My teammate and I summarize our solution in the following perspectives:\n<code>\n- Data\n- Preprocessing/Augmentation methods\n- Models\n- Prediction technique\n- Class imbalance\n- Fine-grained features\n</code></p>\n\n<ul>\n<li><p>Data</p>\n\n<ol><li>We used crops of animals for training provided by the official detector results. Both 64x64 crops and 224x224 crops were tried and 224x224 crops provided better score;</li>\n<li>Some of the original bounding boxes were highly skewed, we tried to extend the skewed bounding boxes to make it quasi-square, and resized them to 224x224;</li></ol></li>\n<li><p>Preprocessing methods</p>\n\n<ol><li>We tried multiple augmentation methods, ranging from traditional physical augmentation to optical techniques. The final augmentation methods employed by us were \"horizontal-flip\" &amp; \" CLAHE/contrast limited adaptive histogram equalization\".</li>\n<li>Augmentations were applied during training on the fly;</li></ol></li>\n<li><p>Models\n<code>\n| Model | Public LB | Private LB |\n|-------------------------------|\n| EfficientNet | 0.835 | 0.823 |\n| ResNet152 | 0.794 | 0.815 |\n| NTS with ResNet50 | 0.837 | 0.841 |\n| NTS with ResNext50 | 0.840 | 0.850 |\n|Final Ensemble Model = 6 NST models + EfficientNet | 0.843 | 0.848 |\n</code>\nAnother trick we'd like to point out is using <code>GridSearch</code> to search for the optimal weights of each weak learner, using the local validation score as the benchmark.</p></li>\n<li><p>Prediction technique\nWe find the test images are sorted in the following order:\n<code>\nSequence ID --&amp;gt; Locations --&amp;gt; Time Stamp --&amp;gt; image dimension\n</code>\nTherefore, test crops were sorted into clips following the above logic. While prediction, a prediction by clip was adopted naturally. Specifically, for each clip, we average predicted probs of all images to calculate the final label; then the final label was shared by all the images under this clip.</p></li>\n<li><p>Class imbalance</p>\n\n<ol><li>We tried multiple ways to counter class imbalance issue. For example,  we used class weights calculated via the effective number pointed out by the paper, \"Class-Balanced Loss Based on Effective Number of Samples\", which didn't work. We also tried BBN Network but failed to make it run. At last, we still used routine class weights calculated by <code>1 / # of images</code> of each class.</li></ol></li>\n<li><p>Fine-grained features</p>\n\n<ol><li>We used a SOTA network, NTS, which consists of a Navigator, a Teacher and a Scrutinizer. </li></ol></li>\n<li><p>Wrap up\nSo, that's basically our team's solution to Rank 3 in this competition. We'd like to learn from the rest teams, especially #1 &amp; #2 team. If any team member from these two teams happens to see my post, please leave me a message so that I can drop you an email/Wechat message to connect. </p></li>\n</ul>\n\n<p>Thanks,\nSteven Yin &amp; Ryan Zheng</p>",
  "messages": [
    {
      "id": 883371,
      "postDate": "2020-06-12T15:14:23.683Z",
      "content": "<p>Hi Kagglers, \nSorry for the late post about our solution. My teammate and I summarize our solution in the following perspectives:\n<code>\n- Data\n- Preprocessing/Augmentation methods\n- Models\n- Prediction technique\n- Class imbalance\n- Fine-grained features\n</code></p>\n\n<ul>\n<li><p>Data</p>\n\n<ol><li>We used crops of animals for training provided by the official detector results. Both 64x64 crops and 224x224 crops were tried and 224x224 crops provided better score;</li>\n<li>Some of the original bounding boxes were highly skewed, we tried to extend the skewed bounding boxes to make it quasi-square, and resized them to 224x224;</li></ol></li>\n<li><p>Preprocessing methods</p>\n\n<ol><li>We tried multiple augmentation methods, ranging from traditional physical augmentation to optical techniques. The final augmentation methods employed by us were \"horizontal-flip\" &amp; \" CLAHE/contrast limited adaptive histogram equalization\".</li>\n<li>Augmentations were applied during training on the fly;</li></ol></li>\n<li><p>Models\n<code>\n| Model | Public LB | Private LB |\n|-------------------------------|\n| EfficientNet | 0.835 | 0.823 |\n| ResNet152 | 0.794 | 0.815 |\n| NTS with ResNet50 | 0.837 | 0.841 |\n| NTS with ResNext50 | 0.840 | 0.850 |\n|Final Ensemble Model = 6 NST models + EfficientNet | 0.843 | 0.848 |\n</code>\nAnother trick we'd like to point out is using <code>GridSearch</code> to search for the optimal weights of each weak learner, using the local validation score as the benchmark.</p></li>\n<li><p>Prediction technique\nWe find the test images are sorted in the following order:\n<code>\nSequence ID --&amp;gt; Locations --&amp;gt; Time Stamp --&amp;gt; image dimension\n</code>\nTherefore, test crops were sorted into clips following the above logic. While prediction, a prediction by clip was adopted naturally. Specifically, for each clip, we average predicted probs of all images to calculate the final label; then the final label was shared by all the images under this clip.</p></li>\n<li><p>Class imbalance</p>\n\n<ol><li>We tried multiple ways to counter class imbalance issue. For example,  we used class weights calculated via the effective number pointed out by the paper, \"Class-Balanced Loss Based on Effective Number of Samples\", which didn't work. We also tried BBN Network but failed to make it run. At last, we still used routine class weights calculated by <code>1 / # of images</code> of each class.</li></ol></li>\n<li><p>Fine-grained features</p>\n\n<ol><li>We used a SOTA network, NTS, which consists of a Navigator, a Teacher and a Scrutinizer. </li></ol></li>\n<li><p>Wrap up\nSo, that's basically our team's solution to Rank 3 in this competition. We'd like to learn from the rest teams, especially #1 &amp; #2 team. If any team member from these two teams happens to see my post, please leave me a message so that I can drop you an email/Wechat message to connect. </p></li>\n</ul>\n\n<p>Thanks,\nSteven Yin &amp; Ryan Zheng</p>",
      "rawMarkdown": "Hi Kagglers, \nSorry for the late post about our solution. My teammate and I summarize our solution in the following perspectives:\n```\n- Data\n- Preprocessing/Augmentation methods\n- Models\n- Prediction technique\n- Class imbalance\n- Fine-grained features\n```\n\n- Data\n1. We used crops of animals for training provided by the official detector results. Both 64x64 crops and 224x224 crops were tried and 224x224 crops provided better score;\n2. Some of the original bounding boxes were highly skewed, we tried to extend the skewed bounding boxes to make it quasi-square, and resized them to 224x224;\n\n- Preprocessing methods\n1. We tried multiple augmentation methods, ranging from traditional physical augmentation to optical techniques. The final augmentation methods employed by us were \"horizontal-flip\" &amp; \" CLAHE/contrast limited adaptive histogram equalization\".\n2. Augmentations were applied during training on the fly;\n\n- Models\n```\n| Model | Public LB | Private LB |\n|-------------------------------|\n| EfficientNet | 0.835 | 0.823 |\n| ResNet152 | 0.794 | 0.815 |\n| NTS with ResNet50 | 0.837 | 0.841 |\n| NTS with ResNext50 | 0.840 | 0.850 |\n|Final Ensemble Model = 6 NST models + EfficientNet | 0.843 | 0.848 |\n```\nAnother trick we'd like to point out is using `GridSearch` to search for the optimal weights of each weak learner, using the local validation score as the benchmark.\n\n- Prediction technique\nWe find the test images are sorted in the following order:\n```\nSequence ID --&gt; Locations --&gt; Time Stamp --&gt; image dimension\n```\nTherefore, test crops were sorted into clips following the above logic. While prediction, a prediction by clip was adopted naturally. Specifically, for each clip, we average predicted probs of all images to calculate the final label; then the final label was shared by all the images under this clip.\n\n- Class imbalance\n1. We tried multiple ways to counter class imbalance issue. For example,  we used class weights calculated via the effective number pointed out by the paper, \"Class-Balanced Loss Based on Effective Number of Samples\", which didn't work. We also tried BBN Network but failed to make it run. At last, we still used routine class weights calculated by `1 / # of images` of each class.\n\n- Fine-grained features\n1. We used a SOTA network, NTS, which consists of a Navigator, a Teacher and a Scrutinizer. \n\n- Wrap up\nSo, that's basically our team's solution to Rank 3 in this competition. We'd like to learn from the rest teams, especially #1 &amp; #2 team. If any team member from these two teams happens to see my post, please leave me a message so that I can drop you an email/Wechat message to connect. \n\nThanks,\nSteven Yin &amp; Ryan Zheng",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "883371": "Hi Kagglers, \nSorry for the late post about our solution. My teammate and I summarize our solution in the following perspectives:\n```\n- Data\n- Preprocessing/Augmentation methods\n- Models\n- Prediction technique\n- Class imbalance\n- Fine-grained features\n```\n\n- Data\n1. We used crops of animals for training provided by the official detector results. Both 64x64 crops and 224x224 crops were tried and 224x224 crops provided better score;\n2. Some of the original bounding boxes were highly skewed, we tried to extend the skewed bounding boxes to make it quasi-square, and resized them to 224x224;\n\n- Preprocessing methods\n1. We tried multiple augmentation methods, ranging from traditional physical augmentation to optical techniques. The final augmentation methods employed by us were \"horizontal-flip\" &amp; \" CLAHE/contrast limited adaptive histogram equalization\".\n2. Augmentations were applied during training on the fly;\n\n- Models\n```\n| Model | Public LB | Private LB |\n|-------------------------------|\n| EfficientNet | 0.835 | 0.823 |\n| ResNet152 | 0.794 | 0.815 |\n| NTS with ResNet50 | 0.837 | 0.841 |\n| NTS with ResNext50 | 0.840 | 0.850 |\n|Final Ensemble Model = 6 NST models + EfficientNet | 0.843 | 0.848 |\n```\nAnother trick we'd like to point out is using `GridSearch` to search for the optimal weights of each weak learner, using the local validation score as the benchmark.\n\n- Prediction technique\nWe find the test images are sorted in the following order:\n```\nSequence ID --&gt; Locations --&gt; Time Stamp --&gt; image dimension\n```\nTherefore, test crops were sorted into clips following the above logic. While prediction, a prediction by clip was adopted naturally. Specifically, for each clip, we average predicted probs of all images to calculate the final label; then the final label was shared by all the images under this clip.\n\n- Class imbalance\n1. We tried multiple ways to counter class imbalance issue. For example,  we used class weights calculated via the effective number pointed out by the paper, \"Class-Balanced Loss Based on Effective Number of Samples\", which didn't work. We also tried BBN Network but failed to make it run. At last, we still used routine class weights calculated by `1 / # of images` of each class.\n\n- Fine-grained features\n1. We used a SOTA network, NTS, which consists of a Navigator, a Teacher and a Scrutinizer. \n\n- Wrap up\nSo, that's basically our team's solution to Rank 3 in this competition. We'd like to learn from the rest teams, especially #1 &amp; #2 team. If any team member from these two teams happens to see my post, please leave me a message so that I can drop you an email/Wechat message to connect. \n\nThanks,\nSteven Yin &amp; Ryan Zheng"
  }
}