{
  "id": 242978,
  "title": "7th Place Solution",
  "url": "/competitions/iwildcam2021-fgvc8/discussion/242978",
  "author_name": "Devashish Prasad",
  "post_date": "2021-05-31T18:06:35.270000",
  "votes": 5,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I created a mini dataset (10% of the original dataset) with the same class proportion as in the original. And carried out many experiments to move quickly through the competition.</p>\n<p>1) Preprocessing -<br>\nIn my initial experiments, I found that instead of directly cropping and resizing the detections, cropping and padding the detections using reflection padding improved the score. Hence my final solution used reflection padding while training and testing. My preprocessing pipeline was as follows</p>\n<p>crop detection -&gt; reflection padding -&gt; augmentations -&gt; standardize</p>\n<p>2) Augmentations -<br>\nIn my experiments adding any type of augmentation reduced my validation score. Hence I kept simple augmentations with very little intensity. The augmentations that I used were as follows </p>\n<p>Random Horizontal Flip (p=0.5)<br>\nColor Jitter : Brightness and Contrast (0.9 to 1.2)<br>\nGrayscale (p=0.5)<br>\nGaussian Blur (sigma = 0.0 to 0.8)</p>\n<p>3) Handle Imbalance - <br>\nHere is how I have handled imbalance: <a href=\"https://www.kaggle.com/c/iwildcam2021-fgvc8/discussion/242455\" target=\"_blank\">https://www.kaggle.com/c/iwildcam2021-fgvc8/discussion/242455</a></p>\n<p>4) Model Architecture -<br>\nFor my mini dataset, I experimented with efficient b2 noisy student. For the final submission, I used the efficient b5 noisy student. I did not use pre-trained image net weights.</p>\n<p>5) Submission pipeline -<br>\nIn a public kernel, I came across the max count logic. In this, the maximum number of detections from an image in a sequence (max count) was considered as the prediction value for that sequence. I modified this logic and also considered maximum frequency. For eg - in a sequence of 9 images having only one specie of animal, the number of detections for each image were - [1,1,1,1,2,2,2,4,2]. So the max count logic will predict 4 while max count + max frequency logic will predict 2. There might be a strong possibility of false detections, and thus image number 8 having 4 detections (out of which 2 could be false detections) would be eliminated. The frequency of both 1 and 2 is the same, so the max count will be considered, that is 2. I saw a validation accuracy boost and public score boost using max count + max frequency logic, and hence used in my final submission.</p>\n<p>Also, a higher megadetector detection confidence threshold worked better, so I used 0.7 as detection confidence for the final submission pipeline.</p>\n<p>6) Training -<br>\nI experimented with mixed-precision training, but training without mixed-precision yielded better results for me. I did not use any LR scheduler for mini dataset experiments. But for the final training of efficient b5 noisy student, I used reduce on plateau. I underestimated the training time required for the efficient b5 noisy student. I spent all remaining 26 hrs of GPU in the last week in efficient b5 noisy student's training and still, it was a little underfit with ~ 78% val accuracy and 82% train accuracy. After exhausting my Kaggle GPU quota I switched to colab and resumed training. This time I changed some hyperparameters like batch-size and LR, but I could only reach 82% val accuracy and 85% train accuracy due to lack of time (deadline was near). But I couldn't make a submission using colab because I later realized that the test set was 30GB while on colab I had 28GB disk space. I couldn't even download the zip there. So my final model is still underfit (i.e. it needs more training).  </p>\n<p><strong>Here is my final submission pipeline <a href=\"https://www.kaggle.com/devashishprasad/iwildcam2021-submission\" target=\"_blank\">notebook</a></strong></p>",
  "messages": [
    {
      "id": 1330366,
      "postDate": "2021-05-31T18:06:35.270Z",
      "content": "<p>I created a mini dataset (10% of the original dataset) with the same class proportion as in the original. And carried out many experiments to move quickly through the competition.</p>\n<p>1) Preprocessing -<br>\nIn my initial experiments, I found that instead of directly cropping and resizing the detections, cropping and padding the detections using reflection padding improved the score. Hence my final solution used reflection padding while training and testing. My preprocessing pipeline was as follows</p>\n<p>crop detection -&gt; reflection padding -&gt; augmentations -&gt; standardize</p>\n<p>2) Augmentations -<br>\nIn my experiments adding any type of augmentation reduced my validation score. Hence I kept simple augmentations with very little intensity. The augmentations that I used were as follows </p>\n<p>Random Horizontal Flip (p=0.5)<br>\nColor Jitter : Brightness and Contrast (0.9 to 1.2)<br>\nGrayscale (p=0.5)<br>\nGaussian Blur (sigma = 0.0 to 0.8)</p>\n<p>3) Handle Imbalance - <br>\nHere is how I have handled imbalance: <a href=\"https://www.kaggle.com/c/iwildcam2021-fgvc8/discussion/242455\" target=\"_blank\">https://www.kaggle.com/c/iwildcam2021-fgvc8/discussion/242455</a></p>\n<p>4) Model Architecture -<br>\nFor my mini dataset, I experimented with efficient b2 noisy student. For the final submission, I used the efficient b5 noisy student. I did not use pre-trained image net weights.</p>\n<p>5) Submission pipeline -<br>\nIn a public kernel, I came across the max count logic. In this, the maximum number of detections from an image in a sequence (max count) was considered as the prediction value for that sequence. I modified this logic and also considered maximum frequency. For eg - in a sequence of 9 images having only one specie of animal, the number of detections for each image were - [1,1,1,1,2,2,2,4,2]. So the max count logic will predict 4 while max count + max frequency logic will predict 2. There might be a strong possibility of false detections, and thus image number 8 having 4 detections (out of which 2 could be false detections) would be eliminated. The frequency of both 1 and 2 is the same, so the max count will be considered, that is 2. I saw a validation accuracy boost and public score boost using max count + max frequency logic, and hence used in my final submission.</p>\n<p>Also, a higher megadetector detection confidence threshold worked better, so I used 0.7 as detection confidence for the final submission pipeline.</p>\n<p>6) Training -<br>\nI experimented with mixed-precision training, but training without mixed-precision yielded better results for me. I did not use any LR scheduler for mini dataset experiments. But for the final training of efficient b5 noisy student, I used reduce on plateau. I underestimated the training time required for the efficient b5 noisy student. I spent all remaining 26 hrs of GPU in the last week in efficient b5 noisy student's training and still, it was a little underfit with ~ 78% val accuracy and 82% train accuracy. After exhausting my Kaggle GPU quota I switched to colab and resumed training. This time I changed some hyperparameters like batch-size and LR, but I could only reach 82% val accuracy and 85% train accuracy due to lack of time (deadline was near). But I couldn't make a submission using colab because I later realized that the test set was 30GB while on colab I had 28GB disk space. I couldn't even download the zip there. So my final model is still underfit (i.e. it needs more training).  </p>\n<p><strong>Here is my final submission pipeline <a href=\"https://www.kaggle.com/devashishprasad/iwildcam2021-submission\" target=\"_blank\">notebook</a></strong></p>",
      "rawMarkdown": "I created a mini dataset (10% of the original dataset) with the same class proportion as in the original. And carried out many experiments to move quickly through the competition.\n\n1) Preprocessing -\nIn my initial experiments, I found that instead of directly cropping and resizing the detections, cropping and padding the detections using reflection padding improved the score. Hence my final solution used reflection padding while training and testing. My preprocessing pipeline was as follows\n\ncrop detection -> reflection padding -> augmentations -> standardize\n\n2) Augmentations -\nIn my experiments adding any type of augmentation reduced my validation score. Hence I kept simple augmentations with very little intensity. The augmentations that I used were as follows \n\nRandom Horizontal Flip (p=0.5)\nColor Jitter : Brightness and Contrast (0.9 to 1.2)\nGrayscale (p=0.5)\nGaussian Blur (sigma = 0.0 to 0.8)\n \n3) Handle Imbalance - \nHere is how I have handled imbalance: https://www.kaggle.com/c/iwildcam2021-fgvc8/discussion/242455\n\n4) Model Architecture -\nFor my mini dataset, I experimented with efficient b2 noisy student. For the final submission, I used the efficient b5 noisy student. I did not use pre-trained image net weights.\n\n5) Submission pipeline -\nIn a public kernel, I came across the max count logic. In this, the maximum number of detections from an image in a sequence (max count) was considered as the prediction value for that sequence. I modified this logic and also considered maximum frequency. For eg - in a sequence of 9 images having only one specie of animal, the number of detections for each image were - [1,1,1,1,2,2,2,4,2]. So the max count logic will predict 4 while max count + max frequency logic will predict 2. There might be a strong possibility of false detections, and thus image number 8 having 4 detections (out of which 2 could be false detections) would be eliminated. The frequency of both 1 and 2 is the same, so the max count will be considered, that is 2. I saw a validation accuracy boost and public score boost using max count + max frequency logic, and hence used in my final submission.\n\nAlso, a higher megadetector detection confidence threshold worked better, so I used 0.7 as detection confidence for the final submission pipeline.\n\n6) Training -\nI experimented with mixed-precision training, but training without mixed-precision yielded better results for me. I did not use any LR scheduler for mini dataset experiments. But for the final training of efficient b5 noisy student, I used reduce on plateau. I underestimated the training time required for the efficient b5 noisy student. I spent all remaining 26 hrs of GPU in the last week in efficient b5 noisy student's training and still, it was a little underfit with ~ 78% val accuracy and 82% train accuracy. After exhausting my Kaggle GPU quota I switched to colab and resumed training. This time I changed some hyperparameters like batch-size and LR, but I could only reach 82% val accuracy and 85% train accuracy due to lack of time (deadline was near). But I couldn't make a submission using colab because I later realized that the test set was 30GB while on colab I had 28GB disk space. I couldn't even download the zip there. So my final model is still underfit (i.e. it needs more training).  \n\n**Here is my final submission pipeline [notebook](https://www.kaggle.com/devashishprasad/iwildcam2021-submission)**",
      "votes": 5
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1330366": "I created a mini dataset (10% of the original dataset) with the same class proportion as in the original. And carried out many experiments to move quickly through the competition.\n\n1) Preprocessing -\nIn my initial experiments, I found that instead of directly cropping and resizing the detections, cropping and padding the detections using reflection padding improved the score. Hence my final solution used reflection padding while training and testing. My preprocessing pipeline was as follows\n\ncrop detection -> reflection padding -> augmentations -> standardize\n\n2) Augmentations -\nIn my experiments adding any type of augmentation reduced my validation score. Hence I kept simple augmentations with very little intensity. The augmentations that I used were as follows \n\nRandom Horizontal Flip (p=0.5)\nColor Jitter : Brightness and Contrast (0.9 to 1.2)\nGrayscale (p=0.5)\nGaussian Blur (sigma = 0.0 to 0.8)\n \n3) Handle Imbalance - \nHere is how I have handled imbalance: https://www.kaggle.com/c/iwildcam2021-fgvc8/discussion/242455\n\n4) Model Architecture -\nFor my mini dataset, I experimented with efficient b2 noisy student. For the final submission, I used the efficient b5 noisy student. I did not use pre-trained image net weights.\n\n5) Submission pipeline -\nIn a public kernel, I came across the max count logic. In this, the maximum number of detections from an image in a sequence (max count) was considered as the prediction value for that sequence. I modified this logic and also considered maximum frequency. For eg - in a sequence of 9 images having only one specie of animal, the number of detections for each image were - [1,1,1,1,2,2,2,4,2]. So the max count logic will predict 4 while max count + max frequency logic will predict 2. There might be a strong possibility of false detections, and thus image number 8 having 4 detections (out of which 2 could be false detections) would be eliminated. The frequency of both 1 and 2 is the same, so the max count will be considered, that is 2. I saw a validation accuracy boost and public score boost using max count + max frequency logic, and hence used in my final submission.\n\nAlso, a higher megadetector detection confidence threshold worked better, so I used 0.7 as detection confidence for the final submission pipeline.\n\n6) Training -\nI experimented with mixed-precision training, but training without mixed-precision yielded better results for me. I did not use any LR scheduler for mini dataset experiments. But for the final training of efficient b5 noisy student, I used reduce on plateau. I underestimated the training time required for the efficient b5 noisy student. I spent all remaining 26 hrs of GPU in the last week in efficient b5 noisy student's training and still, it was a little underfit with ~ 78% val accuracy and 82% train accuracy. After exhausting my Kaggle GPU quota I switched to colab and resumed training. This time I changed some hyperparameters like batch-size and LR, but I could only reach 82% val accuracy and 85% train accuracy due to lack of time (deadline was near). But I couldn't make a submission using colab because I later realized that the test set was 30GB while on colab I had 28GB disk space. I couldn't even download the zip there. So my final model is still underfit (i.e. it needs more training).  \n\n**Here is my final submission pipeline [notebook](https://www.kaggle.com/devashishprasad/iwildcam2021-submission)**"
  }
}