{
  "id": 245460,
  "title": "1st Place Solution",
  "url": "/competitions/iwildcam2021-fgvc8/writeups/ufam-1st-place-solution",
  "author_name": "",
  "post_date": "2021-06-14T12:25:11.287Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Our final solution is divided into two components: a sequence classifier based on an ensemble of EfficientNet-B2 [1] models and a counting heuristic following the one proposed for the competition benchmark.</p>\n<h3>Classifier</h3>\n<p>Following the previous year’s solutions, we trained two EfficientNet-B2 models: one using full images and a second one using square crops from bounding boxes generated by MegaDetectorV4. We used Balanced Group Softmax [2] to handle the class imbalance problem in the dataset. The prediction assigned to each image is the weighted average of the predictions assigned to the bounding box with the highest score and to the full image (0.15 full image + 0.15 full image (flip) + 0.35 bbox + 0.35 bbox (flip)). To classify a sequence, we average the predictions of all nonempty images for each sequence.</p>\n<p><strong>Training procedure</strong>: Before using Balanced Group Softmax, we trained the classifiers using the standard softmax. We initialized EfficientNet-B2 with ImageNet weights and trained it for 20 epochs on the iWildCam2021 training set using the default input resolution (260x260). Then we fine-tuned the last layers for 2 epochs using a higher input resolution (380x380) to fix the train/test resolution [3]. Specifically for the bbox model, we pretrained the classifier layer for 4 epochs before unfreezing all layers to train. During the training, we used label smoothing and SGD with momentum. We also applied warmup and cosine decay to the learning rate. Still for the bbox model, we considered all bounding boxes with confidence &gt; 0.6 as an independent training sample and applied a square crop around it with the size of the largest side. We used the image label as the bounding box label.</p>\n<p><strong>Image preprocessing</strong>: In terms of image preprocessing, we applied random crop, random flip, and RandAugment (N=6, M=2) [4]. During the fix train/test resolution stage, we used the test preprocessing, which consists only of resizing the image/crop to the network input (Keras EfficientNet includes normalization layers).</p>\n<p><strong>Handling the class imbalance</strong>: For balanced group softmax (bags), we grouped classes into 4 softmax according to the number N of training instances: N &lt; 10, 10 &lt; = N &lt; 100, 100 &lt; = N &lt; 1000, N &gt; = 1000. We also included a special softmax for the foreground/background. Following the original bags paper, we also used the class “others” on each softmax to represent instances from all classes not included on that softmax. For the final prediction, we remap all predictions to the original softmax, ignoring the “others” class. As in the original paper, the predictions are not true probabilities, since they do not sum up to one, but we consider the highest value as the bags prediction. We used bags on top of the pretrained models with the standard softmax, which was removed, and we kept all weights frozen. First, we trained the bags classifiers layers for 12 epochs using data augmentation with test input resolution (380x380). Then, we tuned for 2 epochs using test time preprocessing. During the training, we subsampled the “others” category instances to avoid dominating the softmax by sampling a maximum of 8 times the number of instances belonging to categories from that softmax at each batch. The inference used 380x380 as model input resolution.</p>\n<p><strong>Validation</strong>: To validate the models, we split the training set into train/validation according to the locations. We also used the validation set to adjust hyperparameters. For the final submission, we used all training instances to train the model using the tuned hyperparameters.</p>\n<p>Initially, we were using voting among images predictions to classify a sequence, but after adding bags we found out that averaging worked better.</p>\n<h3>Counting heuristic:</h3>\n<p>We followed the competition benchmark counting heuristic: the maximum number of bounding boxes across any image in the sequence. We only counted bounding boxes from MegaDetectorV4 with confidence &gt; 0.8.</p>\n<p>This counting strategy limits us to predict only one species per sequence. We see our solution as a strong baseline for counting animals on camera trap sequences, but we do believe that the best solution should be based on tracking animals across images (Multi-object tracking), classifying each track, and counting them. We tried DeepSORT to track animals, but it was not as good as this heuristics for counting (we describe this procedure below).</p>\n<h3>Solution:</h3>\n<p>Our code, training configs and models are publicly available at <a href=\"https://github.com/alcunha/iwildcam2021ufam\" target=\"_blank\">https://github.com/alcunha/iwildcam2021ufam</a>.</p>\n<h3>Other things that we tried</h3>\n<p><strong>Geo Prior Model</strong>: We tried to use the GPS coordinates and time of year (image timestamps) to train the Geo Prior model [5], but the model overfitted the iWildCam 2021 training set. We tried to add some noise to data by varying the GPS coordinates within 5km radius and timestamp within 10 days. We also replaced the original Geo Prior model loss with the focal loss. However, using only the classifier predictions was better than using them combined with geo priors. We believe GPS coordinates can be useful for the problem (it worked very well for our iNat 2021 solution, for instance), but it's necessary to develop a model to deal with camera trap specificities, such as the fixed position. Our Geo Prior implementation is publicly available at <a href=\"https://github.com/alcunha/geo_prior_tf/\" target=\"_blank\">https://github.com/alcunha/geo_prior_tf/</a>.</p>\n<p><strong>Multi-object tracking with DeepSORT</strong>: We tried to use DeepSORT [6] to generate tracks over sequences and then classify each track by averaging the predictions of its bounding boxes list using the bbox model. We used the feature embedding from EfficientNet-B2 trained on bounding boxes. We also tried to use features from ReID models trained by the winner of the ECCV TAO 2020 challenge, but they performed slightly worse than using EfficientNet-B2 features. Finally, we tried to use DeepSORT removing its dependency on geometry cues as the ECCV winner, but it generated 37000 tracks which is much more than the original DeeepSORT which generated ~4000 or the counting heuristic, ~6000.</p>\n<p>After some analysis on the validation set, we observed that our classifier could be improved, then we got back to it for the final sprint but we had not time to work on DeepSORT. We would like to have tuned better the DeepSORT hyperparameters. For the track classifier, instead of just averaging predictions, we could have trained a classifier using a set of images as inputs. To avoid GPLv3 licensing conflicts we kept our DeepSORT implementation on a separate repository: &nbsp;<a href=\"https://github.com/alcunha/deep_sort_iwildcam2021ufam\" target=\"_blank\">https://github.com/alcunha/deep_sort_iwildcam2021ufam</a>.</p>\n<h4>References</h4>\n<p>[1] Tan, Mingxing, and Quoc Le. \"Efficientnet: Rethinking model scaling for convolutional neural networks.\" International Conference on Machine Learning. PMLR, 2019.</p>\n<p>[2] Li, Yu, et al. \"Overcoming classifier imbalance for long-tail object detection with balanced group softmax.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.</p>\n<p>[3] Touvron, Hugo, et al. \"Fixing the train-test resolution discrepancy.\" arXiv preprint arXiv:1906.06423 (2019).</p>\n<p>[4] Cubuk, Ekin D., et al. \"Randaugment: Practical automated data augmentation with a reduced search space.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 2020.</p>\n<p>[5] Mac Aodha, Oisin, Elijah Cole, and Pietro Perona. \"Presence-only geographical priors for fine-grained image classification.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019.</p>\n<p>[6] Wojke, Nicolai, Alex Bewley, and Dietrich Paulus. \"Simple online and realtime tracking with a deep association metric.\" 2017 IEEE international conference on image processing (ICIP). IEEE, 2017.</p>",
  "messages": [
    {
      "id": "1344527",
      "postDate": "06/11/2021 01:48:45",
      "content": "<p>Our final solution is divided into two components: a sequence classifier based on an ensemble of EfficientNet-B2 [1] models and a counting heuristic following the one proposed for the competition benchmark.</p>\n<h3>Classifier</h3>\n<p>Following the previous year’s solutions, we trained two EfficientNet-B2 models: one using full images and a second one using square crops from bounding boxes generated by MegaDetectorV4. We used Balanced Group Softmax [2] to handle the class imbalance problem in the dataset. The prediction assigned to each image is the weighted average of the predictions assigned to the bounding box with the highest score and to the full image (0.15 full image + 0.15 full image (flip) + 0.35 bbox + 0.35 bbox (flip)). To classify a sequence, we average the predictions of all nonempty images for each sequence.</p>\n<p><strong>Training procedure</strong>: Before using Balanced Group Softmax, we trained the classifiers using the standard softmax. We initialized EfficientNet-B2 with ImageNet weights and trained it for 20 epochs on the iWildCam2021 training set using the default input resolution (260x260). Then we fine-tuned the last layers for 2 epochs using a higher input resolution (380x380) to fix the train/test resolution [3]. Specifically for the bbox model, we pretrained the classifier layer for 4 epochs before unfreezing all layers to train. During the training, we used label smoothing and SGD with momentum. We also applied warmup and cosine decay to the learning rate. Still for the bbox model, we considered all bounding boxes with confidence &gt; 0.6 as an independent training sample and applied a square crop around it with the size of the largest side. We used the image label as the bounding box label.</p>\n<p><strong>Image preprocessing</strong>: In terms of image preprocessing, we applied random crop, random flip, and RandAugment (N=6, M=2) [4]. During the fix train/test resolution stage, we used the test preprocessing, which consists only of resizing the image/crop to the network input (Keras EfficientNet includes normalization layers).</p>\n<p><strong>Handling the class imbalance</strong>: For balanced group softmax (bags), we grouped classes into 4 softmax according to the number N of training instances: N &lt; 10, 10 &lt; = N &lt; 100, 100 &lt; = N &lt; 1000, N &gt; = 1000. We also included a special softmax for the foreground/background. Following the original bags paper, we also used the class “others” on each softmax to represent instances from all classes not included on that softmax. For the final prediction, we remap all predictions to the original softmax, ignoring the “others” class. As in the original paper, the predictions are not true probabilities, since they do not sum up to one, but we consider the highest value as the bags prediction. We used bags on top of the pretrained models with the standard softmax, which was removed, and we kept all weights frozen. First, we trained the bags classifiers layers for 12 epochs using data augmentation with test input resolution (380x380). Then, we tuned for 2 epochs using test time preprocessing. During the training, we subsampled the “others” category instances to avoid dominating the softmax by sampling a maximum of 8 times the number of instances belonging to categories from that softmax at each batch. The inference used 380x380 as model input resolution.</p>\n<p><strong>Validation</strong>: To validate the models, we split the training set into train/validation according to the locations. We also used the validation set to adjust hyperparameters. For the final submission, we used all training instances to train the model using the tuned hyperparameters.</p>\n<p>Initially, we were using voting among images predictions to classify a sequence, but after adding bags we found out that averaging worked better.</p>\n<h3>Counting heuristic:</h3>\n<p>We followed the competition benchmark counting heuristic: the maximum number of bounding boxes across any image in the sequence. We only counted bounding boxes from MegaDetectorV4 with confidence &gt; 0.8.</p>\n<p>This counting strategy limits us to predict only one species per sequence. We see our solution as a strong baseline for counting animals on camera trap sequences, but we do believe that the best solution should be based on tracking animals across images (Multi-object tracking), classifying each track, and counting them. We tried DeepSORT to track animals, but it was not as good as this heuristics for counting (we describe this procedure below).</p>\n<h3>Solution:</h3>\n<p>Our code, training configs and models are publicly available at <a href=\"https://github.com/alcunha/iwildcam2021ufam\" target=\"_blank\">https://github.com/alcunha/iwildcam2021ufam</a>.</p>\n<h3>Other things that we tried</h3>\n<p><strong>Geo Prior Model</strong>: We tried to use the GPS coordinates and time of year (image timestamps) to train the Geo Prior model [5], but the model overfitted the iWildCam 2021 training set. We tried to add some noise to data by varying the GPS coordinates within 5km radius and timestamp within 10 days. We also replaced the original Geo Prior model loss with the focal loss. However, using only the classifier predictions was better than using them combined with geo priors. We believe GPS coordinates can be useful for the problem (it worked very well for our iNat 2021 solution, for instance), but it's necessary to develop a model to deal with camera trap specificities, such as the fixed position. Our Geo Prior implementation is publicly available at <a href=\"https://github.com/alcunha/geo_prior_tf/\" target=\"_blank\">https://github.com/alcunha/geo_prior_tf/</a>.</p>\n<p><strong>Multi-object tracking with DeepSORT</strong>: We tried to use DeepSORT [6] to generate tracks over sequences and then classify each track by averaging the predictions of its bounding boxes list using the bbox model. We used the feature embedding from EfficientNet-B2 trained on bounding boxes. We also tried to use features from ReID models trained by the winner of the ECCV TAO 2020 challenge, but they performed slightly worse than using EfficientNet-B2 features. Finally, we tried to use DeepSORT removing its dependency on geometry cues as the ECCV winner, but it generated 37000 tracks which is much more than the original DeeepSORT which generated ~4000 or the counting heuristic, ~6000.</p>\n<p>After some analysis on the validation set, we observed that our classifier could be improved, then we got back to it for the final sprint but we had not time to work on DeepSORT. We would like to have tuned better the DeepSORT hyperparameters. For the track classifier, instead of just averaging predictions, we could have trained a classifier using a set of images as inputs. To avoid GPLv3 licensing conflicts we kept our DeepSORT implementation on a separate repository: &nbsp;<a href=\"https://github.com/alcunha/deep_sort_iwildcam2021ufam\" target=\"_blank\">https://github.com/alcunha/deep_sort_iwildcam2021ufam</a>.</p>\n<h4>References</h4>\n<p>[1] Tan, Mingxing, and Quoc Le. \"Efficientnet: Rethinking model scaling for convolutional neural networks.\" International Conference on Machine Learning. PMLR, 2019.</p>\n<p>[2] Li, Yu, et al. \"Overcoming classifier imbalance for long-tail object detection with balanced group softmax.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.</p>\n<p>[3] Touvron, Hugo, et al. \"Fixing the train-test resolution discrepancy.\" arXiv preprint arXiv:1906.06423 (2019).</p>\n<p>[4] Cubuk, Ekin D., et al. \"Randaugment: Practical automated data augmentation with a reduced search space.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 2020.</p>\n<p>[5] Mac Aodha, Oisin, Elijah Cole, and Pietro Perona. \"Presence-only geographical priors for fine-grained image classification.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019.</p>\n<p>[6] Wojke, Nicolai, Alex Bewley, and Dietrich Paulus. \"Simple online and realtime tracking with a deep association metric.\" 2017 IEEE international conference on image processing (ICIP). IEEE, 2017.</p>",
      "rawMarkdown": "Our final solution is divided into two components: a sequence classifier based on an ensemble of EfficientNet-B2 [1] models and a counting heuristic following the one proposed for the competition benchmark.\n\n### Classifier\nFollowing the previous year’s solutions, we trained two EfficientNet-B2 models: one using full images and a second one using square crops from bounding boxes generated by MegaDetectorV4. We used Balanced Group Softmax [2] to handle the class imbalance problem in the dataset. The prediction assigned to each image is the weighted average of the predictions assigned to the bounding box with the highest score and to the full image (0.15 full image + 0.15 full image (flip) + 0.35 bbox + 0.35 bbox (flip)). To classify a sequence, we average the predictions of all nonempty images for each sequence.\n\n**Training procedure**: Before using Balanced Group Softmax, we trained the classifiers using the standard softmax. We initialized EfficientNet-B2 with ImageNet weights and trained it for 20 epochs on the iWildCam2021 training set using the default input resolution (260x260). Then we fine-tuned the last layers for 2 epochs using a higher input resolution (380x380) to fix the train/test resolution [3]. Specifically for the bbox model, we pretrained the classifier layer for 4 epochs before unfreezing all layers to train. During the training, we used label smoothing and SGD with momentum. We also applied warmup and cosine decay to the learning rate. Still for the bbox model, we considered all bounding boxes with confidence > 0.6 as an independent training sample and applied a square crop around it with the size of the largest side. We used the image label as the bounding box label.\n\n**Image preprocessing**: In terms of image preprocessing, we applied random crop, random flip, and RandAugment (N=6, M=2) [4]. During the fix train/test resolution stage, we used the test preprocessing, which consists only of resizing the image/crop to the network input (Keras EfficientNet includes normalization layers).\n\n**Handling the class imbalance**: For balanced group softmax (bags), we grouped classes into 4 softmax according to the number N of training instances: N < 10, 10 < = N < 100, 100 < = N < 1000, N > = 1000. We also included a special softmax for the foreground/background. Following the original bags paper, we also used the class “others” on each softmax to represent instances from all classes not included on that softmax. For the final prediction, we remap all predictions to the original softmax, ignoring the “others” class. As in the original paper, the predictions are not true probabilities, since they do not sum up to one, but we consider the highest value as the bags prediction. We used bags on top of the pretrained models with the standard softmax, which was removed, and we kept all weights frozen. First, we trained the bags classifiers layers for 12 epochs using data augmentation with test input resolution (380x380). Then, we tuned for 2 epochs using test time preprocessing. During the training, we subsampled the “others” category instances to avoid dominating the softmax by sampling a maximum of 8 times the number of instances belonging to categories from that softmax at each batch. The inference used 380x380 as model input resolution.\n\n**Validation**: To validate the models, we split the training set into train/validation according to the locations. We also used the validation set to adjust hyperparameters. For the final submission, we used all training instances to train the model using the tuned hyperparameters.\n\nInitially, we were using voting among images predictions to classify a sequence, but after adding bags we found out that averaging worked better.\n\n### Counting heuristic:\nWe followed the competition benchmark counting heuristic: the maximum number of bounding boxes across any image in the sequence. We only counted bounding boxes from MegaDetectorV4 with confidence > 0.8.\n\nThis counting strategy limits us to predict only one species per sequence. We see our solution as a strong baseline for counting animals on camera trap sequences, but we do believe that the best solution should be based on tracking animals across images (Multi-object tracking), classifying each track, and counting them. We tried DeepSORT to track animals, but it was not as good as this heuristics for counting (we describe this procedure below).\n\n### Solution:\nOur code, training configs and models are publicly available at https://github.com/alcunha/iwildcam2021ufam.\n\n### Other things that we tried\n\n**Geo Prior Model**: We tried to use the GPS coordinates and time of year (image timestamps) to train the Geo Prior model [5], but the model overfitted the iWildCam 2021 training set. We tried to add some noise to data by varying the GPS coordinates within 5km radius and timestamp within 10 days. We also replaced the original Geo Prior model loss with the focal loss. However, using only the classifier predictions was better than using them combined with geo priors. We believe GPS coordinates can be useful for the problem (it worked very well for our iNat 2021 solution, for instance), but it's necessary to develop a model to deal with camera trap specificities, such as the fixed position. Our Geo Prior implementation is publicly available at https://github.com/alcunha/geo_prior_tf/.\n\n\n**Multi-object tracking with DeepSORT**: We tried to use DeepSORT [6] to generate tracks over sequences and then classify each track by averaging the predictions of its bounding boxes list using the bbox model. We used the feature embedding from EfficientNet-B2 trained on bounding boxes. We also tried to use features from ReID models trained by the winner of the ECCV TAO 2020 challenge, but they performed slightly worse than using EfficientNet-B2 features. Finally, we tried to use DeepSORT removing its dependency on geometry cues as the ECCV winner, but it generated 37000 tracks which is much more than the original DeeepSORT which generated ~4000 or the counting heuristic, ~6000.\n\nAfter some analysis on the validation set, we observed that our classifier could be improved, then we got back to it for the final sprint but we had not time to work on DeepSORT. We would like to have tuned better the DeepSORT hyperparameters. For the track classifier, instead of just averaging predictions, we could have trained a classifier using a set of images as inputs. To avoid GPLv3 licensing conflicts we kept our DeepSORT implementation on a separate repository:  https://github.com/alcunha/deep_sort_iwildcam2021ufam.\n\n\n#### References\n[1] Tan, Mingxing, and Quoc Le. \"Efficientnet: Rethinking model scaling for convolutional neural networks.\" International Conference on Machine Learning. PMLR, 2019.\n\n[2] Li, Yu, et al. \"Overcoming classifier imbalance for long-tail object detection with balanced group softmax.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.\n\n[3] Touvron, Hugo, et al. \"Fixing the train-test resolution discrepancy.\" arXiv preprint arXiv:1906.06423 (2019).\n\n[4] Cubuk, Ekin D., et al. \"Randaugment: Practical automated data augmentation with a reduced search space.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 2020.\n\n[5] Mac Aodha, Oisin, Elijah Cole, and Pietro Perona. \"Presence-only geographical priors for fine-grained image classification.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019.\n\n[6] Wojke, Nicolai, Alex Bewley, and Dietrich Paulus. \"Simple online and realtime tracking with a deep association metric.\" 2017 IEEE international conference on image processing (ICIP). IEEE, 2017.",
      "votes": null
    },
    {
      "id": "1348438",
      "postDate": "06/14/2021 04:14:01",
      "content": "<p>What an elegant solution, well deserved first place!<br>\nThank you so much for sharing your code, it's a great help to me.</p>",
      "rawMarkdown": "What an elegant solution, well deserved first place!\nThank you so much for sharing your code, it's a great help to me.",
      "votes": null
    },
    {
      "id": "1354471",
      "postDate": "06/17/2021 15:53:49",
      "content": "<p>Thank you so much for sharing the solution!! Congratulations!!</p>",
      "rawMarkdown": "Thank you so much for sharing the solution!! Congratulations!!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1348438,
      "author_name": "jcle609",
      "author_url": "",
      "post_date": "06/14/2021 04:14:01",
      "content": "<p>What an elegant solution, well deserved first place!<br>\nThank you so much for sharing your code, it's a great help to me.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1354471,
      "author_name": "devashishprasad",
      "author_url": "",
      "post_date": "06/17/2021 15:53:49",
      "content": "<p>Thank you so much for sharing the solution!! Congratulations!!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1344527": "Our final solution is divided into two components: a sequence classifier based on an ensemble of EfficientNet-B2 [1] models and a counting heuristic following the one proposed for the competition benchmark.\n\n### Classifier\nFollowing the previous year’s solutions, we trained two EfficientNet-B2 models: one using full images and a second one using square crops from bounding boxes generated by MegaDetectorV4. We used Balanced Group Softmax [2] to handle the class imbalance problem in the dataset. The prediction assigned to each image is the weighted average of the predictions assigned to the bounding box with the highest score and to the full image (0.15 full image + 0.15 full image (flip) + 0.35 bbox + 0.35 bbox (flip)). To classify a sequence, we average the predictions of all nonempty images for each sequence.\n\n**Training procedure**: Before using Balanced Group Softmax, we trained the classifiers using the standard softmax. We initialized EfficientNet-B2 with ImageNet weights and trained it for 20 epochs on the iWildCam2021 training set using the default input resolution (260x260). Then we fine-tuned the last layers for 2 epochs using a higher input resolution (380x380) to fix the train/test resolution [3]. Specifically for the bbox model, we pretrained the classifier layer for 4 epochs before unfreezing all layers to train. During the training, we used label smoothing and SGD with momentum. We also applied warmup and cosine decay to the learning rate. Still for the bbox model, we considered all bounding boxes with confidence > 0.6 as an independent training sample and applied a square crop around it with the size of the largest side. We used the image label as the bounding box label.\n\n**Image preprocessing**: In terms of image preprocessing, we applied random crop, random flip, and RandAugment (N=6, M=2) [4]. During the fix train/test resolution stage, we used the test preprocessing, which consists only of resizing the image/crop to the network input (Keras EfficientNet includes normalization layers).\n\n**Handling the class imbalance**: For balanced group softmax (bags), we grouped classes into 4 softmax according to the number N of training instances: N < 10, 10 < = N < 100, 100 < = N < 1000, N > = 1000. We also included a special softmax for the foreground/background. Following the original bags paper, we also used the class “others” on each softmax to represent instances from all classes not included on that softmax. For the final prediction, we remap all predictions to the original softmax, ignoring the “others” class. As in the original paper, the predictions are not true probabilities, since they do not sum up to one, but we consider the highest value as the bags prediction. We used bags on top of the pretrained models with the standard softmax, which was removed, and we kept all weights frozen. First, we trained the bags classifiers layers for 12 epochs using data augmentation with test input resolution (380x380). Then, we tuned for 2 epochs using test time preprocessing. During the training, we subsampled the “others” category instances to avoid dominating the softmax by sampling a maximum of 8 times the number of instances belonging to categories from that softmax at each batch. The inference used 380x380 as model input resolution.\n\n**Validation**: To validate the models, we split the training set into train/validation according to the locations. We also used the validation set to adjust hyperparameters. For the final submission, we used all training instances to train the model using the tuned hyperparameters.\n\nInitially, we were using voting among images predictions to classify a sequence, but after adding bags we found out that averaging worked better.\n\n### Counting heuristic:\nWe followed the competition benchmark counting heuristic: the maximum number of bounding boxes across any image in the sequence. We only counted bounding boxes from MegaDetectorV4 with confidence > 0.8.\n\nThis counting strategy limits us to predict only one species per sequence. We see our solution as a strong baseline for counting animals on camera trap sequences, but we do believe that the best solution should be based on tracking animals across images (Multi-object tracking), classifying each track, and counting them. We tried DeepSORT to track animals, but it was not as good as this heuristics for counting (we describe this procedure below).\n\n### Solution:\nOur code, training configs and models are publicly available at https://github.com/alcunha/iwildcam2021ufam.\n\n### Other things that we tried\n\n**Geo Prior Model**: We tried to use the GPS coordinates and time of year (image timestamps) to train the Geo Prior model [5], but the model overfitted the iWildCam 2021 training set. We tried to add some noise to data by varying the GPS coordinates within 5km radius and timestamp within 10 days. We also replaced the original Geo Prior model loss with the focal loss. However, using only the classifier predictions was better than using them combined with geo priors. We believe GPS coordinates can be useful for the problem (it worked very well for our iNat 2021 solution, for instance), but it's necessary to develop a model to deal with camera trap specificities, such as the fixed position. Our Geo Prior implementation is publicly available at https://github.com/alcunha/geo_prior_tf/.\n\n\n**Multi-object tracking with DeepSORT**: We tried to use DeepSORT [6] to generate tracks over sequences and then classify each track by averaging the predictions of its bounding boxes list using the bbox model. We used the feature embedding from EfficientNet-B2 trained on bounding boxes. We also tried to use features from ReID models trained by the winner of the ECCV TAO 2020 challenge, but they performed slightly worse than using EfficientNet-B2 features. Finally, we tried to use DeepSORT removing its dependency on geometry cues as the ECCV winner, but it generated 37000 tracks which is much more than the original DeeepSORT which generated ~4000 or the counting heuristic, ~6000.\n\nAfter some analysis on the validation set, we observed that our classifier could be improved, then we got back to it for the final sprint but we had not time to work on DeepSORT. We would like to have tuned better the DeepSORT hyperparameters. For the track classifier, instead of just averaging predictions, we could have trained a classifier using a set of images as inputs. To avoid GPLv3 licensing conflicts we kept our DeepSORT implementation on a separate repository:  https://github.com/alcunha/deep_sort_iwildcam2021ufam.\n\n\n#### References\n[1] Tan, Mingxing, and Quoc Le. \"Efficientnet: Rethinking model scaling for convolutional neural networks.\" International Conference on Machine Learning. PMLR, 2019.\n\n[2] Li, Yu, et al. \"Overcoming classifier imbalance for long-tail object detection with balanced group softmax.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020.\n\n[3] Touvron, Hugo, et al. \"Fixing the train-test resolution discrepancy.\" arXiv preprint arXiv:1906.06423 (2019).\n\n[4] Cubuk, Ekin D., et al. \"Randaugment: Practical automated data augmentation with a reduced search space.\" Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops. 2020.\n\n[5] Mac Aodha, Oisin, Elijah Cole, and Pietro Perona. \"Presence-only geographical priors for fine-grained image classification.\" Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019.\n\n[6] Wojke, Nicolai, Alex Bewley, and Dietrich Paulus. \"Simple online and realtime tracking with a deep association metric.\" 2017 IEEE international conference on image processing (ICIP). IEEE, 2017.",
    "1348438": "What an elegant solution, well deserved first place!\nThank you so much for sharing your code, it's a great help to me.",
    "1354471": "Thank you so much for sharing the solution!! Congratulations!!"
  },
  "source": "meta"
}