{
  "id": 156420,
  "title": "LB #4 Documentation",
  "url": "/competitions/iwildcam-2020-fgvc7/discussion/156420",
  "author_name": "Fagner Cunha",
  "post_date": "2020-06-05T23:21:14.992000",
  "votes": 10,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Our final solution is based on a simple InceptionV3 model pre-trained on Imagenet. This model was trained on the camera trap images whose animals were cropped out using the provided Megadetector bounding boxes as reference. Then, we choose the model that reached the best performance on a validation set, which was split from the training set taking into account the locations. The final prediction assigned for each image is obtained by using the majority voting among images within the same sequence. Due to the fact that the provided sequence data is noisy, we have used location and timestamp information to determine which images should belong to the same sequence.</p>\n\n<h3>Dataset splitting</h3>\n\n<p>Our validation set has approximately 11% of the original images of the training set. We have used a heuristic to try to keep both species and background diversity in the training and validation sets, described as follows. First, all the locations where the number of instances represents more than 1% of the total are added to the training set. This represents around 50% of the original training set. Then, for each class we randomly select locations to compose the validation set, one at once, until at least 5% of instances of the current class are added to the validation set. After that, the remaining locations available for that class are used to compose the new training set. Classes appearing on less than 6 locations or with less than 30 instances were ignored during the splitting process. The classes 'empty', 'start', 'end', 'unknown', and 'unidentifiable' were ignored as well. This heuristic does not produce splits following the original class distribution. However, we do not expect the test set to follow exactly the same distribution either. It also leads many classes not appearing in the validation set, especially those classes with fewer instances. Nonetheless, we have only removed classes that appear in the validation set but no in the training set. At the end, the training set was composed of 211 classes from 275 locations and the validation set was composed of 114 classes from 49 locations.</p>\n\n<h3>Megadetector bounding boxes</h3>\n\n<p>We have used the Megadetector results to crop images out before training. As the image resolutions vary along the dataset, we decided to use a square crop around the bounding box generated by Megadetector. The square size is the same as the largest dimension of the bounding box. Additionally, we have limited the square crop to have at least 450x450 of resolution and the maximum of the height of each image. In addition, the square crop will always keep the bounding box area centered unless the crop gets out the image. In those cases, we move the crop to get valid pixels and the animal is no more centered. For the images without a bounding box, we just crop a centered square with the image height. This procedure was applied on all images before training, including images from the test set.</p>\n\n<h3>Model training</h3>\n\n<p>We have used an InceptionV3 model pre-trained on Imagenet with input size 299x299, which was trained on our training subset for 48 epochs. The best accuracy model among all epochs was selected as the final model. Data augmentation was used during the training, using the following transformations applied to the input images: zoom, rotation, horizontal flip, shearing, horizontal and vertical shifts. The model was trained using the SGD optimizer with momentum of 0.9 and a simple learning rate schedule based on epochs, similar to the procedure employed in the paper “Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning”. The learning rate schedule was 1-5: 0.005, 6-10: 0.001, 11-20: 0.005, 21-27: 0.001, 28-35: 0.0005, 36-onwards:  0.0001.</p>\n\n<h3>Majority voting among images from the same sequence</h3>\n\n<p>To determine the final prediction for each image, we first check for empty images. If our model predicts “empty” and Megadetector does not provide bounding box, the class “empty” is assigned to the input image. Otherwise, majority vote is calculated considering all images from the same sequence, excluding the ones already marked as empty. Images are considered as belonging to the same sequence when they are from the same location and the difference between their timestamps is at most 30 minutes.</p>\n\n<h3>What we would like to have tested</h3>\n\n<p><strong>Big Transfer (BiT)</strong>. This new model seems promising for transfer learning, requiring fewer instances per class to reach high performance. Unfortunately, we were unaware of it until few days before the deadline of the competition. Consequently, we weren’t able to test it in time.</p>\n\n<p><strong>A model employing all images from the same sequence to identify animal species</strong>. Depending on the position of the animal on the scene, it may be very difficult even for experts to identify the species. However, if the same individual can be identified with high confidence on another image of the same capture event, this classification can be used for the whole event. An approach that uses this temporal information is described in the paper “Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection”, but we didn’t try it.</p>\n\n<p><strong>MegadetectorV4 bounding boxes</strong>. We would like to have tested the new version of the Megadetector. We think this new version could have generated bounding boxes more accurately. It also includes a new vehicle class, which could have improved our results for this class.</p>",
  "messages": [
    {
      "id": 875567,
      "postDate": "2020-06-05T23:21:14.993Z",
      "content": "<p>Our final solution is based on a simple InceptionV3 model pre-trained on Imagenet. This model was trained on the camera trap images whose animals were cropped out using the provided Megadetector bounding boxes as reference. Then, we choose the model that reached the best performance on a validation set, which was split from the training set taking into account the locations. The final prediction assigned for each image is obtained by using the majority voting among images within the same sequence. Due to the fact that the provided sequence data is noisy, we have used location and timestamp information to determine which images should belong to the same sequence.</p>\n\n<h3>Dataset splitting</h3>\n\n<p>Our validation set has approximately 11% of the original images of the training set. We have used a heuristic to try to keep both species and background diversity in the training and validation sets, described as follows. First, all the locations where the number of instances represents more than 1% of the total are added to the training set. This represents around 50% of the original training set. Then, for each class we randomly select locations to compose the validation set, one at once, until at least 5% of instances of the current class are added to the validation set. After that, the remaining locations available for that class are used to compose the new training set. Classes appearing on less than 6 locations or with less than 30 instances were ignored during the splitting process. The classes 'empty', 'start', 'end', 'unknown', and 'unidentifiable' were ignored as well. This heuristic does not produce splits following the original class distribution. However, we do not expect the test set to follow exactly the same distribution either. It also leads many classes not appearing in the validation set, especially those classes with fewer instances. Nonetheless, we have only removed classes that appear in the validation set but no in the training set. At the end, the training set was composed of 211 classes from 275 locations and the validation set was composed of 114 classes from 49 locations.</p>\n\n<h3>Megadetector bounding boxes</h3>\n\n<p>We have used the Megadetector results to crop images out before training. As the image resolutions vary along the dataset, we decided to use a square crop around the bounding box generated by Megadetector. The square size is the same as the largest dimension of the bounding box. Additionally, we have limited the square crop to have at least 450x450 of resolution and the maximum of the height of each image. In addition, the square crop will always keep the bounding box area centered unless the crop gets out the image. In those cases, we move the crop to get valid pixels and the animal is no more centered. For the images without a bounding box, we just crop a centered square with the image height. This procedure was applied on all images before training, including images from the test set.</p>\n\n<h3>Model training</h3>\n\n<p>We have used an InceptionV3 model pre-trained on Imagenet with input size 299x299, which was trained on our training subset for 48 epochs. The best accuracy model among all epochs was selected as the final model. Data augmentation was used during the training, using the following transformations applied to the input images: zoom, rotation, horizontal flip, shearing, horizontal and vertical shifts. The model was trained using the SGD optimizer with momentum of 0.9 and a simple learning rate schedule based on epochs, similar to the procedure employed in the paper “Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning”. The learning rate schedule was 1-5: 0.005, 6-10: 0.001, 11-20: 0.005, 21-27: 0.001, 28-35: 0.0005, 36-onwards:  0.0001.</p>\n\n<h3>Majority voting among images from the same sequence</h3>\n\n<p>To determine the final prediction for each image, we first check for empty images. If our model predicts “empty” and Megadetector does not provide bounding box, the class “empty” is assigned to the input image. Otherwise, majority vote is calculated considering all images from the same sequence, excluding the ones already marked as empty. Images are considered as belonging to the same sequence when they are from the same location and the difference between their timestamps is at most 30 minutes.</p>\n\n<h3>What we would like to have tested</h3>\n\n<p><strong>Big Transfer (BiT)</strong>. This new model seems promising for transfer learning, requiring fewer instances per class to reach high performance. Unfortunately, we were unaware of it until few days before the deadline of the competition. Consequently, we weren’t able to test it in time.</p>\n\n<p><strong>A model employing all images from the same sequence to identify animal species</strong>. Depending on the position of the animal on the scene, it may be very difficult even for experts to identify the species. However, if the same individual can be identified with high confidence on another image of the same capture event, this classification can be used for the whole event. An approach that uses this temporal information is described in the paper “Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection”, but we didn’t try it.</p>\n\n<p><strong>MegadetectorV4 bounding boxes</strong>. We would like to have tested the new version of the Megadetector. We think this new version could have generated bounding boxes more accurately. It also includes a new vehicle class, which could have improved our results for this class.</p>",
      "rawMarkdown": "Our final solution is based on a simple InceptionV3 model pre-trained on Imagenet. This model was trained on the camera trap images whose animals were cropped out using the provided Megadetector bounding boxes as reference. Then, we choose the model that reached the best performance on a validation set, which was split from the training set taking into account the locations. The final prediction assigned for each image is obtained by using the majority voting among images within the same sequence. Due to the fact that the provided sequence data is noisy, we have used location and timestamp information to determine which images should belong to the same sequence.\n\n### Dataset splitting\n\nOur validation set has approximately 11% of the original images of the training set. We have used a heuristic to try to keep both species and background diversity in the training and validation sets, described as follows. First, all the locations where the number of instances represents more than 1% of the total are added to the training set. This represents around 50% of the original training set. Then, for each class we randomly select locations to compose the validation set, one at once, until at least 5% of instances of the current class are added to the validation set. After that, the remaining locations available for that class are used to compose the new training set. Classes appearing on less than 6 locations or with less than 30 instances were ignored during the splitting process. The classes 'empty', 'start', 'end', 'unknown', and 'unidentifiable' were ignored as well. This heuristic does not produce splits following the original class distribution. However, we do not expect the test set to follow exactly the same distribution either. It also leads many classes not appearing in the validation set, especially those classes with fewer instances. Nonetheless, we have only removed classes that appear in the validation set but no in the training set. At the end, the training set was composed of 211 classes from 275 locations and the validation set was composed of 114 classes from 49 locations.\n\n### Megadetector bounding boxes\n\nWe have used the Megadetector results to crop images out before training. As the image resolutions vary along the dataset, we decided to use a square crop around the bounding box generated by Megadetector. The square size is the same as the largest dimension of the bounding box. Additionally, we have limited the square crop to have at least 450x450 of resolution and the maximum of the height of each image. In addition, the square crop will always keep the bounding box area centered unless the crop gets out the image. In those cases, we move the crop to get valid pixels and the animal is no more centered. For the images without a bounding box, we just crop a centered square with the image height. This procedure was applied on all images before training, including images from the test set.\n\n \n### Model training\n\nWe have used an InceptionV3 model pre-trained on Imagenet with input size 299x299, which was trained on our training subset for 48 epochs. The best accuracy model among all epochs was selected as the final model. Data augmentation was used during the training, using the following transformations applied to the input images: zoom, rotation, horizontal flip, shearing, horizontal and vertical shifts. The model was trained using the SGD optimizer with momentum of 0.9 and a simple learning rate schedule based on epochs, similar to the procedure employed in the paper “Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning”. The learning rate schedule was 1-5: 0.005, 6-10: 0.001, 11-20: 0.005, 21-27: 0.001, 28-35: 0.0005, 36-onwards:  0.0001.\n\n### Majority voting among images from the same sequence\n\nTo determine the final prediction for each image, we first check for empty images. If our model predicts “empty” and Megadetector does not provide bounding box, the class “empty” is assigned to the input image. Otherwise, majority vote is calculated considering all images from the same sequence, excluding the ones already marked as empty. Images are considered as belonging to the same sequence when they are from the same location and the difference between their timestamps is at most 30 minutes.\n\n### What we would like to have tested\n\n**Big Transfer (BiT)**. This new model seems promising for transfer learning, requiring fewer instances per class to reach high performance. Unfortunately, we were unaware of it until few days before the deadline of the competition. Consequently, we weren’t able to test it in time.\n\n**A model employing all images from the same sequence to identify animal species**. Depending on the position of the animal on the scene, it may be very difficult even for experts to identify the species. However, if the same individual can be identified with high confidence on another image of the same capture event, this classification can be used for the whole event. An approach that uses this temporal information is described in the paper “Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection”, but we didn’t try it.\n\n**MegadetectorV4 bounding boxes**. We would like to have tested the new version of the Megadetector. We think this new version could have generated bounding boxes more accurately. It also includes a new vehicle class, which could have improved our results for this class.",
      "votes": 10
    },
    {
      "id": 887654,
      "postDate": "2020-06-15T19:37:12.027Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 887654,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-15T19:37:12.027000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "875567": "Our final solution is based on a simple InceptionV3 model pre-trained on Imagenet. This model was trained on the camera trap images whose animals were cropped out using the provided Megadetector bounding boxes as reference. Then, we choose the model that reached the best performance on a validation set, which was split from the training set taking into account the locations. The final prediction assigned for each image is obtained by using the majority voting among images within the same sequence. Due to the fact that the provided sequence data is noisy, we have used location and timestamp information to determine which images should belong to the same sequence.\n\n### Dataset splitting\n\nOur validation set has approximately 11% of the original images of the training set. We have used a heuristic to try to keep both species and background diversity in the training and validation sets, described as follows. First, all the locations where the number of instances represents more than 1% of the total are added to the training set. This represents around 50% of the original training set. Then, for each class we randomly select locations to compose the validation set, one at once, until at least 5% of instances of the current class are added to the validation set. After that, the remaining locations available for that class are used to compose the new training set. Classes appearing on less than 6 locations or with less than 30 instances were ignored during the splitting process. The classes 'empty', 'start', 'end', 'unknown', and 'unidentifiable' were ignored as well. This heuristic does not produce splits following the original class distribution. However, we do not expect the test set to follow exactly the same distribution either. It also leads many classes not appearing in the validation set, especially those classes with fewer instances. Nonetheless, we have only removed classes that appear in the validation set but no in the training set. At the end, the training set was composed of 211 classes from 275 locations and the validation set was composed of 114 classes from 49 locations.\n\n### Megadetector bounding boxes\n\nWe have used the Megadetector results to crop images out before training. As the image resolutions vary along the dataset, we decided to use a square crop around the bounding box generated by Megadetector. The square size is the same as the largest dimension of the bounding box. Additionally, we have limited the square crop to have at least 450x450 of resolution and the maximum of the height of each image. In addition, the square crop will always keep the bounding box area centered unless the crop gets out the image. In those cases, we move the crop to get valid pixels and the animal is no more centered. For the images without a bounding box, we just crop a centered square with the image height. This procedure was applied on all images before training, including images from the test set.\n\n \n### Model training\n\nWe have used an InceptionV3 model pre-trained on Imagenet with input size 299x299, which was trained on our training subset for 48 epochs. The best accuracy model among all epochs was selected as the final model. Data augmentation was used during the training, using the following transformations applied to the input images: zoom, rotation, horizontal flip, shearing, horizontal and vertical shifts. The model was trained using the SGD optimizer with momentum of 0.9 and a simple learning rate schedule based on epochs, similar to the procedure employed in the paper “Automatically identifying, counting, and describing wild animals in camera-trap images with deep learning”. The learning rate schedule was 1-5: 0.005, 6-10: 0.001, 11-20: 0.005, 21-27: 0.001, 28-35: 0.0005, 36-onwards:  0.0001.\n\n### Majority voting among images from the same sequence\n\nTo determine the final prediction for each image, we first check for empty images. If our model predicts “empty” and Megadetector does not provide bounding box, the class “empty” is assigned to the input image. Otherwise, majority vote is calculated considering all images from the same sequence, excluding the ones already marked as empty. Images are considered as belonging to the same sequence when they are from the same location and the difference between their timestamps is at most 30 minutes.\n\n### What we would like to have tested\n\n**Big Transfer (BiT)**. This new model seems promising for transfer learning, requiring fewer instances per class to reach high performance. Unfortunately, we were unaware of it until few days before the deadline of the competition. Consequently, we weren’t able to test it in time.\n\n**A model employing all images from the same sequence to identify animal species**. Depending on the position of the animal on the scene, it may be very difficult even for experts to identify the species. However, if the same individual can be identified with high confidence on another image of the same capture event, this classification can be used for the whole event. An approach that uses this temporal information is described in the paper “Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection”, but we didn’t try it.\n\n**MegadetectorV4 bounding boxes**. We would like to have tested the new version of the Megadetector. We think this new version could have generated bounding boxes more accurately. It also includes a new vehicle class, which could have improved our results for this class.",
    "887654": ""
  }
}