{
  "id": 154211,
  "title": "LB #8 Documentation",
  "url": "/competitions/iwildcam-2020-fgvc7/discussion/154211",
  "author_name": "Justin Kay",
  "post_date": "2020-05-27T16:05:24.203000",
  "votes": 9,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi everyone! Thanks to the organizers for putting together such a fun competition and really interesting set of data. Here is our submission (rank #8 or #10 depending on whether you count the [Deleted] entries) and some other things we tried / wanted to try.</p>\n\n<p>Our submission was a pretty simple ensemble of a resnet 152 and a resnet 101.</p>\n\n<p>We used all of the provided Megadetector annotations to create a dataset of cropped images with species labels. We then split the data into a training and validation set (just 2 fold), making sure the validation set came from distinct locations and keeping the class distributions approximately the same. We used a ‘mini’ version of each of these (~5-10% of the images) as a mini-training and mini-validation set, which we did our hyperparameter tuning on. </p>\n\n<p>The resnets were pretrained on ImageNet and finetuned on the camera trap crops with pretty heavy data augmentation (random flip, random rotate, random zoom, random contrast and lighting modification, random warp, and mixup). Both models were further fine-tuned for a couple epochs with a class-balanced subset of crops which contained no more than 100 examples per class.</p>\n\n<p>Within each image, each Megadetector detection is cropped, sent through the classifiers, and then weighted by Megadetector’s detection confidence. We tuned how much to weight this using a simple ‘diminishing factor’ <em>d</em>: (detector_confidence + d) / (d + 1) * crop_prediction.</p>\n\n<p><strong>(Here’s the interesting part that didn’t really go anywhere)</strong>\nWe then multiply each image’s prediction vector by a set of ‘geographical priors’ we constructed in a similar fashion to host Elijah’s paper <a href=\"https://arxiv.org/abs/1906.05272\">Presence-Only Geographical Priors for Fine-Grained Image Classification</a>. Since we didn’t have latitude and longitude information, we used the datetime (encoded to sin/cos as in the paper) and the mean normalized difference vegetation index (basically infrared minus red) of each satellite image to construct the geographical priors model. We used a ‘diminishing factor’ for these priors too, but only ever saw about a ~1% increase in our validation score, and a 0.2% increase in our test score. But we thought this was interesting and are excited to see what others tried with the satellite imagery.</p>\n\n<p>We tried a couple other methods in addition to mean NDVI that didn't work as well. First, we tried just using resnet-extracted image features for each satellite image as features for the model (replacing latitude and longitude in the paper’s model) - but, since the data is ‘presence-only’ the model relies upon generated negative examples, which we couldn’t really generate reasonably for the image features. We tried just selecting random satellite images from the dataset to use as ‘negatives’, but this is such a limited number of data points it didn’t seem to work. And just generating a random length-2048 feature vector as a negative seemed silly, since it seems unlikely that a satellite image would ever generate those features.</p>\n\n<p>Then we tried the mean NDVI thing. Cool. As a sidenote, we eliminated all the cloud, shadow, and 0-value pixels from the satellite images using the QA data, so the NDVI we calculate is only for visible land.</p>\n\n<p>Then we tried making that more complicated. Instead of using the mean NDVI (one value per image), we modeled a normal distribution of pixel-wise NDVI values over each satellite image, created a cumulative distribution function for those values for each image, and sampled 100 points from the CDF for features. That allowed us to generate more complex negative examples by generating random normal distributions with a mean between -1 and 1 (the values that the NDVI can take).  This seemed like it was helping, as the validation loss of the geo priors model went down…but when we actually used it for priors for classification our score went down.</p>\n\n<p>Anyway, this didn’t really help much, but I thought it was cool. Would love to chat more if anyone has any thoughts on this (or maybe the hosts can tell me why my ideas are bad :P )</p>\n\n<p><strong>(Back to the boring part)</strong>\nWe then perform a moving average of the image-wise prediction vectors over each sequence. So, each image’s prediction becomes the average of the images nearest to it. We tuned this window size.</p>\n\n<p>We then tried a majority vote ‘with dissenters’ within sequences (this may have a real name). It’s just a majority vote for the final class of each sequence, but if an image has a differing classification with a confidence higher than our ‘dissenter threshold’, we let it keep its prediction. This was meant to help sequences that were mostly empty save for a few frames. We tuned this threshold as well as a ‘handicap’ for non-empty images, such that a majority vote can only be for ‘empty’ if it has at least ‘handicap’ more empty images than the 2nd place class in the vote.</p>\n\n<p><strong>What we wanted to try</strong>\nSara’s paper! <a href=\"https://arxiv.org/abs/1912.03538\">Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection</a> \nWe actually implemented a modification of this which just uses the memory banks and attention networks during classification (still using Megadetector for detections). Unfortunately we entered this thing pretty late and didn’t have time to train it….but we’ll give it a go as a ‘Late Submission’ :)</p>\n\n<p>Well, that was a spiel. If you read through hopefully it’s either interesting, or you’re reading this a year from now as a starting point for iWildcam 2021. Heh</p>\n\n<p>So long, and thanks for all the ocellated turkey….</p>",
  "messages": [
    {
      "id": 863874,
      "postDate": "2020-05-27T16:05:24.203Z",
      "content": "<p>Hi everyone! Thanks to the organizers for putting together such a fun competition and really interesting set of data. Here is our submission (rank #8 or #10 depending on whether you count the [Deleted] entries) and some other things we tried / wanted to try.</p>\n\n<p>Our submission was a pretty simple ensemble of a resnet 152 and a resnet 101.</p>\n\n<p>We used all of the provided Megadetector annotations to create a dataset of cropped images with species labels. We then split the data into a training and validation set (just 2 fold), making sure the validation set came from distinct locations and keeping the class distributions approximately the same. We used a ‘mini’ version of each of these (~5-10% of the images) as a mini-training and mini-validation set, which we did our hyperparameter tuning on. </p>\n\n<p>The resnets were pretrained on ImageNet and finetuned on the camera trap crops with pretty heavy data augmentation (random flip, random rotate, random zoom, random contrast and lighting modification, random warp, and mixup). Both models were further fine-tuned for a couple epochs with a class-balanced subset of crops which contained no more than 100 examples per class.</p>\n\n<p>Within each image, each Megadetector detection is cropped, sent through the classifiers, and then weighted by Megadetector’s detection confidence. We tuned how much to weight this using a simple ‘diminishing factor’ <em>d</em>: (detector_confidence + d) / (d + 1) * crop_prediction.</p>\n\n<p><strong>(Here’s the interesting part that didn’t really go anywhere)</strong>\nWe then multiply each image’s prediction vector by a set of ‘geographical priors’ we constructed in a similar fashion to host Elijah’s paper <a href=\"https://arxiv.org/abs/1906.05272\">Presence-Only Geographical Priors for Fine-Grained Image Classification</a>. Since we didn’t have latitude and longitude information, we used the datetime (encoded to sin/cos as in the paper) and the mean normalized difference vegetation index (basically infrared minus red) of each satellite image to construct the geographical priors model. We used a ‘diminishing factor’ for these priors too, but only ever saw about a ~1% increase in our validation score, and a 0.2% increase in our test score. But we thought this was interesting and are excited to see what others tried with the satellite imagery.</p>\n\n<p>We tried a couple other methods in addition to mean NDVI that didn't work as well. First, we tried just using resnet-extracted image features for each satellite image as features for the model (replacing latitude and longitude in the paper’s model) - but, since the data is ‘presence-only’ the model relies upon generated negative examples, which we couldn’t really generate reasonably for the image features. We tried just selecting random satellite images from the dataset to use as ‘negatives’, but this is such a limited number of data points it didn’t seem to work. And just generating a random length-2048 feature vector as a negative seemed silly, since it seems unlikely that a satellite image would ever generate those features.</p>\n\n<p>Then we tried the mean NDVI thing. Cool. As a sidenote, we eliminated all the cloud, shadow, and 0-value pixels from the satellite images using the QA data, so the NDVI we calculate is only for visible land.</p>\n\n<p>Then we tried making that more complicated. Instead of using the mean NDVI (one value per image), we modeled a normal distribution of pixel-wise NDVI values over each satellite image, created a cumulative distribution function for those values for each image, and sampled 100 points from the CDF for features. That allowed us to generate more complex negative examples by generating random normal distributions with a mean between -1 and 1 (the values that the NDVI can take).  This seemed like it was helping, as the validation loss of the geo priors model went down…but when we actually used it for priors for classification our score went down.</p>\n\n<p>Anyway, this didn’t really help much, but I thought it was cool. Would love to chat more if anyone has any thoughts on this (or maybe the hosts can tell me why my ideas are bad :P )</p>\n\n<p><strong>(Back to the boring part)</strong>\nWe then perform a moving average of the image-wise prediction vectors over each sequence. So, each image’s prediction becomes the average of the images nearest to it. We tuned this window size.</p>\n\n<p>We then tried a majority vote ‘with dissenters’ within sequences (this may have a real name). It’s just a majority vote for the final class of each sequence, but if an image has a differing classification with a confidence higher than our ‘dissenter threshold’, we let it keep its prediction. This was meant to help sequences that were mostly empty save for a few frames. We tuned this threshold as well as a ‘handicap’ for non-empty images, such that a majority vote can only be for ‘empty’ if it has at least ‘handicap’ more empty images than the 2nd place class in the vote.</p>\n\n<p><strong>What we wanted to try</strong>\nSara’s paper! <a href=\"https://arxiv.org/abs/1912.03538\">Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection</a> \nWe actually implemented a modification of this which just uses the memory banks and attention networks during classification (still using Megadetector for detections). Unfortunately we entered this thing pretty late and didn’t have time to train it….but we’ll give it a go as a ‘Late Submission’ :)</p>\n\n<p>Well, that was a spiel. If you read through hopefully it’s either interesting, or you’re reading this a year from now as a starting point for iWildcam 2021. Heh</p>\n\n<p>So long, and thanks for all the ocellated turkey….</p>",
      "rawMarkdown": "Hi everyone! Thanks to the organizers for putting together such a fun competition and really interesting set of data. Here is our submission (rank #8 or #10 depending on whether you count the [Deleted] entries) and some other things we tried / wanted to try.\n\nOur submission was a pretty simple ensemble of a resnet 152 and a resnet 101.\n\nWe used all of the provided Megadetector annotations to create a dataset of cropped images with species labels. We then split the data into a training and validation set (just 2 fold), making sure the validation set came from distinct locations and keeping the class distributions approximately the same. We used a ‘mini’ version of each of these (~5-10% of the images) as a mini-training and mini-validation set, which we did our hyperparameter tuning on. \n\nThe resnets were pretrained on ImageNet and finetuned on the camera trap crops with pretty heavy data augmentation (random flip, random rotate, random zoom, random contrast and lighting modification, random warp, and mixup). Both models were further fine-tuned for a couple epochs with a class-balanced subset of crops which contained no more than 100 examples per class.\n\nWithin each image, each Megadetector detection is cropped, sent through the classifiers, and then weighted by Megadetector’s detection confidence. We tuned how much to weight this using a simple ‘diminishing factor’ _d_: (detector\\_confidence + d) / (d + 1) * crop_prediction.\n\n**(Here’s the interesting part that didn’t really go anywhere)**\nWe then multiply each image’s prediction vector by a set of ‘geographical priors’ we constructed in a similar fashion to host Elijah’s paper [Presence-Only Geographical Priors for Fine-Grained Image Classification](https://arxiv.org/abs/1906.05272). Since we didn’t have latitude and longitude information, we used the datetime (encoded to sin/cos as in the paper) and the mean normalized difference vegetation index (basically infrared minus red) of each satellite image to construct the geographical priors model. We used a ‘diminishing factor’ for these priors too, but only ever saw about a ~1% increase in our validation score, and a 0.2% increase in our test score. But we thought this was interesting and are excited to see what others tried with the satellite imagery.\n\nWe tried a couple other methods in addition to mean NDVI that didn't work as well. First, we tried just using resnet-extracted image features for each satellite image as features for the model (replacing latitude and longitude in the paper’s model) - but, since the data is ‘presence-only’ the model relies upon generated negative examples, which we couldn’t really generate reasonably for the image features. We tried just selecting random satellite images from the dataset to use as ‘negatives’, but this is such a limited number of data points it didn’t seem to work. And just generating a random length-2048 feature vector as a negative seemed silly, since it seems unlikely that a satellite image would ever generate those features.\n\nThen we tried the mean NDVI thing. Cool. As a sidenote, we eliminated all the cloud, shadow, and 0-value pixels from the satellite images using the QA data, so the NDVI we calculate is only for visible land.\n\nThen we tried making that more complicated. Instead of using the mean NDVI (one value per image), we modeled a normal distribution of pixel-wise NDVI values over each satellite image, created a cumulative distribution function for those values for each image, and sampled 100 points from the CDF for features. That allowed us to generate more complex negative examples by generating random normal distributions with a mean between -1 and 1 (the values that the NDVI can take).  This seemed like it was helping, as the validation loss of the geo priors model went down…but when we actually used it for priors for classification our score went down.\n\nAnyway, this didn’t really help much, but I thought it was cool. Would love to chat more if anyone has any thoughts on this (or maybe the hosts can tell me why my ideas are bad :P )\n\n**(Back to the boring part)**\nWe then perform a moving average of the image-wise prediction vectors over each sequence. So, each image’s prediction becomes the average of the images nearest to it. We tuned this window size.\n\nWe then tried a majority vote ‘with dissenters’ within sequences (this may have a real name). It’s just a majority vote for the final class of each sequence, but if an image has a differing classification with a confidence higher than our ‘dissenter threshold’, we let it keep its prediction. This was meant to help sequences that were mostly empty save for a few frames. We tuned this threshold as well as a ‘handicap’ for non-empty images, such that a majority vote can only be for ‘empty’ if it has at least ‘handicap’ more empty images than the 2nd place class in the vote.\n\n**What we wanted to try**\nSara’s paper! [Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection](https://arxiv.org/abs/1912.03538) \nWe actually implemented a modification of this which just uses the memory banks and attention networks during classification (still using Megadetector for detections). Unfortunately we entered this thing pretty late and didn’t have time to train it….but we’ll give it a go as a ‘Late Submission’ :)\n\nWell, that was a spiel. If you read through hopefully it’s either interesting, or you’re reading this a year from now as a starting point for iWildcam 2021. Heh\n\nSo long, and thanks for all the ocellated turkey….",
      "votes": 9
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "863874": "Hi everyone! Thanks to the organizers for putting together such a fun competition and really interesting set of data. Here is our submission (rank #8 or #10 depending on whether you count the [Deleted] entries) and some other things we tried / wanted to try.\n\nOur submission was a pretty simple ensemble of a resnet 152 and a resnet 101.\n\nWe used all of the provided Megadetector annotations to create a dataset of cropped images with species labels. We then split the data into a training and validation set (just 2 fold), making sure the validation set came from distinct locations and keeping the class distributions approximately the same. We used a ‘mini’ version of each of these (~5-10% of the images) as a mini-training and mini-validation set, which we did our hyperparameter tuning on. \n\nThe resnets were pretrained on ImageNet and finetuned on the camera trap crops with pretty heavy data augmentation (random flip, random rotate, random zoom, random contrast and lighting modification, random warp, and mixup). Both models were further fine-tuned for a couple epochs with a class-balanced subset of crops which contained no more than 100 examples per class.\n\nWithin each image, each Megadetector detection is cropped, sent through the classifiers, and then weighted by Megadetector’s detection confidence. We tuned how much to weight this using a simple ‘diminishing factor’ _d_: (detector\\_confidence + d) / (d + 1) * crop_prediction.\n\n**(Here’s the interesting part that didn’t really go anywhere)**\nWe then multiply each image’s prediction vector by a set of ‘geographical priors’ we constructed in a similar fashion to host Elijah’s paper [Presence-Only Geographical Priors for Fine-Grained Image Classification](https://arxiv.org/abs/1906.05272). Since we didn’t have latitude and longitude information, we used the datetime (encoded to sin/cos as in the paper) and the mean normalized difference vegetation index (basically infrared minus red) of each satellite image to construct the geographical priors model. We used a ‘diminishing factor’ for these priors too, but only ever saw about a ~1% increase in our validation score, and a 0.2% increase in our test score. But we thought this was interesting and are excited to see what others tried with the satellite imagery.\n\nWe tried a couple other methods in addition to mean NDVI that didn't work as well. First, we tried just using resnet-extracted image features for each satellite image as features for the model (replacing latitude and longitude in the paper’s model) - but, since the data is ‘presence-only’ the model relies upon generated negative examples, which we couldn’t really generate reasonably for the image features. We tried just selecting random satellite images from the dataset to use as ‘negatives’, but this is such a limited number of data points it didn’t seem to work. And just generating a random length-2048 feature vector as a negative seemed silly, since it seems unlikely that a satellite image would ever generate those features.\n\nThen we tried the mean NDVI thing. Cool. As a sidenote, we eliminated all the cloud, shadow, and 0-value pixels from the satellite images using the QA data, so the NDVI we calculate is only for visible land.\n\nThen we tried making that more complicated. Instead of using the mean NDVI (one value per image), we modeled a normal distribution of pixel-wise NDVI values over each satellite image, created a cumulative distribution function for those values for each image, and sampled 100 points from the CDF for features. That allowed us to generate more complex negative examples by generating random normal distributions with a mean between -1 and 1 (the values that the NDVI can take).  This seemed like it was helping, as the validation loss of the geo priors model went down…but when we actually used it for priors for classification our score went down.\n\nAnyway, this didn’t really help much, but I thought it was cool. Would love to chat more if anyone has any thoughts on this (or maybe the hosts can tell me why my ideas are bad :P )\n\n**(Back to the boring part)**\nWe then perform a moving average of the image-wise prediction vectors over each sequence. So, each image’s prediction becomes the average of the images nearest to it. We tuned this window size.\n\nWe then tried a majority vote ‘with dissenters’ within sequences (this may have a real name). It’s just a majority vote for the final class of each sequence, but if an image has a differing classification with a confidence higher than our ‘dissenter threshold’, we let it keep its prediction. This was meant to help sequences that were mostly empty save for a few frames. We tuned this threshold as well as a ‘handicap’ for non-empty images, such that a majority vote can only be for ‘empty’ if it has at least ‘handicap’ more empty images than the 2nd place class in the vote.\n\n**What we wanted to try**\nSara’s paper! [Context R-CNN: Long Term Temporal Context for Per-Camera Object Detection](https://arxiv.org/abs/1912.03538) \nWe actually implemented a modification of this which just uses the memory banks and attention networks during classification (still using Megadetector for detections). Unfortunately we entered this thing pretty late and didn’t have time to train it….but we’ll give it a go as a ‘Late Submission’ :)\n\nWell, that was a spiel. If you read through hopefully it’s either interesting, or you’re reading this a year from now as a starting point for iWildcam 2021. Heh\n\nSo long, and thanks for all the ocellated turkey…."
  }
}