{
  "id": 328526,
  "title": "9th place solution",
  "url": "/competitions/iwildcam2022-fgvc9/discussion/328526",
  "author_name": "Charlie Turner",
  "post_date": "2022-06-01T16:58:21.552000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Congratulations to the winners and thank you to the competition hosts, current and past, for continuing to sponsor this important, enjoyable competition.  I confess that it still challenges me even in its reduced animal-only, counting form.  Goats! How can anyone count goats, wandering here and there, staring into the camera?  I would love to hear how the top teams - anybody for that matter - approached this counting problem.</p>\n<p><strong>Outline of my approach</strong></p>\n<ul>\n<li>Use <em>MegaDetector</em> v4 detections to train a 2nd animal detector based on <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\"><em>Yolov5</em></a>.</li>\n<li>Merge MegaDetector and Yolov5 detections using <em>weighted boxes fusion</em> <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">WBF</a>.</li>\n<li>Apply a custom frame-to-frame object tracking algorithm focused on individuals within herds/packs. </li>\n</ul>\n<p><strong>The herd/pack species that I focused on:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Category Id</th>\n<th>Common Name</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2</td>\n<td>white-lipped peccary</td>\n</tr>\n<tr>\n<td>8</td>\n<td>collared peccary</td>\n</tr>\n<tr>\n<td>70</td>\n<td>wild goat</td>\n</tr>\n<tr>\n<td>71</td>\n<td>domesticated cattle</td>\n</tr>\n<tr>\n<td>72</td>\n<td>domestic sheep</td>\n</tr>\n<tr>\n<td>90</td>\n<td>african bush elephant</td>\n</tr>\n<tr>\n<td>96</td>\n<td>impala</td>\n</tr>\n<tr>\n<td>256</td>\n<td>dromedary camel</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Observations</strong></p>\n<ul>\n<li><p>I first learned about WBF in the <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef\" target=\"_blank\">Kaggle COTS starfish competition</a> where it was used by many teams to post process detections. I used it successfully to merge the sponsor-provided MedaDetector detections with my Yolov5 detections and it gave me a public/private LB score of 0.275/0.265 (above Benchmark: iWildCam 2021 winner). I spent some time tuning the algorithm's hyperparameters but I'm not sure it provided any significant benefit beyond my initial success.</p></li>\n<li><p>I found plenty of anecdotal evidence that an inter-frame counting approach could improve on the standard <strong>max</strong> approach of the benchmarks, but I was never able to get my solution to provide any benefit once I scaled up to the whole dataset.  Since the approach works well in video, I hoped it would work with some of our sequences. The first step in my process was to try to estimate the overall direction a herd was moving so that I could eliminate many candidates from the inter-frame matching process. Again, this was easy to do in many cases, but unreliable overall, leading to weaker matching in most cases. I'm still looking into the details of exactly what happened.</p></li>\n<li><p>Additional detail for the inter-frame matching approach:  Herd sequences were identified by a 9-class Yolov5 detector. For any sequence with a preponderance of herd detections, detection-level matching was performed using the following method.  An autoencoder was trained on small chips from the center of mass of the DeepMac mask for each detection. The autoencoder produced a 512-element latent feature vector for each detection. These latent features (along with X,Y, area, delta-T) where used in detection-to-detection matching within the sequence.  Using the DeepMac segmentation masks to select which chips are fed to the autoencoder showed improvement over using the entire detection (scaled) or the center of the detection. Presumably this is due to the elimination of non-individual pixels due to occlusion and/or background between the individual's legs.</p></li>\n<li><p>My largest counting errors in the validation set usually involved domestic cattle. Consequently, I spent a lot of time looking at sequences of domestic cattle and getting a little discouraged about not being able to spend more time with the more <em>exotic animals</em>. Then I heard a story on public radio about how damaging cows can be in many different situations.  This got me thinking about camera traps for habitat destruction monitoring and suddenly counting cattle seemed much more significant. </p></li>\n</ul>",
  "messages": [
    {
      "id": 1808262,
      "postDate": "2022-06-01T16:58:21.553Z",
      "content": "<p>Congratulations to the winners and thank you to the competition hosts, current and past, for continuing to sponsor this important, enjoyable competition.  I confess that it still challenges me even in its reduced animal-only, counting form.  Goats! How can anyone count goats, wandering here and there, staring into the camera?  I would love to hear how the top teams - anybody for that matter - approached this counting problem.</p>\n<p><strong>Outline of my approach</strong></p>\n<ul>\n<li>Use <em>MegaDetector</em> v4 detections to train a 2nd animal detector based on <a href=\"https://github.com/ultralytics/yolov5\" target=\"_blank\"><em>Yolov5</em></a>.</li>\n<li>Merge MegaDetector and Yolov5 detections using <em>weighted boxes fusion</em> <a href=\"https://github.com/ZFTurbo/Weighted-Boxes-Fusion\" target=\"_blank\">WBF</a>.</li>\n<li>Apply a custom frame-to-frame object tracking algorithm focused on individuals within herds/packs. </li>\n</ul>\n<p><strong>The herd/pack species that I focused on:</strong></p>\n<table>\n<thead>\n<tr>\n<th>Category Id</th>\n<th>Common Name</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>2</td>\n<td>white-lipped peccary</td>\n</tr>\n<tr>\n<td>8</td>\n<td>collared peccary</td>\n</tr>\n<tr>\n<td>70</td>\n<td>wild goat</td>\n</tr>\n<tr>\n<td>71</td>\n<td>domesticated cattle</td>\n</tr>\n<tr>\n<td>72</td>\n<td>domestic sheep</td>\n</tr>\n<tr>\n<td>90</td>\n<td>african bush elephant</td>\n</tr>\n<tr>\n<td>96</td>\n<td>impala</td>\n</tr>\n<tr>\n<td>256</td>\n<td>dromedary camel</td>\n</tr>\n</tbody>\n</table>\n<p><strong>Observations</strong></p>\n<ul>\n<li><p>I first learned about WBF in the <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef\" target=\"_blank\">Kaggle COTS starfish competition</a> where it was used by many teams to post process detections. I used it successfully to merge the sponsor-provided MedaDetector detections with my Yolov5 detections and it gave me a public/private LB score of 0.275/0.265 (above Benchmark: iWildCam 2021 winner). I spent some time tuning the algorithm's hyperparameters but I'm not sure it provided any significant benefit beyond my initial success.</p></li>\n<li><p>I found plenty of anecdotal evidence that an inter-frame counting approach could improve on the standard <strong>max</strong> approach of the benchmarks, but I was never able to get my solution to provide any benefit once I scaled up to the whole dataset.  Since the approach works well in video, I hoped it would work with some of our sequences. The first step in my process was to try to estimate the overall direction a herd was moving so that I could eliminate many candidates from the inter-frame matching process. Again, this was easy to do in many cases, but unreliable overall, leading to weaker matching in most cases. I'm still looking into the details of exactly what happened.</p></li>\n<li><p>Additional detail for the inter-frame matching approach:  Herd sequences were identified by a 9-class Yolov5 detector. For any sequence with a preponderance of herd detections, detection-level matching was performed using the following method.  An autoencoder was trained on small chips from the center of mass of the DeepMac mask for each detection. The autoencoder produced a 512-element latent feature vector for each detection. These latent features (along with X,Y, area, delta-T) where used in detection-to-detection matching within the sequence.  Using the DeepMac segmentation masks to select which chips are fed to the autoencoder showed improvement over using the entire detection (scaled) or the center of the detection. Presumably this is due to the elimination of non-individual pixels due to occlusion and/or background between the individual's legs.</p></li>\n<li><p>My largest counting errors in the validation set usually involved domestic cattle. Consequently, I spent a lot of time looking at sequences of domestic cattle and getting a little discouraged about not being able to spend more time with the more <em>exotic animals</em>. Then I heard a story on public radio about how damaging cows can be in many different situations.  This got me thinking about camera traps for habitat destruction monitoring and suddenly counting cattle seemed much more significant. </p></li>\n</ul>",
      "rawMarkdown": "Congratulations to the winners and thank you to the competition hosts, current and past, for continuing to sponsor this important, enjoyable competition.  I confess that it still challenges me even in its reduced animal-only, counting form.  Goats! How can anyone count goats, wandering here and there, staring into the camera?  I would love to hear how the top teams - anybody for that matter - approached this counting problem.\n\n**Outline of my approach**\n* Use *MegaDetector* v4 detections to train a 2nd animal detector based on [*Yolov5*](https://github.com/ultralytics/yolov5).\n* Merge MegaDetector and Yolov5 detections using *weighted boxes fusion* [WBF](https://github.com/ZFTurbo/Weighted-Boxes-Fusion).\n* Apply a custom frame-to-frame object tracking algorithm focused on individuals within herds/packs. \n\n**The herd/pack species that I focused on:**\n| Category Id |Common Name |\n| --- | --- |\n|2|white-lipped peccary|\n|8|collared peccary|\n|70|wild goat|\n|71|domesticated cattle|\n|72|domestic sheep|\n|90|african bush elephant|\n|96|impala|\n|256|dromedary camel|\n\n**Observations**\n\n- I first learned about WBF in the [Kaggle COTS starfish competition](https://www.kaggle.com/c/tensorflow-great-barrier-reef) where it was used by many teams to post process detections. I used it successfully to merge the sponsor-provided MedaDetector detections with my Yolov5 detections and it gave me a public/private LB score of 0.275/0.265 (above Benchmark: iWildCam 2021 winner). I spent some time tuning the algorithm's hyperparameters but I'm not sure it provided any significant benefit beyond my initial success.\n\n- I found plenty of anecdotal evidence that an inter-frame counting approach could improve on the standard **max** approach of the benchmarks, but I was never able to get my solution to provide any benefit once I scaled up to the whole dataset.  Since the approach works well in video, I hoped it would work with some of our sequences. The first step in my process was to try to estimate the overall direction a herd was moving so that I could eliminate many candidates from the inter-frame matching process. Again, this was easy to do in many cases, but unreliable overall, leading to weaker matching in most cases. I'm still looking into the details of exactly what happened.\n\n- Additional detail for the inter-frame matching approach:  Herd sequences were identified by a 9-class Yolov5 detector. For any sequence with a preponderance of herd detections, detection-level matching was performed using the following method.  An autoencoder was trained on small chips from the center of mass of the DeepMac mask for each detection. The autoencoder produced a 512-element latent feature vector for each detection. These latent features (along with X,Y, area, delta-T) where used in detection-to-detection matching within the sequence.  Using the DeepMac segmentation masks to select which chips are fed to the autoencoder showed improvement over using the entire detection (scaled) or the center of the detection. Presumably this is due to the elimination of non-individual pixels due to occlusion and/or background between the individual's legs.\n\n- My largest counting errors in the validation set usually involved domestic cattle. Consequently, I spent a lot of time looking at sequences of domestic cattle and getting a little discouraged about not being able to spend more time with the more *exotic animals*. Then I heard a story on public radio about how damaging cows can be in many different situations.  This got me thinking about camera traps for habitat destruction monitoring and suddenly counting cattle seemed much more significant. \n\n",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1808262": "Congratulations to the winners and thank you to the competition hosts, current and past, for continuing to sponsor this important, enjoyable competition.  I confess that it still challenges me even in its reduced animal-only, counting form.  Goats! How can anyone count goats, wandering here and there, staring into the camera?  I would love to hear how the top teams - anybody for that matter - approached this counting problem.\n\n**Outline of my approach**\n* Use *MegaDetector* v4 detections to train a 2nd animal detector based on [*Yolov5*](https://github.com/ultralytics/yolov5).\n* Merge MegaDetector and Yolov5 detections using *weighted boxes fusion* [WBF](https://github.com/ZFTurbo/Weighted-Boxes-Fusion).\n* Apply a custom frame-to-frame object tracking algorithm focused on individuals within herds/packs. \n\n**The herd/pack species that I focused on:**\n| Category Id |Common Name |\n| --- | --- |\n|2|white-lipped peccary|\n|8|collared peccary|\n|70|wild goat|\n|71|domesticated cattle|\n|72|domestic sheep|\n|90|african bush elephant|\n|96|impala|\n|256|dromedary camel|\n\n**Observations**\n\n- I first learned about WBF in the [Kaggle COTS starfish competition](https://www.kaggle.com/c/tensorflow-great-barrier-reef) where it was used by many teams to post process detections. I used it successfully to merge the sponsor-provided MedaDetector detections with my Yolov5 detections and it gave me a public/private LB score of 0.275/0.265 (above Benchmark: iWildCam 2021 winner). I spent some time tuning the algorithm's hyperparameters but I'm not sure it provided any significant benefit beyond my initial success.\n\n- I found plenty of anecdotal evidence that an inter-frame counting approach could improve on the standard **max** approach of the benchmarks, but I was never able to get my solution to provide any benefit once I scaled up to the whole dataset.  Since the approach works well in video, I hoped it would work with some of our sequences. The first step in my process was to try to estimate the overall direction a herd was moving so that I could eliminate many candidates from the inter-frame matching process. Again, this was easy to do in many cases, but unreliable overall, leading to weaker matching in most cases. I'm still looking into the details of exactly what happened.\n\n- Additional detail for the inter-frame matching approach:  Herd sequences were identified by a 9-class Yolov5 detector. For any sequence with a preponderance of herd detections, detection-level matching was performed using the following method.  An autoencoder was trained on small chips from the center of mass of the DeepMac mask for each detection. The autoencoder produced a 512-element latent feature vector for each detection. These latent features (along with X,Y, area, delta-T) where used in detection-to-detection matching within the sequence.  Using the DeepMac segmentation masks to select which chips are fed to the autoencoder showed improvement over using the entire detection (scaled) or the center of the detection. Presumably this is due to the elimination of non-individual pixels due to occlusion and/or background between the individual's legs.\n\n- My largest counting errors in the validation set usually involved domestic cattle. Consequently, I spent a lot of time looking at sequences of domestic cattle and getting a little discouraged about not being able to spend more time with the more *exotic animals*. Then I heard a story on public radio about how damaging cows can be in many different situations.  This got me thinking about camera traps for habitat destruction monitoring and suddenly counting cattle seemed much more significant. \n\n"
  }
}