{
  "id": 157946,
  "title": "LB #2 Documentation",
  "url": "/competitions/iwildcam-2020-fgvc7/writeups/jon-daly-lb-2-documentation",
  "author_name": "",
  "post_date": "2020-06-12T16:29:16.074755200Z",
  "votes": 14,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Apologies for the late submission documentation, and thanks to the organisers of the competition - this was great fun.</p>\n\n<p>After attempting many different approaches my highest scoring solution was generally as follows:</p>\n\n<ol>\n<li>Removed noisy and/or 'human' class categories - 'start', 'end', 'unknown', and 'unidentifiable'. Retaining the empty class seemed important for reducing activation on background to generalise better.</li>\n<li>Encorporated iNat 2017/2018 data for shared classes.</li>\n<li>Ran MegaDetector v4 on all images.</li>\n<li>Took only animal detection crops with confidence &gt;0.3 and discarded detections which were too small. Also discarded crops within the same image if they were much less confident than the highest.</li>\n<li>Made sure to expand height of crop to be 224 pixels, even if the bounding box was smaller, then used reflection padding to square the crop and resized to 224x224 resolution.</li>\n<li>Applied various data augmentations including CLAHE, hue shift, gaussian noise, cutout, scale/rotate/shift, brightness/contrast adjustments and grayscale.</li>\n<li>Created a custom classifier head which incorporated metadata from the image. This used cos/sin representations of of the time of year and time of day, and information regarding the crop (width, height, pitch, yaw).</li>\n<li>Trained for 7 epochs with a single EfficientNetB4 model pretrained with imagenet noisystudent weights, using label-smoothed cross-entropy focal loss. This was 2 hours of training.</li>\n<li>Used TTA with every training augmentation except for noise to get predictions.</li>\n<li>Combined the predictions for each crop within an image to get a total average for the image.</li>\n<li>Averaged these predictions across images which were clustered by time and location for the final submission. This was only done as the sequence ID labels appeared to be unreliable.</li>\n<li>Achieved 0.930 public, 0.903 private LB score. The highest submission with this approach which I didn't pick was 0.908 private.</li>\n</ol>\n\n<p>This was my first proper experience with computer vision so I unfortunately ran out of time to try a lot of the things I wanted due to my slowness :). I wanted to try:\n- Ensembling different EfficientNets at different input resolutions.\n- Using the location multisat data to create a geo prior with which to augment the prediction. I had initially used this data directly in a single classifier model but it too easily overfit.\n- Creating a more realistic infrared data augmentation rather than relying on grayscale.\n- Using a siamese network to distinguish between the most confused species.\n- Reintroducing the 'vehicle'/'human' megadetector crops with a prior or separate classifier. </p>",
  "messages": [
    {
      "id": "883443",
      "postDate": "06/12/2020 16:29:16",
      "content": "<p>Apologies for the late submission documentation, and thanks to the organisers of the competition - this was great fun.</p>\n\n<p>After attempting many different approaches my highest scoring solution was generally as follows:</p>\n\n<ol>\n<li>Removed noisy and/or 'human' class categories - 'start', 'end', 'unknown', and 'unidentifiable'. Retaining the empty class seemed important for reducing activation on background to generalise better.</li>\n<li>Encorporated iNat 2017/2018 data for shared classes.</li>\n<li>Ran MegaDetector v4 on all images.</li>\n<li>Took only animal detection crops with confidence &gt;0.3 and discarded detections which were too small. Also discarded crops within the same image if they were much less confident than the highest.</li>\n<li>Made sure to expand height of crop to be 224 pixels, even if the bounding box was smaller, then used reflection padding to square the crop and resized to 224x224 resolution.</li>\n<li>Applied various data augmentations including CLAHE, hue shift, gaussian noise, cutout, scale/rotate/shift, brightness/contrast adjustments and grayscale.</li>\n<li>Created a custom classifier head which incorporated metadata from the image. This used cos/sin representations of of the time of year and time of day, and information regarding the crop (width, height, pitch, yaw).</li>\n<li>Trained for 7 epochs with a single EfficientNetB4 model pretrained with imagenet noisystudent weights, using label-smoothed cross-entropy focal loss. This was 2 hours of training.</li>\n<li>Used TTA with every training augmentation except for noise to get predictions.</li>\n<li>Combined the predictions for each crop within an image to get a total average for the image.</li>\n<li>Averaged these predictions across images which were clustered by time and location for the final submission. This was only done as the sequence ID labels appeared to be unreliable.</li>\n<li>Achieved 0.930 public, 0.903 private LB score. The highest submission with this approach which I didn't pick was 0.908 private.</li>\n</ol>\n\n<p>This was my first proper experience with computer vision so I unfortunately ran out of time to try a lot of the things I wanted due to my slowness :). I wanted to try:\n- Ensembling different EfficientNets at different input resolutions.\n- Using the location multisat data to create a geo prior with which to augment the prediction. I had initially used this data directly in a single classifier model but it too easily overfit.\n- Creating a more realistic infrared data augmentation rather than relying on grayscale.\n- Using a siamese network to distinguish between the most confused species.\n- Reintroducing the 'vehicle'/'human' megadetector crops with a prior or separate classifier. </p>",
      "rawMarkdown": "Apologies for the late submission documentation, and thanks to the organisers of the competition - this was great fun.\n\nAfter attempting many different approaches my highest scoring solution was generally as follows:\n\n1. Removed noisy and/or 'human' class categories - 'start', 'end', 'unknown', and 'unidentifiable'. Retaining the empty class seemed important for reducing activation on background to generalise better.\n2. Encorporated iNat 2017/2018 data for shared classes.\n3. Ran MegaDetector v4 on all images.\n4. Took only animal detection crops with confidence &gt;0.3 and discarded detections which were too small. Also discarded crops within the same image if they were much less confident than the highest.\n5. Made sure to expand height of crop to be 224 pixels, even if the bounding box was smaller, then used reflection padding to square the crop and resized to 224x224 resolution.\n6. Applied various data augmentations including CLAHE, hue shift, gaussian noise, cutout, scale/rotate/shift, brightness/contrast adjustments and grayscale.\n7. Created a custom classifier head which incorporated metadata from the image. This used cos/sin representations of of the time of year and time of day, and information regarding the crop (width, height, pitch, yaw).\n8. Trained for 7 epochs with a single EfficientNetB4 model pretrained with imagenet noisystudent weights, using label-smoothed cross-entropy focal loss. This was 2 hours of training.\n9. Used TTA with every training augmentation except for noise to get predictions.\n10. Combined the predictions for each crop within an image to get a total average for the image.\n11. Averaged these predictions across images which were clustered by time and location for the final submission. This was only done as the sequence ID labels appeared to be unreliable.\n12. Achieved 0.930 public, 0.903 private LB score. The highest submission with this approach which I didn't pick was 0.908 private.\n\nThis was my first proper experience with computer vision so I unfortunately ran out of time to try a lot of the things I wanted due to my slowness :). I wanted to try:\n- Ensembling different EfficientNets at different input resolutions.\n- Using the location multisat data to create a geo prior with which to augment the prediction. I had initially used this data directly in a single classifier model but it too easily overfit.\n- Creating a more realistic infrared data augmentation rather than relying on grayscale.\n- Using a siamese network to distinguish between the most confused species.\n- Reintroducing the 'vehicle'/'human' megadetector crops with a prior or separate classifier.",
      "votes": null
    },
    {
      "id": "883446",
      "postDate": "06/12/2020 16:29:42",
      "content": "<p><a href=\"/stevenyin\">@stevenyin</a> here you go :)</p>",
      "rawMarkdown": "stevenyin here you go :)",
      "votes": null
    },
    {
      "id": "884672",
      "postDate": "06/13/2020 14:38:39",
      "content": "<p>Hi Jon, great thanks for your post! Let me read it carefully and go back to you later.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Jon, great thanks for your post! Let me read it carefully and go back to you later.\n\n  Thanks!",
      "votes": null
    },
    {
      "id": "887507",
      "postDate": "06/15/2020 17:48:14",
      "content": "<p>Thanks for sharing, couple questions - \nWhat is the pitch and yaw of the crop? \nHow much did the iNat classes improve your score?</p>",
      "rawMarkdown": "Thanks for sharing, couple questions - \nWhat is the pitch and yaw of the crop? \nHow much did the iNat classes improve your score?",
      "votes": null
    },
    {
      "id": "887629",
      "postDate": "06/15/2020 19:21:26",
      "content": "<p>Hi Justin. Since I'm using a crop from a larger image rather than the original the intuition was that knowing whether a crop is high above the camera's horizon or on the ground, or whether it was cropped at the edge of the image may help in making accurate predictions. The pitch is 0. -&gt; 1. from the bottom of the frame to the top, and the yaw is normalized to be 0. at the center and 1. at the edge. It seemed to give a small bump in validation accuracy so I kept it.</p>\n\n<p>I didn't record/don't remember how big the bump in accuracy from the iNat classes was, sorry. I don't believe it was too significant.</p>",
      "rawMarkdown": "Hi Justin. Since I'm using a crop from a larger image rather than the original the intuition was that knowing whether a crop is high above the camera's horizon or on the ground, or whether it was cropped at the edge of the image may help in making accurate predictions. The pitch is 0. -&gt; 1. from the bottom of the frame to the top, and the yaw is normalized to be 0. at the center and 1. at the edge. It seemed to give a small bump in validation accuracy so I kept it.\n\nI didn't record/don't remember how big the bump in accuracy from the iNat classes was, sorry. I don't believe it was too significant.",
      "votes": null
    },
    {
      "id": "887650",
      "postDate": "06/15/2020 19:36:14",
      "content": "<p>Hey Jon!  Can you make sure to fill out this google form with the info above? <a href=\"https://docs.google.com/forms/d/e/1FAIpQLSeZbowaBJYvRstyO3lgToomOkHuTeooW6qvDzH0LNzZPDnrbg/viewform?usp=sf_link\">https://docs.google.com/forms/d/e/1FAIpQLSeZbowaBJYvRstyO3lgToomOkHuTeooW6qvDzH0LNzZPDnrbg/viewform?usp=sf_link</a></p>",
      "rawMarkdown": "Hey Jon!  Can you make sure to fill out this google form with the info above? https://docs.google.com/forms/d/e/1FAIpQLSeZbowaBJYvRstyO3lgToomOkHuTeooW6qvDzH0LNzZPDnrbg/viewform?usp=sf_link",
      "votes": null
    },
    {
      "id": "887715",
      "postDate": "06/15/2020 20:17:46",
      "content": "<p>Done! :) </p>",
      "rawMarkdown": "Done! :)",
      "votes": null
    },
    {
      "id": "887863",
      "postDate": "06/16/2020 00:22:43",
      "content": "<p>Ah. Cool idea!</p>",
      "rawMarkdown": "Ah. Cool idea!",
      "votes": null
    },
    {
      "id": "930942",
      "postDate": "07/15/2020 20:31:56",
      "content": "<p>Hey Jon, congrats! Can you explain how you designed validation set? Many thanks</p>",
      "rawMarkdown": "Hey Jon, congrats! Can you explain how you designed validation set? Many thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 883446,
      "author_name": "jondaly",
      "author_url": "",
      "post_date": "06/12/2020 16:29:42",
      "content": "<p><a href=\"/stevenyin\">@stevenyin</a> here you go :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 884672,
          "author_name": "",
          "author_url": "",
          "post_date": "06/13/2020 14:38:39",
          "content": "<p>Hi Jon, great thanks for your post! Let me read it carefully and go back to you later.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 887507,
      "author_name": "justinkay92",
      "author_url": "",
      "post_date": "06/15/2020 17:48:14",
      "content": "<p>Thanks for sharing, couple questions - \nWhat is the pitch and yaw of the crop? \nHow much did the iNat classes improve your score?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 887629,
      "author_name": "jondaly",
      "author_url": "",
      "post_date": "06/15/2020 19:21:26",
      "content": "<p>Hi Justin. Since I'm using a crop from a larger image rather than the original the intuition was that knowing whether a crop is high above the camera's horizon or on the ground, or whether it was cropped at the edge of the image may help in making accurate predictions. The pitch is 0. -&gt; 1. from the bottom of the frame to the top, and the yaw is normalized to be 0. at the center and 1. at the edge. It seemed to give a small bump in validation accuracy so I kept it.</p>\n\n<p>I didn't record/don't remember how big the bump in accuracy from the iNat classes was, sorry. I don't believe it was too significant.</p>",
      "votes": null,
      "replies": [
        {
          "id": 887863,
          "author_name": "justinkay92",
          "author_url": "",
          "post_date": "06/16/2020 00:22:43",
          "content": "<p>Ah. Cool idea!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 887650,
      "author_name": "sbeery",
      "author_url": "",
      "post_date": "06/15/2020 19:36:14",
      "content": "<p>Hey Jon!  Can you make sure to fill out this google form with the info above? <a href=\"https://docs.google.com/forms/d/e/1FAIpQLSeZbowaBJYvRstyO3lgToomOkHuTeooW6qvDzH0LNzZPDnrbg/viewform?usp=sf_link\">https://docs.google.com/forms/d/e/1FAIpQLSeZbowaBJYvRstyO3lgToomOkHuTeooW6qvDzH0LNzZPDnrbg/viewform?usp=sf_link</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 887715,
          "author_name": "jondaly",
          "author_url": "",
          "post_date": "06/15/2020 20:17:46",
          "content": "<p>Done! :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 930942,
      "author_name": "lennyom",
      "author_url": "",
      "post_date": "07/15/2020 20:31:56",
      "content": "<p>Hey Jon, congrats! Can you explain how you designed validation set? Many thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "883443": "Apologies for the late submission documentation, and thanks to the organisers of the competition - this was great fun.\n\nAfter attempting many different approaches my highest scoring solution was generally as follows:\n\n1. Removed noisy and/or 'human' class categories - 'start', 'end', 'unknown', and 'unidentifiable'. Retaining the empty class seemed important for reducing activation on background to generalise better.\n2. Encorporated iNat 2017/2018 data for shared classes.\n3. Ran MegaDetector v4 on all images.\n4. Took only animal detection crops with confidence &gt;0.3 and discarded detections which were too small. Also discarded crops within the same image if they were much less confident than the highest.\n5. Made sure to expand height of crop to be 224 pixels, even if the bounding box was smaller, then used reflection padding to square the crop and resized to 224x224 resolution.\n6. Applied various data augmentations including CLAHE, hue shift, gaussian noise, cutout, scale/rotate/shift, brightness/contrast adjustments and grayscale.\n7. Created a custom classifier head which incorporated metadata from the image. This used cos/sin representations of of the time of year and time of day, and information regarding the crop (width, height, pitch, yaw).\n8. Trained for 7 epochs with a single EfficientNetB4 model pretrained with imagenet noisystudent weights, using label-smoothed cross-entropy focal loss. This was 2 hours of training.\n9. Used TTA with every training augmentation except for noise to get predictions.\n10. Combined the predictions for each crop within an image to get a total average for the image.\n11. Averaged these predictions across images which were clustered by time and location for the final submission. This was only done as the sequence ID labels appeared to be unreliable.\n12. Achieved 0.930 public, 0.903 private LB score. The highest submission with this approach which I didn't pick was 0.908 private.\n\nThis was my first proper experience with computer vision so I unfortunately ran out of time to try a lot of the things I wanted due to my slowness :). I wanted to try:\n- Ensembling different EfficientNets at different input resolutions.\n- Using the location multisat data to create a geo prior with which to augment the prediction. I had initially used this data directly in a single classifier model but it too easily overfit.\n- Creating a more realistic infrared data augmentation rather than relying on grayscale.\n- Using a siamese network to distinguish between the most confused species.\n- Reintroducing the 'vehicle'/'human' megadetector crops with a prior or separate classifier.",
    "883446": "stevenyin here you go :)",
    "884672": "Hi Jon, great thanks for your post! Let me read it carefully and go back to you later.\n\n  Thanks!",
    "887507": "Thanks for sharing, couple questions - \nWhat is the pitch and yaw of the crop? \nHow much did the iNat classes improve your score?",
    "887629": "Hi Justin. Since I'm using a crop from a larger image rather than the original the intuition was that knowing whether a crop is high above the camera's horizon or on the ground, or whether it was cropped at the edge of the image may help in making accurate predictions. The pitch is 0. -&gt; 1. from the bottom of the frame to the top, and the yaw is normalized to be 0. at the center and 1. at the edge. It seemed to give a small bump in validation accuracy so I kept it.\n\nI didn't record/don't remember how big the bump in accuracy from the iNat classes was, sorry. I don't believe it was too significant.",
    "887650": "Hey Jon!  Can you make sure to fill out this google form with the info above? https://docs.google.com/forms/d/e/1FAIpQLSeZbowaBJYvRstyO3lgToomOkHuTeooW6qvDzH0LNzZPDnrbg/viewform?usp=sf_link",
    "887715": "Done! :)",
    "887863": "Ah. Cool idea!",
    "930942": "Hey Jon, congrats! Can you explain how you designed validation set? Many thanks"
  },
  "source": "meta"
}