{
  "id": 310101,
  "title": "Let's crowd-source a bounding box dataset !",
  "url": "/competitions/happy-whale-and-dolphin/discussion/310101",
  "author_name": "",
  "post_date": "2022-02-27T15:05:07.104980300Z",
  "votes": 44,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Hi kagglers 👋,</p>\n<p>Now that we know how important it is to properly crop the images in this competition, I've been wondering about a way to make a good quality dataset of bounding boxes. <br>\nNumerous manual annotations are already available in public datasets, but it seems like it's still not enough…</p>\n<p>So I've been working on a way to crowd-source a bounding box dataset. I made and hosted a little app where anybody can come and annotate images or review other people's annotations.</p>\n<h3>Links</h3>\n<p><strong>The app is available here</strong>: <a href=\"http://happywhale.theoboyer.fr\" target=\"_blank\">http://happywhale.theoboyer.fr</a></p>\n<p><strong>The source code and explanations on how the app works are available on this GitHub repo</strong>: <a href=\"https://github.com/TheoBoyer/Manual-bbox-annotation-tool\" target=\"_blank\">https://github.com/TheoBoyer/Manual-bbox-annotation-tool</a></p>\n<h3>Download the dataset</h3>\n<p>On the home page you can find buttons to download directly either the raw annotations and review informations or the final dataset constituted of only manually reviewed samples. You can also download the dataset in a kaggle notebook with these command lines:</p>\n<p><code>!curl http://happywhale.theoboyer.fr/data/annotations.csv &gt; annotations.csv</code></p>\n<p><code>!curl http://happywhale.theoboyer.fr/data/dataset.csv &gt; dataset.csv</code></p>\n<h3>Note</h3>\n<ul>\n<li>The dataset doesn't start from scratch, I merged several public datasets to begin ! The details of sources and how I gathered manual annotations are available in the notebook: [🐳&amp;🐬 - 👨‍🔬 Merging public bounding box datasets](<a href=\"https://www.kaggle.com/wolfy73/merging-public-bounding-box-\" target=\"_blank\">https://www.kaggle.com/wolfy73/merging-public-bounding-box-</a></li>\n<li>Some automatic annotations are already entered and just need a manual review. The annotations were obtained by taking the dataset from <a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5\" target=\"_blank\">Happywhale: Cropped Dataset [YOLOv5] ✂️</a> and filtered using the methods described in <a href=\"https://www.kaggle.com/wolfy73/filter-yolov5-failure-cases\" target=\"_blank\">🐳&amp;🐬 - Filter YOLOv5 failure cases</a></li>\n<li>Only the training images are in the app to comply with the competition's rules </li>\n</ul>\n<p>Do not hesitate to come and annotate some images, or review some annotations (while your models are training for example 😏), it will help everyone here !</p>\n<p>For more information, see the README of the repo or ask your questions on this topic !</p>\n<p>Happy Kaggling everyone 👍</p>",
  "messages": [
    {
      "id": "1706548",
      "postDate": "02/27/2022 15:05:07",
      "content": "<p>Hi kagglers 👋,</p>\n<p>Now that we know how important it is to properly crop the images in this competition, I've been wondering about a way to make a good quality dataset of bounding boxes. <br>\nNumerous manual annotations are already available in public datasets, but it seems like it's still not enough…</p>\n<p>So I've been working on a way to crowd-source a bounding box dataset. I made and hosted a little app where anybody can come and annotate images or review other people's annotations.</p>\n<h3>Links</h3>\n<p><strong>The app is available here</strong>: <a href=\"http://happywhale.theoboyer.fr\" target=\"_blank\">http://happywhale.theoboyer.fr</a></p>\n<p><strong>The source code and explanations on how the app works are available on this GitHub repo</strong>: <a href=\"https://github.com/TheoBoyer/Manual-bbox-annotation-tool\" target=\"_blank\">https://github.com/TheoBoyer/Manual-bbox-annotation-tool</a></p>\n<h3>Download the dataset</h3>\n<p>On the home page you can find buttons to download directly either the raw annotations and review informations or the final dataset constituted of only manually reviewed samples. You can also download the dataset in a kaggle notebook with these command lines:</p>\n<p><code>!curl http://happywhale.theoboyer.fr/data/annotations.csv &gt; annotations.csv</code></p>\n<p><code>!curl http://happywhale.theoboyer.fr/data/dataset.csv &gt; dataset.csv</code></p>\n<h3>Note</h3>\n<ul>\n<li>The dataset doesn't start from scratch, I merged several public datasets to begin ! The details of sources and how I gathered manual annotations are available in the notebook: [🐳&amp;🐬 - 👨‍🔬 Merging public bounding box datasets](<a href=\"https://www.kaggle.com/wolfy73/merging-public-bounding-box-\" target=\"_blank\">https://www.kaggle.com/wolfy73/merging-public-bounding-box-</a></li>\n<li>Some automatic annotations are already entered and just need a manual review. The annotations were obtained by taking the dataset from <a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5\" target=\"_blank\">Happywhale: Cropped Dataset [YOLOv5] ✂️</a> and filtered using the methods described in <a href=\"https://www.kaggle.com/wolfy73/filter-yolov5-failure-cases\" target=\"_blank\">🐳&amp;🐬 - Filter YOLOv5 failure cases</a></li>\n<li>Only the training images are in the app to comply with the competition's rules </li>\n</ul>\n<p>Do not hesitate to come and annotate some images, or review some annotations (while your models are training for example 😏), it will help everyone here !</p>\n<p>For more information, see the README of the repo or ask your questions on this topic !</p>\n<p>Happy Kaggling everyone 👍</p>",
      "rawMarkdown": "Hi kagglers 👋,\n\nNow that we know how important it is to properly crop the images in this competition, I've been wondering about a way to make a good quality dataset of bounding boxes. \nNumerous manual annotations are already available in public datasets, but it seems like it's still not enough...\n\nSo I've been working on a way to crowd-source a bounding box dataset. I made and hosted a little app where anybody can come and annotate images or review other people's annotations.\n\n### Links\n**The app is available here**: [http://happywhale.theoboyer.fr](http://happywhale.theoboyer.fr)\n\n**The source code and explanations on how the app works are available on this GitHub repo**: [https://github.com/TheoBoyer/Manual-bbox-annotation-tool](https://github.com/TheoBoyer/Manual-bbox-annotation-tool)\n\n### Download the dataset\nOn the home page you can find buttons to download directly either the raw annotations and review informations or the final dataset constituted of only manually reviewed samples. You can also download the dataset in a kaggle notebook with these command lines:\n\n`!curl http://happywhale.theoboyer.fr/data/annotations.csv > annotations.csv`\n\n`!curl http://happywhale.theoboyer.fr/data/dataset.csv > dataset.csv`\n\n### Note\n* The dataset doesn't start from scratch, I merged several public datasets to begin ! The details of sources and how I gathered manual annotations are available in the notebook: [🐳&🐬 - 👨‍🔬 Merging public bounding box datasets](https://www.kaggle.com/wolfy73/merging-public-bounding-box-\n* Some automatic annotations are already entered and just need a manual review. The annotations were obtained by taking the dataset from [Happywhale: Cropped Dataset [YOLOv5] ✂️](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5) and filtered using the methods described in [🐳&🐬 - Filter YOLOv5 failure cases](https://www.kaggle.com/wolfy73/filter-yolov5-failure-cases)\n* Only the training images are in the app to comply with the competition's rules \n\nDo not hesitate to come and annotate some images, or review some annotations (while your models are training for example 😏), it will help everyone here !\n\nFor more information, see the README of the repo or ask your questions on this topic !\n\nHappy Kaggling everyone 👍",
      "votes": null
    },
    {
      "id": "1708483",
      "postDate": "03/01/2022 13:39:32",
      "content": "<p>We nearly reached 200 new annotations + 800 reviews in just 2 days, this is amazing !<br>\nA big thank you to all those who participate 👏 </p>",
      "rawMarkdown": "We nearly reached 200 new annotations + 800 reviews in just 2 days, this is amazing !\nA big thank you to all those who participate 👏",
      "votes": null
    },
    {
      "id": "1712241",
      "postDate": "03/04/2022 17:55:24",
      "content": "<p>Amazing work <a href=\"https://www.kaggle.com/wolfy73\" target=\"_blank\">@wolfy73</a>. How many annotations until now ?</p>",
      "rawMarkdown": "Amazing work @wolfy73. How many annotations until now ?",
      "votes": null
    },
    {
      "id": "1712257",
      "postDate": "03/04/2022 18:26:15",
      "content": "<p>I made a public dataset that contains the dataset. It's updated on a daily basis using a scheduled notebook:<br>\n<a href=\"https://www.kaggle.com/wolfy73/wandd-crowed-sourced-bounging-boxes\" target=\"_blank\">🐳&amp;🐬 - Crowdsourced bounding boxes</a></p>",
      "rawMarkdown": "I made a public dataset that contains the dataset. It's updated on a daily basis using a scheduled notebook:\n[🐳&🐬 - Crowdsourced bounding boxes](https://www.kaggle.com/wolfy73/wandd-crowed-sourced-bounging-boxes)",
      "votes": null
    },
    {
      "id": "1712260",
      "postDate": "03/04/2022 18:27:39",
      "content": "<p>The reviewed dataset contains currently 2671 annotated images !</p>",
      "rawMarkdown": "The reviewed dataset contains currently 2671 annotated images !",
      "votes": null
    },
    {
      "id": "1712729",
      "postDate": "03/05/2022 08:48:58",
      "content": "<p>This is a great initiative! <a href=\"https://www.kaggle.com/wolfy73\" target=\"_blank\">@wolfy73</a> </p>\n<p>I am thinking if it is possible to integrate Active Learning into the whole pipeline to ramp up the marginal gain of our annotations.</p>\n<p>My raw idea is to make a notebook that is scheduled to pull our current annotations and train a model, and then rank the unlabelled images by their \"worthiness\" to be annotated. As a step further, the rank can be fed to your app to complete the loop. We keep this loop going until the performance gain diminish at a certain extent.</p>\n<p>One uncertainty is how feasible active learning can be applied on object detection, as it is typically applied on classification problems.</p>\n<p>I am trying to do some research to assess its feasibility and the possible pipeline, let me know if you are interested to also work on this! I believe this is a good chance to apply active learning in real problem and assess its effectiveness!</p>\n<p>FYI, here are some initial finding I have got:</p>\n<ol>\n<li><a href=\"https://github.com/NVlabs/AL-MDN\" target=\"_blank\">Active Learning for Deep Object Detection via Probabilistic Modeling</a></li>\n<li><a href=\"https://github.com/superannotateai/active_learning\" target=\"_blank\">Learning Loss for Active Learning</a></li>\n</ol>",
      "rawMarkdown": "This is a great initiative! @wolfy73 \n\nI am thinking if it is possible to integrate Active Learning into the whole pipeline to ramp up the marginal gain of our annotations.\n\nMy raw idea is to make a notebook that is scheduled to pull our current annotations and train a model, and then rank the unlabelled images by their \"worthiness\" to be annotated. As a step further, the rank can be fed to your app to complete the loop. We keep this loop going until the performance gain diminish at a certain extent.\n\nOne uncertainty is how feasible active learning can be applied on object detection, as it is typically applied on classification problems.\n\nI am trying to do some research to assess its feasibility and the possible pipeline, let me know if you are interested to also work on this! I believe this is a good chance to apply active learning in real problem and assess its effectiveness!\n\nFYI, here are some initial finding I have got:\n1. [Active Learning for Deep Object Detection via Probabilistic Modeling](https://github.com/NVlabs/AL-MDN)\n2. [Learning Loss for Active Learning](https://github.com/superannotateai/active_learning)",
      "votes": null
    },
    {
      "id": "1712735",
      "postDate": "03/05/2022 08:53:07",
      "content": "<p>FYI I think the link is not working</p>",
      "rawMarkdown": "FYI I think the link is not working",
      "votes": null
    },
    {
      "id": "1712745",
      "postDate": "03/05/2022 09:15:12",
      "content": "<p>Thank you for letting me know, it should work now !</p>",
      "rawMarkdown": "Thank you for letting me know, it should work now !",
      "votes": null
    },
    {
      "id": "1712748",
      "postDate": "03/05/2022 09:19:57",
      "content": "<p>I'm interested !</p>",
      "rawMarkdown": "I'm interested !",
      "votes": null
    },
    {
      "id": "1727314",
      "postDate": "03/17/2022 21:22:53",
      "content": "<p><strong>[03/17/2022]</strong><br>\nI support this, if we do 1 annotation per second, it will take an hour per 3.6k images.</p>\n<p>Also it seems that the photos disappear before the bounding box does. Is this a problem? Or is that normal?</p>\n<p>Also if you could drag and drop the bounding box, it would make it faster to annotate. Auto-snapping to the edges/corners AND/OR make the corners redraggable after dropping that would be awesome. I wonder how hard this is to code in.</p>\n<p>Are we bounding boxing the surface above the water? Or above and below? I've been annotating only the above areas of an animal UNLESS the camera is underwater, in which case (usually dolphins), I annotate the whole body.</p>\n<p>Can there be a button to \"suggest for removal\"? ie. A photo of a clipboard, multiple shots of the same animal in the same position at the same time(potential for the AI for overemphasize certain positions/angles), photos where the primary animal isn't clear (ie pod of dolphins swimming by), bad photos of a fluke (underwater or a picture of a side fluke)</p>\n<p>Bounding boxes close to corners should be assumed to be at the corner, as the current implementation doesn't allow for clicking EXACTLY at the corner.</p>\n<p>We should have a FLUKE button, so that any fluke photos can be sent to a separate fluke model imo. (or maybe use the winning model from the 2018 contest). I'm not sure if the current \"top of marine animal\" detector would work for flukes.</p>\n<p><strong>[03/18/2022]</strong><br>\n<strong>When looking through photos, I've seen:</strong><br>\n-Photos of dolphin/whale fins, where the primary one isn't clear. ie. The one is front is clearer, but the one in the back is smaller.<br>\n-Photos of dolphins underwater in a pod<br>\n-Photo of dolphins from above<br>\n-Photos of fins so small, they're about 10 pixels high and 20 pixels wide<br>\n-Photos of crustaceans on the top of a whale, against a cliff that looks nearly the same, telling the difference isn't easy<br>\n -Photo of flukes, sides of flukes, flukes underwater, flukes so small they're about 20 pixels by 20 pixels.<br>\n-Baby whales next to the PRIMARY whale, obstructing the view</p>\n<p><strong>Looking at these photos, I would very surprised if data leakage (shapes of waves, color of water) isn't playing a MAJOR role in correct identification of individuals, so the identifying features in the images is so small, it would be impossible without other details playing a role.</strong><br>\nI wouldn't be surprised if there are duplicate photos' somewhere in the dataset. I'll be checking this later today.</p>",
      "rawMarkdown": "**[03/17/2022]**\nI support this, if we do 1 annotation per second, it will take an hour per 3.6k images.\n\nAlso it seems that the photos disappear before the bounding box does. Is this a problem? Or is that normal?\n\nAlso if you could drag and drop the bounding box, it would make it faster to annotate. Auto-snapping to the edges/corners AND/OR make the corners redraggable after dropping that would be awesome. I wonder how hard this is to code in.\n\nAre we bounding boxing the surface above the water? Or above and below? I've been annotating only the above areas of an animal UNLESS the camera is underwater, in which case (usually dolphins), I annotate the whole body.\n\nCan there be a button to \"suggest for removal\"? ie. A photo of a clipboard, multiple shots of the same animal in the same position at the same time(potential for the AI for overemphasize certain positions/angles), photos where the primary animal isn't clear (ie pod of dolphins swimming by), bad photos of a fluke (underwater or a picture of a side fluke)\n\n\n\nBounding boxes close to corners should be assumed to be at the corner, as the current implementation doesn't allow for clicking EXACTLY at the corner.\n\nWe should have a FLUKE button, so that any fluke photos can be sent to a separate fluke model imo. (or maybe use the winning model from the 2018 contest). I'm not sure if the current \"top of marine animal\" detector would work for flukes.\n\n**[03/18/2022]**\n**When looking through photos, I've seen:**\n-Photos of dolphin/whale fins, where the primary one isn't clear. ie. The one is front is clearer, but the one in the back is smaller.\n-Photos of dolphins underwater in a pod\n-Photo of dolphins from above\n-Photos of fins so small, they're about 10 pixels high and 20 pixels wide\n-Photos of crustaceans on the top of a whale, against a cliff that looks nearly the same, telling the difference isn't easy\n -Photo of flukes, sides of flukes, flukes underwater, flukes so small they're about 20 pixels by 20 pixels.\n-Baby whales next to the PRIMARY whale, obstructing the view\n\n**Looking at these photos, I would very surprised if data leakage (shapes of waves, color of water) isn't playing a MAJOR role in correct identification of individuals, so the identifying features in the images is so small, it would be impossible without other details playing a role.**\nI wouldn't be surprised if there are duplicate photos' somewhere in the dataset. I'll be checking this later today.",
      "votes": null
    },
    {
      "id": "1736205",
      "postDate": "03/27/2022 05:15:19",
      "content": "<p>Cool app! Great idea!</p>",
      "rawMarkdown": "Cool app! Great idea!",
      "votes": null
    },
    {
      "id": "1738064",
      "postDate": "03/29/2022 01:41:22",
      "content": "<p>Maybe we should provide an example which demonstrates that people should provide an extra area around the whale/dolphin. I notice that half of the bboxs are missing some fin because i think people try to get very close and as a result remove part of the fin (like example below)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2022/wrong.png\" alt=\"\"></p>",
      "rawMarkdown": "Maybe we should provide an example which demonstrates that people should provide an extra area around the whale/dolphin. I notice that half of the bboxs are missing some fin because i think people try to get very close and as a result remove part of the fin (like example below)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2022/wrong.png)",
      "votes": null
    },
    {
      "id": "1738069",
      "postDate": "03/29/2022 01:48:07",
      "content": "<p><strong>Chris I have manually annotated a few thousand of these and would love tips.</strong><br>\nHow much space should be given around the fin?<br>\nHow should whale flukes be annotated? (I've seen many so far)<br>\nHow should bubbles be annotated (I had a photo that I annotated that only had bubbles, as if the whale had recently dived)?<br>\n\"Principle marine animal\". I've seen images with a fin and a fluke in equal size and view from the camera. I usually prioritize the fin, but as flukes make up a part of the dataset, I sometimes question my choices</p>",
      "rawMarkdown": "**Chris I have manually annotated a few thousand of these and would love tips.**\nHow much space should be given around the fin?\nHow should whale flukes be annotated? (I've seen many so far)\nHow should bubbles be annotated (I had a photo that I annotated that only had bubbles, as if the whale had recently dived)?\n\"Principle marine animal\". I've seen images with a fin and a fluke in equal size and view from the camera. I usually prioritize the fin, but as flukes make up a part of the dataset, I sometimes question my choices",
      "votes": null
    },
    {
      "id": "1738072",
      "postDate": "03/29/2022 01:58:40",
      "content": "<p>IMHO, I think the important thing is consistency, i.e. that everyone annotates the images the same. So we either (1) clip the tops of all the fins, or (2) we leave a little border around all the fins. Or we (3) only box the fins, Or etc etc. Without instruction, i think everyone will make different boxes and I think that will be a problem.</p>",
      "rawMarkdown": "IMHO, I think the important thing is consistency, i.e. that everyone annotates the images the same. So we either (1) clip the tops of all the fins, or (2) we leave a little border around all the fins. Or we (3) only box the fins, Or etc etc. Without instruction, i think everyone will make different boxes and I think that will be a problem.",
      "votes": null
    },
    {
      "id": "1744183",
      "postDate": "04/03/2022 17:10:38",
      "content": "<p>Hi, part of the annotations are coming from Yolov5 model trained on a previous whale competition. I believe this type of annotation where some parts are cut out are coming from automated annotations and note manual ones</p>",
      "rawMarkdown": "Hi, part of the annotations are coming from Yolov5 model trained on a previous whale competition. I believe this type of annotation where some parts are cut out are coming from automated annotations and note manual ones",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1708483,
      "author_name": "wolfy73",
      "author_url": "",
      "post_date": "03/01/2022 13:39:32",
      "content": "<p>We nearly reached 200 new annotations + 800 reviews in just 2 days, this is amazing !<br>\nA big thank you to all those who participate 👏 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1712241,
      "author_name": "atharvap329",
      "author_url": "",
      "post_date": "03/04/2022 17:55:24",
      "content": "<p>Amazing work <a href=\"https://www.kaggle.com/wolfy73\" target=\"_blank\">@wolfy73</a>. How many annotations until now ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1712260,
          "author_name": "wolfy73",
          "author_url": "",
          "post_date": "03/04/2022 18:27:39",
          "content": "<p>The reviewed dataset contains currently 2671 annotated images !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1712257,
      "author_name": "wolfy73",
      "author_url": "",
      "post_date": "03/04/2022 18:26:15",
      "content": "<p>I made a public dataset that contains the dataset. It's updated on a daily basis using a scheduled notebook:<br>\n<a href=\"https://www.kaggle.com/wolfy73/wandd-crowed-sourced-bounging-boxes\" target=\"_blank\">🐳&amp;🐬 - Crowdsourced bounding boxes</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1712735,
          "author_name": "alexlwh",
          "author_url": "",
          "post_date": "03/05/2022 08:53:07",
          "content": "<p>FYI I think the link is not working</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1712745,
          "author_name": "wolfy73",
          "author_url": "",
          "post_date": "03/05/2022 09:15:12",
          "content": "<p>Thank you for letting me know, it should work now !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1712729,
      "author_name": "alexlwh",
      "author_url": "",
      "post_date": "03/05/2022 08:48:58",
      "content": "<p>This is a great initiative! <a href=\"https://www.kaggle.com/wolfy73\" target=\"_blank\">@wolfy73</a> </p>\n<p>I am thinking if it is possible to integrate Active Learning into the whole pipeline to ramp up the marginal gain of our annotations.</p>\n<p>My raw idea is to make a notebook that is scheduled to pull our current annotations and train a model, and then rank the unlabelled images by their \"worthiness\" to be annotated. As a step further, the rank can be fed to your app to complete the loop. We keep this loop going until the performance gain diminish at a certain extent.</p>\n<p>One uncertainty is how feasible active learning can be applied on object detection, as it is typically applied on classification problems.</p>\n<p>I am trying to do some research to assess its feasibility and the possible pipeline, let me know if you are interested to also work on this! I believe this is a good chance to apply active learning in real problem and assess its effectiveness!</p>\n<p>FYI, here are some initial finding I have got:</p>\n<ol>\n<li><a href=\"https://github.com/NVlabs/AL-MDN\" target=\"_blank\">Active Learning for Deep Object Detection via Probabilistic Modeling</a></li>\n<li><a href=\"https://github.com/superannotateai/active_learning\" target=\"_blank\">Learning Loss for Active Learning</a></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 1712748,
          "author_name": "wolfy73",
          "author_url": "",
          "post_date": "03/05/2022 09:19:57",
          "content": "<p>I'm interested !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1727314,
      "author_name": "dentistdad",
      "author_url": "",
      "post_date": "03/17/2022 21:22:53",
      "content": "<p><strong>[03/17/2022]</strong><br>\nI support this, if we do 1 annotation per second, it will take an hour per 3.6k images.</p>\n<p>Also it seems that the photos disappear before the bounding box does. Is this a problem? Or is that normal?</p>\n<p>Also if you could drag and drop the bounding box, it would make it faster to annotate. Auto-snapping to the edges/corners AND/OR make the corners redraggable after dropping that would be awesome. I wonder how hard this is to code in.</p>\n<p>Are we bounding boxing the surface above the water? Or above and below? I've been annotating only the above areas of an animal UNLESS the camera is underwater, in which case (usually dolphins), I annotate the whole body.</p>\n<p>Can there be a button to \"suggest for removal\"? ie. A photo of a clipboard, multiple shots of the same animal in the same position at the same time(potential for the AI for overemphasize certain positions/angles), photos where the primary animal isn't clear (ie pod of dolphins swimming by), bad photos of a fluke (underwater or a picture of a side fluke)</p>\n<p>Bounding boxes close to corners should be assumed to be at the corner, as the current implementation doesn't allow for clicking EXACTLY at the corner.</p>\n<p>We should have a FLUKE button, so that any fluke photos can be sent to a separate fluke model imo. (or maybe use the winning model from the 2018 contest). I'm not sure if the current \"top of marine animal\" detector would work for flukes.</p>\n<p><strong>[03/18/2022]</strong><br>\n<strong>When looking through photos, I've seen:</strong><br>\n-Photos of dolphin/whale fins, where the primary one isn't clear. ie. The one is front is clearer, but the one in the back is smaller.<br>\n-Photos of dolphins underwater in a pod<br>\n-Photo of dolphins from above<br>\n-Photos of fins so small, they're about 10 pixels high and 20 pixels wide<br>\n-Photos of crustaceans on the top of a whale, against a cliff that looks nearly the same, telling the difference isn't easy<br>\n -Photo of flukes, sides of flukes, flukes underwater, flukes so small they're about 20 pixels by 20 pixels.<br>\n-Baby whales next to the PRIMARY whale, obstructing the view</p>\n<p><strong>Looking at these photos, I would very surprised if data leakage (shapes of waves, color of water) isn't playing a MAJOR role in correct identification of individuals, so the identifying features in the images is so small, it would be impossible without other details playing a role.</strong><br>\nI wouldn't be surprised if there are duplicate photos' somewhere in the dataset. I'll be checking this later today.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1736205,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/27/2022 05:15:19",
      "content": "<p>Cool app! Great idea!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1738064,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/29/2022 01:41:22",
      "content": "<p>Maybe we should provide an example which demonstrates that people should provide an extra area around the whale/dolphin. I notice that half of the bboxs are missing some fin because i think people try to get very close and as a result remove part of the fin (like example below)</p>\n<p><img src=\"https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2022/wrong.png\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 1738069,
          "author_name": "dentistdad",
          "author_url": "",
          "post_date": "03/29/2022 01:48:07",
          "content": "<p><strong>Chris I have manually annotated a few thousand of these and would love tips.</strong><br>\nHow much space should be given around the fin?<br>\nHow should whale flukes be annotated? (I've seen many so far)<br>\nHow should bubbles be annotated (I had a photo that I annotated that only had bubbles, as if the whale had recently dived)?<br>\n\"Principle marine animal\". I've seen images with a fin and a fluke in equal size and view from the camera. I usually prioritize the fin, but as flukes make up a part of the dataset, I sometimes question my choices</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1738072,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/29/2022 01:58:40",
          "content": "<p>IMHO, I think the important thing is consistency, i.e. that everyone annotates the images the same. So we either (1) clip the tops of all the fins, or (2) we leave a little border around all the fins. Or we (3) only box the fins, Or etc etc. Without instruction, i think everyone will make different boxes and I think that will be a problem.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1744183,
          "author_name": "wolfy73",
          "author_url": "",
          "post_date": "04/03/2022 17:10:38",
          "content": "<p>Hi, part of the annotations are coming from Yolov5 model trained on a previous whale competition. I believe this type of annotation where some parts are cut out are coming from automated annotations and note manual ones</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1706548": "Hi kagglers 👋,\n\nNow that we know how important it is to properly crop the images in this competition, I've been wondering about a way to make a good quality dataset of bounding boxes. \nNumerous manual annotations are already available in public datasets, but it seems like it's still not enough...\n\nSo I've been working on a way to crowd-source a bounding box dataset. I made and hosted a little app where anybody can come and annotate images or review other people's annotations.\n\n### Links\n**The app is available here**: [http://happywhale.theoboyer.fr](http://happywhale.theoboyer.fr)\n\n**The source code and explanations on how the app works are available on this GitHub repo**: [https://github.com/TheoBoyer/Manual-bbox-annotation-tool](https://github.com/TheoBoyer/Manual-bbox-annotation-tool)\n\n### Download the dataset\nOn the home page you can find buttons to download directly either the raw annotations and review informations or the final dataset constituted of only manually reviewed samples. You can also download the dataset in a kaggle notebook with these command lines:\n\n`!curl http://happywhale.theoboyer.fr/data/annotations.csv > annotations.csv`\n\n`!curl http://happywhale.theoboyer.fr/data/dataset.csv > dataset.csv`\n\n### Note\n* The dataset doesn't start from scratch, I merged several public datasets to begin ! The details of sources and how I gathered manual annotations are available in the notebook: [🐳&🐬 - 👨‍🔬 Merging public bounding box datasets](https://www.kaggle.com/wolfy73/merging-public-bounding-box-\n* Some automatic annotations are already entered and just need a manual review. The annotations were obtained by taking the dataset from [Happywhale: Cropped Dataset [YOLOv5] ✂️](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5) and filtered using the methods described in [🐳&🐬 - Filter YOLOv5 failure cases](https://www.kaggle.com/wolfy73/filter-yolov5-failure-cases)\n* Only the training images are in the app to comply with the competition's rules \n\nDo not hesitate to come and annotate some images, or review some annotations (while your models are training for example 😏), it will help everyone here !\n\nFor more information, see the README of the repo or ask your questions on this topic !\n\nHappy Kaggling everyone 👍",
    "1708483": "We nearly reached 200 new annotations + 800 reviews in just 2 days, this is amazing !\nA big thank you to all those who participate 👏",
    "1712241": "Amazing work @wolfy73. How many annotations until now ?",
    "1712257": "I made a public dataset that contains the dataset. It's updated on a daily basis using a scheduled notebook:\n[🐳&🐬 - Crowdsourced bounding boxes](https://www.kaggle.com/wolfy73/wandd-crowed-sourced-bounging-boxes)",
    "1712260": "The reviewed dataset contains currently 2671 annotated images !",
    "1712729": "This is a great initiative! @wolfy73 \n\nI am thinking if it is possible to integrate Active Learning into the whole pipeline to ramp up the marginal gain of our annotations.\n\nMy raw idea is to make a notebook that is scheduled to pull our current annotations and train a model, and then rank the unlabelled images by their \"worthiness\" to be annotated. As a step further, the rank can be fed to your app to complete the loop. We keep this loop going until the performance gain diminish at a certain extent.\n\nOne uncertainty is how feasible active learning can be applied on object detection, as it is typically applied on classification problems.\n\nI am trying to do some research to assess its feasibility and the possible pipeline, let me know if you are interested to also work on this! I believe this is a good chance to apply active learning in real problem and assess its effectiveness!\n\nFYI, here are some initial finding I have got:\n1. [Active Learning for Deep Object Detection via Probabilistic Modeling](https://github.com/NVlabs/AL-MDN)\n2. [Learning Loss for Active Learning](https://github.com/superannotateai/active_learning)",
    "1712735": "FYI I think the link is not working",
    "1712745": "Thank you for letting me know, it should work now !",
    "1712748": "I'm interested !",
    "1727314": "**[03/17/2022]**\nI support this, if we do 1 annotation per second, it will take an hour per 3.6k images.\n\nAlso it seems that the photos disappear before the bounding box does. Is this a problem? Or is that normal?\n\nAlso if you could drag and drop the bounding box, it would make it faster to annotate. Auto-snapping to the edges/corners AND/OR make the corners redraggable after dropping that would be awesome. I wonder how hard this is to code in.\n\nAre we bounding boxing the surface above the water? Or above and below? I've been annotating only the above areas of an animal UNLESS the camera is underwater, in which case (usually dolphins), I annotate the whole body.\n\nCan there be a button to \"suggest for removal\"? ie. A photo of a clipboard, multiple shots of the same animal in the same position at the same time(potential for the AI for overemphasize certain positions/angles), photos where the primary animal isn't clear (ie pod of dolphins swimming by), bad photos of a fluke (underwater or a picture of a side fluke)\n\n\n\nBounding boxes close to corners should be assumed to be at the corner, as the current implementation doesn't allow for clicking EXACTLY at the corner.\n\nWe should have a FLUKE button, so that any fluke photos can be sent to a separate fluke model imo. (or maybe use the winning model from the 2018 contest). I'm not sure if the current \"top of marine animal\" detector would work for flukes.\n\n**[03/18/2022]**\n**When looking through photos, I've seen:**\n-Photos of dolphin/whale fins, where the primary one isn't clear. ie. The one is front is clearer, but the one in the back is smaller.\n-Photos of dolphins underwater in a pod\n-Photo of dolphins from above\n-Photos of fins so small, they're about 10 pixels high and 20 pixels wide\n-Photos of crustaceans on the top of a whale, against a cliff that looks nearly the same, telling the difference isn't easy\n -Photo of flukes, sides of flukes, flukes underwater, flukes so small they're about 20 pixels by 20 pixels.\n-Baby whales next to the PRIMARY whale, obstructing the view\n\n**Looking at these photos, I would very surprised if data leakage (shapes of waves, color of water) isn't playing a MAJOR role in correct identification of individuals, so the identifying features in the images is so small, it would be impossible without other details playing a role.**\nI wouldn't be surprised if there are duplicate photos' somewhere in the dataset. I'll be checking this later today.",
    "1736205": "Cool app! Great idea!",
    "1738064": "Maybe we should provide an example which demonstrates that people should provide an extra area around the whale/dolphin. I notice that half of the bboxs are missing some fin because i think people try to get very close and as a result remove part of the fin (like example below)\n\n![](https://raw.githubusercontent.com/cdeotte/Kaggle_Images/main/Mar-2022/wrong.png)",
    "1738069": "**Chris I have manually annotated a few thousand of these and would love tips.**\nHow much space should be given around the fin?\nHow should whale flukes be annotated? (I've seen many so far)\nHow should bubbles be annotated (I had a photo that I annotated that only had bubbles, as if the whale had recently dived)?\n\"Principle marine animal\". I've seen images with a fin and a fluke in equal size and view from the camera. I usually prioritize the fin, but as flukes make up a part of the dataset, I sometimes question my choices",
    "1738072": "IMHO, I think the important thing is consistency, i.e. that everyone annotates the images the same. So we either (1) clip the tops of all the fins, or (2) we leave a little border around all the fins. Or we (3) only box the fins, Or etc etc. Without instruction, i think everyone will make different boxes and I think that will be a problem.",
    "1744183": "Hi, part of the annotations are coming from Yolov5 model trained on a previous whale competition. I believe this type of annotation where some parts are cut out are coming from automated annotations and note manual ones"
  },
  "source": "meta"
}