{
  "id": 16353,
  "title": "Annotated faces for NOAA Right Whale",
  "url": "/competitions/noaa-right-whale-recognition/discussion/16353",
  "author_name": "",
  "post_date": "2015-09-06T22:46:03.157Z",
  "votes": 26,
  "comment_count": 9,
  "views": 3752,
  "content": "<p>I have annotated 300 whale faces for 11 of the most frequent whales in the dataset and released those annotations online for anyone to use.</p>\n\n<p><a href=\"https://github.com/Smerity/right_whale_hunt\">https://github.com/Smerity/right_whale_hunt</a></p>\n\n<p>The annotations were done using <a href=\"http://sloth.readthedocs.org/en/latest/\">Sloth</a> and the Sloth configuration file has been included. There is also simple thumbnail extraction code included which processes the JSON file produced by Sloth. The JSON format should be readily convertible into the format used by Matlab.</p>\n\n<p><img src=\"http://i.imgur.com/o5cf6pd.jpg\" alt=\"Whale 26288\" title></p>\n\n<p>If people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. The thumbnail extraction code could also use some love - the optimal case would be to not pad the images with white but instead to use extensions from the original image.</p>\n\n<p>The annotation task is not too difficult, especially given Sloth. There are some images which are highly ambiguous but the majority are quite clear. I've skipped ambiguous images as we'd likely need clear annotation guidelines to resolve them.</p>\n\n<p>Two notes I've picked up from manually analyzing so many images:</p>\n\n<ul>\n<li>Many of the classifiers will be taking advantage of whale images being taken at similar times, occasionally within seconds of each other. I've tested this by running a CNN on 96x64 thumbnails of the full images. The parable about the US army using neural networks to automatically detect camouflaged enemy tanks, then discovering it only picked out clear and cloudy photos, seems applicable here.</li>\n<li>Some of the cameras have defects that could be used in a similar way. As an example, there is a distinctive black dot in the right bottom of [w_3620.jpg, w_5478.jpg, w_6828.jpg, w_7812.jpg, w_8351.jpg, w_934.jpg, w_9463.jpg] amongst others.</li>\n<li>Specularity and ocean spray make identifying callosity patterns quite difficult.</li>\n<li>Some manner of data augmentation using blurring or obfuscation would be a good idea due to callosity patterns being below the surface of the water.</li>\n</ul>\n\n<p>P.S. I nearly lost all the annotations as I stupidly ran a <code>sed 's/imgs_sub/imgs/g' sloth.json &gt; sloth.json</code> (piping into the same file). It was like a slow motion horror film. My head said &quot;I'll pipe into the same filename in <code>/tmp/</code> to prevent piping into the same file and destroying it&quot; but my fingers said &quot;I'm lazy&quot;. I recovered the file by grepping my hard drive for the distinctive JSON fields and then reassembling the individual JSON records, saving 285 of the 289 then annotated whales. Horror.</p>",
  "messages": [
    {
      "id": "91738",
      "postDate": "09/06/2015 22:46:03",
      "content": "<p>I have annotated 300 whale faces for 11 of the most frequent whales in the dataset and released those annotations online for anyone to use.</p>\n\n<p><a href=\"https://github.com/Smerity/right_whale_hunt\">https://github.com/Smerity/right_whale_hunt</a></p>\n\n<p>The annotations were done using <a href=\"http://sloth.readthedocs.org/en/latest/\">Sloth</a> and the Sloth configuration file has been included. There is also simple thumbnail extraction code included which processes the JSON file produced by Sloth. The JSON format should be readily convertible into the format used by Matlab.</p>\n\n<p><img src=\"http://i.imgur.com/o5cf6pd.jpg\" alt=\"Whale 26288\" title></p>\n\n<p>If people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. The thumbnail extraction code could also use some love - the optimal case would be to not pad the images with white but instead to use extensions from the original image.</p>\n\n<p>The annotation task is not too difficult, especially given Sloth. There are some images which are highly ambiguous but the majority are quite clear. I've skipped ambiguous images as we'd likely need clear annotation guidelines to resolve them.</p>\n\n<p>Two notes I've picked up from manually analyzing so many images:</p>\n\n<ul>\n<li>Many of the classifiers will be taking advantage of whale images being taken at similar times, occasionally within seconds of each other. I've tested this by running a CNN on 96x64 thumbnails of the full images. The parable about the US army using neural networks to automatically detect camouflaged enemy tanks, then discovering it only picked out clear and cloudy photos, seems applicable here.</li>\n<li>Some of the cameras have defects that could be used in a similar way. As an example, there is a distinctive black dot in the right bottom of [w_3620.jpg, w_5478.jpg, w_6828.jpg, w_7812.jpg, w_8351.jpg, w_934.jpg, w_9463.jpg] amongst others.</li>\n<li>Specularity and ocean spray make identifying callosity patterns quite difficult.</li>\n<li>Some manner of data augmentation using blurring or obfuscation would be a good idea due to callosity patterns being below the surface of the water.</li>\n</ul>\n\n<p>P.S. I nearly lost all the annotations as I stupidly ran a <code>sed 's/imgs_sub/imgs/g' sloth.json &gt; sloth.json</code> (piping into the same file). It was like a slow motion horror film. My head said &quot;I'll pipe into the same filename in <code>/tmp/</code> to prevent piping into the same file and destroying it&quot; but my fingers said &quot;I'm lazy&quot;. I recovered the file by grepping my hard drive for the distinctive JSON fields and then reassembling the individual JSON records, saving 285 of the 289 then annotated whales. Horror.</p>",
      "rawMarkdown": "I have annotated 300 whale faces for 11 of the most frequent whales in the dataset and released those annotations online for anyone to use.\r\n\r\n[https://github.com/Smerity/right_whale_hunt][2]\r\n\r\nThe annotations were done using [Sloth](http://sloth.readthedocs.org/en/latest/) and the Sloth configuration file has been included. There is also simple thumbnail extraction code included which processes the JSON file produced by Sloth. The JSON format should be readily convertible into the format used by Matlab.\r\n\r\n![Whale 26288][1]\r\n\r\nIf people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. The thumbnail extraction code could also use some love - the optimal case would be to not pad the images with white but instead to use extensions from the original image.\r\n\r\nThe annotation task is not too difficult, especially given Sloth. There are some images which are highly ambiguous but the majority are quite clear. I've skipped ambiguous images as we'd likely need clear annotation guidelines to resolve them.\r\n\r\nTwo notes I've picked up from manually analyzing so many images:\r\n\r\n - Many of the classifiers will be taking advantage of whale images being taken at similar times, occasionally within seconds of each other. I've tested this by running a CNN on 96x64 thumbnails of the full images. The parable about the US army using neural networks to automatically detect camouflaged enemy tanks, then discovering it only picked out clear and cloudy photos, seems applicable here.\r\n - Some of the cameras have defects that could be used in a similar way. As an example, there is a distinctive black dot in the right bottom of [w_3620.jpg, w_5478.jpg, w_6828.jpg, w_7812.jpg, w_8351.jpg, w_934.jpg, w_9463.jpg] amongst others.\r\n - Specularity and ocean spray make identifying callosity patterns quite difficult.\r\n - Some manner of data augmentation using blurring or obfuscation would be a good idea due to callosity patterns being below the surface of the water.\r\n\r\nP.S. I nearly lost all the annotations as I stupidly ran a `sed 's/imgs_sub/imgs/g' sloth.json > sloth.json` (piping into the same file). It was like a slow motion horror film. My head said \"I'll pipe into the same filename in `/tmp/` to prevent piping into the same file and destroying it\" but my fingers said \"I'm lazy\". I recovered the file by grepping my hard drive for the distinctive JSON fields and then reassembling the individual JSON records, saving 285 of the 289 then annotated whales. Horror.\r\n\r\n  [1]: http://i.imgur.com/o5cf6pd.jpg\r\n  [2]: https://github.com/Smerity/right_whale_hunt",
      "votes": null
    },
    {
      "id": "91739",
      "postDate": "09/06/2015 23:34:51",
      "content": "<p>This is a great idea, and I would like to contribute.</p>\n\n<p>I don't think the kaggle rules against rehosting data apply here since you're only providing json files.</p>\n\n<p>Are there any guidelines we should use for cropping so that we're consistent across whales? For example in some images you can't see the whale's face at all. It might be best to omit these at first.</p>",
      "rawMarkdown": "This is a great idea, and I would like to contribute.\r\n\r\nI don't think the kaggle rules against rehosting data apply here since you're only providing json files.\r\n\r\nAre there any guidelines we should use for cropping so that we're consistent across whales? For example in some images you can't see the whale's face at all. It might be best to omit these at first.",
      "votes": null
    },
    {
      "id": "91740",
      "postDate": "09/07/2015 00:01:00",
      "content": "<p>[quote=Smerity;91738]</p>\n\n<p>If people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. </p>\n\n<p>[/quote]</p>\n\n<p>Fantastic work.</p>\n\n<p>I suppose some redundancy is good. (Correctly me if I'm wrong, but different croppings of the same face shouldn't an issue, right?)</p>\n\n<p>On the other hand, if everyone starts at 1 . . .</p>",
      "rawMarkdown": "[quote=Smerity;91738]\r\n\r\nIf people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. \r\n\r\n[/quote]\r\n\r\nFantastic work.\r\n\r\nI suppose some redundancy is good. (Correctly me if I'm wrong, but different croppings of the same face shouldn't an issue, right?)\r\n\r\nOn the other hand, if everyone starts at 1 . . .",
      "votes": null
    },
    {
      "id": "91745",
      "postDate": "09/07/2015 02:52:37",
      "content": "<p>@James King, I believe that providing just the JSON files should be allowable as they're useless without the data hosted on Kaggle.</p>\n\n<p>Re: guidelines for cropping, I mention the very loose guidelines I had whilst annotating at <a href=\"https://github.com/Smerity/right_whale_hunt#contributing-to-annotations\">https://github.com/Smerity/right_whale_hunt#contributing-to-annotations</a>. I might extend it to include more specific guidelines.</p>\n\n<p>@inversion, some redundancy could indeed be good if someone put the time in to integrate them, though I still think having better coverage of the dataset is likely more beneficial than much overlapping.</p>\n\n<hr>\n\n<p>My proposal for having multiple people contribute well would be to have everyone annotate files randomly, save their annotations as JSON files, then write a script to merge the JSON annotations as desired. This would allow people to sanely add pull requests. For anyone interested in getting started with that (taken partly from the <a href=\"http://sloth.readthedocs.org/en/latest/examples.html#adding-every-nth-image-to-label-file\">Sloth example</a>):</p>\n\n<pre><code>find imgs/*/* -iname &quot;*.jpg&quot; | shuf | xargs sloth appendfiles whale_faces_smerity.json\n</code></pre>\n\n<p>and then start annotating with:</p>\n\n<pre><code>sloth --config slothwhales.py whale_faces_smerity.json\n</code></pre>\n\n<p>Then submit a pull request, replacing <code>smerity</code> with your identifying handle for easy merging.</p>",
      "rawMarkdown": "James King, I believe that providing just the JSON files should be allowable as they're useless without the data hosted on Kaggle.\r\n\r\nRe: guidelines for cropping, I mention the very loose guidelines I had whilst annotating at https://github.com/Smerity/right_whale_hunt#contributing-to-annotations. I might extend it to include more specific guidelines.\r\n\r\n@inversion, some redundancy could indeed be good if someone put the time in to integrate them, though I still think having better coverage of the dataset is likely more beneficial than much overlapping.\r\n\r\n---\r\n\r\nMy proposal for having multiple people contribute well would be to have everyone annotate files randomly, save their annotations as JSON files, then write a script to merge the JSON annotations as desired. This would allow people to sanely add pull requests. For anyone interested in getting started with that (taken partly from the [Sloth example][1]):\r\n\r\n    find imgs/*/* -iname \"*.jpg\" | shuf | xargs sloth appendfiles whale_faces_smerity.json\r\n\r\nand then start annotating with:\r\n\r\n    sloth --config slothwhales.py whale_faces_smerity.json\r\n\r\nThen submit a pull request, replacing `smerity` with your identifying handle for easy merging.\r\n\r\n  [1]: http://sloth.readthedocs.org/en/latest/examples.html#adding-every-nth-image-to-label-file",
      "votes": null
    },
    {
      "id": "91746",
      "postDate": "09/07/2015 04:13:44",
      "content": "<p>OK, I see that by running inversion's python script first we separate out the train images. We need to make sure we don't annotate test images by mistake.</p>",
      "rawMarkdown": "OK, I see that by running inversion's python script first we separate out the train images. We need to make sure we don't annotate test images by mistake.",
      "votes": null
    },
    {
      "id": "91747",
      "postDate": "09/07/2015 04:28:40",
      "content": "<p>@James King, correct. The <code>find imgs/*/* -iname &quot;*.jpg&quot;</code> part of the commands I listed above and on the repository should only add training images to be annotated as only those images with a label are stored in a subdirectory (i.e. the whale's ID).</p>",
      "rawMarkdown": "James King, correct. The `find imgs/*/* -iname \"*.jpg\"` part of the commands I listed above and on the repository should only add training images to be annotated as only those images with a label are stored in a subdirectory (i.e. the whale's ID).",
      "votes": null
    },
    {
      "id": "91804",
      "postDate": "09/07/2015 20:51:25",
      "content": "<p>Smerity, Thanks for the example. inversion, thanks for the script.</p>\n\n<p>I have annotated 300 images (Head only), just submitted a pull request.</p>",
      "rawMarkdown": "Smerity, Thanks for the example. inversion, thanks for the script.\r\n\r\nI have annotated 300 images (Head only), just submitted a pull request.",
      "votes": null
    },
    {
      "id": "91807",
      "postDate": "09/07/2015 23:06:08",
      "content": "<p>Thanks to @James King and @sunil, there are now 630 annotated whale faces. I've updated the extraction script such that it can handle as many JSON annotations as you can throw at it (overwriting any duplicate extractions with the most recent) and also pad to a square instead of the original annotation's dimensions.</p>\n\n<p>Note: the padding does have a small bug (occasional off by one results in one pixel of white) and if it's off the edge of the original photo, the expanded region is black.</p>",
      "rawMarkdown": "Thanks to @James King and @sunil, there are now 630 annotated whale faces. I've updated the extraction script such that it can handle as many JSON annotations as you can throw at it (overwriting any duplicate extractions with the most recent) and also pad to a square instead of the original annotation's dimensions.\r\n\r\nNote: the padding does have a small bug (occasional off by one results in one pixel of white) and if it's off the edge of the original photo, the expanded region is black.",
      "votes": null
    },
    {
      "id": "97379",
      "postDate": "10/26/2015 17:49:31",
      "content": "<p>@Smerity,</p>\n\n<p>don't know if u are still working on this competition, but I have gotten across some problems during the conversion of data from JSON to ROI.</p>\n\n<p>Error messages whlie I do trainCascadeObjectDetector('detectorFile.xml', positiveInstances, 'negativeFolder'... in MATLAB\nError using trainCascadeObjectDetector (line 245)\nError reading instance 1 from image 1, bounding box possibly out of image bounds.</p>\n\n<p>It seems like the image are out of bound if applying the json format data directly to positiveInstances?\nAppreciate if there's any suggestion, thank you!</p>",
      "rawMarkdown": "Smerity,\r\n\r\ndon't know if u are still working on this competition, but I have gotten across some problems during the conversion of data from JSON to ROI.\r\n\r\nError messages whlie I do trainCascadeObjectDetector('detectorFile.xml', positiveInstances, 'negativeFolder'... in MATLAB\r\nError using trainCascadeObjectDetector (line 245)\r\nError reading instance 1 from image 1, bounding box possibly out of image bounds.\r\n\r\nIt seems like the image are out of bound if applying the json format data directly to positiveInstances?\r\nAppreciate if there's any suggestion, thank you!",
      "votes": null
    },
    {
      "id": "97390",
      "postDate": "10/26/2015 20:51:08",
      "content": "<p>How are you guys transforming JSON to matlab compatible ROI? </p>\n\n<p>Waiting for your reply.</p>",
      "rawMarkdown": "How are you guys transforming JSON to matlab compatible ROI? \r\n\r\nWaiting for your reply.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 91739,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "09/06/2015 23:34:51",
      "content": "<p>This is a great idea, and I would like to contribute.</p>\n\n<p>I don't think the kaggle rules against rehosting data apply here since you're only providing json files.</p>\n\n<p>Are there any guidelines we should use for cropping so that we're consistent across whales? For example in some images you can't see the whale's face at all. It might be best to omit these at first.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91740,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "09/07/2015 00:01:00",
      "content": "<p>[quote=Smerity;91738]</p>\n\n<p>If people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. </p>\n\n<p>[/quote]</p>\n\n<p>Fantastic work.</p>\n\n<p>I suppose some redundancy is good. (Correctly me if I'm wrong, but different croppings of the same face shouldn't an issue, right?)</p>\n\n<p>On the other hand, if everyone starts at 1 . . .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91745,
      "author_name": "smerity",
      "author_url": "",
      "post_date": "09/07/2015 02:52:37",
      "content": "<p>@James King, I believe that providing just the JSON files should be allowable as they're useless without the data hosted on Kaggle.</p>\n\n<p>Re: guidelines for cropping, I mention the very loose guidelines I had whilst annotating at <a href=\"https://github.com/Smerity/right_whale_hunt#contributing-to-annotations\">https://github.com/Smerity/right_whale_hunt#contributing-to-annotations</a>. I might extend it to include more specific guidelines.</p>\n\n<p>@inversion, some redundancy could indeed be good if someone put the time in to integrate them, though I still think having better coverage of the dataset is likely more beneficial than much overlapping.</p>\n\n<hr>\n\n<p>My proposal for having multiple people contribute well would be to have everyone annotate files randomly, save their annotations as JSON files, then write a script to merge the JSON annotations as desired. This would allow people to sanely add pull requests. For anyone interested in getting started with that (taken partly from the <a href=\"http://sloth.readthedocs.org/en/latest/examples.html#adding-every-nth-image-to-label-file\">Sloth example</a>):</p>\n\n<pre><code>find imgs/*/* -iname &quot;*.jpg&quot; | shuf | xargs sloth appendfiles whale_faces_smerity.json\n</code></pre>\n\n<p>and then start annotating with:</p>\n\n<pre><code>sloth --config slothwhales.py whale_faces_smerity.json\n</code></pre>\n\n<p>Then submit a pull request, replacing <code>smerity</code> with your identifying handle for easy merging.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91746,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "09/07/2015 04:13:44",
      "content": "<p>OK, I see that by running inversion's python script first we separate out the train images. We need to make sure we don't annotate test images by mistake.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91747,
      "author_name": "smerity",
      "author_url": "",
      "post_date": "09/07/2015 04:28:40",
      "content": "<p>@James King, correct. The <code>find imgs/*/* -iname &quot;*.jpg&quot;</code> part of the commands I listed above and on the repository should only add training images to be annotated as only those images with a label are stored in a subdirectory (i.e. the whale's ID).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91804,
      "author_name": "sunilv",
      "author_url": "",
      "post_date": "09/07/2015 20:51:25",
      "content": "<p>Smerity, Thanks for the example. inversion, thanks for the script.</p>\n\n<p>I have annotated 300 images (Head only), just submitted a pull request.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 91807,
      "author_name": "smerity",
      "author_url": "",
      "post_date": "09/07/2015 23:06:08",
      "content": "<p>Thanks to @James King and @sunil, there are now 630 annotated whale faces. I've updated the extraction script such that it can handle as many JSON annotations as you can throw at it (overwriting any duplicate extractions with the most recent) and also pad to a square instead of the original annotation's dimensions.</p>\n\n<p>Note: the padding does have a small bug (occasional off by one results in one pixel of white) and if it's off the edge of the original photo, the expanded region is black.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 97379,
      "author_name": "luke831215",
      "author_url": "",
      "post_date": "10/26/2015 17:49:31",
      "content": "<p>@Smerity,</p>\n\n<p>don't know if u are still working on this competition, but I have gotten across some problems during the conversion of data from JSON to ROI.</p>\n\n<p>Error messages whlie I do trainCascadeObjectDetector('detectorFile.xml', positiveInstances, 'negativeFolder'... in MATLAB\nError using trainCascadeObjectDetector (line 245)\nError reading instance 1 from image 1, bounding box possibly out of image bounds.</p>\n\n<p>It seems like the image are out of bound if applying the json format data directly to positiveInstances?\nAppreciate if there's any suggestion, thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 97390,
      "author_name": "avi9999",
      "author_url": "",
      "post_date": "10/26/2015 20:51:08",
      "content": "<p>How are you guys transforming JSON to matlab compatible ROI? </p>\n\n<p>Waiting for your reply.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "91738": "I have annotated 300 whale faces for 11 of the most frequent whales in the dataset and released those annotations online for anyone to use.\r\n\r\n[https://github.com/Smerity/right_whale_hunt][2]\r\n\r\nThe annotations were done using [Sloth](http://sloth.readthedocs.org/en/latest/) and the Sloth configuration file has been included. There is also simple thumbnail extraction code included which processes the JSON file produced by Sloth. The JSON format should be readily convertible into the format used by Matlab.\r\n\r\n![Whale 26288][1]\r\n\r\nIf people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. The thumbnail extraction code could also use some love - the optimal case would be to not pad the images with white but instead to use extensions from the original image.\r\n\r\nThe annotation task is not too difficult, especially given Sloth. There are some images which are highly ambiguous but the majority are quite clear. I've skipped ambiguous images as we'd likely need clear annotation guidelines to resolve them.\r\n\r\nTwo notes I've picked up from manually analyzing so many images:\r\n\r\n - Many of the classifiers will be taking advantage of whale images being taken at similar times, occasionally within seconds of each other. I've tested this by running a CNN on 96x64 thumbnails of the full images. The parable about the US army using neural networks to automatically detect camouflaged enemy tanks, then discovering it only picked out clear and cloudy photos, seems applicable here.\r\n - Some of the cameras have defects that could be used in a similar way. As an example, there is a distinctive black dot in the right bottom of [w_3620.jpg, w_5478.jpg, w_6828.jpg, w_7812.jpg, w_8351.jpg, w_934.jpg, w_9463.jpg] amongst others.\r\n - Specularity and ocean spray make identifying callosity patterns quite difficult.\r\n - Some manner of data augmentation using blurring or obfuscation would be a good idea due to callosity patterns being below the surface of the water.\r\n\r\nP.S. I nearly lost all the annotations as I stupidly ran a `sed 's/imgs_sub/imgs/g' sloth.json > sloth.json` (piping into the same file). It was like a slow motion horror film. My head said \"I'll pipe into the same filename in `/tmp/` to prevent piping into the same file and destroying it\" but my fingers said \"I'm lazy\". I recovered the file by grepping my hard drive for the distinctive JSON fields and then reassembling the individual JSON records, saving 285 of the 289 then annotated whales. Horror.\r\n\r\n  [1]: http://i.imgur.com/o5cf6pd.jpg\r\n  [2]: https://github.com/Smerity/right_whale_hunt",
    "91739": "This is a great idea, and I would like to contribute.\r\n\r\nI don't think the kaggle rules against rehosting data apply here since you're only providing json files.\r\n\r\nAre there any guidelines we should use for cropping so that we're consistent across whales? For example in some images you can't see the whale's face at all. It might be best to omit these at first.",
    "91740": "[quote=Smerity;91738]\r\n\r\nIf people would like to contribute, more annotated images would be brilliant, though I'm not sure how we can coordinate that to ensure there isn't large amounts of redundant effort. \r\n\r\n[/quote]\r\n\r\nFantastic work.\r\n\r\nI suppose some redundancy is good. (Correctly me if I'm wrong, but different croppings of the same face shouldn't an issue, right?)\r\n\r\nOn the other hand, if everyone starts at 1 . . .",
    "91745": "James King, I believe that providing just the JSON files should be allowable as they're useless without the data hosted on Kaggle.\r\n\r\nRe: guidelines for cropping, I mention the very loose guidelines I had whilst annotating at https://github.com/Smerity/right_whale_hunt#contributing-to-annotations. I might extend it to include more specific guidelines.\r\n\r\n@inversion, some redundancy could indeed be good if someone put the time in to integrate them, though I still think having better coverage of the dataset is likely more beneficial than much overlapping.\r\n\r\n---\r\n\r\nMy proposal for having multiple people contribute well would be to have everyone annotate files randomly, save their annotations as JSON files, then write a script to merge the JSON annotations as desired. This would allow people to sanely add pull requests. For anyone interested in getting started with that (taken partly from the [Sloth example][1]):\r\n\r\n    find imgs/*/* -iname \"*.jpg\" | shuf | xargs sloth appendfiles whale_faces_smerity.json\r\n\r\nand then start annotating with:\r\n\r\n    sloth --config slothwhales.py whale_faces_smerity.json\r\n\r\nThen submit a pull request, replacing `smerity` with your identifying handle for easy merging.\r\n\r\n  [1]: http://sloth.readthedocs.org/en/latest/examples.html#adding-every-nth-image-to-label-file",
    "91746": "OK, I see that by running inversion's python script first we separate out the train images. We need to make sure we don't annotate test images by mistake.",
    "91747": "James King, correct. The `find imgs/*/* -iname \"*.jpg\"` part of the commands I listed above and on the repository should only add training images to be annotated as only those images with a label are stored in a subdirectory (i.e. the whale's ID).",
    "91804": "Smerity, Thanks for the example. inversion, thanks for the script.\r\n\r\nI have annotated 300 images (Head only), just submitted a pull request.",
    "91807": "Thanks to @James King and @sunil, there are now 630 annotated whale faces. I've updated the extraction script such that it can handle as many JSON annotations as you can throw at it (overwriting any duplicate extractions with the most recent) and also pad to a square instead of the original annotation's dimensions.\r\n\r\nNote: the padding does have a small bug (occasional off by one results in one pixel of white) and if it's off the edge of the original photo, the expanded region is black.",
    "97379": "Smerity,\r\n\r\ndon't know if u are still working on this competition, but I have gotten across some problems during the conversion of data from JSON to ROI.\r\n\r\nError messages whlie I do trainCascadeObjectDetector('detectorFile.xml', positiveInstances, 'negativeFolder'... in MATLAB\r\nError using trainCascadeObjectDetector (line 245)\r\nError reading instance 1 from image 1, bounding box possibly out of image bounds.\r\n\r\nIt seems like the image are out of bound if applying the json format data directly to positiveInstances?\r\nAppreciate if there's any suggestion, thank you!",
    "97390": "How are you guys transforming JSON to matlab compatible ROI? \r\n\r\nWaiting for your reply."
  },
  "source": "meta"
}