{
  "id": 80356,
  "title": "Synthetic Whale Tail Dataset",
  "url": "/competitions/humpback-whale-identification/discussion/80356",
  "author_name": "",
  "post_date": "2019-02-12T22:25:42.359858800Z",
  "votes": 36,
  "comment_count": 33,
  "views": 0,
  "content": "<p>Dear fellow whale observers,</p>\n\n<p>We think that we can all agree that the limiting factor of the dataset is the amount of samples per whale. Therefore we started working on our idea to generate synthetic whale tails. For this we combined our expertise in programming and 3D-graphics to create our own synthetic tail generator. </p>\n\n<p>Since it took quite a while to figure out all the implementation details we had to set the deep learning part aside for the majority of time. We do not think that there is enough time left for us to use this dataset to come up with a winning deep learning algorithm. Thus we hereby share it. The set contains:</p>\n\n<p>100.000 images total (true color and labels) with total size 50GB\n250 whales (tails)\n200 RGB images of each whale with resolution 1024x768\n200 RGB images with whale labels (whale:white, ocean:black, sky:gray)\n200 json files with metadata</p>\n\n<p>Each whale is unique with its own mesh (unique notch, trailing edges, tips, fluke span) and texture. The texture consists of random scratches, barnacle marks and black- and white patches. We also rigged the mesh such that the tail can bend in a natural way. For an idea of what the whales and labels look like see the five examples below.</p>\n\n<p>The data is available via the following link (password=tail):\n<a href=\"https://www.syntheticwhales.nl/s/ZXgQcZKCCks50wH\">Synthetic Whale Tail Dataset</a></p>\n\n<p>Which we will keep alive until the end of the competition. At the moment of writing we are still uploading the set to the cloud which may take at least a full day. It should be done by the weekend.</p>\n\n<p>We appreciate any feedback. If possible we can make adaptations and improve the set!</p>\n\n<p>We wish you all the best!</p>\n\n<p>Erik, Oscar and Tim</p>\n\n<p><img src=\"https://i.ibb.co/Vx91c8n/008885.png\" alt=\"Image whale1\">\n<img src=\"https://i.ibb.co/w7Rs3h7/008885-cs.png\" alt=\"Label whale1\"></p>\n\n<p><img src=\"https://i.ibb.co/fSNzLCp/025359.png\" alt=\"Image whale2\">\n<img src=\"https://i.ibb.co/nCQ4g12/025359-cs.png\" alt=\"Label whale2\"></p>\n\n<p><img src=\"https://i.ibb.co/PhBHV4q/028348.png\" alt=\"Image whale3\">\n<img src=\"https://i.ibb.co/Km38wZs/028348-cs.png\" alt=\"Label whale3\"></p>\n\n<p><img src=\"https://i.ibb.co/WxTphf1/043631.png\" alt=\"Image whale4\">\n<img src=\"https://i.ibb.co/CnBJkbV/043631-cs.png\" alt=\"Label whale4\"></p>\n\n<p><img src=\"https://i.ibb.co/2sbbrFr/047244.png\" alt=\"Image whale5\">\n<img src=\"https://i.ibb.co/xsjv64Y/047244-cs.png\" alt=\"Label whale5\"></p>",
  "messages": [
    {
      "id": "470405",
      "postDate": "02/12/2019 22:25:42",
      "content": "<p>Dear fellow whale observers,</p>\n\n<p>We think that we can all agree that the limiting factor of the dataset is the amount of samples per whale. Therefore we started working on our idea to generate synthetic whale tails. For this we combined our expertise in programming and 3D-graphics to create our own synthetic tail generator. </p>\n\n<p>Since it took quite a while to figure out all the implementation details we had to set the deep learning part aside for the majority of time. We do not think that there is enough time left for us to use this dataset to come up with a winning deep learning algorithm. Thus we hereby share it. The set contains:</p>\n\n<p>100.000 images total (true color and labels) with total size 50GB\n250 whales (tails)\n200 RGB images of each whale with resolution 1024x768\n200 RGB images with whale labels (whale:white, ocean:black, sky:gray)\n200 json files with metadata</p>\n\n<p>Each whale is unique with its own mesh (unique notch, trailing edges, tips, fluke span) and texture. The texture consists of random scratches, barnacle marks and black- and white patches. We also rigged the mesh such that the tail can bend in a natural way. For an idea of what the whales and labels look like see the five examples below.</p>\n\n<p>The data is available via the following link (password=tail):\n<a href=\"https://www.syntheticwhales.nl/s/ZXgQcZKCCks50wH\">Synthetic Whale Tail Dataset</a></p>\n\n<p>Which we will keep alive until the end of the competition. At the moment of writing we are still uploading the set to the cloud which may take at least a full day. It should be done by the weekend.</p>\n\n<p>We appreciate any feedback. If possible we can make adaptations and improve the set!</p>\n\n<p>We wish you all the best!</p>\n\n<p>Erik, Oscar and Tim</p>\n\n<p><img src=\"https://i.ibb.co/Vx91c8n/008885.png\" alt=\"Image whale1\">\n<img src=\"https://i.ibb.co/w7Rs3h7/008885-cs.png\" alt=\"Label whale1\"></p>\n\n<p><img src=\"https://i.ibb.co/fSNzLCp/025359.png\" alt=\"Image whale2\">\n<img src=\"https://i.ibb.co/nCQ4g12/025359-cs.png\" alt=\"Label whale2\"></p>\n\n<p><img src=\"https://i.ibb.co/PhBHV4q/028348.png\" alt=\"Image whale3\">\n<img src=\"https://i.ibb.co/Km38wZs/028348-cs.png\" alt=\"Label whale3\"></p>\n\n<p><img src=\"https://i.ibb.co/WxTphf1/043631.png\" alt=\"Image whale4\">\n<img src=\"https://i.ibb.co/CnBJkbV/043631-cs.png\" alt=\"Label whale4\"></p>\n\n<p><img src=\"https://i.ibb.co/2sbbrFr/047244.png\" alt=\"Image whale5\">\n<img src=\"https://i.ibb.co/xsjv64Y/047244-cs.png\" alt=\"Label whale5\"></p>",
      "rawMarkdown": "Dear fellow whale observers,\n\nWe think that we can all agree that the limiting factor of the dataset is the amount of samples per whale. Therefore we started working on our idea to generate synthetic whale tails. For this we combined our expertise in programming and 3D-graphics to create our own synthetic tail generator. \n\nSince it took quite a while to figure out all the implementation details we had to set the deep learning part aside for the majority of time. We do not think that there is enough time left for us to use this dataset to come up with a winning deep learning algorithm. Thus we hereby share it. The set contains:\n\n100.000 images total (true color and labels) with total size 50GB\n250 whales (tails)\n200 RGB images of each whale with resolution 1024x768\n200 RGB images with whale labels (whale:white, ocean:black, sky:gray)\n200 json files with metadata\n\nEach whale is unique with its own mesh (unique notch, trailing edges, tips, fluke span) and texture. The texture consists of random scratches, barnacle marks and black- and white patches. We also rigged the mesh such that the tail can bend in a natural way. For an idea of what the whales and labels look like see the five examples below.\n\nThe data is available via the following link (password=tail):\n[Synthetic Whale Tail Dataset][1]\n\nWhich we will keep alive until the end of the competition. At the moment of writing we are still uploading the set to the cloud which may take at least a full day. It should be done by the weekend.\n\nWe appreciate any feedback. If possible we can make adaptations and improve the set!\n\nWe wish you all the best!\n\nErik, Oscar and Tim\n\n\n\n![Image whale1][2]\n![Label whale1][3]\n\n![Image whale2][4]\n![Label whale2][5]\n\n![Image whale3][6]\n![Label whale3][7]\n\n![Image whale4][8]\n![Label whale4][9]\n\n![Image whale5][10]\n![Label whale5][11]\n\n\n  [1]: https://www.syntheticwhales.nl/s/ZXgQcZKCCks50wH\n  [2]: https://i.ibb.co/Vx91c8n/008885.png\n  [3]: https://i.ibb.co/w7Rs3h7/008885-cs.png\n  [4]: https://i.ibb.co/fSNzLCp/025359.png\n  [5]: https://i.ibb.co/nCQ4g12/025359-cs.png\n  [6]: https://i.ibb.co/PhBHV4q/028348.png\n  [7]: https://i.ibb.co/Km38wZs/028348-cs.png\n  [8]: https://i.ibb.co/WxTphf1/043631.png\n  [9]: https://i.ibb.co/CnBJkbV/043631-cs.png\n  [10]: https://i.ibb.co/2sbbrFr/047244.png\n  [11]: https://i.ibb.co/xsjv64Y/047244-cs.png",
      "votes": null
    },
    {
      "id": "470486",
      "postDate": "02/13/2019 04:05:39",
      "content": "<p>Wow! Really impressive stuff!</p>",
      "rawMarkdown": "Wow! Really impressive stuff!",
      "votes": null
    },
    {
      "id": "470487",
      "postDate": "02/13/2019 04:07:42",
      "content": "<p>thanks a lot!\nvery superb work!</p>\n\n<p>this is very valuable for domain adaptation work!</p>",
      "rawMarkdown": "thanks a lot!\nvery superb work!\n\nthis is very valuable for domain adaptation work!",
      "votes": null
    },
    {
      "id": "470489",
      "postDate": "02/13/2019 04:19:51",
      "content": "<p>Thanks wise, could I ask that do these 250 whales belong to 5004 classes?</p>",
      "rawMarkdown": "Thanks wise, could I ask that do these 250 whales belong to 5004 classes?",
      "votes": null
    },
    {
      "id": "470490",
      "postDate": "02/13/2019 04:21:19",
      "content": "<p>wow, I see. \"The texture consists of random scratches, barnacle marks and black- and white patches. \"</p>",
      "rawMarkdown": "wow, I see. \"The texture consists of random scratches, barnacle marks and black- and white patches. \"",
      "votes": null
    },
    {
      "id": "470519",
      "postDate": "02/13/2019 05:37:11",
      "content": "<p>If every generated whale is a class then this set has 250 classes :)</p>",
      "rawMarkdown": "If every generated whale is a class then this set has 250 classes :)",
      "votes": null
    },
    {
      "id": "472707",
      "postDate": "02/16/2019 14:15:52",
      "content": "<p>thanks a lot! \nbut there are almost 5000 kinds of whale in the train dataset, your dataset just includes 250 classes.  how you pick up those 250 classes ?  </p>",
      "rawMarkdown": "thanks a lot! \nbut there are almost 5000 kinds of whale in the train dataset, your dataset just includes 250 classes.  how you pick up those 250 classes ?",
      "votes": null
    },
    {
      "id": "472789",
      "postDate": "02/16/2019 17:04:20",
      "content": "<p>Couple of questions I had about the dataset.</p>\n\n<ol>\n<li>By clicking on the \"Download map\" button at the top level directory, the entire dataset will be downloaded?</li>\n<li>Was whale249 successfully uploaded? The folder size only has 900KB worth of data whereas the other whales have ~200 MB of data.</li>\n<li>Can you help explain some of the .json data, it seems that bounding_box, cuboid, and projected_cuboid might have been corrupted? (bounding box is always [0,0], cuboid and projected cuboid always the same value for all points)</li>\n<li>You mentioned about keeping the dataset up only until the competition ends. Are you planning on uploading it as a kaggle dataset?</li>\n<li>What sort of License if any are you releasing this under? If i did some derivative work on this data, could I release that (including the original files) as a kaggle dataset?</li>\n</ol>\n\n<p>Once again, thanks for doing all this hard work. I really think it will be helpful to the community!</p>",
      "rawMarkdown": "Couple of questions I had about the dataset.\n\n1. By clicking on the \"Download map\" button at the top level directory, the entire dataset will be downloaded?\n2. Was whale249 successfully uploaded? The folder size only has 900KB worth of data whereas the other whales have ~200 MB of data.\n3. Can you help explain some of the .json data, it seems that bounding_box, cuboid, and projected_cuboid might have been corrupted? (bounding box is always [0,0], cuboid and projected cuboid always the same value for all points)\n4. You mentioned about keeping the dataset up only until the competition ends. Are you planning on uploading it as a kaggle dataset?\n5. What sort of License if any are you releasing this under? If i did some derivative work on this data, could I release that (including the original files) as a kaggle dataset?\n\nOnce again, thanks for doing all this hard work. I really think it will be helpful to the community!",
      "votes": null
    },
    {
      "id": "472874",
      "postDate": "02/16/2019 20:05:36",
      "content": "<p>Thanks for your feedback! I will respond to your questions to the best of my knowledge:</p>\n\n<ol>\n<li>Yes that is right. Map means folder (Dutch) so \"Download folder\". It is also possible to selectively download the folders of your choice by checkmarking them and then on the right click download.</li>\n<li>No I checked and you correctly mention that there are indeed 249 whales! We stopped one iteration too early.</li>\n<li>We forgot to credit Nvidia for their code on domain randomization:  <a href=\"https://github.com/NVIDIA/Dataset_Synthesizer\">https://github.com/NVIDIA/Dataset_Synthesizer</a>. Along with this wonderful paper: <a href=\"https://arxiv.org/abs/1804.06516\">https://arxiv.org/abs/1804.06516</a>. There may be a bug in their code or in the way we used it. We will check. The top left and lower right can for now be obtained by taking the bounds of the label images.</li>\n<li>We haven't thought about that, but we can for sure :-)</li>\n<li>Also not thought about, but either CC-0 or CC-BY.</li>\n</ol>",
      "rawMarkdown": "Thanks for your feedback! I will respond to your questions to the best of my knowledge:\n\n 1. Yes that is right. Map means folder (Dutch) so \"Download folder\". It is also possible to selectively download the folders of your choice by checkmarking them and then on the right click download.\n 2. No I checked and you correctly mention that there are indeed 249 whales! We stopped one iteration too early.\n 3. We forgot to credit Nvidia for their code on domain randomization:  https://github.com/NVIDIA/Dataset_Synthesizer. Along with this wonderful paper: https://arxiv.org/abs/1804.06516. There may be a bug in their code or in the way we used it. We will check. The top left and lower right can for now be obtained by taking the bounds of the label images.\n 4. We haven't thought about that, but we can for sure :-)\n 5. Also not thought about, but either CC-0 or CC-BY.",
      "votes": null
    },
    {
      "id": "472885",
      "postDate": "02/16/2019 20:37:05",
      "content": "<p>These tails are computer generated and fully independent from the set of real world images. If data transfer was no bottleneck, we could just as easily have created 10000 classes.</p>",
      "rawMarkdown": "These tails are computer generated and fully independent from the set of real world images. If data transfer was no bottleneck, we could just as easily have created 10000 classes.",
      "votes": null
    },
    {
      "id": "472893",
      "postDate": "02/16/2019 21:08:36",
      "content": "<p>Very impressive work. I will use it immediately.</p>",
      "rawMarkdown": "Very impressive work. I will use it immediately.",
      "votes": null
    },
    {
      "id": "472895",
      "postDate": "02/16/2019 21:17:53",
      "content": "<p>Be careful. Keep in mind that far not all models generalise well from synthetic datasets to the real data. From what I see, these images have vivid texture and simple imperfections, which may not be true for the competition data.</p>",
      "rawMarkdown": "Be careful. Keep in mind that far not all models generalise well from synthetic datasets to the real data. From what I see, these images have vivid texture and simple imperfections, which may not be true for the competition data.",
      "votes": null
    },
    {
      "id": "472899",
      "postDate": "02/16/2019 21:24:21",
      "content": "<p>Thank you for your advice.\nOK, I will download it and look at the data for the time being.</p>",
      "rawMarkdown": "Thank you for your advice.\nOK, I will download it and look at the data for the time being.",
      "votes": null
    },
    {
      "id": "473250",
      "postDate": "02/17/2019 16:35:53",
      "content": "<p>One note for everyone trying to get the dataset: </p>\n\n<p>I've been unable to download the entire dataset as one zip file (seems like the download speed is capped at 2.5-3MB/s per connection and something times out before all ~50 GB is completed). What you can do instead is select chunks of folders (say 25 or 50 whale folders) in the UI by holding ctrl and clicking on the item then select 'Downloaden'. (after you select a chunk, you can click 'Deselecteren' to deselect all folders and start on the next chunk)</p>\n\n<p>You can download 4 chunks at a time (each one has a ~3MB/s download speed). It's a little bit more manual but at least it seems to be reliable for downloading smaller (5-10 GB chunks).</p>",
      "rawMarkdown": "One note for everyone trying to get the dataset: \n\nI've been unable to download the entire dataset as one zip file (seems like the download speed is capped at 2.5-3MB/s per connection and something times out before all ~50 GB is completed). What you can do instead is select chunks of folders (say 25 or 50 whale folders) in the UI by holding ctrl and clicking on the item then select 'Downloaden'. (after you select a chunk, you can click 'Deselecteren' to deselect all folders and start on the next chunk)\n\nYou can download 4 chunks at a time (each one has a ~3MB/s download speed). It's a little bit more manual but at least it seems to be reliable for downloading smaller (5-10 GB chunks).",
      "votes": null
    },
    {
      "id": "473708",
      "postDate": "02/18/2019 12:10:07",
      "content": "<p>It would be pretty helpful if anyone can help to upload via google drive : )</p>",
      "rawMarkdown": "It would be pretty helpful if anyone can help to upload via google drive : )",
      "votes": null
    },
    {
      "id": "473817",
      "postDate": "02/18/2019 15:00:23",
      "content": "<p>I started looking through the dataset and noticed some additional anomalies.</p>\n\n<p>Images are numbered <code>000000.*</code> to <code>049999</code></p>\n\n<p>I think what is expected (and is true for some whale folders) is for each whale directory to have exactly 600 files(200 metadata <code>.json</code>, 200 mask <code>.cs.png</code> and 200 images <code>.png</code>) <code>whale0</code> is a good example which has <code>000000{.json|cs.png|png}</code> to <code>000199{json|cs.png|png}</code>.</p>\n\n<p>Some folders (<code>whale15</code> through <code>whale200</code> -- excepting <code>whale195</code> (see below)) have 603 files. It seems as though 600 of these files are the expected 200 instances with 3 images each in the natural ordering (e.g. <code>whale15</code> has files <code>003000*</code> to <code>003199*</code>) in addition it has one extra set of a higher numbered image. This higher numbered sequence seems to increase by one for each folder: (e.g. <code>whale15</code> has <code>049815*</code>, <code>whale16</code> has <code>049816*</code>, etc)) I've verified at least for several instances the higher numbered image is not the same whale as the folder, and the higher number images do not appear to be the same whale (<code>049815*</code> is not the same whale as <code>049816</code>).</p>\n\n<p>Finally it seems that <code>whale195/039120.json</code> is missing (the mask and image files are there) and files <code>049801*</code> through <code>049814*</code> are all missing. I would presume following the above pattern these would have appeared in <code>whale1</code> through <code>whale14</code> and maybe were caught and deleted?</p>",
      "rawMarkdown": "I started looking through the dataset and noticed some additional anomalies.\n\n\nImages are numbered `000000.*` to `049999`\n\n\nI think what is expected (and is true for some whale folders) is for each whale directory to have exactly 600 files(200 metadata `.json`, 200 mask `.cs.png` and 200 images `.png`) `whale0` is a good example which has `000000{.json|cs.png|png}` to `000199{json|cs.png|png}`.\n\n\nSome folders (`whale15` through `whale200` -- excepting `whale195` (see below)) have 603 files. It seems as though 600 of these files are the expected 200 instances with 3 images each in the natural ordering (e.g. `whale15` has files `003000* ` to `003199*`) in addition it has one extra set of a higher numbered image. This higher numbered sequence seems to increase by one for each folder: (e.g. `whale15` has `049815*`, `whale16` has `049816*`, etc)) I've verified at least for several instances the higher numbered image is not the same whale as the folder, and the higher number images do not appear to be the same whale (`049815*` is not the same whale as `049816`).\n\n\nFinally it seems that `whale195/039120.json` is missing (the mask and image files are there) and files `049801*` through `049814*` are all missing. I would presume following the above pattern these would have appeared in `whale1` through `whale14` and maybe were caught and deleted?",
      "votes": null
    },
    {
      "id": "475530",
      "postDate": "02/20/2019 21:28:39",
      "content": "<p>Nice! I'm happy that someone made this :) Any details of implementation?</p>\n\n<p>P.S. However the main issue here as I see it - the water. It doesn't look right. I think that would be better to just place the tails on top of the real sea photos.</p>",
      "rawMarkdown": "Nice! I'm happy that someone made this :) Any details of implementation?\n\nP.S. However the main issue here as I see it - the water. It doesn't look right. I think that would be better to just place the tails on top of the real sea photos.",
      "votes": null
    },
    {
      "id": "475531",
      "postDate": "02/20/2019 21:29:22",
      "content": "<p>Amazing! Thank you for sharing! I spent a lot of time for study how work blander and this was not a success. But you make it perfect.</p>",
      "rawMarkdown": "Amazing! Thank you for sharing! I spent a lot of time for study how work blander and this was not a success. But you make it perfect.",
      "votes": null
    },
    {
      "id": "475536",
      "postDate": "02/20/2019 21:43:22",
      "content": "<p>You don’t believe there are this many whales swimming in the waters of Morrowind?</p>",
      "rawMarkdown": "You don’t believe there are this many whales swimming in the waters of Morrowind?",
      "votes": null
    },
    {
      "id": "475787",
      "postDate": "02/21/2019 07:46:42",
      "content": "<p>Any feedback whether this dataset improved the performance?</p>",
      "rawMarkdown": "Any feedback whether this dataset improved the performance?",
      "votes": null
    },
    {
      "id": "476138",
      "postDate": "02/21/2019 16:49:59",
      "content": "<p>Yes we found the bug and are working on a version 2. The new set will be online by tomorrow evening (UTC+01.00)  </p>\n\n<p>We have looked at all the feedback and made the following changes:\n - Fixed bug with the 2d and 3d bounding boxes, this now works properly.\n - Added domain randomization: light (apart from the changes in time of day we added small random light sources\n - Added domain randomization: objects with random textures location and orientation. (This includes random geometries as well as a boat and eagle mesh.\n - Orientation of the whale tails will be more oriented to the camera. The current set is too random in our opinion.\n - More random water (possibly from real world images)</p>\n\n<p>Furthermore this time we will upload the whole set in parts on separate Google Drive locations. The Transip was not as stable as we had hoped.</p>",
      "rawMarkdown": "Yes we found the bug and are working on a version 2. The new set will be online by tomorrow evening (UTC+01.00)  \n\nWe have looked at all the feedback and made the following changes:\n - Fixed bug with the 2d and 3d bounding boxes, this now works properly.\n - Added domain randomization: light (apart from the changes in time of day we added small random light sources\n - Added domain randomization: objects with random textures location and orientation. (This includes random geometries as well as a boat and eagle mesh.\n - Orientation of the whale tails will be more oriented to the camera. The current set is too random in our opinion.\n - More random water (possibly from real world images)\n\nFurthermore this time we will upload the whole set in parts on separate Google Drive locations. The Transip was not as stable as we had hoped.",
      "votes": null
    },
    {
      "id": "476141",
      "postDate": "02/21/2019 16:50:29",
      "content": "<p>Haha love the Morrowind ref &lt;3</p>",
      "rawMarkdown": "Haha love the Morrowind ref &lt;3",
      "votes": null
    },
    {
      "id": "476144",
      "postDate": "02/21/2019 16:52:01",
      "content": "<p>Working on a version 2, see comment below :-). This time we will upload everything in parts to Drive.</p>",
      "rawMarkdown": "Working on a version 2, see comment below :-). This time we will upload everything in parts to Drive.",
      "votes": null
    },
    {
      "id": "476145",
      "postDate": "02/21/2019 16:53:06",
      "content": "<p>Thanks for your time to look into these details! We assume that this has something to do with the file writing sequencing. For version 2 we will try to fix this!</p>",
      "rawMarkdown": "Thanks for your time to look into these details! We assume that this has something to do with the file writing sequencing. For version 2 we will try to fix this!",
      "votes": null
    },
    {
      "id": "476148",
      "postDate": "02/21/2019 16:55:09",
      "content": "<p>For version 2 (see my comment on Hafizur's post) we will upload the whole set to Google Drive.</p>",
      "rawMarkdown": "For version 2 (see my comment on Hafizur's post) we will upload the whole set to Google Drive.",
      "votes": null
    },
    {
      "id": "476169",
      "postDate": "02/21/2019 17:25:51",
      "content": "<p>So could you share details of implementation? Which software was used, approaches to texture generation, etc. That would be really helpful. Thanks.</p>",
      "rawMarkdown": "So could you share details of implementation? Which software was used, approaches to texture generation, etc. That would be really helpful. Thanks.",
      "votes": null
    },
    {
      "id": "476994",
      "postDate": "02/23/2019 16:32:33",
      "content": "<p>When you planning update your dataset?</p>",
      "rawMarkdown": "When you planning update your dataset?",
      "votes": null
    },
    {
      "id": "477113",
      "postDate": "02/23/2019 23:11:50",
      "content": "<p>We are uploading the whole set as we speak on Google Drive. Accessible via the following link:\n<a href=\"https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing\">Synthetic Whale Tails V2</a></p>",
      "rawMarkdown": "We are uploading the whole set as we speak on Google Drive. Accessible via the following link:\n[Synthetic Whale Tails V2][1]\n\n  [1]: https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing",
      "votes": null
    },
    {
      "id": "477542",
      "postDate": "02/24/2019 20:08:23",
      "content": "<p>Update: the whole set is up and running.\nAs with the previous version, any feedback is welcome!</p>",
      "rawMarkdown": "Update: the whole set is up and running.\nAs with the previous version, any feedback is welcome!",
      "votes": null
    },
    {
      "id": "477543",
      "postDate": "02/24/2019 20:09:57",
      "content": "<p>Thanks for the feedback!</p>",
      "rawMarkdown": "Thanks for the feedback!",
      "votes": null
    },
    {
      "id": "477842",
      "postDate": "02/25/2019 10:52:18",
      "content": "<p>Version 2 has been uploaded on Google Drive. We tried to incorporate all feedback. Accessible via the following link:\n<a href=\"https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing\">Synthetic Whale Tails V2</a></p>\n\n<p>This set contains:\n- Two sets (With and Without distractors)\n- Distractors are randomly placed geometric objects\n- 200 different whales (so keep in mind that although they have the same label, they are different)\n- 300 images of each whale</p>\n\n<p>Major changes are:\n- We used 2k textures for the ocean\n- Tails are oriented more towards the camera\n- Added extra light sources\n- Added distractors\n- Fixed a few bugs (in particular the bounding boxes)</p>\n\n<p>A quick preview below :-)\n<img src=\"https://i.ibb.co/K0LTrZG/019596.png\" alt=\"enter image description here\">\n<img src=\"https://i.ibb.co/3sCPBV5/019596-cs.png\" alt=\"enter image description here\">\n<img src=\"https://i.ibb.co/N2bnVpG/004214.png\" alt=\"enter image description here\">\n<img src=\"https://i.ibb.co/HgwRpKs/004214-cs.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "Version 2 has been uploaded on Google Drive. We tried to incorporate all feedback. Accessible via the following link:\n[Synthetic Whale Tails V2][1]\n\nThis set contains:\n- Two sets (With and Without distractors)\n- Distractors are randomly placed geometric objects\n- 200 different whales (so keep in mind that although they have the same label, they are different)\n- 300 images of each whale\n\nMajor changes are:\n- We used 2k textures for the ocean\n- Tails are oriented more towards the camera\n- Added extra light sources\n- Added distractors\n- Fixed a few bugs (in particular the bounding boxes)\n\nA quick preview below :-)\n![enter image description here][2]\n![enter image description here][3]\n![enter image description here][4]\n![enter image description here][5]\n\n  [1]: https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing\n  [2]: https://i.ibb.co/K0LTrZG/019596.png\n  [3]: https://i.ibb.co/3sCPBV5/019596-cs.png\n  [4]: https://i.ibb.co/N2bnVpG/004214.png\n  [5]: https://i.ibb.co/HgwRpKs/004214-cs.png",
      "votes": null
    },
    {
      "id": "481951",
      "postDate": "03/02/2019 05:04:20",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "482244",
      "postDate": "03/02/2019 15:19:24",
      "content": "<p>Got around to downloading the new dataset. Here's a couple of tips for anyone else attempts to grab it:\n- Google's Download Folder option never worked for me, it would start only about half of the files.\n- each whale.rar contains all of the images, but extracting it will put the images into the current directory (at least with linux's <code>unrar</code> tool). You need to do something like: <code>unrar e whale1.rar -d ./whale1/</code> to extract the contents into a folder called whale1.\n- Magic incantation to do this for all .rars in the current working directory: <code>find . -name \"*.rar\" -exec sh -c 'unrar e $1 -d \"${1%.*}/\"' _ {} \\;</code></p>",
      "rawMarkdown": "Got around to downloading the new dataset. Here's a couple of tips for anyone else attempts to grab it:\n- Google's Download Folder option never worked for me, it would start only about half of the files.\n- each whale.rar contains all of the images, but extracting it will put the images into the current directory (at least with linux's `unrar` tool). You need to do something like: `unrar e whale1.rar -d ./whale1/` to extract the contents into a folder called whale1.\n- Magic incantation to do this for all .rars in the current working directory: `find . -name \"*.rar\" -exec sh -c 'unrar e $1 -d \"${1%.*}/\"' _ {} \\;`",
      "votes": null
    },
    {
      "id": "482511",
      "postDate": "03/03/2019 05:53:08",
      "content": "<p>how about using this and combine with deepmind paper to general novel view from single input image?</p>\n\n<p>train = synthetic image</p>\n\n<p>output = generate-new-view(real image or images)</p>\n\n<p>Original Paper: Neural scene representation and rendering (Eslami, et al., 2018)</p>\n\n<p><a href=\"https://deepmind.com/blog/neural-scene-representation-and-rendering\">https://deepmind.com/blog/neural-scene-representation-and-rendering</a>\n<a href=\"https://www.youtube.com/watch?v=G-kWNQJ4idw\">https://www.youtube.com/watch?v=G-kWNQJ4idw</a>\n<a href=\"https://github.com/musyoku/chainer-gqn\">https://github.com/musyoku/chainer-gqn</a></p>",
      "rawMarkdown": "how about using this and combine with deepmind paper to general novel view from single input image?\n\ntrain = synthetic image\n\noutput = generate-new-view(real image or images)\n\nOriginal Paper: Neural scene representation and rendering (Eslami, et al., 2018)\n\nhttps://deepmind.com/blog/neural-scene-representation-and-rendering\nhttps://www.youtube.com/watch?v=G-kWNQJ4idw\nhttps://github.com/musyoku/chainer-gqn",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 470486,
      "author_name": "oewyn000",
      "author_url": "",
      "post_date": "02/13/2019 04:05:39",
      "content": "<p>Wow! Really impressive stuff!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 470487,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/13/2019 04:07:42",
      "content": "<p>thanks a lot!\nvery superb work!</p>\n\n<p>this is very valuable for domain adaptation work!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 470489,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "02/13/2019 04:19:51",
      "content": "<p>Thanks wise, could I ask that do these 250 whales belong to 5004 classes?</p>",
      "votes": null,
      "replies": [
        {
          "id": 470490,
          "author_name": "yiheng",
          "author_url": "",
          "post_date": "02/13/2019 04:21:19",
          "content": "<p>wow, I see. \"The texture consists of random scratches, barnacle marks and black- and white patches. \"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 470519,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/13/2019 05:37:11",
          "content": "<p>If every generated whale is a class then this set has 250 classes :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472707,
      "author_name": "newfangzl",
      "author_url": "",
      "post_date": "02/16/2019 14:15:52",
      "content": "<p>thanks a lot! \nbut there are almost 5000 kinds of whale in the train dataset, your dataset just includes 250 classes.  how you pick up those 250 classes ?  </p>",
      "votes": null,
      "replies": [
        {
          "id": 472885,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/16/2019 20:37:05",
          "content": "<p>These tails are computer generated and fully independent from the set of real world images. If data transfer was no bottleneck, we could just as easily have created 10000 classes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472789,
      "author_name": "oewyn000",
      "author_url": "",
      "post_date": "02/16/2019 17:04:20",
      "content": "<p>Couple of questions I had about the dataset.</p>\n\n<ol>\n<li>By clicking on the \"Download map\" button at the top level directory, the entire dataset will be downloaded?</li>\n<li>Was whale249 successfully uploaded? The folder size only has 900KB worth of data whereas the other whales have ~200 MB of data.</li>\n<li>Can you help explain some of the .json data, it seems that bounding_box, cuboid, and projected_cuboid might have been corrupted? (bounding box is always [0,0], cuboid and projected cuboid always the same value for all points)</li>\n<li>You mentioned about keeping the dataset up only until the competition ends. Are you planning on uploading it as a kaggle dataset?</li>\n<li>What sort of License if any are you releasing this under? If i did some derivative work on this data, could I release that (including the original files) as a kaggle dataset?</li>\n</ol>\n\n<p>Once again, thanks for doing all this hard work. I really think it will be helpful to the community!</p>",
      "votes": null,
      "replies": [
        {
          "id": 472874,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/16/2019 20:05:36",
          "content": "<p>Thanks for your feedback! I will respond to your questions to the best of my knowledge:</p>\n\n<ol>\n<li>Yes that is right. Map means folder (Dutch) so \"Download folder\". It is also possible to selectively download the folders of your choice by checkmarking them and then on the right click download.</li>\n<li>No I checked and you correctly mention that there are indeed 249 whales! We stopped one iteration too early.</li>\n<li>We forgot to credit Nvidia for their code on domain randomization:  <a href=\"https://github.com/NVIDIA/Dataset_Synthesizer\">https://github.com/NVIDIA/Dataset_Synthesizer</a>. Along with this wonderful paper: <a href=\"https://arxiv.org/abs/1804.06516\">https://arxiv.org/abs/1804.06516</a>. There may be a bug in their code or in the way we used it. We will check. The top left and lower right can for now be obtained by taking the bounds of the label images.</li>\n<li>We haven't thought about that, but we can for sure :-)</li>\n<li>Also not thought about, but either CC-0 or CC-BY.</li>\n</ol>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472893,
      "author_name": "kerukun",
      "author_url": "",
      "post_date": "02/16/2019 21:08:36",
      "content": "<p>Very impressive work. I will use it immediately.</p>",
      "votes": null,
      "replies": [
        {
          "id": 472895,
          "author_name": "vshakhray",
          "author_url": "",
          "post_date": "02/16/2019 21:17:53",
          "content": "<p>Be careful. Keep in mind that far not all models generalise well from synthetic datasets to the real data. From what I see, these images have vivid texture and simple imperfections, which may not be true for the competition data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 472899,
          "author_name": "kerukun",
          "author_url": "",
          "post_date": "02/16/2019 21:24:21",
          "content": "<p>Thank you for your advice.\nOK, I will download it and look at the data for the time being.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477543,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/24/2019 20:09:57",
          "content": "<p>Thanks for the feedback!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 473250,
      "author_name": "oewyn000",
      "author_url": "",
      "post_date": "02/17/2019 16:35:53",
      "content": "<p>One note for everyone trying to get the dataset: </p>\n\n<p>I've been unable to download the entire dataset as one zip file (seems like the download speed is capped at 2.5-3MB/s per connection and something times out before all ~50 GB is completed). What you can do instead is select chunks of folders (say 25 or 50 whale folders) in the UI by holding ctrl and clicking on the item then select 'Downloaden'. (after you select a chunk, you can click 'Deselecteren' to deselect all folders and start on the next chunk)</p>\n\n<p>You can download 4 chunks at a time (each one has a ~3MB/s download speed). It's a little bit more manual but at least it seems to be reliable for downloading smaller (5-10 GB chunks).</p>",
      "votes": null,
      "replies": [
        {
          "id": 476148,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/21/2019 16:55:09",
          "content": "<p>For version 2 (see my comment on Hafizur's post) we will upload the whole set to Google Drive.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 473708,
      "author_name": "yiheng",
      "author_url": "",
      "post_date": "02/18/2019 12:10:07",
      "content": "<p>It would be pretty helpful if anyone can help to upload via google drive : )</p>",
      "votes": null,
      "replies": [
        {
          "id": 476144,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/21/2019 16:52:01",
          "content": "<p>Working on a version 2, see comment below :-). This time we will upload everything in parts to Drive.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 473817,
      "author_name": "oewyn000",
      "author_url": "",
      "post_date": "02/18/2019 15:00:23",
      "content": "<p>I started looking through the dataset and noticed some additional anomalies.</p>\n\n<p>Images are numbered <code>000000.*</code> to <code>049999</code></p>\n\n<p>I think what is expected (and is true for some whale folders) is for each whale directory to have exactly 600 files(200 metadata <code>.json</code>, 200 mask <code>.cs.png</code> and 200 images <code>.png</code>) <code>whale0</code> is a good example which has <code>000000{.json|cs.png|png}</code> to <code>000199{json|cs.png|png}</code>.</p>\n\n<p>Some folders (<code>whale15</code> through <code>whale200</code> -- excepting <code>whale195</code> (see below)) have 603 files. It seems as though 600 of these files are the expected 200 instances with 3 images each in the natural ordering (e.g. <code>whale15</code> has files <code>003000*</code> to <code>003199*</code>) in addition it has one extra set of a higher numbered image. This higher numbered sequence seems to increase by one for each folder: (e.g. <code>whale15</code> has <code>049815*</code>, <code>whale16</code> has <code>049816*</code>, etc)) I've verified at least for several instances the higher numbered image is not the same whale as the folder, and the higher number images do not appear to be the same whale (<code>049815*</code> is not the same whale as <code>049816</code>).</p>\n\n<p>Finally it seems that <code>whale195/039120.json</code> is missing (the mask and image files are there) and files <code>049801*</code> through <code>049814*</code> are all missing. I would presume following the above pattern these would have appeared in <code>whale1</code> through <code>whale14</code> and maybe were caught and deleted?</p>",
      "votes": null,
      "replies": [
        {
          "id": 476145,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/21/2019 16:53:06",
          "content": "<p>Thanks for your time to look into these details! We assume that this has something to do with the file writing sequencing. For version 2 we will try to fix this!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 475530,
      "author_name": "dene33",
      "author_url": "",
      "post_date": "02/20/2019 21:28:39",
      "content": "<p>Nice! I'm happy that someone made this :) Any details of implementation?</p>\n\n<p>P.S. However the main issue here as I see it - the water. It doesn't look right. I think that would be better to just place the tails on top of the real sea photos.</p>",
      "votes": null,
      "replies": [
        {
          "id": 475536,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "02/20/2019 21:43:22",
          "content": "<p>You don’t believe there are this many whales swimming in the waters of Morrowind?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 476141,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/21/2019 16:50:29",
          "content": "<p>Haha love the Morrowind ref &lt;3</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 476169,
          "author_name": "dene33",
          "author_url": "",
          "post_date": "02/21/2019 17:25:51",
          "content": "<p>So could you share details of implementation? Which software was used, approaches to texture generation, etc. That would be really helpful. Thanks.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 475531,
      "author_name": "i7p9h9",
      "author_url": "",
      "post_date": "02/20/2019 21:29:22",
      "content": "<p>Amazing! Thank you for sharing! I spent a lot of time for study how work blander and this was not a success. But you make it perfect.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 475787,
      "author_name": "hafizurrahman",
      "author_url": "",
      "post_date": "02/21/2019 07:46:42",
      "content": "<p>Any feedback whether this dataset improved the performance?</p>",
      "votes": null,
      "replies": [
        {
          "id": 476138,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/21/2019 16:49:59",
          "content": "<p>Yes we found the bug and are working on a version 2. The new set will be online by tomorrow evening (UTC+01.00)  </p>\n\n<p>We have looked at all the feedback and made the following changes:\n - Fixed bug with the 2d and 3d bounding boxes, this now works properly.\n - Added domain randomization: light (apart from the changes in time of day we added small random light sources\n - Added domain randomization: objects with random textures location and orientation. (This includes random geometries as well as a boat and eagle mesh.\n - Orientation of the whale tails will be more oriented to the camera. The current set is too random in our opinion.\n - More random water (possibly from real world images)</p>\n\n<p>Furthermore this time we will upload the whole set in parts on separate Google Drive locations. The Transip was not as stable as we had hoped.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 476994,
          "author_name": "i7p9h9",
          "author_url": "",
          "post_date": "02/23/2019 16:32:33",
          "content": "<p>When you planning update your dataset?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477113,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/23/2019 23:11:50",
          "content": "<p>We are uploading the whole set as we speak on Google Drive. Accessible via the following link:\n<a href=\"https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing\">Synthetic Whale Tails V2</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477542,
          "author_name": "erikoudejans",
          "author_url": "",
          "post_date": "02/24/2019 20:08:23",
          "content": "<p>Update: the whole set is up and running.\nAs with the previous version, any feedback is welcome!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 477842,
      "author_name": "erikoudejans",
      "author_url": "",
      "post_date": "02/25/2019 10:52:18",
      "content": "<p>Version 2 has been uploaded on Google Drive. We tried to incorporate all feedback. Accessible via the following link:\n<a href=\"https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing\">Synthetic Whale Tails V2</a></p>\n\n<p>This set contains:\n- Two sets (With and Without distractors)\n- Distractors are randomly placed geometric objects\n- 200 different whales (so keep in mind that although they have the same label, they are different)\n- 300 images of each whale</p>\n\n<p>Major changes are:\n- We used 2k textures for the ocean\n- Tails are oriented more towards the camera\n- Added extra light sources\n- Added distractors\n- Fixed a few bugs (in particular the bounding boxes)</p>\n\n<p>A quick preview below :-)\n<img src=\"https://i.ibb.co/K0LTrZG/019596.png\" alt=\"enter image description here\">\n<img src=\"https://i.ibb.co/3sCPBV5/019596-cs.png\" alt=\"enter image description here\">\n<img src=\"https://i.ibb.co/N2bnVpG/004214.png\" alt=\"enter image description here\">\n<img src=\"https://i.ibb.co/HgwRpKs/004214-cs.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 482244,
          "author_name": "oewyn000",
          "author_url": "",
          "post_date": "03/02/2019 15:19:24",
          "content": "<p>Got around to downloading the new dataset. Here's a couple of tips for anyone else attempts to grab it:\n- Google's Download Folder option never worked for me, it would start only about half of the files.\n- each whale.rar contains all of the images, but extracting it will put the images into the current directory (at least with linux's <code>unrar</code> tool). You need to do something like: <code>unrar e whale1.rar -d ./whale1/</code> to extract the contents into a folder called whale1.\n- Magic incantation to do this for all .rars in the current working directory: <code>find . -name \"*.rar\" -exec sh -c 'unrar e $1 -d \"${1%.*}/\"' _ {} \\;</code></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 481951,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/02/2019 05:04:20",
      "content": "",
      "votes": null,
      "replies": []
    },
    {
      "id": 482511,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "03/03/2019 05:53:08",
      "content": "<p>how about using this and combine with deepmind paper to general novel view from single input image?</p>\n\n<p>train = synthetic image</p>\n\n<p>output = generate-new-view(real image or images)</p>\n\n<p>Original Paper: Neural scene representation and rendering (Eslami, et al., 2018)</p>\n\n<p><a href=\"https://deepmind.com/blog/neural-scene-representation-and-rendering\">https://deepmind.com/blog/neural-scene-representation-and-rendering</a>\n<a href=\"https://www.youtube.com/watch?v=G-kWNQJ4idw\">https://www.youtube.com/watch?v=G-kWNQJ4idw</a>\n<a href=\"https://github.com/musyoku/chainer-gqn\">https://github.com/musyoku/chainer-gqn</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "470405": "Dear fellow whale observers,\n\nWe think that we can all agree that the limiting factor of the dataset is the amount of samples per whale. Therefore we started working on our idea to generate synthetic whale tails. For this we combined our expertise in programming and 3D-graphics to create our own synthetic tail generator. \n\nSince it took quite a while to figure out all the implementation details we had to set the deep learning part aside for the majority of time. We do not think that there is enough time left for us to use this dataset to come up with a winning deep learning algorithm. Thus we hereby share it. The set contains:\n\n100.000 images total (true color and labels) with total size 50GB\n250 whales (tails)\n200 RGB images of each whale with resolution 1024x768\n200 RGB images with whale labels (whale:white, ocean:black, sky:gray)\n200 json files with metadata\n\nEach whale is unique with its own mesh (unique notch, trailing edges, tips, fluke span) and texture. The texture consists of random scratches, barnacle marks and black- and white patches. We also rigged the mesh such that the tail can bend in a natural way. For an idea of what the whales and labels look like see the five examples below.\n\nThe data is available via the following link (password=tail):\n[Synthetic Whale Tail Dataset][1]\n\nWhich we will keep alive until the end of the competition. At the moment of writing we are still uploading the set to the cloud which may take at least a full day. It should be done by the weekend.\n\nWe appreciate any feedback. If possible we can make adaptations and improve the set!\n\nWe wish you all the best!\n\nErik, Oscar and Tim\n\n\n\n![Image whale1][2]\n![Label whale1][3]\n\n![Image whale2][4]\n![Label whale2][5]\n\n![Image whale3][6]\n![Label whale3][7]\n\n![Image whale4][8]\n![Label whale4][9]\n\n![Image whale5][10]\n![Label whale5][11]\n\n\n  [1]: https://www.syntheticwhales.nl/s/ZXgQcZKCCks50wH\n  [2]: https://i.ibb.co/Vx91c8n/008885.png\n  [3]: https://i.ibb.co/w7Rs3h7/008885-cs.png\n  [4]: https://i.ibb.co/fSNzLCp/025359.png\n  [5]: https://i.ibb.co/nCQ4g12/025359-cs.png\n  [6]: https://i.ibb.co/PhBHV4q/028348.png\n  [7]: https://i.ibb.co/Km38wZs/028348-cs.png\n  [8]: https://i.ibb.co/WxTphf1/043631.png\n  [9]: https://i.ibb.co/CnBJkbV/043631-cs.png\n  [10]: https://i.ibb.co/2sbbrFr/047244.png\n  [11]: https://i.ibb.co/xsjv64Y/047244-cs.png",
    "470486": "Wow! Really impressive stuff!",
    "470487": "thanks a lot!\nvery superb work!\n\nthis is very valuable for domain adaptation work!",
    "470489": "Thanks wise, could I ask that do these 250 whales belong to 5004 classes?",
    "470490": "wow, I see. \"The texture consists of random scratches, barnacle marks and black- and white patches. \"",
    "470519": "If every generated whale is a class then this set has 250 classes :)",
    "472707": "thanks a lot! \nbut there are almost 5000 kinds of whale in the train dataset, your dataset just includes 250 classes.  how you pick up those 250 classes ?",
    "472789": "Couple of questions I had about the dataset.\n\n1. By clicking on the \"Download map\" button at the top level directory, the entire dataset will be downloaded?\n2. Was whale249 successfully uploaded? The folder size only has 900KB worth of data whereas the other whales have ~200 MB of data.\n3. Can you help explain some of the .json data, it seems that bounding_box, cuboid, and projected_cuboid might have been corrupted? (bounding box is always [0,0], cuboid and projected cuboid always the same value for all points)\n4. You mentioned about keeping the dataset up only until the competition ends. Are you planning on uploading it as a kaggle dataset?\n5. What sort of License if any are you releasing this under? If i did some derivative work on this data, could I release that (including the original files) as a kaggle dataset?\n\nOnce again, thanks for doing all this hard work. I really think it will be helpful to the community!",
    "472874": "Thanks for your feedback! I will respond to your questions to the best of my knowledge:\n\n 1. Yes that is right. Map means folder (Dutch) so \"Download folder\". It is also possible to selectively download the folders of your choice by checkmarking them and then on the right click download.\n 2. No I checked and you correctly mention that there are indeed 249 whales! We stopped one iteration too early.\n 3. We forgot to credit Nvidia for their code on domain randomization:  https://github.com/NVIDIA/Dataset_Synthesizer. Along with this wonderful paper: https://arxiv.org/abs/1804.06516. There may be a bug in their code or in the way we used it. We will check. The top left and lower right can for now be obtained by taking the bounds of the label images.\n 4. We haven't thought about that, but we can for sure :-)\n 5. Also not thought about, but either CC-0 or CC-BY.",
    "472885": "These tails are computer generated and fully independent from the set of real world images. If data transfer was no bottleneck, we could just as easily have created 10000 classes.",
    "472893": "Very impressive work. I will use it immediately.",
    "472895": "Be careful. Keep in mind that far not all models generalise well from synthetic datasets to the real data. From what I see, these images have vivid texture and simple imperfections, which may not be true for the competition data.",
    "472899": "Thank you for your advice.\nOK, I will download it and look at the data for the time being.",
    "473250": "One note for everyone trying to get the dataset: \n\nI've been unable to download the entire dataset as one zip file (seems like the download speed is capped at 2.5-3MB/s per connection and something times out before all ~50 GB is completed). What you can do instead is select chunks of folders (say 25 or 50 whale folders) in the UI by holding ctrl and clicking on the item then select 'Downloaden'. (after you select a chunk, you can click 'Deselecteren' to deselect all folders and start on the next chunk)\n\nYou can download 4 chunks at a time (each one has a ~3MB/s download speed). It's a little bit more manual but at least it seems to be reliable for downloading smaller (5-10 GB chunks).",
    "473708": "It would be pretty helpful if anyone can help to upload via google drive : )",
    "473817": "I started looking through the dataset and noticed some additional anomalies.\n\n\nImages are numbered `000000.*` to `049999`\n\n\nI think what is expected (and is true for some whale folders) is for each whale directory to have exactly 600 files(200 metadata `.json`, 200 mask `.cs.png` and 200 images `.png`) `whale0` is a good example which has `000000{.json|cs.png|png}` to `000199{json|cs.png|png}`.\n\n\nSome folders (`whale15` through `whale200` -- excepting `whale195` (see below)) have 603 files. It seems as though 600 of these files are the expected 200 instances with 3 images each in the natural ordering (e.g. `whale15` has files `003000* ` to `003199*`) in addition it has one extra set of a higher numbered image. This higher numbered sequence seems to increase by one for each folder: (e.g. `whale15` has `049815*`, `whale16` has `049816*`, etc)) I've verified at least for several instances the higher numbered image is not the same whale as the folder, and the higher number images do not appear to be the same whale (`049815*` is not the same whale as `049816`).\n\n\nFinally it seems that `whale195/039120.json` is missing (the mask and image files are there) and files `049801*` through `049814*` are all missing. I would presume following the above pattern these would have appeared in `whale1` through `whale14` and maybe were caught and deleted?",
    "475530": "Nice! I'm happy that someone made this :) Any details of implementation?\n\nP.S. However the main issue here as I see it - the water. It doesn't look right. I think that would be better to just place the tails on top of the real sea photos.",
    "475531": "Amazing! Thank you for sharing! I spent a lot of time for study how work blander and this was not a success. But you make it perfect.",
    "475536": "You don’t believe there are this many whales swimming in the waters of Morrowind?",
    "475787": "Any feedback whether this dataset improved the performance?",
    "476138": "Yes we found the bug and are working on a version 2. The new set will be online by tomorrow evening (UTC+01.00)  \n\nWe have looked at all the feedback and made the following changes:\n - Fixed bug with the 2d and 3d bounding boxes, this now works properly.\n - Added domain randomization: light (apart from the changes in time of day we added small random light sources\n - Added domain randomization: objects with random textures location and orientation. (This includes random geometries as well as a boat and eagle mesh.\n - Orientation of the whale tails will be more oriented to the camera. The current set is too random in our opinion.\n - More random water (possibly from real world images)\n\nFurthermore this time we will upload the whole set in parts on separate Google Drive locations. The Transip was not as stable as we had hoped.",
    "476141": "Haha love the Morrowind ref &lt;3",
    "476144": "Working on a version 2, see comment below :-). This time we will upload everything in parts to Drive.",
    "476145": "Thanks for your time to look into these details! We assume that this has something to do with the file writing sequencing. For version 2 we will try to fix this!",
    "476148": "For version 2 (see my comment on Hafizur's post) we will upload the whole set to Google Drive.",
    "476169": "So could you share details of implementation? Which software was used, approaches to texture generation, etc. That would be really helpful. Thanks.",
    "476994": "When you planning update your dataset?",
    "477113": "We are uploading the whole set as we speak on Google Drive. Accessible via the following link:\n[Synthetic Whale Tails V2][1]\n\n  [1]: https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing",
    "477542": "Update: the whole set is up and running.\nAs with the previous version, any feedback is welcome!",
    "477543": "Thanks for the feedback!",
    "477842": "Version 2 has been uploaded on Google Drive. We tried to incorporate all feedback. Accessible via the following link:\n[Synthetic Whale Tails V2][1]\n\nThis set contains:\n- Two sets (With and Without distractors)\n- Distractors are randomly placed geometric objects\n- 200 different whales (so keep in mind that although they have the same label, they are different)\n- 300 images of each whale\n\nMajor changes are:\n- We used 2k textures for the ocean\n- Tails are oriented more towards the camera\n- Added extra light sources\n- Added distractors\n- Fixed a few bugs (in particular the bounding boxes)\n\nA quick preview below :-)\n![enter image description here][2]\n![enter image description here][3]\n![enter image description here][4]\n![enter image description here][5]\n\n  [1]: https://drive.google.com/drive/folders/1jCtlJvR1Dmoxi6mSWo_Dtkbu_rgWPEyV?usp=sharing\n  [2]: https://i.ibb.co/K0LTrZG/019596.png\n  [3]: https://i.ibb.co/3sCPBV5/019596-cs.png\n  [4]: https://i.ibb.co/N2bnVpG/004214.png\n  [5]: https://i.ibb.co/HgwRpKs/004214-cs.png",
    "481951": "",
    "482244": "Got around to downloading the new dataset. Here's a couple of tips for anyone else attempts to grab it:\n- Google's Download Folder option never worked for me, it would start only about half of the files.\n- each whale.rar contains all of the images, but extracting it will put the images into the current directory (at least with linux's `unrar` tool). You need to do something like: `unrar e whale1.rar -d ./whale1/` to extract the contents into a folder called whale1.\n- Magic incantation to do this for all .rars in the current working directory: `find . -name \"*.rar\" -exec sh -c 'unrar e $1 -d \"${1%.*}/\"' _ {} \\;`",
    "482511": "how about using this and combine with deepmind paper to general novel view from single input image?\n\ntrain = synthetic image\n\noutput = generate-new-view(real image or images)\n\nOriginal Paper: Neural scene representation and rendering (Eslami, et al., 2018)\n\nhttps://deepmind.com/blog/neural-scene-representation-and-rendering\nhttps://www.youtube.com/watch?v=G-kWNQJ4idw\nhttps://github.com/musyoku/chainer-gqn"
  },
  "source": "meta"
}