{
  "id": 309691,
  "title": "Dataset Dataset Dataset",
  "url": "/competitions/happy-whale-and-dolphin/discussion/309691",
  "author_name": "",
  "post_date": "2022-02-24T21:37:38.121183400Z",
  "votes": 29,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The more you invest in this competition, the more you will realize that the dataset and ways you are handling it will play a crucial role. </p>\n<p>A lot of great work has already been done and this post is a way to appreciate their works and amass them here. </p>\n<h2>Raw Images Resized</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304686\" target=\"_blank\">Reduced Resolution Image Data (128 x 128, 256 x 256, 384 x 384) 🐋</a> by <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> - <a href=\"https://www.kaggle.com/rdizzl3/jpeg-happywhale-128x128\" target=\"_blank\">128x128</a>, <a href=\"https://www.kaggle.com/rdizzl3/jpeg-happywhale-256x256\" target=\"_blank\">256x256</a>, <a href=\"https://www.kaggle.com/rdizzl3/jpeg-happywhale-384x384\" target=\"_blank\">384x384</a></li>\n</ul>\n<h2>TFRecords of Raw Images</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/ks2019/happywhale-tfrecords\" target=\"_blank\">HappyWhale TFRecords</a> by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> - <a href=\"https://www.kaggle.com/ks2019/happywhale-tfrecords-v1\" target=\"_blank\">Dataset</a></li>\n</ul>\n<h2>Bounding Box Based Crops</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5/\" target=\"_blank\">Happywhale: Cropped Dataset [YOLOv5] ✂️</a> by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> - <a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5-ds\" target=\"_blank\">Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503\" target=\"_blank\">cropped&amp;resized(512x512) dataset using detic</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> - <a href=\"https://www.kaggle.com/phalanx/whale2-cropped-dataset\" target=\"_blank\">Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309611\" target=\"_blank\">TF records with Detic Crop</a> by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> </li>\n</ul>\n<h2>Augmented Images</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/ruchi798/happywhale-augmented\" target=\"_blank\">🐋🐬Happywhale: Augmented</a> by <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </li>\n</ul>\n<h2>Segmentation Masks</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/shubhambaid/happywhales-labelme-segmentation-dataset\" target=\"_blank\">HappyWhales LabelMe Segmentation Dataset</a> by <a href=\"https://www.kaggle.com/shubhambaid\" target=\"_blank\">@shubhambaid</a> (not full dataset)</li>\n</ul>\n<h2>Teased us :P</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214\" target=\"_blank\">🔥 DATASET - dorsal fins for all IDs without background 🔥</a> by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </li>\n</ul>\n<h2>Dataset Painpoints and EDA</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026\" target=\"_blank\">🐧 Things to know before starting image preprocessing</a> by <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>.</li>\n</ul>\n<h2>Camera View Point</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints\" target=\"_blank\">Happywhale Camera Viewpoints Dataset</a></li>\n</ul>",
  "messages": [
    {
      "id": "1703798",
      "postDate": "02/24/2022 21:37:38",
      "content": "<p>The more you invest in this competition, the more you will realize that the dataset and ways you are handling it will play a crucial role. </p>\n<p>A lot of great work has already been done and this post is a way to appreciate their works and amass them here. </p>\n<h2>Raw Images Resized</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304686\" target=\"_blank\">Reduced Resolution Image Data (128 x 128, 256 x 256, 384 x 384) 🐋</a> by <a href=\"https://www.kaggle.com/rdizzl3\" target=\"_blank\">@rdizzl3</a> - <a href=\"https://www.kaggle.com/rdizzl3/jpeg-happywhale-128x128\" target=\"_blank\">128x128</a>, <a href=\"https://www.kaggle.com/rdizzl3/jpeg-happywhale-256x256\" target=\"_blank\">256x256</a>, <a href=\"https://www.kaggle.com/rdizzl3/jpeg-happywhale-384x384\" target=\"_blank\">384x384</a></li>\n</ul>\n<h2>TFRecords of Raw Images</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/ks2019/happywhale-tfrecords\" target=\"_blank\">HappyWhale TFRecords</a> by <a href=\"https://www.kaggle.com/ks2019\" target=\"_blank\">@ks2019</a> - <a href=\"https://www.kaggle.com/ks2019/happywhale-tfrecords-v1\" target=\"_blank\">Dataset</a></li>\n</ul>\n<h2>Bounding Box Based Crops</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5/\" target=\"_blank\">Happywhale: Cropped Dataset [YOLOv5] ✂️</a> by <a href=\"https://www.kaggle.com/awsaf49\" target=\"_blank\">@awsaf49</a> - <a href=\"https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5-ds\" target=\"_blank\">Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503\" target=\"_blank\">cropped&amp;resized(512x512) dataset using detic</a> by <a href=\"https://www.kaggle.com/phalanx\" target=\"_blank\">@phalanx</a> - <a href=\"https://www.kaggle.com/phalanx/whale2-cropped-dataset\" target=\"_blank\">Dataset</a></li>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309611\" target=\"_blank\">TF records with Detic Crop</a> by <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> </li>\n</ul>\n<h2>Augmented Images</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/ruchi798/happywhale-augmented\" target=\"_blank\">🐋🐬Happywhale: Augmented</a> by <a href=\"https://www.kaggle.com/ruchi798\" target=\"_blank\">@ruchi798</a> </li>\n</ul>\n<h2>Segmentation Masks</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/shubhambaid/happywhales-labelme-segmentation-dataset\" target=\"_blank\">HappyWhales LabelMe Segmentation Dataset</a> by <a href=\"https://www.kaggle.com/shubhambaid\" target=\"_blank\">@shubhambaid</a> (not full dataset)</li>\n</ul>\n<h2>Teased us :P</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214\" target=\"_blank\">🔥 DATASET - dorsal fins for all IDs without background 🔥</a> by <a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> </li>\n</ul>\n<h2>Dataset Painpoints and EDA</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026\" target=\"_blank\">🐧 Things to know before starting image preprocessing</a> by <a href=\"https://www.kaggle.com/andradaolteanu\" target=\"_blank\">@andradaolteanu</a>.</li>\n</ul>\n<h2>Camera View Point</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints\" target=\"_blank\">Happywhale Camera Viewpoints Dataset</a></li>\n</ul>",
      "rawMarkdown": "The more you invest in this competition, the more you will realize that the dataset and ways you are handling it will play a crucial role. \n\nA lot of great work has already been done and this post is a way to appreciate their works and amass them here. \n\n## Raw Images Resized\n\n* [Reduced Resolution Image Data (128 x 128, 256 x 256, 384 x 384) 🐋](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304686) by @rdizzl3 - [128x128](https://www.kaggle.com/rdizzl3/jpeg-happywhale-128x128), [256x256](https://www.kaggle.com/rdizzl3/jpeg-happywhale-256x256), [384x384](https://www.kaggle.com/rdizzl3/jpeg-happywhale-384x384)\n\n## TFRecords of Raw Images\n\n* [HappyWhale TFRecords](https://www.kaggle.com/ks2019/happywhale-tfrecords) by @ks2019 - [Dataset](https://www.kaggle.com/ks2019/happywhale-tfrecords-v1)\n\n## Bounding Box Based Crops\n\n* [Happywhale: Cropped Dataset [YOLOv5] ✂️](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5/) by @awsaf49 - [Dataset](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5-ds)\n* [cropped&resized(512x512) dataset using detic](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503) by @phalanx - [Dataset](https://www.kaggle.com/phalanx/whale2-cropped-dataset)\n* [TF records with Detic Crop](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309611) by @ragnar123 \n\n## Augmented Images\n\n* [🐋🐬Happywhale: Augmented](https://www.kaggle.com/ruchi798/happywhale-augmented) by @ruchi798 \n\n## Segmentation Masks\n\n* [HappyWhales LabelMe Segmentation Dataset](https://www.kaggle.com/shubhambaid/happywhales-labelme-segmentation-dataset) by @shubhambaid (not full dataset)\n\n## Teased us :P\n\n* [🔥 DATASET - dorsal fins for all IDs without background 🔥](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214) by @remekkinas \n\n## Dataset Painpoints and EDA\n\n* [🐧 Things to know before starting image preprocessing](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026) by @andradaolteanu.\n\n## Camera View Point\n\n* [Happywhale Camera Viewpoints Dataset](https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints)",
      "votes": null
    },
    {
      "id": "1703813",
      "postDate": "02/24/2022 22:22:12",
      "content": "<p><a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> thank you for mentioning my work. I am still working on better solution - removing background (now it is almost acceptable - almost). To improve it even more I need retrain model - it requires annotations like in semantic segmentation. It is time consuming but … now testing simillarity model performance (using my dataset). I am almost sure that there is no huge benefit from removing background - still testing. If you prepare good data there is no need to remove background. I will let you know when finish testing.</p>",
      "rawMarkdown": "ayuraj thank you for mentioning my work. I am still working on better solution - removing background (now it is almost acceptable - almost). To improve it even more I need retrain model - it requires annotations like in semantic segmentation. It is time consuming but ... now testing simillarity model performance (using my dataset). I am almost sure that there is no huge benefit from removing background - still testing. If you prepare good data there is no need to remove background. I will let you know when finish testing.",
      "votes": null
    },
    {
      "id": "1703825",
      "postDate": "02/24/2022 23:19:30",
      "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I found this segmentation dataset, you may be able to use it.<br>\nI haven't checked it yet. I think mostly dolphins, but I am not sure.</p>\n<p><a href=\"https://data.ncl.ac.uk/articles/dataset/NDD20_zip/12357383\" target=\"_blank\">https://data.ncl.ac.uk/articles/dataset/NDD20_zip/12357383</a></p>",
      "rawMarkdown": "remekkinas I found this segmentation dataset, you may be able to use it.\nI haven't checked it yet. I think mostly dolphins, but I am not sure.\n\nhttps://data.ncl.ac.uk/articles/dataset/NDD20_zip/12357383",
      "votes": null
    },
    {
      "id": "1704305",
      "postDate": "02/25/2022 11:50:37",
      "content": "<p>Thank you for putting this list together! I have also manually labelled 1,680 images with camera viewpoints in case it helps!</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints\" target=\"_blank\">Happywhale Camera Viewpoints Dataset</a></li>\n</ul>",
      "rawMarkdown": "Thank you for putting this list together! I have also manually labelled 1,680 images with camera viewpoints in case it helps!\n- [Happywhale Camera Viewpoints Dataset](https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints)",
      "votes": null
    },
    {
      "id": "1704812",
      "postDate": "02/25/2022 20:53:49",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/frlemarchand\" target=\"_blank\">@frlemarchand</a>. Updated the list.</p>",
      "rawMarkdown": "Thanks for sharing @frlemarchand. Updated the list.",
      "votes": null
    },
    {
      "id": "1704844",
      "postDate": "02/25/2022 22:01:31",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>! <br>\nSegmentation should work in my case. I have to improve NN a little bit and had to take decision -&gt; do it manually or find dataset. Now problem looks easier for me :) Thank you!</p>",
      "rawMarkdown": "Thank you @pestipeti! \nSegmentation should work in my case. I have to improve NN a little bit and had to take decision -> do it manually or find dataset. Now problem looks easier for me :) Thank you!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1703813,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "02/24/2022 22:22:12",
      "content": "<p><a href=\"https://www.kaggle.com/ayuraj\" target=\"_blank\">@ayuraj</a> thank you for mentioning my work. I am still working on better solution - removing background (now it is almost acceptable - almost). To improve it even more I need retrain model - it requires annotations like in semantic segmentation. It is time consuming but … now testing simillarity model performance (using my dataset). I am almost sure that there is no huge benefit from removing background - still testing. If you prepare good data there is no need to remove background. I will let you know when finish testing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1703825,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "02/24/2022 23:19:30",
          "content": "<p><a href=\"https://www.kaggle.com/remekkinas\" target=\"_blank\">@remekkinas</a> I found this segmentation dataset, you may be able to use it.<br>\nI haven't checked it yet. I think mostly dolphins, but I am not sure.</p>\n<p><a href=\"https://data.ncl.ac.uk/articles/dataset/NDD20_zip/12357383\" target=\"_blank\">https://data.ncl.ac.uk/articles/dataset/NDD20_zip/12357383</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1704844,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "02/25/2022 22:01:31",
          "content": "<p>Thank you <a href=\"https://www.kaggle.com/pestipeti\" target=\"_blank\">@pestipeti</a>! <br>\nSegmentation should work in my case. I have to improve NN a little bit and had to take decision -&gt; do it manually or find dataset. Now problem looks easier for me :) Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1704305,
      "author_name": "frlemarchand",
      "author_url": "",
      "post_date": "02/25/2022 11:50:37",
      "content": "<p>Thank you for putting this list together! I have also manually labelled 1,680 images with camera viewpoints in case it helps!</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints\" target=\"_blank\">Happywhale Camera Viewpoints Dataset</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1704812,
          "author_name": "ayuraj",
          "author_url": "",
          "post_date": "02/25/2022 20:53:49",
          "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/frlemarchand\" target=\"_blank\">@frlemarchand</a>. Updated the list.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1703798": "The more you invest in this competition, the more you will realize that the dataset and ways you are handling it will play a crucial role. \n\nA lot of great work has already been done and this post is a way to appreciate their works and amass them here. \n\n## Raw Images Resized\n\n* [Reduced Resolution Image Data (128 x 128, 256 x 256, 384 x 384) 🐋](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/304686) by @rdizzl3 - [128x128](https://www.kaggle.com/rdizzl3/jpeg-happywhale-128x128), [256x256](https://www.kaggle.com/rdizzl3/jpeg-happywhale-256x256), [384x384](https://www.kaggle.com/rdizzl3/jpeg-happywhale-384x384)\n\n## TFRecords of Raw Images\n\n* [HappyWhale TFRecords](https://www.kaggle.com/ks2019/happywhale-tfrecords) by @ks2019 - [Dataset](https://www.kaggle.com/ks2019/happywhale-tfrecords-v1)\n\n## Bounding Box Based Crops\n\n* [Happywhale: Cropped Dataset [YOLOv5] ✂️](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5/) by @awsaf49 - [Dataset](https://www.kaggle.com/awsaf49/happywhale-cropped-dataset-yolov5-ds)\n* [cropped&resized(512x512) dataset using detic](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/305503) by @phalanx - [Dataset](https://www.kaggle.com/phalanx/whale2-cropped-dataset)\n* [TF records with Detic Crop](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309611) by @ragnar123 \n\n## Augmented Images\n\n* [🐋🐬Happywhale: Augmented](https://www.kaggle.com/ruchi798/happywhale-augmented) by @ruchi798 \n\n## Segmentation Masks\n\n* [HappyWhales LabelMe Segmentation Dataset](https://www.kaggle.com/shubhambaid/happywhales-labelme-segmentation-dataset) by @shubhambaid (not full dataset)\n\n## Teased us :P\n\n* [🔥 DATASET - dorsal fins for all IDs without background 🔥](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/309214) by @remekkinas \n\n## Dataset Painpoints and EDA\n\n* [🐧 Things to know before starting image preprocessing](https://www.kaggle.com/c/happy-whale-and-dolphin/discussion/308026) by @andradaolteanu.\n\n## Camera View Point\n\n* [Happywhale Camera Viewpoints Dataset](https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints)",
    "1703813": "ayuraj thank you for mentioning my work. I am still working on better solution - removing background (now it is almost acceptable - almost). To improve it even more I need retrain model - it requires annotations like in semantic segmentation. It is time consuming but ... now testing simillarity model performance (using my dataset). I am almost sure that there is no huge benefit from removing background - still testing. If you prepare good data there is no need to remove background. I will let you know when finish testing.",
    "1703825": "remekkinas I found this segmentation dataset, you may be able to use it.\nI haven't checked it yet. I think mostly dolphins, but I am not sure.\n\nhttps://data.ncl.ac.uk/articles/dataset/NDD20_zip/12357383",
    "1704305": "Thank you for putting this list together! I have also manually labelled 1,680 images with camera viewpoints in case it helps!\n- [Happywhale Camera Viewpoints Dataset](https://www.kaggle.com/frlemarchand/happywhale-camera-viewpoints)",
    "1704812": "Thanks for sharing @frlemarchand. Updated the list.",
    "1704844": "Thank you @pestipeti! \nSegmentation should work in my case. I have to improve NN a little bit and had to take decision -> do it manually or find dataset. Now problem looks easier for me :) Thank you!"
  },
  "source": "meta"
}