{
  "id": 81475,
  "title": "Flukes Only Dataset(No water background) ",
  "url": "/competitions/humpback-whale-identification/discussion/81475",
  "author_name": "",
  "post_date": "2019-02-21T19:41:37.150463700Z",
  "votes": 18,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>This is my first ever Kaggle Competition which I started working on about 2 weeks ago. After submitting a few times I realized that I don't have the knowledge/skills to do extremely well in this competition.</p>\n\n<p>My knowledge and understanding has progressed extremely well since I joined this competition mainly due to the very helpful discussion boards of this competition and other competitions.</p>\n\n<p>During the last 2 weeks I created a dataset of the same images as provided with this competition but without the water background. The images in this dataset contain the exact same images as the original dataset but with only the flukes and data augmentation applied to copies of the same images.</p>\n\n<p><strong>Train Dataset:</strong> <a href=\"https://drive.google.com/open?id=1JOLbVY3SMdBV52Iegc9P-VnQejGa4brD\">https://drive.google.com/open?id=1JOLbVY3SMdBV52Iegc9P-VnQejGa4brD</a></p>\n\n<ul>\n<li>Number of images: 211,529</li>\n<li>The dataset is sorted into 5005 folders where each folder is named as its label class.</li>\n<li>Minimum number of images per class is 13. i.e. 1 original image and 12 augmented images (also without water background)</li>\n<li>Images are named exactly the same as in the original competition dataset except for the augmented images which have a suffix \"EDITED\" appended to their name. </li>\n<li>~12 augmentations were applied to each image from the original dataset hence the imbalance is the almost the same as that of the original dataset.</li>\n</ul>\n\n<p><strong>Test Dataset:</strong> <a href=\"https://drive.google.com/open?id=14Yv8pc9bg72cvPocnTLJi4TkUbtwn0Cu\">https://drive.google.com/open?id=14Yv8pc9bg72cvPocnTLJi4TkUbtwn0Cu</a>\n - </p>\n\n<ul>\n<li><p>Number of images: 7960</p></li>\n<li><p>Names of each image is exactly the same as in the original test dataset (public).</p></li>\n<li>No augmentations except for water background deletion were applied. You can use this dataset to make predictions for submission in exactly the same way as you did for the original test dataset.</li>\n</ul>\n\n<p><strong>NOTE</strong>: Some of the images in both the train and test datasets have a minor element of water background in them though images like these occur very infrequently.</p>\n\n<p><strong>Downloading Dataset:</strong>\nI recommend downloading the zip files rather than using the google drive link. If you choose to use the google drive link, the preview will not work because the dataset is quite large, just click download instead of previewing.</p>\n\n<p>I hope this is helpful to anyone in this competition.\nI have also attached both the datasets as zip files.</p>",
  "messages": [
    {
      "id": "476215",
      "postDate": "02/21/2019 19:41:37",
      "content": "<p>Hi all,</p>\n\n<p>This is my first ever Kaggle Competition which I started working on about 2 weeks ago. After submitting a few times I realized that I don't have the knowledge/skills to do extremely well in this competition.</p>\n\n<p>My knowledge and understanding has progressed extremely well since I joined this competition mainly due to the very helpful discussion boards of this competition and other competitions.</p>\n\n<p>During the last 2 weeks I created a dataset of the same images as provided with this competition but without the water background. The images in this dataset contain the exact same images as the original dataset but with only the flukes and data augmentation applied to copies of the same images.</p>\n\n<p><strong>Train Dataset:</strong> <a href=\"https://drive.google.com/open?id=1JOLbVY3SMdBV52Iegc9P-VnQejGa4brD\">https://drive.google.com/open?id=1JOLbVY3SMdBV52Iegc9P-VnQejGa4brD</a></p>\n\n<ul>\n<li>Number of images: 211,529</li>\n<li>The dataset is sorted into 5005 folders where each folder is named as its label class.</li>\n<li>Minimum number of images per class is 13. i.e. 1 original image and 12 augmented images (also without water background)</li>\n<li>Images are named exactly the same as in the original competition dataset except for the augmented images which have a suffix \"EDITED\" appended to their name. </li>\n<li>~12 augmentations were applied to each image from the original dataset hence the imbalance is the almost the same as that of the original dataset.</li>\n</ul>\n\n<p><strong>Test Dataset:</strong> <a href=\"https://drive.google.com/open?id=14Yv8pc9bg72cvPocnTLJi4TkUbtwn0Cu\">https://drive.google.com/open?id=14Yv8pc9bg72cvPocnTLJi4TkUbtwn0Cu</a>\n - </p>\n\n<ul>\n<li><p>Number of images: 7960</p></li>\n<li><p>Names of each image is exactly the same as in the original test dataset (public).</p></li>\n<li>No augmentations except for water background deletion were applied. You can use this dataset to make predictions for submission in exactly the same way as you did for the original test dataset.</li>\n</ul>\n\n<p><strong>NOTE</strong>: Some of the images in both the train and test datasets have a minor element of water background in them though images like these occur very infrequently.</p>\n\n<p><strong>Downloading Dataset:</strong>\nI recommend downloading the zip files rather than using the google drive link. If you choose to use the google drive link, the preview will not work because the dataset is quite large, just click download instead of previewing.</p>\n\n<p>I hope this is helpful to anyone in this competition.\nI have also attached both the datasets as zip files.</p>",
      "rawMarkdown": "Hi all,\n\nThis is my first ever Kaggle Competition which I started working on about 2 weeks ago. After submitting a few times I realized that I don't have the knowledge/skills to do extremely well in this competition.\n\nMy knowledge and understanding has progressed extremely well since I joined this competition mainly due to the very helpful discussion boards of this competition and other competitions.\n\nDuring the last 2 weeks I created a dataset of the same images as provided with this competition but without the water background. The images in this dataset contain the exact same images as the original dataset but with only the flukes and data augmentation applied to copies of the same images.\n\n**Train Dataset:** https://drive.google.com/open?id=1JOLbVY3SMdBV52Iegc9P-VnQejGa4brD\n\n - Number of images: 211,529\n - The dataset is sorted into 5005 folders where each folder is named as its label class.\n - Minimum number of images per class is 13. i.e. 1 original image and 12 augmented images (also without water background)\n - Images are named exactly the same as in the original competition dataset except for the augmented images which have a suffix \"EDITED\" appended to their name. \n - ~12 augmentations were applied to each image from the original dataset hence the imbalance is the almost the same as that of the original dataset.\n\n**Test Dataset:** https://drive.google.com/open?id=14Yv8pc9bg72cvPocnTLJi4TkUbtwn0Cu\n - \n\n - Number of images: 7960\n\n - Names of each image is exactly the same as in the original test dataset (public).\n - No augmentations except for water background deletion were applied. You can use this dataset to make predictions for submission in exactly the same way as you did for the original test dataset.\n\n**NOTE**: Some of the images in both the train and test datasets have a minor element of water background in them though images like these occur very infrequently.\n\n**Downloading Dataset:**\nI recommend downloading the zip files rather than using the google drive link. If you choose to use the google drive link, the preview will not work because the dataset is quite large, just click download instead of previewing.\n\nI hope this is helpful to anyone in this competition.\nI have also attached both the datasets as zip files.",
      "votes": null
    },
    {
      "id": "476278",
      "postDate": "02/21/2019 22:40:47",
      "content": "<p>Thank you for sharing this Alamjeet. I was wondering how you removed the water from the background. Did you use another model to do this?</p>",
      "rawMarkdown": "Thank you for sharing this Alamjeet. I was wondering how you removed the water from the background. Did you use another model to do this?",
      "votes": null
    },
    {
      "id": "476285",
      "postDate": "02/21/2019 22:54:59",
      "content": "<p>I wanted to use another model initially but I realized it would be easier to do in Photoshop by using the Automation and Batch tools.</p>",
      "rawMarkdown": "I wanted to use another model initially but I realized it would be easier to do in Photoshop by using the Automation and Batch tools.",
      "votes": null
    },
    {
      "id": "476291",
      "postDate": "02/21/2019 23:07:32",
      "content": "<p>Thanks for sharing! How accurate are your masks? Did you run something like U-net for binary segmentation?</p>",
      "rawMarkdown": "Thanks for sharing! How accurate are your masks? Did you run something like U-net for binary segmentation?",
      "votes": null
    },
    {
      "id": "476332",
      "postDate": "02/22/2019 01:41:43",
      "content": "<p>I did not use any models or masks, just Photoshop. In most of the images the background has been removed quite accurately while there are some where some background is still present.</p>",
      "rawMarkdown": "I did not use any models or masks, just Photoshop. In most of the images the background has been removed quite accurately while there are some where some background is still present.",
      "votes": null
    },
    {
      "id": "477007",
      "postDate": "02/23/2019 17:14:11",
      "content": "<p>Thank you for your nice work.\nIf you do not mind, is it possible to attach several sample images to this discussion?\nI just download it and see the image, but the file size is a little big, so I am glad if you can attach a sample image. \nThat will help us understand :)</p>",
      "rawMarkdown": "Thank you for your nice work.\nIf you do not mind, is it possible to attach several sample images to this discussion?\nI just download it and see the image, but the file size is a little big, so I am glad if you can attach a sample image. \nThat will help us understand :)",
      "votes": null
    },
    {
      "id": "477049",
      "postDate": "02/23/2019 18:52:46",
      "content": "<p>Attached</p>",
      "rawMarkdown": "Attached",
      "votes": null
    },
    {
      "id": "477092",
      "postDate": "02/23/2019 21:13:50",
      "content": "<p>Is it legal to use a manually enhanced test dataset?</p>",
      "rawMarkdown": "Is it legal to use a manually enhanced test dataset?",
      "votes": null
    },
    {
      "id": "477196",
      "postDate": "02/24/2019 05:12:57",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "477294",
      "postDate": "02/24/2019 09:57:58",
      "content": "<p>This is great and after removing the background in images, it might help to improve the results. I would give it a try and thanks for sharing.</p>",
      "rawMarkdown": "This is great and after removing the background in images, it might help to improve the results. I would give it a try and thanks for sharing.",
      "votes": null
    },
    {
      "id": "477410",
      "postDate": "02/24/2019 13:49:10",
      "content": "<p>I tried using the test image on a single model with LB 0.865, but the result was LB 0.843. \nHowever, I got a score of +0.004 (LB 0.869) when I tried ensemble with the result(LB 0.843 with No water background) and the 0.865 model.</p>\n\n<p>However, although @lytic also points out, it is doubtful whether you can use the test annotated test image. \nHow is it actually?</p>",
      "rawMarkdown": "I tried using the test image on a single model with LB 0.865, but the result was LB 0.843. \nHowever, I got a score of +0.004 (LB 0.869) when I tried ensemble with the result(LB 0.843 with No water background) and the 0.865 model.\n\nHowever, although @lytic also points out, it is doubtful whether you can use the test annotated test image. \nHow is it actually?",
      "votes": null
    },
    {
      "id": "477436",
      "postDate": "02/24/2019 14:51:02",
      "content": "<p>In general, hand annotation for test images are not allowed. If test dataset are made by automatically (ML or some algorithms), it is allowed to use it.</p>",
      "rawMarkdown": "In general, hand annotation for test images are not allowed. If test dataset are made by automatically (ML or some algorithms), it is allowed to use it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 476278,
      "author_name": "zjimmy",
      "author_url": "",
      "post_date": "02/21/2019 22:40:47",
      "content": "<p>Thank you for sharing this Alamjeet. I was wondering how you removed the water from the background. Did you use another model to do this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 476285,
          "author_name": "alamjs",
          "author_url": "",
          "post_date": "02/21/2019 22:54:59",
          "content": "<p>I wanted to use another model initially but I realized it would be easier to do in Photoshop by using the Automation and Batch tools.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 476291,
      "author_name": "asanakoev",
      "author_url": "",
      "post_date": "02/21/2019 23:07:32",
      "content": "<p>Thanks for sharing! How accurate are your masks? Did you run something like U-net for binary segmentation?</p>",
      "votes": null,
      "replies": [
        {
          "id": 476332,
          "author_name": "alamjs",
          "author_url": "",
          "post_date": "02/22/2019 01:41:43",
          "content": "<p>I did not use any models or masks, just Photoshop. In most of the images the background has been removed quite accurately while there are some where some background is still present.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 477007,
      "author_name": "kerukun",
      "author_url": "",
      "post_date": "02/23/2019 17:14:11",
      "content": "<p>Thank you for your nice work.\nIf you do not mind, is it possible to attach several sample images to this discussion?\nI just download it and see the image, but the file size is a little big, so I am glad if you can attach a sample image. \nThat will help us understand :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 477049,
          "author_name": "alamjs",
          "author_url": "",
          "post_date": "02/23/2019 18:52:46",
          "content": "<p>Attached</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477196,
          "author_name": "kerukun",
          "author_url": "",
          "post_date": "02/24/2019 05:12:57",
          "content": "<p>Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 477092,
      "author_name": "sorokin",
      "author_url": "",
      "post_date": "02/23/2019 21:13:50",
      "content": "<p>Is it legal to use a manually enhanced test dataset?</p>",
      "votes": null,
      "replies": [
        {
          "id": 477436,
          "author_name": "toshik",
          "author_url": "",
          "post_date": "02/24/2019 14:51:02",
          "content": "<p>In general, hand annotation for test images are not allowed. If test dataset are made by automatically (ML or some algorithms), it is allowed to use it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 477294,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "02/24/2019 09:57:58",
      "content": "<p>This is great and after removing the background in images, it might help to improve the results. I would give it a try and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 477410,
      "author_name": "kerukun",
      "author_url": "",
      "post_date": "02/24/2019 13:49:10",
      "content": "<p>I tried using the test image on a single model with LB 0.865, but the result was LB 0.843. \nHowever, I got a score of +0.004 (LB 0.869) when I tried ensemble with the result(LB 0.843 with No water background) and the 0.865 model.</p>\n\n<p>However, although @lytic also points out, it is doubtful whether you can use the test annotated test image. \nHow is it actually?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "476215": "Hi all,\n\nThis is my first ever Kaggle Competition which I started working on about 2 weeks ago. After submitting a few times I realized that I don't have the knowledge/skills to do extremely well in this competition.\n\nMy knowledge and understanding has progressed extremely well since I joined this competition mainly due to the very helpful discussion boards of this competition and other competitions.\n\nDuring the last 2 weeks I created a dataset of the same images as provided with this competition but without the water background. The images in this dataset contain the exact same images as the original dataset but with only the flukes and data augmentation applied to copies of the same images.\n\n**Train Dataset:** https://drive.google.com/open?id=1JOLbVY3SMdBV52Iegc9P-VnQejGa4brD\n\n - Number of images: 211,529\n - The dataset is sorted into 5005 folders where each folder is named as its label class.\n - Minimum number of images per class is 13. i.e. 1 original image and 12 augmented images (also without water background)\n - Images are named exactly the same as in the original competition dataset except for the augmented images which have a suffix \"EDITED\" appended to their name. \n - ~12 augmentations were applied to each image from the original dataset hence the imbalance is the almost the same as that of the original dataset.\n\n**Test Dataset:** https://drive.google.com/open?id=14Yv8pc9bg72cvPocnTLJi4TkUbtwn0Cu\n - \n\n - Number of images: 7960\n\n - Names of each image is exactly the same as in the original test dataset (public).\n - No augmentations except for water background deletion were applied. You can use this dataset to make predictions for submission in exactly the same way as you did for the original test dataset.\n\n**NOTE**: Some of the images in both the train and test datasets have a minor element of water background in them though images like these occur very infrequently.\n\n**Downloading Dataset:**\nI recommend downloading the zip files rather than using the google drive link. If you choose to use the google drive link, the preview will not work because the dataset is quite large, just click download instead of previewing.\n\nI hope this is helpful to anyone in this competition.\nI have also attached both the datasets as zip files.",
    "476278": "Thank you for sharing this Alamjeet. I was wondering how you removed the water from the background. Did you use another model to do this?",
    "476285": "I wanted to use another model initially but I realized it would be easier to do in Photoshop by using the Automation and Batch tools.",
    "476291": "Thanks for sharing! How accurate are your masks? Did you run something like U-net for binary segmentation?",
    "476332": "I did not use any models or masks, just Photoshop. In most of the images the background has been removed quite accurately while there are some where some background is still present.",
    "477007": "Thank you for your nice work.\nIf you do not mind, is it possible to attach several sample images to this discussion?\nI just download it and see the image, but the file size is a little big, so I am glad if you can attach a sample image. \nThat will help us understand :)",
    "477049": "Attached",
    "477092": "Is it legal to use a manually enhanced test dataset?",
    "477196": "Thanks",
    "477294": "This is great and after removing the background in images, it might help to improve the results. I would give it a try and thanks for sharing.",
    "477410": "I tried using the test image on a single model with LB 0.865, but the result was LB 0.843. \nHowever, I got a score of +0.004 (LB 0.869) when I tried ensemble with the result(LB 0.843 with No water background) and the 0.865 model.\n\nHowever, although @lytic also points out, it is doubtful whether you can use the test annotated test image. \nHow is it actually?",
    "477436": "In general, hand annotation for test images are not allowed. If test dataset are made by automatically (ML or some algorithms), it is allowed to use it."
  },
  "source": "meta"
}