{
  "id": 162281,
  "title": "Balanced and Segmented dataset.",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/162281",
  "author_name": "",
  "post_date": "2020-06-28T08:59:47.137620100Z",
  "votes": 11,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I was searching about how dermatologists identify the melanoma lesion's and then I found The <a href=\"https://www.skincancer.org/skin-cancer-information/melanoma/melanoma-warning-signs-and-images/\">ABCDEs Rule</a> which helps them in the recognition of melanoma lesions. \nI found this <a href=\"https://web.stanford.edu/~kalouche/docs/Vision_Based_Classification_of_Skin_Cancer_using_Deep_Learning_%28Kalouche%29.pdf\">paper</a> which use segmented images of the lesion for training. I think it is a good idea to enforce our model to learn about the ABCDE rule properly. Based on this idea and his <a href=\"https://github.com/skalouche/Vision-based-Melanoma-Classification\">code</a> I have prepared a segmented, balanced dataset of 224x224 image size. I have made sure that the distribution of the training and the testing dataset is same.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3530433%2Fdc45565c660aaa7c2d53ff7601875f4e%2F2.png?generation=1593333290099622&amp;alt=media\" alt=\"\">\nNot all the images in the dataset are this much clean which is because of the asymmetric shape of the lesion (I tried to clean them but on doing that images were losing the asymmetric property. ) \n<strong>I got a boost of 4% after training my efficient net b0 on this segmented dataset.</strong>\n*<em>I have published the dataset <a href=\"https://www.kaggle.com/prateek0x/siimisic-segmented-and-balanced-dataset\">here</a></em>*</p>",
  "messages": [
    {
      "id": "905115",
      "postDate": "06/28/2020 08:59:47",
      "content": "<p>I was searching about how dermatologists identify the melanoma lesion's and then I found The <a href=\"https://www.skincancer.org/skin-cancer-information/melanoma/melanoma-warning-signs-and-images/\">ABCDEs Rule</a> which helps them in the recognition of melanoma lesions. \nI found this <a href=\"https://web.stanford.edu/~kalouche/docs/Vision_Based_Classification_of_Skin_Cancer_using_Deep_Learning_%28Kalouche%29.pdf\">paper</a> which use segmented images of the lesion for training. I think it is a good idea to enforce our model to learn about the ABCDE rule properly. Based on this idea and his <a href=\"https://github.com/skalouche/Vision-based-Melanoma-Classification\">code</a> I have prepared a segmented, balanced dataset of 224x224 image size. I have made sure that the distribution of the training and the testing dataset is same.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3530433%2Fdc45565c660aaa7c2d53ff7601875f4e%2F2.png?generation=1593333290099622&amp;alt=media\" alt=\"\">\nNot all the images in the dataset are this much clean which is because of the asymmetric shape of the lesion (I tried to clean them but on doing that images were losing the asymmetric property. ) \n<strong>I got a boost of 4% after training my efficient net b0 on this segmented dataset.</strong>\n*<em>I have published the dataset <a href=\"https://www.kaggle.com/prateek0x/siimisic-segmented-and-balanced-dataset\">here</a></em>*</p>",
      "rawMarkdown": "I was searching about how dermatologists identify the melanoma lesion's and then I found The [ABCDEs Rule](https://www.skincancer.org/skin-cancer-information/melanoma/melanoma-warning-signs-and-images/) which helps them in the recognition of melanoma lesions. \nI found this [paper](https://web.stanford.edu/~kalouche/docs/Vision_Based_Classification_of_Skin_Cancer_using_Deep_Learning_(Kalouche).pdf) which use segmented images of the lesion for training. I think it is a good idea to enforce our model to learn about the ABCDE rule properly. Based on this idea and his [code](https://github.com/skalouche/Vision-based-Melanoma-Classification) I have prepared a segmented, balanced dataset of 224x224 image size. I have made sure that the distribution of the training and the testing dataset is same.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3530433%2Fdc45565c660aaa7c2d53ff7601875f4e%2F2.png?generation=1593333290099622&amp;alt=media)\nNot all the images in the dataset are this much clean which is because of the asymmetric shape of the lesion (I tried to clean them but on doing that images were losing the asymmetric property. ) \n**I got a boost of 4% after training my efficient net b0 on this segmented dataset.**\n**I have published the dataset [here](https://www.kaggle.com/prateek0x/siimisic-segmented-and-balanced-dataset)**",
      "votes": null
    },
    {
      "id": "905890",
      "postDate": "06/28/2020 23:00:39",
      "content": "<p>thanks man , would use it</p>",
      "rawMarkdown": "thanks man , would use it",
      "votes": null
    },
    {
      "id": "905912",
      "postDate": "06/28/2020 23:57:28",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks",
      "votes": null
    },
    {
      "id": "906803",
      "postDate": "06/29/2020 14:43:46",
      "content": "<p>You’re welcome.</p>",
      "rawMarkdown": "You’re welcome.",
      "votes": null
    },
    {
      "id": "906804",
      "postDate": "06/29/2020 14:43:53",
      "content": "<p>You’re welcome.</p>",
      "rawMarkdown": "You’re welcome.",
      "votes": null
    },
    {
      "id": "910744",
      "postDate": "07/01/2020 10:38:30",
      "content": "<p>Thanks for sharing. \nI tried out the coding in github you link and saw that you must have done some processing not covered by the current version. What processing have you included? I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?</p>",
      "rawMarkdown": "Thanks for sharing. \nI tried out the coding in github you link and saw that you must have done some processing not covered by the current version. What processing have you included? I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?",
      "votes": null
    },
    {
      "id": "911076",
      "postDate": "07/01/2020 14:55:39",
      "content": "<p>&gt; What processing have you included? \nHis pre-processing techniques were not that good. It required some changes like ;\n- He was doing normal thresholding which does not work well and I changed it to otsu's thresholding which worked well.\n- In most of the images, I observed brightness in the vertical sides, which i cropped.\n- His canny edge detection parameters were not tuned.</p>\n\n<p>&gt; I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?\nThe existing dataset is very imbalanced and I just downsampled it. while down sampling  i made sure that test and train distribution is the same.\nThis data contains approx 600 images of each class.</p>",
      "rawMarkdown": "&gt; What processing have you included? \nHis pre-processing techniques were not that good. It required some changes like ;\n- He was doing normal thresholding which does not work well and I changed it to otsu's thresholding which worked well.\n- In most of the images, I observed brightness in the vertical sides, which i cropped.\n- His canny edge detection parameters were not tuned.\n\n&gt; I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?\nThe existing dataset is very imbalanced and I just downsampled it. while down sampling  i made sure that test and train distribution is the same.\nThis data contains approx 600 images of each class.",
      "votes": null
    },
    {
      "id": "920690",
      "postDate": "07/08/2020 18:50:10",
      "content": "<p>Nice works! Thanks</p>",
      "rawMarkdown": "Nice works! Thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 905890,
      "author_name": "yash612",
      "author_url": "",
      "post_date": "06/28/2020 23:00:39",
      "content": "<p>thanks man , would use it</p>",
      "votes": null,
      "replies": [
        {
          "id": 906803,
          "author_name": "prateek0x",
          "author_url": "",
          "post_date": "06/29/2020 14:43:46",
          "content": "<p>You’re welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 905912,
      "author_name": "felipejardimf",
      "author_url": "",
      "post_date": "06/28/2020 23:57:28",
      "content": "<p>Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 906804,
          "author_name": "prateek0x",
          "author_url": "",
          "post_date": "06/29/2020 14:43:53",
          "content": "<p>You’re welcome.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 910744,
      "author_name": "helgith",
      "author_url": "",
      "post_date": "07/01/2020 10:38:30",
      "content": "<p>Thanks for sharing. \nI tried out the coding in github you link and saw that you must have done some processing not covered by the current version. What processing have you included? I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?</p>",
      "votes": null,
      "replies": [
        {
          "id": 911076,
          "author_name": "prateek0x",
          "author_url": "",
          "post_date": "07/01/2020 14:55:39",
          "content": "<p>&gt; What processing have you included? \nHis pre-processing techniques were not that good. It required some changes like ;\n- He was doing normal thresholding which does not work well and I changed it to otsu's thresholding which worked well.\n- In most of the images, I observed brightness in the vertical sides, which i cropped.\n- His canny edge detection parameters were not tuned.</p>\n\n<p>&gt; I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?\nThe existing dataset is very imbalanced and I just downsampled it. while down sampling  i made sure that test and train distribution is the same.\nThis data contains approx 600 images of each class.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 920690,
      "author_name": "doanquanvietnamca",
      "author_url": "",
      "post_date": "07/08/2020 18:50:10",
      "content": "<p>Nice works! Thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "905115": "I was searching about how dermatologists identify the melanoma lesion's and then I found The [ABCDEs Rule](https://www.skincancer.org/skin-cancer-information/melanoma/melanoma-warning-signs-and-images/) which helps them in the recognition of melanoma lesions. \nI found this [paper](https://web.stanford.edu/~kalouche/docs/Vision_Based_Classification_of_Skin_Cancer_using_Deep_Learning_(Kalouche).pdf) which use segmented images of the lesion for training. I think it is a good idea to enforce our model to learn about the ABCDE rule properly. Based on this idea and his [code](https://github.com/skalouche/Vision-based-Melanoma-Classification) I have prepared a segmented, balanced dataset of 224x224 image size. I have made sure that the distribution of the training and the testing dataset is same.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3530433%2Fdc45565c660aaa7c2d53ff7601875f4e%2F2.png?generation=1593333290099622&amp;alt=media)\nNot all the images in the dataset are this much clean which is because of the asymmetric shape of the lesion (I tried to clean them but on doing that images were losing the asymmetric property. ) \n**I got a boost of 4% after training my efficient net b0 on this segmented dataset.**\n**I have published the dataset [here](https://www.kaggle.com/prateek0x/siimisic-segmented-and-balanced-dataset)**",
    "905890": "thanks man , would use it",
    "905912": "Thanks",
    "906803": "You’re welcome.",
    "906804": "You’re welcome.",
    "910744": "Thanks for sharing. \nI tried out the coding in github you link and saw that you must have done some processing not covered by the current version. What processing have you included? I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?",
    "911076": "&gt; What processing have you included? \nHis pre-processing techniques were not that good. It required some changes like ;\n- He was doing normal thresholding which does not work well and I changed it to otsu's thresholding which worked well.\n- In most of the images, I observed brightness in the vertical sides, which i cropped.\n- His canny edge detection parameters were not tuned.\n\n&gt; I also saw that the dataset linked only includes 1178 images for train and not 33126 as the complete train dataset. Is this intentional?\nThe existing dataset is very imbalanced and I just downsampled it. while down sampling  i made sure that test and train distribution is the same.\nThis data contains approx 600 images of each class.",
    "920690": "Nice works! Thanks"
  },
  "source": "meta"
}