{
  "id": 159101,
  "title": "Yet another dataset with external data",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/159101",
  "author_name": "",
  "post_date": "2020-06-16T12:58:05.159352700Z",
  "votes": 15,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone.</p>\n\n<p>I have <a href=\"https://www.kaggle.com/nroman/melanoma-external-malignant-256\">created a dataset</a> which (partially) solves a class imbalance of the provided data.\nAs you know the original competition data has ~98% of benign cases and only ~2% of malignant.</p>\n\n<p>So I have scrapped only malignant melanoma images from the <a href=\"https://challenge2019.isic-archive.com/data.html\">2019 ISIC competition</a> and added them to original dataset, which leaves us with ~86% benign and ~14% malignant cases.</p>\n\n<p>Then all images were center cropped and resized to 256x256.\nI also have a <a href=\"https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet\">kernel</a> which uses this dataset and score 0.925 on the LB.</p>",
  "messages": [
    {
      "id": "888591",
      "postDate": "06/16/2020 12:58:05",
      "content": "<p>Hello everyone.</p>\n\n<p>I have <a href=\"https://www.kaggle.com/nroman/melanoma-external-malignant-256\">created a dataset</a> which (partially) solves a class imbalance of the provided data.\nAs you know the original competition data has ~98% of benign cases and only ~2% of malignant.</p>\n\n<p>So I have scrapped only malignant melanoma images from the <a href=\"https://challenge2019.isic-archive.com/data.html\">2019 ISIC competition</a> and added them to original dataset, which leaves us with ~86% benign and ~14% malignant cases.</p>\n\n<p>Then all images were center cropped and resized to 256x256.\nI also have a <a href=\"https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet\">kernel</a> which uses this dataset and score 0.925 on the LB.</p>",
      "rawMarkdown": "Hello everyone.\n\nI have [created a dataset](https://www.kaggle.com/nroman/melanoma-external-malignant-256) which (partially) solves a class imbalance of the provided data.\nAs you know the original competition data has ~98% of benign cases and only ~2% of malignant.\n\nSo I have scrapped only malignant melanoma images from the [2019 ISIC competition](https://challenge2019.isic-archive.com/data.html) and added them to original dataset, which leaves us with ~86% benign and ~14% malignant cases.\n\nThen all images were center cropped and resized to 256x256.\nI also have a [kernel](https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet) which uses this dataset and score 0.925 on the LB.",
      "votes": null
    },
    {
      "id": "889334",
      "postDate": "06/16/2020 23:13:35",
      "content": "<p>I downloaded the data and used it in colab but many images are missing from the folder but are there in the csv</p>",
      "rawMarkdown": "I downloaded the data and used it in colab but many images are missing from the folder but are there in the csv",
      "votes": null
    },
    {
      "id": "891702",
      "postDate": "06/18/2020 11:54:24",
      "content": "<p><a href=\"/nroman\">@nroman</a>  i have used the above dataset that you mentioned without TTA. I  have trained  a efficientnet-b1.  The validation score is around is quite good around 0.95 on 5 fold. But the LB score is not at all good.</p>",
      "rawMarkdown": "nroman  i have used the above dataset that you mentioned without TTA. I  have trained  a efficientnet-b1.  The validation score is around is quite good around 0.95 on 5 fold. But the LB score is not at all good.",
      "votes": null
    },
    {
      "id": "891850",
      "postDate": "06/18/2020 13:48:16",
      "content": "<p>Same experience, using EfficientNet B6 with Metadata. Validation 0.9+ but LB 0.85-.87 range.</p>",
      "rawMarkdown": "Same experience, using EfficientNet B6 with Metadata. Validation 0.9+ but LB 0.85-.87 range.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 889334,
      "author_name": "rajprakrit",
      "author_url": "",
      "post_date": "06/16/2020 23:13:35",
      "content": "<p>I downloaded the data and used it in colab but many images are missing from the folder but are there in the csv</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 891702,
      "author_name": "ayushpatidar",
      "author_url": "",
      "post_date": "06/18/2020 11:54:24",
      "content": "<p><a href=\"/nroman\">@nroman</a>  i have used the above dataset that you mentioned without TTA. I  have trained  a efficientnet-b1.  The validation score is around is quite good around 0.95 on 5 fold. But the LB score is not at all good.</p>",
      "votes": null,
      "replies": [
        {
          "id": 891850,
          "author_name": "richardepstein",
          "author_url": "",
          "post_date": "06/18/2020 13:48:16",
          "content": "<p>Same experience, using EfficientNet B6 with Metadata. Validation 0.9+ but LB 0.85-.87 range.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "888591": "Hello everyone.\n\nI have [created a dataset](https://www.kaggle.com/nroman/melanoma-external-malignant-256) which (partially) solves a class imbalance of the provided data.\nAs you know the original competition data has ~98% of benign cases and only ~2% of malignant.\n\nSo I have scrapped only malignant melanoma images from the [2019 ISIC competition](https://challenge2019.isic-archive.com/data.html) and added them to original dataset, which leaves us with ~86% benign and ~14% malignant cases.\n\nThen all images were center cropped and resized to 256x256.\nI also have a [kernel](https://www.kaggle.com/nroman/melanoma-pytorch-starter-efficientnet) which uses this dataset and score 0.925 on the LB.",
    "889334": "I downloaded the data and used it in colab but many images are missing from the folder but are there in the csv",
    "891702": "nroman  i have used the above dataset that you mentioned without TTA. I  have trained  a efficientnet-b1.  The validation score is around is quite good around 0.95 on 5 fold. But the LB score is not at all good.",
    "891850": "Same experience, using EfficientNet B6 with Metadata. Validation 0.9+ but LB 0.85-.87 range."
  },
  "source": "meta"
}