{
  "id": 154271,
  "title": "Welcome!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/154271",
  "author_name": "Phil Culliton",
  "post_date": "2020-05-27T22:01:30.323000",
  "votes": 15,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Welcome to the SIIM-ISIC Melanoma Classification challenge!</p>\n\n<p>In this competition, you’ll identify melanoma / malignancy in images of skin lesions. Using patient-level contextual information is provided, which may help with use of this technology in the field.</p>\n\n<p>Note that there are a couple of differences between this and normal competitions:</p>\n\n<ul>\n<li>We are allowing 3 final selected submissions, to allow for more eligible submissions towards special prizes.</li>\n<li>All external data is allowed, but only those using publicly, freely available data (for use that includes academic or research purposes) are eligible for prizes.</li>\n<li>There are some extra winners’ obligations - please check the Prizes tab for details.</li>\n</ul>\n\n<p>Additionally, we've provided the data in DICOM format (as the preferred format by clinicians/researchers/radiologists), as well as TFRecord (for ease of use with accelerators) and plain JPEG.</p>\n\n<p>Enjoy, and good luck!</p>\n\n<blockquote>\n  <p><strong>Remember</strong>: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\">Kaggle community guidelines</a>.</p>\n</blockquote>",
  "messages": [
    {
      "id": 864224,
      "postDate": "2020-05-27T22:01:30.323Z",
      "content": "<p>Welcome to the SIIM-ISIC Melanoma Classification challenge!</p>\n\n<p>In this competition, you’ll identify melanoma / malignancy in images of skin lesions. Using patient-level contextual information is provided, which may help with use of this technology in the field.</p>\n\n<p>Note that there are a couple of differences between this and normal competitions:</p>\n\n<ul>\n<li>We are allowing 3 final selected submissions, to allow for more eligible submissions towards special prizes.</li>\n<li>All external data is allowed, but only those using publicly, freely available data (for use that includes academic or research purposes) are eligible for prizes.</li>\n<li>There are some extra winners’ obligations - please check the Prizes tab for details.</li>\n</ul>\n\n<p>Additionally, we've provided the data in DICOM format (as the preferred format by clinicians/researchers/radiologists), as well as TFRecord (for ease of use with accelerators) and plain JPEG.</p>\n\n<p>Enjoy, and good luck!</p>\n\n<blockquote>\n  <p><strong>Remember</strong>: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our <a href=\"https://www.kaggle.com/community-guidelines\">Kaggle community guidelines</a>.</p>\n</blockquote>",
      "rawMarkdown": "Welcome to the SIIM-ISIC Melanoma Classification challenge!\n\nIn this competition, you’ll identify melanoma / malignancy in images of skin lesions. Using patient-level contextual information is provided, which may help with use of this technology in the field.\n\nNote that there are a couple of differences between this and normal competitions:\n\n- We are allowing 3 final selected submissions, to allow for more eligible submissions towards special prizes.\n- All external data is allowed, but only those using publicly, freely available data (for use that includes academic or research purposes) are eligible for prizes.\n- There are some extra winners’ obligations - please check the Prizes tab for details.\n\nAdditionally, we've provided the data in DICOM format (as the preferred format by clinicians/researchers/radiologists), as well as TFRecord (for ease of use with accelerators) and plain JPEG.\n\nEnjoy, and good luck!\n\n&gt; **Remember**: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).\n",
      "votes": 14
    },
    {
      "id": 932650,
      "postDate": "2020-07-17T07:27:21.553Z",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> </p>\n\n<p>Hi\nI just thought about, why there are no images of ppl with black skin. Is there a special reason? Our models will not work well on samples with black skin. Aren't there enough images? Is it too difficult to mix? I think there should be a competition focused on that. Are there datasets with these kinds of images? </p>\n\n<p>best Roman</p>",
      "rawMarkdown": "@philculliton \n\nHi\nI just thought about, why there are no images of ppl with black skin. Is there a special reason? Our models will not work well on samples with black skin. Aren't there enough images? Is it too difficult to mix? I think there should be a competition focused on that. Are there datasets with these kinds of images? \n\nbest Roman",
      "votes": 3,
      "replies": [
        {
          "id": 936757,
          "postDate": "2020-07-20T13:56:35.263Z",
          "content": "<p>Thanks <a href=\"/romanweilguny\">@romanweilguny</a>. It is part of the ongoing mission of the ISIC group to collect images of diverse skin types. There are many reasons that we are still discovering and researching why these images have not been easy to include, but we hope that in future years with future challenges, we will be able to have more diverse datasets. </p>",
          "rawMarkdown": "Thanks @romanweilguny. It is part of the ongoing mission of the ISIC group to collect images of diverse skin types. There are many reasons that we are still discovering and researching why these images have not been easy to include, but we hope that in future years with future challenges, we will be able to have more diverse datasets. "
        }
      ]
    },
    {
      "id": 868063,
      "postDate": "2020-05-30T21:36:02.747Z",
      "content": "<p><a href=\"/veronicarotemberg\">@veronicarotemberg</a> I am very confused with this dataset as it differs quite a lot from a \"golden\" ISIC standard. Over 80% of images have \"unknown\" diagnosis. And i do not see any carcinomas sarcomas or other malignant manifestations. Is this dataset cherry picked not to contain carcinomas&amp;sarcomas? Are we 100% certain that all \"unknown\" diagnosis are of benign origin?</p>\n\n<p>IMHO, having only melanoma as malignant samples would make all our modelling efforts useless.</p>",
      "rawMarkdown": "@veronicarotemberg I am very confused with this dataset as it differs quite a lot from a \"golden\" ISIC standard. Over 80% of images have \"unknown\" diagnosis. And i do not see any carcinomas sarcomas or other malignant manifestations. Is this dataset cherry picked not to contain carcinomas&amp;sarcomas? Are we 100% certain that all \"unknown\" diagnosis are of benign origin?\n\nIMHO, having only melanoma as malignant samples would make all our modelling efforts useless.\n\n",
      "votes": 3
    },
    {
      "id": 865355,
      "postDate": "2020-05-28T15:28:05.287Z",
      "content": "<p>We saw a question about External Data and I wanted to post here as well - we've done our best to make this dataset completely unique, meaning that other images at <a href=\"https://isic-archive.com/\">https://isic-archive.com/</a> or via the ISIC API are not expected to overlap with the ones provided here. </p>\n\n<p>Due to the retrospective nature of our database queries, there may be a few <em>lesions</em> imaged with different photographic equipment that are repeated in the train set here and in train of prior ISIC Challenges, but those are expected to be very few in number. </p>\n\n<p>We did this so as to generate a completely new dataset with patient-contextual information (aka multiple images of lesions on the same patient), a type of labeling that hasn't been performed on the previous datasets and which we suspect may improve performance (as it does for clinicians who are trying to identify melanomas). </p>",
      "rawMarkdown": "We saw a question about External Data and I wanted to post here as well - we've done our best to make this dataset completely unique, meaning that other images at https://isic-archive.com/ or via the ISIC API are not expected to overlap with the ones provided here. \n\nDue to the retrospective nature of our database queries, there may be a few *lesions* imaged with different photographic equipment that are repeated in the train set here and in train of prior ISIC Challenges, but those are expected to be very few in number. \n\nWe did this so as to generate a completely new dataset with patient-contextual information (aka multiple images of lesions on the same patient), a type of labeling that hasn't been performed on the previous datasets and which we suspect may improve performance (as it does for clinicians who are trying to identify melanomas). ",
      "votes": 3
    },
    {
      "id": 962019,
      "postDate": "2020-08-07T18:03:45.583Z",
      "content": "<p>Thanks to all for your submissions to this competition - we are so grateful for your efforts and excited about the potential contributions to melanoma diagnosis in a clinical setting! </p>\n<p>For the purposes of analyzing the clinical implications of your submissions, we are planning to review and analyze the CSVs for submissions of the top 100 participants. </p>\n<p>Please let <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> know (julia@kaggle.com) if you would like to OPT OUT of sharing your submission CSV with us. We will not publicize any data relating to this challenge that includes team names. </p>\n<p>Thanks so much again,</p>\n<p>Veronica</p>",
      "rawMarkdown": "Thanks to all for your submissions to this competition - we are so grateful for your efforts and excited about the potential contributions to melanoma diagnosis in a clinical setting! \n\nFor the purposes of analyzing the clinical implications of your submissions, we are planning to review and analyze the CSVs for submissions of the top 100 participants. \n\nPlease let @juliaelliott know (julia@kaggle.com) if you would like to OPT OUT of sharing your submission CSV with us. We will not publicize any data relating to this challenge that includes team names. \n\nThanks so much again,\n\nVeronica",
      "votes": 1
    },
    {
      "id": 864259,
      "postDate": "2020-05-27T23:02:26.270Z",
      "content": "<ol>\n<li>Are TPUs allowed?</li>\n<li>Open Internet Kernels?</li>\n</ol>",
      "rawMarkdown": "1. Are TPUs allowed?\n2. Open Internet Kernels?",
      "votes": 1,
      "replies": [
        {
          "id": 864287,
          "postDate": "2020-05-27T23:53:03.487Z",
          "content": "<p><a href=\"/yeayates21\">@yeayates21</a> \n1. Yes, TPUs can be used. \n2. This is not a code competition; it is a prediction submission competition. So no specific constraints apply to how you submit those predictions, including allowing your notebook/code to access the internet. However, standard rules around not hand-labeling the test set apply.</p>",
          "rawMarkdown": "@yeayates21 \n1. Yes, TPUs can be used. \n2. This is not a code competition; it is a prediction submission competition. So no specific constraints apply to how you submit those predictions, including allowing your notebook/code to access the internet. However, standard rules around not hand-labeling the test set apply.",
          "votes": 2
        },
        {
          "id": 864417,
          "postDate": "2020-05-28T01:54:00.477Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> - Thanks so much! 👍 🤘 </p>",
          "rawMarkdown": "@juliaelliott - Thanks so much! 👍 🤘 "
        }
      ]
    },
    {
      "id": 973862,
      "postDate": "2020-08-17T15:09:57.013Z",
      "content": "<p>If I want to use these data in a publication: 1) is it allowed? and 2) Footnote, acknowledgement and/or reference?</p>",
      "rawMarkdown": "If I want to use these data in a publication: 1) is it allowed? and 2) Footnote, acknowledgement and/or reference?",
      "replies": [
        {
          "id": 977321,
          "postDate": "2020-08-19T12:15:02.390Z",
          "content": "<p>So, just a reference to <a href=\"https://doi.org/10.34970/2020-ds01\" target=\"_blank\">https://doi.org/10.34970/2020-ds01</a> is sufficient?</p>",
          "rawMarkdown": "So, just a reference to https://doi.org/10.34970/2020-ds01 is sufficient?"
        },
        {
          "id": 979567,
          "postDate": "2020-08-21T00:33:39.870Z",
          "content": "<p>Yes, please reference <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\" target=\"_blank\">section 7A of the Competition Rules</a> for details about attribution required when using the dataset.</p>",
          "rawMarkdown": "Yes, please reference [section 7A of the Competition Rules](https://www.kaggle.com/c/siim-isic-melanoma-classification/rules) for details about attribution required when using the dataset.",
          "votes": 1
        }
      ]
    },
    {
      "id": 973575,
      "postDate": "2020-08-17T11:28:29.190Z",
      "content": "<p>Hi          </p>",
      "rawMarkdown": "Hi          "
    },
    {
      "id": 943233,
      "postDate": "2020-07-24T08:12:20.790Z",
      "content": "<p><a href=\"/philculliton\">@philculliton</a>\nAs submission can be done through direct upload on current test data which is just 30% of all test data.\nWhen will we have rest 70% and will submission be allowed as we can do now.\nThanks for feedback.</p>",
      "rawMarkdown": "@philculliton\nAs submission can be done through direct upload on current test data which is just 30% of all test data.\nWhen will we have rest 70% and will submission be allowed as we can do now.\nThanks for feedback.",
      "replies": [
        {
          "id": 943624,
          "postDate": "2020-07-24T13:30:26.930Z",
          "content": "<p>The test data in the test.csv file and image directories is 100% of the data. Your public Leaderboard score is only on 30% of that data. This is to prevent you from overfitting your model to do better on the test data than it would on unseen data.</p>\n\n<p>The final score is based on the other 70% of the test data. So you are already predicting the private leaderboard data, but you don't get to see your score.</p>\n\n<p>Your score on the private leaderboard isn't reveled until the end of the competition.</p>\n\n<p>So, while you could have a great score on the public Leaderboard because you tune your model and parameters and send in a lot of submissions, you might not be improving on the real private leaderboard, which determines the final ranking.</p>\n\n<p>Other competitions have different structures. Some require you to upload your model and release new test data in the last two weeks and you re-run your models on the new test data.</p>\n\n<p>Other competitions require you to run your model in a notebook and you never see the test data (the current OSIC Pulmonary Fibrosis competition uses this method).</p>\n\n<p>Hope this helps.</p>\n\n<p>-Rich</p>",
          "rawMarkdown": "The test data in the test.csv file and image directories is 100% of the data. Your public Leaderboard score is only on 30% of that data. This is to prevent you from overfitting your model to do better on the test data than it would on unseen data.\n\nThe final score is based on the other 70% of the test data. So you are already predicting the private leaderboard data, but you don't get to see your score.\n\nYour score on the private leaderboard isn't reveled until the end of the competition.\n\nSo, while you could have a great score on the public Leaderboard because you tune your model and parameters and send in a lot of submissions, you might not be improving on the real private leaderboard, which determines the final ranking.\n\nOther competitions have different structures. Some require you to upload your model and release new test data in the last two weeks and you re-run your models on the new test data.\n\nOther competitions require you to run your model in a notebook and you never see the test data (the current OSIC Pulmonary Fibrosis competition uses this method).\n\nHope this helps.\n\n-Rich",
          "votes": 1
        }
      ]
    },
    {
      "id": 925155,
      "postDate": "2020-07-11T21:08:25.353Z",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> \nCan we use test data as well for pseudo labeling ?</p>",
      "rawMarkdown": "@juliaelliott \nCan we use test data as well for pseudo labeling ?",
      "replies": [
        {
          "id": 930993,
          "postDate": "2020-07-15T21:46:32.947Z",
          "content": "<p>You may pseudolabel the test data, but you cannot hand-label the test data. So if you pseudolabel, it should not make use of any test set hand-labeling.</p>",
          "rawMarkdown": "You may pseudolabel the test data, but you cannot hand-label the test data. So if you pseudolabel, it should not make use of any test set hand-labeling."
        },
        {
          "id": 938837,
          "postDate": "2020-07-21T19:38:43.167Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 938901,
          "postDate": "2020-07-21T20:54:51.673Z",
          "content": "<p><a href=\"https://www.kaggle.com/rajnishc3\" target=\"_blank\">@rajnishc3</a> - Because this is a Predictions competition, this is the full set of images for the test set, which has been split for the public and private leaderboard.</p>",
          "rawMarkdown": "@rajnishc3 - Because this is a Predictions competition, this is the full set of images for the test set, which has been split for the public and private leaderboard."
        }
      ]
    },
    {
      "id": 916807,
      "postDate": "2020-07-06T03:58:52.733Z",
      "content": "<p>Hi</p>",
      "rawMarkdown": "Hi"
    },
    {
      "id": 880704,
      "postDate": "2020-06-10T14:05:38.590Z",
      "content": "<p>The rules say:</p>\n\n<blockquote>\n  <p>To receive a Prize, a potential winner must have used External Data available to use by all participants of the competition for purposes of the competition, for use that includes research or academic purposes, at no cost to the other participants (\"Public External Data\").</p>\n</blockquote>\n\n<p>Do I understand correctly that submissions that do not use any Public External Data but rather only the competition data, are not eligible to win prizes?</p>",
      "rawMarkdown": "The rules say:\n\n&gt; To receive a Prize, a potential winner must have used External Data available to use by all participants of the competition for purposes of the competition, for use that includes research or academic purposes, at no cost to the other participants (\"Public External Data\").\n\nDo I understand correctly that submissions that do not use any Public External Data but rather only the competition data, are not eligible to win prizes?\n\n\n",
      "replies": [
        {
          "id": 880855,
          "postDate": "2020-06-10T15:50:47.023Z",
          "content": "<p><a href=\"/byoussin\">@byoussin</a> Not quite. This means that you may use external data, but it must be publicly and freely available (for research or academic use) in order to be eligible for prizes. Publicly and freely available external data is not limited to the competition’s data, as you’ll see many datasets shared in the external data thread which are also available that meet this requirement. But if you use data that is proprietary, is gated to any participants, and/or comes at a cost, then you will not be eligible for prizes.</p>",
          "rawMarkdown": "@byoussin Not quite. This means that you may use external data, but it must be publicly and freely available (for research or academic use) in order to be eligible for prizes. Publicly and freely available external data is not limited to the competition’s data, as you’ll see many datasets shared in the external data thread which are also available that meet this requirement. But if you use data that is proprietary, is gated to any participants, and/or comes at a cost, then you will not be eligible for prizes."
        }
      ]
    },
    {
      "id": 868237,
      "postDate": "2020-05-31T04:52:32.373Z",
      "content": "<p>Hi, what if I use only image processing algorithms, to make predictions, i mean, no models generated at all, just predictions. it is allowed?</p>",
      "rawMarkdown": "Hi, what if I use only image processing algorithms, to make predictions, i mean, no models generated at all, just predictions. it is allowed?\n",
      "replies": [
        {
          "id": 880859,
          "postDate": "2020-06-10T15:53:23.333Z",
          "content": "<p><a href=\"/mithos\">@mithos</a> If those algorithms are created by you or licensed in a way where you are permitted to use them (public pre-trained models, for example), then sure.</p>",
          "rawMarkdown": "@mithos If those algorithms are created by you or licensed in a way where you are permitted to use them (public pre-trained models, for example), then sure."
        }
      ]
    },
    {
      "id": 865300,
      "postDate": "2020-05-28T14:38:09.583Z",
      "content": "<p>Hello,\nTwo questions:\nIn the sentence:\n'Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma.'\nDo you mean that current models do not take into account the localization of the skin lezion on the body ?\nIs it also all what you mean by \"Existing AI approaches have not adequately considered this clinical frame of reference\"  ?</p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "Hello,\nTwo questions:\nIn the sentence:\n'Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma.'\nDo you mean that current models do not take into account the localization of the skin lezion on the body ?\nIs it also all what you mean by \"Existing AI approaches have not adequately considered this clinical frame of reference\"  ?\n\nThank you!",
      "replies": [
        {
          "id": 865360,
          "postDate": "2020-05-28T15:29:37.273Z",
          "content": "<p>We mean that prior algorithms do not take into account multiple lesions or images that are from the same patient, while clinicians evaluate all lesions on one patient when making a malignancy determination. </p>",
          "rawMarkdown": "We mean that prior algorithms do not take into account multiple lesions or images that are from the same patient, while clinicians evaluate all lesions on one patient when making a malignancy determination. ",
          "votes": 1
        },
        {
          "id": 866168,
          "postDate": "2020-05-29T07:10:40.163Z",
          "content": "<p>Perfect. Thank you!</p>",
          "rawMarkdown": "Perfect. Thank you!"
        },
        {
          "id": 866588,
          "postDate": "2020-05-29T14:22:58.837Z",
          "content": "<p>A related question: were additional patients information such as age and sex taken into account in prior algorithms ?</p>",
          "rawMarkdown": "A related question: were additional patients information such as age and sex taken into account in prior algorithms ?"
        },
        {
          "id": 875343,
          "postDate": "2020-06-05T17:19:09.367Z",
          "content": "<p>Yes - in the 2019 challenge:</p>\n\n<p><a href=\"https://challenge2019.isic-archive.com/leaderboard.html\">https://challenge2019.isic-archive.com/leaderboard.html</a></p>\n\n<p>Recommend taking a look at the excellent thread on using both images and tabular data for approaches taken this year!</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155251\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155251</a></p>\n\n<p>Clinically, we expect that tabular data and imaging data will contain some overlapping and some independent information, and that seems to be borne out this year and last year as well. </p>",
          "rawMarkdown": "Yes - in the 2019 challenge:\n\nhttps://challenge2019.isic-archive.com/leaderboard.html\n\nRecommend taking a look at the excellent thread on using both images and tabular data for approaches taken this year!\n\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155251\n\nClinically, we expect that tabular data and imaging data will contain some overlapping and some independent information, and that seems to be borne out this year and last year as well. "
        }
      ]
    },
    {
      "id": 864843,
      "postDate": "2020-05-28T08:17:00.417Z",
      "content": "<p>A little bit confusing:</p>\n\n<p>Reading competition description on the competition Data page we can see:</p>\n\n<p>What am I predicting?\nYou are predicting a binary target for each image. Your model should predict the probability (floating point) that the lesion in the image is malignant (the target). In the training data, train.csv, the value 0 denotes benign, and 1 indicates malignant. Predictions should be floating point values between 0.0 and 1.0, with 0.5 as a binary decision threshold.</p>\n\n<p>But the actual metric is AUC, which is invariant to predicted probability distribution. </p>",
      "rawMarkdown": "A little bit confusing:\n\nReading competition description on the competition Data page we can see:\n\nWhat am I predicting?\nYou are predicting a binary target for each image. Your model should predict the probability (floating point) that the lesion in the image is malignant (the target). In the training data, train.csv, the value 0 denotes benign, and 1 indicates malignant. Predictions should be floating point values between 0.0 and 1.0, with 0.5 as a binary decision threshold.\n\n\nBut the actual metric is AUC, which is invariant to predicted probability distribution. \n",
      "replies": [
        {
          "id": 875295,
          "postDate": "2020-06-05T16:26:15.107Z",
          "content": "<p><a href=\"/raddar\">@raddar</a> This description has since been updated, as we recognized it was confusing.</p>",
          "rawMarkdown": "@raddar This description has since been updated, as we recognized it was confusing."
        }
      ]
    },
    {
      "id": 941794,
      "postDate": "2020-07-23T12:17:24.327Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 941839,
          "postDate": "2020-07-23T12:48:56.380Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 948454,
          "postDate": "2020-07-28T01:01:27.607Z",
          "content": "<p><a href=\"/usergupsak\">@usergupsak</a> <a href=\"/synked\">@synked</a> Are you referring to this notebook? <a href=\"https://www.kaggle.com/mahmoodhoseini/siim-isic-notebook-0-951-top-6\">https://www.kaggle.com/mahmoodhoseini/siim-isic-notebook-0-951-top-6</a> -- Because it's still 21 days until the end of the competition, there is no prohibition against public notebook sharing. </p>",
          "rawMarkdown": "@usergupsak @synked Are you referring to this notebook? https://www.kaggle.com/mahmoodhoseini/siim-isic-notebook-0-951-top-6 -- Because it's still 21 days until the end of the competition, there is no prohibition against public notebook sharing. "
        }
      ]
    },
    {
      "id": 864442,
      "postDate": "2020-05-28T02:13:46.973Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 932650,
      "author_name": "Roman Weilguny",
      "author_url": "",
      "post_date": "2020-07-17T07:27:21.553000",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> </p>\n\n<p>Hi\nI just thought about, why there are no images of ppl with black skin. Is there a special reason? Our models will not work well on samples with black skin. Aren't there enough images? Is it too difficult to mix? I think there should be a competition focused on that. Are there datasets with these kinds of images? </p>\n\n<p>best Roman</p>",
      "votes": 3,
      "replies": [
        {
          "id": 936757,
          "author_name": "Veronica Rotemberg",
          "author_url": "",
          "post_date": "2020-07-20T13:56:35.263000",
          "content": "<p>Thanks <a href=\"/romanweilguny\">@romanweilguny</a>. It is part of the ongoing mission of the ISIC group to collect images of diverse skin types. There are many reasons that we are still discovering and researching why these images have not been easy to include, but we hope that in future years with future challenges, we will be able to have more diverse datasets. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 868063,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2020-05-30T21:36:02.747000",
      "content": "<p><a href=\"/veronicarotemberg\">@veronicarotemberg</a> I am very confused with this dataset as it differs quite a lot from a \"golden\" ISIC standard. Over 80% of images have \"unknown\" diagnosis. And i do not see any carcinomas sarcomas or other malignant manifestations. Is this dataset cherry picked not to contain carcinomas&amp;sarcomas? Are we 100% certain that all \"unknown\" diagnosis are of benign origin?</p>\n\n<p>IMHO, having only melanoma as malignant samples would make all our modelling efforts useless.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 865355,
      "author_name": "Veronica Rotemberg",
      "author_url": "",
      "post_date": "2020-05-28T15:28:05.287000",
      "content": "<p>We saw a question about External Data and I wanted to post here as well - we've done our best to make this dataset completely unique, meaning that other images at <a href=\"https://isic-archive.com/\">https://isic-archive.com/</a> or via the ISIC API are not expected to overlap with the ones provided here. </p>\n\n<p>Due to the retrospective nature of our database queries, there may be a few <em>lesions</em> imaged with different photographic equipment that are repeated in the train set here and in train of prior ISIC Challenges, but those are expected to be very few in number. </p>\n\n<p>We did this so as to generate a completely new dataset with patient-contextual information (aka multiple images of lesions on the same patient), a type of labeling that hasn't been performed on the previous datasets and which we suspect may improve performance (as it does for clinicians who are trying to identify melanomas). </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 962019,
      "author_name": "Veronica Rotemberg",
      "author_url": "",
      "post_date": "2020-08-07T18:03:45.583000",
      "content": "<p>Thanks to all for your submissions to this competition - we are so grateful for your efforts and excited about the potential contributions to melanoma diagnosis in a clinical setting! </p>\n<p>For the purposes of analyzing the clinical implications of your submissions, we are planning to review and analyze the CSVs for submissions of the top 100 participants. </p>\n<p>Please let <a href=\"https://www.kaggle.com/juliaelliott\" target=\"_blank\">@juliaelliott</a> know (julia@kaggle.com) if you would like to OPT OUT of sharing your submission CSV with us. We will not publicize any data relating to this challenge that includes team names. </p>\n<p>Thanks so much again,</p>\n<p>Veronica</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 864259,
      "author_name": "Matt Yates",
      "author_url": "",
      "post_date": "2020-05-27T23:02:26.270000",
      "content": "<ol>\n<li>Are TPUs allowed?</li>\n<li>Open Internet Kernels?</li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 864287,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-05-27T23:53:03.487000",
          "content": "<p><a href=\"/yeayates21\">@yeayates21</a> \n1. Yes, TPUs can be used. \n2. This is not a code competition; it is a prediction submission competition. So no specific constraints apply to how you submit those predictions, including allowing your notebook/code to access the internet. However, standard rules around not hand-labeling the test set apply.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 864417,
          "author_name": "Matt Yates",
          "author_url": "",
          "post_date": "2020-05-28T01:54:00.477000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> - Thanks so much! 👍 🤘 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 973862,
      "author_name": "Tord Malmgren",
      "author_url": "",
      "post_date": "2020-08-17T15:09:57.013000",
      "content": "<p>If I want to use these data in a publication: 1) is it allowed? and 2) Footnote, acknowledgement and/or reference?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 977321,
          "author_name": "Tord Malmgren",
          "author_url": "",
          "post_date": "2020-08-19T12:15:02.390000",
          "content": "<p>So, just a reference to <a href=\"https://doi.org/10.34970/2020-ds01\" target=\"_blank\">https://doi.org/10.34970/2020-ds01</a> is sufficient?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 979567,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-08-21T00:33:39.870000",
          "content": "<p>Yes, please reference <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\" target=\"_blank\">section 7A of the Competition Rules</a> for details about attribution required when using the dataset.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 973575,
      "author_name": "JiaCheng121211",
      "author_url": "",
      "post_date": "2020-08-17T11:28:29.190000",
      "content": "<p>Hi          </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 943233,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2020-07-24T08:12:20.790000",
      "content": "<p><a href=\"/philculliton\">@philculliton</a>\nAs submission can be done through direct upload on current test data which is just 30% of all test data.\nWhen will we have rest 70% and will submission be allowed as we can do now.\nThanks for feedback.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 943624,
          "author_name": "quadcore/Richard Epstein",
          "author_url": "",
          "post_date": "2020-07-24T13:30:26.930000",
          "content": "<p>The test data in the test.csv file and image directories is 100% of the data. Your public Leaderboard score is only on 30% of that data. This is to prevent you from overfitting your model to do better on the test data than it would on unseen data.</p>\n\n<p>The final score is based on the other 70% of the test data. So you are already predicting the private leaderboard data, but you don't get to see your score.</p>\n\n<p>Your score on the private leaderboard isn't reveled until the end of the competition.</p>\n\n<p>So, while you could have a great score on the public Leaderboard because you tune your model and parameters and send in a lot of submissions, you might not be improving on the real private leaderboard, which determines the final ranking.</p>\n\n<p>Other competitions have different structures. Some require you to upload your model and release new test data in the last two weeks and you re-run your models on the new test data.</p>\n\n<p>Other competitions require you to run your model in a notebook and you never see the test data (the current OSIC Pulmonary Fibrosis competition uses this method).</p>\n\n<p>Hope this helps.</p>\n\n<p>-Rich</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 925155,
      "author_name": "Rajnish Chauhan",
      "author_url": "",
      "post_date": "2020-07-11T21:08:25.353000",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> \nCan we use test data as well for pseudo labeling ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 930993,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-07-15T21:46:32.947000",
          "content": "<p>You may pseudolabel the test data, but you cannot hand-label the test data. So if you pseudolabel, it should not make use of any test set hand-labeling.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938837,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-21T19:38:43.167000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 938901,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-07-21T20:54:51.673000",
          "content": "<p><a href=\"https://www.kaggle.com/rajnishc3\" target=\"_blank\">@rajnishc3</a> - Because this is a Predictions competition, this is the full set of images for the test set, which has been split for the public and private leaderboard.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 916807,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-06T03:58:52.733000",
      "content": "<p>Hi</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 880704,
      "author_name": "byoussin",
      "author_url": "",
      "post_date": "2020-06-10T14:05:38.590000",
      "content": "<p>The rules say:</p>\n\n<blockquote>\n  <p>To receive a Prize, a potential winner must have used External Data available to use by all participants of the competition for purposes of the competition, for use that includes research or academic purposes, at no cost to the other participants (\"Public External Data\").</p>\n</blockquote>\n\n<p>Do I understand correctly that submissions that do not use any Public External Data but rather only the competition data, are not eligible to win prizes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 880855,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-10T15:50:47.023000",
          "content": "<p><a href=\"/byoussin\">@byoussin</a> Not quite. This means that you may use external data, but it must be publicly and freely available (for research or academic use) in order to be eligible for prizes. Publicly and freely available external data is not limited to the competition’s data, as you’ll see many datasets shared in the external data thread which are also available that meet this requirement. But if you use data that is proprietary, is gated to any participants, and/or comes at a cost, then you will not be eligible for prizes.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 868237,
      "author_name": "mithos",
      "author_url": "",
      "post_date": "2020-05-31T04:52:32.373000",
      "content": "<p>Hi, what if I use only image processing algorithms, to make predictions, i mean, no models generated at all, just predictions. it is allowed?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 880859,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-10T15:53:23.333000",
          "content": "<p><a href=\"/mithos\">@mithos</a> If those algorithms are created by you or licensed in a way where you are permitted to use them (public pre-trained models, for example), then sure.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 865300,
      "author_name": "ChristianColot",
      "author_url": "",
      "post_date": "2020-05-28T14:38:09.583000",
      "content": "<p>Hello,\nTwo questions:\nIn the sentence:\n'Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma.'\nDo you mean that current models do not take into account the localization of the skin lezion on the body ?\nIs it also all what you mean by \"Existing AI approaches have not adequately considered this clinical frame of reference\"  ?</p>\n\n<p>Thank you!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 865360,
          "author_name": "Veronica Rotemberg",
          "author_url": "",
          "post_date": "2020-05-28T15:29:37.273000",
          "content": "<p>We mean that prior algorithms do not take into account multiple lesions or images that are from the same patient, while clinicians evaluate all lesions on one patient when making a malignancy determination. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 866168,
          "author_name": "ChristianColot",
          "author_url": "",
          "post_date": "2020-05-29T07:10:40.163000",
          "content": "<p>Perfect. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 866588,
          "author_name": "ChristianColot",
          "author_url": "",
          "post_date": "2020-05-29T14:22:58.837000",
          "content": "<p>A related question: were additional patients information such as age and sex taken into account in prior algorithms ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 875343,
          "author_name": "Veronica Rotemberg",
          "author_url": "",
          "post_date": "2020-06-05T17:19:09.367000",
          "content": "<p>Yes - in the 2019 challenge:</p>\n\n<p><a href=\"https://challenge2019.isic-archive.com/leaderboard.html\">https://challenge2019.isic-archive.com/leaderboard.html</a></p>\n\n<p>Recommend taking a look at the excellent thread on using both images and tabular data for approaches taken this year!</p>\n\n<p><a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155251\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155251</a></p>\n\n<p>Clinically, we expect that tabular data and imaging data will contain some overlapping and some independent information, and that seems to be borne out this year and last year as well. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 864843,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "2020-05-28T08:17:00.417000",
      "content": "<p>A little bit confusing:</p>\n\n<p>Reading competition description on the competition Data page we can see:</p>\n\n<p>What am I predicting?\nYou are predicting a binary target for each image. Your model should predict the probability (floating point) that the lesion in the image is malignant (the target). In the training data, train.csv, the value 0 denotes benign, and 1 indicates malignant. Predictions should be floating point values between 0.0 and 1.0, with 0.5 as a binary decision threshold.</p>\n\n<p>But the actual metric is AUC, which is invariant to predicted probability distribution. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 875295,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-06-05T16:26:15.107000",
          "content": "<p><a href=\"/raddar\">@raddar</a> This description has since been updated, as we recognized it was confusing.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 941794,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-23T12:17:24.327000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 941839,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-07-23T12:48:56.380000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 948454,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-07-28T01:01:27.607000",
          "content": "<p><a href=\"/usergupsak\">@usergupsak</a> <a href=\"/synked\">@synked</a> Are you referring to this notebook? <a href=\"https://www.kaggle.com/mahmoodhoseini/siim-isic-notebook-0-951-top-6\">https://www.kaggle.com/mahmoodhoseini/siim-isic-notebook-0-951-top-6</a> -- Because it's still 21 days until the end of the competition, there is no prohibition against public notebook sharing. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 864442,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-28T02:13:46.973000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "864224": "Welcome to the SIIM-ISIC Melanoma Classification challenge!\n\nIn this competition, you’ll identify melanoma / malignancy in images of skin lesions. Using patient-level contextual information is provided, which may help with use of this technology in the field.\n\nNote that there are a couple of differences between this and normal competitions:\n\n- We are allowing 3 final selected submissions, to allow for more eligible submissions towards special prizes.\n- All external data is allowed, but only those using publicly, freely available data (for use that includes academic or research purposes) are eligible for prizes.\n- There are some extra winners’ obligations - please check the Prizes tab for details.\n\nAdditionally, we've provided the data in DICOM format (as the preferred format by clinicians/researchers/radiologists), as well as TFRecord (for ease of use with accelerators) and plain JPEG.\n\nEnjoy, and good luck!\n\n&gt; **Remember**: Kaggle is for everyone. Whether you're teaming up or sharing tips in the competition forum, we expect everyone to follow our [Kaggle community guidelines](https://www.kaggle.com/community-guidelines).\n",
    "932650": "@philculliton \n\nHi\nI just thought about, why there are no images of ppl with black skin. Is there a special reason? Our models will not work well on samples with black skin. Aren't there enough images? Is it too difficult to mix? I think there should be a competition focused on that. Are there datasets with these kinds of images? \n\nbest Roman",
    "868063": "@veronicarotemberg I am very confused with this dataset as it differs quite a lot from a \"golden\" ISIC standard. Over 80% of images have \"unknown\" diagnosis. And i do not see any carcinomas sarcomas or other malignant manifestations. Is this dataset cherry picked not to contain carcinomas&amp;sarcomas? Are we 100% certain that all \"unknown\" diagnosis are of benign origin?\n\nIMHO, having only melanoma as malignant samples would make all our modelling efforts useless.\n\n",
    "865355": "We saw a question about External Data and I wanted to post here as well - we've done our best to make this dataset completely unique, meaning that other images at https://isic-archive.com/ or via the ISIC API are not expected to overlap with the ones provided here. \n\nDue to the retrospective nature of our database queries, there may be a few *lesions* imaged with different photographic equipment that are repeated in the train set here and in train of prior ISIC Challenges, but those are expected to be very few in number. \n\nWe did this so as to generate a completely new dataset with patient-contextual information (aka multiple images of lesions on the same patient), a type of labeling that hasn't been performed on the previous datasets and which we suspect may improve performance (as it does for clinicians who are trying to identify melanomas). ",
    "962019": "Thanks to all for your submissions to this competition - we are so grateful for your efforts and excited about the potential contributions to melanoma diagnosis in a clinical setting! \n\nFor the purposes of analyzing the clinical implications of your submissions, we are planning to review and analyze the CSVs for submissions of the top 100 participants. \n\nPlease let @juliaelliott know (julia@kaggle.com) if you would like to OPT OUT of sharing your submission CSV with us. We will not publicize any data relating to this challenge that includes team names. \n\nThanks so much again,\n\nVeronica",
    "864259": "1. Are TPUs allowed?\n2. Open Internet Kernels?",
    "973862": "If I want to use these data in a publication: 1) is it allowed? and 2) Footnote, acknowledgement and/or reference?",
    "973575": "Hi          ",
    "943233": "@philculliton\nAs submission can be done through direct upload on current test data which is just 30% of all test data.\nWhen will we have rest 70% and will submission be allowed as we can do now.\nThanks for feedback.",
    "925155": "@juliaelliott \nCan we use test data as well for pseudo labeling ?",
    "916807": "Hi",
    "880704": "The rules say:\n\n&gt; To receive a Prize, a potential winner must have used External Data available to use by all participants of the competition for purposes of the competition, for use that includes research or academic purposes, at no cost to the other participants (\"Public External Data\").\n\nDo I understand correctly that submissions that do not use any Public External Data but rather only the competition data, are not eligible to win prizes?\n\n\n",
    "868237": "Hi, what if I use only image processing algorithms, to make predictions, i mean, no models generated at all, just predictions. it is allowed?\n",
    "865300": "Hello,\nTwo questions:\nIn the sentence:\n'Dermatologists could enhance their diagnostic accuracy if detection algorithms take into account “contextual” images within the same patient to determine which images represent a melanoma.'\nDo you mean that current models do not take into account the localization of the skin lezion on the body ?\nIs it also all what you mean by \"Existing AI approaches have not adequately considered this clinical frame of reference\"  ?\n\nThank you!",
    "864843": "A little bit confusing:\n\nReading competition description on the competition Data page we can see:\n\nWhat am I predicting?\nYou are predicting a binary target for each image. Your model should predict the probability (floating point) that the lesion in the image is malignant (the target). In the training data, train.csv, the value 0 denotes benign, and 1 indicates malignant. Predictions should be floating point values between 0.0 and 1.0, with 0.5 as a binary decision threshold.\n\n\nBut the actual metric is AUC, which is invariant to predicted probability distribution. \n",
    "941794": "",
    "864442": ""
  }
}