{
  "id": 167215,
  "title": "[Updated] Distribution of test data 2.367%",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/167215",
  "author_name": "",
  "post_date": "2020-07-15T16:41:05.775555200Z",
  "votes": 24,
  "comment_count": 11,
  "views": 0,
  "content": "<p>With the availability of an <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167210\">additional digit</a> in the Public Leaderboard,\nmy previous post on <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497\">Distribution of test data</a> should be revised as follows:</p>\n\n<p>Number of MMs in Public LB = 78</p>\n\n<p>Total number of MMs in 10982 test set = 260 = 78/0.3 (assuming that exactly 30% data is used)</p>",
  "messages": [
    {
      "id": "930686",
      "postDate": "07/15/2020 16:41:05",
      "content": "<p>With the availability of an <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167210\">additional digit</a> in the Public Leaderboard,\nmy previous post on <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497\">Distribution of test data</a> should be revised as follows:</p>\n\n<p>Number of MMs in Public LB = 78</p>\n\n<p>Total number of MMs in 10982 test set = 260 = 78/0.3 (assuming that exactly 30% data is used)</p>",
      "rawMarkdown": "With the availability of an [additional digit](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167210) in the Public Leaderboard,\nmy previous post on [Distribution of test data](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497) should be revised as follows:\n\nNumber of MMs in Public LB = 78\n\nTotal number of MMs in 10982 test set = 260 = 78/0.3 (assuming that exactly 30% data is used)",
      "votes": null
    },
    {
      "id": "930718",
      "postDate": "07/15/2020 17:06:23",
      "content": "<p>Can you please clarify, why do you state that there are 260 MMs in the test? Basically, what came first for you, that there are 78 in the public test or that there are 260 in the whole test? I guess it is the former, how have you found it, if it is not a secret?</p>",
      "rawMarkdown": "Can you please clarify, why do you state that there are 260 MMs in the test? Basically, what came first for you, that there are 78 in the public test or that there are 260 in the whole test? I guess it is the former, how have you found it, if it is not a secret?",
      "votes": null
    },
    {
      "id": "930723",
      "postDate": "07/15/2020 17:08:33",
      "content": "<p>Updated my post. (Initially I mentioned the total number first because that is more important!)</p>\n\n<p><strong>Secret Recipe</strong> below:\nSet all predictions to 0, except <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">one known Malignant image</a> to 1, and submit!\n(and then a little back-of-the-envelope calculation 😏 to estimate the 2-class distribution)</p>",
      "rawMarkdown": "Updated my post. (Initially I mentioned the total number first because that is more important!)\n\n**Secret Recipe** below:\nSet all predictions to 0, except [one known Malignant image](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943) to 1, and submit!\n(and then a little back-of-the-envelope calculation 😏 to estimate the 2-class distribution)",
      "votes": null
    },
    {
      "id": "931496",
      "postDate": "07/16/2020 08:29:49",
      "content": "<p><a href=\"/sirishks\">@sirishks</a> is there any duplicate from train to test? I read that there isn't so how do you find that one malignant (which also need to be in the public LB)?</p>\n\n<p>Could you describe back of the envelope? Once you get that public malignant example doing what you describe will give you an uplift of 1/M (where M is the number of malignant examples in public test set). So what is the AUC you get after doing this experiment? 0.5128?</p>",
      "rawMarkdown": "sirishks is there any duplicate from train to test? I read that there isn't so how do you find that one malignant (which also need to be in the public LB)?\n\nCould you describe back of the envelope? Once you get that public malignant example doing what you describe will give you an uplift of 1/M (where M is the number of malignant examples in public test set). So what is the AUC you get after doing this experiment? 0.5128?",
      "votes": null
    },
    {
      "id": "931533",
      "postDate": "07/16/2020 08:55:04",
      "content": "<p><a href=\"/optimo\">@optimo</a> For example, test image 'ISIC_9207777.jpg' in Public LB is a <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414\">known MM</a>.</p>",
      "rawMarkdown": "optimo For example, test image 'ISIC_9207777.jpg' in Public LB is a [known MM](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414).",
      "votes": null
    },
    {
      "id": "931741",
      "postDate": "07/16/2020 12:20:21",
      "content": "<p><a href=\"/sirishks\">@sirishks</a> ISIC_9207777.jpg is in private LB you owe me a submission haha!</p>",
      "rawMarkdown": "sirishks ISIC_9207777.jpg is in private LB you owe me a submission haha!",
      "votes": null
    },
    {
      "id": "931963",
      "postDate": "07/16/2020 15:29:01",
      "content": "<p>The duplicates that Sirish is referring to are duplicates between test images and 2019 external data images. There are no duplicates between test images and 2020 train images.</p>",
      "rawMarkdown": "The duplicates that Sirish is referring to are duplicates between test images and 2019 external data images. There are no duplicates between test images and 2020 train images.",
      "votes": null
    },
    {
      "id": "932194",
      "postDate": "07/16/2020 19:50:15",
      "content": "<p>yep <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> but it's enough to do LB probing!</p>\n<p>Just to avoid more work on this for other people : </p>\n<ul>\n<li>ISIC<em>9207777 and  ISIC</em>5224960 are both on private LB</li>\n<li>ISIC_6457527 however is in public LB</li>\n<li>if you submit only 0s but ISIC_6457527 set to 1 you have an AUC=0.5064</li>\n<li>what this implies is that if you have M malignant lesions in public LB, if you set them all one by one to 1 you'll get closer to AUC=1 with the same upflift (0.0064 each time) so you know that M*0.0064=0.5 (because the base score is 0.5, so your uplift is 1-0.5) which gives you M = 78,125. This is how <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> is assuming that M=78!</li>\n</ul>\n<p>I feel like it's better to be transparent on that kind of things as it's not a secret nor a secret recipe, just a few LB probing and basic math. <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> I hope this fully answer your initial question!</p>\n<p>However I'm still not sure how this information is useful, it does not say anything about private LB and it does not help classifying samples correctly. But 78 positive examples is quite easy to overfit indeed…</p>",
      "rawMarkdown": "yep @cdeotte but it's enough to do LB probing!\n\nJust to avoid more work on this for other people : \n-  ISIC_9207777 and  ISIC_5224960 are both on private LB\n- ISIC_6457527 however is in public LB\n- if you submit only 0s but ISIC_6457527 set to 1 you have an AUC=0.5064\n- what this implies is that if you have M malignant lesions in public LB, if you set them all one by one to 1 you'll get closer to AUC=1 with the same upflift (0.0064 each time) so you know that M*0.0064=0.5 (because the base score is 0.5, so your uplift is 1-0.5) which gives you M = 78,125. This is how @sirishks is assuming that M=78!\n\nI feel like it's better to be transparent on that kind of things as it's not a secret nor a secret recipe, just a few LB probing and basic math. @zaharch I hope this fully answer your initial question!\n\nHowever I'm still not sure how this information is useful, it does not say anything about private LB and it does not help classifying samples correctly. But 78 positive examples is quite easy to overfit indeed...",
      "votes": null
    },
    {
      "id": "932733",
      "postDate": "07/17/2020 08:43:31",
      "content": "<p>77 or 78 in public LB right? (considering 0.50649 is rounded to 0.5064). Also, do they mention anywhere that the competition that data is stratified split to 70 and 30 and that we can expect the equal percent of melanoma images on public and private LBs?</p>",
      "rawMarkdown": "77 or 78 in public LB right? (considering 0.50649 is rounded to 0.5064). Also, do they mention anywhere that the competition that data is stratified split to 70 and 30 and that we can expect the equal percent of melanoma images on public and private LBs?",
      "votes": null
    },
    {
      "id": "932742",
      "postDate": "07/17/2020 08:48:22",
      "content": "<p>If you assume 77, but reality is 78, then you will end up with a <strong>False Negative</strong>.\nHowever, if you assume 78, but really there are only 77, the penalty is much smaller.</p>",
      "rawMarkdown": "If you assume 77, but reality is 78, then you will end up with a **False Negative**.\nHowever, if you assume 78, but really there are only 77, the penalty is much smaller.",
      "votes": null
    },
    {
      "id": "932750",
      "postDate": "07/17/2020 08:54:06",
      "content": "<p>Oh thanks!!! About the other question? Can we assume a similar 2.367% on Private test set too?</p>",
      "rawMarkdown": "Oh thanks!!! About the other question? Can we assume a similar 2.367% on Private test set too?",
      "votes": null
    },
    {
      "id": "932757",
      "postDate": "07/17/2020 08:58:56",
      "content": "<p>If the organizers did not use exactly 30%, but say 25% or 35% data for Private LB, then the ratio changes back to my <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497\">original post</a>, where I mentioned 2%-3% range (i.e. 220-330 MM images) in test set.</p>",
      "rawMarkdown": "If the organizers did not use exactly 30%, but say 25% or 35% data for Private LB, then the ratio changes back to my [original post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497), where I mentioned 2%-3% range (i.e. 220-330 MM images) in test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 930718,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "07/15/2020 17:06:23",
      "content": "<p>Can you please clarify, why do you state that there are 260 MMs in the test? Basically, what came first for you, that there are 78 in the public test or that there are 260 in the whole test? I guess it is the former, how have you found it, if it is not a secret?</p>",
      "votes": null,
      "replies": [
        {
          "id": 930723,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/15/2020 17:08:33",
          "content": "<p>Updated my post. (Initially I mentioned the total number first because that is more important!)</p>\n\n<p><strong>Secret Recipe</strong> below:\nSet all predictions to 0, except <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943\">one known Malignant image</a> to 1, and submit!\n(and then a little back-of-the-envelope calculation 😏 to estimate the 2-class distribution)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931496,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/16/2020 08:29:49",
          "content": "<p><a href=\"/sirishks\">@sirishks</a> is there any duplicate from train to test? I read that there isn't so how do you find that one malignant (which also need to be in the public LB)?</p>\n\n<p>Could you describe back of the envelope? Once you get that public malignant example doing what you describe will give you an uplift of 1/M (where M is the number of malignant examples in public test set). So what is the AUC you get after doing this experiment? 0.5128?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931533,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/16/2020 08:55:04",
          "content": "<p><a href=\"/optimo\">@optimo</a> For example, test image 'ISIC_9207777.jpg' in Public LB is a <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414\">known MM</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931741,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/16/2020 12:20:21",
          "content": "<p><a href=\"/sirishks\">@sirishks</a> ISIC_9207777.jpg is in private LB you owe me a submission haha!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931963,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "07/16/2020 15:29:01",
          "content": "<p>The duplicates that Sirish is referring to are duplicates between test images and 2019 external data images. There are no duplicates between test images and 2020 train images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932194,
          "author_name": "optimo",
          "author_url": "",
          "post_date": "07/16/2020 19:50:15",
          "content": "<p>yep <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> but it's enough to do LB probing!</p>\n<p>Just to avoid more work on this for other people : </p>\n<ul>\n<li>ISIC<em>9207777 and  ISIC</em>5224960 are both on private LB</li>\n<li>ISIC_6457527 however is in public LB</li>\n<li>if you submit only 0s but ISIC_6457527 set to 1 you have an AUC=0.5064</li>\n<li>what this implies is that if you have M malignant lesions in public LB, if you set them all one by one to 1 you'll get closer to AUC=1 with the same upflift (0.0064 each time) so you know that M*0.0064=0.5 (because the base score is 0.5, so your uplift is 1-0.5) which gives you M = 78,125. This is how <a href=\"https://www.kaggle.com/sirishks\" target=\"_blank\">@sirishks</a> is assuming that M=78!</li>\n</ul>\n<p>I feel like it's better to be transparent on that kind of things as it's not a secret nor a secret recipe, just a few LB probing and basic math. <a href=\"https://www.kaggle.com/zaharch\" target=\"_blank\">@zaharch</a> I hope this fully answer your initial question!</p>\n<p>However I'm still not sure how this information is useful, it does not say anything about private LB and it does not help classifying samples correctly. But 78 positive examples is quite easy to overfit indeed…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932733,
          "author_name": "josealways123",
          "author_url": "",
          "post_date": "07/17/2020 08:43:31",
          "content": "<p>77 or 78 in public LB right? (considering 0.50649 is rounded to 0.5064). Also, do they mention anywhere that the competition that data is stratified split to 70 and 30 and that we can expect the equal percent of melanoma images on public and private LBs?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932742,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/17/2020 08:48:22",
          "content": "<p>If you assume 77, but reality is 78, then you will end up with a <strong>False Negative</strong>.\nHowever, if you assume 78, but really there are only 77, the penalty is much smaller.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932750,
          "author_name": "josealways123",
          "author_url": "",
          "post_date": "07/17/2020 08:54:06",
          "content": "<p>Oh thanks!!! About the other question? Can we assume a similar 2.367% on Private test set too?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932757,
          "author_name": "sirishks",
          "author_url": "",
          "post_date": "07/17/2020 08:58:56",
          "content": "<p>If the organizers did not use exactly 30%, but say 25% or 35% data for Private LB, then the ratio changes back to my <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497\">original post</a>, where I mentioned 2%-3% range (i.e. 220-330 MM images) in test set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "930686": "With the availability of an [additional digit](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/167210) in the Public Leaderboard,\nmy previous post on [Distribution of test data](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497) should be revised as follows:\n\nNumber of MMs in Public LB = 78\n\nTotal number of MMs in 10982 test set = 260 = 78/0.3 (assuming that exactly 30% data is used)",
    "930718": "Can you please clarify, why do you state that there are 260 MMs in the test? Basically, what came first for you, that there are 78 in the public test or that there are 260 in the whole test? I guess it is the former, how have you found it, if it is not a secret?",
    "930723": "Updated my post. (Initially I mentioned the total number first because that is more important!)\n\n**Secret Recipe** below:\nSet all predictions to 0, except [one known Malignant image](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161943) to 1, and submit!\n(and then a little back-of-the-envelope calculation 😏 to estimate the 2-class distribution)",
    "931496": "sirishks is there any duplicate from train to test? I read that there isn't so how do you find that one malignant (which also need to be in the public LB)?\n\nCould you describe back of the envelope? Once you get that public malignant example doing what you describe will give you an uplift of 1/M (where M is the number of malignant examples in public test set). So what is the AUC you get after doing this experiment? 0.5128?",
    "931533": "optimo For example, test image 'ISIC_9207777.jpg' in Public LB is a [known MM](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/158414).",
    "931741": "sirishks ISIC_9207777.jpg is in private LB you owe me a submission haha!",
    "931963": "The duplicates that Sirish is referring to are duplicates between test images and 2019 external data images. There are no duplicates between test images and 2020 train images.",
    "932194": "yep @cdeotte but it's enough to do LB probing!\n\nJust to avoid more work on this for other people : \n-  ISIC_9207777 and  ISIC_5224960 are both on private LB\n- ISIC_6457527 however is in public LB\n- if you submit only 0s but ISIC_6457527 set to 1 you have an AUC=0.5064\n- what this implies is that if you have M malignant lesions in public LB, if you set them all one by one to 1 you'll get closer to AUC=1 with the same upflift (0.0064 each time) so you know that M*0.0064=0.5 (because the base score is 0.5, so your uplift is 1-0.5) which gives you M = 78,125. This is how @sirishks is assuming that M=78!\n\nI feel like it's better to be transparent on that kind of things as it's not a secret nor a secret recipe, just a few LB probing and basic math. @zaharch I hope this fully answer your initial question!\n\nHowever I'm still not sure how this information is useful, it does not say anything about private LB and it does not help classifying samples correctly. But 78 positive examples is quite easy to overfit indeed...",
    "932733": "77 or 78 in public LB right? (considering 0.50649 is rounded to 0.5064). Also, do they mention anywhere that the competition that data is stratified split to 70 and 30 and that we can expect the equal percent of melanoma images on public and private LBs?",
    "932742": "If you assume 77, but reality is 78, then you will end up with a **False Negative**.\nHowever, if you assume 78, but really there are only 77, the penalty is much smaller.",
    "932750": "Oh thanks!!! About the other question? Can we assume a similar 2.367% on Private test set too?",
    "932757": "If the organizers did not use exactly 30%, but say 25% or 35% data for Private LB, then the ratio changes back to my [original post](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/161497), where I mentioned 2%-3% range (i.e. 220-330 MM images) in test set."
  },
  "source": "meta"
}