{
  "id": 161497,
  "title": "Distribution of test data 2%-3% (0.970)",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/161497",
  "author_name": "Sirish Somanchi",
  "post_date": "2020-06-25T04:40:58.508000",
  "votes": 71,
  "comment_count": 23,
  "views": 0,
  "content": "<p>My observations:\n1. Based on recent analysis, I was able to figure out that the number of Malignant Melanomas in test data (10982 images) is in the 2%-3% range i.e. only top 220-330 images are important and rest are benign lesions.\n2. My AUC metric understanding is that we can get a high score in spite of a <strong>lot of false positives</strong> (i.e. benign images incorrectly classified as malignant) as long as we can get all malignant images into the top 1000 (ranking out of 10982).\n3. So, I am primarily focusing on top 1000 (based on my prediction scores) and actually started ignoring bottom 3982 rows because they are <strong>very likely</strong> benign anyway. This halves the rest of the test set into the middle 6000 images (range 1000 to 7000) where some <strong>Amelanotic Malignant Melanomas</strong> (the very difficult to diagnose kind) are probably hidden.</p>\n\n<p>This helped me <strong>temporarily</strong> reach 0.970 (Public LB #1 position) today.\nPS: I am <strong>really</strong> overfitting to the Public LB, and my CV scores (Private LB) are actually pretty bad 😪 </p>",
  "messages": [
    {
      "id": 900811,
      "postDate": "2020-06-25T04:40:58.507Z",
      "content": "<p>My observations:\n1. Based on recent analysis, I was able to figure out that the number of Malignant Melanomas in test data (10982 images) is in the 2%-3% range i.e. only top 220-330 images are important and rest are benign lesions.\n2. My AUC metric understanding is that we can get a high score in spite of a <strong>lot of false positives</strong> (i.e. benign images incorrectly classified as malignant) as long as we can get all malignant images into the top 1000 (ranking out of 10982).\n3. So, I am primarily focusing on top 1000 (based on my prediction scores) and actually started ignoring bottom 3982 rows because they are <strong>very likely</strong> benign anyway. This halves the rest of the test set into the middle 6000 images (range 1000 to 7000) where some <strong>Amelanotic Malignant Melanomas</strong> (the very difficult to diagnose kind) are probably hidden.</p>\n\n<p>This helped me <strong>temporarily</strong> reach 0.970 (Public LB #1 position) today.\nPS: I am <strong>really</strong> overfitting to the Public LB, and my CV scores (Private LB) are actually pretty bad 😪 </p>",
      "rawMarkdown": "My observations:\n1. Based on recent analysis, I was able to figure out that the number of Malignant Melanomas in test data (10982 images) is in the 2%-3% range i.e. only top 220-330 images are important and rest are benign lesions.\n2. My AUC metric understanding is that we can get a high score in spite of a **lot of false positives** (i.e. benign images incorrectly classified as malignant) as long as we can get all malignant images into the top 1000 (ranking out of 10982).\n3. So, I am primarily focusing on top 1000 (based on my prediction scores) and actually started ignoring bottom 3982 rows because they are **very likely** benign anyway. This halves the rest of the test set into the middle 6000 images (range 1000 to 7000) where some **Amelanotic Malignant Melanomas** (the very difficult to diagnose kind) are probably hidden.\n\nThis helped me **temporarily** reach 0.970 (Public LB #1 position) today.\nPS: I am **really** overfitting to the Public LB, and my CV scores (Private LB) are actually pretty bad 😪 ",
      "votes": 70
    },
    {
      "id": 901831,
      "postDate": "2020-06-25T18:16:18.957Z",
      "content": "<p>Always nice to see a method that cannot be used in real life to improve a model.</p>\n\n<p>Even better when it promises to make the leader board useless.</p>",
      "rawMarkdown": "Always nice to see a method that cannot be used in real life to improve a model.\n\nEven better when it promises to make the leader board useless.",
      "votes": 16
    },
    {
      "id": 900995,
      "postDate": "2020-06-25T07:39:16.697Z",
      "content": "<p>I saw your score... stared at the screen for ages... threw the mouse... well no... but it was a shock. 😲 </p>\n\n<p>Congrats and thanks for your comments. 😄 </p>\n\n<p>I will now continue on with my ideas (if I can work out how to actually implement them, and that takes time.)</p>",
      "rawMarkdown": "I saw your score... stared at the screen for ages... threw the mouse... well no... but it was a shock. 😲 \n\nCongrats and thanks for your comments. 😄 \n\nI will now continue on with my ideas (if I can work out how to actually implement them, and that takes time.)\n\n",
      "votes": 5,
      "replies": [
        {
          "id": 901086,
          "postDate": "2020-06-25T08:49:31.360Z",
          "content": "<p><a href=\"/mutantspore\">@mutantspore</a> We are not even at half-time and you have nearly 8 more weeks starting today to try all your ideas. Best wishes 👍 </p>",
          "rawMarkdown": "@mutantspore We are not even at half-time and you have nearly 8 more weeks starting today to try all your ideas. Best wishes 👍 ",
          "votes": 2
        }
      ]
    },
    {
      "id": 900851,
      "postDate": "2020-06-25T05:30:45.783Z",
      "content": "<p>Thanks for sharing. I was wondering how you jumped so high  so fast. Is it LB overfitting or we are missing something big. You saved me some time in a wild goose chase </p>\n\n<p>As the number of positive cases is so small in LB we are all overfitting it </p>",
      "rawMarkdown": "Thanks for sharing. I was wondering how you jumped so high  so fast. Is it LB overfitting or we are missing something big. You saved me some time in a wild goose chase \n\nAs the number of positive cases is so small in LB we are all overfitting it ",
      "votes": 4,
      "replies": [
        {
          "id": 900907,
          "postDate": "2020-06-25T06:21:11.350Z",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> 😃 I did not find/exploit any <strong>LEAK</strong>. I can understand your concern, especially because recently top teams lost their ranking (and prize money), for example during Liverpool competition <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/153824\">link1</a> <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/154355\">link2</a></p>",
          "rawMarkdown": "@yuval6967 😃 I did not find/exploit any **LEAK**. I can understand your concern, especially because recently top teams lost their ranking (and prize money), for example during Liverpool competition [link1](https://www.kaggle.com/c/liverpool-ion-switching/discussion/153824) [link2](https://www.kaggle.com/c/liverpool-ion-switching/discussion/154355)",
          "votes": 5
        }
      ]
    },
    {
      "id": 901160,
      "postDate": "2020-06-25T09:27:10.637Z",
      "content": "<p>Hello <a href=\"/sirishks\">@sirishks</a> congratulations for this amazing score!</p>\n\n<p>I won't ask you for the secret recipe but I'm curious to understand how you evaluate number of positive targets in the test set ? I've read this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154624#866439\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154624#866439</a> but could not understand the demonstration.</p>\n\n<p>Here is what I would do : \n- take a ranked prediction with unique predictions values (score is auc0)\n- change only one prediction- the smallest one-and set it to the higher score possible\n- if this prediction is in the public set you should see a difference in your AUC score (auc1) (your score should get worst if this example is benign - ie 0)\n- BUT if p is the proportion of positive samples in public set your delta=auc0-auc1=p/((1-p)*N) where N is the total number of samples in public test set. Let's say N=11K*30%=3300 and p=2% then your delta=6.18e-6 which is impossible to see in the leaderboard.\n- so in order to see a change in the leaderboard you'll need to do this for the 170 lowest scores at the same time (hoping that all of them are actually 0 and all in the public test set)  you'll get a delta of 0.001051329, if p=3% you'll see a delta of 0.001030 if you change the 110 lowest scores. So you can change different number of scores and see when something changes.\n- BUT it's going to take a long time before finding 170 exact 0s that are actually in the public test set, how would you do that?</p>\n\n<p>Since you have 93 submissions I guess that you might have done something similar but could you share it with us? You still have no guaranty that this proportion will hold within the private test set.</p>\n\n<p>I'm not sure to understand what's the point of knowing this anyway. ^^</p>",
      "rawMarkdown": "Hello @sirishks congratulations for this amazing score!\n\nI won't ask you for the secret recipe but I'm curious to understand how you evaluate number of positive targets in the test set ? I've read this https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154624#866439 but could not understand the demonstration.\n\nHere is what I would do : \n- take a ranked prediction with unique predictions values (score is auc0)\n- change only one prediction- the smallest one-and set it to the higher score possible\n- if this prediction is in the public set you should see a difference in your AUC score (auc1) (your score should get worst if this example is benign - ie 0)\n- BUT if p is the proportion of positive samples in public set your delta=auc0-auc1=p/((1-p)*N) where N is the total number of samples in public test set. Let's say N=11K*30%=3300 and p=2% then your delta=6.18e-6 which is impossible to see in the leaderboard.\n- so in order to see a change in the leaderboard you'll need to do this for the 170 lowest scores at the same time (hoping that all of them are actually 0 and all in the public test set)  you'll get a delta of 0.001051329, if p=3% you'll see a delta of 0.001030 if you change the 110 lowest scores. So you can change different number of scores and see when something changes.\n- BUT it's going to take a long time before finding 170 exact 0s that are actually in the public test set, how would you do that?\n\nSince you have 93 submissions I guess that you might have done something similar but could you share it with us? You still have no guaranty that this proportion will hold within the private test set.\n\nI'm not sure to understand what's the point of knowing this anyway. ^^",
      "votes": 1,
      "replies": [
        {
          "id": 901185,
          "postDate": "2020-06-25T09:50:43.030Z",
          "content": "<p><a href=\"/optimo\">@optimo</a> As I have already mentioned above, my CV score (Private LB) is actually not great. So, I don't recommend re-creating my overfitting experiments.</p>\n\n<p><strong>Secret Recipe</strong> below:\nIf you really insist, you could take submission.csv of one of the public kernels with 0.946 score, sort DESC by target, <strong>freeze</strong> bottom 7982 rows, and experiment only with the top 3000 rows ... In my experience, the top 3k rows contain 95% of MMs.\nBeware, this is a bad approach which will <strong>NOT</strong> help either CV or Private LB later.</p>\n\n<p>The <strong>key take-away</strong> in my original post above is the count of MMs in test data i.e. 220-330 range. This is the information you should really focus on.</p>",
          "rawMarkdown": "@optimo As I have already mentioned above, my CV score (Private LB) is actually not great. So, I don't recommend re-creating my overfitting experiments.\n\n**Secret Recipe** below:\nIf you really insist, you could take submission.csv of one of the public kernels with 0.946 score, sort DESC by target, **freeze** bottom 7982 rows, and experiment only with the top 3000 rows ... In my experience, the top 3k rows contain 95% of MMs.\nBeware, this is a bad approach which will **NOT** help either CV or Private LB later.\n\nThe **key take-away** in my original post above is the count of MMs in test data i.e. 220-330 range. This is the information you should really focus on.",
          "votes": 6
        },
        {
          "id": 901199,
          "postDate": "2020-06-25T10:01:29.980Z",
          "content": "<p><a href=\"/optimo\">@optimo</a> FYI - Please see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">Rules</a>, where <strong>EXTERNAL DATA</strong> section clearly mentions <strong>No hand-labeling of the test set.</strong></p>",
          "rawMarkdown": "@optimo FYI - Please see [Rules](https://www.kaggle.com/c/siim-isic-melanoma-classification/rules), where **EXTERNAL DATA** section clearly mentions **No hand-labeling of the test set.**",
          "votes": 1
        },
        {
          "id": 901329,
          "postDate": "2020-06-25T11:47:59.120Z",
          "content": "<p>Thanks for your answer, I still don't understand how you came up with \"MMs in test data i.e. 220-330 range\"</p>\n\n<p>I feel a bit sad that secret recipe consists in trying random order of the top 3K predictions from public submission, since fairly soon the public leaderboard will become completely irrelevant.</p>\n\n<p>My take away from this : I need to focus really hard on my CV scheme!</p>\n\n<p>Thanks for sharing how you got to 0.97!</p>",
          "rawMarkdown": "Thanks for your answer, I still don't understand how you came up with \"MMs in test data i.e. 220-330 range\"\n\nI feel a bit sad that secret recipe consists in trying random order of the top 3K predictions from public submission, since fairly soon the public leaderboard will become completely irrelevant.\n\nMy take away from this : I need to focus really hard on my CV scheme!\n\nThanks for sharing how you got to 0.97!",
          "votes": -1
        }
      ]
    },
    {
      "id": 900947,
      "postDate": "2020-06-25T06:51:04.900Z",
      "content": "<p>Nice observation and thanks for sharing , congrats for first place though.</p>",
      "rawMarkdown": "Nice observation and thanks for sharing , congrats for first place though.",
      "votes": 2
    },
    {
      "id": 2675495,
      "postDate": "2024-03-01T00:15:50.190Z",
      "content": "<p>Hi,</p>\n<p>I'm looking for guidance on how to obtain the test data labels for evaluating my model. Despite the competition ending four years ago, and the metadata not containing a column for test data labels, I'm eager to find a solution to effectively evaluate my model's performance. Your help in this regard would be greatly appreciated. </p>\n<p>Thank you.</p>",
      "rawMarkdown": "Hi,\n\nI'm looking for guidance on how to obtain the test data labels for evaluating my model. Despite the competition ending four years ago, and the metadata not containing a column for test data labels, I'm eager to find a solution to effectively evaluate my model's performance. Your help in this regard would be greatly appreciated. \n\nThank you."
    },
    {
      "id": 939231,
      "postDate": "2020-07-22T05:37:52.513Z",
      "content": "<p>Thanks so much for sharing <a href=\"/sirishks\">@sirishks</a> \nthis is helpful to understand the data </p>",
      "rawMarkdown": "Thanks so much for sharing @sirishks \nthis is helpful to understand the data "
    },
    {
      "id": 912611,
      "postDate": "2020-07-02T15:55:08.540Z",
      "content": "<p>amazing,you got 0.98👍 </p>",
      "rawMarkdown": "amazing,you got 0.98👍 ",
      "replies": [
        {
          "id": 912648,
          "postDate": "2020-07-02T16:16:10.197Z",
          "content": "<p>Thx <a href=\"/bixiaopeng\">@bixiaopeng</a> </p>\n\n<p>However, since I am overfitting to public LB, but not private LB, the score is <strong>not an actual indication</strong> of the performance of my model.</p>\n\n<p>More details: The 30% data used in Public LB will <strong>not be used later</strong> in Private LB (70% <strong>other</strong> data) and so any gains that I have made in identifying Ms here will not be useful because those images will not be present in the Private LB, and I have no idea how my model will perform on those unknown images.</p>\n\n<p>I <strong>know</strong> that I will face a <strong>shake-down</strong> when Private LB is revealed in the end. This is because of the nature of class imbalance where very few images will decide the winner.\n10/220=0.045 and 10/330=0.03 i.e. missing just 10 images (False Negatives) will decrease the score by 3%-4.5%</p>",
          "rawMarkdown": "Thx @bixiaopeng \n\nHowever, since I am overfitting to public LB, but not private LB, the score is **not an actual indication** of the performance of my model.\n\nMore details: The 30% data used in Public LB will **not be used later** in Private LB (70% **other** data) and so any gains that I have made in identifying Ms here will not be useful because those images will not be present in the Private LB, and I have no idea how my model will perform on those unknown images.\n\nI **know** that I will face a **shake-down** when Private LB is revealed in the end. This is because of the nature of class imbalance where very few images will decide the winner.\n10/220=0.045 and 10/330=0.03 i.e. missing just 10 images (False Negatives) will decrease the score by 3%-4.5%",
          "votes": 4
        },
        {
          "id": 912670,
          "postDate": "2020-07-02T16:30:30.887Z",
          "content": "<p>😳 😲 😯 </p>\n\n<p>Starting to sound like you're trying to lull us into not worrying about your score. 😉 😃 </p>\n\n<p>Unfortunately, I cannot unsee it now and that peak too must be climbed... because it's there...</p>\n\n<p>Just kidding... you seem to have a firm grasp of the issues involved and I'm sure you have a plan B for the Private leaderboard. </p>\n\n<p>Good Luck 😁 </p>",
          "rawMarkdown": "😳 😲 😯 \n\nStarting to sound like you're trying to lull us into not worrying about your score. 😉 😃 \n\nUnfortunately, I cannot unsee it now and that peak too must be climbed... because it's there...\n\nJust kidding... you seem to have a firm grasp of the issues involved and I'm sure you have a plan B for the Private leaderboard. \n\nGood Luck 😁 ",
          "votes": 4
        }
      ]
    },
    {
      "id": 901949,
      "postDate": "2020-06-25T20:05:52.283Z",
      "content": "<p>Congrats for first place !! I hope you will solve the problem of overfitting .</p>",
      "rawMarkdown": "Congrats for first place !! I hope you will solve the problem of overfitting .",
      "replies": [
        {
          "id": 902243,
          "postDate": "2020-06-26T02:55:00.033Z",
          "content": "<p>Thank you <a href=\"/syphax93\">@syphax93</a> </p>\n\n<p>What I learned by participating in Kaggle competitions:\nStep1: Solve Overfitting, LB probing, Leak hunting\nStep2: Scrape Discussions and Notebooks for great ideas\nStep3: Solve the <strong>actual problem</strong> of <strong>reducing False Negatives</strong> i.e. distinguishing Amelanotic MMs from Benign\n            and pray that the final solution is good enough!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F656212%2Fad87938764cdcf91c7c5b67450f170b8%2FKaggle_PublicLB_not_PrivateLB.jpg?generation=1593139886609279&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Thank you @syphax93 \n\nWhat I learned by participating in Kaggle competitions:\nStep1: Solve Overfitting, LB probing, Leak hunting\nStep2: Scrape Discussions and Notebooks for great ideas\nStep3: Solve the **actual problem** of **reducing False Negatives** i.e. distinguishing Amelanotic MMs from Benign\n            and pray that the final solution is good enough!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F656212%2Fad87938764cdcf91c7c5b67450f170b8%2FKaggle_PublicLB_not_PrivateLB.jpg?generation=1593139886609279&amp;alt=media)\n ",
          "votes": 16
        }
      ]
    },
    {
      "id": 901103,
      "postDate": "2020-06-25T08:58:49.683Z",
      "content": "<p>Seems like a kind of Balance Class Sampler ,right?\nDid you use metadata ?</p>",
      "rawMarkdown": "Seems like a kind of Balance Class Sampler ,right?\nDid you use metadata ?"
    },
    {
      "id": 972190,
      "postDate": "2020-08-16T10:09:24.540Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 972186,
      "postDate": "2020-08-16T10:07:30.667Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 900918,
      "postDate": "2020-06-25T06:29:00.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 972254,
      "postDate": "2020-08-16T11:45:51.277Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 960296,
      "postDate": "2020-08-06T09:38:40.723Z",
      "content": "<p>Thanks <a href=\"/sirishks\">@sirishks</a> for sharing this.</p>",
      "rawMarkdown": "Thanks @sirishks for sharing this."
    }
  ],
  "comments": [
    {
      "id": 901831,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2020-06-25T18:16:18.957000",
      "content": "<p>Always nice to see a method that cannot be used in real life to improve a model.</p>\n\n<p>Even better when it promises to make the leader board useless.</p>",
      "votes": 16,
      "replies": []
    },
    {
      "id": 900995,
      "author_name": "Bruce Young",
      "author_url": "",
      "post_date": "2020-06-25T07:39:16.697000",
      "content": "<p>I saw your score... stared at the screen for ages... threw the mouse... well no... but it was a shock. 😲 </p>\n\n<p>Congrats and thanks for your comments. 😄 </p>\n\n<p>I will now continue on with my ideas (if I can work out how to actually implement them, and that takes time.)</p>",
      "votes": 5,
      "replies": [
        {
          "id": 901086,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-06-25T08:49:31.360000",
          "content": "<p><a href=\"/mutantspore\">@mutantspore</a> We are not even at half-time and you have nearly 8 more weeks starting today to try all your ideas. Best wishes 👍 </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 900851,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2020-06-25T05:30:45.783000",
      "content": "<p>Thanks for sharing. I was wondering how you jumped so high  so fast. Is it LB overfitting or we are missing something big. You saved me some time in a wild goose chase </p>\n\n<p>As the number of positive cases is so small in LB we are all overfitting it </p>",
      "votes": 4,
      "replies": [
        {
          "id": 900907,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-06-25T06:21:11.350000",
          "content": "<p><a href=\"/yuval6967\">@yuval6967</a> 😃 I did not find/exploit any <strong>LEAK</strong>. I can understand your concern, especially because recently top teams lost their ranking (and prize money), for example during Liverpool competition <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/153824\">link1</a> <a href=\"https://www.kaggle.com/c/liverpool-ion-switching/discussion/154355\">link2</a></p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 901160,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2020-06-25T09:27:10.637000",
      "content": "<p>Hello <a href=\"/sirishks\">@sirishks</a> congratulations for this amazing score!</p>\n\n<p>I won't ask you for the secret recipe but I'm curious to understand how you evaluate number of positive targets in the test set ? I've read this <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154624#866439\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154624#866439</a> but could not understand the demonstration.</p>\n\n<p>Here is what I would do : \n- take a ranked prediction with unique predictions values (score is auc0)\n- change only one prediction- the smallest one-and set it to the higher score possible\n- if this prediction is in the public set you should see a difference in your AUC score (auc1) (your score should get worst if this example is benign - ie 0)\n- BUT if p is the proportion of positive samples in public set your delta=auc0-auc1=p/((1-p)*N) where N is the total number of samples in public test set. Let's say N=11K*30%=3300 and p=2% then your delta=6.18e-6 which is impossible to see in the leaderboard.\n- so in order to see a change in the leaderboard you'll need to do this for the 170 lowest scores at the same time (hoping that all of them are actually 0 and all in the public test set)  you'll get a delta of 0.001051329, if p=3% you'll see a delta of 0.001030 if you change the 110 lowest scores. So you can change different number of scores and see when something changes.\n- BUT it's going to take a long time before finding 170 exact 0s that are actually in the public test set, how would you do that?</p>\n\n<p>Since you have 93 submissions I guess that you might have done something similar but could you share it with us? You still have no guaranty that this proportion will hold within the private test set.</p>\n\n<p>I'm not sure to understand what's the point of knowing this anyway. ^^</p>",
      "votes": 1,
      "replies": [
        {
          "id": 901185,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-06-25T09:50:43.030000",
          "content": "<p><a href=\"/optimo\">@optimo</a> As I have already mentioned above, my CV score (Private LB) is actually not great. So, I don't recommend re-creating my overfitting experiments.</p>\n\n<p><strong>Secret Recipe</strong> below:\nIf you really insist, you could take submission.csv of one of the public kernels with 0.946 score, sort DESC by target, <strong>freeze</strong> bottom 7982 rows, and experiment only with the top 3000 rows ... In my experience, the top 3k rows contain 95% of MMs.\nBeware, this is a bad approach which will <strong>NOT</strong> help either CV or Private LB later.</p>\n\n<p>The <strong>key take-away</strong> in my original post above is the count of MMs in test data i.e. 220-330 range. This is the information you should really focus on.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 901199,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-06-25T10:01:29.980000",
          "content": "<p><a href=\"/optimo\">@optimo</a> FYI - Please see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/rules\">Rules</a>, where <strong>EXTERNAL DATA</strong> section clearly mentions <strong>No hand-labeling of the test set.</strong></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 901329,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2020-06-25T11:47:59.120000",
          "content": "<p>Thanks for your answer, I still don't understand how you came up with \"MMs in test data i.e. 220-330 range\"</p>\n\n<p>I feel a bit sad that secret recipe consists in trying random order of the top 3K predictions from public submission, since fairly soon the public leaderboard will become completely irrelevant.</p>\n\n<p>My take away from this : I need to focus really hard on my CV scheme!</p>\n\n<p>Thanks for sharing how you got to 0.97!</p>",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 900947,
      "author_name": "yash chaudhary",
      "author_url": "",
      "post_date": "2020-06-25T06:51:04.900000",
      "content": "<p>Nice observation and thanks for sharing , congrats for first place though.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2675495,
      "author_name": "hrithik sarda",
      "author_url": "",
      "post_date": "2024-03-01T00:15:50.190000",
      "content": "<p>Hi,</p>\n<p>I'm looking for guidance on how to obtain the test data labels for evaluating my model. Despite the competition ending four years ago, and the metadata not containing a column for test data labels, I'm eager to find a solution to effectively evaluate my model's performance. Your help in this regard would be greatly appreciated. </p>\n<p>Thank you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 939231,
      "author_name": "Kamal Das",
      "author_url": "",
      "post_date": "2020-07-22T05:37:52.513000",
      "content": "<p>Thanks so much for sharing <a href=\"/sirishks\">@sirishks</a> \nthis is helpful to understand the data </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 912611,
      "author_name": "xiaopeng",
      "author_url": "",
      "post_date": "2020-07-02T15:55:08.540000",
      "content": "<p>amazing,you got 0.98👍 </p>",
      "votes": 0,
      "replies": [
        {
          "id": 912648,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-07-02T16:16:10.197000",
          "content": "<p>Thx <a href=\"/bixiaopeng\">@bixiaopeng</a> </p>\n\n<p>However, since I am overfitting to public LB, but not private LB, the score is <strong>not an actual indication</strong> of the performance of my model.</p>\n\n<p>More details: The 30% data used in Public LB will <strong>not be used later</strong> in Private LB (70% <strong>other</strong> data) and so any gains that I have made in identifying Ms here will not be useful because those images will not be present in the Private LB, and I have no idea how my model will perform on those unknown images.</p>\n\n<p>I <strong>know</strong> that I will face a <strong>shake-down</strong> when Private LB is revealed in the end. This is because of the nature of class imbalance where very few images will decide the winner.\n10/220=0.045 and 10/330=0.03 i.e. missing just 10 images (False Negatives) will decrease the score by 3%-4.5%</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 912670,
          "author_name": "Bruce Young",
          "author_url": "",
          "post_date": "2020-07-02T16:30:30.887000",
          "content": "<p>😳 😲 😯 </p>\n\n<p>Starting to sound like you're trying to lull us into not worrying about your score. 😉 😃 </p>\n\n<p>Unfortunately, I cannot unsee it now and that peak too must be climbed... because it's there...</p>\n\n<p>Just kidding... you seem to have a firm grasp of the issues involved and I'm sure you have a plan B for the Private leaderboard. </p>\n\n<p>Good Luck 😁 </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 901949,
      "author_name": "Syphax",
      "author_url": "",
      "post_date": "2020-06-25T20:05:52.283000",
      "content": "<p>Congrats for first place !! I hope you will solve the problem of overfitting .</p>",
      "votes": 0,
      "replies": [
        {
          "id": 902243,
          "author_name": "Sirish Somanchi",
          "author_url": "",
          "post_date": "2020-06-26T02:55:00.033000",
          "content": "<p>Thank you <a href=\"/syphax93\">@syphax93</a> </p>\n\n<p>What I learned by participating in Kaggle competitions:\nStep1: Solve Overfitting, LB probing, Leak hunting\nStep2: Scrape Discussions and Notebooks for great ideas\nStep3: Solve the <strong>actual problem</strong> of <strong>reducing False Negatives</strong> i.e. distinguishing Amelanotic MMs from Benign\n            and pray that the final solution is good enough!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F656212%2Fad87938764cdcf91c7c5b67450f170b8%2FKaggle_PublicLB_not_PrivateLB.jpg?generation=1593139886609279&amp;alt=media\" alt=\"\"></p>",
          "votes": 16,
          "replies": []
        }
      ]
    },
    {
      "id": 901103,
      "author_name": "fate",
      "author_url": "",
      "post_date": "2020-06-25T08:58:49.683000",
      "content": "<p>Seems like a kind of Balance Class Sampler ,right?\nDid you use metadata ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972190,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-16T10:09:24.540000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972186,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-16T10:07:30.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 900918,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-25T06:29:00.580000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 972254,
      "author_name": "Alan Tang",
      "author_url": "",
      "post_date": "2020-08-16T11:45:51.277000",
      "content": "<p>Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 960296,
      "author_name": "Anders Ericsson Gnosco",
      "author_url": "",
      "post_date": "2020-08-06T09:38:40.723000",
      "content": "<p>Thanks <a href=\"/sirishks\">@sirishks</a> for sharing this.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "900811": "My observations:\n1. Based on recent analysis, I was able to figure out that the number of Malignant Melanomas in test data (10982 images) is in the 2%-3% range i.e. only top 220-330 images are important and rest are benign lesions.\n2. My AUC metric understanding is that we can get a high score in spite of a **lot of false positives** (i.e. benign images incorrectly classified as malignant) as long as we can get all malignant images into the top 1000 (ranking out of 10982).\n3. So, I am primarily focusing on top 1000 (based on my prediction scores) and actually started ignoring bottom 3982 rows because they are **very likely** benign anyway. This halves the rest of the test set into the middle 6000 images (range 1000 to 7000) where some **Amelanotic Malignant Melanomas** (the very difficult to diagnose kind) are probably hidden.\n\nThis helped me **temporarily** reach 0.970 (Public LB #1 position) today.\nPS: I am **really** overfitting to the Public LB, and my CV scores (Private LB) are actually pretty bad 😪 ",
    "901831": "Always nice to see a method that cannot be used in real life to improve a model.\n\nEven better when it promises to make the leader board useless.",
    "900995": "I saw your score... stared at the screen for ages... threw the mouse... well no... but it was a shock. 😲 \n\nCongrats and thanks for your comments. 😄 \n\nI will now continue on with my ideas (if I can work out how to actually implement them, and that takes time.)\n\n",
    "900851": "Thanks for sharing. I was wondering how you jumped so high  so fast. Is it LB overfitting or we are missing something big. You saved me some time in a wild goose chase \n\nAs the number of positive cases is so small in LB we are all overfitting it ",
    "901160": "Hello @sirishks congratulations for this amazing score!\n\nI won't ask you for the secret recipe but I'm curious to understand how you evaluate number of positive targets in the test set ? I've read this https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/154624#866439 but could not understand the demonstration.\n\nHere is what I would do : \n- take a ranked prediction with unique predictions values (score is auc0)\n- change only one prediction- the smallest one-and set it to the higher score possible\n- if this prediction is in the public set you should see a difference in your AUC score (auc1) (your score should get worst if this example is benign - ie 0)\n- BUT if p is the proportion of positive samples in public set your delta=auc0-auc1=p/((1-p)*N) where N is the total number of samples in public test set. Let's say N=11K*30%=3300 and p=2% then your delta=6.18e-6 which is impossible to see in the leaderboard.\n- so in order to see a change in the leaderboard you'll need to do this for the 170 lowest scores at the same time (hoping that all of them are actually 0 and all in the public test set)  you'll get a delta of 0.001051329, if p=3% you'll see a delta of 0.001030 if you change the 110 lowest scores. So you can change different number of scores and see when something changes.\n- BUT it's going to take a long time before finding 170 exact 0s that are actually in the public test set, how would you do that?\n\nSince you have 93 submissions I guess that you might have done something similar but could you share it with us? You still have no guaranty that this proportion will hold within the private test set.\n\nI'm not sure to understand what's the point of knowing this anyway. ^^",
    "900947": "Nice observation and thanks for sharing , congrats for first place though.",
    "2675495": "Hi,\n\nI'm looking for guidance on how to obtain the test data labels for evaluating my model. Despite the competition ending four years ago, and the metadata not containing a column for test data labels, I'm eager to find a solution to effectively evaluate my model's performance. Your help in this regard would be greatly appreciated. \n\nThank you.",
    "939231": "Thanks so much for sharing @sirishks \nthis is helpful to understand the data ",
    "912611": "amazing,you got 0.98👍 ",
    "901949": "Congrats for first place !! I hope you will solve the problem of overfitting .",
    "901103": "Seems like a kind of Balance Class Sampler ,right?\nDid you use metadata ?",
    "972190": "",
    "972186": "",
    "900918": "",
    "972254": "Thanks for sharing.",
    "960296": "Thanks @sirishks for sharing this."
  }
}