{
  "id": 172950,
  "title": "96.19LB score around 100 people achieved .... Really !!! What are you talking? ",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/172950",
  "author_name": "",
  "post_date": "2020-08-07T05:31:52.173754100Z",
  "votes": 5,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I found that a lot of people got 96.19 with a few submissions. It took months to reach that stage but few people able to achieve it within multiple submissions... </p>\n\n<p>Check Link below </p>",
  "messages": [
    {
      "id": "961346",
      "postDate": "08/07/2020 05:31:52",
      "content": "<p>I found that a lot of people got 96.19 with a few submissions. It took months to reach that stage but few people able to achieve it within multiple submissions... </p>\n\n<p>Check Link below </p>",
      "rawMarkdown": "I found that a lot of people got 96.19 with a few submissions. It took months to reach that stage but few people able to achieve it within multiple submissions... \n\nCheck Link below",
      "votes": null
    },
    {
      "id": "961392",
      "postDate": "08/07/2020 06:29:58",
      "content": "<p>Maybe someone has shared their solution which got 0.9619, also, of course many of them would be having there own submissions, otherwise its just a coincidence, it may happen at the top since the leaderboard gets denser upwards! :)</p>",
      "rawMarkdown": "Maybe someone has shared their solution which got 0.9619, also, of course many of them would be having there own submissions, otherwise its just a coincidence, it may happen at the top since the leaderboard gets denser upwards! :)",
      "votes": null
    },
    {
      "id": "961395",
      "postDate": "08/07/2020 06:38:41",
      "content": "<p>don't worry too much. </p>",
      "rawMarkdown": "don't worry too much.",
      "votes": null
    },
    {
      "id": "961413",
      "postDate": "08/07/2020 07:12:03",
      "content": "<p>There is nothing to worry about the public LB. The submission that got 0.9619 score has almost 0 malignant predictions. Not just that, other public high scoring NBs have around 40 malignant prediction. That is too low from what I understand so far. So, do not worry about the public score. In private LB, we will have shake up for sure. </p>",
      "rawMarkdown": "There is nothing to worry about the public LB. The submission that got 0.9619 score has almost 0 malignant predictions. Not just that, other public high scoring NBs have around 40 malignant prediction. That is too low from what I understand so far. So, do not worry about the public score. In private LB, we will have shake up for sure.",
      "votes": null
    },
    {
      "id": "961420",
      "postDate": "08/07/2020 07:18:12",
      "content": "<p>How do you know whether a prediction is malignant? The competition metric is AUC. So you just have to make sure that your malignant predictions get higher predictions/scores than the benign ones.</p>\n\n<p>I.e. if you predict 0.0000001 for all malignant images and 0 for all benign, your AUC=1.0</p>",
      "rawMarkdown": "How do you know whether a prediction is malignant? The competition metric is AUC. So you just have to make sure that your malignant predictions get higher predictions/scores than the benign ones.\n\nI.e. if you predict 0.0000001 for all malignant images and 0 for all benign, your AUC=1.0",
      "votes": null
    },
    {
      "id": "961439",
      "postDate": "08/07/2020 07:29:04",
      "content": "<p>I actually round the target column to 0 and 1 and counting it. something like this <br>\n<code>df = pd.read_csv('sub.csv')</code><br>\n<code>df.target.round().value_counts()</code></p>\n<p>this gave 0s and 1s count. For most of the public kernel, it is around 40 for 1s and all other 0s. So I am just saying that 40 1s is too low. In one of the discussion, the first place holder said that he thinks there might be around 27x malignant in the private LB. In one notebook <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> showed there are 77 or 78 malignant images in public LB.  </p>\n<p>I think what I said and what you interpreted are two different things. </p>",
      "rawMarkdown": "I actually round the target column to 0 and 1 and counting it. something like this \n`df = pd.read_csv('sub.csv')`\n`df.target.round().value_counts()`\n\nthis gave 0s and 1s count. For most of the public kernel, it is around 40 for 1s and all other 0s. So I am just saying that 40 1s is too low. In one of the discussion, the first place holder said that he thinks there might be around 27x malignant in the private LB. In one notebook @cpmpml showed there are 77 or 78 malignant images in public LB.  \n\nI think what I said and what you interpreted are two different things.",
      "votes": null
    },
    {
      "id": "961462",
      "postDate": "08/07/2020 07:45:08",
      "content": "<p>By rounding, you will only assign 1 to a prediction if it is &gt; 0.5.</p>\n\n<p>This doesn't matter for AUC. So again, if we would assign those 77 or 78 images a score of 0.000000000000001 and all others 0, then your AUC is 1.0. They don't even need to be in the range 0-1. If you multiply all your predictions with, for example, 1000000, your AUC will be identical.</p>\n\n<p>So the low counts you report for malignant predictions are <strong>irrelevant</strong> for AUC. They are relevant for other metrics such as accuracy, precision and recall.</p>",
      "rawMarkdown": "By rounding, you will only assign 1 to a prediction if it is &gt; 0.5.\n\nThis doesn't matter for AUC. So again, if we would assign those 77 or 78 images a score of 0.000000000000001 and all others 0, then your AUC is 1.0. They don't even need to be in the range 0-1. If you multiply all your predictions with, for example, 1000000, your AUC will be identical.\n\nSo the low counts you report for malignant predictions are **irrelevant** for AUC. They are relevant for other metrics such as accuracy, precision and recall.",
      "votes": null
    },
    {
      "id": "961465",
      "postDate": "08/07/2020 07:47:25",
      "content": "<p>Yes someone did it..</p>",
      "rawMarkdown": "Yes someone did it..",
      "votes": null
    },
    {
      "id": "961468",
      "postDate": "08/07/2020 07:48:30",
      "content": "<p><a href=\"/group16\">@group16</a> Great piece of advice..</p>",
      "rawMarkdown": "group16 Great piece of advice..",
      "votes": null
    },
    {
      "id": "961470",
      "postDate": "08/07/2020 07:50:38",
      "content": "<p>But this will disrupt the ranking and ruins many people's hard work...</p>",
      "rawMarkdown": "But this will disrupt the ranking and ruins many people's hard work...",
      "votes": null
    },
    {
      "id": "961472",
      "postDate": "08/07/2020 07:56:49",
      "content": "<p>Yeah, only the order matters for AUC.</p>",
      "rawMarkdown": "Yeah, only the order matters for AUC.",
      "votes": null
    },
    {
      "id": "961480",
      "postDate": "08/07/2020 08:01:19",
      "content": "<p>Thanks for your insight <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>. I understand what you mean. Appreciate it. </p>",
      "rawMarkdown": "Thanks for your insight @group16. I understand what you mean. Appreciate it.",
      "votes": null
    },
    {
      "id": "961489",
      "postDate": "08/07/2020 08:23:08",
      "content": "<p>That's why in some ensembles, we can see people using ranking, because the order matters for AUC and not the value itself! </p>",
      "rawMarkdown": "That's why in some ensembles, we can see people using ranking, because the order matters for AUC and not the value itself!",
      "votes": null
    },
    {
      "id": "961692",
      "postDate": "08/07/2020 12:22:24",
      "content": "<p><a href=\"/philippsinger\">@philippsinger</a> <a href=\"/group16\">@group16</a> Could you please explain what is this rank and how it does affect the AUC and all..I'm still confused..</p>",
      "rawMarkdown": "philippsinger @group16 Could you please explain what is this rank and how it does affect the AUC and all..I'm still confused..",
      "votes": null
    },
    {
      "id": "961695",
      "postDate": "08/07/2020 12:27:45",
      "content": "<p><a href=\"https://www.kaggle.com/msharuk589\" target=\"_blank\">@msharuk589</a> why dont you just read up on what AUC is? AUC is solely defined by how the predictions are ranked.</p>",
      "rawMarkdown": "msharuk589 why dont you just read up on what AUC is? AUC is solely defined by how the predictions are ranked.",
      "votes": null
    },
    {
      "id": "961697",
      "postDate": "08/07/2020 12:28:44",
      "content": "<p>Of course. The easiest interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".</p>\n\n<p>The important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is <strong>irrelevant</strong> to AUC.</p>\n\n<p>Hope this helps, else I would google for some blog that explains AUC intuitively :)</p>",
      "rawMarkdown": "Of course. The easiest interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".\n\nThe important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is **irrelevant** to AUC.\n\nHope this helps, else I would google for some blog that explains AUC intuitively :)",
      "votes": null
    },
    {
      "id": "961699",
      "postDate": "08/07/2020 12:30:36",
      "content": "<p>Another way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.</p>\n\n<p>If you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0</p>\n\n<p>Here's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F3164e370157caf9928bd84ef807cd05a%2F0_eM6jPtzsMCg3He0_.png?generation=1596803709911134&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Another way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.\n\nIf you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0\n\nHere's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F3164e370157caf9928bd84ef807cd05a%2F0_eM6jPtzsMCg3He0_.png?generation=1596803709911134&amp;alt=media)",
      "votes": null
    },
    {
      "id": "961704",
      "postDate": "08/07/2020 12:37:51",
      "content": "<p><a href=\"/group16\">@group16</a> Thanks for the info..I understood now\n( <a href=\"/philippsinger\">@philippsinger</a> ...🙄 )</p>",
      "rawMarkdown": "group16 Thanks for the info..I understood now\n( @philippsinger ...🙄 )",
      "votes": null
    },
    {
      "id": "961743",
      "postDate": "08/07/2020 13:19:06",
      "content": "<p><a href=\"/group16\">@group16</a>  Wow .. you completely nailed the explanation. I also planning to design something like this but you did it...</p>",
      "rawMarkdown": "group16  Wow .. you completely nailed the explanation. I also planning to design something like this but you did it...",
      "votes": null
    },
    {
      "id": "961755",
      "postDate": "08/07/2020 13:30:22",
      "content": "<p>Public LB doesn’t mean anything at this stage  … After competition is over we will know real standings:) Hopefully your hard work will pay off. </p>",
      "rawMarkdown": "Public LB doesn’t mean anything at this stage  ... After competition is over we will know real standings:) Hopefully your hard work will pay off.",
      "votes": null
    },
    {
      "id": "962164",
      "postDate": "08/07/2020 21:30:50",
      "content": "<p>I also participated in global wheat, just wait for the massive shakeup. I expect to see +/- 500+ for tons of people. Don't worry, if you have good cv, you will do well. </p>",
      "rawMarkdown": "I also participated in global wheat, just wait for the massive shakeup. I expect to see +/- 500+ for tons of people. Don't worry, if you have good cv, you will do well.",
      "votes": null
    },
    {
      "id": "962276",
      "postDate": "08/08/2020 01:51:55",
      "content": "<p>Offcourse <a href=\"/drhabib\">@drhabib</a>, will wait to witness that story too..  </p>",
      "rawMarkdown": "Offcourse @drhabib, will wait to witness that story too..",
      "votes": null
    },
    {
      "id": "963700",
      "postDate": "08/09/2020 08:12:21",
      "content": "<p>I'm relatively new to kaggle so correct me if i'm wrong here. But isn't these high public LB blend very likely to overfit? The blends seems to be using LB score to tune.</p>",
      "rawMarkdown": "I'm relatively new to kaggle so correct me if i'm wrong here. But isn't these high public LB blend very likely to overfit? The blends seems to be using LB score to tune.",
      "votes": null
    },
    {
      "id": "964982",
      "postDate": "08/10/2020 09:59:15",
      "content": "<p>yeah someone shared his notebook with submission csv file which got LB 0.9619. At this point they should share ideas only rather than csv in case of such high scores. yes it indeed ruins fellow kagglers' months of hardwork.</p>",
      "rawMarkdown": "yeah someone shared his notebook with submission csv file which got LB 0.9619. At this point they should share ideas only rather than csv in case of such high scores. yes it indeed ruins fellow kagglers' months of hardwork.",
      "votes": null
    },
    {
      "id": "965374",
      "postDate": "08/10/2020 15:31:01",
      "content": "<p>Yep It's true. It took me nearly 1.5 months and more than 65  submissions to reach 0.9643.</p>",
      "rawMarkdown": "Yep It's true. It took me nearly 1.5 months and more than 65  submissions to reach 0.9643.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961392,
      "author_name": "sarques",
      "author_url": "",
      "post_date": "08/07/2020 06:29:58",
      "content": "<p>Maybe someone has shared their solution which got 0.9619, also, of course many of them would be having there own submissions, otherwise its just a coincidence, it may happen at the top since the leaderboard gets denser upwards! :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 961465,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "08/07/2020 07:47:25",
          "content": "<p>Yes someone did it..</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 961413,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "08/07/2020 07:12:03",
      "content": "<p>There is nothing to worry about the public LB. The submission that got 0.9619 score has almost 0 malignant predictions. Not just that, other public high scoring NBs have around 40 malignant prediction. That is too low from what I understand so far. So, do not worry about the public score. In private LB, we will have shake up for sure. </p>",
      "votes": null,
      "replies": [
        {
          "id": 961420,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/07/2020 07:18:12",
          "content": "<p>How do you know whether a prediction is malignant? The competition metric is AUC. So you just have to make sure that your malignant predictions get higher predictions/scores than the benign ones.</p>\n\n<p>I.e. if you predict 0.0000001 for all malignant images and 0 for all benign, your AUC=1.0</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961439,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "08/07/2020 07:29:04",
          "content": "<p>I actually round the target column to 0 and 1 and counting it. something like this <br>\n<code>df = pd.read_csv('sub.csv')</code><br>\n<code>df.target.round().value_counts()</code></p>\n<p>this gave 0s and 1s count. For most of the public kernel, it is around 40 for 1s and all other 0s. So I am just saying that 40 1s is too low. In one of the discussion, the first place holder said that he thinks there might be around 27x malignant in the private LB. In one notebook <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> showed there are 77 or 78 malignant images in public LB.  </p>\n<p>I think what I said and what you interpreted are two different things. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961462,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/07/2020 07:45:08",
          "content": "<p>By rounding, you will only assign 1 to a prediction if it is &gt; 0.5.</p>\n\n<p>This doesn't matter for AUC. So again, if we would assign those 77 or 78 images a score of 0.000000000000001 and all others 0, then your AUC is 1.0. They don't even need to be in the range 0-1. If you multiply all your predictions with, for example, 1000000, your AUC will be identical.</p>\n\n<p>So the low counts you report for malignant predictions are <strong>irrelevant</strong> for AUC. They are relevant for other metrics such as accuracy, precision and recall.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961468,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "08/07/2020 07:48:30",
          "content": "<p><a href=\"/group16\">@group16</a> Great piece of advice..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961472,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/07/2020 07:56:49",
          "content": "<p>Yeah, only the order matters for AUC.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961480,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "08/07/2020 08:01:19",
          "content": "<p>Thanks for your insight <a href=\"https://www.kaggle.com/group16\" target=\"_blank\">@group16</a>. I understand what you mean. Appreciate it. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961489,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/07/2020 08:23:08",
          "content": "<p>That's why in some ensembles, we can see people using ranking, because the order matters for AUC and not the value itself! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961692,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "08/07/2020 12:22:24",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> <a href=\"/group16\">@group16</a> Could you please explain what is this rank and how it does affect the AUC and all..I'm still confused..</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961695,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "08/07/2020 12:27:45",
          "content": "<p><a href=\"https://www.kaggle.com/msharuk589\" target=\"_blank\">@msharuk589</a> why dont you just read up on what AUC is? AUC is solely defined by how the predictions are ranked.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961697,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/07/2020 12:28:44",
          "content": "<p>Of course. The easiest interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".</p>\n\n<p>The important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is <strong>irrelevant</strong> to AUC.</p>\n\n<p>Hope this helps, else I would google for some blog that explains AUC intuitively :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961699,
          "author_name": "group16",
          "author_url": "",
          "post_date": "08/07/2020 12:30:36",
          "content": "<p>Another way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.</p>\n\n<p>If you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0</p>\n\n<p>Here's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F3164e370157caf9928bd84ef807cd05a%2F0_eM6jPtzsMCg3He0_.png?generation=1596803709911134&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961704,
          "author_name": "msharuk589",
          "author_url": "",
          "post_date": "08/07/2020 12:37:51",
          "content": "<p><a href=\"/group16\">@group16</a> Thanks for the info..I understood now\n( <a href=\"/philippsinger\">@philippsinger</a> ...🙄 )</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961743,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "08/07/2020 13:19:06",
          "content": "<p><a href=\"/group16\">@group16</a>  Wow .. you completely nailed the explanation. I also planning to design something like this but you did it...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 961755,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "08/07/2020 13:30:22",
      "content": "<p>Public LB doesn’t mean anything at this stage  … After competition is over we will know real standings:) Hopefully your hard work will pay off. </p>",
      "votes": null,
      "replies": [
        {
          "id": 962276,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "08/08/2020 01:51:55",
          "content": "<p>Offcourse <a href=\"/drhabib\">@drhabib</a>, will wait to witness that story too..  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 965374,
      "author_name": "",
      "author_url": "",
      "post_date": "08/10/2020 15:31:01",
      "content": "<p>Yep It's true. It took me nearly 1.5 months and more than 65  submissions to reach 0.9643.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 961395,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "08/07/2020 06:38:41",
      "content": "<p>don't worry too much. </p>",
      "votes": null,
      "replies": [
        {
          "id": 961470,
          "author_name": "vin1234",
          "author_url": "",
          "post_date": "08/07/2020 07:50:38",
          "content": "<p>But this will disrupt the ranking and ruins many people's hard work...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 962164,
          "author_name": "stanleyjzheng",
          "author_url": "",
          "post_date": "08/07/2020 21:30:50",
          "content": "<p>I also participated in global wheat, just wait for the massive shakeup. I expect to see +/- 500+ for tons of people. Don't worry, if you have good cv, you will do well. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 963700,
      "author_name": "brachester",
      "author_url": "",
      "post_date": "08/09/2020 08:12:21",
      "content": "<p>I'm relatively new to kaggle so correct me if i'm wrong here. But isn't these high public LB blend very likely to overfit? The blends seems to be using LB score to tune.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 964982,
      "author_name": "pawankumarsahu",
      "author_url": "",
      "post_date": "08/10/2020 09:59:15",
      "content": "<p>yeah someone shared his notebook with submission csv file which got LB 0.9619. At this point they should share ideas only rather than csv in case of such high scores. yes it indeed ruins fellow kagglers' months of hardwork.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "961346": "I found that a lot of people got 96.19 with a few submissions. It took months to reach that stage but few people able to achieve it within multiple submissions... \n\nCheck Link below",
    "961392": "Maybe someone has shared their solution which got 0.9619, also, of course many of them would be having there own submissions, otherwise its just a coincidence, it may happen at the top since the leaderboard gets denser upwards! :)",
    "961395": "don't worry too much.",
    "961413": "There is nothing to worry about the public LB. The submission that got 0.9619 score has almost 0 malignant predictions. Not just that, other public high scoring NBs have around 40 malignant prediction. That is too low from what I understand so far. So, do not worry about the public score. In private LB, we will have shake up for sure.",
    "961420": "How do you know whether a prediction is malignant? The competition metric is AUC. So you just have to make sure that your malignant predictions get higher predictions/scores than the benign ones.\n\nI.e. if you predict 0.0000001 for all malignant images and 0 for all benign, your AUC=1.0",
    "961439": "I actually round the target column to 0 and 1 and counting it. something like this \n`df = pd.read_csv('sub.csv')`\n`df.target.round().value_counts()`\n\nthis gave 0s and 1s count. For most of the public kernel, it is around 40 for 1s and all other 0s. So I am just saying that 40 1s is too low. In one of the discussion, the first place holder said that he thinks there might be around 27x malignant in the private LB. In one notebook @cpmpml showed there are 77 or 78 malignant images in public LB.  \n\nI think what I said and what you interpreted are two different things.",
    "961462": "By rounding, you will only assign 1 to a prediction if it is &gt; 0.5.\n\nThis doesn't matter for AUC. So again, if we would assign those 77 or 78 images a score of 0.000000000000001 and all others 0, then your AUC is 1.0. They don't even need to be in the range 0-1. If you multiply all your predictions with, for example, 1000000, your AUC will be identical.\n\nSo the low counts you report for malignant predictions are **irrelevant** for AUC. They are relevant for other metrics such as accuracy, precision and recall.",
    "961465": "Yes someone did it..",
    "961468": "group16 Great piece of advice..",
    "961470": "But this will disrupt the ranking and ruins many people's hard work...",
    "961472": "Yeah, only the order matters for AUC.",
    "961480": "Thanks for your insight @group16. I understand what you mean. Appreciate it.",
    "961489": "That's why in some ensembles, we can see people using ranking, because the order matters for AUC and not the value itself!",
    "961692": "philippsinger @group16 Could you please explain what is this rank and how it does affect the AUC and all..I'm still confused..",
    "961695": "msharuk589 why dont you just read up on what AUC is? AUC is solely defined by how the predictions are ranked.",
    "961697": "Of course. The easiest interpretation is the following: \"if I would take 1 positive sample and 1 negative sample, what is the probability that our model assigned a higher score to the positive sample\".\n\nThe important detail here is, is that the only thing that matters is that the positive samples get higher scores. How high these scores or how much higher these scores are, is **irrelevant** to AUC.\n\nHope this helps, else I would google for some blog that explains AUC intuitively :)",
    "961699": "Another way of looking at AUC is that it measures the degree of separation between predictions of positive samples and negative samples.\n\nIf you plot a histogram of negative samples and a histogram of positive samples, then you want the histogram of the positive to be to the right of the negative. If there is no overlap between the histograms, then your AUC is 1.0\n\nHere's an image to make this more clear: assume that the red are predictions of positive samples and blue predictions of negative samples. 1-AUC corresponds to the amount of overlap between these two histograms/kde\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F443651%2F3164e370157caf9928bd84ef807cd05a%2F0_eM6jPtzsMCg3He0_.png?generation=1596803709911134&amp;alt=media)",
    "961704": "group16 Thanks for the info..I understood now\n( @philippsinger ...🙄 )",
    "961743": "group16  Wow .. you completely nailed the explanation. I also planning to design something like this but you did it...",
    "961755": "Public LB doesn’t mean anything at this stage  ... After competition is over we will know real standings:) Hopefully your hard work will pay off.",
    "962164": "I also participated in global wheat, just wait for the massive shakeup. I expect to see +/- 500+ for tons of people. Don't worry, if you have good cv, you will do well.",
    "962276": "Offcourse @drhabib, will wait to witness that story too..",
    "963700": "I'm relatively new to kaggle so correct me if i'm wrong here. But isn't these high public LB blend very likely to overfit? The blends seems to be using LB score to tune.",
    "964982": "yeah someone shared his notebook with submission csv file which got LB 0.9619. At this point they should share ideas only rather than csv in case of such high scores. yes it indeed ruins fellow kagglers' months of hardwork.",
    "965374": "Yep It's true. It took me nearly 1.5 months and more than 65  submissions to reach 0.9643."
  },
  "source": "meta"
}