{
  "id": 273655,
  "title": "Thoughts on Expected AUC for Random classifier on public test set",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655",
  "author_name": "",
  "post_date": "2021-09-22T02:57:07.896908200Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi All, <br>\nI wanted to share some thoughts on Expected AUC for Random classifier on public test set. </p>\n<ol>\n<li>I think a Random classifier <strong>will have 0.5 AUC on public test</strong>. I tried 4 submissions to verify:  </li>\n<li>1st submission: all 0.500; the public test AUC is exactly 0.500.</li>\n<li>2nd submission: all 0.000; the public test AUC is exactly 0.500.</li>\n<li>3rd submission: all 1.000; the public test AUC is exactly 0.500. </li>\n<li>4th submission: all 0 and only one 1; the public test AUC was 0.500. </li>\n<li>I saw a very nice simulation notebook here: </li>\n<li><a href=\"https://www.kaggle.com/osciiart/public-lb-simulation\" target=\"_blank\">https://www.kaggle.com/osciiart/public-lb-simulation</a> . the conclusion in it is: </li>\n<li>\"The simulations above shows that a score of about 0.65 is almost by chance\"</li>\n<li>I believe this is incorrect; in this notebook if we run 100 simulations, max AUC is 0.65</li>\n<li>Same notebook if you run 1000 simulations the max AUC is 0.68 and 0.71 for 10000 sims. </li>\n<li>I think the simulation means that there is <strong>1/100 chance of random classifier getting 0.65</strong>.</li>\n<li>and <strong>1/10000 chance of a random classifier getting AUC 0.71</strong> on public test set. </li>\n</ol>\n<p>I will be very happy to hear yours comments on this topic. </p>",
  "messages": [
    {
      "id": "1519896",
      "postDate": "09/22/2021 02:57:07",
      "content": "<p>Hi All, <br>\nI wanted to share some thoughts on Expected AUC for Random classifier on public test set. </p>\n<ol>\n<li>I think a Random classifier <strong>will have 0.5 AUC on public test</strong>. I tried 4 submissions to verify:  </li>\n<li>1st submission: all 0.500; the public test AUC is exactly 0.500.</li>\n<li>2nd submission: all 0.000; the public test AUC is exactly 0.500.</li>\n<li>3rd submission: all 1.000; the public test AUC is exactly 0.500. </li>\n<li>4th submission: all 0 and only one 1; the public test AUC was 0.500. </li>\n<li>I saw a very nice simulation notebook here: </li>\n<li><a href=\"https://www.kaggle.com/osciiart/public-lb-simulation\" target=\"_blank\">https://www.kaggle.com/osciiart/public-lb-simulation</a> . the conclusion in it is: </li>\n<li>\"The simulations above shows that a score of about 0.65 is almost by chance\"</li>\n<li>I believe this is incorrect; in this notebook if we run 100 simulations, max AUC is 0.65</li>\n<li>Same notebook if you run 1000 simulations the max AUC is 0.68 and 0.71 for 10000 sims. </li>\n<li>I think the simulation means that there is <strong>1/100 chance of random classifier getting 0.65</strong>.</li>\n<li>and <strong>1/10000 chance of a random classifier getting AUC 0.71</strong> on public test set. </li>\n</ol>\n<p>I will be very happy to hear yours comments on this topic. </p>",
      "rawMarkdown": "Hi All, \nI wanted to share some thoughts on Expected AUC for Random classifier on public test set. \n1. I think a Random classifier **will have 0.5 AUC on public test**. I tried 4 submissions to verify:  \n2. 1st submission: all 0.500; the public test AUC is exactly 0.500.\n3. 2nd submission: all 0.000; the public test AUC is exactly 0.500.\n4. 3rd submission: all 1.000; the public test AUC is exactly 0.500. \n5. 4th submission: all 0 and only one 1; the public test AUC was 0.500. \n6. I saw a very nice simulation notebook here: \n7. https://www.kaggle.com/osciiart/public-lb-simulation . the conclusion in it is: \n8. \"The simulations above shows that a score of about 0.65 is almost by chance\"\n9.  I believe this is incorrect; in this notebook if we run 100 simulations, max AUC is 0.65\n10. Same notebook if you run 1000 simulations the max AUC is 0.68 and 0.71 for 10000 sims. \n11. I think the simulation means that there is **1/100 chance of random classifier getting 0.65**.\n12. and **1/10000 chance of a random classifier getting AUC 0.71** on public test set. \n\nI will be very happy to hear yours comments on this topic.",
      "votes": null
    },
    {
      "id": "1519899",
      "postDate": "09/22/2021 02:59:49",
      "content": "<p><a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> thank you for sharing the very nice simulation notebook. it was very helpful. <br>\nI have some thoughts on what a random classifier AUC would be on public test set. <br>\nCould you please see them when you have a chance ? Please correct me if I am wrong.</p>",
      "rawMarkdown": "osciiart thank you for sharing the very nice simulation notebook. it was very helpful. \nI have some thoughts on what a random classifier AUC would be on public test set. \nCould you please see them when you have a chance ? Please correct me if I am wrong.",
      "votes": null
    },
    {
      "id": "1519900",
      "postDate": "09/22/2021 03:03:27",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> just curious to hear your thoughts on this topic :) </p>",
      "rawMarkdown": "dschettler8845 just curious to hear your thoughts on this topic :)",
      "votes": null
    },
    {
      "id": "1519965",
      "postDate": "09/22/2021 04:29:34",
      "content": "<p>I found some papers that have tackled this issue in the past. <br>\nkaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273671<br>\nSome of the best AUC is around 0.85. </p>",
      "rawMarkdown": "I found some papers that have tackled this issue in the past. \nkaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273671\nSome of the best AUC is around 0.85.",
      "votes": null
    },
    {
      "id": "1522690",
      "postDate": "09/24/2021 13:14:52",
      "content": "<p>I think <strong>11.</strong> and <strong>12.</strong> are about right for a single random submit. I ran the same simulations 1M times and got very similar statistics:</p>\n<pre><code>(1 random sub) percentile-50 : 0.5\n(1 random sub) percentile-90 : 0.5804232804232804\n(1 random sub) percentile-95 : 0.6026455026455027\n(1 random sub) percentile-99 : 0.6439153439153439\n(1 random sub) percentile-99.9 : 0.687831216931219\n(1 random sub) percentile-99.99 : 0.7285714814814763\n</code></pre>\n<p>However, the leaderboard will show your highest scoring submit, so if you submit random values 10 times, you are likely to have LB over 0.59</p>\n<pre><code>(10 random subs) percentile-50 : 0.5941798941798941\n(10 random subs) percentile-90 : 0.6423280423280423\n(10 random subs) percentile-95 : 0.6587301587301587\n(10 random subs) percentile-99 : 0.6862433862433863\n(10 random subs) percentile-99.9 : 0.7285714285714286\n(10 random subs) percentile-99.99 : 0.764021164021164\n</code></pre>\n<p>…and 20 subs will likely get you above 0.61</p>\n<pre><code>(20 random subs) percentile-50 : 0.6132275132275132\n(20 random subs) percentile-90 : 0.6566137566137566\n(20 random subs) percentile-95 : 0.6714285714285714\n(20 random subs) percentile-99 : 0.7005291005291006\n(20 random subs) percentile-99.9 : 0.7391534391534391\n(20 random subs) percentile-99.99 : 0.764021164021164\n</code></pre>\n<p>By chance, my first random-value submit to this competition got an LB score of 0.65</p>",
      "rawMarkdown": "I think **11.** and **12.** are about right for a single random submit. I ran the same simulations 1M times and got very similar statistics:\n\n```\n(1 random sub) percentile-50 : 0.5\n(1 random sub) percentile-90 : 0.5804232804232804\n(1 random sub) percentile-95 : 0.6026455026455027\n(1 random sub) percentile-99 : 0.6439153439153439\n(1 random sub) percentile-99.9 : 0.687831216931219\n(1 random sub) percentile-99.99 : 0.7285714814814763\n``` \n\nHowever, the leaderboard will show your highest scoring submit, so if you submit random values 10 times, you are likely to have LB over 0.59\n\n```\n(10 random subs) percentile-50 : 0.5941798941798941\n(10 random subs) percentile-90 : 0.6423280423280423\n(10 random subs) percentile-95 : 0.6587301587301587\n(10 random subs) percentile-99 : 0.6862433862433863\n(10 random subs) percentile-99.9 : 0.7285714285714286\n(10 random subs) percentile-99.99 : 0.764021164021164\n```\n...and 20 subs will likely get you above 0.61\n\n```\n(20 random subs) percentile-50 : 0.6132275132275132\n(20 random subs) percentile-90 : 0.6566137566137566\n(20 random subs) percentile-95 : 0.6714285714285714\n(20 random subs) percentile-99 : 0.7005291005291006\n(20 random subs) percentile-99.9 : 0.7391534391534391\n(20 random subs) percentile-99.99 : 0.764021164021164\n```\n\nBy chance, my first random-value submit to this competition got an LB score of 0.65",
      "votes": null
    },
    {
      "id": "1522894",
      "postDate": "09/24/2021 17:25:41",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/qitvision\" target=\"_blank\">@qitvision</a> ! thank you so much for your feedback! <br>\nIf possible, could you share the code for how you did this calculation: <br>\n\"(20 random subs) percentile-50 : 0.6132275132275132\"</p>\n<p>Or just the algorithm steps to come to this conclusion. </p>",
      "rawMarkdown": "Hi @qitvision ! thank you so much for your feedback! \nIf possible, could you share the code for how you did this calculation: \n\"(20 random subs) percentile-50 : 0.6132275132275132\"\n\nOr just the algorithm steps to come to this conclusion.",
      "votes": null
    },
    {
      "id": "1522987",
      "postDate": "09/24/2021 19:55:56",
      "content": "<p>Sure, here's my fork of the public LB simulation notebook: <a href=\"https://www.kaggle.com/qitvision/public-lb-simulation/\" target=\"_blank\">https://www.kaggle.com/qitvision/public-lb-simulation/</a></p>\n<p>I picked k samples randomly from 1M simulation run score-pool and kept maximum scores. Did it enough times (10K) and calculated percentiles.</p>",
      "rawMarkdown": "Sure, here's my fork of the public LB simulation notebook: https://www.kaggle.com/qitvision/public-lb-simulation/\n\nI picked k samples randomly from 1M simulation run score-pool and kept maximum scores. Did it enough times (10K) and calculated percentiles.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1519899,
      "author_name": "mpsampat",
      "author_url": "",
      "post_date": "09/22/2021 02:59:49",
      "content": "<p><a href=\"https://www.kaggle.com/osciiart\" target=\"_blank\">@osciiart</a> thank you for sharing the very nice simulation notebook. it was very helpful. <br>\nI have some thoughts on what a random classifier AUC would be on public test set. <br>\nCould you please see them when you have a chance ? Please correct me if I am wrong.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1519900,
      "author_name": "mpsampat",
      "author_url": "",
      "post_date": "09/22/2021 03:03:27",
      "content": "<p><a href=\"https://www.kaggle.com/dschettler8845\" target=\"_blank\">@dschettler8845</a> just curious to hear your thoughts on this topic :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1519965,
      "author_name": "mpsampat",
      "author_url": "",
      "post_date": "09/22/2021 04:29:34",
      "content": "<p>I found some papers that have tackled this issue in the past. <br>\nkaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273671<br>\nSome of the best AUC is around 0.85. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1522690,
      "author_name": "qitvision",
      "author_url": "",
      "post_date": "09/24/2021 13:14:52",
      "content": "<p>I think <strong>11.</strong> and <strong>12.</strong> are about right for a single random submit. I ran the same simulations 1M times and got very similar statistics:</p>\n<pre><code>(1 random sub) percentile-50 : 0.5\n(1 random sub) percentile-90 : 0.5804232804232804\n(1 random sub) percentile-95 : 0.6026455026455027\n(1 random sub) percentile-99 : 0.6439153439153439\n(1 random sub) percentile-99.9 : 0.687831216931219\n(1 random sub) percentile-99.99 : 0.7285714814814763\n</code></pre>\n<p>However, the leaderboard will show your highest scoring submit, so if you submit random values 10 times, you are likely to have LB over 0.59</p>\n<pre><code>(10 random subs) percentile-50 : 0.5941798941798941\n(10 random subs) percentile-90 : 0.6423280423280423\n(10 random subs) percentile-95 : 0.6587301587301587\n(10 random subs) percentile-99 : 0.6862433862433863\n(10 random subs) percentile-99.9 : 0.7285714285714286\n(10 random subs) percentile-99.99 : 0.764021164021164\n</code></pre>\n<p>…and 20 subs will likely get you above 0.61</p>\n<pre><code>(20 random subs) percentile-50 : 0.6132275132275132\n(20 random subs) percentile-90 : 0.6566137566137566\n(20 random subs) percentile-95 : 0.6714285714285714\n(20 random subs) percentile-99 : 0.7005291005291006\n(20 random subs) percentile-99.9 : 0.7391534391534391\n(20 random subs) percentile-99.99 : 0.764021164021164\n</code></pre>\n<p>By chance, my first random-value submit to this competition got an LB score of 0.65</p>",
      "votes": null,
      "replies": [
        {
          "id": 1522894,
          "author_name": "mpsampat",
          "author_url": "",
          "post_date": "09/24/2021 17:25:41",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/qitvision\" target=\"_blank\">@qitvision</a> ! thank you so much for your feedback! <br>\nIf possible, could you share the code for how you did this calculation: <br>\n\"(20 random subs) percentile-50 : 0.6132275132275132\"</p>\n<p>Or just the algorithm steps to come to this conclusion. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1522987,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "09/24/2021 19:55:56",
          "content": "<p>Sure, here's my fork of the public LB simulation notebook: <a href=\"https://www.kaggle.com/qitvision/public-lb-simulation/\" target=\"_blank\">https://www.kaggle.com/qitvision/public-lb-simulation/</a></p>\n<p>I picked k samples randomly from 1M simulation run score-pool and kept maximum scores. Did it enough times (10K) and calculated percentiles.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1519896": "Hi All, \nI wanted to share some thoughts on Expected AUC for Random classifier on public test set. \n1. I think a Random classifier **will have 0.5 AUC on public test**. I tried 4 submissions to verify:  \n2. 1st submission: all 0.500; the public test AUC is exactly 0.500.\n3. 2nd submission: all 0.000; the public test AUC is exactly 0.500.\n4. 3rd submission: all 1.000; the public test AUC is exactly 0.500. \n5. 4th submission: all 0 and only one 1; the public test AUC was 0.500. \n6. I saw a very nice simulation notebook here: \n7. https://www.kaggle.com/osciiart/public-lb-simulation . the conclusion in it is: \n8. \"The simulations above shows that a score of about 0.65 is almost by chance\"\n9.  I believe this is incorrect; in this notebook if we run 100 simulations, max AUC is 0.65\n10. Same notebook if you run 1000 simulations the max AUC is 0.68 and 0.71 for 10000 sims. \n11. I think the simulation means that there is **1/100 chance of random classifier getting 0.65**.\n12. and **1/10000 chance of a random classifier getting AUC 0.71** on public test set. \n\nI will be very happy to hear yours comments on this topic.",
    "1519899": "osciiart thank you for sharing the very nice simulation notebook. it was very helpful. \nI have some thoughts on what a random classifier AUC would be on public test set. \nCould you please see them when you have a chance ? Please correct me if I am wrong.",
    "1519900": "dschettler8845 just curious to hear your thoughts on this topic :)",
    "1519965": "I found some papers that have tackled this issue in the past. \nkaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273671\nSome of the best AUC is around 0.85.",
    "1522690": "I think **11.** and **12.** are about right for a single random submit. I ran the same simulations 1M times and got very similar statistics:\n\n```\n(1 random sub) percentile-50 : 0.5\n(1 random sub) percentile-90 : 0.5804232804232804\n(1 random sub) percentile-95 : 0.6026455026455027\n(1 random sub) percentile-99 : 0.6439153439153439\n(1 random sub) percentile-99.9 : 0.687831216931219\n(1 random sub) percentile-99.99 : 0.7285714814814763\n``` \n\nHowever, the leaderboard will show your highest scoring submit, so if you submit random values 10 times, you are likely to have LB over 0.59\n\n```\n(10 random subs) percentile-50 : 0.5941798941798941\n(10 random subs) percentile-90 : 0.6423280423280423\n(10 random subs) percentile-95 : 0.6587301587301587\n(10 random subs) percentile-99 : 0.6862433862433863\n(10 random subs) percentile-99.9 : 0.7285714285714286\n(10 random subs) percentile-99.99 : 0.764021164021164\n```\n...and 20 subs will likely get you above 0.61\n\n```\n(20 random subs) percentile-50 : 0.6132275132275132\n(20 random subs) percentile-90 : 0.6566137566137566\n(20 random subs) percentile-95 : 0.6714285714285714\n(20 random subs) percentile-99 : 0.7005291005291006\n(20 random subs) percentile-99.9 : 0.7391534391534391\n(20 random subs) percentile-99.99 : 0.764021164021164\n```\n\nBy chance, my first random-value submit to this competition got an LB score of 0.65",
    "1522894": "Hi @qitvision ! thank you so much for your feedback! \nIf possible, could you share the code for how you did this calculation: \n\"(20 random subs) percentile-50 : 0.6132275132275132\"\n\nOr just the algorithm steps to come to this conclusion.",
    "1522987": "Sure, here's my fork of the public LB simulation notebook: https://www.kaggle.com/qitvision/public-lb-simulation/\n\nI picked k samples randomly from 1M simulation run score-pool and kept maximum scores. Did it enough times (10K) and calculated percentiles."
  },
  "source": "meta"
}