{
  "id": 280296,
  "title": "Expected AUC for \"N\" submissions for a Random Classifier ",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/280296",
  "author_name": "",
  "post_date": "2021-10-20T23:33:47.479768100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<ol>\n<li>In the past, I had made this discussion thread to discuss random classifier (<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655\" target=\"_blank\">link</a>)</li>\n<li>there are some notebook links in it.</li>\n<li>I received great feedback from <a href=\"https://www.kaggle.com/qitvision\" target=\"_blank\">@qitvision</a> on it. </li>\n<li>To summarize that discussion thread, there are two main points: </li>\n<li>Score after 1 random submission versus score after 10,20, N submissions: for example: </li>\n<li>There is \"1/100 chance of random classifier getting 0.65\" for one random submission; </li>\n<li>for 10 random submissions (<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655\" target=\"_blank\">link</a>)<br>\n(10 random subs) percentile-50 : 0.5941798941798941<br>\n(10 random subs) percentile-90 : 0.6423280423280423<br>\n(10 random subs) percentile-95 : 0.6587301587301587<br>\n(10 random subs) percentile-99 : 0.6862433862433863<br>\n(10 random subs) percentile-99.9 : 0.7285714285714286<br>\n(10 random subs) percentile-99.99 : 0.764021164021164</li>\n<li>for 20 random submissions:(<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655\" target=\"_blank\">link</a>)<br>\n(20 random subs) percentile-50 : 0.6132275132275132<br>\n(20 random subs) percentile-90 : 0.6566137566137566<br>\n(20 random subs) percentile-95 : 0.6714285714285714<br>\n(20 random subs) percentile-99 : 0.7005291005291006<br>\n(20 random subs) percentile-99.9 : 0.7391534391534391<br>\n(20 random subs) percentile-99.99 : 0.764021164021164</li>\n</ol>",
  "messages": [
    {
      "id": "1551875",
      "postDate": "10/20/2021 23:33:47",
      "content": "<ol>\n<li>In the past, I had made this discussion thread to discuss random classifier (<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655\" target=\"_blank\">link</a>)</li>\n<li>there are some notebook links in it.</li>\n<li>I received great feedback from <a href=\"https://www.kaggle.com/qitvision\" target=\"_blank\">@qitvision</a> on it. </li>\n<li>To summarize that discussion thread, there are two main points: </li>\n<li>Score after 1 random submission versus score after 10,20, N submissions: for example: </li>\n<li>There is \"1/100 chance of random classifier getting 0.65\" for one random submission; </li>\n<li>for 10 random submissions (<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655\" target=\"_blank\">link</a>)<br>\n(10 random subs) percentile-50 : 0.5941798941798941<br>\n(10 random subs) percentile-90 : 0.6423280423280423<br>\n(10 random subs) percentile-95 : 0.6587301587301587<br>\n(10 random subs) percentile-99 : 0.6862433862433863<br>\n(10 random subs) percentile-99.9 : 0.7285714285714286<br>\n(10 random subs) percentile-99.99 : 0.764021164021164</li>\n<li>for 20 random submissions:(<a href=\"https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655\" target=\"_blank\">link</a>)<br>\n(20 random subs) percentile-50 : 0.6132275132275132<br>\n(20 random subs) percentile-90 : 0.6566137566137566<br>\n(20 random subs) percentile-95 : 0.6714285714285714<br>\n(20 random subs) percentile-99 : 0.7005291005291006<br>\n(20 random subs) percentile-99.9 : 0.7391534391534391<br>\n(20 random subs) percentile-99.99 : 0.764021164021164</li>\n</ol>",
      "rawMarkdown": "1. In the past, I had made this discussion thread to discuss random classifier ([link]( https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655))\n2. there are some notebook links in it.\n3. I received great feedback from @qitvision on it. \n4. To summarize that discussion thread, there are two main points: \n5. Score after 1 random submission versus score after 10,20, N submissions: for example: \n5. There is \"1/100 chance of random classifier getting 0.65\" for one random submission; \n6. for 10 random submissions ([link]( https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655))\n(10 random subs) percentile-50 : 0.5941798941798941\n(10 random subs) percentile-90 : 0.6423280423280423\n(10 random subs) percentile-95 : 0.6587301587301587\n(10 random subs) percentile-99 : 0.6862433862433863\n(10 random subs) percentile-99.9 : 0.7285714285714286\n(10 random subs) percentile-99.99 : 0.764021164021164\n7. for 20 random submissions:([link]( https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655))\n(20 random subs) percentile-50 : 0.6132275132275132\n(20 random subs) percentile-90 : 0.6566137566137566\n(20 random subs) percentile-95 : 0.6714285714285714\n(20 random subs) percentile-99 : 0.7005291005291006\n(20 random subs) percentile-99.9 : 0.7391534391534391\n(20 random subs) percentile-99.99 : 0.764021164021164",
      "votes": null
    },
    {
      "id": "1551880",
      "postDate": "10/20/2021 23:39:57",
      "content": "<p>In this competition 1,555 teams made 27,466 submissions. <br>\nI was wondering if anyone had thought about a random classifier with 27,466 submissions ? <br>\nHow should one account for 1,555 teams ? <br>\nI am curious how would one simulate 27,466 random submission from 1,555 teams ? <br>\n(I will check with some statisticians and reply back here if I get some insightful answers)</p>",
      "rawMarkdown": "In this competition 1,555 teams made 27,466 submissions. \nI was wondering if anyone had thought about a random classifier with 27,466 submissions ? \nHow should one account for 1,555 teams ? \nI am curious how would one simulate 27,466 random submission from 1,555 teams ? \n(I will check with some statisticians and reply back here if I get some insightful answers)",
      "votes": null
    },
    {
      "id": "1552167",
      "postDate": "10/21/2021 07:49:57",
      "content": "<p><code>There is \"1/100 chance of random classifier getting 0.65\" for one random submission;</code> for public or private score?</p>",
      "rawMarkdown": "`There is \"1/100 chance of random classifier getting 0.65\" for one random submission;` for public or private score?",
      "votes": null
    },
    {
      "id": "1552174",
      "postDate": "10/21/2021 08:02:12",
      "content": "<p>Good point . I believe that was for the public test set; will try and recompute it for the private test set and report back here. </p>",
      "rawMarkdown": "Good point . I believe that was for the public test set; will try and recompute it for the private test set and report back here.",
      "votes": null
    },
    {
      "id": "1552479",
      "postDate": "10/21/2021 13:16:52",
      "content": "<p>Yes, that was for the public test set. <a href=\"https://www.kaggle.com/anlthms\" target=\"_blank\">@anlthms</a> created this <a href=\"https://www.kaggle.com/anlthms/leaderboard-simulation\" target=\"_blank\">private score simulation notebook</a> that shows that even if predictions of all teams were shuffled, the top-10 scores would remain in the same range.</p>\n<p>I was also curious about simulating private scores so I modified the previous LB simulation notebook to <a href=\"https://www.kaggle.com/qitvision/private-lb-simulation\" target=\"_blank\">simulate private score</a> (sorry, I beat you to it <a href=\"https://www.kaggle.com/mpsampat\" target=\"_blank\">@mpsampat</a>).</p>\n<p>Note that this used the training set and not the actual private set distribution of MGMT values so the results might vary. It shows that one is likely to get over <code>AUC 0.51</code> with two all-random submits.</p>\n<p>If everyone would predict random values, 1 out of 100 teams would get <code>AUC 0.57</code> and 1 out of 1000 <code>AUC 0.60</code>.</p>\n<pre><code>(2 random subs) percentile-50 : 0.5172779028200715\n(2 random subs) percentile-90 : 0.5502482457301735\n(2 random subs) percentile-95 : 0.5608367536078379\n(2 random subs) percentile-99 : 0.5782470541506686\n(2 random subs) percentile-99.9 : 0.6034704753078249\n(2 random subs) percentile-99.99 : 0.625480759300934\n</code></pre>",
      "rawMarkdown": "Yes, that was for the public test set. @anlthms created this [private score simulation notebook](https://www.kaggle.com/anlthms/leaderboard-simulation) that shows that even if predictions of all teams were shuffled, the top-10 scores would remain in the same range.\n\nI was also curious about simulating private scores so I modified the previous LB simulation notebook to [simulate private score](https://www.kaggle.com/qitvision/private-lb-simulation) (sorry, I beat you to it @mpsampat).\n\nNote that this used the training set and not the actual private set distribution of MGMT values so the results might vary. It shows that one is likely to get over `AUC 0.51` with two all-random submits.\n\nIf everyone would predict random values, 1 out of 100 teams would get `AUC 0.57` and 1 out of 1000 `AUC 0.60`.\n\n```\n(2 random subs) percentile-50 : 0.5172779028200715\n(2 random subs) percentile-90 : 0.5502482457301735\n(2 random subs) percentile-95 : 0.5608367536078379\n(2 random subs) percentile-99 : 0.5782470541506686\n(2 random subs) percentile-99.9 : 0.6034704753078249\n(2 random subs) percentile-99.99 : 0.625480759300934\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1551880,
      "author_name": "mpsampat",
      "author_url": "",
      "post_date": "10/20/2021 23:39:57",
      "content": "<p>In this competition 1,555 teams made 27,466 submissions. <br>\nI was wondering if anyone had thought about a random classifier with 27,466 submissions ? <br>\nHow should one account for 1,555 teams ? <br>\nI am curious how would one simulate 27,466 random submission from 1,555 teams ? <br>\n(I will check with some statisticians and reply back here if I get some insightful answers)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1552167,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "10/21/2021 07:49:57",
      "content": "<p><code>There is \"1/100 chance of random classifier getting 0.65\" for one random submission;</code> for public or private score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1552174,
          "author_name": "mpsampat",
          "author_url": "",
          "post_date": "10/21/2021 08:02:12",
          "content": "<p>Good point . I believe that was for the public test set; will try and recompute it for the private test set and report back here. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1552479,
          "author_name": "qitvision",
          "author_url": "",
          "post_date": "10/21/2021 13:16:52",
          "content": "<p>Yes, that was for the public test set. <a href=\"https://www.kaggle.com/anlthms\" target=\"_blank\">@anlthms</a> created this <a href=\"https://www.kaggle.com/anlthms/leaderboard-simulation\" target=\"_blank\">private score simulation notebook</a> that shows that even if predictions of all teams were shuffled, the top-10 scores would remain in the same range.</p>\n<p>I was also curious about simulating private scores so I modified the previous LB simulation notebook to <a href=\"https://www.kaggle.com/qitvision/private-lb-simulation\" target=\"_blank\">simulate private score</a> (sorry, I beat you to it <a href=\"https://www.kaggle.com/mpsampat\" target=\"_blank\">@mpsampat</a>).</p>\n<p>Note that this used the training set and not the actual private set distribution of MGMT values so the results might vary. It shows that one is likely to get over <code>AUC 0.51</code> with two all-random submits.</p>\n<p>If everyone would predict random values, 1 out of 100 teams would get <code>AUC 0.57</code> and 1 out of 1000 <code>AUC 0.60</code>.</p>\n<pre><code>(2 random subs) percentile-50 : 0.5172779028200715\n(2 random subs) percentile-90 : 0.5502482457301735\n(2 random subs) percentile-95 : 0.5608367536078379\n(2 random subs) percentile-99 : 0.5782470541506686\n(2 random subs) percentile-99.9 : 0.6034704753078249\n(2 random subs) percentile-99.99 : 0.625480759300934\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1551875": "1. In the past, I had made this discussion thread to discuss random classifier ([link]( https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655))\n2. there are some notebook links in it.\n3. I received great feedback from @qitvision on it. \n4. To summarize that discussion thread, there are two main points: \n5. Score after 1 random submission versus score after 10,20, N submissions: for example: \n5. There is \"1/100 chance of random classifier getting 0.65\" for one random submission; \n6. for 10 random submissions ([link]( https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655))\n(10 random subs) percentile-50 : 0.5941798941798941\n(10 random subs) percentile-90 : 0.6423280423280423\n(10 random subs) percentile-95 : 0.6587301587301587\n(10 random subs) percentile-99 : 0.6862433862433863\n(10 random subs) percentile-99.9 : 0.7285714285714286\n(10 random subs) percentile-99.99 : 0.764021164021164\n7. for 20 random submissions:([link]( https://www.kaggle.com/c/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/273655))\n(20 random subs) percentile-50 : 0.6132275132275132\n(20 random subs) percentile-90 : 0.6566137566137566\n(20 random subs) percentile-95 : 0.6714285714285714\n(20 random subs) percentile-99 : 0.7005291005291006\n(20 random subs) percentile-99.9 : 0.7391534391534391\n(20 random subs) percentile-99.99 : 0.764021164021164",
    "1551880": "In this competition 1,555 teams made 27,466 submissions. \nI was wondering if anyone had thought about a random classifier with 27,466 submissions ? \nHow should one account for 1,555 teams ? \nI am curious how would one simulate 27,466 random submission from 1,555 teams ? \n(I will check with some statisticians and reply back here if I get some insightful answers)",
    "1552167": "`There is \"1/100 chance of random classifier getting 0.65\" for one random submission;` for public or private score?",
    "1552174": "Good point . I believe that was for the public test set; will try and recompute it for the private test set and report back here.",
    "1552479": "Yes, that was for the public test set. @anlthms created this [private score simulation notebook](https://www.kaggle.com/anlthms/leaderboard-simulation) that shows that even if predictions of all teams were shuffled, the top-10 scores would remain in the same range.\n\nI was also curious about simulating private scores so I modified the previous LB simulation notebook to [simulate private score](https://www.kaggle.com/qitvision/private-lb-simulation) (sorry, I beat you to it @mpsampat).\n\nNote that this used the training set and not the actual private set distribution of MGMT values so the results might vary. It shows that one is likely to get over `AUC 0.51` with two all-random submits.\n\nIf everyone would predict random values, 1 out of 100 teams would get `AUC 0.57` and 1 out of 1000 `AUC 0.60`.\n\n```\n(2 random subs) percentile-50 : 0.5172779028200715\n(2 random subs) percentile-90 : 0.5502482457301735\n(2 random subs) percentile-95 : 0.5608367536078379\n(2 random subs) percentile-99 : 0.5782470541506686\n(2 random subs) percentile-99.9 : 0.6034704753078249\n(2 random subs) percentile-99.99 : 0.625480759300934\n```"
  },
  "source": "meta"
}