{
  "id": 162936,
  "title": "An estimation of public/private LB",
  "url": "/competitions/alaska2-image-steganalysis/discussion/162936",
  "author_name": "Johnny Lee",
  "post_date": "2020-06-30T13:52:09.192000",
  "votes": 11,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I made an estimation of public/private LB.</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 20% of the test data.\n  The final results will be based on the other 80%, so the final standings may be different.</p>\n</blockquote>\n\n<ol>\n<li>Use a model to predict the train dataset (75000 x 4 = 300000 images).</li>\n<li>Random choice 5000 from  300000</li>\n<li>Random choice 1000 from 5000</li>\n<li>Calculate the wAUC of 1000 and other 4000</li>\n<li>Loop 2~4</li>\n</ol>\n\n<p>The result like below\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F37ec4663ee99805a86fbc99d530d6f1a%2Fdownload.png?generation=1593524284516478&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fa00cd757cf05668b6807ab2997191b32%2Fdownload%20(1\" alt=\"\">.png?generation=1593524356349332&amp;alt=media)</p>\n\n<p>If this is correct, does it mean the public LB is too small. How do you think about it. </p>",
  "messages": [
    {
      "id": 908303,
      "postDate": "2020-06-30T13:52:09.193Z",
      "content": "<p>I made an estimation of public/private LB.</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 20% of the test data.\n  The final results will be based on the other 80%, so the final standings may be different.</p>\n</blockquote>\n\n<ol>\n<li>Use a model to predict the train dataset (75000 x 4 = 300000 images).</li>\n<li>Random choice 5000 from  300000</li>\n<li>Random choice 1000 from 5000</li>\n<li>Calculate the wAUC of 1000 and other 4000</li>\n<li>Loop 2~4</li>\n</ol>\n\n<p>The result like below\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F37ec4663ee99805a86fbc99d530d6f1a%2Fdownload.png?generation=1593524284516478&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fa00cd757cf05668b6807ab2997191b32%2Fdownload%20(1\" alt=\"\">.png?generation=1593524356349332&amp;alt=media)</p>\n\n<p>If this is correct, does it mean the public LB is too small. How do you think about it. </p>",
      "rawMarkdown": "I made an estimation of public/private LB.\n\n&gt; This leaderboard is calculated with approximately 20% of the test data.\nThe final results will be based on the other 80%, so the final standings may be different.\n\n1. Use a model to predict the train dataset (75000 x 4 = 300000 images).\n2. Random choice 5000 from  300000\n3. Random choice 1000 from 5000\n4. Calculate the wAUC of 1000 and other 4000\n5. Loop 2~4\n\nThe result like below\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F37ec4663ee99805a86fbc99d530d6f1a%2Fdownload.png?generation=1593524284516478&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fa00cd757cf05668b6807ab2997191b32%2Fdownload%20(1).png?generation=1593524356349332&amp;alt=media)\n\nIf this is correct, does it mean the public LB is too small. How do you think about it. ",
      "votes": 11
    },
    {
      "id": 909366,
      "postDate": "2020-06-30T14:42:55.137Z",
      "content": "<p>Great approach! It should give a rough idea about the shake-up that might come.</p>\n\n<p>I think you should choose the random 5000 from your validation or OOF set. That way there's a guarantee that the model will process only unseen data.\nThen you can again randomly choose 1k from the above 5k to simulate the public/private LB</p>",
      "rawMarkdown": "Great approach! It should give a rough idea about the shake-up that might come.\n\nI think you should choose the random 5000 from your validation or OOF set. That way there's a guarantee that the model will process only unseen data.\nThen you can again randomly choose 1k from the above 5k to simulate the public/private LB",
      "votes": 3
    },
    {
      "id": 921343,
      "postDate": "2020-07-09T08:17:02.413Z",
      "content": "<p>Even with 60,000 - the size of most val sets during training, the range is significant at .004 - be careful even with local cv choice for final submission.</p>\n\n<p>Random 60000, from an example 300000 CV set, 10000 times below.\nCould the metric be tuned for less random impact?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582365%2F2d14539e28329ae503c41a5f4f5fec76%2FScreenshot%202020-07-09%2009.16.07.png?generation=1594282614100250&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Even with 60,000 - the size of most val sets during training, the range is significant at .004 - be careful even with local cv choice for final submission.\n\nRandom 60000, from an example 300000 CV set, 10000 times below.\nCould the metric be tuned for less random impact?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582365%2F2d14539e28329ae503c41a5f4f5fec76%2FScreenshot%202020-07-09%2009.16.07.png?generation=1594282614100250&amp;alt=media)"
    },
    {
      "id": 916137,
      "postDate": "2020-07-05T12:02:39.407Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 909366,
      "author_name": "Mighty Rains",
      "author_url": "",
      "post_date": "2020-06-30T14:42:55.137000",
      "content": "<p>Great approach! It should give a rough idea about the shake-up that might come.</p>\n\n<p>I think you should choose the random 5000 from your validation or OOF set. That way there's a guarantee that the model will process only unseen data.\nThen you can again randomly choose 1k from the above 5k to simulate the public/private LB</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 921343,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2020-07-09T08:17:02.413000",
      "content": "<p>Even with 60,000 - the size of most val sets during training, the range is significant at .004 - be careful even with local cv choice for final submission.</p>\n\n<p>Random 60000, from an example 300000 CV set, 10000 times below.\nCould the metric be tuned for less random impact?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582365%2F2d14539e28329ae503c41a5f4f5fec76%2FScreenshot%202020-07-09%2009.16.07.png?generation=1594282614100250&amp;alt=media\" alt=\"\"></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 916137,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-07-05T12:02:39.407000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "908303": "I made an estimation of public/private LB.\n\n&gt; This leaderboard is calculated with approximately 20% of the test data.\nThe final results will be based on the other 80%, so the final standings may be different.\n\n1. Use a model to predict the train dataset (75000 x 4 = 300000 images).\n2. Random choice 5000 from  300000\n3. Random choice 1000 from 5000\n4. Calculate the wAUC of 1000 and other 4000\n5. Loop 2~4\n\nThe result like below\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2F37ec4663ee99805a86fbc99d530d6f1a%2Fdownload.png?generation=1593524284516478&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2006644%2Fa00cd757cf05668b6807ab2997191b32%2Fdownload%20(1).png?generation=1593524356349332&amp;alt=media)\n\nIf this is correct, does it mean the public LB is too small. How do you think about it. ",
    "909366": "Great approach! It should give a rough idea about the shake-up that might come.\n\nI think you should choose the random 5000 from your validation or OOF set. That way there's a guarantee that the model will process only unseen data.\nThen you can again randomly choose 1k from the above 5k to simulate the public/private LB",
    "921343": "Even with 60,000 - the size of most val sets during training, the range is significant at .004 - be careful even with local cv choice for final submission.\n\nRandom 60000, from an example 300000 CV set, 10000 times below.\nCould the metric be tuned for less random impact?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1582365%2F2d14539e28329ae503c41a5f4f5fec76%2FScreenshot%202020-07-09%2009.16.07.png?generation=1594282614100250&amp;alt=media)",
    "916137": ""
  }
}