{
  "id": 255915,
  "title": "Sampling technique - Public LB dataset",
  "url": "/competitions/siim-covid19-detection/discussion/255915",
  "author_name": "Mikołaj Glinka",
  "post_date": "2021-07-29T19:25:14.144000",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi all,</p>\n<p>the public LB is calculated based on approx. 29% of test data while the private LB will be scored based on the rest of test data. I've got a question regarding the technique used for sampling this public LB subset from all test data. Could you please provide some info on how it was done? </p>\n<p>Cheers,<br>\nMikołaj</p>",
  "messages": [
    {
      "id": 1404364,
      "postDate": "2021-07-29T19:25:14.143Z",
      "content": "<p>Hi all,</p>\n<p>the public LB is calculated based on approx. 29% of test data while the private LB will be scored based on the rest of test data. I've got a question regarding the technique used for sampling this public LB subset from all test data. Could you please provide some info on how it was done? </p>\n<p>Cheers,<br>\nMikołaj</p>",
      "rawMarkdown": "Hi all,\n\nthe public LB is calculated based on approx. 29% of test data while the private LB will be scored based on the rest of test data. I've got a question regarding the technique used for sampling this public LB subset from all test data. Could you please provide some info on how it was done? \n\nCheers,\nMikołaj",
      "votes": 2
    },
    {
      "id": 1442490,
      "postDate": "2021-08-03T23:03:16.687Z",
      "content": "<p>Hmm that is 100,000$ question :))) usually hosts doesn't reveal such info (at least from my kaggle experience so far) <br>\nMaybe if you tag hosts/kaggle team you will get some feedback</p>",
      "rawMarkdown": "Hmm that is 100,000$ question :))) usually hosts doesn't reveal such info (at least from my kaggle experience so far) \nMaybe if you tag hosts/kaggle team you will get some feedback"
    },
    {
      "id": 1404840,
      "postDate": "2021-07-30T09:21:05.213Z",
      "content": "<p>How did you know public data is 29% of all test dataset while the rest is private data?</p>",
      "rawMarkdown": "How did you know public data is 29% of all test dataset while the rest is private data?",
      "replies": [
        {
          "id": 1405035,
          "postDate": "2021-07-30T12:41:47.430Z",
          "content": "<p>If you click on the \"Leaderboard\" tab you will see:<br>\n\"This leaderboard is calculated with approximately 29% of the test data.<br>\nThe final results will be based on the other 71%, so the final standings may be different.\"</p>",
          "rawMarkdown": "If you click on the \"Leaderboard\" tab you will see:\n\"This leaderboard is calculated with approximately 29% of the test data.\nThe final results will be based on the other 71%, so the final standings may be different.\""
        }
      ]
    },
    {
      "id": 1405377,
      "postDate": "2021-07-30T19:12:54.507Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1442490,
      "author_name": "Ioannis M",
      "author_url": "",
      "post_date": "2021-08-03T23:03:16.687000",
      "content": "<p>Hmm that is 100,000$ question :))) usually hosts doesn't reveal such info (at least from my kaggle experience so far) <br>\nMaybe if you tag hosts/kaggle team you will get some feedback</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1404840,
      "author_name": "liuneng12",
      "author_url": "",
      "post_date": "2021-07-30T09:21:05.213000",
      "content": "<p>How did you know public data is 29% of all test dataset while the rest is private data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1405035,
          "author_name": "Mikołaj Glinka",
          "author_url": "",
          "post_date": "2021-07-30T12:41:47.430000",
          "content": "<p>If you click on the \"Leaderboard\" tab you will see:<br>\n\"This leaderboard is calculated with approximately 29% of the test data.<br>\nThe final results will be based on the other 71%, so the final standings may be different.\"</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1405377,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-07-30T19:12:54.507000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1404364": "Hi all,\n\nthe public LB is calculated based on approx. 29% of test data while the private LB will be scored based on the rest of test data. I've got a question regarding the technique used for sampling this public LB subset from all test data. Could you please provide some info on how it was done? \n\nCheers,\nMikołaj",
    "1442490": "Hmm that is 100,000$ question :))) usually hosts doesn't reveal such info (at least from my kaggle experience so far) \nMaybe if you tag hosts/kaggle team you will get some feedback",
    "1404840": "How did you know public data is 29% of all test dataset while the rest is private data?",
    "1405377": ""
  }
}