{
  "id": 174545,
  "title": "Novice question about test set",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/174545",
  "author_name": "",
  "post_date": "2020-08-14T03:15:15.389970400Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Are we being evaluated on just 5 patients' data? That's weird. If one or two of those patients are statistical outliers, that will throw potentially accurate models off the leaderboard. Or is it that our submitted notebooks will be run on a larger private evaluation dataset? </p>",
  "messages": [
    {
      "id": "969886",
      "postDate": "08/14/2020 03:15:15",
      "content": "<p>Are we being evaluated on just 5 patients' data? That's weird. If one or two of those patients are statistical outliers, that will throw potentially accurate models off the leaderboard. Or is it that our submitted notebooks will be run on a larger private evaluation dataset? </p>",
      "rawMarkdown": "Are we being evaluated on just 5 patients' data? That's weird. If one or two of those patients are statistical outliers, that will throw potentially accurate models off the leaderboard. Or is it that our submitted notebooks will be run on a larger private evaluation dataset?",
      "votes": null
    },
    {
      "id": "969957",
      "postDate": "08/14/2020 04:56:56",
      "content": "<p>I think you might have missed this thread: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930\" target=\"_blank\">Thread Link</a></p>\n<p>The hosts had pointed out this earlier:</p>\n<blockquote>\n  <p>This is caused by all the Ignored rows being counted as \"not public\" (we show it this way in normal competitions to not leak the actual number of ignored rows). I'll manually override this to show the actual value, which is approximately 15%.</p>\n</blockquote>",
      "rawMarkdown": "I think you might have missed this thread: [Thread Link](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930)\n\nThe hosts had pointed out this earlier:\n> This is caused by all the Ignored rows being counted as \"not public\" (we show it this way in normal competitions to not leak the actual number of ignored rows). I'll manually override this to show the actual value, which is approximately 15%.",
      "votes": null
    },
    {
      "id": "969982",
      "postDate": "08/14/2020 05:24:34",
      "content": "<p>They hide the actually test set from us to prevent manually labeling or leaking of the data. Therefor, to submit you must create an inference kernel where you perform the prediction on the smaller public test set. When you submit, your code will be rerun on the hidden test set and that is where you get the LB score.</p>",
      "rawMarkdown": "They hide the actually test set from us to prevent manually labeling or leaking of the data. Therefor, to submit you must create an inference kernel where you perform the prediction on the smaller public test set. When you submit, your code will be rerun on the hidden test set and that is where you get the LB score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 969957,
      "author_name": "aadhavvignesh",
      "author_url": "",
      "post_date": "08/14/2020 04:56:56",
      "content": "<p>I think you might have missed this thread: <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930\" target=\"_blank\">Thread Link</a></p>\n<p>The hosts had pointed out this earlier:</p>\n<blockquote>\n  <p>This is caused by all the Ignored rows being counted as \"not public\" (we show it this way in normal competitions to not leak the actual number of ignored rows). I'll manually override this to show the actual value, which is approximately 15%.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 969982,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "08/14/2020 05:24:34",
      "content": "<p>They hide the actually test set from us to prevent manually labeling or leaking of the data. Therefor, to submit you must create an inference kernel where you perform the prediction on the smaller public test set. When you submit, your code will be rerun on the hidden test set and that is where you get the LB score.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "969886": "Are we being evaluated on just 5 patients' data? That's weird. If one or two of those patients are statistical outliers, that will throw potentially accurate models off the leaderboard. Or is it that our submitted notebooks will be run on a larger private evaluation dataset?",
    "969957": "I think you might have missed this thread: [Thread Link](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/164930)\n\nThe hosts had pointed out this earlier:\n> This is caused by all the Ignored rows being counted as \"not public\" (we show it this way in normal competitions to not leak the actual number of ignored rows). I'll manually override this to show the actual value, which is approximately 15%.",
    "969982": "They hide the actually test set from us to prevent manually labeling or leaking of the data. Therefor, to submit you must create an inference kernel where you perform the prediction on the smaller public test set. When you submit, your code will be rerun on the hidden test set and that is where you get the LB score."
  },
  "source": "meta"
}