{
  "id": 171919,
  "title": "How many patients are there in public/private datasets?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/171919",
  "author_name": "",
  "post_date": "2020-08-02T23:50:25.427988900Z",
  "votes": 8,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I could'nt find how public/private sets divided, then how many patients are there in public/private sets?\nAnyway, we know that only 3 data points per patients are used for scoring, so there will be one of the biggest shake in history I think.</p>",
  "messages": [
    {
      "id": "955773",
      "postDate": "08/02/2020 23:50:25",
      "content": "<p>I could'nt find how public/private sets divided, then how many patients are there in public/private sets?\nAnyway, we know that only 3 data points per patients are used for scoring, so there will be one of the biggest shake in history I think.</p>",
      "rawMarkdown": "I could'nt find how public/private sets divided, then how many patients are there in public/private sets?\nAnyway, we know that only 3 data points per patients are used for scoring, so there will be one of the biggest shake in history I think.",
      "votes": null
    },
    {
      "id": "956156",
      "postDate": "08/03/2020 09:50:47",
      "content": "<p>Oh I found it. The host said that there are around 200 in public+private.\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723</a></p>\n\n<p>Then public is around 30 patients * 3 weeks = 90 points and private is around 170 patients * 3 weeks = 510 points?\nLB seems too shaky so we need concrete CV.</p>",
      "rawMarkdown": "Oh I found it. The host said that there are around 200 in public+private.\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723\n\nThen public is around 30 patients * 3 weeks = 90 points and private is around 170 patients * 3 weeks = 510 points?\nLB seems too shaky so we need concrete CV.",
      "votes": null
    },
    {
      "id": "960158",
      "postDate": "08/06/2020 07:19:54",
      "content": "<p>But there is only 5 patients in the test set - isn't it 15 cases for public?</p>",
      "rawMarkdown": "But there is only 5 patients in the test set - isn't it 15 cases for public?",
      "votes": null
    },
    {
      "id": "960424",
      "postDate": "08/06/2020 11:47:19",
      "content": "<p>Public and private sets are entirely hidden from us and when we finish a commit with test set the script will be rerun with public and private sets.\nThe test set is not used for scoring, but for sanity check.</p>",
      "rawMarkdown": "Public and private sets are entirely hidden from us and when we finish a commit with test set the script will be rerun with public and private sets.\nThe test set is not used for scoring, but for sanity check.",
      "votes": null
    },
    {
      "id": "961041",
      "postDate": "08/06/2020 21:40:06",
      "content": "<p>Thanks for the explanation; I didn’t realise this. </p>",
      "rawMarkdown": "Thanks for the explanation; I didn’t realise this.",
      "votes": null
    },
    {
      "id": "961284",
      "postDate": "08/07/2020 04:41:11",
      "content": "<p>I'm expecting very few patients in public test set compared to private test set because every model is oversensitive right now. A single parameter change can even return ±0.1 change in score.</p>",
      "rawMarkdown": "I'm expecting very few patients in public test set compared to private test set because every model is oversensitive right now. A single parameter change can even return ±0.1 change in score.",
      "votes": null
    },
    {
      "id": "961303",
      "postDate": "08/07/2020 04:58:59",
      "content": "<p>Yeah this would make sense. I'm trying to create a reliable cross validation for this competition so that my LB score is close to what I think it will be but that may be difficult. </p>",
      "rawMarkdown": "Yeah this would make sense. I'm trying to create a reliable cross validation for this competition so that my LB score is close to what I think it will be but that may be difficult.",
      "votes": null
    },
    {
      "id": "961624",
      "postDate": "08/07/2020 10:54:48",
      "content": "<p>±0.1 change? That's big. Maybe different seeds will yield similar result.<br>\nGiven that the top score is -6.7599 and the best notebook score is -6.8322, it seems that no one has achieved meaningful result yet.</p>\n<p>I think we can detect the size of public set by LB probing, can't we?  </p>",
      "rawMarkdown": "±0.1 change? That's big. Maybe different seeds will yield similar result.\nGiven that the top score is -6.7599 and the best notebook score is -6.8322, it seems that no one has achieved meaningful result yet.\n\nI think we can detect the size of public set by LB probing, can't we?",
      "votes": null
    },
    {
      "id": "961640",
      "postDate": "08/07/2020 11:05:36",
      "content": "<p>You can basically create a generator object with <code>df_test.groupby('Patient')</code> and predict all data points of  1 patient, then do that for 2 patients and go on like this. This will consume too many submissions and there must be a more efficient way to do this.</p>",
      "rawMarkdown": "You can basically create a generator object with `df_test.groupby('Patient')` and predict all data points of  1 patient, then do that for 2 patients and go on like this. This will consume too many submissions and there must be a more efficient way to do this.",
      "votes": null
    },
    {
      "id": "961650",
      "postDate": "08/07/2020 11:11:55",
      "content": "<ol>\n<li>predict all data points with (10000, 50)</li>\n<li>predict all data points with (10000, 100)</li>\n<li>predict 1 data points with (10000, 50) and others with (10000, 100)</li>\n</ol>\n<p>I thought this is enough.</p>",
      "rawMarkdown": "1. predict all data points with (10000, 50)\n2. predict all data points with (10000, 100)\n3. predict 1 data points with (10000, 50) and others with (10000, 100)\n\nI thought this is enough.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 961284,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "08/07/2020 04:41:11",
      "content": "<p>I'm expecting very few patients in public test set compared to private test set because every model is oversensitive right now. A single parameter change can even return ±0.1 change in score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 961303,
          "author_name": "lgreig",
          "author_url": "",
          "post_date": "08/07/2020 04:58:59",
          "content": "<p>Yeah this would make sense. I'm trying to create a reliable cross validation for this competition so that my LB score is close to what I think it will be but that may be difficult. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961624,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "08/07/2020 10:54:48",
          "content": "<p>±0.1 change? That's big. Maybe different seeds will yield similar result.<br>\nGiven that the top score is -6.7599 and the best notebook score is -6.8322, it seems that no one has achieved meaningful result yet.</p>\n<p>I think we can detect the size of public set by LB probing, can't we?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961640,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "08/07/2020 11:05:36",
          "content": "<p>You can basically create a generator object with <code>df_test.groupby('Patient')</code> and predict all data points of  1 patient, then do that for 2 patients and go on like this. This will consume too many submissions and there must be a more efficient way to do this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961650,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "08/07/2020 11:11:55",
          "content": "<ol>\n<li>predict all data points with (10000, 50)</li>\n<li>predict all data points with (10000, 100)</li>\n<li>predict 1 data points with (10000, 50) and others with (10000, 100)</li>\n</ol>\n<p>I thought this is enough.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 956156,
      "author_name": "resistance0108",
      "author_url": "",
      "post_date": "08/03/2020 09:50:47",
      "content": "<p>Oh I found it. The host said that there are around 200 in public+private.\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723</a></p>\n\n<p>Then public is around 30 patients * 3 weeks = 90 points and private is around 170 patients * 3 weeks = 510 points?\nLB seems too shaky so we need concrete CV.</p>",
      "votes": null,
      "replies": [
        {
          "id": 960158,
          "author_name": "lgreig",
          "author_url": "",
          "post_date": "08/06/2020 07:19:54",
          "content": "<p>But there is only 5 patients in the test set - isn't it 15 cases for public?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960424,
          "author_name": "resistance0108",
          "author_url": "",
          "post_date": "08/06/2020 11:47:19",
          "content": "<p>Public and private sets are entirely hidden from us and when we finish a commit with test set the script will be rerun with public and private sets.\nThe test set is not used for scoring, but for sanity check.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 961041,
          "author_name": "lgreig",
          "author_url": "",
          "post_date": "08/06/2020 21:40:06",
          "content": "<p>Thanks for the explanation; I didn’t realise this. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "955773": "I could'nt find how public/private sets divided, then how many patients are there in public/private sets?\nAnyway, we know that only 3 data points per patients are used for scoring, so there will be one of the biggest shake in history I think.",
    "956156": "Oh I found it. The host said that there are around 200 in public+private.\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723\n\nThen public is around 30 patients * 3 weeks = 90 points and private is around 170 patients * 3 weeks = 510 points?\nLB seems too shaky so we need concrete CV.",
    "960158": "But there is only 5 patients in the test set - isn't it 15 cases for public?",
    "960424": "Public and private sets are entirely hidden from us and when we finish a commit with test set the script will be rerun with public and private sets.\nThe test set is not used for scoring, but for sanity check.",
    "961041": "Thanks for the explanation; I didn’t realise this.",
    "961284": "I'm expecting very few patients in public test set compared to private test set because every model is oversensitive right now. A single parameter change can even return ±0.1 change in score.",
    "961303": "Yeah this would make sense. I'm trying to create a reliable cross validation for this competition so that my LB score is close to what I think it will be but that may be difficult.",
    "961624": "±0.1 change? That's big. Maybe different seeds will yield similar result.\nGiven that the top score is -6.7599 and the best notebook score is -6.8322, it seems that no one has achieved meaningful result yet.\n\nI think we can detect the size of public set by LB probing, can't we?",
    "961640": "You can basically create a generator object with `df_test.groupby('Patient')` and predict all data points of  1 patient, then do that for 2 patients and go on like this. This will consume too many submissions and there must be a more efficient way to do this.",
    "961650": "1. predict all data points with (10000, 50)\n2. predict all data points with (10000, 100)\n3. predict 1 data points with (10000, 50) and others with (10000, 100)\n\nI thought this is enough."
  },
  "source": "meta"
}