{
  "id": 183248,
  "title": "concerns about finetuning the parameters carefully",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/183248",
  "author_name": "",
  "post_date": "2020-09-16T03:12:52.201090900Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I notice that some top solutions in the PL choose to finetune the parameters carefully, such as the dropout_model, FVC_weight, and confidence_weight, but it will not lead to overfitting on the PL? I am afraid about that and if anyone can tell me how to balance finetune and overfitting on the PL?</p>",
  "messages": [
    {
      "id": "1012336",
      "postDate": "09/16/2020 03:12:52",
      "content": "<p>I notice that some top solutions in the PL choose to finetune the parameters carefully, such as the dropout_model, FVC_weight, and confidence_weight, but it will not lead to overfitting on the PL? I am afraid about that and if anyone can tell me how to balance finetune and overfitting on the PL?</p>",
      "rawMarkdown": "I notice that some top solutions in the PL choose to finetune the parameters carefully, such as the dropout_model, FVC_weight, and confidence_weight, but it will not lead to overfitting on the PL? I am afraid about that and if anyone can tell me how to balance finetune and overfitting on the PL?",
      "votes": null
    },
    {
      "id": "1012379",
      "postDate": "09/16/2020 03:58:19",
      "content": "<p>The public leaderboard contains very few data points: 30 patients x 3 last weeks scored, so a total of 90 samples only.<br>\nIf you tune your model based on the public leaderboard, it is probably you will be overfitting to it. In private leaderboard, there would be 170 patients x 3 last weeks scored, so a total of 510 samples.</p>\n<p>It's better to trust your CV score (if how you do CV is reasonable) and tune your models based on CV given we have much more data point in CV set (at least 176 patients x 3 last weeks, so at least 528 samples) than public leaderboard.</p>",
      "rawMarkdown": "The public leaderboard contains very few data points: 30 patients x 3 last weeks scored, so a total of 90 samples only.\nIf you tune your model based on the public leaderboard, it is probably you will be overfitting to it. In private leaderboard, there would be 170 patients x 3 last weeks scored, so a total of 510 samples.\n\nIt's better to trust your CV score (if how you do CV is reasonable) and tune your models based on CV given we have much more data point in CV set (at least 176 patients x 3 last weeks, so at least 528 samples) than public leaderboard.",
      "votes": null
    },
    {
      "id": "1012668",
      "postDate": "09/16/2020 08:06:23",
      "content": "<p>i doubt there are only 17 patients in in public lb. I found it takes time eq to that of 400 patients-450 patients ,15 pct of which is 50 to 60</p>",
      "rawMarkdown": "i doubt there are only 17 patients in in public lb. I found it takes time eq to that of 400 patients-450 patients ,15 pct of which is 50 to 60",
      "votes": null
    },
    {
      "id": "1012752",
      "postDate": "09/16/2020 09:14:35",
      "content": "<p>Please see the reply from the competition host, I'm wrong, let me update the number above.<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723#948386\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723#948386</a></p>",
      "rawMarkdown": "Please see the reply from the competition host, I'm wrong, let me update the number above.\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723#948386",
      "votes": null
    },
    {
      "id": "1012757",
      "postDate": "09/16/2020 09:17:31",
      "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> The competition organizers mentioned that the total number of unique patients in the hidden private test set is around 200 patients. So the time the notebook takes to execute on the hidden test set is for all 200 patients and the public LB score is only for about 30 patients.</p>\n<p>Edit: <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> has already replied with the discussion link while i was typing my reply :)</p>",
      "rawMarkdown": "jaideepvalani The competition organizers mentioned that the total number of unique patients in the hidden private test set is around 200 patients. So the time the notebook takes to execute on the hidden test set is for all 200 patients and the public LB score is only for about 30 patients.\n\nEdit: @khyeh0719 has already replied with the discussion link while i was typing my reply :)",
      "votes": null
    },
    {
      "id": "1012884",
      "postDate": "09/16/2020 11:20:07",
      "content": "<p>thank you so much! I will be careful to finetune the parameters.</p>",
      "rawMarkdown": "thank you so much! I will be careful to finetune the parameters.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1012379,
      "author_name": "khyeh0719",
      "author_url": "",
      "post_date": "09/16/2020 03:58:19",
      "content": "<p>The public leaderboard contains very few data points: 30 patients x 3 last weeks scored, so a total of 90 samples only.<br>\nIf you tune your model based on the public leaderboard, it is probably you will be overfitting to it. In private leaderboard, there would be 170 patients x 3 last weeks scored, so a total of 510 samples.</p>\n<p>It's better to trust your CV score (if how you do CV is reasonable) and tune your models based on CV given we have much more data point in CV set (at least 176 patients x 3 last weeks, so at least 528 samples) than public leaderboard.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1012668,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "09/16/2020 08:06:23",
          "content": "<p>i doubt there are only 17 patients in in public lb. I found it takes time eq to that of 400 patients-450 patients ,15 pct of which is 50 to 60</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012752,
          "author_name": "khyeh0719",
          "author_url": "",
          "post_date": "09/16/2020 09:14:35",
          "content": "<p>Please see the reply from the competition host, I'm wrong, let me update the number above.<br>\n<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723#948386\" target=\"_blank\">https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723#948386</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012757,
          "author_name": "yovinyahathugoda",
          "author_url": "",
          "post_date": "09/16/2020 09:17:31",
          "content": "<p><a href=\"https://www.kaggle.com/jaideepvalani\" target=\"_blank\">@jaideepvalani</a> The competition organizers mentioned that the total number of unique patients in the hidden private test set is around 200 patients. So the time the notebook takes to execute on the hidden test set is for all 200 patients and the public LB score is only for about 30 patients.</p>\n<p>Edit: <a href=\"https://www.kaggle.com/khyeh0719\" target=\"_blank\">@khyeh0719</a> has already replied with the discussion link while i was typing my reply :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1012884,
          "author_name": "yolo1996",
          "author_url": "",
          "post_date": "09/16/2020 11:20:07",
          "content": "<p>thank you so much! I will be careful to finetune the parameters.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1012336": "I notice that some top solutions in the PL choose to finetune the parameters carefully, such as the dropout_model, FVC_weight, and confidence_weight, but it will not lead to overfitting on the PL? I am afraid about that and if anyone can tell me how to balance finetune and overfitting on the PL?",
    "1012379": "The public leaderboard contains very few data points: 30 patients x 3 last weeks scored, so a total of 90 samples only.\nIf you tune your model based on the public leaderboard, it is probably you will be overfitting to it. In private leaderboard, there would be 170 patients x 3 last weeks scored, so a total of 510 samples.\n\nIt's better to trust your CV score (if how you do CV is reasonable) and tune your models based on CV given we have much more data point in CV set (at least 176 patients x 3 last weeks, so at least 528 samples) than public leaderboard.",
    "1012668": "i doubt there are only 17 patients in in public lb. I found it takes time eq to that of 400 patients-450 patients ,15 pct of which is 50 to 60",
    "1012752": "Please see the reply from the competition host, I'm wrong, let me update the number above.\nhttps://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/165723#948386",
    "1012757": "jaideepvalani The competition organizers mentioned that the total number of unique patients in the hidden private test set is around 200 patients. So the time the notebook takes to execute on the hidden test set is for all 200 patients and the public LB score is only for about 30 patients.\n\nEdit: @khyeh0719 has already replied with the discussion link while i was typing my reply :)",
    "1012884": "thank you so much! I will be careful to finetune the parameters."
  },
  "source": "meta"
}