{
  "id": 170328,
  "title": "Some confusion about percent",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/170328",
  "author_name": "",
  "post_date": "2020-07-27T08:57:37.075722200Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p>In this notebook <a href=\"https://www.kaggle.com/piantic/osic-pulmonary-fibrosis-progression-basic-eda/notebook\">OSIC Pulmonary Fibrosis Progression: Basic EDA!</a></p>\n\n<p>FVC seems to related Percent linearly. Makes sense as both terms are proportional.</p>\n\n<p>However,some baseline use a fixed percent as feature to make prediction</p>\n\n<p>Will this influence the model ?</p>",
  "messages": [
    {
      "id": "947416",
      "postDate": "07/27/2020 08:57:37",
      "content": "<p>In this notebook <a href=\"https://www.kaggle.com/piantic/osic-pulmonary-fibrosis-progression-basic-eda/notebook\">OSIC Pulmonary Fibrosis Progression: Basic EDA!</a></p>\n\n<p>FVC seems to related Percent linearly. Makes sense as both terms are proportional.</p>\n\n<p>However,some baseline use a fixed percent as feature to make prediction</p>\n\n<p>Will this influence the model ?</p>",
      "rawMarkdown": "In this notebook [OSIC Pulmonary Fibrosis Progression: Basic EDA!](https://www.kaggle.com/piantic/osic-pulmonary-fibrosis-progression-basic-eda/notebook)\n\nFVC seems to related Percent linearly. Makes sense as both terms are proportional.\n\nHowever,some baseline use a fixed percent as feature to make prediction\n\nWill this influence the model ?",
      "votes": null
    },
    {
      "id": "947993",
      "postDate": "07/27/2020 15:40:55",
      "content": "<p>I tried the \"percent\" feature with each patient's initial percent value (for the training samples), but the result was worse than simply using the actual percent values, which is observed for both the local CV and the public LB. </p>",
      "rawMarkdown": "I tried the \"percent\" feature with each patient's initial percent value (for the training samples), but the result was worse than simply using the actual percent values, which is observed for both the local CV and the public LB.",
      "votes": null
    },
    {
      "id": "951202",
      "postDate": "07/30/2020 01:37:22",
      "content": "<p>Thanks for your answer!\nIn my comprehend, the 'percent' feature can only set as a fixed value(for the test samples)to predict every week of patient's FVC in submission.csv.Is it that right?</p>",
      "rawMarkdown": "Thanks for your answer!\nIn my comprehend, the 'percent' feature can only set as a fixed value(for the test samples)to predict every week of patient's FVC in submission.csv.Is it that right?",
      "votes": null
    },
    {
      "id": "951586",
      "postDate": "07/30/2020 08:34:06",
      "content": "<p>Yes, the percent for each week is only used for training data, and the test samples use only the first week's value, but the CV and LB are better.</p>",
      "rawMarkdown": "Yes, the percent for each week is only used for training data, and the test samples use only the first week's value, but the CV and LB are better.",
      "votes": null
    },
    {
      "id": "953217",
      "postDate": "07/31/2020 15:56:51",
      "content": "<p>I have a question: is using percent for each week in training  worse than use only baseline percent? Using only baseline percent we have better representation of test set situation, right? I don't understand why performance are better in LB.. In Local CV it was predictable , but I don't know if use all percent is the right way for private and public leaderboard</p>",
      "rawMarkdown": "I have a question: is using percent for each week in training  worse than use only baseline percent? Using only baseline percent we have better representation of test set situation, right? I don't understand why performance are better in LB.. In Local CV it was predictable , but I don't know if use all percent is the right way for private and public leaderboard",
      "votes": null
    },
    {
      "id": "953251",
      "postDate": "07/31/2020 16:15:41",
      "content": "<p>I think the reason is models are very weak at this stage with only tabular features. Predictions of the test set probably look like this when I don't use <code>Percent</code> as a feature (I use <code>Percent_Baseline</code>). </p>\n\n<p><img src=\"https://i.ibb.co/WPymFQd/ddd.jpg\" alt=\"\"></p>\n\n<p>It looks like it is very hard to predict this non-stationary data but <code>Percent</code> has the strongest relationship with <code>FVC</code>, so there is no doubt that it will increase cv score. Even if the <code>Percent</code> values are all same for test set patients, models can still make better predictions compared to models that are only using <code>Percent_Baseline</code>.</p>",
      "rawMarkdown": "I think the reason is models are very weak at this stage with only tabular features. Predictions of the test set probably look like this when I don't use `Percent` as a feature (I use `Percent_Baseline`). \n\n![](https://i.ibb.co/WPymFQd/ddd.jpg)\n\nIt looks like it is very hard to predict this non-stationary data but `Percent` has the strongest relationship with `FVC`, so there is no doubt that it will increase cv score. Even if the `Percent` values are all same for test set patients, models can still make better predictions compared to models that are only using `Percent_Baseline`.",
      "votes": null
    },
    {
      "id": "970459",
      "postDate": "08/14/2020 13:26:53",
      "content": "<p>As I understand, it doesn't make sense to use percent in the model. FVC is a multiplication of percent with an Expected value of FVC given the patient demographics.  Hence, percent is like FVC</p>",
      "rawMarkdown": "As I understand, it doesn't make sense to use percent in the model. FVC is a multiplication of percent with an Expected value of FVC given the patient demographics.  Hence, percent is like FVC",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 970459,
      "author_name": "regivm",
      "author_url": "",
      "post_date": "08/14/2020 13:26:53",
      "content": "<p>As I understand, it doesn't make sense to use percent in the model. FVC is a multiplication of percent with an Expected value of FVC given the patient demographics.  Hence, percent is like FVC</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 947993,
      "author_name": "yalickj",
      "author_url": "",
      "post_date": "07/27/2020 15:40:55",
      "content": "<p>I tried the \"percent\" feature with each patient's initial percent value (for the training samples), but the result was worse than simply using the actual percent values, which is observed for both the local CV and the public LB. </p>",
      "votes": null,
      "replies": [
        {
          "id": 951202,
          "author_name": "z526830652",
          "author_url": "",
          "post_date": "07/30/2020 01:37:22",
          "content": "<p>Thanks for your answer!\nIn my comprehend, the 'percent' feature can only set as a fixed value(for the test samples)to predict every week of patient's FVC in submission.csv.Is it that right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 951586,
          "author_name": "yalickj",
          "author_url": "",
          "post_date": "07/30/2020 08:34:06",
          "content": "<p>Yes, the percent for each week is only used for training data, and the test samples use only the first week's value, but the CV and LB are better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 953217,
      "author_name": "agostinodorano",
      "author_url": "",
      "post_date": "07/31/2020 15:56:51",
      "content": "<p>I have a question: is using percent for each week in training  worse than use only baseline percent? Using only baseline percent we have better representation of test set situation, right? I don't understand why performance are better in LB.. In Local CV it was predictable , but I don't know if use all percent is the right way for private and public leaderboard</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 953251,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "07/31/2020 16:15:41",
      "content": "<p>I think the reason is models are very weak at this stage with only tabular features. Predictions of the test set probably look like this when I don't use <code>Percent</code> as a feature (I use <code>Percent_Baseline</code>). </p>\n\n<p><img src=\"https://i.ibb.co/WPymFQd/ddd.jpg\" alt=\"\"></p>\n\n<p>It looks like it is very hard to predict this non-stationary data but <code>Percent</code> has the strongest relationship with <code>FVC</code>, so there is no doubt that it will increase cv score. Even if the <code>Percent</code> values are all same for test set patients, models can still make better predictions compared to models that are only using <code>Percent_Baseline</code>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "947416": "In this notebook [OSIC Pulmonary Fibrosis Progression: Basic EDA!](https://www.kaggle.com/piantic/osic-pulmonary-fibrosis-progression-basic-eda/notebook)\n\nFVC seems to related Percent linearly. Makes sense as both terms are proportional.\n\nHowever,some baseline use a fixed percent as feature to make prediction\n\nWill this influence the model ?",
    "947993": "I tried the \"percent\" feature with each patient's initial percent value (for the training samples), but the result was worse than simply using the actual percent values, which is observed for both the local CV and the public LB.",
    "951202": "Thanks for your answer!\nIn my comprehend, the 'percent' feature can only set as a fixed value(for the test samples)to predict every week of patient's FVC in submission.csv.Is it that right?",
    "951586": "Yes, the percent for each week is only used for training data, and the test samples use only the first week's value, but the CV and LB are better.",
    "953217": "I have a question: is using percent for each week in training  worse than use only baseline percent? Using only baseline percent we have better representation of test set situation, right? I don't understand why performance are better in LB.. In Local CV it was predictable , but I don't know if use all percent is the right way for private and public leaderboard",
    "953251": "I think the reason is models are very weak at this stage with only tabular features. Predictions of the test set probably look like this when I don't use `Percent` as a feature (I use `Percent_Baseline`). \n\n![](https://i.ibb.co/WPymFQd/ddd.jpg)\n\nIt looks like it is very hard to predict this non-stationary data but `Percent` has the strongest relationship with `FVC`, so there is no doubt that it will increase cv score. Even if the `Percent` values are all same for test set patients, models can still make better predictions compared to models that are only using `Percent_Baseline`.",
    "970459": "As I understand, it doesn't make sense to use percent in the model. FVC is a multiplication of percent with an Expected value of FVC given the patient demographics.  Hence, percent is like FVC"
  },
  "source": "meta"
}