{
  "id": 187625,
  "title": "What's the recipe for sucess?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/187625",
  "author_name": "from coffee import *",
  "post_date": "2020-09-29T16:47:56.511000",
  "votes": 23,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Dear fellow Kagglers,</p>\n<p>let me start with a quick summary of <strong>what we know so far</strong>:</p>\n<ul>\n<li><p>We know so far that<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186683\" target=\"_blank\"> CV is definitely the way to go</a> instead of relying purely on the public LB-score, which is only based on 15% of the data.</p></li>\n<li><p>GroupKFold strongly beats the usage of basic KFold, as it reduces leakage.</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077\" target=\"_blank\">IMG data only grants very limited additional value (except for tissue/lung-area feature)</a> and it's hard to engineer useful features out of the few pictures we have.</p></li>\n</ul>\n<h3>What is not completely known by now is the following:</h3>\n<ul>\n<li>Is using PERCENT as a feature reasonable? <br>\nDefintion:</li>\n</ul>\n<blockquote>\n  <p>Percent- a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics</p>\n</blockquote>\n<p>Let me quickly kickstart the debate: <br>\nAs I explained <a href=\"https://www.kaggle.com/ChristianDenich/quantile-reg-lr-schedulers-checkpoints\" target=\"_blank\">here </a> NOT using percent at all seems kind of wrong to me. Using it for all training will introduce leakage, as the combination of AGE and PERCENT is enough to numerically calculate the correct FVC. So my idea is to only use it for the first row for each patient, as it's also available in the test data and gives us insigts.</p>\n<ul>\n<li>What's your preferred way of calculating the confidence/sigma? So far very different approaches have been introduced.</li>\n</ul>",
  "messages": [
    {
      "id": 1031738,
      "postDate": "2020-09-29T16:47:56.510Z",
      "content": "<p>Dear fellow Kagglers,</p>\n<p>let me start with a quick summary of <strong>what we know so far</strong>:</p>\n<ul>\n<li><p>We know so far that<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186683\" target=\"_blank\"> CV is definitely the way to go</a> instead of relying purely on the public LB-score, which is only based on 15% of the data.</p></li>\n<li><p>GroupKFold strongly beats the usage of basic KFold, as it reduces leakage.</p></li>\n<li><p><a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077\" target=\"_blank\">IMG data only grants very limited additional value (except for tissue/lung-area feature)</a> and it's hard to engineer useful features out of the few pictures we have.</p></li>\n</ul>\n<h3>What is not completely known by now is the following:</h3>\n<ul>\n<li>Is using PERCENT as a feature reasonable? <br>\nDefintion:</li>\n</ul>\n<blockquote>\n  <p>Percent- a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics</p>\n</blockquote>\n<p>Let me quickly kickstart the debate: <br>\nAs I explained <a href=\"https://www.kaggle.com/ChristianDenich/quantile-reg-lr-schedulers-checkpoints\" target=\"_blank\">here </a> NOT using percent at all seems kind of wrong to me. Using it for all training will introduce leakage, as the combination of AGE and PERCENT is enough to numerically calculate the correct FVC. So my idea is to only use it for the first row for each patient, as it's also available in the test data and gives us insigts.</p>\n<ul>\n<li>What's your preferred way of calculating the confidence/sigma? So far very different approaches have been introduced.</li>\n</ul>",
      "rawMarkdown": "Dear fellow Kagglers,\n\n let me start with a quick summary of **what we know so far**:\n- We know so far that[ CV is definitely the way to go](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186683) instead of relying purely on the public LB-score, which is only based on 15% of the data.\n\n- GroupKFold strongly beats the usage of basic KFold, as it reduces leakage.\n\n- [IMG data only grants very limited additional value (except for tissue/lung-area feature)](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077) and it's hard to engineer useful features out of the few pictures we have.\n\n\n\n### What is not completely known by now is the following:\n- Is using PERCENT as a feature reasonable? \nDefintion:\n> Percent- a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics\n\nLet me quickly kickstart the debate: \nAs I explained [here ](https://www.kaggle.com/ChristianDenich/quantile-reg-lr-schedulers-checkpoints) NOT using percent at all seems kind of wrong to me. Using it for all training will introduce leakage, as the combination of AGE and PERCENT is enough to numerically calculate the correct FVC. So my idea is to only use it for the first row for each patient, as it's also available in the test data and gives us insigts.\n\n- What's your preferred way of calculating the confidence/sigma? So far very different approaches have been introduced.\n\n\n",
      "votes": 22
    },
    {
      "id": 1033335,
      "postDate": "2020-09-30T21:26:04.267Z",
      "content": "<p><a href=\"https://www.kaggle.com/ChristianDenich\" target=\"_blank\">@ChristianDenich</a> About PERCENT: As per my analysis, two people of the same gender and age don't have standard FVC values as same (if you divide the patient-specific FVC by % and find out 'standard' FVC for that person/gender/age combination). Standard FVC seems to be the function of body weight/height/build/BMI/race/ethnicity etc. also. In that sense % in incomplete information as provided here.</p>",
      "rawMarkdown": "@ChristianDenich About PERCENT: As per my analysis, two people of the same gender and age don't have standard FVC values as same (if you divide the patient-specific FVC by % and find out 'standard' FVC for that person/gender/age combination). Standard FVC seems to be the function of body weight/height/build/BMI/race/ethnicity etc. also. In that sense % in incomplete information as provided here.",
      "votes": 1
    },
    {
      "id": 1033862,
      "postDate": "2020-10-01T10:18:38.033Z",
      "content": "<p>If the Percent feature is wrong to use, a shakeup is inevitable. There is a big gap in CV/LB when it's not included.</p>",
      "rawMarkdown": "If the Percent feature is wrong to use, a shakeup is inevitable. There is a big gap in CV/LB when it's not included.",
      "votes": 2,
      "replies": [
        {
          "id": 1034675,
          "postDate": "2020-10-02T04:31:31.657Z",
          "content": "<p>Problem is how do you compute the Percent for each week? All I can say is for each Patient the typical FVC = FVC / Percent * 100 is constant. Maybe the typical FVC can provide some useful hidden information since Patients with same Age, Sex, Smoking Status do not necessary have the same typical FVC. </p>",
          "rawMarkdown": "Problem is how do you compute the Percent for each week? All I can say is for each Patient the typical FVC = FVC / Percent * 100 is constant. Maybe the typical FVC can provide some useful hidden information since Patients with same Age, Sex, Smoking Status do not necessary have the same typical FVC. ",
          "votes": 1
        },
        {
          "id": 1034853,
          "postDate": "2020-10-02T08:39:00.937Z",
          "content": "<p>You are right that it could contain hidden info about a patient.  I believe you can use either the typical FVC or the base Percent of a patient to capture that</p>",
          "rawMarkdown": "You are right that it could contain hidden info about a patient.  I believe you can use either the typical FVC or the base Percent of a patient to capture that"
        }
      ]
    },
    {
      "id": 1032637,
      "postDate": "2020-09-30T10:28:21.483Z",
      "content": "<p><a href=\"https://www.kaggle.com/gannadolinska\" target=\"_blank\">@gannadolinska</a>, <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a>, <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> any insights you want to share?</p>",
      "rawMarkdown": "@gannadolinska, @ulrich07, @carlossouza any insights you want to share?",
      "replies": [
        {
          "id": 1032967,
          "postDate": "2020-09-30T15:03:25.890Z",
          "content": "<p>So far, our best submission uses GroupKFold, uses the images to predict the betas (slopes for the FVC declines), uses all features including Percent (although I personally think it's a form of data leak), and blends models that predict the confidence with (very different) methods…<br>\nWe tried hundreds of different approaches: some that we thought it would work failed miserably, and some others that we didn't thought much surprised us… go figure :)</p>\n<p>Reflecting on your (very good) question, I learned 2 \"recipes\" in this competition:</p>\n<ol>\n<li><strong>The power of a team</strong>. My team members are amazing, very smart… I love to discuss new approaches with them, I always learn. This is my first competition that I work in a team. It's an amazing experience! Really, it's another ball game: very fun! Getting to know &amp; work with them was the competition's key prize imho!</li>\n<li><strong>Keep trying new approaches, and don't ever stop submitting</strong>. We pursued hundreds of approaches, and most of them ended up in \"nothing\". But we learned a lot from every failure. This reminds me of sales: 99% of the time a salesman hears \"no\", but when he/she gets a \"yes\", it pays out beautifully :)</li>\n</ol>",
          "rawMarkdown": "So far, our best submission uses GroupKFold, uses the images to predict the betas (slopes for the FVC declines), uses all features including Percent (although I personally think it's a form of data leak), and blends models that predict the confidence with (very different) methods...\nWe tried hundreds of different approaches: some that we thought it would work failed miserably, and some others that we didn't thought much surprised us... go figure :)\n\nReflecting on your (very good) question, I learned 2 \"recipes\" in this competition:\n1. **The power of a team**. My team members are amazing, very smart... I love to discuss new approaches with them, I always learn. This is my first competition that I work in a team. It's an amazing experience! Really, it's another ball game: very fun! Getting to know & work with them was the competition's key prize imho!\n2. **Keep trying new approaches, and don't ever stop submitting**. We pursued hundreds of approaches, and most of them ended up in \"nothing\". But we learned a lot from every failure. This reminds me of sales: 99% of the time a salesman hears \"no\", but when he/she gets a \"yes\", it pays out beautifully :)",
          "votes": 8
        },
        {
          "id": 1033106,
          "postDate": "2020-09-30T17:02:06.370Z",
          "content": "<p>Wow, it seems that our team is perfectly aligned with <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>'s team: our <strong>\"recipes\"</strong> are basically the same with some tweaks, we too experimented with several approaches, most of them resulting in no to very minor changes in the model's performance.</p>\n<p>We have learnt a lot from our submissions, and we haven't stopped yet. With each submission, we're getting new insights, some contradicting our hypothesis for selecting which one would be the ideal submission.</p>",
          "rawMarkdown": "Wow, it seems that our team is perfectly aligned with @carlossouza's team: our **\"recipes\"** are basically the same with some tweaks, we too experimented with several approaches, most of them resulting in no to very minor changes in the model's performance.\n\nWe have learnt a lot from our submissions, and we haven't stopped yet. With each submission, we're getting new insights, some contradicting our hypothesis for selecting which one would be the ideal submission.",
          "votes": 1
        },
        {
          "id": 1033158,
          "postDate": "2020-09-30T17:54:52.627Z",
          "content": "<p><a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a> We are using the same recipes as well and some additional things as well ,<br>\n<a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> I too have experimented on a lot of things that haven't work well for me though on lb , but I am very much trusting the ideas to work on private . We are yet to ensemble models as we were waiting for the last week 😉😉😉 . The only thing was we didn't have the full support of the team as they had prior commitments but nevertheless I have learned a lot here because all the pipelines I have written myself , I hope it ends well for us <br>\nThanks Chris for this post</p>",
          "rawMarkdown": "@aadhavvignesh We are using the same recipes as well and some additional things as well ,\n@carlossouza I too have experimented on a lot of things that haven't work well for me though on lb , but I am very much trusting the ideas to work on private . We are yet to ensemble models as we were waiting for the last week 😉😉😉 . The only thing was we didn't have the full support of the team as they had prior commitments but nevertheless I have learned a lot here because all the pipelines I have written myself , I hope it ends well for us \nThanks Chris for this post"
        }
      ]
    },
    {
      "id": 1032585,
      "postDate": "2020-09-30T09:57:56.777Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1033335,
      "author_name": "PrasunMishra",
      "author_url": "",
      "post_date": "2020-09-30T21:26:04.267000",
      "content": "<p><a href=\"https://www.kaggle.com/ChristianDenich\" target=\"_blank\">@ChristianDenich</a> About PERCENT: As per my analysis, two people of the same gender and age don't have standard FVC values as same (if you divide the patient-specific FVC by % and find out 'standard' FVC for that person/gender/age combination). Standard FVC seems to be the function of body weight/height/build/BMI/race/ethnicity etc. also. In that sense % in incomplete information as provided here.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1033862,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2020-10-01T10:18:38.033000",
      "content": "<p>If the Percent feature is wrong to use, a shakeup is inevitable. There is a big gap in CV/LB when it's not included.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1034675,
          "author_name": "Quan",
          "author_url": "",
          "post_date": "2020-10-02T04:31:31.657000",
          "content": "<p>Problem is how do you compute the Percent for each week? All I can say is for each Patient the typical FVC = FVC / Percent * 100 is constant. Maybe the typical FVC can provide some useful hidden information since Patients with same Age, Sex, Smoking Status do not necessary have the same typical FVC. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1034853,
          "author_name": "Rafi Hai",
          "author_url": "",
          "post_date": "2020-10-02T08:39:00.937000",
          "content": "<p>You are right that it could contain hidden info about a patient.  I believe you can use either the typical FVC or the base Percent of a patient to capture that</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1032637,
      "author_name": "from coffee import *",
      "author_url": "",
      "post_date": "2020-09-30T10:28:21.483000",
      "content": "<p><a href=\"https://www.kaggle.com/gannadolinska\" target=\"_blank\">@gannadolinska</a>, <a href=\"https://www.kaggle.com/ulrich07\" target=\"_blank\">@ulrich07</a>, <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> any insights you want to share?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1032967,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-09-30T15:03:25.890000",
          "content": "<p>So far, our best submission uses GroupKFold, uses the images to predict the betas (slopes for the FVC declines), uses all features including Percent (although I personally think it's a form of data leak), and blends models that predict the confidence with (very different) methods…<br>\nWe tried hundreds of different approaches: some that we thought it would work failed miserably, and some others that we didn't thought much surprised us… go figure :)</p>\n<p>Reflecting on your (very good) question, I learned 2 \"recipes\" in this competition:</p>\n<ol>\n<li><strong>The power of a team</strong>. My team members are amazing, very smart… I love to discuss new approaches with them, I always learn. This is my first competition that I work in a team. It's an amazing experience! Really, it's another ball game: very fun! Getting to know &amp; work with them was the competition's key prize imho!</li>\n<li><strong>Keep trying new approaches, and don't ever stop submitting</strong>. We pursued hundreds of approaches, and most of them ended up in \"nothing\". But we learned a lot from every failure. This reminds me of sales: 99% of the time a salesman hears \"no\", but when he/she gets a \"yes\", it pays out beautifully :)</li>\n</ol>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 1033106,
          "author_name": "Aadhav Vignesh",
          "author_url": "",
          "post_date": "2020-09-30T17:02:06.370000",
          "content": "<p>Wow, it seems that our team is perfectly aligned with <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a>'s team: our <strong>\"recipes\"</strong> are basically the same with some tweaks, we too experimented with several approaches, most of them resulting in no to very minor changes in the model's performance.</p>\n<p>We have learnt a lot from our submissions, and we haven't stopped yet. With each submission, we're getting new insights, some contradicting our hypothesis for selecting which one would be the ideal submission.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1033158,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-09-30T17:54:52.627000",
          "content": "<p><a href=\"https://www.kaggle.com/aadhavvignesh\" target=\"_blank\">@aadhavvignesh</a> We are using the same recipes as well and some additional things as well ,<br>\n<a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> I too have experimented on a lot of things that haven't work well for me though on lb , but I am very much trusting the ideas to work on private . We are yet to ensemble models as we were waiting for the last week 😉😉😉 . The only thing was we didn't have the full support of the team as they had prior commitments but nevertheless I have learned a lot here because all the pipelines I have written myself , I hope it ends well for us <br>\nThanks Chris for this post</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1032585,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-09-30T09:57:56.777000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1031738": "Dear fellow Kagglers,\n\n let me start with a quick summary of **what we know so far**:\n- We know so far that[ CV is definitely the way to go](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/186683) instead of relying purely on the public LB-score, which is only based on 15% of the data.\n\n- GroupKFold strongly beats the usage of basic KFold, as it reduces leakage.\n\n- [IMG data only grants very limited additional value (except for tissue/lung-area feature)](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/185077) and it's hard to engineer useful features out of the few pictures we have.\n\n\n\n### What is not completely known by now is the following:\n- Is using PERCENT as a feature reasonable? \nDefintion:\n> Percent- a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics\n\nLet me quickly kickstart the debate: \nAs I explained [here ](https://www.kaggle.com/ChristianDenich/quantile-reg-lr-schedulers-checkpoints) NOT using percent at all seems kind of wrong to me. Using it for all training will introduce leakage, as the combination of AGE and PERCENT is enough to numerically calculate the correct FVC. So my idea is to only use it for the first row for each patient, as it's also available in the test data and gives us insigts.\n\n- What's your preferred way of calculating the confidence/sigma? So far very different approaches have been introduced.\n\n\n",
    "1033335": "@ChristianDenich About PERCENT: As per my analysis, two people of the same gender and age don't have standard FVC values as same (if you divide the patient-specific FVC by % and find out 'standard' FVC for that person/gender/age combination). Standard FVC seems to be the function of body weight/height/build/BMI/race/ethnicity etc. also. In that sense % in incomplete information as provided here.",
    "1033862": "If the Percent feature is wrong to use, a shakeup is inevitable. There is a big gap in CV/LB when it's not included.",
    "1032637": "@gannadolinska, @ulrich07, @carlossouza any insights you want to share?",
    "1032585": ""
  }
}