{
  "id": 189705,
  "title": "Thanks to everyone and it's time to give back!",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/189705",
  "author_name": "",
  "post_date": "2020-10-08T12:23:30.008496700Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, I would like to thank competition host and every wonderful kagglers whoever worked really hard with a lot of passions. I ended up with 156th place in Private LB, and fortunately got a bronze medal (first medal in competition!). </p>\n<p>Now, it's time for me to give back something I have done. I think a lot of top-ranked kagglers had already given great overall solutions, so I would like to share about <strong>data preprocessing and CV construction</strong>, which saved me from selecting low scored notebook for private dataset. </p>\n<p>Before diving into main discussion, I would like to mention that I <strong>did not use image info at all</strong>. There is a lot of reason behind this (like a vast loads of work in uni), but would like to note this first. </p>\n<h2>Data preprocessing</h2>\n<p>Okay, let's move on. For preprocessing, as a lot of kagglers noticed, the <strong>\"percentage\" data was totally unnecessary data</strong>. It will be a good data for predicting FVC in train data, but it is only a noise for test prediction, since it is totally estimated (or say, systematically calculated). Otherwise, I just did min-max scaling for continuous numerical data. </p>\n<h2>CV construction</h2>\n<p>Since this competition measured public LB score with extremely small size of data, construction of reliable CV was crucial. Since this competition dataset (without images) can be seen as a hybrid of time series and tabular data, <strong>I assumed that CV should reflect both properties</strong>. Precisely, I thought we need to add one single property to GroupKFold, which reflects time series of data. When we look at sample submission data, we can only observe <code>Patient_ID</code> and <code>Week</code> properties. Thus, I used GroupStratifiedKFold CV, with grouping <code>Patient_ID</code> and stratifying with <code>Week</code>. This reliable k fold CV gave me a good estimate of actual score of model (although there was some deviation between scores since some week only appears once). Also this saved me from choosing notebook which received low score on private LB.</p>\n<p>Thanks for reading! If you have any questions or thoughts, feel free to comment!</p>",
  "messages": [
    {
      "id": "1042715",
      "postDate": "10/08/2020 12:23:30",
      "content": "<p>First of all, I would like to thank competition host and every wonderful kagglers whoever worked really hard with a lot of passions. I ended up with 156th place in Private LB, and fortunately got a bronze medal (first medal in competition!). </p>\n<p>Now, it's time for me to give back something I have done. I think a lot of top-ranked kagglers had already given great overall solutions, so I would like to share about <strong>data preprocessing and CV construction</strong>, which saved me from selecting low scored notebook for private dataset. </p>\n<p>Before diving into main discussion, I would like to mention that I <strong>did not use image info at all</strong>. There is a lot of reason behind this (like a vast loads of work in uni), but would like to note this first. </p>\n<h2>Data preprocessing</h2>\n<p>Okay, let's move on. For preprocessing, as a lot of kagglers noticed, the <strong>\"percentage\" data was totally unnecessary data</strong>. It will be a good data for predicting FVC in train data, but it is only a noise for test prediction, since it is totally estimated (or say, systematically calculated). Otherwise, I just did min-max scaling for continuous numerical data. </p>\n<h2>CV construction</h2>\n<p>Since this competition measured public LB score with extremely small size of data, construction of reliable CV was crucial. Since this competition dataset (without images) can be seen as a hybrid of time series and tabular data, <strong>I assumed that CV should reflect both properties</strong>. Precisely, I thought we need to add one single property to GroupKFold, which reflects time series of data. When we look at sample submission data, we can only observe <code>Patient_ID</code> and <code>Week</code> properties. Thus, I used GroupStratifiedKFold CV, with grouping <code>Patient_ID</code> and stratifying with <code>Week</code>. This reliable k fold CV gave me a good estimate of actual score of model (although there was some deviation between scores since some week only appears once). Also this saved me from choosing notebook which received low score on private LB.</p>\n<p>Thanks for reading! If you have any questions or thoughts, feel free to comment!</p>",
      "rawMarkdown": "First of all, I would like to thank competition host and every wonderful kagglers whoever worked really hard with a lot of passions. I ended up with 156th place in Private LB, and fortunately got a bronze medal (first medal in competition!). \n\nNow, it's time for me to give back something I have done. I think a lot of top-ranked kagglers had already given great overall solutions, so I would like to share about **data preprocessing and CV construction**, which saved me from selecting low scored notebook for private dataset. \n\nBefore diving into main discussion, I would like to mention that I **did not use image info at all**. There is a lot of reason behind this (like a vast loads of work in uni), but would like to note this first. \n\n## Data preprocessing \nOkay, let's move on. For preprocessing, as a lot of kagglers noticed, the **\"percentage\" data was totally unnecessary data**. It will be a good data for predicting FVC in train data, but it is only a noise for test prediction, since it is totally estimated (or say, systematically calculated). Otherwise, I just did min-max scaling for continuous numerical data. \n\n## CV construction\nSince this competition measured public LB score with extremely small size of data, construction of reliable CV was crucial. Since this competition dataset (without images) can be seen as a hybrid of time series and tabular data, **I assumed that CV should reflect both properties**. Precisely, I thought we need to add one single property to GroupKFold, which reflects time series of data. When we look at sample submission data, we can only observe `Patient_ID` and `Week` properties. Thus, I used GroupStratifiedKFold CV, with grouping `Patient_ID` and stratifying with `Week`. This reliable k fold CV gave me a good estimate of actual score of model (although there was some deviation between scores since some week only appears once). Also this saved me from choosing notebook which received low score on private LB.\n\nThanks for reading! If you have any questions or thoughts, feel free to comment!",
      "votes": null
    },
    {
      "id": "1044369",
      "postDate": "10/09/2020 17:40:08",
      "content": "<p>That's a great observation you did in your kfolding. I only grouped by <code>Patient_ID</code> and hoped for the k-fold to do the heavy job ^^'</p>\n<p>About the percentage I must disagree. As some other kaggler noted, the percentage is computed from some data that is not present in the tabular data, as height or etnicity, so it could be encoding hidden useful features. I didn't have time to dive much into that, but I did discover that <strong>there's a constant relation between a patient's FVC and its corresponding percentage</strong>. This relationship is almost unique per patient <strong>and is linear</strong>. I noted that I could compute that relationship factor from the first given FVC and Percent and predicting both with separatedly fitted modules helped in my CV score (I shifted the predicted FVC-Percent value pairs following the shortest path towards the line <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F809080%2F908d88694cb7356cf9da996f2f34568d%2Fcao.png?generation=1602265190174042&amp;alt=media\" alt=\"\">.</p>",
      "rawMarkdown": "That's a great observation you did in your kfolding. I only grouped by ```Patient_ID``` and hoped for the k-fold to do the heavy job ^^'\n\nAbout the percentage I must disagree. As some other kaggler noted, the percentage is computed from some data that is not present in the tabular data, as height or etnicity, so it could be encoding hidden useful features. I didn't have time to dive much into that, but I did discover that **there's a constant relation between a patient's FVC and its corresponding percentage**. This relationship is almost unique per patient **and is linear**. I noted that I could compute that relationship factor from the first given FVC and Percent and predicting both with separatedly fitted modules helped in my CV score (I shifted the predicted FVC-Percent value pairs following the shortest path towards the line ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F809080%2F908d88694cb7356cf9da996f2f34568d%2Fcao.png?generation=1602265190174042&alt=media).",
      "votes": null
    },
    {
      "id": "1044385",
      "postDate": "10/09/2020 17:56:18",
      "content": "<p>I used GroupKFold as well, which really makes sense, and I was surprised to not see it in public notebooks. I didn't think about stratifying week though. Good call! </p>\n<p>BTW - I was wondering regarding the use of images, and whether or not it proved beneficial at all in <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/190016\" target=\"_blank\">this tread</a>. Would be glad to hear your thoughts. </p>",
      "rawMarkdown": "I used GroupKFold as well, which really makes sense, and I was surprised to not see it in public notebooks. I didn't think about stratifying week though. Good call! \n\nBTW - I was wondering regarding the use of images, and whether or not it proved beneficial at all in [this tread](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/190016). Would be glad to hear your thoughts.",
      "votes": null
    },
    {
      "id": "1045342",
      "postDate": "10/10/2020 14:34:19",
      "content": "<p>Thanks for your comment! </p>\n<p>Ah, honestly, the use of image was not as much beneficial in terms of laplace log likelihood metric in my opinion. Well, this is only a consequence of this single competition, so I cannot tell whether it influences prediction accuracy in positive way or no in total. Might need to check for it with different metrics, or even different cases, and average out the thoughts. </p>",
      "rawMarkdown": "Thanks for your comment! \n\nAh, honestly, the use of image was not as much beneficial in terms of laplace log likelihood metric in my opinion. Well, this is only a consequence of this single competition, so I cannot tell whether it influences prediction accuracy in positive way or no in total. Might need to check for it with different metrics, or even different cases, and average out the thoughts.",
      "votes": null
    },
    {
      "id": "1045358",
      "postDate": "10/10/2020 14:48:25",
      "content": "<p>Thanks for your comment! </p>\n<p>First of all, I did know that percent has certain encoded features, and also has linear constant relationship with FVC, which is actually denoted in official data explanation: <br>\n<code>a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics</code>. <br>\nIn fact, I tried to extract some features, but I couldn't since I had no time. Instead of including these, I just dropped it, in order to avoid a leak. However, I didn't have that idea of \"shifting pair of predicted data with computed factors\". YES, that post processing will definitely work with lower risk of leakage. Thanks for sharing a great thoughts! </p>",
      "rawMarkdown": "Thanks for your comment! \n\nFirst of all, I did know that percent has certain encoded features, and also has linear constant relationship with FVC, which is actually denoted in official data explanation: \n`a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics`. \nIn fact, I tried to extract some features, but I couldn't since I had no time. Instead of including these, I just dropped it, in order to avoid a leak. However, I didn't have that idea of \"shifting pair of predicted data with computed factors\". YES, that post processing will definitely work with lower risk of leakage. Thanks for sharing a great thoughts!",
      "votes": null
    },
    {
      "id": "1045427",
      "postDate": "10/10/2020 16:14:36",
      "content": "<p>I believe that the notable bit of all this is that the typical FVC didn't seem to change across all captured patient's measures, which I thought could happen if it was grounded on age, weight or such, and that's what made the shifting postprocessing possible without risking leaking data. I also found non-trivial to extract those hidden features.</p>",
      "rawMarkdown": "I believe that the notable bit of all this is that the typical FVC didn't seem to change across all captured patient's measures, which I thought could happen if it was grounded on age, weight or such, and that's what made the shifting postprocessing possible without risking leaking data. I also found non-trivial to extract those hidden features.",
      "votes": null
    },
    {
      "id": "1045713",
      "postDate": "10/10/2020 23:49:50",
      "content": "<p>Definitely true. Actually, to get some hidden features, we had too small amount of features. Thus it might also required the use of external data. </p>",
      "rawMarkdown": "Definitely true. Actually, to get some hidden features, we had too small amount of features. Thus it might also required the use of external data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1044369,
      "author_name": "dcasbol",
      "author_url": "",
      "post_date": "10/09/2020 17:40:08",
      "content": "<p>That's a great observation you did in your kfolding. I only grouped by <code>Patient_ID</code> and hoped for the k-fold to do the heavy job ^^'</p>\n<p>About the percentage I must disagree. As some other kaggler noted, the percentage is computed from some data that is not present in the tabular data, as height or etnicity, so it could be encoding hidden useful features. I didn't have time to dive much into that, but I did discover that <strong>there's a constant relation between a patient's FVC and its corresponding percentage</strong>. This relationship is almost unique per patient <strong>and is linear</strong>. I noted that I could compute that relationship factor from the first given FVC and Percent and predicting both with separatedly fitted modules helped in my CV score (I shifted the predicted FVC-Percent value pairs following the shortest path towards the line <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F809080%2F908d88694cb7356cf9da996f2f34568d%2Fcao.png?generation=1602265190174042&amp;alt=media\" alt=\"\">.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1045358,
          "author_name": "harutot",
          "author_url": "",
          "post_date": "10/10/2020 14:48:25",
          "content": "<p>Thanks for your comment! </p>\n<p>First of all, I did know that percent has certain encoded features, and also has linear constant relationship with FVC, which is actually denoted in official data explanation: <br>\n<code>a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics</code>. <br>\nIn fact, I tried to extract some features, but I couldn't since I had no time. Instead of including these, I just dropped it, in order to avoid a leak. However, I didn't have that idea of \"shifting pair of predicted data with computed factors\". YES, that post processing will definitely work with lower risk of leakage. Thanks for sharing a great thoughts! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045427,
          "author_name": "dcasbol",
          "author_url": "",
          "post_date": "10/10/2020 16:14:36",
          "content": "<p>I believe that the notable bit of all this is that the typical FVC didn't seem to change across all captured patient's measures, which I thought could happen if it was grounded on age, weight or such, and that's what made the shifting postprocessing possible without risking leaking data. I also found non-trivial to extract those hidden features.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1045713,
          "author_name": "harutot",
          "author_url": "",
          "post_date": "10/10/2020 23:49:50",
          "content": "<p>Definitely true. Actually, to get some hidden features, we had too small amount of features. Thus it might also required the use of external data. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1044385,
      "author_name": "shovalt",
      "author_url": "",
      "post_date": "10/09/2020 17:56:18",
      "content": "<p>I used GroupKFold as well, which really makes sense, and I was surprised to not see it in public notebooks. I didn't think about stratifying week though. Good call! </p>\n<p>BTW - I was wondering regarding the use of images, and whether or not it proved beneficial at all in <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/190016\" target=\"_blank\">this tread</a>. Would be glad to hear your thoughts. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1045342,
          "author_name": "harutot",
          "author_url": "",
          "post_date": "10/10/2020 14:34:19",
          "content": "<p>Thanks for your comment! </p>\n<p>Ah, honestly, the use of image was not as much beneficial in terms of laplace log likelihood metric in my opinion. Well, this is only a consequence of this single competition, so I cannot tell whether it influences prediction accuracy in positive way or no in total. Might need to check for it with different metrics, or even different cases, and average out the thoughts. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1042715": "First of all, I would like to thank competition host and every wonderful kagglers whoever worked really hard with a lot of passions. I ended up with 156th place in Private LB, and fortunately got a bronze medal (first medal in competition!). \n\nNow, it's time for me to give back something I have done. I think a lot of top-ranked kagglers had already given great overall solutions, so I would like to share about **data preprocessing and CV construction**, which saved me from selecting low scored notebook for private dataset. \n\nBefore diving into main discussion, I would like to mention that I **did not use image info at all**. There is a lot of reason behind this (like a vast loads of work in uni), but would like to note this first. \n\n## Data preprocessing \nOkay, let's move on. For preprocessing, as a lot of kagglers noticed, the **\"percentage\" data was totally unnecessary data**. It will be a good data for predicting FVC in train data, but it is only a noise for test prediction, since it is totally estimated (or say, systematically calculated). Otherwise, I just did min-max scaling for continuous numerical data. \n\n## CV construction\nSince this competition measured public LB score with extremely small size of data, construction of reliable CV was crucial. Since this competition dataset (without images) can be seen as a hybrid of time series and tabular data, **I assumed that CV should reflect both properties**. Precisely, I thought we need to add one single property to GroupKFold, which reflects time series of data. When we look at sample submission data, we can only observe `Patient_ID` and `Week` properties. Thus, I used GroupStratifiedKFold CV, with grouping `Patient_ID` and stratifying with `Week`. This reliable k fold CV gave me a good estimate of actual score of model (although there was some deviation between scores since some week only appears once). Also this saved me from choosing notebook which received low score on private LB.\n\nThanks for reading! If you have any questions or thoughts, feel free to comment!",
    "1044369": "That's a great observation you did in your kfolding. I only grouped by ```Patient_ID``` and hoped for the k-fold to do the heavy job ^^'\n\nAbout the percentage I must disagree. As some other kaggler noted, the percentage is computed from some data that is not present in the tabular data, as height or etnicity, so it could be encoding hidden useful features. I didn't have time to dive much into that, but I did discover that **there's a constant relation between a patient's FVC and its corresponding percentage**. This relationship is almost unique per patient **and is linear**. I noted that I could compute that relationship factor from the first given FVC and Percent and predicting both with separatedly fitted modules helped in my CV score (I shifted the predicted FVC-Percent value pairs following the shortest path towards the line ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F809080%2F908d88694cb7356cf9da996f2f34568d%2Fcao.png?generation=1602265190174042&alt=media).",
    "1044385": "I used GroupKFold as well, which really makes sense, and I was surprised to not see it in public notebooks. I didn't think about stratifying week though. Good call! \n\nBTW - I was wondering regarding the use of images, and whether or not it proved beneficial at all in [this tread](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/190016). Would be glad to hear your thoughts.",
    "1045342": "Thanks for your comment! \n\nAh, honestly, the use of image was not as much beneficial in terms of laplace log likelihood metric in my opinion. Well, this is only a consequence of this single competition, so I cannot tell whether it influences prediction accuracy in positive way or no in total. Might need to check for it with different metrics, or even different cases, and average out the thoughts.",
    "1045358": "Thanks for your comment! \n\nFirst of all, I did know that percent has certain encoded features, and also has linear constant relationship with FVC, which is actually denoted in official data explanation: \n`a computed field which approximates the patient's FVC as a percent of the typical FVC for a person of similar characteristics`. \nIn fact, I tried to extract some features, but I couldn't since I had no time. Instead of including these, I just dropped it, in order to avoid a leak. However, I didn't have that idea of \"shifting pair of predicted data with computed factors\". YES, that post processing will definitely work with lower risk of leakage. Thanks for sharing a great thoughts!",
    "1045427": "I believe that the notable bit of all this is that the typical FVC didn't seem to change across all captured patient's measures, which I thought could happen if it was grounded on age, weight or such, and that's what made the shifting postprocessing possible without risking leaking data. I also found non-trivial to extract those hidden features.",
    "1045713": "Definitely true. Actually, to get some hidden features, we had too small amount of features. Thus it might also required the use of external data."
  },
  "source": "meta"
}