{
  "id": 189201,
  "title": "Best score using only tabular data",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/189201",
  "author_name": "",
  "post_date": "2020-10-07T01:19:08.975392100Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>The disease progress is very unpredictable as we can see by ploting each <a href=\"https://www.kaggle.com/jsaguiar/fvc-curve-for-each-patient-and-outliers\" target=\"_blank\">FVC curve</a>. I guess it is quite difficult to consistently beat simple linear models using NN and a single CT scan. That said, it would be interesting to see:</p>\n<ol>\n<li><strong>What kind of model did you use?</strong></li>\n<li><strong>Did you use any CT scans (images)?</strong></li>\n<li><strong>How did you get the confidence value?</strong></li>\n</ol>\n<p>I used a very similar Ridge model to the one from <a href=\"https://www.kaggle.com/yasufuminakama/osic-ridge-baseline\" target=\"_blank\">this public notebook</a> (without images), which btw would get a bronze medal by just forking and making a submission (around 140th PL)…</p>",
  "messages": [
    {
      "id": "1040089",
      "postDate": "10/07/2020 01:19:08",
      "content": "<p>The disease progress is very unpredictable as we can see by ploting each <a href=\"https://www.kaggle.com/jsaguiar/fvc-curve-for-each-patient-and-outliers\" target=\"_blank\">FVC curve</a>. I guess it is quite difficult to consistently beat simple linear models using NN and a single CT scan. That said, it would be interesting to see:</p>\n<ol>\n<li><strong>What kind of model did you use?</strong></li>\n<li><strong>Did you use any CT scans (images)?</strong></li>\n<li><strong>How did you get the confidence value?</strong></li>\n</ol>\n<p>I used a very similar Ridge model to the one from <a href=\"https://www.kaggle.com/yasufuminakama/osic-ridge-baseline\" target=\"_blank\">this public notebook</a> (without images), which btw would get a bronze medal by just forking and making a submission (around 140th PL)…</p>",
      "rawMarkdown": "The disease progress is very unpredictable as we can see by ploting each [FVC curve](https://www.kaggle.com/jsaguiar/fvc-curve-for-each-patient-and-outliers). I guess it is quite difficult to consistently beat simple linear models using NN and a single CT scan. That said, it would be interesting to see:\n\n1. **What kind of model did you use?**\n2. **Did you use any CT scans (images)?**\n3. **How did you get the confidence value?**\n\nI used a very similar Ridge model to the one from [this public notebook] (https://www.kaggle.com/yasufuminakama/osic-ridge-baseline) (without images), which btw would get a bronze medal by just forking and making a submission (around 140th PL)...",
      "votes": null
    },
    {
      "id": "1040146",
      "postDate": "10/07/2020 02:06:30",
      "content": "<p>My current LB position is based only on tabular data. I didn't use any image data for the submission. I shall post the notebook public. The private score was -6.8438 and public lb score was -6.8504.. Mine was  CNN + quantile regression ( 5 different models), Bayesian Ridge and Ridge ensemble..</p>",
      "rawMarkdown": "My current LB position is based only on tabular data. I didn't use any image data for the submission. I shall post the notebook public. The private score was -6.8438 and public lb score was -6.8504.. Mine was  CNN + quantile regression ( 5 different models), Bayesian Ridge and Ridge ensemble..",
      "votes": null
    },
    {
      "id": "1040552",
      "postDate": "10/07/2020 08:09:01",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "1040576",
      "postDate": "10/07/2020 08:25:50",
      "content": "<p>Glad to know my notebook helped :)</p>",
      "rawMarkdown": "Glad to know my notebook helped :)",
      "votes": null
    },
    {
      "id": "1040610",
      "postDate": "10/07/2020 08:56:26",
      "content": "<p>We only used tabular data, the write up is coming soon</p>",
      "rawMarkdown": "We only used tabular data, the write up is coming soon",
      "votes": null
    },
    {
      "id": "1040624",
      "postDate": "10/07/2020 09:04:28",
      "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> - My part of code was also based on your notebook.. Thanks for sharing :).  </p>",
      "rawMarkdown": "yasufuminakama - My part of code was also based on your notebook.. Thanks for sharing :).",
      "votes": null
    },
    {
      "id": "1040761",
      "postDate": "10/07/2020 10:35:02",
      "content": "<ol>\n<li><p>One of my best results was based just on fitting a decay slope from the smoking status, so I totally understand what you say about difficulty (or easiness I'd say. I observed mostly linear progressions). However, my best result was based on directly predicting the FVC via a randomised ensemble of MLPs.</p></li>\n<li><p>I tried to, but it was both very time consuming, hard to implement/make it work and gave little return, so I finally focused on tabular data.</p></li>\n<li><p>I got the confidence value by computing the uncertainty via MC-dropout plus expected error variance. Then I took the standard deviation of all that.</p></li>\n</ol>\n<p>I summarised everything in <a href=\"https://www.kaggle.com/dcasbol/osic-my-best-solution-6-8348\" target=\"_blank\">this notebook</a>.</p>",
      "rawMarkdown": "1. One of my best results was based just on fitting a decay slope from the smoking status, so I totally understand what you say about difficulty (or easiness I'd say. I observed mostly linear progressions). However, my best result was based on directly predicting the FVC via a randomised ensemble of MLPs.\n\n2. I tried to, but it was both very time consuming, hard to implement/make it work and gave little return, so I finally focused on tabular data.\n\n3. I got the confidence value by computing the uncertainty via MC-dropout plus expected error variance. Then I took the standard deviation of all that.\n\nI summarised everything in [this notebook](https://www.kaggle.com/dcasbol/osic-my-best-solution-6-8348).",
      "votes": null
    },
    {
      "id": "1040814",
      "postDate": "10/07/2020 11:13:44",
      "content": "<p>In terms of individual models:</p>\n<ul>\n<li>I tried using 3 separate LightGBMs for different quantiles (0.158655, 0.5 and  0.8413447 based on theoretical considerations), but with the hyperparameters fixed to be the same (I also tried completely separate hyperparameters, but that seemed to somehow be less stable in CV), which got a private LB of -6.8520 and public-6.9485. The public LB score put us off this solution, so it was just one of many models in our final stacking.</li>\n<li>One of the most interesting models I had was a pystan Bayesian model (private LB -6.9489, public LB -6.9742), notebook with explanation is <a href=\"https://www.kaggle.com/bjoernholzhauer/bayes-decision-theory-using-pystan\" target=\"_blank\">here</a> that applied a decision theoretic framework to pick FVC (unsurprisingly usually about the posterior median) and Confidence (more interesting).</li>\n<li>I applied the same approach to Carlos Souza's Bayesian notebook (private LB: -6.8898, public LB -6.9208 when using their priors, private LB -6.8876, public -6.9248 when going with something I believed in more). Looks like me making my pystan model more complicated by trying to find what explains between patient variation just made the private LB worse, even if it looked good in CV.</li>\n<li>For some of the models that do not easily give you uncertainty (such as one elastic net that achieved private LB -6.8969, public LB -6.9674), I tried using the out-of-fold MAE and finding the single constant value for Confidence that maximizes the competition metric. That felt too dangerous though, if you then want to stack afterwards, as the solution may overfit the out-of-fold values. </li>\n<li>I tried pretty simple tabular NNs (with the loss function from <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">the famous public notebook</a> so many people used), a very basic one in keras (private: -6.9225, public: -7.0509) and a fancier multi-output one using fastai (private -6.9516, public: -6.9671).</li>\n</ul>\n<p>Our final stacking as a team did include one model CT scans, but otherwise 6 tabular models, as base models.</p>",
      "rawMarkdown": "In terms of individual models:\n* I tried using 3 separate LightGBMs for different quantiles (0.158655, 0.5 and  0.8413447 based on theoretical considerations), but with the hyperparameters fixed to be the same (I also tried completely separate hyperparameters, but that seemed to somehow be less stable in CV), which got a private LB of -6.8520 and public-6.9485. The public LB score put us off this solution, so it was just one of many models in our final stacking.\n* One of the most interesting models I had was a pystan Bayesian model (private LB -6.9489, public LB -6.9742), notebook with explanation is [here](https://www.kaggle.com/bjoernholzhauer/bayes-decision-theory-using-pystan) that applied a decision theoretic framework to pick FVC (unsurprisingly usually about the posterior median) and Confidence (more interesting).\n* I applied the same approach to Carlos Souza's Bayesian notebook (private LB: -6.8898, public LB -6.9208 when using their priors, private LB -6.8876, public -6.9248 when going with something I believed in more). Looks like me making my pystan model more complicated by trying to find what explains between patient variation just made the private LB worse, even if it looked good in CV.\n* For some of the models that do not easily give you uncertainty (such as one elastic net that achieved private LB -6.8969, public LB -6.9674), I tried using the out-of-fold MAE and finding the single constant value for Confidence that maximizes the competition metric. That felt too dangerous though, if you then want to stack afterwards, as the solution may overfit the out-of-fold values. \n* I tried pretty simple tabular NNs (with the loss function from [the famous public notebook](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter) so many people used), a very basic one in keras (private: -6.9225, public: -7.0509) and a fancier multi-output one using fastai (private -6.9516, public: -6.9671).\n\nOur final stacking as a team did include one model CT scans, but otherwise 6 tabular models, as base models.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1040146,
      "author_name": "manojprabhaakr",
      "author_url": "",
      "post_date": "10/07/2020 02:06:30",
      "content": "<p>My current LB position is based only on tabular data. I didn't use any image data for the submission. I shall post the notebook public. The private score was -6.8438 and public lb score was -6.8504.. Mine was  CNN + quantile regression ( 5 different models), Bayesian Ridge and Ridge ensemble..</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040576,
      "author_name": "yasufuminakama",
      "author_url": "",
      "post_date": "10/07/2020 08:25:50",
      "content": "<p>Glad to know my notebook helped :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1040624,
          "author_name": "manojprabhaakr",
          "author_url": "",
          "post_date": "10/07/2020 09:04:28",
          "content": "<p><a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> - My part of code was also based on your notebook.. Thanks for sharing :).  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1040610,
      "author_name": "lukereijnen",
      "author_url": "",
      "post_date": "10/07/2020 08:56:26",
      "content": "<p>We only used tabular data, the write up is coming soon</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040761,
      "author_name": "dcasbol",
      "author_url": "",
      "post_date": "10/07/2020 10:35:02",
      "content": "<ol>\n<li><p>One of my best results was based just on fitting a decay slope from the smoking status, so I totally understand what you say about difficulty (or easiness I'd say. I observed mostly linear progressions). However, my best result was based on directly predicting the FVC via a randomised ensemble of MLPs.</p></li>\n<li><p>I tried to, but it was both very time consuming, hard to implement/make it work and gave little return, so I finally focused on tabular data.</p></li>\n<li><p>I got the confidence value by computing the uncertainty via MC-dropout plus expected error variance. Then I took the standard deviation of all that.</p></li>\n</ol>\n<p>I summarised everything in <a href=\"https://www.kaggle.com/dcasbol/osic-my-best-solution-6-8348\" target=\"_blank\">this notebook</a>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040814,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "10/07/2020 11:13:44",
      "content": "<p>In terms of individual models:</p>\n<ul>\n<li>I tried using 3 separate LightGBMs for different quantiles (0.158655, 0.5 and  0.8413447 based on theoretical considerations), but with the hyperparameters fixed to be the same (I also tried completely separate hyperparameters, but that seemed to somehow be less stable in CV), which got a private LB of -6.8520 and public-6.9485. The public LB score put us off this solution, so it was just one of many models in our final stacking.</li>\n<li>One of the most interesting models I had was a pystan Bayesian model (private LB -6.9489, public LB -6.9742), notebook with explanation is <a href=\"https://www.kaggle.com/bjoernholzhauer/bayes-decision-theory-using-pystan\" target=\"_blank\">here</a> that applied a decision theoretic framework to pick FVC (unsurprisingly usually about the posterior median) and Confidence (more interesting).</li>\n<li>I applied the same approach to Carlos Souza's Bayesian notebook (private LB: -6.8898, public LB -6.9208 when using their priors, private LB -6.8876, public -6.9248 when going with something I believed in more). Looks like me making my pystan model more complicated by trying to find what explains between patient variation just made the private LB worse, even if it looked good in CV.</li>\n<li>For some of the models that do not easily give you uncertainty (such as one elastic net that achieved private LB -6.8969, public LB -6.9674), I tried using the out-of-fold MAE and finding the single constant value for Confidence that maximizes the competition metric. That felt too dangerous though, if you then want to stack afterwards, as the solution may overfit the out-of-fold values. </li>\n<li>I tried pretty simple tabular NNs (with the loss function from <a href=\"https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter\" target=\"_blank\">the famous public notebook</a> so many people used), a very basic one in keras (private: -6.9225, public: -7.0509) and a fancier multi-output one using fastai (private -6.9516, public: -6.9671).</li>\n</ul>\n<p>Our final stacking as a team did include one model CT scans, but otherwise 6 tabular models, as base models.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1040552,
      "author_name": "deepika1212",
      "author_url": "",
      "post_date": "10/07/2020 08:09:01",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1040089": "The disease progress is very unpredictable as we can see by ploting each [FVC curve](https://www.kaggle.com/jsaguiar/fvc-curve-for-each-patient-and-outliers). I guess it is quite difficult to consistently beat simple linear models using NN and a single CT scan. That said, it would be interesting to see:\n\n1. **What kind of model did you use?**\n2. **Did you use any CT scans (images)?**\n3. **How did you get the confidence value?**\n\nI used a very similar Ridge model to the one from [this public notebook] (https://www.kaggle.com/yasufuminakama/osic-ridge-baseline) (without images), which btw would get a bronze medal by just forking and making a submission (around 140th PL)...",
    "1040146": "My current LB position is based only on tabular data. I didn't use any image data for the submission. I shall post the notebook public. The private score was -6.8438 and public lb score was -6.8504.. Mine was  CNN + quantile regression ( 5 different models), Bayesian Ridge and Ridge ensemble..",
    "1040552": "Thanks!",
    "1040576": "Glad to know my notebook helped :)",
    "1040610": "We only used tabular data, the write up is coming soon",
    "1040624": "yasufuminakama - My part of code was also based on your notebook.. Thanks for sharing :).",
    "1040761": "1. One of my best results was based just on fitting a decay slope from the smoking status, so I totally understand what you say about difficulty (or easiness I'd say. I observed mostly linear progressions). However, my best result was based on directly predicting the FVC via a randomised ensemble of MLPs.\n\n2. I tried to, but it was both very time consuming, hard to implement/make it work and gave little return, so I finally focused on tabular data.\n\n3. I got the confidence value by computing the uncertainty via MC-dropout plus expected error variance. Then I took the standard deviation of all that.\n\nI summarised everything in [this notebook](https://www.kaggle.com/dcasbol/osic-my-best-solution-6-8348).",
    "1040814": "In terms of individual models:\n* I tried using 3 separate LightGBMs for different quantiles (0.158655, 0.5 and  0.8413447 based on theoretical considerations), but with the hyperparameters fixed to be the same (I also tried completely separate hyperparameters, but that seemed to somehow be less stable in CV), which got a private LB of -6.8520 and public-6.9485. The public LB score put us off this solution, so it was just one of many models in our final stacking.\n* One of the most interesting models I had was a pystan Bayesian model (private LB -6.9489, public LB -6.9742), notebook with explanation is [here](https://www.kaggle.com/bjoernholzhauer/bayes-decision-theory-using-pystan) that applied a decision theoretic framework to pick FVC (unsurprisingly usually about the posterior median) and Confidence (more interesting).\n* I applied the same approach to Carlos Souza's Bayesian notebook (private LB: -6.8898, public LB -6.9208 when using their priors, private LB -6.8876, public -6.9248 when going with something I believed in more). Looks like me making my pystan model more complicated by trying to find what explains between patient variation just made the private LB worse, even if it looked good in CV.\n* For some of the models that do not easily give you uncertainty (such as one elastic net that achieved private LB -6.8969, public LB -6.9674), I tried using the out-of-fold MAE and finding the single constant value for Confidence that maximizes the competition metric. That felt too dangerous though, if you then want to stack afterwards, as the solution may overfit the out-of-fold values. \n* I tried pretty simple tabular NNs (with the loss function from [the famous public notebook](https://www.kaggle.com/ulrich07/osic-multiple-quantile-regression-starter) so many people used), a very basic one in keras (private: -6.9225, public: -7.0509) and a fancier multi-output one using fastai (private -6.9516, public: -6.9671).\n\nOur final stacking as a team did include one model CT scans, but otherwise 6 tabular models, as base models."
  },
  "source": "meta"
}