{
  "id": 181981,
  "title": "Practicality of a Time-Series Model? ",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/181981",
  "author_name": "",
  "post_date": "2020-09-10T23:39:51.129530600Z",
  "votes": 7,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Initially, I was very surprised that many of the high-performing public notebooks do not utilize a time-series regression approach in the traditional sense. Instead, it seems as if the most common model is a vanilla neural network + multiple quantile regression to process the tabular data. </p>\n<p>I was wondering if it's practical to use a model such as LSTM which is specifically designed to process and extract useful information from sequential data or use a random forest with lagged timesteps. If so, why haven't more people used these types of models for this competition? (If there is a public notebook available that makes use of the fact that the data is time-series, could you please post a link on this thread?)</p>\n<p>If not, then why is it not applicable to this contest? Is it because of the inconsistent time intervals of each patient (FVC recorded on different weeks) or the small number of training samples per patient? </p>\n<p>Thank you for your comments and insights. I would appreciate anyone who could help clarify my understanding of this topic. </p>",
  "messages": [
    {
      "id": "1005967",
      "postDate": "09/10/2020 23:39:51",
      "content": "<p>Initially, I was very surprised that many of the high-performing public notebooks do not utilize a time-series regression approach in the traditional sense. Instead, it seems as if the most common model is a vanilla neural network + multiple quantile regression to process the tabular data. </p>\n<p>I was wondering if it's practical to use a model such as LSTM which is specifically designed to process and extract useful information from sequential data or use a random forest with lagged timesteps. If so, why haven't more people used these types of models for this competition? (If there is a public notebook available that makes use of the fact that the data is time-series, could you please post a link on this thread?)</p>\n<p>If not, then why is it not applicable to this contest? Is it because of the inconsistent time intervals of each patient (FVC recorded on different weeks) or the small number of training samples per patient? </p>\n<p>Thank you for your comments and insights. I would appreciate anyone who could help clarify my understanding of this topic. </p>",
      "rawMarkdown": "Initially, I was very surprised that many of the high-performing public notebooks do not utilize a time-series regression approach in the traditional sense. Instead, it seems as if the most common model is a vanilla neural network + multiple quantile regression to process the tabular data. \n\nI was wondering if it's practical to use a model such as LSTM which is specifically designed to process and extract useful information from sequential data or use a random forest with lagged timesteps. If so, why haven't more people used these types of models for this competition? (If there is a public notebook available that makes use of the fact that the data is time-series, could you please post a link on this thread?)\n\nIf not, then why is it not applicable to this contest? Is it because of the inconsistent time intervals of each patient (FVC recorded on different weeks) or the small number of training samples per patient? \n\nThank you for your comments and insights. I would appreciate anyone who could help clarify my understanding of this topic.",
      "votes": null
    },
    {
      "id": "1006307",
      "postDate": "09/11/2020 07:04:18",
      "content": "<p>Good question. Personally, I did not invest time in that direction, because I thought it was not so likely to work. I have two main reasons for this (but they are speculation, so I could be wrong). </p>\n<ol>\n<li>We will not have a time series as an input on the test data, but rather only a single time. That made me suspect that a lot of the benefits one gains from treating it as a time series might be lost (i.e. because LSTMs, transformers etc. are good at finding/exploiting the patterns in the times series inputs - something that takes a lot of careful feature engineering for a human - but here, there would be no such patterns to use, at least not on the test data). </li>\n<li>We have very few time series available on the training data (176 patients), so I was mostly tryint to put more assumptions into my models (at least for the tabular data) instead of letting a neural network figure it out \"on its own\". But, perhaps the right kind of regularization might overcome that issue.</li>\n</ol>\n<p>It could of course be that a time series model could learn from the training data out to sequentially predict the whole time series starting from just one observation. That might be worth trying. However, you'd also have to come up with a good way of handling the weeks in which there is no FVC available - e.g. setting FVC to 0 and then predicting each week in turn would likely not be a good idea, because then the model might also predict 0 for the weeks in the test set. Perhaps some smooth interpolation with noise added (could also act as a data augmentation?)?</p>",
      "rawMarkdown": "Good question. Personally, I did not invest time in that direction, because I thought it was not so likely to work. I have two main reasons for this (but they are speculation, so I could be wrong). \n1. We will not have a time series as an input on the test data, but rather only a single time. That made me suspect that a lot of the benefits one gains from treating it as a time series might be lost (i.e. because LSTMs, transformers etc. are good at finding/exploiting the patterns in the times series inputs - something that takes a lot of careful feature engineering for a human - but here, there would be no such patterns to use, at least not on the test data). \n2. We have very few time series available on the training data (176 patients), so I was mostly tryint to put more assumptions into my models (at least for the tabular data) instead of letting a neural network figure it out \"on its own\". But, perhaps the right kind of regularization might overcome that issue.\n\nIt could of course be that a time series model could learn from the training data out to sequentially predict the whole time series starting from just one observation. That might be worth trying. However, you'd also have to come up with a good way of handling the weeks in which there is no FVC available - e.g. setting FVC to 0 and then predicting each week in turn would likely not be a good idea, because then the model might also predict 0 for the weeks in the test set. Perhaps some smooth interpolation with noise added (could also act as a data augmentation?)?",
      "votes": null
    },
    {
      "id": "1006868",
      "postDate": "09/11/2020 15:48:10",
      "content": "<p>You could look at this: <a href=\"https://www.kaggle.com/eladwar/conditional-rnn?scriptVersionId=42399034\">https://www.kaggle.com/eladwar/conditional-rnn?scriptVersionId=42399034</a></p>",
      "rawMarkdown": "You could look at this: https://www.kaggle.com/eladwar/conditional-rnn?scriptVersionId=42399034",
      "votes": null
    },
    {
      "id": "1007703",
      "postDate": "09/12/2020 12:06:17",
      "content": "<p>Probably it's possible to design a [CT scan, tabular feature input] -&gt; CNN feature extractor -&gt; (CT scan features + tabular features) merged -&gt; LSTM -&gt; predict a time series -12 to 133 weeks.</p>\n<p>But this seems like a heavy model, trained on a really small data (not enough to validate).</p>",
      "rawMarkdown": "Probably it's possible to design a [CT scan, tabular feature input] -> CNN feature extractor -> (CT scan features + tabular features) merged -> LSTM -> predict a time series -12 to 133 weeks.\n\nBut this seems like a heavy model, trained on a really small data (not enough to validate).",
      "votes": null
    },
    {
      "id": "1015358",
      "postDate": "09/18/2020 06:22:50",
      "content": "<p>I've had a go at it, and it was fun and educational to try, but I think for the reasons mentioned by others here it's not quite viable…but then again it seems like a logical approach and I definitely don't know everything!</p>",
      "rawMarkdown": "I've had a go at it, and it was fun and educational to try, but I think for the reasons mentioned by others here it's not quite viable...but then again it seems like a logical approach and I definitely don't know everything!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1006307,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "09/11/2020 07:04:18",
      "content": "<p>Good question. Personally, I did not invest time in that direction, because I thought it was not so likely to work. I have two main reasons for this (but they are speculation, so I could be wrong). </p>\n<ol>\n<li>We will not have a time series as an input on the test data, but rather only a single time. That made me suspect that a lot of the benefits one gains from treating it as a time series might be lost (i.e. because LSTMs, transformers etc. are good at finding/exploiting the patterns in the times series inputs - something that takes a lot of careful feature engineering for a human - but here, there would be no such patterns to use, at least not on the test data). </li>\n<li>We have very few time series available on the training data (176 patients), so I was mostly tryint to put more assumptions into my models (at least for the tabular data) instead of letting a neural network figure it out \"on its own\". But, perhaps the right kind of regularization might overcome that issue.</li>\n</ol>\n<p>It could of course be that a time series model could learn from the training data out to sequentially predict the whole time series starting from just one observation. That might be worth trying. However, you'd also have to come up with a good way of handling the weeks in which there is no FVC available - e.g. setting FVC to 0 and then predicting each week in turn would likely not be a good idea, because then the model might also predict 0 for the weeks in the test set. Perhaps some smooth interpolation with noise added (could also act as a data augmentation?)?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1007703,
      "author_name": "furcifer",
      "author_url": "",
      "post_date": "09/12/2020 12:06:17",
      "content": "<p>Probably it's possible to design a [CT scan, tabular feature input] -&gt; CNN feature extractor -&gt; (CT scan features + tabular features) merged -&gt; LSTM -&gt; predict a time series -12 to 133 weeks.</p>\n<p>But this seems like a heavy model, trained on a really small data (not enough to validate).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1015358,
      "author_name": "geoffcc",
      "author_url": "",
      "post_date": "09/18/2020 06:22:50",
      "content": "<p>I've had a go at it, and it was fun and educational to try, but I think for the reasons mentioned by others here it's not quite viable…but then again it seems like a logical approach and I definitely don't know everything!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1006868,
      "author_name": "eladwar",
      "author_url": "",
      "post_date": "09/11/2020 15:48:10",
      "content": "<p>You could look at this: <a href=\"https://www.kaggle.com/eladwar/conditional-rnn?scriptVersionId=42399034\">https://www.kaggle.com/eladwar/conditional-rnn?scriptVersionId=42399034</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1005967": "Initially, I was very surprised that many of the high-performing public notebooks do not utilize a time-series regression approach in the traditional sense. Instead, it seems as if the most common model is a vanilla neural network + multiple quantile regression to process the tabular data. \n\nI was wondering if it's practical to use a model such as LSTM which is specifically designed to process and extract useful information from sequential data or use a random forest with lagged timesteps. If so, why haven't more people used these types of models for this competition? (If there is a public notebook available that makes use of the fact that the data is time-series, could you please post a link on this thread?)\n\nIf not, then why is it not applicable to this contest? Is it because of the inconsistent time intervals of each patient (FVC recorded on different weeks) or the small number of training samples per patient? \n\nThank you for your comments and insights. I would appreciate anyone who could help clarify my understanding of this topic.",
    "1006307": "Good question. Personally, I did not invest time in that direction, because I thought it was not so likely to work. I have two main reasons for this (but they are speculation, so I could be wrong). \n1. We will not have a time series as an input on the test data, but rather only a single time. That made me suspect that a lot of the benefits one gains from treating it as a time series might be lost (i.e. because LSTMs, transformers etc. are good at finding/exploiting the patterns in the times series inputs - something that takes a lot of careful feature engineering for a human - but here, there would be no such patterns to use, at least not on the test data). \n2. We have very few time series available on the training data (176 patients), so I was mostly tryint to put more assumptions into my models (at least for the tabular data) instead of letting a neural network figure it out \"on its own\". But, perhaps the right kind of regularization might overcome that issue.\n\nIt could of course be that a time series model could learn from the training data out to sequentially predict the whole time series starting from just one observation. That might be worth trying. However, you'd also have to come up with a good way of handling the weeks in which there is no FVC available - e.g. setting FVC to 0 and then predicting each week in turn would likely not be a good idea, because then the model might also predict 0 for the weeks in the test set. Perhaps some smooth interpolation with noise added (could also act as a data augmentation?)?",
    "1006868": "You could look at this: https://www.kaggle.com/eladwar/conditional-rnn?scriptVersionId=42399034",
    "1007703": "Probably it's possible to design a [CT scan, tabular feature input] -> CNN feature extractor -> (CT scan features + tabular features) merged -> LSTM -> predict a time series -12 to 133 weeks.\n\nBut this seems like a heavy model, trained on a really small data (not enough to validate).",
    "1015358": "I've had a go at it, and it was fun and educational to try, but I think for the reasons mentioned by others here it's not quite viable...but then again it seems like a logical approach and I definitely don't know everything!"
  },
  "source": "meta"
}