{
  "id": 400311,
  "title": "On the use of 'Time' as a feature",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/400311",
  "author_name": "",
  "post_date": "2023-04-07T17:52:26.704168600Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hey everyone -- I've been thinking about the time column in the data lately and wanted to share a few thoughts.  This competition scores users/teams on their prediction of FoG events in the tDCS and DeFOG datasets, which are comprised of accelerometer data during a (standard?) series of FoG-provoking tasks.  If folks take a <em>roughly</em> similar amount of time to complete the tasks and certain tasks elicit certain FoG events, then perhaps time is a reasonable feature to use when modeling the data in these datasets.  </p>\n<p>However, it seems to me that the ultimate goal or aspiration behind the competition is to identify (or even better predict) FoG events in <em>any</em> setting at <em>any</em> time. ** In this case, time would likely be a poor feature choice, even if the competition scoring (might?) incentivize the use of time as a feature due to the (standard?) series of tasks.  **</p>\n<p>Here's a lightly edited excerpt from the competition page for reference:</p>\n<blockquote>\n  <p>[The] Competition host … aims to improve the personalized treatment of age-related movement, cognition, and mobility disorders and to alleviate the associated burden. They leverage a combination of clinical, engineering, and neuroscience expertise to: 1) Gain new understandings into the physiologic and pathophysiologic mechanisms that contribute to cognitive and motor function, the factors that influence these functions, and their changes with aging and disease … 2) Develop new methods … for the early detection and tracking of cognitive and motor decline. A major focus is on using leveraging wearable devices and digital technologies; and 3) Develop and evaluate novel methods for the prevention and treatment of gait, falls, and cognitive function. … Your work will help advance the evaluation, understanding and treatment of FOG, improving the lives of the many people who suffer from this … symptom.</p>\n</blockquote>\n<p>I've seen a number of notebooks using time as a feature and am wondering if anyone has gleaned any insights that they'd be willing to share.  Do your models return reasonable results on the daily living data?  My hypothesis is no, but wanted to share nonetheless. </p>",
  "messages": [
    {
      "id": "2213613",
      "postDate": "04/07/2023 17:52:26",
      "content": "<p>Hey everyone -- I've been thinking about the time column in the data lately and wanted to share a few thoughts.  This competition scores users/teams on their prediction of FoG events in the tDCS and DeFOG datasets, which are comprised of accelerometer data during a (standard?) series of FoG-provoking tasks.  If folks take a <em>roughly</em> similar amount of time to complete the tasks and certain tasks elicit certain FoG events, then perhaps time is a reasonable feature to use when modeling the data in these datasets.  </p>\n<p>However, it seems to me that the ultimate goal or aspiration behind the competition is to identify (or even better predict) FoG events in <em>any</em> setting at <em>any</em> time. ** In this case, time would likely be a poor feature choice, even if the competition scoring (might?) incentivize the use of time as a feature due to the (standard?) series of tasks.  **</p>\n<p>Here's a lightly edited excerpt from the competition page for reference:</p>\n<blockquote>\n  <p>[The] Competition host … aims to improve the personalized treatment of age-related movement, cognition, and mobility disorders and to alleviate the associated burden. They leverage a combination of clinical, engineering, and neuroscience expertise to: 1) Gain new understandings into the physiologic and pathophysiologic mechanisms that contribute to cognitive and motor function, the factors that influence these functions, and their changes with aging and disease … 2) Develop new methods … for the early detection and tracking of cognitive and motor decline. A major focus is on using leveraging wearable devices and digital technologies; and 3) Develop and evaluate novel methods for the prevention and treatment of gait, falls, and cognitive function. … Your work will help advance the evaluation, understanding and treatment of FOG, improving the lives of the many people who suffer from this … symptom.</p>\n</blockquote>\n<p>I've seen a number of notebooks using time as a feature and am wondering if anyone has gleaned any insights that they'd be willing to share.  Do your models return reasonable results on the daily living data?  My hypothesis is no, but wanted to share nonetheless. </p>",
      "rawMarkdown": "Hey everyone -- I've been thinking about the time column in the data lately and wanted to share a few thoughts.  This competition scores users/teams on their prediction of FoG events in the tDCS and DeFOG datasets, which are comprised of accelerometer data during a (standard?) series of FoG-provoking tasks.  If folks take a *roughly* similar amount of time to complete the tasks and certain tasks elicit certain FoG events, then perhaps time is a reasonable feature to use when modeling the data in these datasets.  \n\nHowever, it seems to me that the ultimate goal or aspiration behind the competition is to identify (or even better predict) FoG events in *any* setting at *any* time. ** In this case, time would likely be a poor feature choice, even if the competition scoring (might?) incentivize the use of time as a feature due to the (standard?) series of tasks.  **\n\nHere's a lightly edited excerpt from the competition page for reference:\n>[The] Competition host ... aims to improve the personalized treatment of age-related movement, cognition, and mobility disorders and to alleviate the associated burden. They leverage a combination of clinical, engineering, and neuroscience expertise to: 1) Gain new understandings into the physiologic and pathophysiologic mechanisms that contribute to cognitive and motor function, the factors that influence these functions, and their changes with aging and disease ... 2) Develop new methods ... for the early detection and tracking of cognitive and motor decline. A major focus is on using leveraging wearable devices and digital technologies; and 3) Develop and evaluate novel methods for the prevention and treatment of gait, falls, and cognitive function. ... Your work will help advance the evaluation, understanding and treatment of FOG, improving the lives of the many people who suffer from this ... symptom.\n\nI've seen a number of notebooks using time as a feature and am wondering if anyone has gleaned any insights that they'd be willing to share.  Do your models return reasonable results on the daily living data?  My hypothesis is no, but wanted to share nonetheless.",
      "votes": null
    },
    {
      "id": "2216137",
      "postDate": "04/09/2023 19:53:56",
      "content": "<p>In my opinion, time itself is somewhat useless.<br>\nBut, using the time to build some features can be very useful.</p>\n<p>One approach is to take the lag values (time-based)</p>\n<p>But this competition is a unique type of time series, so you could also take future \"lag\" features (which worked for me).</p>",
      "rawMarkdown": "In my opinion, time itself is somewhat useless.\nBut, using the time to build some features can be very useful.\n\nOne approach is to take the lag values (time-based)\n\nBut this competition is a unique type of time series, so you could also take future \"lag\" features (which worked for me).",
      "votes": null
    },
    {
      "id": "2216715",
      "postDate": "04/10/2023 09:44:04",
      "content": "<p><a href=\"https://www.kaggle.com/austinhinkel\" target=\"_blank\">@austinhinkel</a> - great post. I've been thinking on this, myself.</p>\n<p><a href=\"https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets\" target=\"_blank\">This notebook</a> shows the target distributions across time, clearly illustrating that there's an advantage to using explicit time encodings as a feature. As you suggest, that's not at all helpful for generalisation to a useful deployment setting.</p>\n<p>I think this is another neat example of the King Midas problem in data science, and how it's not just about setting an objective function, but also the framing of the problem, and the data that evaluation will be conducted on.</p>\n<p>That's not to say that this is a meaningless competition, it's just likely to dramatically overestimate real-world usefulness. If instead it  were evaluated on the daily living dataset, then the problem would probably be much harder, require considerably more labelling effort, but also be much more useful.</p>",
      "rawMarkdown": "austinhinkel - great post. I've been thinking on this, myself.\n\n[This notebook](https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets) shows the target distributions across time, clearly illustrating that there's an advantage to using explicit time encodings as a feature. As you suggest, that's not at all helpful for generalisation to a useful deployment setting.\n\nI think this is another neat example of the King Midas problem in data science, and how it's not just about setting an objective function, but also the framing of the problem, and the data that evaluation will be conducted on.\n\nThat's not to say that this is a meaningless competition, it's just likely to dramatically overestimate real-world usefulness. If instead it  were evaluated on the daily living dataset, then the problem would probably be much harder, require considerably more labelling effort, but also be much more useful.",
      "votes": null
    },
    {
      "id": "2221661",
      "postDate": "04/14/2023 12:56:41",
      "content": "<p>For sure it helps predictions because training and testing timeseries are not random chunks.<br>\nAnyway, I see it as unethical and I won't use it. <br>\nIn my opinion the host should forbid the use of that feature for submissions.</p>",
      "rawMarkdown": "For sure it helps predictions because training and testing timeseries are not random chunks.\nAnyway, I see it as unethical and I won't use it. \nIn my opinion the host should forbid the use of that feature for submissions.",
      "votes": null
    },
    {
      "id": "2226062",
      "postDate": "04/18/2023 16:09:56",
      "content": "<p>Hey Alberto Annoni,</p>\n<p>I suppose it depends on the context.   If the models we develop are used to analyze data taken during the standard set of tasks, then perhaps it is still useful for that context.</p>\n<p>That said, many of us are likely using moving window analyses of some type in at least some subset of our feature engineering efforts.  The moving window approach can implicitly build in a similar dependence on the relative time, as times near the beginning and the end are often padded with (e.g.) zeros.  </p>",
      "rawMarkdown": "Hey Alberto Annoni,\n\nI suppose it depends on the context.   If the models we develop are used to analyze data taken during the standard set of tasks, then perhaps it is still useful for that context.\n\nThat said, many of us are likely using moving window analyses of some type in at least some subset of our feature engineering efforts.  The moving window approach can implicitly build in a similar dependence on the relative time, as times near the beginning and the end are often padded with (e.g.) zeros.",
      "votes": null
    },
    {
      "id": "2292420",
      "postDate": "06/08/2023 10:34:35",
      "content": "<p>To fix it, organizators could use some sliced amounts of timeline dataset, not full length of a whole exercise, or real world data (I don't think it is so expensive to label, it would significantly improve possible solutions). But it is another question - can anyone use winner models in real life as they are.</p>",
      "rawMarkdown": "To fix it, organizators could use some sliced amounts of timeline dataset, not full length of a whole exercise, or real world data (I don't think it is so expensive to label, it would significantly improve possible solutions). But it is another question - can anyone use winner models in real life as they are.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2216137,
      "author_name": "avivlevi815",
      "author_url": "",
      "post_date": "04/09/2023 19:53:56",
      "content": "<p>In my opinion, time itself is somewhat useless.<br>\nBut, using the time to build some features can be very useful.</p>\n<p>One approach is to take the lag values (time-based)</p>\n<p>But this competition is a unique type of time series, so you could also take future \"lag\" features (which worked for me).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2216715,
      "author_name": "jagofc",
      "author_url": "",
      "post_date": "04/10/2023 09:44:04",
      "content": "<p><a href=\"https://www.kaggle.com/austinhinkel\" target=\"_blank\">@austinhinkel</a> - great post. I've been thinking on this, myself.</p>\n<p><a href=\"https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets\" target=\"_blank\">This notebook</a> shows the target distributions across time, clearly illustrating that there's an advantage to using explicit time encodings as a feature. As you suggest, that's not at all helpful for generalisation to a useful deployment setting.</p>\n<p>I think this is another neat example of the King Midas problem in data science, and how it's not just about setting an objective function, but also the framing of the problem, and the data that evaluation will be conducted on.</p>\n<p>That's not to say that this is a meaningless competition, it's just likely to dramatically overestimate real-world usefulness. If instead it  were evaluated on the daily living dataset, then the problem would probably be much harder, require considerably more labelling effort, but also be much more useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2221661,
      "author_name": "albertoannoni",
      "author_url": "",
      "post_date": "04/14/2023 12:56:41",
      "content": "<p>For sure it helps predictions because training and testing timeseries are not random chunks.<br>\nAnyway, I see it as unethical and I won't use it. <br>\nIn my opinion the host should forbid the use of that feature for submissions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2226062,
          "author_name": "austinhinkel",
          "author_url": "",
          "post_date": "04/18/2023 16:09:56",
          "content": "<p>Hey Alberto Annoni,</p>\n<p>I suppose it depends on the context.   If the models we develop are used to analyze data taken during the standard set of tasks, then perhaps it is still useful for that context.</p>\n<p>That said, many of us are likely using moving window analyses of some type in at least some subset of our feature engineering efforts.  The moving window approach can implicitly build in a similar dependence on the relative time, as times near the beginning and the end are often padded with (e.g.) zeros.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2292420,
      "author_name": "sdelnikovaleksandr",
      "author_url": "",
      "post_date": "06/08/2023 10:34:35",
      "content": "<p>To fix it, organizators could use some sliced amounts of timeline dataset, not full length of a whole exercise, or real world data (I don't think it is so expensive to label, it would significantly improve possible solutions). But it is another question - can anyone use winner models in real life as they are.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2213613": "Hey everyone -- I've been thinking about the time column in the data lately and wanted to share a few thoughts.  This competition scores users/teams on their prediction of FoG events in the tDCS and DeFOG datasets, which are comprised of accelerometer data during a (standard?) series of FoG-provoking tasks.  If folks take a *roughly* similar amount of time to complete the tasks and certain tasks elicit certain FoG events, then perhaps time is a reasonable feature to use when modeling the data in these datasets.  \n\nHowever, it seems to me that the ultimate goal or aspiration behind the competition is to identify (or even better predict) FoG events in *any* setting at *any* time. ** In this case, time would likely be a poor feature choice, even if the competition scoring (might?) incentivize the use of time as a feature due to the (standard?) series of tasks.  **\n\nHere's a lightly edited excerpt from the competition page for reference:\n>[The] Competition host ... aims to improve the personalized treatment of age-related movement, cognition, and mobility disorders and to alleviate the associated burden. They leverage a combination of clinical, engineering, and neuroscience expertise to: 1) Gain new understandings into the physiologic and pathophysiologic mechanisms that contribute to cognitive and motor function, the factors that influence these functions, and their changes with aging and disease ... 2) Develop new methods ... for the early detection and tracking of cognitive and motor decline. A major focus is on using leveraging wearable devices and digital technologies; and 3) Develop and evaluate novel methods for the prevention and treatment of gait, falls, and cognitive function. ... Your work will help advance the evaluation, understanding and treatment of FOG, improving the lives of the many people who suffer from this ... symptom.\n\nI've seen a number of notebooks using time as a feature and am wondering if anyone has gleaned any insights that they'd be willing to share.  Do your models return reasonable results on the daily living data?  My hypothesis is no, but wanted to share nonetheless.",
    "2216137": "In my opinion, time itself is somewhat useless.\nBut, using the time to build some features can be very useful.\n\nOne approach is to take the lag values (time-based)\n\nBut this competition is a unique type of time series, so you could also take future \"lag\" features (which worked for me).",
    "2216715": "austinhinkel - great post. I've been thinking on this, myself.\n\n[This notebook](https://www.kaggle.com/code/xzj19013742/simple-eda-on-time-for-targets) shows the target distributions across time, clearly illustrating that there's an advantage to using explicit time encodings as a feature. As you suggest, that's not at all helpful for generalisation to a useful deployment setting.\n\nI think this is another neat example of the King Midas problem in data science, and how it's not just about setting an objective function, but also the framing of the problem, and the data that evaluation will be conducted on.\n\nThat's not to say that this is a meaningless competition, it's just likely to dramatically overestimate real-world usefulness. If instead it  were evaluated on the daily living dataset, then the problem would probably be much harder, require considerably more labelling effort, but also be much more useful.",
    "2221661": "For sure it helps predictions because training and testing timeseries are not random chunks.\nAnyway, I see it as unethical and I won't use it. \nIn my opinion the host should forbid the use of that feature for submissions.",
    "2226062": "Hey Alberto Annoni,\n\nI suppose it depends on the context.   If the models we develop are used to analyze data taken during the standard set of tasks, then perhaps it is still useful for that context.\n\nThat said, many of us are likely using moving window analyses of some type in at least some subset of our feature engineering efforts.  The moving window approach can implicitly build in a similar dependence on the relative time, as times near the beginning and the end are often padded with (e.g.) zeros.",
    "2292420": "To fix it, organizators could use some sliced amounts of timeline dataset, not full length of a whole exercise, or real world data (I don't think it is so expensive to label, it would significantly improve possible solutions). But it is another question - can anyone use winner models in real life as they are."
  },
  "source": "meta"
}