{
  "id": 403009,
  "title": "Over fitting and daily living data",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/403009",
  "author_name": "",
  "post_date": "2023-04-20T16:32:42.154439Z",
  "votes": 13,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Great to see the results to date. Curious if any teams are using the 24/7 daily living data; using it might help to avoid over fitting and to enhance generalizability. </p>\n<p>While labels for FOG events do not exist for the daily living data, there are some ways of evaluating if the model results applied to that data set \"make sense.\"<br>\nFor example, among non-freezers, the number of FOG events and the total FOG duration (eg over the week) should be 0 or close to that and much larger in the freezers. </p>\n<p>it would also be interesting to see if total FOG duration (summed across the week, for example) differs between those with and without FOG and how big the effect size is<br>\ne.g., Hedges's g (<a href=\"https://rowannicholls.github.io/python/statistics/effect_size.html#hedgess-g-1\" target=\"_blank\">https://rowannicholls.github.io/python/statistics/effect_size.html#hedgess-g-1</a>)<br>\nThanks!  Jeff</p>",
  "messages": [
    {
      "id": "2228601",
      "postDate": "04/20/2023 16:32:42",
      "content": "<p>Great to see the results to date. Curious if any teams are using the 24/7 daily living data; using it might help to avoid over fitting and to enhance generalizability. </p>\n<p>While labels for FOG events do not exist for the daily living data, there are some ways of evaluating if the model results applied to that data set \"make sense.\"<br>\nFor example, among non-freezers, the number of FOG events and the total FOG duration (eg over the week) should be 0 or close to that and much larger in the freezers. </p>\n<p>it would also be interesting to see if total FOG duration (summed across the week, for example) differs between those with and without FOG and how big the effect size is<br>\ne.g., Hedges's g (<a href=\"https://rowannicholls.github.io/python/statistics/effect_size.html#hedgess-g-1\" target=\"_blank\">https://rowannicholls.github.io/python/statistics/effect_size.html#hedgess-g-1</a>)<br>\nThanks!  Jeff</p>",
      "rawMarkdown": "Great to see the results to date. Curious if any teams are using the 24/7 daily living data; using it might help to avoid over fitting and to enhance generalizability. \n\nWhile labels for FOG events do not exist for the daily living data, there are some ways of evaluating if the model results applied to that data set \"make sense.\"\nFor example, among non-freezers, the number of FOG events and the total FOG duration (eg over the week) should be 0 or close to that and much larger in the freezers. \n\nit would also be interesting to see if total FOG duration (summed across the week, for example) differs between those with and without FOG and how big the effect size is\ne.g., Hedges's g (https://rowannicholls.github.io/python/statistics/effect_size.html#hedgess-g-1)\nThanks!  Jeff",
      "votes": null
    },
    {
      "id": "2231755",
      "postDate": "04/23/2023 16:13:32",
      "content": "<p>Interesting post <a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> - is predicting on the daily living dataset one of the ways you'd be hoping to test/use successful models?</p>\n<p>There's been some discussion <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/400311\" target=\"_blank\">in this thread</a> about the legitimacy of explicit time features. If we use these in modelling then performance on the defog/tdcsfog datasets will likely be better, but generalisation to less contrived situations (such as daily living) will likely suffer.</p>\n<p><a href=\"https://www.kaggle.com/code/lingduowang/feature-importance-visualize-in-lgbm\" target=\"_blank\">This notebook</a> by <a href=\"https://www.kaggle.com/lingduowang\" target=\"_blank\">@lingduowang</a> shows the feature importance of one of the better performing public models. <code>Time_frac</code> - an explicit time feature - dominates the feature importance.</p>\n<p>Any thoughts on this?</p>",
      "rawMarkdown": "Interesting post @jeffhausdorff - is predicting on the daily living dataset one of the ways you'd be hoping to test/use successful models?\n\nThere's been some discussion [in this thread](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/400311) about the legitimacy of explicit time features. If we use these in modelling then performance on the defog/tdcsfog datasets will likely be better, but generalisation to less contrived situations (such as daily living) will likely suffer.\n\n[This notebook](https://www.kaggle.com/code/lingduowang/feature-importance-visualize-in-lgbm) by @lingduowang shows the feature importance of one of the better performing public models. `Time_frac` - an explicit time feature - dominates the feature importance.\n\nAny thoughts on this?",
      "votes": null
    },
    {
      "id": "2234098",
      "postDate": "04/24/2023 21:50:40",
      "content": "<p>Good points.   Yes, a long-term goal is to predict on daily living datasets.  That would be very useful and powerful. The more generalizable and robust the models are based on tDCS and DeFOG, the more likely they will do well on other data, collected in other patients with FOG (and more likely to do well on the test data here).  And more general approaches and models are more likely to be successfully applied to daily-living data.  We included tDCS and DeFOG data - two similar, but different FOG-provoking protocols, to try to capture a broad range of freezing (among a relatively broad range of patients).</p>",
      "rawMarkdown": "Good points.   Yes, a long-term goal is to predict on daily living datasets.  That would be very useful and powerful. The more generalizable and robust the models are based on tDCS and DeFOG, the more likely they will do well on other data, collected in other patients with FOG (and more likely to do well on the test data here).  And more general approaches and models are more likely to be successfully applied to daily-living data.  We included tDCS and DeFOG data - two similar, but different FOG-provoking protocols, to try to capture a broad range of freezing (among a relatively broad range of patients).",
      "votes": null
    },
    {
      "id": "2238153",
      "postDate": "04/28/2023 08:48:58",
      "content": "<p>I don't quite see the usefulness of the variable \"Time_frac\" for a real model. From what I see in the code, it is the most representative variable of the model, but it doesn't make sense in the real life. The episodes cannot be a function of time as they can happen any time. In the episodes for model training, it is clear that they can be related since specific episodes are chosen, the data is cut, and the episode is left within this data (and it has a sort of time relation). But if we want to obtain a model for real life continuous use, \"Time_frac\" should not be in the equation since an episode will never depend on time. I understand that it may be important during model training and for scoring, but ideally, we should obtain models independent of time.</p>",
      "rawMarkdown": "I don't quite see the usefulness of the variable \"Time_frac\" for a real model. From what I see in the code, it is the most representative variable of the model, but it doesn't make sense in the real life. The episodes cannot be a function of time as they can happen any time. In the episodes for model training, it is clear that they can be related since specific episodes are chosen, the data is cut, and the episode is left within this data (and it has a sort of time relation). But if we want to obtain a model for real life continuous use, \"Time_frac\" should not be in the equation since an episode will never depend on time. I understand that it may be important during model training and for scoring, but ideally, we should obtain models independent of time.",
      "votes": null
    },
    {
      "id": "2238335",
      "postDate": "04/28/2023 12:34:16",
      "content": "<p>I fully agree. Using Time_frac or similar approaches likely will not lead to a robust, generalizable solution. Also, I may be wrong, but my guess is that even for this specific competition, Time_frac and other uses of time along those lines might help to get in the general area of a FOG event, but it won't be very accurate or precise.  </p>",
      "rawMarkdown": "I fully agree. Using Time_frac or similar approaches likely will not lead to a robust, generalizable solution. Also, I may be wrong, but my guess is that even for this specific competition, Time_frac and other uses of time along those lines might help to get in the general area of a FOG event, but it won't be very accurate or precise.",
      "votes": null
    },
    {
      "id": "2240503",
      "postDate": "04/30/2023 15:37:34",
      "content": "<p>It seems to me that the only way to fight this would be if the time-series were completely random splits of data, also with already started or unfinished events: this way the time passed since the start of the time-series would give no information on the targets. </p>\n<p>I'm surprised the host is not much worried about this problem.</p>",
      "rawMarkdown": "It seems to me that the only way to fight this would be if the time-series were completely random splits of data, also with already started or unfinished events: this way the time passed since the start of the time-series would give no information on the targets. \n\nI'm surprised the host is not much worried about this problem.",
      "votes": null
    },
    {
      "id": "2242327",
      "postDate": "05/02/2023 07:31:34",
      "content": "<p>This is a concern. However, we anticipate that the best models will be generalizable, will need to be trained on a relatively larger dataset, and will work similarly on all datasets. In that case, too much reliance on the structure of a specific dataset structure won't be so helpful. A model that trains on all datasets will see a richer and more fully representative sample of FOG episodes; freezing can be highly variable across and within subjects. </p>",
      "rawMarkdown": "This is a concern. However, we anticipate that the best models will be generalizable, will need to be trained on a relatively larger dataset, and will work similarly on all datasets. In that case, too much reliance on the structure of a specific dataset structure won't be so helpful. A model that trains on all datasets will see a richer and more fully representative sample of FOG episodes; freezing can be highly variable across and within subjects.",
      "votes": null
    },
    {
      "id": "2242419",
      "postDate": "05/02/2023 09:05:34",
      "content": "<p>Are the time-series where you are checking the generalizability built in the same way the ones for training are, i.e. the events are generally in the middle of the time-series?</p>\n<p>If you already accounted for that sorry for this further message.</p>",
      "rawMarkdown": "Are the time-series where you are checking the generalizability built in the same way the ones for training are, i.e. the events are generally in the middle of the time-series?\n\nIf you already accounted for that sorry for this further message.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2231755,
      "author_name": "jagofc",
      "author_url": "",
      "post_date": "04/23/2023 16:13:32",
      "content": "<p>Interesting post <a href=\"https://www.kaggle.com/jeffhausdorff\" target=\"_blank\">@jeffhausdorff</a> - is predicting on the daily living dataset one of the ways you'd be hoping to test/use successful models?</p>\n<p>There's been some discussion <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/400311\" target=\"_blank\">in this thread</a> about the legitimacy of explicit time features. If we use these in modelling then performance on the defog/tdcsfog datasets will likely be better, but generalisation to less contrived situations (such as daily living) will likely suffer.</p>\n<p><a href=\"https://www.kaggle.com/code/lingduowang/feature-importance-visualize-in-lgbm\" target=\"_blank\">This notebook</a> by <a href=\"https://www.kaggle.com/lingduowang\" target=\"_blank\">@lingduowang</a> shows the feature importance of one of the better performing public models. <code>Time_frac</code> - an explicit time feature - dominates the feature importance.</p>\n<p>Any thoughts on this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2234098,
          "author_name": "jeffhausdorff",
          "author_url": "",
          "post_date": "04/24/2023 21:50:40",
          "content": "<p>Good points.   Yes, a long-term goal is to predict on daily living datasets.  That would be very useful and powerful. The more generalizable and robust the models are based on tDCS and DeFOG, the more likely they will do well on other data, collected in other patients with FOG (and more likely to do well on the test data here).  And more general approaches and models are more likely to be successfully applied to daily-living data.  We included tDCS and DeFOG data - two similar, but different FOG-provoking protocols, to try to capture a broad range of freezing (among a relatively broad range of patients).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2238153,
          "author_name": "quimquadrada",
          "author_url": "",
          "post_date": "04/28/2023 08:48:58",
          "content": "<p>I don't quite see the usefulness of the variable \"Time_frac\" for a real model. From what I see in the code, it is the most representative variable of the model, but it doesn't make sense in the real life. The episodes cannot be a function of time as they can happen any time. In the episodes for model training, it is clear that they can be related since specific episodes are chosen, the data is cut, and the episode is left within this data (and it has a sort of time relation). But if we want to obtain a model for real life continuous use, \"Time_frac\" should not be in the equation since an episode will never depend on time. I understand that it may be important during model training and for scoring, but ideally, we should obtain models independent of time.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2238335,
              "author_name": "jeffhausdorff",
              "author_url": "",
              "post_date": "04/28/2023 12:34:16",
              "content": "<p>I fully agree. Using Time_frac or similar approaches likely will not lead to a robust, generalizable solution. Also, I may be wrong, but my guess is that even for this specific competition, Time_frac and other uses of time along those lines might help to get in the general area of a FOG event, but it won't be very accurate or precise.  </p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 2240503,
              "author_name": "albertoannoni",
              "author_url": "",
              "post_date": "04/30/2023 15:37:34",
              "content": "<p>It seems to me that the only way to fight this would be if the time-series were completely random splits of data, also with already started or unfinished events: this way the time passed since the start of the time-series would give no information on the targets. </p>\n<p>I'm surprised the host is not much worried about this problem.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2242327,
                  "author_name": "jeffhausdorff",
                  "author_url": "",
                  "post_date": "05/02/2023 07:31:34",
                  "content": "<p>This is a concern. However, we anticipate that the best models will be generalizable, will need to be trained on a relatively larger dataset, and will work similarly on all datasets. In that case, too much reliance on the structure of a specific dataset structure won't be so helpful. A model that trains on all datasets will see a richer and more fully representative sample of FOG episodes; freezing can be highly variable across and within subjects. </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2242419,
                      "author_name": "albertoannoni",
                      "author_url": "",
                      "post_date": "05/02/2023 09:05:34",
                      "content": "<p>Are the time-series where you are checking the generalizability built in the same way the ones for training are, i.e. the events are generally in the middle of the time-series?</p>\n<p>If you already accounted for that sorry for this further message.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2228601": "Great to see the results to date. Curious if any teams are using the 24/7 daily living data; using it might help to avoid over fitting and to enhance generalizability. \n\nWhile labels for FOG events do not exist for the daily living data, there are some ways of evaluating if the model results applied to that data set \"make sense.\"\nFor example, among non-freezers, the number of FOG events and the total FOG duration (eg over the week) should be 0 or close to that and much larger in the freezers. \n\nit would also be interesting to see if total FOG duration (summed across the week, for example) differs between those with and without FOG and how big the effect size is\ne.g., Hedges's g (https://rowannicholls.github.io/python/statistics/effect_size.html#hedgess-g-1)\nThanks!  Jeff",
    "2231755": "Interesting post @jeffhausdorff - is predicting on the daily living dataset one of the ways you'd be hoping to test/use successful models?\n\nThere's been some discussion [in this thread](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/400311) about the legitimacy of explicit time features. If we use these in modelling then performance on the defog/tdcsfog datasets will likely be better, but generalisation to less contrived situations (such as daily living) will likely suffer.\n\n[This notebook](https://www.kaggle.com/code/lingduowang/feature-importance-visualize-in-lgbm) by @lingduowang shows the feature importance of one of the better performing public models. `Time_frac` - an explicit time feature - dominates the feature importance.\n\nAny thoughts on this?",
    "2234098": "Good points.   Yes, a long-term goal is to predict on daily living datasets.  That would be very useful and powerful. The more generalizable and robust the models are based on tDCS and DeFOG, the more likely they will do well on other data, collected in other patients with FOG (and more likely to do well on the test data here).  And more general approaches and models are more likely to be successfully applied to daily-living data.  We included tDCS and DeFOG data - two similar, but different FOG-provoking protocols, to try to capture a broad range of freezing (among a relatively broad range of patients).",
    "2238153": "I don't quite see the usefulness of the variable \"Time_frac\" for a real model. From what I see in the code, it is the most representative variable of the model, but it doesn't make sense in the real life. The episodes cannot be a function of time as they can happen any time. In the episodes for model training, it is clear that they can be related since specific episodes are chosen, the data is cut, and the episode is left within this data (and it has a sort of time relation). But if we want to obtain a model for real life continuous use, \"Time_frac\" should not be in the equation since an episode will never depend on time. I understand that it may be important during model training and for scoring, but ideally, we should obtain models independent of time.",
    "2238335": "I fully agree. Using Time_frac or similar approaches likely will not lead to a robust, generalizable solution. Also, I may be wrong, but my guess is that even for this specific competition, Time_frac and other uses of time along those lines might help to get in the general area of a FOG event, but it won't be very accurate or precise.",
    "2240503": "It seems to me that the only way to fight this would be if the time-series were completely random splits of data, also with already started or unfinished events: this way the time passed since the start of the time-series would give no information on the targets. \n\nI'm surprised the host is not much worried about this problem.",
    "2242327": "This is a concern. However, we anticipate that the best models will be generalizable, will need to be trained on a relatively larger dataset, and will work similarly on all datasets. In that case, too much reliance on the structure of a specific dataset structure won't be so helpful. A model that trains on all datasets will see a richer and more fully representative sample of FOG episodes; freezing can be highly variable across and within subjects.",
    "2242419": "Are the time-series where you are checking the generalizability built in the same way the ones for training are, i.e. the events are generally in the middle of the time-series?\n\nIf you already accounted for that sorry for this further message."
  },
  "source": "meta"
}