{
  "id": 399764,
  "title": "Handling dataset ",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/399764",
  "author_name": "",
  "post_date": "2023-04-05T13:05:01.844557500Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>How are you approaching the scale of this dataset? So far I have thought of a few approaches, but I've only had time to test downsampling - seems to lead to better convergence than naively training for some cases. I can't verify yet if it results in better lb/local correlation, but convergence time is definitely better. </p>",
  "messages": [
    {
      "id": "2210535",
      "postDate": "04/05/2023 13:05:01",
      "content": "<p>How are you approaching the scale of this dataset? So far I have thought of a few approaches, but I've only had time to test downsampling - seems to lead to better convergence than naively training for some cases. I can't verify yet if it results in better lb/local correlation, but convergence time is definitely better. </p>",
      "rawMarkdown": "How are you approaching the scale of this dataset? So far I have thought of a few approaches, but I've only had time to test downsampling - seems to lead to better convergence than naively training for some cases. I can't verify yet if it results in better lb/local correlation, but convergence time is definitely better.",
      "votes": null
    },
    {
      "id": "2211359",
      "postDate": "04/06/2023 01:40:18",
      "content": "<p>I haven't had any need to downsample or do anything particularly special to cope with the scale of the training dataset. I just process the training data in big batches on the GPU. Training times are perfectly tolerable, around 35 minutes on an RTX 2080 Ti to produce a model that ranks very well on the public LB.</p>\n<p>Downsampling is an interesting idea. I would worry about it crippling the model's ability to detect high frequency tremors, but <a href=\"https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores\" target=\"_blank\">https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores</a> seems to indicate that the highest frequency characteristics of FOG events are in the ballpark of 8 Hz, so I imagine the 100+ Hz sample rates in the raw data are probably excessive. You could probably get away with downsampling quite a bit without causing major issues. Thanks for sharing your observation that it seems to lead to better convergence sometimes… I wonder if maybe it could help to prevent overfitting by eliminating super high frequency noise that's useless for FOG detection? I added that to my list of things to investigate :-)</p>",
      "rawMarkdown": "I haven't had any need to downsample or do anything particularly special to cope with the scale of the training dataset. I just process the training data in big batches on the GPU. Training times are perfectly tolerable, around 35 minutes on an RTX 2080 Ti to produce a model that ranks very well on the public LB.\n\nDownsampling is an interesting idea. I would worry about it crippling the model's ability to detect high frequency tremors, but [https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores](https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores) seems to indicate that the highest frequency characteristics of FOG events are in the ballpark of 8 Hz, so I imagine the 100+ Hz sample rates in the raw data are probably excessive. You could probably get away with downsampling quite a bit without causing major issues. Thanks for sharing your observation that it seems to lead to better convergence sometimes... I wonder if maybe it could help to prevent overfitting by eliminating super high frequency noise that's useless for FOG detection? I added that to my list of things to investigate :-)",
      "votes": null
    },
    {
      "id": "2211541",
      "postDate": "04/06/2023 05:31:34",
      "content": "<p>I agree that 100+ Hz sampling is probably excessive -- typical human movement is on the order of a few Hz, while muscle spasms might top a couple 10's of Hz (<a href=\"https://www.sciencedirect.com/topics/medicine-and-dentistry/muscle-action-potential\" target=\"_blank\">Salvage et al.</a>).  Doing anything 100 times a second seems inhuman. </p>\n<p>That said, I've played with <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/395143#2188535\" target=\"_blank\">low pass filtering</a> the TDCS data and the frequency domain features I tested it on did not benefit from changing the cutoff frequency between 6 and 20 Hz.  As for the time domain features I tested, low pass filtering did seem to help a bit for a feature centered on AP movement, but I've only tried a 10 Hz cutoff so far.</p>\n<p>These results may be specific to my model/subset of features, but thought I'd share nonetheless.  Please let me know if you find something different! </p>",
      "rawMarkdown": "I agree that 100+ Hz sampling is probably excessive -- typical human movement is on the order of a few Hz, while muscle spasms might top a couple 10's of Hz ([Salvage et al.](https://www.sciencedirect.com/topics/medicine-and-dentistry/muscle-action-potential)).  Doing anything 100 times a second seems inhuman. \n\nThat said, I've played with [low pass filtering](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/395143#2188535) the TDCS data and the frequency domain features I tested it on did not benefit from changing the cutoff frequency between 6 and 20 Hz.  As for the time domain features I tested, low pass filtering did seem to help a bit for a feature centered on AP movement, but I've only tried a 10 Hz cutoff so far.\n\nThese results may be specific to my model/subset of features, but thought I'd share nonetheless.  Please let me know if you find something different!",
      "votes": null
    },
    {
      "id": "2211543",
      "postDate": "04/06/2023 05:32:06",
      "content": "<p>I see. I'm not going with deep learning approaches at the moment, but I think I'll try those eventually. That being said, how are you approaching stride/window for this? I figured setting stride would be helpful to act as a downsampling counterpart. </p>",
      "rawMarkdown": "I see. I'm not going with deep learning approaches at the moment, but I think I'll try those eventually. That being said, how are you approaching stride/window for this? I figured setting stride would be helpful to act as a downsampling counterpart.",
      "votes": null
    },
    {
      "id": "2211969",
      "postDate": "04/06/2023 11:59:26",
      "content": "<p>\"Doing anything 100 times a second seems inhuman\" - One thing to watch out for is that the max frequency that can be accurately resolved without aliasing artifacts is half the sample rate. So, the sample rate of the accelerometers used to collect the raw data is probably less excessive than one might intuitively expect. See <a href=\"https://en.wikipedia.org/wiki/Nyquist_frequency\" target=\"_blank\">https://en.wikipedia.org/wiki/Nyquist_frequency</a> and <a href=\"https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem\" target=\"_blank\">https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem</a> for details.</p>\n<p>I just tried adding random low pass filtering as a form of data augmentation. Preliminary local cross validation results are not good, but not totally hopeless either. It <em>might</em> be possible to get this working well with additional hyperparameter tuning (adjusting more than just data augmentation).</p>",
      "rawMarkdown": "\"Doing anything 100 times a second seems inhuman\" - One thing to watch out for is that the max frequency that can be accurately resolved without aliasing artifacts is half the sample rate. So, the sample rate of the accelerometers used to collect the raw data is probably less excessive than one might intuitively expect. See [https://en.wikipedia.org/wiki/Nyquist_frequency](https://en.wikipedia.org/wiki/Nyquist_frequency) and [https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem](https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem) for details.\n\nI just tried adding random low pass filtering as a form of data augmentation. Preliminary local cross validation results are not good, but not totally hopeless either. It *might* be possible to get this working well with additional hyperparameter tuning (adjusting more than just data augmentation).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2211359,
      "author_name": "jsday96",
      "author_url": "",
      "post_date": "04/06/2023 01:40:18",
      "content": "<p>I haven't had any need to downsample or do anything particularly special to cope with the scale of the training dataset. I just process the training data in big batches on the GPU. Training times are perfectly tolerable, around 35 minutes on an RTX 2080 Ti to produce a model that ranks very well on the public LB.</p>\n<p>Downsampling is an interesting idea. I would worry about it crippling the model's ability to detect high frequency tremors, but <a href=\"https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores\" target=\"_blank\">https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores</a> seems to indicate that the highest frequency characteristics of FOG events are in the ballpark of 8 Hz, so I imagine the 100+ Hz sample rates in the raw data are probably excessive. You could probably get away with downsampling quite a bit without causing major issues. Thanks for sharing your observation that it seems to lead to better convergence sometimes… I wonder if maybe it could help to prevent overfitting by eliminating super high frequency noise that's useless for FOG detection? I added that to my list of things to investigate :-)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2211541,
          "author_name": "austinhinkel",
          "author_url": "",
          "post_date": "04/06/2023 05:31:34",
          "content": "<p>I agree that 100+ Hz sampling is probably excessive -- typical human movement is on the order of a few Hz, while muscle spasms might top a couple 10's of Hz (<a href=\"https://www.sciencedirect.com/topics/medicine-and-dentistry/muscle-action-potential\" target=\"_blank\">Salvage et al.</a>).  Doing anything 100 times a second seems inhuman. </p>\n<p>That said, I've played with <a href=\"https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/395143#2188535\" target=\"_blank\">low pass filtering</a> the TDCS data and the frequency domain features I tested it on did not benefit from changing the cutoff frequency between 6 and 20 Hz.  As for the time domain features I tested, low pass filtering did seem to help a bit for a feature centered on AP movement, but I've only tried a 10 Hz cutoff so far.</p>\n<p>These results may be specific to my model/subset of features, but thought I'd share nonetheless.  Please let me know if you find something different! </p>",
          "votes": null,
          "replies": [
            {
              "id": 2211969,
              "author_name": "jsday96",
              "author_url": "",
              "post_date": "04/06/2023 11:59:26",
              "content": "<p>\"Doing anything 100 times a second seems inhuman\" - One thing to watch out for is that the max frequency that can be accurately resolved without aliasing artifacts is half the sample rate. So, the sample rate of the accelerometers used to collect the raw data is probably less excessive than one might intuitively expect. See <a href=\"https://en.wikipedia.org/wiki/Nyquist_frequency\" target=\"_blank\">https://en.wikipedia.org/wiki/Nyquist_frequency</a> and <a href=\"https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem\" target=\"_blank\">https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem</a> for details.</p>\n<p>I just tried adding random low pass filtering as a form of data augmentation. Preliminary local cross validation results are not good, but not totally hopeless either. It <em>might</em> be possible to get this working well with additional hyperparameter tuning (adjusting more than just data augmentation).</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 2211543,
          "author_name": "phantnguyen",
          "author_url": "",
          "post_date": "04/06/2023 05:32:06",
          "content": "<p>I see. I'm not going with deep learning approaches at the moment, but I think I'll try those eventually. That being said, how are you approaching stride/window for this? I figured setting stride would be helpful to act as a downsampling counterpart. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2210535": "How are you approaching the scale of this dataset? So far I have thought of a few approaches, but I've only had time to test downsampling - seems to lead to better convergence than naively training for some cases. I can't verify yet if it results in better lb/local correlation, but convergence time is definitely better.",
    "2211359": "I haven't had any need to downsample or do anything particularly special to cope with the scale of the training dataset. I just process the training data in big batches on the GPU. Training times are perfectly tolerable, around 35 minutes on an RTX 2080 Ti to produce a model that ranks very well on the public LB.\n\nDownsampling is an interesting idea. I would worry about it crippling the model's ability to detect high frequency tremors, but [https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores](https://www.researchgate.net/publication/313454263_Freezing_of_Gait_Detection_in_Parkinson's_Disease_A_Subject-Independent_Detector_Using_Anomaly_Scores) seems to indicate that the highest frequency characteristics of FOG events are in the ballpark of 8 Hz, so I imagine the 100+ Hz sample rates in the raw data are probably excessive. You could probably get away with downsampling quite a bit without causing major issues. Thanks for sharing your observation that it seems to lead to better convergence sometimes... I wonder if maybe it could help to prevent overfitting by eliminating super high frequency noise that's useless for FOG detection? I added that to my list of things to investigate :-)",
    "2211541": "I agree that 100+ Hz sampling is probably excessive -- typical human movement is on the order of a few Hz, while muscle spasms might top a couple 10's of Hz ([Salvage et al.](https://www.sciencedirect.com/topics/medicine-and-dentistry/muscle-action-potential)).  Doing anything 100 times a second seems inhuman. \n\nThat said, I've played with [low pass filtering](https://www.kaggle.com/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/395143#2188535) the TDCS data and the frequency domain features I tested it on did not benefit from changing the cutoff frequency between 6 and 20 Hz.  As for the time domain features I tested, low pass filtering did seem to help a bit for a feature centered on AP movement, but I've only tried a 10 Hz cutoff so far.\n\nThese results may be specific to my model/subset of features, but thought I'd share nonetheless.  Please let me know if you find something different!",
    "2211543": "I see. I'm not going with deep learning approaches at the moment, but I think I'll try those eventually. That being said, how are you approaching stride/window for this? I figured setting stride would be helpful to act as a downsampling counterpart.",
    "2211969": "\"Doing anything 100 times a second seems inhuman\" - One thing to watch out for is that the max frequency that can be accurately resolved without aliasing artifacts is half the sample rate. So, the sample rate of the accelerometers used to collect the raw data is probably less excessive than one might intuitively expect. See [https://en.wikipedia.org/wiki/Nyquist_frequency](https://en.wikipedia.org/wiki/Nyquist_frequency) and [https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem](https://en.wikipedia.org/wiki/Nyquist%E2%80%93Shannon_sampling_theorem) for details.\n\nI just tried adding random low pass filtering as a form of data augmentation. Preliminary local cross validation results are not good, but not totally hopeless either. It *might* be possible to get this working well with additional hyperparameter tuning (adjusting more than just data augmentation)."
  },
  "source": "meta"
}