{
  "id": 416797,
  "title": "52nd Place Solution + Code",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/writeups/abandura-52nd-place-solution-code",
  "author_name": "",
  "post_date": "2023-06-13T03:21:25.805201900Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to the hosts for providing such an interesting/worthwhile problem to work on, and to everyone who shared their code and insights throughout the competition! This is my first completed Kaggle competition; I feel like I spent as much time learning about the Kaggle ecosystem as I did about the FOG data 😂 but overall it was a fun experience.</p>\n<p>My solution was a fairly straightforward 1D CNN ensemble, but I still wanted to share it!</p>\n<p>My code: <a href=\"https://www.kaggle.com/code/abandura/fog-1d-cnn\" target=\"_blank\">https://www.kaggle.com/code/abandura/fog-1d-cnn</a></p>\n<p>I referenced <a href=\"https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model\" target=\"_blank\">this</a> public notebook when I began the comp, and wanted to thank the author for sharing their work.</p>\n<h2>Solution Overview</h2>\n<h3>Preprocessing</h3>\n<p>TS window dimensions:</p>\n<pre><code>window_size = \nwindow_future = \nwindow_past = window_size - window_future\n</code></pre>\n<p>TS Features (standardized for each file):</p>\n<pre><code>Raw 3D accelerometer data\nSeperate 3D acc data into high freq  low freq components  MA \nSpectrogram of 3D accelerometer data (NFFT = )\nTemporal location of sequence**: time_index/total_time\n</code></pre>\n<p>**Some people feel that using temporal features reduces model usability, but I disagree. These models could be useful for auto-labeling more data from the same FOG-inducing protocols, in which case temporal features would be consistent to the data we used in this competition. On the flip side, I doubt these models would generalize well to something like the daily living dataset regardless of whether temporal features were leveraged. Just my two cents.</p>\n<p>CV strategy: 6 fold on patients, separating patients so each fold has roughly the same amount of data</p>\n<h3>Model Architecture</h3>\n<ul>\n<li>The best models based on CV for each fold were ensembled using a weighted average.</li>\n<li>There are batch normalization layers between each convolutional layer even though they aren't included in the diagram.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13474157%2F6746c1774c63e26d52986939d718f401%2FFOG%20CNN%20Architecture.jpeg?generation=1686624959127555&amp;alt=media\"></p>\n<h3>Training Parameters</h3>\n<ul>\n<li>batch_size = 1024</li>\n<li>learning_rate = 1e-4</li>\n<li>num_epochs = 1</li>\n<li>max_batches_per_epoch = 12000</li>\n<li>max_val_batches = 500</li>\n<li>val_freq = 500</li>\n<li>loss function:<br>\nFor tdcsfog I used cross entropy loss with inverse class frequency weights<br>\nFor defog I used focal loss with gamma = 5</li>\n</ul>\n<p><strong>Public LB</strong>: 0.362<br>\n<strong>Private LB</strong>: 0.305</p>\n<h4>Other Ideas I tried that didn't improve accuracy:</h4>\n<ul>\n<li>Combined models: single CNN model for both tdcsfog and defog after unifying sampling rates and units. </li>\n<li>Multi-head CNN model: Add a second loss function for presence/absence of FOG event so that the notype data could be leveraged.</li>\n<li>CNN features as inputs for a Light GBM model: Use the output of the convolutional blocks as input to a GBM model, along with other temporal and patient features.</li>\n</ul>",
  "messages": [
    {
      "id": "2300104",
      "postDate": "06/13/2023 03:21:25",
      "content": "<p>Thanks to the hosts for providing such an interesting/worthwhile problem to work on, and to everyone who shared their code and insights throughout the competition! This is my first completed Kaggle competition; I feel like I spent as much time learning about the Kaggle ecosystem as I did about the FOG data 😂 but overall it was a fun experience.</p>\n<p>My solution was a fairly straightforward 1D CNN ensemble, but I still wanted to share it!</p>\n<p>My code: <a href=\"https://www.kaggle.com/code/abandura/fog-1d-cnn\" target=\"_blank\">https://www.kaggle.com/code/abandura/fog-1d-cnn</a></p>\n<p>I referenced <a href=\"https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model\" target=\"_blank\">this</a> public notebook when I began the comp, and wanted to thank the author for sharing their work.</p>\n<h2>Solution Overview</h2>\n<h3>Preprocessing</h3>\n<p>TS window dimensions:</p>\n<pre><code>window_size = \nwindow_future = \nwindow_past = window_size - window_future\n</code></pre>\n<p>TS Features (standardized for each file):</p>\n<pre><code>Raw 3D accelerometer data\nSeperate 3D acc data into high freq  low freq components  MA \nSpectrogram of 3D accelerometer data (NFFT = )\nTemporal location of sequence**: time_index/total_time\n</code></pre>\n<p>**Some people feel that using temporal features reduces model usability, but I disagree. These models could be useful for auto-labeling more data from the same FOG-inducing protocols, in which case temporal features would be consistent to the data we used in this competition. On the flip side, I doubt these models would generalize well to something like the daily living dataset regardless of whether temporal features were leveraged. Just my two cents.</p>\n<p>CV strategy: 6 fold on patients, separating patients so each fold has roughly the same amount of data</p>\n<h3>Model Architecture</h3>\n<ul>\n<li>The best models based on CV for each fold were ensembled using a weighted average.</li>\n<li>There are batch normalization layers between each convolutional layer even though they aren't included in the diagram.</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13474157%2F6746c1774c63e26d52986939d718f401%2FFOG%20CNN%20Architecture.jpeg?generation=1686624959127555&amp;alt=media\"></p>\n<h3>Training Parameters</h3>\n<ul>\n<li>batch_size = 1024</li>\n<li>learning_rate = 1e-4</li>\n<li>num_epochs = 1</li>\n<li>max_batches_per_epoch = 12000</li>\n<li>max_val_batches = 500</li>\n<li>val_freq = 500</li>\n<li>loss function:<br>\nFor tdcsfog I used cross entropy loss with inverse class frequency weights<br>\nFor defog I used focal loss with gamma = 5</li>\n</ul>\n<p><strong>Public LB</strong>: 0.362<br>\n<strong>Private LB</strong>: 0.305</p>\n<h4>Other Ideas I tried that didn't improve accuracy:</h4>\n<ul>\n<li>Combined models: single CNN model for both tdcsfog and defog after unifying sampling rates and units. </li>\n<li>Multi-head CNN model: Add a second loss function for presence/absence of FOG event so that the notype data could be leveraged.</li>\n<li>CNN features as inputs for a Light GBM model: Use the output of the convolutional blocks as input to a GBM model, along with other temporal and patient features.</li>\n</ul>",
      "rawMarkdown": "Thanks to the hosts for providing such an interesting/worthwhile problem to work on, and to everyone who shared their code and insights throughout the competition! This is my first completed Kaggle competition; I feel like I spent as much time learning about the Kaggle ecosystem as I did about the FOG data 😂 but overall it was a fun experience.\n\nMy solution was a fairly straightforward 1D CNN ensemble, but I still wanted to share it!\n\nMy code: https://www.kaggle.com/code/abandura/fog-1d-cnn\n\nI referenced [this](https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model) public notebook when I began the comp, and wanted to thank the author for sharing their work.\n\n## Solution Overview\n\n### Preprocessing\n\nTS window dimensions:\n```python\nwindow_size = 200\nwindow_future = 75\nwindow_past = window_size - window_future\n```\n\nTS Features (standardized for each file):\n```python\nRaw 3D accelerometer data\nSeperate 3D acc data into high freq and low freq components with MA filter\nSpectrogram of 3D accelerometer data (NFFT = 8)\nTemporal location of sequence**: time_index/total_time\n```\n\t\n**Some people feel that using temporal features reduces model usability, but I disagree. These models could be useful for auto-labeling more data from the same FOG-inducing protocols, in which case temporal features would be consistent to the data we used in this competition. On the flip side, I doubt these models would generalize well to something like the daily living dataset regardless of whether temporal features were leveraged. Just my two cents.\n\nCV strategy: 6 fold on patients, separating patients so each fold has roughly the same amount of data\n\n###Model Architecture\n- The best models based on CV for each fold were ensembled using a weighted average.\n- There are batch normalization layers between each convolutional layer even though they aren't included in the diagram.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13474157%2F6746c1774c63e26d52986939d718f401%2FFOG%20CNN%20Architecture.jpeg?generation=1686624959127555&alt=media\" width=\"400\">\n\n###Training Parameters\n- batch_size = 1024\n- learning_rate = 1e-4\n- num_epochs = 1\n- max_batches_per_epoch = 12000\n- max_val_batches = 500\n- val_freq = 500\n- loss function:\n\tFor tdcsfog I used cross entropy loss with inverse class frequency weights\n\tFor defog I used focal loss with gamma = 5\n\n**Public LB**: 0.362\n**Private LB**: 0.305\n\n#### Other Ideas I tried that didn't improve accuracy:\n- Combined models: single CNN model for both tdcsfog and defog after unifying sampling rates and units. \n- Multi-head CNN model: Add a second loss function for presence/absence of FOG event so that the notype data could be leveraged.\n- CNN features as inputs for a Light GBM model: Use the output of the convolutional blocks as input to a GBM model, along with other temporal and patient features.",
      "votes": null
    },
    {
      "id": "2303298",
      "postDate": "06/15/2023 08:16:38",
      "content": "<p>Nice work! I noticed you didnt use GroupKFold, but manually defined folds at the patient level, which I think is the same? Did you observe any systematic differences between LB and CV?</p>",
      "rawMarkdown": "Nice work! I noticed you didnt use GroupKFold, but manually defined folds at the patient level, which I think is the same? Did you observe any systematic differences between LB and CV?",
      "votes": null
    },
    {
      "id": "2305393",
      "postDate": "06/16/2023 17:03:35",
      "content": "<p>Thanks! I think you're right that GroupKFold would be similar to the way I manually split up the folds--I did that so I'd have a little more control over the process. My CV tended to be lower than LB, but they were well correlated so still useful for selecting models. I think I saw discussions of other folks seeing a similar trend?</p>",
      "rawMarkdown": "Thanks! I think you're right that GroupKFold would be similar to the way I manually split up the folds--I did that so I'd have a little more control over the process. My CV tended to be lower than LB, but they were well correlated so still useful for selecting models. I think I saw discussions of other folks seeing a similar trend?",
      "votes": null
    },
    {
      "id": "2479910",
      "postDate": "10/13/2023 02:00:44",
      "content": "<p>thanks for sharing，really learned a lot. simple solution also leads to a good result!😃</p>",
      "rawMarkdown": "thanks for sharing，really learned a lot. simple solution also leads to a good result!😃",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2303298,
      "author_name": "exjustice",
      "author_url": "",
      "post_date": "06/15/2023 08:16:38",
      "content": "<p>Nice work! I noticed you didnt use GroupKFold, but manually defined folds at the patient level, which I think is the same? Did you observe any systematic differences between LB and CV?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2305393,
          "author_name": "abandura",
          "author_url": "",
          "post_date": "06/16/2023 17:03:35",
          "content": "<p>Thanks! I think you're right that GroupKFold would be similar to the way I manually split up the folds--I did that so I'd have a little more control over the process. My CV tended to be lower than LB, but they were well correlated so still useful for selecting models. I think I saw discussions of other folks seeing a similar trend?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2479910,
      "author_name": "roger92",
      "author_url": "",
      "post_date": "10/13/2023 02:00:44",
      "content": "<p>thanks for sharing，really learned a lot. simple solution also leads to a good result!😃</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2300104": "Thanks to the hosts for providing such an interesting/worthwhile problem to work on, and to everyone who shared their code and insights throughout the competition! This is my first completed Kaggle competition; I feel like I spent as much time learning about the Kaggle ecosystem as I did about the FOG data 😂 but overall it was a fun experience.\n\nMy solution was a fairly straightforward 1D CNN ensemble, but I still wanted to share it!\n\nMy code: https://www.kaggle.com/code/abandura/fog-1d-cnn\n\nI referenced [this](https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model) public notebook when I began the comp, and wanted to thank the author for sharing their work.\n\n## Solution Overview\n\n### Preprocessing\n\nTS window dimensions:\n```python\nwindow_size = 200\nwindow_future = 75\nwindow_past = window_size - window_future\n```\n\nTS Features (standardized for each file):\n```python\nRaw 3D accelerometer data\nSeperate 3D acc data into high freq and low freq components with MA filter\nSpectrogram of 3D accelerometer data (NFFT = 8)\nTemporal location of sequence**: time_index/total_time\n```\n\t\n**Some people feel that using temporal features reduces model usability, but I disagree. These models could be useful for auto-labeling more data from the same FOG-inducing protocols, in which case temporal features would be consistent to the data we used in this competition. On the flip side, I doubt these models would generalize well to something like the daily living dataset regardless of whether temporal features were leveraged. Just my two cents.\n\nCV strategy: 6 fold on patients, separating patients so each fold has roughly the same amount of data\n\n###Model Architecture\n- The best models based on CV for each fold were ensembled using a weighted average.\n- There are batch normalization layers between each convolutional layer even though they aren't included in the diagram.\n\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13474157%2F6746c1774c63e26d52986939d718f401%2FFOG%20CNN%20Architecture.jpeg?generation=1686624959127555&alt=media\" width=\"400\">\n\n###Training Parameters\n- batch_size = 1024\n- learning_rate = 1e-4\n- num_epochs = 1\n- max_batches_per_epoch = 12000\n- max_val_batches = 500\n- val_freq = 500\n- loss function:\n\tFor tdcsfog I used cross entropy loss with inverse class frequency weights\n\tFor defog I used focal loss with gamma = 5\n\n**Public LB**: 0.362\n**Private LB**: 0.305\n\n#### Other Ideas I tried that didn't improve accuracy:\n- Combined models: single CNN model for both tdcsfog and defog after unifying sampling rates and units. \n- Multi-head CNN model: Add a second loss function for presence/absence of FOG event so that the notype data could be leveraged.\n- CNN features as inputs for a Light GBM model: Use the output of the convolutional blocks as input to a GBM model, along with other temporal and patient features.",
    "2303298": "Nice work! I noticed you didnt use GroupKFold, but manually defined folds at the patient level, which I think is the same? Did you observe any systematic differences between LB and CV?",
    "2305393": "Thanks! I think you're right that GroupKFold would be similar to the way I manually split up the folds--I did that so I'd have a little more control over the process. My CV tended to be lower than LB, but they were well correlated so still useful for selecting models. I think I saw discussions of other folks seeing a similar trend?",
    "2479910": "thanks for sharing，really learned a lot. simple solution also leads to a good result!😃"
  },
  "source": "meta"
}