{
  "id": 416248,
  "title": "12th place solution: Simple features and LSTM.",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/writeups/mrt-12th-place-solution-simple-features-and-lstm",
  "author_name": "",
  "post_date": "2023-06-10T12:01:46.550Z",
  "votes": 14,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I would like to thank host and kaggle for organizing such interesting competition.<br>\nI believe that the method proposed by the winner and prizewinners who received high scores in both the public and private tests will be helpful to host.</p>\n<h2>Summary</h2>\n<ul>\n<li>Ensamble of 4 model based by LSTM</li>\n<li>Use 4 features. mean, std, max, min, median.</li>\n<li>Notype data is used as validation data.</li>\n<li>Selected model separately for tDCS FOG and DeFOG.</li>\n</ul>\n<h2>Data Preprocessing</h2>\n<p>I compressed all data into fixed length. Length is 2048 because shortest training data is 2359.</p>\n<p>Compression Method</p>\n<ol>\n<li>Expand the original data to a smallest multiple of the target length that exceeds the original data length. (ch, target_size * n)</li>\n<li>Fold the data into (ch, target_size, n).</li>\n</ol>\n<p>As feature, I used the mean, std, max, min, and median within this N.<br>\nIn addition to this, it was also useful to add each percentile point (15~90 percentile).<br>\nI rarely used the information on each ID and Subjects, but only \"Visit\" was used as a label for multitasking.</p>\n<pre><code> ():\n    ch = x.shape[] \n    input_size = x.shape[]\n\n    pad = target_size - input_size % target_size\n    factor = (input_size + pad) / input_size\n\n    x = np.array([ndi.zoom(xi, zoom=factor, mode=)  xi  x])\n    x = x.reshape((ch, target_size, -))\n\n    res = {} \n    res[] = np.mean(x, axis=).reshape(ch, -)\n    res[] = np.(x, axis=).reshape(ch, -)\n    res[] = np.(x, axis=).reshape(ch, -)\n    res[] = np.median(x, axis=).reshape(ch, -)\n    res[] = np.sqrt(np.var(x, axis=).reshape(ch, -))\n     use_percentile_feat:\n         p  [, , , , , ]:\n            res[] = np.percentile(x, [p], axis=).reshape(ch, -)\n\n     res\n</code></pre>\n<h2>Model</h2>\n<p>The model used was LSTM: a simple 4-block model and a model that were <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285330\" target=\"_blank\">3rd place of similar competitions in the past</a>. For DeFOG, I also used <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112\" target=\"_blank\">Transformer</a>.</p>\n<h2>CV Strategy</h2>\n<ul>\n<li><p>tDCS FOG<br>\nIn tDCS FOG, I used 4fold StratifiedGroupKfold with Event label and Subject as group.<br>\nHowever, CV was quite high for certain folds.<br>\nThis is probably due to the large number of examples of difficult classes such as StartHesitation and Walking appearing for a long time in Subejt \"2d57c2\", which reduced the percentage of False Positives.<br>\nTherefore, I had selected a model based on a score that excluded this Subject.<br>\nPersonally, I think that the reason why Shake occurred was also due to the existence of such a Subject in the Public data, and that many teams overfit to it.</p></li>\n<li><p>DeFOG<br>\nBecause DeFOG has few labeled data, I used all labeled data as training data.<br>\nSince there was a correlation between the loss of the three labels and the loss of the Event label only, I used Nototype as the validation data.<br>\nAlthough it was not possible to make a final selection, I trained on all data including tdcsfog, and the model selected based on the Notype data was the my best private score(<a href=\"https://www.kaggle.com/code/mrt0933/keras-lstm-trained-by-all-labeled-data\" target=\"_blank\">this notebook</a>).</p></li>\n</ul>\n<h2>Final submission</h2>\n<p>I selected 4 models for each of tdcsfog and defog and ensemble them.<br>\nFor tDCS FOG, I chose the one with the highest cv for each fold.<br>\nThe best public score had a high enough cv, so I selected this score as the final submission (I'm glad I didn't get caught in the shaking).</p>\n<ul>\n<li>tDCS FOG. CV:0.28</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>Model</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 (w/o \"2d57c2”)</td>\n<td>LSTM (multitasking)</td>\n<td>0.179</td>\n</tr>\n<tr>\n<td>1</td>\n<td>LSTM (1d CNN head)</td>\n<td>0.337</td>\n</tr>\n<tr>\n<td>2</td>\n<td>LSTM (add percentile features)</td>\n<td>0.278</td>\n</tr>\n<tr>\n<td>3</td>\n<td>LSTM (1d CNN head)</td>\n<td>0.293</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>DeFOG. CV: 0.381(Event)</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Model</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>LSTM</td>\n<td>0.292</td>\n</tr>\n<tr>\n<td>1</td>\n<td>LSTM (add percentile features)</td>\n<td>0.318</td>\n</tr>\n<tr>\n<td>2</td>\n<td>LSTM (long 4096)</td>\n<td>0.286</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Transformer</td>\n<td>0.286</td>\n</tr>\n</tbody>\n</table>\n<p>Public LB: 0.444<br>\nPrivate LB: 0.341</p>\n<p>Finally, I am happy to become a kaggle master with this gold medal.<br>\nHowever, although I got a gold medal, I am disappointed that I could not design a model that could capture FOG events well both quantitatively and qualitatively.<br>\nI would love to participate in the second and third competitions if they are held. Thank you very much.</p>",
  "messages": [
    {
      "id": "2294826",
      "postDate": "06/10/2023 11:11:44",
      "content": "<p>I would like to thank host and kaggle for organizing such interesting competition.<br>\nI believe that the method proposed by the winner and prizewinners who received high scores in both the public and private tests will be helpful to host.</p>\n<h2>Summary</h2>\n<ul>\n<li>Ensamble of 4 model based by LSTM</li>\n<li>Use 4 features. mean, std, max, min, median.</li>\n<li>Notype data is used as validation data.</li>\n<li>Selected model separately for tDCS FOG and DeFOG.</li>\n</ul>\n<h2>Data Preprocessing</h2>\n<p>I compressed all data into fixed length. Length is 2048 because shortest training data is 2359.</p>\n<p>Compression Method</p>\n<ol>\n<li>Expand the original data to a smallest multiple of the target length that exceeds the original data length. (ch, target_size * n)</li>\n<li>Fold the data into (ch, target_size, n).</li>\n</ol>\n<p>As feature, I used the mean, std, max, min, and median within this N.<br>\nIn addition to this, it was also useful to add each percentile point (15~90 percentile).<br>\nI rarely used the information on each ID and Subjects, but only \"Visit\" was used as a label for multitasking.</p>\n<pre><code> ():\n    ch = x.shape[] \n    input_size = x.shape[]\n\n    pad = target_size - input_size % target_size\n    factor = (input_size + pad) / input_size\n\n    x = np.array([ndi.zoom(xi, zoom=factor, mode=)  xi  x])\n    x = x.reshape((ch, target_size, -))\n\n    res = {} \n    res[] = np.mean(x, axis=).reshape(ch, -)\n    res[] = np.(x, axis=).reshape(ch, -)\n    res[] = np.(x, axis=).reshape(ch, -)\n    res[] = np.median(x, axis=).reshape(ch, -)\n    res[] = np.sqrt(np.var(x, axis=).reshape(ch, -))\n     use_percentile_feat:\n         p  [, , , , , ]:\n            res[] = np.percentile(x, [p], axis=).reshape(ch, -)\n\n     res\n</code></pre>\n<h2>Model</h2>\n<p>The model used was LSTM: a simple 4-block model and a model that were <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285330\" target=\"_blank\">3rd place of similar competitions in the past</a>. For DeFOG, I also used <a href=\"https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112\" target=\"_blank\">Transformer</a>.</p>\n<h2>CV Strategy</h2>\n<ul>\n<li><p>tDCS FOG<br>\nIn tDCS FOG, I used 4fold StratifiedGroupKfold with Event label and Subject as group.<br>\nHowever, CV was quite high for certain folds.<br>\nThis is probably due to the large number of examples of difficult classes such as StartHesitation and Walking appearing for a long time in Subejt \"2d57c2\", which reduced the percentage of False Positives.<br>\nTherefore, I had selected a model based on a score that excluded this Subject.<br>\nPersonally, I think that the reason why Shake occurred was also due to the existence of such a Subject in the Public data, and that many teams overfit to it.</p></li>\n<li><p>DeFOG<br>\nBecause DeFOG has few labeled data, I used all labeled data as training data.<br>\nSince there was a correlation between the loss of the three labels and the loss of the Event label only, I used Nototype as the validation data.<br>\nAlthough it was not possible to make a final selection, I trained on all data including tdcsfog, and the model selected based on the Notype data was the my best private score(<a href=\"https://www.kaggle.com/code/mrt0933/keras-lstm-trained-by-all-labeled-data\" target=\"_blank\">this notebook</a>).</p></li>\n</ul>\n<h2>Final submission</h2>\n<p>I selected 4 models for each of tdcsfog and defog and ensemble them.<br>\nFor tDCS FOG, I chose the one with the highest cv for each fold.<br>\nThe best public score had a high enough cv, so I selected this score as the final submission (I'm glad I didn't get caught in the shaking).</p>\n<ul>\n<li>tDCS FOG. CV:0.28</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>Fold</th>\n<th>Model</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0 (w/o \"2d57c2”)</td>\n<td>LSTM (multitasking)</td>\n<td>0.179</td>\n</tr>\n<tr>\n<td>1</td>\n<td>LSTM (1d CNN head)</td>\n<td>0.337</td>\n</tr>\n<tr>\n<td>2</td>\n<td>LSTM (add percentile features)</td>\n<td>0.278</td>\n</tr>\n<tr>\n<td>3</td>\n<td>LSTM (1d CNN head)</td>\n<td>0.293</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>DeFOG. CV: 0.381(Event)</li>\n</ul>\n<table>\n<thead>\n<tr>\n<th>No</th>\n<th>Model</th>\n<th>CV</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>LSTM</td>\n<td>0.292</td>\n</tr>\n<tr>\n<td>1</td>\n<td>LSTM (add percentile features)</td>\n<td>0.318</td>\n</tr>\n<tr>\n<td>2</td>\n<td>LSTM (long 4096)</td>\n<td>0.286</td>\n</tr>\n<tr>\n<td>3</td>\n<td>Transformer</td>\n<td>0.286</td>\n</tr>\n</tbody>\n</table>\n<p>Public LB: 0.444<br>\nPrivate LB: 0.341</p>\n<p>Finally, I am happy to become a kaggle master with this gold medal.<br>\nHowever, although I got a gold medal, I am disappointed that I could not design a model that could capture FOG events well both quantitatively and qualitatively.<br>\nI would love to participate in the second and third competitions if they are held. Thank you very much.</p>",
      "rawMarkdown": "I would like to thank host and kaggle for organizing such interesting competition.\nI believe that the method proposed by the winner and prizewinners who received high scores in both the public and private tests will be helpful to host.\n\n## Summary\n- Ensamble of 4 model based by LSTM\n- Use 4 features. mean, std, max, min, median.\n- Notype data is used as validation data.\n- Selected model separately for tDCS FOG and DeFOG.\n\n## Data Preprocessing\nI compressed all data into fixed length. Length is 2048 because shortest training data is 2359.\n\nCompression Method\n1. Expand the original data to a smallest multiple of the target length that exceeds the original data length. (ch, target_size * n)\n2. Fold the data into (ch, target_size, n).\n\nAs feature, I used the mean, std, max, min, and median within this N.\nIn addition to this, it was also useful to add each percentile point (15~90 percentile).\nI rarely used the information on each ID and Subjects, but only \"Visit\" was used as a label for multitasking.\n\n```python\ndef resize_func(x, target_size=2048, use_percentile_feat=False):\n    ch = x.shape[0] \n    input_size = x.shape[1]\n    \n    pad = target_size - input_size % target_size\n    factor = (input_size + pad) / input_size\n\n    x = np.array([ndi.zoom(xi, zoom=factor, mode='reflect') for xi in x])\n    x = x.reshape((ch, target_size, -1))\n\n    res = {} \n    res['mean'] = np.mean(x, axis=2).reshape(ch, -1)\n    res['max'] = np.max(x, axis=2).reshape(ch, -1)\n    res['min'] = np.min(x, axis=2).reshape(ch, -1)\n    res['med'] = np.median(x, axis=2).reshape(ch, -1)\n    res['std'] = np.sqrt(np.var(x, axis=2).reshape(ch, -1))\n    if use_percentile_feat:\n        for p in [15, 30, 45, 60, 75, 90]:\n            res[f\"p{p}\"] = np.percentile(x, [p], axis=2).reshape(ch, -1)\n\n    return res\n```\n\n## Model\nThe model used was LSTM: a simple 4-block model and a model that were [3rd place of similar competitions in the past](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285330). For DeFOG, I also used [Transformer](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112).\n\n\n## CV Strategy\n- tDCS FOG\nIn tDCS FOG, I used 4fold StratifiedGroupKfold with Event label and Subject as group.\nHowever, CV was quite high for certain folds.\nThis is probably due to the large number of examples of difficult classes such as StartHesitation and Walking appearing for a long time in Subejt \"2d57c2\", which reduced the percentage of False Positives.\nTherefore, I had selected a model based on a score that excluded this Subject.\nPersonally, I think that the reason why Shake occurred was also due to the existence of such a Subject in the Public data, and that many teams overfit to it.\n\n- DeFOG\n Because DeFOG has few labeled data, I used all labeled data as training data.\nSince there was a correlation between the loss of the three labels and the loss of the Event label only, I used Nototype as the validation data.\nAlthough it was not possible to make a final selection, I trained on all data including tdcsfog, and the model selected based on the Notype data was the my best private score([this notebook](https://www.kaggle.com/code/mrt0933/keras-lstm-trained-by-all-labeled-data)).\n\n## Final submission\nI selected 4 models for each of tdcsfog and defog and ensemble them.\nFor tDCS FOG, I chose the one with the highest cv for each fold.\nThe best public score had a high enough cv, so I selected this score as the final submission (I'm glad I didn't get caught in the shaking).\n\n- tDCS FOG. CV:0.28\n\n| Fold | Model | CV |\n|:-----------|------------:|:------------:|\n| 0 (w/o \"2d57c2”) |  LSTM (multitasking) | 0.179  |\n| 1 |  LSTM (1d CNN head) | 0.337  |\n| 2 |  LSTM (add percentile features) | 0.278 |\n| 3 |  LSTM (1d CNN head) | 0.293 |\n\n- DeFOG. CV: 0.381(Event)\n\n| No | Model | CV |\n|:-----------|------------:|:------------:|\n| 0 |  LSTM | 0.292  |\n| 1 |  LSTM (add percentile features) | 0.318  |\n| 2 |  LSTM (long 4096) | 0.286 |\n| 3 |  Transformer | 0.286 |\n\nPublic LB: 0.444\nPrivate LB: 0.341\n\n\nFinally, I am happy to become a kaggle master with this gold medal.\nHowever, although I got a gold medal, I am disappointed that I could not design a model that could capture FOG events well both quantitatively and qualitatively.\nI would love to participate in the second and third competitions if they are held. Thank you very much.",
      "votes": null
    },
    {
      "id": "2302368",
      "postDate": "06/14/2023 13:53:58",
      "content": "<p>hello,congratulations on your medal，i am a beginner and have a question that why need to fold the data and set the target_size 2048 ？i can't comprehend the shortest training data is 2359</p>",
      "rawMarkdown": "hello,congratulations on your medal，i am a beginner and have a question that why need to fold the data and set the target_size 2048 ？i can't comprehend the shortest training data is 2359",
      "votes": null
    },
    {
      "id": "2303081",
      "postDate": "06/15/2023 04:41:23",
      "content": "<p>thank you for your comment. <br>\nWhy I compress data by using such scheme is to keep information even with long sequence. I want to tried other method like cropping but it was little bit difficult for me, so first I adopted the simplest one I could imagine (but the competition ended before I could try other methods😅).<br>\nOther teams had tried various methods, such as Crop, it is useful for you and me I think.<br>\nAnd 'shortest training data' that I wrote meant the shortest length of \"len(df)\".</p>",
      "rawMarkdown": "thank you for your comment. \nWhy I compress data by using such scheme is to keep information even with long sequence. I want to tried other method like cropping but it was little bit difficult for me, so first I adopted the simplest one I could imagine (but the competition ended before I could try other methods😅).\nOther teams had tried various methods, such as Crop, it is useful for you and me I think.\nAnd 'shortest training data' that I wrote meant the shortest length of \"len(df)\".",
      "votes": null
    },
    {
      "id": "2304770",
      "postDate": "06/16/2023 08:19:50",
      "content": "<p>Congratulations! Why did you go with the Transformer component for DeFog?</p>",
      "rawMarkdown": "Congratulations! Why did you go with the Transformer component for DeFog?",
      "votes": null
    },
    {
      "id": "2305347",
      "postDate": "06/16/2023 16:06:03",
      "content": "<p>I used it because it was effective in past similar <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion?sort=recent-comments\" target=\"_blank\">competitions</a>.<br>\n I also tried the transformer to tDCS but couldn't get a good cv, so I applied it to DeFOG only.</p>",
      "rawMarkdown": "I used it because it was effective in past similar [competitions](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion?sort=recent-comments).\n I also tried the transformer to tDCS but couldn't get a good cv, so I applied it to DeFOG only.",
      "votes": null
    },
    {
      "id": "2306051",
      "postDate": "06/17/2023 04:41:14",
      "content": "<p>Congrats on the rank/medal!  Can I ask about the second and third competitions you mentioned?  Are those mentioned by the hosts as a possibility? </p>",
      "rawMarkdown": "Congrats on the rank/medal!  Can I ask about the second and third competitions you mentioned?  Are those mentioned by the hosts as a possibility?",
      "votes": null
    },
    {
      "id": "2307960",
      "postDate": "06/18/2023 14:47:17",
      "content": "<p>thank you! <br>\nBy the way, sorry,  I couldn't understand a point of your question.<br>\nCompetition that I mentioned is <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction\" target=\"_blank\">Google Brain - Ventilator Pressure Prediction</a>. And I just modified and used model that from the 3rd place team of this competition .</p>",
      "rawMarkdown": "thank you! \nBy the way, sorry,  I couldn't understand a point of your question.\nCompetition that I mentioned is [Google Brain - Ventilator Pressure Prediction](https://www.kaggle.com/competitions/ventilator-pressure-prediction). And I just modified and used model that from the 3rd place team of this competition .",
      "votes": null
    },
    {
      "id": "2309072",
      "postDate": "06/19/2023 11:30:45",
      "content": "<p>ok,thanks for your reply.</p>",
      "rawMarkdown": "ok,thanks for your reply.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2302368,
      "author_name": "thewuuuu",
      "author_url": "",
      "post_date": "06/14/2023 13:53:58",
      "content": "<p>hello,congratulations on your medal，i am a beginner and have a question that why need to fold the data and set the target_size 2048 ？i can't comprehend the shortest training data is 2359</p>",
      "votes": null,
      "replies": [
        {
          "id": 2303081,
          "author_name": "mrt0933",
          "author_url": "",
          "post_date": "06/15/2023 04:41:23",
          "content": "<p>thank you for your comment. <br>\nWhy I compress data by using such scheme is to keep information even with long sequence. I want to tried other method like cropping but it was little bit difficult for me, so first I adopted the simplest one I could imagine (but the competition ended before I could try other methods😅).<br>\nOther teams had tried various methods, such as Crop, it is useful for you and me I think.<br>\nAnd 'shortest training data' that I wrote meant the shortest length of \"len(df)\".</p>",
          "votes": null,
          "replies": [
            {
              "id": 2309072,
              "author_name": "thewuuuu",
              "author_url": "",
              "post_date": "06/19/2023 11:30:45",
              "content": "<p>ok,thanks for your reply.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2304770,
      "author_name": "exjustice",
      "author_url": "",
      "post_date": "06/16/2023 08:19:50",
      "content": "<p>Congratulations! Why did you go with the Transformer component for DeFog?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2305347,
          "author_name": "mrt0933",
          "author_url": "",
          "post_date": "06/16/2023 16:06:03",
          "content": "<p>I used it because it was effective in past similar <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion?sort=recent-comments\" target=\"_blank\">competitions</a>.<br>\n I also tried the transformer to tDCS but couldn't get a good cv, so I applied it to DeFOG only.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2306051,
      "author_name": "austinhinkel",
      "author_url": "",
      "post_date": "06/17/2023 04:41:14",
      "content": "<p>Congrats on the rank/medal!  Can I ask about the second and third competitions you mentioned?  Are those mentioned by the hosts as a possibility? </p>",
      "votes": null,
      "replies": [
        {
          "id": 2307960,
          "author_name": "mrt0933",
          "author_url": "",
          "post_date": "06/18/2023 14:47:17",
          "content": "<p>thank you! <br>\nBy the way, sorry,  I couldn't understand a point of your question.<br>\nCompetition that I mentioned is <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction\" target=\"_blank\">Google Brain - Ventilator Pressure Prediction</a>. And I just modified and used model that from the 3rd place team of this competition .</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2294826": "I would like to thank host and kaggle for organizing such interesting competition.\nI believe that the method proposed by the winner and prizewinners who received high scores in both the public and private tests will be helpful to host.\n\n## Summary\n- Ensamble of 4 model based by LSTM\n- Use 4 features. mean, std, max, min, median.\n- Notype data is used as validation data.\n- Selected model separately for tDCS FOG and DeFOG.\n\n## Data Preprocessing\nI compressed all data into fixed length. Length is 2048 because shortest training data is 2359.\n\nCompression Method\n1. Expand the original data to a smallest multiple of the target length that exceeds the original data length. (ch, target_size * n)\n2. Fold the data into (ch, target_size, n).\n\nAs feature, I used the mean, std, max, min, and median within this N.\nIn addition to this, it was also useful to add each percentile point (15~90 percentile).\nI rarely used the information on each ID and Subjects, but only \"Visit\" was used as a label for multitasking.\n\n```python\ndef resize_func(x, target_size=2048, use_percentile_feat=False):\n    ch = x.shape[0] \n    input_size = x.shape[1]\n    \n    pad = target_size - input_size % target_size\n    factor = (input_size + pad) / input_size\n\n    x = np.array([ndi.zoom(xi, zoom=factor, mode='reflect') for xi in x])\n    x = x.reshape((ch, target_size, -1))\n\n    res = {} \n    res['mean'] = np.mean(x, axis=2).reshape(ch, -1)\n    res['max'] = np.max(x, axis=2).reshape(ch, -1)\n    res['min'] = np.min(x, axis=2).reshape(ch, -1)\n    res['med'] = np.median(x, axis=2).reshape(ch, -1)\n    res['std'] = np.sqrt(np.var(x, axis=2).reshape(ch, -1))\n    if use_percentile_feat:\n        for p in [15, 30, 45, 60, 75, 90]:\n            res[f\"p{p}\"] = np.percentile(x, [p], axis=2).reshape(ch, -1)\n\n    return res\n```\n\n## Model\nThe model used was LSTM: a simple 4-block model and a model that were [3rd place of similar competitions in the past](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285330). For DeFOG, I also used [Transformer](https://www.kaggle.com/code/cdeotte/tensorflow-transformer-0-112).\n\n\n## CV Strategy\n- tDCS FOG\nIn tDCS FOG, I used 4fold StratifiedGroupKfold with Event label and Subject as group.\nHowever, CV was quite high for certain folds.\nThis is probably due to the large number of examples of difficult classes such as StartHesitation and Walking appearing for a long time in Subejt \"2d57c2\", which reduced the percentage of False Positives.\nTherefore, I had selected a model based on a score that excluded this Subject.\nPersonally, I think that the reason why Shake occurred was also due to the existence of such a Subject in the Public data, and that many teams overfit to it.\n\n- DeFOG\n Because DeFOG has few labeled data, I used all labeled data as training data.\nSince there was a correlation between the loss of the three labels and the loss of the Event label only, I used Nototype as the validation data.\nAlthough it was not possible to make a final selection, I trained on all data including tdcsfog, and the model selected based on the Notype data was the my best private score([this notebook](https://www.kaggle.com/code/mrt0933/keras-lstm-trained-by-all-labeled-data)).\n\n## Final submission\nI selected 4 models for each of tdcsfog and defog and ensemble them.\nFor tDCS FOG, I chose the one with the highest cv for each fold.\nThe best public score had a high enough cv, so I selected this score as the final submission (I'm glad I didn't get caught in the shaking).\n\n- tDCS FOG. CV:0.28\n\n| Fold | Model | CV |\n|:-----------|------------:|:------------:|\n| 0 (w/o \"2d57c2”) |  LSTM (multitasking) | 0.179  |\n| 1 |  LSTM (1d CNN head) | 0.337  |\n| 2 |  LSTM (add percentile features) | 0.278 |\n| 3 |  LSTM (1d CNN head) | 0.293 |\n\n- DeFOG. CV: 0.381(Event)\n\n| No | Model | CV |\n|:-----------|------------:|:------------:|\n| 0 |  LSTM | 0.292  |\n| 1 |  LSTM (add percentile features) | 0.318  |\n| 2 |  LSTM (long 4096) | 0.286 |\n| 3 |  Transformer | 0.286 |\n\nPublic LB: 0.444\nPrivate LB: 0.341\n\n\nFinally, I am happy to become a kaggle master with this gold medal.\nHowever, although I got a gold medal, I am disappointed that I could not design a model that could capture FOG events well both quantitatively and qualitatively.\nI would love to participate in the second and third competitions if they are held. Thank you very much.",
    "2302368": "hello,congratulations on your medal，i am a beginner and have a question that why need to fold the data and set the target_size 2048 ？i can't comprehend the shortest training data is 2359",
    "2303081": "thank you for your comment. \nWhy I compress data by using such scheme is to keep information even with long sequence. I want to tried other method like cropping but it was little bit difficult for me, so first I adopted the simplest one I could imagine (but the competition ended before I could try other methods😅).\nOther teams had tried various methods, such as Crop, it is useful for you and me I think.\nAnd 'shortest training data' that I wrote meant the shortest length of \"len(df)\".",
    "2304770": "Congratulations! Why did you go with the Transformer component for DeFog?",
    "2305347": "I used it because it was effective in past similar [competitions](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion?sort=recent-comments).\n I also tried the transformer to tDCS but couldn't get a good cv, so I applied it to DeFOG only.",
    "2306051": "Congrats on the rank/medal!  Can I ask about the second and third competitions you mentioned?  Are those mentioned by the hosts as a possibility?",
    "2307960": "thank you! \nBy the way, sorry,  I couldn't understand a point of your question.\nCompetition that I mentioned is [Google Brain - Ventilator Pressure Prediction](https://www.kaggle.com/competitions/ventilator-pressure-prediction). And I just modified and used model that from the 3rd place team of this competition .",
    "2309072": "ok,thanks for your reply."
  },
  "source": "meta"
}