{
  "id": 415975,
  "title": "21st place solution: Conv1d with denoising",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/writeups/oom-21st-place-solution-conv1d-with-denoising",
  "author_name": "",
  "post_date": "2023-06-09T01:46:19.760Z",
  "votes": 24,
  "comment_count": 10,
  "views": 0,
  "content": "<p>My 21st solution is based on a great public notebook <a href=\"https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model\" target=\"_blank\">here</a>. The chosen one of my submissions is an ensemble of 5 similar models(<strong>public LB: 0.433, private LB: 0.324</strong>). My best private LB is a single model(<strong>public LB: 0.421, private LB: 0.327</strong>), I publish it <a href=\"https://www.kaggle.com/code/takanashihumbert/gait-single-models-inference/notebook\" target=\"_blank\">here</a>.</p>\n<h2>model</h2>\n<p>The basic structure of my model composes of 3 different <code>Conv1d</code> blocks, then <code>concatenate</code> and <code>flatten</code>, finally with a 4 labels output.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3110858%2F0662912f80e49cb1bb8a7fb91a77364c%2FWX20230609-0907112x.png?generation=1686272859248587&amp;alt=media\" alt=\"\"><br>\nThe features are <strong>Time_frac</strong>, <strong>AccV</strong>, <strong>AccML</strong>, <strong>AccAP</strong>, <strong>V_ML</strong>, <strong>V_AP</strong>, <strong>ML_AP</strong>(the last 3 means the difference value between two of AccV, AccML, and AccAP). And I know <strong>Time_frac</strong> is a powerful(improve nearly 0.1) but controversial one. I use multi-class, so the labels are <strong>StartHesitation</strong>, <strong>Turn</strong>, <strong>Walking</strong>, and <strong>Normal</strong>.</p>\n<h2>denoising</h2>\n<p>My denoising code is very simple:</p>\n<pre><code>def wavelet_denoising_2(x, wavelet=):\n    coeffs = pywt.wavedec(x, wavelet, mode=)\n    coeffs[(coeffs)] *= \n    coeffs[(coeffs)] *= \n     = pywt.waverec(coeffs, wavelet, mode=)\n     (x)%==:\n         = [:]\n     \n</code></pre>\n<h2>window size</h2>\n<pre><code> = \n = \n = \n = \n</code></pre>\n<p>I think that's all, nothing special. Thanks. :)<br>\nLooking forward to top and interesting solutions！</p>",
  "messages": [
    {
      "id": "2293132",
      "postDate": "06/09/2023 01:05:30",
      "content": "<p>My 21st solution is based on a great public notebook <a href=\"https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model\" target=\"_blank\">here</a>. The chosen one of my submissions is an ensemble of 5 similar models(<strong>public LB: 0.433, private LB: 0.324</strong>). My best private LB is a single model(<strong>public LB: 0.421, private LB: 0.327</strong>), I publish it <a href=\"https://www.kaggle.com/code/takanashihumbert/gait-single-models-inference/notebook\" target=\"_blank\">here</a>.</p>\n<h2>model</h2>\n<p>The basic structure of my model composes of 3 different <code>Conv1d</code> blocks, then <code>concatenate</code> and <code>flatten</code>, finally with a 4 labels output.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3110858%2F0662912f80e49cb1bb8a7fb91a77364c%2FWX20230609-0907112x.png?generation=1686272859248587&amp;alt=media\" alt=\"\"><br>\nThe features are <strong>Time_frac</strong>, <strong>AccV</strong>, <strong>AccML</strong>, <strong>AccAP</strong>, <strong>V_ML</strong>, <strong>V_AP</strong>, <strong>ML_AP</strong>(the last 3 means the difference value between two of AccV, AccML, and AccAP). And I know <strong>Time_frac</strong> is a powerful(improve nearly 0.1) but controversial one. I use multi-class, so the labels are <strong>StartHesitation</strong>, <strong>Turn</strong>, <strong>Walking</strong>, and <strong>Normal</strong>.</p>\n<h2>denoising</h2>\n<p>My denoising code is very simple:</p>\n<pre><code>def wavelet_denoising_2(x, wavelet=):\n    coeffs = pywt.wavedec(x, wavelet, mode=)\n    coeffs[(coeffs)] *= \n    coeffs[(coeffs)] *= \n     = pywt.waverec(coeffs, wavelet, mode=)\n     (x)%==:\n         = [:]\n     \n</code></pre>\n<h2>window size</h2>\n<pre><code> = \n = \n = \n = \n</code></pre>\n<p>I think that's all, nothing special. Thanks. :)<br>\nLooking forward to top and interesting solutions！</p>",
      "rawMarkdown": "My 21st solution is based on a great public notebook [here](https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model). The chosen one of my submissions is an ensemble of 5 similar models(**public LB: 0.433, private LB: 0.324**). My best private LB is a single model(**public LB: 0.421, private LB: 0.327**), I publish it [here](https://www.kaggle.com/code/takanashihumbert/gait-single-models-inference/notebook).\n## model\nThe basic structure of my model composes of 3 different `Conv1d` blocks, then `concatenate` and `flatten`, finally with a 4 labels output.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3110858%2F0662912f80e49cb1bb8a7fb91a77364c%2FWX20230609-0907112x.png?generation=1686272859248587&alt=media)\nThe features are **Time_frac**, **AccV**, **AccML**, **AccAP**, **V_ML**, **V_AP**, **ML_AP**(the last 3 means the difference value between two of AccV, AccML, and AccAP). And I know **Time_frac** is a powerful(improve nearly 0.1) but controversial one. I use multi-class, so the labels are **StartHesitation**, **Turn**, **Walking**, and **Normal**.\n## denoising\nMy denoising code is very simple:\n```\ndef wavelet_denoising_2(x, wavelet='db4'):\n    coeffs = pywt.wavedec(x, wavelet, mode=\"per\")\n    coeffs[len(coeffs)-1] *= 0\n    coeffs[len(coeffs)-2] *= 0\n    result = pywt.waverec(coeffs, wavelet, mode='per')\n    if len(x)%2==1:\n        result = result[:-1]\n    return result\n```\n## window size\n```\ndefog_window_size = 200\ndefog_window_future = 50\ntdcsfog_window_size = 256\ntdcsfog_window_future = 64\n```\nI think that's all, nothing special. Thanks. :)\nLooking forward to top and interesting solutions！",
      "votes": null
    },
    {
      "id": "2293332",
      "postDate": "06/09/2023 05:19:45",
      "content": "<p>What were your CV values? Did you do GroupKFold?</p>",
      "rawMarkdown": "What were your CV values? Did you do GroupKFold?",
      "votes": null
    },
    {
      "id": "2293351",
      "postDate": "06/09/2023 05:56:56",
      "content": "<p>Nice solution! By the way, you have two \"Conv1d_2\" on your diagram - probably one should be \"Conv1d_3\"?</p>",
      "rawMarkdown": "Nice solution! By the way, you have two \"Conv1d_2\" on your diagram - probably one should be \"Conv1d_3\"?",
      "votes": null
    },
    {
      "id": "2293355",
      "postDate": "06/09/2023 05:58:06",
      "content": "<p>Yes, my mistake.😮</p>",
      "rawMarkdown": "Yes, my mistake.😮",
      "votes": null
    },
    {
      "id": "2293356",
      "postDate": "06/09/2023 06:00:09",
      "content": "<p>I use <code>StratifiedGroupKFold</code>(y=labels, group=Subject). CV is around 0.32-0.34.</p>",
      "rawMarkdown": "I use `StratifiedGroupKFold`(y=labels, group=Subject). CV is around 0.32-0.34.",
      "votes": null
    },
    {
      "id": "2293667",
      "postDate": "06/09/2023 11:51:20",
      "content": "<p>Thank for referencing my notebook. Great work on your solution 👍</p>",
      "rawMarkdown": "Thank for referencing my notebook. Great work on your solution 👍",
      "votes": null
    },
    {
      "id": "2294793",
      "postDate": "06/10/2023 10:31:48",
      "content": "<p>Thanks, your notebook is well-structured and amazing.👍</p>",
      "rawMarkdown": "Thanks, your notebook is well-structured and amazing.👍",
      "votes": null
    },
    {
      "id": "2324011",
      "postDate": "06/30/2023 10:12:35",
      "content": "<p>Hi Joseph, I'm really impressed by the simplicity of your solution!<br>\nI'm trying to understand the structure of your input windows.<br>\nWhat do you mean by \"_future\"? are you predecting the next 50 points? what's the use of it?<br>\nThank you</p>",
      "rawMarkdown": "Hi Joseph, I'm really impressed by the simplicity of your solution!\nI'm trying to understand the structure of your input windows.\nWhat do you mean by \"_future\"? are you predecting the next 50 points? what's the use of it?\nThank you",
      "votes": null
    },
    {
      "id": "2324032",
      "postDate": "06/30/2023 10:35:20",
      "content": "<p>Hi Alberto, we can use the 'future' points to predict a 'history' one. For example, if we want to predict the time <code>T</code>, we can use the time from <code>T-50</code> to <code>T+50</code> to generate features. It seems to be a leakage, but in this comp it's fine, since we get the whole series of data.</p>",
      "rawMarkdown": "Hi Alberto, we can use the 'future' points to predict a 'history' one. For example, if we want to predict the time `T`, we can use the time from `T-50` to `T+50` to generate features. It seems to be a leakage, but in this comp it's fine, since we get the whole series of data.",
      "votes": null
    },
    {
      "id": "2324127",
      "postDate": "06/30/2023 12:04:04",
      "content": "<p>so you mean e.g. for tdcs that you are using 200 timesteps before the target event and 50 after it?</p>",
      "rawMarkdown": "so you mean e.g. for tdcs that you are using 200 timesteps before the target event and 50 after it?",
      "votes": null
    },
    {
      "id": "2324292",
      "postDate": "06/30/2023 14:07:39",
      "content": "<p>That's right.</p>",
      "rawMarkdown": "That's right.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2293332,
      "author_name": "exjustice",
      "author_url": "",
      "post_date": "06/09/2023 05:19:45",
      "content": "<p>What were your CV values? Did you do GroupKFold?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2293356,
          "author_name": "takanashihumbert",
          "author_url": "",
          "post_date": "06/09/2023 06:00:09",
          "content": "<p>I use <code>StratifiedGroupKFold</code>(y=labels, group=Subject). CV is around 0.32-0.34.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2293351,
      "author_name": "atamazian",
      "author_url": "",
      "post_date": "06/09/2023 05:56:56",
      "content": "<p>Nice solution! By the way, you have two \"Conv1d_2\" on your diagram - probably one should be \"Conv1d_3\"?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2293355,
          "author_name": "takanashihumbert",
          "author_url": "",
          "post_date": "06/09/2023 05:58:06",
          "content": "<p>Yes, my mistake.😮</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2293667,
      "author_name": "coderrkj",
      "author_url": "",
      "post_date": "06/09/2023 11:51:20",
      "content": "<p>Thank for referencing my notebook. Great work on your solution 👍</p>",
      "votes": null,
      "replies": [
        {
          "id": 2294793,
          "author_name": "takanashihumbert",
          "author_url": "",
          "post_date": "06/10/2023 10:31:48",
          "content": "<p>Thanks, your notebook is well-structured and amazing.👍</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2324011,
      "author_name": "albertoannoni",
      "author_url": "",
      "post_date": "06/30/2023 10:12:35",
      "content": "<p>Hi Joseph, I'm really impressed by the simplicity of your solution!<br>\nI'm trying to understand the structure of your input windows.<br>\nWhat do you mean by \"_future\"? are you predecting the next 50 points? what's the use of it?<br>\nThank you</p>",
      "votes": null,
      "replies": [
        {
          "id": 2324032,
          "author_name": "takanashihumbert",
          "author_url": "",
          "post_date": "06/30/2023 10:35:20",
          "content": "<p>Hi Alberto, we can use the 'future' points to predict a 'history' one. For example, if we want to predict the time <code>T</code>, we can use the time from <code>T-50</code> to <code>T+50</code> to generate features. It seems to be a leakage, but in this comp it's fine, since we get the whole series of data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 2324127,
              "author_name": "albertoannoni",
              "author_url": "",
              "post_date": "06/30/2023 12:04:04",
              "content": "<p>so you mean e.g. for tdcs that you are using 200 timesteps before the target event and 50 after it?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2324292,
                  "author_name": "takanashihumbert",
                  "author_url": "",
                  "post_date": "06/30/2023 14:07:39",
                  "content": "<p>That's right.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2293132": "My 21st solution is based on a great public notebook [here](https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model). The chosen one of my submissions is an ensemble of 5 similar models(**public LB: 0.433, private LB: 0.324**). My best private LB is a single model(**public LB: 0.421, private LB: 0.327**), I publish it [here](https://www.kaggle.com/code/takanashihumbert/gait-single-models-inference/notebook).\n## model\nThe basic structure of my model composes of 3 different `Conv1d` blocks, then `concatenate` and `flatten`, finally with a 4 labels output.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3110858%2F0662912f80e49cb1bb8a7fb91a77364c%2FWX20230609-0907112x.png?generation=1686272859248587&alt=media)\nThe features are **Time_frac**, **AccV**, **AccML**, **AccAP**, **V_ML**, **V_AP**, **ML_AP**(the last 3 means the difference value between two of AccV, AccML, and AccAP). And I know **Time_frac** is a powerful(improve nearly 0.1) but controversial one. I use multi-class, so the labels are **StartHesitation**, **Turn**, **Walking**, and **Normal**.\n## denoising\nMy denoising code is very simple:\n```\ndef wavelet_denoising_2(x, wavelet='db4'):\n    coeffs = pywt.wavedec(x, wavelet, mode=\"per\")\n    coeffs[len(coeffs)-1] *= 0\n    coeffs[len(coeffs)-2] *= 0\n    result = pywt.waverec(coeffs, wavelet, mode='per')\n    if len(x)%2==1:\n        result = result[:-1]\n    return result\n```\n## window size\n```\ndefog_window_size = 200\ndefog_window_future = 50\ntdcsfog_window_size = 256\ntdcsfog_window_future = 64\n```\nI think that's all, nothing special. Thanks. :)\nLooking forward to top and interesting solutions！",
    "2293332": "What were your CV values? Did you do GroupKFold?",
    "2293351": "Nice solution! By the way, you have two \"Conv1d_2\" on your diagram - probably one should be \"Conv1d_3\"?",
    "2293355": "Yes, my mistake.😮",
    "2293356": "I use `StratifiedGroupKFold`(y=labels, group=Subject). CV is around 0.32-0.34.",
    "2293667": "Thank for referencing my notebook. Great work on your solution 👍",
    "2294793": "Thanks, your notebook is well-structured and amazing.👍",
    "2324011": "Hi Joseph, I'm really impressed by the simplicity of your solution!\nI'm trying to understand the structure of your input windows.\nWhat do you mean by \"_future\"? are you predecting the next 50 points? what's the use of it?\nThank you",
    "2324032": "Hi Alberto, we can use the 'future' points to predict a 'history' one. For example, if we want to predict the time `T`, we can use the time from `T-50` to `T+50` to generate features. It seems to be a leakage, but in this comp it's fine, since we get the whole series of data.",
    "2324127": "so you mean e.g. for tdcs that you are using 200 timesteps before the target event and 50 after it?",
    "2324292": "That's right."
  },
  "source": "meta"
}