{
  "id": 416100,
  "title": "57th place solution: Conv1d model with aux head",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/416100",
  "author_name": "Andrіі Didenko",
  "post_date": "2023-06-09T14:45:23.857000",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Hi! 👋 I want to share my solution with you.<br>\nFirst of all, I want to thank the organizers of this competition!</p>\n<p>Then I want to thank the participants who made useful EDA notebooks and notebooks with baselines.<br>\n<em>This solution was inspired by these two notebooks: <a href=\"https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model\" target=\"_blank\">this one</a> and <a href=\"https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254\" target=\"_blank\">this one</a>.</em></p>\n<h1>Data preprocessing</h1>\n<p>For train and inference, I used the rolling window technique. The window was made so the target timestep is in the middle of this window.<br>\nSize of the window for training: 256<br>\nSize of the window during inference: 1224</p>\n<p><strong>AccV, AccML, AccAP</strong> features were preprocessed with savgol filter (kernel=21, n=3).<br>\nFor the <strong>Time</strong> feature I used <code>pandas.qcut</code> to discretize it with 10 bins.</p>\n<h1>Model architecture</h1>\n<p>I used Conv1D model with two heads.<br>\nThe first head (multilabel) is the main one, which performs multilabel classification (<em>StartHesitation, Turn, Walking</em>).<br>\nThe second head (binary) is aux head, which classifies into two classes: <em>event</em>, <em>no event</em>.</p>\n<p>Below is the model architecture.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4268817%2Fa84cc1f97065442d4183a84de2bdf98f%2Fpfogp_solution.drawio.png?generation=1686319440790344&amp;alt=media\" alt=\"\"></p>\n<p>In this model, I also used Conv1D layer with kernel size 8 and stride 8. This helps to reduce input size while keeping important information as much as possible.</p>\n<h1>Training</h1>\n<p>I trained <strong>5 models</strong> on split dataset (by subject_id) using <code>GroupKFold</code>.<br>\n<strong>Optimizer</strong>: AdamW + LookAhead<br>\n<strong>LR</strong>: 0.0002<br>\n<strong>Scheduler</strong>: OneCycleLR<br>\n<strong>Loss</strong>: FocalLoss (binary head) + BCEWithLogitsLoss (multilabel head)<br>\n<strong>Epochs</strong>: 20</p>\n<h1>Inference</h1>\n<p>Inference was performed using an ensemble of 5 models.<br>\nCV scores for each of the model: 0.16, 0.175, 0.198, 0.254, 0.317</p>\n<p>Then I found that using a bigger window size during inference gives better results than using a window size of 256 (which the model was trained on).<br>\nAlso, using an ensemble of different window sizes (3 groups of 5 models) also gives better results.</p>\n<p>Here is the table of some of my submissions</p>\n<table>\n<thead>\n<tr>\n<th>N of models</th>\n<th>Window size</th>\n<th>Public score</th>\n<th>Private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>10 (5 + 5)</td>\n<td>1024 + 1224</td>\n<td>0.37</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>5</td>\n<td>1024</td>\n<td>0.352</td>\n<td>0.307</td>\n</tr>\n<tr>\n<td>15 (5 + 5 + 5)</td>\n<td>256 + 512 + 1024</td>\n<td>0.351</td>\n<td>0.311</td>\n</tr>\n</tbody>\n</table>\n<p>For the final submission, I have chosen the first one (😭).</p>\n<h1>Interesting takeaways</h1>\n<p>During doing experiments I found several interesting takeaways:</p>\n<ul>\n<li>Aux stuff (like heads or losses) can improve performance score.</li>\n<li>Smaller models perform better than bigger ones</li>\n<li>ConvNets perform better than transformers with approximately the same number of parameters</li>\n<li>Using a bigger window size for training reduces LB score, but using a bigger window size during inference (when the model was trained on a smaller window size) increases LB.</li>\n</ul>\n<h1>Conclusions</h1>\n<p>It was an interesting experience for me to participate in this competition. I learned several useful optimization and training tricks (e.g. using aux head). I hope you will find my solution interesting too!<br>\nIf you have any questions about my solution, I would be happy to answer them! 😊</p>",
  "messages": [
    {
      "id": 2293859,
      "postDate": "2023-06-09T14:45:23.857Z",
      "content": "<p>Hi! 👋 I want to share my solution with you.<br>\nFirst of all, I want to thank the organizers of this competition!</p>\n<p>Then I want to thank the participants who made useful EDA notebooks and notebooks with baselines.<br>\n<em>This solution was inspired by these two notebooks: <a href=\"https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model\" target=\"_blank\">this one</a> and <a href=\"https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254\" target=\"_blank\">this one</a>.</em></p>\n<h1>Data preprocessing</h1>\n<p>For train and inference, I used the rolling window technique. The window was made so the target timestep is in the middle of this window.<br>\nSize of the window for training: 256<br>\nSize of the window during inference: 1224</p>\n<p><strong>AccV, AccML, AccAP</strong> features were preprocessed with savgol filter (kernel=21, n=3).<br>\nFor the <strong>Time</strong> feature I used <code>pandas.qcut</code> to discretize it with 10 bins.</p>\n<h1>Model architecture</h1>\n<p>I used Conv1D model with two heads.<br>\nThe first head (multilabel) is the main one, which performs multilabel classification (<em>StartHesitation, Turn, Walking</em>).<br>\nThe second head (binary) is aux head, which classifies into two classes: <em>event</em>, <em>no event</em>.</p>\n<p>Below is the model architecture.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4268817%2Fa84cc1f97065442d4183a84de2bdf98f%2Fpfogp_solution.drawio.png?generation=1686319440790344&amp;alt=media\" alt=\"\"></p>\n<p>In this model, I also used Conv1D layer with kernel size 8 and stride 8. This helps to reduce input size while keeping important information as much as possible.</p>\n<h1>Training</h1>\n<p>I trained <strong>5 models</strong> on split dataset (by subject_id) using <code>GroupKFold</code>.<br>\n<strong>Optimizer</strong>: AdamW + LookAhead<br>\n<strong>LR</strong>: 0.0002<br>\n<strong>Scheduler</strong>: OneCycleLR<br>\n<strong>Loss</strong>: FocalLoss (binary head) + BCEWithLogitsLoss (multilabel head)<br>\n<strong>Epochs</strong>: 20</p>\n<h1>Inference</h1>\n<p>Inference was performed using an ensemble of 5 models.<br>\nCV scores for each of the model: 0.16, 0.175, 0.198, 0.254, 0.317</p>\n<p>Then I found that using a bigger window size during inference gives better results than using a window size of 256 (which the model was trained on).<br>\nAlso, using an ensemble of different window sizes (3 groups of 5 models) also gives better results.</p>\n<p>Here is the table of some of my submissions</p>\n<table>\n<thead>\n<tr>\n<th>N of models</th>\n<th>Window size</th>\n<th>Public score</th>\n<th>Private score</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>10 (5 + 5)</td>\n<td>1024 + 1224</td>\n<td>0.37</td>\n<td>0.303</td>\n</tr>\n<tr>\n<td>5</td>\n<td>1024</td>\n<td>0.352</td>\n<td>0.307</td>\n</tr>\n<tr>\n<td>15 (5 + 5 + 5)</td>\n<td>256 + 512 + 1024</td>\n<td>0.351</td>\n<td>0.311</td>\n</tr>\n</tbody>\n</table>\n<p>For the final submission, I have chosen the first one (😭).</p>\n<h1>Interesting takeaways</h1>\n<p>During doing experiments I found several interesting takeaways:</p>\n<ul>\n<li>Aux stuff (like heads or losses) can improve performance score.</li>\n<li>Smaller models perform better than bigger ones</li>\n<li>ConvNets perform better than transformers with approximately the same number of parameters</li>\n<li>Using a bigger window size for training reduces LB score, but using a bigger window size during inference (when the model was trained on a smaller window size) increases LB.</li>\n</ul>\n<h1>Conclusions</h1>\n<p>It was an interesting experience for me to participate in this competition. I learned several useful optimization and training tricks (e.g. using aux head). I hope you will find my solution interesting too!<br>\nIf you have any questions about my solution, I would be happy to answer them! 😊</p>",
      "rawMarkdown": "Hi! 👋 I want to share my solution with you.\nFirst of all, I want to thank the organizers of this competition!\n\nThen I want to thank the participants who made useful EDA notebooks and notebooks with baselines.\n*This solution was inspired by these two notebooks: [this one](https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model) and [this one](https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254).*\n\n# Data preprocessing\nFor train and inference, I used the rolling window technique. The window was made so the target timestep is in the middle of this window.\nSize of the window for training: 256\nSize of the window during inference: 1224\n\n**AccV, AccML, AccAP** features were preprocessed with savgol filter (kernel=21, n=3).\nFor the **Time** feature I used `pandas.qcut` to discretize it with 10 bins.\n\n# Model architecture\nI used Conv1D model with two heads.\nThe first head (multilabel) is the main one, which performs multilabel classification (*StartHesitation, Turn, Walking*).\nThe second head (binary) is aux head, which classifies into two classes: *event*, *no event*.\n\nBelow is the model architecture.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4268817%2Fa84cc1f97065442d4183a84de2bdf98f%2Fpfogp_solution.drawio.png?generation=1686319440790344&alt=media)\n\nIn this model, I also used Conv1D layer with kernel size 8 and stride 8. This helps to reduce input size while keeping important information as much as possible.\n\n# Training\nI trained **5 models** on split dataset (by subject_id) using `GroupKFold`.\n**Optimizer**: AdamW + LookAhead\n**LR**: 0.0002\n**Scheduler**: OneCycleLR\n**Loss**: FocalLoss (binary head) + BCEWithLogitsLoss (multilabel head)\n**Epochs**: 20\n\n# Inference\nInference was performed using an ensemble of 5 models.\nCV scores for each of the model: 0.16, 0.175, 0.198, 0.254, 0.317\n\nThen I found that using a bigger window size during inference gives better results than using a window size of 256 (which the model was trained on).\nAlso, using an ensemble of different window sizes (3 groups of 5 models) also gives better results.\n\nHere is the table of some of my submissions\n| N of models | Window size      | Public score | Private score |\n| ----------- | ---------------- | ------------ | ------------- |\n| 10 (5 + 5)           | 1024 + 1224      | 0.37         | 0.303         |\n| 5           | 1024             | 0.352        | 0.307         |\n| 15 (5 + 5 + 5)           | 256 + 512 + 1024 | 0.351        | 0.311         |\n\nFor the final submission, I have chosen the first one (😭).\n\n# Interesting takeaways\nDuring doing experiments I found several interesting takeaways:\n- Aux stuff (like heads or losses) can improve performance score.\n- Smaller models perform better than bigger ones\n- ConvNets perform better than transformers with approximately the same number of parameters\n- Using a bigger window size for training reduces LB score, but using a bigger window size during inference (when the model was trained on a smaller window size) increases LB.\n\n# Conclusions\nIt was an interesting experience for me to participate in this competition. I learned several useful optimization and training tricks (e.g. using aux head). I hope you will find my solution interesting too!\nIf you have any questions about my solution, I would be happy to answer them! 😊",
      "votes": 9
    },
    {
      "id": 2321371,
      "postDate": "2023-06-28T13:57:14.120Z",
      "content": "<p>Hi Andrii and thank you for sharing.<br>\nCould you help me understand how could you use two different sizes of window for training and inference?</p>",
      "rawMarkdown": "Hi Andrii and thank you for sharing.\nCould you help me understand how could you use two different sizes of window for training and inference?",
      "votes": 1,
      "replies": [
        {
          "id": 2321389,
          "postDate": "2023-06-28T14:20:56.353Z",
          "content": "<p>Hi!<br>\nI wrote a function that splits input sequence into windows. Then I created train set where window size is 256 and trained model. For inference, I took this trained model and created inference dataset with window size of 512 (and more). This is possible, because my model can process input with variable length.<br>\nIn pseudo code, traning and inference look something like this:</p>\n<pre><code>model = create_model()\ntrain_dataset = create_dataset(input_file=, window_size=)\nfit(model, train_dataset)\n\n...\n\ntest_dataset = create_dataset(input_file=, window_size=)\ny_preds = predict(model, test_dataset)\n</code></pre>\n<p>This is my model in PyTorch:</p>\n<pre><code> (nn.Module):\n     ():\n        (BiHeadConvFOGModel, self).__init__()\n        self.in_layer = nn.Linear(config.in_features, dim)\n        self.reduction_layer = nn.Conv1d(dim, dim, kernel_size=, stride=)\n        self.blocks = nn.Sequential(*[_conv_block(dim, dim, p)  _  (nblocks)])\n        self.avg_pool = nn.AdaptiveAvgPool1d()\n\n        self.multilabel_head = nn.Linear(dim, config.out_classes)\n        self.binary_head = nn.Linear(dim, )\n\n     ():\n        features = self.in_layer(x)\n        features = features.permute(, , )\n        features = self.reduction_layer(features)\n         block  self.blocks:\n            features = block(features)\n        features = self.avg_pool(features)\n        features = features.squeeze()\n\n        multilabel_output = self.multilabel_head(features)\n        binary_output = self.binary_head(features)\n\n         multilabel_output, binary_output\n</code></pre>",
          "rawMarkdown": "Hi!\nI wrote a function that splits input sequence into windows. Then I created train set where window size is 256 and trained model. For inference, I took this trained model and created inference dataset with window size of 512 (and more). This is possible, because my model can process input with variable length.\nIn pseudo code, traning and inference look something like this:\n\n```python\nmodel = create_model()\ntrain_dataset = create_dataset(input_file=\"train_files.npy\", window_size=256)\nfit(model, train_dataset)\n\n...\n\ntest_dataset = create_dataset(input_file=\"test_files.npy\", window_size=512)\ny_preds = predict(model, test_dataset)\n```\n\nThis is my model in PyTorch:\n\n```python\nclass BiHeadConvFOGModel(nn.Module):\n    def __init__(self, p, dim, nblocks):\n        super(BiHeadConvFOGModel, self).__init__()\n        self.in_layer = nn.Linear(config.in_features, dim)\n        self.reduction_layer = nn.Conv1d(dim, dim, kernel_size=8, stride=8)\n        self.blocks = nn.Sequential(*[_conv_block(dim, dim, p) for _ in range(nblocks)])\n        self.avg_pool = nn.AdaptiveAvgPool1d(1)\n\n        self.multilabel_head = nn.Linear(dim, config.out_classes)\n        self.binary_head = nn.Linear(dim, 1)\n\n    def forward(self, x):\n        features = self.in_layer(x)\n        features = features.permute(0, 2, 1)\n        features = self.reduction_layer(features)\n        for block in self.blocks:\n            features = block(features)\n        features = self.avg_pool(features)\n        features = features.squeeze(2)\n\n        multilabel_output = self.multilabel_head(features)\n        binary_output = self.binary_head(features)\n\n        return multilabel_output, binary_output\n```",
          "replies": [
            {
              "id": 2321448,
              "postDate": "2023-06-28T15:26:58.383Z",
              "content": "<p>Thank you again but I still don't understand how this model can handle sequences of different sizes, indeed in the input layer you have to set the shapes, as can be seen in the code you pasted. Can you help me understand in this code where the magic happens? :-)</p>",
              "rawMarkdown": "Thank you again but I still don't understand how this model can handle sequences of different sizes, indeed in the input layer you have to set the shapes, as can be seen in the code you pasted. Can you help me understand in this code where the magic happens? :-)",
              "votes": 1
            },
            {
              "id": 2321589,
              "postDate": "2023-06-28T17:25:46.470Z",
              "content": "<p>The model receives input of shape (B, S, 4), where B - batch size, S - sequence length. Below I am showing output shape of each layer.</p>\n<pre><code> ():\n    features = self.in_layer(x) \n    features = features.permute(, , ) \n    features = self.reduction_layer(features) \n     block  self.blocks:\n        features = block(features) \n    features = self.avg_pool(features) \n    features = features.squeeze() \n\n    multilabel_output = self.multilabel_head(features) \n    binary_output = self.binary_head(features) \n\n     multilabel_output, binary_output \n</code></pre>\n<p>See, after swapping two last axis, I can perform 1D convolution over sequence length. This is because Conv1D in PyTorch performs convolution over last axis. Almost at the end I am using average pooling. And that's why I don't care about sequence length.<br>\nHope this helps. Ask if you have any questions 🙃</p>",
              "rawMarkdown": "The model receives input of shape (B, S, 4), where B - batch size, S - sequence length. Below I am showing output shape of each layer.\n\n```python\ndef forward(self, x):\n    features = self.in_layer(x) # -> (B, S, 128) \n    features = features.permute(0, 2, 1) # -> (B, 128, S) \n    features = self.reduction_layer(features) # -> (B, 128, S/8) \n    for block in self.blocks:\n        features = block(features) # -> (B, 128, S/8)\n    features = self.avg_pool(features) # -> (B, 128, 1) \n    features = features.squeeze(2) # -> (B, 128) \n\n    multilabel_output = self.multilabel_head(features) # -> (B, 4) \n    binary_output = self.binary_head(features) # -> (B, 1) \n\n    return multilabel_output, binary_output # -> [(B, 4), (B, 1)] \n```\n\nSee, after swapping two last axis, I can perform 1D convolution over sequence length. This is because Conv1D in PyTorch performs convolution over last axis. Almost at the end I am using average pooling. And that's why I don't care about sequence length.\nHope this helps. Ask if you have any questions 🙃"
            },
            {
              "id": 2321591,
              "postDate": "2023-06-28T17:26:41.773Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2322834,
              "postDate": "2023-06-29T14:17:56.850Z",
              "content": "<p>Ok so it's the AdaptiveAveragePooling that makes the magic: I normally use Keras and I have not seen such an operation.<br>\nStill I'm quite surprised all the parameters learnt in previous layers works fine with sequences of different lengths! :-)</p>",
              "rawMarkdown": "Ok so it's the AdaptiveAveragePooling that makes the magic: I normally use Keras and I have not seen such an operation.\nStill I'm quite surprised all the parameters learnt in previous layers works fine with sequences of different lengths! :-)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2296226,
      "postDate": "2023-06-11T16:40:25.213Z",
      "content": "<p>Hey Andrii,</p>\n<p>very nice solution! Very similar to my approach, actually much more sophisticated in my opinion - I will definitely try some of your ideas next time I use a 1D Convnet, thank you for sharing!</p>\n<p>Best,<br>\nJan</p>",
      "rawMarkdown": "Hey Andrii,\n\nvery nice solution! Very similar to my approach, actually much more sophisticated in my opinion - I will definitely try some of your ideas next time I use a 1D Convnet, thank you for sharing!\n\nBest,\nJan",
      "votes": 1
    },
    {
      "id": 2294222,
      "postDate": "2023-06-09T22:19:15.793Z",
      "content": "<p>Nice work! Thanks for sharing. First time I am hearing about Lookahead Optimizer, it seems interesting.</p>\n<p>Nice to know about functions like <code>pandas.qcut</code>.</p>",
      "rawMarkdown": "Nice work! Thanks for sharing. First time I am hearing about Lookahead Optimizer, it seems interesting.\n\nNice to know about functions like `pandas.qcut`.",
      "votes": 1
    },
    {
      "id": 2294023,
      "postDate": "2023-06-09T17:00:31.323Z",
      "content": "<p>Nice! What were your CV scores for the models?</p>",
      "rawMarkdown": "Nice! What were your CV scores for the models?",
      "votes": 1,
      "replies": [
        {
          "id": 2294064,
          "postDate": "2023-06-09T17:41:14.957Z",
          "content": "<p>Thank you!<br>\nCV scores for each model (5 folds): 0.16, 0.175, 0.198, 0.254, 0.317. I will edit my post and add them to the summary.<br>\nThen I do experiments with these models. So experiments with window size during inference is actually a postprocessing.<br>\nI should have measured CV scores for each of these experiments, but I didn't, because it was almost the end of competition.</p>",
          "rawMarkdown": "Thank you!\nCV scores for each model (5 folds): 0.16, 0.175, 0.198, 0.254, 0.317. I will edit my post and add them to the summary.\nThen I do experiments with these models. So experiments with window size during inference is actually a postprocessing.\nI should have measured CV scores for each of these experiments, but I didn't, because it was almost the end of competition."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2321371,
      "author_name": "Alberto Annoni",
      "author_url": "",
      "post_date": "2023-06-28T13:57:14.120000",
      "content": "<p>Hi Andrii and thank you for sharing.<br>\nCould you help me understand how could you use two different sizes of window for training and inference?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2321389,
          "author_name": "Andrіі Didenko",
          "author_url": "",
          "post_date": "2023-06-28T14:20:56.353000",
          "content": "<p>Hi!<br>\nI wrote a function that splits input sequence into windows. Then I created train set where window size is 256 and trained model. For inference, I took this trained model and created inference dataset with window size of 512 (and more). This is possible, because my model can process input with variable length.<br>\nIn pseudo code, traning and inference look something like this:</p>\n<pre><code>model = create_model()\ntrain_dataset = create_dataset(input_file=, window_size=)\nfit(model, train_dataset)\n\n...\n\ntest_dataset = create_dataset(input_file=, window_size=)\ny_preds = predict(model, test_dataset)\n</code></pre>\n<p>This is my model in PyTorch:</p>\n<pre><code> (nn.Module):\n     ():\n        (BiHeadConvFOGModel, self).__init__()\n        self.in_layer = nn.Linear(config.in_features, dim)\n        self.reduction_layer = nn.Conv1d(dim, dim, kernel_size=, stride=)\n        self.blocks = nn.Sequential(*[_conv_block(dim, dim, p)  _  (nblocks)])\n        self.avg_pool = nn.AdaptiveAvgPool1d()\n\n        self.multilabel_head = nn.Linear(dim, config.out_classes)\n        self.binary_head = nn.Linear(dim, )\n\n     ():\n        features = self.in_layer(x)\n        features = features.permute(, , )\n        features = self.reduction_layer(features)\n         block  self.blocks:\n            features = block(features)\n        features = self.avg_pool(features)\n        features = features.squeeze()\n\n        multilabel_output = self.multilabel_head(features)\n        binary_output = self.binary_head(features)\n\n         multilabel_output, binary_output\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2321448,
              "author_name": "Alberto Annoni",
              "author_url": "",
              "post_date": "2023-06-28T15:26:58.383000",
              "content": "<p>Thank you again but I still don't understand how this model can handle sequences of different sizes, indeed in the input layer you have to set the shapes, as can be seen in the code you pasted. Can you help me understand in this code where the magic happens? :-)</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2321589,
              "author_name": "Andrіі Didenko",
              "author_url": "",
              "post_date": "2023-06-28T17:25:46.470000",
              "content": "<p>The model receives input of shape (B, S, 4), where B - batch size, S - sequence length. Below I am showing output shape of each layer.</p>\n<pre><code> ():\n    features = self.in_layer(x) \n    features = features.permute(, , ) \n    features = self.reduction_layer(features) \n     block  self.blocks:\n        features = block(features) \n    features = self.avg_pool(features) \n    features = features.squeeze() \n\n    multilabel_output = self.multilabel_head(features) \n    binary_output = self.binary_head(features) \n\n     multilabel_output, binary_output \n</code></pre>\n<p>See, after swapping two last axis, I can perform 1D convolution over sequence length. This is because Conv1D in PyTorch performs convolution over last axis. Almost at the end I am using average pooling. And that's why I don't care about sequence length.<br>\nHope this helps. Ask if you have any questions 🙃</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2321591,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-06-28T17:26:41.773000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2322834,
              "author_name": "Alberto Annoni",
              "author_url": "",
              "post_date": "2023-06-29T14:17:56.850000",
              "content": "<p>Ok so it's the AdaptiveAveragePooling that makes the magic: I normally use Keras and I have not seen such an operation.<br>\nStill I'm quite surprised all the parameters learnt in previous layers works fine with sequences of different lengths! :-)</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2296226,
      "author_name": "Jan Brederecke",
      "author_url": "",
      "post_date": "2023-06-11T16:40:25.213000",
      "content": "<p>Hey Andrii,</p>\n<p>very nice solution! Very similar to my approach, actually much more sophisticated in my opinion - I will definitely try some of your ideas next time I use a 1D Convnet, thank you for sharing!</p>\n<p>Best,<br>\nJan</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2294222,
      "author_name": "coderRKJ",
      "author_url": "",
      "post_date": "2023-06-09T22:19:15.793000",
      "content": "<p>Nice work! Thanks for sharing. First time I am hearing about Lookahead Optimizer, it seems interesting.</p>\n<p>Nice to know about functions like <code>pandas.qcut</code>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2294023,
      "author_name": "Yijie Xu",
      "author_url": "",
      "post_date": "2023-06-09T17:00:31.323000",
      "content": "<p>Nice! What were your CV scores for the models?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2294064,
          "author_name": "Andrіі Didenko",
          "author_url": "",
          "post_date": "2023-06-09T17:41:14.957000",
          "content": "<p>Thank you!<br>\nCV scores for each model (5 folds): 0.16, 0.175, 0.198, 0.254, 0.317. I will edit my post and add them to the summary.<br>\nThen I do experiments with these models. So experiments with window size during inference is actually a postprocessing.<br>\nI should have measured CV scores for each of these experiments, but I didn't, because it was almost the end of competition.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2293859": "Hi! 👋 I want to share my solution with you.\nFirst of all, I want to thank the organizers of this competition!\n\nThen I want to thank the participants who made useful EDA notebooks and notebooks with baselines.\n*This solution was inspired by these two notebooks: [this one](https://www.kaggle.com/code/coderrkj/parkinson-fog-pred-conv1d-separate-tf-model) and [this one](https://www.kaggle.com/code/mayukh18/pytorch-fog-end-to-end-baseline-lb-0-254).*\n\n# Data preprocessing\nFor train and inference, I used the rolling window technique. The window was made so the target timestep is in the middle of this window.\nSize of the window for training: 256\nSize of the window during inference: 1224\n\n**AccV, AccML, AccAP** features were preprocessed with savgol filter (kernel=21, n=3).\nFor the **Time** feature I used `pandas.qcut` to discretize it with 10 bins.\n\n# Model architecture\nI used Conv1D model with two heads.\nThe first head (multilabel) is the main one, which performs multilabel classification (*StartHesitation, Turn, Walking*).\nThe second head (binary) is aux head, which classifies into two classes: *event*, *no event*.\n\nBelow is the model architecture.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4268817%2Fa84cc1f97065442d4183a84de2bdf98f%2Fpfogp_solution.drawio.png?generation=1686319440790344&alt=media)\n\nIn this model, I also used Conv1D layer with kernel size 8 and stride 8. This helps to reduce input size while keeping important information as much as possible.\n\n# Training\nI trained **5 models** on split dataset (by subject_id) using `GroupKFold`.\n**Optimizer**: AdamW + LookAhead\n**LR**: 0.0002\n**Scheduler**: OneCycleLR\n**Loss**: FocalLoss (binary head) + BCEWithLogitsLoss (multilabel head)\n**Epochs**: 20\n\n# Inference\nInference was performed using an ensemble of 5 models.\nCV scores for each of the model: 0.16, 0.175, 0.198, 0.254, 0.317\n\nThen I found that using a bigger window size during inference gives better results than using a window size of 256 (which the model was trained on).\nAlso, using an ensemble of different window sizes (3 groups of 5 models) also gives better results.\n\nHere is the table of some of my submissions\n| N of models | Window size      | Public score | Private score |\n| ----------- | ---------------- | ------------ | ------------- |\n| 10 (5 + 5)           | 1024 + 1224      | 0.37         | 0.303         |\n| 5           | 1024             | 0.352        | 0.307         |\n| 15 (5 + 5 + 5)           | 256 + 512 + 1024 | 0.351        | 0.311         |\n\nFor the final submission, I have chosen the first one (😭).\n\n# Interesting takeaways\nDuring doing experiments I found several interesting takeaways:\n- Aux stuff (like heads or losses) can improve performance score.\n- Smaller models perform better than bigger ones\n- ConvNets perform better than transformers with approximately the same number of parameters\n- Using a bigger window size for training reduces LB score, but using a bigger window size during inference (when the model was trained on a smaller window size) increases LB.\n\n# Conclusions\nIt was an interesting experience for me to participate in this competition. I learned several useful optimization and training tricks (e.g. using aux head). I hope you will find my solution interesting too!\nIf you have any questions about my solution, I would be happy to answer them! 😊",
    "2321371": "Hi Andrii and thank you for sharing.\nCould you help me understand how could you use two different sizes of window for training and inference?",
    "2296226": "Hey Andrii,\n\nvery nice solution! Very similar to my approach, actually much more sophisticated in my opinion - I will definitely try some of your ideas next time I use a 1D Convnet, thank you for sharing!\n\nBest,\nJan",
    "2294222": "Nice work! Thanks for sharing. First time I am hearing about Lookahead Optimizer, it seems interesting.\n\nNice to know about functions like `pandas.qcut`.",
    "2294023": "Nice! What were your CV scores for the models?"
  }
}