{
  "id": 416410,
  "title": "4th Place Solution: a MultiLayer Bidirectional GRU with Residual Connections",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/writeups/zinxira-4th-place-solution-a-multilayer-bidirectio",
  "author_name": "",
  "post_date": "2023-06-18T18:40:37.783Z",
  "votes": 37,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the competition hosts for this competition, and congratulations to the other teams! This was for me the first Kaggle competition in which I invested myself, and it was an awesome experience, I truly learnt a lot. </p>\n<p>The model that performed the best for me is a variant of a multi-layer GRU model in which some residual connections and fully connected layers have been added between the GRU layers:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1969534%2F18364882e67c41284297f4eaed4ddb47%2Fmodel.svg?generation=1686319620125192&amp;alt=media\" alt=\"\"></p>\n<p>Here are the PyTorch classes corresponding to this model: </p>\n<pre><code> (nn.Module):\n     ():\n        (ResidualBiGRU, self).__init__()\n\n        self.hidden_size = hidden_size\n        self.n_layers = n_layers\n\n        self.gru = nn.GRU(\n            hidden_size,\n            hidden_size,\n            n_layers,\n            batch_first=,\n            bidirectional=bidir,\n        )\n        dir_factor =   bidir  \n        self.fc1 = nn.Linear(\n            hidden_size * dir_factor, hidden_size * dir_factor * \n        )\n        self.ln1 = nn.LayerNorm(hidden_size * dir_factor * )\n        self.fc2 = nn.Linear(hidden_size * dir_factor * , hidden_size)\n        self.ln2 = nn.LayerNorm(hidden_size)\n\n     ():\n        res, new_h = self.gru(x, h)\n        \n\n        res = self.fc1(res)\n        res = self.ln1(res)\n        res = nn.functional.relu(res)\n\n        res = self.fc2(res)\n        res = self.ln2(res)\n        res = nn.functional.relu(res)\n\n        \n        res = res + x\n\n         res, new_h\n\n (nn.Module):\n     ():\n        (MultiResidualBiGRU, self).__init__()\n\n        self.input_size = input_size\n        self.hidden_size = hidden_size\n        self.out_size = out_size\n        self.n_layers = n_layers\n\n        self.fc_in = nn.Linear(input_size, hidden_size)\n        self.ln = nn.LayerNorm(hidden_size)\n        self.res_bigrus = nn.ModuleList(\n            [\n                ResidualBiGRU(hidden_size, n_layers=, bidir=bidir)\n                 _  (n_layers)\n            ]\n        )\n        self.fc_out = nn.Linear(hidden_size, out_size)\n\n     ():\n        \n         h  :\n            \n            h = [  _  (self.n_layers)]\n\n        x = self.fc_in(x)\n        x = self.ln(x)\n        x = nn.functional.relu(x)\n\n        new_h = []\n         i, res_bigru  (self.res_bigrus):\n            x, new_hi = res_bigru(x, h[i])\n            new_h.append(new_hi)\n\n        x = self.fc_out(x)\n\n         x, new_h  \n</code></pre>\n<p>Note that the code can be simplified: for my best model which performed a private lb score of 0.417, the \"h\" was actually always initialized with None. </p>\n<h2>Preprocessing</h2>\n<p>In terms of data, the choice I made for my model is very simplistic: consider only the accelerometer data (AccV, AccML, AccAP), merge the data from tdcsfog and defog together and train a single model on it. The main steps of my preprocessing pipeline are the following ones: </p>\n<ul>\n<li>downsample each sequence from their initial frequency (resp. 128 and 100 Hz) to 50Hz;</li>\n<li>for defog: <ul>\n<li>convert from g units to m/s^2 units;</li>\n<li>build a mask using \"Valid\" and \"Task\" to know which time steps are labeled during the training. The unlabeled time steps are fed to the model to get the full sequence context, but they are masked during the loss computation.</li></ul></li>\n<li>add a 4th \"no-activity\" class: the model is trained to recognize this class in the same way as the other classes. Outside of the loss, during the validation, i only use the 3 other classes to compute my metrics;</li>\n<li>per-sequence standard normalization (StandardScaler). </li>\n</ul>\n<p>Outside of the unlabeled time steps coming from defog, I did not use any other unlabeled data. </p>\n<p>For the prototype of another model i did not have the time to finish, I also began to consider some of the characteristics associated to the person who was producing the sequence, in particular I used \"Visit\", \"Age\", \"Sex\", \"YearsSinceDx\", \"UPDRSIII_On\" and \"NFOGQ\". This prototype was roughly following the same architecture as my best model ; the main idea was to initialize the GRU's hidden states with these characteristics, after using some fully connected layers to project them in the dimension of the hidden states. This prototype was also using 1D convolutions to extract features from the accelerometer data before passing them to the GRU layers, and I also considered adding dropout. I think that with more time for me to tune it, it would have beaten my current best model. The first version achieved a private lb score of 0.398. </p>\n<h2>Training details</h2>\n<p>My best model - the one which performed 0.417 on the private leaderboard - has been trained without any form of cross-validation, only with a train/validation split of 80% / 20%. To be honest, this model appeared as a prototype in my early experimentation process, and I considered stratified cross-validation only after. </p>\n<p>In this solution, I fed <strong>each whole downsampled (50Hz)</strong> sequence to my model, one after the other ie with <strong>a batch size of 1</strong>. Note that I would have been unable to feed some of the sequences to my model without downsampling them. I tried multiple different approaches with this architecture, but was unable to produce a better score when increasing the batch size. I tried multiple window sizes for my sequences ; however as I am pretty new in time series and as I also arrived pretty late in the competition, I did not implement any form of overlap and only thought about it too late. This could have probably been a key. Also when increasing the batch size, it seemed apparent that batch normalization was better than layer normalization. </p>\n<p>For the loss, I used a simple cross entropy. As the classes are pretty imbalanced (in particular with my 4th artificial one), I also considered using a weighted cross-entropy, using the inverse frequency of each class as a weight. I also considered trying a focal loss ; but these initial tests seemed unable to perform better than the cross entropy in my case. Despite these negative experiments, I still think that dealing with the imbalance nature of the problem in a better way than I did is important. </p>\n<p>In terms of optimizer, I used Ranger. I also tried Adam and AdamW and honestly i don't think this choice mattered too much. With Ranger I used a learning rate of 1e-3 with 20 epochs, with a cosine annealing schedule starting at 15. </p>\n<p>Note that I also used mixed precision training and gradient clipping. </p>\n<p>The best parameters I found for the architecture of my model are:</p>\n<ul>\n<li>hidden_size: 128;</li>\n<li>n_layers: 3;</li>\n<li>bidir: True.</li>\n</ul>\n<p>Later on, I also tried a stratified k-fold cross-validation in which I stacked the predictions of my k models via a simple average. The architecture and the training details for each fold were the same as for my 0.417 lb model, and this stacking process led to my 2nd best model, performing a score of 0.415 on the private leaderboard (with k=5). I also tried retraining my model on the whole dataset, but this approach did not improve my lb score. </p>\n<p>In no particular order, here are a few other architectures that I also tried but did not improve my score: </p>\n<ul>\n<li>replacing GRU by LSTM in my model: with my architecture, GRUs outperformed LSTMs in all the tests I've realized;</li>\n<li>multiple variants of <a href=\"https://www.researchgate.net/publication/349964066_Multi-input_CNN-GRU_based_human_activity_recognition_using_wearable_sensors\" target=\"_blank\">https://www.researchgate.net/publication/349964066_Multi-input_CNN-GRU_based_human_activity_recognition_using_wearable_sensors</a> ;</li>\n<li>a classic multi-layers bidirectional GRU followed by one or more fully connected layers, also with layer normalization and ReLUs. </li>\n</ul>\n<p>Edit: <br>\nSubmission Notebook: <br>\n<a href=\"https://www.kaggle.com/zinxira/parkinson-fog-pred-4th-place-submission-notebook\" target=\"_blank\">https://www.kaggle.com/zinxira/parkinson-fog-pred-4th-place-submission-notebook</a></p>\n<p>Pretrained models \"dataset\": <br>\n<a href=\"https://www.kaggle.com/datasets/zinxira/models\" target=\"_blank\">https://www.kaggle.com/datasets/zinxira/models</a></p>\n<p>Full open-source code: <br>\n<a href=\"https://github.com/Zinxira/tlvmc-parkinsons-fog-prediction-4th-place-solution\" target=\"_blank\">https://github.com/Zinxira/tlvmc-parkinsons-fog-prediction-4th-place-solution</a></p>",
  "messages": [
    {
      "id": "2295803",
      "postDate": "06/11/2023 08:53:43",
      "content": "<p>Thanks to Kaggle and the competition hosts for this competition, and congratulations to the other teams! This was for me the first Kaggle competition in which I invested myself, and it was an awesome experience, I truly learnt a lot. </p>\n<p>The model that performed the best for me is a variant of a multi-layer GRU model in which some residual connections and fully connected layers have been added between the GRU layers:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1969534%2F18364882e67c41284297f4eaed4ddb47%2Fmodel.svg?generation=1686319620125192&amp;alt=media\" alt=\"\"></p>\n<p>Here are the PyTorch classes corresponding to this model: </p>\n<pre><code> (nn.Module):\n     ():\n        (ResidualBiGRU, self).__init__()\n\n        self.hidden_size = hidden_size\n        self.n_layers = n_layers\n\n        self.gru = nn.GRU(\n            hidden_size,\n            hidden_size,\n            n_layers,\n            batch_first=,\n            bidirectional=bidir,\n        )\n        dir_factor =   bidir  \n        self.fc1 = nn.Linear(\n            hidden_size * dir_factor, hidden_size * dir_factor * \n        )\n        self.ln1 = nn.LayerNorm(hidden_size * dir_factor * )\n        self.fc2 = nn.Linear(hidden_size * dir_factor * , hidden_size)\n        self.ln2 = nn.LayerNorm(hidden_size)\n\n     ():\n        res, new_h = self.gru(x, h)\n        \n\n        res = self.fc1(res)\n        res = self.ln1(res)\n        res = nn.functional.relu(res)\n\n        res = self.fc2(res)\n        res = self.ln2(res)\n        res = nn.functional.relu(res)\n\n        \n        res = res + x\n\n         res, new_h\n\n (nn.Module):\n     ():\n        (MultiResidualBiGRU, self).__init__()\n\n        self.input_size = input_size\n        self.hidden_size = hidden_size\n        self.out_size = out_size\n        self.n_layers = n_layers\n\n        self.fc_in = nn.Linear(input_size, hidden_size)\n        self.ln = nn.LayerNorm(hidden_size)\n        self.res_bigrus = nn.ModuleList(\n            [\n                ResidualBiGRU(hidden_size, n_layers=, bidir=bidir)\n                 _  (n_layers)\n            ]\n        )\n        self.fc_out = nn.Linear(hidden_size, out_size)\n\n     ():\n        \n         h  :\n            \n            h = [  _  (self.n_layers)]\n\n        x = self.fc_in(x)\n        x = self.ln(x)\n        x = nn.functional.relu(x)\n\n        new_h = []\n         i, res_bigru  (self.res_bigrus):\n            x, new_hi = res_bigru(x, h[i])\n            new_h.append(new_hi)\n\n        x = self.fc_out(x)\n\n         x, new_h  \n</code></pre>\n<p>Note that the code can be simplified: for my best model which performed a private lb score of 0.417, the \"h\" was actually always initialized with None. </p>\n<h2>Preprocessing</h2>\n<p>In terms of data, the choice I made for my model is very simplistic: consider only the accelerometer data (AccV, AccML, AccAP), merge the data from tdcsfog and defog together and train a single model on it. The main steps of my preprocessing pipeline are the following ones: </p>\n<ul>\n<li>downsample each sequence from their initial frequency (resp. 128 and 100 Hz) to 50Hz;</li>\n<li>for defog: <ul>\n<li>convert from g units to m/s^2 units;</li>\n<li>build a mask using \"Valid\" and \"Task\" to know which time steps are labeled during the training. The unlabeled time steps are fed to the model to get the full sequence context, but they are masked during the loss computation.</li></ul></li>\n<li>add a 4th \"no-activity\" class: the model is trained to recognize this class in the same way as the other classes. Outside of the loss, during the validation, i only use the 3 other classes to compute my metrics;</li>\n<li>per-sequence standard normalization (StandardScaler). </li>\n</ul>\n<p>Outside of the unlabeled time steps coming from defog, I did not use any other unlabeled data. </p>\n<p>For the prototype of another model i did not have the time to finish, I also began to consider some of the characteristics associated to the person who was producing the sequence, in particular I used \"Visit\", \"Age\", \"Sex\", \"YearsSinceDx\", \"UPDRSIII_On\" and \"NFOGQ\". This prototype was roughly following the same architecture as my best model ; the main idea was to initialize the GRU's hidden states with these characteristics, after using some fully connected layers to project them in the dimension of the hidden states. This prototype was also using 1D convolutions to extract features from the accelerometer data before passing them to the GRU layers, and I also considered adding dropout. I think that with more time for me to tune it, it would have beaten my current best model. The first version achieved a private lb score of 0.398. </p>\n<h2>Training details</h2>\n<p>My best model - the one which performed 0.417 on the private leaderboard - has been trained without any form of cross-validation, only with a train/validation split of 80% / 20%. To be honest, this model appeared as a prototype in my early experimentation process, and I considered stratified cross-validation only after. </p>\n<p>In this solution, I fed <strong>each whole downsampled (50Hz)</strong> sequence to my model, one after the other ie with <strong>a batch size of 1</strong>. Note that I would have been unable to feed some of the sequences to my model without downsampling them. I tried multiple different approaches with this architecture, but was unable to produce a better score when increasing the batch size. I tried multiple window sizes for my sequences ; however as I am pretty new in time series and as I also arrived pretty late in the competition, I did not implement any form of overlap and only thought about it too late. This could have probably been a key. Also when increasing the batch size, it seemed apparent that batch normalization was better than layer normalization. </p>\n<p>For the loss, I used a simple cross entropy. As the classes are pretty imbalanced (in particular with my 4th artificial one), I also considered using a weighted cross-entropy, using the inverse frequency of each class as a weight. I also considered trying a focal loss ; but these initial tests seemed unable to perform better than the cross entropy in my case. Despite these negative experiments, I still think that dealing with the imbalance nature of the problem in a better way than I did is important. </p>\n<p>In terms of optimizer, I used Ranger. I also tried Adam and AdamW and honestly i don't think this choice mattered too much. With Ranger I used a learning rate of 1e-3 with 20 epochs, with a cosine annealing schedule starting at 15. </p>\n<p>Note that I also used mixed precision training and gradient clipping. </p>\n<p>The best parameters I found for the architecture of my model are:</p>\n<ul>\n<li>hidden_size: 128;</li>\n<li>n_layers: 3;</li>\n<li>bidir: True.</li>\n</ul>\n<p>Later on, I also tried a stratified k-fold cross-validation in which I stacked the predictions of my k models via a simple average. The architecture and the training details for each fold were the same as for my 0.417 lb model, and this stacking process led to my 2nd best model, performing a score of 0.415 on the private leaderboard (with k=5). I also tried retraining my model on the whole dataset, but this approach did not improve my lb score. </p>\n<p>In no particular order, here are a few other architectures that I also tried but did not improve my score: </p>\n<ul>\n<li>replacing GRU by LSTM in my model: with my architecture, GRUs outperformed LSTMs in all the tests I've realized;</li>\n<li>multiple variants of <a href=\"https://www.researchgate.net/publication/349964066_Multi-input_CNN-GRU_based_human_activity_recognition_using_wearable_sensors\" target=\"_blank\">https://www.researchgate.net/publication/349964066_Multi-input_CNN-GRU_based_human_activity_recognition_using_wearable_sensors</a> ;</li>\n<li>a classic multi-layers bidirectional GRU followed by one or more fully connected layers, also with layer normalization and ReLUs. </li>\n</ul>\n<p>Edit: <br>\nSubmission Notebook: <br>\n<a href=\"https://www.kaggle.com/zinxira/parkinson-fog-pred-4th-place-submission-notebook\" target=\"_blank\">https://www.kaggle.com/zinxira/parkinson-fog-pred-4th-place-submission-notebook</a></p>\n<p>Pretrained models \"dataset\": <br>\n<a href=\"https://www.kaggle.com/datasets/zinxira/models\" target=\"_blank\">https://www.kaggle.com/datasets/zinxira/models</a></p>\n<p>Full open-source code: <br>\n<a href=\"https://github.com/Zinxira/tlvmc-parkinsons-fog-prediction-4th-place-solution\" target=\"_blank\">https://github.com/Zinxira/tlvmc-parkinsons-fog-prediction-4th-place-solution</a></p>",
      "rawMarkdown": "Thanks to Kaggle and the competition hosts for this competition, and congratulations to the other teams! This was for me the first Kaggle competition in which I invested myself, and it was an awesome experience, I truly learnt a lot. \n\nThe model that performed the best for me is a variant of a multi-layer GRU model in which some residual connections and fully connected layers have been added between the GRU layers:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1969534%2F18364882e67c41284297f4eaed4ddb47%2Fmodel.svg?generation=1686319620125192&alt=media)\n\nHere are the PyTorch classes corresponding to this model: \n\n```\nclass ResidualBiGRU(nn.Module):\n    def __init__(self, hidden_size, n_layers=1, bidir=True):\n        super(ResidualBiGRU, self).__init__()\n\n        self.hidden_size = hidden_size\n        self.n_layers = n_layers\n\n        self.gru = nn.GRU(\n            hidden_size,\n            hidden_size,\n            n_layers,\n            batch_first=True,\n            bidirectional=bidir,\n        )\n        dir_factor = 2 if bidir else 1\n        self.fc1 = nn.Linear(\n            hidden_size * dir_factor, hidden_size * dir_factor * 2\n        )\n        self.ln1 = nn.LayerNorm(hidden_size * dir_factor * 2)\n        self.fc2 = nn.Linear(hidden_size * dir_factor * 2, hidden_size)\n        self.ln2 = nn.LayerNorm(hidden_size)\n\n    def forward(self, x, h=None):\n        res, new_h = self.gru(x, h)\n        # res.shape = (batch_size, sequence_size, 2*hidden_size)\n\n        res = self.fc1(res)\n        res = self.ln1(res)\n        res = nn.functional.relu(res)\n\n        res = self.fc2(res)\n        res = self.ln2(res)\n        res = nn.functional.relu(res)\n\n        # skip connection\n        res = res + x\n\n        return res, new_h\n\nclass MultiResidualBiGRU(nn.Module):\n    def __init__(self, input_size, hidden_size, out_size, n_layers, bidir=True):\n        super(MultiResidualBiGRU, self).__init__()\n\n        self.input_size = input_size\n        self.hidden_size = hidden_size\n        self.out_size = out_size\n        self.n_layers = n_layers\n\n        self.fc_in = nn.Linear(input_size, hidden_size)\n        self.ln = nn.LayerNorm(hidden_size)\n        self.res_bigrus = nn.ModuleList(\n            [\n                ResidualBiGRU(hidden_size, n_layers=1, bidir=bidir)\n                for _ in range(n_layers)\n            ]\n        )\n        self.fc_out = nn.Linear(hidden_size, out_size)\n\n    def forward(self, x, h=None):\n        # if we are at the beginning of a sequence (no hidden state)\n        if h is None:\n            # (re)initialize the hidden state\n            h = [None for _ in range(self.n_layers)]\n\n        x = self.fc_in(x)\n        x = self.ln(x)\n        x = nn.functional.relu(x)\n\n        new_h = []\n        for i, res_bigru in enumerate(self.res_bigrus):\n            x, new_hi = res_bigru(x, h[i])\n            new_h.append(new_hi)\n\n        x = self.fc_out(x)\n\n        return x, new_h  # log probabilities + hidden states\n```\n\nNote that the code can be simplified: for my best model which performed a private lb score of 0.417, the \"h\" was actually always initialized with None. \n\n## Preprocessing\nIn terms of data, the choice I made for my model is very simplistic: consider only the accelerometer data (AccV, AccML, AccAP), merge the data from tdcsfog and defog together and train a single model on it. The main steps of my preprocessing pipeline are the following ones: \n- downsample each sequence from their initial frequency (resp. 128 and 100 Hz) to 50Hz;\n- for defog: \n  - convert from g units to m/s^2 units;\n  - build a mask using \"Valid\" and \"Task\" to know which time steps are labeled during the training. The unlabeled time steps are fed to the model to get the full sequence context, but they are masked during the loss computation.\n- add a 4th \"no-activity\" class: the model is trained to recognize this class in the same way as the other classes. Outside of the loss, during the validation, i only use the 3 other classes to compute my metrics;\n- per-sequence standard normalization (StandardScaler). \n\nOutside of the unlabeled time steps coming from defog, I did not use any other unlabeled data. \n\nFor the prototype of another model i did not have the time to finish, I also began to consider some of the characteristics associated to the person who was producing the sequence, in particular I used \"Visit\", \"Age\", \"Sex\", \"YearsSinceDx\", \"UPDRSIII_On\" and \"NFOGQ\". This prototype was roughly following the same architecture as my best model ; the main idea was to initialize the GRU's hidden states with these characteristics, after using some fully connected layers to project them in the dimension of the hidden states. This prototype was also using 1D convolutions to extract features from the accelerometer data before passing them to the GRU layers, and I also considered adding dropout. I think that with more time for me to tune it, it would have beaten my current best model. The first version achieved a private lb score of 0.398. \n\n## Training details\nMy best model - the one which performed 0.417 on the private leaderboard - has been trained without any form of cross-validation, only with a train/validation split of 80% / 20%. To be honest, this model appeared as a prototype in my early experimentation process, and I considered stratified cross-validation only after. \n\nIn this solution, I fed **each whole downsampled (50Hz)** sequence to my model, one after the other ie with **a batch size of 1**. Note that I would have been unable to feed some of the sequences to my model without downsampling them. I tried multiple different approaches with this architecture, but was unable to produce a better score when increasing the batch size. I tried multiple window sizes for my sequences ; however as I am pretty new in time series and as I also arrived pretty late in the competition, I did not implement any form of overlap and only thought about it too late. This could have probably been a key. Also when increasing the batch size, it seemed apparent that batch normalization was better than layer normalization. \n\nFor the loss, I used a simple cross entropy. As the classes are pretty imbalanced (in particular with my 4th artificial one), I also considered using a weighted cross-entropy, using the inverse frequency of each class as a weight. I also considered trying a focal loss ; but these initial tests seemed unable to perform better than the cross entropy in my case. Despite these negative experiments, I still think that dealing with the imbalance nature of the problem in a better way than I did is important. \n\nIn terms of optimizer, I used Ranger. I also tried Adam and AdamW and honestly i don't think this choice mattered too much. With Ranger I used a learning rate of 1e-3 with 20 epochs, with a cosine annealing schedule starting at 15. \n\nNote that I also used mixed precision training and gradient clipping. \n\nThe best parameters I found for the architecture of my model are:\n- hidden_size: 128;\n- n_layers: 3;\n- bidir: True.\n\nLater on, I also tried a stratified k-fold cross-validation in which I stacked the predictions of my k models via a simple average. The architecture and the training details for each fold were the same as for my 0.417 lb model, and this stacking process led to my 2nd best model, performing a score of 0.415 on the private leaderboard (with k=5). I also tried retraining my model on the whole dataset, but this approach did not improve my lb score. \n\nIn no particular order, here are a few other architectures that I also tried but did not improve my score: \n- replacing GRU by LSTM in my model: with my architecture, GRUs outperformed LSTMs in all the tests I've realized;\n- multiple variants of https://www.researchgate.net/publication/349964066_Multi-input_CNN-GRU_based_human_activity_recognition_using_wearable_sensors ;\n- a classic multi-layers bidirectional GRU followed by one or more fully connected layers, also with layer normalization and ReLUs. \n\nEdit: \nSubmission Notebook: \nhttps://www.kaggle.com/zinxira/parkinson-fog-pred-4th-place-submission-notebook\n\nPretrained models \"dataset\": \nhttps://www.kaggle.com/datasets/zinxira/models\n\nFull open-source code: \nhttps://github.com/Zinxira/tlvmc-parkinsons-fog-prediction-4th-place-solution",
      "votes": null
    },
    {
      "id": "2297020",
      "postDate": "06/12/2023 10:25:56",
      "content": "<p>Such a clean solution and explanation! Thank you for sharing.</p>",
      "rawMarkdown": "Such a clean solution and explanation! Thank you for sharing.",
      "votes": null
    },
    {
      "id": "2321381",
      "postDate": "06/28/2023 14:11:24",
      "content": "<p>Hi and thank you, in particular for the well written code!!<br>\nIt seems like a residual network but with Recurrent instead of Convolutional layers<br>\nDid you have a reference for such architecture?</p>",
      "rawMarkdown": "Hi and thank you, in particular for the well written code!!\nIt seems like a residual network but with Recurrent instead of Convolutional layers\nDid you have a reference for such architecture?",
      "votes": null
    },
    {
      "id": "2321413",
      "postDate": "06/28/2023 14:46:56",
      "content": "<p>Hi, thank you :) I do not have a reference for this architecture in particular, but I inspired myself from a part of <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285965\" target=\"_blank\">the #1 solution from the \"Google Brain - Ventilator Pressure Prediction\" competition</a>. In <a href=\"https://www.kaggle.com/code/shujun717/1-solution-lstm-cnn-transformer-1-fold/notebook\" target=\"_blank\">their code</a>, they are using a ResidualLSTM module where they also use fully connected layers between the LSTMs. Their architecture is complex, and the ResidualLSTM only represents a small part of it, but that was one of my main starting points in my brainstorming process to build my model. <br>\nAlso note that I did not have the time to explore all the possible variations one can think about when looking at my architecture: some design choices I made can be seen as arbitrary, and the architecture could probably be refined by a large margin</p>",
      "rawMarkdown": "Hi, thank you :) I do not have a reference for this architecture in particular, but I inspired myself from a part of [the #1 solution from the \"Google Brain - Ventilator Pressure Prediction\" competition](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285965). In [their code](https://www.kaggle.com/code/shujun717/1-solution-lstm-cnn-transformer-1-fold/notebook), they are using a ResidualLSTM module where they also use fully connected layers between the LSTMs. Their architecture is complex, and the ResidualLSTM only represents a small part of it, but that was one of my main starting points in my brainstorming process to build my model. \nAlso note that I did not have the time to explore all the possible variations one can think about when looking at my architecture: some design choices I made can be seen as arbitrary, and the architecture could probably be refined by a large margin",
      "votes": null
    },
    {
      "id": "2473608",
      "postDate": "10/08/2023 12:13:20",
      "content": "<p>can you explain how you did best_model and simple_model</p>\n<p>how you created data card  this is :<br>\ntlvmc-fog-pred-4th-place-pretrained-models</p>",
      "rawMarkdown": "can you explain how you did best_model and simple_model\n\n\nhow you created data card  this is :\ntlvmc-fog-pred-4th-place-pretrained-models",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2297020,
      "author_name": "zikate",
      "author_url": "",
      "post_date": "06/12/2023 10:25:56",
      "content": "<p>Such a clean solution and explanation! Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2321381,
      "author_name": "albertoannoni",
      "author_url": "",
      "post_date": "06/28/2023 14:11:24",
      "content": "<p>Hi and thank you, in particular for the well written code!!<br>\nIt seems like a residual network but with Recurrent instead of Convolutional layers<br>\nDid you have a reference for such architecture?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2321413,
          "author_name": "zinxira",
          "author_url": "",
          "post_date": "06/28/2023 14:46:56",
          "content": "<p>Hi, thank you :) I do not have a reference for this architecture in particular, but I inspired myself from a part of <a href=\"https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285965\" target=\"_blank\">the #1 solution from the \"Google Brain - Ventilator Pressure Prediction\" competition</a>. In <a href=\"https://www.kaggle.com/code/shujun717/1-solution-lstm-cnn-transformer-1-fold/notebook\" target=\"_blank\">their code</a>, they are using a ResidualLSTM module where they also use fully connected layers between the LSTMs. Their architecture is complex, and the ResidualLSTM only represents a small part of it, but that was one of my main starting points in my brainstorming process to build my model. <br>\nAlso note that I did not have the time to explore all the possible variations one can think about when looking at my architecture: some design choices I made can be seen as arbitrary, and the architecture could probably be refined by a large margin</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2473608,
      "author_name": "satheeshbhukya1",
      "author_url": "",
      "post_date": "10/08/2023 12:13:20",
      "content": "<p>can you explain how you did best_model and simple_model</p>\n<p>how you created data card  this is :<br>\ntlvmc-fog-pred-4th-place-pretrained-models</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2295803": "Thanks to Kaggle and the competition hosts for this competition, and congratulations to the other teams! This was for me the first Kaggle competition in which I invested myself, and it was an awesome experience, I truly learnt a lot. \n\nThe model that performed the best for me is a variant of a multi-layer GRU model in which some residual connections and fully connected layers have been added between the GRU layers:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1969534%2F18364882e67c41284297f4eaed4ddb47%2Fmodel.svg?generation=1686319620125192&alt=media)\n\nHere are the PyTorch classes corresponding to this model: \n\n```\nclass ResidualBiGRU(nn.Module):\n    def __init__(self, hidden_size, n_layers=1, bidir=True):\n        super(ResidualBiGRU, self).__init__()\n\n        self.hidden_size = hidden_size\n        self.n_layers = n_layers\n\n        self.gru = nn.GRU(\n            hidden_size,\n            hidden_size,\n            n_layers,\n            batch_first=True,\n            bidirectional=bidir,\n        )\n        dir_factor = 2 if bidir else 1\n        self.fc1 = nn.Linear(\n            hidden_size * dir_factor, hidden_size * dir_factor * 2\n        )\n        self.ln1 = nn.LayerNorm(hidden_size * dir_factor * 2)\n        self.fc2 = nn.Linear(hidden_size * dir_factor * 2, hidden_size)\n        self.ln2 = nn.LayerNorm(hidden_size)\n\n    def forward(self, x, h=None):\n        res, new_h = self.gru(x, h)\n        # res.shape = (batch_size, sequence_size, 2*hidden_size)\n\n        res = self.fc1(res)\n        res = self.ln1(res)\n        res = nn.functional.relu(res)\n\n        res = self.fc2(res)\n        res = self.ln2(res)\n        res = nn.functional.relu(res)\n\n        # skip connection\n        res = res + x\n\n        return res, new_h\n\nclass MultiResidualBiGRU(nn.Module):\n    def __init__(self, input_size, hidden_size, out_size, n_layers, bidir=True):\n        super(MultiResidualBiGRU, self).__init__()\n\n        self.input_size = input_size\n        self.hidden_size = hidden_size\n        self.out_size = out_size\n        self.n_layers = n_layers\n\n        self.fc_in = nn.Linear(input_size, hidden_size)\n        self.ln = nn.LayerNorm(hidden_size)\n        self.res_bigrus = nn.ModuleList(\n            [\n                ResidualBiGRU(hidden_size, n_layers=1, bidir=bidir)\n                for _ in range(n_layers)\n            ]\n        )\n        self.fc_out = nn.Linear(hidden_size, out_size)\n\n    def forward(self, x, h=None):\n        # if we are at the beginning of a sequence (no hidden state)\n        if h is None:\n            # (re)initialize the hidden state\n            h = [None for _ in range(self.n_layers)]\n\n        x = self.fc_in(x)\n        x = self.ln(x)\n        x = nn.functional.relu(x)\n\n        new_h = []\n        for i, res_bigru in enumerate(self.res_bigrus):\n            x, new_hi = res_bigru(x, h[i])\n            new_h.append(new_hi)\n\n        x = self.fc_out(x)\n\n        return x, new_h  # log probabilities + hidden states\n```\n\nNote that the code can be simplified: for my best model which performed a private lb score of 0.417, the \"h\" was actually always initialized with None. \n\n## Preprocessing\nIn terms of data, the choice I made for my model is very simplistic: consider only the accelerometer data (AccV, AccML, AccAP), merge the data from tdcsfog and defog together and train a single model on it. The main steps of my preprocessing pipeline are the following ones: \n- downsample each sequence from their initial frequency (resp. 128 and 100 Hz) to 50Hz;\n- for defog: \n  - convert from g units to m/s^2 units;\n  - build a mask using \"Valid\" and \"Task\" to know which time steps are labeled during the training. The unlabeled time steps are fed to the model to get the full sequence context, but they are masked during the loss computation.\n- add a 4th \"no-activity\" class: the model is trained to recognize this class in the same way as the other classes. Outside of the loss, during the validation, i only use the 3 other classes to compute my metrics;\n- per-sequence standard normalization (StandardScaler). \n\nOutside of the unlabeled time steps coming from defog, I did not use any other unlabeled data. \n\nFor the prototype of another model i did not have the time to finish, I also began to consider some of the characteristics associated to the person who was producing the sequence, in particular I used \"Visit\", \"Age\", \"Sex\", \"YearsSinceDx\", \"UPDRSIII_On\" and \"NFOGQ\". This prototype was roughly following the same architecture as my best model ; the main idea was to initialize the GRU's hidden states with these characteristics, after using some fully connected layers to project them in the dimension of the hidden states. This prototype was also using 1D convolutions to extract features from the accelerometer data before passing them to the GRU layers, and I also considered adding dropout. I think that with more time for me to tune it, it would have beaten my current best model. The first version achieved a private lb score of 0.398. \n\n## Training details\nMy best model - the one which performed 0.417 on the private leaderboard - has been trained without any form of cross-validation, only with a train/validation split of 80% / 20%. To be honest, this model appeared as a prototype in my early experimentation process, and I considered stratified cross-validation only after. \n\nIn this solution, I fed **each whole downsampled (50Hz)** sequence to my model, one after the other ie with **a batch size of 1**. Note that I would have been unable to feed some of the sequences to my model without downsampling them. I tried multiple different approaches with this architecture, but was unable to produce a better score when increasing the batch size. I tried multiple window sizes for my sequences ; however as I am pretty new in time series and as I also arrived pretty late in the competition, I did not implement any form of overlap and only thought about it too late. This could have probably been a key. Also when increasing the batch size, it seemed apparent that batch normalization was better than layer normalization. \n\nFor the loss, I used a simple cross entropy. As the classes are pretty imbalanced (in particular with my 4th artificial one), I also considered using a weighted cross-entropy, using the inverse frequency of each class as a weight. I also considered trying a focal loss ; but these initial tests seemed unable to perform better than the cross entropy in my case. Despite these negative experiments, I still think that dealing with the imbalance nature of the problem in a better way than I did is important. \n\nIn terms of optimizer, I used Ranger. I also tried Adam and AdamW and honestly i don't think this choice mattered too much. With Ranger I used a learning rate of 1e-3 with 20 epochs, with a cosine annealing schedule starting at 15. \n\nNote that I also used mixed precision training and gradient clipping. \n\nThe best parameters I found for the architecture of my model are:\n- hidden_size: 128;\n- n_layers: 3;\n- bidir: True.\n\nLater on, I also tried a stratified k-fold cross-validation in which I stacked the predictions of my k models via a simple average. The architecture and the training details for each fold were the same as for my 0.417 lb model, and this stacking process led to my 2nd best model, performing a score of 0.415 on the private leaderboard (with k=5). I also tried retraining my model on the whole dataset, but this approach did not improve my lb score. \n\nIn no particular order, here are a few other architectures that I also tried but did not improve my score: \n- replacing GRU by LSTM in my model: with my architecture, GRUs outperformed LSTMs in all the tests I've realized;\n- multiple variants of https://www.researchgate.net/publication/349964066_Multi-input_CNN-GRU_based_human_activity_recognition_using_wearable_sensors ;\n- a classic multi-layers bidirectional GRU followed by one or more fully connected layers, also with layer normalization and ReLUs. \n\nEdit: \nSubmission Notebook: \nhttps://www.kaggle.com/zinxira/parkinson-fog-pred-4th-place-submission-notebook\n\nPretrained models \"dataset\": \nhttps://www.kaggle.com/datasets/zinxira/models\n\nFull open-source code: \nhttps://github.com/Zinxira/tlvmc-parkinsons-fog-prediction-4th-place-solution",
    "2297020": "Such a clean solution and explanation! Thank you for sharing.",
    "2321381": "Hi and thank you, in particular for the well written code!!\nIt seems like a residual network but with Recurrent instead of Convolutional layers\nDid you have a reference for such architecture?",
    "2321413": "Hi, thank you :) I do not have a reference for this architecture in particular, but I inspired myself from a part of [the #1 solution from the \"Google Brain - Ventilator Pressure Prediction\" competition](https://www.kaggle.com/competitions/ventilator-pressure-prediction/discussion/285965). In [their code](https://www.kaggle.com/code/shujun717/1-solution-lstm-cnn-transformer-1-fold/notebook), they are using a ResidualLSTM module where they also use fully connected layers between the LSTMs. Their architecture is complex, and the ResidualLSTM only represents a small part of it, but that was one of my main starting points in my brainstorming process to build my model. \nAlso note that I did not have the time to explore all the possible variations one can think about when looking at my architecture: some design choices I made can be seen as arbitrary, and the architecture could probably be refined by a large margin",
    "2473608": "can you explain how you did best_model and simple_model\n\n\nhow you created data card  this is :\ntlvmc-fog-pred-4th-place-pretrained-models"
  },
  "source": "meta"
}