{
  "id": 418275,
  "title": "5th Place Training and Inference",
  "url": "/competitions/tlvmc-parkinsons-freezing-gait-prediction/discussion/418275",
  "author_name": "InnerVoice",
  "post_date": "2023-06-19T18:25:01.984000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<h2><strong>Overall Pipeline</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F699965e32e4940ab317dbe0f7cbe95d1%2Fpiplinev1.PNG?generation=1687193496974817&amp;alt=media\" alt=\"\"></p>\n<h2><strong>Training Methodology</strong></h2>\n<p><strong><em>DataSet</em></strong></p>\n<ul>\n<li>For Pretraining divide each unlabeled series into segment with length 100000</li>\n<li>For Training</li>\n<li>- Divide each time series with window size of 2000 observation with overlapping of 500</li>\n<li>- For defog randomly select 4 window from the above set. </li>\n<li>- For tdcsfog select 1 window from the above set. tdcsfog data is resampled 10 100 Hz using librosa (v 0.9.2). when I ported my training on Kaggle code then i found that training is not converging fast enough when I did resampled using librosa( v 0.10.0). Maybe difference is due to default resampling technique.</li>\n<li>- Set length of dataloader as 8 times number of time series each fold<br>\n<strong><em>Folds</em></strong></li>\n<li>Create 5 Folds using GroupKFold with Subject as groups</li>\n<li>- This could be improved by creating by carefully selecting subjects so that there is similar representation of each target type in each fold</li>\n</ul>\n<p><strong><em>Network Architecture</em></strong><br>\nAll models has following architecture<br>\n`class Wave_Block(nn.Module):</p>\n<pre><code>def :\n    super(Wave_Block, self).\n    self.num_rates = dilation_rates\n    self.convs = nn.\n    self.filter_convs = nn.\n    self.gate_convs = nn.\n\n    self.convs.append(nn.)\n    dilation_rates = \n     dilation_rate  dilation_rates:\n        self.filter_convs.append(\n            nn.)/), dilation=dilation_rate))\n        self.gate_convs.append(\n            nn.)/), dilation=dilation_rate))\n        self.convs.append(nn.)\n\ndef forward(self, x):\n    x = self.convs(x)\n    res = x\n     i  range(self.num_rates):\n        x = torch.tanh(self.filter_convs(x))torch.sigmoid(self.gate_convs(x))\n        x = self.convs(x)\n        res = res + x\n    return res\n</code></pre>\n<p><code>\n</code>class Classifier(nn.Module):<br>\n    def <strong>init</strong>(self, inch=3, kernel_size=3):<br>\n        super().<strong>init</strong>()<br>\n        self.LSTM = nn.GRU(input_size=128, hidden_size=128, num_layers=4, <br>\n                           batch_first=True, bidirectional=True)</p>\n<pre><code>    #self.wave_block1 = \n    self.wave_block2 = \n    self.wave_block3 = \n    self.wave_block4 = \n    self.fc1 = nn.\n\ndef forward(self, x):\n    x = x.permute(, , )\n    #x = self.wave\n    x = self.wave\n    x = self.wave\n\n    x = self.wave\n    x = x.permute(, , )\n    x, h = self.\n    x = self.fc1(x)\n\n\n    return x`\n</code></pre>\n<h3>Different Models</h3>\n<h4>WaveNet-GRU-v1</h4>\n<p>Training Notebook is found at following link</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-4096-v6/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-4096-v6/notebook</a><br>\nThe model is using all available data in training and validation irrespective of <em>Valid</em> column value is True and False and only best weight are saved. Last 2 best weights for each fold are used for inference.</li>\n</ul>\n<h4>WaveNet-GRU-v2</h4>\n<p>Training Notebook is found at following link</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-v6-public/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-v6-public/notebook</a><br>\nThe model is using all available data in training and for validation data <em>Valid</em> column value as True is selected.<br>\nAll the weights with average precision score &gt; 0.25 are saved. Best 2 weights, based on average precision score, are selected.</li>\n</ul>\n<h4>WaveNet-GRU-v3</h4>\n<p>This notebook is based on pre-training on unlabeled data. In pretraining target is to predict next value in the time series. For data creation each unlabeled series is divided into segment of length 100000.</p>\n<p><strong>Data Creation Notebook is available at</strong> : <a href=\"https://www.kaggle.com/code/adityakumarsinha/unlabeled-data-creation/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/unlabeled-data-creation/notebook</a>. <br>\n<em>Note</em>: This notebook will fail in Kaggle kernel as it requires more disk space as default available for kaggle notebooks. Please run on different PC/ Server/ VM<br>\n<strong>PreTraining Notebook is available at</strong> : <a href=\"https://www.kaggle.com/code/adityakumarsinha/pretrain-wavenet-4096-v1/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/pretrain-wavenet-4096-v1/notebook</a><br>\nThe training will be executed for single fold and best weight will be used as initial weights (without LSTM layer) <br>\nfor WaveNet-GRU-v3<br>\n<em>Note</em>: Singe epoch takes around 1-1:30 hours on RTX 3090 so the kernel will timeout. Please run on different PC/ Server/ VM<br>\n<strong>WaveNet-GRU-v3</strong>: **Training notebook is available at **: <a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-from-pretrain/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-from-pretrain/notebook</a>.<br>\nUse best weight for each fold in final inference. <br>\nCV score for this notebook is low as compared to WaveNet-GRU-v1 and WaveNet-GRU-v2 but it improves the final ensemble ( During competition time it improved CV score but due to some bug in inferencing code the final private leader-board score as come down. I will explain this in inferencing section.</p>\n<h2>** Inference Methodology**</h2>\n<ul>\n<li>Each series is predicted independently</li>\n<li>For inference, each series are divided into segments of size 16000 or 20000 and the last segment is comprised of last 16000/20000 data points of the series. It is possible that with this size complete tdcsfog series is predicted in single step.</li>\n<li>tdcsfog data are resampled at 100 Hz and prediction are restored back to 128 Hz.</li>\n<li>librosa 0.10.0, is used for resampling. After competition, I found that librosa 0.9.2 is improves score a bit. This is miss from my side (as i did training using librosa 0.9.2) but it has not much impact on the final score.</li>\n<li>Prediction of all the models are ensembled using simple mean.</li>\n</ul>\n<h3><strong>CPU based Inference Methodology</strong></h3>\n<p>As during last week my GPU quota has been exhausted so i need to use CPU for inference. Simple CPU based pytorch inference was exceeding the time limit of 9 hours. So I need to convert pytorch models into ONNX model. <em>Please refer following notebook for model conversion</em>: <a href=\"https://www.kaggle.com/code/adityakumarsinha/openvino-model-converter-all-models-v3/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/openvino-model-converter-all-models-v3/notebook</a></p>\n<p>The converted models are used in final inference. One of the final inference notebook is available at:<br>\n<a href=\"https://www.kaggle.com/adityakumarsinha/gait-openvion-bunch-v2\" target=\"_blank\">https://www.kaggle.com/adityakumarsinha/gait-openvion-bunch-v2</a>.</p>\n<p>After competition I found that, in ensemble, WaveNet-GRU-v3 (model that uses pretrained weight) is overfitting on public leaderboard and in private leaderboard its inclusion had decreased the score. While in local CV ensemble inclusion of this model was increasing  the CV score. </p>\n<p>So i debugged more and I found that with GPU based inference WaveNet-GRU-v3 is indeed increasing the score. in face simple ensemble of WaveNet-GRU-v1 and WaveNet-GRU-v3 has private leaderboard score of 0.437.  More than third position score.</p>\n<p>The best GPU based inference notebook is available at<br>\n<a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-subm-focal-v2/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-subm-focal-v2/notebook</a></p>\n<p>Regards<br>\nAditya</p>",
  "messages": [
    {
      "id": 2309605,
      "postDate": "2023-06-19T18:25:01.983Z",
      "content": "<h2><strong>Overall Pipeline</strong></h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F699965e32e4940ab317dbe0f7cbe95d1%2Fpiplinev1.PNG?generation=1687193496974817&amp;alt=media\" alt=\"\"></p>\n<h2><strong>Training Methodology</strong></h2>\n<p><strong><em>DataSet</em></strong></p>\n<ul>\n<li>For Pretraining divide each unlabeled series into segment with length 100000</li>\n<li>For Training</li>\n<li>- Divide each time series with window size of 2000 observation with overlapping of 500</li>\n<li>- For defog randomly select 4 window from the above set. </li>\n<li>- For tdcsfog select 1 window from the above set. tdcsfog data is resampled 10 100 Hz using librosa (v 0.9.2). when I ported my training on Kaggle code then i found that training is not converging fast enough when I did resampled using librosa( v 0.10.0). Maybe difference is due to default resampling technique.</li>\n<li>- Set length of dataloader as 8 times number of time series each fold<br>\n<strong><em>Folds</em></strong></li>\n<li>Create 5 Folds using GroupKFold with Subject as groups</li>\n<li>- This could be improved by creating by carefully selecting subjects so that there is similar representation of each target type in each fold</li>\n</ul>\n<p><strong><em>Network Architecture</em></strong><br>\nAll models has following architecture<br>\n`class Wave_Block(nn.Module):</p>\n<pre><code>def :\n    super(Wave_Block, self).\n    self.num_rates = dilation_rates\n    self.convs = nn.\n    self.filter_convs = nn.\n    self.gate_convs = nn.\n\n    self.convs.append(nn.)\n    dilation_rates = \n     dilation_rate  dilation_rates:\n        self.filter_convs.append(\n            nn.)/), dilation=dilation_rate))\n        self.gate_convs.append(\n            nn.)/), dilation=dilation_rate))\n        self.convs.append(nn.)\n\ndef forward(self, x):\n    x = self.convs(x)\n    res = x\n     i  range(self.num_rates):\n        x = torch.tanh(self.filter_convs(x))torch.sigmoid(self.gate_convs(x))\n        x = self.convs(x)\n        res = res + x\n    return res\n</code></pre>\n<p><code>\n</code>class Classifier(nn.Module):<br>\n    def <strong>init</strong>(self, inch=3, kernel_size=3):<br>\n        super().<strong>init</strong>()<br>\n        self.LSTM = nn.GRU(input_size=128, hidden_size=128, num_layers=4, <br>\n                           batch_first=True, bidirectional=True)</p>\n<pre><code>    #self.wave_block1 = \n    self.wave_block2 = \n    self.wave_block3 = \n    self.wave_block4 = \n    self.fc1 = nn.\n\ndef forward(self, x):\n    x = x.permute(, , )\n    #x = self.wave\n    x = self.wave\n    x = self.wave\n\n    x = self.wave\n    x = x.permute(, , )\n    x, h = self.\n    x = self.fc1(x)\n\n\n    return x`\n</code></pre>\n<h3>Different Models</h3>\n<h4>WaveNet-GRU-v1</h4>\n<p>Training Notebook is found at following link</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-4096-v6/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-4096-v6/notebook</a><br>\nThe model is using all available data in training and validation irrespective of <em>Valid</em> column value is True and False and only best weight are saved. Last 2 best weights for each fold are used for inference.</li>\n</ul>\n<h4>WaveNet-GRU-v2</h4>\n<p>Training Notebook is found at following link</p>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-v6-public/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-v6-public/notebook</a><br>\nThe model is using all available data in training and for validation data <em>Valid</em> column value as True is selected.<br>\nAll the weights with average precision score &gt; 0.25 are saved. Best 2 weights, based on average precision score, are selected.</li>\n</ul>\n<h4>WaveNet-GRU-v3</h4>\n<p>This notebook is based on pre-training on unlabeled data. In pretraining target is to predict next value in the time series. For data creation each unlabeled series is divided into segment of length 100000.</p>\n<p><strong>Data Creation Notebook is available at</strong> : <a href=\"https://www.kaggle.com/code/adityakumarsinha/unlabeled-data-creation/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/unlabeled-data-creation/notebook</a>. <br>\n<em>Note</em>: This notebook will fail in Kaggle kernel as it requires more disk space as default available for kaggle notebooks. Please run on different PC/ Server/ VM<br>\n<strong>PreTraining Notebook is available at</strong> : <a href=\"https://www.kaggle.com/code/adityakumarsinha/pretrain-wavenet-4096-v1/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/pretrain-wavenet-4096-v1/notebook</a><br>\nThe training will be executed for single fold and best weight will be used as initial weights (without LSTM layer) <br>\nfor WaveNet-GRU-v3<br>\n<em>Note</em>: Singe epoch takes around 1-1:30 hours on RTX 3090 so the kernel will timeout. Please run on different PC/ Server/ VM<br>\n<strong>WaveNet-GRU-v3</strong>: **Training notebook is available at **: <a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-from-pretrain/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-from-pretrain/notebook</a>.<br>\nUse best weight for each fold in final inference. <br>\nCV score for this notebook is low as compared to WaveNet-GRU-v1 and WaveNet-GRU-v2 but it improves the final ensemble ( During competition time it improved CV score but due to some bug in inferencing code the final private leader-board score as come down. I will explain this in inferencing section.</p>\n<h2>** Inference Methodology**</h2>\n<ul>\n<li>Each series is predicted independently</li>\n<li>For inference, each series are divided into segments of size 16000 or 20000 and the last segment is comprised of last 16000/20000 data points of the series. It is possible that with this size complete tdcsfog series is predicted in single step.</li>\n<li>tdcsfog data are resampled at 100 Hz and prediction are restored back to 128 Hz.</li>\n<li>librosa 0.10.0, is used for resampling. After competition, I found that librosa 0.9.2 is improves score a bit. This is miss from my side (as i did training using librosa 0.9.2) but it has not much impact on the final score.</li>\n<li>Prediction of all the models are ensembled using simple mean.</li>\n</ul>\n<h3><strong>CPU based Inference Methodology</strong></h3>\n<p>As during last week my GPU quota has been exhausted so i need to use CPU for inference. Simple CPU based pytorch inference was exceeding the time limit of 9 hours. So I need to convert pytorch models into ONNX model. <em>Please refer following notebook for model conversion</em>: <a href=\"https://www.kaggle.com/code/adityakumarsinha/openvino-model-converter-all-models-v3/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/openvino-model-converter-all-models-v3/notebook</a></p>\n<p>The converted models are used in final inference. One of the final inference notebook is available at:<br>\n<a href=\"https://www.kaggle.com/adityakumarsinha/gait-openvion-bunch-v2\" target=\"_blank\">https://www.kaggle.com/adityakumarsinha/gait-openvion-bunch-v2</a>.</p>\n<p>After competition I found that, in ensemble, WaveNet-GRU-v3 (model that uses pretrained weight) is overfitting on public leaderboard and in private leaderboard its inclusion had decreased the score. While in local CV ensemble inclusion of this model was increasing  the CV score. </p>\n<p>So i debugged more and I found that with GPU based inference WaveNet-GRU-v3 is indeed increasing the score. in face simple ensemble of WaveNet-GRU-v1 and WaveNet-GRU-v3 has private leaderboard score of 0.437.  More than third position score.</p>\n<p>The best GPU based inference notebook is available at<br>\n<a href=\"https://www.kaggle.com/code/adityakumarsinha/wavenet-subm-focal-v2/notebook\" target=\"_blank\">https://www.kaggle.com/code/adityakumarsinha/wavenet-subm-focal-v2/notebook</a></p>\n<p>Regards<br>\nAditya</p>",
      "rawMarkdown": "## **Overall Pipeline**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F699965e32e4940ab317dbe0f7cbe95d1%2Fpiplinev1.PNG?generation=1687193496974817&alt=media)\n## **Training Methodology**\n***DataSet***\n- For Pretraining divide each unlabeled series into segment with length 100000\n- For Training\n- - Divide each time series with window size of 2000 observation with overlapping of 500\n- - For defog randomly select 4 window from the above set. \n- - For tdcsfog select 1 window from the above set. tdcsfog data is resampled 10 100 Hz using librosa (v 0.9.2). when I ported my training on Kaggle code then i found that training is not converging fast enough when I did resampled using librosa( v 0.10.0). Maybe difference is due to default resampling technique.\n- - Set length of dataloader as 8 times number of time series each fold\n***Folds***\n- Create 5 Folds using GroupKFold with Subject as groups\n- - This could be improved by creating by carefully selecting subjects so that there is similar representation of each target type in each fold\n\n***Network Architecture***\nAll models has following architecture\n`class Wave_Block(nn.Module):\n\n    def __init__(self, in_channels, out_channels, dilation_rates, kernel_size):\n        super(Wave_Block, self).__init__()\n        self.num_rates = dilation_rates\n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=1))\n        dilation_rates = [2 ** i for i in range(dilation_rates)]\n        for dilation_rate in dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=1))\n\n    def forward(self, x):\n        x = self.convs[0](x)\n        res = x\n        for i in range(self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x))\n            x = self.convs[i + 1](x)\n            res = res + x\n        return res\n`\n`class Classifier(nn.Module):\n    def __init__(self, inch=3, kernel_size=3):\n        super().__init__()\n        self.LSTM = nn.GRU(input_size=128, hidden_size=128, num_layers=4, \n                           batch_first=True, bidirectional=True)\n        \n        #self.wave_block1 = Wave_Block(inch, 16, 12, kernel_size)\n        self.wave_block2 = Wave_Block(inch, 32, 8, kernel_size)\n        self.wave_block3 = Wave_Block(32, 64, 4, kernel_size)\n        self.wave_block4 = Wave_Block(64, 128, 1, kernel_size)\n        self.fc1 = nn.Linear(256, 3)\n\n    def forward(self, x):\n        x = x.permute(0, 2, 1)\n        #x = self.wave_block1(x)\n        x = self.wave_block2(x)\n        x = self.wave_block3(x)\n\n        x = self.wave_block4(x)\n        x = x.permute(0, 2, 1)\n        x, h = self.LSTM(x)\n        x = self.fc1(x)\n    \n        \n        return x`\n### Different Models \n#### WaveNet-GRU-v1\nTraining Notebook is found at following link\n- https://www.kaggle.com/code/adityakumarsinha/wavenet-4096-v6/notebook\nThe model is using all available data in training and validation irrespective of *Valid* column value is True and False and only best weight are saved. Last 2 best weights for each fold are used for inference.\n  \n#### WaveNet-GRU-v2\nTraining Notebook is found at following link\n- https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-v6-public/notebook\nThe model is using all available data in training and for validation data *Valid* column value as True is selected.\nAll the weights with average precision score > 0.25 are saved. Best 2 weights, based on average precision score, are selected.\n\n#### WaveNet-GRU-v3\nThis notebook is based on pre-training on unlabeled data. In pretraining target is to predict next value in the time series. For data creation each unlabeled series is divided into segment of length 100000.\n\n**Data Creation Notebook is available at** : https://www.kaggle.com/code/adityakumarsinha/unlabeled-data-creation/notebook. \n*Note*: This notebook will fail in Kaggle kernel as it requires more disk space as default available for kaggle notebooks. Please run on different PC/ Server/ VM\n**PreTraining Notebook is available at** : https://www.kaggle.com/code/adityakumarsinha/pretrain-wavenet-4096-v1/notebook\nThe training will be executed for single fold and best weight will be used as initial weights (without LSTM layer) \nfor WaveNet-GRU-v3\n*Note*: Singe epoch takes around 1-1:30 hours on RTX 3090 so the kernel will timeout. Please run on different PC/ Server/ VM\n**WaveNet-GRU-v3**: **Training notebook is available at **: https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-from-pretrain/notebook.\nUse best weight for each fold in final inference. \nCV score for this notebook is low as compared to WaveNet-GRU-v1 and WaveNet-GRU-v2 but it improves the final ensemble ( During competition time it improved CV score but due to some bug in inferencing code the final private leader-board score as come down. I will explain this in inferencing section.\n\n## ** Inference Methodology**\n- Each series is predicted independently\n- For inference, each series are divided into segments of size 16000 or 20000 and the last segment is comprised of last 16000/20000 data points of the series. It is possible that with this size complete tdcsfog series is predicted in single step.\n- tdcsfog data are resampled at 100 Hz and prediction are restored back to 128 Hz.\n- librosa 0.10.0, is used for resampling. After competition, I found that librosa 0.9.2 is improves score a bit. This is miss from my side (as i did training using librosa 0.9.2) but it has not much impact on the final score.\n- Prediction of all the models are ensembled using simple mean.\n\n### **CPU based Inference Methodology**\nAs during last week my GPU quota has been exhausted so i need to use CPU for inference. Simple CPU based pytorch inference was exceeding the time limit of 9 hours. So I need to convert pytorch models into ONNX model. *Please refer following notebook for model conversion*: https://www.kaggle.com/code/adityakumarsinha/openvino-model-converter-all-models-v3/notebook\n\nThe converted models are used in final inference. One of the final inference notebook is available at:\nhttps://www.kaggle.com/adityakumarsinha/gait-openvion-bunch-v2.\n\nAfter competition I found that, in ensemble, WaveNet-GRU-v3 (model that uses pretrained weight) is overfitting on public leaderboard and in private leaderboard its inclusion had decreased the score. While in local CV ensemble inclusion of this model was increasing  the CV score. \n\nSo i debugged more and I found that with GPU based inference WaveNet-GRU-v3 is indeed increasing the score. in face simple ensemble of WaveNet-GRU-v1 and WaveNet-GRU-v3 has private leaderboard score of 0.437.  More than third position score.\n\nThe best GPU based inference notebook is available at\nhttps://www.kaggle.com/code/adityakumarsinha/wavenet-subm-focal-v2/notebook\n\n\nRegards\nAditya\n",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2309605": "## **Overall Pipeline**\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F651378%2F699965e32e4940ab317dbe0f7cbe95d1%2Fpiplinev1.PNG?generation=1687193496974817&alt=media)\n## **Training Methodology**\n***DataSet***\n- For Pretraining divide each unlabeled series into segment with length 100000\n- For Training\n- - Divide each time series with window size of 2000 observation with overlapping of 500\n- - For defog randomly select 4 window from the above set. \n- - For tdcsfog select 1 window from the above set. tdcsfog data is resampled 10 100 Hz using librosa (v 0.9.2). when I ported my training on Kaggle code then i found that training is not converging fast enough when I did resampled using librosa( v 0.10.0). Maybe difference is due to default resampling technique.\n- - Set length of dataloader as 8 times number of time series each fold\n***Folds***\n- Create 5 Folds using GroupKFold with Subject as groups\n- - This could be improved by creating by carefully selecting subjects so that there is similar representation of each target type in each fold\n\n***Network Architecture***\nAll models has following architecture\n`class Wave_Block(nn.Module):\n\n    def __init__(self, in_channels, out_channels, dilation_rates, kernel_size):\n        super(Wave_Block, self).__init__()\n        self.num_rates = dilation_rates\n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=1))\n        dilation_rates = [2 ** i for i in range(dilation_rates)]\n        for dilation_rate in dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=1))\n\n    def forward(self, x):\n        x = self.convs[0](x)\n        res = x\n        for i in range(self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x))\n            x = self.convs[i + 1](x)\n            res = res + x\n        return res\n`\n`class Classifier(nn.Module):\n    def __init__(self, inch=3, kernel_size=3):\n        super().__init__()\n        self.LSTM = nn.GRU(input_size=128, hidden_size=128, num_layers=4, \n                           batch_first=True, bidirectional=True)\n        \n        #self.wave_block1 = Wave_Block(inch, 16, 12, kernel_size)\n        self.wave_block2 = Wave_Block(inch, 32, 8, kernel_size)\n        self.wave_block3 = Wave_Block(32, 64, 4, kernel_size)\n        self.wave_block4 = Wave_Block(64, 128, 1, kernel_size)\n        self.fc1 = nn.Linear(256, 3)\n\n    def forward(self, x):\n        x = x.permute(0, 2, 1)\n        #x = self.wave_block1(x)\n        x = self.wave_block2(x)\n        x = self.wave_block3(x)\n\n        x = self.wave_block4(x)\n        x = x.permute(0, 2, 1)\n        x, h = self.LSTM(x)\n        x = self.fc1(x)\n    \n        \n        return x`\n### Different Models \n#### WaveNet-GRU-v1\nTraining Notebook is found at following link\n- https://www.kaggle.com/code/adityakumarsinha/wavenet-4096-v6/notebook\nThe model is using all available data in training and validation irrespective of *Valid* column value is True and False and only best weight are saved. Last 2 best weights for each fold are used for inference.\n  \n#### WaveNet-GRU-v2\nTraining Notebook is found at following link\n- https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-v6-public/notebook\nThe model is using all available data in training and for validation data *Valid* column value as True is selected.\nAll the weights with average precision score > 0.25 are saved. Best 2 weights, based on average precision score, are selected.\n\n#### WaveNet-GRU-v3\nThis notebook is based on pre-training on unlabeled data. In pretraining target is to predict next value in the time series. For data creation each unlabeled series is divided into segment of length 100000.\n\n**Data Creation Notebook is available at** : https://www.kaggle.com/code/adityakumarsinha/unlabeled-data-creation/notebook. \n*Note*: This notebook will fail in Kaggle kernel as it requires more disk space as default available for kaggle notebooks. Please run on different PC/ Server/ VM\n**PreTraining Notebook is available at** : https://www.kaggle.com/code/adityakumarsinha/pretrain-wavenet-4096-v1/notebook\nThe training will be executed for single fold and best weight will be used as initial weights (without LSTM layer) \nfor WaveNet-GRU-v3\n*Note*: Singe epoch takes around 1-1:30 hours on RTX 3090 so the kernel will timeout. Please run on different PC/ Server/ VM\n**WaveNet-GRU-v3**: **Training notebook is available at **: https://www.kaggle.com/code/adityakumarsinha/wavenet-2000-from-pretrain/notebook.\nUse best weight for each fold in final inference. \nCV score for this notebook is low as compared to WaveNet-GRU-v1 and WaveNet-GRU-v2 but it improves the final ensemble ( During competition time it improved CV score but due to some bug in inferencing code the final private leader-board score as come down. I will explain this in inferencing section.\n\n## ** Inference Methodology**\n- Each series is predicted independently\n- For inference, each series are divided into segments of size 16000 or 20000 and the last segment is comprised of last 16000/20000 data points of the series. It is possible that with this size complete tdcsfog series is predicted in single step.\n- tdcsfog data are resampled at 100 Hz and prediction are restored back to 128 Hz.\n- librosa 0.10.0, is used for resampling. After competition, I found that librosa 0.9.2 is improves score a bit. This is miss from my side (as i did training using librosa 0.9.2) but it has not much impact on the final score.\n- Prediction of all the models are ensembled using simple mean.\n\n### **CPU based Inference Methodology**\nAs during last week my GPU quota has been exhausted so i need to use CPU for inference. Simple CPU based pytorch inference was exceeding the time limit of 9 hours. So I need to convert pytorch models into ONNX model. *Please refer following notebook for model conversion*: https://www.kaggle.com/code/adityakumarsinha/openvino-model-converter-all-models-v3/notebook\n\nThe converted models are used in final inference. One of the final inference notebook is available at:\nhttps://www.kaggle.com/adityakumarsinha/gait-openvion-bunch-v2.\n\nAfter competition I found that, in ensemble, WaveNet-GRU-v3 (model that uses pretrained weight) is overfitting on public leaderboard and in private leaderboard its inclusion had decreased the score. While in local CV ensemble inclusion of this model was increasing  the CV score. \n\nSo i debugged more and I found that with GPU based inference WaveNet-GRU-v3 is indeed increasing the score. in face simple ensemble of WaveNet-GRU-v1 and WaveNet-GRU-v3 has private leaderboard score of 0.437.  More than third position score.\n\nThe best GPU based inference notebook is available at\nhttps://www.kaggle.com/code/adityakumarsinha/wavenet-subm-focal-v2/notebook\n\n\nRegards\nAditya\n"
  }
}