{
  "id": 471205,
  "title": "What should the WaveNet version of Pytorch look like?",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/471205",
  "author_name": "",
  "post_date": "2024-01-27T11:33:10.862675100Z",
  "votes": 5,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am trying to train <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's  WaveNet network (<a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52#Data-Loader-with-Butter-Low-Pass-Filter\" target=\"_blank\">WaveNet Starter - [LB 0.52]</a>) using PyTorch and <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>'s code <a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/\" target=\"_blank\">HMS-HBAC: ResNet34d Baseline [Training]</a>. However, during the training process, a tensor size of [128, 8, 2000] data takes 13 s to pass through the network, while in the following example (<a href=\"https://www.kaggle.com/code/cswwp347724/wavenet-pytorch/notebook\" target=\"_blank\">Wavenet pytorch</a>), it only takes 1s.  Although I have tried some methods to optimize the speed of PyTorch, they have not worked. I think there is some error in my code that I am not aware of, which makes the training speed slow down. If someone can give me some relevant suggestions, I would be grateful.</p>\n<pre><code> (nn.Module):\n\n     ():\n        (Wave_Block, self).__init__()\n        self.num_rates = n  \n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=))  \n        dilation_rates = [ ** i  i  (n)]  \n         dilation_rate  dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=((dilation_rate*(kernel_size-))/), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=((dilation_rate*(kernel_size-))/), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=))\n\n     ():\n        x = self.convs[](x)\n        res = x\n         i  (self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x)) \n            x = self.convs[i + ](x)\n            res = res + x\n         res\n\n\n (nn.Module):\n\n     ():\n        ().__init__()\n        self.feature_extract = nn.Sequential(\n            Wave_Block(in_channels=in_channels, out_channels=, n=, kernel_size=),\n            Wave_Block(, , n=, kernel_size=),\n            Wave_Block(, , n=, kernel_size=),\n            Wave_Block(, out_channels=out_channels, n=, kernel_size=),\n        )\n\n     (): \n         self.feature_extract(x)\n\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.feature_model = nn.Sequential(\n            WaveBranch(, ) \n        )\n        self.combine_chains = nn.Sequential(\n            nn.Linear(, ),\n            nn.ReLU(inplace=),\n            nn.Linear(, ),\n        )\n\n\n     (): \n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z1 = torch.mean(x3, dim=-)\n\n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z2 = torch.mean(x3, dim=-)\n\n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z3 = torch.mean(x3, dim=-)\n\n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z4 = torch.mean(x3, dim=-)\n\n        \n        y = torch.cat([z1, z2, z3, z4], dim=-)\n        y = self.combine_chains(y)\n\n         y\n\n\nepoch_start = time()\n\ndevice = torch.device()\nwave_model = Classifier().to(device)\ndata = torch.rand(, , ).to(device)\nout = wave_model(data) \n(out.shape)\nelapsed_time = time() - epoch_start\n(elapsed_time)\n</code></pre>",
  "messages": [
    {
      "id": "2622260",
      "postDate": "01/27/2024 11:33:10",
      "content": "<p>I am trying to train <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> 's  WaveNet network (<a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52#Data-Loader-with-Butter-Low-Pass-Filter\" target=\"_blank\">WaveNet Starter - [LB 0.52]</a>) using PyTorch and <a href=\"https://www.kaggle.com/ttahara\" target=\"_blank\">@ttahara</a>'s code <a href=\"https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/\" target=\"_blank\">HMS-HBAC: ResNet34d Baseline [Training]</a>. However, during the training process, a tensor size of [128, 8, 2000] data takes 13 s to pass through the network, while in the following example (<a href=\"https://www.kaggle.com/code/cswwp347724/wavenet-pytorch/notebook\" target=\"_blank\">Wavenet pytorch</a>), it only takes 1s.  Although I have tried some methods to optimize the speed of PyTorch, they have not worked. I think there is some error in my code that I am not aware of, which makes the training speed slow down. If someone can give me some relevant suggestions, I would be grateful.</p>\n<pre><code> (nn.Module):\n\n     ():\n        (Wave_Block, self).__init__()\n        self.num_rates = n  \n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=))  \n        dilation_rates = [ ** i  i  (n)]  \n         dilation_rate  dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=((dilation_rate*(kernel_size-))/), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=((dilation_rate*(kernel_size-))/), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=))\n\n     ():\n        x = self.convs[](x)\n        res = x\n         i  (self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x)) \n            x = self.convs[i + ](x)\n            res = res + x\n         res\n\n\n (nn.Module):\n\n     ():\n        ().__init__()\n        self.feature_extract = nn.Sequential(\n            Wave_Block(in_channels=in_channels, out_channels=, n=, kernel_size=),\n            Wave_Block(, , n=, kernel_size=),\n            Wave_Block(, , n=, kernel_size=),\n            Wave_Block(, out_channels=out_channels, n=, kernel_size=),\n        )\n\n     (): \n         self.feature_extract(x)\n\n\n (nn.Module):\n     ():\n        ().__init__()\n        self.feature_model = nn.Sequential(\n            WaveBranch(, ) \n        )\n        self.combine_chains = nn.Sequential(\n            nn.Linear(, ),\n            nn.ReLU(inplace=),\n            nn.Linear(, ),\n        )\n\n\n     (): \n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z1 = torch.mean(x3, dim=-)\n\n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z2 = torch.mean(x3, dim=-)\n\n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z3 = torch.mean(x3, dim=-)\n\n        \n        x1 = self.feature_model(x[:, :, :])\n        x1 = F.adaptive_avg_pool1d(x1, )\n        x2 = self.feature_model(x[:, :, :])\n        x2 = F.adaptive_avg_pool1d(x2, )\n        x3 = torch.cat([x1, x2], dim=)\n        z4 = torch.mean(x3, dim=-)\n\n        \n        y = torch.cat([z1, z2, z3, z4], dim=-)\n        y = self.combine_chains(y)\n\n         y\n\n\nepoch_start = time()\n\ndevice = torch.device()\nwave_model = Classifier().to(device)\ndata = torch.rand(, , ).to(device)\nout = wave_model(data) \n(out.shape)\nelapsed_time = time() - epoch_start\n(elapsed_time)\n</code></pre>",
      "rawMarkdown": "I am trying to train @cdeotte 's  WaveNet network ([WaveNet Starter - [LB 0.52]](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52#Data-Loader-with-Butter-Low-Pass-Filter)) using PyTorch and @ttahara's code [HMS-HBAC: ResNet34d Baseline [Training]](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/). However, during the training process, a tensor size of [128, 8, 2000] data takes 13 s to pass through the network, while in the following example ([Wavenet pytorch](https://www.kaggle.com/code/cswwp347724/wavenet-pytorch/notebook)), it only takes 1s.  Although I have tried some methods to optimize the speed of PyTorch, they have not worked. I think there is some error in my code that I am not aware of, which makes the training speed slow down. If someone can give me some relevant suggestions, I would be grateful.\n``` python\nclass Wave_Block(nn.Module):\n\n    def __init__(self, in_channels, out_channels, n, kernel_size):\n        super(Wave_Block, self).__init__()\n        self.num_rates = n  # 扩张次数\n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=1))  # 1x1 的一维卷积\n        dilation_rates = [2 ** i for i in range(n)]  # 扩张率为 2 的扩张次数的平方\n        for dilation_rate in dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=1))\n\n    def forward(self, x):\n        x = self.convs[0](x)\n        res = x\n        for i in range(self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x)) # [-1, 1] * (0，1) 信息乘以权重\n            x = self.convs[i + 1](x)\n            res = res + x\n        return res\n    \n\nclass WaveBranch(nn.Module):\n    \n    def __init__(self, in_channels=1, out_channels=64):\n        super().__init__()\n        self.feature_extract = nn.Sequential(\n            Wave_Block(in_channels=in_channels, out_channels=8, n=12, kernel_size=3),\n            Wave_Block(8, 16, n=8, kernel_size=3),\n            Wave_Block(16, 32, n=4, kernel_size=3),\n            Wave_Block(32, out_channels=out_channels, n=1, kernel_size=3),\n        )\n\n    def forward(self, x): # x shape (B, channel, dim)\n        return self.feature_extract(x)\n    \n\nclass Classifier(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.feature_model = nn.Sequential(\n            WaveBranch(1, 64) \n        )\n        self.combine_chains = nn.Sequential(\n            nn.Linear(256, 64),\n            nn.ReLU(inplace=True),\n            nn.Linear(64, 6),\n        )\n\n\n    def forward(self, x): # x shape (B, channel, dim) -> (32, 8, 2000)\n        # LT\n        x1 = self.feature_model(x[:, 0:1, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 1:2, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z1 = torch.mean(x3, dim=-1)\n        \n        # LP\n        x1 = self.feature_model(x[:, 2:3, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 3:4, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z2 = torch.mean(x3, dim=-1)\n        \n        # RP\n        x1 = self.feature_model(x[:, 4:5, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 5:6, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z3 = torch.mean(x3, dim=-1)\n        \n        # RT\n        x1 = self.feature_model(x[:, 6:7, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 7:8, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z4 = torch.mean(x3, dim=-1)\n        \n        # Combine Chain\n        y = torch.cat([z1, z2, z3, z4], dim=-1)\n        y = self.combine_chains(y)\n        \n        return y\n    \n\nepoch_start = time()\n\ndevice = torch.device(\"cuda\")\nwave_model = Classifier().to(device)\ndata = torch.rand(128, 8, 2000).to(device)\nout = wave_model(data) \nprint(out.shape)\nelapsed_time = time() - epoch_start\nprint(elapsed_time)\n```",
      "votes": null
    },
    {
      "id": "2624251",
      "postDate": "01/28/2024 17:13:43",
      "content": "<p>I'm translating wavenet from tensorflow to pytorch version too, but it has different validate loss result (tensorflow version is dwon after each training epoch, but my version looks like fluctuate), this is my version:</p>\n<pre><code>import torch\nimport torch.nn  nn\nimport torch.nn.functional  F\n\n\n :\n\n    def :\n        super(WaveBlock, self).\n        self.num_rates = dilation_rates\n        self.convs = nn.\n        self.filter_convs = nn.\n        self.gate_convs = nn.\n\n        self.convs.append(nn.)\n        dilation_rates = \n         dilation_rate  dilation_rates:\n            self.filter_convs.append(\n                nn.)/), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.)/), dilation=dilation_rate))\n            self.convs.append(nn.)\n\n    def forward(self, x):\n        x = self.convs(x)\n        res = x\n         i  range(self.num_rates):\n            x = torch.tanh(self.filter_convs(x))torch.sigmoid(self.gate_convs(x))\n            x = self.convs(x)\n            res = res + x\n        return res\n\n :\n\n    def :\n        super(WaveNet, self).\n        self.wave_block1 = \n        self.wave_block2 = \n        self.wave_block3 = \n        self.wave_block4 = \n        self.avg_pool = nn.\n        self.fc1 = nn.\n        self.fc2 = nn.\n\n    def feature:\n        x = self.wave\n        x = self.wave\n        x = self.wave\n        x = self.wave\n        return x\n\n    def forward(self, x):\n\n        # LEFT TEMPORAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z1 = torch.mean(torch.stack(), dim=)\n\n        # LEFT PARASAGITTAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z2 = torch.mean(torch.stack(), dim=)\n\n        # RIGHT PARASAGITTAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z3 = torch.mean(torch.stack(), dim=)\n\n        # RIGHT TEMPORAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z4 = torch.mean(torch.stack(), dim=)\n\n        y = torch.cat(, dim=)\n        y = y.view(y.size(), -)\n        y = self.fc1(y)\n        y = torch.relu(y)\n        y = self.fc2(y)\n        y = torch.softmax(y, dim=)\n\n        return y\n</code></pre>",
      "rawMarkdown": "I'm translating wavenet from tensorflow to pytorch version too, but it has different validate loss result (tensorflow version is dwon after each training epoch, but my version looks like fluctuate), this is my version:\n```\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\n\nclass WaveBlock(nn.Module):\n\n    def __init__(self, in_channels, out_channels, kernel_size, dilation_rates):\n        super(WaveBlock, self).__init__()\n        self.num_rates = dilation_rates\n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=1))\n        dilation_rates = [2 ** i for i in range(dilation_rates)]\n        for dilation_rate in dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=1))\n\n    def forward(self, x):\n        x = self.convs[0](x)\n        res = x\n        for i in range(self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x))\n            x = self.convs[i + 1](x)\n            res = res + x\n        return res\n\nclass WaveNet(nn.Module):\n\n    def __init__(self, chn, n_class):\n        super(WaveNet, self).__init__()\n        self.wave_block1 = WaveBlock(chn, 8, 3, 12)\n        self.wave_block2 = WaveBlock(8, 16, 3, 8)\n        self.wave_block3 = WaveBlock(16, 32, 3, 4)\n        self.wave_block4 = WaveBlock(32, 64, 3, 1)\n        self.avg_pool = nn.AdaptiveAvgPool1d(1)\n        self.fc1 = nn.Linear(256, 64)\n        self.fc2 = nn.Linear(64, n_class)\n\n    def feature_extra(self, x):\n        x = self.wave_block1(x)\n        x = self.wave_block2(x)\n        x = self.wave_block3(x)\n        x = self.wave_block4(x)\n        return x\n\n    def forward(self, x):\n\n        # LEFT TEMPORAL CHAIN\n        x1 = self.feature_extra(x[:,:,0:1])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,1:2])\n        x2 = self.avg_pool(x2)\n        z1 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        # LEFT PARASAGITTAL CHAIN\n        x1 = self.feature_extra(x[:,:,2:3])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,3:4])\n        x2 = self.avg_pool(x2)\n        z2 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        # RIGHT PARASAGITTAL CHAIN\n        x1 = self.feature_extra(x[:,:,4:5])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,5:6])\n        x2 = self.avg_pool(x2)\n        z3 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        # RIGHT TEMPORAL CHAIN\n        x1 = self.feature_extra(x[:,:,6:7])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,7:8])\n        x2 = self.avg_pool(x2)\n        z4 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        y = torch.cat([z1,z2,z3,z4], dim=1)\n        y = y.view(y.size(0), -1)\n        y = self.fc1(y)\n        y = torch.relu(y)\n        y = self.fc2(y)\n        y = torch.softmax(y, dim=1)\n\n        return y\n```",
      "votes": null
    },
    {
      "id": "2624947",
      "postDate": "01/29/2024 05:15:11",
      "content": "<p>I remember that the loss function of PytTorch will be processed by softmax itself, I think it is not necessary to use softmax in the network. In addition, I would like to ask how long it takes you to train one epoch, my one epoch takes more than ten minutes, it seems that there is a problem with my training code rather than the model code. Thank you very much for your answer.☺️</p>",
      "rawMarkdown": "I remember that the loss function of PytTorch will be processed by softmax itself, I think it is not necessary to use softmax in the network. In addition, I would like to ask how long it takes you to train one epoch, my one epoch takes more than ten minutes, it seems that there is a problem with my training code rather than the model code. Thank you very much for your answer.☺️",
      "votes": null
    },
    {
      "id": "2625458",
      "postDate": "01/29/2024 12:37:17",
      "content": "<p>If I remember correctly, with dilated convolutions you should set <code>torch.backends.cudnn.benchmark = True</code>, <br>\notherwise inefficient kernel would be chosen</p>",
      "rawMarkdown": "If I remember correctly, with dilated convolutions you should set `torch.backends.cudnn.benchmark = True`, \notherwise inefficient kernel would be chosen",
      "votes": null
    },
    {
      "id": "2625614",
      "postDate": "01/29/2024 14:13:30",
      "content": "<p>Thank you very much for your suggestion. I have tried it but it doesn't work.😭</p>",
      "rawMarkdown": "Thank you very much for your suggestion. I have tried it but it doesn't work.😭",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2624251,
      "author_name": "ax12han9",
      "author_url": "",
      "post_date": "01/28/2024 17:13:43",
      "content": "<p>I'm translating wavenet from tensorflow to pytorch version too, but it has different validate loss result (tensorflow version is dwon after each training epoch, but my version looks like fluctuate), this is my version:</p>\n<pre><code>import torch\nimport torch.nn  nn\nimport torch.nn.functional  F\n\n\n :\n\n    def :\n        super(WaveBlock, self).\n        self.num_rates = dilation_rates\n        self.convs = nn.\n        self.filter_convs = nn.\n        self.gate_convs = nn.\n\n        self.convs.append(nn.)\n        dilation_rates = \n         dilation_rate  dilation_rates:\n            self.filter_convs.append(\n                nn.)/), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.)/), dilation=dilation_rate))\n            self.convs.append(nn.)\n\n    def forward(self, x):\n        x = self.convs(x)\n        res = x\n         i  range(self.num_rates):\n            x = torch.tanh(self.filter_convs(x))torch.sigmoid(self.gate_convs(x))\n            x = self.convs(x)\n            res = res + x\n        return res\n\n :\n\n    def :\n        super(WaveNet, self).\n        self.wave_block1 = \n        self.wave_block2 = \n        self.wave_block3 = \n        self.wave_block4 = \n        self.avg_pool = nn.\n        self.fc1 = nn.\n        self.fc2 = nn.\n\n    def feature:\n        x = self.wave\n        x = self.wave\n        x = self.wave\n        x = self.wave\n        return x\n\n    def forward(self, x):\n\n        # LEFT TEMPORAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z1 = torch.mean(torch.stack(), dim=)\n\n        # LEFT PARASAGITTAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z2 = torch.mean(torch.stack(), dim=)\n\n        # RIGHT PARASAGITTAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z3 = torch.mean(torch.stack(), dim=)\n\n        # RIGHT TEMPORAL CHAIN\n        x1 = self.feature\n        x1 = self.avg\n        x2 = self.feature\n        x2 = self.avg\n        z4 = torch.mean(torch.stack(), dim=)\n\n        y = torch.cat(, dim=)\n        y = y.view(y.size(), -)\n        y = self.fc1(y)\n        y = torch.relu(y)\n        y = self.fc2(y)\n        y = torch.softmax(y, dim=)\n\n        return y\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 2624947,
          "author_name": "orzlala",
          "author_url": "",
          "post_date": "01/29/2024 05:15:11",
          "content": "<p>I remember that the loss function of PytTorch will be processed by softmax itself, I think it is not necessary to use softmax in the network. In addition, I would like to ask how long it takes you to train one epoch, my one epoch takes more than ten minutes, it seems that there is a problem with my training code rather than the model code. Thank you very much for your answer.☺️</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2625458,
      "author_name": "martynoveduard",
      "author_url": "",
      "post_date": "01/29/2024 12:37:17",
      "content": "<p>If I remember correctly, with dilated convolutions you should set <code>torch.backends.cudnn.benchmark = True</code>, <br>\notherwise inefficient kernel would be chosen</p>",
      "votes": null,
      "replies": [
        {
          "id": 2625614,
          "author_name": "orzlala",
          "author_url": "",
          "post_date": "01/29/2024 14:13:30",
          "content": "<p>Thank you very much for your suggestion. I have tried it but it doesn't work.😭</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2622260": "I am trying to train @cdeotte 's  WaveNet network ([WaveNet Starter - [LB 0.52]](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52#Data-Loader-with-Butter-Low-Pass-Filter)) using PyTorch and @ttahara's code [HMS-HBAC: ResNet34d Baseline [Training]](https://www.kaggle.com/code/ttahara/hms-hbac-resnet34d-baseline-training/). However, during the training process, a tensor size of [128, 8, 2000] data takes 13 s to pass through the network, while in the following example ([Wavenet pytorch](https://www.kaggle.com/code/cswwp347724/wavenet-pytorch/notebook)), it only takes 1s.  Although I have tried some methods to optimize the speed of PyTorch, they have not worked. I think there is some error in my code that I am not aware of, which makes the training speed slow down. If someone can give me some relevant suggestions, I would be grateful.\n``` python\nclass Wave_Block(nn.Module):\n\n    def __init__(self, in_channels, out_channels, n, kernel_size):\n        super(Wave_Block, self).__init__()\n        self.num_rates = n  # 扩张次数\n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=1))  # 1x1 的一维卷积\n        dilation_rates = [2 ** i for i in range(n)]  # 扩张率为 2 的扩张次数的平方\n        for dilation_rate in dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=1))\n\n    def forward(self, x):\n        x = self.convs[0](x)\n        res = x\n        for i in range(self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x)) # [-1, 1] * (0，1) 信息乘以权重\n            x = self.convs[i + 1](x)\n            res = res + x\n        return res\n    \n\nclass WaveBranch(nn.Module):\n    \n    def __init__(self, in_channels=1, out_channels=64):\n        super().__init__()\n        self.feature_extract = nn.Sequential(\n            Wave_Block(in_channels=in_channels, out_channels=8, n=12, kernel_size=3),\n            Wave_Block(8, 16, n=8, kernel_size=3),\n            Wave_Block(16, 32, n=4, kernel_size=3),\n            Wave_Block(32, out_channels=out_channels, n=1, kernel_size=3),\n        )\n\n    def forward(self, x): # x shape (B, channel, dim)\n        return self.feature_extract(x)\n    \n\nclass Classifier(nn.Module):\n    def __init__(self):\n        super().__init__()\n        self.feature_model = nn.Sequential(\n            WaveBranch(1, 64) \n        )\n        self.combine_chains = nn.Sequential(\n            nn.Linear(256, 64),\n            nn.ReLU(inplace=True),\n            nn.Linear(64, 6),\n        )\n\n\n    def forward(self, x): # x shape (B, channel, dim) -> (32, 8, 2000)\n        # LT\n        x1 = self.feature_model(x[:, 0:1, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 1:2, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z1 = torch.mean(x3, dim=-1)\n        \n        # LP\n        x1 = self.feature_model(x[:, 2:3, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 3:4, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z2 = torch.mean(x3, dim=-1)\n        \n        # RP\n        x1 = self.feature_model(x[:, 4:5, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 5:6, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z3 = torch.mean(x3, dim=-1)\n        \n        # RT\n        x1 = self.feature_model(x[:, 6:7, :])\n        x1 = F.adaptive_avg_pool1d(x1, 1)\n        x2 = self.feature_model(x[:, 7:8, :])\n        x2 = F.adaptive_avg_pool1d(x2, 1)\n        x3 = torch.cat([x1, x2], dim=2)\n        z4 = torch.mean(x3, dim=-1)\n        \n        # Combine Chain\n        y = torch.cat([z1, z2, z3, z4], dim=-1)\n        y = self.combine_chains(y)\n        \n        return y\n    \n\nepoch_start = time()\n\ndevice = torch.device(\"cuda\")\nwave_model = Classifier().to(device)\ndata = torch.rand(128, 8, 2000).to(device)\nout = wave_model(data) \nprint(out.shape)\nelapsed_time = time() - epoch_start\nprint(elapsed_time)\n```",
    "2624251": "I'm translating wavenet from tensorflow to pytorch version too, but it has different validate loss result (tensorflow version is dwon after each training epoch, but my version looks like fluctuate), this is my version:\n```\nimport torch\nimport torch.nn as nn\nimport torch.nn.functional as F\n\n\nclass WaveBlock(nn.Module):\n\n    def __init__(self, in_channels, out_channels, kernel_size, dilation_rates):\n        super(WaveBlock, self).__init__()\n        self.num_rates = dilation_rates\n        self.convs = nn.ModuleList()\n        self.filter_convs = nn.ModuleList()\n        self.gate_convs = nn.ModuleList()\n\n        self.convs.append(nn.Conv1d(in_channels, out_channels, kernel_size=1))\n        dilation_rates = [2 ** i for i in range(dilation_rates)]\n        for dilation_rate in dilation_rates:\n            self.filter_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.gate_convs.append(\n                nn.Conv1d(out_channels, out_channels, kernel_size=kernel_size, padding=int((dilation_rate*(kernel_size-1))/2), dilation=dilation_rate))\n            self.convs.append(nn.Conv1d(out_channels, out_channels, kernel_size=1))\n\n    def forward(self, x):\n        x = self.convs[0](x)\n        res = x\n        for i in range(self.num_rates):\n            x = torch.tanh(self.filter_convs[i](x)) * torch.sigmoid(self.gate_convs[i](x))\n            x = self.convs[i + 1](x)\n            res = res + x\n        return res\n\nclass WaveNet(nn.Module):\n\n    def __init__(self, chn, n_class):\n        super(WaveNet, self).__init__()\n        self.wave_block1 = WaveBlock(chn, 8, 3, 12)\n        self.wave_block2 = WaveBlock(8, 16, 3, 8)\n        self.wave_block3 = WaveBlock(16, 32, 3, 4)\n        self.wave_block4 = WaveBlock(32, 64, 3, 1)\n        self.avg_pool = nn.AdaptiveAvgPool1d(1)\n        self.fc1 = nn.Linear(256, 64)\n        self.fc2 = nn.Linear(64, n_class)\n\n    def feature_extra(self, x):\n        x = self.wave_block1(x)\n        x = self.wave_block2(x)\n        x = self.wave_block3(x)\n        x = self.wave_block4(x)\n        return x\n\n    def forward(self, x):\n\n        # LEFT TEMPORAL CHAIN\n        x1 = self.feature_extra(x[:,:,0:1])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,1:2])\n        x2 = self.avg_pool(x2)\n        z1 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        # LEFT PARASAGITTAL CHAIN\n        x1 = self.feature_extra(x[:,:,2:3])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,3:4])\n        x2 = self.avg_pool(x2)\n        z2 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        # RIGHT PARASAGITTAL CHAIN\n        x1 = self.feature_extra(x[:,:,4:5])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,5:6])\n        x2 = self.avg_pool(x2)\n        z3 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        # RIGHT TEMPORAL CHAIN\n        x1 = self.feature_extra(x[:,:,6:7])\n        x1 = self.avg_pool(x1)\n        x2 = self.feature_extra(x[:,:,7:8])\n        x2 = self.avg_pool(x2)\n        z4 = torch.mean(torch.stack([x1, x2]), dim=0)\n\n        y = torch.cat([z1,z2,z3,z4], dim=1)\n        y = y.view(y.size(0), -1)\n        y = self.fc1(y)\n        y = torch.relu(y)\n        y = self.fc2(y)\n        y = torch.softmax(y, dim=1)\n\n        return y\n```",
    "2624947": "I remember that the loss function of PytTorch will be processed by softmax itself, I think it is not necessary to use softmax in the network. In addition, I would like to ask how long it takes you to train one epoch, my one epoch takes more than ten minutes, it seems that there is a problem with my training code rather than the model code. Thank you very much for your answer.☺️",
    "2625458": "If I remember correctly, with dilated convolutions you should set `torch.backends.cudnn.benchmark = True`, \notherwise inefficient kernel would be chosen",
    "2625614": "Thank you very much for your suggestion. I have tried it but it doesn't work.😭"
  },
  "source": "meta"
}