{
  "id": 476899,
  "title": "Reproducibility of Chris' WaveNet Baseline in torch and Randomness Study",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/476899",
  "author_name": "",
  "post_date": "2024-02-14T00:22:03.812709400Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>I joined this competition late, and started from reproducing the amazing <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter</a> shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In this topic, I want to share some of my findings, related forums and extended randomness study.</p>\n<h2>Overview</h2>\n<p>In the past few days, I found it hard to reproduce the CV score of Chris' WaveNet baseline implemented in <code>torch</code>. After many trials and checking, I thought that all the settings were exactly the same. But, my best CV score was around 0.91, which looks far from Chris' 0.81. Then, I saw the forum posted by <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a> <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/476550\" target=\"_blank\">here</a>, discussing the runtime performance drop of dilated convolution in <code>torch</code>. Finally, he provided me with a <a href=\"https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764\" target=\"_blank\">solution</a>, which helped me succeed to reproduce Chris' CV score! However, the CV scheme (gpkf on <code>patient_id</code>) seems unstable here…</p>\n<h2>Issues and Solutions</h2>\n<h3><em>Runtime Performance Drop</em></h3>\n<p>As experienced by many, I also find that runtime performance drops a lot for dilated convolution in <code>torch</code>. A quick workaround is to use the following setting,</p>\n<pre><code> = \n</code></pre>\n<p>, which is related to <a href=\"https://pytorch.org/docs/stable/notes/randomness.html#cuda-convolution-determinism\" target=\"_blank\">convolution determinism</a>.</p>\n<h3><em>CV Score Gap</em></h3>\n<p>As mentioned above, I couldn't get CV score close to the one shown in keras implementation. After changing weight initialization of convolutions in each <code>wave_block</code> following <a href=\"https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764\" target=\"_blank\">this solution</a>, I succeeded to get CV score close to Chris'. The code snippet is provided as follows,</p>\n<pre><code>\n\nnn.init.xavier_uniform_(self.in_conv.weight, gain=nn.init.calculate_gain())\nnn.init.zeros_(self.in_conv.bias)\n\n\n i  ((self.filts)):\n    nn.init.xavier_uniform_(self.filts[i].weight, gain=nn.init.calculate_gain())\n    nn.init.zeros_(self.filts[i].bias)\n\n\n i  ((self.gates)):\n    nn.init.xavier_uniform_(self.gates[i].weight, gain=nn.init.calculate_gain())\n    nn.init.zeros_(self.gates[i].bias)\n\n\n i  ((self.skip_convs)):\n    nn.init.xavier_uniform_(self.skip_convs[i].weight, gain=nn.init.calculate_gain())\n    nn.init.zeros_(self.skip_convs[i].bias)\n</code></pre>\n<h3><em>Randomness Study</em></h3>\n<p>To see if CV score is stable, I run experiments with exactly the same settings but alter the random seed (10 seeds in total). Following illustrates the experimental results,<br>\n<a href=\"https://postimg.cc/CBxbXg2S\" target=\"_blank\"><img src=\"https://i.postimg.cc/G2kzB3J8/Screenshot-2024-02-14-at-12-34-44-AM.png\" alt=\"Screenshot-2024-02-14-at-12-34-44-AM.png\"></a> <br>\nAs can be seen, CV score is unstable from two perspective,</p>\n<ol>\n<li>Data sampling<br>\nValidation performance varies fold-by-fold and fold2 performs much better than others.</li>\n<li>Training process<br>\nConsidering different random seeds lead to different weight initialization, dataloader iteration, etc., we also observe large variation of validation scores within each single fold (fold score standard deviation marked in red), especially fold1.</li>\n</ol>\n<p>Hence, we can take this randomness factor into account when considering CV improvement. Furthermore, trying to find a stable CV scheme is also important.</p>\n<p>It's the first time I try to do this kind of study. If there's any mistake, please let me know. Thanks for reading!</p>",
  "messages": [
    {
      "id": "2651122",
      "postDate": "02/14/2024 00:22:03",
      "content": "<p>Hi everyone,</p>\n<p>I joined this competition late, and started from reproducing the amazing <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52\" target=\"_blank\">WaveNet Starter</a> shared by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a>. In this topic, I want to share some of my findings, related forums and extended randomness study.</p>\n<h2>Overview</h2>\n<p>In the past few days, I found it hard to reproduce the CV score of Chris' WaveNet baseline implemented in <code>torch</code>. After many trials and checking, I thought that all the settings were exactly the same. But, my best CV score was around 0.91, which looks far from Chris' 0.81. Then, I saw the forum posted by <a href=\"https://www.kaggle.com/chaudharypriyanshu\" target=\"_blank\">@chaudharypriyanshu</a> <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/476550\" target=\"_blank\">here</a>, discussing the runtime performance drop of dilated convolution in <code>torch</code>. Finally, he provided me with a <a href=\"https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764\" target=\"_blank\">solution</a>, which helped me succeed to reproduce Chris' CV score! However, the CV scheme (gpkf on <code>patient_id</code>) seems unstable here…</p>\n<h2>Issues and Solutions</h2>\n<h3><em>Runtime Performance Drop</em></h3>\n<p>As experienced by many, I also find that runtime performance drops a lot for dilated convolution in <code>torch</code>. A quick workaround is to use the following setting,</p>\n<pre><code> = \n</code></pre>\n<p>, which is related to <a href=\"https://pytorch.org/docs/stable/notes/randomness.html#cuda-convolution-determinism\" target=\"_blank\">convolution determinism</a>.</p>\n<h3><em>CV Score Gap</em></h3>\n<p>As mentioned above, I couldn't get CV score close to the one shown in keras implementation. After changing weight initialization of convolutions in each <code>wave_block</code> following <a href=\"https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764\" target=\"_blank\">this solution</a>, I succeeded to get CV score close to Chris'. The code snippet is provided as follows,</p>\n<pre><code>\n\nnn.init.xavier_uniform_(self.in_conv.weight, gain=nn.init.calculate_gain())\nnn.init.zeros_(self.in_conv.bias)\n\n\n i  ((self.filts)):\n    nn.init.xavier_uniform_(self.filts[i].weight, gain=nn.init.calculate_gain())\n    nn.init.zeros_(self.filts[i].bias)\n\n\n i  ((self.gates)):\n    nn.init.xavier_uniform_(self.gates[i].weight, gain=nn.init.calculate_gain())\n    nn.init.zeros_(self.gates[i].bias)\n\n\n i  ((self.skip_convs)):\n    nn.init.xavier_uniform_(self.skip_convs[i].weight, gain=nn.init.calculate_gain())\n    nn.init.zeros_(self.skip_convs[i].bias)\n</code></pre>\n<h3><em>Randomness Study</em></h3>\n<p>To see if CV score is stable, I run experiments with exactly the same settings but alter the random seed (10 seeds in total). Following illustrates the experimental results,<br>\n<a href=\"https://postimg.cc/CBxbXg2S\" target=\"_blank\"><img src=\"https://i.postimg.cc/G2kzB3J8/Screenshot-2024-02-14-at-12-34-44-AM.png\" alt=\"Screenshot-2024-02-14-at-12-34-44-AM.png\"></a> <br>\nAs can be seen, CV score is unstable from two perspective,</p>\n<ol>\n<li>Data sampling<br>\nValidation performance varies fold-by-fold and fold2 performs much better than others.</li>\n<li>Training process<br>\nConsidering different random seeds lead to different weight initialization, dataloader iteration, etc., we also observe large variation of validation scores within each single fold (fold score standard deviation marked in red), especially fold1.</li>\n</ol>\n<p>Hence, we can take this randomness factor into account when considering CV improvement. Furthermore, trying to find a stable CV scheme is also important.</p>\n<p>It's the first time I try to do this kind of study. If there's any mistake, please let me know. Thanks for reading!</p>",
      "rawMarkdown": "Hi everyone,\n\nI joined this competition late, and started from reproducing the amazing [WaveNet Starter](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52) shared by @cdeotte. In this topic, I want to share some of my findings, related forums and extended randomness study.\n\n## Overview\nIn the past few days, I found it hard to reproduce the CV score of Chris' WaveNet baseline implemented in `torch`. After many trials and checking, I thought that all the settings were exactly the same. But, my best CV score was around 0.91, which looks far from Chris' 0.81. Then, I saw the forum posted by @chaudharypriyanshu [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/476550), discussing the runtime performance drop of dilated convolution in `torch`. Finally, he provided me with a [solution](https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764), which helped me succeed to reproduce Chris' CV score! However, the CV scheme (gpkf on `patient_id`) seems unstable here...\n\n## Issues and Solutions\n### *Runtime Performance Drop*\nAs experienced by many, I also find that runtime performance drops a lot for dilated convolution in `torch`. A quick workaround is to use the following setting,\n```\ntorch.backends.cudnn.deterministic = False\n```\n, which is related to [convolution determinism](https://pytorch.org/docs/stable/notes/randomness.html#cuda-convolution-determinism).\n\n### *CV Score Gap*\nAs mentioned above, I couldn't get CV score close to the one shown in keras implementation. After changing weight initialization of convolutions in each `wave_block` following [this solution](https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764), I succeeded to get CV score close to Chris'. The code snippet is provided as follows,\n```python\n# For each _WaveBlock\n# Input convolution\nnn.init.xavier_uniform_(self.in_conv.weight, gain=nn.init.calculate_gain('relu'))\nnn.init.zeros_(self.in_conv.bias)\n\n# Filters (tanh)\nfor i in range(len(self.filts)):\n    nn.init.xavier_uniform_(self.filts[i].weight, gain=nn.init.calculate_gain('relu'))\n    nn.init.zeros_(self.filts[i].bias)\n\n# Gates (sigmoid)\nfor i in range(len(self.gates)):\n    nn.init.xavier_uniform_(self.gates[i].weight, gain=nn.init.calculate_gain('relu'))\n    nn.init.zeros_(self.gates[i].bias)\n\n# Skip-connection\nfor i in range(len(self.skip_convs)):\n    nn.init.xavier_uniform_(self.skip_convs[i].weight, gain=nn.init.calculate_gain('relu'))\n    nn.init.zeros_(self.skip_convs[i].bias)\n```\n\n### *Randomness Study*\nTo see if CV score is stable, I run experiments with exactly the same settings but alter the random seed (10 seeds in total). Following illustrates the experimental results,\n[![Screenshot-2024-02-14-at-12-34-44-AM.png](https://i.postimg.cc/G2kzB3J8/Screenshot-2024-02-14-at-12-34-44-AM.png)](https://postimg.cc/CBxbXg2S) \nAs can be seen, CV score is unstable from two perspective,\n1. Data sampling\nValidation performance varies fold-by-fold and fold2 performs much better than others.\n2. Training process\nConsidering different random seeds lead to different weight initialization, dataloader iteration, etc., we also observe large variation of validation scores within each single fold (fold score standard deviation marked in red), especially fold1.\n\nHence, we can take this randomness factor into account when considering CV improvement. Furthermore, trying to find a stable CV scheme is also important.\n\nIt's the first time I try to do this kind of study. If there's any mistake, please let me know. Thanks for reading!",
      "votes": null
    },
    {
      "id": "2651137",
      "postDate": "02/14/2024 01:17:15",
      "content": "<p>Thanks! I've also been having trouble replicating Chris' WaveNet model in torch. Initial weights is one of the only things I haven't tried.</p>",
      "rawMarkdown": "Thanks! I've also been having trouble replicating Chris' WaveNet model in torch. Initial weights is one of the only things I haven't tried.",
      "votes": null
    },
    {
      "id": "2651143",
      "postDate": "02/14/2024 01:22:56",
      "content": "<p>I've also noticed instability in CV score (also if you use diff seeds to produce different folds). Considering that LB scores seem to generally be much lower, I wonder how generally reliable differences in CV scores are in predicting differences in LB scores.</p>",
      "rawMarkdown": "I've also noticed instability in CV score (also if you use diff seeds to produce different folds). Considering that LB scores seem to generally be much lower, I wonder how generally reliable differences in CV scores are in predicting differences in LB scores.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2651137,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "02/14/2024 01:17:15",
      "content": "<p>Thanks! I've also been having trouble replicating Chris' WaveNet model in torch. Initial weights is one of the only things I haven't tried.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2651143,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "02/14/2024 01:22:56",
      "content": "<p>I've also noticed instability in CV score (also if you use diff seeds to produce different folds). Considering that LB scores seem to generally be much lower, I wonder how generally reliable differences in CV scores are in predicting differences in LB scores.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2651122": "Hi everyone,\n\nI joined this competition late, and started from reproducing the amazing [WaveNet Starter](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-52) shared by @cdeotte. In this topic, I want to share some of my findings, related forums and extended randomness study.\n\n## Overview\nIn the past few days, I found it hard to reproduce the CV score of Chris' WaveNet baseline implemented in `torch`. After many trials and checking, I thought that all the settings were exactly the same. But, my best CV score was around 0.91, which looks far from Chris' 0.81. Then, I saw the forum posted by @chaudharypriyanshu [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/476550), discussing the runtime performance drop of dilated convolution in `torch`. Finally, he provided me with a [solution](https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764), which helped me succeed to reproduce Chris' CV score! However, the CV scheme (gpkf on `patient_id`) seems unstable here...\n\n## Issues and Solutions\n### *Runtime Performance Drop*\nAs experienced by many, I also find that runtime performance drops a lot for dilated convolution in `torch`. A quick workaround is to use the following setting,\n```\ntorch.backends.cudnn.deterministic = False\n```\n, which is related to [convolution determinism](https://pytorch.org/docs/stable/notes/randomness.html#cuda-convolution-determinism).\n\n### *CV Score Gap*\nAs mentioned above, I couldn't get CV score close to the one shown in keras implementation. After changing weight initialization of convolutions in each `wave_block` following [this solution](https://www.kaggle.com/competitions/liverpool-ion-switching/discussion/145256#863764), I succeeded to get CV score close to Chris'. The code snippet is provided as follows,\n```python\n# For each _WaveBlock\n# Input convolution\nnn.init.xavier_uniform_(self.in_conv.weight, gain=nn.init.calculate_gain('relu'))\nnn.init.zeros_(self.in_conv.bias)\n\n# Filters (tanh)\nfor i in range(len(self.filts)):\n    nn.init.xavier_uniform_(self.filts[i].weight, gain=nn.init.calculate_gain('relu'))\n    nn.init.zeros_(self.filts[i].bias)\n\n# Gates (sigmoid)\nfor i in range(len(self.gates)):\n    nn.init.xavier_uniform_(self.gates[i].weight, gain=nn.init.calculate_gain('relu'))\n    nn.init.zeros_(self.gates[i].bias)\n\n# Skip-connection\nfor i in range(len(self.skip_convs)):\n    nn.init.xavier_uniform_(self.skip_convs[i].weight, gain=nn.init.calculate_gain('relu'))\n    nn.init.zeros_(self.skip_convs[i].bias)\n```\n\n### *Randomness Study*\nTo see if CV score is stable, I run experiments with exactly the same settings but alter the random seed (10 seeds in total). Following illustrates the experimental results,\n[![Screenshot-2024-02-14-at-12-34-44-AM.png](https://i.postimg.cc/G2kzB3J8/Screenshot-2024-02-14-at-12-34-44-AM.png)](https://postimg.cc/CBxbXg2S) \nAs can be seen, CV score is unstable from two perspective,\n1. Data sampling\nValidation performance varies fold-by-fold and fold2 performs much better than others.\n2. Training process\nConsidering different random seeds lead to different weight initialization, dataloader iteration, etc., we also observe large variation of validation scores within each single fold (fold score standard deviation marked in red), especially fold1.\n\nHence, we can take this randomness factor into account when considering CV improvement. Furthermore, trying to find a stable CV scheme is also important.\n\nIt's the first time I try to do this kind of study. If there's any mistake, please let me know. Thanks for reading!",
    "2651137": "Thanks! I've also been having trouble replicating Chris' WaveNet model in torch. Initial weights is one of the only things I haven't tried.",
    "2651143": "I've also noticed instability in CV score (also if you use diff seeds to produce different folds). Considering that LB scores seem to generally be much lower, I wonder how generally reliable differences in CV scores are in predicting differences in LB scores."
  },
  "source": "meta"
}