{
  "id": 245757,
  "title": "Single Fold 0.97+: Custom Head + GradualWarmup",
  "url": "/competitions/seti-breakthrough-listen/discussion/245757",
  "author_name": "gao-hongnan",
  "post_date": "2021-06-12T08:15:30.850000",
  "votes": 11,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Here is my faithful <a href=\"https://www.kaggle.com/reighns/custom-head-gradual-warmup-single-fold-0-97\" target=\"_blank\">PyTorch Pipeline</a> that I use for the past half a year. It has served me well because I can literally plug and play in the <code>config</code> file (or as a dict in notebook). Note that I set <code>DEBUG</code> in the <code>config</code> to be <code>True</code>. Please turn it off and change <code>EPOCHS</code> when you want to train.</p>\n<p>This is just a single fold for <code>efficientnet_b0</code>, trained for 20 epochs. I did not wait for it to converge due to limited resources. This single-fold is also top of the public baseline for the 0.97. I think if you train all folds and wait for more epochs it will converge to a better score. If you ensemble all folds it should reach 0.98.</p>\n<p>Something to highlight here is I used a custom head with <code>swish</code> activation, I do not think this additional layer will make the network any more complicated than it is, but it does provide me with a stable training process across models.</p>\n<pre><code>        self.single_head_fc = torch.nn.Sequential(\n            torch.nn.Linear(self.in_features, self.in_features),\n            self.activation,\n            torch.nn.Dropout(p=0.5),\n            torch.nn.Linear(self.in_features, self.config[\"DATA\"][\"NUM_CLASSES\"]),\n        )\n</code></pre>\n<p>I also used a custom scheduler <code>GradualWarmupSchedulerV2</code> that <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> and (qishen)'s used before.</p>",
  "messages": [
    {
      "id": 1346238,
      "postDate": "2021-06-12T08:15:30.850Z",
      "content": "<p>Here is my faithful <a href=\"https://www.kaggle.com/reighns/custom-head-gradual-warmup-single-fold-0-97\" target=\"_blank\">PyTorch Pipeline</a> that I use for the past half a year. It has served me well because I can literally plug and play in the <code>config</code> file (or as a dict in notebook). Note that I set <code>DEBUG</code> in the <code>config</code> to be <code>True</code>. Please turn it off and change <code>EPOCHS</code> when you want to train.</p>\n<p>This is just a single fold for <code>efficientnet_b0</code>, trained for 20 epochs. I did not wait for it to converge due to limited resources. This single-fold is also top of the public baseline for the 0.97. I think if you train all folds and wait for more epochs it will converge to a better score. If you ensemble all folds it should reach 0.98.</p>\n<p>Something to highlight here is I used a custom head with <code>swish</code> activation, I do not think this additional layer will make the network any more complicated than it is, but it does provide me with a stable training process across models.</p>\n<pre><code>        self.single_head_fc = torch.nn.Sequential(\n            torch.nn.Linear(self.in_features, self.in_features),\n            self.activation,\n            torch.nn.Dropout(p=0.5),\n            torch.nn.Linear(self.in_features, self.config[\"DATA\"][\"NUM_CLASSES\"]),\n        )\n</code></pre>\n<p>I also used a custom scheduler <code>GradualWarmupSchedulerV2</code> that <a href=\"https://www.kaggle.com/underwearfitting\" target=\"_blank\">@underwearfitting</a> and (qishen)'s used before.</p>",
      "rawMarkdown": "Here is my faithful [PyTorch Pipeline](https://www.kaggle.com/reighns/custom-head-gradual-warmup-single-fold-0-97) that I use for the past half a year. It has served me well because I can literally plug and play in the `config` file (or as a dict in notebook). Note that I set `DEBUG` in the `config` to be `True`. Please turn it off and change `EPOCHS` when you want to train.\n\nThis is just a single fold for `efficientnet_b0`, trained for 20 epochs. I did not wait for it to converge due to limited resources. This single-fold is also top of the public baseline for the 0.97. I think if you train all folds and wait for more epochs it will converge to a better score. If you ensemble all folds it should reach 0.98.\n\nSomething to highlight here is I used a custom head with `swish` activation, I do not think this additional layer will make the network any more complicated than it is, but it does provide me with a stable training process across models.\n\n```\n        self.single_head_fc = torch.nn.Sequential(\n            torch.nn.Linear(self.in_features, self.in_features),\n            self.activation,\n            torch.nn.Dropout(p=0.5),\n            torch.nn.Linear(self.in_features, self.config[\"DATA\"][\"NUM_CLASSES\"]),\n        )\n```\n\nI also used a custom scheduler `GradualWarmupSchedulerV2` that @underwearfitting and (qishen)'s used before.",
      "votes": 11
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1346238": "Here is my faithful [PyTorch Pipeline](https://www.kaggle.com/reighns/custom-head-gradual-warmup-single-fold-0-97) that I use for the past half a year. It has served me well because I can literally plug and play in the `config` file (or as a dict in notebook). Note that I set `DEBUG` in the `config` to be `True`. Please turn it off and change `EPOCHS` when you want to train.\n\nThis is just a single fold for `efficientnet_b0`, trained for 20 epochs. I did not wait for it to converge due to limited resources. This single-fold is also top of the public baseline for the 0.97. I think if you train all folds and wait for more epochs it will converge to a better score. If you ensemble all folds it should reach 0.98.\n\nSomething to highlight here is I used a custom head with `swish` activation, I do not think this additional layer will make the network any more complicated than it is, but it does provide me with a stable training process across models.\n\n```\n        self.single_head_fc = torch.nn.Sequential(\n            torch.nn.Linear(self.in_features, self.in_features),\n            self.activation,\n            torch.nn.Dropout(p=0.5),\n            torch.nn.Linear(self.in_features, self.config[\"DATA\"][\"NUM_CLASSES\"]),\n        )\n```\n\nI also used a custom scheduler `GradualWarmupSchedulerV2` that @underwearfitting and (qishen)'s used before."
  }
}