{
  "id": 229666,
  "title": "Marginal Improvements using \"Iterative Training\" of Resnet18",
  "url": "/competitions/herbarium-2021-fgvc8/discussion/229666",
  "author_name": "Saurav Maheshkar ☕️",
  "post_date": "2021-03-31T06:25:17.002000",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I had published a kernel titled <a href=\"https://www.kaggle.com/sauravmaheshkar/herbarium-2021-pytorch-starter-weights-biases\" target=\"_blank\">\"Herbarium 2021:Pytorch 🔥 Starter + Weights&amp;Biases\"</a>, in which I'd used a Resnet18 for training using a GPU and then a separate kernel for running Inference. </p>\n<p>I'd also experimented with using the old weights as checkpoints to continue training (\"Iterative Training\"), and obtained a significant jump in the 2nd iteration while the subsequent provided marginal improvements in the f1 score. This was done because even with <code>n_epochs</code> set to <code>1</code>, the kernel took 6 hrs on a GPU. </p>\n<p>Have a look at these metrics for reference. Or have a look at this <a href=\"https://wandb.ai/sauravmaheshkar/Herbarium%202021\" target=\"_blank\">Weights and Biases Project Page</a></p>\n<p><img src=\"https://imgur.com/3CaTjYa.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/voiTQnQ.png\" alt=\"\"></p>\n<p>Seems like this is as far as I can go with a simple architecture like Resnet18😅. The next step would be to work with either EfficientNet or Resnet101. It'd also be worth trying out a Bottleneck Layer followed by a Dense Layer instead of an Average Pooling layer followed by a Dense Layer.</p>\n<p>PS: Is this technique actually called Iterative Training ?? </p>",
  "messages": [
    {
      "id": 1258182,
      "postDate": "2021-03-31T12:10:46.630Z",
      "content": "<p>Maybe \"cold restart\" training? Though that doesn't seem to be a term used either. In the same vein as warm restarts for LR, maybe you didn't load your optimizer state from the old weights and it \"mimic'd\" what warm restarting your learning rate would do? This is just speculation on my part though since I don't know much of the details.</p>\n<p>If they're weights from other datasets though that would be closer to transfer learning/finetuning (I'm sure you already know that).</p>\n<p>Also, yes I think you'll get quite the improvement on EFNet even with the smaller versions.</p>",
      "rawMarkdown": "Maybe \"cold restart\" training? Though that doesn't seem to be a term used either. In the same vein as warm restarts for LR, maybe you didn't load your optimizer state from the old weights and it \"mimic'd\" what warm restarting your learning rate would do? This is just speculation on my part though since I don't know much of the details.\n\nIf they're weights from other datasets though that would be closer to transfer learning/finetuning (I'm sure you already know that).\n\nAlso, yes I think you'll get quite the improvement on EFNet even with the smaller versions.",
      "votes": 1
    },
    {
      "id": 1257841,
      "postDate": "2021-03-31T06:25:17.003Z",
      "content": "<p>I had published a kernel titled <a href=\"https://www.kaggle.com/sauravmaheshkar/herbarium-2021-pytorch-starter-weights-biases\" target=\"_blank\">\"Herbarium 2021:Pytorch 🔥 Starter + Weights&amp;Biases\"</a>, in which I'd used a Resnet18 for training using a GPU and then a separate kernel for running Inference. </p>\n<p>I'd also experimented with using the old weights as checkpoints to continue training (\"Iterative Training\"), and obtained a significant jump in the 2nd iteration while the subsequent provided marginal improvements in the f1 score. This was done because even with <code>n_epochs</code> set to <code>1</code>, the kernel took 6 hrs on a GPU. </p>\n<p>Have a look at these metrics for reference. Or have a look at this <a href=\"https://wandb.ai/sauravmaheshkar/Herbarium%202021\" target=\"_blank\">Weights and Biases Project Page</a></p>\n<p><img src=\"https://imgur.com/3CaTjYa.png\" alt=\"\"></p>\n<p><img src=\"https://i.imgur.com/voiTQnQ.png\" alt=\"\"></p>\n<p>Seems like this is as far as I can go with a simple architecture like Resnet18😅. The next step would be to work with either EfficientNet or Resnet101. It'd also be worth trying out a Bottleneck Layer followed by a Dense Layer instead of an Average Pooling layer followed by a Dense Layer.</p>\n<p>PS: Is this technique actually called Iterative Training ?? </p>",
      "rawMarkdown": "I had published a kernel titled [\"Herbarium 2021:Pytorch 🔥 Starter + Weights&Biases\"](https://www.kaggle.com/sauravmaheshkar/herbarium-2021-pytorch-starter-weights-biases), in which I'd used a Resnet18 for training using a GPU and then a separate kernel for running Inference. \n\nI'd also experimented with using the old weights as checkpoints to continue training (\"Iterative Training\"), and obtained a significant jump in the 2nd iteration while the subsequent provided marginal improvements in the f1 score. This was done because even with `n_epochs` set to `1`, the kernel took 6 hrs on a GPU. \n\nHave a look at these metrics for reference. Or have a look at this [Weights and Biases Project Page](https://wandb.ai/sauravmaheshkar/Herbarium%202021)\n\n![](https://imgur.com/3CaTjYa.png)\n\n![](https://i.imgur.com/voiTQnQ.png)\n\nSeems like this is as far as I can go with a simple architecture like Resnet18😅. The next step would be to work with either EfficientNet or Resnet101. It'd also be worth trying out a Bottleneck Layer followed by a Dense Layer instead of an Average Pooling layer followed by a Dense Layer.\n\nPS: Is this technique actually called Iterative Training ?? ",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 1258182,
      "author_name": "Dax Ledesma",
      "author_url": "",
      "post_date": "2021-03-31T12:10:46.630000",
      "content": "<p>Maybe \"cold restart\" training? Though that doesn't seem to be a term used either. In the same vein as warm restarts for LR, maybe you didn't load your optimizer state from the old weights and it \"mimic'd\" what warm restarting your learning rate would do? This is just speculation on my part though since I don't know much of the details.</p>\n<p>If they're weights from other datasets though that would be closer to transfer learning/finetuning (I'm sure you already know that).</p>\n<p>Also, yes I think you'll get quite the improvement on EFNet even with the smaller versions.</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1258182": "Maybe \"cold restart\" training? Though that doesn't seem to be a term used either. In the same vein as warm restarts for LR, maybe you didn't load your optimizer state from the old weights and it \"mimic'd\" what warm restarting your learning rate would do? This is just speculation on my part though since I don't know much of the details.\n\nIf they're weights from other datasets though that would be closer to transfer learning/finetuning (I'm sure you already know that).\n\nAlso, yes I think you'll get quite the improvement on EFNet even with the smaller versions.",
    "1257841": "I had published a kernel titled [\"Herbarium 2021:Pytorch 🔥 Starter + Weights&Biases\"](https://www.kaggle.com/sauravmaheshkar/herbarium-2021-pytorch-starter-weights-biases), in which I'd used a Resnet18 for training using a GPU and then a separate kernel for running Inference. \n\nI'd also experimented with using the old weights as checkpoints to continue training (\"Iterative Training\"), and obtained a significant jump in the 2nd iteration while the subsequent provided marginal improvements in the f1 score. This was done because even with `n_epochs` set to `1`, the kernel took 6 hrs on a GPU. \n\nHave a look at these metrics for reference. Or have a look at this [Weights and Biases Project Page](https://wandb.ai/sauravmaheshkar/Herbarium%202021)\n\n![](https://imgur.com/3CaTjYa.png)\n\n![](https://i.imgur.com/voiTQnQ.png)\n\nSeems like this is as far as I can go with a simple architecture like Resnet18😅. The next step would be to work with either EfficientNet or Resnet101. It'd also be worth trying out a Bottleneck Layer followed by a Dense Layer instead of an Average Pooling layer followed by a Dense Layer.\n\nPS: Is this technique actually called Iterative Training ?? "
  }
}