{
  "id": 204951,
  "title": "How many epochs in one fold you use?",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/204951",
  "author_name": "",
  "post_date": "2020-12-17T16:48:22.646387500Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi,everyone use how many epochs in one fold? 10 or other?</p>",
  "messages": [
    {
      "id": "1117023",
      "postDate": "12/17/2020 16:48:22",
      "content": "<p>Hi,everyone use how many epochs in one fold? 10 or other?</p>",
      "rawMarkdown": "Hi,everyone use how many epochs in one fold? 10 or other?",
      "votes": null
    },
    {
      "id": "1117588",
      "postDate": "12/18/2020 08:28:19",
      "content": "<p>For my best LB score (0.945), I <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">used (link to shared notebook)</a> 10 epochs and it actually looks like it is a little underfit. I'm mostly saying that because training loss was still a good bit lower than the validation loss. So, perhaps it would have been better to do 15 to 20 (that's just a guess based on other competitions, but 20 was used in <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\" target=\"_blank\">another high scoring notebook</a>). The optimal number of epochs of course depends on a lot of things: whether you use a pre-trained model (otherwise more, not using a pre-trained model is probably not a good idea here though), how many layers you put on top of a pre-trained model, how much augmentation you use (heavy augmentation and e.g. Mix-up - not that that necessarily makes sense here - tends to need more epochs), how much normalization one uses (e.g. larger and more final layers with lots of drop-out presumably need a bit longer to train), and the learning rate schedule. In the end, it's a hyper-parameter that we can take a good guess on by looking at how training vs. validation loss and other metrics develop when fitting and then we can confirm the choice vs. other choices using CV (just like for any other hyper parameter).</p>",
      "rawMarkdown": "For my best LB score (0.945), I [used (link to shared notebook)](https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb) 10 epochs and it actually looks like it is a little underfit. I'm mostly saying that because training loss was still a good bit lower than the validation loss. So, perhaps it would have been better to do 15 to 20 (that's just a guess based on other competitions, but 20 was used in [another high scoring notebook](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training)). The optimal number of epochs of course depends on a lot of things: whether you use a pre-trained model (otherwise more, not using a pre-trained model is probably not a good idea here though), how many layers you put on top of a pre-trained model, how much augmentation you use (heavy augmentation and e.g. Mix-up - not that that necessarily makes sense here - tends to need more epochs), how much normalization one uses (e.g. larger and more final layers with lots of drop-out presumably need a bit longer to train), and the learning rate schedule. In the end, it's a hyper-parameter that we can take a good guess on by looking at how training vs. validation loss and other metrics develop when fitting and then we can confirm the choice vs. other choices using CV (just like for any other hyper parameter).",
      "votes": null
    },
    {
      "id": "1120440",
      "postDate": "12/20/2020 20:11:38",
      "content": "<p>I have used Early Stopping and it mostly stops around 15 and it seems good enough for my model.</p>",
      "rawMarkdown": "I have used Early Stopping and it mostly stops around 15 and it seems good enough for my model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1117588,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "12/18/2020 08:28:19",
      "content": "<p>For my best LB score (0.945), I <a href=\"https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb\" target=\"_blank\">used (link to shared notebook)</a> 10 epochs and it actually looks like it is a little underfit. I'm mostly saying that because training loss was still a good bit lower than the validation loss. So, perhaps it would have been better to do 15 to 20 (that's just a guess based on other competitions, but 20 was used in <a href=\"https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training\" target=\"_blank\">another high scoring notebook</a>). The optimal number of epochs of course depends on a lot of things: whether you use a pre-trained model (otherwise more, not using a pre-trained model is probably not a good idea here though), how many layers you put on top of a pre-trained model, how much augmentation you use (heavy augmentation and e.g. Mix-up - not that that necessarily makes sense here - tends to need more epochs), how much normalization one uses (e.g. larger and more final layers with lots of drop-out presumably need a bit longer to train), and the learning rate schedule. In the end, it's a hyper-parameter that we can take a good guess on by looking at how training vs. validation loss and other metrics develop when fitting and then we can confirm the choice vs. other choices using CV (just like for any other hyper parameter).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1120440,
      "author_name": "harshsdw",
      "author_url": "",
      "post_date": "12/20/2020 20:11:38",
      "content": "<p>I have used Early Stopping and it mostly stops around 15 and it seems good enough for my model.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1117023": "Hi,everyone use how many epochs in one fold? 10 or other?",
    "1117588": "For my best LB score (0.945), I [used (link to shared notebook)](https://www.kaggle.com/bjoernholzhauer/fastai-how-to-set-up-efficientnet-b4-0-945-lb) 10 epochs and it actually looks like it is a little underfit. I'm mostly saying that because training loss was still a good bit lower than the validation loss. So, perhaps it would have been better to do 15 to 20 (that's just a guess based on other competitions, but 20 was used in [another high scoring notebook](https://www.kaggle.com/xhlulu/ranzcr-efficientnet-tpu-training)). The optimal number of epochs of course depends on a lot of things: whether you use a pre-trained model (otherwise more, not using a pre-trained model is probably not a good idea here though), how many layers you put on top of a pre-trained model, how much augmentation you use (heavy augmentation and e.g. Mix-up - not that that necessarily makes sense here - tends to need more epochs), how much normalization one uses (e.g. larger and more final layers with lots of drop-out presumably need a bit longer to train), and the learning rate schedule. In the end, it's a hyper-parameter that we can take a good guess on by looking at how training vs. validation loss and other metrics develop when fitting and then we can confirm the choice vs. other choices using CV (just like for any other hyper parameter).",
    "1120440": "I have used Early Stopping and it mostly stops around 15 and it seems good enough for my model."
  },
  "source": "meta"
}