{
  "id": 301295,
  "title": "How long do you train? - Discussion on epoch number (or training resource for poor)",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/301295",
  "author_name": "",
  "post_date": "2022-01-16T23:59:32.105896300Z",
  "votes": 8,
  "comment_count": 15,
  "views": 0,
  "content": "<p>According to <a href=\"https://github.com/ultralytics/yolov5/wiki/Train-Custom-Data\" target=\"_blank\">YOLOv5 official</a>, 300 is recommended for default epoch number.<br>\nBut, partly due to data size or training setting, training saturates early. For some simple situation, epoch number ~30 is enough to get a best validation result.</p>\n<p>How do you set the training schedule including epoch number?<br>\nAlthough \"small epoch number is better.\" sounds like gospel for guys with poor resource (like me), it should be considered carefully.</p>\n<p>Thanks.</p>",
  "messages": [
    {
      "id": "1652730",
      "postDate": "01/16/2022 23:59:32",
      "content": "<p>According to <a href=\"https://github.com/ultralytics/yolov5/wiki/Train-Custom-Data\" target=\"_blank\">YOLOv5 official</a>, 300 is recommended for default epoch number.<br>\nBut, partly due to data size or training setting, training saturates early. For some simple situation, epoch number ~30 is enough to get a best validation result.</p>\n<p>How do you set the training schedule including epoch number?<br>\nAlthough \"small epoch number is better.\" sounds like gospel for guys with poor resource (like me), it should be considered carefully.</p>\n<p>Thanks.</p>",
      "rawMarkdown": "According to [YOLOv5 official](https://github.com/ultralytics/yolov5/wiki/Train-Custom-Data), 300 is recommended for default epoch number.\nBut, partly due to data size or training setting, training saturates early. For some simple situation, epoch number ~30 is enough to get a best validation result.\n\nHow do you set the training schedule including epoch number?\nAlthough \"small epoch number is better.\" sounds like gospel for guys with poor resource (like me), it should be considered carefully.\n\nThanks.",
      "votes": null
    },
    {
      "id": "1652874",
      "postDate": "01/17/2022 04:28:25",
      "content": "<p>with pretrained weight,  fine tuning around 10 epochs would get pretty good result.<br>\nfor better result,  GPU resource is a must. </p>",
      "rawMarkdown": "with pretrained weight,  fine tuning around 10 epochs would get pretty good result.\nfor better result,  GPU resource is a must.",
      "votes": null
    },
    {
      "id": "1652916",
      "postDate": "01/17/2022 05:24:32",
      "content": "<p>Thank you for your reply. I have a similar feel with you.</p>\n<p>I wonder: If we have much more epochs (100~), a model just gets overfitted even if along with strong data-augmentation methods.</p>",
      "rawMarkdown": "Thank you for your reply. I have a similar feel with you.\n\nI wonder: If we have much more epochs (100~), a model just gets overfitted even if along with strong data-augmentation methods.",
      "votes": null
    },
    {
      "id": "1652984",
      "postDate": "01/17/2022 06:35:50",
      "content": "<p>Keep in mind that the 300-epoch recommendation is for 80 objects in COCO. For 1 object, you don't need as many epochs.</p>\n<p>Train until validation loss does not improve with further cycles. It is hard to say how many is correct, as heavier augmentation or lower learning rate will lead to slower learning but may result in better mAP.</p>\n<p>Overtraining the model may result in higher precision and lower recall, which is bad for this competition as the F2 metric favors recall over precision.</p>",
      "rawMarkdown": "Keep in mind that the 300-epoch recommendation is for 80 objects in COCO. For 1 object, you don't need as many epochs.\n\nTrain until validation loss does not improve with further cycles. It is hard to say how many is correct, as heavier augmentation or lower learning rate will lead to slower learning but may result in better mAP.\n\nOvertraining the model may result in higher precision and lower recall, which is bad for this competition as the F2 metric favors recall over precision.",
      "votes": null
    },
    {
      "id": "1653018",
      "postDate": "01/17/2022 07:09:50",
      "content": "<p>Thanks a lot. Makes sense to me.<br>\nBecause I was skeptical about my understanding, your comment made me trustable. :)</p>\n<p>It is very hard to determine training schedule, such as epoch number, learning rate, and heaviness of DA method. To estimate delicate balance among these, I think trust CV method is most important.</p>",
      "rawMarkdown": "Thanks a lot. Makes sense to me.\nBecause I was skeptical about my understanding, your comment made me trustable. :)\n\nIt is very hard to determine training schedule, such as epoch number, learning rate, and heaviness of DA method. To estimate delicate balance among these, I think trust CV method is most important.",
      "votes": null
    },
    {
      "id": "1653076",
      "postDate": "01/17/2022 08:27:41",
      "content": "<p>How much epoch? It depends ….. You should look into metrics … and choose correct model (according to metrics and your goal). Just train yolov5 with <code>--save-period 1</code> and choose model.</p>",
      "rawMarkdown": "How much epoch? It depends ..... You should look into metrics ... and choose correct model (according to metrics and your goal). Just train yolov5 with `--save-period 1` and choose model.",
      "votes": null
    },
    {
      "id": "1653291",
      "postDate": "01/17/2022 12:13:27",
      "content": "<p>I am grateful for your kind advice. I agree with you.<br>\nI'll try using --save-period option.</p>",
      "rawMarkdown": "I am grateful for your kind advice. I agree with you.\nI'll try using --save-period option.",
      "votes": null
    },
    {
      "id": "1653298",
      "postDate": "01/17/2022 12:22:55",
      "content": "<p>Second one in fitness strategy in yolov5. You can automate process of \"best.pt\" model (<a href=\"https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/utils/metrics.py#L15-L19\" target=\"_blank\">https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/utils/metrics.py#L15-L19</a>)</p>\n<pre><code>def fitness(x):\n    w = [0.0, 0.0, 0.1, 0.9]  # weights for [P, R, mAP@0.5, mAP@0.5:0.95]\n    return (x[:, :4] * w).sum(1)\n</code></pre>\n<p>for really custom things you can modify code here: <a href=\"https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/train.py#L376-L395\" target=\"_blank\">https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/train.py#L376-L395</a></p>",
      "rawMarkdown": "Second one in fitness strategy in yolov5. You can automate process of \"best.pt\" model (https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/utils/metrics.py#L15-L19)\n\n```python\ndef fitness(x):\n    w = [0.0, 0.0, 0.1, 0.9]  # weights for [P, R, mAP@0.5, mAP@0.5:0.95]\n    return (x[:, :4] * w).sum(1)\n```\n\nfor really custom things you can modify code here: https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/train.py#L376-L395",
      "votes": null
    },
    {
      "id": "1653487",
      "postDate": "01/17/2022 15:19:37",
      "content": "<p>You can use YOLO's --evolve parameter to tune hyperparameters.</p>",
      "rawMarkdown": "You can use YOLO's --evolve parameter to tune hyperparameters.",
      "votes": null
    },
    {
      "id": "1653520",
      "postDate": "01/17/2022 15:55:49",
      "content": "<p>Great information.<br>\nAre there some recommendations about these weights?<br>\nIn GBR condition, recall may be worth to focus on, but it's hard to define better ratio;(</p>",
      "rawMarkdown": "Great information.\nAre there some recommendations about these weights?\nIn GBR condition, recall may be worth to focus on, but it's hard to define better ratio;(",
      "votes": null
    },
    {
      "id": "1653526",
      "postDate": "01/17/2022 16:01:54",
      "content": "<p>Thank you for your suggestion.<br>\nWould you tell me how expensive is it in computational cost?</p>",
      "rawMarkdown": "Thank you for your suggestion.\nWould you tell me how expensive is it in computational cost?",
      "votes": null
    },
    {
      "id": "1653605",
      "postDate": "01/17/2022 17:24:45",
      "content": "<p>It's the same compute as training. It just tries to tune params to the 'fitness' metrics.</p>\n<p>Here's the overview from Ultralytics -&gt; <a href=\"https://docs.ultralytics.com/tutorials/hyperparameter-evolution/\" target=\"_blank\">https://docs.ultralytics.com/tutorials/hyperparameter-evolution/</a></p>",
      "rawMarkdown": "It's the same compute as training. It just tries to tune params to the 'fitness' metrics.\n\nHere's the overview from Ultralytics -> https://docs.ultralytics.com/tutorials/hyperparameter-evolution/",
      "votes": null
    },
    {
      "id": "1653611",
      "postDate": "01/17/2022 17:31:00",
      "content": "<p>But …. evolve requires GPU resources (a lot) and good baseline model. You do not improve poor model (maybe a little bit) - it is bed strategy. So firstly work on dataset, create strong model and then tune.</p>",
      "rawMarkdown": "But …. evolve requires GPU resources (a lot) and good baseline model. You do not improve poor model (maybe a little bit) - it is bed strategy. So firstly work on dataset, create strong model and then tune.",
      "votes": null
    },
    {
      "id": "1653858",
      "postDate": "01/17/2022 23:40:02",
      "content": "<p>I thought evolve option was so expensive!<br>\nI'm thinking of challenge it if I get good models.</p>",
      "rawMarkdown": "I thought evolve option was so expensive!\nI'm thinking of challenge it if I get good models.",
      "votes": null
    },
    {
      "id": "1653868",
      "postDate": "01/17/2022 23:46:54",
      "content": "<p>Thank you for giving me a lot of advice on my question.<br>\nI say.. First thing's first. We have to hurry! 🔥</p>",
      "rawMarkdown": "Thank you for giving me a lot of advice on my question.\nI say.. First thing's first. We have to hurry! 🔥",
      "votes": null
    },
    {
      "id": "1653871",
      "postDate": "01/17/2022 23:55:15",
      "content": "<p>Evolve doesn't require more GPU memory, just more time, because it does multiple runs. I found it easier to smooth out my training with evolve than by manually trying random hyperparameters. Of course if you're training on Kaggle or Colab, time is an issue.</p>",
      "rawMarkdown": "Evolve doesn't require more GPU memory, just more time, because it does multiple runs. I found it easier to smooth out my training with evolve than by manually trying random hyperparameters. Of course if you're training on Kaggle or Colab, time is an issue.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1652874,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "01/17/2022 04:28:25",
      "content": "<p>with pretrained weight,  fine tuning around 10 epochs would get pretty good result.<br>\nfor better result,  GPU resource is a must. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1652916,
          "author_name": "okunot",
          "author_url": "",
          "post_date": "01/17/2022 05:24:32",
          "content": "<p>Thank you for your reply. I have a similar feel with you.</p>\n<p>I wonder: If we have much more epochs (100~), a model just gets overfitted even if along with strong data-augmentation methods.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1652984,
      "author_name": "alexchwong",
      "author_url": "",
      "post_date": "01/17/2022 06:35:50",
      "content": "<p>Keep in mind that the 300-epoch recommendation is for 80 objects in COCO. For 1 object, you don't need as many epochs.</p>\n<p>Train until validation loss does not improve with further cycles. It is hard to say how many is correct, as heavier augmentation or lower learning rate will lead to slower learning but may result in better mAP.</p>\n<p>Overtraining the model may result in higher precision and lower recall, which is bad for this competition as the F2 metric favors recall over precision.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1653018,
          "author_name": "okunot",
          "author_url": "",
          "post_date": "01/17/2022 07:09:50",
          "content": "<p>Thanks a lot. Makes sense to me.<br>\nBecause I was skeptical about my understanding, your comment made me trustable. :)</p>\n<p>It is very hard to determine training schedule, such as epoch number, learning rate, and heaviness of DA method. To estimate delicate balance among these, I think trust CV method is most important.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1653487,
          "author_name": "davidbroberts",
          "author_url": "",
          "post_date": "01/17/2022 15:19:37",
          "content": "<p>You can use YOLO's --evolve parameter to tune hyperparameters.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1653526,
          "author_name": "okunot",
          "author_url": "",
          "post_date": "01/17/2022 16:01:54",
          "content": "<p>Thank you for your suggestion.<br>\nWould you tell me how expensive is it in computational cost?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1653605,
          "author_name": "davidbroberts",
          "author_url": "",
          "post_date": "01/17/2022 17:24:45",
          "content": "<p>It's the same compute as training. It just tries to tune params to the 'fitness' metrics.</p>\n<p>Here's the overview from Ultralytics -&gt; <a href=\"https://docs.ultralytics.com/tutorials/hyperparameter-evolution/\" target=\"_blank\">https://docs.ultralytics.com/tutorials/hyperparameter-evolution/</a></p>",
          "votes": null,
          "replies": [
            {
              "id": 1653858,
              "author_name": "okunot",
              "author_url": "",
              "post_date": "01/17/2022 23:40:02",
              "content": "<p>I thought evolve option was so expensive!<br>\nI'm thinking of challenge it if I get good models.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 1653611,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "01/17/2022 17:31:00",
          "content": "<p>But …. evolve requires GPU resources (a lot) and good baseline model. You do not improve poor model (maybe a little bit) - it is bed strategy. So firstly work on dataset, create strong model and then tune.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1653868,
          "author_name": "okunot",
          "author_url": "",
          "post_date": "01/17/2022 23:46:54",
          "content": "<p>Thank you for giving me a lot of advice on my question.<br>\nI say.. First thing's first. We have to hurry! 🔥</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1653871,
          "author_name": "davidbroberts",
          "author_url": "",
          "post_date": "01/17/2022 23:55:15",
          "content": "<p>Evolve doesn't require more GPU memory, just more time, because it does multiple runs. I found it easier to smooth out my training with evolve than by manually trying random hyperparameters. Of course if you're training on Kaggle or Colab, time is an issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1653076,
      "author_name": "remekkinas",
      "author_url": "",
      "post_date": "01/17/2022 08:27:41",
      "content": "<p>How much epoch? It depends ….. You should look into metrics … and choose correct model (according to metrics and your goal). Just train yolov5 with <code>--save-period 1</code> and choose model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1653291,
          "author_name": "okunot",
          "author_url": "",
          "post_date": "01/17/2022 12:13:27",
          "content": "<p>I am grateful for your kind advice. I agree with you.<br>\nI'll try using --save-period option.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1653298,
          "author_name": "remekkinas",
          "author_url": "",
          "post_date": "01/17/2022 12:22:55",
          "content": "<p>Second one in fitness strategy in yolov5. You can automate process of \"best.pt\" model (<a href=\"https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/utils/metrics.py#L15-L19\" target=\"_blank\">https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/utils/metrics.py#L15-L19</a>)</p>\n<pre><code>def fitness(x):\n    w = [0.0, 0.0, 0.1, 0.9]  # weights for [P, R, mAP@0.5, mAP@0.5:0.95]\n    return (x[:, :4] * w).sum(1)\n</code></pre>\n<p>for really custom things you can modify code here: <a href=\"https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/train.py#L376-L395\" target=\"_blank\">https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/train.py#L376-L395</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1653520,
          "author_name": "okunot",
          "author_url": "",
          "post_date": "01/17/2022 15:55:49",
          "content": "<p>Great information.<br>\nAre there some recommendations about these weights?<br>\nIn GBR condition, recall may be worth to focus on, but it's hard to define better ratio;(</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1652730": "According to [YOLOv5 official](https://github.com/ultralytics/yolov5/wiki/Train-Custom-Data), 300 is recommended for default epoch number.\nBut, partly due to data size or training setting, training saturates early. For some simple situation, epoch number ~30 is enough to get a best validation result.\n\nHow do you set the training schedule including epoch number?\nAlthough \"small epoch number is better.\" sounds like gospel for guys with poor resource (like me), it should be considered carefully.\n\nThanks.",
    "1652874": "with pretrained weight,  fine tuning around 10 epochs would get pretty good result.\nfor better result,  GPU resource is a must.",
    "1652916": "Thank you for your reply. I have a similar feel with you.\n\nI wonder: If we have much more epochs (100~), a model just gets overfitted even if along with strong data-augmentation methods.",
    "1652984": "Keep in mind that the 300-epoch recommendation is for 80 objects in COCO. For 1 object, you don't need as many epochs.\n\nTrain until validation loss does not improve with further cycles. It is hard to say how many is correct, as heavier augmentation or lower learning rate will lead to slower learning but may result in better mAP.\n\nOvertraining the model may result in higher precision and lower recall, which is bad for this competition as the F2 metric favors recall over precision.",
    "1653018": "Thanks a lot. Makes sense to me.\nBecause I was skeptical about my understanding, your comment made me trustable. :)\n\nIt is very hard to determine training schedule, such as epoch number, learning rate, and heaviness of DA method. To estimate delicate balance among these, I think trust CV method is most important.",
    "1653076": "How much epoch? It depends ..... You should look into metrics ... and choose correct model (according to metrics and your goal). Just train yolov5 with `--save-period 1` and choose model.",
    "1653291": "I am grateful for your kind advice. I agree with you.\nI'll try using --save-period option.",
    "1653298": "Second one in fitness strategy in yolov5. You can automate process of \"best.pt\" model (https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/utils/metrics.py#L15-L19)\n\n```python\ndef fitness(x):\n    w = [0.0, 0.0, 0.1, 0.9]  # weights for [P, R, mAP@0.5, mAP@0.5:0.95]\n    return (x[:, :4] * w).sum(1)\n```\n\nfor really custom things you can modify code here: https://github.com/ultralytics/yolov5/blob/affa284352fa6d094d32fe2be69dbffe36bd20f8/train.py#L376-L395",
    "1653487": "You can use YOLO's --evolve parameter to tune hyperparameters.",
    "1653520": "Great information.\nAre there some recommendations about these weights?\nIn GBR condition, recall may be worth to focus on, but it's hard to define better ratio;(",
    "1653526": "Thank you for your suggestion.\nWould you tell me how expensive is it in computational cost?",
    "1653605": "It's the same compute as training. It just tries to tune params to the 'fitness' metrics.\n\nHere's the overview from Ultralytics -> https://docs.ultralytics.com/tutorials/hyperparameter-evolution/",
    "1653611": "But …. evolve requires GPU resources (a lot) and good baseline model. You do not improve poor model (maybe a little bit) - it is bed strategy. So firstly work on dataset, create strong model and then tune.",
    "1653858": "I thought evolve option was so expensive!\nI'm thinking of challenge it if I get good models.",
    "1653868": "Thank you for giving me a lot of advice on my question.\nI say.. First thing's first. We have to hurry! 🔥",
    "1653871": "Evolve doesn't require more GPU memory, just more time, because it does multiple runs. I found it easier to smooth out my training with evolve than by manually trying random hyperparameters. Of course if you're training on Kaggle or Colab, time is an issue."
  },
  "source": "meta"
}