{
  "id": 302384,
  "title": "How many epochs should we train with?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/302384",
  "author_name": "Wonjun Park",
  "post_date": "2022-01-22T08:53:58.237000",
  "votes": 6,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I completed training by 300 epoch using yolov5, yoloX, and yoloR models.</p>\n<p>Their total loss fell to an amazing degree!!</p>\n<p>However, the results were disastrous at the time of inference.</p>\n<p>In my opinion, since there is a lot of noise in the water, learning a lot makes it difficult to distinguish clear objects.</p>\n<p>How much do you think it's best to train?</p>",
  "messages": [
    {
      "id": 1659950,
      "postDate": "2022-01-22T08:53:58.237Z",
      "content": "<p>I completed training by 300 epoch using yolov5, yoloX, and yoloR models.</p>\n<p>Their total loss fell to an amazing degree!!</p>\n<p>However, the results were disastrous at the time of inference.</p>\n<p>In my opinion, since there is a lot of noise in the water, learning a lot makes it difficult to distinguish clear objects.</p>\n<p>How much do you think it's best to train?</p>",
      "rawMarkdown": "I completed training by 300 epoch using yolov5, yoloX, and yoloR models.\n\nTheir total loss fell to an amazing degree!!\n\nHowever, the results were disastrous at the time of inference.\n\nIn my opinion, since there is a lot of noise in the water, learning a lot makes it difficult to distinguish clear objects.\n\nHow much do you think it's best to train?",
      "votes": 5
    },
    {
      "id": 1663243,
      "postDate": "2022-01-24T22:46:48.910Z",
      "content": "<p>300 epochs is waaay too many! This is for COCO dataset which has 80 different objects.</p>\n<p>In my experience, 4 epochs of YOLOX-L is enough to get a \"quick and dirty\" model with LB &gt; 0.4</p>",
      "rawMarkdown": "300 epochs is waaay too many! This is for COCO dataset which has 80 different objects.\n\nIn my experience, 4 epochs of YOLOX-L is enough to get a \"quick and dirty\" model with LB > 0.4",
      "votes": 3,
      "replies": [
        {
          "id": 1663397,
          "postDate": "2022-01-25T04:08:55.613Z",
          "content": "<p>That's right, we are currently conducting repetitive experiments with a small epoch, and we are achieving high performance in a low time. 👍</p>",
          "rawMarkdown": "That's right, we are currently conducting repetitive experiments with a small epoch, and we are achieving high performance in a low time. 👍"
        }
      ]
    },
    {
      "id": 1660027,
      "postDate": "2022-01-22T10:26:10.613Z",
      "content": "<p>300 epochs, must be a long time to train🙌<br>\nBTW, I just use 15.. maybe more epoch will give better result</p>",
      "rawMarkdown": "300 epochs, must be a long time to train🙌\nBTW, I just use 15.. maybe more epoch will give better result",
      "votes": 1,
      "replies": [
        {
          "id": 1660327,
          "postDate": "2022-01-22T15:41:57.390Z",
          "content": "<p>Thank you for your answer!</p>",
          "rawMarkdown": "Thank you for your answer!"
        }
      ]
    },
    {
      "id": 1660577,
      "postDate": "2022-01-22T19:07:00.043Z",
      "content": "<p>I think you can check the competition metrics during training. with too small epochs, a model might underfit, with too many epochs, a model might overfit </p>",
      "rawMarkdown": "I think you can check the competition metrics during training. with too small epochs, a model might underfit, with too many epochs, a model might overfit ",
      "votes": 2,
      "replies": [
        {
          "id": 1660846,
          "postDate": "2022-01-23T02:47:58.077Z",
          "content": "<p>That's right. I think it's important to get the balance right.😂</p>",
          "rawMarkdown": "That's right. I think it's important to get the balance right.😂",
          "votes": 1
        }
      ]
    },
    {
      "id": 1660208,
      "postDate": "2022-01-22T13:45:27.263Z",
      "content": "<p><code>Their total loss fell to an amazing degree!!</code> on …. training or validation dataset? I assume training … it means that your NN is not able to generalize (is overfitted). Second I assume that you used checkpoint (yolov5 base model) - it is not needed to train network more then …. 20 epochs in this case (it is hard to say how many - you have to look into training metrices). </p>",
      "rawMarkdown": "`Their total loss fell to an amazing degree!!` on .... training or validation dataset? I assume training ... it means that your NN is not able to generalize (is overfitted). Second I assume that you used checkpoint (yolov5 base model) - it is not needed to train network more then .... 20 epochs in this case (it is hard to say how many - you have to look into training metrices). ",
      "votes": 2,
      "replies": [
        {
          "id": 1660331,
          "postDate": "2022-01-22T15:43:29.500Z",
          "content": "<p>However, the only metrics I can check during learning is loss.  Which metrics should I watch?</p>",
          "rawMarkdown": "However, the only metrics I can check during learning is loss.  Which metrics should I watch?"
        }
      ]
    },
    {
      "id": 1660087,
      "postDate": "2022-01-22T11:43:20.477Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 1660328,
          "postDate": "2022-01-22T15:42:18.890Z",
          "content": "<p>That's right. I think it's overfitting.</p>",
          "rawMarkdown": "That's right. I think it's overfitting."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1663243,
      "author_name": "Alex Wong",
      "author_url": "",
      "post_date": "2022-01-24T22:46:48.910000",
      "content": "<p>300 epochs is waaay too many! This is for COCO dataset which has 80 different objects.</p>\n<p>In my experience, 4 epochs of YOLOX-L is enough to get a \"quick and dirty\" model with LB &gt; 0.4</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1663397,
          "author_name": "Wonjun Park",
          "author_url": "",
          "post_date": "2022-01-25T04:08:55.613000",
          "content": "<p>That's right, we are currently conducting repetitive experiments with a small epoch, and we are achieving high performance in a low time. 👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1660027,
      "author_name": "Hao Chen",
      "author_url": "",
      "post_date": "2022-01-22T10:26:10.613000",
      "content": "<p>300 epochs, must be a long time to train🙌<br>\nBTW, I just use 15.. maybe more epoch will give better result</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1660327,
          "author_name": "Wonjun Park",
          "author_url": "",
          "post_date": "2022-01-22T15:41:57.390000",
          "content": "<p>Thank you for your answer!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1660577,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2022-01-22T19:07:00.043000",
      "content": "<p>I think you can check the competition metrics during training. with too small epochs, a model might underfit, with too many epochs, a model might overfit </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1660846,
          "author_name": "Wonjun Park",
          "author_url": "",
          "post_date": "2022-01-23T02:47:58.077000",
          "content": "<p>That's right. I think it's important to get the balance right.😂</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1660208,
      "author_name": "Remek Kinas",
      "author_url": "",
      "post_date": "2022-01-22T13:45:27.263000",
      "content": "<p><code>Their total loss fell to an amazing degree!!</code> on …. training or validation dataset? I assume training … it means that your NN is not able to generalize (is overfitted). Second I assume that you used checkpoint (yolov5 base model) - it is not needed to train network more then …. 20 epochs in this case (it is hard to say how many - you have to look into training metrices). </p>",
      "votes": 2,
      "replies": [
        {
          "id": 1660331,
          "author_name": "Wonjun Park",
          "author_url": "",
          "post_date": "2022-01-22T15:43:29.500000",
          "content": "<p>However, the only metrics I can check during learning is loss.  Which metrics should I watch?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1660087,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-01-22T11:43:20.477000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 1660328,
          "author_name": "Wonjun Park",
          "author_url": "",
          "post_date": "2022-01-22T15:42:18.890000",
          "content": "<p>That's right. I think it's overfitting.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1659950": "I completed training by 300 epoch using yolov5, yoloX, and yoloR models.\n\nTheir total loss fell to an amazing degree!!\n\nHowever, the results were disastrous at the time of inference.\n\nIn my opinion, since there is a lot of noise in the water, learning a lot makes it difficult to distinguish clear objects.\n\nHow much do you think it's best to train?",
    "1663243": "300 epochs is waaay too many! This is for COCO dataset which has 80 different objects.\n\nIn my experience, 4 epochs of YOLOX-L is enough to get a \"quick and dirty\" model with LB > 0.4",
    "1660027": "300 epochs, must be a long time to train🙌\nBTW, I just use 15.. maybe more epoch will give better result",
    "1660577": "I think you can check the competition metrics during training. with too small epochs, a model might underfit, with too many epochs, a model might overfit ",
    "1660208": "`Their total loss fell to an amazing degree!!` on .... training or validation dataset? I assume training ... it means that your NN is not able to generalize (is overfitted). Second I assume that you used checkpoint (yolov5 base model) - it is not needed to train network more then .... 20 epochs in this case (it is hard to say how many - you have to look into training metrices). ",
    "1660087": ""
  }
}