{
  "id": 77836,
  "title": "Do we agree that val loss increase is overfitting?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77836",
  "author_name": "",
  "post_date": "2019-01-17T02:46:33.644059800Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>After 4~5 epochs the val loss starts to increase, this happens in almost all public kernels. Do we agree that is a sign of overfitting? I've seen similar posts but didn't get a clear answer. Is that the optimal point? should we stop there? If the local cv is the golden rule, why public kernels all train 5 epochs instead of stop at 3.</p>\n\n<p>The hold-out(or local test set) approach doesn't make sense to me either. To me, it is just a bigger val set, which may give you more robust val score but you are losing training cases for it.</p>",
  "messages": [
    {
      "id": "457155",
      "postDate": "01/17/2019 02:46:33",
      "content": "<p>After 4~5 epochs the val loss starts to increase, this happens in almost all public kernels. Do we agree that is a sign of overfitting? I've seen similar posts but didn't get a clear answer. Is that the optimal point? should we stop there? If the local cv is the golden rule, why public kernels all train 5 epochs instead of stop at 3.</p>\n\n<p>The hold-out(or local test set) approach doesn't make sense to me either. To me, it is just a bigger val set, which may give you more robust val score but you are losing training cases for it.</p>",
      "rawMarkdown": "After 4~5 epochs the val loss starts to increase, this happens in almost all public kernels. Do we agree that is a sign of overfitting? I've seen similar posts but didn't get a clear answer. Is that the optimal point? should we stop there? If the local cv is the golden rule, why public kernels all train 5 epochs instead of stop at 3.\n\nThe hold-out(or local test set) approach doesn't make sense to me either. To me, it is just a bigger val set, which may give you more robust val score but you are losing training cases for it.",
      "votes": null
    },
    {
      "id": "457342",
      "postDate": "01/17/2019 08:57:16",
      "content": "<p>Hi. \nYes it could be a sign of overfitting but I wouldn't depend much on it for this task. I have found validation loss a bit noisy for NLP tasks.  I would focus more on the metric of interest (in this case the F1 score) in order to get a better estimate on the effect of overfitting (i.e. F1 starts to decrease significantly). In my experience training for 4-5 epochs seem enough for this task.</p>\n\n<p>As for your second point, cross-validation is preferable to a holdout test set but it costs more. I would go for cross-validation</p>",
      "rawMarkdown": "Hi. \nYes it could be a sign of overfitting but I wouldn't depend much on it for this task. I have found validation loss a bit noisy for NLP tasks.  I would focus more on the metric of interest (in this case the F1 score) in order to get a better estimate on the effect of overfitting (i.e. F1 starts to decrease significantly). In my experience training for 4-5 epochs seem enough for this task.\n\nAs for your second point, cross-validation is preferable to a holdout test set but it costs more. I would go for cross-validation",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 457342,
      "author_name": "erwtokritos",
      "author_url": "",
      "post_date": "01/17/2019 08:57:16",
      "content": "<p>Hi. \nYes it could be a sign of overfitting but I wouldn't depend much on it for this task. I have found validation loss a bit noisy for NLP tasks.  I would focus more on the metric of interest (in this case the F1 score) in order to get a better estimate on the effect of overfitting (i.e. F1 starts to decrease significantly). In my experience training for 4-5 epochs seem enough for this task.</p>\n\n<p>As for your second point, cross-validation is preferable to a holdout test set but it costs more. I would go for cross-validation</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "457155": "After 4~5 epochs the val loss starts to increase, this happens in almost all public kernels. Do we agree that is a sign of overfitting? I've seen similar posts but didn't get a clear answer. Is that the optimal point? should we stop there? If the local cv is the golden rule, why public kernels all train 5 epochs instead of stop at 3.\n\nThe hold-out(or local test set) approach doesn't make sense to me either. To me, it is just a bigger val set, which may give you more robust val score but you are losing training cases for it.",
    "457342": "Hi. \nYes it could be a sign of overfitting but I wouldn't depend much on it for this task. I have found validation loss a bit noisy for NLP tasks.  I would focus more on the metric of interest (in this case the F1 score) in order to get a better estimate on the effect of overfitting (i.e. F1 starts to decrease significantly). In my experience training for 4-5 epochs seem enough for this task.\n\nAs for your second point, cross-validation is preferable to a holdout test set but it costs more. I would go for cross-validation"
  },
  "source": "meta"
}