{
  "id": 91514,
  "title": "is there a general rule on training-validation loss gap?",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/91514",
  "author_name": "Massoud Hosseinali",
  "post_date": "2019-05-05T23:27:48.961000",
  "votes": 11,
  "comment_count": 15,
  "views": 0,
  "content": "<p>consider two hypothetical cases of training and validation loss below:\n- training MAE: 1.10  | validation MAE: 1.40\n- training MAE: 0.80 | validation MAE: 1.15</p>\n\n<p>which one would you prefer to go with? On one hand, first option seems to generalize better because the difference is about 20% as opposed to the second option where the difference is about 30%, so the LB score could be expected to be closer to the validation MAE. On the other hand, despite the bigger train-validation gap in the second option, it has performed better than option one, on the unseen (validation) set and I am wondering if despite second option has less generalization ability and I can expect a bigger LB score could it still be less than the LB score of first option? My experience with this competition is that the first option gave me a better LB score but is it just an accident or is there a general rule? (assume the learning curves of training and validation loss in both options were monotonically decreasing)</p>\n\n<p>It's been a while since I started being active on Kaggle and so far it's been a wonderful experience because I learned a lot from experts here. I appreciate if you share your thoughts on this too.</p>",
  "messages": [
    {
      "id": 527608,
      "postDate": "2019-05-05T23:27:48.963Z",
      "content": "<p>consider two hypothetical cases of training and validation loss below:\n- training MAE: 1.10  | validation MAE: 1.40\n- training MAE: 0.80 | validation MAE: 1.15</p>\n\n<p>which one would you prefer to go with? On one hand, first option seems to generalize better because the difference is about 20% as opposed to the second option where the difference is about 30%, so the LB score could be expected to be closer to the validation MAE. On the other hand, despite the bigger train-validation gap in the second option, it has performed better than option one, on the unseen (validation) set and I am wondering if despite second option has less generalization ability and I can expect a bigger LB score could it still be less than the LB score of first option? My experience with this competition is that the first option gave me a better LB score but is it just an accident or is there a general rule? (assume the learning curves of training and validation loss in both options were monotonically decreasing)</p>\n\n<p>It's been a while since I started being active on Kaggle and so far it's been a wonderful experience because I learned a lot from experts here. I appreciate if you share your thoughts on this too.</p>",
      "rawMarkdown": "consider two hypothetical cases of training and validation loss below:\n- training MAE: 1.10  | validation MAE: 1.40\n- training MAE: 0.80 | validation MAE: 1.15\n\nwhich one would you prefer to go with? On one hand, first option seems to generalize better because the difference is about 20% as opposed to the second option where the difference is about 30%, so the LB score could be expected to be closer to the validation MAE. On the other hand, despite the bigger train-validation gap in the second option, it has performed better than option one, on the unseen (validation) set and I am wondering if despite second option has less generalization ability and I can expect a bigger LB score could it still be less than the LB score of first option? My experience with this competition is that the first option gave me a better LB score but is it just an accident or is there a general rule? (assume the learning curves of training and validation loss in both options were monotonically decreasing)\n\nIt's been a while since I started being active on Kaggle and so far it's been a wonderful experience because I learned a lot from experts here. I appreciate if you share your thoughts on this too.",
      "votes": 11
    },
    {
      "id": 527722,
      "postDate": "2019-05-06T06:54:49.097Z",
      "content": "<p>I'm glad to see this discussed.  I usually watch train_validation gap, and try to get it as low as possible.  In your example above, the difference in validation score is large enough to prefer the second one IMHO, but in something like below I would consider the first:</p>\n\n<ul>\n<li>training MAE: 1.10 | validation MAE: 1.40</li>\n<li>training MAE: 0.80 | validation MAE: 1.35</li>\n</ul>",
      "rawMarkdown": "I'm glad to see this discussed.  I usually watch train_validation gap, and try to get it as low as possible.  In your example above, the difference in validation score is large enough to prefer the second one IMHO, but in something like below I would consider the first:\n\n\n-    training MAE: 1.10 | validation MAE: 1.40\n-    training MAE: 0.80 | validation MAE: 1.35\n\n",
      "votes": 5,
      "replies": [
        {
          "id": 527887,
          "postDate": "2019-05-06T14:24:39.887Z",
          "content": "<p>Thank you <a href=\"/cpmpml\">@cpmpml</a>. I know the general rule is to sorta stop early if the validation loss is leveling off and training loss is still going down. What if both of them are going down but the gap is increasing too (which happens most of the times)? do you let them model run up to a point that validation loss levels off or would you stop early based on the gap too? if it's the latter, what is your rule to stop early? how do you determine which iteration is good enough and you wouldn't want to go further than that?</p>\n\n<p>PS. I know in the hypothetical instance that I brought up I could use regularization to reduce the gap. I kinda exaggerated there to make my point clearer.</p>",
          "rawMarkdown": "Thank you @cpmpml. I know the general rule is to sorta stop early if the validation loss is leveling off and training loss is still going down. What if both of them are going down but the gap is increasing too (which happens most of the times)? do you let them model run up to a point that validation loss levels off or would you stop early based on the gap too? if it's the latter, what is your rule to stop early? how do you determine which iteration is good enough and you wouldn't want to go further than that?\n\nPS. I know in the hypothetical instance that I brought up I could use regularization to reduce the gap. I kinda exaggerated there to make my point clearer.",
          "votes": 1
        },
        {
          "id": 527910,
          "postDate": "2019-05-06T16:01:41.017Z",
          "content": "<p>I don't do early stopping on the gap, but that's an interesting idea to experiment with.  However, it is not supported by popular models like XGBoost or LightGBM, hence would require significant coding effort to implement.</p>",
          "rawMarkdown": "I don't do early stopping on the gap, but that's an interesting idea to experiment with.  However, it is not supported by popular models like XGBoost or LightGBM, hence would require significant coding effort to implement.",
          "votes": 1
        },
        {
          "id": 527935,
          "postDate": "2019-05-06T16:33:56.663Z",
          "content": "<p>Probably not so <a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html\">significant</a></p>",
          "rawMarkdown": "Probably not so [significant](https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html)",
          "votes": 1
        },
        {
          "id": 527939,
          "postDate": "2019-05-06T16:47:59.023Z",
          "content": "<p>Waiting for your code then :)</p>",
          "rawMarkdown": "Waiting for your code then :)",
          "votes": 2
        },
        {
          "id": 528220,
          "postDate": "2019-05-07T09:21:27.217Z",
          "content": "<p>Haha, I feel like I'm being manipulated :)\nCheck implementation <a href=\"https://www.kaggle.com/alexfir/lgbm-early-stopping-based-on-train-validation-gap\">here</a>.</p>\n\n<p>Just pass it to callback like below:\n<code>model = lgb.train(..., callbacks=[TrainValidationGapEarlyStopping(2.3)])</code></p>",
          "rawMarkdown": "Haha, I feel like I'm being manipulated :)\nCheck implementation [here](https://www.kaggle.com/alexfir/lgbm-early-stopping-based-on-train-validation-gap).\n\nJust pass it to callback like below:\n`model = lgb.train(..., callbacks=[TrainValidationGapEarlyStopping(2.3)])`",
          "votes": 7
        },
        {
          "id": 528242,
          "postDate": "2019-05-07T10:35:39.553Z",
          "content": "<p>Thanks for the code!  I'll report if it is helping.</p>",
          "rawMarkdown": "Thanks for the code!  I'll report if it is helping."
        }
      ]
    },
    {
      "id": 528202,
      "postDate": "2019-05-07T08:53:17.307Z",
      "content": "<p>How you divide these two conditions, yeah your assumption is hypothetical but how you got exact values? Please explain so that I can understand your words. BTW great start!!! </p>",
      "rawMarkdown": "How you divide these two conditions, yeah your assumption is hypothetical but how you got exact values? Please explain so that I can understand your words. BTW great start!!! ",
      "votes": 2
    },
    {
      "id": 527873,
      "postDate": "2019-05-06T13:52:58.437Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 527886,
          "postDate": "2019-05-06T14:24:30.017Z",
          "content": "<p>they're <em>hypothetical</em> values</p>",
          "rawMarkdown": "they're *hypothetical* values"
        },
        {
          "id": 527897,
          "postDate": "2019-05-06T15:03:40.823Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 527936,
          "postDate": "2019-05-06T16:35:05.010Z",
          "content": "<p>Why is the training MAE weak? What does that even mean?</p>",
          "rawMarkdown": "Why is the training MAE weak? What does that even mean?"
        },
        {
          "id": 527983,
          "postDate": "2019-05-06T18:39:51.773Z",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I guess he meant \"low\" by \"weak\".</p>",
          "rawMarkdown": "@philippsinger I guess he meant \"low\" by \"weak\"."
        },
        {
          "id": 527988,
          "postDate": "2019-05-06T18:50:55.993Z",
          "content": "<p>Makes more sense ;)</p>",
          "rawMarkdown": "Makes more sense ;)"
        },
        {
          "id": 527990,
          "postDate": "2019-05-06T18:58:13.620Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 527722,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2019-05-06T06:54:49.097000",
      "content": "<p>I'm glad to see this discussed.  I usually watch train_validation gap, and try to get it as low as possible.  In your example above, the difference in validation score is large enough to prefer the second one IMHO, but in something like below I would consider the first:</p>\n\n<ul>\n<li>training MAE: 1.10 | validation MAE: 1.40</li>\n<li>training MAE: 0.80 | validation MAE: 1.35</li>\n</ul>",
      "votes": 5,
      "replies": [
        {
          "id": 527887,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-05-06T14:24:39.887000",
          "content": "<p>Thank you <a href=\"/cpmpml\">@cpmpml</a>. I know the general rule is to sorta stop early if the validation loss is leveling off and training loss is still going down. What if both of them are going down but the gap is increasing too (which happens most of the times)? do you let them model run up to a point that validation loss levels off or would you stop early based on the gap too? if it's the latter, what is your rule to stop early? how do you determine which iteration is good enough and you wouldn't want to go further than that?</p>\n\n<p>PS. I know in the hypothetical instance that I brought up I could use regularization to reduce the gap. I kinda exaggerated there to make my point clearer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527910,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-06T16:01:41.017000",
          "content": "<p>I don't do early stopping on the gap, but that's an interesting idea to experiment with.  However, it is not supported by popular models like XGBoost or LightGBM, hence would require significant coding effort to implement.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527935,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2019-05-06T16:33:56.663000",
          "content": "<p>Probably not so <a href=\"https://lightgbm.readthedocs.io/en/latest/_modules/lightgbm/callback.html\">significant</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 527939,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-06T16:47:59.023000",
          "content": "<p>Waiting for your code then :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 528220,
          "author_name": "Alexander Firsov",
          "author_url": "",
          "post_date": "2019-05-07T09:21:27.217000",
          "content": "<p>Haha, I feel like I'm being manipulated :)\nCheck implementation <a href=\"https://www.kaggle.com/alexfir/lgbm-early-stopping-based-on-train-validation-gap\">here</a>.</p>\n\n<p>Just pass it to callback like below:\n<code>model = lgb.train(..., callbacks=[TrainValidationGapEarlyStopping(2.3)])</code></p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 528242,
          "author_name": "CPMP",
          "author_url": "",
          "post_date": "2019-05-07T10:35:39.553000",
          "content": "<p>Thanks for the code!  I'll report if it is helping.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 528202,
      "author_name": "Himanshu Soni",
      "author_url": "",
      "post_date": "2019-05-07T08:53:17.307000",
      "content": "<p>How you divide these two conditions, yeah your assumption is hypothetical but how you got exact values? Please explain so that I can understand your words. BTW great start!!! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 527873,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-06T13:52:58.437000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 527886,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-05-06T14:24:30.017000",
          "content": "<p>they're <em>hypothetical</em> values</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527897,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-06T15:03:40.823000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527936,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-06T16:35:05.010000",
          "content": "<p>Why is the training MAE weak? What does that even mean?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527983,
          "author_name": "Massoud Hosseinali",
          "author_url": "",
          "post_date": "2019-05-06T18:39:51.773000",
          "content": "<p><a href=\"/philippsinger\">@philippsinger</a> I guess he meant \"low\" by \"weak\".</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527988,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-05-06T18:50:55.993000",
          "content": "<p>Makes more sense ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 527990,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-06T18:58:13.620000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "527608": "consider two hypothetical cases of training and validation loss below:\n- training MAE: 1.10  | validation MAE: 1.40\n- training MAE: 0.80 | validation MAE: 1.15\n\nwhich one would you prefer to go with? On one hand, first option seems to generalize better because the difference is about 20% as opposed to the second option where the difference is about 30%, so the LB score could be expected to be closer to the validation MAE. On the other hand, despite the bigger train-validation gap in the second option, it has performed better than option one, on the unseen (validation) set and I am wondering if despite second option has less generalization ability and I can expect a bigger LB score could it still be less than the LB score of first option? My experience with this competition is that the first option gave me a better LB score but is it just an accident or is there a general rule? (assume the learning curves of training and validation loss in both options were monotonically decreasing)\n\nIt's been a while since I started being active on Kaggle and so far it's been a wonderful experience because I learned a lot from experts here. I appreciate if you share your thoughts on this too.",
    "527722": "I'm glad to see this discussed.  I usually watch train_validation gap, and try to get it as low as possible.  In your example above, the difference in validation score is large enough to prefer the second one IMHO, but in something like below I would consider the first:\n\n\n-    training MAE: 1.10 | validation MAE: 1.40\n-    training MAE: 0.80 | validation MAE: 1.35\n\n",
    "528202": "How you divide these two conditions, yeah your assumption is hypothetical but how you got exact values? Please explain so that I can understand your words. BTW great start!!! ",
    "527873": ""
  }
}