{
  "id": 335689,
  "title": "Overfitting vs overfitting",
  "url": "/competitions/amex-default-prediction/discussion/335689",
  "author_name": "",
  "post_date": "2022-07-07T11:58:19.990232Z",
  "votes": 20,
  "comment_count": 8,
  "views": 0,
  "content": "<p>This topic is inspired on this twitter <a href=\"https://twitter.com/giffmana/status/1542241800647282688?t=I1IwilMHlKg9RJfdaQuUgA&amp;s=09\" target=\"_blank\">poll</a> from Lucas Beyer about what people have in mind when they think about overfitting.</p>\n<p>I've always thought about overfitting when the spread between train and validation metric increases. However, IIRC <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> said in one of the fast.ai videos that (in the context of NNs training) as long as your validation metric continues improving, you are good to go.</p>\n<p>Let's take a concrete example to illustrate this: <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963\" target=\"_blank\">notebook</a>. The gap between CV train/val scores is much higher than the previous highest public notebooks (e.g. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> or <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> or <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> notebooks), however the CV val (even running the CV for several seeds) and Public LB scores are also better. </p>\n<p>The question is, if you have to choose between this two options, what would you choose?:</p>\n<ol>\n<li>A model with higher CV train/val scores gap but better CV val/Public LB scores</li>\n<li>A model with lower CV train/val scores gap but worse CV val/Public LB scores</li>\n</ol>\n<p>Probably it depends on how much the gap increases vs how much CV/LB improves, and that choice is a little bit of an art. However, I would be very happy to know your opinion based on your experience on this topic.</p>\n<p>PS: In one of the Kaggle Days Paris <a href=\"https://www.youtube.com/watch?v=VC8Jc9_lNoY&amp;t=710s\" target=\"_blank\">talks</a>, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> talked about this topic.</p>",
  "messages": [
    {
      "id": "1846849",
      "postDate": "07/07/2022 11:58:19",
      "content": "<p>This topic is inspired on this twitter <a href=\"https://twitter.com/giffmana/status/1542241800647282688?t=I1IwilMHlKg9RJfdaQuUgA&amp;s=09\" target=\"_blank\">poll</a> from Lucas Beyer about what people have in mind when they think about overfitting.</p>\n<p>I've always thought about overfitting when the spread between train and validation metric increases. However, IIRC <a href=\"https://www.kaggle.com/jhoward\" target=\"_blank\">@jhoward</a> said in one of the fast.ai videos that (in the context of NNs training) as long as your validation metric continues improving, you are good to go.</p>\n<p>Let's take a concrete example to illustrate this: <a href=\"https://www.kaggle.com/ragnar123\" target=\"_blank\">@ragnar123</a> <a href=\"https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963\" target=\"_blank\">notebook</a>. The gap between CV train/val scores is much higher than the previous highest public notebooks (e.g. <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> or <a href=\"https://www.kaggle.com/jiweiliu\" target=\"_blank\">@jiweiliu</a> or <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> notebooks), however the CV val (even running the CV for several seeds) and Public LB scores are also better. </p>\n<p>The question is, if you have to choose between this two options, what would you choose?:</p>\n<ol>\n<li>A model with higher CV train/val scores gap but better CV val/Public LB scores</li>\n<li>A model with lower CV train/val scores gap but worse CV val/Public LB scores</li>\n</ol>\n<p>Probably it depends on how much the gap increases vs how much CV/LB improves, and that choice is a little bit of an art. However, I would be very happy to know your opinion based on your experience on this topic.</p>\n<p>PS: In one of the Kaggle Days Paris <a href=\"https://www.youtube.com/watch?v=VC8Jc9_lNoY&amp;t=710s\" target=\"_blank\">talks</a>, <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> talked about this topic.</p>",
      "rawMarkdown": "This topic is inspired on this twitter [poll](https://twitter.com/giffmana/status/1542241800647282688?t=I1IwilMHlKg9RJfdaQuUgA&s=09) from Lucas Beyer about what people have in mind when they think about overfitting.\n\nI've always thought about overfitting when the spread between train and validation metric increases. However, IIRC @jhoward said in one of the fast.ai videos that (in the context of NNs training) as long as your validation metric continues improving, you are good to go.\n\nLet's take a concrete example to illustrate this: @ragnar123 [notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963). The gap between CV train/val scores is much higher than the previous highest public notebooks (e.g. @cdeotte or @jiweiliu or @ambrosm notebooks), however the CV val (even running the CV for several seeds) and Public LB scores are also better. \n\nThe question is, if you have to choose between this two options, what would you choose?:\n\n1. A model with higher CV train/val scores gap but better CV val/Public LB scores\n2. A model with lower CV train/val scores gap but worse CV val/Public LB scores\n\nProbably it depends on how much the gap increases vs how much CV/LB improves, and that choice is a little bit of an art. However, I would be very happy to know your opinion based on your experience on this topic.\n\nPS: In one of the Kaggle Days Paris [talks](https://www.youtube.com/watch?v=VC8Jc9_lNoY&t=710s), @cpmpml talked about this topic.",
      "votes": null
    },
    {
      "id": "1846897",
      "postDate": "07/07/2022 12:43:31",
      "content": "<p>I favor 1. (simply optimize CV score) by a wide margin. The concept of doing \"too well\" on training data in isolation doesn't make sense -- the best example is random forest, where the standard RF will almost perfectly predict the training data but of course have a much lower val score. But that's a quirk of the model design (ensembling many low bias, high variance learners)   and not proof of poor generalization behavior.</p>\n<p>Stepping back, we want our models to have the best generalization performance possible, which is what hold-out sets directly proxy. A train-val gap is a much worse, model-dependent proxy. In my experience, the gap is most useful for model refinement on a relative basis -- for example, if I see that increasing <code>max_depth</code> and <code>num_leaves</code> in lgbm increases my train score a lot and my val score somewhat, that suggests that my model may still be underfitting. A more complex model might generalize better, even if the score gap is growing worse.</p>",
      "rawMarkdown": "I favor 1. (simply optimize CV score) by a wide margin. The concept of doing \"too well\" on training data in isolation doesn't make sense -- the best example is random forest, where the standard RF will almost perfectly predict the training data but of course have a much lower val score. But that's a quirk of the model design (ensembling many low bias, high variance learners)   and not proof of poor generalization behavior.\n\nStepping back, we want our models to have the best generalization performance possible, which is what hold-out sets directly proxy. A train-val gap is a much worse, model-dependent proxy. In my experience, the gap is most useful for model refinement on a relative basis -- for example, if I see that increasing `max_depth` and `num_leaves` in lgbm increases my train score a lot and my val score somewhat, that suggests that my model may still be underfitting. A more complex model might generalize better, even if the score gap is growing worse.",
      "votes": null
    },
    {
      "id": "1847193",
      "postDate": "07/07/2022 17:39:25",
      "content": "<p>I only care about how my model will perform on unseen data. In my experience on kaggle, the model with the best validation score does the best on the test set, regardless of the training/validation gap. However, if two models have the same validation score, I prefer the one with a small train/validation gap. The best way to test out this hypothesis is by doing nested cross validation. This is when you hold out a true test set that is not seen during training, as opposed to a \"validation\" set that we are using for early stopping. </p>",
      "rawMarkdown": "I only care about how my model will perform on unseen data. In my experience on kaggle, the model with the best validation score does the best on the test set, regardless of the training/validation gap. However, if two models have the same validation score, I prefer the one with a small train/validation gap. The best way to test out this hypothesis is by doing nested cross validation. This is when you hold out a true test set that is not seen during training, as opposed to a \"validation\" set that we are using for early stopping.",
      "votes": null
    },
    {
      "id": "1847469",
      "postDate": "07/08/2022 00:35:46",
      "content": "<p>Great question. I think it really depends on the situation. <br>\nIf you have a model with a higher CV train/val score gap, but the CV val/Public LB scores are better, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. <br>\nHowever, if you have a model with a lower CV train/val score gap, but the CV val/Public LB scores are worse, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. <br>\nHope this helps. Thanks!</p>",
      "rawMarkdown": "Great question. I think it really depends on the situation. \nIf you have a model with a higher CV train/val score gap, but the CV val/Public LB scores are better, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. \nHowever, if you have a model with a lower CV train/val score gap, but the CV val/Public LB scores are worse, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. \nHope this helps. Thanks!",
      "votes": null
    },
    {
      "id": "1849239",
      "postDate": "07/09/2022 10:29:01",
      "content": "<p>Thanks for the answers. I think that it depends on the situation and therefore is tricky. All what you said totally makes sense to me guys, but from the other point of view, CPMP says in the video:</p>\n<blockquote>\n  <p>What I am doing when I am building new models, when I have a new CV score, I want it to be better but I also look at the training score, and if the training score improves too much, then my model is just overfitting and probably it won't fare well on the test data. So I always watch the gap between train and validation, and I avoid features that provide a very small CV improvement at the expense of very large train improvement. I don't even submit it.</p>\n</blockquote>",
      "rawMarkdown": "Thanks for the answers. I think that it depends on the situation and therefore is tricky. All what you said totally makes sense to me guys, but from the other point of view, CPMP says in the video:\n\n> What I am doing when I am building new models, when I have a new CV score, I want it to be better but I also look at the training score, and if the training score improves too much, then my model is just overfitting and probably it won't fare well on the test data. So I always watch the gap between train and validation, and I avoid features that provide a very small CV improvement at the expense of very large train improvement. I don't even submit it.",
      "votes": null
    },
    {
      "id": "1852512",
      "postDate": "07/12/2022 05:48:27",
      "content": "<p>Very good question. As long as the Eval improves, I push the training further, even if the train/Eval gap is high. Sometimes I see the training log loss function geting better and the eval amex-metric getting worse for several iterations until the eval suddenly jumps better. This is why I use a large early stopping. The best indicator of overfiting is not the train/Eval gap but the deterioration of the eval score.</p>",
      "rawMarkdown": "Very good question. As long as the Eval improves, I push the training further, even if the train/Eval gap is high. Sometimes I see the training log loss function geting better and the eval amex-metric getting worse for several iterations until the eval suddenly jumps better. This is why I use a large early stopping. The best indicator of overfiting is not the train/Eval gap but the deterioration of the eval score.",
      "votes": null
    },
    {
      "id": "1853598",
      "postDate": "07/13/2022 01:52:48",
      "content": "<p>I generally ignore the train/val CV gap and just focus on doing anything that improves my overall CV score. If you want to close the train/val CV gap, then you can add regularization to training, like data augmentation or model regularization. Making it harder for your model to memorize all the train data shrinks the gap and generally boosts your overall CV score.</p>\n<p>In every competition I pay very close attention to the CV/LB gap. Both CV val and Public LB are unseen data. And if test data and train data come from the same distribution, then these two scores (CV and LB) should be the same. When they are not, it is upmost important to discover why. Then redesign your CV scheme to shrink the CV/LB gap, OR adjust training and/or post process to change LB predictions and shrink the gap. In most competitions, discovering why CV and LB have a gap, and discovering how to shrink it will significantly boost your LB.</p>",
      "rawMarkdown": "I generally ignore the train/val CV gap and just focus on doing anything that improves my overall CV score. If you want to close the train/val CV gap, then you can add regularization to training, like data augmentation or model regularization. Making it harder for your model to memorize all the train data shrinks the gap and generally boosts your overall CV score.\n\nIn every competition I pay very close attention to the CV/LB gap. Both CV val and Public LB are unseen data. And if test data and train data come from the same distribution, then these two scores (CV and LB) should be the same. When they are not, it is upmost important to discover why. Then redesign your CV scheme to shrink the CV/LB gap, OR adjust training and/or post process to change LB predictions and shrink the gap. In most competitions, discovering why CV and LB have a gap, and discovering how to shrink it will significantly boost your LB.",
      "votes": null
    },
    {
      "id": "1853740",
      "postDate": "07/13/2022 05:28:48",
      "content": "<p>Basically, the normal and the theoretical is 100% #1, if it does better on unseen data it's not overfitting. Models are designed to find the optimal trade-off for you, hence early stopping and all that.</p>\n<p>That said, the true test is how the model does on \"unseen data\", any caveats to favoring #1 100% is if you think there's any sort of leakage between your train set and your validation set. The most trivial example is if observations (rows) aren't fully independent. Time series stock market data would be an example of that. Hundreds of thousands of credit card customers sounds pretty independent, though even there the train data (and thus the CV data used for early stopping) is from one time period, and the LB is from two other time periods. So it would be technically possible to overfit while seeing your CV score go up, and possibly finding a way of getting closer to #2 could help.</p>\n<p>In practice, for something like this: I'm thinking #1 with no real worries.</p>",
      "rawMarkdown": "Basically, the normal and the theoretical is 100% #1, if it does better on unseen data it's not overfitting. Models are designed to find the optimal trade-off for you, hence early stopping and all that.\n\nThat said, the true test is how the model does on \"unseen data\", any caveats to favoring #1 100% is if you think there's any sort of leakage between your train set and your validation set. The most trivial example is if observations (rows) aren't fully independent. Time series stock market data would be an example of that. Hundreds of thousands of credit card customers sounds pretty independent, though even there the train data (and thus the CV data used for early stopping) is from one time period, and the LB is from two other time periods. So it would be technically possible to overfit while seeing your CV score go up, and possibly finding a way of getting closer to #2 could help.\n\nIn practice, for something like this: I'm thinking #1 with no real worries.",
      "votes": null
    },
    {
      "id": "1896884",
      "postDate": "08/13/2022 08:39:36",
      "content": "<p>00000000000000</p>",
      "rawMarkdown": "00000000000000",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1846897,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "07/07/2022 12:43:31",
      "content": "<p>I favor 1. (simply optimize CV score) by a wide margin. The concept of doing \"too well\" on training data in isolation doesn't make sense -- the best example is random forest, where the standard RF will almost perfectly predict the training data but of course have a much lower val score. But that's a quirk of the model design (ensembling many low bias, high variance learners)   and not proof of poor generalization behavior.</p>\n<p>Stepping back, we want our models to have the best generalization performance possible, which is what hold-out sets directly proxy. A train-val gap is a much worse, model-dependent proxy. In my experience, the gap is most useful for model refinement on a relative basis -- for example, if I see that increasing <code>max_depth</code> and <code>num_leaves</code> in lgbm increases my train score a lot and my val score somewhat, that suggests that my model may still be underfitting. A more complex model might generalize better, even if the score gap is growing worse.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1847193,
      "author_name": "chrisrichardmiles",
      "author_url": "",
      "post_date": "07/07/2022 17:39:25",
      "content": "<p>I only care about how my model will perform on unseen data. In my experience on kaggle, the model with the best validation score does the best on the test set, regardless of the training/validation gap. However, if two models have the same validation score, I prefer the one with a small train/validation gap. The best way to test out this hypothesis is by doing nested cross validation. This is when you hold out a true test set that is not seen during training, as opposed to a \"validation\" set that we are using for early stopping. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1847469,
      "author_name": "thedevastator",
      "author_url": "",
      "post_date": "07/08/2022 00:35:46",
      "content": "<p>Great question. I think it really depends on the situation. <br>\nIf you have a model with a higher CV train/val score gap, but the CV val/Public LB scores are better, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. <br>\nHowever, if you have a model with a lower CV train/val score gap, but the CV val/Public LB scores are worse, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. <br>\nHope this helps. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1849239,
      "author_name": "delai50",
      "author_url": "",
      "post_date": "07/09/2022 10:29:01",
      "content": "<p>Thanks for the answers. I think that it depends on the situation and therefore is tricky. All what you said totally makes sense to me guys, but from the other point of view, CPMP says in the video:</p>\n<blockquote>\n  <p>What I am doing when I am building new models, when I have a new CV score, I want it to be better but I also look at the training score, and if the training score improves too much, then my model is just overfitting and probably it won't fare well on the test data. So I always watch the gap between train and validation, and I avoid features that provide a very small CV improvement at the expense of very large train improvement. I don't even submit it.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1852512,
      "author_name": "gehallak",
      "author_url": "",
      "post_date": "07/12/2022 05:48:27",
      "content": "<p>Very good question. As long as the Eval improves, I push the training further, even if the train/Eval gap is high. Sometimes I see the training log loss function geting better and the eval amex-metric getting worse for several iterations until the eval suddenly jumps better. This is why I use a large early stopping. The best indicator of overfiting is not the train/Eval gap but the deterioration of the eval score.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1853598,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "07/13/2022 01:52:48",
      "content": "<p>I generally ignore the train/val CV gap and just focus on doing anything that improves my overall CV score. If you want to close the train/val CV gap, then you can add regularization to training, like data augmentation or model regularization. Making it harder for your model to memorize all the train data shrinks the gap and generally boosts your overall CV score.</p>\n<p>In every competition I pay very close attention to the CV/LB gap. Both CV val and Public LB are unseen data. And if test data and train data come from the same distribution, then these two scores (CV and LB) should be the same. When they are not, it is upmost important to discover why. Then redesign your CV scheme to shrink the CV/LB gap, OR adjust training and/or post process to change LB predictions and shrink the gap. In most competitions, discovering why CV and LB have a gap, and discovering how to shrink it will significantly boost your LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1853740,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/13/2022 05:28:48",
      "content": "<p>Basically, the normal and the theoretical is 100% #1, if it does better on unseen data it's not overfitting. Models are designed to find the optimal trade-off for you, hence early stopping and all that.</p>\n<p>That said, the true test is how the model does on \"unseen data\", any caveats to favoring #1 100% is if you think there's any sort of leakage between your train set and your validation set. The most trivial example is if observations (rows) aren't fully independent. Time series stock market data would be an example of that. Hundreds of thousands of credit card customers sounds pretty independent, though even there the train data (and thus the CV data used for early stopping) is from one time period, and the LB is from two other time periods. So it would be technically possible to overfit while seeing your CV score go up, and possibly finding a way of getting closer to #2 could help.</p>\n<p>In practice, for something like this: I'm thinking #1 with no real worries.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1896884,
      "author_name": "kaiserguo",
      "author_url": "",
      "post_date": "08/13/2022 08:39:36",
      "content": "<p>00000000000000</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1846849": "This topic is inspired on this twitter [poll](https://twitter.com/giffmana/status/1542241800647282688?t=I1IwilMHlKg9RJfdaQuUgA&s=09) from Lucas Beyer about what people have in mind when they think about overfitting.\n\nI've always thought about overfitting when the spread between train and validation metric increases. However, IIRC @jhoward said in one of the fast.ai videos that (in the context of NNs training) as long as your validation metric continues improving, you are good to go.\n\nLet's take a concrete example to illustrate this: @ragnar123 [notebook](https://www.kaggle.com/code/ragnar123/amex-lgbm-dart-cv-0-7963). The gap between CV train/val scores is much higher than the previous highest public notebooks (e.g. @cdeotte or @jiweiliu or @ambrosm notebooks), however the CV val (even running the CV for several seeds) and Public LB scores are also better. \n\nThe question is, if you have to choose between this two options, what would you choose?:\n\n1. A model with higher CV train/val scores gap but better CV val/Public LB scores\n2. A model with lower CV train/val scores gap but worse CV val/Public LB scores\n\nProbably it depends on how much the gap increases vs how much CV/LB improves, and that choice is a little bit of an art. However, I would be very happy to know your opinion based on your experience on this topic.\n\nPS: In one of the Kaggle Days Paris [talks](https://www.youtube.com/watch?v=VC8Jc9_lNoY&t=710s), @cpmpml talked about this topic.",
    "1846897": "I favor 1. (simply optimize CV score) by a wide margin. The concept of doing \"too well\" on training data in isolation doesn't make sense -- the best example is random forest, where the standard RF will almost perfectly predict the training data but of course have a much lower val score. But that's a quirk of the model design (ensembling many low bias, high variance learners)   and not proof of poor generalization behavior.\n\nStepping back, we want our models to have the best generalization performance possible, which is what hold-out sets directly proxy. A train-val gap is a much worse, model-dependent proxy. In my experience, the gap is most useful for model refinement on a relative basis -- for example, if I see that increasing `max_depth` and `num_leaves` in lgbm increases my train score a lot and my val score somewhat, that suggests that my model may still be underfitting. A more complex model might generalize better, even if the score gap is growing worse.",
    "1847193": "I only care about how my model will perform on unseen data. In my experience on kaggle, the model with the best validation score does the best on the test set, regardless of the training/validation gap. However, if two models have the same validation score, I prefer the one with a small train/validation gap. The best way to test out this hypothesis is by doing nested cross validation. This is when you hold out a true test set that is not seen during training, as opposed to a \"validation\" set that we are using for early stopping.",
    "1847469": "Great question. I think it really depends on the situation. \nIf you have a model with a higher CV train/val score gap, but the CV val/Public LB scores are better, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. \nHowever, if you have a model with a lower CV train/val score gap, but the CV val/Public LB scores are worse, I would choose that model. I think it really depends on how much the gap increases vs. how much the CV/LB score improves. \nHope this helps. Thanks!",
    "1849239": "Thanks for the answers. I think that it depends on the situation and therefore is tricky. All what you said totally makes sense to me guys, but from the other point of view, CPMP says in the video:\n\n> What I am doing when I am building new models, when I have a new CV score, I want it to be better but I also look at the training score, and if the training score improves too much, then my model is just overfitting and probably it won't fare well on the test data. So I always watch the gap between train and validation, and I avoid features that provide a very small CV improvement at the expense of very large train improvement. I don't even submit it.",
    "1852512": "Very good question. As long as the Eval improves, I push the training further, even if the train/Eval gap is high. Sometimes I see the training log loss function geting better and the eval amex-metric getting worse for several iterations until the eval suddenly jumps better. This is why I use a large early stopping. The best indicator of overfiting is not the train/Eval gap but the deterioration of the eval score.",
    "1853598": "I generally ignore the train/val CV gap and just focus on doing anything that improves my overall CV score. If you want to close the train/val CV gap, then you can add regularization to training, like data augmentation or model regularization. Making it harder for your model to memorize all the train data shrinks the gap and generally boosts your overall CV score.\n\nIn every competition I pay very close attention to the CV/LB gap. Both CV val and Public LB are unseen data. And if test data and train data come from the same distribution, then these two scores (CV and LB) should be the same. When they are not, it is upmost important to discover why. Then redesign your CV scheme to shrink the CV/LB gap, OR adjust training and/or post process to change LB predictions and shrink the gap. In most competitions, discovering why CV and LB have a gap, and discovering how to shrink it will significantly boost your LB.",
    "1853740": "Basically, the normal and the theoretical is 100% #1, if it does better on unseen data it's not overfitting. Models are designed to find the optimal trade-off for you, hence early stopping and all that.\n\nThat said, the true test is how the model does on \"unseen data\", any caveats to favoring #1 100% is if you think there's any sort of leakage between your train set and your validation set. The most trivial example is if observations (rows) aren't fully independent. Time series stock market data would be an example of that. Hundreds of thousands of credit card customers sounds pretty independent, though even there the train data (and thus the CV data used for early stopping) is from one time period, and the LB is from two other time periods. So it would be technically possible to overfit while seeing your CV score go up, and possibly finding a way of getting closer to #2 could help.\n\nIn practice, for something like this: I'm thinking #1 with no real worries.",
    "1896884": "00000000000000"
  },
  "source": "meta"
}