{
  "id": 54673,
  "title": "Importance of CV - LB constant?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54673",
  "author_name": "",
  "post_date": "2018-04-16T16:42:46.559127900Z",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>Participating in past competitions like Porto and Mercari I have heard about importance of CV -LB constant from experienced kaggler's(for ex. see this <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41037\">thread</a>) and have used it as one of the criteria to decide if my model is overfitting or not, and it has helped me selecting better submission. But in this competition from the beginning I am not able to keep CV - LB constant. If CV - LB is important, can anyone help me understand the importance of this constant scientifically?<br> And have you managed to keep your CV - LB a constant?<br>\nThanks<br>\nP.S: I tried googling related to this topic but couldn't find a good read, please point me to one if you know any. </p>",
  "messages": [
    {
      "id": "315026",
      "postDate": "04/16/2018 16:42:46",
      "content": "<p>Participating in past competitions like Porto and Mercari I have heard about importance of CV -LB constant from experienced kaggler's(for ex. see this <a href=\"https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41037\">thread</a>) and have used it as one of the criteria to decide if my model is overfitting or not, and it has helped me selecting better submission. But in this competition from the beginning I am not able to keep CV - LB constant. If CV - LB is important, can anyone help me understand the importance of this constant scientifically?<br> And have you managed to keep your CV - LB a constant?<br>\nThanks<br>\nP.S: I tried googling related to this topic but couldn't find a good read, please point me to one if you know any. </p>",
      "rawMarkdown": "Participating in past competitions like Porto and Mercari I have heard about importance of CV -LB constant from experienced kaggler's(for ex. see this [thread][1]) and have used it as one of the criteria to decide if my model is overfitting or not, and it has helped me selecting better submission. But in this competition from the beginning I am not able to keep CV - LB constant. If CV - LB is important, can anyone help me understand the importance of this constant scientifically?<br> And have you managed to keep your CV - LB a constant?<br>\nThanks<br>\nP.S: I tried googling related to this topic but couldn't find a good read, please point me to one if you know any. \n\n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41037",
      "votes": null
    },
    {
      "id": "315040",
      "postDate": "04/16/2018 17:03:17",
      "content": "<p>Is there an intuitive reason to want this \"CV minus LB\" to be constant? I guess this quantifies the generalisation loss from training data to a subset of the evaluation data, but why does it need to be constant? and what it is the meaning of this constant value?</p>",
      "rawMarkdown": "Is there an intuitive reason to want this \"CV minus LB\" to be constant? I guess this quantifies the generalisation loss from training data to a subset of the evaluation data, but why does it need to be constant? and what it is the meaning of this constant value?",
      "votes": null
    },
    {
      "id": "315049",
      "postDate": "04/16/2018 17:27:00",
      "content": "<p>I don't really understand why it should be constant either.</p>",
      "rawMarkdown": "I don't really understand why it should be constant either.",
      "votes": null
    },
    {
      "id": "315069",
      "postDate": "04/16/2018 17:58:58",
      "content": "<p>Exactly, I have been trying to get the same thing.</p>",
      "rawMarkdown": "Exactly, I have been trying to get the same thing.",
      "votes": null
    },
    {
      "id": "315073",
      "postDate": "04/16/2018 18:02:43",
      "content": "<p>I think constant CV/V - LB is a completely non-scientific luxury, not a necessity. I would never actively seek it out, because it could be a mistake to essentially overfit to a particular CV/V scheme just because it shows a lot of consistency with the public LB. That would mean your model feedback is too directly influenced by the particular public LB data, rather than being an undistorted reflection of how your model will generalize to actually unseen data. A validation scheme should be chosen because you can expect the validation sets to be statistically similar to the unseen data you want to predict on, not chosen based on seen data that you already have information about.</p>\n\n<p>I think that's true in general, but especially true of this competition where we should expect there to be statistical differences between hour 4 and the other test hours. Fishing for the right scheme/seed here to work with test hour 4 here would likely be actively detrimental given that test hour 4 will ultimately not matter for the private score.</p>\n\n<p>CV - LB constant might be nice for peace of mind, but it has 0 statistical rigor. What's more meaningful and useful is directional consistency in CV / LB improvements - when this happens, you're properly using the LB as an additional validation set to support your CV without sidetracking or overriding it. You can think of the LB as a less useful, but still relevant version of your local validation. </p>\n\n<p>For what it's worth, since I've been seeing validation hour 4 / public LB differences of less than .0004, I'd bet that there will be similar consistency between other validation test hours and private LB. I think that validation scheme (day 9) will tell you pretty much everything you need to know.  </p>",
      "rawMarkdown": "I think constant CV/V - LB is a completely non-scientific luxury, not a necessity. I would never actively seek it out, because it could be a mistake to essentially overfit to a particular CV/V scheme just because it shows a lot of consistency with the public LB. That would mean your model feedback is too directly influenced by the particular public LB data, rather than being an undistorted reflection of how your model will generalize to actually unseen data. A validation scheme should be chosen because you can expect the validation sets to be statistically similar to the unseen data you want to predict on, not chosen based on seen data that you already have information about.\n\nI think that's true in general, but especially true of this competition where we should expect there to be statistical differences between hour 4 and the other test hours. Fishing for the right scheme/seed here to work with test hour 4 here would likely be actively detrimental given that test hour 4 will ultimately not matter for the private score.\n\nCV - LB constant might be nice for peace of mind, but it has 0 statistical rigor. What's more meaningful and useful is directional consistency in CV / LB improvements - when this happens, you're properly using the LB as an additional validation set to support your CV without sidetracking or overriding it. You can think of the LB as a less useful, but still relevant version of your local validation. \n\nFor what it's worth, since I've been seeing validation hour 4 / public LB differences of less than .0004, I'd bet that there will be similar consistency between other validation test hours and private LB. I think that validation scheme (day 9) will tell you pretty much everything you need to know.",
      "votes": null
    },
    {
      "id": "315106",
      "postDate": "04/16/2018 18:50:12",
      "content": "<blockquote>\n  <p>If CV - LB is important, can anyone help me understand the importance of this constant scientifically?</p>\n</blockquote>\n\n<p>It is not a matter of science or any written rule. It is a matter of practical thinking when it comes to making final submission decisions. Remember, we get very limited feedback from the public LB - less than 20% of total test data.</p>\n\n<p>Let's say you develop a validation scheme and get your first local CV score. Now you submit and get your first LB feedback. So far so good, as there is nothing to be decided about one model. With each model after that you will get another local CV score and another LB score. The more scores you get, the more likely it is that some of the information will be conflicting. Sometimes your local CV jumps quite a bit but your LB score remains unchanged. Is that because you were overfitting locally, or because the subset of public LB data is not representative of the whole and is therefore underestimating your model? You will not get reliable answers to any of these questions from a single model, or even couple of them.</p>\n\n<p>It is a matter of finding a pattern that makes sense. Many competitors will tell you empirically that a constant (CV - LB) fits that pattern. That constant doesn't have to be a defined and strict number - it can be a range. For example, most of my (CV - LB) fits into 0.005-0.007 range. Few models are in 0.009-0.01 range.</p>\n\n<p>If CV and LB go up or down in similar fashion with different features and different ML approaches, that provides three very important points: 1) it increases our confidence that the distribution of public test data is similar to our train data; 2) it makes a case that we are making models that generalize well and are robust to different features and ML methods; 3) when we decide to ensemble models, only submissions that behave similarly in terms of (CV - LB) will mix with each other in predictable fashion. <strong>How are you going to get any of this info from submissions where (CV - LB) jumps all over the place?</strong></p>\n\n<p>To be fair, there is always that other option: we forget about all this (CV - LB) nonsense and pick a model with highest LB score. I recommend that to those who are fans of diving roller coasters.</p>",
      "rawMarkdown": "&gt; If CV - LB is important, can anyone help me understand the importance of this constant scientifically?\n\nIt is not a matter of science or any written rule. It is a matter of practical thinking when it comes to making final submission decisions. Remember, we get very limited feedback from the public LB - less than 20% of total test data.\n\nLet's say you develop a validation scheme and get your first local CV score. Now you submit and get your first LB feedback. So far so good, as there is nothing to be decided about one model. With each model after that you will get another local CV score and another LB score. The more scores you get, the more likely it is that some of the information will be conflicting. Sometimes your local CV jumps quite a bit but your LB score remains unchanged. Is that because you were overfitting locally, or because the subset of public LB data is not representative of the whole and is therefore underestimating your model? You will not get reliable answers to any of these questions from a single model, or even couple of them.\n\nIt is a matter of finding a pattern that makes sense. Many competitors will tell you empirically that a constant (CV - LB) fits that pattern. That constant doesn't have to be a defined and strict number - it can be a range. For example, most of my (CV - LB) fits into 0.005-0.007 range. Few models are in 0.009-0.01 range.\n\nIf CV and LB go up or down in similar fashion with different features and different ML approaches, that provides three very important points: 1) it increases our confidence that the distribution of public test data is similar to our train data; 2) it makes a case that we are making models that generalize well and are robust to different features and ML methods; 3) when we decide to ensemble models, only submissions that behave similarly in terms of (CV - LB) will mix with each other in predictable fashion. __How are you going to get any of this info from submissions where (CV - LB) jumps all over the place?__\n\nTo be fair, there is always that other option: we forget about all this (CV - LB) nonsense and pick a model with highest LB score. I recommend that to those who are fans of diving roller coasters.",
      "votes": null
    },
    {
      "id": "315108",
      "postDate": "04/16/2018 18:55:53",
      "content": "<p>I'm happy when I find a CV setting such that CV  and LB are evolving the same way: a progress on CV results in a progress in LB.  In that case I can ignore the LB and perform as many experiments I want instead of being limited by the daily submission limit.</p>\n\n<p>Having a constant difference does not matter much.</p>",
      "rawMarkdown": "I'm happy when I find a CV setting such that CV  and LB are evolving the same way: a progress on CV results in a progress in LB.  In that case I can ignore the LB and perform as many experiments I want instead of being limited by the daily submission limit.\n\nHaving a constant difference does not matter much.",
      "votes": null
    },
    {
      "id": "315122",
      "postDate": "04/16/2018 19:13:19",
      "content": "<p>Thanks @Joe for great explanation.</p>",
      "rawMarkdown": "Thanks @Joe for great explanation.",
      "votes": null
    },
    {
      "id": "315124",
      "postDate": "04/16/2018 19:16:24",
      "content": "<p>This totally makes sense, thank you very much @tilii for explaining this.</p>",
      "rawMarkdown": "This totally makes sense, thank you very much @tilii for explaining this.",
      "votes": null
    },
    {
      "id": "315126",
      "postDate": "04/16/2018 19:18:05",
      "content": "<p>So far I can predict my public LB seeing my validation score, I hope it stays like this.</p>",
      "rawMarkdown": "So far I can predict my public LB seeing my validation score, I hope it stays like this.",
      "votes": null
    },
    {
      "id": "315936",
      "postDate": "04/17/2018 20:35:57",
      "content": "<p>As others have said, the important thing is that changes in the LB score move in the same direction as changes in the CV score.  It's not reasonable to expect the difference to be constant.  I would argue that constancy might be a sign of underfitting.  You are making improvements to your model based on feedback from your CV.  It's very hard, often impossible, to know which of these improvements will generalize beyond your CV.  The LB is a sanity check on that process but not a perfect guide.  </p>\n\n<p>Typically you will be both underfitting and overfitting your CV data in different respects.  If you make several changes that improve your CV score but find that your LB score has not improved, that may be an indication that <em>on balance</em> the changes are overfitting your CV data.  </p>\n\n<p>On the other hand, if you make several changes that improve your CV score and find that your LB score has also improved but not as much as your CV score, you don't know which improvements are generalizing well and which aren't.  But at least you know that <em>on balance</em> you're not <em>severely</em> overfitting your CV data.  </p>\n\n<p>If you make changes that improve your CV score, and your LB score goes up by exactly the same amount, it would be very optimistic to interpret that to mean that your improvements are generalizing perfectly.  More likely some of them are underfitting and some are overfitting.  But you know that <em>on balance</em> you're not overfitting.  On the margin, your model would probably benefit from more aggressive optimization.</p>",
      "rawMarkdown": "As others have said, the important thing is that changes in the LB score move in the same direction as changes in the CV score.  It's not reasonable to expect the difference to be constant.  I would argue that constancy might be a sign of underfitting.  You are making improvements to your model based on feedback from your CV.  It's very hard, often impossible, to know which of these improvements will generalize beyond your CV.  The LB is a sanity check on that process but not a perfect guide.  \n\nTypically you will be both underfitting and overfitting your CV data in different respects.  If you make several changes that improve your CV score but find that your LB score has not improved, that may be an indication that *on balance* the changes are overfitting your CV data.  \n\nOn the other hand, if you make several changes that improve your CV score and find that your LB score has also improved but not as much as your CV score, you don't know which improvements are generalizing well and which aren't.  But at least you know that *on balance* you're not *severely* overfitting your CV data.  \n\nIf you make changes that improve your CV score, and your LB score goes up by exactly the same amount, it would be very optimistic to interpret that to mean that your improvements are generalizing perfectly.  More likely some of them are underfitting and some are overfitting.  But you know that *on balance* you're not overfitting.  On the margin, your model would probably benefit from more aggressive optimization.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 315040,
      "author_name": "gimunu",
      "author_url": "",
      "post_date": "04/16/2018 17:03:17",
      "content": "<p>Is there an intuitive reason to want this \"CV minus LB\" to be constant? I guess this quantifies the generalisation loss from training data to a subset of the evaluation data, but why does it need to be constant? and what it is the meaning of this constant value?</p>",
      "votes": null,
      "replies": [
        {
          "id": 315069,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/16/2018 17:58:58",
          "content": "<p>Exactly, I have been trying to get the same thing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 315049,
      "author_name": "david26694",
      "author_url": "",
      "post_date": "04/16/2018 17:27:00",
      "content": "<p>I don't really understand why it should be constant either.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 315073,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/16/2018 18:02:43",
      "content": "<p>I think constant CV/V - LB is a completely non-scientific luxury, not a necessity. I would never actively seek it out, because it could be a mistake to essentially overfit to a particular CV/V scheme just because it shows a lot of consistency with the public LB. That would mean your model feedback is too directly influenced by the particular public LB data, rather than being an undistorted reflection of how your model will generalize to actually unseen data. A validation scheme should be chosen because you can expect the validation sets to be statistically similar to the unseen data you want to predict on, not chosen based on seen data that you already have information about.</p>\n\n<p>I think that's true in general, but especially true of this competition where we should expect there to be statistical differences between hour 4 and the other test hours. Fishing for the right scheme/seed here to work with test hour 4 here would likely be actively detrimental given that test hour 4 will ultimately not matter for the private score.</p>\n\n<p>CV - LB constant might be nice for peace of mind, but it has 0 statistical rigor. What's more meaningful and useful is directional consistency in CV / LB improvements - when this happens, you're properly using the LB as an additional validation set to support your CV without sidetracking or overriding it. You can think of the LB as a less useful, but still relevant version of your local validation. </p>\n\n<p>For what it's worth, since I've been seeing validation hour 4 / public LB differences of less than .0004, I'd bet that there will be similar consistency between other validation test hours and private LB. I think that validation scheme (day 9) will tell you pretty much everything you need to know.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 315122,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/16/2018 19:13:19",
          "content": "<p>Thanks @Joe for great explanation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 315106,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/16/2018 18:50:12",
      "content": "<blockquote>\n  <p>If CV - LB is important, can anyone help me understand the importance of this constant scientifically?</p>\n</blockquote>\n\n<p>It is not a matter of science or any written rule. It is a matter of practical thinking when it comes to making final submission decisions. Remember, we get very limited feedback from the public LB - less than 20% of total test data.</p>\n\n<p>Let's say you develop a validation scheme and get your first local CV score. Now you submit and get your first LB feedback. So far so good, as there is nothing to be decided about one model. With each model after that you will get another local CV score and another LB score. The more scores you get, the more likely it is that some of the information will be conflicting. Sometimes your local CV jumps quite a bit but your LB score remains unchanged. Is that because you were overfitting locally, or because the subset of public LB data is not representative of the whole and is therefore underestimating your model? You will not get reliable answers to any of these questions from a single model, or even couple of them.</p>\n\n<p>It is a matter of finding a pattern that makes sense. Many competitors will tell you empirically that a constant (CV - LB) fits that pattern. That constant doesn't have to be a defined and strict number - it can be a range. For example, most of my (CV - LB) fits into 0.005-0.007 range. Few models are in 0.009-0.01 range.</p>\n\n<p>If CV and LB go up or down in similar fashion with different features and different ML approaches, that provides three very important points: 1) it increases our confidence that the distribution of public test data is similar to our train data; 2) it makes a case that we are making models that generalize well and are robust to different features and ML methods; 3) when we decide to ensemble models, only submissions that behave similarly in terms of (CV - LB) will mix with each other in predictable fashion. <strong>How are you going to get any of this info from submissions where (CV - LB) jumps all over the place?</strong></p>\n\n<p>To be fair, there is always that other option: we forget about all this (CV - LB) nonsense and pick a model with highest LB score. I recommend that to those who are fans of diving roller coasters.</p>",
      "votes": null,
      "replies": [
        {
          "id": 315124,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/16/2018 19:16:24",
          "content": "<p>This totally makes sense, thank you very much @tilii for explaining this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 315108,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "04/16/2018 18:55:53",
      "content": "<p>I'm happy when I find a CV setting such that CV  and LB are evolving the same way: a progress on CV results in a progress in LB.  In that case I can ignore the LB and perform as many experiments I want instead of being limited by the daily submission limit.</p>\n\n<p>Having a constant difference does not matter much.</p>",
      "votes": null,
      "replies": [
        {
          "id": 315126,
          "author_name": "sohaibomar",
          "author_url": "",
          "post_date": "04/16/2018 19:18:05",
          "content": "<p>So far I can predict my public LB seeing my validation score, I hope it stays like this.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 315936,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "04/17/2018 20:35:57",
      "content": "<p>As others have said, the important thing is that changes in the LB score move in the same direction as changes in the CV score.  It's not reasonable to expect the difference to be constant.  I would argue that constancy might be a sign of underfitting.  You are making improvements to your model based on feedback from your CV.  It's very hard, often impossible, to know which of these improvements will generalize beyond your CV.  The LB is a sanity check on that process but not a perfect guide.  </p>\n\n<p>Typically you will be both underfitting and overfitting your CV data in different respects.  If you make several changes that improve your CV score but find that your LB score has not improved, that may be an indication that <em>on balance</em> the changes are overfitting your CV data.  </p>\n\n<p>On the other hand, if you make several changes that improve your CV score and find that your LB score has also improved but not as much as your CV score, you don't know which improvements are generalizing well and which aren't.  But at least you know that <em>on balance</em> you're not <em>severely</em> overfitting your CV data.  </p>\n\n<p>If you make changes that improve your CV score, and your LB score goes up by exactly the same amount, it would be very optimistic to interpret that to mean that your improvements are generalizing perfectly.  More likely some of them are underfitting and some are overfitting.  But you know that <em>on balance</em> you're not overfitting.  On the margin, your model would probably benefit from more aggressive optimization.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "315026": "Participating in past competitions like Porto and Mercari I have heard about importance of CV -LB constant from experienced kaggler's(for ex. see this [thread][1]) and have used it as one of the criteria to decide if my model is overfitting or not, and it has helped me selecting better submission. But in this competition from the beginning I am not able to keep CV - LB constant. If CV - LB is important, can anyone help me understand the importance of this constant scientifically?<br> And have you managed to keep your CV - LB a constant?<br>\nThanks<br>\nP.S: I tried googling related to this topic but couldn't find a good read, please point me to one if you know any. \n\n\n\n  [1]: https://www.kaggle.com/c/porto-seguro-safe-driver-prediction/discussion/41037",
    "315040": "Is there an intuitive reason to want this \"CV minus LB\" to be constant? I guess this quantifies the generalisation loss from training data to a subset of the evaluation data, but why does it need to be constant? and what it is the meaning of this constant value?",
    "315049": "I don't really understand why it should be constant either.",
    "315069": "Exactly, I have been trying to get the same thing.",
    "315073": "I think constant CV/V - LB is a completely non-scientific luxury, not a necessity. I would never actively seek it out, because it could be a mistake to essentially overfit to a particular CV/V scheme just because it shows a lot of consistency with the public LB. That would mean your model feedback is too directly influenced by the particular public LB data, rather than being an undistorted reflection of how your model will generalize to actually unseen data. A validation scheme should be chosen because you can expect the validation sets to be statistically similar to the unseen data you want to predict on, not chosen based on seen data that you already have information about.\n\nI think that's true in general, but especially true of this competition where we should expect there to be statistical differences between hour 4 and the other test hours. Fishing for the right scheme/seed here to work with test hour 4 here would likely be actively detrimental given that test hour 4 will ultimately not matter for the private score.\n\nCV - LB constant might be nice for peace of mind, but it has 0 statistical rigor. What's more meaningful and useful is directional consistency in CV / LB improvements - when this happens, you're properly using the LB as an additional validation set to support your CV without sidetracking or overriding it. You can think of the LB as a less useful, but still relevant version of your local validation. \n\nFor what it's worth, since I've been seeing validation hour 4 / public LB differences of less than .0004, I'd bet that there will be similar consistency between other validation test hours and private LB. I think that validation scheme (day 9) will tell you pretty much everything you need to know.",
    "315106": "&gt; If CV - LB is important, can anyone help me understand the importance of this constant scientifically?\n\nIt is not a matter of science or any written rule. It is a matter of practical thinking when it comes to making final submission decisions. Remember, we get very limited feedback from the public LB - less than 20% of total test data.\n\nLet's say you develop a validation scheme and get your first local CV score. Now you submit and get your first LB feedback. So far so good, as there is nothing to be decided about one model. With each model after that you will get another local CV score and another LB score. The more scores you get, the more likely it is that some of the information will be conflicting. Sometimes your local CV jumps quite a bit but your LB score remains unchanged. Is that because you were overfitting locally, or because the subset of public LB data is not representative of the whole and is therefore underestimating your model? You will not get reliable answers to any of these questions from a single model, or even couple of them.\n\nIt is a matter of finding a pattern that makes sense. Many competitors will tell you empirically that a constant (CV - LB) fits that pattern. That constant doesn't have to be a defined and strict number - it can be a range. For example, most of my (CV - LB) fits into 0.005-0.007 range. Few models are in 0.009-0.01 range.\n\nIf CV and LB go up or down in similar fashion with different features and different ML approaches, that provides three very important points: 1) it increases our confidence that the distribution of public test data is similar to our train data; 2) it makes a case that we are making models that generalize well and are robust to different features and ML methods; 3) when we decide to ensemble models, only submissions that behave similarly in terms of (CV - LB) will mix with each other in predictable fashion. __How are you going to get any of this info from submissions where (CV - LB) jumps all over the place?__\n\nTo be fair, there is always that other option: we forget about all this (CV - LB) nonsense and pick a model with highest LB score. I recommend that to those who are fans of diving roller coasters.",
    "315108": "I'm happy when I find a CV setting such that CV  and LB are evolving the same way: a progress on CV results in a progress in LB.  In that case I can ignore the LB and perform as many experiments I want instead of being limited by the daily submission limit.\n\nHaving a constant difference does not matter much.",
    "315122": "Thanks @Joe for great explanation.",
    "315124": "This totally makes sense, thank you very much @tilii for explaining this.",
    "315126": "So far I can predict my public LB seeing my validation score, I hope it stays like this.",
    "315936": "As others have said, the important thing is that changes in the LB score move in the same direction as changes in the CV score.  It's not reasonable to expect the difference to be constant.  I would argue that constancy might be a sign of underfitting.  You are making improvements to your model based on feedback from your CV.  It's very hard, often impossible, to know which of these improvements will generalize beyond your CV.  The LB is a sanity check on that process but not a perfect guide.  \n\nTypically you will be both underfitting and overfitting your CV data in different respects.  If you make several changes that improve your CV score but find that your LB score has not improved, that may be an indication that *on balance* the changes are overfitting your CV data.  \n\nOn the other hand, if you make several changes that improve your CV score and find that your LB score has also improved but not as much as your CV score, you don't know which improvements are generalizing well and which aren't.  But at least you know that *on balance* you're not *severely* overfitting your CV data.  \n\nIf you make changes that improve your CV score, and your LB score goes up by exactly the same amount, it would be very optimistic to interpret that to mean that your improvements are generalizing perfectly.  More likely some of them are underfitting and some are overfitting.  But you know that *on balance* you're not overfitting.  On the margin, your model would probably benefit from more aggressive optimization."
  },
  "source": "meta"
}