{
  "id": 202002,
  "title": "CV improvements stop translating to LB improvements",
  "url": "/competitions/predict-volcanic-eruptions-ingv-oe/discussion/202002",
  "author_name": "",
  "post_date": "2020-12-07T19:54:04.284235Z",
  "votes": 1,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Recently, I've noticed that improvements in my CV score do not translate to improvements in my LB score any more. Up until ~ 4.7M LB error, significant improvements in CV scores led to better LB scores . But after that, material improvements in my CV have resulted in worse LB performance. I tried to use repeated 5-fold CV with different seeds to have a more stable validation and to avoid overfitting to fixed validation folds, but I've still been experiencing this problem. This occurs when I add promising features to the existing set of ~120 features using my best LGBM model.</p>\n<p>Has anyone experienced similar behavior here or in other competitions? Can anyone suggest what might be the reason for this phenomenon?</p>",
  "messages": [
    {
      "id": "1105346",
      "postDate": "12/07/2020 19:54:04",
      "content": "<p>Recently, I've noticed that improvements in my CV score do not translate to improvements in my LB score any more. Up until ~ 4.7M LB error, significant improvements in CV scores led to better LB scores . But after that, material improvements in my CV have resulted in worse LB performance. I tried to use repeated 5-fold CV with different seeds to have a more stable validation and to avoid overfitting to fixed validation folds, but I've still been experiencing this problem. This occurs when I add promising features to the existing set of ~120 features using my best LGBM model.</p>\n<p>Has anyone experienced similar behavior here or in other competitions? Can anyone suggest what might be the reason for this phenomenon?</p>",
      "rawMarkdown": "Recently, I've noticed that improvements in my CV score do not translate to improvements in my LB score any more. Up until ~ 4.7M LB error, significant improvements in CV scores led to better LB scores . But after that, material improvements in my CV have resulted in worse LB performance. I tried to use repeated 5-fold CV with different seeds to have a more stable validation and to avoid overfitting to fixed validation folds, but I've still been experiencing this problem. This occurs when I add promising features to the existing set of ~120 features using my best LGBM model.\n\nHas anyone experienced similar behavior here or in other competitions? Can anyone suggest what might be the reason for this phenomenon?",
      "votes": null
    },
    {
      "id": "1112467",
      "postDate": "12/14/2020 16:21:04",
      "content": "<p>Hi. Yes. Many have this problem. <br>\nPlot a histogram of the predictions of the target variable from your best prediction trial for the LB and compare it to an histogram of your CV predictions from the train set. <br>\nYou will notice a huge difference in the distributions (train is almost an uniform distribution, test set predictions will have more of a shape). If you repeat that for some of your features with the most explanatory power and run KS tests vs the feature of the train set you will come to the conclusion that the test set has a different distribution than the train set.</p>\n<p>I trained a simple model and realized that quite early and stopped submitting. Not sure if I want keep competing.  </p>",
      "rawMarkdown": "Hi. Yes. Many have this problem. \nPlot a histogram of the predictions of the target variable from your best prediction trial for the LB and compare it to an histogram of your CV predictions from the train set. \nYou will notice a huge difference in the distributions (train is almost an uniform distribution, test set predictions will have more of a shape). If you repeat that for some of your features with the most explanatory power and run KS tests vs the feature of the train set you will come to the conclusion that the test set has a different distribution than the train set.\n\nI trained a simple model and realized that quite early and stopped submitting. Not sure if I want keep competing.",
      "votes": null
    },
    {
      "id": "1112521",
      "postDate": "12/14/2020 17:13:45",
      "content": "<p>Thanks for your input, I also noticed that. That's why I only update my best model if it has a significant increase both in terms of CV and LB. In the beginning, it was easy as CV improvements translated to LB improvements, though there has been a huge gap between the two. Recently, I've found that the improvements in CV makes my LB worse, I might just overfit to the training data.</p>",
      "rawMarkdown": "Thanks for your input, I also noticed that. That's why I only update my best model if it has a significant increase both in terms of CV and LB. In the beginning, it was easy as CV improvements translated to LB improvements, though there has been a huge gap between the two. Recently, I've found that the improvements in CV makes my LB worse, I might just overfit to the training data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1112467,
      "author_name": "pldatacybernetics",
      "author_url": "",
      "post_date": "12/14/2020 16:21:04",
      "content": "<p>Hi. Yes. Many have this problem. <br>\nPlot a histogram of the predictions of the target variable from your best prediction trial for the LB and compare it to an histogram of your CV predictions from the train set. <br>\nYou will notice a huge difference in the distributions (train is almost an uniform distribution, test set predictions will have more of a shape). If you repeat that for some of your features with the most explanatory power and run KS tests vs the feature of the train set you will come to the conclusion that the test set has a different distribution than the train set.</p>\n<p>I trained a simple model and realized that quite early and stopped submitting. Not sure if I want keep competing.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1112521,
          "author_name": "leventelippenszky",
          "author_url": "",
          "post_date": "12/14/2020 17:13:45",
          "content": "<p>Thanks for your input, I also noticed that. That's why I only update my best model if it has a significant increase both in terms of CV and LB. In the beginning, it was easy as CV improvements translated to LB improvements, though there has been a huge gap between the two. Recently, I've found that the improvements in CV makes my LB worse, I might just overfit to the training data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1105346": "Recently, I've noticed that improvements in my CV score do not translate to improvements in my LB score any more. Up until ~ 4.7M LB error, significant improvements in CV scores led to better LB scores . But after that, material improvements in my CV have resulted in worse LB performance. I tried to use repeated 5-fold CV with different seeds to have a more stable validation and to avoid overfitting to fixed validation folds, but I've still been experiencing this problem. This occurs when I add promising features to the existing set of ~120 features using my best LGBM model.\n\nHas anyone experienced similar behavior here or in other competitions? Can anyone suggest what might be the reason for this phenomenon?",
    "1112467": "Hi. Yes. Many have this problem. \nPlot a histogram of the predictions of the target variable from your best prediction trial for the LB and compare it to an histogram of your CV predictions from the train set. \nYou will notice a huge difference in the distributions (train is almost an uniform distribution, test set predictions will have more of a shape). If you repeat that for some of your features with the most explanatory power and run KS tests vs the feature of the train set you will come to the conclusion that the test set has a different distribution than the train set.\n\nI trained a simple model and realized that quite early and stopped submitting. Not sure if I want keep competing.",
    "1112521": "Thanks for your input, I also noticed that. That's why I only update my best model if it has a significant increase both in terms of CV and LB. In the beginning, it was easy as CV improvements translated to LB improvements, though there has been a huge gap between the two. Recently, I've found that the improvements in CV makes my LB worse, I might just overfit to the training data."
  },
  "source": "meta"
}