{
  "id": 90416,
  "title": "Differences in CV score when using XGBoost",
  "url": "/competitions/LANL-Earthquake-Prediction/discussion/90416",
  "author_name": "",
  "post_date": "2019-04-23T16:49:33.018546600Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Recently, I have been formalizing the way I am validating my models when I found out something that is puzzling me.</p>\n\n<p>Essentially, I use a KFold with k=5 and shuffling to compute my CV mean and std score. I am using <code>xgboost.cv</code> within a loop that searcher the best set of hyperparameters.</p>\n\n<p>The code looks something like this:\n<code>python\nresult = xgb.cv(params,\n    dataset,\n    nfold=5,\n    num_boost_round=20000,\n    early_stopping_rounds=200,\n    stratified=False)\n</code></p>\n\n<p>To retrieve the CV score, I just get the values from <code>result</code>:\n<code>python\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n</code></p>\n\n<p>My surprise comes when I use a more standard <code>KFold</code> in combination with</p>\n\n<p><code>\nmodel = xgb.XGBRegressor(**best_params)\nmodel.fit(...)\n</code>\nor even\n<code>\nmodel = xgb.train(..., params=best_params)\n</code>\nThen, I compute the score by training and predicting on the validation set:\n<code>\nscore = metrics.mean_absolute_error(val_y, val_pred)\n</code>\nand I obtain a lower std.</p>\n\n<p>Here you have the plot with some of the search I've done:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/521938/13054/download-1.png\" alt=\"scatterplot\"></p>\n\n<p>How would you explain the difference between the yellow and green dot? Where do you think that this difference comes from? Could this difference be attributed only to the shuffling?</p>",
  "messages": [
    {
      "id": "521938",
      "postDate": "04/23/2019 16:49:33",
      "content": "<p>Recently, I have been formalizing the way I am validating my models when I found out something that is puzzling me.</p>\n\n<p>Essentially, I use a KFold with k=5 and shuffling to compute my CV mean and std score. I am using <code>xgboost.cv</code> within a loop that searcher the best set of hyperparameters.</p>\n\n<p>The code looks something like this:\n<code>python\nresult = xgb.cv(params,\n    dataset,\n    nfold=5,\n    num_boost_round=20000,\n    early_stopping_rounds=200,\n    stratified=False)\n</code></p>\n\n<p>To retrieve the CV score, I just get the values from <code>result</code>:\n<code>python\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n</code></p>\n\n<p>My surprise comes when I use a more standard <code>KFold</code> in combination with</p>\n\n<p><code>\nmodel = xgb.XGBRegressor(**best_params)\nmodel.fit(...)\n</code>\nor even\n<code>\nmodel = xgb.train(..., params=best_params)\n</code>\nThen, I compute the score by training and predicting on the validation set:\n<code>\nscore = metrics.mean_absolute_error(val_y, val_pred)\n</code>\nand I obtain a lower std.</p>\n\n<p>Here you have the plot with some of the search I've done:\n<img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/521938/13054/download-1.png\" alt=\"scatterplot\"></p>\n\n<p>How would you explain the difference between the yellow and green dot? Where do you think that this difference comes from? Could this difference be attributed only to the shuffling?</p>",
      "rawMarkdown": "Recently, I have been formalizing the way I am validating my models when I found out something that is puzzling me.\n\nEssentially, I use a KFold with k=5 and shuffling to compute my CV mean and std score. I am using `xgboost.cv` within a loop that searcher the best set of hyperparameters.\n\nThe code looks something like this:\n```python\nresult = xgb.cv(params,\n    dataset,\n    nfold=5,\n    num_boost_round=20000,\n    early_stopping_rounds=200,\n    stratified=False)\n```\n\nTo retrieve the CV score, I just get the values from `result`:\n```python\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n```\n\nMy surprise comes when I use a more standard `KFold` in combination with\n\n```\nmodel = xgb.XGBRegressor(**best_params)\nmodel.fit(...)\n```\nor even\n```\nmodel = xgb.train(..., params=best_params)\n```\nThen, I compute the score by training and predicting on the validation set:\n```\nscore = metrics.mean_absolute_error(val_y, val_pred)\n```\nand I obtain a lower std.\n\nHere you have the plot with some of the search I've done:\n![scatterplot](https://storage.googleapis.com/kaggle-forum-message-attachments/521938/13054/download-1.png)\n\nHow would you explain the difference between the yellow and green dot? Where do you think that this difference comes from? Could this difference be attributed only to the shuffling?",
      "votes": null
    },
    {
      "id": "521982",
      "postDate": "04/23/2019 18:02:31",
      "content": "<p>Maybe because it is using the model of the best iteration for prediction, not the last one before stopping</p>",
      "rawMarkdown": "Maybe because it is using the model of the best iteration for prediction, not the last one before stopping",
      "votes": null
    },
    {
      "id": "522060",
      "postDate": "04/23/2019 20:09:04",
      "content": "<p>Do you mean that the lines\n<code>\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n</code>\nare just returning the score of the last fold?</p>",
      "rawMarkdown": "Do you mean that the lines\n```\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n```\nare just returning the score of the last fold?",
      "votes": null
    },
    {
      "id": "522099",
      "postDate": "04/23/2019 21:27:34",
      "content": "<p>No, the last iteration, not last fold. </p>",
      "rawMarkdown": "No, the last iteration, not last fold.",
      "votes": null
    },
    {
      "id": "588808",
      "postDate": "07/31/2019 04:09:19",
      "content": "<p>The code</p>\n\n<p>result['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]</p>\n\n<p>must return the scores for best iteration.</p>",
      "rawMarkdown": "The code\n \nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n\nmust return the scores for best iteration.",
      "votes": null
    },
    {
      "id": "588813",
      "postDate": "07/31/2019 04:19:59",
      "content": "<p>It's hard to answer without seeing the implementation itself. I can assume that the case in shuffling because by default \"shuffle=False\" for KFold while \"shuffle=True\" for xgb.cv . By the way we can use sklearn.model_selection.cross_val_score to join splitting (KFold in your case) and  actually validation. How the results may differ (due to shuffling ) from xgb.cv and sklearn.model_selection.cross_val_score is shown by link: <a href=\"https://stackoverflow.com/questions/41135987/why-xgboost-cv-and-sklearn-cross-val-score-give-different-results\">https://stackoverflow.com/questions/41135987/why-xgboost-cv-and-sklearn-cross-val-score-give-different-results</a></p>",
      "rawMarkdown": "It's hard to answer without seeing the implementation itself. I can assume that the case in shuffling because by default \"shuffle=False\" for KFold while \"shuffle=True\" for xgb.cv . By the way we can use sklearn.model_selection.cross_val_score to join splitting (KFold in your case) and  actually validation. How the results may differ (due to shuffling ) from xgb.cv and sklearn.model_selection.cross_val_score is shown by link: https://stackoverflow.com/questions/41135987/why-xgboost-cv-and-sklearn-cross-val-score-give-different-results",
      "votes": null
    },
    {
      "id": "646057",
      "postDate": "10/10/2019 19:57:13",
      "content": "<p>Why do we use \"iloc[-1]\" from the cv result that is the MINIMUM instead of finding the MEAN over all the folds to compare RMSE while hyperparameter optimisation. </p>",
      "rawMarkdown": "Why do we use \"iloc[-1]\" from the cv result that is the MINIMUM instead of finding the MEAN over all the folds to compare RMSE while hyperparameter optimisation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 521982,
      "author_name": "amjad85",
      "author_url": "",
      "post_date": "04/23/2019 18:02:31",
      "content": "<p>Maybe because it is using the model of the best iteration for prediction, not the last one before stopping</p>",
      "votes": null,
      "replies": [
        {
          "id": 522060,
          "author_name": "ricarddelgado",
          "author_url": "",
          "post_date": "04/23/2019 20:09:04",
          "content": "<p>Do you mean that the lines\n<code>\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n</code>\nare just returning the score of the last fold?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 522099,
          "author_name": "amjad85",
          "author_url": "",
          "post_date": "04/23/2019 21:27:34",
          "content": "<p>No, the last iteration, not last fold. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 588808,
          "author_name": "dushepa",
          "author_url": "",
          "post_date": "07/31/2019 04:09:19",
          "content": "<p>The code</p>\n\n<p>result['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]</p>\n\n<p>must return the scores for best iteration.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 646057,
          "author_name": "kevinvkasundra",
          "author_url": "",
          "post_date": "10/10/2019 19:57:13",
          "content": "<p>Why do we use \"iloc[-1]\" from the cv result that is the MINIMUM instead of finding the MEAN over all the folds to compare RMSE while hyperparameter optimisation. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 588813,
      "author_name": "dushepa",
      "author_url": "",
      "post_date": "07/31/2019 04:19:59",
      "content": "<p>It's hard to answer without seeing the implementation itself. I can assume that the case in shuffling because by default \"shuffle=False\" for KFold while \"shuffle=True\" for xgb.cv . By the way we can use sklearn.model_selection.cross_val_score to join splitting (KFold in your case) and  actually validation. How the results may differ (due to shuffling ) from xgb.cv and sklearn.model_selection.cross_val_score is shown by link: <a href=\"https://stackoverflow.com/questions/41135987/why-xgboost-cv-and-sklearn-cross-val-score-give-different-results\">https://stackoverflow.com/questions/41135987/why-xgboost-cv-and-sklearn-cross-val-score-give-different-results</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "521938": "Recently, I have been formalizing the way I am validating my models when I found out something that is puzzling me.\n\nEssentially, I use a KFold with k=5 and shuffling to compute my CV mean and std score. I am using `xgboost.cv` within a loop that searcher the best set of hyperparameters.\n\nThe code looks something like this:\n```python\nresult = xgb.cv(params,\n    dataset,\n    nfold=5,\n    num_boost_round=20000,\n    early_stopping_rounds=200,\n    stratified=False)\n```\n\nTo retrieve the CV score, I just get the values from `result`:\n```python\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n```\n\nMy surprise comes when I use a more standard `KFold` in combination with\n\n```\nmodel = xgb.XGBRegressor(**best_params)\nmodel.fit(...)\n```\nor even\n```\nmodel = xgb.train(..., params=best_params)\n```\nThen, I compute the score by training and predicting on the validation set:\n```\nscore = metrics.mean_absolute_error(val_y, val_pred)\n```\nand I obtain a lower std.\n\nHere you have the plot with some of the search I've done:\n![scatterplot](https://storage.googleapis.com/kaggle-forum-message-attachments/521938/13054/download-1.png)\n\nHow would you explain the difference between the yellow and green dot? Where do you think that this difference comes from? Could this difference be attributed only to the shuffling?",
    "521982": "Maybe because it is using the model of the best iteration for prediction, not the last one before stopping",
    "522060": "Do you mean that the lines\n```\nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n```\nare just returning the score of the last fold?",
    "522099": "No, the last iteration, not last fold.",
    "588808": "The code\n \nresult['test-mae-mean'].iloc[-1]\nresult['test-mae-std'].iloc[-1]\n\nmust return the scores for best iteration.",
    "588813": "It's hard to answer without seeing the implementation itself. I can assume that the case in shuffling because by default \"shuffle=False\" for KFold while \"shuffle=True\" for xgb.cv . By the way we can use sklearn.model_selection.cross_val_score to join splitting (KFold in your case) and  actually validation. How the results may differ (due to shuffling ) from xgb.cv and sklearn.model_selection.cross_val_score is shown by link: https://stackoverflow.com/questions/41135987/why-xgboost-cv-and-sklearn-cross-val-score-give-different-results",
    "646057": "Why do we use \"iloc[-1]\" from the cv result that is the MINIMUM instead of finding the MEAN over all the folds to compare RMSE while hyperparameter optimisation."
  },
  "source": "meta"
}