{
  "id": 59335,
  "title": "Xgboost!",
  "url": "/competitions/avito-demand-prediction/discussion/59335",
  "author_name": "",
  "post_date": "2018-06-21T06:45:02.481765100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I ran xgboost with the same feature pack as my lgbm (2200 LB) and got 2160 CV but 2300 LB. What might be the reason of such huge difference between LB and CV? The correlation between xgb and lgbm submission is about 0.97 but 100 points diff in LB. Has anybody solved this issue?</p>",
  "messages": [
    {
      "id": "346142",
      "postDate": "06/21/2018 06:45:02",
      "content": "<p>I ran xgboost with the same feature pack as my lgbm (2200 LB) and got 2160 CV but 2300 LB. What might be the reason of such huge difference between LB and CV? The correlation between xgb and lgbm submission is about 0.97 but 100 points diff in LB. Has anybody solved this issue?</p>",
      "rawMarkdown": "I ran xgboost with the same feature pack as my lgbm (2200 LB) and got 2160 CV but 2300 LB. What might be the reason of such huge difference between LB and CV? The correlation between xgb and lgbm submission is about 0.97 but 100 points diff in LB. Has anybody solved this issue?",
      "votes": null
    },
    {
      "id": "346300",
      "postDate": "06/21/2018 12:41:56",
      "content": "<p>When I read the title, I thought you got big success with Xgboost.... Anyway, good luck.</p>",
      "rawMarkdown": "When I read the title, I thought you got big success with Xgboost.... Anyway, good luck.",
      "votes": null
    },
    {
      "id": "346321",
      "postDate": "06/21/2018 13:48:11",
      "content": "<p>Sound like just overfit. There are some parameters controlling overfitting.\n<a href=\"https://xgboost.readthedocs.io/en/latest/how_to/param_tuning.html\">https://xgboost.readthedocs.io/en/latest/how_to/param_tuning.html</a></p>",
      "rawMarkdown": "Sound like just overfit. There are some parameters controlling overfitting.\nhttps://xgboost.readthedocs.io/en/latest/how_to/param_tuning.html",
      "votes": null
    },
    {
      "id": "346455",
      "postDate": "06/21/2018 18:10:23",
      "content": "<p>Are you sure you did not mix the item_id in your submission? I had the same problem twice due to my own bug in pandas dataframe index usage. My predictions were not matching the correct item_id on submission file generation.</p>",
      "rawMarkdown": "Are you sure you did not mix the item_id in your submission? I had the same problem twice due to my own bug in pandas dataframe index usage. My predictions were not matching the correct item_id on submission file generation.",
      "votes": null
    },
    {
      "id": "347369",
      "postDate": "06/24/2018 04:28:32",
      "content": "<p>@Oleg, what is the validation score of your lgbm model that got 0.2200 on LB. My guess is your xgb may be over-fitting. Pay attention to your parameters and make sure the gap between train and validation is in a reasonable range. My models at the beginning of this contest were over-fitting really badly until I fixed it, it was had to do FE. </p>\n\n<p>Just as an example, there is a public kernel with a validation score of about 0.0005 better than my best lgbm model yet scores way worse (-0.001) on the LB than mine. Hopefully I am not over-fitting as well. So \"over-fitting\" is  something to keep in mind.</p>",
      "rawMarkdown": "Oleg, what is the validation score of your lgbm model that got 0.2200 on LB. My guess is your xgb may be over-fitting. Pay attention to your parameters and make sure the gap between train and validation is in a reasonable range. My models at the beginning of this contest were over-fitting really badly until I fixed it, it was had to do FE. \n\nJust as an example, there is a public kernel with a validation score of about 0.0005 better than my best lgbm model yet scores way worse (-0.001) on the LB than mine. Hopefully I am not over-fitting as well. So \"over-fitting\" is  something to keep in mind.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 346300,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "06/21/2018 12:41:56",
      "content": "<p>When I read the title, I thought you got big success with Xgboost.... Anyway, good luck.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 346321,
      "author_name": "fujihiro",
      "author_url": "",
      "post_date": "06/21/2018 13:48:11",
      "content": "<p>Sound like just overfit. There are some parameters controlling overfitting.\n<a href=\"https://xgboost.readthedocs.io/en/latest/how_to/param_tuning.html\">https://xgboost.readthedocs.io/en/latest/how_to/param_tuning.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 346455,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "06/21/2018 18:10:23",
      "content": "<p>Are you sure you did not mix the item_id in your submission? I had the same problem twice due to my own bug in pandas dataframe index usage. My predictions were not matching the correct item_id on submission file generation.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 347369,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "06/24/2018 04:28:32",
      "content": "<p>@Oleg, what is the validation score of your lgbm model that got 0.2200 on LB. My guess is your xgb may be over-fitting. Pay attention to your parameters and make sure the gap between train and validation is in a reasonable range. My models at the beginning of this contest were over-fitting really badly until I fixed it, it was had to do FE. </p>\n\n<p>Just as an example, there is a public kernel with a validation score of about 0.0005 better than my best lgbm model yet scores way worse (-0.001) on the LB than mine. Hopefully I am not over-fitting as well. So \"over-fitting\" is  something to keep in mind.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "346142": "I ran xgboost with the same feature pack as my lgbm (2200 LB) and got 2160 CV but 2300 LB. What might be the reason of such huge difference between LB and CV? The correlation between xgb and lgbm submission is about 0.97 but 100 points diff in LB. Has anybody solved this issue?",
    "346300": "When I read the title, I thought you got big success with Xgboost.... Anyway, good luck.",
    "346321": "Sound like just overfit. There are some parameters controlling overfitting.\nhttps://xgboost.readthedocs.io/en/latest/how_to/param_tuning.html",
    "346455": "Are you sure you did not mix the item_id in your submission? I had the same problem twice due to my own bug in pandas dataframe index usage. My predictions were not matching the correct item_id on submission file generation.",
    "347369": "Oleg, what is the validation score of your lgbm model that got 0.2200 on LB. My guess is your xgb may be over-fitting. Pay attention to your parameters and make sure the gap between train and validation is in a reasonable range. My models at the beginning of this contest were over-fitting really badly until I fixed it, it was had to do FE. \n\nJust as an example, there is a public kernel with a validation score of about 0.0005 better than my best lgbm model yet scores way worse (-0.001) on the LB than mine. Hopefully I am not over-fitting as well. So \"over-fitting\" is  something to keep in mind."
  },
  "source": "meta"
}