{
  "id": 57250,
  "title": "Boosting overfits by a great amount",
  "url": "/competitions/avito-demand-prediction/discussion/57250",
  "author_name": "",
  "post_date": "2018-05-21T16:01:08.602003500Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hello</p>\n\n<p>Does anybody know why boosting algorithms overfit by that much?</p>\n\n<p>I am extremely new to Python, but I had to try Lightgbm there since it throws too many errors in R. \nSo my cross validation is around 0.225, my test_validation is around 0.226, but when I submit I get 0.29, and I am talking about fairly simple models with num_bust_rounds =1000, learning rate = 0.01. Also, I did not have to work on my deal_probability in my submission for handling values outside[0 1].</p>\n\n<p>The same thing happened in R with xgboost, overfitting by 0.2-0.3. </p>\n\n<p>I know I have no experience in Python, but this it still very odd.</p>\n\n<p>Thank you</p>",
  "messages": [
    {
      "id": "331633",
      "postDate": "05/21/2018 16:01:08",
      "content": "<p>Hello</p>\n\n<p>Does anybody know why boosting algorithms overfit by that much?</p>\n\n<p>I am extremely new to Python, but I had to try Lightgbm there since it throws too many errors in R. \nSo my cross validation is around 0.225, my test_validation is around 0.226, but when I submit I get 0.29, and I am talking about fairly simple models with num_bust_rounds =1000, learning rate = 0.01. Also, I did not have to work on my deal_probability in my submission for handling values outside[0 1].</p>\n\n<p>The same thing happened in R with xgboost, overfitting by 0.2-0.3. </p>\n\n<p>I know I have no experience in Python, but this it still very odd.</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hello\n\nDoes anybody know why boosting algorithms overfit by that much?\n\nI am extremely new to Python, but I had to try Lightgbm there since it throws too many errors in R. \nSo my cross validation is around 0.225, my test_validation is around 0.226, but when I submit I get 0.29, and I am talking about fairly simple models with num_bust_rounds =1000, learning rate = 0.01. Also, I did not have to work on my deal_probability in my submission for handling values outside[0 1].\n\nThe same thing happened in R with xgboost, overfitting by 0.2-0.3. \n\nI know I have no experience in Python, but this it still very odd.\n\nThank you",
      "votes": null
    },
    {
      "id": "331640",
      "postDate": "05/21/2018 16:09:50",
      "content": "<p>Based on that information alone, I would guess the problem is leakage in your feature engineering procedure: with my lgbm models, I am also observing a cv - lb gap, but it's much smaller (~ 0.005). </p>",
      "rawMarkdown": "Based on that information alone, I would guess the problem is leakage in your feature engineering procedure: with my lgbm models, I am also observing a cv - lb gap, but it's much smaller (~ 0.005).",
      "votes": null
    },
    {
      "id": "331705",
      "postDate": "05/21/2018 18:10:49",
      "content": "<p>Hint - look at how many potential parameters there are compared to rows of data ;)</p>",
      "rawMarkdown": "Hint - look at how many potential parameters there are compared to rows of data ;)",
      "votes": null
    },
    {
      "id": "331760",
      "postDate": "05/21/2018 20:21:37",
      "content": "<p>Thank you very much for your answer, Konrad. I have already pursued a complete recheck of all my feature engineering</p>",
      "rawMarkdown": "Thank you very much for your answer, Konrad. I have already pursued a complete recheck of all my feature engineering",
      "votes": null
    },
    {
      "id": "332915",
      "postDate": "05/24/2018 03:18:01",
      "content": "<p>how can you achieve that? do you using N folder cv?</p>",
      "rawMarkdown": "how can you achieve that? do you using N folder cv?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 331640,
      "author_name": "konradb",
      "author_url": "",
      "post_date": "05/21/2018 16:09:50",
      "content": "<p>Based on that information alone, I would guess the problem is leakage in your feature engineering procedure: with my lgbm models, I am also observing a cv - lb gap, but it's much smaller (~ 0.005). </p>",
      "votes": null,
      "replies": [
        {
          "id": 332915,
          "author_name": "liuhdsgoal",
          "author_url": "",
          "post_date": "05/24/2018 03:18:01",
          "content": "<p>how can you achieve that? do you using N folder cv?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331705,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "05/21/2018 18:10:49",
      "content": "<p>Hint - look at how many potential parameters there are compared to rows of data ;)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 331760,
      "author_name": "titamarius",
      "author_url": "",
      "post_date": "05/21/2018 20:21:37",
      "content": "<p>Thank you very much for your answer, Konrad. I have already pursued a complete recheck of all my feature engineering</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "331633": "Hello\n\nDoes anybody know why boosting algorithms overfit by that much?\n\nI am extremely new to Python, but I had to try Lightgbm there since it throws too many errors in R. \nSo my cross validation is around 0.225, my test_validation is around 0.226, but when I submit I get 0.29, and I am talking about fairly simple models with num_bust_rounds =1000, learning rate = 0.01. Also, I did not have to work on my deal_probability in my submission for handling values outside[0 1].\n\nThe same thing happened in R with xgboost, overfitting by 0.2-0.3. \n\nI know I have no experience in Python, but this it still very odd.\n\nThank you",
    "331640": "Based on that information alone, I would guess the problem is leakage in your feature engineering procedure: with my lgbm models, I am also observing a cv - lb gap, but it's much smaller (~ 0.005).",
    "331705": "Hint - look at how many potential parameters there are compared to rows of data ;)",
    "331760": "Thank you very much for your answer, Konrad. I have already pursued a complete recheck of all my feature engineering",
    "332915": "how can you achieve that? do you using N folder cv?"
  },
  "source": "meta"
}