{
  "id": 56915,
  "title": "LightGMB and categorical features - am I overfitting?",
  "url": "/competitions/avito-demand-prediction/discussion/56915",
  "author_name": "",
  "post_date": "2018-05-16T16:42:56.074965300Z",
  "votes": 1,
  "comment_count": 11,
  "views": 0,
  "content": "<p>LightGBM supports the use of categorical features that are first encoded into integer features. From  the documentation: </p>\n\n<blockquote>\n  <p>LightGBM can offer a good accuracy when using native categorical features. Not like simply one-hot coding, LightGBM can find the optimal split of categorical features. Such an optimal split can provide the much better accuracy than one-hot coding solution.</p>\n</blockquote>\n\n<p>I started in this competition by extracting some simple features, and running a Random Forrest model. I'm setting aside 15% of the training data at random as a validation set, and I got a validation score of .27, and a public leaderboard score of .28. </p>\n\n<p>I then tried a LightGBM classifier. This time, I passed categorical features - including some of the more sparse ones like <code>category</code> and <code>city</code>, to the classifier as categorical features. </p>\n\n<p>On my validation set, I got something like .256 RMSE, but when I publish to the leaderboard I get .30 - much much worse. </p>\n\n<p>As the documentation about how categorical features is a bit opaque, and so I'm not sure what it does with integer-encoded categorical features. If these features are sparse, is there a danger of overfitting? I suspect so, since there are parameters <code>max_cat_threshold</code>, <code>cat_smooth</code> etc which deal with the partitioning of categorical features. </p>\n\n<p>But then, why would this overfitting not become evident in my validation set, but only in the pubic leaderboard? Any ideas? </p>",
  "messages": [
    {
      "id": "329528",
      "postDate": "05/16/2018 16:42:56",
      "content": "<p>LightGBM supports the use of categorical features that are first encoded into integer features. From  the documentation: </p>\n\n<blockquote>\n  <p>LightGBM can offer a good accuracy when using native categorical features. Not like simply one-hot coding, LightGBM can find the optimal split of categorical features. Such an optimal split can provide the much better accuracy than one-hot coding solution.</p>\n</blockquote>\n\n<p>I started in this competition by extracting some simple features, and running a Random Forrest model. I'm setting aside 15% of the training data at random as a validation set, and I got a validation score of .27, and a public leaderboard score of .28. </p>\n\n<p>I then tried a LightGBM classifier. This time, I passed categorical features - including some of the more sparse ones like <code>category</code> and <code>city</code>, to the classifier as categorical features. </p>\n\n<p>On my validation set, I got something like .256 RMSE, but when I publish to the leaderboard I get .30 - much much worse. </p>\n\n<p>As the documentation about how categorical features is a bit opaque, and so I'm not sure what it does with integer-encoded categorical features. If these features are sparse, is there a danger of overfitting? I suspect so, since there are parameters <code>max_cat_threshold</code>, <code>cat_smooth</code> etc which deal with the partitioning of categorical features. </p>\n\n<p>But then, why would this overfitting not become evident in my validation set, but only in the pubic leaderboard? Any ideas? </p>",
      "rawMarkdown": "LightGBM supports the use of categorical features that are first encoded into integer features. From  the documentation: \n\n&gt; LightGBM can offer a good accuracy when using native categorical features. Not like simply one-hot coding, LightGBM can find the optimal split of categorical features. Such an optimal split can provide the much better accuracy than one-hot coding solution.\n\nI started in this competition by extracting some simple features, and running a Random Forrest model. I'm setting aside 15% of the training data at random as a validation set, and I got a validation score of .27, and a public leaderboard score of .28. \n\nI then tried a LightGBM classifier. This time, I passed categorical features - including some of the more sparse ones like `category` and `city`, to the classifier as categorical features. \n\nOn my validation set, I got something like .256 RMSE, but when I publish to the leaderboard I get .30 - much much worse. \n\nAs the documentation about how categorical features is a bit opaque, and so I'm not sure what it does with integer-encoded categorical features. If these features are sparse, is there a danger of overfitting? I suspect so, since there are parameters `max_cat_threshold`, `cat_smooth` etc which deal with the partitioning of categorical features. \n\nBut then, why would this overfitting not become evident in my validation set, but only in the pubic leaderboard? Any ideas?",
      "votes": null
    },
    {
      "id": "329534",
      "postDate": "05/16/2018 16:59:38",
      "content": "<p>LightGBM is prone to overfitting in general. It may be some of your other parameters rather than how you handle categorical features.</p>",
      "rawMarkdown": "LightGBM is prone to overfitting in general. It may be some of your other parameters rather than how you handle categorical features.",
      "votes": null
    },
    {
      "id": "329536",
      "postDate": "05/16/2018 17:01:53",
      "content": "<p>generally this happens due to uneven validation dataset that might differ from train data to a large extent, so I would suggest you to use K-Fold validation (try 5 fold).</p>",
      "rawMarkdown": "generally this happens due to uneven validation dataset that might differ from train data to a large extent, so I would suggest you to use K-Fold validation (try 5 fold).",
      "votes": null
    },
    {
      "id": "329546",
      "postDate": "05/16/2018 17:32:14",
      "content": "<p>Yes, this is what I suspected as well. But since the Random Forrest seems to generalize as expected - and I used the same train/validation splits for the Random Forrest and LGBM model, I don't think it's an issue of a misrepresentative validation set.</p>",
      "rawMarkdown": "Yes, this is what I suspected as well. But since the Random Forrest seems to generalize as expected - and I used the same train/validation splits for the Random Forrest and LGBM model, I don't think it's an issue of a misrepresentative validation set.",
      "votes": null
    },
    {
      "id": "329548",
      "postDate": "05/16/2018 17:34:16",
      "content": "<p>Yes, that's true. But if my model is overfitting, I\"m not sure why I wouldn't see it in my validation RMSE, but in the public LB RMSE. </p>",
      "rawMarkdown": "Yes, that's true. But if my model is overfitting, I\"m not sure why I wouldn't see it in my validation RMSE, but in the public LB RMSE.",
      "votes": null
    },
    {
      "id": "329550",
      "postDate": "05/16/2018 17:36:23",
      "content": "<p>Another piece of information - I dropped from my training set all rows that had <code>NA</code> in the column <code>description</code> . I did this because all the <code>description</code> values are not <code>NA</code> in the test set, and so I didn't want my model to learn patterns regarding the missingness of <code>description</code> - as it would not generalize to the test set. </p>\n\n<p>Do you think this is causing my public LB score to be so poor? Perhaps removing the rows with <code>NA</code> in <code>description</code> makes the training set not representative of the test set. </p>",
      "rawMarkdown": "Another piece of information - I dropped from my training set all rows that had `NA` in the column `description` . I did this because all the `description` values are not `NA` in the test set, and so I didn't want my model to learn patterns regarding the missingness of `description` - as it would not generalize to the test set. \n\nDo you think this is causing my public LB score to be so poor? Perhaps removing the rows with `NA` in `description` makes the training set not representative of the test set.",
      "votes": null
    },
    {
      "id": "329554",
      "postDate": "05/16/2018 17:54:29",
      "content": "<p>How consistent is your train-val rmse? If there's a large gap there it's a sign of overfitting.</p>",
      "rawMarkdown": "How consistent is your train-val rmse? If there's a large gap there it's a sign of overfitting.",
      "votes": null
    },
    {
      "id": "329575",
      "postDate": "05/16/2018 18:40:30",
      "content": "<p>I'm using an early stopping scheme - stopping when the validation RMSE has not improved for 10 consecutive boosting rounds. My training RMSE is .24 while the validation is .26. So yes – I'm overfitting my training data – but this is usually the case. I didn't see this as a red flag. </p>\n\n<p>Have you had any experience with passing categorical variables with many levels to LGBM? Did things work fine? I was inspired by <a href=\"https://www.kaggle.com/tunguz/bow-meta-text-and-dense-features-lb-0-2241?scriptVersionId=3616587\">this notebook</a>, which passes <code>image_top_1</code> and even <code>user_id</code> as categorical features.. I find this unintuitive. </p>",
      "rawMarkdown": "I'm using an early stopping scheme - stopping when the validation RMSE has not improved for 10 consecutive boosting rounds. My training RMSE is .24 while the validation is .26. So yes – I'm overfitting my training data – but this is usually the case. I didn't see this as a red flag. \n\nHave you had any experience with passing categorical variables with many levels to LGBM? Did things work fine? I was inspired by [this notebook](https://www.kaggle.com/tunguz/bow-meta-text-and-dense-features-lb-0-2241?scriptVersionId=3616587), which passes `image_top_1` and even `user_id` as categorical features.. I find this unintuitive.",
      "votes": null
    },
    {
      "id": "329600",
      "postDate": "05/16/2018 19:35:20",
      "content": "<p>Did you do this before creating your validation set? The large gap between your validation and leaderboard score suggests that there may be a very substantive difference between the dataset you're working with locally and the test set. </p>\n\n<p>I would be very careful about this. Missing description is likely correlated with other relevant facets of the data, so dropping those rows may make your local data less similar to test rather than more similar.</p>",
      "rawMarkdown": "Did you do this before creating your validation set? The large gap between your validation and leaderboard score suggests that there may be a very substantive difference between the dataset you're working with locally and the test set. \n\nI would be very careful about this. Missing description is likely correlated with other relevant facets of the data, so dropping those rows may make your local data less similar to test rather than more similar.",
      "votes": null
    },
    {
      "id": "329607",
      "postDate": "05/16/2018 19:55:07",
      "content": "<p>Yeah I dropped them before making my validation set - but I think you're right. I'll try re-training without omitting any rows, and let you know how it goes!</p>",
      "rawMarkdown": "Yeah I dropped them before making my validation set - but I think you're right. I'll try re-training without omitting any rows, and let you know how it goes!",
      "votes": null
    },
    {
      "id": "329630",
      "postDate": "05/16/2018 20:33:29",
      "content": "<p>Yes there could be problems with high cardinality, but you can vectorize it like the notebook. This usually isn't a problem with the way LightGBM handles categorical features.</p>",
      "rawMarkdown": "Yes there could be problems with high cardinality, but you can vectorize it like the notebook. This usually isn't a problem with the way LightGBM handles categorical features.",
      "votes": null
    },
    {
      "id": "331553",
      "postDate": "05/21/2018 13:32:26",
      "content": "<p>It turns out the problem was not my use of categorical features, or dropping rows without comments! I had simply scrambled the <code>item_id</code> column in the test set, and so my leaderboard submission were scrambled... smh. </p>\n\n<p>Thanks for the help to those who commented. And if someone has a good explanation of how LGBM works with categorical features, I'd love to hear it!</p>",
      "rawMarkdown": "It turns out the problem was not my use of categorical features, or dropping rows without comments! I had simply scrambled the `item_id` column in the test set, and so my leaderboard submission were scrambled... smh. \n\nThanks for the help to those who commented. And if someone has a good explanation of how LGBM works with categorical features, I'd love to hear it!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 329534,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "05/16/2018 16:59:38",
      "content": "<p>LightGBM is prone to overfitting in general. It may be some of your other parameters rather than how you handle categorical features.</p>",
      "votes": null,
      "replies": [
        {
          "id": 329548,
          "author_name": "timib1203",
          "author_url": "",
          "post_date": "05/16/2018 17:34:16",
          "content": "<p>Yes, that's true. But if my model is overfitting, I\"m not sure why I wouldn't see it in my validation RMSE, but in the public LB RMSE. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329554,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "05/16/2018 17:54:29",
          "content": "<p>How consistent is your train-val rmse? If there's a large gap there it's a sign of overfitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329575,
          "author_name": "timib1203",
          "author_url": "",
          "post_date": "05/16/2018 18:40:30",
          "content": "<p>I'm using an early stopping scheme - stopping when the validation RMSE has not improved for 10 consecutive boosting rounds. My training RMSE is .24 while the validation is .26. So yes – I'm overfitting my training data – but this is usually the case. I didn't see this as a red flag. </p>\n\n<p>Have you had any experience with passing categorical variables with many levels to LGBM? Did things work fine? I was inspired by <a href=\"https://www.kaggle.com/tunguz/bow-meta-text-and-dense-features-lb-0-2241?scriptVersionId=3616587\">this notebook</a>, which passes <code>image_top_1</code> and even <code>user_id</code> as categorical features.. I find this unintuitive. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329630,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "05/16/2018 20:33:29",
          "content": "<p>Yes there could be problems with high cardinality, but you can vectorize it like the notebook. This usually isn't a problem with the way LightGBM handles categorical features.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 329536,
      "author_name": "mk9440",
      "author_url": "",
      "post_date": "05/16/2018 17:01:53",
      "content": "<p>generally this happens due to uneven validation dataset that might differ from train data to a large extent, so I would suggest you to use K-Fold validation (try 5 fold).</p>",
      "votes": null,
      "replies": [
        {
          "id": 329546,
          "author_name": "timib1203",
          "author_url": "",
          "post_date": "05/16/2018 17:32:14",
          "content": "<p>Yes, this is what I suspected as well. But since the Random Forrest seems to generalize as expected - and I used the same train/validation splits for the Random Forrest and LGBM model, I don't think it's an issue of a misrepresentative validation set.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 329550,
      "author_name": "timib1203",
      "author_url": "",
      "post_date": "05/16/2018 17:36:23",
      "content": "<p>Another piece of information - I dropped from my training set all rows that had <code>NA</code> in the column <code>description</code> . I did this because all the <code>description</code> values are not <code>NA</code> in the test set, and so I didn't want my model to learn patterns regarding the missingness of <code>description</code> - as it would not generalize to the test set. </p>\n\n<p>Do you think this is causing my public LB score to be so poor? Perhaps removing the rows with <code>NA</code> in <code>description</code> makes the training set not representative of the test set. </p>",
      "votes": null,
      "replies": [
        {
          "id": 329600,
          "author_name": "aquatic",
          "author_url": "",
          "post_date": "05/16/2018 19:35:20",
          "content": "<p>Did you do this before creating your validation set? The large gap between your validation and leaderboard score suggests that there may be a very substantive difference between the dataset you're working with locally and the test set. </p>\n\n<p>I would be very careful about this. Missing description is likely correlated with other relevant facets of the data, so dropping those rows may make your local data less similar to test rather than more similar.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 329607,
          "author_name": "timib1203",
          "author_url": "",
          "post_date": "05/16/2018 19:55:07",
          "content": "<p>Yeah I dropped them before making my validation set - but I think you're right. I'll try re-training without omitting any rows, and let you know how it goes!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 331553,
      "author_name": "timib1203",
      "author_url": "",
      "post_date": "05/21/2018 13:32:26",
      "content": "<p>It turns out the problem was not my use of categorical features, or dropping rows without comments! I had simply scrambled the <code>item_id</code> column in the test set, and so my leaderboard submission were scrambled... smh. </p>\n\n<p>Thanks for the help to those who commented. And if someone has a good explanation of how LGBM works with categorical features, I'd love to hear it!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "329528": "LightGBM supports the use of categorical features that are first encoded into integer features. From  the documentation: \n\n&gt; LightGBM can offer a good accuracy when using native categorical features. Not like simply one-hot coding, LightGBM can find the optimal split of categorical features. Such an optimal split can provide the much better accuracy than one-hot coding solution.\n\nI started in this competition by extracting some simple features, and running a Random Forrest model. I'm setting aside 15% of the training data at random as a validation set, and I got a validation score of .27, and a public leaderboard score of .28. \n\nI then tried a LightGBM classifier. This time, I passed categorical features - including some of the more sparse ones like `category` and `city`, to the classifier as categorical features. \n\nOn my validation set, I got something like .256 RMSE, but when I publish to the leaderboard I get .30 - much much worse. \n\nAs the documentation about how categorical features is a bit opaque, and so I'm not sure what it does with integer-encoded categorical features. If these features are sparse, is there a danger of overfitting? I suspect so, since there are parameters `max_cat_threshold`, `cat_smooth` etc which deal with the partitioning of categorical features. \n\nBut then, why would this overfitting not become evident in my validation set, but only in the pubic leaderboard? Any ideas?",
    "329534": "LightGBM is prone to overfitting in general. It may be some of your other parameters rather than how you handle categorical features.",
    "329536": "generally this happens due to uneven validation dataset that might differ from train data to a large extent, so I would suggest you to use K-Fold validation (try 5 fold).",
    "329546": "Yes, this is what I suspected as well. But since the Random Forrest seems to generalize as expected - and I used the same train/validation splits for the Random Forrest and LGBM model, I don't think it's an issue of a misrepresentative validation set.",
    "329548": "Yes, that's true. But if my model is overfitting, I\"m not sure why I wouldn't see it in my validation RMSE, but in the public LB RMSE.",
    "329550": "Another piece of information - I dropped from my training set all rows that had `NA` in the column `description` . I did this because all the `description` values are not `NA` in the test set, and so I didn't want my model to learn patterns regarding the missingness of `description` - as it would not generalize to the test set. \n\nDo you think this is causing my public LB score to be so poor? Perhaps removing the rows with `NA` in `description` makes the training set not representative of the test set.",
    "329554": "How consistent is your train-val rmse? If there's a large gap there it's a sign of overfitting.",
    "329575": "I'm using an early stopping scheme - stopping when the validation RMSE has not improved for 10 consecutive boosting rounds. My training RMSE is .24 while the validation is .26. So yes – I'm overfitting my training data – but this is usually the case. I didn't see this as a red flag. \n\nHave you had any experience with passing categorical variables with many levels to LGBM? Did things work fine? I was inspired by [this notebook](https://www.kaggle.com/tunguz/bow-meta-text-and-dense-features-lb-0-2241?scriptVersionId=3616587), which passes `image_top_1` and even `user_id` as categorical features.. I find this unintuitive.",
    "329600": "Did you do this before creating your validation set? The large gap between your validation and leaderboard score suggests that there may be a very substantive difference between the dataset you're working with locally and the test set. \n\nI would be very careful about this. Missing description is likely correlated with other relevant facets of the data, so dropping those rows may make your local data less similar to test rather than more similar.",
    "329607": "Yeah I dropped them before making my validation set - but I think you're right. I'll try re-training without omitting any rows, and let you know how it goes!",
    "329630": "Yes there could be problems with high cardinality, but you can vectorize it like the notebook. This usually isn't a problem with the way LightGBM handles categorical features.",
    "331553": "It turns out the problem was not my use of categorical features, or dropping rows without comments! I had simply scrambled the `item_id` column in the test set, and so my leaderboard submission were scrambled... smh. \n\nThanks for the help to those who commented. And if someone has a good explanation of how LGBM works with categorical features, I'd love to hear it!"
  },
  "source": "meta"
}