{
  "id": 208203,
  "title": "How do you determine features that cause overfitting in LGBM ?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/208203",
  "author_name": "",
  "post_date": "2021-01-02T11:42:40.594471400Z",
  "votes": 10,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The other day, <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> brought up a very important subject in this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206434\" target=\"_blank\">thread</a> in which he emphasized the importance of reducing features number to to improve score, stating that he used <a href=\"https://shap.readthedocs.io/en/latest/\" target=\"_blank\">Shap</a> library to do that.</p>\n<p>For a competition with such a big dataset, every feature would matter given the space it would take in ram (Specially with only 16GB ram). Also, a noisy feature might lead to overfitting and limit the generalization of the model. I noticed that many features that would make sense intuitively, wouldn't necessarily result in any significant gain (such as lecture features). So assessing how prone features are to overfitting is crucial for success here.</p>\n<p>This leads me to the creation of this thread where we can share tips on determining which features are the ones that cause overfitting.</p>",
  "messages": [
    {
      "id": "1135626",
      "postDate": "01/02/2021 11:42:40",
      "content": "<p>The other day, <a href=\"https://www.kaggle.com/bowaka\" target=\"_blank\">@bowaka</a> brought up a very important subject in this <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206434\" target=\"_blank\">thread</a> in which he emphasized the importance of reducing features number to to improve score, stating that he used <a href=\"https://shap.readthedocs.io/en/latest/\" target=\"_blank\">Shap</a> library to do that.</p>\n<p>For a competition with such a big dataset, every feature would matter given the space it would take in ram (Specially with only 16GB ram). Also, a noisy feature might lead to overfitting and limit the generalization of the model. I noticed that many features that would make sense intuitively, wouldn't necessarily result in any significant gain (such as lecture features). So assessing how prone features are to overfitting is crucial for success here.</p>\n<p>This leads me to the creation of this thread where we can share tips on determining which features are the ones that cause overfitting.</p>",
      "rawMarkdown": "The other day, @bowaka brought up a very important subject in this [thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206434) in which he emphasized the importance of reducing features number to to improve score, stating that he used [Shap](https://shap.readthedocs.io/en/latest/) library to do that.\n\nFor a competition with such a big dataset, every feature would matter given the space it would take in ram (Specially with only 16GB ram). Also, a noisy feature might lead to overfitting and limit the generalization of the model. I noticed that many features that would make sense intuitively, wouldn't necessarily result in any significant gain (such as lecture features). So assessing how prone features are to overfitting is crucial for success here.\n\nThis leads me to the creation of this thread where we can share tips on determining which features are the ones that cause overfitting.",
      "votes": null
    },
    {
      "id": "1136575",
      "postDate": "01/03/2021 08:05:32",
      "content": "<p>I have created <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207307\" target=\"_blank\">a thread earlier</a> on how to use feature importance plots to remove over-fitting features. Here is my current understanding:</p>\n<ol>\n<li>Categorical features: This is one major source of over fitting. The number of possible splits possible is <code>2 ^ categories</code>. For features which have high number of possible categories (ex: user_id, content_id, tag_encoding), its best not to use them or use some kind of target encoding for the same.</li>\n<li>Look at the feature importance plots. It is 2 fold (splits vs gain). You could try to remove the features which are not coming in the feature importance or have very low gain. </li>\n</ol>\n<p>These points are basic but they could help remove some over-fitting features.</p>",
      "rawMarkdown": "I have created [a thread earlier](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207307) on how to use feature importance plots to remove over-fitting features. Here is my current understanding:\n1. Categorical features: This is one major source of over fitting. The number of possible splits possible is `2 ^ categories`. For features which have high number of possible categories (ex: user_id, content_id, tag_encoding), its best not to use them or use some kind of target encoding for the same.\n2. Look at the feature importance plots. It is 2 fold (splits vs gain). You could try to remove the features which are not coming in the feature importance or have very low gain. \n\nThese points are basic but they could help remove some over-fitting features.",
      "votes": null
    },
    {
      "id": "1137917",
      "postDate": "01/04/2021 09:30:15",
      "content": "<p>Personnally right now I am still sticking with my approach of train/test on 1M data, test on 200 000, check shap values, and remove very low features… It brings me down to 34 features. </p>\n<p>Regarding the extra features I am including in my model, I think the conditionnal_probabilities that I was discussing <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148\" target=\"_blank\">here </a> is giving a nice boost to my score. </p>\n<p>Also, I am using the elo score shown in <a href=\"https://www.kaggle.com/stevemju/riiid-simple-elo-rating\" target=\"_blank\">that great notebook</a> that was of great help and deserve more visibility!</p>\n<p>I start nethertheless to lack of ideads now, I don't think i'll be able to push my model way further in the last few days ! </p>",
      "rawMarkdown": "Personnally right now I am still sticking with my approach of train/test on 1M data, test on 200 000, check shap values, and remove very low features... It brings me down to 34 features. \n\nRegarding the extra features I am including in my model, I think the conditionnal_probabilities that I was discussing [here ](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148) is giving a nice boost to my score. \n\nAlso, I am using the elo score shown in [that great notebook](https://www.kaggle.com/stevemju/riiid-simple-elo-rating) that was of great help and deserve more visibility!\n\nI start nethertheless to lack of ideads now, I don't think i'll be able to push my model way further in the last few days !",
      "votes": null
    },
    {
      "id": "1137974",
      "postDate": "01/04/2021 10:21:13",
      "content": "<p>How much boost do you get from conditional probabilities?</p>",
      "rawMarkdown": "How much boost do you get from conditional probabilities?",
      "votes": null
    },
    {
      "id": "1137983",
      "postDate": "01/04/2021 10:28:30",
      "content": "<p>You can make a lightweight version of your model or subsample the data and then run LOFO: <a href=\"https://github.com/aerdem4/lofo-importance\" target=\"_blank\">https://github.com/aerdem4/lofo-importance</a> </p>",
      "rawMarkdown": "You can make a lightweight version of your model or subsample the data and then run LOFO: https://github.com/aerdem4/lofo-importance",
      "votes": null
    },
    {
      "id": "1138229",
      "postDate": "01/04/2021 14:03:53",
      "content": "<p>elo method is great, half of my features (about 20) on based on elo scores of questions and users. My simple elo model scored 0.763 lb.</p>",
      "rawMarkdown": "elo method is great, half of my features (about 20) on based on elo scores of questions and users. My simple elo model scored 0.763 lb.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1136575,
      "author_name": "manikanthr5",
      "author_url": "",
      "post_date": "01/03/2021 08:05:32",
      "content": "<p>I have created <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207307\" target=\"_blank\">a thread earlier</a> on how to use feature importance plots to remove over-fitting features. Here is my current understanding:</p>\n<ol>\n<li>Categorical features: This is one major source of over fitting. The number of possible splits possible is <code>2 ^ categories</code>. For features which have high number of possible categories (ex: user_id, content_id, tag_encoding), its best not to use them or use some kind of target encoding for the same.</li>\n<li>Look at the feature importance plots. It is 2 fold (splits vs gain). You could try to remove the features which are not coming in the feature importance or have very low gain. </li>\n</ol>\n<p>These points are basic but they could help remove some over-fitting features.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1137917,
      "author_name": "bowaka",
      "author_url": "",
      "post_date": "01/04/2021 09:30:15",
      "content": "<p>Personnally right now I am still sticking with my approach of train/test on 1M data, test on 200 000, check shap values, and remove very low features… It brings me down to 34 features. </p>\n<p>Regarding the extra features I am including in my model, I think the conditionnal_probabilities that I was discussing <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148\" target=\"_blank\">here </a> is giving a nice boost to my score. </p>\n<p>Also, I am using the elo score shown in <a href=\"https://www.kaggle.com/stevemju/riiid-simple-elo-rating\" target=\"_blank\">that great notebook</a> that was of great help and deserve more visibility!</p>\n<p>I start nethertheless to lack of ideads now, I don't think i'll be able to push my model way further in the last few days ! </p>",
      "votes": null,
      "replies": [
        {
          "id": 1137974,
          "author_name": "nicohrubec",
          "author_url": "",
          "post_date": "01/04/2021 10:21:13",
          "content": "<p>How much boost do you get from conditional probabilities?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1138229,
          "author_name": "lrtmonkey",
          "author_url": "",
          "post_date": "01/04/2021 14:03:53",
          "content": "<p>elo method is great, half of my features (about 20) on based on elo scores of questions and users. My simple elo model scored 0.763 lb.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1137983,
      "author_name": "aerdem4",
      "author_url": "",
      "post_date": "01/04/2021 10:28:30",
      "content": "<p>You can make a lightweight version of your model or subsample the data and then run LOFO: <a href=\"https://github.com/aerdem4/lofo-importance\" target=\"_blank\">https://github.com/aerdem4/lofo-importance</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1135626": "The other day, @bowaka brought up a very important subject in this [thread](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/206434) in which he emphasized the importance of reducing features number to to improve score, stating that he used [Shap](https://shap.readthedocs.io/en/latest/) library to do that.\n\nFor a competition with such a big dataset, every feature would matter given the space it would take in ram (Specially with only 16GB ram). Also, a noisy feature might lead to overfitting and limit the generalization of the model. I noticed that many features that would make sense intuitively, wouldn't necessarily result in any significant gain (such as lecture features). So assessing how prone features are to overfitting is crucial for success here.\n\nThis leads me to the creation of this thread where we can share tips on determining which features are the ones that cause overfitting.",
    "1136575": "I have created [a thread earlier](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207307) on how to use feature importance plots to remove over-fitting features. Here is my current understanding:\n1. Categorical features: This is one major source of over fitting. The number of possible splits possible is `2 ^ categories`. For features which have high number of possible categories (ex: user_id, content_id, tag_encoding), its best not to use them or use some kind of target encoding for the same.\n2. Look at the feature importance plots. It is 2 fold (splits vs gain). You could try to remove the features which are not coming in the feature importance or have very low gain. \n\nThese points are basic but they could help remove some over-fitting features.",
    "1137917": "Personnally right now I am still sticking with my approach of train/test on 1M data, test on 200 000, check shap values, and remove very low features... It brings me down to 34 features. \n\nRegarding the extra features I am including in my model, I think the conditionnal_probabilities that I was discussing [here ](https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/207148) is giving a nice boost to my score. \n\nAlso, I am using the elo score shown in [that great notebook](https://www.kaggle.com/stevemju/riiid-simple-elo-rating) that was of great help and deserve more visibility!\n\nI start nethertheless to lack of ideads now, I don't think i'll be able to push my model way further in the last few days !",
    "1137974": "How much boost do you get from conditional probabilities?",
    "1137983": "You can make a lightweight version of your model or subsample the data and then run LOFO: https://github.com/aerdem4/lofo-importance",
    "1138229": "elo method is great, half of my features (about 20) on based on elo scores of questions and users. My simple elo model scored 0.763 lb."
  },
  "source": "meta"
}