{
  "id": 388201,
  "title": "Easy way to increase CV and decrease training time!",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388201",
  "author_name": "",
  "post_date": "2023-02-16T12:48:59.839581900Z",
  "votes": 16,
  "comment_count": 18,
  "views": 0,
  "content": "<p>We can apply the following approach to increase CV score and decrease the training time of our GBT model.<br>\nNote: The idea comes from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> he mentions it somewhere.<br>\nFor me the workflow was as follows:</p>\n<ul>\n<li>Create about 800 features. Use these features to classify answer for each question. (-&gt; 18 different models)</li>\n<li>Use feature importance to determine the 150 most important feature for each model</li>\n<li>Iterate but this times only use the 150 most important features for this model.<br>\nThis boosted my CV from 0.68394 to 0.68473 and significantly reduced training time.</li>\n</ul>",
  "messages": [
    {
      "id": "2147156",
      "postDate": "02/16/2023 12:48:59",
      "content": "<p>We can apply the following approach to increase CV score and decrease the training time of our GBT model.<br>\nNote: The idea comes from <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> he mentions it somewhere.<br>\nFor me the workflow was as follows:</p>\n<ul>\n<li>Create about 800 features. Use these features to classify answer for each question. (-&gt; 18 different models)</li>\n<li>Use feature importance to determine the 150 most important feature for each model</li>\n<li>Iterate but this times only use the 150 most important features for this model.<br>\nThis boosted my CV from 0.68394 to 0.68473 and significantly reduced training time.</li>\n</ul>",
      "rawMarkdown": "We can apply the following approach to increase CV score and decrease the training time of our GBT model.\nNote: The idea comes from @cdeotte he mentions it somewhere.\nFor me the workflow was as follows:\n* Create about 800 features. Use these features to classify answer for each question. (-> 18 different models)\n* Use feature importance to determine the 150 most important feature for each model\n* Iterate but this times only use the 150 most important features for this model.\nThis boosted my CV from 0.68394 to 0.68473 and significantly reduced training time.",
      "votes": null
    },
    {
      "id": "2147211",
      "postDate": "02/16/2023 13:23:38",
      "content": "<p>if you do it recursively  you can call it <a href=\"https://github.com/Yimeng-Zhang/feature-engineering-and-feature-selection/blob/master/A%20Short%20Guide%20for%20Feature%20Engineering%20and%20Feature%20Selection.md#451-recursive-feature-elimination\" target=\"_blank\">Recursive Feature Elimination</a></p>",
      "rawMarkdown": "if you do it recursively  you can call it [Recursive Feature Elimination](https://github.com/Yimeng-Zhang/feature-engineering-and-feature-selection/blob/master/A%20Short%20Guide%20for%20Feature%20Engineering%20and%20Feature%20Selection.md#451-recursive-feature-elimination)",
      "votes": null
    },
    {
      "id": "2147215",
      "postDate": "02/16/2023 13:26:33",
      "content": "<p>that looks similar to the process i applied! tnx for the reference</p>",
      "rawMarkdown": "that looks similar to the process i applied! tnx for the reference",
      "votes": null
    },
    {
      "id": "2147437",
      "postDate": "02/16/2023 16:02:37",
      "content": "<p>I am curious to know why this would work better than adding L1 reg. (Maybe I am doing it poorly)</p>",
      "rawMarkdown": "I am curious to know why this would work better than adding L1 reg. (Maybe I am doing it poorly)",
      "votes": null
    },
    {
      "id": "2147527",
      "postDate": "02/16/2023 17:18:05",
      "content": "<p>I didn't add l1 reg. Will try this out. Currently I use baseline notebook from Chris. Could you explain on l1? I don't know about this concept</p>",
      "rawMarkdown": "I didn't add l1 reg. Will try this out. Currently I use baseline notebook from Chris. Could you explain on l1? I don't know about this concept",
      "votes": null
    },
    {
      "id": "2147692",
      "postDate": "02/16/2023 19:55:39",
      "content": "<p>Here are the two XGB parameters we can add to our XGB models. Docs are <a href=\"https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster\" target=\"_blank\">here</a></p>\n<ul>\n<li>lambda [default=1, alias: reg_lambda]<br>\nL2 regularization term on weights. Increasing this value will make model more conservative.</li>\n<li>alpha [default=0, alias: reg_alpha]<br>\nL1 regularization term on weights. Increasing this value will make model more conservative.</li>\n</ul>\n<p>Using L1 regularization is done with parameter <code>alpha</code>. This is the same technique that Lasso classification uses to remove features (i.e. force trained weights to zero). So presumably using L1 with XGB would be similar to removing features like you have done via feature importance. (But i haven't tried so don't know). Using L2 regularization is like Ridge classification. And Elastic classification is combination of L1 and L2.</p>",
      "rawMarkdown": "Here are the two XGB parameters we can add to our XGB models. Docs are [here][1]\n\n* lambda [default=1, alias: reg_lambda]\nL2 regularization term on weights. Increasing this value will make model more conservative.\n* alpha [default=0, alias: reg_alpha]\nL1 regularization term on weights. Increasing this value will make model more conservative.\n\nUsing L1 regularization is done with parameter `alpha`. This is the same technique that Lasso classification uses to remove features (i.e. force trained weights to zero). So presumably using L1 with XGB would be similar to removing features like you have done via feature importance. (But i haven't tried so don't know). Using L2 regularization is like Ridge classification. And Elastic classification is combination of L1 and L2.\n\n[1]: https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster",
      "votes": null
    },
    {
      "id": "2147698",
      "postDate": "02/16/2023 19:57:44",
      "content": "<p>So perhaps we can try <code>lambda=0</code> and <code>alpha=1</code> to change the default L2 to L1</p>",
      "rawMarkdown": "So perhaps we can try `lambda=0` and `alpha=1` to change the default L2 to L1",
      "votes": null
    },
    {
      "id": "2147705",
      "postDate": "02/16/2023 20:03:12",
      "content": "<p>tnx. will try out. what is the advantage of using L1? When should it be preferred to L2?</p>",
      "rawMarkdown": "tnx. will try out. what is the advantage of using L1? When should it be preferred to L2?",
      "votes": null
    },
    {
      "id": "2147739",
      "postDate": "02/16/2023 20:45:17",
      "content": "<p>The best advice in data science is always try both and see what produces better CV and LB.</p>\n<p>I think L1 shines when we approach \"the curse of dimensionality\" (i.e. when the number of train rows is small compared with number of train columns). I think a rule of thumb is when number of features (i.e. columns) is anywhere near 1/10th (or larger) the number of training samples (i.e. rows) we approach the \"the curse of dimensionality\".</p>\n<p>So if we use 5-Fold, then we have approximately 11k * 0.8 = 9k users per fold model. So when our number of features is anywhere near 900 we start to see a problem. Using L1 is a natural way for the model to exclude features (because it forces their weight to zero thus eliminating them) and perform feature selection.</p>\n<p>So it doesn't surprise me that when you use 150 versus 800 features the model does better. You safely move away from \"the curse of dimensionality\".</p>",
      "rawMarkdown": "The best advice in data science is always try both and see what produces better CV and LB.\n\nI think L1 shines when we approach \"the curse of dimensionality\" (i.e. when the number of train rows is small compared with number of train columns). I think a rule of thumb is when number of features (i.e. columns) is anywhere near 1/10th (or larger) the number of training samples (i.e. rows) we approach the \"the curse of dimensionality\".\n\nSo if we use 5-Fold, then we have approximately 11k * 0.8 = 9k users per fold model. So when our number of features is anywhere near 900 we start to see a problem. Using L1 is a natural way for the model to exclude features (because it forces their weight to zero thus eliminating them) and perform feature selection.\n\nSo it doesn't surprise me that when you use 150 versus 800 features the model does better. You safely move away from \"the curse of dimensionality\".",
      "votes": null
    },
    {
      "id": "2148268",
      "postDate": "02/17/2023 09:44:46",
      "content": "<p>Dang you are right. L2 is enabled by default. That might be causing a lot of problems.</p>",
      "rawMarkdown": "Dang you are right. L2 is enabled by default. That might be causing a lot of problems.",
      "votes": null
    },
    {
      "id": "2149158",
      "postDate": "02/18/2023 02:30:38",
      "content": "<p>If we assume that using fewer features is good due to dimensionality as started, and thus want L1 regularization, and include it… does that potentially create a reason to <em>avoid</em> L2 reg? Or there's no reason to think they can't \"play nice\" together, and usually it will work well to use both together?</p>",
      "rawMarkdown": "If we assume that using fewer features is good due to dimensionality as started, and thus want L1 regularization, and include it... does that potentially create a reason to *avoid* L2 reg? Or there's no reason to think they can't \"play nice\" together, and usually it will work well to use both together?",
      "votes": null
    },
    {
      "id": "2149560",
      "postDate": "02/18/2023 12:41:04",
      "content": "<p>I can confirm that Chris Suggestion to use lambda=0 and than setting alpha parameter works. For me parameter alpha=8 is good choice. </p>",
      "rawMarkdown": "I can confirm that Chris Suggestion to use lambda=0 and than setting alpha parameter works. For me parameter alpha=8 is good choice.",
      "votes": null
    },
    {
      "id": "2150167",
      "postDate": "02/19/2023 02:20:08",
      "content": "<p>But my LB drops <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">Here</a>🤕. May I know how you implement feature importance?</p>",
      "rawMarkdown": "But my LB drops [Here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682)🤕. May I know how you implement feature importance?",
      "votes": null
    },
    {
      "id": "2150463",
      "postDate": "02/19/2023 09:35:49",
      "content": "<p>I see Chris already answered to ur question. Another Option that could work is using L1 regularization (for example lambda=0, alpha=8)</p>",
      "rawMarkdown": "I see Chris already answered to ur question. Another Option that could work is using L1 regularization (for example lambda=0, alpha=8)",
      "votes": null
    },
    {
      "id": "2160323",
      "postDate": "02/26/2023 15:19:58",
      "content": "<p>Isn't this method similar to what described in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">this discussion</a>, which Chris said it was introducing data leaks?</p>",
      "rawMarkdown": "Isn't this method similar to what described in [this discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682), which Chris said it was introducing data leaks?",
      "votes": null
    },
    {
      "id": "2160331",
      "postDate": "02/26/2023 15:32:50",
      "content": "<p>Would you mind sharing more about your methods please <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>? According to the linked doc, if we have <code>N</code> features, we'd re-train up to <code>N-1</code> times, each time removing a feature, and observe the drop in CV. However, my two concerns are:</p>\n<ol>\n<li>When we have 400-500 features, retraining them for 5 folds and 18 questions means we are training 36,000+ models, which seems easy to exceed Kaggle's 9-hour limit. Do you drop 5-10 instead of 1 feature in each feature elimination round?</li>\n<li>Will it introduce data leaks like Chris's discussed in the <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">linked thread</a>.</li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Would you mind sharing more about your methods please @steubk @simonveitner? According to the linked doc, if we have `N` features, we'd re-train up to `N-1` times, each time removing a feature, and observe the drop in CV. However, my two concerns are:\n1. When we have 400-500 features, retraining them for 5 folds and 18 questions means we are training 36,000+ models, which seems easy to exceed Kaggle's 9-hour limit. Do you drop 5-10 instead of 1 feature in each feature elimination round?\n2. Will it introduce data leaks like Chris's discussed in the [linked thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682).\n\nThanks!",
      "votes": null
    },
    {
      "id": "2160439",
      "postDate": "02/26/2023 17:18:03",
      "content": "<ol>\n<li>I don't understand. You will train 5*18 models.</li>\n<li>To avoid leaks you should find best features for each fold and question. You determine 5*18 lists of features which you use in the above models. Than your CV is reliable.</li>\n</ol>",
      "rawMarkdown": "1. I don't understand. You will train 5*18 models.\n2. To avoid leaks you should find best features for each fold and question. You determine 5*18 lists of features which you use in the above models. Than your CV is reliable.",
      "votes": null
    },
    {
      "id": "2160855",
      "postDate": "02/27/2023 03:23:44",
      "content": "<p>Recursive Feature Elimination, as mentioned in the doc linked by <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a>, consists of the following steps:</p>\n<blockquote>\n  <ol>\n  <li>Rank the features according to their importance</li>\n  <li>Remove one feature -the least important- and build a machine learning algorithm utilizing the remaining features.</li>\n  <li>Calculate a performance metric of your choice</li>\n  <li>If the metric decreases by more of an arbitrarily set threshold, then that feature is important and should be kept. Otherwise, we can remove that feature.</li>\n  <li>Repeat steps 2-4 until all features have been removed (and therefore evaluated) and the drop in performance assessed.</li>\n  </ol>\n</blockquote>\n<p>which means if we have <code>N</code> features, we will remove them one by one, each time re-training the model to examine CV score changes. However, I think your Feature Elimination process consists of only one training round, after which you'll remove all except the top 150 features right? And then in inference you'll predict using all 5 folds' models?</p>",
      "rawMarkdown": "Recursive Feature Elimination, as mentioned in the doc linked by @steubk, consists of the following steps:\n> 1. Rank the features according to their importance\n> 2. Remove one feature -the least important- and build a machine learning algorithm utilizing the remaining features.\n> 3. Calculate a performance metric of your choice\n> 4. If the metric decreases by more of an arbitrarily set threshold, then that feature is important and should be kept. Otherwise, we can remove that feature.\n> 5. Repeat steps 2-4 until all features have been removed (and therefore evaluated) and the drop in performance assessed.\n\nwhich means if we have `N` features, we will remove them one by one, each time re-training the model to examine CV score changes. However, I think your Feature Elimination process consists of only one training round, after which you'll remove all except the top 150 features right? And then in inference you'll predict using all 5 folds' models?",
      "votes": null
    },
    {
      "id": "2161146",
      "postDate": "02/27/2023 09:13:56",
      "content": "<p>Correct. That's how I did (other process to time consuming)</p>",
      "rawMarkdown": "Correct. That's how I did (other process to time consuming)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2147211,
      "author_name": "steubk",
      "author_url": "",
      "post_date": "02/16/2023 13:23:38",
      "content": "<p>if you do it recursively  you can call it <a href=\"https://github.com/Yimeng-Zhang/feature-engineering-and-feature-selection/blob/master/A%20Short%20Guide%20for%20Feature%20Engineering%20and%20Feature%20Selection.md#451-recursive-feature-elimination\" target=\"_blank\">Recursive Feature Elimination</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2147215,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "02/16/2023 13:26:33",
          "content": "<p>that looks similar to the process i applied! tnx for the reference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2160323,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "02/26/2023 15:19:58",
          "content": "<p>Isn't this method similar to what described in <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">this discussion</a>, which Chris said it was introducing data leaks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2160331,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "02/26/2023 15:32:50",
          "content": "<p>Would you mind sharing more about your methods please <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a> <a href=\"https://www.kaggle.com/simonveitner\" target=\"_blank\">@simonveitner</a>? According to the linked doc, if we have <code>N</code> features, we'd re-train up to <code>N-1</code> times, each time removing a feature, and observe the drop in CV. However, my two concerns are:</p>\n<ol>\n<li>When we have 400-500 features, retraining them for 5 folds and 18 questions means we are training 36,000+ models, which seems easy to exceed Kaggle's 9-hour limit. Do you drop 5-10 instead of 1 feature in each feature elimination round?</li>\n<li>Will it introduce data leaks like Chris's discussed in the <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">linked thread</a>.</li>\n</ol>\n<p>Thanks!</p>",
          "votes": null,
          "replies": [
            {
              "id": 2160439,
              "author_name": "simonveitner",
              "author_url": "",
              "post_date": "02/26/2023 17:18:03",
              "content": "<ol>\n<li>I don't understand. You will train 5*18 models.</li>\n<li>To avoid leaks you should find best features for each fold and question. You determine 5*18 lists of features which you use in the above models. Than your CV is reliable.</li>\n</ol>",
              "votes": null,
              "replies": [
                {
                  "id": 2160855,
                  "author_name": "hoangnguyen719",
                  "author_url": "",
                  "post_date": "02/27/2023 03:23:44",
                  "content": "<p>Recursive Feature Elimination, as mentioned in the doc linked by <a href=\"https://www.kaggle.com/steubk\" target=\"_blank\">@steubk</a>, consists of the following steps:</p>\n<blockquote>\n  <ol>\n  <li>Rank the features according to their importance</li>\n  <li>Remove one feature -the least important- and build a machine learning algorithm utilizing the remaining features.</li>\n  <li>Calculate a performance metric of your choice</li>\n  <li>If the metric decreases by more of an arbitrarily set threshold, then that feature is important and should be kept. Otherwise, we can remove that feature.</li>\n  <li>Repeat steps 2-4 until all features have been removed (and therefore evaluated) and the drop in performance assessed.</li>\n  </ol>\n</blockquote>\n<p>which means if we have <code>N</code> features, we will remove them one by one, each time re-training the model to examine CV score changes. However, I think your Feature Elimination process consists of only one training round, after which you'll remove all except the top 150 features right? And then in inference you'll predict using all 5 folds' models?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2161146,
                      "author_name": "simonveitner",
                      "author_url": "",
                      "post_date": "02/27/2023 09:13:56",
                      "content": "<p>Correct. That's how I did (other process to time consuming)</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2147437,
      "author_name": "lucasmorin",
      "author_url": "",
      "post_date": "02/16/2023 16:02:37",
      "content": "<p>I am curious to know why this would work better than adding L1 reg. (Maybe I am doing it poorly)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2147527,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "02/16/2023 17:18:05",
          "content": "<p>I didn't add l1 reg. Will try this out. Currently I use baseline notebook from Chris. Could you explain on l1? I don't know about this concept</p>",
          "votes": null,
          "replies": [
            {
              "id": 2147692,
              "author_name": "cdeotte",
              "author_url": "",
              "post_date": "02/16/2023 19:55:39",
              "content": "<p>Here are the two XGB parameters we can add to our XGB models. Docs are <a href=\"https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster\" target=\"_blank\">here</a></p>\n<ul>\n<li>lambda [default=1, alias: reg_lambda]<br>\nL2 regularization term on weights. Increasing this value will make model more conservative.</li>\n<li>alpha [default=0, alias: reg_alpha]<br>\nL1 regularization term on weights. Increasing this value will make model more conservative.</li>\n</ul>\n<p>Using L1 regularization is done with parameter <code>alpha</code>. This is the same technique that Lasso classification uses to remove features (i.e. force trained weights to zero). So presumably using L1 with XGB would be similar to removing features like you have done via feature importance. (But i haven't tried so don't know). Using L2 regularization is like Ridge classification. And Elastic classification is combination of L1 and L2.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2147698,
                  "author_name": "cdeotte",
                  "author_url": "",
                  "post_date": "02/16/2023 19:57:44",
                  "content": "<p>So perhaps we can try <code>lambda=0</code> and <code>alpha=1</code> to change the default L2 to L1</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2147705,
                      "author_name": "simonveitner",
                      "author_url": "",
                      "post_date": "02/16/2023 20:03:12",
                      "content": "<p>tnx. will try out. what is the advantage of using L1? When should it be preferred to L2?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2147739,
                          "author_name": "cdeotte",
                          "author_url": "",
                          "post_date": "02/16/2023 20:45:17",
                          "content": "<p>The best advice in data science is always try both and see what produces better CV and LB.</p>\n<p>I think L1 shines when we approach \"the curse of dimensionality\" (i.e. when the number of train rows is small compared with number of train columns). I think a rule of thumb is when number of features (i.e. columns) is anywhere near 1/10th (or larger) the number of training samples (i.e. rows) we approach the \"the curse of dimensionality\".</p>\n<p>So if we use 5-Fold, then we have approximately 11k * 0.8 = 9k users per fold model. So when our number of features is anywhere near 900 we start to see a problem. Using L1 is a natural way for the model to exclude features (because it forces their weight to zero thus eliminating them) and perform feature selection.</p>\n<p>So it doesn't surprise me that when you use 150 versus 800 features the model does better. You safely move away from \"the curse of dimensionality\".</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2149158,
                              "author_name": "roberthatch",
                              "author_url": "",
                              "post_date": "02/18/2023 02:30:38",
                              "content": "<p>If we assume that using fewer features is good due to dimensionality as started, and thus want L1 regularization, and include it… does that potentially create a reason to <em>avoid</em> L2 reg? Or there's no reason to think they can't \"play nice\" together, and usually it will work well to use both together?</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    },
                    {
                      "id": 2148268,
                      "author_name": "lucasmorin",
                      "author_url": "",
                      "post_date": "02/17/2023 09:44:46",
                      "content": "<p>Dang you are right. L2 is enabled by default. That might be causing a lot of problems.</p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2149560,
      "author_name": "simonveitner",
      "author_url": "",
      "post_date": "02/18/2023 12:41:04",
      "content": "<p>I can confirm that Chris Suggestion to use lambda=0 and than setting alpha parameter works. For me parameter alpha=8 is good choice. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2150167,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "02/19/2023 02:20:08",
      "content": "<p>But my LB drops <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682\" target=\"_blank\">Here</a>🤕. May I know how you implement feature importance?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2150463,
          "author_name": "simonveitner",
          "author_url": "",
          "post_date": "02/19/2023 09:35:49",
          "content": "<p>I see Chris already answered to ur question. Another Option that could work is using L1 regularization (for example lambda=0, alpha=8)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2147156": "We can apply the following approach to increase CV score and decrease the training time of our GBT model.\nNote: The idea comes from @cdeotte he mentions it somewhere.\nFor me the workflow was as follows:\n* Create about 800 features. Use these features to classify answer for each question. (-> 18 different models)\n* Use feature importance to determine the 150 most important feature for each model\n* Iterate but this times only use the 150 most important features for this model.\nThis boosted my CV from 0.68394 to 0.68473 and significantly reduced training time.",
    "2147211": "if you do it recursively  you can call it [Recursive Feature Elimination](https://github.com/Yimeng-Zhang/feature-engineering-and-feature-selection/blob/master/A%20Short%20Guide%20for%20Feature%20Engineering%20and%20Feature%20Selection.md#451-recursive-feature-elimination)",
    "2147215": "that looks similar to the process i applied! tnx for the reference",
    "2147437": "I am curious to know why this would work better than adding L1 reg. (Maybe I am doing it poorly)",
    "2147527": "I didn't add l1 reg. Will try this out. Currently I use baseline notebook from Chris. Could you explain on l1? I don't know about this concept",
    "2147692": "Here are the two XGB parameters we can add to our XGB models. Docs are [here][1]\n\n* lambda [default=1, alias: reg_lambda]\nL2 regularization term on weights. Increasing this value will make model more conservative.\n* alpha [default=0, alias: reg_alpha]\nL1 regularization term on weights. Increasing this value will make model more conservative.\n\nUsing L1 regularization is done with parameter `alpha`. This is the same technique that Lasso classification uses to remove features (i.e. force trained weights to zero). So presumably using L1 with XGB would be similar to removing features like you have done via feature importance. (But i haven't tried so don't know). Using L2 regularization is like Ridge classification. And Elastic classification is combination of L1 and L2.\n\n[1]: https://xgboost.readthedocs.io/en/stable/parameter.html#parameters-for-tree-booster",
    "2147698": "So perhaps we can try `lambda=0` and `alpha=1` to change the default L2 to L1",
    "2147705": "tnx. will try out. what is the advantage of using L1? When should it be preferred to L2?",
    "2147739": "The best advice in data science is always try both and see what produces better CV and LB.\n\nI think L1 shines when we approach \"the curse of dimensionality\" (i.e. when the number of train rows is small compared with number of train columns). I think a rule of thumb is when number of features (i.e. columns) is anywhere near 1/10th (or larger) the number of training samples (i.e. rows) we approach the \"the curse of dimensionality\".\n\nSo if we use 5-Fold, then we have approximately 11k * 0.8 = 9k users per fold model. So when our number of features is anywhere near 900 we start to see a problem. Using L1 is a natural way for the model to exclude features (because it forces their weight to zero thus eliminating them) and perform feature selection.\n\nSo it doesn't surprise me that when you use 150 versus 800 features the model does better. You safely move away from \"the curse of dimensionality\".",
    "2148268": "Dang you are right. L2 is enabled by default. That might be causing a lot of problems.",
    "2149158": "If we assume that using fewer features is good due to dimensionality as started, and thus want L1 regularization, and include it... does that potentially create a reason to *avoid* L2 reg? Or there's no reason to think they can't \"play nice\" together, and usually it will work well to use both together?",
    "2149560": "I can confirm that Chris Suggestion to use lambda=0 and than setting alpha parameter works. For me parameter alpha=8 is good choice.",
    "2150167": "But my LB drops [Here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682)🤕. May I know how you implement feature importance?",
    "2150463": "I see Chris already answered to ur question. Another Option that could work is using L1 regularization (for example lambda=0, alpha=8)",
    "2160323": "Isn't this method similar to what described in [this discussion](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682), which Chris said it was introducing data leaks?",
    "2160331": "Would you mind sharing more about your methods please @steubk @simonveitner? According to the linked doc, if we have `N` features, we'd re-train up to `N-1` times, each time removing a feature, and observe the drop in CV. However, my two concerns are:\n1. When we have 400-500 features, retraining them for 5 folds and 18 questions means we are training 36,000+ models, which seems easy to exceed Kaggle's 9-hour limit. Do you drop 5-10 instead of 1 feature in each feature elimination round?\n2. Will it introduce data leaks like Chris's discussed in the [linked thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388682).\n\nThanks!",
    "2160439": "1. I don't understand. You will train 5*18 models.\n2. To avoid leaks you should find best features for each fold and question. You determine 5*18 lists of features which you use in the above models. Than your CV is reliable.",
    "2160855": "Recursive Feature Elimination, as mentioned in the doc linked by @steubk, consists of the following steps:\n> 1. Rank the features according to their importance\n> 2. Remove one feature -the least important- and build a machine learning algorithm utilizing the remaining features.\n> 3. Calculate a performance metric of your choice\n> 4. If the metric decreases by more of an arbitrarily set threshold, then that feature is important and should be kept. Otherwise, we can remove that feature.\n> 5. Repeat steps 2-4 until all features have been removed (and therefore evaluated) and the drop in performance assessed.\n\nwhich means if we have `N` features, we will remove them one by one, each time re-training the model to examine CV score changes. However, I think your Feature Elimination process consists of only one training round, after which you'll remove all except the top 150 features right? And then in inference you'll predict using all 5 folds' models?",
    "2161146": "Correct. That's how I did (other process to time consuming)"
  },
  "source": "meta"
}