{
  "id": 53820,
  "title": "Feature engineering - how to determine good features?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/53820",
  "author_name": "",
  "post_date": "2018-04-05T14:31:49.313007100Z",
  "votes": null,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I have a simple question... I developed a lot of new features, but I'm not sure how to evaluate them. I created some graphs they show a correlation with the number of downloads.</p>\n\n<p>Now I would like to make sure they have a real impact on the LightGBM model. This new feature for instance might not be orthogonal with some existing feature, and have no benefits. </p>\n\n<p>So how do you ensure those features are good ? </p>\n\n<ul>\n<li>Do you add them one by one, and check the improvement in the training\nset / validation set error ? </li>\n<li>Or add as many features as you can and\ncheck the \"gain\" from the output of lightGBM ?</li>\n</ul>\n\n<p>Thanks for the help !</p>",
  "messages": [
    {
      "id": "309529",
      "postDate": "04/05/2018 14:31:49",
      "content": "<p>Hello,</p>\n\n<p>I have a simple question... I developed a lot of new features, but I'm not sure how to evaluate them. I created some graphs they show a correlation with the number of downloads.</p>\n\n<p>Now I would like to make sure they have a real impact on the LightGBM model. This new feature for instance might not be orthogonal with some existing feature, and have no benefits. </p>\n\n<p>So how do you ensure those features are good ? </p>\n\n<ul>\n<li>Do you add them one by one, and check the improvement in the training\nset / validation set error ? </li>\n<li>Or add as many features as you can and\ncheck the \"gain\" from the output of lightGBM ?</li>\n</ul>\n\n<p>Thanks for the help !</p>",
      "rawMarkdown": "Hello,\n\nI have a simple question... I developed a lot of new features, but I'm not sure how to evaluate them. I created some graphs they show a correlation with the number of downloads.\n\nNow I would like to make sure they have a real impact on the LightGBM model. This new feature for instance might not be orthogonal with some existing feature, and have no benefits. \n\nSo how do you ensure those features are good ? \n\n - Do you add them one by one, and check the improvement in the training\n   set / validation set error ? \n - Or add as many features as you can and\n   check the \"gain\" from the output of lightGBM ?\n\nThanks for the help !",
      "votes": null
    },
    {
      "id": "309569",
      "postDate": "04/05/2018 15:41:26",
      "content": "<p>You could use the leaderboard (Bad)</p>\n\n<p>Or Cross Validation (Good)</p>",
      "rawMarkdown": "You could use the leaderboard (Bad)\n\nOr Cross Validation (Good)",
      "votes": null
    },
    {
      "id": "309587",
      "postDate": "04/05/2018 16:25:16",
      "content": "<p>So I guess I have to use smaller training set and add the features 1 by 1 to see the impact. </p>\n\n<p>I was wondering if it was more reliable to use cross validation or the model gain given by LightGBM. </p>\n\n<p>Thanks a lot !</p>",
      "rawMarkdown": "So I guess I have to use smaller training set and add the features 1 by 1 to see the impact. \n\nI was wondering if it was more reliable to use cross validation or the model gain given by LightGBM. \n\nThanks a lot !",
      "votes": null
    },
    {
      "id": "309592",
      "postDate": "04/05/2018 16:31:07",
      "content": "<p>Hi Scirpus, do you use cross validation for feature selection? I think this data is time-wised and using cross validation might not be a good idea.</p>",
      "rawMarkdown": "Hi Scirpus, do you use cross validation for feature selection? I think this data is time-wised and using cross validation might not be a good idea.",
      "votes": null
    },
    {
      "id": "309615",
      "postDate": "04/05/2018 17:24:06",
      "content": "<p>It is but we are tryng to use training data to predict for +1 day</p>\n\n<p>So you can train on 7,8 days and see how it performs on day 9</p>",
      "rawMarkdown": "It is but we are tryng to use training data to predict for +1 day\n\nSo you can train on 7,8 days and see how it performs on day 9",
      "votes": null
    },
    {
      "id": "309623",
      "postDate": "04/05/2018 17:33:32",
      "content": "<p>Make sense... very helpful thanks ! This comment was mentioning the same idea :\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634#308329\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634#308329</a></p>",
      "rawMarkdown": "Make sense... very helpful thanks ! This comment was mentioning the same idea :\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634#308329",
      "votes": null
    },
    {
      "id": "309649",
      "postDate": "04/05/2018 18:24:59",
      "content": "<p>I would say try it both ways.  The validation score (with occasional leaderboard scores for confirmation) is your ultimate guide, but there are many ways to go.  Using gain runs a danger of overfitting, but adding features one-by-one runs a danger of missing useful combinations of features.  So do some of both, and maybe also apply some intuition (or domain knowledge) about which features might work well together.</p>",
      "rawMarkdown": "I would say try it both ways.  The validation score (with occasional leaderboard scores for confirmation) is your ultimate guide, but there are many ways to go.  Using gain runs a danger of overfitting, but adding features one-by-one runs a danger of missing useful combinations of features.  So do some of both, and maybe also apply some intuition (or domain knowledge) about which features might work well together.",
      "votes": null
    },
    {
      "id": "309659",
      "postDate": "04/05/2018 18:56:00",
      "content": "<p>I like to narrow down my full feature set to some \"admissible\" feature set first, and then I do experiments modeling with subsets of those features (with some random/not-so-random selection). I use some basic heuristics to get my admissible set. It may be far from optimal and predicated on many assumptions, but it has the added benefit of being easier to check.</p>",
      "rawMarkdown": "I like to narrow down my full feature set to some \"admissible\" feature set first, and then I do experiments modeling with subsets of those features (with some random/not-so-random selection). I use some basic heuristics to get my admissible set. It may be far from optimal and predicated on many assumptions, but it has the added benefit of being easier to check.",
      "votes": null
    },
    {
      "id": "309665",
      "postDate": "04/05/2018 19:13:19",
      "content": "<blockquote>\n  <p>So how do you ensure those features are good ?</p>\n</blockquote>\n\n<p>Feature selection is a well-tractable problem in general terms, but some avenues are more practical than others when it comes to this particular dataset.</p>\n\n<blockquote>\n  <p>Do you add them one by one, and check the improvement in the training set / validation set error ?</p>\n</blockquote>\n\n<p>One way to select features is by applying greedy techniques, which can be done in forward or backward direction. That means you start with 1 feature and keep others as long as their addition improves the evaluation score (forward); or start with all the features and eliminate features one by one, keeping those whose elimination lowered the evaluation score (backward). I am partial to recursive feature elimination with cross-validation (<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFECV.html\"><strong>RFECV</strong></a>), but that is really not feasible for a dataset of this size because you have to do the whole procedure N times (N is number of folds).</p>\n\n<blockquote>\n  <p>Or add as many features as you can and check the \"gain\" from the output of lightGBM ?</p>\n</blockquote>\n\n<p>That's another way of doing it, and probably the most practical one. You've been warned already about the possibility of overfitting when doing it this way, and that's a real problem. I wouldn't worry about that too much if you are using all the data, and especially if you make sure only shallow trees (max_dept = 3-5) and do a limited number of iterations ( 100-200, maybe 500).</p>",
      "rawMarkdown": "&gt; So how do you ensure those features are good ?\n\nFeature selection is a well-tractable problem in general terms, but some avenues are more practical than others when it comes to this particular dataset.\n\n&gt; Do you add them one by one, and check the improvement in the training set / validation set error ?\n\nOne way to select features is by applying greedy techniques, which can be done in forward or backward direction. That means you start with 1 feature and keep others as long as their addition improves the evaluation score (forward); or start with all the features and eliminate features one by one, keeping those whose elimination lowered the evaluation score (backward). I am partial to recursive feature elimination with cross-validation ([__RFECV__](http://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFECV.html)), but that is really not feasible for a dataset of this size because you have to do the whole procedure N times (N is number of folds).\n\n&gt; Or add as many features as you can and check the \"gain\" from the output of lightGBM ?\n\nThat's another way of doing it, and probably the most practical one. You've been warned already about the possibility of overfitting when doing it this way, and that's a real problem. I wouldn't worry about that too much if you are using all the data, and especially if you make sure only shallow trees (max_dept = 3-5) and do a limited number of iterations ( 100-200, maybe 500).",
      "votes": null
    },
    {
      "id": "310019",
      "postDate": "04/06/2018 12:12:06",
      "content": "<p>I didn't know about the forward backward features selection.</p>\n\n<p>The problem with this huge dataset, is that it's very hard to keep all data / features in memory, and train the models with different predictors.</p>\n\n<p>I used a lot the \"gain\" output previously, but with all my fancy features I was never able to go over a few % gain. 70% of the gain was coming from the 'app' category... So I guess my feature were just not good enough.</p>\n\n<p>Thanks for your input !</p>",
      "rawMarkdown": "I didn't know about the forward backward features selection.\n\nThe problem with this huge dataset, is that it's very hard to keep all data / features in memory, and train the models with different predictors.\n\nI used a lot the \"gain\" output previously, but with all my fancy features I was never able to go over a few % gain. 70% of the gain was coming from the 'app' category... So I guess my feature were just not good enough.\n\nThanks for your input !",
      "votes": null
    },
    {
      "id": "310020",
      "postDate": "04/06/2018 12:15:04",
      "content": "<p>Ok I see thanks... My problem is probably coming from my lack of intuition ;) So I wanted to know if there was a \"recommended\" way to do feature selection. I'll continue testing the different approaches !</p>",
      "rawMarkdown": "Ok I see thanks... My problem is probably coming from my lack of intuition ;) So I wanted to know if there was a \"recommended\" way to do feature selection. I'll continue testing the different approaches !",
      "votes": null
    },
    {
      "id": "310041",
      "postDate": "04/06/2018 13:00:11",
      "content": "<p>I have around 30 features. My approach:</p>\n\n<ul>\n<li>Run the model by selecting 5 features randomly from the set. I do this 50 times in a loop. </li>\n<li>For every iteration (each run of the full model), store the score and the list of features used in that iteration. </li>\n<li>After the 50 runs of the model, split the features into single strings and take an avg of the score for each feature. \nFor example: ['F1','F2','F3','F4','F5'] has a score of 0.9543. I split the list into distinct rows with the same score. Concatenate the results similarly from all 50 iterations.</li>\n<li>Pick the top n features according to their avg score.</li>\n</ul>\n\n<p>Even though it took a day to run, it was certainly better than doing it manually. Recursive, forward, backward would take even more time to run.</p>\n\n<p>Edit: After doing the above, I observed that the approach may be picking some highly correlated variables as well. You can choose one of them if the corr is very high</p>",
      "rawMarkdown": "I have around 30 features. My approach:\n\n - Run the model by selecting 5 features randomly from the set. I do this 50 times in a loop. \n - For every iteration (each run of the full model), store the score and the list of features used in that iteration. \n - After the 50 runs of the model, split the features into single strings and take an avg of the score for each feature. \nFor example: ['F1','F2','F3','F4','F5'] has a score of 0.9543. I split the list into distinct rows with the same score. Concatenate the results similarly from all 50 iterations.\n - Pick the top n features according to their avg score.\n\nEven though it took a day to run, it was certainly better than doing it manually. Recursive, forward, backward would take even more time to run.\n\nEdit: After doing the above, I observed that the approach may be picking some highly correlated variables as well. You can choose one of them if the corr is very high",
      "votes": null
    },
    {
      "id": "310154",
      "postDate": "04/06/2018 17:33:50",
      "content": "<p>Hi, @Antoine.</p>\n\n<p>If you have RAM constraints, I recommend the following workflow:</p>\n\n<h1>Make Data Sets</h1>\n\n<p>Break up your data sets how you'd like them. Possibly:</p>\n\n<ol>\n<li>The training data over all of the days you will use for training.</li>\n<li>The validation data over all of the days/hour you will validate.</li>\n<li>Smaller samples of those for <em>sanity checks</em>. *You'd rather not check an idea on a huge set.</li>\n</ol>\n\n<h1>For-Loop (over data sets):</h1>\n\n<h2>Setup</h2>\n\n<ol>\n<li>Get baseline data (training set 1, 2, 3, etc., validation set 1 2, 3, etc., test set).</li>\n<li>Data Prep / Type Down-casting / Etc.</li>\n</ol>\n\n<h2>For-Loop (for Feature Engineering):</h2>\n\n<ol>\n<li>Create Feature.</li>\n<li>Merge to Original data set.</li>\n<li>Make any addition features reliant upon the newly created one.</li>\n<li>Serialize the new features.\n<ul><li>Remove from Original data set.</li>\n<li>Garbage collect.</li></ul></li>\n</ol>\n\n<p>If you also have compute/storage constraints, then you may need to be really economical in your design.</p>\n\n<p>Hope that helps. Good luck!</p>",
      "rawMarkdown": "Hi, @Antoine.\n\nIf you have RAM constraints, I recommend the following workflow:\n\n# Make Data Sets\n\nBreak up your data sets how you'd like them. Possibly:\n\n 1. The training data over all of the days you will use for training.\n 2. The validation data over all of the days/hour you will validate.\n 3. Smaller samples of those for *sanity checks*. *You'd rather not check an idea on a huge set.\n\n# For-Loop (over data sets):\n\n## Setup\n 1. Get baseline data (training set 1, 2, 3, etc., validation set 1 2, 3, etc., test set).\n 2. Data Prep / Type Down-casting / Etc.\n\n## For-Loop (for Feature Engineering):\n 1. Create Feature.\n 2. Merge to Original data set.\n 3. Make any addition features reliant upon the newly created one.\n 4. Serialize the new features.\n     * Remove from Original data set.\n     * Garbage collect.\n\nIf you also have compute/storage constraints, then you may need to be really economical in your design.\n\nHope that helps. Good luck!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 309569,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "04/05/2018 15:41:26",
      "content": "<p>You could use the leaderboard (Bad)</p>\n\n<p>Or Cross Validation (Good)</p>",
      "votes": null,
      "replies": [
        {
          "id": 309592,
          "author_name": "wythhh",
          "author_url": "",
          "post_date": "04/05/2018 16:31:07",
          "content": "<p>Hi Scirpus, do you use cross validation for feature selection? I think this data is time-wised and using cross validation might not be a good idea.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309587,
      "author_name": "areveillon",
      "author_url": "",
      "post_date": "04/05/2018 16:25:16",
      "content": "<p>So I guess I have to use smaller training set and add the features 1 by 1 to see the impact. </p>\n\n<p>I was wondering if it was more reliable to use cross validation or the model gain given by LightGBM. </p>\n\n<p>Thanks a lot !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 309615,
      "author_name": "scirpus",
      "author_url": "",
      "post_date": "04/05/2018 17:24:06",
      "content": "<p>It is but we are tryng to use training data to predict for +1 day</p>\n\n<p>So you can train on 7,8 days and see how it performs on day 9</p>",
      "votes": null,
      "replies": [
        {
          "id": 309623,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "04/05/2018 17:33:32",
          "content": "<p>Make sense... very helpful thanks ! This comment was mentioning the same idea :\n<a href=\"https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634#308329\">https://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634#308329</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309649,
      "author_name": "aharless",
      "author_url": "",
      "post_date": "04/05/2018 18:24:59",
      "content": "<p>I would say try it both ways.  The validation score (with occasional leaderboard scores for confirmation) is your ultimate guide, but there are many ways to go.  Using gain runs a danger of overfitting, but adding features one-by-one runs a danger of missing useful combinations of features.  So do some of both, and maybe also apply some intuition (or domain knowledge) about which features might work well together.</p>",
      "votes": null,
      "replies": [
        {
          "id": 310020,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "04/06/2018 12:15:04",
          "content": "<p>Ok I see thanks... My problem is probably coming from my lack of intuition ;) So I wanted to know if there was a \"recommended\" way to do feature selection. I'll continue testing the different approaches !</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 309659,
      "author_name": "puremath86",
      "author_url": "",
      "post_date": "04/05/2018 18:56:00",
      "content": "<p>I like to narrow down my full feature set to some \"admissible\" feature set first, and then I do experiments modeling with subsets of those features (with some random/not-so-random selection). I use some basic heuristics to get my admissible set. It may be far from optimal and predicated on many assumptions, but it has the added benefit of being easier to check.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 309665,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "04/05/2018 19:13:19",
      "content": "<blockquote>\n  <p>So how do you ensure those features are good ?</p>\n</blockquote>\n\n<p>Feature selection is a well-tractable problem in general terms, but some avenues are more practical than others when it comes to this particular dataset.</p>\n\n<blockquote>\n  <p>Do you add them one by one, and check the improvement in the training set / validation set error ?</p>\n</blockquote>\n\n<p>One way to select features is by applying greedy techniques, which can be done in forward or backward direction. That means you start with 1 feature and keep others as long as their addition improves the evaluation score (forward); or start with all the features and eliminate features one by one, keeping those whose elimination lowered the evaluation score (backward). I am partial to recursive feature elimination with cross-validation (<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFECV.html\"><strong>RFECV</strong></a>), but that is really not feasible for a dataset of this size because you have to do the whole procedure N times (N is number of folds).</p>\n\n<blockquote>\n  <p>Or add as many features as you can and check the \"gain\" from the output of lightGBM ?</p>\n</blockquote>\n\n<p>That's another way of doing it, and probably the most practical one. You've been warned already about the possibility of overfitting when doing it this way, and that's a real problem. I wouldn't worry about that too much if you are using all the data, and especially if you make sure only shallow trees (max_dept = 3-5) and do a limited number of iterations ( 100-200, maybe 500).</p>",
      "votes": null,
      "replies": [
        {
          "id": 310019,
          "author_name": "areveillon",
          "author_url": "",
          "post_date": "04/06/2018 12:12:06",
          "content": "<p>I didn't know about the forward backward features selection.</p>\n\n<p>The problem with this huge dataset, is that it's very hard to keep all data / features in memory, and train the models with different predictors.</p>\n\n<p>I used a lot the \"gain\" output previously, but with all my fancy features I was never able to go over a few % gain. 70% of the gain was coming from the 'app' category... So I guess my feature were just not good enough.</p>\n\n<p>Thanks for your input !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 310154,
          "author_name": "puremath86",
          "author_url": "",
          "post_date": "04/06/2018 17:33:50",
          "content": "<p>Hi, @Antoine.</p>\n\n<p>If you have RAM constraints, I recommend the following workflow:</p>\n\n<h1>Make Data Sets</h1>\n\n<p>Break up your data sets how you'd like them. Possibly:</p>\n\n<ol>\n<li>The training data over all of the days you will use for training.</li>\n<li>The validation data over all of the days/hour you will validate.</li>\n<li>Smaller samples of those for <em>sanity checks</em>. *You'd rather not check an idea on a huge set.</li>\n</ol>\n\n<h1>For-Loop (over data sets):</h1>\n\n<h2>Setup</h2>\n\n<ol>\n<li>Get baseline data (training set 1, 2, 3, etc., validation set 1 2, 3, etc., test set).</li>\n<li>Data Prep / Type Down-casting / Etc.</li>\n</ol>\n\n<h2>For-Loop (for Feature Engineering):</h2>\n\n<ol>\n<li>Create Feature.</li>\n<li>Merge to Original data set.</li>\n<li>Make any addition features reliant upon the newly created one.</li>\n<li>Serialize the new features.\n<ul><li>Remove from Original data set.</li>\n<li>Garbage collect.</li></ul></li>\n</ol>\n\n<p>If you also have compute/storage constraints, then you may need to be really economical in your design.</p>\n\n<p>Hope that helps. Good luck!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310041,
      "author_name": "ug2409",
      "author_url": "",
      "post_date": "04/06/2018 13:00:11",
      "content": "<p>I have around 30 features. My approach:</p>\n\n<ul>\n<li>Run the model by selecting 5 features randomly from the set. I do this 50 times in a loop. </li>\n<li>For every iteration (each run of the full model), store the score and the list of features used in that iteration. </li>\n<li>After the 50 runs of the model, split the features into single strings and take an avg of the score for each feature. \nFor example: ['F1','F2','F3','F4','F5'] has a score of 0.9543. I split the list into distinct rows with the same score. Concatenate the results similarly from all 50 iterations.</li>\n<li>Pick the top n features according to their avg score.</li>\n</ul>\n\n<p>Even though it took a day to run, it was certainly better than doing it manually. Recursive, forward, backward would take even more time to run.</p>\n\n<p>Edit: After doing the above, I observed that the approach may be picking some highly correlated variables as well. You can choose one of them if the corr is very high</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "309529": "Hello,\n\nI have a simple question... I developed a lot of new features, but I'm not sure how to evaluate them. I created some graphs they show a correlation with the number of downloads.\n\nNow I would like to make sure they have a real impact on the LightGBM model. This new feature for instance might not be orthogonal with some existing feature, and have no benefits. \n\nSo how do you ensure those features are good ? \n\n - Do you add them one by one, and check the improvement in the training\n   set / validation set error ? \n - Or add as many features as you can and\n   check the \"gain\" from the output of lightGBM ?\n\nThanks for the help !",
    "309569": "You could use the leaderboard (Bad)\n\nOr Cross Validation (Good)",
    "309587": "So I guess I have to use smaller training set and add the features 1 by 1 to see the impact. \n\nI was wondering if it was more reliable to use cross validation or the model gain given by LightGBM. \n\nThanks a lot !",
    "309592": "Hi Scirpus, do you use cross validation for feature selection? I think this data is time-wised and using cross validation might not be a good idea.",
    "309615": "It is but we are tryng to use training data to predict for +1 day\n\nSo you can train on 7,8 days and see how it performs on day 9",
    "309623": "Make sense... very helpful thanks ! This comment was mentioning the same idea :\nhttps://www.kaggle.com/c/talkingdata-adtracking-fraud-detection/discussion/53634#308329",
    "309649": "I would say try it both ways.  The validation score (with occasional leaderboard scores for confirmation) is your ultimate guide, but there are many ways to go.  Using gain runs a danger of overfitting, but adding features one-by-one runs a danger of missing useful combinations of features.  So do some of both, and maybe also apply some intuition (or domain knowledge) about which features might work well together.",
    "309659": "I like to narrow down my full feature set to some \"admissible\" feature set first, and then I do experiments modeling with subsets of those features (with some random/not-so-random selection). I use some basic heuristics to get my admissible set. It may be far from optimal and predicated on many assumptions, but it has the added benefit of being easier to check.",
    "309665": "&gt; So how do you ensure those features are good ?\n\nFeature selection is a well-tractable problem in general terms, but some avenues are more practical than others when it comes to this particular dataset.\n\n&gt; Do you add them one by one, and check the improvement in the training set / validation set error ?\n\nOne way to select features is by applying greedy techniques, which can be done in forward or backward direction. That means you start with 1 feature and keep others as long as their addition improves the evaluation score (forward); or start with all the features and eliminate features one by one, keeping those whose elimination lowered the evaluation score (backward). I am partial to recursive feature elimination with cross-validation ([__RFECV__](http://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.RFECV.html)), but that is really not feasible for a dataset of this size because you have to do the whole procedure N times (N is number of folds).\n\n&gt; Or add as many features as you can and check the \"gain\" from the output of lightGBM ?\n\nThat's another way of doing it, and probably the most practical one. You've been warned already about the possibility of overfitting when doing it this way, and that's a real problem. I wouldn't worry about that too much if you are using all the data, and especially if you make sure only shallow trees (max_dept = 3-5) and do a limited number of iterations ( 100-200, maybe 500).",
    "310019": "I didn't know about the forward backward features selection.\n\nThe problem with this huge dataset, is that it's very hard to keep all data / features in memory, and train the models with different predictors.\n\nI used a lot the \"gain\" output previously, but with all my fancy features I was never able to go over a few % gain. 70% of the gain was coming from the 'app' category... So I guess my feature were just not good enough.\n\nThanks for your input !",
    "310020": "Ok I see thanks... My problem is probably coming from my lack of intuition ;) So I wanted to know if there was a \"recommended\" way to do feature selection. I'll continue testing the different approaches !",
    "310041": "I have around 30 features. My approach:\n\n - Run the model by selecting 5 features randomly from the set. I do this 50 times in a loop. \n - For every iteration (each run of the full model), store the score and the list of features used in that iteration. \n - After the 50 runs of the model, split the features into single strings and take an avg of the score for each feature. \nFor example: ['F1','F2','F3','F4','F5'] has a score of 0.9543. I split the list into distinct rows with the same score. Concatenate the results similarly from all 50 iterations.\n - Pick the top n features according to their avg score.\n\nEven though it took a day to run, it was certainly better than doing it manually. Recursive, forward, backward would take even more time to run.\n\nEdit: After doing the above, I observed that the approach may be picking some highly correlated variables as well. You can choose one of them if the corr is very high",
    "310154": "Hi, @Antoine.\n\nIf you have RAM constraints, I recommend the following workflow:\n\n# Make Data Sets\n\nBreak up your data sets how you'd like them. Possibly:\n\n 1. The training data over all of the days you will use for training.\n 2. The validation data over all of the days/hour you will validate.\n 3. Smaller samples of those for *sanity checks*. *You'd rather not check an idea on a huge set.\n\n# For-Loop (over data sets):\n\n## Setup\n 1. Get baseline data (training set 1, 2, 3, etc., validation set 1 2, 3, etc., test set).\n 2. Data Prep / Type Down-casting / Etc.\n\n## For-Loop (for Feature Engineering):\n 1. Create Feature.\n 2. Merge to Original data set.\n 3. Make any addition features reliant upon the newly created one.\n 4. Serialize the new features.\n     * Remove from Original data set.\n     * Garbage collect.\n\nIf you also have compute/storage constraints, then you may need to be really economical in your design.\n\nHope that helps. Good luck!"
  },
  "source": "meta"
}