{
  "id": 334139,
  "title": "Should one tune hyperparameters before feature engineering?",
  "url": "/competitions/amex-default-prediction/discussion/334139",
  "author_name": "Man of the year",
  "post_date": "2022-06-30T01:04:16.518000",
  "votes": 12,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I had this question. Let's say one use LGBM without any feature engineering, there are many hyperparameter combinations and one finds the optimal params for best performance. Then let's say there is some feature engineering, some magic features that definitely improves score. Will it be necessary to tune hyperparameters again after adding this feature to find new optimal ones?</p>\n<p>Also similar question. Let's say we have some set of hyperparameters and we want to test a new feature, is it helpful or no. Can it be so that on this exact set of hyperparams quality will decrease, however, this feature is very helpful and on some different set of hyperparams it might increase quality significantly? If it is so, how should new features or tricks be tested?</p>",
  "messages": [
    {
      "id": 1837818,
      "postDate": "2022-06-30T01:04:16.520Z",
      "content": "<p>I had this question. Let's say one use LGBM without any feature engineering, there are many hyperparameter combinations and one finds the optimal params for best performance. Then let's say there is some feature engineering, some magic features that definitely improves score. Will it be necessary to tune hyperparameters again after adding this feature to find new optimal ones?</p>\n<p>Also similar question. Let's say we have some set of hyperparameters and we want to test a new feature, is it helpful or no. Can it be so that on this exact set of hyperparams quality will decrease, however, this feature is very helpful and on some different set of hyperparams it might increase quality significantly? If it is so, how should new features or tricks be tested?</p>",
      "rawMarkdown": "I had this question. Let's say one use LGBM without any feature engineering, there are many hyperparameter combinations and one finds the optimal params for best performance. Then let's say there is some feature engineering, some magic features that definitely improves score. Will it be necessary to tune hyperparameters again after adding this feature to find new optimal ones?\n\nAlso similar question. Let's say we have some set of hyperparameters and we want to test a new feature, is it helpful or no. Can it be so that on this exact set of hyperparams quality will decrease, however, this feature is very helpful and on some different set of hyperparams it might increase quality significantly? If it is so, how should new features or tricks be tested?",
      "votes": 12
    },
    {
      "id": 1838791,
      "postDate": "2022-06-30T20:18:12.500Z",
      "content": "<p>Hyperparameters optimization should be the last step, if not, you might overfit your validation set (or multiple K validation sets).</p>\n<p>This is what I normally do:</p>\n<ul>\n<li>Select model(s) and use the base parameters (without tuning).</li>\n<li>Select features and decide what features you want to use and how to extract them (Bin, One-hot, Frequency, Normalization, etc…).</li>\n<li>Train and Validate with each Feature selection and finally choose the best feature set.</li>\n<li>Tune hyperparameters for best feature set</li>\n</ul>",
      "rawMarkdown": "Hyperparameters optimization should be the last step, if not, you might overfit your validation set (or multiple K validation sets).\n\nThis is what I normally do:\n\n- Select model(s) and use the base parameters (without tuning).\n- Select features and decide what features you want to use and how to extract them (Bin, One-hot, Frequency, Normalization, etc...).\n- Train and Validate with each Feature selection and finally choose the best feature set.\n- Tune hyperparameters for best feature set\n",
      "votes": 9,
      "replies": [
        {
          "id": 1838878,
          "postDate": "2022-06-30T23:23:07.633Z",
          "content": "<p>Wouldnt it be better to HPO first, so that the model is familiar with the type of data you are using? And then doing feature selection? Also is there any where we can see that HPO then FS is overfitting?</p>",
          "rawMarkdown": "Wouldnt it be better to HPO first, so that the model is familiar with the type of data you are using? And then doing feature selection? Also is there any where we can see that HPO then FS is overfitting?"
        },
        {
          "id": 1838882,
          "postDate": "2022-06-30T23:37:25.317Z",
          "content": "<p>I think hyperparams optimization might be useless, if you later add or delete features. As I understood it's easier to check features usefulness on default params and later tune them, when you choose features. About overfitting validation I also do not really understood this, I mean, if you do cross validation on same hyperparams, you'll see what are the best for all folds, not just for the one of them. And then you have public lb score to test final quality and see, if you overfitted due cross validation or not.</p>",
          "rawMarkdown": "I think hyperparams optimization might be useless, if you later add or delete features. As I understood it's easier to check features usefulness on default params and later tune them, when you choose features. About overfitting validation I also do not really understood this, I mean, if you do cross validation on same hyperparams, you'll see what are the best for all folds, not just for the one of them. And then you have public lb score to test final quality and see, if you overfitted due cross validation or not.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1838081,
      "postDate": "2022-06-30T07:36:58.490Z",
      "content": "<p>Changed or added features definetly affect the model directly by changing the values or adding new dimensions to the feature space, thats why you should tune the parameters after you have done your feature engineering.</p>",
      "rawMarkdown": "Changed or added features definetly affect the model directly by changing the values or adding new dimensions to the feature space, thats why you should tune the parameters after you have done your feature engineering.",
      "votes": 6
    },
    {
      "id": 1838910,
      "postDate": "2022-07-01T02:01:18.850Z",
      "content": "<p>As everyone suggested the first step is FE and the last step is stratified k-fold grid search.<br>\nHere are two interesting papers for FE:</p>\n<p><a href=\"https://towardsdatascience.com/automated-feature-engineering-using-neural-networks-5310d6d4280a#:~:text=The%20Concept,certain%20combinations%20by%20engineering%20them.\" target=\"_blank\">Automated Feature Engineering Using Neural Networks</a><br>\n<a href=\"https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716\" target=\"_blank\">Why You Should Always Use Feature Embeddings With Structured Datasets</a></p>",
      "rawMarkdown": "As everyone suggested the first step is FE and the last step is stratified k-fold grid search.\nHere are two interesting papers for FE:\n\n[Automated Feature Engineering Using Neural Networks](https://towardsdatascience.com/automated-feature-engineering-using-neural-networks-5310d6d4280a#:~:text=The%20Concept,certain%20combinations%20by%20engineering%20them.)\n[Why You Should Always Use Feature Embeddings With Structured Datasets](https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716)\n",
      "votes": 3,
      "replies": [
        {
          "id": 1846415,
          "postDate": "2022-07-07T04:26:04.597Z",
          "content": "<p>There's recursive feature elimination (RFE) as well. <br>\nGood share with the links 👍</p>",
          "rawMarkdown": "There's recursive feature elimination (RFE) as well. \nGood share with the links 👍",
          "votes": 2
        }
      ]
    },
    {
      "id": 1837870,
      "postDate": "2022-06-30T02:35:39.460Z",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a>, I usually reserve hyperparam tunning as the last step of my workflow; I feel that feature engineering should receive all my focus on the first iterations </p>",
      "rawMarkdown": "Hello @manwithaflower, I usually reserve hyperparam tunning as the last step of my workflow; I feel that feature engineering should receive all my focus on the first iterations ",
      "votes": 3
    },
    {
      "id": 1838747,
      "postDate": "2022-06-30T19:09:36.867Z",
      "content": "<p>Hi o/<br>\nHyperparameter tuning is the last step in the process, so if you've added a or modified a feature, then you should redo the entire tuning from start to get the best benefit from new feature.</p>\n<p>I usually create a baseline model with increased n_estimators to 500-4000 on CPU, or 5000-10000 on GPU - in the initial stage you want the model to train quickly so you can get the results faster.</p>\n<p>If you want to know how is your new feature performing then train the original data on a model with preset random_state, then add the feature, use same random_state and compare the results<br>\n<code>model = LGBMRegressor(random_state=RS, n_estimators=2000)</code></p>\n<p>To know how the feature compares to others you can use <code>model.feature_importances_</code><br>\nor even better use <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html\" target=\"_blank\">SequentialFeatureSelector</a> to sequentially add features and see which actually add value.</p>",
      "rawMarkdown": "Hi o/\nHyperparameter tuning is the last step in the process, so if you've added a or modified a feature, then you should redo the entire tuning from start to get the best benefit from new feature.\n\nI usually create a baseline model with increased n_estimators to 500-4000 on CPU, or 5000-10000 on GPU - in the initial stage you want the model to train quickly so you can get the results faster.\n\nIf you want to know how is your new feature performing then train the original data on a model with preset random_state, then add the feature, use same random_state and compare the results\n```model = LGBMRegressor(random_state=RS, n_estimators=2000)```\n\nTo know how the feature compares to others you can use ```model.feature_importances_```\nor even better use [SequentialFeatureSelector](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html) to sequentially add features and see which actually add value.\n\n",
      "votes": 4
    },
    {
      "id": 1839958,
      "postDate": "2022-07-01T20:27:20.843Z",
      "content": "<p>You should carry out feature engineering before model training almost always. I don't think one should tune parameters before feature engineering as this could cause improper model fitting. </p>",
      "rawMarkdown": "You should carry out feature engineering before model training almost always. I don't think one should tune parameters before feature engineering as this could cause improper model fitting. ",
      "votes": 2
    },
    {
      "id": 1837928,
      "postDate": "2022-06-30T04:06:52.997Z",
      "content": "<p>I like to build a tuning process for the model - run it once per week.  In general we are looking for 0.00X kind of leaps to win this competition.  I think for kaggle you need to balance model tuning and feature engineering.</p>",
      "rawMarkdown": "I like to build a tuning process for the model - run it once per week.  In general we are looking for 0.00X kind of leaps to win this competition.  I think for kaggle you need to balance model tuning and feature engineering.",
      "votes": 2
    },
    {
      "id": 1845123,
      "postDate": "2022-07-06T04:45:07.520Z",
      "content": "<p>Feature engineering first for sure!</p>\n<p>Here is my rough workflow for a project like this:</p>\n<ol>\n<li>Import and clean data</li>\n<li>Aggregate values by target index (in this case customer_ID)</li>\n<li>Exploratory data analysis (EDA)</li>\n<li>The EDA should then inform the feature engineering</li>\n<li>Train model to get baseline score and feature importance (return to feature engineering as necessary)</li>\n<li>Apply hyperparameter tuning across some search space -&gt; assess performance increase</li>\n<li>Assess whether the hyperparameter search space is worth increasing/decreasing based on computation time and the range of parameter values searched over. </li>\n</ol>",
      "rawMarkdown": "Feature engineering first for sure!\n\nHere is my rough workflow for a project like this:\n1. Import and clean data\n2. Aggregate values by target index (in this case customer_ID)\n3. Exploratory data analysis (EDA)\n4. The EDA should then inform the feature engineering\n5. Train model to get baseline score and feature importance (return to feature engineering as necessary)\n6. Apply hyperparameter tuning across some search space -> assess performance increase\n7. Assess whether the hyperparameter search space is worth increasing/decreasing based on computation time and the range of parameter values searched over. "
    },
    {
      "id": 1845977,
      "postDate": "2022-07-06T18:23:39.360Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1838791,
      "author_name": "The Devastator",
      "author_url": "",
      "post_date": "2022-06-30T20:18:12.500000",
      "content": "<p>Hyperparameters optimization should be the last step, if not, you might overfit your validation set (or multiple K validation sets).</p>\n<p>This is what I normally do:</p>\n<ul>\n<li>Select model(s) and use the base parameters (without tuning).</li>\n<li>Select features and decide what features you want to use and how to extract them (Bin, One-hot, Frequency, Normalization, etc…).</li>\n<li>Train and Validate with each Feature selection and finally choose the best feature set.</li>\n<li>Tune hyperparameters for best feature set</li>\n</ul>",
      "votes": 9,
      "replies": [
        {
          "id": 1838878,
          "author_name": "valindor",
          "author_url": "",
          "post_date": "2022-06-30T23:23:07.633000",
          "content": "<p>Wouldnt it be better to HPO first, so that the model is familiar with the type of data you are using? And then doing feature selection? Also is there any where we can see that HPO then FS is overfitting?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1838882,
          "author_name": "Man of the year",
          "author_url": "",
          "post_date": "2022-06-30T23:37:25.317000",
          "content": "<p>I think hyperparams optimization might be useless, if you later add or delete features. As I understood it's easier to check features usefulness on default params and later tune them, when you choose features. About overfitting validation I also do not really understood this, I mean, if you do cross validation on same hyperparams, you'll see what are the best for all folds, not just for the one of them. And then you have public lb score to test final quality and see, if you overfitted due cross validation or not.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1838081,
      "author_name": "Ali Abdin",
      "author_url": "",
      "post_date": "2022-06-30T07:36:58.490000",
      "content": "<p>Changed or added features definetly affect the model directly by changing the values or adding new dimensions to the feature space, thats why you should tune the parameters after you have done your feature engineering.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 1838910,
      "author_name": "1110Ra",
      "author_url": "",
      "post_date": "2022-07-01T02:01:18.850000",
      "content": "<p>As everyone suggested the first step is FE and the last step is stratified k-fold grid search.<br>\nHere are two interesting papers for FE:</p>\n<p><a href=\"https://towardsdatascience.com/automated-feature-engineering-using-neural-networks-5310d6d4280a#:~:text=The%20Concept,certain%20combinations%20by%20engineering%20them.\" target=\"_blank\">Automated Feature Engineering Using Neural Networks</a><br>\n<a href=\"https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716\" target=\"_blank\">Why You Should Always Use Feature Embeddings With Structured Datasets</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 1846415,
          "author_name": "Ranja Sarkar (She/Her)",
          "author_url": "",
          "post_date": "2022-07-07T04:26:04.597000",
          "content": "<p>There's recursive feature elimination (RFE) as well. <br>\nGood share with the links 👍</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 1837870,
      "author_name": "C4rl05/V",
      "author_url": "",
      "post_date": "2022-06-30T02:35:39.460000",
      "content": "<p>Hello <a href=\"https://www.kaggle.com/manwithaflower\" target=\"_blank\">@manwithaflower</a>, I usually reserve hyperparam tunning as the last step of my workflow; I feel that feature engineering should receive all my focus on the first iterations </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1838747,
      "author_name": "Waldemar",
      "author_url": "",
      "post_date": "2022-06-30T19:09:36.867000",
      "content": "<p>Hi o/<br>\nHyperparameter tuning is the last step in the process, so if you've added a or modified a feature, then you should redo the entire tuning from start to get the best benefit from new feature.</p>\n<p>I usually create a baseline model with increased n_estimators to 500-4000 on CPU, or 5000-10000 on GPU - in the initial stage you want the model to train quickly so you can get the results faster.</p>\n<p>If you want to know how is your new feature performing then train the original data on a model with preset random_state, then add the feature, use same random_state and compare the results<br>\n<code>model = LGBMRegressor(random_state=RS, n_estimators=2000)</code></p>\n<p>To know how the feature compares to others you can use <code>model.feature_importances_</code><br>\nor even better use <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html\" target=\"_blank\">SequentialFeatureSelector</a> to sequentially add features and see which actually add value.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1839958,
      "author_name": "Ravi Ramakrishnan",
      "author_url": "",
      "post_date": "2022-07-01T20:27:20.843000",
      "content": "<p>You should carry out feature engineering before model training almost always. I don't think one should tune parameters before feature engineering as this could cause improper model fitting. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1837928,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2022-06-30T04:06:52.997000",
      "content": "<p>I like to build a tuning process for the model - run it once per week.  In general we are looking for 0.00X kind of leaps to win this competition.  I think for kaggle you need to balance model tuning and feature engineering.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1845123,
      "author_name": "Xavier R Nogueira",
      "author_url": "",
      "post_date": "2022-07-06T04:45:07.520000",
      "content": "<p>Feature engineering first for sure!</p>\n<p>Here is my rough workflow for a project like this:</p>\n<ol>\n<li>Import and clean data</li>\n<li>Aggregate values by target index (in this case customer_ID)</li>\n<li>Exploratory data analysis (EDA)</li>\n<li>The EDA should then inform the feature engineering</li>\n<li>Train model to get baseline score and feature importance (return to feature engineering as necessary)</li>\n<li>Apply hyperparameter tuning across some search space -&gt; assess performance increase</li>\n<li>Assess whether the hyperparameter search space is worth increasing/decreasing based on computation time and the range of parameter values searched over. </li>\n</ol>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1845977,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-07-06T18:23:39.360000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1837818": "I had this question. Let's say one use LGBM without any feature engineering, there are many hyperparameter combinations and one finds the optimal params for best performance. Then let's say there is some feature engineering, some magic features that definitely improves score. Will it be necessary to tune hyperparameters again after adding this feature to find new optimal ones?\n\nAlso similar question. Let's say we have some set of hyperparameters and we want to test a new feature, is it helpful or no. Can it be so that on this exact set of hyperparams quality will decrease, however, this feature is very helpful and on some different set of hyperparams it might increase quality significantly? If it is so, how should new features or tricks be tested?",
    "1838791": "Hyperparameters optimization should be the last step, if not, you might overfit your validation set (or multiple K validation sets).\n\nThis is what I normally do:\n\n- Select model(s) and use the base parameters (without tuning).\n- Select features and decide what features you want to use and how to extract them (Bin, One-hot, Frequency, Normalization, etc...).\n- Train and Validate with each Feature selection and finally choose the best feature set.\n- Tune hyperparameters for best feature set\n",
    "1838081": "Changed or added features definetly affect the model directly by changing the values or adding new dimensions to the feature space, thats why you should tune the parameters after you have done your feature engineering.",
    "1838910": "As everyone suggested the first step is FE and the last step is stratified k-fold grid search.\nHere are two interesting papers for FE:\n\n[Automated Feature Engineering Using Neural Networks](https://towardsdatascience.com/automated-feature-engineering-using-neural-networks-5310d6d4280a#:~:text=The%20Concept,certain%20combinations%20by%20engineering%20them.)\n[Why You Should Always Use Feature Embeddings With Structured Datasets](https://towardsdatascience.com/why-you-should-always-use-feature-embeddings-with-structured-datasets-7f280b40e716)\n",
    "1837870": "Hello @manwithaflower, I usually reserve hyperparam tunning as the last step of my workflow; I feel that feature engineering should receive all my focus on the first iterations ",
    "1838747": "Hi o/\nHyperparameter tuning is the last step in the process, so if you've added a or modified a feature, then you should redo the entire tuning from start to get the best benefit from new feature.\n\nI usually create a baseline model with increased n_estimators to 500-4000 on CPU, or 5000-10000 on GPU - in the initial stage you want the model to train quickly so you can get the results faster.\n\nIf you want to know how is your new feature performing then train the original data on a model with preset random_state, then add the feature, use same random_state and compare the results\n```model = LGBMRegressor(random_state=RS, n_estimators=2000)```\n\nTo know how the feature compares to others you can use ```model.feature_importances_```\nor even better use [SequentialFeatureSelector](https://scikit-learn.org/stable/modules/generated/sklearn.feature_selection.SequentialFeatureSelector.html) to sequentially add features and see which actually add value.\n\n",
    "1839958": "You should carry out feature engineering before model training almost always. I don't think one should tune parameters before feature engineering as this could cause improper model fitting. ",
    "1837928": "I like to build a tuning process for the model - run it once per week.  In general we are looking for 0.00X kind of leaps to win this competition.  I think for kaggle you need to balance model tuning and feature engineering.",
    "1845123": "Feature engineering first for sure!\n\nHere is my rough workflow for a project like this:\n1. Import and clean data\n2. Aggregate values by target index (in this case customer_ID)\n3. Exploratory data analysis (EDA)\n4. The EDA should then inform the feature engineering\n5. Train model to get baseline score and feature importance (return to feature engineering as necessary)\n6. Apply hyperparameter tuning across some search space -> assess performance increase\n7. Assess whether the hyperparameter search space is worth increasing/decreasing based on computation time and the range of parameter values searched over. ",
    "1845977": ""
  }
}