{
  "id": 330381,
  "title": "How to do feature engineering  and adjust parameters?",
  "url": "/competitions/amex-default-prediction/discussion/330381",
  "author_name": "",
  "post_date": "2022-06-12T02:44:13.932499900Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hey guy, Can you share some insights about feature engineering and adjust parameters in your train models?<br>\nl saw some tips about feature engineering like:<br>\nnum features will be enlarge to mean, std, min, max, last…<br>\ncat features will be enlarge to unique, count, last…</p>\n<p>Do you have some margin ways to deal will feature engineering? </p>\n<p>Do you select model's parameters by your experience or your train?</p>\n<p>l hope your guys can share some insights and thanks a lot~~~</p>",
  "messages": [
    {
      "id": "1817905",
      "postDate": "06/12/2022 02:44:13",
      "content": "<p>Hey guy, Can you share some insights about feature engineering and adjust parameters in your train models?<br>\nl saw some tips about feature engineering like:<br>\nnum features will be enlarge to mean, std, min, max, last…<br>\ncat features will be enlarge to unique, count, last…</p>\n<p>Do you have some margin ways to deal will feature engineering? </p>\n<p>Do you select model's parameters by your experience or your train?</p>\n<p>l hope your guys can share some insights and thanks a lot~~~</p>",
      "rawMarkdown": "Hey guy, Can you share some insights about feature engineering and adjust parameters in your train models?\nl saw some tips about feature engineering like:\nnum features will be enlarge to mean, std, min, max, last...\ncat features will be enlarge to unique, count, last...\n\nDo you have some margin ways to deal will feature engineering? \n\nDo you select model's parameters by your experience or your train?\n\nl hope your guys can share some insights and thanks a lot~~~",
      "votes": null
    },
    {
      "id": "1817957",
      "postDate": "06/12/2022 04:26:20",
      "content": "<p>Most efficient way to find the right parameters for your model is <em>Hyperparameter Tuning</em>, but they can take quite a lot of time find them for you. So you might want to explore different algorithms to do so. For sklearn you can find bunch of them <a href=\"https://scikit-learn.org/stable/modules/grid_search.html#tuning-the-hyper-parameters-of-an-estimator\" target=\"_blank\">here</a></p>\n<p>Feature engineering is quite subjective to the task you are working with, but you could do something like:</p>\n<ul>\n<li>Imputation: Where you fill the missing values in your dataset. There are lot's of methods to it!</li>\n<li>Categorical Encoding: This is also important because you can't just <em>Hot-Encode</em> or <em>Label-Encode</em> everything. (Read more <a href=\"https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor\" target=\"_blank\">here</a>)</li>\n<li>Handling outliers</li>\n<li>Creating new features (PCA can help with this)</li>\n<li>etc.<br>\nI learned a lot of <em>Feature Engineering</em> techniques by reading people's notebooks.</li>\n</ul>",
      "rawMarkdown": "Most efficient way to find the right parameters for your model is *Hyperparameter Tuning*, but they can take quite a lot of time find them for you. So you might want to explore different algorithms to do so. For sklearn you can find bunch of them [here](https://scikit-learn.org/stable/modules/grid_search.html#tuning-the-hyper-parameters-of-an-estimator)\n\nFeature engineering is quite subjective to the task you are working with, but you could do something like:\n- Imputation: Where you fill the missing values in your dataset. There are lot's of methods to it!\n- Categorical Encoding: This is also important because you can't just *Hot-Encode* or *Label-Encode* everything. (Read more [here](https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor))\n- Handling outliers\n- Creating new features (PCA can help with this)\n- etc.\nI learned a lot of *Feature Engineering* techniques by reading people's notebooks.",
      "votes": null
    },
    {
      "id": "1818026",
      "postDate": "06/12/2022 06:32:11",
      "content": "<p>Hi, how to use PCA to create new features?</p>",
      "rawMarkdown": "Hi, how to use PCA to create new features?",
      "votes": null
    },
    {
      "id": "1818087",
      "postDate": "06/12/2022 08:22:20",
      "content": "<p>This is good question! And l saw some answers from <a href=\"https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor/40908#40908?newreg=76dcabdedb1e4ac1beac4fa7e989f43\" target=\"_blank\">here</a><br>\n(from <em>Keivan loch Hagh</em> share~).</p>\n<p>Using PCA because of curse of dimensionality by <em>one-hot Encoding.</em></p>\n<p>So, when you use <em>one-hot Encoding</em> or your dataset have too much features, maybe you can take PCA in your consideration. </p>",
      "rawMarkdown": "This is good question! And l saw some answers from [here](https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor/40908#40908?newreg=76dcabdedb1e4ac1beac4fa7e989f43)\n(from *Keivan loch Hagh* share~).\n\nUsing PCA because of curse of dimensionality by *one-hot Encoding.*\n\nSo, when you use *one-hot Encoding* or your dataset have too much features, maybe you can take PCA in your consideration.",
      "votes": null
    },
    {
      "id": "1818089",
      "postDate": "06/12/2022 08:24:37",
      "content": "<p>l learn much in your sharing~  Thanks a lot👍</p>\n<p>Is tree models such as <em>Gradient Boosted Decision Trees</em> and <em>Random Forests</em> have good functions to deal with original categorical features? And Do l need to use <em>one-hot Encoding</em> to deal with categorical features?</p>",
      "rawMarkdown": "l learn much in your sharing~  Thanks a lot👍\n\nIs tree models such as *Gradient Boosted Decision Trees* and *Random Forests* have good functions to deal with original categorical features? And Do l need to use *one-hot Encoding* to deal with categorical features?",
      "votes": null
    },
    {
      "id": "1818241",
      "postDate": "06/12/2022 12:23:59",
      "content": "<p>I have to say that I haven't modified any hyperparameters.<br>\nRegardless of the model, the default parameters are good enough~ (Just for now) 😆</p>",
      "rawMarkdown": "I have to say that I haven't modified any hyperparameters.\nRegardless of the model, the default parameters are good enough~ (Just for now) 😆",
      "votes": null
    },
    {
      "id": "1818617",
      "postDate": "06/13/2022 01:14:41",
      "content": "<p>thanks a lot~<br>\nAnd what about feature engineering?👀</p>",
      "rawMarkdown": "thanks a lot~\nAnd what about feature engineering?👀",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1817957,
      "author_name": "keivanipchihagh",
      "author_url": "",
      "post_date": "06/12/2022 04:26:20",
      "content": "<p>Most efficient way to find the right parameters for your model is <em>Hyperparameter Tuning</em>, but they can take quite a lot of time find them for you. So you might want to explore different algorithms to do so. For sklearn you can find bunch of them <a href=\"https://scikit-learn.org/stable/modules/grid_search.html#tuning-the-hyper-parameters-of-an-estimator\" target=\"_blank\">here</a></p>\n<p>Feature engineering is quite subjective to the task you are working with, but you could do something like:</p>\n<ul>\n<li>Imputation: Where you fill the missing values in your dataset. There are lot's of methods to it!</li>\n<li>Categorical Encoding: This is also important because you can't just <em>Hot-Encode</em> or <em>Label-Encode</em> everything. (Read more <a href=\"https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor\" target=\"_blank\">here</a>)</li>\n<li>Handling outliers</li>\n<li>Creating new features (PCA can help with this)</li>\n<li>etc.<br>\nI learned a lot of <em>Feature Engineering</em> techniques by reading people's notebooks.</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 1818026,
          "author_name": "denisplaj",
          "author_url": "",
          "post_date": "06/12/2022 06:32:11",
          "content": "<p>Hi, how to use PCA to create new features?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818087,
          "author_name": "shanggangli",
          "author_url": "",
          "post_date": "06/12/2022 08:22:20",
          "content": "<p>This is good question! And l saw some answers from <a href=\"https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor/40908#40908?newreg=76dcabdedb1e4ac1beac4fa7e989f43\" target=\"_blank\">here</a><br>\n(from <em>Keivan loch Hagh</em> share~).</p>\n<p>Using PCA because of curse of dimensionality by <em>one-hot Encoding.</em></p>\n<p>So, when you use <em>one-hot Encoding</em> or your dataset have too much features, maybe you can take PCA in your consideration. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1818089,
          "author_name": "shanggangli",
          "author_url": "",
          "post_date": "06/12/2022 08:24:37",
          "content": "<p>l learn much in your sharing~  Thanks a lot👍</p>\n<p>Is tree models such as <em>Gradient Boosted Decision Trees</em> and <em>Random Forests</em> have good functions to deal with original categorical features? And Do l need to use <em>one-hot Encoding</em> to deal with categorical features?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1818241,
      "author_name": "xiaowangiiiii",
      "author_url": "",
      "post_date": "06/12/2022 12:23:59",
      "content": "<p>I have to say that I haven't modified any hyperparameters.<br>\nRegardless of the model, the default parameters are good enough~ (Just for now) 😆</p>",
      "votes": null,
      "replies": [
        {
          "id": 1818617,
          "author_name": "shanggangli",
          "author_url": "",
          "post_date": "06/13/2022 01:14:41",
          "content": "<p>thanks a lot~<br>\nAnd what about feature engineering?👀</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1817905": "Hey guy, Can you share some insights about feature engineering and adjust parameters in your train models?\nl saw some tips about feature engineering like:\nnum features will be enlarge to mean, std, min, max, last...\ncat features will be enlarge to unique, count, last...\n\nDo you have some margin ways to deal will feature engineering? \n\nDo you select model's parameters by your experience or your train?\n\nl hope your guys can share some insights and thanks a lot~~~",
    "1817957": "Most efficient way to find the right parameters for your model is *Hyperparameter Tuning*, but they can take quite a lot of time find them for you. So you might want to explore different algorithms to do so. For sklearn you can find bunch of them [here](https://scikit-learn.org/stable/modules/grid_search.html#tuning-the-hyper-parameters-of-an-estimator)\n\nFeature engineering is quite subjective to the task you are working with, but you could do something like:\n- Imputation: Where you fill the missing values in your dataset. There are lot's of methods to it!\n- Categorical Encoding: This is also important because you can't just *Hot-Encode* or *Label-Encode* everything. (Read more [here](https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor))\n- Handling outliers\n- Creating new features (PCA can help with this)\n- etc.\nI learned a lot of *Feature Engineering* techniques by reading people's notebooks.",
    "1818026": "Hi, how to use PCA to create new features?",
    "1818087": "This is good question! And l saw some answers from [here](https://datascience.stackexchange.com/questions/9443/when-to-use-one-hot-encoding-vs-labelencoder-vs-dictvectorizor/40908#40908?newreg=76dcabdedb1e4ac1beac4fa7e989f43)\n(from *Keivan loch Hagh* share~).\n\nUsing PCA because of curse of dimensionality by *one-hot Encoding.*\n\nSo, when you use *one-hot Encoding* or your dataset have too much features, maybe you can take PCA in your consideration.",
    "1818089": "l learn much in your sharing~  Thanks a lot👍\n\nIs tree models such as *Gradient Boosted Decision Trees* and *Random Forests* have good functions to deal with original categorical features? And Do l need to use *one-hot Encoding* to deal with categorical features?",
    "1818241": "I have to say that I haven't modified any hyperparameters.\nRegardless of the model, the default parameters are good enough~ (Just for now) 😆",
    "1818617": "thanks a lot~\nAnd what about feature engineering?👀"
  },
  "source": "meta"
}