{
  "id": 71272,
  "title": "anyone using PCA to reduce correlated features?",
  "url": "/competitions/PLAsTiCC-2018/discussion/71272",
  "author_name": "",
  "post_date": "2018-11-12T05:41:46.045271400Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "419524",
      "postDate": "11/12/2018 05:41:46",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "419556",
      "postDate": "11/12/2018 07:28:04",
      "content": "<p>I tried but it gives me worse results.</p>",
      "rawMarkdown": "I tried but it gives me worse results.",
      "votes": null
    },
    {
      "id": "419597",
      "postDate": "11/12/2018 09:24:08",
      "content": "<p>same for me</p>",
      "rawMarkdown": "same for me",
      "votes": null
    },
    {
      "id": "419603",
      "postDate": "11/12/2018 09:29:46",
      "content": "<p>Highly correlated features are not an issue for gbms (eg xgboost, lightgbm, catboost, gbdt).  Lots of the advice you get in literature about feature engineering and feature selection is motivated by general linear models, and often they are detrimental when using gbms.  For instance, although it is not applicable here, one hot encoding of categorical variables is usually a bad idea when using gbms.  Same for feature normalization.  And so on.</p>",
      "rawMarkdown": "Highly correlated features are not an issue for gbms (eg xgboost, lightgbm, catboost, gbdt).  Lots of the advice you get in literature about feature engineering and feature selection is motivated by general linear models, and often they are detrimental when using gbms.  For instance, although it is not applicable here, one hot encoding of categorical variables is usually a bad idea when using gbms.  Same for feature normalization.  And so on.",
      "votes": null
    },
    {
      "id": "419607",
      "postDate": "11/12/2018 09:36:52",
      "content": "<p>thanks for the clarification. this explains why when i normalized my features the results for lightgbm went bad but the nn performed well</p>",
      "rawMarkdown": "thanks for the clarification. this explains why when i normalized my features the results for lightgbm went bad but the nn performed well",
      "votes": null
    },
    {
      "id": "419611",
      "postDate": "11/12/2018 09:40:41",
      "content": "<p>Exactly, each model class has its own preferred feature engineering.  Adding noise, one hot encoding, and normalizing are very useful for NNs.</p>",
      "rawMarkdown": "Exactly, each model class has its own preferred feature engineering.  Adding noise, one hot encoding, and normalizing are very useful for NNs.",
      "votes": null
    },
    {
      "id": "419833",
      "postDate": "11/12/2018 16:00:36",
      "content": "<p><a href=\"/niclasdoce\">@niclasdoce</a> <code>this explains why when i normalized my features the results for lightgbm went bad but the nn performed well</code> -- could you explain a little more, I am new to ML and did not get it</p>",
      "rawMarkdown": "niclasdoce ```this explains why when i normalized my features the results for lightgbm went bad but the nn performed well ``` -- could you explain a little more, I am new to ML and did not get it",
      "votes": null
    },
    {
      "id": "419973",
      "postDate": "11/12/2018 21:25:40",
      "content": "<p>@Blonde, it means just as CPMP said. normalizing features will work on NN most of the time but not on gradient boosting</p>",
      "rawMarkdown": "Blonde, it means just as CPMP said. normalizing features will work on NN most of the time but not on gradient boosting",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 419556,
      "author_name": "taniaj",
      "author_url": "",
      "post_date": "11/12/2018 07:28:04",
      "content": "<p>I tried but it gives me worse results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 419597,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "11/12/2018 09:24:08",
          "content": "<p>same for me</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 419603,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "11/12/2018 09:29:46",
      "content": "<p>Highly correlated features are not an issue for gbms (eg xgboost, lightgbm, catboost, gbdt).  Lots of the advice you get in literature about feature engineering and feature selection is motivated by general linear models, and often they are detrimental when using gbms.  For instance, although it is not applicable here, one hot encoding of categorical variables is usually a bad idea when using gbms.  Same for feature normalization.  And so on.</p>",
      "votes": null,
      "replies": [
        {
          "id": 419607,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "11/12/2018 09:36:52",
          "content": "<p>thanks for the clarification. this explains why when i normalized my features the results for lightgbm went bad but the nn performed well</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419611,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/12/2018 09:40:41",
          "content": "<p>Exactly, each model class has its own preferred feature engineering.  Adding noise, one hot encoding, and normalizing are very useful for NNs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419833,
          "author_name": "blondinka",
          "author_url": "",
          "post_date": "11/12/2018 16:00:36",
          "content": "<p><a href=\"/niclasdoce\">@niclasdoce</a> <code>this explains why when i normalized my features the results for lightgbm went bad but the nn performed well</code> -- could you explain a little more, I am new to ML and did not get it</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419973,
          "author_name": "niclasdoce",
          "author_url": "",
          "post_date": "11/12/2018 21:25:40",
          "content": "<p>@Blonde, it means just as CPMP said. normalizing features will work on NN most of the time but not on gradient boosting</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "419524": "",
    "419556": "I tried but it gives me worse results.",
    "419597": "same for me",
    "419603": "Highly correlated features are not an issue for gbms (eg xgboost, lightgbm, catboost, gbdt).  Lots of the advice you get in literature about feature engineering and feature selection is motivated by general linear models, and often they are detrimental when using gbms.  For instance, although it is not applicable here, one hot encoding of categorical variables is usually a bad idea when using gbms.  Same for feature normalization.  And so on.",
    "419607": "thanks for the clarification. this explains why when i normalized my features the results for lightgbm went bad but the nn performed well",
    "419611": "Exactly, each model class has its own preferred feature engineering.  Adding noise, one hot encoding, and normalizing are very useful for NNs.",
    "419833": "niclasdoce ```this explains why when i normalized my features the results for lightgbm went bad but the nn performed well ``` -- could you explain a little more, I am new to ML and did not get it",
    "419973": "Blonde, it means just as CPMP said. normalizing features will work on NN most of the time but not on gradient boosting"
  },
  "source": "meta"
}