{
  "id": 55685,
  "title": "Lightgbm categorical features transformation",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55685",
  "author_name": "",
  "post_date": "2018-04-30T17:11:24.946896400Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>As part of the data preprocessing work, it is common to transform highly skewed features in order to reduce skewness in the data. </p>\n\n<p>For positively skewed features we commonly use transformations such as log or sqrt. I usually set a threshold on skewness of 0.75.</p>\n\n<p>Does this apply to lightgbm as well?</p>",
  "messages": [
    {
      "id": "321126",
      "postDate": "04/30/2018 17:11:24",
      "content": "<p>As part of the data preprocessing work, it is common to transform highly skewed features in order to reduce skewness in the data. </p>\n\n<p>For positively skewed features we commonly use transformations such as log or sqrt. I usually set a threshold on skewness of 0.75.</p>\n\n<p>Does this apply to lightgbm as well?</p>",
      "rawMarkdown": "As part of the data preprocessing work, it is common to transform highly skewed features in order to reduce skewness in the data. \n\nFor positively skewed features we commonly use transformations such as log or sqrt. I usually set a threshold on skewness of 0.75.\n\nDoes this apply to lightgbm as well?",
      "votes": null
    },
    {
      "id": "321133",
      "postDate": "04/30/2018 17:25:35",
      "content": "<p>It doesn't apply, which is actually one of the really nice things about the algorithm. Tree-based methods including gradient boosting are scale-independent and invariant to monotonic transformations of the features. Both log and sqrt are monotonic, so in theory this type of data preparation should have no impact on gradient boosting results.</p>\n\n<p>The reason this is true is that trees are grown by considering splits along <strong>sorted</strong> feature values. As long as a transformation does not change the sorting order (i.e. monotonic), the choices of splitting point and resulting separations of the data will have equivalent effects.     </p>",
      "rawMarkdown": "It doesn't apply, which is actually one of the really nice things about the algorithm. Tree-based methods including gradient boosting are scale-independent and invariant to monotonic transformations of the features. Both log and sqrt are monotonic, so in theory this type of data preparation should have no impact on gradient boosting results.\n\nThe reason this is true is that trees are grown by considering splits along **sorted** feature values. As long as a transformation does not change the sorting order (i.e. monotonic), the choices of splitting point and resulting separations of the data will have equivalent effects.",
      "votes": null
    },
    {
      "id": "321135",
      "postDate": "04/30/2018 17:34:54",
      "content": "<p>What Joe said :-). You might get a nice boost (or a bad drop?) in your score though if you bin on a transformation where you lose information e.g. <code>np.log2(1+feature).astype(np.uint32)</code> depending on if that feature was already over/under fitting.</p>",
      "rawMarkdown": "What Joe said :-). You might get a nice boost (or a bad drop?) in your score though if you bin on a transformation where you lose information e.g. `np.log2(1+feature).astype(np.uint32)` depending on if that feature was already over/under fitting.",
      "votes": null
    },
    {
      "id": "321138",
      "postDate": "04/30/2018 17:38:24",
      "content": "<p>Thanks I will try without transforming and check the accuracy.</p>",
      "rawMarkdown": "Thanks I will try without transforming and check the accuracy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 321133,
      "author_name": "aquatic",
      "author_url": "",
      "post_date": "04/30/2018 17:25:35",
      "content": "<p>It doesn't apply, which is actually one of the really nice things about the algorithm. Tree-based methods including gradient boosting are scale-independent and invariant to monotonic transformations of the features. Both log and sqrt are monotonic, so in theory this type of data preparation should have no impact on gradient boosting results.</p>\n\n<p>The reason this is true is that trees are grown by considering splits along <strong>sorted</strong> feature values. As long as a transformation does not change the sorting order (i.e. monotonic), the choices of splitting point and resulting separations of the data will have equivalent effects.     </p>",
      "votes": null,
      "replies": [
        {
          "id": 321135,
          "author_name": "authman",
          "author_url": "",
          "post_date": "04/30/2018 17:34:54",
          "content": "<p>What Joe said :-). You might get a nice boost (or a bad drop?) in your score though if you bin on a transformation where you lose information e.g. <code>np.log2(1+feature).astype(np.uint32)</code> depending on if that feature was already over/under fitting.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321138,
          "author_name": "rdekou",
          "author_url": "",
          "post_date": "04/30/2018 17:38:24",
          "content": "<p>Thanks I will try without transforming and check the accuracy.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "321126": "As part of the data preprocessing work, it is common to transform highly skewed features in order to reduce skewness in the data. \n\nFor positively skewed features we commonly use transformations such as log or sqrt. I usually set a threshold on skewness of 0.75.\n\nDoes this apply to lightgbm as well?",
    "321133": "It doesn't apply, which is actually one of the really nice things about the algorithm. Tree-based methods including gradient boosting are scale-independent and invariant to monotonic transformations of the features. Both log and sqrt are monotonic, so in theory this type of data preparation should have no impact on gradient boosting results.\n\nThe reason this is true is that trees are grown by considering splits along **sorted** feature values. As long as a transformation does not change the sorting order (i.e. monotonic), the choices of splitting point and resulting separations of the data will have equivalent effects.",
    "321135": "What Joe said :-). You might get a nice boost (or a bad drop?) in your score though if you bin on a transformation where you lose information e.g. `np.log2(1+feature).astype(np.uint32)` depending on if that feature was already over/under fitting.",
    "321138": "Thanks I will try without transforming and check the accuracy."
  },
  "source": "meta"
}