{
  "id": 55398,
  "title": "LightGBM Bosses",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55398",
  "author_name": "عثمان",
  "post_date": "2018-04-26T05:10:11.266000",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I'm trying to hit ~0.9805 before utilizing mean encoding features.</p>\n\n<p>Those of you rank 150 or above, is my goal realistic---what was your highest non-target-encoding-using single model scores?</p>\n\n<p>Also, how exactly does LightGBM deal with variables? Based on LGBM's <code>max_bin</code> parameter, as well as <a href=\"https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree.pdf\">the original abstract</a> released NIPS2017 (wuh? LightGBM has only been around that long--lol?) which has mention of histogram based splits:</p>\n\n<blockquote>\n  <p>histogram-based algorithm buckets continuous feature values into\n  discrete bins and uses these bins to construct feature histograms\n  during training...</p>\n</blockquote>\n\n<p>...my understanding is LightGBM doesn't care about variable type. Categorical variables are respected, unless their cardinality &gt; <code>max_bins</code>, and continuous variables are descritizes first. Then, either type of var is evaluated for the best split using lgbm's histogram based method. Is that accurate?? I find it fascinating if so, because that would mean the 'relationship' between high and low extremes of a continuous var are essentially treated as different categories.</p>\n\n<p>Finally, I took one of my model runs and <em>completely replaced</em> a categorical features, e.g. <code>os</code>, with it's mean encoded value. The resulting model scored within ~0.0005 on my val, compared to just leaving the categorical variable in directly. That got me thinking---these \"best\" DT splits the algo chooses are, after all, based upon the target we're trying to predict. Therefore, isn't using a var in lghm (especially categorical variables) somewhat similar to mean encoding them?</p>",
  "messages": [
    {
      "id": 319478,
      "postDate": "2018-04-26T06:33:29.947Z",
      "content": "<p>could I tell me what mean encode features are</p>",
      "rawMarkdown": "could I tell me what mean encode features are",
      "votes": 1,
      "replies": [
        {
          "id": 319490,
          "postDate": "2018-04-26T07:09:21.647Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 319465,
      "postDate": "2018-04-26T05:10:11.267Z",
      "content": "<p>I'm trying to hit ~0.9805 before utilizing mean encoding features.</p>\n\n<p>Those of you rank 150 or above, is my goal realistic---what was your highest non-target-encoding-using single model scores?</p>\n\n<p>Also, how exactly does LightGBM deal with variables? Based on LGBM's <code>max_bin</code> parameter, as well as <a href=\"https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree.pdf\">the original abstract</a> released NIPS2017 (wuh? LightGBM has only been around that long--lol?) which has mention of histogram based splits:</p>\n\n<blockquote>\n  <p>histogram-based algorithm buckets continuous feature values into\n  discrete bins and uses these bins to construct feature histograms\n  during training...</p>\n</blockquote>\n\n<p>...my understanding is LightGBM doesn't care about variable type. Categorical variables are respected, unless their cardinality &gt; <code>max_bins</code>, and continuous variables are descritizes first. Then, either type of var is evaluated for the best split using lgbm's histogram based method. Is that accurate?? I find it fascinating if so, because that would mean the 'relationship' between high and low extremes of a continuous var are essentially treated as different categories.</p>\n\n<p>Finally, I took one of my model runs and <em>completely replaced</em> a categorical features, e.g. <code>os</code>, with it's mean encoded value. The resulting model scored within ~0.0005 on my val, compared to just leaving the categorical variable in directly. That got me thinking---these \"best\" DT splits the algo chooses are, after all, based upon the target we're trying to predict. Therefore, isn't using a var in lghm (especially categorical variables) somewhat similar to mean encoding them?</p>",
      "rawMarkdown": "I'm trying to hit ~0.9805 before utilizing mean encoding features.\n\nThose of you rank 150 or above, is my goal realistic---what was your highest non-target-encoding-using single model scores?\n\nAlso, how exactly does LightGBM deal with variables? Based on LGBM's `max_bin` parameter, as well as [the original abstract][1] released NIPS2017 (wuh? LightGBM has only been around that long--lol?) which has mention of histogram based splits:\n\n&gt; histogram-based algorithm buckets continuous feature values into\n&gt; discrete bins and uses these bins to construct feature histograms\n&gt; during training...\n\n...my understanding is LightGBM doesn't care about variable type. Categorical variables are respected, unless their cardinality &gt; `max_bins`, and continuous variables are descritizes first. Then, either type of var is evaluated for the best split using lgbm's histogram based method. Is that accurate?? I find it fascinating if so, because that would mean the 'relationship' between high and low extremes of a continuous var are essentially treated as different categories.\n\nFinally, I took one of my model runs and *completely replaced* a categorical features, e.g. `os`, with it's mean encoded value. The resulting model scored within ~0.0005 on my val, compared to just leaving the categorical variable in directly. That got me thinking---these \"best\" DT splits the algo chooses are, after all, based upon the target we're trying to predict. Therefore, isn't using a var in lghm (especially categorical variables) somewhat similar to mean encoding them?\n\n  [1]: https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree.pdf",
      "votes": 2
    },
    {
      "id": 319534,
      "postDate": "2018-04-26T08:47:24.420Z",
      "content": "<p>0.9805 (for now :P), so I guess it s possible.</p>",
      "rawMarkdown": "0.9805 (for now :P), so I guess it s possible."
    },
    {
      "id": 319481,
      "postDate": "2018-04-26T06:48:26.220Z",
      "content": "<p>May I ask what do you mean by mean encoding?</p>",
      "rawMarkdown": "May I ask what do you mean by mean encoding?",
      "replies": [
        {
          "id": 319489,
          "postDate": "2018-04-26T07:07:45.760Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 319476,
      "postDate": "2018-04-26T06:29:00.310Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 319478,
      "author_name": "MengYe",
      "author_url": "",
      "post_date": "2018-04-26T06:33:29.947000",
      "content": "<p>could I tell me what mean encode features are</p>",
      "votes": 1,
      "replies": [
        {
          "id": 319490,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T07:09:21.647000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 319534,
      "author_name": "Yair Beer",
      "author_url": "",
      "post_date": "2018-04-26T08:47:24.420000",
      "content": "<p>0.9805 (for now :P), so I guess it s possible.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 319481,
      "author_name": "Sohaib Omar",
      "author_url": "",
      "post_date": "2018-04-26T06:48:26.220000",
      "content": "<p>May I ask what do you mean by mean encoding?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 319489,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-04-26T07:07:45.760000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 319476,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-04-26T06:29:00.310000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "319478": "could I tell me what mean encode features are",
    "319465": "I'm trying to hit ~0.9805 before utilizing mean encoding features.\n\nThose of you rank 150 or above, is my goal realistic---what was your highest non-target-encoding-using single model scores?\n\nAlso, how exactly does LightGBM deal with variables? Based on LGBM's `max_bin` parameter, as well as [the original abstract][1] released NIPS2017 (wuh? LightGBM has only been around that long--lol?) which has mention of histogram based splits:\n\n&gt; histogram-based algorithm buckets continuous feature values into\n&gt; discrete bins and uses these bins to construct feature histograms\n&gt; during training...\n\n...my understanding is LightGBM doesn't care about variable type. Categorical variables are respected, unless their cardinality &gt; `max_bins`, and continuous variables are descritizes first. Then, either type of var is evaluated for the best split using lgbm's histogram based method. Is that accurate?? I find it fascinating if so, because that would mean the 'relationship' between high and low extremes of a continuous var are essentially treated as different categories.\n\nFinally, I took one of my model runs and *completely replaced* a categorical features, e.g. `os`, with it's mean encoded value. The resulting model scored within ~0.0005 on my val, compared to just leaving the categorical variable in directly. That got me thinking---these \"best\" DT splits the algo chooses are, after all, based upon the target we're trying to predict. Therefore, isn't using a var in lghm (especially categorical variables) somewhat similar to mean encoding them?\n\n  [1]: https://papers.nips.cc/paper/6907-lightgbm-a-highly-efficient-gradient-boosting-decision-tree.pdf",
    "319534": "0.9805 (for now :P), so I guess it s possible.",
    "319481": "May I ask what do you mean by mean encoding?",
    "319476": ""
  }
}