{
  "id": 55227,
  "title": "weird variable importance in lightGBM",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/55227",
  "author_name": "",
  "post_date": "2018-04-23T22:31:08.866175700Z",
  "votes": 3,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I notice sometimes I got very weird variable importance in lightGBM even with same feature (e.g. one variable get 99.99% gain), but the performance still looks reasonable. Has anyone experienced similar issue? </p>",
  "messages": [
    {
      "id": "318474",
      "postDate": "04/23/2018 22:31:08",
      "content": "<p>I notice sometimes I got very weird variable importance in lightGBM even with same feature (e.g. one variable get 99.99% gain), but the performance still looks reasonable. Has anyone experienced similar issue? </p>",
      "rawMarkdown": "I notice sometimes I got very weird variable importance in lightGBM even with same feature (e.g. one variable get 99.99% gain), but the performance still looks reasonable. Has anyone experienced similar issue?",
      "votes": null
    },
    {
      "id": "318489",
      "postDate": "04/23/2018 23:40:46",
      "content": "<p>Thanks for posting this @Cheng.  </p>\n\n<p>I can confirm this as I also have few LGB models of similar nature :)  They are scoring between 0.9806 - 0.9808 on LB but only one feature with almost 99.98% gain in feature imp.  <code>¯\\_(ツ)_/¯</code> </p>",
      "rawMarkdown": "Thanks for posting this @Cheng.  \n\nI can confirm this as I also have few LGB models of similar nature :)  They are scoring between 0.9806 - 0.9808 on LB but only one feature with almost 99.98% gain in feature imp.  `¯\\_(ツ)_/¯`",
      "votes": null
    },
    {
      "id": "318490",
      "postDate": "04/23/2018 23:44:00",
      "content": "<p>When I first saw it, I was super excited because I thought I could achieve 98XX with one feature =)</p>",
      "rawMarkdown": "When I first saw it, I was super excited because I thought I could achieve 98XX with one feature =)",
      "votes": null
    },
    {
      "id": "318928",
      "postDate": "04/24/2018 19:18:33",
      "content": "<p>Did you try this, Cheng? Can one basic feature indeed show that much importance and generalize well?</p>\n\n<p>In models I've seen, feature dominance is much less pronounced - e.g. categorical variable 'app' often dominates up to 80% of total gain, which I interpret as 'some apps are just more popular for downloads, all other signals being equal'.</p>",
      "rawMarkdown": "Did you try this, Cheng? Can one basic feature indeed show that much importance and generalize well?\n\nIn models I've seen, feature dominance is much less pronounced - e.g. categorical variable 'app' often dominates up to 80% of total gain, which I interpret as 'some apps are just more popular for downloads, all other signals being equal'.",
      "votes": null
    },
    {
      "id": "318931",
      "postDate": "04/24/2018 19:32:06",
      "content": "<p>From my understanding it is impossible for one variable to take 100% gain (in my strange results)... I also observed dominance of app and I agreed with your interpretation =)</p>",
      "rawMarkdown": "From my understanding it is impossible for one variable to take 100% gain (in my strange results)... I also observed dominance of app and I agreed with your interpretation =)",
      "votes": null
    },
    {
      "id": "318976",
      "postDate": "04/25/2018 00:14:44",
      "content": "<p>I have separately noticed that total gain may be significantly different between scenarios when very different variables are used, yet very similar scores are produced.</p>\n\n<p>For example, when you take output of GBM model and feed it as a single variable into second stage GBM model of the same type, the second stage model is expected to be trivial (y = x), yet GBM still needs to work a bit to setup this trivial model as a combination of multiple trees. In this setup, I observed that total gain is much smaller when you give it one variable, yet AUROC and presumably output is essentially the same.</p>\n\n<p>Based on the above observation, I could raise a hypothesis: </p>\n\n<p><em>\"High variable importance for some models could be attributed to back-and-forth adjustments of the model scores via large number of splits in the GBM model.\"</em></p>\n\n<p>Please let me know if this rings a bell, or if your data supports it.</p>\n\n<p>P.S. After reading up on definition of variable importance in lightgbm, I would be surprised if my above hypothesis would be correct - perhaps alternatively in the test above, AUROC is similar yet output is quite different (e.g. I would get more narrow distribution in the case of using one variable in the second-stage model).</p>",
      "rawMarkdown": "I have separately noticed that total gain may be significantly different between scenarios when very different variables are used, yet very similar scores are produced.\n\nFor example, when you take output of GBM model and feed it as a single variable into second stage GBM model of the same type, the second stage model is expected to be trivial (y = x), yet GBM still needs to work a bit to setup this trivial model as a combination of multiple trees. In this setup, I observed that total gain is much smaller when you give it one variable, yet AUROC and presumably output is essentially the same.\n\nBased on the above observation, I could raise a hypothesis: \n\n*\"High variable importance for some models could be attributed to back-and-forth adjustments of the model scores via large number of splits in the GBM model.\"*\n\nPlease let me know if this rings a bell, or if your data supports it.\n\nP.S. After reading up on definition of variable importance in lightgbm, I would be surprised if my above hypothesis would be correct - perhaps alternatively in the test above, AUROC is similar yet output is quite different (e.g. I would get more narrow distribution in the case of using one variable in the second-stage model).",
      "votes": null
    },
    {
      "id": "320178",
      "postDate": "04/27/2018 19:02:36",
      "content": "<p>Does this happen only for features that are fundamentally relying on target variable information (e.g. aggregated on relatively small domain)? \nThis would explain their significance, at the same time I presume you would not be able to re-create them consistently on the test set.</p>",
      "rawMarkdown": "Does this happen only for features that are fundamentally relying on target variable information (e.g. aggregated on relatively small domain)? \nThis would explain their significance, at the same time I presume you would not be able to re-create them consistently on the test set.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 318489,
      "author_name": "pranav84",
      "author_url": "",
      "post_date": "04/23/2018 23:40:46",
      "content": "<p>Thanks for posting this @Cheng.  </p>\n\n<p>I can confirm this as I also have few LGB models of similar nature :)  They are scoring between 0.9806 - 0.9808 on LB but only one feature with almost 99.98% gain in feature imp.  <code>¯\\_(ツ)_/¯</code> </p>",
      "votes": null,
      "replies": [
        {
          "id": 318490,
          "author_name": "chengju",
          "author_url": "",
          "post_date": "04/23/2018 23:44:00",
          "content": "<p>When I first saw it, I was super excited because I thought I could achieve 98XX with one feature =)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318928,
          "author_name": "pgyrya",
          "author_url": "",
          "post_date": "04/24/2018 19:18:33",
          "content": "<p>Did you try this, Cheng? Can one basic feature indeed show that much importance and generalize well?</p>\n\n<p>In models I've seen, feature dominance is much less pronounced - e.g. categorical variable 'app' often dominates up to 80% of total gain, which I interpret as 'some apps are just more popular for downloads, all other signals being equal'.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 318931,
          "author_name": "chengju",
          "author_url": "",
          "post_date": "04/24/2018 19:32:06",
          "content": "<p>From my understanding it is impossible for one variable to take 100% gain (in my strange results)... I also observed dominance of app and I agreed with your interpretation =)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 318976,
      "author_name": "pgyrya",
      "author_url": "",
      "post_date": "04/25/2018 00:14:44",
      "content": "<p>I have separately noticed that total gain may be significantly different between scenarios when very different variables are used, yet very similar scores are produced.</p>\n\n<p>For example, when you take output of GBM model and feed it as a single variable into second stage GBM model of the same type, the second stage model is expected to be trivial (y = x), yet GBM still needs to work a bit to setup this trivial model as a combination of multiple trees. In this setup, I observed that total gain is much smaller when you give it one variable, yet AUROC and presumably output is essentially the same.</p>\n\n<p>Based on the above observation, I could raise a hypothesis: </p>\n\n<p><em>\"High variable importance for some models could be attributed to back-and-forth adjustments of the model scores via large number of splits in the GBM model.\"</em></p>\n\n<p>Please let me know if this rings a bell, or if your data supports it.</p>\n\n<p>P.S. After reading up on definition of variable importance in lightgbm, I would be surprised if my above hypothesis would be correct - perhaps alternatively in the test above, AUROC is similar yet output is quite different (e.g. I would get more narrow distribution in the case of using one variable in the second-stage model).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 320178,
      "author_name": "pgyrya",
      "author_url": "",
      "post_date": "04/27/2018 19:02:36",
      "content": "<p>Does this happen only for features that are fundamentally relying on target variable information (e.g. aggregated on relatively small domain)? \nThis would explain their significance, at the same time I presume you would not be able to re-create them consistently on the test set.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "318474": "I notice sometimes I got very weird variable importance in lightGBM even with same feature (e.g. one variable get 99.99% gain), but the performance still looks reasonable. Has anyone experienced similar issue?",
    "318489": "Thanks for posting this @Cheng.  \n\nI can confirm this as I also have few LGB models of similar nature :)  They are scoring between 0.9806 - 0.9808 on LB but only one feature with almost 99.98% gain in feature imp.  `¯\\_(ツ)_/¯`",
    "318490": "When I first saw it, I was super excited because I thought I could achieve 98XX with one feature =)",
    "318928": "Did you try this, Cheng? Can one basic feature indeed show that much importance and generalize well?\n\nIn models I've seen, feature dominance is much less pronounced - e.g. categorical variable 'app' often dominates up to 80% of total gain, which I interpret as 'some apps are just more popular for downloads, all other signals being equal'.",
    "318931": "From my understanding it is impossible for one variable to take 100% gain (in my strange results)... I also observed dominance of app and I agreed with your interpretation =)",
    "318976": "I have separately noticed that total gain may be significantly different between scenarios when very different variables are used, yet very similar scores are produced.\n\nFor example, when you take output of GBM model and feed it as a single variable into second stage GBM model of the same type, the second stage model is expected to be trivial (y = x), yet GBM still needs to work a bit to setup this trivial model as a combination of multiple trees. In this setup, I observed that total gain is much smaller when you give it one variable, yet AUROC and presumably output is essentially the same.\n\nBased on the above observation, I could raise a hypothesis: \n\n*\"High variable importance for some models could be attributed to back-and-forth adjustments of the model scores via large number of splits in the GBM model.\"*\n\nPlease let me know if this rings a bell, or if your data supports it.\n\nP.S. After reading up on definition of variable importance in lightgbm, I would be surprised if my above hypothesis would be correct - perhaps alternatively in the test above, AUROC is similar yet output is quite different (e.g. I would get more narrow distribution in the case of using one variable in the second-stage model).",
    "320178": "Does this happen only for features that are fundamentally relying on target variable information (e.g. aggregated on relatively small domain)? \nThis would explain their significance, at the same time I presume you would not be able to re-create them consistently on the test set."
  },
  "source": "meta"
}