{
  "id": 329094,
  "title": "Bypass the features anonymization with LightGBM's feature interactions",
  "url": "/competitions/amex-default-prediction/discussion/329094",
  "author_name": "",
  "post_date": "2022-06-04T18:28:47.498411300Z",
  "votes": 63,
  "comment_count": 5,
  "views": 0,
  "content": "<h2>Bypass the features anonymization using LightGBM's feature interactions</h2>\n<p>In this competition the features are anonymized, this makes it challenging to engineer meaningful features.</p>\n<p>However, we can still use feature interactions learned by a model (for example LightGBM) to find out meaningful feature pairs and use this information to later on handcraft features to make our models more robust.</p>\n<p><strong>How does this work?</strong></p>\n<p>Lightgbm can also be used to automatically detect useful feature interactions in a dataset and these can be then used to create valuable features.</p>\n<p>Using the splits of the trained LightGBM trees we can find pairs of features and generate new features that are very valuable for training the next models.</p>\n<p>The intuition behind this is that the features used in most of the trees on higher levels are better if they would be preprocessed before entering our models so that the models \"won't waste\" splits on them and the splits are more likely to be used in the next level.</p>\n<p><strong>Code Example</strong></p>\n<p>Code Example for plotting LightGBM feature interactions:</p>\n<p>I have merged some of the code from <a href=\"https://www.kaggle.com/vishalbajaj2000/santander-lightgbm-xgb-feature-interactions\" target=\"_blank\">here</a> (credit to the original author!) and modified it to be simpler.<br>\nThe usage is as simple as: <code>plot_feat_interaction(model_lgb)</code>. </p>\n<p>Hope this helps!</p>\n<pre><code>import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\ndef get_splits_gain(tree_num = 0, parent = -1, tree = None, lev = 0, node_name = None, split_gain = None, reclimit = 50000):\n    if tree == None:\n        raise Exception('No tree present to analyze!')\n    for k, v in tree.items():\n        if type(v) != dict and k in ['split_feature']:\n            old_parent = parent\n            parent = v\n            tag = k\n            yield tree_num, tag, old_parent, parent, lev, node_name, split_gain\n        elif isinstance(v, dict):\n            if v.get('split_gain') == None:\n                continue\n            else:\n                tree = v\n                lev_inc = lev + 1\n                node_name = k\n                split_gain = v['split_gain']\n                for result in get_splits_gain(tree_num, parent, tree, lev_inc, node_name, split_gain):\n                    yield result\n        else: continue\n\ndef plot_feat_interaction(model):\n    dumped_model = model.booster_.dump_model()\n    tree_info = []\n    for j in range(0, len(dumped_model['tree_info'])):\n        for i in get_splits_gain(tree_num = j, tree = dumped_model['tree_info'][j]):\n            tree_info.append(list(i))\n    tree_info_df = pd.DataFrame(tree_info, columns = ['TreeNo', 'Type', 'ParentFeature', 'SplitOnfeature', 'Level', 'TreePos', 'Gain'])\n    lgbm_feat_dict = dict(enumerate(dumped_model['feature_names']))\n    lgbm_feat_dict[-1] = 'base'\n    tree_info_df['ParentFeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['SplitOnfeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['Interactions'] = tree_info_df['ParentFeature'].map(str) + ' - ' + tree_info_df['SplitOnfeature'].map(str)\n    tree_info_df = round(tree_info_df, 2)\n    lgb_inter_calc = tree_info_df.groupby('Interactions')['Gain'].agg(['count','sum','min','max','mean','std']).sort_values(by='sum', ascending=False).reset_index('Interactions').fillna(0)\n    lgb_inter_calc = round(lgb_inter_calc, 2)\n    lgb_inter_calc_nobase = lgb_inter_calc[lgb_inter_calc['Interactions'].str.contains('base')==False]\n    data = lgb_inter_calc_nobase.sort_values('sum', ascending=False).iloc[0:75].reset_index(drop=True)\n    plt.figure(figsize=(20, 14))\n    ax = plt.subplot(121)\n    sns.barplot(x='sum', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('Total Gain for Feature Interaction', fontweight='bold', fontsize=14)\n    ax = plt.subplot(122)\n    sns.barplot(x='count', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('No. of times Feature interacted', fontweight='bold', fontsize=14)\n    plt.tight_layout()\n\n\nplot_feat_interaction(model_lgb)\n</code></pre>\n<p>.   .  ….. .</p>",
  "messages": [
    {
      "id": "1811497",
      "postDate": "06/04/2022 18:28:47",
      "content": "<h2>Bypass the features anonymization using LightGBM's feature interactions</h2>\n<p>In this competition the features are anonymized, this makes it challenging to engineer meaningful features.</p>\n<p>However, we can still use feature interactions learned by a model (for example LightGBM) to find out meaningful feature pairs and use this information to later on handcraft features to make our models more robust.</p>\n<p><strong>How does this work?</strong></p>\n<p>Lightgbm can also be used to automatically detect useful feature interactions in a dataset and these can be then used to create valuable features.</p>\n<p>Using the splits of the trained LightGBM trees we can find pairs of features and generate new features that are very valuable for training the next models.</p>\n<p>The intuition behind this is that the features used in most of the trees on higher levels are better if they would be preprocessed before entering our models so that the models \"won't waste\" splits on them and the splits are more likely to be used in the next level.</p>\n<p><strong>Code Example</strong></p>\n<p>Code Example for plotting LightGBM feature interactions:</p>\n<p>I have merged some of the code from <a href=\"https://www.kaggle.com/vishalbajaj2000/santander-lightgbm-xgb-feature-interactions\" target=\"_blank\">here</a> (credit to the original author!) and modified it to be simpler.<br>\nThe usage is as simple as: <code>plot_feat_interaction(model_lgb)</code>. </p>\n<p>Hope this helps!</p>\n<pre><code>import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\ndef get_splits_gain(tree_num = 0, parent = -1, tree = None, lev = 0, node_name = None, split_gain = None, reclimit = 50000):\n    if tree == None:\n        raise Exception('No tree present to analyze!')\n    for k, v in tree.items():\n        if type(v) != dict and k in ['split_feature']:\n            old_parent = parent\n            parent = v\n            tag = k\n            yield tree_num, tag, old_parent, parent, lev, node_name, split_gain\n        elif isinstance(v, dict):\n            if v.get('split_gain') == None:\n                continue\n            else:\n                tree = v\n                lev_inc = lev + 1\n                node_name = k\n                split_gain = v['split_gain']\n                for result in get_splits_gain(tree_num, parent, tree, lev_inc, node_name, split_gain):\n                    yield result\n        else: continue\n\ndef plot_feat_interaction(model):\n    dumped_model = model.booster_.dump_model()\n    tree_info = []\n    for j in range(0, len(dumped_model['tree_info'])):\n        for i in get_splits_gain(tree_num = j, tree = dumped_model['tree_info'][j]):\n            tree_info.append(list(i))\n    tree_info_df = pd.DataFrame(tree_info, columns = ['TreeNo', 'Type', 'ParentFeature', 'SplitOnfeature', 'Level', 'TreePos', 'Gain'])\n    lgbm_feat_dict = dict(enumerate(dumped_model['feature_names']))\n    lgbm_feat_dict[-1] = 'base'\n    tree_info_df['ParentFeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['SplitOnfeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['Interactions'] = tree_info_df['ParentFeature'].map(str) + ' - ' + tree_info_df['SplitOnfeature'].map(str)\n    tree_info_df = round(tree_info_df, 2)\n    lgb_inter_calc = tree_info_df.groupby('Interactions')['Gain'].agg(['count','sum','min','max','mean','std']).sort_values(by='sum', ascending=False).reset_index('Interactions').fillna(0)\n    lgb_inter_calc = round(lgb_inter_calc, 2)\n    lgb_inter_calc_nobase = lgb_inter_calc[lgb_inter_calc['Interactions'].str.contains('base')==False]\n    data = lgb_inter_calc_nobase.sort_values('sum', ascending=False).iloc[0:75].reset_index(drop=True)\n    plt.figure(figsize=(20, 14))\n    ax = plt.subplot(121)\n    sns.barplot(x='sum', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('Total Gain for Feature Interaction', fontweight='bold', fontsize=14)\n    ax = plt.subplot(122)\n    sns.barplot(x='count', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('No. of times Feature interacted', fontweight='bold', fontsize=14)\n    plt.tight_layout()\n\n\nplot_feat_interaction(model_lgb)\n</code></pre>\n<p>.   .  ….. .</p>",
      "rawMarkdown": "## Bypass the features anonymization using LightGBM's feature interactions\n\nIn this competition the features are anonymized, this makes it challenging to engineer meaningful features.\n\nHowever, we can still use feature interactions learned by a model (for example LightGBM) to find out meaningful feature pairs and use this information to later on handcraft features to make our models more robust.\n\n**How does this work?**\n\nLightgbm can also be used to automatically detect useful feature interactions in a dataset and these can be then used to create valuable features.\n\nUsing the splits of the trained LightGBM trees we can find pairs of features and generate new features that are very valuable for training the next models.\n\nThe intuition behind this is that the features used in most of the trees on higher levels are better if they would be preprocessed before entering our models so that the models \"won't waste\" splits on them and the splits are more likely to be used in the next level.\n\n**Code Example**\n\nCode Example for plotting LightGBM feature interactions:\n\nI have merged some of the code from [here](https://www.kaggle.com/vishalbajaj2000/santander-lightgbm-xgb-feature-interactions) (credit to the original author!) and modified it to be simpler.\nThe usage is as simple as: `plot_feat_interaction(model_lgb)`. \n\nHope this helps!\n\n```python\n\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\ndef get_splits_gain(tree_num = 0, parent = -1, tree = None, lev = 0, node_name = None, split_gain = None, reclimit = 50000):\n    if tree == None:\n        raise Exception('No tree present to analyze!')\n    for k, v in tree.items():\n        if type(v) != dict and k in ['split_feature']:\n            old_parent = parent\n            parent = v\n            tag = k\n            yield tree_num, tag, old_parent, parent, lev, node_name, split_gain\n        elif isinstance(v, dict):\n            if v.get('split_gain') == None:\n                continue\n            else:\n                tree = v\n                lev_inc = lev + 1\n                node_name = k\n                split_gain = v['split_gain']\n                for result in get_splits_gain(tree_num, parent, tree, lev_inc, node_name, split_gain):\n                    yield result\n        else: continue\n\ndef plot_feat_interaction(model):\n    dumped_model = model.booster_.dump_model()\n    tree_info = []\n    for j in range(0, len(dumped_model['tree_info'])):\n        for i in get_splits_gain(tree_num = j, tree = dumped_model['tree_info'][j]):\n            tree_info.append(list(i))\n    tree_info_df = pd.DataFrame(tree_info, columns = ['TreeNo', 'Type', 'ParentFeature', 'SplitOnfeature', 'Level', 'TreePos', 'Gain'])\n    lgbm_feat_dict = dict(enumerate(dumped_model['feature_names']))\n    lgbm_feat_dict[-1] = 'base'\n    tree_info_df['ParentFeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['SplitOnfeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['Interactions'] = tree_info_df['ParentFeature'].map(str) + ' - ' + tree_info_df['SplitOnfeature'].map(str)\n    tree_info_df = round(tree_info_df, 2)\n    lgb_inter_calc = tree_info_df.groupby('Interactions')['Gain'].agg(['count','sum','min','max','mean','std']).sort_values(by='sum', ascending=False).reset_index('Interactions').fillna(0)\n    lgb_inter_calc = round(lgb_inter_calc, 2)\n    lgb_inter_calc_nobase = lgb_inter_calc[lgb_inter_calc['Interactions'].str.contains('base')==False]\n    data = lgb_inter_calc_nobase.sort_values('sum', ascending=False).iloc[0:75].reset_index(drop=True)\n    plt.figure(figsize=(20, 14))\n    ax = plt.subplot(121)\n    sns.barplot(x='sum', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('Total Gain for Feature Interaction', fontweight='bold', fontsize=14)\n    ax = plt.subplot(122)\n    sns.barplot(x='count', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('No. of times Feature interacted', fontweight='bold', fontsize=14)\n    plt.tight_layout()\n\n\nplot_feat_interaction(model_lgb)\n\n```\n.   .  ..... .",
      "votes": null
    },
    {
      "id": "1811539",
      "postDate": "06/04/2022 19:42:05",
      "content": "<p>Great suggestion. Thanks!</p>",
      "rawMarkdown": "Great suggestion. Thanks!",
      "votes": null
    },
    {
      "id": "1836846",
      "postDate": "06/29/2022 06:30:10",
      "content": "<p>great suggestion!</p>",
      "rawMarkdown": "great suggestion!",
      "votes": null
    },
    {
      "id": "1838936",
      "postDate": "07/01/2022 02:30:48",
      "content": "<p>Great work! Thanks for sharing your work and snippet. </p>",
      "rawMarkdown": "Great work! Thanks for sharing your work and snippet.",
      "votes": null
    },
    {
      "id": "1855391",
      "postDate": "07/14/2022 14:40:29",
      "content": "<p>Nice idea to now was checking via catboost fimp but this is better 👍</p>",
      "rawMarkdown": "Nice idea to now was checking via catboost fimp but this is better 👍",
      "votes": null
    },
    {
      "id": "1873981",
      "postDate": "07/28/2022 01:58:23",
      "content": "<p>Does the same apply to xgboost? 👀 </p>",
      "rawMarkdown": "Does the same apply to xgboost? 👀",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1811539,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "06/04/2022 19:42:05",
      "content": "<p>Great suggestion. Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836846,
      "author_name": "mx1989",
      "author_url": "",
      "post_date": "06/29/2022 06:30:10",
      "content": "<p>great suggestion!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1838936,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "07/01/2022 02:30:48",
      "content": "<p>Great work! Thanks for sharing your work and snippet. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1855391,
      "author_name": "gauravbrills",
      "author_url": "",
      "post_date": "07/14/2022 14:40:29",
      "content": "<p>Nice idea to now was checking via catboost fimp but this is better 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1873981,
      "author_name": "kgxiao",
      "author_url": "",
      "post_date": "07/28/2022 01:58:23",
      "content": "<p>Does the same apply to xgboost? 👀 </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1811497": "## Bypass the features anonymization using LightGBM's feature interactions\n\nIn this competition the features are anonymized, this makes it challenging to engineer meaningful features.\n\nHowever, we can still use feature interactions learned by a model (for example LightGBM) to find out meaningful feature pairs and use this information to later on handcraft features to make our models more robust.\n\n**How does this work?**\n\nLightgbm can also be used to automatically detect useful feature interactions in a dataset and these can be then used to create valuable features.\n\nUsing the splits of the trained LightGBM trees we can find pairs of features and generate new features that are very valuable for training the next models.\n\nThe intuition behind this is that the features used in most of the trees on higher levels are better if they would be preprocessed before entering our models so that the models \"won't waste\" splits on them and the splits are more likely to be used in the next level.\n\n**Code Example**\n\nCode Example for plotting LightGBM feature interactions:\n\nI have merged some of the code from [here](https://www.kaggle.com/vishalbajaj2000/santander-lightgbm-xgb-feature-interactions) (credit to the original author!) and modified it to be simpler.\nThe usage is as simple as: `plot_feat_interaction(model_lgb)`. \n\nHope this helps!\n\n```python\n\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\ndef get_splits_gain(tree_num = 0, parent = -1, tree = None, lev = 0, node_name = None, split_gain = None, reclimit = 50000):\n    if tree == None:\n        raise Exception('No tree present to analyze!')\n    for k, v in tree.items():\n        if type(v) != dict and k in ['split_feature']:\n            old_parent = parent\n            parent = v\n            tag = k\n            yield tree_num, tag, old_parent, parent, lev, node_name, split_gain\n        elif isinstance(v, dict):\n            if v.get('split_gain') == None:\n                continue\n            else:\n                tree = v\n                lev_inc = lev + 1\n                node_name = k\n                split_gain = v['split_gain']\n                for result in get_splits_gain(tree_num, parent, tree, lev_inc, node_name, split_gain):\n                    yield result\n        else: continue\n\ndef plot_feat_interaction(model):\n    dumped_model = model.booster_.dump_model()\n    tree_info = []\n    for j in range(0, len(dumped_model['tree_info'])):\n        for i in get_splits_gain(tree_num = j, tree = dumped_model['tree_info'][j]):\n            tree_info.append(list(i))\n    tree_info_df = pd.DataFrame(tree_info, columns = ['TreeNo', 'Type', 'ParentFeature', 'SplitOnfeature', 'Level', 'TreePos', 'Gain'])\n    lgbm_feat_dict = dict(enumerate(dumped_model['feature_names']))\n    lgbm_feat_dict[-1] = 'base'\n    tree_info_df['ParentFeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['SplitOnfeature'].replace(lgbm_feat_dict, inplace = True)\n    tree_info_df['Interactions'] = tree_info_df['ParentFeature'].map(str) + ' - ' + tree_info_df['SplitOnfeature'].map(str)\n    tree_info_df = round(tree_info_df, 2)\n    lgb_inter_calc = tree_info_df.groupby('Interactions')['Gain'].agg(['count','sum','min','max','mean','std']).sort_values(by='sum', ascending=False).reset_index('Interactions').fillna(0)\n    lgb_inter_calc = round(lgb_inter_calc, 2)\n    lgb_inter_calc_nobase = lgb_inter_calc[lgb_inter_calc['Interactions'].str.contains('base')==False]\n    data = lgb_inter_calc_nobase.sort_values('sum', ascending=False).iloc[0:75].reset_index(drop=True)\n    plt.figure(figsize=(20, 14))\n    ax = plt.subplot(121)\n    sns.barplot(x='sum', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('Total Gain for Feature Interaction', fontweight='bold', fontsize=14)\n    ax = plt.subplot(122)\n    sns.barplot(x='count', y='Interactions', data=data.sort_values('sum', ascending=False), ax=ax)\n    ax.set_title('No. of times Feature interacted', fontweight='bold', fontsize=14)\n    plt.tight_layout()\n\n\nplot_feat_interaction(model_lgb)\n\n```\n.   .  ..... .",
    "1811539": "Great suggestion. Thanks!",
    "1836846": "great suggestion!",
    "1838936": "Great work! Thanks for sharing your work and snippet.",
    "1855391": "Nice idea to now was checking via catboost fimp but this is better 👍",
    "1873981": "Does the same apply to xgboost? 👀"
  },
  "source": "meta"
}