{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"### To discover features that should be prioritized, we visualize the feature importance(split and gain).\n\n### 優先するべき特徴量を発見するため、特徴量重要度(splitおよびgain)を可視化します。","metadata":{}},{"cell_type":"markdown","source":"# Step 1 Train LightGBM model / ステップ1 LightGBMモデルの学習\nThe LightGBM model was trained using the following very helpful notebook.\n\nLightGBMモデルの学習は、下記の非常に参考になるnotebookを用いて行いました。\n\n**amex-features-the-best-of-both-worlds** by The Devastator\n\nhttps://www.kaggle.com/code/thedevastator/amex-features-the-best-of-both-worlds","metadata":{}},{"cell_type":"markdown","source":"# Step 2 Plot Importance","metadata":{"execution":{"iopub.status.busy":"2022-08-15T13:53:32.804511Z","iopub.execute_input":"2022-08-15T13:53:32.804919Z","iopub.status.idle":"2022-08-15T13:53:32.812027Z","shell.execute_reply.started":"2022-08-15T13:53:32.804885Z","shell.execute_reply":"2022-08-15T13:53:32.810513Z"}}},{"cell_type":"markdown","source":" We can plot 2 types of feature importance.\n \n 2種類の特徴量重要度を可視化します。\n \n\n 1. “split”: numbers of times the feature is used in a model. / モデル内で特徴量が使われた回数。\n 2. “gain”: total gains of splits which use the feature. / 特長量によって改善できた目的関数の合計。\n \n  https://lightgbm.readthedocs.io/en/latest/pythonapi/lightgbm.plot_importance.html","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nfrom tqdm import tqdm\nimport lightgbm as lgb\nimport joblib","metadata":{"execution":{"iopub.status.busy":"2022-08-15T15:55:36.869346Z","iopub.execute_input":"2022-08-15T15:55:36.869781Z","iopub.status.idle":"2022-08-15T15:55:36.875103Z","shell.execute_reply.started":"2022-08-15T15:55:36.869724Z","shell.execute_reply":"2022-08-15T15:55:36.874038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Read pre-trained model of each fold.\n\n各Foldの学習済みモデルを読み込みます。","metadata":{}},{"cell_type":"code","source":"model0 = joblib.load(\"../input/amexpretrainedlgbm/lgbm_fold0_seed42.pkl\")\nmodel1 = joblib.load(\"../input/amexpretrainedlgbm/lgbm_fold1_seed42.pkl\")\nmodel2 = joblib.load(\"../input/amexpretrainedlgbm/lgbm_fold2_seed42.pkl\")\nmodel3 = joblib.load(\"../input/amexpretrainedlgbm/lgbm_fold3_seed42.pkl\")\nmodel4 = joblib.load(\"../input/amexpretrainedlgbm/lgbm_fold4_seed42.pkl\")","metadata":{"execution":{"iopub.status.busy":"2022-08-15T15:55:36.883790Z","iopub.execute_input":"2022-08-15T15:55:36.884146Z","iopub.status.idle":"2022-08-15T15:55:41.205903Z","shell.execute_reply.started":"2022-08-15T15:55:36.884115Z","shell.execute_reply":"2022-08-15T15:55:41.204850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Get feature list. There are 2,177 features.\n\n特徴量の一覧を取得します。2,177種類の特徴量があります。","metadata":{}},{"cell_type":"code","source":"features = model1.feature_name()\nprint(\"num features:\", len(features))","metadata":{"execution":{"iopub.status.busy":"2022-08-15T15:55:41.207892Z","iopub.execute_input":"2022-08-15T15:55:41.208228Z","iopub.status.idle":"2022-08-15T15:55:41.218046Z","shell.execute_reply.started":"2022-08-15T15:55:41.208196Z","shell.execute_reply":"2022-08-15T15:55:41.216899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Visualize the feature importance pre-trained model of each fold.\n\n各Foldの学習済みモデルの特徴量重要度を可視化します。","metadata":{}},{"cell_type":"markdown","source":"# 1. importance type: split","metadata":{}},{"cell_type":"code","source":"for model in [model0, model1, model2, model3, model4]:\n    lgb.plot_importance(model, max_num_features=20, importance_type='split')","metadata":{"execution":{"iopub.status.busy":"2022-08-15T16:27:18.069416Z","iopub.execute_input":"2022-08-15T16:27:18.070092Z","iopub.status.idle":"2022-08-15T16:27:20.036485Z","shell.execute_reply.started":"2022-08-15T16:27:18.070049Z","shell.execute_reply":"2022-08-15T16:27:20.035366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Importance varies slightly by Fold. So we plot mean of the importance.\n\nP2, D_39, S_3, B_4 and D_43 appear to be important features for these models.\n\nFoldによって、特徴量重要度が若干異なるので、平均を可視化します。\n\nP2、D_39、S_3、B_4、D_43が、これらのモデルにとって重要な特徴量のようです。","metadata":{}},{"cell_type":"markdown","source":"Top 100","metadata":{}},{"cell_type":"code","source":"first_flg = True\nfor model in [model0, model1, model2, model3, model4]:\n    if first_flg:\n        importance = pd.DataFrame(model.feature_importance(importance_type='split'), index=features, columns=['importance'])\n        first_flg=False\n    else:\n        importance = pd.concat([importance, pd.DataFrame(model.feature_importance(importance_type='split'), index=features, columns=['importance'])],axis=1)\nimportance[\"mean\"] = importance.mean(axis=1)\nimportance.sort_values(\"mean\", ascending=True)[\"mean\"][-100:].plot(figsize=(15, 20), kind=\"barh\")","metadata":{"execution":{"iopub.status.busy":"2022-08-15T16:29:12.954906Z","iopub.execute_input":"2022-08-15T16:29:12.955326Z","iopub.status.idle":"2022-08-15T16:29:14.440794Z","shell.execute_reply.started":"2022-08-15T16:29:12.955294Z","shell.execute_reply":"2022-08-15T16:29:14.439617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. importance type: gain","metadata":{}},{"cell_type":"markdown","source":"In case of importance_type='gain', P_2, B_9, D_48, B_1 and D_44 are the upper levels.\n\n\nimportance_type='gain'の場合、P_2、B_9、D_48、B_1、D_44 が上位となります。","metadata":{}},{"cell_type":"code","source":"for model in [model0, model1, model2, model3, model4]:\n    lgb.plot_importance(model, max_num_features=20, importance_type='gain')","metadata":{"execution":{"iopub.status.busy":"2022-08-15T16:29:38.647250Z","iopub.execute_input":"2022-08-15T16:29:38.647791Z","iopub.status.idle":"2022-08-15T16:29:40.696392Z","shell.execute_reply.started":"2022-08-15T16:29:38.647753Z","shell.execute_reply":"2022-08-15T16:29:40.695292Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"first_flg = True\nfor model in [model0, model1, model2, model3, model4]:\n    if first_flg:\n        importance = pd.DataFrame(model.feature_importance(importance_type='gain'), index=features, columns=['importance'])\n        first_flg=False\n    else:\n        importance = pd.concat([importance, pd.DataFrame(model.feature_importance(importance_type='gain'), index=features, columns=['importance'])],axis=1)\nimportance[\"mean\"] = importance.mean(axis=1)\nimportance.sort_values(\"mean\", ascending=True)[\"mean\"][-100:].plot(figsize=(15, 20), kind=\"barh\")","metadata":{"execution":{"iopub.status.busy":"2022-08-15T16:29:40.698711Z","iopub.execute_input":"2022-08-15T16:29:40.699853Z","iopub.status.idle":"2022-08-15T16:29:42.546880Z","shell.execute_reply.started":"2022-08-15T16:29:40.699800Z","shell.execute_reply":"2022-08-15T16:29:42.545836Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}