{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# LOFO Feature Importance","metadata":{}},{"cell_type":"markdown","source":"From [aerdem4](https://www.kaggle.com/aerdem4/) notebook \n\nSee the git repo and the other notebooks by [aerdem4](https://www.kaggle.com/aerdem4/)","metadata":{}},{"cell_type":"markdown","source":"> LOFO (Leave One Feature Out) Importance calculates the importances of a set of features based on a metric of choice, for a model of choice, by iteratively removing each feature from the set, and evaluating the performance of the model, with a validation scheme of choice, based on the chosen metric.\nLOFO first evaluates the performance of the model with all the input features included, then iteratively removes one feature at a time, retrains the model, and evaluates its performance on a validation set. The mean and standard deviation (across the folds) of the importance of each feature is then reported.\n\n> While other feature importance methods usually calculate how much a feature is used by the model, LOFO estimates how much a feature can make a difference by itself given that we have the other features. \n\nHere are some advantages of LOFO:\n- It generalises well to unseen test sets since it uses a validation scheme.\n- It is model agnostic.\n- It gives negative importance to features that hurt performance upon inclusion.\n- It can group the features. Especially useful for high dimensional features like TFIDF or OHE features. It is also good practice to group very correlated features to avoid misleading results.\n- It can automatically group highly correlated features to avoid underestimating their importance.\n","metadata":{}},{"cell_type":"markdown","source":"## Refs: \n\n- git repo: https://github.com/aerdem4/lofo-importance\n- similar notebooks: https://www.kaggle.com/code/aerdem4/google-ventilator-lofo-feature-importance","metadata":{}},{"cell_type":"code","source":"!pip install lofo-importance -q","metadata":{"_kg_hide-output":true,"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-01-05T16:55:36.770886Z","iopub.execute_input":"2023-01-05T16:55:36.772110Z","iopub.status.idle":"2023-01-05T16:55:48.116206Z","shell.execute_reply.started":"2023-01-05T16:55:36.771920Z","shell.execute_reply":"2023-01-05T16:55:48.115275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nfrom tqdm import tqdm\nimport os, sys\nimport torch\n\n\ndf = pd.read_csv(f\"/kaggle/input/playground-series-s3e1/train.csv\")\ntarget = 'MedHouseVal'\n\nprint(df.shape)\ndf.head()","metadata":{"execution":{"iopub.status.busy":"2023-01-05T16:55:48.118405Z","iopub.execute_input":"2023-01-05T16:55:48.118650Z","iopub.status.idle":"2023-01-05T16:55:49.562628Z","shell.execute_reply.started":"2023-01-05T16:55:48.118619Z","shell.execute_reply":"2023-01-05T16:55:49.561774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Minimal FE","metadata":{}},{"cell_type":"code","source":"## from: https://www.kaggle.com/code/dmitryuarov/ps-s3e1-coordinates-key-to-victory\n\ndef rt_crds(df): \n    df['rot_15_x'] = (np.cos(np.radians(15)) * df['Longitude']) + \\\n                      (np.sin(np.radians(15)) * df['Latitude'])\n    \n    df['rot_15_y'] = (np.cos(np.radians(15)) * df['Latitude']) - \\\n                      (np.sin(np.radians(15)) * df['Longitude'])\n    \n    df['rot_30_x'] = (np.cos(np.radians(30)) * df['Longitude']) + \\\n                      (np.sin(np.radians(30)) * df['Latitude'])\n    \n    df['rot_30_y'] = (np.cos(np.radians(30)) * df['Latitude']) - \\\n                      (np.sin(np.radians(30)) * df['Longitude'])\n    \n    df['rot_45_x'] = (np.cos(np.radians(45)) * df['Longitude']) + \\\n                      (np.sin(np.radians(45)) * df['Latitude'])\n    \n    df['rot_45_y'] = (np.cos(np.radians(45)) * df['Latitude']) - \\\n                      (np.sin(np.radians(45)) * df['Longitude'])\n    return df\n\ndf = rt_crds(df)","metadata":{"execution":{"iopub.status.busy":"2023-01-05T16:55:49.571210Z","iopub.execute_input":"2023-01-05T16:55:49.571493Z","iopub.status.idle":"2023-01-05T16:55:49.600217Z","shell.execute_reply.started":"2023-01-05T16:55:49.571440Z","shell.execute_reply":"2023-01-05T16:55:49.599096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# cat_cols = df.select_dtypes('category').columns.tolist()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# LOFO Importance","metadata":{}},{"cell_type":"code","source":"from lofo import Dataset, LOFOImportance, plot_importance\nfrom sklearn.model_selection import StratifiedKFold, KFold\n\n\ncv = list(KFold(n_splits=10, shuffle=True).split(df))\nfeatures = [c for c in df.columns if c not in ['id', target] ]\n\nds = Dataset(df, target=target, features=features, feature_groups=None, auto_group_threshold=0.9)","metadata":{"execution":{"iopub.status.busy":"2023-01-05T16:55:49.601531Z","iopub.execute_input":"2023-01-05T16:55:49.601797Z","iopub.status.idle":"2023-01-05T16:55:52.473046Z","shell.execute_reply.started":"2023-01-05T16:55:49.601767Z","shell.execute_reply":"2023-01-05T16:55:52.471951Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lofo_imp = LOFOImportance(ds, cv=cv, scoring='neg_root_mean_squared_error')  ### sorted(sklearn.metrics.SCORERS.keys())\n\nimportance_df = lofo_imp.get_importance()\nimportance_df","metadata":{"execution":{"iopub.status.busy":"2023-01-05T17:01:43.431120Z","iopub.execute_input":"2023-01-05T17:01:43.432132Z","iopub.status.idle":"2023-01-05T17:02:07.555437Z","shell.execute_reply.started":"2023-01-05T17:01:43.432080Z","shell.execute_reply":"2023-01-05T17:02:07.554492Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_importance(importance_df, figsize=(10, 10)) ## kind=\"box\"","metadata":{"execution":{"iopub.status.busy":"2023-01-05T17:02:23.074797Z","iopub.execute_input":"2023-01-05T17:02:23.075158Z","iopub.status.idle":"2023-01-05T17:02:23.427097Z","shell.execute_reply.started":"2023-01-05T17:02:23.075112Z","shell.execute_reply":"2023-01-05T17:02:23.426019Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### let's increase the threshold for grouping similar cols..\n\nds = Dataset(df, target=target, features=features, feature_groups=None, auto_group_threshold=0.97)\n\nlofo_imp = LOFOImportance(ds, cv=cv, scoring='neg_root_mean_squared_error')  ### sorted(sklearn.metrics.SCORERS.keys())\n\nimportance_df = lofo_imp.get_importance()\nimportance_df","metadata":{"execution":{"iopub.status.busy":"2023-01-05T17:03:36.754536Z","iopub.execute_input":"2023-01-05T17:03:36.754840Z","iopub.status.idle":"2023-01-05T17:04:06.372775Z","shell.execute_reply.started":"2023-01-05T17:03:36.754810Z","shell.execute_reply":"2023-01-05T17:04:06.371892Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_importance(importance_df, figsize=(10, 10))","metadata":{"execution":{"iopub.status.busy":"2023-01-05T17:08:56.271876Z","iopub.execute_input":"2023-01-05T17:08:56.272322Z","iopub.status.idle":"2023-01-05T17:08:56.639589Z","shell.execute_reply.started":"2023-01-05T17:08:56.272286Z","shell.execute_reply":"2023-01-05T17:08:56.638188Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"### alternative vis\n\nplot_importance(importance_df, figsize=(12, 10), kind=\"box\")","metadata":{"execution":{"iopub.status.busy":"2023-01-05T17:09:25.522513Z","iopub.execute_input":"2023-01-05T17:09:25.522831Z","iopub.status.idle":"2023-01-05T17:09:25.996328Z","shell.execute_reply.started":"2023-01-05T17:09:25.522799Z","shell.execute_reply":"2023-01-05T17:09:25.994811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ds = Dataset(df, target=target, features=features, feature_groups=None, auto_group_threshold=0.99)\n\nlofo_imp = LOFOImportance(ds, cv=cv, scoring='neg_root_mean_squared_error')  \n\nimportance_df = lofo_imp.get_importance()\nimportance_df","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_importance(importance_df, figsize=(10, 10))","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Fork and add more features to see their ranking.. ","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}