{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import pandas as pd\nimport itertools\nimport numpy as np\n\ndef read_data(cols):\n    \n    print('Reading data...')\n    \n    df = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet', columns=cols)\n    \n    # simplify cus_id\n    unique_cus_ids = df.customer_ID.unique()\n    assignment     = dict(zip(unique_cus_ids, list(range(len(unique_cus_ids)))))\n    df.customer_ID = df.customer_ID.apply(lambda x: assignment[x]).astype('int32')\n    \n    print('shape of data:', df.shape)\n    \n    return df","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-05T11:27:37.710468Z","iopub.execute_input":"2022-08-05T11:27:37.710873Z","iopub.status.idle":"2022-08-05T11:27:37.719009Z","shell.execute_reply.started":"2022-08-05T11:27:37.710841Z","shell.execute_reply":"2022-08-05T11:27:37.717756Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"By now most of you probably noted that the following variables seem to be the most important in their groups. At least my permutation importance tells me so :) Note they are all numerical. \n\nBut how can we use this knowledge to further compute features in a FAST manner, WITHOUT running all computed variables through a long permutation imprtance process? In this notebook I will provide you a fast approach using \"Information Values\", which are often used in \"Weights od Evidence\" for credit scoring tasks like ours.\n\nThe final result - the additional features - can be found here:<br> https://www.kaggle.com/datasets/gzguevara/amex-crossnon-linear-features","metadata":{}},{"cell_type":"code","source":"high = ['P_2', 'D_48', 'R_1', 'S_3', 'B_7']","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:50:18.804375Z","iopub.execute_input":"2022-08-05T11:50:18.804814Z","iopub.status.idle":"2022-08-05T11:50:18.810342Z","shell.execute_reply.started":"2022-08-05T11:50:18.804772Z","shell.execute_reply":"2022-08-05T11:50:18.808898Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# read only categorial  features\ndata  = read_data(['customer_ID'] + high)\n# read targets\ntarget = pd.read_csv('../input/amex-default-prediction/train_labels.csv', usecols=['target'])","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:50:20.176242Z","iopub.execute_input":"2022-08-05T11:50:20.176639Z","iopub.status.idle":"2022-08-05T11:50:25.088878Z","shell.execute_reply.started":"2022-08-05T11:50:20.176605Z","shell.execute_reply":"2022-08-05T11:50:25.087639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's see what this \"Information Value\" is.\n\nInformation value is one of the most useful technique to select important variables in a predictive model. It helps to rank variables on the basis of their importance. \n\nIf you want to read more about it:\nhttps://www.listendata.com/2015/03/weight-of-evidence-woe-and-information.html\n\nFrom the same link I got the snippet in the following hidden cell. ","metadata":{}},{"cell_type":"code","source":"def iv_woe(data, target, bins=20, show_woe=False):\n    \n    #Empty Dataframe\n    newDF,woeDF = pd.DataFrame(), pd.DataFrame()\n    \n    #Extract Column Names\n    cols = data.columns\n    \n    #Run WOE and IV on all the independent variables\n    for ivars in cols[~cols.isin([target])]:\n\n        print(ivars)\n\n        if (data[ivars].dtype.kind in 'bifc') and (len(np.unique(data[ivars]))>3):\n            binned_x = pd.qcut(data[ivars], bins,  duplicates='drop')\n            d0 = pd.DataFrame({'x': binned_x, 'y': data[target]})\n        else:\n            d0 = pd.DataFrame({'x': data[ivars], 'y': data[target]})\n            \n        d0 = d0.astype({\"x\": str})\n        d = d0.groupby(\"x\", as_index=False, dropna=False).agg({\"y\": [\"count\", \"sum\"]})\n        d.columns = ['Cutoff', 'N', 'Events']\n        d['% of Events'] = np.maximum(d['Events'], 0.5) / d['Events'].sum()\n        d['Non-Events'] = d['N'] - d['Events']\n        d['% of Non-Events'] = np.maximum(d['Non-Events'], 0.5) / d['Non-Events'].sum()\n        d['WoE'] = np.log(d['% of Non-Events']/d['% of Events'])\n        d['IV'] = d['WoE'] * (d['% of Non-Events']-d['% of Events'])\n        d.insert(loc=0, column='Variable', value=ivars)\n        print(\"Information value of \" + ivars + \" is \" + str(round(d['IV'].sum(),6)))\n        temp =pd.DataFrame({\"Variable\" : [ivars], \"IV\" : [d['IV'].sum()]}, columns = [\"Variable\", \"IV\"])\n        newDF=pd.concat([newDF,temp], axis=0)\n        woeDF=pd.concat([woeDF,d], axis=0)\n\n        #Show WOE Table\n        if show_woe == True:\n            print(d)\n    return newDF, woeDF","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:25:26.544926Z","iopub.execute_input":"2022-08-05T11:25:26.545935Z","iopub.status.idle":"2022-08-05T11:25:26.567198Z","shell.execute_reply.started":"2022-08-05T11:25:26.545891Z","shell.execute_reply":"2022-08-05T11:25:26.565773Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's now create some basic data and check the information value.","metadata":{}},{"cell_type":"code","source":"aggs = data.groupby('customer_ID').agg(['last'])\naggs.columns = ['_'.join(x) for x in aggs.columns]\naggs['target'] = target.target","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:50:28.738231Z","iopub.execute_input":"2022-08-05T11:50:28.738700Z","iopub.status.idle":"2022-08-05T11:50:29.240183Z","shell.execute_reply.started":"2022-08-05T11:50:28.738660Z","shell.execute_reply":"2022-08-05T11:50:29.239226Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's see how it works!\n\nSo we need to provide data and indicate which of your column is the target.","metadata":{}},{"cell_type":"code","source":"info, _ = iv_woe(aggs, 'target')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:50:31.728552Z","iopub.execute_input":"2022-08-05T11:50:31.728968Z","iopub.status.idle":"2022-08-05T11:50:32.868287Z","shell.execute_reply.started":"2022-08-05T11:50:31.728930Z","shell.execute_reply":"2022-08-05T11:50:32.867197Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Ok! so now we know aproximately how much information those variables contain on how a good split might looks like.\n\nIf we now get all possible pairs of those high information features, we can add, substract, multiply and dive them, aggregate those features with 'last', 'mean, 'max' etc. and see how much information we get - WITHOUT first implementing them into our model!","metadata":{}},{"cell_type":"code","source":"all_pairs = []\nfor i in range(len(high) -1):\n\n    all_pairs.extend(list(itertools.product([high[i]], high[i+1:])))\n\nlen(all_pairs)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:50:36.040604Z","iopub.execute_input":"2022-08-05T11:50:36.041788Z","iopub.status.idle":"2022-08-05T11:50:36.050783Z","shell.execute_reply.started":"2022-08-05T11:50:36.041742Z","shell.execute_reply":"2022-08-05T11:50:36.049494Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can 10 pairs. Let's compute all possible combinations of those pairs and check the information value we get! NOTE we only save features with high information value! more than 2.5.","metadata":{}},{"cell_type":"code","source":"all_features = pd.DataFrame()\n\nfor pair in all_pairs:\n\n    aggregates = pd.DataFrame()\n    aggregates['customer_ID'] = data.customer_ID\n    \n    #compute new features\n    aggregates[f'{pair[0]}_t_{pair[1]}'] = data[pair[0]] * data[pair[1]]\n    aggregates[f'{pair[0]}_d_{pair[1]}'] = data[pair[0]] / data[pair[1]]\n    aggregates[f'{pair[0]}_p_{pair[1]}'] = data[pair[0]] + data[pair[1]]\n    aggregates[f'{pair[0]}_m_{pair[1]}'] = data[pair[0]] - data[pair[1]]\n    \n    # compute aggregation\n    aggregates = aggregates.groupby('customer_ID').agg(['last', 'first', 'mean', 'std', 'max', 'min'])\n    aggregates.columns = ['_'.join(x) for x in aggregates.columns]\n    aggregates = aggregates.fillna(0)\n    aggregates.replace([np.inf, -np.inf], 0, inplace=True)\n    \n    # add target\n    aggregates['target'] = target.target\n    \n    # Get Information Value\n    a, b = iv_woe(aggregates, 'target')\n    \n    # Select good features\n    good_ones = a.loc[a.IV > 2.5].Variable.values\n    all_features[good_ones] = aggregates[good_ones]\n    \n    print('\\n', 'current shape', all_features.shape, '\\n')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:50:45.403724Z","iopub.execute_input":"2022-08-05T11:50:45.404165Z","iopub.status.idle":"2022-08-05T11:52:21.475029Z","shell.execute_reply.started":"2022-08-05T11:50:45.404126Z","shell.execute_reply":"2022-08-05T11:52:21.473885Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Amazing! we have computed a 37 new features, and without implementing them into our model, we can confidently say that they will be useful! :) In my permutation importance those features always score high!","metadata":{}},{"cell_type":"code","source":"all_features","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:52:55.521739Z","iopub.execute_input":"2022-08-05T11:52:55.522240Z","iopub.status.idle":"2022-08-05T11:52:55.655019Z","shell.execute_reply.started":"2022-08-05T11:52:55.522197Z","shell.execute_reply":"2022-08-05T11:52:55.653874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"I hope this will help you!\n\nHere is the link to the data:\nhttps://www.kaggle.com/datasets/gzguevara/amex-crossnon-linear-features\n\nHere I used much more features incially! - Take a look!","metadata":{}},{"cell_type":"code","source":"pd.read_parquet('../input/amex-crossnon-linear-features').columns","metadata":{"execution":{"iopub.status.busy":"2022-08-05T11:55:33.734497Z","iopub.execute_input":"2022-08-05T11:55:33.734929Z","iopub.status.idle":"2022-08-05T11:55:34.291478Z","shell.execute_reply.started":"2022-08-05T11:55:33.734893Z","shell.execute_reply":"2022-08-05T11:55:34.290295Z"},"trusted":true},"execution_count":null,"outputs":[]}]}