{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import pandas as pd\nimport itertools\nimport numpy as np\n\ndef read_data(cols):\n    \n    print('Reading data...')\n    \n    df = pd.read_parquet('../input/amex-data-integer-dtypes-parquet-format/train.parquet', columns=cols)\n    \n    # simplify cus_id\n    unique_cus_ids = df.customer_ID.unique()\n    assignment     = dict(zip(unique_cus_ids, list(range(len(unique_cus_ids)))))\n    df.customer_ID = df.customer_ID.apply(lambda x: assignment[x]).astype('int32')\n    \n    print('shape of data:', df.shape)\n    \n    return df\n\n# method for Information Value\ndef iv_woe(data, target, bins=20, show_woe=False):\n    \n    #Empty Dataframe\n    newDF,woeDF = pd.DataFrame(), pd.DataFrame()\n    \n    #Extract Column Names\n    cols = data.columns\n    \n    #Run WOE and IV on all the independent variables\n    for ivars in cols[~cols.isin([target])]:\n\n        print(ivars)\n\n        if (data[ivars].dtype.kind in 'bifc') and (len(np.unique(data[ivars]))>3):\n            binned_x = pd.qcut(data[ivars], bins,  duplicates='drop')\n            d0 = pd.DataFrame({'x': binned_x, 'y': data[target]})\n        else:\n            d0 = pd.DataFrame({'x': data[ivars], 'y': data[target]})\n            \n        d0 = d0.astype({\"x\": str})\n        d = d0.groupby(\"x\", as_index=False, dropna=False).agg({\"y\": [\"count\", \"sum\"]})\n        d.columns = ['Cutoff', 'N', 'Events']\n        d['% of Events'] = np.maximum(d['Events'], 0.5) / d['Events'].sum()\n        d['Non-Events'] = d['N'] - d['Events']\n        d['% of Non-Events'] = np.maximum(d['Non-Events'], 0.5) / d['Non-Events'].sum()\n        d['WoE'] = np.log(d['% of Non-Events']/d['% of Events'])\n        d['IV'] = d['WoE'] * (d['% of Non-Events']-d['% of Events'])\n        d.insert(loc=0, column='Variable', value=ivars)\n        print(\"Information value of \" + ivars + \" is \" + str(round(d['IV'].sum(),6)))\n        temp =pd.DataFrame({\"Variable\" : [ivars], \"IV\" : [d['IV'].sum()]}, columns = [\"Variable\", \"IV\"])\n        newDF=pd.concat([newDF,temp], axis=0)\n        woeDF=pd.concat([woeDF,d], axis=0)\n\n        #Show WOE Table\n        if show_woe == True:\n            print(d)\n    return newDF, woeDF","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-06T10:42:55.078032Z","iopub.execute_input":"2022-08-06T10:42:55.078418Z","iopub.status.idle":"2022-08-06T10:42:55.101766Z","shell.execute_reply.started":"2022-08-06T10:42:55.078386Z","shell.execute_reply":"2022-08-06T10:42:55.100482Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Smart Brute Force Feature Engineering","metadata":{}},{"cell_type":"markdown","source":"Brute force feature engineering is a effective approach to extract new information from data. [thedevastator](https://www.kaggle.com/thedevastator) has shown in his [notebook](https://www.kaggle.com/code/thedevastator/amex-bruteforce-feature-engineering) that this method indeed is fast and improves the score of your model! However, I want to claim that this method is not efficient. Of course, you can compute all kind of features like:\n\n- `P_2_Last -  P_2_Mean`\n- `B_6_Last /  B_6_Mean`\n- `S_14_Last  +  S_14_Mean`\n- `S_14_Last   /  S_14_Mean`\n- `P_2_Last  +  P_2_Mean`\n\nYou can even try to compute new features based on different feature aggregation like:\n\n- `P_2_First  -  B_3_Last`\n- `S_14_First /  B_6_Last`\n- `P_2_Mean   +  B_6_std`\n- `B_3_Max    *  S_14_Min`\n- `B_3_Last   +  P_2_Mean^2`\n\n**BUT HOW DO YOU KNOW WHETHER THIS IS NOT JUST NOISE?!?!?!**\n\nWell the is where the 'smart' comes in! In my previous [notebook](https://www.kaggle.com/code/gzguevara/new-features-based-on-information-value) I have introduced the usage of \"Information Values\" to get a grasp of whether your feature contains good new information. I propose to compute all kind of crazy features, of which we cannot really know, whether they are of good information, and then apply the same approach, based on the \"Information Value\". This way we can select only those crazy features, which indeed contain useful information. ","metadata":{}},{"cell_type":"code","source":"# Features which are know to have high information\nhigh = ['P_2', 'P_3', 'P_4', \n        'D_48', 'D_42', 'D_44', 'D_61',\n        'R_1', 'R_3', 'R_10', 'R_5', 'R_16',\n        'S_3', 'S_7', 'S_15', 'S_22', 'S_8',\n        'B_7', 'B_23', 'B_9', 'B_10', 'B_2']\n\n# Get all possible combination of high information features\nall_pairs = []\nfor i in range(len(high) -1): all_pairs.extend(list(itertools.product([high[i]], high[i+1:])))","metadata":{"execution":{"iopub.status.busy":"2022-08-06T10:29:23.897070Z","iopub.execute_input":"2022-08-06T10:29:23.897501Z","iopub.status.idle":"2022-08-06T10:29:23.904871Z","shell.execute_reply.started":"2022-08-06T10:29:23.897469Z","shell.execute_reply":"2022-08-06T10:29:23.903860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we have all possible pairs, based on which we can compute all kind of crazy features. As soon as they are computed, we evalute them based on their Information Value and if the Information Value is high enough, we keep this feature!","metadata":{}},{"cell_type":"code","source":"# read only high-information features\ndata  = read_data(['customer_ID'] + high)\n# read targets\ntarget = pd.read_csv('../input/amex-default-prediction/train_labels.csv', usecols=['target'])\n\ndef smart_brute_force(info_cutoff):\n\n    all_features = pd.DataFrame()\n\n    for pair in all:\n\n        # Basic aggregations\n        group_a = train[['customer_ID', pair[0]]].groupby('customer_ID').agg(['last', 'first', 'median', 'mean', 'std', 'max', 'min'])\n        group_b = train[['customer_ID', pair[1]]].groupby('customer_ID').agg(['last', 'first', 'median', 'mean', 'std', 'max', 'min'])\n        group_a.columns = [x[1] + '_' for x in group_a.columns]\n        group_b.columns = [x[1] + '_' for x in group_b.columns]\n\n        # Combinations\n        new_features = pd.DataFrame()\n\n        # Crazy Features and much more are possible if you want - try you own!\n        new_features[f'{pair[0]}_last_t_{pair[1]}_std']  = group_a.last_ * group_b.std_\n        new_features[f'{pair[0]}_last_d_{pair[1]}_mean'] = group_a.last_ / group_b.mean_\n        new_features[f'{pair[0]}_last_p_{pair[1]}_max']  = group_a.last_ + group_b.max_\n        new_features[f'{pair[0]}_last_m_{pair[1]}_min']  = group_a.last_ - group_b.min_\n        new_features[f'{pair[0]}_last_t_{pair[1]}_median']  = group_a.last_ * group_b.median_\n        new_features[f'{pair[0]}_last_t_{pair[1]}_first']  = group_a.last_ * group_b.first_\n        new_features[f'{pair[0]}_last_t_{pair[1]}_last']  = group_a.last_ * group_b.last_\n\n        new_features[f'{pair[0]}_mean_t_{pair[1]}_std']  = group_a.mean_ * group_b.std_\n        new_features[f'{pair[0]}_mean_d_{pair[1]}_mean'] = group_a.mean_ / group_b.mean_\n        new_features[f'{pair[0]}_mean_p_{pair[1]}_max']  = group_a.mean_ + group_b.max_\n        new_features[f'{pair[0]}_mean_m_{pair[1]}_min']  = group_a.mean_ - group_b.min_\n        new_features[f'{pair[0]}_mean_t_{pair[1]}_median']  = group_a.mean_ * group_b.median_\n        new_features[f'{pair[0]}_mean_t_{pair[1]}_first']  = group_a.mean_ * group_b.first_\n        new_features[f'{pair[0]}_mean_t_{pair[1]}_last']  = group_a.mean_ * group_b.last_\n\n        # Clean possible missing values and inf's\n        new_features = new_features.fillna(0)\n        new_features.replace([np.inf, -np.inf], 0, inplace=True)\n\n        # Get Information Value\n        new_features['target'] = targets.target\n        a, b = iv_woe(new_features, 'target')\n\n        # Select only new features with high information!\n        good_ones = a.loc[a.IV > info_cutoff].Variable.values\n\n        # Save new high-information features\n        all_features[good_ones] = new_features[good_ones]\n\n        print('\\n', all_features.shape, '\\n')\n        \n        return all_features\n        \n#new_features = smart_brute_force(2.65)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The result of this method is a collection of over 200 features, from which you know, that they will contribute new & high information to your model! You can find dataset resulting from the method [here](https://www.kaggle.com/datasets/gzguevara/amex-smart-brute-force-features). If you have any further questions - let me know! \n\nNOTE! including 20 features, which are base on P_2_last is useless! The information contained in those 20 features will be too similar! \n\nImagine you are at police. Mrs. Smith has been found dead in the forrest. You are telling the police that you saw Mr. Smith on the day before in the forrest. On the next day you go to the police again and tell them: \"I saw Mr. Smith and he had a black shoes\". On the next day you go to the police again and tell them: \"I saw Mr. Smith and he had green pants\". etc... The information you are givin to the police is helpfull, yes! But all those little variation do not add significant *new* information! \n\n**You need to find features with high information, but also with different information!**","metadata":{}},{"cell_type":"markdown","source":"Hee are the resulting new features:","metadata":{}},{"cell_type":"code","source":"pd.read_parquet('../input/amex-smart-brute-force-features')","metadata":{"execution":{"iopub.status.busy":"2022-08-06T10:52:57.733016Z","iopub.execute_input":"2022-08-06T10:52:57.733772Z","iopub.status.idle":"2022-08-06T10:53:00.761656Z","shell.execute_reply.started":"2022-08-06T10:52:57.733734Z","shell.execute_reply":"2022-08-06T10:53:00.760042Z"},"trusted":true},"execution_count":null,"outputs":[]}]}