{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Intorduction","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"This notebook aims at getting a good insight in the data for the PorteSeguro competition. Besides that it gives some tips and tricks to prepare your data for modeling. the notebook consists of the following main sections:\n","metadata":{}},{"cell_type":"markdown","source":"The notebook consists of the following main section.","metadata":{}},{"cell_type":"markdown","source":"# Loading apckages","metadata":{"execution":{"iopub.status.busy":"2022-07-05T05:56:09.235336Z","iopub.execute_input":"2022-07-05T05:56:09.235858Z","iopub.status.idle":"2022-07-05T05:56:09.260286Z","shell.execute_reply.started":"2022-07-05T05:56:09.23576Z","shell.execute_reply":"2022-07-05T05:56:09.259476Z"}}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import PolynomialFeatures\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.feature_selection import VarianceThreshold\nfrom sklearn.feature_selection import SelectFromModel\nfrom sklearn.utils import shuffle\nfrom sklearn.ensemble import RandomForestClassifier\n\npd.set_option('display.max_columns', 100)\npd.set_option('display.max_rows', 100)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:37.060414Z","iopub.execute_input":"2022-07-07T06:09:37.060908Z","iopub.status.idle":"2022-07-07T06:09:38.101205Z","shell.execute_reply.started":"2022-07-07T06:09:37.060810Z","shell.execute_reply":"2022-07-07T06:09:38.100179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Loading data","metadata":{}},{"cell_type":"code","source":"# DEBUG = True","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:38.480426Z","iopub.execute_input":"2022-07-07T06:09:38.482928Z","iopub.status.idle":"2022-07-07T06:09:38.488380Z","shell.execute_reply.started":"2022-07-07T06:09:38.482854Z","shell.execute_reply":"2022-07-07T06:09:38.487408Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# if DEBUG:\n#     NROWS = 500000\n# else:\n#     NROWS = None","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:38.917559Z","iopub.execute_input":"2022-07-07T06:09:38.918360Z","iopub.status.idle":"2022-07-07T06:09:38.922996Z","shell.execute_reply.started":"2022-07-07T06:09:38.918312Z","shell.execute_reply":"2022-07-07T06:09:38.921871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* train: 600,000\n* test: 900,000","metadata":{}},{"cell_type":"code","source":"'''\n# 데이터가 많을 때, 모델만들때 조금만 사용해서 빨리 돌아가게\n%%time\ntrain= pd.read_csv('../input/porto-seguro-safe-driver-prediction/train.csv', nrows= NROWS)\ntest= pd.read_csv('../input/porto-seguro-safe-driver-prediction/test.csv', nrows= NROWS)\ntrain= train.sample(frac= 0.1)   # 랜덤 샘플링 10%\n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:39.820202Z","iopub.execute_input":"2022-07-07T06:09:39.820897Z","iopub.status.idle":"2022-07-07T06:09:39.830643Z","shell.execute_reply.started":"2022-07-07T06:09:39.820862Z","shell.execute_reply":"2022-07-07T06:09:39.829392Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\ntrain= pd.read_csv('../input/porto-seguro-safe-driver-prediction/train.csv')\ntest= pd.read_csv('../input/porto-seguro-safe-driver-prediction/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:40.422630Z","iopub.execute_input":"2022-07-07T06:09:40.423072Z","iopub.status.idle":"2022-07-07T06:09:52.914294Z","shell.execute_reply.started":"2022-07-07T06:09:40.422986Z","shell.execute_reply":"2022-07-07T06:09:52.910136Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data at first sight","metadata":{}},{"cell_type":"markdown","source":"Here is an excerpt of the the data description for the competition:\n\n   * Features that belong to **similar groupings are tagged** as such in the feature names (e.g., ind, reg, car, calc).\n   * Feature names include the postfix **bin** to indicate binary features and **cat** to indicate categorical features.\n   * Features **without these designations are either continuous or ordinal**.\n   * Values of **-1** indicate that the feature was **missing** from the observation.\n   * The **target** columns signifies whether or not a claim was filed for that policy holder.\n   \nOk, that's important information to get us started. Let's have a quick look at the first and last rows to confirm all of this.","metadata":{}},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:52.920411Z","iopub.execute_input":"2022-07-07T06:09:52.921974Z","iopub.status.idle":"2022-07-07T06:09:53.052903Z","shell.execute_reply.started":"2022-07-07T06:09:52.921778Z","shell.execute_reply":"2022-07-07T06:09:53.049083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.tail()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:53.057888Z","iopub.execute_input":"2022-07-07T06:09:53.060487Z","iopub.status.idle":"2022-07-07T06:09:53.201446Z","shell.execute_reply.started":"2022-07-07T06:09:53.060337Z","shell.execute_reply":"2022-07-07T06:09:53.198121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We indeed see the following\n\n   * binary variables\n   * categorical variables of which the category values are integers\n   * other variables with integer or float values\n   * variables with -1 representing missing values\n   * the target variable and an ID variable\n   ","metadata":{}},{"cell_type":"markdown","source":"Let's look at the number of rows and columns in the train data.","metadata":{}},{"cell_type":"code","source":"train.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:53.209683Z","iopub.execute_input":"2022-07-07T06:09:53.213547Z","iopub.status.idle":"2022-07-07T06:09:53.239891Z","shell.execute_reply.started":"2022-07-07T06:09:53.213312Z","shell.execute_reply":"2022-07-07T06:09:53.237235Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols = [col for col in train.columns if 'cat' in col]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:53.243979Z","iopub.execute_input":"2022-07-07T06:09:53.245688Z","iopub.status.idle":"2022-07-07T06:09:53.260382Z","shell.execute_reply.started":"2022-07-07T06:09:53.245502Z","shell.execute_reply":"2022-07-07T06:09:53.257184Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_cols","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:53.264205Z","iopub.execute_input":"2022-07-07T06:09:53.265937Z","iopub.status.idle":"2022-07-07T06:09:53.287917Z","shell.execute_reply.started":"2022-07-07T06:09:53.265793Z","shell.execute_reply":"2022-07-07T06:09:53.284413Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in cat_cols:\n    print(col, train[col].nunique())","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:53.291723Z","iopub.execute_input":"2022-07-07T06:09:53.294949Z","iopub.status.idle":"2022-07-07T06:09:53.458692Z","shell.execute_reply.started":"2022-07-07T06:09:53.294835Z","shell.execute_reply":"2022-07-07T06:09:53.454133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:53.466296Z","iopub.execute_input":"2022-07-07T06:09:53.468192Z","iopub.status.idle":"2022-07-07T06:09:53.494199Z","shell.execute_reply.started":"2022-07-07T06:09:53.467997Z","shell.execute_reply":"2022-07-07T06:09:53.489422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.drop_duplicates()\ntrain.shape","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:53.498766Z","iopub.execute_input":"2022-07-07T06:09:53.501769Z","iopub.status.idle":"2022-07-07T06:09:55.282913Z","shell.execute_reply.started":"2022-07-07T06:09:53.501482Z","shell.execute_reply":"2022-07-07T06:09:55.281781Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.306231Z","iopub.execute_input":"2022-07-07T06:09:55.308409Z","iopub.status.idle":"2022-07-07T06:09:55.415713Z","shell.execute_reply.started":"2022-07-07T06:09:55.308364Z","shell.execute_reply":"2022-07-07T06:09:55.414410Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Metadata\n\nTo facilitate the data management, we'll store meta-information about the variables in a DataFrame. This will be helpful when we want to select specific variables for analysis, visualization, modeling, ...\n\nConcretely we will store:\n\n* **role**: input, ID, target\n* **level**: nominal, interval, ordinal, binary\n* **keep**: True or False\n* **dtype**: int, float, str","metadata":{}},{"cell_type":"code","source":"data = []\nfor f in train.columns:\n    # Defining the role\n    if f == 'target':\n        role = 'target'\n    elif f == 'id':\n        role = 'id'\n    else:\n        role = 'input'\n         \n    # Defining the level\n    if 'bin' in f or f == 'target':\n        level = 'binary'\n    elif 'cat' in f or f == 'id':\n        level = 'nominal'\n    elif train[f].dtype == float:\n        level = 'interval'\n    elif train[f].dtype == int:\n        level = 'ordinal'\n        \n    # Initialize keep to True for all variables except for id\n    keep = True\n    if f == 'id':\n        keep = False\n    \n    # Defining the data type \n    dtype = train[f].dtype\n    \n    # Creating a Dict that contains all the metadata for the variable\n    f_dict = {\n        'varname': f,\n        'role': role,\n        'level': level,\n        'keep': keep,\n        'dtype': dtype\n    }\n    data.append(f_dict)\n    \nmeta = pd.DataFrame(data, columns=['varname', 'role', 'level', 'keep', 'dtype'])\nmeta.set_index('varname', inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.417458Z","iopub.execute_input":"2022-07-07T06:09:55.418260Z","iopub.status.idle":"2022-07-07T06:09:55.432109Z","shell.execute_reply.started":"2022-07-07T06:09:55.418213Z","shell.execute_reply":"2022-07-07T06:09:55.431122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.433932Z","iopub.execute_input":"2022-07-07T06:09:55.434455Z","iopub.status.idle":"2022-07-07T06:09:55.463534Z","shell.execute_reply.started":"2022-07-07T06:09:55.434424Z","shell.execute_reply":"2022-07-07T06:09:55.462658Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Example to extract all nominal variables that are not dropped","metadata":{}},{"cell_type":"code","source":"meta.loc[(meta.level== 'nominal') & (meta.keep)].index","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.464814Z","iopub.execute_input":"2022-07-07T06:09:55.465845Z","iopub.status.idle":"2022-07-07T06:09:55.477043Z","shell.execute_reply.started":"2022-07-07T06:09:55.465799Z","shell.execute_reply":"2022-07-07T06:09:55.476145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below the number of variables per role and level are displayed","metadata":{}},{"cell_type":"code","source":"pd.DataFrame({'count': meta.groupby(['role', 'level'])['role'].size()}).reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.478450Z","iopub.execute_input":"2022-07-07T06:09:55.479143Z","iopub.status.idle":"2022-07-07T06:09:55.498503Z","shell.execute_reply.started":"2022-07-07T06:09:55.479112Z","shell.execute_reply":"2022-07-07T06:09:55.496840Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta.groupby(['role', 'level'])['role'].size()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.500874Z","iopub.execute_input":"2022-07-07T06:09:55.501741Z","iopub.status.idle":"2022-07-07T06:09:55.518350Z","shell.execute_reply.started":"2022-07-07T06:09:55.501683Z","shell.execute_reply":"2022-07-07T06:09:55.517094Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Descriptive statistics","metadata":{}},{"cell_type":"markdown","source":"We can also apply the descirbe method on the dataframe. However, it doesn't make much sense to calculate the mean, std, ... on categorical variables and the id variable. We'll explore the categorical variables visually later.\n\nThanks to our meta file we can easily select the variables on which we want to compute the decriptive statistics. To keep things clear, we'll do this per data type.","metadata":{}},{"cell_type":"markdown","source":"### Interval variables","metadata":{}},{"cell_type":"code","source":"categorical_feats= [col for col in train.columns if 'cat' in col]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.520325Z","iopub.execute_input":"2022-07-07T06:09:55.520938Z","iopub.status.idle":"2022-07-07T06:09:55.526691Z","shell.execute_reply.started":"2022-07-07T06:09:55.520885Z","shell.execute_reply":"2022-07-07T06:09:55.525718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"v= meta[(meta.level == 'interval') & (meta.keep)].index\ntrain[v].describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.529230Z","iopub.execute_input":"2022-07-07T06:09:55.529828Z","iopub.status.idle":"2022-07-07T06:09:55.878391Z","shell.execute_reply.started":"2022-07-07T06:09:55.529789Z","shell.execute_reply":"2022-07-07T06:09:55.877244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### reg variables\n\n* only ps_reg_03 has missing values\n* the range (min to max) differs between the variables. We could apply scaling (e.g. StandardSclaer), but it depends on the classifier\nwe will ant to use.\n\n#### car variables\n\n* ps_car_12 and ps_car_15 have missing values\n* again, the range differs and we could apply sacling.\n\n#### calc variables\n\n* no missing values\n* this seems to be some kind of ratio as the maximum in 0.9\n* all three _calc variables have very similar distributions\n\n**Overall**, we can see that the range of the interval variables is rather small.\nPerhaps some transformation (e.g. log) is already applied in order to anonymize the data?","metadata":{}},{"cell_type":"markdown","source":"### Ordinal variables","metadata":{}},{"cell_type":"code","source":"v= meta[(meta.level== 'ordinal') & (meta.keep)].index\ntrain[v].describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:55.879675Z","iopub.execute_input":"2022-07-07T06:09:55.879977Z","iopub.status.idle":"2022-07-07T06:09:56.312141Z","shell.execute_reply.started":"2022-07-07T06:09:55.879949Z","shell.execute_reply":"2022-07-07T06:09:56.310876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* Only one missing variable: ps_car_11\n* We could apply scaling to deal with the different ranges","metadata":{}},{"cell_type":"markdown","source":"### Binary variables","metadata":{}},{"cell_type":"code","source":"v= meta[(meta.level== 'binary') & (meta.keep)].index\ntrain[v].describe()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:56.313873Z","iopub.execute_input":"2022-07-07T06:09:56.315115Z","iopub.status.idle":"2022-07-07T06:09:56.748997Z","shell.execute_reply.started":"2022-07-07T06:09:56.315063Z","shell.execute_reply":"2022-07-07T06:09:56.747811Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* A priori in the train data is 3.645%, which is **strongly imbalanced.**\n* From the means we can conclude that for most variables the value is zero in most cases.","metadata":{}},{"cell_type":"markdown","source":"### Handling Imbalanced classes\n\nAs we mentioned above  the proportion of recored with target= 1 is far less than target= 0. This can lead to a model that has great accuracy but does have any added value in practice. Two possible strategies to deal with this problem are:\n\n* oversampling records with target= 1\n* undersampling records with target= 0\n\nThere are many more strategies of course and\nMachineLearningMastery.com gives a nice [overview](https://www.kaggleusercontent.com/kf/1676042/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..L0qPOakH4SvEVtrGYi_tHA.YF_t3f_NbfEkaQqBdT-zVShCDOiT8Mle9L7iAEa8VrTYvnjSCwKT7pP3PXBlUTCjbkrOWg2xNZOCz-S1h_oA-ugGuUVamq_O_DTayK5bUUCgkoqRm4HWGBWUzG3Rs_DgXSqG0_8eeWfqTVFsGzRTOHDFDdmKTASUJ35HOZoi2LngduOhTkw_nEDhlBgAVtNab2I4MN8LW95uM7yI-rWTjSz3r7Mp4z3SnEFZFOA8AvsSJU9ybxk6gFrq5JHKfUoC4a2RH2sG6nQxLW7E-VgWTLmUoRVBnOYRLh1Rke3g8_XtvQKoHzTH2JSl3La51d2YGo2NNHqbiDwhLm3dQpPVYtI3TnkxRkDzlfavgcV-GpJhuUqa8JiR1SKbh-y0NeefmltqlAZl6rYeSyN5E0CWTgMWZ9yFKkKiGyLs5vyni-g40KuE8jfQSfnP7drBsZ0L94EAAgdOLTp2-C7-Ck0j-ZR35ynwKWtCiV8pWhNrQWGOxZ8xYsJS001-Y5NvR9hFSCxbEi1jyiXH2Uo-Q33rVtBANVbNo5VXn4Is4OHeyHE3Yp-SQkBZ2VPa99MXPOpIsIKLyi6eSTrA9Qn1XiFvyK8f4LZCW5wAkI_PoRuz9fS3gq8o24X9g7yyZDRdtp7L1P9grYBWDfsVZjtKsJMxzjlmOjJ1iX4AxROIBNp7O5K4P3rK4czdmADuL3Pc0CtQ.3vibrqjVqSf-Sq6ntyVSOQ/(https://machinelearningmastery.com/tactics-to-combat-imbalanced-classes-in-your-machine-learning-dataset/). As we have a rather large training set, we can go for **undersampling**.","metadata":{}},{"cell_type":"code","source":"desired_apriori= 0.10\n\n# Get the indices per target value\nidx_0 = train[train.target == 0].index\nidx_1 = train[train.target == 1].index\n\n# Get original number of records per target value\nnb_0 = len(train.loc[idx_0])\nnb_1 = len(train.loc[idx_1])\n\n# Calculate the undersampling rate and resulting nuber of records with target= 0\nundersampling_rate = ((1 - desired_apriori)* nb_1)/ (nb_0* desired_apriori)\nundersampling_nb_0 = int(undersampling_rate* nb_0)\nprint('Rate to undersampli records with target= 0: {}'.format(undersampling_rate))\nprint('Number of records with target= 0 after undersampling: {}'.format(undersampling_nb_0))\n\n# Randomly select records with target = 0 to get at the desired apriori\nusdersampling_idx = shuffle(idx_0, random_state= 37, n_samples= undersampling_nb_0)\n\n# Construct list with remaining indices\nidx_list= list(usdersampling_idx) + list(idx_1)\n\n# Return undersample data frame\ntrain= train.loc[idx_list].reset_index(drop= True)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:56.750572Z","iopub.execute_input":"2022-07-07T06:09:56.750916Z","iopub.status.idle":"2022-07-07T06:09:57.526483Z","shell.execute_reply.started":"2022-07-07T06:09:56.750886Z","shell.execute_reply":"2022-07-07T06:09:57.525453Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"train","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:57.528200Z","iopub.execute_input":"2022-07-07T06:09:57.528580Z","iopub.status.idle":"2022-07-07T06:09:57.597328Z","shell.execute_reply.started":"2022-07-07T06:09:57.528538Z","shell.execute_reply":"2022-07-07T06:09:57.596272Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Data Quality Checks","metadata":{}},{"cell_type":"markdown","source":"#### Checking missing values\nMissing are represented as -1","metadata":{}},{"cell_type":"code","source":"vars_with_missing= []\n\nfor f in train.columns:\n    missings = train.loc[train[f] == -1][f].count()\n    if missings > 0:\n        vars_with_missing.append(f)\n        missings_perc = missings / train.shape[0]\n        \n        print('Variable {} has {} records ({:.2%}) with missing values'.\n             format(f, missings, missings_perc))\n        \nprint('In total, there are {} variables with missing values'. format(len(vars_with_missing)))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:57.598510Z","iopub.execute_input":"2022-07-07T06:09:57.599435Z","iopub.status.idle":"2022-07-07T06:09:58.140869Z","shell.execute_reply.started":"2022-07-07T06:09:57.599401Z","shell.execute_reply":"2022-07-07T06:09:58.139312Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* **ps_car_03_cat and ps_car_05_cat** have a large proportion of records with missing values. Remove these variables\n* For the other categorical variables with missing values, we can leave the missing value -1 as such.\n* **ps_reg_03** (continuous) has missing values for 18% of all records. Replace by the mean.\n* **ps_car_11** (ordinary) has only 6 records with missing values. Replace by the mode.\n* **ps_car_12** (continuous) has only 1 records with missing value. Replace bt the mode.\n* **ps_car_14** (continuous) has missing values for 7% of all records. Replace by the mean.","metadata":{}},{"cell_type":"code","source":"train[['ps_car_03_cat', 'target']].groupby('ps_car_03_cat').mean()\n# 학인해보니 지워도 되겠다","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:58.142672Z","iopub.execute_input":"2022-07-07T06:09:58.143027Z","iopub.status.idle":"2022-07-07T06:09:58.165073Z","shell.execute_reply.started":"2022-07-07T06:09:58.142972Z","shell.execute_reply":"2022-07-07T06:09:58.163936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Dropping the variavles with too many missing values\nvars_to_drop = ['ps_car_03_cat', 'ps_car_05_cat']\ntrain.drop(vars_to_drop, inplace= True, axis = 1)\nmeta.loc[(vars_to_drop), 'keep'] = False  # Updating the meta","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:58.166481Z","iopub.execute_input":"2022-07-07T06:09:58.167472Z","iopub.status.idle":"2022-07-07T06:09:58.423998Z","shell.execute_reply.started":"2022-07-07T06:09:58.167437Z","shell.execute_reply":"2022-07-07T06:09:58.422740Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Imputing with the mean or mode\nmean_imp = SimpleImputer(missing_values= -1, strategy= 'mean')\nmode_imp = SimpleImputer(missing_values= -1, strategy= 'most_frequent')\ntrain['ps_reg_03']= mean_imp.fit_transform(train[['ps_reg_03']]).ravel()\ntrain['ps_car_12']= mean_imp.fit_transform(train[['ps_car_12']]).ravel()\ntrain['ps_car_14']= mean_imp.fit_transform(train[['ps_car_14']]).ravel()\ntrain['ps_car_11']= mode_imp.fit_transform(train[['ps_car_11']]).ravel()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:58.425446Z","iopub.execute_input":"2022-07-07T06:09:58.426263Z","iopub.status.idle":"2022-07-07T06:09:58.479039Z","shell.execute_reply.started":"2022-07-07T06:09:58.426217Z","shell.execute_reply":"2022-07-07T06:09:58.478142Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Checking the cardinality of the categorical variables\n\nCardinality refers to the number of different values in a variable. As we will create dummy variables from the categorical variables later on, we need to check whether there are variables with many distinct values. We should handle these variables differently as they as they would result in many dummy variables","metadata":{}},{"cell_type":"code","source":"v = meta[(meta.level== 'nominal') & (meta.keep)].index\n\nfor f in v:\n    dist_values = train[f].value_counts().shape[0]   # == train[f].nunique()\n    print('Variable {} has {} distinct values'.format(f, dist_values))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:58.480557Z","iopub.execute_input":"2022-07-07T06:09:58.481160Z","iopub.status.idle":"2022-07-07T06:09:58.516273Z","shell.execute_reply.started":"2022-07-07T06:09:58.481128Z","shell.execute_reply":"2022-07-07T06:09:58.515456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Only **ps_car_11_cat** has many distinct values, although it is still reasonable.\n\n**EDIT**: nickycan made an excellent remark on the fact that my first solution could lead to data leakage. He also pointed me to another kernel made by oliver which deals with that. I therefore replaced this part with the kernel of oliver. All credits go to him. It si so great what you can learn by participating in the Kaggle competitions.","metadata":{}},{"cell_type":"code","source":"# * 어려운 것 만날때 하나하나 뜯어서 찾아보자\n'''\ntrn_series = train[\"ps_car_11_cat\"]\ntst_series = test[\"ps_car_11_cat\"]\ntarget = train.target\nmin_samples_leaf = 100\nsmoothing = 10\nnoise_level = 0.01\n\nassert len(trn_series) == len(target)\nassert trn_series.name == tst_series.name\ntemp = pd.concat([trn_series, target], axis =1)\ntemp.groupby(by=trn_series.name)[target.name].agg(['mean', 'count'])\n\n'''\n# 이런것도 가능하다 중요함! 많이사용!\n# df max_min(x):\n#     return x.max()- x.min()\n\n# temp.groupby(by=trn_series.name)[target.name].agg(['mean', 'count', max_min])\n'''\n# Compute target mean \naverages = temp.groupby(by= trn_series.name)[target.name].agg(['mean', 'count'])\n\n# Compute smoothing\nsmoothing = 1 / (1 + np.exp(-(averages[\"count\"] - min_samples_leaf) / smoothing))\n\n# Apply average function to all target data\nprior = target.mean()\n\n # The bigger the count the less full_avg is taken into account\naverages[target.name] = prior * (1 - smoothing) + averages[\"mean\"] * smoothing\naverages.drop(['mean', 'count'], axis = 1, inplace= True)\n\n# Apply averages to trn and tst series\nft_trn_series = pd.merge(trn_series.to_frame(trn_series.name),\n        averages.reset_index().\n         rename(columns= {'index': target.name, target.name: 'average'}),\n        on = trn_series.name, how = 'left')['average'].rename(trn_series.name + '_mean').fillna(prior)\n\n# pd.merge does not keep the index so restore it\nft_trn_series.index = trn_series.index\nft_tst_series = pd.merge(tst_series.to_frame(trn_series.name),\n        averages.reset_index().\n         rename(columns= {'index': target.name, target.name: 'average'}),\n        on = tst_series.name, how = 'left')['average'].rename(trn_series.name + '_mean').fillna(prior)\nft_tst_series.index= tst_series.index\n\n# 노이즈를 준다 1이면 너무크니깐 \n0.01 * np.random.randn(ft_trn_series.shape[0])\n'''","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:58.517743Z","iopub.execute_input":"2022-07-07T06:09:58.518375Z","iopub.status.idle":"2022-07-07T06:09:58.527633Z","shell.execute_reply.started":"2022-07-07T06:09:58.518342Z","shell.execute_reply":"2022-07-07T06:09:58.526432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Script by https://www.kaggle.com/ogrellier\n# Code: https://www.kaggle.com/ogrellier/python-target-encoding-for-categorical-features\ndef add_noise(series, noise_level):\n    return series * (1 + noise_level * np.random.randn(len(series)))\n\ndef target_encode(trn_series=None, \n                  tst_series=None, \n                  target=None, \n                  min_samples_leaf=1, \n                  smoothing=1,\n                  noise_level=0):\n    \"\"\"\n    Smoothing is computed like in the following paper by Daniele Micci-Barreca\n    https://kaggle2.blob.core.windows.net/forum-message-attachments/225952/7441/high%20cardinality%20categoricals.pdf\n    \n    trn_series : training categorical feature as a pd.Series\n    tst_series : test categorical feature as a pd.Series\n    target : target data as a pd.Series\n    min_samples_leaf (int) : minimum samples to take category average into account\n    smoothing (int) : smoothing effect to balance categorical average vs prior  \n    \"\"\" \n    \n    assert len(trn_series) == len(target)\n    assert trn_series.name == tst_series.name\n    temp = pd.concat([trn_series, target], axis = 1)\n    \n    # Compute target mean\n    averages= temp.groupby(by= trn_series.name)[target.name].agg(['mean', 'count'])\n    \n    # Compute smoothing\n    smoothing = 1 / (1 + np.exp(-(averages['count'] - min_samples_leaf) / smoothing))\n    \n    # Apply average function to all target data\n    prior = target.mean()\n    \n    # The bigger the count the less full_vag is taken into account\n    averages[target.name] = prior * (1 - smoothing) + averages['mean'] * smoothing\n    averages.drop(['mean', 'count'], axis= 1, inplace= True)\n    \n    # Apply averages to trn and tst seires\n    ft_trn_series = pd.merge(\n            trn_series.to_frame(trn_series.name),\n            averages.reset_index().rename(columns = \n                                         {'index': target.name, target.name: 'average'}),\n            on = trn_series.name,\n            how = 'left')['average'].rename(trn_series.name + '_mean').fillna(prior)\n    # pd.merge does not keep the index so restore it\n    ft_trn_series.index = trn_series.index\n    \n    ft_tst_series = pd.merge(\n            tst_series.to_frame(tst_series.name),\n            averages.reset_index().rename(columns = \n                                         {'index': target.name, target.name: 'average'}),\n            on = tst_series.name,\n            how = 'left')['average'].rename(trn_series.name + '_mean').fillna(prior)\n    # pd.merge does no keep the index so restore it\n    ft_tst_series.index = tst_series.index\n    return add_noise(ft_trn_series, noise_level), add_noise(ft_tst_series, noise_level)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:58.536556Z","iopub.execute_input":"2022-07-07T06:09:58.537029Z","iopub.status.idle":"2022-07-07T06:09:58.554868Z","shell.execute_reply.started":"2022-07-07T06:09:58.536979Z","shell.execute_reply":"2022-07-07T06:09:58.553563Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_encoded, test_encoded = target_encode(train['ps_car_11_cat'],\n                                           test['ps_car_11_cat'],\n                                           target = train.target,\n                                           min_samples_leaf= 100,\n                                           smoothing = 10,\n                                           noise_level= 0.01)\n\ntrain['ps_car_11_cat_te'] = train_encoded\ntrain.drop('ps_car_11_cat', axis= 1, inplace= True)\nmeta.loc['ps_car_11_cat', 'keep'] = False  # Updating the meta\ntest['ps_car_11_cat_te'] = test_encoded\ntest.drop('ps_car_11_cat', axis = 1, inplace= True)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:09:58.556331Z","iopub.execute_input":"2022-07-07T06:09:58.557165Z","iopub.status.idle":"2022-07-07T06:10:00.029221Z","shell.execute_reply.started":"2022-07-07T06:09:58.557133Z","shell.execute_reply":"2022-07-07T06:10:00.027996Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Exploratory Data Visualization","metadata":{}},{"cell_type":"markdown","source":"#### Categorical variables\nLet's look into the categorical variables and the proportion of customers with target = 1","metadata":{}},{"cell_type":"code","source":"v = meta[(meta.level == 'nominal') & (meta.keep)].index\n\nfor f in v:\n    plt.figure()\n    fig, ax = plt.subplots(figsize=(20,10))\n    \n    # Calculate the percentage of target=1 per category value\n    cat_perc = train[[f, 'target']].groupby([f],as_index=False).mean()\n    cat_perc.sort_values(by='target', ascending=False, inplace=True)\n    \n    # Bar plot\n    # Order the bars descending on target mean\n    sns.barplot(ax=ax, x=f, y='target', data=cat_perc, order=cat_perc[f])\n    sns.set(font_scale= 3)\n    plt.ylabel('% target')\n    plt.xlabel(f)\n    plt.tick_params(axis='both', which='major')\n    plt.show();","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:00.030876Z","iopub.execute_input":"2022-07-07T06:10:00.031487Z","iopub.status.idle":"2022-07-07T06:10:03.066211Z","shell.execute_reply.started":"2022-07-07T06:10:00.031453Z","shell.execute_reply":"2022-07-07T06:10:03.065387Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"* sample 수 확인 필요하다!","metadata":{}},{"cell_type":"markdown","source":"As we can see from the variables **with missing values**, it is a good idea to keep the missing valuse as a separate category value, instead of replacing them by the mode for instance. The customers with a missing value appear to have a much higher (in some cases much lower probability to ask for an insurance claim.","metadata":{}},{"cell_type":"markdown","source":"#### Interval variables\n\nChecking the correlations between inerval variables. A heatmap is a good way to visualize the correlation between variables. The code below is based on [an example by Michael Waskom](http://seaborn.pydata.org/examples/many_pairwise_correlations.html)","metadata":{}},{"cell_type":"code","source":"def corr_heatmap(v):\n    correlations = train[v].corr()\n    \n    # Create color map ranging between two colors\n    cmap = sns.diverging_palette(220, 10, as_cmap= True)\n    \n    fig, ax= plt.subplots(figsize= (10,10))\n    sns.heatmap(correlations, cmap= cmap, vmax= 1.0, center= 0, fmt= '.2f',\n               square = True, linewidths= .5, annot= True,\n               cbar_kws= {'shrink': .75})\n    sns.set(font_scale= 1)\n    plt.show()\n    \nv= meta[(meta.level== 'interval') & (meta.keep)].index\ncorr_heatmap(v)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:03.067405Z","iopub.execute_input":"2022-07-07T06:10:03.068088Z","iopub.status.idle":"2022-07-07T06:10:04.142852Z","shell.execute_reply.started":"2022-07-07T06:10:03.068050Z","shell.execute_reply":"2022-07-07T06:10:04.141774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are a strong correlations between the variables:\n\n* ps_reg_02 and ps_reg_03 (0.7)\n* ps_car_12 and ps_car13 (0.67)\n* ps_car_12 and ps_car14 (0.58)\n* ps_car_13 and ps_car15 (0.53)\n\nSeaborn has some handy plots to visualize the (linear) relationship\nbetween variables. We could use a pariplot to visualize the relationship\nbetween the variables. But because the heatmap already showed the limited number of correlated variables, we'll look at each of the highly correlated variables separately.\n\n**NOTE**: I take a sample of the train data to speed up the process.","metadata":{}},{"cell_type":"code","source":"s = train.sample(frac= 0.1)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:04.144280Z","iopub.execute_input":"2022-07-07T06:10:04.145316Z","iopub.status.idle":"2022-07-07T06:10:04.181203Z","shell.execute_reply.started":"2022-07-07T06:10:04.145273Z","shell.execute_reply":"2022-07-07T06:10:04.180214Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### ps_reg_02 and ps_reg_03\n\nAs the regression line shows, there is a linear relationship between these variables. Thanks to the hue parameter we can see that the regression lines for target= 0 and target= 1 are the same.","metadata":{}},{"cell_type":"code","source":"sns.lmplot(x= 'ps_reg_02', y= 'ps_reg_03', data= s, hue= 'target', palette= 'Set1', \n          scatter_kws= {'alpha': 0.3})\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:04.182647Z","iopub.execute_input":"2022-07-07T06:10:04.182984Z","iopub.status.idle":"2022-07-07T06:10:06.325500Z","shell.execute_reply.started":"2022-07-07T06:10:04.182955Z","shell.execute_reply":"2022-07-07T06:10:06.324565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### ps_car_12 and ps_car_13","metadata":{}},{"cell_type":"code","source":"sns.lmplot(x= 'ps_car_12', y= 'ps_car_13', data= s, hue= 'target', palette= 'Set1',\n          scatter_kws= {'alpha': 0.3})\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:06.326571Z","iopub.execute_input":"2022-07-07T06:10:06.327421Z","iopub.status.idle":"2022-07-07T06:10:08.319289Z","shell.execute_reply.started":"2022-07-07T06:10:06.327390Z","shell.execute_reply":"2022-07-07T06:10:08.318086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### ps_car_12 and ps_car_14","metadata":{}},{"cell_type":"code","source":"sns.lmplot(x= 'ps_car_12', y= 'ps_car_14', data= s, hue= 'target', palette= 'Set1',\n          scatter_kws= {'alpha': 0.3})\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:08.320965Z","iopub.execute_input":"2022-07-07T06:10:08.321301Z","iopub.status.idle":"2022-07-07T06:10:10.371863Z","shell.execute_reply.started":"2022-07-07T06:10:08.321254Z","shell.execute_reply":"2022-07-07T06:10:10.370759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"##### ps_car_13 and ps_car_15","metadata":{}},{"cell_type":"code","source":"sns.lmplot(x= 'ps_car_15', y= 'ps_car_13', data= s, hue= 'target', palette= 'Set1',\n          scatter_kws= {'alpha': 0.3})\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:10.373498Z","iopub.execute_input":"2022-07-07T06:10:10.374582Z","iopub.status.idle":"2022-07-07T06:10:12.463577Z","shell.execute_reply.started":"2022-07-07T06:10:10.374536Z","shell.execute_reply":"2022-07-07T06:10:12.462383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Allright, so now what? How can we decide which of the correlated variablese to keep?\nWe could perform Principal Component Analysis (PCA) on the variables to reduce the dimensions. In the AllState Claims Severity Competition I made [this kernel](https://www.kaggle.com/bertcarremans/reducing-number-of-numerical-features-with-pca) to do that. But as the number of correlated variables is rather low, we will let the model do the heavy-lifting.","metadata":{}},{"cell_type":"markdown","source":"#### Checking the correlations between ordinal variables","metadata":{}},{"cell_type":"code","source":"v= meta[(meta.level == 'ordinal') & (meta.keep)].index\ncorr_heatmap(v)","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:12.464772Z","iopub.execute_input":"2022-07-07T06:10:12.465087Z","iopub.status.idle":"2022-07-07T06:10:13.961868Z","shell.execute_reply.started":"2022-07-07T06:10:12.465056Z","shell.execute_reply":"2022-07-07T06:10:13.961056Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For the ordinal variables we do not see many correlationss. We could, on the other hand, look at how the distributions are when grouping by the target value.","metadata":{}},{"cell_type":"markdown","source":"### Feature engineering\n","metadata":{}},{"cell_type":"markdown","source":"#### Creating dummy variables\n\nThe values of the categorical variables do not represent any order or magnitude. For instance, category 2 is not twice the value of category 1.\nTherefore we can create dummy variables to deal with that. We drop the first dummy variables as this information can be derived from the other dummy variables generated for the categories of the original variable.","metadata":{}},{"cell_type":"code","source":"v = meta[(meta.level == 'nominal') & (meta.keep)].index\nprint('Before dummification we have {} variables in train'.format(train.shape[1]))\ntrain = pd.get_dummies(train, columns=v, drop_first=True)\nprint('After dummification we have {} variables in train'.format(train.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:10:13.963081Z","iopub.execute_input":"2022-07-07T06:10:13.964008Z","iopub.status.idle":"2022-07-07T06:10:14.260193Z","shell.execute_reply.started":"2022-07-07T06:10:13.963955Z","shell.execute_reply":"2022-07-07T06:10:14.259045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So creating dummy vairables adds 52 variables to the training set.","metadata":{}},{"cell_type":"markdown","source":"#### Creating interaction variables","metadata":{}},{"cell_type":"code","source":"v = meta[(meta.level == 'interval') & (meta.keep)].index\npoly = PolynomialFeatures(degree= 2, interaction_only= False, include_bias= False)\ninteractions = pd.DataFrame(data= poly.fit_transform(train[v]),\n                           columns= poly.get_feature_names(v))\ninteractions.drop(v, axis= 1, inplace= True)  # Remove the original columns\n\n# Concat the interaction variables to the train data\nprint('Before creating interactions we have {} variables in train'.format(train.shape[1]))\ntrain = pd.concat([train, interactions], axis= 1)\nprint('Before creating interactions we have {} variables in train'.format(train.shape[1]))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:27:03.827591Z","iopub.execute_input":"2022-07-07T06:27:03.828003Z","iopub.status.idle":"2022-07-07T06:27:04.304232Z","shell.execute_reply.started":"2022-07-07T06:27:03.827968Z","shell.execute_reply":"2022-07-07T06:27:04.303037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This adds extra interaction variables to the train data. Thanks to the get_feature_names method we can assign column names to these new variables","metadata":{}},{"cell_type":"markdown","source":"### Feature selection","metadata":{}},{"cell_type":"markdown","source":"#### Removing features with low or zero variance","metadata":{}},{"cell_type":"markdown","source":"Personally, I prefer to let the classifier algorithm chose which features to keep.\nBut there is one thing that we can do ourselves. Taht is removing features with no or a very low variance. Sklearn has as handy method to do that: **VarianceThreshold**. By defauly it removes features with zero variance. This will not be applicable for this competition as we saw there are no zero-varaince in the previous steps. But if we would remove features with less than 1% variance, we would remove 31 vairables.","metadata":{}},{"cell_type":"code","source":"selector = VarianceThreshold(threshold= .01)\nselector.fit(train.drop(['id', 'target'], axis= 1))  # Fit to train without id and target variables\n\nf= np.vectorize(lambda x: not x)  # Functioin to toggle boolean array elements\n# same method\n# ~selector.get_support()\n\nv = train.drop(['id', 'target'], axis= 1).columns[f(selector.get_support())]\nprint('{} variables have too low variance.'.format(len(v)))\nprint('These variables are {}'.format(list(v)))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:35:42.613765Z","iopub.execute_input":"2022-07-07T06:35:42.614166Z","iopub.status.idle":"2022-07-07T06:35:43.462609Z","shell.execute_reply.started":"2022-07-07T06:35:42.614136Z","shell.execute_reply":"2022-07-07T06:35:43.461416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We would lose rather many variables if we would select based on variance. But because we do nothave so many variables, we'll let the classifier chose. For data sets with many more vairables this could reduce the processing time.\n\nSklearn also comes with other [feature selection methods](http://scikit-learn.org/stable/modules/feature_selection.html). One of these methods is SelectFromModel in which you let another classifier select the best features and contunue with these. Below I'll show you how to do that with a Random Forest.","metadata":{}},{"cell_type":"markdown","source":"* 천기누설\n- 1000개의 피쳐가 있다.\n- 1개씩 넣고 빼고해서, 피처를 성능보고 셀렉션하면 졸업못해\n- 그래서 할수 있는것은 block을 쓰면 됨\n- 10개씩 또는 20개씩 넣었다 뺐다 함","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:43:54.668493Z","iopub.execute_input":"2022-07-07T06:43:54.668931Z","iopub.status.idle":"2022-07-07T06:43:54.675361Z","shell.execute_reply.started":"2022-07-07T06:43:54.668893Z","shell.execute_reply":"2022-07-07T06:43:54.674328Z"}}},{"cell_type":"markdown","source":"#### 이유한님 방법  (ecursive feature selection)\n- 20 model baseline\n- random choosing 20\n- 40 model train -> baseline compare\n- if performance improvement -> feature importance top 10%\n- new model -> choosing\n- no new model -> another randomly choosing\n- repeat!!","metadata":{}},{"cell_type":"markdown","source":"#### Selecting features with a Random Froest and SelectFromModel","metadata":{}},{"cell_type":"markdown","source":"Here we'll base feature selection on the feature importances of a random forest. With Sklearn's SelectFromModel you can then specify how many variables you want to keep. You can set a threshold on the level of feature importance manually. But we'll simply select the top 50% best variables.\n\nThe code in the cell below is borrowed from the [GitHub repo of Sebastian Raschka](https://github.com/rasbt/python-machine-learning-book/blob/master/code/ch04/ch04.ipynb). This repo contains code samples of his book Python Machine Learning, which is an absolute must to read.","metadata":{}},{"cell_type":"code","source":"X_train = train.drop(['id', 'target'], axis= 1)\ny_train = train['target']\n\nfeat_labels = X_train.columns\n\nrf = RandomForestClassifier(n_estimators= 1000, random_state= 0, n_jobs= -1)\nrf.fit(X_train, y_train)\n\nimportances= rf.feature_importances_\n\nindices = np.argsort(rf.feature_importances_)[::-1]\n\nfor f in range(X_train.shape[1]):\n    print('%df) %-*s %f' % (f + 1, 30, feat_labels[indices[f]], importances[indices[f]]))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T06:59:27.998912Z","iopub.execute_input":"2022-07-07T06:59:27.999585Z","iopub.status.idle":"2022-07-07T07:10:21.519806Z","shell.execute_reply.started":"2022-07-07T06:59:27.999547Z","shell.execute_reply":"2022-07-07T07:10:21.518574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"With SelctFromModel we can specify which prefit classifier to use and what the threshold is for the feature importances. Wiht the get_support method we can then limit the number of vairalbes in the train data.","metadata":{}},{"cell_type":"code","source":"sfm = SelectFromModel(rf, threshold= 'median', prefit= True)\nprint('Number of features before selection: {}'.format(X_train.shape[1]))\n\nn_features= sfm.transform(X_train).shape[1]\nprint('Number of features after selection: {}'.format(n_features))\n\nselected_vars = list(feat_labels[sfm.get_support()])","metadata":{"execution":{"iopub.status.busy":"2022-07-07T07:10:51.147116Z","iopub.execute_input":"2022-07-07T07:10:51.147571Z","iopub.status.idle":"2022-07-07T07:10:53.056125Z","shell.execute_reply.started":"2022-07-07T07:10:51.147535Z","shell.execute_reply":"2022-07-07T07:10:53.055185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train= train[selected_vars + ['target']]","metadata":{"execution":{"iopub.status.busy":"2022-07-07T07:11:05.274327Z","iopub.execute_input":"2022-07-07T07:11:05.275147Z","iopub.status.idle":"2022-07-07T07:11:05.330218Z","shell.execute_reply.started":"2022-07-07T07:11:05.275102Z","shell.execute_reply":"2022-07-07T07:11:05.328267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Feature scaling\n\nAs mentioned before, we can apply standard scaling to the training data.\nsome classifiers perform better when this done.","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()\nscaler.fit_transform(train.drop(['target'], axis= 1))","metadata":{"execution":{"iopub.status.busy":"2022-07-07T07:13:29.705380Z","iopub.execute_input":"2022-07-07T07:13:29.705777Z","iopub.status.idle":"2022-07-07T07:13:30.098115Z","shell.execute_reply.started":"2022-07-07T07:13:29.705744Z","shell.execute_reply":"2022-07-07T07:13:30.096920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### 아주 핵심인것\n- take home message","metadata":{}},{"cell_type":"markdown","source":"- scaling - train으로 먼저 fitting 시키고\n- 그다음에, test transform\n\n**scaler = StandardScaler()**\n\n**df_train = scaler.fit_transform(df_train)**\n\n\n**df_test = scaler.transform(df_test)**","metadata":{}},{"cell_type":"markdown","source":"\n### Conclusion\n\nHopefully this notebook helped you with some tips on how to start with this competition.\nFeel free to vote for it. And if you have questions, post a comment","metadata":{}}]}