{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# **Introduction**\n\nThe objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\nwith the following features being categorical:\n\n['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\n\nTask is to predict, for each customer_ID, the probability of a future payment default (target = 1).","metadata":{}},{"cell_type":"markdown","source":"### **Import libraries**","metadata":{}},{"cell_type":"code","source":"import vaex\nvaex.multithreading.thread_count_default = 8\nimport vaex.ml\n\nimport pandas as pd\nimport numpy  as np \n \nimport os\nimport gc\nimport psutil\nimport glob\nimport tensorflow as tf ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:25:09.294278Z","iopub.execute_input":"2022-08-23T20:25:09.295533Z","iopub.status.idle":"2022-08-23T20:25:22.665500Z","shell.execute_reply.started":"2022-08-23T20:25:09.295364Z","shell.execute_reply":"2022-08-23T20:25:22.664198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def amex_metric(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n\n    def top_four_percent_captured(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        four_pct_cutoff = int(0.04 * df['weight'].sum())\n        df['weight_cumsum'] = df['weight'].cumsum()\n        df_cutoff = df.loc[df['weight_cumsum'] <= four_pct_cutoff]\n        return (df_cutoff['target'] == 1).sum() / (df['target'] == 1).sum()\n        \n    def weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        df = (pd.concat([y_true, y_pred], axis='columns')\n              .sort_values('prediction', ascending=False))\n        df['weight'] = df['target'].apply(lambda x: 20 if x==0 else 1)\n        df['random'] = (df['weight'] / df['weight'].sum()).cumsum()\n        total_pos = (df['target'] * df['weight']).sum()\n        df['cum_pos_found'] = (df['target'] * df['weight']).cumsum()\n        df['lorentz'] = df['cum_pos_found'] / total_pos\n        df['gini'] = (df['lorentz'] - df['random']) * df['weight']\n        return df['gini'].sum()\n\n    def normalized_weighted_gini(y_true: pd.DataFrame, y_pred: pd.DataFrame) -> float:\n        y_true_pred = y_true.rename(columns={'target': 'prediction'})\n        return weighted_gini(y_true, y_pred) / weighted_gini(y_true, y_true_pred)\n\n    g = normalized_weighted_gini(y_true, y_pred)\n    d = top_four_percent_captured(y_true, y_pred)\n\n    return 0.5 * (g + d)\n\ndef amex_metric_np(y_true, y_pred):\n    \n    if type(y_true) != np.ndarray:\n        try:\n            y_true = y_true.numpy()\n        except:\n            y_true = y_true.eval(session=tf.compat.v1.Session())\n            \n    if type(y_pred) != np.ndarray:\n        try:\n            y_pred = y_pred.numpy()\n        except:\n            y_pred = y_pred.eval(session=tf.compat.v1.Session())\n        \n    labels = np.transpose(np.array([y_true, y_pred]))\n    labels = labels[labels[:, 1].argsort()[::-1]]\n    weights = np.where(labels[:,0]==0, 20, 1)\n    cut_vals = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n    gini = [0,0]\n    for i in [1,0]:\n        labels = np.transpose(np.array([y_true, y_pred]))\n        labels = labels[labels[:, i].argsort()[::-1]]\n        weight = np.where(labels[:,0]==0, 20, 1)\n        weight_random = np.cumsum(weight / np.sum(weight))\n        total_pos = np.sum(labels[:, 0] *  weight)\n        cum_pos_found = np.cumsum(labels[:, 0] * weight)\n        lorentz = cum_pos_found / total_pos\n        gini[i] = np.sum((lorentz - weight_random) * weight)\n    return 0.5 * (gini[1]/gini[0] + top_four)\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-08-24T11:27:08.090272Z","iopub.execute_input":"2022-08-24T11:27:08.090833Z","iopub.status.idle":"2022-08-24T11:27:08.110471Z","shell.execute_reply.started":"2022-08-24T11:27:08.090797Z","shell.execute_reply":"2022-08-24T11:27:08.109043Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Last statement data analysis**","metadata":{}},{"cell_type":"markdown","source":"Here we will analyse data that was prepared using https://www.kaggle.com/code/mirfanazam/amex-prediction-starter\n","metadata":{}},{"cell_type":"markdown","source":"### Loading prepared data \n\nPrevious data preporations:\n\n* S_2 date feature is changed to float32\n* All float64 features are converted to float32\n* Missing values in all numeric features are set to 0.0\n* Categorical features D_63 and D_64 are encoded - This was required to extract last statement\n* Categorical features B_31 is converted to float32 - This was required to extract last statement\n* For each customer only last statement is kept as available in each chunk\n\n* Keep only last statement for each customer\n* Remove date feature S_2\n* Encode Customer ID\n","metadata":{"execution":{"iopub.status.busy":"2022-07-20T16:18:23.422588Z","iopub.execute_input":"2022-07-20T16:18:23.422907Z","iopub.status.idle":"2022-07-20T16:18:23.430175Z","shell.execute_reply.started":"2022-07-20T16:18:23.422882Z","shell.execute_reply":"2022-07-20T16:18:23.428336Z"}}},{"cell_type":"code","source":"df_train = vaex.open('../input/amex-prediction-starter-level-2/train_datav2.hdf5')\ndf_test = vaex.open('../input/amex-prediction-starter-level-2/test_datav2.hdf5')\ndf_train_map = vaex.open('../input/amex-prediction-starter-level-2/train_data_customer_map.hdf5')\ndf_test_map = vaex.open('../input/amex-prediction-starter-level-2/test_data_customer_map.hdf5')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:25:39.259012Z","iopub.execute_input":"2022-08-23T20:25:39.259636Z","iopub.status.idle":"2022-08-23T20:25:46.685308Z","shell.execute_reply.started":"2022-08-23T20:25:39.259512Z","shell.execute_reply":"2022-08-23T20:25:46.683924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train_labels = vaex.open('../input/amex-default-prediction/train_labels.csv')\ndf_train_labels = df_train_labels.join(df_train_map, how=\"inner\", on=\"customer_ID\")\n\nall_features = [col for col in df_train]\n\ndf_customer = df_train[all_features]\ndf_customer = df_customer.join(df_train_labels, left_on='customer_ID', right_on='label_encoded_customer_ID', how='inner')\ndf_customer.drop(['label_encoded_customer_ID'], inplace=True)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:25:46.688000Z","iopub.execute_input":"2022-08-23T20:25:46.688319Z","iopub.status.idle":"2022-08-23T20:25:49.627027Z","shell.execute_reply.started":"2022-08-23T20:25:46.688291Z","shell.execute_reply":"2022-08-23T20:25:49.625570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df_train = df_train.to_pandas_df()\ndf_test = df_test.to_pandas_df()\ndf_train_map = df_train_map.to_pandas_df()\ndf_test_map = df_test_map.to_pandas_df()\n\ndf_train_labels = df_train_labels.to_pandas_df()\ndf_customer = df_customer.to_pandas_df()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:25:49.629019Z","iopub.execute_input":"2022-08-23T20:25:49.629531Z","iopub.status.idle":"2022-08-23T20:26:03.341555Z","shell.execute_reply.started":"2022-08-23T20:25:49.629484Z","shell.execute_reply":"2022-08-23T20:26:03.337556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Review data** ","metadata":{}},{"cell_type":"code","source":"# Check na values\nfor col in df_train.columns:\n    na_values = df_train[col].isin([\"NA\", \"\", None, np.NaN]).sum().sum()\n    if na_values >0:\n        print(col, na_values)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:03.346561Z","iopub.execute_input":"2022-08-23T20:26:03.347567Z","iopub.status.idle":"2022-08-23T20:26:19.704006Z","shell.execute_reply.started":"2022-08-23T20:26:03.347505Z","shell.execute_reply":"2022-08-23T20:26:19.701434Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's check significance of linear correlation to target in data","metadata":{"execution":{"iopub.status.busy":"2022-07-20T17:17:56.735580Z","iopub.execute_input":"2022-07-20T17:17:56.736207Z","iopub.status.idle":"2022-07-20T17:17:56.742796Z","shell.execute_reply.started":"2022-07-20T17:17:56.736173Z","shell.execute_reply":"2022-07-20T17:17:56.741446Z"}}},{"cell_type":"code","source":"\ntrain_for_analysis = df_customer.copy() ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:19.707222Z","iopub.execute_input":"2022-08-23T20:26:19.707731Z","iopub.status.idle":"2022-08-23T20:26:19.787087Z","shell.execute_reply.started":"2022-08-23T20:26:19.707682Z","shell.execute_reply":"2022-08-23T20:26:19.785709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lst_keys, lst_corrs_pearson = [], []\n\nfor col in train_for_analysis.drop(['target','customer_ID'],axis=1).keys():\n    \n    corr_pearson = float(train_for_analysis[['target', col]].corr(method ='pearson')[col][:1]) \n    lst_keys.append(col)\n    \n    lst_corrs_pearson.append(corr_pearson)\n\nframe_corr_target = pd.DataFrame(lst_corrs_pearson,  lst_keys).sort_values(by=0, ascending = False)\npd.concat([frame_corr_target[:10],frame_corr_target[-10:]],axis=0)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:19.789046Z","iopub.execute_input":"2022-08-23T20:26:19.790136Z","iopub.status.idle":"2022-08-23T20:26:22.399618Z","shell.execute_reply.started":"2022-08-23T20:26:19.790100Z","shell.execute_reply":"2022-08-23T20:26:22.397949Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sklearn\nfrom sklearn import preprocessing\n\ndef normalization (data, scaler=[]):\n    '''\n    normalization (change distribution of values to range (0,1)) \n    of data using created scaler, or creating new one \n    '''\n    \n    if type(scaler) == sklearn.preprocessing._data.MinMaxScaler:\n        min_max_scaler = scaler\n    else: \n        min_max_scaler = preprocessing.MinMaxScaler(feature_range=(0,1)) \n\n    X_nrm = min_max_scaler.fit_transform(data)\n    X_nrm = pd.DataFrame(X_nrm)\n    headers = list(data.columns.values)\n    X_nrm.columns = headers\n    \n    return X_nrm, min_max_scaler \n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:22.401895Z","iopub.execute_input":"2022-08-23T20:26:22.402316Z","iopub.status.idle":"2022-08-23T20:26:23.102005Z","shell.execute_reply.started":"2022-08-23T20:26:22.402280Z","shell.execute_reply":"2022-08-23T20:26:23.100412Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\nX = train_for_analysis.drop(['target','customer_ID'], axis=1)\ny = train_for_analysis[['target','customer_ID']]\n\nX_nrm,scaler = normalization(X)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:23.104264Z","iopub.execute_input":"2022-08-23T20:26:23.104710Z","iopub.status.idle":"2022-08-23T20:26:23.993303Z","shell.execute_reply.started":"2022-08-23T20:26:23.104674Z","shell.execute_reply":"2022-08-23T20:26:23.991832Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we will check distribution of data that has significant differences by statistics by target column","metadata":{}},{"cell_type":"markdown","source":"## **Assosiation Rules**","metadata":{}},{"cell_type":"code","source":"# Filter significantly differences column by target\n\nX_nrm_y = pd.concat([X_nrm, y['target']], axis =1 ) \n\nsignificant_differnces = (X_nrm_y[X_nrm_y['target']==0].describe() \\\n                          - X_nrm_y[X_nrm_y['target']==1].describe()\n                         ).drop('target',axis=1)\n\nlist_top_corr_columns =  significant_differnces[significant_differnces.index.isin(['mean','min','25%','50%','75%' ])\n                      ].abs().sum().sort_values(ascending=False)[0:10].index.to_list()\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:23.995462Z","iopub.execute_input":"2022-08-23T20:26:23.996186Z","iopub.status.idle":"2022-08-23T20:26:29.089453Z","shell.execute_reply.started":"2022-08-23T20:26:23.996139Z","shell.execute_reply":"2022-08-23T20:26:29.088186Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_top_corr_columns  ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:29.092914Z","iopub.execute_input":"2022-08-23T20:26:29.093259Z","iopub.status.idle":"2022-08-23T20:26:29.100639Z","shell.execute_reply.started":"2022-08-23T20:26:29.093228Z","shell.execute_reply":"2022-08-23T20:26:29.099538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\n\nf, axes = plt.subplots(nrows=3, ncols=2, figsize=(30,30))\n\n# Анализируем данные по достигшим цели\nsns.histplot(data=train_for_analysis, x=list_top_corr_columns[0], hue=\"target\", multiple=\"stack\",stat='percent', kde=True,element=\"step\", ax=axes[0][0]) #, bins =150\naxes[0][0].set_title(f\" Column {list_top_corr_columns[0]}\" )\n\naxes[0][0].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[0] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution Default\", c='red')\naxes[0][0].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[0] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution do not Default\", c='blue')\naxes[0][0].axvline(x=(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[0] ]) \\\n                   +np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[0] ]) )/2, \n                   linestyle='--', linewidth=2.5, label=\"between averages of distribution Default\", c='green')\n\n\n\nsns.histplot(data=train_for_analysis, x=list_top_corr_columns[1], hue=\"target\", multiple=\"stack\", bins =50, kde=True,element=\"step\", legend=True, ax=axes[0][1])\naxes[0][1].set_title(f\"Column {list_top_corr_columns[1]}\")\n\naxes[0][1].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[1] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution Default\", c='red')\naxes[0][1].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[1] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution do not Default\", c='blue')\naxes[0][1].axvline(x=(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[1] ]) \\\n                   +np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[1] ]) )/2, \n                   linestyle='--', linewidth=2.5, label=\"between averages of distribution Default\", c='green')\n\n\naxes[0][1].legend()\n\n# ------------------------------------------- 2 row\nsns.histplot(data=train_for_analysis, x=list_top_corr_columns[2], hue=\"target\", multiple=\"stack\", bins =50, kde=True,element=\"step\", ax=axes[1][0])\naxes[1][0].set_title(f\"Distribution of Sample Means {list_top_corr_columns[2]}\")\n\naxes[1][0].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[2] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution Default\", c='red')\naxes[1][0].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[2] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution do not Default\", c='blue')\naxes[1][0].axvline(x=(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[2] ]) \\\n                   +np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[2] ]) )/2, \n                   linestyle='--', linewidth=2.5, label=\"between averages of distribution Default\", c='green')\n\n\n\nsns.histplot(data=train_for_analysis, x=list_top_corr_columns[3], hue=\"target\", multiple=\"stack\", bins =50, kde=True,element=\"step\", ax=axes[1][1])\naxes[1][1].set_title(f\"Distribution of Sample Means {list_top_corr_columns[3]}\" )\n\naxes[1][1].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[3] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution Default\", c='red')\naxes[1][1].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[3] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution do not Default\", c='blue')\naxes[1][1].axvline(x=(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[3] ]) \\\n                   +np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[3] ]) )/2, \n                   linestyle='--', linewidth=2.5, label=\"between averages of distribution Default\", c='green')\n\n\naxes[1][1].legend()\n\n\n# ------------------------------------------- 3 row\nsns.histplot(data=train_for_analysis, x=list_top_corr_columns[4], hue=\"target\", multiple=\"stack\", bins =50, kde=True,element=\"step\", ax=axes[2][0])\naxes[2][0].set_title(f\"Distribution of Sample Means {list_top_corr_columns[4]}\")\n\naxes[2][0].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[4] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution Default\", c='red')\naxes[2][0].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[4] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution do not Default\", c='blue')\naxes[2][0].axvline(x=(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[4] ]) \\\n                   +np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[4] ]) )/2, \n                   linestyle='--', linewidth=2.5, label=\"between averages of distribution Default\", c='green')\n\n\n\nsns.histplot(data=train_for_analysis, x=list_top_corr_columns[5], hue=\"target\", multiple=\"stack\", bins =50, kde=True,element=\"step\", ax=axes[2][1])\naxes[2][1].set_title(f\"Distribution of Sample Means {list_top_corr_columns[5]}\")\n\naxes[2][1].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[5] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution Default\", c='red')\naxes[2][1].axvline(x=np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[5] ]), \n                   linestyle='--', linewidth=2.5, label=\"average of distribution do not Default\", c='blue')\naxes[2][1].axvline(x=(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,list_top_corr_columns[5] ]) \\\n                   +np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,list_top_corr_columns[5] ]) )/2, \n                   linestyle='--', linewidth=2.5, label=\"between averages of distribution Default\", c='green')\n\n\naxes[2][1].legend()\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:29.102317Z","iopub.execute_input":"2022-08-23T20:26:29.102777Z","iopub.status.idle":"2022-08-23T20:26:42.422099Z","shell.execute_reply.started":"2022-08-23T20:26:29.102733Z","shell.execute_reply":"2022-08-23T20:26:42.420677Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can see that in some ranges by metrics distribution is difference significantly. For better representation we can look for association rules by ranges of these metrics and categorical metrics","metadata":{"execution":{"iopub.status.busy":"2022-07-24T11:05:57.440626Z","iopub.execute_input":"2022-07-24T11:05:57.441132Z","iopub.status.idle":"2022-07-24T11:05:57.449992Z","shell.execute_reply.started":"2022-07-24T11:05:57.441092Z","shell.execute_reply":"2022-07-24T11:05:57.448227Z"}}},{"cell_type":"code","source":"#Calculating intervals\npd.options.mode.chained_assignment = None\n\ncategorical_cls=['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\ncorr_columns = train_for_analysis.columns[(train_for_analysis.columns.isin(list_top_corr_columns+categorical_cls))]\ndf_assc_rules = train_for_analysis[corr_columns]\n\nintervals_dict = {}\n\nfor n in corr_columns[~corr_columns.isin(categorical_cls)]:\n    \n    intervals = []\n    intervals.append(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,n ]))\n    intervals.append(np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,n ]))\n    intervals.append(\n       ( np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 0 ,n ]) \\\n    + np.mean(train_for_analysis.loc[ train_for_analysis['target'] == 1 ,n ]) )/2 \n    )  # between means of two distributions \n    intervals.sort()\n    intervals_dict[n] = intervals\n    df_assc_rules.loc[df_assc_rules[n] <= intervals[0], n+'_classes'] =  '0' # n+'_1'\n    df_assc_rules.loc[(df_assc_rules[n] >= intervals[0]) & (df_assc_rules[n] <= intervals[1]), n+'_classes'] = '1' #n+'_2'\n    df_assc_rules.loc[(df_assc_rules[n] >= intervals[1]) & (df_assc_rules[n] <= intervals[2]), n+'_classes'] = '2' #n+'_3'\n    df_assc_rules.loc[(df_assc_rules[n] >= intervals[2]), n+'_classes'] = '3' #n+'_4'\n\nfor cat in categorical_cls:\n    df_assc_rules[cat+'_classes'] = df_assc_rules[cat].astype(int) # f'{cat}' + df_assc_rules[cat].astype(str)\n    \ndf_assc_rules = df_assc_rules[[n for n in df_assc_rules.columns if 'classes' in n]]# filter cols ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:26:42.424159Z","iopub.execute_input":"2022-08-23T20:26:42.425017Z","iopub.status.idle":"2022-08-23T20:26:44.119290Z","shell.execute_reply.started":"2022-08-23T20:26:42.424977Z","shell.execute_reply":"2022-08-23T20:26:44.117950Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test = pd.concat([df_assc_rules,train_for_analysis['target']], axis=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:27:16.751797Z","iopub.execute_input":"2022-08-23T20:27:16.752206Z","iopub.status.idle":"2022-08-23T20:27:16.802524Z","shell.execute_reply.started":"2022-08-23T20:27:16.752175Z","shell.execute_reply":"2022-08-23T20:27:16.801061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\n\n# take a small sample \nX_train,X_test, y_train,y_test= train_test_split(test.drop(['target'],axis=1) ,test['target'], test_size = 0.01)\n\nsample_test = pd.concat([X_test,y_test], axis=1) ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:27:16.804484Z","iopub.execute_input":"2022-08-23T20:27:16.805047Z","iopub.status.idle":"2022-08-23T20:27:17.937420Z","shell.execute_reply.started":"2022-08-23T20:27:16.805012Z","shell.execute_reply":"2022-08-23T20:27:17.935464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# exclude binary columns \nbinary_cols = []\nframe_encoded = pd.DataFrame()\n\n# Encoding the datasets\nfor col in sample_test.columns: \n    lst_vals = list(sample_test[col].unique())\n    lst_vals.sort()\n    if lst_vals == [0,1]:\n        binary_cols.append(col)\n    elif frame_encoded.shape == (0,0): \n        col_encoded = pd.get_dummies(sample_test[col])\n        col_encoded = col_encoded.add_prefix(col+'_')\n        frame_encoded = col_encoded\n    else:\n        col_encoded = pd.get_dummies(sample_test[col])\n        col_encoded = col_encoded.add_prefix(col+'_')\n        frame_encoded = pd.concat([frame_encoded,col_encoded],axis=1) \nfor col_binary in binary_cols:\n    frame_encoded = pd.concat([frame_encoded,sample_test[col_binary]], axis=1)\n\nframe_encoded\n","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:27:17.939456Z","iopub.execute_input":"2022-08-23T20:27:17.940409Z","iopub.status.idle":"2022-08-23T20:27:18.016604Z","shell.execute_reply.started":"2022-08-23T20:27:17.940339Z","shell.execute_reply":"2022-08-23T20:27:18.015042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have cutted columns by classes of distribution and encoded them by binary for training association rules model","metadata":{}},{"cell_type":"code","source":"from mlxtend.frequent_patterns import apriori, association_rules\n\n# Building the model \n# Using 2 different slices because RAM threshold \n\nlst_cols = list(frame_encoded.drop(['target'], axis=1).columns[:32])\nlst_cols.append('target')\n\nlst_cols_2 = list(frame_encoded.drop(['target'], axis=1).columns[32:])\nlst_cols_2.append('target')\n\nfrq_items_1 = apriori(frame_encoded[lst_cols].astype('bool') , min_support = 0.05, use_colnames = True)\nfrq_items_2 = apriori(frame_encoded[lst_cols_2].astype('bool') , min_support = 0.05, use_colnames = True)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:27:24.067499Z","iopub.execute_input":"2022-08-23T20:27:24.067908Z","iopub.status.idle":"2022-08-23T20:27:24.452697Z","shell.execute_reply.started":"2022-08-23T20:27:24.067874Z","shell.execute_reply":"2022-08-23T20:27:24.451540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Collecting the inferred rules in a dataframe\n\n# Confidence  an indicator of how often our rule works for the entire dataset, how many cases of joint two events, relative to the occurrence of the first event\n# Support an indicator of the \"frequency\" of this item set in all analyzed transactions, how many such cases are relative to the entire data set\n\nrules = association_rules(pd.concat([frq_items_1,frq_items_2],axis=0), metric =\"lift\", min_threshold = 1)\nrules = rules.sort_values(['confidence', 'lift'], ascending =[False, False])\nrules.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:40:53.338982Z","iopub.execute_input":"2022-08-23T20:40:53.339499Z","iopub.status.idle":"2022-08-23T20:40:56.746807Z","shell.execute_reply.started":"2022-08-23T20:40:53.339462Z","shell.execute_reply":"2022-08-23T20:40:56.745518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# decoding and checking on a big frame (train_for_analysis)\n\n# filtering by 0.6 confidence threshold\ntop_confidence = list(rules[(rules.consequents.isin(\n    [n for n in rules['consequents'] if 'target' in n])\n      ) &\n     (rules.confidence>.6)\n     ]['antecedents'].unique()[:250])  # we take top 250 rules\n\nnum = 0\nrules_dict = {}\n\nfor rule in top_confidence:\n    rules_dict[num] = {'cols':[],'intervals':[]}\n    for n in rule:\n        \n        colname = n[:n.find(\"_classes\")]\n        rules_dict[num]['cols'].append(colname)\n        \n        if n in binary_cols:\n            rules_dict[num]['intervals'].append(1)\n            continue \n            \n        else:\n            interval_value = int(n[-1])\n            rules_dict[num]['intervals'].append(interval_value)\n            \n    num+=1","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:40:59.996859Z","iopub.execute_input":"2022-08-23T20:40:59.997382Z","iopub.status.idle":"2022-08-23T20:41:00.275103Z","shell.execute_reply.started":"2022-08-23T20:40:59.997325Z","shell.execute_reply":"2022-08-23T20:41:00.273693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def extracting_and_using_rules (source_frame: pd.DataFrame(), rules_dict : dict, binary_cols: list): \n    summary_frame = {}\n    # we also will record all the unique columns \n    unique_columns = set()\n\n    for rule in rules_dict:\n\n        big_frame = source_frame.copy()\n\n        number_of_antecedents_cols = len(rules_dict[rule]['cols'])\n        frame_rule = rules_dict[rule]\n        summary_frame[rule] = {}\n        for n in range(number_of_antecedents_cols):\n            col = frame_rule['cols'][n]\n            interval = frame_rule['intervals'][n]\n\n            summary_frame[rule][col] = {'start':[],'end':[]}\n            \n            if col in binary_cols:\n                summary_frame[rule][col]['start'] = 1\n                summary_frame[rule][col]['end'] = 1\n            else:    \n                if interval == 0: \n                    big_frame = big_frame[big_frame[col] <= intervals_dict[col][interval]]\n                    summary_frame[rule][col]['end'] = intervals_dict[col][interval]\n\n                elif interval in [1,2]:\n                    big_frame = big_frame[(big_frame[col] >= intervals_dict[col][interval-1])&\n                                         (big_frame[col] <= intervals_dict[col][interval])]\n                    summary_frame[rule][col]['start'] = intervals_dict[col][interval-1]\n                    summary_frame[rule][col]['end'] = intervals_dict[col][interval]\n\n                else:\n                    big_frame = big_frame[big_frame[col] >= intervals_dict[col][interval-1]]\n                    summary_frame[rule][col]['start'] = intervals_dict[col][interval-1]\n\n            unique_columns.add(col)\n\n        summary_frame[rule]['precision'] = round(len(big_frame[big_frame['target']==1]) \n                /len(big_frame) * 100,2)\n\n    return summary_frame, unique_columns","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:41:55.941595Z","iopub.execute_input":"2022-08-23T20:41:55.942108Z","iopub.status.idle":"2022-08-23T20:41:55.959836Z","shell.execute_reply.started":"2022-08-23T20:41:55.942071Z","shell.execute_reply":"2022-08-23T20:41:55.958061Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rules_dict_filtered = rules_dict.copy() \n\nfor n in rules_dict:\n    to_drop = 0\n    for col in rules_dict[n]['cols']:\n        if col not in list(intervals_dict.keys())+binary_cols:\n            to_drop = 1\n    if to_drop == 1: \n        rules_dict_filtered.pop(n)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:41:58.614497Z","iopub.execute_input":"2022-08-23T20:41:58.615419Z","iopub.status.idle":"2022-08-23T20:41:58.624453Z","shell.execute_reply.started":"2022-08-23T20:41:58.615369Z","shell.execute_reply":"2022-08-23T20:41:58.622837Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frame_rules_encoded = extracting_and_using_rules(train_for_analysis, rules_dict_filtered, binary_cols)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:42:01.160572Z","iopub.execute_input":"2022-08-23T20:42:01.161405Z","iopub.status.idle":"2022-08-23T20:44:07.735603Z","shell.execute_reply.started":"2022-08-23T20:42:01.161363Z","shell.execute_reply":"2022-08-23T20:44:07.733523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We have 250 frames of rules that look like this: ","metadata":{}},{"cell_type":"code","source":"pd.DataFrame(frame_rules_encoded[0][0])","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:44:07.738276Z","iopub.execute_input":"2022-08-23T20:44:07.738707Z","iopub.status.idle":"2022-08-23T20:44:07.758864Z","shell.execute_reply.started":"2022-08-23T20:44:07.738671Z","shell.execute_reply":"2022-08-23T20:44:07.756983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lst_precisions = [frame_rules_encoded[0][n]['precision'] for n in frame_rules_encoded[0]]","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:44:07.760165Z","iopub.execute_input":"2022-08-23T20:44:07.760559Z","iopub.status.idle":"2022-08-23T20:44:07.771301Z","shell.execute_reply.started":"2022-08-23T20:44:07.760525Z","shell.execute_reply":"2022-08-23T20:44:07.769808Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"And we have set of columns that was used for rules the most often","metadata":{}},{"cell_type":"code","source":"print('min, max precision of these rules is:',min(lst_precisions), max(lst_precisions)) ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:44:07.774157Z","iopub.execute_input":"2022-08-23T20:44:07.774709Z","iopub.status.idle":"2022-08-23T20:44:07.790217Z","shell.execute_reply.started":"2022-08-23T20:44:07.774659Z","shell.execute_reply":"2022-08-23T20:44:07.788894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frame_rules_encoded[1]","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:44:07.791899Z","iopub.execute_input":"2022-08-23T20:44:07.792600Z","iopub.status.idle":"2022-08-23T20:44:07.809189Z","shell.execute_reply.started":"2022-08-23T20:44:07.792559Z","shell.execute_reply":"2022-08-23T20:44:07.807774Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Conclusion \nWe have 250 association rules that we can for analysis and making decision about status of client. \nWe can also use this information to prepare data for machine learning","metadata":{}},{"cell_type":"markdown","source":"## **Removing outliers by filtered columns**","metadata":{}},{"cell_type":"markdown","source":"we will remove outliers by the most important columns that we found while making assosiation rules, we will use Interquartile range for it. We will exclude categorical columns and add additional important features using Kbest Feature selection","metadata":{}},{"cell_type":"code","source":"def interquantile_remove_outliners( frame_important_cols, data_orig=pd.DataFrame(), factor = 3, only_show = True ):\n    \n    '''\n    function return current presence of outliers or dataframe with removed outliers using quantile range method \n    '''\n    a = []\n    percs = []\n    report = []\n    for k, v in frame_important_cols.items():\n        q1 = v.quantile(0.25)\n        q3 = v.quantile(0.75)\n        irq = q3 - q1 # interquartile range\n        # outliers are beyond these points\n        v_col = v[(v <= q1 - factor * irq) | (v >= q3 + factor * irq)]\n        perc = np.shape(v_col)[0] * 100.0 / np.shape(frame_important_cols)[0] # we calculate the relative indicator for each column \n        if only_show == True:\n            percs.append(perc)\n            report.append(\"Outlinters by %s = %.2f%%\" % (k, perc))\n        elif len(v_col) > 0 and only_show == False:\n            for i in list(v_col.index):\n                a.append(i)\n                \n    if only_show == True:\n        return report , percs\n    else:\n        return data_orig.drop(list(set(a)))","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:07.764673Z","iopub.execute_input":"2022-08-23T20:45:07.765202Z","iopub.status.idle":"2022-08-23T20:45:07.779612Z","shell.execute_reply.started":"2022-08-23T20:45:07.765160Z","shell.execute_reply":"2022-08-23T20:45:07.778153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# use oversampling to normilized train data and split in to train/valid\n\nX = train_for_analysis.drop(['target','customer_ID'],axis=1)\nY = train_for_analysis['target'] \n\n# X_scaled = min_max_scaler.fit_transform(X)\n# X_scaled = pd.DataFrame(X_scaled)\n# headers = list(X.columns.values)\n# X_scaled.columns = headers\n\nX_train_orig,X_valid_orig, y_train_orig,y_valid_orig= train_test_split(X, Y, shuffle = True, test_size = 0.2)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:07.781590Z","iopub.execute_input":"2022-08-23T20:45:07.783057Z","iopub.status.idle":"2022-08-23T20:45:08.475702Z","shell.execute_reply.started":"2022-08-23T20:45:07.783016Z","shell.execute_reply":"2022-08-23T20:45:08.474355Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.feature_selection import SelectKBest, chi2\n\n\nframe_train_summary =  pd.concat([X_train_orig, y_train_orig],axis = 1)\nrules_cols = list(frame_rules_encoded[1])\n\n# Excluding categorical columns and found using association rules columns\nX_train_for_selection =  X_train_orig[X_train_orig.columns[~X_train_orig.columns.isin(rules_cols+categorical_cls)]]\n\nselect = SelectKBest(chi2, k=20) # take 20 most important columns \n\nX_new = select.fit_transform(\n    normalization(X_train_for_selection,scaler )[0],y_train_orig)\nX_new = pd.DataFrame(X_new)\nfilter = select.get_support()\nselected_cols = X_train_for_selection[X_train_for_selection.columns[filter]].columns","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:08.477854Z","iopub.execute_input":"2022-08-23T20:45:08.478220Z","iopub.status.idle":"2022-08-23T20:45:10.424887Z","shell.execute_reply.started":"2022-08-23T20:45:08.478188Z","shell.execute_reply":"2022-08-23T20:45:10.423259Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that by some columns have distribution by different classes is cross together significantly (B20, B17), but others we can cut carefully","metadata":{"execution":{"iopub.status.busy":"2022-08-10T17:24:02.388702Z","iopub.execute_input":"2022-08-10T17:24:02.389219Z","iopub.status.idle":"2022-08-10T17:24:02.397124Z","shell.execute_reply.started":"2022-08-10T17:24:02.389177Z","shell.execute_reply":"2022-08-10T17:24:02.395657Z"}}},{"cell_type":"code","source":"frame_important_cols = frame_train_summary.loc[frame_train_summary['target'] == 1,  list(selected_cols)+rules_cols]\nshow_otliners = interquantile_remove_outliners(frame_important_cols)\nshow_otliners[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:10.427075Z","iopub.execute_input":"2022-08-23T20:45:10.427775Z","iopub.status.idle":"2022-08-23T20:45:10.689685Z","shell.execute_reply.started":"2022-08-23T20:45:10.427740Z","shell.execute_reply":"2022-08-23T20:45:10.688558Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We don't have any outliners by some columns, but it's too significant otliners by some of them. But we can't be sure that it's outliners, because it's too big part of our distribution.\nSo we will remove otliners where it is less than 7% ","metadata":{}},{"cell_type":"code","source":"frame_important_cols_prc = pd.DataFrame(  show_otliners[1], frame_important_cols.keys(), columns = ['data_part_percent'])\nframe_important_cols_prc = frame_important_cols_prc[(frame_important_cols_prc['data_part_percent']>0)& (frame_important_cols_prc['data_part_percent']<=7 )].index","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:10.692923Z","iopub.execute_input":"2022-08-23T20:45:10.693705Z","iopub.status.idle":"2022-08-23T20:45:10.700766Z","shell.execute_reply.started":"2022-08-23T20:45:10.693658Z","shell.execute_reply":"2022-08-23T20:45:10.699802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"frame_important_cols_prc","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:10.701909Z","iopub.execute_input":"2022-08-23T20:45:10.702487Z","iopub.status.idle":"2022-08-23T20:45:10.715080Z","shell.execute_reply.started":"2022-08-23T20:45:10.702454Z","shell.execute_reply":"2022-08-23T20:45:10.714187Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f, axes = plt.subplots(nrows =2 , ncols=5, figsize=(40,16))\n\n# Positive correlations (The higher the feature the probability increases that it will be a fraud transaction)\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[0]}\", data=frame_train_summary, palette=\"Set1\",  ax=axes[0][0])\naxes[0][0].set_title(f'{frame_important_cols_prc[0]} by target')\naxes[0][0].set(ylabel=None)\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[1]}\", data=frame_train_summary, palette=\"Set1\", ax=axes[0][1])\naxes[0][1].set_title(f'{frame_important_cols_prc[1]} by target')\naxes[0][1].set(ylabel=None)\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[2]}\", data=frame_train_summary,palette=\"Set1\", ax=axes[0][2])\naxes[0][2].set_title(f'{frame_important_cols_prc[2]} by target')\naxes[0][2].set(ylabel=None)\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[3]}\", data=frame_train_summary, palette=\"Set1\",  ax=axes[0][3])\naxes[0][3].set_title(f'{frame_important_cols_prc[3]} by target')\naxes[0][3].set(ylabel=None)\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[4]}\", data=frame_train_summary, palette=\"Set1\", ax=axes[0][4])\naxes[0][4].set_title(f'{frame_important_cols_prc[4]} by target')\naxes[0][4].set(ylabel=None)\n\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[5]}\", data=frame_train_summary, palette=\"Set1\", ax=axes[1][0])\naxes[1][0].set_title(f'{frame_important_cols_prc[5]} by target')\naxes[1][0].set(ylabel=None)\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[6]}\", data=frame_train_summary, palette=\"Set1\", ax=axes[1][1])\naxes[1][1].set_title(f'{frame_important_cols_prc[6]} by target')\naxes[1][1].set(ylabel=None)\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[7]}\", data=frame_train_summary, palette=\"Set1\", ax=axes[1][2])\naxes[1][2].set_title(f'{frame_important_cols_prc[7]} by target')\naxes[1][2].set(ylabel=None)\n\nsns.boxplot(x=\"target\", y=f\"{frame_important_cols_prc[8]}\", data=frame_train_summary,palette=\"Set1\", ax=axes[1][3])\naxes[1][3].set_title(f'{frame_important_cols_prc[8]} by target')\naxes[1][3].set(ylabel=None)\n\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:10.716210Z","iopub.execute_input":"2022-08-23T20:45:10.716953Z","iopub.status.idle":"2022-08-23T20:45:12.827879Z","shell.execute_reply.started":"2022-08-23T20:45:10.716919Z","shell.execute_reply":"2022-08-23T20:45:12.827005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_outliners_removed =  interquantile_remove_outliners( frame_important_cols[frame_important_cols_prc], frame_train_summary, only_show = False )","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:12.829202Z","iopub.execute_input":"2022-08-23T20:45:12.829837Z","iopub.status.idle":"2022-08-23T20:45:13.186932Z","shell.execute_reply.started":"2022-08-23T20:45:12.829800Z","shell.execute_reply":"2022-08-23T20:45:13.186023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"interquantile_remove_outliners( X_train_outliners_removed.loc[X_train_outliners_removed['target'] == 1, frame_important_cols_prc])[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:45:13.188236Z","iopub.execute_input":"2022-08-23T20:45:13.188768Z","iopub.status.idle":"2022-08-23T20:45:13.327733Z","shell.execute_reply.started":"2022-08-23T20:45:13.188735Z","shell.execute_reply":"2022-08-23T20:45:13.326459Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Conclusion \n\nWe have removed outliners by several columns and got frame without them ","metadata":{}},{"cell_type":"markdown","source":"## **Resampling data**","metadata":{}},{"cell_type":"code","source":"X_train_outliners_removed['target'].plot(kind='hist')","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:52:42.799766Z","iopub.execute_input":"2022-08-23T20:52:42.800532Z","iopub.status.idle":"2022-08-23T20:52:43.080866Z","shell.execute_reply.started":"2022-08-23T20:52:42.800470Z","shell.execute_reply":"2022-08-23T20:52:43.079303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Target classes represented in differently in data, we will resample it. Let's compare different approaches of sampling. In our case we will test only upsampling because we have just part of data for education and we should not waste any line of it","metadata":{}},{"cell_type":"code","source":"#we will use small sample of data to represent different approaches result \n\nX = train_for_analysis.drop(['target','customer_ID'],axis=1)\nY = train_for_analysis['target'] \n\nX_train,X_test, y_train,y_test= train_test_split( X, Y, test_size = 0.01) # get 1% sample of data ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:52:43.083577Z","iopub.execute_input":"2022-08-23T20:52:43.084089Z","iopub.status.idle":"2022-08-23T20:52:43.810867Z","shell.execute_reply.started":"2022-08-23T20:52:43.084042Z","shell.execute_reply":"2022-08-23T20:52:43.809554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.manifold import TSNE\n\n#reducing dimension size of slice to 2-d \n\nX_st = normalization(X_test,scaler)[0]\nX_tsne = TSNE(init='pca', learning_rate='auto',).fit_transform(X_st.values)","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:52:43.812202Z","iopub.execute_input":"2022-08-23T20:52:43.812573Z","iopub.status.idle":"2022-08-23T20:53:09.201060Z","shell.execute_reply.started":"2022-08-23T20:52:43.812543Z","shell.execute_reply":"2022-08-23T20:53:09.200046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import seaborn as sns\n\ndf_tsne_wlabels = pd.concat([pd.DataFrame(X_tsne,columns=['X','Y'],index=X_test.index), y_test],axis=1)\n\n# current dstribution \nsns.set(rc = {'figure.figsize':(20,10)})\nsns.scatterplot(data=df_tsne_wlabels, x=\"X\", y=\"Y\", hue='target')","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:53:09.203395Z","iopub.execute_input":"2022-08-23T20:53:09.204244Z","iopub.status.idle":"2022-08-23T20:53:09.691585Z","shell.execute_reply.started":"2022-08-23T20:53:09.204206Z","shell.execute_reply":"2022-08-23T20:53:09.690550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from imblearn.over_sampling import  RandomOverSampler, SMOTE, ADASYN, BorderlineSMOTE, KMeansSMOTE, SVMSMOTE\n\n\n# get small slice at the center of distribution \nslice_reprs = df_tsne_wlabels[(df_tsne_wlabels['X']>-10) & (df_tsne_wlabels['X']<-2)]\n\n\nx_ros, y_ros = RandomOverSampler().fit_resample(slice_reprs.drop(['target'],axis=1), slice_reprs['target'])\nx_smote, y_smote = SMOTE().fit_resample(slice_reprs.drop(['target'],axis=1), slice_reprs['target'])\nx_bordsm, y_bordsm = BorderlineSMOTE().fit_resample(slice_reprs.drop(['target'],axis=1), slice_reprs['target'])\nx_svsm, y_svsm = SVMSMOTE().fit_resample(slice_reprs.drop(['target'],axis=1), slice_reprs['target'])","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:53:09.693135Z","iopub.execute_input":"2022-08-23T20:53:09.693525Z","iopub.status.idle":"2022-08-23T20:53:09.910025Z","shell.execute_reply.started":"2022-08-23T20:53:09.693490Z","shell.execute_reply":"2022-08-23T20:53:09.908711Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\nimport seaborn as sns\n\nfig, axs = plt.subplots(nrows=2, ncols=3)\n\nsns.scatterplot(data=slice_reprs, x=\"X\", y=\"Y\", hue='target', ax=axs[0][0]).set(title='Original')\nsns.scatterplot(data=x_ros, x=\"X\", y=\"Y\", hue=y_ros, ax=axs[0][1]).set(title='Random') \nsns.scatterplot(data=x_smote, x=\"X\", y=\"Y\", hue=y_smote, ax=axs[0][2]).set(title='SMOTE')\nsns.scatterplot(data=x_bordsm, x=\"X\", y=\"Y\", hue=y_bordsm, ax=axs[1][0]).set(title='BorderlineSMOTE')\nsns.scatterplot(data=x_svsm, x=\"X\", y=\"Y\", hue=y_svsm, ax=axs[1][1]).set(title='SVMSMOTE') ","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:53:09.911512Z","iopub.execute_input":"2022-08-23T20:53:09.911958Z","iopub.status.idle":"2022-08-23T20:53:11.480118Z","shell.execute_reply.started":"2022-08-23T20:53:09.911914Z","shell.execute_reply":"2022-08-23T20:53:11.478942Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, Borderline and SVMSMOTE put a lot of emphasis on some areas, this can help in training our model, but it skews the original distribution significantly. In our example it is better to use SMOTE","metadata":{}},{"cell_type":"code","source":"#Resampling train data \n\n# Original data\nX_train_rsm,y_train_rsm = SMOTE().fit_resample(X_train_orig,y_train_orig)\nX_train_rsm = normalization(X_train_rsm, scaler)[0]\n\n# Removed outliners data  \nX_out = X_train_outliners_removed.drop(['target'],axis=1)\nY_out = X_train_outliners_removed['target'] \n\nX_out_sm,Y_out_sm = SMOTE().fit_resample(X_out,Y_out)\nX_out_sm = normalization(X_out_sm, scaler)[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-23T20:53:11.481916Z","iopub.execute_input":"2022-08-23T20:53:11.482250Z","iopub.status.idle":"2022-08-23T21:02:54.534168Z","shell.execute_reply.started":"2022-08-23T20:53:11.482220Z","shell.execute_reply":"2022-08-23T21:02:54.532839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.set(rc = {'figure.figsize':(10,5)})\ny_train_rsm.plot(kind='hist')","metadata":{"execution":{"iopub.status.busy":"2022-08-23T21:02:54.536055Z","iopub.execute_input":"2022-08-23T21:02:54.536451Z","iopub.status.idle":"2022-08-23T21:02:54.793441Z","shell.execute_reply.started":"2022-08-23T21:02:54.536417Z","shell.execute_reply":"2022-08-23T21:02:54.792511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Data is balanced now ","metadata":{}},{"cell_type":"markdown","source":"# **Fast education and chosing ml algorithm**\n","metadata":{}},{"cell_type":"markdown","source":"We gonna train and test algorithms on small slice ","metadata":{"execution":{"iopub.status.busy":"2022-08-11T13:59:26.578445Z","iopub.execute_input":"2022-08-11T13:59:26.578896Z","iopub.status.idle":"2022-08-11T13:59:26.586156Z","shell.execute_reply.started":"2022-08-11T13:59:26.578860Z","shell.execute_reply":"2022-08-11T13:59:26.584896Z"}}},{"cell_type":"markdown","source":"### classification algorithms load","metadata":{"execution":{"iopub.status.busy":"2022-08-19T08:29:54.465339Z","iopub.execute_input":"2022-08-19T08:29:54.465857Z","iopub.status.idle":"2022-08-19T08:29:54.471027Z","shell.execute_reply.started":"2022-08-19T08:29:54.465819Z","shell.execute_reply":"2022-08-19T08:29:54.469932Z"}}},{"cell_type":"code","source":"##Подберём модель для классификации \nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.metrics import classification_report\n\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.svm import SVC\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.ensemble import AdaBoostClassifier\nfrom sklearn.ensemble import GradientBoostingClassifier\nfrom lightgbm import LGBMClassifier\n\nfrom sklearn.metrics import f1_score\n\nmodels = {\n#     \"LogisiticRegression\": LogisticRegression(solver='liblinear', dual=True),\n    \"RandomForestClassifier\" : RandomForestClassifier(),\n    \"DecisionTreeClassifier\" : DecisionTreeClassifier(), \n    \"KNeighborsClassifier\" : KNeighborsClassifier(),\n    \"Support Vector Classifier\":SVC(), \n    \"GaussianNB\" : GaussianNB(),\n    \"AdaBoostClassifier\" : AdaBoostClassifier(),\n    \"GradientBoostingClassifier\" : GradientBoostingClassifier(),\n    \"LGBMClassifier\" : LGBMClassifier()\n}\n\n\n# amex_metric\n# amex_metric_mod\n\ndef model(name, model, X_train, y_train, X_valid, y_valid):\n    \n    model.fit(X_train, y_train)\n    score = model.score(X_train, y_train)\n    prediction  = model.predict(X_valid)\n    cv_score = cross_val_score(model,X_train, y_train,cv=5)\n    f1_valid = f1_score(y_valid, prediction, average='weighted')\n    amex_metric = amex_metric_np( y_valid.to_numpy() if type(y_valid) != np.ndarray else y_valid, prediction)\n    return { \"acc\" :score,  \"cv_score\" : np.mean(cv_score), \"f1_valid\" :  f1_valid, 'model': model, 'amex_metric': amex_metric} ","metadata":{"execution":{"iopub.status.busy":"2022-08-24T10:06:13.423151Z","iopub.execute_input":"2022-08-24T10:06:13.423557Z","iopub.status.idle":"2022-08-24T10:06:13.434443Z","shell.execute_reply.started":"2022-08-24T10:06:13.423523Z","shell.execute_reply":"2022-08-24T10:06:13.433282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def fast_testing(X_train,y_train,X_valid,y_valid, models: dict(), min_max_scaler=[], one_model = False, norm=True):\n    ''' \n    training testing and validation models on small slices of dataset for find the \n    '''\n    X_train_small, X_valid_small, y_train_small, y_valid_small = train_test_split(X_valid, y_valid , test_size = 0.05) # clear orig data\n    X_train_small, X_test_small, y_train_small, y_test_small = train_test_split(X_train,y_train,  test_size = 0.05) # train resampled data \n\n    #normalization valid data \n    if norm:\n        X_valid_nrm = normalization(X_valid_small, min_max_scaler)[0]\n    \n    else:\n        X_valid_nrm = X_valid_small\n    \n    acc = []\n    scores = []\n    f1 = []\n    models_fitted = []\n    amex_metric = []\n    \n    if one_model==False:\n\n        for name, clf in models.items():\n            model_fitted = model(name, clf, X_test_small, y_test_small, X_valid_nrm, y_valid_small)\n            acc.append(model_fitted['acc'])\n            scores.append(model_fitted['cv_score'])\n            f1.append(model_fitted['f1_valid'])\n            models_fitted.append(model_fitted['model'])\n            amex_metric.append(model_fitted['amex_metric'])\n            index=list(models.keys())\n        dff = pd.DataFrame({'Accuracy': acc,\n                'Cross Val Score': scores,\n                    'f1_valid':  f1, \n                        'amex_metric':  amex_metric\n                       }, index=index)\n        dff = dff.sort_values(by=['amex_metric','Accuracy','f1_valid','Cross Val Score'], ascending=False)\n            \n    else:\n        model_fitted = model('model', models, X_test_small, y_test_small, X_valid_nrm, y_valid_small)\n        acc.append(model_fitted['acc'])\n        scores.append(model_fitted['cv_score'])\n        f1.append(model_fitted['f1_valid'])\n        models_fitted.append(model_fitted['model'])\n        amex_metric.append(model_fitted['amex_metric'])\n        index = 'model'\n\n        dff = pd.DataFrame({'Accuracy': acc,\n                    'Cross Val Score': scores,\n                        'f1_valid':  f1, \n                            'amex_metric':  amex_metric\n                           })\n        dff = dff.sort_values(by=['amex_metric','Accuracy','f1_valid','Cross Val Score'], ascending=False)\n    \n    return dff,models_fitted ","metadata":{"execution":{"iopub.status.busy":"2022-08-24T09:39:38.189081Z","iopub.execute_input":"2022-08-24T09:39:38.189455Z","iopub.status.idle":"2022-08-24T09:39:38.202424Z","shell.execute_reply.started":"2022-08-24T09:39:38.189406Z","shell.execute_reply":"2022-08-24T09:39:38.201489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"results_orig, models_fitted_orig = fast_testing(X_train_rsm, y_train_rsm, X_valid_orig,y_valid_orig, models, scaler) # Original data\nresults_flt, models_fitted_flt = fast_testing(X_out_sm,Y_out_sm, X_valid_orig,y_valid_orig, models, scaler) # Removed outliners data  ","metadata":{"execution":{"iopub.status.busy":"2022-08-12T16:44:53.818711Z","iopub.execute_input":"2022-08-12T16:44:53.819746Z","iopub.status.idle":"2022-08-12T17:59:18.031292Z","shell.execute_reply.started":"2022-08-12T16:44:53.819704Z","shell.execute_reply":"2022-08-12T17:59:18.030016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Original data results:**","metadata":{}},{"cell_type":"code","source":"results_orig","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:06:46.219345Z","iopub.execute_input":"2022-08-12T18:06:46.220061Z","iopub.status.idle":"2022-08-12T18:06:46.233855Z","shell.execute_reply.started":"2022-08-12T18:06:46.220015Z","shell.execute_reply":"2022-08-12T18:06:46.232733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Removed outliners data fast testing results:**","metadata":{}},{"cell_type":"code","source":"results_flt","metadata":{"execution":{"iopub.status.busy":"2022-08-12T18:00:59.212245Z","iopub.execute_input":"2022-08-12T18:00:59.213113Z","iopub.status.idle":"2022-08-12T18:00:59.226460Z","shell.execute_reply.started":"2022-08-12T18:00:59.213070Z","shell.execute_reply":"2022-08-12T18:00:59.225273Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We see that for some of the algorithms it's better to remove outliners, but generally Support Vector Classifier with outliner is better for our case, but it's the slowest algorithm","metadata":{}},{"cell_type":"code","source":"# let's create a dictionary of features and their importance values\n\nfeat_dict = {}\nmodel_fitted = data[1][0]\n\nfor col, val in sorted(zip(X_train.columns, model_fitted.feature_importances_),key=lambda x:x[1],reverse=True):\n    feat_dict[col]=val","metadata":{"execution":{"iopub.status.busy":"2022-08-17T13:07:08.858023Z","iopub.execute_input":"2022-08-17T13:07:08.858761Z","iopub.status.idle":"2022-08-17T13:07:08.883820Z","shell.execute_reply.started":"2022-08-17T13:07:08.858677Z","shell.execute_reply":"2022-08-17T13:07:08.882678Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feat_df = pd.DataFrame({'Feature':feat_dict.keys(),'Importance':feat_dict.values()})","metadata":{"execution":{"iopub.status.busy":"2022-08-17T13:07:08.885050Z","iopub.execute_input":"2022-08-17T13:07:08.885501Z","iopub.status.idle":"2022-08-17T13:07:08.891762Z","shell.execute_reply.started":"2022-08-17T13:07:08.885469Z","shell.execute_reply":"2022-08-17T13:07:08.890552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"feat_df.plot()","metadata":{"execution":{"iopub.status.busy":"2022-08-17T13:07:08.893732Z","iopub.execute_input":"2022-08-17T13:07:08.894272Z","iopub.status.idle":"2022-08-17T13:07:09.176431Z","shell.execute_reply.started":"2022-08-17T13:07:08.894192Z","shell.execute_reply":"2022-08-17T13:07:09.175168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After 75 feature importance is changing much worse, but we will take 160 features and it will be enough to build quality model ","metadata":{}},{"cell_type":"markdown","source":"### Conclusion \n\nFast testing of algorithms using small data shows that we are pretty far from significant high amex metric level (even using more data and changing settings does not help). We will generate more features using historical changes by each customer   ","metadata":{}},{"cell_type":"markdown","source":"# **XGBoost**","metadata":{"execution":{"iopub.status.busy":"2022-08-19T08:30:17.790263Z","iopub.execute_input":"2022-08-19T08:30:17.790724Z","iopub.status.idle":"2022-08-19T08:30:17.796416Z","shell.execute_reply.started":"2022-08-19T08:30:17.790683Z","shell.execute_reply":"2022-08-19T08:30:17.794885Z"}}},{"cell_type":"code","source":"# LOAD LIBRARIES\nimport pandas as pd, numpy as np # CPU libraries\nimport cupy, cudf # GPU libraries\nimport matplotlib.pyplot as plt, gc, os\n\n\n# VERSION NAME FOR SAVED MODEL FILES\nVER = 1\n\n# TRAIN RANDOM SEED\nSEED = 42\n\n# FILL NAN VALUE\nNAN_VALUE = -127 # will fit in int8\n\n# FOLDS PER MODEL\nFOLDS = 5","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:05:55.919906Z","iopub.execute_input":"2022-08-24T18:05:55.920949Z","iopub.status.idle":"2022-08-24T18:05:58.873893Z","shell.execute_reply.started":"2022-08-24T18:05:55.920897Z","shell.execute_reply":"2022-08-24T18:05:58.872723Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# data from \n# https://www.kaggle.com/datasets/raddar/amex-data-integer-dtypes-parquet-format\n\n\n# approach to load data\n# https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793#XGBoost-Starter---LB-0.793\n\nprint('RAPIDS version',cudf.__version__)\n\ndef read_file(path = '', usecols = None):\n    # LOAD DATAFRAME\n    if usecols is not None: df = cudf.read_parquet(path, columns=usecols)\n    else: df = cudf.read_parquet(path)\n    # REDUCE DTYPE FOR CUSTOMER AND DATE\n    df['customer_ID'] = df['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\n    df.S_2 = cudf.to_datetime( df.S_2 )\n    # SORT BY CUSTOMER AND DATE (so agg('last') works correctly)\n    #df = df.sort_values(['customer_ID','S_2'])\n    #df = df.reset_index(drop=True)\n    # FILL NAN\n    df = df.fillna(NAN_VALUE) \n    print('shape of data:', df.shape)\n    \n    return df\n\nprint('Reading train data...')\nTRAIN_PATH = '../input/amex-data-integer-dtypes-parquet-format/train.parquet'\ntrain = read_file(path = TRAIN_PATH)\n\n# ADD TARGETS\ntargets = cudf.read_csv('../input/amex-default-prediction/train_labels.csv')\ntargets['customer_ID'] = targets['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\ntargets = targets.set_index('customer_ID')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:05:58.875667Z","iopub.execute_input":"2022-08-24T18:05:58.876026Z","iopub.status.idle":"2022-08-24T18:06:23.434273Z","shell.execute_reply.started":"2022-08-24T18:05:58.875972Z","shell.execute_reply":"2022-08-24T18:06:23.433306Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data preporation functions ","metadata":{}},{"cell_type":"code","source":"import sklearn\nfrom sklearn import preprocessing\nfrom sklearn.feature_selection import SelectKBest, chi2\n\ndef normalization (data, scaler=[]):\n    '''\n    normalization (change distribution of values to range (0,1)) \n    of data using created scaler, or creating new one \n    '''\n    \n    if type(scaler) == sklearn.preprocessing._data.MinMaxScaler:\n        min_max_scaler = scaler\n    else: \n        min_max_scaler = preprocessing.MinMaxScaler(feature_range=(0,1)) \n\n    X_nrm = min_max_scaler.fit_transform(data)\n    X_nrm = pd.DataFrame(X_nrm)\n    headers = list(data.columns.values)\n    X_nrm.columns = headers\n    X_nrm.index = data.index\n    \n    del data\n    return X_nrm, min_max_scaler \n\n\ndef feauters_selection_Kbest(data, koef = 0.5, scaler=[]):\n   \n    select = SelectKBest(chi2, k=int(len(data.columns)*koef))\n\n    X, scaler = normalization (data.drop(['target'],axis=1).to_pandas(), scaler)\n    y = data['target'].to_pandas()\n\n    X_new = select.fit_transform( X, y)\n    X_new = pd.DataFrame(X_new)\n    filter = select.get_support()\n    X_train_1 = X[X.columns[filter]]\n\n    del X, y , X_new\n    return X_train_1, scaler","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def process_and_feature_engineer(df, targets=cudf.DataFrame(), is_train_line=True, scaler_df=[], scaler_fr=[], columns_filtered=[]):\n    # FEATURE ENGINEERING PARTIALLY FROM \n    # https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793\n    \n    # sort values by date \n    df=df.sort_values(by='S_2')\n    \n    all_cols = [c for c in list(df.columns) if c not in ['customer_ID','S_2']]\n    cat_features = [\"B_30\",\"B_38\",\"D_114\",\"D_116\",\"D_117\",\"D_120\",\"D_126\",\"D_63\",\"D_64\",\"D_66\",\"D_68\"]\n    num_features = [col for col in all_cols if col not in cat_features and '_ma' not in col]\n    \n    \n     # Calculating features by difference with shifted data \n    fr = df.sort_values(['customer_ID', 'S_2'], \\\n             ascending=[False, True])\\\n                 - \\\n        df.sort_values(['customer_ID', 'S_2'], \\\n                     ascending=[False, True]).shift(1)\n\n    fr.drop('customer_ID',axis=1,inplace=True)\n    fr.rename(columns={\"S_2\": \"dif_S_2\"},inplace=True)\n\n    filter_first = df[['customer_ID', 'S_2']].groupby('customer_ID').agg(['first', 'last', 'count'])\n    filter_first.columns = ['_filter_'.join(x) for x in  filter_first.columns ]\n    filter_first.rename(columns={\"S_2_filter_first\": \"S_2\"},inplace=True)\n    filter_first['life_time'] = (filter_first['S_2_filter_last'].astype(np.int64)   -filter_first['S_2'].astype(np.int64)) // 10 ** 9 \n\n    d_reference = df.sort_values(by =['customer_ID', 'S_2'], \\\n                     ascending=[False, True])[['customer_ID', 'S_2']]\n\n    fr = cudf.concat([fr,d_reference],axis=1) \n\n    del d_reference\n\n    fr = fr.merge(filter_first[['S_2', 'S_2_filter_count']], on=['customer_ID','S_2'], how='left')\n    fr['S_2_filter_count'] = fr['S_2_filter_count'].fillna(0)\n    fr = fr[fr['S_2_filter_count']<=1]\n\n    fr_1_time_ind = np.arange(len(fr[fr['S_2_filter_count']==1 ]))\n    fr_1_time_ind_customer = fr[fr['S_2_filter_count']==1]['customer_ID']\n\n    fr = fr[fr['S_2_filter_count']==0]\n\n    fr['dif_S_2'] = fr['dif_S_2'].astype(np.int64) // 10 ** 9 \n    fr.drop(['S_2', 'S_2_filter_count'],axis=1,inplace=True)\n\n    fr = fr.groupby('customer_ID')[fr.columns[:189]].agg(['sum', 'std','mean', 'last'])\n    fr.columns = ['_shifted_'.join(x) for x in  fr.columns]\n\n    fr_1_time = pd.DataFrame(0, index=fr_1_time_ind, columns=fr.columns)\n    fr_1_time.index = list(fr_1_time_ind_customer.to_pandas())\n    fr_1_time = cudf.DataFrame(fr_1_time)\n\n    fr = cudf.concat([fr,fr_1_time],axis=0) \n    fr = fr.merge(filter_first['life_time'], left_index=True, right_index=True)\n    \n    del filter_first, fr_1_time\n    \n    \n    test_num_agg = df.groupby(\"customer_ID\")[num_features].agg(['mean', 'std', 'min', 'max', 'last'])\n    test_num_agg.columns = ['_'.join(x) for x in test_num_agg.columns]\n\n    test_cat_agg = df.groupby(\"customer_ID\")[cat_features].agg(['count', 'last','first', 'nunique'])\n    test_cat_agg.columns = ['_'.join(x) for x in test_cat_agg.columns]\n    \n    df = cudf.concat([test_num_agg, test_cat_agg], axis=1)\n    \n    del test_num_agg, test_cat_agg\n    \n    df, fr = df.fillna(0.0).astype('float32'), fr.fillna(0.0).astype('float32')\n    \n    # if train then need targets etc. if test needs list of cols and normalization scaler\n    \n    if is_train_line:\n        df, fr = df.merge(targets, left_index=True, right_index=True, how='left'), fr.merge(targets, left_index=True, right_index=True, how='left')\n        df.target, fr.target = df.target.astype('int8'), fr.target.astype('int8')\n        \n#         df, scaler_df = feauters_selection_Kbest(df)\n        df, scaler_df = normalization (df.drop(['target'],axis=1).to_pandas())\n        scaler_df = []\n        fr, scaler_fr = feauters_selection_Kbest(fr)\n        df, fr = cudf.DataFrame(df), cudf.DataFrame(fr)\n    else:\n        df, fr = df[df.columns[df.columns.isin(columns_filtered)]], fr[fr.columns[fr.columns.isin(columns_filtered)]]\n        df, fr = normalization (df.to_pandas())[0], normalization (fr.to_pandas())[0]\n        df, fr = cudf.DataFrame(df), cudf.DataFrame(fr)\n        scaler_df, scaler_fr = [], []\n    \n    df = cudf.concat([df, fr], axis=1) \n    del fr\n    \n    print('summary number of features after', df.shape[1])\n    return df,  scaler_df, scaler_fr\n","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:28:41.851241Z","iopub.execute_input":"2022-08-24T18:28:41.851687Z","iopub.status.idle":"2022-08-24T18:28:41.884483Z","shell.execute_reply.started":"2022-08-24T18:28:41.851645Z","shell.execute_reply":"2022-08-24T18:28:41.883251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ntrain = process_and_feature_engineer(train, targets)[0]","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:06:23.476238Z","iopub.execute_input":"2022-08-24T18:06:23.476592Z","iopub.status.idle":"2022-08-24T18:06:53.783220Z","shell.execute_reply.started":"2022-08-24T18:06:23.476557Z","shell.execute_reply":"2022-08-24T18:06:53.782133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = train.merge(targets, left_index=True, right_index=True, how='left')\ntrain.target = train.target.astype('int8')\ndel targets","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:07:20.223780Z","iopub.execute_input":"2022-08-24T18:07:20.224798Z","iopub.status.idle":"2022-08-24T18:07:20.677352Z","shell.execute_reply.started":"2022-08-24T18:07:20.224750Z","shell.execute_reply":"2022-08-24T18:07:20.676377Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = train.sort_index().reset_index()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:07:22.293527Z","iopub.execute_input":"2022-08-24T18:07:22.294652Z","iopub.status.idle":"2022-08-24T18:07:23.379897Z","shell.execute_reply.started":"2022-08-24T18:07:22.294603Z","shell.execute_reply":"2022-08-24T18:07:23.378899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"rows = train.shape[0]\n\n# train, valid = train[:int(rows*.85)], train[int(rows*.85):]","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:07:25.494340Z","iopub.execute_input":"2022-08-24T18:07:25.494778Z","iopub.status.idle":"2022-08-24T18:07:25.500237Z","shell.execute_reply.started":"2022-08-24T18:07:25.494743Z","shell.execute_reply":"2022-08-24T18:07:25.499227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"type(train)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T17:30:47.586046Z","iopub.execute_input":"2022-08-24T17:30:47.587108Z","iopub.status.idle":"2022-08-24T17:30:47.597275Z","shell.execute_reply.started":"2022-08-24T17:30:47.587076Z","shell.execute_reply":"2022-08-24T17:30:47.595689Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Training ","metadata":{}},{"cell_type":"code","source":"# approach to train model from\n# https://www.kaggle.com/code/cdeotte/xgboost-starter-0-793#XGBoost-Starter---LB-0.793\n\n\n# LOAD XGB LIBRARY\nfrom sklearn.model_selection import KFold\nimport xgboost as xgb\nprint('XGB Version',xgb.__version__)\n\n# XGB MODEL PARAMETERS\nxgb_parms = { \n    'max_depth':4, \n    'learning_rate':0.05, \n    'subsample':0.8,\n    'colsample_bytree':0.6, \n    'eval_metric':'logloss',\n    'objective':'binary:logistic',\n    'tree_method':'gpu_hist',\n    'predictor':'gpu_predictor',\n    'random_state':SEED\n}\n\nFEATURES = train.columns[1:-1]","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:07:37.462756Z","iopub.execute_input":"2022-08-24T18:07:37.463140Z","iopub.status.idle":"2022-08-24T18:07:37.555016Z","shell.execute_reply.started":"2022-08-24T18:07:37.463106Z","shell.execute_reply":"2022-08-24T18:07:37.553506Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# NEEDED WITH DeviceQuantileDMatrix BELOW\nclass IterLoadForDMatrix(xgb.core.DataIter):\n    def __init__(self, df=None, features=None, target=None, batch_size=256*1024):\n        self.features = features\n        self.target = target\n        self.df = df\n        self.it = 0 # set iterator to 0\n        self.batch_size = batch_size\n        self.batches = int( np.ceil( len(df) / self.batch_size ) )\n        super().__init__()\n\n    def reset(self):\n        '''Reset the iterator'''\n        self.it = 0\n\n    def next(self, input_data):\n        '''Yield next batch of data.'''\n        if self.it == self.batches:\n            return 0 # Return 0 when there's no more batch.\n        \n        a = self.it * self.batch_size\n        b = min( (self.it + 1) * self.batch_size, len(self.df) )\n        dt = cudf.DataFrame(self.df.iloc[a:b])\n        input_data(data=dt[self.features], label=dt[self.target]) #, weight=dt['weight'])\n        self.it += 1\n        return 1","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:07:42.170640Z","iopub.execute_input":"2022-08-24T18:07:42.171659Z","iopub.status.idle":"2022-08-24T18:07:42.181190Z","shell.execute_reply.started":"2022-08-24T18:07:42.171611Z","shell.execute_reply":"2022-08-24T18:07:42.179927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def amex_metric_np(y_true, y_pred):\n    \n    if type(y_true) != np.ndarray:\n        try:\n            y_true = y_true.numpy()\n        except:\n            y_true = y_true.eval(session=tf.compat.v1.Session())\n            \n    if type(y_pred) != np.ndarray:\n        try:\n            y_pred = y_pred.numpy()\n        except:\n            y_pred = y_pred.eval(session=tf.compat.v1.Session())\n        \n    labels = np.transpose(np.array([y_true, y_pred]))\n    labels = labels[labels[:, 1].argsort()[::-1]]\n    weights = np.where(labels[:,0]==0, 20, 1)\n    cut_vals = labels[np.cumsum(weights) <= int(0.04 * np.sum(weights))]\n    top_four = np.sum(cut_vals[:,0]) / np.sum(labels[:,0])\n    gini = [0,0]\n    for i in [1,0]:\n        labels = np.transpose(np.array([y_true, y_pred]))\n        labels = labels[labels[:, i].argsort()[::-1]]\n        weight = np.where(labels[:,0]==0, 20, 1)\n        weight_random = np.cumsum(weight / np.sum(weight))\n        total_pos = np.sum(labels[:, 0] *  weight)\n        cum_pos_found = np.cumsum(labels[:, 0] * weight)\n        lorentz = cum_pos_found / total_pos\n        gini[i] = np.sum((lorentz - weight_random) * weight)\n    return 0.5 * (gini[1]/gini[0] + top_four)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:07:46.539706Z","iopub.execute_input":"2022-08-24T18:07:46.540136Z","iopub.status.idle":"2022-08-24T18:07:46.559223Z","shell.execute_reply.started":"2022-08-24T18:07:46.540099Z","shell.execute_reply":"2022-08-24T18:07:46.558329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"importances = []\noof = []\ntrain = train.to_pandas() # free GPU memory\nTRAIN_SUBSAMPLE = 1.0\ngc.collect()\n\nskf = KFold(n_splits=FOLDS, shuffle=True, random_state=SEED)\nfor fold,(train_idx, valid_idx) in enumerate(skf.split(\n            train, train.target )):\n    \n    # TRAIN WITH SUBSAMPLE OF TRAIN FOLD DATA\n    if TRAIN_SUBSAMPLE<1.0:\n        np.random.seed(SEED)\n        train_idx = np.random.choice(train_idx, \n                       int(len(train_idx)*TRAIN_SUBSAMPLE), replace=False)\n        np.random.seed(None)\n    \n    print('#'*25)\n    print('### Fold',fold+1)\n    print('### Train size',len(train_idx),'Valid size',len(valid_idx))\n    print(f'### Training with {int(TRAIN_SUBSAMPLE*100)}% fold data...')\n    print('#'*25)\n    \n    # TRAIN, VALID, TEST FOR FOLD K\n    Xy_train = IterLoadForDMatrix(train.loc[train_idx], FEATURES, 'target')\n    X_valid = train.loc[valid_idx, FEATURES]\n    y_valid = train.loc[valid_idx, 'target']\n    \n    dtrain = xgb.DeviceQuantileDMatrix(Xy_train, max_bin=256)\n    dvalid = xgb.DMatrix(data=X_valid, label=y_valid)\n    \n    # TRAIN MODEL FOLD K\n    model = xgb.train(xgb_parms, \n                dtrain=dtrain,\n                evals=[(dtrain,'train'),(dvalid,'valid')],\n                num_boost_round=9999,\n                early_stopping_rounds=100,\n                verbose_eval=100) \n    model.save_model(f'XGB_v{VER}_fold{fold}.xgb')\n    \n    # GET FEATURE IMPORTANCE FOR FOLD K\n    dd = model.get_score(importance_type='weight')\n    df = pd.DataFrame({'feature':dd.keys(),f'importance_{fold}':dd.values()})\n    importances.append(df)\n            \n    # INFER OOF FOLD K\n    oof_preds = model.predict(dvalid)\n    acc = amex_metric_np(y_valid.values, oof_preds)\n    print('Kaggle Metric =',acc,'\\n')\n    \n    # SAVE OOF\n    df = train.loc[valid_idx, ['customer_ID','target'] ].copy()\n    df['oof_pred'] = oof_preds\n    oof.append( df )\n    \n    del dtrain, Xy_train, dd, df\n    del X_valid, y_valid, dvalid, model\n    _ = gc.collect()\n    \nprint('#'*25)\noof = pd.concat(oof,axis=0,ignore_index=True).set_index('customer_ID')\nacc = amex_metric_np(oof.target.values, oof.oof_pred.values)\nprint('OVERALL CV Kaggle Metric =',acc)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:07:56.317614Z","iopub.execute_input":"2022-08-24T18:07:56.318109Z","iopub.status.idle":"2022-08-24T18:18:17.582926Z","shell.execute_reply.started":"2022-08-24T18:07:56.318066Z","shell.execute_reply":"2022-08-24T18:18:17.581863Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# TEST DATA FOR XGB\nX_valid = valid[FEATURES]\ndvalid = xgb.DMatrix(data=X_valid)\n# valid = valid[['P_2_mean']] # reduce memory\ndel X_valid\ngc.collect()\n\n# INFER XGB MODELS ON TEST DATA\nmodel = xgb.Booster()\nmodel.load_model('./XGB_v1_fold4.xgb')\npreds = model.predict(dvalid)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T12:34:15.595449Z","iopub.execute_input":"2022-08-24T12:34:15.596203Z","iopub.status.idle":"2022-08-24T12:34:21.300962Z","shell.execute_reply.started":"2022-08-24T12:34:15.596153Z","shell.execute_reply":"2022-08-24T12:34:21.299366Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.metrics import classification_report\n\nprint(classification_report(valid['target'].to_pandas() , preds , target_names=['0','1'])) ","metadata":{"execution":{"iopub.status.busy":"2022-08-24T12:32:26.200991Z","iopub.execute_input":"2022-08-24T12:32:26.201680Z","iopub.status.idle":"2022-08-24T12:32:26.521076Z","shell.execute_reply.started":"2022-08-24T12:32:26.201634Z","shell.execute_reply":"2022-08-24T12:32:26.519511Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"amex_metric_np( valid['target'].to_pandas().to_numpy() , preds)","metadata":{"execution":{"iopub.status.busy":"2022-08-24T12:34:23.872184Z","iopub.execute_input":"2022-08-24T12:34:23.873842Z","iopub.status.idle":"2022-08-24T12:34:23.912638Z","shell.execute_reply.started":"2022-08-24T12:34:23.873809Z","shell.execute_reply":"2022-08-24T12:34:23.911216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## **Load test data**","metadata":{}},{"cell_type":"code","source":"del train\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:18:17.585239Z","iopub.execute_input":"2022-08-24T18:18:17.585659Z","iopub.status.idle":"2022-08-24T18:18:17.735139Z","shell.execute_reply.started":"2022-08-24T18:18:17.585619Z","shell.execute_reply":"2022-08-24T18:18:17.733860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pickle\n\nopen_file = open('../input/cols-train-list/cols_train.pkl', \"rb\") # cols filter for model \ncolumns = pickle.load(open_file)\nopen_file.close()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:18:17.737058Z","iopub.execute_input":"2022-08-24T18:18:17.737857Z","iopub.status.idle":"2022-08-24T18:18:17.757202Z","shell.execute_reply.started":"2022-08-24T18:18:17.737815Z","shell.execute_reply":"2022-08-24T18:18:17.756364Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# CALCULATE SIZE OF EACH SEPARATE TEST PART\ndef get_rows(customers, test, NUM_PARTS = 4, verbose = ''):\n    chunk = len(customers)//NUM_PARTS\n    if verbose != '':\n        print(f'We will process {verbose} data as {NUM_PARTS} separate parts.')\n        print(f'There will be {chunk} customers in each part (except the last part).')\n        print('Below are number of rows in each part:')\n    rows = []\n\n    for k in range(NUM_PARTS):\n        if k==NUM_PARTS-1: cc = customers[k*chunk:]\n        else: cc = customers[k*chunk:(k+1)*chunk]\n        s = test.loc[test.customer_ID.isin(cc)].shape[0]\n        rows.append(s)\n    if verbose != '': print( rows )\n    return rows,chunk\n\n# COMPUTE SIZE OF 4 PARTS FOR TEST DATA\nNUM_PARTS = 4\nTEST_PATH = '../input/amex-data-integer-dtypes-parquet-format/test.parquet'\n\nprint(f'Reading test data...')\ntest = read_file(path = TEST_PATH, usecols = ['customer_ID','S_2'])\ncustomers = test[['customer_ID']].drop_duplicates().sort_index().values.flatten()\nrows,num_cust = get_rows(customers, test[['customer_ID']], NUM_PARTS = NUM_PARTS, verbose = 'test')","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:20:49.293826Z","iopub.execute_input":"2022-08-24T18:20:49.294387Z","iopub.status.idle":"2022-08-24T18:20:50.370102Z","shell.execute_reply.started":"2022-08-24T18:20:49.294342Z","shell.execute_reply":"2022-08-24T18:20:50.369045Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"del test\n_ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:33:39.585508Z","iopub.execute_input":"2022-08-24T18:33:39.585859Z","iopub.status.idle":"2022-08-24T18:33:40.086037Z","shell.execute_reply.started":"2022-08-24T18:33:39.585828Z","shell.execute_reply":"2022-08-24T18:33:40.084948Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def read_file(path = '', usecols = None):\n    # LOAD DATAFRAME\n    if usecols is not None: df = cudf.read_parquet(path, columns=usecols)\n    else: df = cudf.read_parquet(path)\n    # REDUCE DTYPE FOR CUSTOMER AND DATE\n    df['customer_ID'] = df['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\n    df.S_2 = cudf.to_datetime( df.S_2 )\n    # SORT BY CUSTOMER AND DATE (so agg('last') works correctly)\n    #df = df.sort_values(['customer_ID','S_2'])\n    #df = df.reset_index(drop=True)\n    # FILL NAN\n    df = df.fillna(NAN_VALUE) \n    print('shape of data:', df.shape)\n    \n    return df\n\nprint('Reading train data...')\nTRAIN_PATH = '../input/amex-data-integer-dtypes-parquet-format/train.parquet'\ntrain = read_file(path = TRAIN_PATH)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# INFER TEST DATA IN PARTS\nskip_rows = 0\nskip_cust = 0\ntest_preds = []\n\nfor k in range(NUM_PARTS):\n    \n    # READ PART OF TEST DATA\n    print(f'\\nReading test data...')\n    test = read_file(path = TEST_PATH)\n    test = test.iloc[skip_rows:skip_rows+rows[k]]\n    skip_rows += rows[k]\n    print(f'=> Test part {k+1} has shape', test.shape )\n    \n    # PROCESS AND FEATURE ENGINEER PART OF TEST DATA\n    test = process_and_feature_engineer(test, is_train_line=False, columns_filtered=columns)[0]\n    if k==NUM_PARTS-1: test = test.loc[customers[skip_cust:]]\n    else: test = test.loc[customers[skip_cust:skip_cust+num_cust]]\n    skip_cust += num_cust\n    \n    # TEST DATA FOR XGB\n    X_test = test[columns]\n    dtest = xgb.DMatrix(data=X_test)\n    test = test[['P_2_mean']] # reduce memory\n    del X_test\n    gc.collect()\n\n    # INFER XGB MODELS ON TEST DATA\n    model = xgb.Booster()\n    model.load_model(f'XGB_v{VER}_fold0.xgb')\n    preds = model.predict(dtest)\n    for f in range(1,FOLDS):\n        model.load_model(f'XGB_v{VER}_fold{f}.xgb')\n        preds += model.predict(dtest)\n    preds /= FOLDS\n    test_preds.append(preds)\n\n    # CLEAN MEMORY\n    del dtest, model\n    _ = gc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:33:43.901768Z","iopub.execute_input":"2022-08-24T18:33:43.902153Z","iopub.status.idle":"2022-08-24T18:37:19.980357Z","shell.execute_reply.started":"2022-08-24T18:33:43.902123Z","shell.execute_reply":"2022-08-24T18:37:19.979329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# WRITE SUBMISSION FILE\ntest_preds = np.concatenate(test_preds)\ntest = cudf.DataFrame(index=customers,data={'prediction':test_preds})\nsub = cudf.read_csv('../input/amex-default-prediction/sample_submission.csv')[['customer_ID']]\nsub['customer_ID_hash'] = sub['customer_ID'].str[-16:].str.hex_to_int().astype('int64')\nsub = sub.set_index('customer_ID_hash')\nsub = sub.merge(test[['prediction']], left_index=True, right_index=True, how='left')\nsub = sub.reset_index(drop=True)\n\n# DISPLAY PREDICTIONS\nsub.to_csv(f'submission_xgb_v{VER}.csv',index=False)\nprint('Submission file shape is', sub.shape )\nsub.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-24T18:39:48.417481Z","iopub.execute_input":"2022-08-24T18:39:48.418180Z","iopub.status.idle":"2022-08-24T18:39:49.354158Z","shell.execute_reply.started":"2022-08-24T18:39:48.418143Z","shell.execute_reply":"2022-08-24T18:39:49.353025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# **Summary conclusion**\n\n**Last statement data analysis done:**\n\n* Association rules found\n* Outliers found and removed\n* Oversampling approaches compared and data resampled\n\n**Fast education and chosing ml algorithm**\n\n* Built pipeline of fast training and comparing of ml algorithms\n* Algorithms compared on data with outliers and without them\n\n**XGBoost**\n\n* Built pipeline of data preparation and load using GPU\n* Model trained and validated \n* Made prediction \n\n\n\n\n","metadata":{}}]}