{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# American Express Credit Card Default Dataset Overview\n\nAmerican Express is a globally integrated payments company. The largest payment card issuer in the world, they provide customers with access to products, insights, and experiences that enrich lives and build business success.\n\nThe objective of this competition is to predict the probability that a  customer does not pay back their credit card balance amount in the  future based on their monthly customer profile. The target binary  variable is calculated by observing 18 months performance window after  the latest credit card statement, and if the customer does not pay the due  amount in 120 days after their latest statement date it is considered a  default event.\n\nThe dataset contains aggregated profile features for each customer at  each statement date. Features are anonymized and normalized, and fall  into the following general categories:\n\nD_* = Delinquency variables\nS_* = Spend variables\nP_* = Payment variables\nB_* = Balance variables\nR_* = Risk variables\n\nwith the following features being categorical:","metadata":{}},{"cell_type":"markdown","source":"# Importing packages","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport seaborn as sns\nimport matplotlib.pyplot as plt\n\n%matplotlib inline\n\nimport warnings\nwarnings.filterwarnings(\"ignore\")","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-08-18T16:22:20.545185Z","iopub.execute_input":"2023-08-18T16:22:20.546299Z","iopub.status.idle":"2023-08-18T16:22:22.223162Z","shell.execute_reply.started":"2023-08-18T16:22:20.546256Z","shell.execute_reply":"2023-08-18T16:22:22.222132Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing the train data and labels and merging them","metadata":{}},{"cell_type":"code","source":"train_data = pd.read_csv('/kaggle/input/amex-default-prediction/train_data.csv', nrows=20000)\ntrain_labels = pd.read_csv('/kaggle/input/amex-default-prediction/train_labels.csv')","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:22.225198Z","iopub.execute_input":"2023-08-18T16:22:22.225866Z","iopub.status.idle":"2023-08-18T16:22:25.368740Z","shell.execute_reply.started":"2023-08-18T16:22:22.225832Z","shell.execute_reply":"2023-08-18T16:22:25.367608Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df = pd.merge(train_data, train_labels, how='inner', left_on=['customer_ID'], right_on=['customer_ID'])","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:25.370260Z","iopub.execute_input":"2023-08-18T16:22:25.370597Z","iopub.status.idle":"2023-08-18T16:22:25.686425Z","shell.execute_reply.started":"2023-08-18T16:22:25.370562Z","shell.execute_reply":"2023-08-18T16:22:25.685201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Top 5 rows of the dataset","metadata":{}},{"cell_type":"code","source":"df.head()","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:25.689115Z","iopub.execute_input":"2023-08-18T16:22:25.689483Z","iopub.status.idle":"2023-08-18T16:22:25.726254Z","shell.execute_reply.started":"2023-08-18T16:22:25.689450Z","shell.execute_reply":"2023-08-18T16:22:25.724965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.info()","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:25.727727Z","iopub.execute_input":"2023-08-18T16:22:25.728115Z","iopub.status.idle":"2023-08-18T16:22:25.755534Z","shell.execute_reply.started":"2023-08-18T16:22:25.728083Z","shell.execute_reply":"2023-08-18T16:22:25.754353Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:25.757179Z","iopub.execute_input":"2023-08-18T16:22:25.757936Z","iopub.status.idle":"2023-08-18T16:22:25.814172Z","shell.execute_reply.started":"2023-08-18T16:22:25.757892Z","shell.execute_reply":"2023-08-18T16:22:25.812665Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Removing the columns with more than 80% NaN values","metadata":{}},{"cell_type":"code","source":"mask = df.isnull().mean() > 0.8\ndf.drop(df.columns[mask], axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:25.815822Z","iopub.execute_input":"2023-08-18T16:22:25.816774Z","iopub.status.idle":"2023-08-18T16:22:25.872730Z","shell.execute_reply.started":"2023-08-18T16:22:25.816725Z","shell.execute_reply":"2023-08-18T16:22:25.871240Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Descriptive Statistics of the DatasetSeparating integer and float columns as mask_int and mask_float","metadata":{}},{"cell_type":"code","source":"df.describe()","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:25.874362Z","iopub.execute_input":"2023-08-18T16:22:25.874724Z","iopub.status.idle":"2023-08-18T16:22:26.556817Z","shell.execute_reply.started":"2023-08-18T16:22:25.874676Z","shell.execute_reply":"2023-08-18T16:22:26.555616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.shape","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:26.558280Z","iopub.execute_input":"2023-08-18T16:22:26.559187Z","iopub.status.idle":"2023-08-18T16:22:26.566518Z","shell.execute_reply.started":"2023-08-18T16:22:26.559152Z","shell.execute_reply":"2023-08-18T16:22:26.565274Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Separating integer and float columns as mask_int and mask_float","metadata":{}},{"cell_type":"code","source":"mask_int = df.dtypes == int\ndf_int = df.columns[mask_int]\nprint(df_int)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:26.571963Z","iopub.execute_input":"2023-08-18T16:22:26.572583Z","iopub.status.idle":"2023-08-18T16:22:26.582094Z","shell.execute_reply.started":"2023-08-18T16:22:26.572532Z","shell.execute_reply":"2023-08-18T16:22:26.580409Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask_float = df.dtypes == float\ndf_cols = df.columns[mask_float]\nprint(df_cols)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:26.583699Z","iopub.execute_input":"2023-08-18T16:22:26.584186Z","iopub.status.idle":"2023-08-18T16:22:26.597501Z","shell.execute_reply.started":"2023-08-18T16:22:26.584142Z","shell.execute_reply":"2023-08-18T16:22:26.596086Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Top 10 columns with highest correlations","metadata":{}},{"cell_type":"code","source":"corr = df[df_cols].corrwith(df.target)\ncorr = corr.abs()\ncorr.sort_values(inplace=True, ascending=False)\ncorr[:10]","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:26.599281Z","iopub.execute_input":"2023-08-18T16:22:26.599991Z","iopub.status.idle":"2023-08-18T16:22:26.741987Z","shell.execute_reply.started":"2023-08-18T16:22:26.599948Z","shell.execute_reply":"2023-08-18T16:22:26.740803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"high_corr = corr[:10].index\nhigh_corr  = list(high_corr)\nhigh_corr.append('target')","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:26.743587Z","iopub.execute_input":"2023-08-18T16:22:26.744137Z","iopub.status.idle":"2023-08-18T16:22:26.750389Z","shell.execute_reply.started":"2023-08-18T16:22:26.744095Z","shell.execute_reply":"2023-08-18T16:22:26.749074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Pairplot of the columns\n","metadata":{}},{"cell_type":"code","source":"sns.pairplot(df[high_corr], hue='target')","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:22:26.751845Z","iopub.execute_input":"2023-08-18T16:22:26.752274Z","iopub.status.idle":"2023-08-18T16:25:55.198695Z","shell.execute_reply.started":"2023-08-18T16:22:26.752235Z","shell.execute_reply":"2023-08-18T16:25:55.197349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"mask = df.dtypes == object\nmask['customer_ID'] = False\ndf_cat = df.columns[mask]\nprint(df_cat)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:55.201453Z","iopub.execute_input":"2023-08-18T16:25:55.202073Z","iopub.status.idle":"2023-08-18T16:25:55.211284Z","shell.execute_reply.started":"2023-08-18T16:25:55.201999Z","shell.execute_reply":"2023-08-18T16:25:55.210234Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Number of columns to create for One Hot Encoding","metadata":{}},{"cell_type":"code","source":"# Determine how many extra columns would be created\nnum_ohc_cols = (df[df_cat].apply(lambda x: x.nunique()).sort_values(ascending=False))\n\n# No need to encode if there is only one value\nsmall_num_ohc_cols = num_ohc_cols.loc[num_ohc_cols>1]\n\n# Number of one-hot columns is one less than the number of categories\nsmall_num_ohc_cols -= 1\n\nsmall_num_ohc_cols.sum()","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:55.212824Z","iopub.execute_input":"2023-08-18T16:25:55.213201Z","iopub.status.idle":"2023-08-18T16:25:55.245950Z","shell.execute_reply.started":"2023-08-18T16:25:55.213169Z","shell.execute_reply":"2023-08-18T16:25:55.244521Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# One Hot Encoding and Label Encoding","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import OneHotEncoder, LabelEncoder","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:55.248191Z","iopub.execute_input":"2023-08-18T16:25:55.248909Z","iopub.status.idle":"2023-08-18T16:25:55.355495Z","shell.execute_reply.started":"2023-08-18T16:25:55.248860Z","shell.execute_reply":"2023-08-18T16:25:55.354289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# The encoders\nle = LabelEncoder()\nohc = OneHotEncoder()\nfor col in num_ohc_cols.index:\n    \n    # Integer encode the string categories\n    dat = le.fit_transform(df[col]).astype(int)\n    \n    # One hot encode the data--this returns a sparse array\n    new_dat = ohc.fit_transform(dat.reshape(-1,1))\n\n    # Create unique column names\n    n_cols = new_dat.shape[1]\n    col_names = ['_cat_'.join([col, str(x)]) for x in range(n_cols)]\n\n    # Create the new dataframe\n    new_df = pd.DataFrame(new_dat.toarray(), \n                          index=df.index, \n                          columns=col_names)\n    \n    # Append the new data to the dataframe\n    df = pd.concat([df, new_df], axis=1)\n    \n    # Remove the original column from the dataframe\n    df = df.drop(col, axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:55.357340Z","iopub.execute_input":"2023-08-18T16:25:55.357822Z","iopub.status.idle":"2023-08-18T16:25:55.801177Z","shell.execute_reply.started":"2023-08-18T16:25:55.357776Z","shell.execute_reply":"2023-08-18T16:25:55.800178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.head().T","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:55.802354Z","iopub.execute_input":"2023-08-18T16:25:55.802682Z","iopub.status.idle":"2023-08-18T16:25:55.820977Z","shell.execute_reply.started":"2023-08-18T16:25:55.802653Z","shell.execute_reply":"2023-08-18T16:25:55.819998Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Filling up the NaN cells with mean values","metadata":{}},{"cell_type":"code","source":"for col in df.columns:\n    if col in df_int or col in df_cols:\n        mean_value = df[col].mean()\n        print('Filling NAN of {} with mean value of {}'.format(col, mean_value))\n        df[col].fillna(value=mean_value, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:55.822388Z","iopub.execute_input":"2023-08-18T16:25:55.822929Z","iopub.status.idle":"2023-08-18T16:25:55.921421Z","shell.execute_reply.started":"2023-08-18T16:25:55.822898Z","shell.execute_reply":"2023-08-18T16:25:55.920378Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Scaling the data using Min Max Scaler","metadata":{}},{"cell_type":"code","source":"from sklearn.preprocessing import MinMaxScaler","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.049139Z","iopub.execute_input":"2023-08-18T16:25:56.049498Z","iopub.status.idle":"2023-08-18T16:25:56.054646Z","shell.execute_reply.started":"2023-08-18T16:25:56.049463Z","shell.execute_reply":"2023-08-18T16:25:56.053333Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Assigning the independent X and dependent y variables","metadata":{}},{"cell_type":"code","source":"X = df.drop(['customer_ID','target'], axis=1)\ny=df['target']","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.056382Z","iopub.execute_input":"2023-08-18T16:25:56.057003Z","iopub.status.idle":"2023-08-18T16:25:56.095214Z","shell.execute_reply.started":"2023-08-18T16:25:56.056959Z","shell.execute_reply":"2023-08-18T16:25:56.094084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Scaling the independent variables","metadata":{}},{"cell_type":"code","source":"mm_scaler = MinMaxScaler()\nX = pd.DataFrame(mm_scaler.fit_transform(X), columns=X.columns)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.096666Z","iopub.execute_input":"2023-08-18T16:25:56.097107Z","iopub.status.idle":"2023-08-18T16:25:56.357374Z","shell.execute_reply.started":"2023-08-18T16:25:56.097067Z","shell.execute_reply":"2023-08-18T16:25:56.356064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing train test split","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split","metadata":{"execution":{"iopub.status.busy":"2023-08-18T17:51:54.445068Z","iopub.execute_input":"2023-08-18T17:51:54.445432Z","iopub.status.idle":"2023-08-18T17:51:54.450939Z","shell.execute_reply.started":"2023-08-18T17:51:54.445402Z","shell.execute_reply":"2023-08-18T17:51:54.449777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.392796Z","iopub.execute_input":"2023-08-18T16:25:56.393231Z","iopub.status.idle":"2023-08-18T16:25:56.470647Z","shell.execute_reply.started":"2023-08-18T16:25:56.393192Z","shell.execute_reply":"2023-08-18T16:25:56.469530Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting the count of the target variable","metadata":{}},{"cell_type":"code","source":"y_train.value_counts().plot.bar(color=['green', 'red'])","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.472002Z","iopub.execute_input":"2023-08-18T16:25:56.472426Z","iopub.status.idle":"2023-08-18T16:25:56.737445Z","shell.execute_reply.started":"2023-08-18T16:25:56.472393Z","shell.execute_reply":"2023-08-18T16:25:56.736006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing various metrics to assess out model\n","metadata":{}},{"cell_type":"code","source":"from sklearn import metrics\nfrom sklearn.metrics import classification_report, accuracy_score, precision_recall_fscore_support, confusion_matrix, precision_score, recall_score, roc_auc_score, f1_score,r2_score\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.metrics import ConfusionMatrixDisplay","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.745621Z","iopub.execute_input":"2023-08-18T16:25:56.746039Z","iopub.status.idle":"2023-08-18T16:25:56.752582Z","shell.execute_reply.started":"2023-08-18T16:25:56.745988Z","shell.execute_reply":"2023-08-18T16:25:56.751084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Logistic regression","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.753962Z","iopub.execute_input":"2023-08-18T16:25:56.754446Z","iopub.status.idle":"2023-08-18T16:25:56.877315Z","shell.execute_reply.started":"2023-08-18T16:25:56.754405Z","shell.execute_reply":"2023-08-18T16:25:56.876080Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"lr=LogisticRegression()\nlr.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:56.878771Z","iopub.execute_input":"2023-08-18T16:25:56.879493Z","iopub.status.idle":"2023-08-18T16:25:57.878394Z","shell.execute_reply.started":"2023-08-18T16:25:56.879460Z","shell.execute_reply":"2023-08-18T16:25:57.876833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_lr=lr.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:57.880776Z","iopub.execute_input":"2023-08-18T16:25:57.881790Z","iopub.status.idle":"2023-08-18T16:25:57.914743Z","shell.execute_reply.started":"2023-08-18T16:25:57.881732Z","shell.execute_reply":"2023-08-18T16:25:57.913135Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm_lr=confusion_matrix(y_test,pred_lr)\nprint(\"Confusion Matrix   :\", cm_lr)\nr2_lr=r2_score(y_test,pred_lr)\nprint(\"R Squared   :\", r2_lr)\naccuracy_lr = accuracy_score(y_test, pred_lr)\nprint(\"Accuracy   :\", accuracy_lr)\nprecision_lr = precision_score(y_test, pred_lr)\nprint(\"Precision :\", precision_lr)\nrecall_lr = recall_score(y_test, pred_lr)\nprint(\"Recall    :\", recall_lr)\nF1_score_lr = f1_score(y_test, pred_lr)\nprint(\"F1-score  :\", F1_score_lr)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:57.917576Z","iopub.execute_input":"2023-08-18T16:25:57.918766Z","iopub.status.idle":"2023-08-18T16:25:57.966250Z","shell.execute_reply.started":"2023-08-18T16:25:57.918704Z","shell.execute_reply":"2023-08-18T16:25:57.964798Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result1 = [\"1.\",\"Logistic Regression\"]\nresult1.append(round(r2_lr,2))\nresult1.append(round(accuracy_lr,2))\nresult1.append(round(precision_lr,2))\nresult1.append(round(recall_lr,2))\nresult1.append(round(F1_score_lr,2))\nprint(result1)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:57.968849Z","iopub.execute_input":"2023-08-18T16:25:57.969957Z","iopub.status.idle":"2023-08-18T16:25:57.981570Z","shell.execute_reply.started":"2023-08-18T16:25:57.969899Z","shell.execute_reply":"2023-08-18T16:25:57.979914Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Random Forest Classifier","metadata":{}},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:57.984347Z","iopub.execute_input":"2023-08-18T16:25:57.985489Z","iopub.status.idle":"2023-08-18T16:25:58.303824Z","shell.execute_reply.started":"2023-08-18T16:25:57.985432Z","shell.execute_reply":"2023-08-18T16:25:58.302515Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_rf = RandomForestClassifier(n_estimators=1000 , oob_score = True, n_jobs = -1, max_features = \"auto\", class_weight=\"balanced\", max_leaf_nodes = 30)\nmodel_rf.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:25:58.305304Z","iopub.execute_input":"2023-08-18T16:25:58.305646Z","iopub.status.idle":"2023-08-18T16:26:29.248366Z","shell.execute_reply.started":"2023-08-18T16:25:58.305617Z","shell.execute_reply":"2023-08-18T16:26:29.247199Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predict_rf = model_rf.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:26:29.249786Z","iopub.execute_input":"2023-08-18T16:26:29.250187Z","iopub.status.idle":"2023-08-18T16:26:29.861365Z","shell.execute_reply.started":"2023-08-18T16:26:29.250156Z","shell.execute_reply":"2023-08-18T16:26:29.860155Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm_rf=confusion_matrix(y_test,predict_rf)\nprint(\"Confusion Matrix   :\", cm_rf)\nr2_rf=r2_score(y_test,pred_lr)\nprint(\"R Squared   :\", r2_rf)\naccuracy_rf = accuracy_score(y_test, predict_rf)\nprint(\"Accuracy   :\", accuracy_rf)\nprecision_rf = precision_score(y_test, predict_rf)\nprint(\"Precision :\", precision_rf)\nrecall_rf = recall_score(y_test, predict_rf)\nprint(\"Recall    :\", recall_rf)\nF1_score_rf = f1_score(y_test, predict_rf)\nprint(\"F1-score  :\", F1_score_rf)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:26:29.862779Z","iopub.execute_input":"2023-08-18T16:26:29.863240Z","iopub.status.idle":"2023-08-18T16:26:29.891965Z","shell.execute_reply.started":"2023-08-18T16:26:29.863201Z","shell.execute_reply":"2023-08-18T16:26:29.890654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result2 = [\"2.\",\"Random Forest Classifier\"]\nresult2.append(round(r2_rf,2))\nresult2.append(round(accuracy_rf,2))\nresult2.append(round(precision_rf,2))\nresult2.append(round(recall_rf,2))\nresult2.append(round(F1_score_rf,2))\nprint(result2)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:26:29.893489Z","iopub.execute_input":"2023-08-18T16:26:29.893937Z","iopub.status.idle":"2023-08-18T16:26:29.901955Z","shell.execute_reply.started":"2023-08-18T16:26:29.893904Z","shell.execute_reply":"2023-08-18T16:26:29.900753Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Support Vector Machine","metadata":{}},{"cell_type":"code","source":"from sklearn.svm import SVC","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:26:29.903615Z","iopub.execute_input":"2023-08-18T16:26:29.904012Z","iopub.status.idle":"2023-08-18T16:26:29.914578Z","shell.execute_reply.started":"2023-08-18T16:26:29.903950Z","shell.execute_reply":"2023-08-18T16:26:29.913336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_svm = SVC(kernel='linear')\nmodel_svm.fit(X_train,y_train)\npred_svm = model_svm.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:26:29.916078Z","iopub.execute_input":"2023-08-18T16:26:29.916425Z","iopub.status.idle":"2023-08-18T16:27:14.626969Z","shell.execute_reply.started":"2023-08-18T16:26:29.916395Z","shell.execute_reply":"2023-08-18T16:27:14.625565Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm_svm=confusion_matrix(y_test,pred_svm)\nprint(\"Confusion Matrix   :\", cm_svm)\nr2_svm=r2_score(y_test,pred_lr)\nprint(\"R Squared   :\", r2_svm)\naccuracy_svm = accuracy_score(y_test, pred_svm)\nprint(\"Accuracy   :\", accuracy_svm)\nprecision_svm = precision_score(y_test, pred_svm)\nprint(\"Precision :\", precision_svm)\nrecall_svm = recall_score(y_test, pred_svm)\nprint(\"Recall    :\", recall_svm)\nF1_score_svm = f1_score(y_test, pred_svm)\nprint(\"F1-score  :\", F1_score_svm)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:14.628322Z","iopub.execute_input":"2023-08-18T16:27:14.628650Z","iopub.status.idle":"2023-08-18T16:27:14.656380Z","shell.execute_reply.started":"2023-08-18T16:27:14.628622Z","shell.execute_reply":"2023-08-18T16:27:14.655174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result3 = [\"3.\",\"Support Vector Machine\"]\nresult3.append(round(r2_svm,2))\nresult3.append(round(accuracy_svm,2))\nresult3.append(round(precision_svm,2))\nresult3.append(round(recall_svm,2))\nresult3.append(round(F1_score_svm,2))\nprint(result3)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:14.660251Z","iopub.execute_input":"2023-08-18T16:27:14.660618Z","iopub.status.idle":"2023-08-18T16:27:14.668248Z","shell.execute_reply.started":"2023-08-18T16:27:14.660590Z","shell.execute_reply":"2023-08-18T16:27:14.666839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Extreme Gradient Boosting (XG Boost)","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBClassifier","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:14.669891Z","iopub.execute_input":"2023-08-18T16:27:14.670346Z","iopub.status.idle":"2023-08-18T16:27:14.861916Z","shell.execute_reply.started":"2023-08-18T16:27:14.670306Z","shell.execute_reply":"2023-08-18T16:27:14.860867Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_xg = XGBClassifier()\nmodel_xg.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:14.863482Z","iopub.execute_input":"2023-08-18T16:27:14.863842Z","iopub.status.idle":"2023-08-18T16:27:47.064832Z","shell.execute_reply.started":"2023-08-18T16:27:14.863812Z","shell.execute_reply":"2023-08-18T16:27:47.063921Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pred_xg = model_xg.predict(X_test)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:47.066160Z","iopub.execute_input":"2023-08-18T16:27:47.066710Z","iopub.status.idle":"2023-08-18T16:27:47.156078Z","shell.execute_reply.started":"2023-08-18T16:27:47.066679Z","shell.execute_reply":"2023-08-18T16:27:47.155148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cm_xg=confusion_matrix(y_test,pred_xg)\nprint(\"Confusion Matrix   :\", cm_xg)\nr2_xg=r2_score(y_test,pred_lr)\nprint(\"R Squared   :\", r2_xg)\naccuracy_xg = accuracy_score(y_test, pred_xg)\nprint(\"Accuracy   :\", accuracy_xg)\nprecision_xg = precision_score(y_test, pred_xg)\nprint(\"Precision :\", precision_xg)\nrecall_xg = recall_score(y_test, pred_xg)\nprint(\"Recall    :\", recall_xg)\nF1_score_xg = f1_score(y_test, pred_xg)\nprint(\"F1-score  :\", F1_score_xg)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:47.157073Z","iopub.execute_input":"2023-08-18T16:27:47.157391Z","iopub.status.idle":"2023-08-18T16:27:47.184201Z","shell.execute_reply.started":"2023-08-18T16:27:47.157363Z","shell.execute_reply":"2023-08-18T16:27:47.183222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result4 = [\"4.\",\"Extreme Gradient Boosting\"]\nresult4.append(round(r2_xg,2))\nresult4.append(round(accuracy_xg,2))\nresult4.append(round(precision_xg,2))\nresult4.append(round(recall_xg,2))\nresult4.append(round(F1_score_xg,2))\nprint(result4)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:47.187752Z","iopub.execute_input":"2023-08-18T16:27:47.188142Z","iopub.status.idle":"2023-08-18T16:27:47.204724Z","shell.execute_reply.started":"2023-08-18T16:27:47.188110Z","shell.execute_reply":"2023-08-18T16:27:47.201201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Result Table of all Models","metadata":{}},{"cell_type":"code","source":"from prettytable import PrettyTable","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:47.206549Z","iopub.execute_input":"2023-08-18T16:27:47.207427Z","iopub.status.idle":"2023-08-18T16:27:47.214068Z","shell.execute_reply.started":"2023-08-18T16:27:47.207359Z","shell.execute_reply":"2023-08-18T16:27:47.212993Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Result_table = PrettyTable([\"S.No.\",\"Model\",\"R Squared Score\",\"Accuracy\",\"Precison\",\"Recall\",\"F1 Score\"])\nResult_table.add_row(result1)\nResult_table.add_row(result2)\nResult_table.add_row(result3)\nResult_table.add_row(result4)\nprint(Result_table)","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:47.215606Z","iopub.execute_input":"2023-08-18T16:27:47.216330Z","iopub.status.idle":"2023-08-18T16:27:47.229096Z","shell.execute_reply.started":"2023-08-18T16:27:47.216298Z","shell.execute_reply":"2023-08-18T16:27:47.227638Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"result = pd.DataFrame(y_test, pred_xg)\nresult.T","metadata":{"execution":{"iopub.status.busy":"2023-08-18T16:27:47.232097Z","iopub.execute_input":"2023-08-18T16:27:47.232749Z","iopub.status.idle":"2023-08-18T16:27:47.265896Z","shell.execute_reply.started":"2023-08-18T16:27:47.232716Z","shell.execute_reply":"2023-08-18T16:27:47.265068Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Conclusion\n\nBy this we have been able to detect fraudulent transactions. We have seen that Extreme Gradient Boosting (XG Boost) has been so far the best model in form of metrics to be able to classify our models.\n\nXG Boost has an accuracy of 97% showing that we can identify 97% of the transactions as whether its fraudulent or not with 94% recall showing the True Positives which is backed by a high F1 score of 94%.\n\nThus we have been able to build a model which suits the business needs and diminishes error.","metadata":{}}]}