{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"In this notebook, we will work on IEEE-CIS Fraud Detection competition. Our aim is to train machine learning model to predict the probability that an online transaction is fraudulent.<br>\n\nThe data is broken into two files identity and transaction, which are joined by TransactionID. Not all transactions have corresponding identity information.\n\n#### Categorical Features - Transaction\nProductCD<br>\ncard1 - card6<br>\naddr1, addr2<br>\nP_emaildomain<br>\nR_emaildomain<br>\nM1 - M9<br>\n\n#### Categorical Features - Identity\nDeviceType<br>\nDeviceInfo<br>\nid_12 - id_38<br>\n\nThe TransactionDT feature is a timedelta from a given reference datetime (not an actual timestamp).\n\n# Agenda\n1. Importing Necessary Libraries<br>\n2. Data Loading, Understanding, and Cleaning<br>\n3. Data Preprocessing<br>\n4. ML Modeling<br>\n5. Prediction Submission<br>\n\nI highly recommend you to go through the following kernels as well. These have helped in shaping my work.<br>\n[EDA and models](https://www.kaggle.com/code/artgor/eda-and-models)<br>\n[Extensive EDA and Modeling XGB Hyperopt](https://www.kaggle.com/code/kabure/extensive-eda-and-modeling-xgb-hyperopt)","metadata":{}},{"cell_type":"markdown","source":"# Importing Necessary Libraries","metadata":{}},{"cell_type":"code","source":"# Data Analysis\nimport pandas as pd\nimport numpy as np\n\n# Data Visualization\nfrom matplotlib import pyplot as plt\nimport seaborn as sns\n\n# Machine Learning\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom imblearn.over_sampling import SMOTE\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.ensemble import RandomForestClassifier\n\n# Warnings\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:35:30.664024Z","iopub.execute_input":"2022-08-05T06:35:30.664383Z","iopub.status.idle":"2022-08-05T06:35:30.671310Z","shell.execute_reply.started":"2022-08-05T06:35:30.664354Z","shell.execute_reply":"2022-08-05T06:35:30.669850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Loading, Understanding, and Cleaning\nSince the data is big in size, we will use function to reduce its memory for fast processing and consuming less storage.","metadata":{}},{"cell_type":"code","source":"# Helper function\ndef reduce_mem_usage(df):\n    \"\"\" iterate through all the columns of a dataframe and modify the data type\n        to reduce memory usage.        \n    \"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        \n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n        else:\n            df[col] = df[col].astype('category')\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:35:30.699307Z","iopub.execute_input":"2022-08-05T06:35:30.701140Z","iopub.status.idle":"2022-08-05T06:35:30.714087Z","shell.execute_reply.started":"2022-08-05T06:35:30.701098Z","shell.execute_reply":"2022-08-05T06:35:30.713059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# loading train_transaction data\ntrain_transaction = pd.read_csv('../input/ieee-fraud-detection/train_transaction.csv')\nprint(train_transaction.shape)\ntrain_transaction = reduce_mem_usage(train_transaction)\ntrain_transaction.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:35:30.740455Z","iopub.execute_input":"2022-08-05T06:35:30.742212Z","iopub.status.idle":"2022-08-05T06:37:39.093122Z","shell.execute_reply.started":"2022-08-05T06:35:30.742182Z","shell.execute_reply":"2022-08-05T06:37:39.091954Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading train_identity data\ntrain_identity = pd.read_csv('../input/ieee-fraud-detection/train_identity.csv')\nprint(train_identity.shape)\ntrain_identity = reduce_mem_usage(train_identity)\ntrain_identity.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:37:39.095461Z","iopub.execute_input":"2022-08-05T06:37:39.096287Z","iopub.status.idle":"2022-08-05T06:37:40.184087Z","shell.execute_reply.started":"2022-08-05T06:37:39.096244Z","shell.execute_reply":"2022-08-05T06:37:40.182667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merging transaction and identity train data\ntrain_df = pd.merge(train_transaction, train_identity, how='left')\nprint(train_df.shape)\nlen_train_df = len(train_df)\ndel train_transaction, train_identity\ntrain_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:37:40.185960Z","iopub.execute_input":"2022-08-05T06:37:40.186693Z","iopub.status.idle":"2022-08-05T06:37:42.433925Z","shell.execute_reply.started":"2022-08-05T06:37:40.186651Z","shell.execute_reply":"2022-08-05T06:37:42.432805Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Loading test data\ntest_transaction = pd.read_csv('../input/ieee-fraud-detection/test_transaction.csv')\nprint(test_transaction.shape)\ntest_transaction = reduce_mem_usage(test_transaction)\n\ntest_identity = pd.read_csv('../input/ieee-fraud-detection/train_identity.csv')\nprint(test_identity.shape)\ntest_identity = reduce_mem_usage(test_identity)\n\ntest_df = pd.merge(test_transaction, test_identity, how='left')\ntest_df.columns = train_df.drop('isFraud', axis=1).columns\nprint(test_df.shape)\ndel test_transaction, test_identity\ntest_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:37:42.437157Z","iopub.execute_input":"2022-08-05T06:37:42.437826Z","iopub.status.idle":"2022-08-05T06:39:35.764052Z","shell.execute_reply.started":"2022-08-05T06:37:42.437789Z","shell.execute_reply":"2022-08-05T06:39:35.762944Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Creating a submission file\nsubmission = pd.DataFrame({'TransactionID':test_df.TransactionID})\nprint(submission.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:39:35.765830Z","iopub.execute_input":"2022-08-05T06:39:35.766223Z","iopub.status.idle":"2022-08-05T06:39:35.776639Z","shell.execute_reply.started":"2022-08-05T06:39:35.766186Z","shell.execute_reply":"2022-08-05T06:39:35.775663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Duplicates check in train data\ntrain_df.duplicated().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:39:35.778049Z","iopub.execute_input":"2022-08-05T06:39:35.779374Z","iopub.status.idle":"2022-08-05T06:39:43.554102Z","shell.execute_reply.started":"2022-08-05T06:39:35.779336Z","shell.execute_reply":"2022-08-05T06:39:43.552121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"So, we have total 434 columns including the dependent variable as 'isFraud'. The training data consists of 5,90,540 samples and test data has 5,06,691 records. We also checked that there are no duplicate records. Let us investigate distribution of the dependent variable.","metadata":{}},{"cell_type":"code","source":"# Class imbalance check\nplt.pie(train_df.isFraud.value_counts(), labels=['Not Fraud', 'Fraud'], autopct='%0.1f%%')\nplt.axis('equal')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:39:43.555848Z","iopub.execute_input":"2022-08-05T06:39:43.556674Z","iopub.status.idle":"2022-08-05T06:39:43.692291Z","shell.execute_reply.started":"2022-08-05T06:39:43.556631Z","shell.execute_reply":"2022-08-05T06:39:43.691050Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As one can expect, this is a class imbalance problem. We will apply SMOTE (Synthetic Minority Over-sampling Technique) to deal with class imbalance in later steps. Let us understand the distribution of the timestamp column.","metadata":{}},{"cell_type":"code","source":"# Timestamp of train and test data\nplt.figure(figsize=(8, 4))\nplt.hist(train_df['TransactionDT'], label='Train')\nplt.hist(test_df['TransactionDT'], label='Test')\nplt.ylabel('Count')\nplt.title('Transaction Timestamp')\nplt.legend()\nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:39:43.698307Z","iopub.execute_input":"2022-08-05T06:39:43.701591Z","iopub.status.idle":"2022-08-05T06:39:44.085931Z","shell.execute_reply.started":"2022-08-05T06:39:43.701553Z","shell.execute_reply":"2022-08-05T06:39:44.084872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can notice that the timestamp of the test data is ahead of the timestamp of the train data. Therefore, while training machine learning model, we need to perform time-based splitting to create training and validation sets. Let us deal with the missing values first.<br> \nThere are considerable number of columns with high missing values. We'll use only those columns that has at least 80% data which leaves 20% to the missing values that can be fillied.","metadata":{}},{"cell_type":"code","source":"# Missing values check\ncombined_df = pd.concat([train_df.drop(columns=['isFraud', 'TransactionID']), test_df.drop(columns='TransactionID')])\nprint(combined_df.shape)\n\n# Dependent variable\ny = train_df['isFraud']\nprint(y.shape)\n\n# Dropping columns with more than 20% missing values \nmv = combined_df.isnull().sum()/len(combined_df)\ncombined_mv_df = combined_df.drop(columns=mv[mv>0.2].index)\ndel combined_df, train_df, test_df\nprint(combined_mv_df.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:39:44.087485Z","iopub.execute_input":"2022-08-05T06:39:44.088089Z","iopub.status.idle":"2022-08-05T06:39:52.425904Z","shell.execute_reply.started":"2022-08-05T06:39:44.088051Z","shell.execute_reply":"2022-08-05T06:39:52.424751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We are left with 180 columns out of 432 after removing features with more than 20% missing values. We also have removed 'TransactionID' column as it does not hold any importance in the prediction. Let us now fill all the missing values. For numerical columns, we will use median value and for categorical column, we will use the most frequent category to fill the missing values.","metadata":{}},{"cell_type":"code","source":"# Filtering numerical data\nnum_mv_df = combined_mv_df.select_dtypes(include=np.number)\nprint(num_mv_df.shape)\n\n# Filtering categorical data\ncat_mv_df = combined_mv_df.select_dtypes(exclude=np.number)\nprint(cat_mv_df.shape)\ndel combined_mv_df\n\n# Filling missing values by median for numerical columns \nimp_median = SimpleImputer(missing_values=np.nan, strategy='median')\nnum_df = pd.DataFrame(imp_median.fit_transform(num_mv_df), columns=num_mv_df.columns)\ndel num_mv_df\nprint(num_df.shape)\n\n# Filling missing values by most frequent value for categorical columns\nimp_max = SimpleImputer(missing_values=np.nan, strategy='most_frequent')\ncat_df = pd.DataFrame(imp_max.fit_transform(cat_mv_df), columns=cat_mv_df.columns)\ndel cat_mv_df\nprint(cat_df.shape)\n\n# Concatinating numerical and categorical data\ncombined_df_cleaned = pd.concat([num_df, cat_df], axis=1)\ndel num_df, cat_df\n\n# Verifying missing values\nprint(f'Total missing values: {combined_df_cleaned.isnull().sum().sum()}')\nprint(combined_df_cleaned.shape)\ncombined_df_cleaned.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:39:52.429756Z","iopub.execute_input":"2022-08-05T06:39:52.430061Z","iopub.status.idle":"2022-08-05T06:40:22.873805Z","shell.execute_reply.started":"2022-08-05T06:39:52.430033Z","shell.execute_reply":"2022-08-05T06:40:22.872642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Preprocessing\nNow, we have dealt with the missing values. Let us perform categorical encoding.","metadata":{}},{"cell_type":"code","source":"# One-hot encoding\ncombined_df_encoded = pd.get_dummies(combined_df_cleaned, drop_first=True)\nprint(combined_df_encoded.shape)\ndel combined_df_cleaned\ncombined_df_encoded.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:40:36.468138Z","iopub.execute_input":"2022-08-05T06:40:36.468487Z","iopub.status.idle":"2022-08-05T06:40:38.186148Z","shell.execute_reply.started":"2022-08-05T06:40:36.468458Z","shell.execute_reply":"2022-08-05T06:40:38.184919Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Separating train and test data\nX = combined_df_encoded.iloc[:len_train_df]\nprint(X.shape)\ntest = combined_df_encoded.iloc[len_train_df:]\nprint(test.shape)\ndel combined_df_encoded","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:40:46.367870Z","iopub.execute_input":"2022-08-05T06:40:46.368251Z","iopub.status.idle":"2022-08-05T06:40:46.375110Z","shell.execute_reply.started":"2022-08-05T06:40:46.368219Z","shell.execute_reply":"2022-08-05T06:40:46.374076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Time-based train validation splitting with 20% data in validation set\ntrain = pd.concat([X, y], axis=1)\ntrain.sort_values('TransactionDT', inplace=True)\nX = train.drop(['isFraud'], axis=1)\ny = train['isFraud']\nsplitting_index = int(0.8*len(X))\nX_train = X.iloc[:splitting_index].values\nX_val = X.iloc[splitting_index:].values\ny_train = y.iloc[:splitting_index].values\ny_val = y.iloc[splitting_index:].values\ntest = test.values\nprint(X_train.shape, X_val.shape, y_train.shape, y_val.shape)\ndel y, train","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:40:49.970336Z","iopub.execute_input":"2022-08-05T06:40:49.970757Z","iopub.status.idle":"2022-08-05T06:40:52.043226Z","shell.execute_reply.started":"2022-08-05T06:40:49.970723Z","shell.execute_reply":"2022-08-05T06:40:52.042088Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Standardization\nscaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_val_scaled = scaler.transform(X_val)\ntest_scaled = scaler.transform(test)\ndel X_train, X_val, test\n\n# Class imbalance check\npd.value_counts(y_train)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:40:52.276883Z","iopub.execute_input":"2022-08-05T06:40:52.277569Z","iopub.status.idle":"2022-08-05T06:40:55.120815Z","shell.execute_reply.started":"2022-08-05T06:40:52.277533Z","shell.execute_reply":"2022-08-05T06:40:55.119773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Applying SMOTE to deal with the class imbalance by oversampling\nsmote = SMOTE()\nX_train_smote, y_train_smote = smote.fit_resample(X_train_scaled, y_train)\nprint(X_train_smote.shape, y_train_smote.shape)\ndel X_train_scaled, y_train\npd.value_counts(y_train_smote)","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:40:56.357795Z","iopub.execute_input":"2022-08-05T06:40:56.358465Z","iopub.status.idle":"2022-08-05T06:41:09.177822Z","shell.execute_reply.started":"2022-08-05T06:40:56.358430Z","shell.execute_reply":"2022-08-05T06:41:09.176866Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ML Modeling\nI tried PCA to reduce dimension but it resulted in relatively poor performance so we will go with this dimensions only.<br>\nAfter hyperparameter tuning, following parameters are selected.","metadata":{}},{"cell_type":"code","source":"# Random Forest Classifier\nrfc = RandomForestClassifier(criterion='entropy', max_features='sqrt', max_samples=0.5, min_samples_split=80)\nrfc.fit(X_train_smote, y_train_smote)\ny_predproba = rfc.predict_proba(X_val_scaled)\nprint(f'Validation AUC={roc_auc_score(y_val, y_predproba[:, 1])}')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:41:09.181258Z","iopub.execute_input":"2022-08-05T06:41:09.184916Z","iopub.status.idle":"2022-08-05T06:45:24.084089Z","shell.execute_reply.started":"2022-08-05T06:41:09.184859Z","shell.execute_reply":"2022-08-05T06:45:24.083038Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Feature importances\npd.Series(rfc.feature_importances_, index=X.columns).nlargest(15).plot(kind='barh')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:45:24.085873Z","iopub.execute_input":"2022-08-05T06:45:24.087045Z","iopub.status.idle":"2022-08-05T06:45:24.132643Z","shell.execute_reply.started":"2022-08-05T06:45:24.087004Z","shell.execute_reply":"2022-08-05T06:45:24.131217Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prediction Submission","metadata":{}},{"cell_type":"code","source":"# Predicting for the test data \npredictions = rfc.predict_proba(test_scaled)\nsubmission['isFraud'] = predictions[:, 1]\nprint(submission.shape)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:45:24.133597Z","iopub.status.idle":"2022-08-05T06:45:24.134257Z","shell.execute_reply.started":"2022-08-05T06:45:24.133966Z","shell.execute_reply":"2022-08-05T06:45:24.133995Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Submitting the predictions\nsubmission.to_csv('submission.csv', index=False)\nprint('Submission is successful!')","metadata":{"execution":{"iopub.status.busy":"2022-08-05T06:40:24.616818Z","iopub.status.idle":"2022-08-05T06:40:24.617322Z","shell.execute_reply.started":"2022-08-05T06:40:24.617083Z","shell.execute_reply":"2022-08-05T06:40:24.617106Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Thank you so much for reading. I hope it was worth your time.<br>\nKindly consider upvoting. It means a lot. :)<br>\nBeing a beginner, I genuinely encourage you to share any views or suggestions you might have on my work.<br>\n\n\nDo check out my earlier work as well.\n* [Customer Segmentation: K-Means and Hierarchical](https://www.kaggle.com/code/pradneshlachake/customer-segmentation-k-means-and-hierarchical)<br>\n* [Time Series Analysis: ETS, ARIMA, SARIMA, Prophet](https://www.kaggle.com/code/pradneshlachake/time-series-analysis-ets-arima-sarima-prophet)<br>\n* [Multiple Linear Regression and Regularization](https://www.kaggle.com/code/pradneshlachake/multiple-linear-regression-and-regularization)<br>\n* [My First Kaggle Project: Titanic Disaster](https://www.kaggle.com/code/pradneshlachake/my-first-kaggle-project-titanic-disaster)","metadata":{}}]}