{"cells":[{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport seaborn as sns\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\n# import os\n# for dirname, _, filenames in os.walk('/kaggle/input'):\n#     for filename in filenames:\n#         print(os.path.join(dirname, filename))\n\n# !nvidia-smi\n# !gcc --version\n# !nvcc -V\n# !ls ../input\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"import matplotlib.pyplot as plt\n%matplotlib inline\n!ls ../input\ndata=pd.read_csv('../input/creditcardfraud/creditcard.csv')\ndata.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Get the general number of fraud and non-fraud transactions.\ncount_class = pd.value_counts(data['Class'], sort=True).sort_index()\nprint(count_class)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Describe the data using box graph.\ndata.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# data.isnull() no missing blank\ndata.isnull().sum().max()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"data.columns","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# print(data['Class'].value_counts()[0]/len(data) * 100)\n# print (len(data))\n# data['Class'].value_counts()和dp.value_counts(data['Class']的效果是一样的)\n# print(data['Class'].value_counts())    \nprint(\"The class of no fraud is： \",round(data['Class'].value_counts()[0]/len(data)*100,2), \"%\")\nprint(\"The class of fraud is： \",round(data['Class'].value_counts()[1]/len(data)*100,2), \"%\")","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## The general understanding of dataframe\n\n**Note:** The dataframe we use is so unbanlanced that we can not get the suitable model through it.\n> Notice how imbalanced is our original dataset! Most of the transactions are non-fraud. If we use this dataframe as the base for our predictive models and analysis we might get a lot of errors and our algorithms will probably overfit since it will \"assume\" that most transactions are not fraud. But we don't want our model to assume, we want our model to detect patterns that give signs of fraud!"},{"metadata":{"trusted":true},"cell_type":"code","source":"\ncolors = [\"#DF0101\", \"#0101DF\"]\n\nsns.countplot('Class', data=data, palette=colors)\nplt.title('Class Distributions \\n (0: No Fraud || 1: Fraud)', fontsize=14)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"* * * * "},{"metadata":{"trusted":true},"cell_type":"code","source":"fig, ax = plt.subplots(1, 2, figsize=(18,4))\n# print(data['Amount'].values)\n# print(data['Time'].values)\n\namount_val = data['Amount'].values\ntime_val = data['Time'].values\nsns.distplot(amount_val, ax=ax[0], color='r')\nax[0].set_title('Distribution of Transaction Amount', fontsize=14)\nax[0].set_xlim(min(amount_val), max(amount_val))\n# print(max(data['Amount']), min(data['Amount']))\n\nsns.distplot(time_val, ax=ax[1], color='g')\nax[1].set_title('Distribution of Transaction Time', fontsize=14)\nax[1].set_xlim(min(time_val), max(time_val))\nplt.show()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"# Since most of our data has already been scaled we should scale the columns that are left to scale (Amount and Time)\nfrom sklearn.preprocessing import StandardScaler, RobustScaler\n# RobustScaler is less prone to outliers.\n\nstd_scaler = StandardScaler()\nprint(std_scaler)\nrob_scaler = RobustScaler()\nprint(rob_scaler)\ndata['scaled_amount'] = rob_scaler.fit_transform(data['Amount'].values.reshape(-1,1))\ndata['scaled_time'] = rob_scaler.fit_transform(data['Time'].values.reshape(-1,1))\n# print(data['scaled_amount'])\ndata.drop(['Time','Amount'], axis=1, inplace=True)","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"\n## Why do we create a sub-Sample?\nIn the beginning of this notebook we saw that the original dataframe was **heavily imbalanced**! Using the original dataframe  will cause the following issues:\n\n* **Overfitting:** Our classification models will assume that in most cases there are no frauds! What we want for our model is to be certain when a fraud occurs. \n* **Wrong Correlations:** Although we don't know what the \"V\" features stand for, **it will be useful to understand how each of this features influence the result** (Fraud or No Fraud) *by having an imbalance dataframe we are not able to see the true correlations between the class and features.* \n\n\nSummary: \nScaled amount  and scaled time  are the columns with scaled values. \n There are 492 casesof fraud in our dataset so we can randomly get 492 cases of non-fraud to create our new sub dataframe. \nWe concat the 492 cases of fraud and non fraud, creating a new sub-sample. \n\n## Scaler and standardization\nStandardization of a dataset is a common requirement for many machine learning estimators. Typically this is done by removing the mean and scaling to unit variance. However, outliers can often influence the sample mean / variance in a negative way. In such cases, the median and the interquartile range often give better results.[The link to scalers](https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.RobustScaler.html)"},{"metadata":{"trusted":true},"cell_type":"code","source":"scaled_amount=data['scaled_amount']\nscaled_time=data['scaled_time']\ndata.drop(['scaled_amount', 'scaled_time'], axis = 1, inplace = True)\ndata.insert(0, 'scaled_time', scaled_time)\ndata.insert(1, 'scaled_amount', scaled_amount)\ndata.head()","execution_count":null,"outputs":[]},{"metadata":{},"cell_type":"markdown","source":"## Splitting the Data (Original DataFrame)\n<a id=\"splitting\"></a>\nBefore proceeding with the <b> Random UnderSampling technique</b> we have to separate the orginal dataframe. <b> Why? for testing purposes, remember although we are splitting the data when implementing Random UnderSampling or OverSampling techniques, we want to test our models on the original testing set not on the testing set created by either of these techniques.</b> The main goal is to fit the model either with the dataframes that were undersample and oversample (in order for our models to detect the patterns), and test it on the original testing set.  \n\n#### Summary\nTo fit the model and recongnize the pattern, two techniques are provided whihc is Random UnderSampling and Random OverSampling. However, to test it, we use the original testing set. The steps of doing is as follows."},{"metadata":{"trusted":true},"cell_type":"code","source":"from sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import StratifiedShuffleSplit, StratifiedKFold\nprint('No Frauds', round(data['Class'].value_counts()[0]/len(data) * 100,2), '% of the dataset')\nprint('Frauds', round(data['Class'].value_counts()[1]/len(data) * 100,2), '% of the dataset')\n\nX = data.drop('Class', axis=1)\ny = data['Class']\n# X.head()\nprint(len(X), len(y), count_class[0], count_class[1])\n\nsss = StratifiedKFold(n_splits=5, random_state=None, shuffle=False)\n\nfor train_index, test_index in sss.split(X, y):\n    print(\"Train:\", train_index, len(train_index), \"Test:\", test_index, len(test_index))\n    original_Xtrain, original_Xtest = X.iloc[train_index], X.iloc[test_index]\n    original_ytrain, original_ytest = y.iloc[train_index], y.iloc[test_index]\n    print(\"The Frauds in y\", round(original_ytrain.value_counts()[1]/len(original_ytrain)*100, 2), \"%\")\n    \n# We already have X_train and y_train for undersample data thats why I am using original to distinguish and to not overwrite these variables.\n# original_Xtrain, original_Xtest, original_ytrain, original_ytest = train_test_split(X, y, test_size=0.2, random_state=42)\n\n# Check the Distribution of the labels\n\n\n# Turn into an array\noriginal_Xtrain = original_Xtrain.values\noriginal_Xtest = original_Xtest.values\noriginal_ytrain = original_ytrain.values\noriginal_ytest = original_ytest.values\n\n# See if both the train and test label distribution are similarly distributed\ntrain_unique_label, train_counts_label = np.unique(original_ytrain, return_counts=True)\ntest_unique_label, test_counts_label = np.unique(original_ytest, return_counts=True)\nprint(train_counts_label)\nprint('-' * 100)\n\nprint('Label Distributions: \\n')\nprint(train_counts_label/ len(original_ytrain))\nprint(test_counts_label/ len(original_ytest))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true},"cell_type":"code","source":"    ","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat":4,"nbformat_minor":1}