{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# What is Logistic Regression\n\nLogistic regression is a Supervised Learning Classification Algorithm. It is used to predict a binary outcome based on a set of independent variables.\n\nA binary outcome is one where there are only two possible scenarios:\n* Positive: the event happens (1)\n* Negative: the event does not happen (0). \n\nIndependent variables are those variables or factors which may influence the outcome (or dependent variable).\n\n\n## Example: Predicting if a tumor is benign or malignant\n![](https://i.postimg.cc/YqYmq1S5/tumn.jpg)\n### In the above example you can see that the markers at the level of y = 1 (yes) demonstrate examples of a malignant tumor while the examples labeled at y = 0(no) demonstrate benign tumors. Logistic Regression aims to classify an input tumor marker from the user as either 1 (malignan) or 0 (benign)\n\n\n# What is Linear Regression?\n\nLinear Regression is a Supervised Learning Regression Algorithm that takes an input x and outputs a prediction y utilizing a trendline within a set of Data Points.\n\n* Supervised learning Algorithm learns through being given a dataset which represents accurate data or \"the right answers\".\n\n\n![](https://qph.fs.quoracdn.net/main-qimg-37d126042d314752b8642f742d0bc109)\n\n\n\n\n\n\n\n\n\n\n# The difference between Linear and Logistic Regression\n![](https://www.machinelearningplus.com/wp-content/uploads/2017/09/linear_vs_logistic_regression.jpg)\n\n\n\n# The Similarities with Linear Regression and Logistic Regression \n\nWhile Logistic Regression is a classification system, it shares many similarities with that of the Linear Regression Model\n\n* Line of Best Fit/Trendline/Regression Line/ Cost Function\n* Equation for Cost Function\n* Implimentation of Gradient Descent\n\n# Line of Best fit\n## To Understand the Logistic Regression, One must first understand Linear Regression and its parameters\n​\nWhen conducting our linear regression there is room for error. That error falls within our line of best fit.\n​\nThe line of best fit has a target x1 and y1. This means that for every x value there is a y value that corresponds, however where the value of x1 meets the trendline is where the prediction is executed.\n​\n![](https://i.postimg.cc/50XBNQQT/error.png)\n​\n​\n![](https://i.postimg.cc/hvS7QJXK/cost.png)\n​\n​\n​\n# Tuning Line of Best fit\n​\nThe importance of the Cost Function is to find the parameters for the equation of our line of best fit. With these new parameters we would reduce our total squared error (Cost) significantly and thus have more accurate predictions. \n​\n![](https://i.postimg.cc/zGt3XzZR/descent.png)\n​\n# Gradient Descent Algorithm\n![](https://i.postimg.cc/gjzPZh3c/algo.png)\n\n\n\n## The above Gradient Descent Algoritm is to be run simultaneously to ensure that the Cost Function is updating with new unique parameters. This is important because any deviation will be detrimental to the Cost Function. \n\n\n\n# Cost Function : Linear Regression vs Logistic Regression\n![](https://i.postimg.cc/wjgsG7kD/main-qimg-d551e49dd975703f260152e111883af9-lq.jpg)\n\n# The f(x) function is the major difference between calculating the cost function of the Linear Regression Model vs Logistic Regression Model\nIn the above function:\n  * Linear Regression Function : y\n  * Logistic Regression Function: p\n  \nThe Logistic Regression model calculates the sigmoid curve of the Linear Regression function\n\n\n\n\n\n​","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"# Credit Card Fault Prediction with Logistic Regression\n\nWe will be conducting a Logistic Regression model utilizing the sklearn library. We will also run the models of Support Vector Machine and MLPClassifier to compare. The primary focus will be Logistic Regression, with a manual example of gradient descent below that. \n\n# Import Libraries and Data\n\n\n","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.svm import SVC\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.model_selection import train_test_split\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom sklearn.metrics import roc_curve, auc\n","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:47:30.959023Z","iopub.execute_input":"2022-11-06T18:47:30.959801Z","iopub.status.idle":"2022-11-06T18:47:30.966663Z","shell.execute_reply.started":"2022-11-06T18:47:30.959756Z","shell.execute_reply":"2022-11-06T18:47:30.965554Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data = pd.read_csv('../input/default-of-credit-card-clients-dataset/UCI_Credit_Card.csv')\ndata","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:26.589422Z","iopub.execute_input":"2022-11-06T18:04:26.589757Z","iopub.status.idle":"2022-11-06T18:04:26.797994Z","shell.execute_reply.started":"2022-11-06T18:04:26.589727Z","shell.execute_reply":"2022-11-06T18:04:26.796938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.info()\n# There are no Object collumns or NA colums thus we can move on to Cleansing the Data\n#It also helps that these values are all numerical as we can issue a correlation heatmap","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:26.799341Z","iopub.execute_input":"2022-11-06T18:04:26.800128Z","iopub.status.idle":"2022-11-06T18:04:26.825570Z","shell.execute_reply.started":"2022-11-06T18:04:26.800082Z","shell.execute_reply":"2022-11-06T18:04:26.824480Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualization\nA large part of Regression Models has to do with correlations. Correlations help the algorithms to detect waht features influence each other or the desired outcomes. Heatmaps are a very impactful way to visualize these correlations","metadata":{}},{"cell_type":"code","source":"#Find the correlation of the dataset\n\ncorr= data.corr()\n\n#Plot heatmap\nplt.figure(figsize=(18,15))\nsns.heatmap(corr, annot=True, vmin=-1.0, cmap ='mako')\nplt.title('Correlation')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:26.827205Z","iopub.execute_input":"2022-11-06T18:04:26.827530Z","iopub.status.idle":"2022-11-06T18:04:29.806905Z","shell.execute_reply.started":"2022-11-06T18:04:26.827502Z","shell.execute_reply":"2022-11-06T18:04:29.805874Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Cleanse Data\n\nAn important factor to our regression model is the type of our features.\n\nA feature could be: \n* Nominal:\n\nHas no difference in  magnitude dictated by the columns, ex:sex, marriage. We dont want our algorithm to think that there is an ordered difference of sex or education. This means if 1 denotes male and 2 denotes female we dont want our model to think of this as an actual numerical difference. Thus we would choose to separate the features\n* Ordinal:\n\nUsed to demonstrate the order or rank of elements in a collection\n    \nTherefore, Nominal = identity, Ordinal =  magnitude\n\n# One hot encoding\nOne hot encoding is one method of converting data to prepare it for an algorithm and get a better prediction. With one-hot, we convert each categorical value into a new categorical column and assign a binary value of 1 or 0 to those columns. Each integer value is represented as a binary vector.\n    ","metadata":{}},{"cell_type":"code","source":"#Setup dummies to split nominal features into separate columns\ndef onehotencode(df, columndict):\n    df= df.copy()\n    \n    for column, prefix in columndict.items():\n        dummies= pd.get_dummies(df[column], prefix = prefix)\n        #concat to original dataframe \n        df = pd.concat([df, dummies], axis=1)\n        df = df.drop(column, axis=1)\n        \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:29.810456Z","iopub.execute_input":"2022-11-06T18:04:29.811317Z","iopub.status.idle":"2022-11-06T18:04:29.818555Z","shell.execute_reply.started":"2022-11-06T18:04:29.811270Z","shell.execute_reply":"2022-11-06T18:04:29.817369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Cleanse Data\n\ndef cleaner(df):\n    df = df.copy()\n    \n    #We do not need the ID column index\n    df=df.drop('ID', axis=1)\n    \n    df=onehotencode(\n        df,\n        {\n            'EDUCATION': 'EDU',\n            'MARRIAGE': 'MAR'\n        }\n    )\n    \n    #Turn Data Set into X and y sets\n    #Use the copy function so we can keep our original data frames at all times\n    y = df['default.payment.next.month'].copy()\n    X = df.drop('default.payment.next.month', axis=1).copy()\n    \n    \n    #Scale our Data \n    scaler = StandardScaler()\n    X = pd.DataFrame(scaler.fit_transform(X), columns = X.columns)\n    \n    \n    \n    return X, y","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:29.820131Z","iopub.execute_input":"2022-11-06T18:04:29.820485Z","iopub.status.idle":"2022-11-06T18:04:29.840090Z","shell.execute_reply.started":"2022-11-06T18:04:29.820455Z","shell.execute_reply":"2022-11-06T18:04:29.839122Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X,y =  cleaner(data)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:29.841662Z","iopub.execute_input":"2022-11-06T18:04:29.842362Z","iopub.status.idle":"2022-11-06T18:04:29.913675Z","shell.execute_reply.started":"2022-11-06T18:04:29.842320Z","shell.execute_reply":"2022-11-06T18:04:29.912532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:29.915544Z","iopub.execute_input":"2022-11-06T18:04:29.915885Z","iopub.status.idle":"2022-11-06T18:04:29.948373Z","shell.execute_reply.started":"2022-11-06T18:04:29.915855Z","shell.execute_reply":"2022-11-06T18:04:29.947084Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Train the Models\nYou want to train your models against others to see their performance and what models could be done better","metadata":{}},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, train_size=0.7, random_state=1234) #Set state to be apple to reproduce values automatically","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:29.952056Z","iopub.execute_input":"2022-11-06T18:04:29.952467Z","iopub.status.idle":"2022-11-06T18:04:29.968878Z","shell.execute_reply.started":"2022-11-06T18:04:29.952432Z","shell.execute_reply":"2022-11-06T18:04:29.967700Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"models = {\n    LogisticRegression(): \"Logistic Regression\",\n    SVC():                \"Support Vector Machine\",\n    MLPClassifier():      \"Neural Network\"\n}\n\nfor model in models.keys():\n    model.fit(X_train, y_train)","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:04:29.970813Z","iopub.execute_input":"2022-11-06T18:04:29.971287Z","iopub.status.idle":"2022-11-06T18:05:37.683327Z","shell.execute_reply.started":"2022-11-06T18:04:29.971243Z","shell.execute_reply":"2022-11-06T18:05:37.681705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for model, name in models.items():\n    print(name+ \": {:.2f}%\".format(model.score(X_test, y_test)*100))","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:49:40.613318Z","iopub.execute_input":"2022-11-06T18:49:40.613690Z","iopub.status.idle":"2022-11-06T18:49:46.405760Z","shell.execute_reply.started":"2022-11-06T18:49:40.613661Z","shell.execute_reply":"2022-11-06T18:49:46.404289Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As seen above Support Vector Machine was the most accurate in predicting the defaults on credit card clients. However as this is a Logistic Regression Model we will now go into the actual evaluation of this dataset.","metadata":{}},{"cell_type":"markdown","source":"# ROC and AUC\nThe Receiver Operating Characteristics(ROC) is a measure of a classification model's performance over various thresholds. This is accomplished using the parameters of a Confusion Matrix.\nThe Area Under the Curve is a metric to evaulate the ROC of various classification models.\n\n## ROC measures model's accuracy\n## AUC measures model's accuracy vs other models\n\n# Confusion Matrix\n![](https://ekababisong.org/assets/ieee_ompi/confusion_matrix.png)","metadata":{}},{"cell_type":"markdown","source":"# Compare SVM and Logistic regression for ROC and AUC","metadata":{}},{"cell_type":"code","source":"\nmodel_SVC = SVC(kernel = 'rbf', random_state = 1234)\nmodel_SVC.fit(X_train, y_train)\ny_pred_svm = model_SVC.decision_function(X_test)\n\n\nmodel_logistic = LogisticRegression()\nmodel_logistic.fit(X_train, y_train)\ny_pred_logistic = model_logistic.decision_function(X_test)\n\n\n\n","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:50:13.674958Z","iopub.execute_input":"2022-11-06T18:50:13.675372Z","iopub.status.idle":"2022-11-06T18:50:37.978814Z","shell.execute_reply.started":"2022-11-06T18:50:13.675337Z","shell.execute_reply":"2022-11-06T18:50:37.977248Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Plot ROC and compare AUC\nlogistic_fpr, logistic_tpr, threshold = roc_curve(y_test, y_pred_logistic)\nauc_logistic = auc(logistic_fpr, logistic_tpr)\n\nsvm_fpr, svm_tpr, threshold = roc_curve(y_test, y_pred_svm)\nauc_svm = auc(svm_fpr, svm_tpr)\n\nplt.figure(figsize=(5, 5), dpi=100)\nplt.plot(svm_fpr, svm_tpr, linestyle='-', label='SVM (auc = %0.3f)' % auc_svm)\nplt.plot(logistic_fpr, logistic_tpr, marker='.', label='Logistic (auc = %0.3f)' % auc_logistic)\n\nplt.xlabel('False Positive Rate -->')\nplt.ylabel('True Positive Rate -->')\n\nplt.legend()\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-11-06T18:50:38.248441Z","iopub.execute_input":"2022-11-06T18:50:38.249237Z","iopub.status.idle":"2022-11-06T18:50:38.487630Z","shell.execute_reply.started":"2022-11-06T18:50:38.249193Z","shell.execute_reply":"2022-11-06T18:50:38.486593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above we see that in this score the Logistic Regression has the larger AUC and could be seen as a stronger model regarding this metric.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\nThe Logistic Regression is an amazing tool that aids in the creation of classification. There are various means casses for utilization for the Logistic Regression model as there are various evaluation metrics to judge model performance.  I hope this tutorial helped.","metadata":{}}]}