{"cells":[{"metadata":{"_uuid":"2582ebec91bfbc35a4792cdc4bc17529d60a71ba"},"cell_type":"markdown","source":"**Confusion Matrix**\n\nConfusion matrix is an important tool to evaluate our classifier performance. It provides us a clear picture of the performance of our classifier. It creates a matrix where we can find the frequency of hits and misses of each labels.\nTo compute the confusion matrix, you need to have a set of predictions, so they can be compared to actual targets.\n\nConfusion matrix can be used to compute following parameters:\n* Precission\n* Recall\n* F1-Score\n\nNote: This kernel intend to help newbies to understand confusion metrics and make sense of Precision, Recall and F1 Score.\n\n**For demonstrating the confusion matrix we will make use of the Diabetes Classification dataset and will perform following steps.**\n* Split the data set into training and test sets.\n* Train the model using RandomForestClassifier.\n* Predict on the test set.\n* Generate confusion Matrix using sklearns confusion_matrix.\n*  Calculate Precision, Recall and F1-Score"},{"metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true},"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load in \n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix, classification_report\n\n# Input data files are available in the \"../input/\" directory.\n# For example, running this (by clicking run or pressing Shift+Enter) will list the files in the input directory\n\nimport os\nprint(os.listdir(\"../input\"))\n\n# Any results you write to the current directory are saved as output.","execution_count":null,"outputs":[]},{"metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","trusted":true},"cell_type":"code","source":"data = pd.read_csv(\"../input/train.csv\")","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"bbb3255f0b7645ccd3862cb59b32df05e617923b"},"cell_type":"code","source":"#Lets have a look into the matadata \ndata.info()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"2c9e25bf1bd5d775dab53a5125f9c26582141086"},"cell_type":"code","source":"data.describe()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ab6047a9b2fd098c04b9f3e655442008af76d9e7"},"cell_type":"code","source":"#Lets have a look into some sample data\ndata.head()","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"08783424d0766586c576d3c186506eef0dc65559"},"cell_type":"code","source":"#Provide features for X and label for y\nX = data.drop('diabetes',axis=1)\ny = data['diabetes']\n# Split the data into traing and test sets\nX_train, X_test, y_train, y_test = train_test_split(\n                                    X, y, random_state=42, test_size=.33)\n#Initialize Random Forest Classifier\nrfc = RandomForestClassifier()\n#Fit model on the training Data\nrfc.fit(X_train,y_train)\n#Make prediction\npredictions = rfc.predict(X_test)","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"879ecdb9351040149c925d0e90eb3e053711dd2b"},"cell_type":"code","source":"#Generate Confusion Matrix\nconf_matrix = confusion_matrix(predictions,y_test)\nprint(conf_matrix)","execution_count":null,"outputs":[]},{"metadata":{"_uuid":"bd71e2e62ca9c28d2920e251e79be23072a9c70b"},"cell_type":"markdown","source":"![](https://cdncontribute.geeksforgeeks.org/wp-content/uploads/Confusion_Matrix1_1.png)\n\nWe can visualize the confusion matrix that is generated for our model's prediction as below: \n\n           0       1\n    0     117     34\n    1     14      38\n    \n   Here in this matrix we have:\n   \n   For 0:\n   True Positive(TP)     = 117,\n   False Positive(FP)    = 14,\n   False Negative(FN)  = 34\n   \n   For 1:\n   True Positive(TP) = 38,\n   False Positive(FP) = 34,\n   False Negative(FN) = 14"},{"metadata":{"_uuid":"8ec9cc5692173fde70cd29da98d148464bcb5310"},"cell_type":"markdown","source":"Precision, Recall and F1-Score can be calculated as fellow:\n* Precision  = TP / (TP + FP)\n* Recall       = TP / (TP + FN)\n* F1-Score  = 2 ( Precision * Recall)/(Precision + Recall)\n"},{"metadata":{"trusted":true,"_uuid":"36a73776f14434e9cc07a2a659a598dc53b902e6"},"cell_type":"code","source":"#Lets calculate Precision, Recall and F1 score for label 0 and 1\n#For Label 0\ntp = conf_matrix[0,0]\nfp = conf_matrix[1,0]\nfn = conf_matrix[0,1]\n\nprecision  = tp / (tp + fp)\nrecall     = tp / (tp + fn)\nf1_score   = 2*( precision * recall)/(precision + recall)\n\nprint('precision, recall and f1-score for label 0')\nprint('The precision for label 0 is: {0:.2f}'.format(precision))\nprint('The recall for label 0 is: {0:.2f}'.format(recall))\nprint('The f1-score for label 0 is: {0:.2f}'.format(f1_score))\nprint('\\n')\n\n#For Label 1 \n\ntp = conf_matrix[1,1]\nfp = conf_matrix[0,1]\nfn = conf_matrix[1,0]\n\nprecision  = tp / (tp + fp)\nrecall     = tp / (tp + fn)\nf1_score   = 2*( precision * recall)/(precision + recall)\n\nprint('precision, recall and f1-score for label 1')\nprint('The precision for label 1 is: {0:.2f}'.format(precision))\nprint('The recall for label 1 is: {0:.2f}'.format(recall))\nprint('The f1-score for label 1 is: {0:.2f}'.format(f1_score))\n","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"1190bd17ca46d6e27ec5eadcc96c130d1f26f9fe"},"cell_type":"markdown","source":"We can Validate our calculation by comparing it with sklearn Metrics classification_report."},{"metadata":{"trusted":true,"_uuid":"11cb595af0899ca4518aa604d9ba6ae72aec910e"},"cell_type":"code","source":"print(classification_report(predictions,y_test))","execution_count":null,"outputs":[]},{"metadata":{"trusted":true,"_uuid":"ddc20b1b8e8c5b6d168e6f5f4f5075f9807af151"},"cell_type":"code","source":"","execution_count":null,"outputs":[]}],"metadata":{"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"}},"nbformat":4,"nbformat_minor":1}