{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Introduction","metadata":{"execution":{"iopub.status.busy":"2022-02-11T15:39:48.159314Z","iopub.execute_input":"2022-02-11T15:39:48.159594Z","iopub.status.idle":"2022-02-11T15:39:48.164941Z","shell.execute_reply.started":"2022-02-11T15:39:48.159562Z","shell.execute_reply":"2022-02-11T15:39:48.164008Z"}}},{"cell_type":"markdown","source":"MNIST (\"Modified National Institute of Standards and Technology\") is one of the basic dataset of computer vision. Since 1999, this classic dataset of handwritten images has served as the basis for benchmarking classification algorithms. A new machine learning techniques emerge, MNIST remains a reliable resource for researchers and learners alike. In this code, I will try to correctly identify digits from a dataset of tens of thousands of handwritten images. The data is simple and identification of digits is really accurate and powerful with the power of CV.","metadata":{}},{"cell_type":"markdown","source":"**Let's import the necessary libraries**","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")\n\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.image as mpimg\nimport seaborn as sns\n%matplotlib inline\n\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.metrics import confusion_matrix\nimport itertools\n\nfrom keras.utils.np_utils import to_categorical # convert to one-hot-encoding\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten, Conv2D, MaxPool2D, BatchNormalization\nfrom tensorflow.keras.optimizers import RMSprop\nfrom keras.preprocessing.image import ImageDataGenerator\nfrom keras.callbacks import ReduceLROnPlateau\nfrom keras.datasets import mnist\nimport tensorflow as tf\n\nsns.set(style='white', context='notebook', palette='deep')","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:46.152203Z","iopub.execute_input":"2022-02-11T17:55:46.152516Z","iopub.status.idle":"2022-02-11T17:55:46.168554Z","shell.execute_reply.started":"2022-02-11T17:55:46.15247Z","shell.execute_reply":"2022-02-11T17:55:46.167542Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Data preparation","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Load data","metadata":{}},{"cell_type":"markdown","source":"**First we'll upload our data into Python**","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('../input/digit-recognizer/train.csv')\ntest = pd.read_csv('../input/digit-recognizer/test.csv')\nsub = pd.read_csv('../input/digit-recognizer/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:46.170062Z","iopub.execute_input":"2022-02-11T17:55:46.170878Z","iopub.status.idle":"2022-02-11T17:55:51.890417Z","shell.execute_reply.started":"2022-02-11T17:55:46.170829Z","shell.execute_reply":"2022-02-11T17:55:51.889524Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Let's check the shapes**","metadata":{}},{"cell_type":"code","source":"print(train.shape)\nprint(test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:51.891828Z","iopub.execute_input":"2022-02-11T17:55:51.892053Z","iopub.status.idle":"2022-02-11T17:55:51.897684Z","shell.execute_reply.started":"2022-02-11T17:55:51.892026Z","shell.execute_reply":"2022-02-11T17:55:51.896538Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We can easily see that our train dataset has more inputs than the test dataset. We should always divide our data into train and test while keeping train bigger than test, in this case, data is divided by 60 to 40.**","metadata":{}},{"cell_type":"markdown","source":"**We need to set data features and trget labels for our train data**","metadata":{}},{"cell_type":"code","source":"Y_train = train[\"label\"]\nX_train = train.drop(labels = [\"label\"], axis = 1) ","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:51.899089Z","iopub.execute_input":"2022-02-11T17:55:51.899322Z","iopub.status.idle":"2022-02-11T17:55:52.029305Z","shell.execute_reply.started":"2022-02-11T17:55:51.899295Z","shell.execute_reply":"2022-02-11T17:55:52.02843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Lets download data from mnist to ","metadata":{}},{"cell_type":"code","source":"(x_train1, y_train1), (x_test1, y_test1) = mnist.load_data()\n\ntrain1 = np.concatenate([x_train1, x_test1], axis=0)\ny_train1 = np.concatenate([y_train1, y_test1], axis=0)\n\nY_train1 = y_train1\nX_train1 = train1.reshape(-1, 28*28)","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:52.031748Z","iopub.execute_input":"2022-02-11T17:55:52.031985Z","iopub.status.idle":"2022-02-11T17:55:52.511627Z","shell.execute_reply.started":"2022-02-11T17:55:52.031957Z","shell.execute_reply":"2022-02-11T17:55:52.510683Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sns.countplot(Y_train);","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:52.512908Z","iopub.execute_input":"2022-02-11T17:55:52.513251Z","iopub.status.idle":"2022-02-11T17:55:52.772391Z","shell.execute_reply.started":"2022-02-11T17:55:52.513222Z","shell.execute_reply":"2022-02-11T17:55:52.771746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**As you can see the our labels pretty balanced.**","metadata":{}},{"cell_type":"markdown","source":"## 2.2 Normalization","metadata":{}},{"cell_type":"markdown","source":"**We will grayscale the data to make CNN faster because if we use the original format we have inputs from [1-255] but when we normalize we will have [0-1].**","metadata":{}},{"cell_type":"code","source":"X_train = X_train / 255.0\ntest = test / 255.0\n\nX_train1 = X_train1 / 255.0","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:52.773396Z","iopub.execute_input":"2022-02-11T17:55:52.773727Z","iopub.status.idle":"2022-02-11T17:55:53.042942Z","shell.execute_reply.started":"2022-02-11T17:55:52.773699Z","shell.execute_reply":"2022-02-11T17:55:53.042133Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We will merge all the data we have.**","metadata":{}},{"cell_type":"code","source":"X_train = np.concatenate((X_train.values, X_train1))\nY_train = np.concatenate((Y_train, Y_train1))","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:53.044405Z","iopub.execute_input":"2022-02-11T17:55:53.044942Z","iopub.status.idle":"2022-02-11T17:55:53.395501Z","shell.execute_reply.started":"2022-02-11T17:55:53.044898Z","shell.execute_reply":"2022-02-11T17:55:53.394848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 Reshape","metadata":{}},{"cell_type":"markdown","source":"**We will reshape image in 3D (height = 28px, width = 28px , canal = 1)**\n\n**canal is setted 1 for gray scale**","metadata":{}},{"cell_type":"code","source":"X_train = X_train.reshape(-1,28,28,1)\ntest = test.values.reshape(-1,28,28,1)","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:53.396795Z","iopub.execute_input":"2022-02-11T17:55:53.397258Z","iopub.status.idle":"2022-02-11T17:55:53.402735Z","shell.execute_reply.started":"2022-02-11T17:55:53.397218Z","shell.execute_reply":"2022-02-11T17:55:53.401693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## One-Hot Encoding","metadata":{"execution":{"iopub.status.busy":"2022-02-11T16:44:02.971922Z","iopub.execute_input":"2022-02-11T16:44:02.972807Z","iopub.status.idle":"2022-02-11T16:44:03.003109Z","shell.execute_reply.started":"2022-02-11T16:44:02.972676Z","shell.execute_reply":"2022-02-11T16:44:03.001856Z"}}},{"cell_type":"markdown","source":"**Our labels are from 0 to 9, 10 digits. We need to encode these lables to vectors like 2 is => [0,0,1,0,0,0,0,0,0,0].**","metadata":{}},{"cell_type":"code","source":"Y_train = to_categorical(Y_train, num_classes = 10)","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:53.404837Z","iopub.execute_input":"2022-02-11T17:55:53.405337Z","iopub.status.idle":"2022-02-11T17:55:53.418538Z","shell.execute_reply.started":"2022-02-11T17:55:53.405296Z","shell.execute_reply":"2022-02-11T17:55:53.417609Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.5 Split training and valdiation set","metadata":{}},{"cell_type":"markdown","source":"**We create validation for our training because validation help us identify if the model is working well or not.**","metadata":{}},{"cell_type":"code","source":"X_train, X_val, Y_train, Y_val = train_test_split(X_train, Y_train, test_size = 0.1, random_state=2)\nX_train.shape, X_val.shape, Y_train.shape, Y_val.shape","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:55:53.419905Z","iopub.execute_input":"2022-02-11T17:55:53.420972Z","iopub.status.idle":"2022-02-11T17:55:53.713439Z","shell.execute_reply.started":"2022-02-11T17:55:53.420924Z","shell.execute_reply":"2022-02-11T17:55:53.712535Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**We can choose different ratios for splitting the train set. In this case, a small portion of 8% choosen to be the validation set which the model is evaluated and 92% is used to actually train the model.**","metadata":{}},{"cell_type":"markdown","source":"**Lets visualize one of our datapoint**","metadata":{}},{"cell_type":"code","source":"g = plt.imshow(X_train[12][:,:,0])","metadata":{"execution":{"iopub.status.busy":"2022-02-11T17:56:23.660042Z","iopub.execute_input":"2022-02-11T17:56:23.66045Z","iopub.status.idle":"2022-02-11T17:56:23.856385Z","shell.execute_reply.started":"2022-02-11T17:56:23.660421Z","shell.execute_reply":"2022-02-11T17:56:23.855662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Convolutional Neural Network (CNN)\n## 3.1 Define the model","metadata":{}},{"cell_type":"markdown","source":"We can use the most known Keras Sequential, just need to add one layer at a time, starting from the input layer. It is like a set of learnable filters where after each filter our model will learn more.\n\nThe first is the convolutional (Conv2D) layer. I choose to set 64 filters for the first two conv2D layers and 64 filters for the two second layers and 64 filters for one third layers and 256 for the last one. Each filter transforms a part of the image (defined by the kernel size) using the kernel filter. The kernel filter matrix is applied on the whole image. Filters can be seen as a transformation of the image.\n\nThe CNN can isolate features that are useful everywhere from these transformed images (feature maps).\n\nThe second important layer in CNN is the pooling (MaxPool2D) layer. This layer simply acts as a downsampling filter. It looks at the 2 neighboring pixels and picks the maximal value. These are used to reduce computational cost, and to some extent also reduce overfitting. We have to choose the pooling size (i.e the area size pooled each time) more the pooling dimension is high, more the downsampling is important.\n\nCombining convolutional and pooling layers, CNN are able to combine local features and learn more global features of the image.\n\n'relu' is the rectifier (activation function max(0,x). The rectifier activation function is used to add non linearity to the network.\n\nThe Flatten layer is use to convert the final feature maps into a one single 1D vector. This flattening step is needed so that you can make use of fully connected layers after some convolutional/maxpool layers. It combines all the found local features of the previous convolutional layers.\n\nIn the end i used the features in two fully-connected (Dense) layers which is just artificial an neural networks (ANN) classifier. In the last layer(Dense(10,activation=\"softmax\")) the net outputs distribution of probability of each class.","metadata":{}}]}