{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# My 2nd Attempt at MNIST","metadata":{}},{"cell_type":"markdown","source":"### Jarred Priester\n### 7/8/22","metadata":{}},{"cell_type":"markdown","source":"1. Introduction\n2. Downloading the data\n3. Processing the images\n4. Model\n5. Predictions\n6. Conclusion","metadata":{}},{"cell_type":"markdown","source":"# 1. Introduction","metadata":{}},{"cell_type":"markdown","source":"In my first submission for the Kaggle Digit Recognizer competition I submitted a simple yet affective model with the accuracy score of 98%. In this notebook I will attempt to beat that score by adding more layers to the first model.\n\nThe first submission can be found here for comparison:\nhttps://www.kaggle.com/code/jarredpriester/my-first-cnn","metadata":{}},{"cell_type":"markdown","source":"# 2. Downloading the data","metadata":{}},{"cell_type":"code","source":"#importing libaries\nimport numpy as np \nimport pandas as pd\nimport random as rd\n\nimport matplotlib.pyplot as plt\n%matplotlib inline\n\nfrom PIL import Image\n\nfrom sklearn.model_selection import train_test_split\n\nimport tensorflow as tf\nfrom tensorflow import keras\nfrom tensorflow.keras import layers\nfrom tensorflow.keras.layers.experimental import preprocessing\nfrom keras.applications.vgg16 import VGG16\n\n#setting seed for reproducability\nfrom numpy.random import seed\nseed(10)\ntf.random.set_seed(20)\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n        ","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:34.008521Z","iopub.execute_input":"2022-07-09T02:36:34.009173Z","iopub.status.idle":"2022-07-09T02:36:41.006367Z","shell.execute_reply.started":"2022-07-09T02:36:34.009112Z","shell.execute_reply":"2022-07-09T02:36:41.005448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#downloading the sample submission\nsample_sub = pd.read_csv(\"../input/digit-recognizer/sample_submission.csv\")\nsample_sub.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:41.011455Z","iopub.execute_input":"2022-07-09T02:36:41.012077Z","iopub.status.idle":"2022-07-09T02:36:41.051707Z","shell.execute_reply.started":"2022-07-09T02:36:41.012045Z","shell.execute_reply":"2022-07-09T02:36:41.050883Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#downloading the training data\ntrain = pd.read_csv(\"../input/digit-recognizer/train.csv\")\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:41.053443Z","iopub.execute_input":"2022-07-09T02:36:41.053803Z","iopub.status.idle":"2022-07-09T02:36:44.124369Z","shell.execute_reply.started":"2022-07-09T02:36:41.053772Z","shell.execute_reply":"2022-07-09T02:36:44.123591Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#downloading the test data\ntest = pd.read_csv(\"../input/digit-recognizer/test.csv\")\ntest.head()","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:44.126183Z","iopub.execute_input":"2022-07-09T02:36:44.126738Z","iopub.status.idle":"2022-07-09T02:36:45.938583Z","shell.execute_reply.started":"2022-07-09T02:36:44.126703Z","shell.execute_reply":"2022-07-09T02:36:45.937952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. Processing the images","metadata":{}},{"cell_type":"markdown","source":"In the training data set we can see that the first column is 'label' and the rest are pixels. We will use the label column as our Y and the rest as our X","metadata":{}},{"cell_type":"code","source":"Y_train = train[\"label\"]\n\nX_train = train.drop(labels = [\"label\"],axis = 1) ","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:45.939531Z","iopub.execute_input":"2022-07-09T02:36:45.940450Z","iopub.status.idle":"2022-07-09T02:36:46.061015Z","shell.execute_reply.started":"2022-07-09T02:36:45.940406Z","shell.execute_reply":"2022-07-09T02:36:46.059938Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Next we need to scale the data from 0-255 to 0-1. This will make things easier to work with the neural network because it will allow the nn to converge faster.","metadata":{}},{"cell_type":"code","source":"X_train = X_train / 255.0\ntest = test / 255.0","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.062330Z","iopub.execute_input":"2022-07-09T02:36:46.062810Z","iopub.status.idle":"2022-07-09T02:36:46.171902Z","shell.execute_reply.started":"2022-07-09T02:36:46.062765Z","shell.execute_reply":"2022-07-09T02:36:46.171058Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The images will need to be reshaped in order feed into our model. Our images will be 28x28 and since we will be using grayscale the color channel will be 1.","metadata":{}},{"cell_type":"code","source":"X_train = X_train.values.reshape(-1,28,28,1)\ntest = test.values.reshape(-1,28,28,1)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.172716Z","iopub.execute_input":"2022-07-09T02:36:46.173046Z","iopub.status.idle":"2022-07-09T02:36:46.181406Z","shell.execute_reply.started":"2022-07-09T02:36:46.173020Z","shell.execute_reply":"2022-07-09T02:36:46.179642Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train.shape)\nprint(test.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.182959Z","iopub.execute_input":"2022-07-09T02:36:46.183575Z","iopub.status.idle":"2022-07-09T02:36:46.191842Z","shell.execute_reply.started":"2022-07-09T02:36:46.183501Z","shell.execute_reply":"2022-07-09T02:36:46.190610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now that the data has been scaled and reshaped we can now split our training data into training and validation sets. The split will be 90% to train on and 10% to use as validation. Which leaves us with the test data which will we be making our predictions on.","metadata":{}},{"cell_type":"code","source":"X_train, X_val, Y_train, Y_val = train_test_split(X_train, Y_train, test_size = 0.1, random_state=5)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.193650Z","iopub.execute_input":"2022-07-09T02:36:46.194367Z","iopub.status.idle":"2022-07-09T02:36:46.620951Z","shell.execute_reply.started":"2022-07-09T02:36:46.194311Z","shell.execute_reply":"2022-07-09T02:36:46.619929Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(X_train.shape)\nprint(Y_train.shape)\nprint(X_val.shape)\nprint(Y_val.shape)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.622183Z","iopub.execute_input":"2022-07-09T02:36:46.622921Z","iopub.status.idle":"2022-07-09T02:36:46.628602Z","shell.execute_reply.started":"2022-07-09T02:36:46.622876Z","shell.execute_reply":"2022-07-09T02:36:46.627574Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let us take a look at an image from the training set so we have a visual of the images we are working with","metadata":{}},{"cell_type":"code","source":"plt.imshow(X_train[18])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.630344Z","iopub.execute_input":"2022-07-09T02:36:46.631131Z","iopub.status.idle":"2022-07-09T02:36:46.841709Z","shell.execute_reply.started":"2022-07-09T02:36:46.631087Z","shell.execute_reply":"2022-07-09T02:36:46.840793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Model","metadata":{}},{"cell_type":"markdown","source":"The orginal plan for this model was to use a pretrained model. That plan did not work, I was unable to find a pretrained model that fit these images. If you know of any please feel free to let me know in the comments. The second plan was to use data augmentation but to my surprise the training results were worst than the orginal model. The third plan was to add additonal layers to the orginal model. Adding batch normalization and dropout layers to the orginal model did improve the validaton loss and validation accuracy from the orginal model.","metadata":{}},{"cell_type":"code","source":"model = keras.Sequential([\n    \n    layers.BatchNormalization(),\n    layers.Conv2D(filters=32, kernel_size=(5,5), activation=\"relu\", padding='same',\n                  input_shape=[28, 28, 1]),\n    layers.MaxPool2D(),\n    layers.Dropout(.1),\n    \n    layers.BatchNormalization(),\n    layers.Conv2D(filters=64, kernel_size=(3,3), activation=\"relu\", padding='same'),\n    layers.MaxPool2D(),\n    layers.Dropout(.1),\n    \n    layers.BatchNormalization(),\n    layers.Conv2D(filters=128, kernel_size=(3,3), activation=\"relu\", padding='same'),\n    layers.MaxPool2D(),\n    layers.Dropout(.1),\n    \n    layers.BatchNormalization(),\n    layers.Conv2D(filters=128, kernel_size=(3,3), activation=\"relu\", padding='same'),\n    layers.MaxPool2D(),\n    layers.Dropout(.1),\n    \n    layers.Flatten(),\n    layers.Dense(units=256, activation=\"relu\"),\n    layers.Dense(units=10, activation=\"softmax\"),\n])","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.843232Z","iopub.execute_input":"2022-07-09T02:36:46.843608Z","iopub.status.idle":"2022-07-09T02:36:46.944992Z","shell.execute_reply.started":"2022-07-09T02:36:46.843544Z","shell.execute_reply":"2022-07-09T02:36:46.943927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(\n    optimizer=tf.keras.optimizers.Adam(epsilon=0.01),\n    loss='sparse_categorical_crossentropy',\n    metrics=['accuracy']\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.948364Z","iopub.execute_input":"2022-07-09T02:36:46.948826Z","iopub.status.idle":"2022-07-09T02:36:46.973887Z","shell.execute_reply.started":"2022-07-09T02:36:46.948791Z","shell.execute_reply":"2022-07-09T02:36:46.972501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(\n    x = X_train,\n    y = Y_train, \n    validation_data= (X_val,Y_val),\n    batch_size = 128,\n    epochs=20,\n    verbose=2,\n)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:36:46.977787Z","iopub.execute_input":"2022-07-09T02:36:46.978931Z","iopub.status.idle":"2022-07-09T02:47:15.919760Z","shell.execute_reply.started":"2022-07-09T02:36:46.978879Z","shell.execute_reply":"2022-07-09T02:47:15.917956Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history_frame = pd.DataFrame(history.history)\nhistory_frame.loc[:, ['loss', 'val_loss']].plot()\nhistory_frame.loc[:, ['accuracy', 'val_accuracy']].plot();","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:47:15.922532Z","iopub.execute_input":"2022-07-09T02:47:15.923002Z","iopub.status.idle":"2022-07-09T02:47:16.331454Z","shell.execute_reply.started":"2022-07-09T02:47:15.922960Z","shell.execute_reply":"2022-07-09T02:47:16.330398Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5. Predictions","metadata":{}},{"cell_type":"code","source":"predictions = model.predict(test)\n\npredictions = np.argmax(predictions,axis = 1)\n\nsubmissions=pd.DataFrame({\"ImageId\": list(range(1,len(predictions)+1)),\n                         \"Label\": predictions})\nsubmissions.to_csv(\"submissions.csv\", index=False, header=True)\n\nsubmissions.head(20)","metadata":{"execution":{"iopub.status.busy":"2022-07-09T02:47:16.333244Z","iopub.execute_input":"2022-07-09T02:47:16.333730Z","iopub.status.idle":"2022-07-09T02:47:23.879377Z","shell.execute_reply.started":"2022-07-09T02:47:16.333687Z","shell.execute_reply":"2022-07-09T02:47:23.878342Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Conclusion","metadata":{}},{"cell_type":"markdown","source":"We started with the task of improving on the result from our first model, a result of 98% accuracy. By applying batch normalization and dropout layers to the orginal model we were able to improve the validation loss to .029 and the validation accuracy to 99%, achieving our objective of improving our first model.","metadata":{}}]}