{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"## Enhancing Deep Learning Model Performance\n\nIn this notebook, we will be generating samples from existing data, monitor learning rate and decay it over time. The dataset we will be working with is Fashion MNIST. Fashion-MNIST is a dataset consisting of a training set of 60,000 examples and a test set of 10,000 examples. Each example is a 28x28 grayscale image, associated with a label from 10 classes.\n\n<center>\n<img src=\"https://machinelearningmastery.com/wp-content/uploads/2019/02/Plot-of-a-Subset-of-Images-from-the-Fashion-MNIST-Dataset-1024x768.png\", width=600>\n</center>\n\n\n**Content**\n\n- Each image is 28 pixels in height and 28 pixels in width, for a total of 784 pixels in total. \n- Each pixel has a single pixel-value associated with it, indicating the lightness or darkness of that pixel, with higher numbers meaning darker. \n- This pixel-value is an integer between 0 and 255. The training and test data sets have 785 columns. \n- The first column consists of the class labels, and represents the article of clothing. The rest of the columns contain the pixel-values of the associated image.\n- To locate a pixel on the image, suppose that we have decomposed x as x = i * 28 + j, where i and j are integers between 0 and 27. The pixel is located on row i and column j of a 28 x 28 matrix.\n\n*For example*, pixel31 indicates the pixel that is in the fourth column from the left, and the second row from the top, as in the ascii-diagram below.\n\n**Labels**\n\nEach training and test example is assigned to one of the following labels:\n\n- 0 T-shirt/top\n- 1 Trouser\n- 2 Pullover\n- 3 Dress\n- 4 Coat\n- 5 Sandal\n- 6 Shirt\n- 7 Sneaker\n- 8 Bag\n- 9 Ankle boot\n\n\n**TL;DR**\n\n- Each row is a separate image\n- Column 1 is the class label.\n- Remaining columns are pixel numbers (784 total).\n- Each value is the darkness of the pixel (1 to 255)","metadata":{}},{"cell_type":"code","source":"product = ['T-shirt/top', 'Trouser', 'Pullover', 'Dress', 'Coat', 'Sandal', 'Shirt', 'Sneaker', 'Bag', 'Ankle boot']\nlabels = [0,1,2,3,4,5,6,7,8,9]","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:52:45.252690Z","iopub.execute_input":"2023-10-07T03:52:45.253135Z","iopub.status.idle":"2023-10-07T03:52:45.279221Z","shell.execute_reply.started":"2023-10-07T03:52:45.253032Z","shell.execute_reply":"2023-10-07T03:52:45.278158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Imports","metadata":{}},{"cell_type":"code","source":"import warnings\nwarnings.filterwarnings(\"ignore\")\nimport seaborn as sns\nimport numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\n%matplotlib inline\nplt.rcParams['font.size'] = 12\nplt.rcParams['figure.figsize'] = (22, 5)\nplt.rcParams['figure.dpi'] = 100\ncustom_params = {\"axes.spines.right\": False, \"axes.spines.top\": False}\nsns.set_theme(style=\"ticks\", context='notebook', rc=custom_params)\n\nimport tensorflow\nfrom tensorflow import keras\nfrom tensorflow.keras import layers, Sequential\nfrom tensorflow.keras.layers import Dense, Dropout, Flatten, MaxPool2D, Conv2D, BatchNormalization\nfrom tensorflow.keras.callbacks import EarlyStopping, LearningRateScheduler\nfrom sklearn.metrics import confusion_matrix, classification_report","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2023-10-07T05:15:34.621505Z","iopub.execute_input":"2023-10-07T05:15:34.621894Z","iopub.status.idle":"2023-10-07T05:15:34.633635Z","shell.execute_reply.started":"2023-10-07T05:15:34.621859Z","shell.execute_reply":"2023-10-07T05:15:34.632649Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Reading the data","metadata":{}},{"cell_type":"code","source":"train =  pd.read_csv('../input/fashionmnist/fashion-mnist_train.csv')\ntest= pd.read_csv('../input/fashionmnist/fashion-mnist_test.csv')","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:52:51.757881Z","iopub.execute_input":"2023-10-07T03:52:51.758422Z","iopub.status.idle":"2023-10-07T03:52:57.357350Z","shell.execute_reply.started":"2023-10-07T03:52:51.758394Z","shell.execute_reply":"2023-10-07T03:52:57.356361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:53:11.522959Z","iopub.execute_input":"2023-10-07T03:53:11.523689Z","iopub.status.idle":"2023-10-07T03:53:11.547479Z","shell.execute_reply.started":"2023-10-07T03:53:11.523651Z","shell.execute_reply":"2023-10-07T03:53:11.546441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test.head()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:53:12.414304Z","iopub.execute_input":"2023-10-07T03:53:12.415047Z","iopub.status.idle":"2023-10-07T03:53:12.432091Z","shell.execute_reply.started":"2023-10-07T03:53:12.415005Z","shell.execute_reply":"2023-10-07T03:53:12.431174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(f'There are {train.shape[0]} samples in the training data and {train.shape[1]} columns.')\nprint(f'There are {test.shape[0]} samples in the test data and {test.shape[1]} columns.')","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:53:15.918728Z","iopub.execute_input":"2023-10-07T03:53:15.919066Z","iopub.status.idle":"2023-10-07T03:53:15.925145Z","shell.execute_reply.started":"2023-10-07T03:53:15.919038Z","shell.execute_reply":"2023-10-07T03:53:15.924037Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Exploratory Data Analysis\n\nLet's look at the images that we are dealing with and the target classes countplot","metadata":{}},{"cell_type":"code","source":"ax = sns.countplot(train.label)\nplt.xticks(ticks=labels, labels = product)\nplt.title('Distribution of Target Classes, Train', fontsize=18, y=1.02)\n# Loops through each bar in the plot and adds an annotation with the height of the bar\nfor bar in ax.patches:\n    ax.annotate(format(bar.get_height(), '.0f'),(bar.get_x() + bar.get_width() / 2,bar.get_height()), \n                 ha='center', va='center',size=15, xytext=(0, 8),textcoords='offset points')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:15:38.077433Z","iopub.execute_input":"2023-10-07T05:15:38.077787Z","iopub.status.idle":"2023-10-07T05:15:38.349068Z","shell.execute_reply.started":"2023-10-07T05:15:38.077757Z","shell.execute_reply":"2023-10-07T05:15:38.348179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.countplot(test.label)\nplt.xticks(ticks=labels, labels = product)\nfor bar in ax.patches:\n    ax.annotate(format(bar.get_height(), '.0f'),(bar.get_x() + bar.get_width() / 2,bar.get_height()), \n                 ha='center', va='center',size=15, xytext=(0, 8),textcoords='offset points')\nplt.title('Distribution of Target Classes, Test', fontsize=18, y=1.02)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:15:44.805192Z","iopub.execute_input":"2023-10-07T05:15:44.805559Z","iopub.status.idle":"2023-10-07T05:15:45.066550Z","shell.execute_reply.started":"2023-10-07T05:15:44.805530Z","shell.execute_reply":"2023-10-07T05:15:45.065607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-info\">\n<b>INFO:</b> Notice the classes have the equal number of instances for each class which means we don't need to think about class imbalance here. It's always a good idea to look at the countplot of the classes before going ahead to train the model.\n</div>","metadata":{}},{"cell_type":"markdown","source":"## Seperating and Scaling\n\nLet's seperate our images from the labels and scale the inputs so that our model converges easily.","metadata":{}},{"cell_type":"code","source":"y_train = train['label'].values\nX_train = train[train.columns[1:]].values/255\ny_test = test['label'].values\nX_test = test[test.columns[1:]].values/255\nprint(X_train.shape[0], \"train samples\")\nprint(X_test.shape[0], \"test samples\")","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:58:22.957212Z","iopub.execute_input":"2023-10-07T03:58:22.957569Z","iopub.status.idle":"2023-10-07T03:58:23.315373Z","shell.execute_reply.started":"2023-10-07T03:58:22.957539Z","shell.execute_reply":"2023-10-07T03:58:23.313841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Convolutional Neural Network** has become very common in the field of computer vision in recent years. But it comes with a severe restriction regarding the size of the input image. Most convolutional neural networks are designed in a way so that they can only accept images of a fixed size. \n\nThis creates several challenges during data acquisition and model deployment. The common practice to overcome this limitation is to reshape the input images so that they can be fed into the networks.","metadata":{}},{"cell_type":"code","source":"X_train = X_train.reshape(60000,28,28,1)\nX_test = X_test.reshape(10000,28,28,1)\nprint(f'X_train shape {X_train.shape}')\nprint(f'X_test shape {X_test.shape}')\nprint(f'There are {len(y_train)} image labels in training dataset')\nprint(f'There are {len(y_test)} image labels in test dataset')","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:59:19.120812Z","iopub.execute_input":"2023-10-07T03:59:19.121183Z","iopub.status.idle":"2023-10-07T03:59:19.127766Z","shell.execute_reply.started":"2023-10-07T03:59:19.121150Z","shell.execute_reply":"2023-10-07T03:59:19.126617Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_, axes = plt.subplots(nrows=1, ncols=4, figsize=(20,15))\nfor ax, image, label in zip(axes, X_train, y_train):\n    ax.set_axis_off()\n    image = image.reshape(28, 28)\n    ax.imshow(image, cmap=plt.cm.gray_r, interpolation=\"nearest\")\n    ax.set_title(f\"Training image Label: {label}\")    ","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:59:28.278381Z","iopub.execute_input":"2023-10-07T03:59:28.278735Z","iopub.status.idle":"2023-10-07T03:59:28.581345Z","shell.execute_reply.started":"2023-10-07T03:59:28.278706Z","shell.execute_reply":"2023-10-07T03:59:28.580372Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_, axes = plt.subplots(nrows=1, ncols=4, figsize=(20,15))\nfor ax, image,label in zip(axes, X_test, y_test):\n    ax.set_axis_off()\n    image = image.reshape(28, 28)\n    ax.imshow(image, cmap=plt.cm.gray_r, interpolation=\"nearest\")\n    ax.set_title(f\"Test image Label: {label}\")    ","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:59:32.132596Z","iopub.execute_input":"2023-10-07T03:59:32.132944Z","iopub.status.idle":"2023-10-07T03:59:32.417582Z","shell.execute_reply.started":"2023-10-07T03:59:32.132913Z","shell.execute_reply":"2023-10-07T03:59:32.416639Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('There are {} unique image labels in the training dataset'.format(len(train['label'].unique())))\nprint('{} are unique image labels in the training dataset'.format(np.unique(train['label'])))","metadata":{"execution":{"iopub.status.busy":"2023-10-07T03:59:33.592965Z","iopub.execute_input":"2023-10-07T03:59:33.594048Z","iopub.status.idle":"2023-10-07T03:59:33.603930Z","shell.execute_reply.started":"2023-10-07T03:59:33.594003Z","shell.execute_reply":"2023-10-07T03:59:33.602679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data augmentation\n\n**Data Augmentation** is a technique that can be used to artificially expand the size of a training set by creating modified data from the existing one. It is a good practice to use DA if you want to prevent overfitting, or the initial dataset is too small to train on, or even if you want to squeeze better performance from your model. In general, DA is frequently used when building a DL model. Although the augmentation is done in Deep learning, you can also augment the data for the ML problems as well.\n\nYou can augment:\n\n1. *Audio*\n2. *Text*\n3. *Images*\n4. *Any other types of data*\n\nWe will focus on image augmentations as those are the most popular ones. Nevertheless, augmenting other types of data is as efficient and easy. That is why it’s good to remember some common techniques which can be performed to augment the data.\n\n### Data Augmentation techniques\n\nWe can apply various changes to the initial data. \n\n**For example**, for images we can use:\n\n1. Geometric transformations – you can randomly flip, crop, rotate or translate images, and that is just the tip of the iceberg\n2. Color space transformations – change RGB color channels, intensify any color\n3. Kernel filters – sharpen or blur an image \n4. Random Erasing – delete a part of the initial image\n5. Mixing images – basically, mix images with one another. Might be counterintuitive but it works\n\nMoreover, the greatest advantage of the augmentation techniques is that you may use all of them at once. Thus, you may get plenty of unique samples of data from the initial one. Below are some of the ways that you can augment data. The layers apply random augmentation transforms to a batch of images. They are only active during training.\n\n- tf.keras.layers.RandomCrop\n- tf.keras.layers.RandomFlip\n- tf.keras.layers.RandomTranslation\n- tf.keras.layers.RandomRotation\n- tf.keras.layers.RandomZoom\n- tf.keras.layers.RandomHeight\n- tf.keras.layers.RandomWidth\n- tf.keras.layers.RandomContrast","metadata":{}},{"cell_type":"markdown","source":"We are using the `RandomFlip` because our data is usually grayscale and is not rotated or has any other contrast. However we do have specified the both Horizontal and Vertical Flip","metadata":{}},{"cell_type":"code","source":"data_augmentation = keras.Sequential(\n  [\n    layers.experimental.preprocessing.RandomFlip(\"horizontal_and_vertical\", \n                                                 input_shape=(28, \n                                                              28,\n                                                              1))\n  ]\n)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:00:10.229222Z","iopub.execute_input":"2023-10-07T04:00:10.229588Z","iopub.status.idle":"2023-10-07T04:00:12.918542Z","shell.execute_reply.started":"2023-10-07T04:00:10.229560Z","shell.execute_reply":"2023-10-07T04:00:12.917283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below is the original Image from the training dataset:-","metadata":{}},{"cell_type":"code","source":"plt.imshow(X_train[4]);","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:00:20.946901Z","iopub.execute_input":"2023-10-07T04:00:20.947615Z","iopub.status.idle":"2023-10-07T04:00:21.141791Z","shell.execute_reply.started":"2023-10-07T04:00:20.947577Z","shell.execute_reply":"2023-10-07T04:00:21.140848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here is an augmented image:-","metadata":{}},{"cell_type":"code","source":"plt.imshow(data_augmentation(X_train)[4]);","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:00:24.948028Z","iopub.execute_input":"2023-10-07T04:00:24.949090Z","iopub.status.idle":"2023-10-07T04:00:25.808346Z","shell.execute_reply.started":"2023-10-07T04:00:24.949043Z","shell.execute_reply":"2023-10-07T04:00:25.807422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Model Creation and Training\n\nThere are a few steps involved in training a deep learning model that we will be going through\n\n1. Choose a model architecture.\n2. Choose a set of training data.\n3. Train the model.\n4. Evaluate the model.\n\nWith Keras and Tensorflow, it's relatively straightforward and you essentially don't end up thinking about creating classes, your forward passes etc.","metadata":{}},{"cell_type":"code","source":"model = Sequential()\n\n#Augmentation Layer\nmodel.add(data_augmentation)\n\n#Input Layer\nmodel.add(Conv2D(filters=64, kernel_size=(3, 3), activation='relu', padding = 'Same', input_shape=(28,28,1))),\nmodel.add(MaxPool2D(pool_size=(2, 2), strides=2)),\nmodel.add(BatchNormalization())\n\n#Hidden Layer\nmodel.add(Conv2D(filters = 64, kernel_size = (3,3),padding = 'Same', activation ='relu'))\nmodel.add(MaxPool2D(pool_size=(2,2), strides=(2,2)))\nmodel.add(BatchNormalization())\n\n#Dropout Layer\nmodel.add(Dropout(0.5))\n\n#Hidden Layer\nmodel.add(Conv2D(filters = 64, kernel_size = (3,3),padding = 'Same', activation ='relu'))\nmodel.add(BatchNormalization())\n\n#Hidden Layer\nmodel.add(Conv2D(filters = 64, kernel_size = (3,3),padding = 'Same', activation ='relu'))\nmodel.add(BatchNormalization())\n\n#Hidden Layer\nmodel.add(Conv2D(filters = 64, kernel_size = (3,3),padding = 'Same', activation ='relu'))\nmodel.add(MaxPool2D(pool_size=(2,2), strides=(2,2)))\nmodel.add(BatchNormalization())\n\n#Fully Connected Layer\nmodel.add(Flatten())\nmodel.add(BatchNormalization())\n\n#Fully Connected Layer\nmodel.add(Dense(30, activation = \"softmax\"))\nmodel.add(BatchNormalization())\n\nmodel.add(Dense(10, activation = \"softmax\"))\n\n#Model Summary\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:00:50.819774Z","iopub.execute_input":"2023-10-07T04:00:50.820165Z","iopub.status.idle":"2023-10-07T04:00:51.009112Z","shell.execute_reply.started":"2023-10-07T04:00:50.820123Z","shell.execute_reply":"2023-10-07T04:00:51.008153Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Model architecture diagram","metadata":{}},{"cell_type":"code","source":"tensorflow.keras.utils.plot_model(\n    model,\n    show_shapes=True,\n    show_dtype=False,\n    show_layer_names=True,\n    rankdir='TB',\n    expand_nested=False,\n    dpi=96,\n    layer_range=None\n)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:00:56.294461Z","iopub.execute_input":"2023-10-07T04:00:56.294806Z","iopub.status.idle":"2023-10-07T04:00:57.339160Z","shell.execute_reply.started":"2023-10-07T04:00:56.294777Z","shell.execute_reply":"2023-10-07T04:00:57.338004Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compiling the model and Learning Rate Decay\n\n**Compiling Model**\n\nCompiling a model in TensorFlow is a crucial step in the process of preparing a neural network for training. During compilation, you specify various settings that guide how the model should be trained. This includes choosing an optimizer (such as Adam or SGD), a loss function (e.g., mean squared error for regression or categorical cross-entropy for classification), and optional metrics for evaluation (like accuracy or F1 score). Additionally, you can define other configurations like learning rate and regularization parameters. Once compiled, the model is ready for training using the specified settings. \n\n**Learning Rate Decay**\n\nLearning rate decay is a technique commonly used in machine learning and deep learning to optimize the training of neural networks. It involves gradually reducing the learning rate, which determines the step size during the model's weight updates. The idea behind learning rate decay is to start with a relatively high learning rate to make quick progress in the initial training stages, allowing the model to escape local minima. \n\nAs training progresses and the model approaches convergence, the learning rate is lowered to fine-tune the model's parameters more precisely and prevent overshooting the optimal solution. This dynamic adjustment of the learning rate helps improve training stability and convergence, making it a valuable tool in the process of training complex machine learning models.","metadata":{}},{"cell_type":"code","source":"def get_lr_metric(optimizer):\n    '''\n    This function defines a learning rate metric to be used in a model's compile step\n    '''\n    # This is an inner function that takes y_true and y_pred as arguments\n    def lr(y_true, y_pred):\n        \n        # This line returns the learning rate from the optimizer object\n        return optimizer.lr\n    # This line returns the lr function as the metric\n    return lr","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:01:06.906336Z","iopub.execute_input":"2023-10-07T04:01:06.907174Z","iopub.status.idle":"2023-10-07T04:01:06.913423Z","shell.execute_reply.started":"2023-10-07T04:01:06.907127Z","shell.execute_reply":"2023-10-07T04:01:06.911918Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Setting the optimizations for the training","metadata":{}},{"cell_type":"code","source":"#setting the Adam optimizer with a learning rate of 0.1\noptimizer = keras.optimizers.Adam(learning_rate=0.1)\n\n#Optimising the model to stop if val_loss is not changed consecutively for 5 epochs\ncallback = EarlyStopping(\n    monitor=\"val_loss\",\n    min_delta=0.001,\n    patience=5,\n    verbose=1,\n    mode=\"auto\",\n    baseline=None,\n    restore_best_weights=False\n)\n\n#Heplful for tracking learning rate\nlr_metric = get_lr_metric(optimizer)\n\n\n#Decaying the learning rate 10% after 2 epochs\nLearningRate = [LearningRateScheduler(lambda epochs: 0.01 * 0.1 ** (epochs // 2))]","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:01:15.768504Z","iopub.execute_input":"2023-10-07T04:01:15.768866Z","iopub.status.idle":"2023-10-07T04:01:15.787768Z","shell.execute_reply.started":"2023-10-07T04:01:15.768834Z","shell.execute_reply":"2023-10-07T04:01:15.786848Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the LearningRateScheduler, we used something known as Step decay schedule which drops the learning rate by a factor every few epochs. We start with an initial LR of 0.01 and decrease it by a factor of 10% every two epochs during training.\n\nA typical way is to to drop the learning rate by half every 10 epochs. Below is the mathematical formula for our Step Decay:- \n\n$$lr = {lr_0} * drop^{\\frac{epoch} {epochs\\space{drop}}} $$\n","metadata":{}},{"cell_type":"code","source":"model.compile(loss='sparse_categorical_crossentropy',optimizer=optimizer,metrics=['accuracy', lr_metric], \n              steps_per_execution=50)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the `model.compile` function, the `steps_per_execution` parameter specifies how many batches of data should be processed before calculating and updating the metrics and losses during training. It affects the granularity at which training metrics are reported and updated.","metadata":{}},{"cell_type":"markdown","source":"## Training the CNN Model","metadata":{}},{"cell_type":"markdown","source":"We are using a batch size of 500 samples for both training and validation however for callbacks, Early Stopping and Learning Rate Scheduler is aligned.","metadata":{}},{"cell_type":"code","source":"history = model.fit(x=X_train,y=y_train,epochs=100,batch_size=500,validation_data = (X_test,y_test), \n                    validation_batch_size=500, callbacks=[callback, LearningRate])","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:01:21.824322Z","iopub.execute_input":"2023-10-07T04:01:21.824686Z","iopub.status.idle":"2023-10-07T04:02:02.258531Z","shell.execute_reply.started":"2023-10-07T04:01:21.824655Z","shell.execute_reply":"2023-10-07T04:02:02.257419Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Predictions\n\nNow that the model is trained, its time to make predictions on the training and test data both and judge the performance of the model.","metadata":{}},{"cell_type":"code","source":"#Making Prediction\ny_prob = model.predict(X_test)\ny_pred = y_prob.argmax(axis=1)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:09:01.997304Z","iopub.execute_input":"2023-10-07T04:09:01.997699Z","iopub.status.idle":"2023-10-07T04:09:02.951771Z","shell.execute_reply.started":"2023-10-07T04:09:01.997666Z","shell.execute_reply":"2023-10-07T04:09:02.950566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Evaluate Model\n\nThere are several essential metrics for evaluating the performance of a deep learning model:\n\n1. **Accuracy**: This metric assesses how well the model correctly predicts the class labels for a given dataset. It represents the overall correctness of predictions.\n\n2. **Recall**: Recall, also known as sensitivity or true positive rate, measures how effectively the model identifies the relevant instances within a class. It focuses on minimizing false negatives, ensuring that the model captures as many true positives as possible.\n\n3. **Precision**: Precision quantifies how often the model's positive predictions are correct. It is crucial for scenarios where false positives are costly, as it aims to minimize them.\n\n4. **F1 Score**: The F1 score is a balance between precision and recall. It provides a single metric that considers both false positives and false negatives, making it a useful measure for models that need to find a trade-off between precision and recall.\n\n5. **Confusion Matrix**: A confusion matrix provides a detailed breakdown of a model's performance. It not only highlights accuracy but also shows how well the model discriminates between true positives and true negatives. The diagonal elements indicate correct predictions, while off-diagonal elements reveal prediction errors.\n\nThese metrics offer a comprehensive view of a deep learning model's performance, addressing different aspects like overall correctness, ability to capture relevant instances, precision, and the trade-off between false positives and false negatives. The confusion matrix provides a detailed breakdown, making it especially valuable for in-depth analysis.","metadata":{}},{"cell_type":"code","source":"def plot_confusion_matrix(cm, classes,\n                          normalize=True,\n                          title='Confusion matrix',\n                          cmap=plt.cm.Blues):\n    \"\"\"\n    This function prints and plots the confusion matrix.\n    Normalization can be applied by setting `normalize=True`.\n    \"\"\"\n    import itertools\n    plt.figure(figsize=(10,10))\n    plt.imshow(cm, interpolation='nearest', cmap=cmap)\n    plt.title(title)\n    plt.colorbar()\n    tick_marks = np.arange(len(classes))\n    plt.xticks(tick_marks, product, rotation=45)\n    plt.yticks(tick_marks, product)\n\n    if normalize:\n        cm = cm.astype('float') / cm.sum(axis=1)[:, np.newaxis]\n        print(\"Normalized confusion matrix\")\n    else:\n        print('Confusion matrix, without normalization')\n\n    thresh = cm.max() / 2.\n    for i, j in itertools.product(range(cm.shape[0]), range(cm.shape[1])):\n        plt.text(j, i, cm[i, j],horizontalalignment=\"center\",color=\"white\" if cm[i, j] > thresh else \"black\")\n    plt.tight_layout()\n    plt.ylabel('True label')\n    plt.xlabel('Predicted label')\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:09:12.674087Z","iopub.execute_input":"2023-10-07T04:09:12.674692Z","iopub.status.idle":"2023-10-07T04:09:12.683486Z","shell.execute_reply.started":"2023-10-07T04:09:12.674660Z","shell.execute_reply":"2023-10-07T04:09:12.682140Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Confusion Matrix\n\n- The confusion matrix is a tool used in machine learning to help analyze the performance of a learning algorithm. It is a two-by-two matrix where each column corresponds to a training dataset and each row corresponds to a classifier. \n\n\n<center>\n<img src=\"https://miro.medium.com/v2/resize:fit:712/1*Z54JgbS4DUwWSknhDCvNTQ.png\">\n</center>\n\n\n\n- The rows and columns are filled with numbers that represent the number of times each classifier classified a sample as belonging to a particular class.\n\n- The confusion matrix can be used to measure the performance of a learning algorithm by determining which classifiers are most confused by the data. It can also be used to determine which classes are most difficult for a learning algorithm to distinguish.","metadata":{}},{"cell_type":"code","source":"#Making the Confusion Matrix\nconfusionmatrix = np.around(confusion_matrix(y_test, y_pred, normalize='true'),3)\nplot_confusion_matrix(confusionmatrix, np.unique(train.label))","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:09:14.708679Z","iopub.execute_input":"2023-10-07T04:09:14.709045Z","iopub.status.idle":"2023-10-07T04:09:15.441213Z","shell.execute_reply.started":"2023-10-07T04:09:14.709013Z","shell.execute_reply":"2023-10-07T04:09:15.440335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Plotting the figure\nplt.plot(range(len(history.history['loss'])),np.around(history.history['loss'],4), marker='o' )\nplt.plot(range(len(history.history['val_loss'])),np.around(history.history['val_loss'],4), marker='o' )\n\nxpoints =  range(len(history.history['val_loss']))\nypoints =  np.around(history.history['val_loss'],4)\n\n#Annotating\nfor x, y in zip(xpoints, ypoints):\n    plt.annotate(y,xy=(x, y), xytext=(-2, 8),textcoords='offset points', ha='center', va='bottom')\n\n#Setting the ticks, title, legend and x and y label\nplt.xticks(ticks=range(len(history.history['loss'])))\nplt.title('Training Loss Plot', fontsize=18, y=1.02)\nplt.legend(['Loss', 'Validation Loss'])\nplt.xlabel('Epochs')\nplt.ylabel('Training Loss')\n\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-07T05:16:14.011862Z","iopub.execute_input":"2023-10-07T05:16:14.012248Z","iopub.status.idle":"2023-10-07T05:16:14.334358Z","shell.execute_reply.started":"2023-10-07T05:16:14.012217Z","shell.execute_reply":"2023-10-07T05:16:14.333465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(range(len(history.history['accuracy'])),np.around(history.history['accuracy'],4), marker='o' )\nplt.plot(range(len(history.history['val_accuracy'])),np.around(history.history['val_accuracy'],4), marker='o' )\n\nxpoints =  range(len(history.history['val_accuracy']))\nypoints =  np.around(history.history['val_accuracy'],4)\nfor x, y in zip(xpoints, ypoints):\n    plt.annotate(y,xy=(x, y), xytext=(-2, 5),textcoords='offset points', ha='center', va='bottom')\n    \nplt.xticks(ticks=range(len(history.history['val_accuracy'])))\nplt.title('Training Accuracy Plot', fontsize=18, y=1.1)\nplt.legend(['Accuracy', 'Validation Accuracy'])\nplt.xlabel('Epochs')\nplt.ylabel('Training Accuracy')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-07T05:16:20.635581Z","iopub.execute_input":"2023-10-07T05:16:20.635931Z","iopub.status.idle":"2023-10-07T05:16:20.982022Z","shell.execute_reply.started":"2023-10-07T05:16:20.635901Z","shell.execute_reply":"2023-10-07T05:16:20.981114Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(range(len(history.history['lr'])),np.around(history.history['lr'],4), marker='o' )\nxpoints =  range(len(history.history['lr']))\nypoints =  history.history['lr']\nfor x, y in zip(xpoints, ypoints):\n    plt.annotate(y,xy=(x, y), xytext=(-2, 10),textcoords='offset points', ha='center', va='bottom')\n    \nplt.xticks(ticks=range(len(history.history['lr'])))\nplt.title('LR Decay Plot', fontsize=18, y=1.1)\nplt.xlabel('Epochs')\nplt.ylabel('Learning Rate')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2023-10-07T05:16:26.112078Z","iopub.execute_input":"2023-10-07T05:16:26.112648Z","iopub.status.idle":"2023-10-07T05:16:26.417821Z","shell.execute_reply.started":"2023-10-07T05:16:26.112606Z","shell.execute_reply":"2023-10-07T05:16:26.416891Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we can see, the learning rate decayed over the epochs and this helped our model converge better in a way that we don't want to be speeding as quickly as we did initially which will result in overshooting and the recalibrating ourselves to reach the global minima.","metadata":{}},{"cell_type":"code","source":"#Getting the probablities\ny_prob_train = model.predict(X_train)\ntrain_pred = y_prob_train.argmax(axis=1)\nprint('=== TRAIN CLASSIFICATION REPORT ===')\nprint(classification_report(y_train, train_pred))","metadata":{"execution":{"iopub.status.busy":"2023-10-07T06:02:40.504414Z","iopub.execute_input":"2023-10-07T06:02:40.504794Z","iopub.status.idle":"2023-10-07T06:02:44.452595Z","shell.execute_reply.started":"2023-10-07T06:02:40.504762Z","shell.execute_reply":"2023-10-07T06:02:44.451576Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('=== TEST CLASSIFICATION REPORT ===')\nprint(classification_report(y_test, y_pred))","metadata":{"execution":{"iopub.status.busy":"2023-10-07T06:01:26.737826Z","iopub.execute_input":"2023-10-07T06:01:26.738256Z","iopub.status.idle":"2023-10-07T06:01:26.769445Z","shell.execute_reply.started":"2023-10-07T06:01:26.738223Z","shell.execute_reply":"2023-10-07T06:01:26.768348Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Misclassification \n\nUsing a misclassification plot, we aim to find out the class where the model is making the most mistakes.","metadata":{}},{"cell_type":"code","source":"count = {int(value): list(y_train[y_train != train_pred]).count(value) for value in set(y_train[y_train != train_pred])}\nmaxvalue = max(count, key=count.get)\n\n#Setting the color of the most misqualified labels to 'indianred'\ntr_colors = {}\nkeys = range(len(np.unique(y_train)))\nfor i in keys:\n    for x in ['lightgray']:\n        for j in ['red']:\n            tr_colors[i] = x\n            tr_colors[maxvalue] = j\nprint(tr_colors)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T06:06:18.572078Z","iopub.execute_input":"2023-10-07T06:06:18.573206Z","iopub.status.idle":"2023-10-07T06:06:18.592283Z","shell.execute_reply.started":"2023-10-07T06:06:18.573159Z","shell.execute_reply":"2023-10-07T06:06:18.591066Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.countplot(y_train[y_train != train_pred], palette=tr_colors)\nplt.xticks(ticks=np.unique(train.label), labels= product)\nfor bar in ax.patches:\n    ax.annotate(format(bar.get_height(), '.0f'),(bar.get_x() + bar.get_width() / 2,bar.get_height()), \n                 ha='center', va='center',size=15, xytext=(0, 8),textcoords='offset points')\nplt.title('Misclassification Plot, Train', fontsize=18, y=1.1)\nplt.xlabel('Classes')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T06:06:32.474396Z","iopub.execute_input":"2023-10-07T06:06:32.475185Z","iopub.status.idle":"2023-10-07T06:06:32.785309Z","shell.execute_reply.started":"2023-10-07T06:06:32.475139Z","shell.execute_reply":"2023-10-07T06:06:32.784354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"count = {int(value): list(y_test[y_test != y_pred]).count(value) for value in set(y_test[y_test != y_pred])}\nmaxvalue = max(count, key=count.get)\n\n#Setting the color of the most misqualified labels to 'indianred'\ncolors = {}\nkeys = range(len(np.unique(y_train)))\nfor i in keys:\n    for x in ['lightgray']:\n        for j in ['red']:\n            colors[i] = x\n            colors[maxvalue] = j\nprint(colors)","metadata":{"execution":{"iopub.status.busy":"2023-10-07T04:09:42.008903Z","iopub.execute_input":"2023-10-07T04:09:42.009284Z","iopub.status.idle":"2023-10-07T04:09:42.022141Z","shell.execute_reply.started":"2023-10-07T04:09:42.009253Z","shell.execute_reply":"2023-10-07T04:09:42.021068Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = sns.countplot(y_test[y_test != y_pred], palette=colors)\nplt.xticks(ticks=np.unique(train.label), labels= product)\nfor bar in ax.patches:\n    ax.annotate(format(bar.get_height(), '.0f'),(bar.get_x() + bar.get_width() / 2,bar.get_height()), \n                 ha='center', va='center',size=15, xytext=(0, 8),textcoords='offset points')\nplt.title('Misclassification Plot, Test', fontsize=18, y=1.1)\nplt.xlabel('Classes')\nplt.ylabel('Count')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2023-10-07T05:16:32.597807Z","iopub.execute_input":"2023-10-07T05:16:32.598188Z","iopub.status.idle":"2023-10-07T05:16:32.949772Z","shell.execute_reply.started":"2023-10-07T05:16:32.598158Z","shell.execute_reply":"2023-10-07T05:16:32.948830Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observation**\n\n- Model is making the highest error on shirt which means it's probably getting confused between Shirt and a Pullover.","metadata":{}},{"cell_type":"markdown","source":"### Misclassification Error Rate","metadata":{}},{"cell_type":"code","source":"# Calculate the number of misclassified cases in the training data\nmisclassified_train_cases = len(y_train[y_train != train_pred])\n\n# Calculate the total number of cases in the training data\ntotal_train_cases = len(y_train)\n\n# Calculate the error rate as a percentage\nerror_rate = np.around((misclassified_train_cases / total_train_cases) * 100, 4)\n\n# Print the results\nprint(f\"Misclassified cases in training data: {misclassified_train_cases} out of {total_train_cases} cases\")\nprint(f\"Error rate: {error_rate}%\")","metadata":{"execution":{"iopub.status.busy":"2023-10-07T06:10:54.214502Z","iopub.execute_input":"2023-10-07T06:10:54.215559Z","iopub.status.idle":"2023-10-07T06:10:54.223129Z","shell.execute_reply.started":"2023-10-07T06:10:54.215510Z","shell.execute_reply":"2023-10-07T06:10:54.221945Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Calculate the number of misclassified cases in the testing data\nmisclassified = len(y_test[y_test != y_pred])\n\n# Calculate the total number of cases in the testing data\ntotal_test_cases = len(X_test)\n\n# Calculate the error rate as a percentage\nerror_rate = np.around((misclassified / total_test_cases) * 100, 4)\n\nprint(f\"Misclassified cases in testing data: {misclassified} out of {total_test_cases} cases\")\nprint(f\"Error rate: {error_rate}%\")","metadata":{"execution":{"iopub.status.busy":"2023-10-07T06:16:26.233154Z","iopub.execute_input":"2023-10-07T06:16:26.233544Z","iopub.status.idle":"2023-10-07T06:16:26.241201Z","shell.execute_reply.started":"2023-10-07T06:16:26.233513Z","shell.execute_reply":"2023-10-07T06:16:26.239934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Conclusion\n\n1. **Convolutional Neural Networks (CNNs):** We explored the power of CNNs for image classification on the MNIST dataset. CNNs are specifically designed for image-related tasks, utilizing convolutional layers to extract meaningful features from images. Our CNN model demonstrated strong performance in digit recognition.\n\n2. **Data Augmentation:** Data Augmentation proved to be a vital technique for enhancing model generalization. By generating augmented images through various transformations, such as rotations and flips, we expanded our training dataset, making our model more robust to variations in input data.\n\n3. **Learning Rate (LR) Tracking and Decay:** In our journey to improve the performance of our model, we found that implementing learning rate (LR) decay played a pivotal role. \n\n    - By experimenting with LR decay schedule, we fine-tuned our LR reduction strategy resulting in faster model convergence and higher accuracy. \n\n    - Alongside this, closely monitoring the Learning Rate throughout training was essential for gaining valuable insights into our model's learning dynamics.\n    \n4. **Image Classification:** Our primary objective was image classification, specifically recognizing fashion in the Fashion MNIST dataset. The CNN model demonstrated impressive accuracy in correctly classifying digits, showcasing the effectiveness of CNNs in image classification tasks.\n\nThis project highlighted the effectiveness of CNNs for image classification, emphasizing the importance of Data Augmentation, LR tracking, and LR decay as essential components of training deep learning models. These techniques not only enhanced the model's performance but also provided valuable insights into the training process. When working on image classification tasks, leveraging CNNs and these techniques can lead to robust and accurate models.\n\nIf you liked the notebook, consider upvoting and follow for more!","metadata":{}},{"cell_type":"markdown","source":"## Further Reading\n\n1. Digit Classification - ANNs vs CNNs https://bit.ly/3pC4xec\n2. Image augmentation and overfitting - YouTube Video https://bit.ly/3QJT5Jp\n3. Google Colab Notebook - https://bit.ly/3Ac5xe8\n4. Data Augmentation Jupyter Notebook https://bit.ly/3Kew0Mx","metadata":{}}]}