{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Machine Learning Engineer Nanodegree\n\n## Capstone Project\n\n## Project: Write an Algorithm for Distracted Driver Detection \n\n---\n\n\n\n>**Note:** Code and Markdown cells can be executed using the **Shift + Enter** keyboard shortcut.  Markdown cells can be edited by double-clicking the cell to enter edit mode.\n\n\n\n---\n\n### The Road Ahead\n\nThe notebook is broken into separate steps as shown below.\n\n* [Step 0](#step0): Import Datasets\n* [Step 1](#step1): Create and train a CNN to Classify Driver Images (from Scratch)\n* [Step 2](#step3): Train a CNN with Transfer Learning (Using Fine-tuned VGG16)\n* [Step 3](#step4): Kaggle Results\n\n\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","trusted":true}},{"cell_type":"code","source":"import numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport os\n\nimport keras\nimport numpy\nfrom keras.preprocessing.image import ImageDataGenerator","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:12.593115Z","iopub.execute_input":"2023-01-04T18:00:12.593447Z","iopub.status.idle":"2023-01-04T18:00:12.597882Z","shell.execute_reply.started":"2023-01-04T18:00:12.593390Z","shell.execute_reply":"2023-01-04T18:00:12.597143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import keras\nimport numpy\nfrom keras.preprocessing.image import ImageDataGenerator","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:12.603065Z","iopub.execute_input":"2023-01-04T18:00:12.603455Z","iopub.status.idle":"2023-01-04T18:00:12.607856Z","shell.execute_reply.started":"2023-01-04T18:00:12.603400Z","shell.execute_reply":"2023-01-04T18:00:12.606901Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<a id=\"step0\"></a>\n## Step 0: Import Datasets\n\n### Import Driver Dataset\n\nIn the following code cell, we create a instance of ImageDataGenerator which does all preprocessing operations on images that we are going to feed to our CNN","metadata":{}},{"cell_type":"code","source":"train_datagen = ImageDataGenerator(\n        rescale=1./255, validation_split=0.2)\n#         shear_range=0.2,\n#         zoom_range=0.2,\n#         horizontal_flip=True)\n","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","execution":{"iopub.status.busy":"2023-01-04T18:00:12.609024Z","iopub.execute_input":"2023-01-04T18:00:12.609510Z","iopub.status.idle":"2023-01-04T18:00:12.617215Z","shell.execute_reply.started":"2023-01-04T18:00:12.609460Z","shell.execute_reply":"2023-01-04T18:00:12.616477Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"In the following code, we are loading only 32 images at once into memory and performing all the preprocessing operations on loaded images. The flow_from_directory loads a defined set of images from the location of images, instead of loading all images at once into memory.\n\n- train_generator contains training set\n- val_generator contains validation set","metadata":{}},{"cell_type":"code","source":"train_data = '../input/state-farm-distracted-driver-detection/imgs/train'\ntest_data = '../input/state-farm-distracted-driver-detection/imgs/test'\ntrain_generator = train_datagen.flow_from_directory(\n        train_data,\n        target_size=(224, 224),\n        batch_size=32,\n        class_mode='categorical',\n        subset='training')\n\nval_generator = train_datagen.flow_from_directory(\n        train_data,\n        target_size=(224,224),\n        batch_size=32,\n        class_mode='categorical',\n        subset='validation')\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:12.618415Z","iopub.execute_input":"2023-01-04T18:00:12.618930Z","iopub.status.idle":"2023-01-04T18:00:47.456404Z","shell.execute_reply.started":"2023-01-04T18:00:12.618881Z","shell.execute_reply":"2023-01-04T18:00:47.455479Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from PIL import Image\n\nac_labels=  [\"c0: safe driving\",\n\"c1: texting - right\",\n\"c2: talking on the phone - right\",\n\"c3: texting - left\",\n\"c4: talking on the phone - left\",\n\"c5: operating the radio\",\n\"c6: drinking\",\n\"c7: reaching behind\",\n\"c8: hair and makeup\",\n\"c9: talking to passenger\"]\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:47.457516Z","iopub.execute_input":"2023-01-04T18:00:47.457770Z","iopub.status.idle":"2023-01-04T18:00:47.462801Z","shell.execute_reply.started":"2023-01-04T18:00:47.457724Z","shell.execute_reply":"2023-01-04T18:00:47.461772Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"imgs, labels = next(train_generator)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:47.464119Z","iopub.execute_input":"2023-01-04T18:00:47.464691Z","iopub.status.idle":"2023-01-04T18:00:47.853644Z","shell.execute_reply.started":"2023-01-04T18:00:47.464641Z","shell.execute_reply":"2023-01-04T18:00:47.852965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Dataset Exploration\n- Following code block explores and reveals the dataset","metadata":{}},{"cell_type":"code","source":"import functools\n\ndef list_counts(start_dir):\n    lst = sorted(os.listdir(start_dir))\n    out = [(fil, len(os.listdir( os.path.join(start_dir, fil)))) for fil in lst if os.path.isdir(os.path.join(start_dir,fil))]\n    return out\n\nout = list_counts(train_data)\nlabels, counts = zip(*out)\nprint(\"Total number of images : \",functools.reduce(lambda a,b : a+b, counts))\nout","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:47.856132Z","iopub.execute_input":"2023-01-04T18:00:47.856598Z","iopub.status.idle":"2023-01-04T18:00:47.881656Z","shell.execute_reply.started":"2023-01-04T18:00:47.856545Z","shell.execute_reply":"2023-01-04T18:00:47.880920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Visualization\nFollowing code block. displays images and their respective labels of training/ validation dataset","metadata":{}},{"cell_type":"code","source":"\nimport matplotlib.pyplot as plt\nimport pandas as pd\n# Pretty display for notebooks\n%matplotlib inline\n\n\ny = np.array(counts)\nwidth = 1/1.5\nN = len(y)\nx = range(N)\n\nfig = plt.figure(figsize=(20,15))\nay = fig.add_subplot(211)\n\nplt.xticks(x, labels, size=15)\nplt.yticks(size=15)\n\nay.bar(x, y, width, color=\"blue\")\n\nplt.title('Bar Chart',size=25)\nplt.xlabel('classname',size=15)\nplt.ylabel('Count',size=15)\n\nplt.show()\n\n\n\ndef showImages(imgs ,inlabels=None, single=True):\n    if single:\n        aim = (imgs * 255 ).astype(np.uint8)\n        img = Image.fromarray(aim)\n        if labels is not None:\n            print(\"Label : \", ac_labels[np.argmax(inlabels)])\n        plt.imshow(img)\n        plt.show()\n    else:\n        for i,img in enumerate(imgs):\n            lbl = None\n            if inlabels is not None:\n                lbl = labels[i]\n            showImages(img, lbl)\n\nind = 1\nshowImages(imgs[:ind], inlabels=labels[:ind], single = False)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:47.882856Z","iopub.execute_input":"2023-01-04T18:00:47.883094Z","iopub.status.idle":"2023-01-04T18:00:48.724997Z","shell.execute_reply.started":"2023-01-04T18:00:47.883051Z","shell.execute_reply":"2023-01-04T18:00:48.724158Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n<a id='step1'></a>\n## Step 1: Create a CNN to Classify Driver Images (from Scratch)\n\n\n\nA CNN is created to classify driver images.  At the end of the code cell block, the layers of the model are summarized by executing the line:\n    \n        model.summary()\n\nWe have created 6 convolutional layers with 1 max pooling layer and 1 GlobalAveragePooling in between. Filters were increased from 8 to 512 in total convolutional layers. Also dropout was used along with Global average pooling layer before using the fully connected layer. Number of nodes in the last fully connected layer were setup as 10 along with softmax activation function. ReLU activation function was used for all other layers.\n\n6 convolutional layers were used to learn hierarchy of high level features. Max pooling layer is added to reduce the dimensionality. Global Average Pooling layer is added to reduce the dimensionality as well as the matrix to row vector. This is because fully connected layer only accepts row vector. Dropout layers were added to reduce overfitting and ensure that the network generalizes well. The last fully connected layer with softmax activation function is added to obtain probabilities of the prediction.\n","metadata":{}},{"cell_type":"code","source":"from keras.layers import ZeroPadding2D, Conv2D, MaxPooling2D, Flatten, Dense, Dropout, Input\nfrom keras.layers import GlobalAveragePooling2D, MaxPooling2D\nfrom keras.models import Model, Sequential\nfrom keras.callbacks import ModelCheckpoint\nfrom keras import regularizers","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:48.726551Z","iopub.execute_input":"2023-01-04T18:00:48.727029Z","iopub.status.idle":"2023-01-04T18:00:48.732579Z","shell.execute_reply.started":"2023-01-04T18:00:48.726813Z","shell.execute_reply":"2023-01-04T18:00:48.731659Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"input_layer = Input(shape=(224,224, 3))\n\nconv = Conv2D(filters=8, kernel_size=2)(input_layer)\nconv = Conv2D(filters=16, kernel_size=2, activation='relu')(conv)\nconv = Conv2D(filters=32, kernel_size=2, activation='relu')(conv)\nconv = MaxPooling2D()(conv)\n\nconv = Conv2D(filters=64, kernel_size=2, activation='relu')(conv)\nconv = Conv2D(filters=128, kernel_size=2, activation='relu')(conv)\nconv = Conv2D(filters=512, kernel_size=2, activation='relu')(conv)\n\nconv = GlobalAveragePooling2D()(conv)\ndense = Dense(units=500, activation='relu')(conv)\ndense = Dropout(0.1)(dense)\ndense = Dense(units=100, activation='relu')(dense)\ndense = Dropout(0.1)(dense)\noutput = Dense(units=10, activation='softmax')(dense)\n\nmodel = Model(inputs=input_layer, outputs = output)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:48.734063Z","iopub.execute_input":"2023-01-04T18:00:48.734625Z","iopub.status.idle":"2023-01-04T18:00:48.925319Z","shell.execute_reply.started":"2023-01-04T18:00:48.734498Z","shell.execute_reply":"2023-01-04T18:00:48.924552Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.summary()","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:48.926575Z","iopub.execute_input":"2023-01-04T18:00:48.926844Z","iopub.status.idle":"2023-01-04T18:00:48.937116Z","shell.execute_reply.started":"2023-01-04T18:00:48.926789Z","shell.execute_reply":"2023-01-04T18:00:48.936456Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compile CNN from scratch Model","metadata":{}},{"cell_type":"code","source":"model.compile(optimizer='adam', loss='categorical_crossentropy', metrics=['accuracy'])","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:48.938401Z","iopub.execute_input":"2023-01-04T18:00:48.938641Z","iopub.status.idle":"2023-01-04T18:00:48.981480Z","shell.execute_reply.started":"2023-01-04T18:00:48.938597Z","shell.execute_reply":"2023-01-04T18:00:48.980820Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### (IMPLEMENTATION) Train the Model\n\nThe model is trained in the code cell below. Model checkpointing is used to save the model that attains the best validation loss.\n","metadata":{}},{"cell_type":"code","source":"#model.load_weights('best_model_1.hdf5')\ncheckpoint = ModelCheckpoint('best_model_1.hdf5', save_best_only=True, verbose=1)\n\nhistory = model.fit_generator(train_generator, steps_per_epoch=len(train_generator),\n                    epochs=10,\n                   validation_data = val_generator,\n                   validation_steps=len(val_generator),\n                    callbacks=[checkpoint] )","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:00:48.982530Z","iopub.execute_input":"2023-01-04T18:00:48.982768Z","iopub.status.idle":"2023-01-04T18:31:04.117480Z","shell.execute_reply.started":"2023-01-04T18:00:48.982724Z","shell.execute_reply":"2023-01-04T18:31:04.116726Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- In the following code block, we are plotting the loss value of training and validation to see whether the model is learning or not","metadata":{}},{"cell_type":"code","source":"plt.plot(history.history['val_loss'])\nplt.show()\nplt.plot(history.history['loss'])","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:31:04.121380Z","iopub.execute_input":"2023-01-04T18:31:04.124372Z","iopub.status.idle":"2023-01-04T18:31:04.751380Z","shell.execute_reply.started":"2023-01-04T18:31:04.124312Z","shell.execute_reply":"2023-01-04T18:31:04.750297Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#### Loading test dataset into memory and predict on test images\n- Following code block contains a method, that loads all test images into memory by batches, each batch with 32 images when it is called each time","metadata":{}},{"cell_type":"code","source":"\n!ls ../input/state-farm-distracted-driver-detection/\nimport os\n# model.load_weights('../input/best_model_1.hdf5')\n#Test Images\nbatch_index = 0\nfiles_list = os.listdir(\"../input/state-farm-distracted-driver-detection/imgs/test/\")\ndef load_test_images(batch_size=32, src='../input/state-farm-distracted-driver-detection/imgs/test/'):\n    global batch_index, files_list\n    imgs_list = files_list[batch_index: batch_index+batch_size]\n    batch_index += len(imgs_list)\n    batch_imgs = []\n    for img_name in imgs_list:\n        img = Image.open(src+img_name)\n        im = img.resize((224,224))\n        batch_imgs.append(np.array(im)/255.)\n#     plt.imshow()\n#     plt.show()\n    return np.array(batch_imgs)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:32:13.714488Z","iopub.execute_input":"2023-01-04T18:32:13.714792Z","iopub.status.idle":"2023-01-04T18:32:18.209557Z","shell.execute_reply.started":"2023-01-04T18:32:13.714737Z","shell.execute_reply":"2023-01-04T18:32:18.208673Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Following code block predicts on test images by loading 32 images a batch using load_test_images function","metadata":{}},{"cell_type":"code","source":"#Test Images write\nimport sys\npreds_list = np.array([])\nbatch_index=0\nbatch_size = 32\nwhile True:\n    tst_imgs = load_test_images(batch_size=batch_size)\n    if(tst_imgs.shape[0] <= 0  ):\n        print(\"Batchsize is less : \",batch_index)\n        break\n    preds = model.predict(tst_imgs)\n    print(\"\\r {},  batch_size : {}, nth_batch/all_batch : {}/{}\".format(preds_list.shape,batch_size, batch_index, len(files_list)),end=\"\")    \n    sys.stdout.flush()\n    if len(preds_list) == 0:\n        preds_list = np.array(preds)\n    else:\n        preds_list = np.append(preds_list, preds, axis=0)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T18:32:31.244698Z","iopub.execute_input":"2023-01-04T18:32:31.245002Z","iopub.status.idle":"2023-01-04T18:56:59.618301Z","shell.execute_reply.started":"2023-01-04T18:32:31.244946Z","shell.execute_reply":"2023-01-04T18:56:59.617422Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Convert the predicted output to submittable output file","metadata":{}},{"cell_type":"code","source":"titles = \"img,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9\".split(\",\")\nnames = pd.DataFrame(files_list[:len(preds_list)])\nnames.columns=[\"img\"]\ndf = pd.DataFrame(preds_list)\ndf.columns=titles[1:]\ndf['img']=names['img']\ndf = df[titles]\ndf.tail()\ndf.to_csv('sub.csv',index=False)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:01:04.670481Z","iopub.execute_input":"2023-01-04T19:01:04.670772Z","iopub.status.idle":"2023-01-04T19:01:05.534101Z","shell.execute_reply.started":"2023-01-04T19:01:04.670715Z","shell.execute_reply":"2023-01-04T19:01:05.533369Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Predictions by scratch CNN on some sample images\n- Seems like the predictions are not so accurate because of less training","metadata":{}},{"cell_type":"code","source":"indices = [1,24]\nfor index in indices:\n#     display(df.iloc[index])\n    cls = np.argmax(list(df.iloc[index][1:]))\n    print(\"label : \",ac_labels[cls])\n    im_test = Image.open('../input/state-farm-distracted-driver-detection/imgs/test/'+df.iloc[index]['img'])\n    plt.imshow(np.array(im_test))\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:01:32.185381Z","iopub.execute_input":"2023-01-04T19:01:32.185668Z","iopub.status.idle":"2023-01-04T19:01:32.838384Z","shell.execute_reply.started":"2023-01-04T19:01:32.185617Z","shell.execute_reply":"2023-01-04T19:01:32.837709Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"---\n<a id=\"step2\"></a>\n##  Step 2: Train a CNN with Transfer Learning (Using Fine-tuned VGG16 Model)\n\n- In the following steps, we are going to load a VGG16 model, without top Fully connected layers, with its trained weights\n- Then we are going add Fully Connected Layers on top of GlobalAveragePooling layer with Dropout layers in between\n- The following code block contains a method load_VGG16, which loads VGG16 model with weights loaded when weights file location is given","metadata":{}},{"cell_type":"code","source":"\ndef load_VGG16(weights_path=None, no_top=True):\n\n    input_shape = (224, 224, 3)\n\n    #Instantiate an empty model\n    img_input = Input(shape=input_shape)   # Block 1\n    x = Conv2D(64, (3, 3), activation='relu', padding='same', name='block1_conv1')(img_input)\n    x = Conv2D(64, (3, 3), activation='relu', padding='same', name='block1_conv2')(x)\n    x = MaxPooling2D((2, 2), strides=(2, 2), name='block1_pool')(x)\n\n    # Block 2\n    x = Conv2D(128, (3, 3), activation='relu', padding='same', name='block2_conv1')(x)\n    x = Conv2D(128, (3, 3), activation='relu', padding='same', name='block2_conv2')(x)\n    x = MaxPooling2D((2, 2), strides=(2, 2), name='block2_pool')(x)\n\n    # Block 3\n    x = Conv2D(256, (3, 3), activation='relu', padding='same', name='block3_conv1')(x)\n    x = Conv2D(256, (3, 3), activation='relu', padding='same', name='block3_conv2')(x)\n    x = Conv2D(256, (3, 3), activation='relu', padding='same', name='block3_conv3')(x)\n    x = MaxPooling2D((2, 2), strides=(2, 2), name='block3_pool')(x)\n\n    # Block 4\n    x = Conv2D(512, (3, 3), activation='relu', padding='same', name='block4_conv1')(x)\n    x = Conv2D(512, (3, 3), activation='relu', padding='same', name='block4_conv2')(x)\n    x = Conv2D(512, (3, 3), activation='relu', padding='same', name='block4_conv3')(x)\n    x = MaxPooling2D((2, 2), strides=(2, 2), name='block4_pool')(x)\n\n    # Block 5\n    x = Conv2D(512, (3, 3), activation='relu', padding='same', name='block5_conv1')(x)\n    x = Conv2D(512, (3, 3), activation='relu', padding='same', name='block5_conv2')(x)\n    x = Conv2D(512, (3, 3), activation='relu', padding='same', name='block5_conv3')(x)\n    x = MaxPooling2D((2, 2), strides=(2, 2), name='block5_pool')(x)\n    x = GlobalAveragePooling2D()(x)\n    vmodel = Model(img_input, x, name='vgg16')\n    if weights_path is not None:\n        print(\"Weights have been loaded.\")\n        vmodel.load_weights(weights_path)\n\n    return vmodel\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:03:26.075599Z","iopub.execute_input":"2023-01-04T19:03:26.075914Z","iopub.status.idle":"2023-01-04T19:03:26.088379Z","shell.execute_reply.started":"2023-01-04T19:03:26.075853Z","shell.execute_reply":"2023-01-04T19:03:26.087705Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- In the following code block, we have added 3 Dense layers with dropout layers on top of VGG16 CNN model.","metadata":{}},{"cell_type":"code","source":"vgg_model_raw = load_VGG16('/kaggle/input/vgg16-weight/vgg16_weights_tf_dim_ordering_tf_kernels_notop.h5')\n\nvgg_model = vgg_model_raw.output\n#vgg_model = Flatten()(vgg_model)\nvgg_model = Dense(5000, activation='relu',kernel_regularizer=regularizers.l2(0.00001))(vgg_model)\n#vgg_model = Dropout(0.1)(vgg_model)\n#vgg_model = Dense(1000, activation='relu')(vgg_model)\nvgg_model = Dropout(0.1)(vgg_model)\nvgg_model = Dense(500, activation='relu',kernel_regularizer=regularizers.l2(0.00001))(vgg_model)\nvgg_model = Dropout(0.1)(vgg_model)\nvgg_model = Dense(10, activation='softmax')(vgg_model)\nvgg_m = Model(inputs=vgg_model_raw.input, outputs= vgg_model)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:31:23.764916Z","iopub.execute_input":"2023-01-04T19:31:23.765196Z","iopub.status.idle":"2023-01-04T19:31:24.823650Z","shell.execute_reply.started":"2023-01-04T19:31:23.765144Z","shell.execute_reply":"2023-01-04T19:31:24.822755Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vgg_m.layers[16].get_weights()","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:31:38.903723Z","iopub.execute_input":"2023-01-04T19:31:38.904161Z","iopub.status.idle":"2023-01-04T19:31:39.110031Z","shell.execute_reply.started":"2023-01-04T19:31:38.904094Z","shell.execute_reply":"2023-01-04T19:31:39.109288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Compile the vgg16 model with categorical crossentropy as loss function and SGD as optimizer","metadata":{}},{"cell_type":"code","source":"vgg_m.compile(loss='categorical_crossentropy', optimizer=keras.optimizers.SGD(0.001), metrics=['accuracy'])\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:31:51.915568Z","iopub.execute_input":"2023-01-04T19:31:51.915841Z","iopub.status.idle":"2023-01-04T19:31:51.956440Z","shell.execute_reply.started":"2023-01-04T19:31:51.915795Z","shell.execute_reply":"2023-01-04T19:31:51.955660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"vgg_m.summary()","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:32:01.463029Z","iopub.execute_input":"2023-01-04T19:32:01.463333Z","iopub.status.idle":"2023-01-04T19:32:01.477594Z","shell.execute_reply.started":"2023-01-04T19:32:01.463274Z","shell.execute_reply":"2023-01-04T19:32:01.476727Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Create a Checkpoint to save best weights, when loss in improvised\n- Train the model for 6 epochs","metadata":{}},{"cell_type":"code","source":"\ncheckpoint = ModelCheckpoint('vgg_model.h5', save_best_only=True, verbose=1)\n\nhistory = vgg_m.fit_generator(train_generator, steps_per_epoch=len(train_generator),\n                   epochs=6,\n                   validation_data = val_generator,\n                   validation_steps=len(val_generator),\n                   callbacks=[checkpoint] )\n\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:32:07.630125Z","iopub.execute_input":"2023-01-04T19:32:07.630453Z","iopub.status.idle":"2023-01-04T19:50:12.298860Z","shell.execute_reply.started":"2023-01-04T19:32:07.630394Z","shell.execute_reply":"2023-01-04T19:50:12.297912Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.plot(history.history['val_loss'])\nplt.show()\nplt.plot(history.history['loss'])","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:50:31.379038Z","iopub.execute_input":"2023-01-04T19:50:31.379347Z","iopub.status.idle":"2023-01-04T19:50:31.883355Z","shell.execute_reply.started":"2023-01-04T19:50:31.379290Z","shell.execute_reply":"2023-01-04T19:50:31.882390Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- Following method gives a batch of 32 images at each call, just like a generator","metadata":{}},{"cell_type":"code","source":"!ls\nimport os\n#model.load_weights('../input/best_model_1.hdf5')\n#Test Images\nbatch_index = 0\n#files_lst = os.listdir(\"../input/state-farm-distracted-driver-detection/test\")\nfiles_list = os.listdir(\"../input/state-farm-distracted-driver-detection/imgs/test\")\ndef load_test_images(batch_size=32, src='../input/state-farm-distracted-driver-detection/imgs/test/'):\n    global batch_index, files_list\n    imgs_list = files_list[batch_index: batch_index+batch_size]\n    batch_index += len(imgs_list)\n    batch_imgs = []\n    for img_name in imgs_list:\n        img = Image.open(src+img_name)\n        im = img.resize((224,224))\n        batch_imgs.append(np.array(im)/255.)\n#     plt.imshow()\n#     plt.show()\n    return np.array(batch_imgs)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:51:09.559816Z","iopub.execute_input":"2023-01-04T19:51:09.560141Z","iopub.status.idle":"2023-01-04T19:51:10.585652Z","shell.execute_reply.started":"2023-01-04T19:51:09.560087Z","shell.execute_reply":"2023-01-04T19:51:10.584696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Test Images write\nimport sys\npreds_list = np.array([])\nbatch_index=0\nbatch_size = 32\n\nmm_raw = load_VGG16()\n\nmm_model = mm_raw.output\nmm_model = Dense(5000, activation='relu',kernel_regularizer=regularizers.l2(0.00001))(mm_model)\nmm_model = Dropout(0.1)(mm_model)\nmm_model = Dense(500, activation='relu',kernel_regularizer=regularizers.l2(0.00001))(mm_model)\nmm_model = Dropout(0.1)(mm_model)\nmm_model = Dense(10, activation='softmax')(mm_model)\nmm = Model(inputs=vgg_model_raw.input, outputs= vgg_model)\n\nmm.load_weights('vgg_model.h5')\n\n\n\nwhile True:\n    tst_imgs = load_test_images(batch_size=batch_size)\n    if(tst_imgs.shape[0] <= 0  ):\n        print(\"Batchsize is less : \",batch_index)\n        break\n    preds = mm.predict(tst_imgs)\n    print(\"\\r {},  batch_size : {}, nth_batch/all_batch : {}/{}\".format(preds_list.shape,batch_size, batch_index, len(files_list)),end=\"\")    \n    sys.stdout.flush()\n    if len(preds_list) == 0:\n        preds_list = np.array(preds)\n    else:\n        preds_list = np.append(preds_list, preds, axis=0)\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T19:51:33.992233Z","iopub.execute_input":"2023-01-04T19:51:33.992567Z","iopub.status.idle":"2023-01-04T20:04:22.418607Z","shell.execute_reply.started":"2023-01-04T19:51:33.992513Z","shell.execute_reply":"2023-01-04T20:04:22.417729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"\ntitles = \"img,c0,c1,c2,c3,c4,c5,c6,c7,c8,c9\".split(\",\")\nnames = pd.DataFrame(files_list[:len(preds_list)])\nnames.columns=[\"img\"]\ndf = pd.DataFrame(preds_list)\ndf.columns=titles[1:]\ndf['img']=names['img']\ndf = df[titles]\ndf.tail()","metadata":{"execution":{"iopub.status.busy":"2023-01-04T20:04:32.960786Z","iopub.execute_input":"2023-01-04T20:04:32.961085Z","iopub.status.idle":"2023-01-04T20:04:33.013116Z","shell.execute_reply.started":"2023-01-04T20:04:32.961030Z","shell.execute_reply":"2023-01-04T20:04:33.012354Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.to_csv('sub_VGG16.csv',index=False)","metadata":{"execution":{"iopub.status.busy":"2023-01-04T20:04:40.123650Z","iopub.execute_input":"2023-01-04T20:04:40.123936Z","iopub.status.idle":"2023-01-04T20:04:41.181061Z","shell.execute_reply.started":"2023-01-04T20:04:40.123883Z","shell.execute_reply":"2023-01-04T20:04:41.180267Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"indices = [10,24]\nfor index in indices:\n#     display(df.iloc[index])\n    cls = np.argmax(list(df.iloc[index][1:]))\n    print(\"label : \",ac_labels[cls])\n    im_test = Image.open('../input/state-farm-distracted-driver-detection/imgs/test/'+df.iloc[index]['img'])\n    plt.imshow(np.array(im_test))\n    plt.show()\n","metadata":{"execution":{"iopub.status.busy":"2023-01-04T20:05:04.709256Z","iopub.execute_input":"2023-01-04T20:05:04.709575Z","iopub.status.idle":"2023-01-04T20:05:05.382838Z","shell.execute_reply.started":"2023-01-04T20:05:04.709521Z","shell.execute_reply":"2023-01-04T20:05:05.381759Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"!cat sub_VGG16.csv | head -10","metadata":{"execution":{"iopub.status.busy":"2023-01-04T20:05:15.947283Z","iopub.execute_input":"2023-01-04T20:05:15.947583Z","iopub.status.idle":"2023-01-04T20:05:16.924904Z","shell.execute_reply.started":"2023-01-04T20:05:15.947532Z","shell.execute_reply":"2023-01-04T20:05:16.924016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]}]}