{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"nvidiaTeslaT4","dataSources":[{"sourceId":5048,"databundleVersionId":868335,"sourceType":"competition"}],"dockerImageVersionId":30699,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# DISTRACTION DRIVER DETECTION\nWe've all been there: a light turns green and the car in front of you doesn't budge. Or, a previously unremarkable vehicle suddenly slows and starts swerving from side-to-side.\n\nWhen you pass the offending driver, what do you expect to see? You certainly aren't surprised when you spot a driver who is texting, seemingly enraptured by social media, or in a lively hand-held conversation on their phone.\n\nWe are using [State Farm Distracted Driver Detection](https://www.kaggle.com/competitions/state-farm-distracted-driver-detection/overview) Dataset. This dataset includes various images labelled as 10 different classes given below:\n- c0: normal driving\n- c1: texting - right\n- c2: talking on the phone - right\n- c3: texting - left\n- c4: talking on the phone - left\n- c5: operating the radio\n- c6: drinking\n- c7: reaching behind\n- c8: hair and makeup\n- c9: talking to passenger\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19"}},{"cell_type":"markdown","source":"## Importing required Library\n- **numpy**: Numerical computing library for array operations.  \n- **pandas**: Data manipulation and analysis library for tabular data.  \n- **os**: Operating system interface for interacting with the file system.  \n- **shutil**: High-level file operations utility for file copying and removal.  \n- **time**: Time-related functions for measuring code execution time or delaying execution.  \n- **tqdm**: Progress bar utility for tracking iterations.  \n- **random**: Module for generating pseudo-random numbers and performing random selections.  \n- **matplotlib**: Plotting library for creating static, interactive, and animated visualizations.  ","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport shutil\nimport time\nfrom tqdm import tqdm\nimport random\nimport matplotlib.pyplot as plt","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:52:55.640259Z","iopub.execute_input":"2024-04-23T01:52:55.640608Z","iopub.status.idle":"2024-04-23T01:52:56.746802Z","shell.execute_reply.started":"2024-04-23T01:52:55.640579Z","shell.execute_reply":"2024-04-23T01:52:56.745817Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **PIL.Image**: Python Imaging Library (PIL) for opening, manipulating, and saving many different image file formats.\n- **from IPython.display import Image**: IPython utility for displaying images inline within Jupyter notebooks or IPython environments.","metadata":{}},{"cell_type":"code","source":"import PIL.Image\nfrom IPython.display import Image","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:52:57.648034Z","iopub.execute_input":"2024-04-23T01:52:57.648500Z","iopub.status.idle":"2024-04-23T01:52:57.653254Z","shell.execute_reply.started":"2024-04-23T01:52:57.648469Z","shell.execute_reply":"2024-04-23T01:52:57.652181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- **torch**: PyTorch, a popular deep learning framework for building and training neural networks.\n- **torch.nn as nn**: PyTorch's neural network module containing various layers, loss functions, and utilities for constructing neural network architectures.\n- **torchvision**: PyTorch's computer vision library containing datasets, models, and transformations for image-related tasks.\n- **from torchvision import models, transforms, datasets**: Specific modules within torchvision, including pre-trained models, image transformations, and common datasets for computer vision tasks.","metadata":{}},{"cell_type":"code","source":"import torch\nimport torch.nn as nn\nimport torchvision\nfrom torchvision import models,transforms,datasets","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:01.739161Z","iopub.execute_input":"2024-04-23T01:53:01.739769Z","iopub.status.idle":"2024-04-23T01:53:08.838534Z","shell.execute_reply.started":"2024-04-23T01:53:01.739725Z","shell.execute_reply":"2024-04-23T01:53:08.837654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Loading Dataset","metadata":{}},{"cell_type":"code","source":"path_train = \"/kaggle/input/state-farm-distracted-driver-detection/imgs/train\"","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:08.840340Z","iopub.execute_input":"2024-04-23T01:53:08.840835Z","iopub.status.idle":"2024-04-23T01:53:08.845108Z","shell.execute_reply.started":"2024-04-23T01:53:08.840805Z","shell.execute_reply":"2024-04-23T01:53:08.844182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Below mentioned are 10 different classes in which images are divided","metadata":{}},{"cell_type":"code","source":"classes = [c for c in os.listdir(path_train) if not c.startswith(\".\")]\nclasses.sort()\nprint(classes)","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:08.846854Z","iopub.execute_input":"2024-04-23T01:53:08.847203Z","iopub.status.idle":"2024-04-23T01:53:08.870490Z","shell.execute_reply.started":"2024-04-23T01:53:08.847171Z","shell.execute_reply":"2024-04-23T01:53:08.869650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"class_dict = {0 : \"safe driving\",\n              1 : \"texting - right\",\n              2 : \"talking on the phone - right\",\n              3 : \"texting - left\",\n              4 : \"talking on the phone - left\",\n              5 : \"operating the radio\",\n              6 : \"drinking\",\n              7 : \"reaching behind\",\n              8 : \"hair and makeup\",\n              9 : \"talking to passenger\"}","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:08.871875Z","iopub.execute_input":"2024-04-23T01:53:08.872107Z","iopub.status.idle":"2024-04-23T01:53:08.876958Z","shell.execute_reply.started":"2024-04-23T01:53:08.872087Z","shell.execute_reply":"2024-04-23T01:53:08.876066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code written below defines a series of image transformations using PyTorch's `transforms.Compose` function. \n- `transforms.Resize((400, 400))`: Resizes the image to have a height and width of 400 pixels. Aspect ratio is maintained.\n- `transforms.RandomRotation(10)`: Randomly rotates the image by a maximum of 10 degrees clockwise or counterclockwise.\n- `transforms.ToTensor()`: Converts the image to a PyTorch tensor. \n- `transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5))`: Normalizes the tensor by subtracting the mean (0.5, 0.5, 0.5) and dividing by the standard deviation (0.5, 0.5, 0.5) for each channel. This operation shifts the pixel values to be centered around 0 and scales them to be within the range [-1, 1].","metadata":{}},{"cell_type":"code","source":"transform = transforms.Compose([transforms.Resize((400, 400)),\n                           transforms.RandomRotation(10),\n                           transforms.ToTensor(),\n                           transforms.Normalize((0.5, 0.5, 0.5), (0.5, 0.5, 0.5)) \n                           #Scale image pixel value in range [-1, 1]\n                          ])","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:08.933554Z","iopub.execute_input":"2024-04-23T01:53:08.933904Z","iopub.status.idle":"2024-04-23T01:53:08.939476Z","shell.execute_reply.started":"2024-04-23T01:53:08.933876Z","shell.execute_reply":"2024-04-23T01:53:08.938381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet is used to create training and testing datasets using PyTorch's `ImageFolder` dataset class and the `random_split` function from `torch.utils.data`. We will use 80 percent of data as training data and rest for testing. For splitting images into training and testing data we will use `torch.utils.data.random_split` ","metadata":{}},{"cell_type":"code","source":"data = datasets.ImageFolder(root = path_train, transform = transform)\ntotal_len = len(data)\ntraining_len = int(0.8*total_len)\ntesting_len = total_len - training_len\n\ntraining_data,testing_data = torch.utils.data.random_split(data,(training_len,testing_len))","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:11.944103Z","iopub.execute_input":"2024-04-23T01:53:11.944483Z","iopub.status.idle":"2024-04-23T01:53:33.308025Z","shell.execute_reply.started":"2024-04-23T01:53:11.944454Z","shell.execute_reply":"2024-04-23T01:53:33.306715Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now we are creating a data loader for the training and testing dataset. It specifies a `batch size of 32`, meaning each iteration will provide 32 samples to the model. The `shuffle=True` argument shuffles the dataset before each epoch, ensuring the model sees data in different orders during training. The `drop_last=False` argument ensures that all samples are included, even if the dataset size is not perfectly divisible by the batch size.","metadata":{}},{"cell_type":"code","source":"train_loader = torch.utils.data.DataLoader(dataset=training_data,\n                                           batch_size=32,\n                                           shuffle=True,\n                                           drop_last=False)\ntest_loader = torch.utils.data.DataLoader(dataset=testing_data,\n                                          batch_size=32,\n                                          shuffle=False,\n                                          drop_last=False)","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:33.309861Z","iopub.execute_input":"2024-04-23T01:53:33.310199Z","iopub.status.idle":"2024-04-23T01:53:33.316135Z","shell.execute_reply.started":"2024-04-23T01:53:33.310171Z","shell.execute_reply":"2024-04-23T01:53:33.315049Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"img,c = data[0]\nprint(img.shape)\nprint(\"Label:\", classes[c], f\"({class_dict[c]})\")\n# imshow expects (height,width,channels)\n# but image tensors have (channels,width,height), impermute is used to map these properly\nplt.imshow(img.permute(1,2,0))  \nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:33.317746Z","iopub.execute_input":"2024-04-23T01:53:33.318126Z","iopub.status.idle":"2024-04-23T01:53:33.739381Z","shell.execute_reply.started":"2024-04-23T01:53:33.318098Z","shell.execute_reply":"2024-04-23T01:53:33.738476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code sets up the PyTorch device to utilize CUDA, which is the parallel computing platform and application programming interface model created by NVIDIA.","metadata":{}},{"cell_type":"code","source":"device = torch.device(\"cuda:0\")\nprint(device)\nprint(torch.cuda.get_device_name(device))","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:33.741498Z","iopub.execute_input":"2024-04-23T01:53:33.741796Z","iopub.status.idle":"2024-04-23T01:53:33.818296Z","shell.execute_reply.started":"2024-04-23T01:53:33.741771Z","shell.execute_reply":"2024-04-23T01:53:33.817148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Defining a function train_model with parameters: model, criterion, optimizer, scheduler, n_epochs\nto train model","metadata":{}},{"cell_type":"code","source":"def train_model(model, criterion, optimizer, scheduler, n_epochs = 5):\n    \n    losses = []\n    accuracies = []\n    test_accuracies = []\n    # set the model to train mode initially\n    model.train()\n    for epoch in tqdm(range(n_epochs)):\n        since = time.time()\n        running_loss = 0.0\n        running_correct = 0.0\n        for data in train_loader:\n\n            # get the inputs and assign them to cuda\n            inputs, labels = data\n            inputs = inputs.to(device)\n            labels = labels.to(device)\n            optimizer.zero_grad() # Clears the gradients of all optimized tensors. \n            \n            # forward + backward + optimize\n            # output contains probability of each class\n            outputs = model(inputs)\n            _, predicted = torch.max(outputs.data, 1) # select max prob - label\n            loss = criterion(outputs, labels) # compute loss\n            loss.backward() # compute gradients of loss\n            optimizer.step() # update parameters based on gradients\n            \n            # calculate the loss/acc later\n            running_loss += loss.item() # accumulates loss across batch\n            running_correct += (labels==predicted).sum().item() #no of correct predictions\n\n        epoch_duration = time.time()-since\n        epoch_loss = running_loss/len(train_loader)\n        epoch_acc = 100/32*running_correct/len(train_loader)\n\n        print(\"Epoch %s, duration: %d s, loss: %.4f, acc: %.4f\" % (epoch+1, epoch_duration, epoch_loss, epoch_acc))\n        \n        losses.append(epoch_loss)\n        accuracies.append(epoch_acc)\n        \n        # switch the model to eval mode to evaluate on test data\n        model.eval() # dropout layers and batch normalization layers are disabled\n        test_acc = eval_model(model)\n        test_accuracies.append(test_acc)\n        \n        # re-set the model to train mode after validating\n        model.train()\n        scheduler.step(test_acc) # adjust learining rates based on test_acc\n        since = time.time()\n    print('Finished Training')\n    return model, losses, accuracies, test_accuracies","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:33.819683Z","iopub.execute_input":"2024-04-23T01:53:33.820074Z","iopub.status.idle":"2024-04-23T01:53:33.831946Z","shell.execute_reply.started":"2024-04-23T01:53:33.820039Z","shell.execute_reply":"2024-04-23T01:53:33.830986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def eval_model(model):\n    correct = 0.0\n    total = 0.0\n    with torch.no_grad(): # disables gradient calculation during evaluation, no parameter updation \n        for i, data in enumerate(test_loader, 0):\n            images, labels = data\n            images = images.to(device)\n            labels = labels.to(device)\n            \n            outputs = model_ft(images)\n            _, predicted = torch.max(outputs.data, 1)\n            \n            total += labels.size(0)\n            correct += (predicted == labels).sum().item()\n\n    test_acc = 100.0 * correct / total\n    print('Accuracy of the network on the test images: %d %%' % (\n        test_acc))\n    return test_acc","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:33.833056Z","iopub.execute_input":"2024-04-23T01:53:33.833319Z","iopub.status.idle":"2024-04-23T01:53:33.848721Z","shell.execute_reply.started":"2024-04-23T01:53:33.833287Z","shell.execute_reply":"2024-04-23T01:53:33.847860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## ResNet-50 \nResNet-50 is a 50-layer convolutional neural network (48 convolutional layers, one MaxPool layer, and one average pool layer). ","metadata":{}},{"cell_type":"markdown","source":"Loads the `pre-trained ResNet-50 model` from torchvision's models module. The pretrained=True argument indicates that the model should be initialized with weights pre-trained on the ImageNet dataset.","metadata":{}},{"cell_type":"code","source":"model_ft = models.resnet50(pretrained=True)\nnum_ftrs = model_ft.fc.in_features # retrieves number of input features\n\nmodel_ft.fc = nn.Linear(num_ftrs, 10) #add a new layer for no. of output classes = 10\nmodel_ft = model_ft.to(device) # move model to device\n\ncriterion = nn.CrossEntropyLoss() # loss function used Cross entropy loss\n\n# optimizer is Stochastic Gradient Descent (SGD) with learning rate = 0.01 and momentum = 0.9 to accelerate convergence\noptimizer = torch.optim.SGD(model_ft.parameters(), lr=0.01, momentum=0.9)\n\n# define learning rate scheduler that reduce learning rate after 3 epochs if there is no improvement,by maximum 0.9 i.e threshold\nlrscheduler = torch.optim.lr_scheduler.ReduceLROnPlateau(optimizer, mode='max', patience=3, threshold = 0.9)","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:33.849854Z","iopub.execute_input":"2024-04-23T01:53:33.850133Z","iopub.status.idle":"2024-04-23T01:53:35.484267Z","shell.execute_reply.started":"2024-04-23T01:53:33.850102Z","shell.execute_reply":"2024-04-23T01:53:35.482804Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.cuda.empty_cache()","metadata":{"execution":{"iopub.status.busy":"2024-04-22T14:56:59.904933Z","iopub.execute_input":"2024-04-22T14:56:59.905789Z","iopub.status.idle":"2024-04-22T14:56:59.909696Z","shell.execute_reply.started":"2024-04-22T14:56:59.905755Z","shell.execute_reply":"2024-04-22T14:56:59.908778Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model_ft, training_losses, training_accs, test_accs = train_model(model_ft, criterion, optimizer, lrscheduler, n_epochs=3)","metadata":{"execution":{"iopub.status.busy":"2024-04-23T01:53:35.485674Z","iopub.execute_input":"2024-04-23T01:53:35.486062Z","iopub.status.idle":"2024-04-23T02:37:34.803404Z","shell.execute_reply.started":"2024-04-23T01:53:35.486026Z","shell.execute_reply":"2024-04-23T02:37:34.802455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.title('Training losses')\nplt.plot(training_losses)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:37:37.641316Z","iopub.execute_input":"2024-04-23T02:37:37.641706Z","iopub.status.idle":"2024-04-23T02:37:37.869797Z","shell.execute_reply.started":"2024-04-23T02:37:37.641668Z","shell.execute_reply":"2024-04-23T02:37:37.868925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.title('Training Accuracy')\nplt.plot(training_accs)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:37:37.871221Z","iopub.execute_input":"2024-04-23T02:37:37.871606Z","iopub.status.idle":"2024-04-23T02:37:38.171492Z","shell.execute_reply.started":"2024-04-23T02:37:37.871569Z","shell.execute_reply":"2024-04-23T02:37:38.170545Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.title('Testing Accuracy')\nplt.plot(test_accs)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:37:38.172931Z","iopub.execute_input":"2024-04-23T02:37:38.173549Z","iopub.status.idle":"2024-04-23T02:37:38.460615Z","shell.execute_reply.started":"2024-04-23T02:37:38.173514Z","shell.execute_reply":"2024-04-23T02:37:38.459650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"torch.save(model_ft.state_dict(), \"/kaggle/working/model-driver\")","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:37:34.804763Z","iopub.execute_input":"2024-04-23T02:37:34.805149Z","iopub.status.idle":"2024-04-23T02:37:34.974589Z","shell.execute_reply.started":"2024-04-23T02:37:34.805114Z","shell.execute_reply":"2024-04-23T02:37:34.973741Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model = models.resnet50()\nnum_ftrs = model.fc.in_features\nmodel.fc = nn.Linear(num_ftrs, 10)\nmodel.load_state_dict(torch.load(\"/kaggle/working/model-driver\"))\nmodel.eval()\nmodel.cuda()","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-04-23T02:37:34.977931Z","iopub.execute_input":"2024-04-23T02:37:34.978224Z","iopub.status.idle":"2024-04-23T02:37:35.592031Z","shell.execute_reply.started":"2024-04-23T02:37:34.978198Z","shell.execute_reply":"2024-04-23T02:37:35.591048Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"path_test = \"/kaggle/input/state-farm-distracted-driver-detection/imgs/test\"\nlist_img_test = [img for img in os.listdir(path_test) if not img.startswith(\".\")]\nlist_img_test.sort()","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:37:35.593336Z","iopub.execute_input":"2024-04-23T02:37:35.593702Z","iopub.status.idle":"2024-04-23T02:37:37.639802Z","shell.execute_reply.started":"2024-04-23T02:37:35.593670Z","shell.execute_reply":"2024-04-23T02:37:37.638556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"file = random.choice(list_img_test)\nim_path = os.path.join(path_test,file)\ndisplay(Image(filename=im_path))\nwith PIL.Image.open(im_path) as im:\n    im = transform(im)\n    im = im.unsqueeze(0) # Adds a batch dimension to the image tensor\n    output = model(im.cuda())\n    proba = nn.Softmax(dim=1)(output) # to convert output into probability\n    proba = [round(float(elem),4) for elem in proba[0]]\n    print(proba)\n    print(\"Predicted class:\",class_dict[proba.index(max(proba))])\n    print(\"Confidence:\",max(proba))","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:37:50.028989Z","iopub.execute_input":"2024-04-23T02:37:50.029656Z","iopub.status.idle":"2024-04-23T02:37:50.140101Z","shell.execute_reply.started":"2024-04-23T02:37:50.029620Z","shell.execute_reply":"2024-04-23T02:37:50.139023Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"list_img_test = os.listdir(path_test)\nselected_files = random.sample(list_img_test, 10)\nfor file in selected_files:\n    im_path = os.path.join(path_test, file)\n    display(Image(filename=im_path))\n    \n    with PIL.Image.open(im_path) as im:\n        im = transform(im)\n        im = im.unsqueeze(0)  # Adds a batch dimension to the image tensor\n        im = im.cuda()  # Move tensor to GPU\n        output = model(im)\n        proba = nn.Softmax(dim=1)(output)  # Convert output into probability\n        proba = [round(float(elem), 4) for elem in proba[0]]\n        print(proba)\n        print(\"Predicted class:\", class_dict[proba.index(max(proba))])\n        print(\"Confidence:\", max(proba))","metadata":{"execution":{"iopub.status.busy":"2024-04-23T02:39:46.825247Z","iopub.execute_input":"2024-04-23T02:39:46.825629Z","iopub.status.idle":"2024-04-23T02:39:47.255836Z","shell.execute_reply.started":"2024-04-23T02:39:46.825599Z","shell.execute_reply":"2024-04-23T02:39:47.254785Z"},"trusted":true},"execution_count":null,"outputs":[]}]}