{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-07-11T18:16:10.869879Z","iopub.execute_input":"2022-07-11T18:16:10.870437Z","iopub.status.idle":"2022-07-11T18:16:10.882920Z","shell.execute_reply.started":"2022-07-11T18:16:10.870381Z","shell.execute_reply":"2022-07-11T18:16:10.881578Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1 style=\"font-family:arial;\"> <center> Using CNN to predict dog or cat in any image </center> </h1>\n<h3><center>The main goal of this tutorial is to develop a system that can identify images of cats and dogs.</h3></center>","metadata":{}},{"cell_type":"markdown","source":"# Introduction\n\nIn this notebook, we will learn how to classify images of dog and cat by building a simple Neural Network. What we will learn:\n\n- Load the image\n- Visualize the Data distribution of all data\n- Visualizing some of the images\n- How CNN works\n- Use Image Data Generator\n- Graph the training loss and validation loss\n- Model  Evaluation\n- Confusion Matrix","metadata":{}},{"cell_type":"code","source":"# Importing the libraries\n\nimport zipfile\nimport os\nimport tensorflow as tf\nfrom tensorflow import keras\nimport matplotlib.pyplot as plt\nfrom sklearn.metrics import confusion_matrix\nimport plotly.graph_objects as go\nimport plotly.express as px\nimport matplotlib.image as img\nfrom keras.models import Sequential\nfrom keras.layers import Dense, Dropout, Flatten, Conv2D, MaxPooling2D\nfrom sklearn.model_selection import train_test_split\nfrom tensorflow.keras.preprocessing.image import ImageDataGenerator\n%matplotlib inline","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:15:56.554853Z","iopub.execute_input":"2022-07-11T18:15:56.555413Z","iopub.status.idle":"2022-07-11T18:16:10.866366Z","shell.execute_reply.started":"2022-07-11T18:15:56.555365Z","shell.execute_reply":"2022-07-11T18:16:10.865025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"'''setting seed'''\nseed = 0\nnp.random.seed(seed)\ntf.random.set_seed(3)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:16:10.886808Z","iopub.execute_input":"2022-07-11T18:16:10.888103Z","iopub.status.idle":"2022-07-11T18:16:10.913240Z","shell.execute_reply.started":"2022-07-11T18:16:10.888004Z","shell.execute_reply":"2022-07-11T18:16:10.911827Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import zipfile\n\nzip_files = ['test1', 'train']\n\nfor zip_file in zip_files:\n    with zipfile.ZipFile(\"../input/dogs-vs-cats/{}.zip\".format(zip_file),\"r\") as z:\n        z.extractall(\".\")\n        print(\"{} unzipped\".format(zip_file))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:14:50.208285Z","iopub.execute_input":"2022-07-11T18:14:50.208809Z","iopub.status.idle":"2022-07-11T18:15:14.331010Z","shell.execute_reply.started":"2022-07-11T18:14:50.208777Z","shell.execute_reply":"2022-07-11T18:15:14.329178Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(os.listdir('../working'))","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:16:10.915731Z","iopub.execute_input":"2022-07-11T18:16:10.916150Z","iopub.status.idle":"2022-07-11T18:16:10.935669Z","shell.execute_reply.started":"2022-07-11T18:16:10.916099Z","shell.execute_reply":"2022-07-11T18:16:10.934301Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing Data","metadata":{}},{"cell_type":"code","source":"IMAGE_FOLDER_PATH = \"../working/train\"\nFILE_NAMES = os.listdir(IMAGE_FOLDER_PATH)\nWIDTH = 150\nHEIGHT = 150","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:16:55.641813Z","iopub.execute_input":"2022-07-11T18:16:55.642273Z","iopub.status.idle":"2022-07-11T18:16:55.663360Z","shell.execute_reply.started":"2022-07-11T18:16:55.642236Z","shell.execute_reply":"2022-07-11T18:16:55.662083Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"FILE_NAMES[0:5]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:17:04.653094Z","iopub.execute_input":"2022-07-11T18:17:04.653490Z","iopub.status.idle":"2022-07-11T18:17:04.663862Z","shell.execute_reply.started":"2022-07-11T18:17:04.653455Z","shell.execute_reply":"2022-07-11T18:17:04.662507Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"labels = []\nfor i in os.listdir(IMAGE_FOLDER_PATH):\n    labels+=[i]","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:17:23.907859Z","iopub.execute_input":"2022-07-11T18:17:23.908344Z","iopub.status.idle":"2022-07-11T18:17:23.937122Z","shell.execute_reply.started":"2022-07-11T18:17:23.908308Z","shell.execute_reply":"2022-07-11T18:17:23.935719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data distribuation of all data","metadata":{}},{"cell_type":"code","source":"# empty list\ntargets = list()\nfull_paths = list()\ntrain_cats_dir = list()\ntrain_dogs_dir = list()\n\n# finding each file's target\nfor file_name in FILE_NAMES:\n    target = file_name.split(\".\")[0] # target name\n    full_path = os.path.join(IMAGE_FOLDER_PATH, file_name)\n    \n    if(target == \"dog\"):\n        train_dogs_dir.append(full_path)\n    if(target == \"cat\"):\n        train_cats_dir.append(full_path)\n    \n    full_paths.append(full_path)\n    targets.append(target)\n\ndataset = pd.DataFrame() # make dataframe\ndataset['image_path'] = full_paths # file path\ndataset['target'] = targets # file's target","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:19:40.455997Z","iopub.execute_input":"2022-07-11T18:19:40.456660Z","iopub.status.idle":"2022-07-11T18:19:40.578460Z","shell.execute_reply.started":"2022-07-11T18:19:40.456602Z","shell.execute_reply":"2022-07-11T18:19:40.577394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"dataset","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:21:21.309590Z","iopub.execute_input":"2022-07-11T18:21:21.310042Z","iopub.status.idle":"2022-07-11T18:21:21.343637Z","shell.execute_reply.started":"2022-07-11T18:21:21.310003Z","shell.execute_reply":"2022-07-11T18:21:21.342347Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"total data counts:\", dataset['target'].count())\ncounts = dataset['target'].value_counts()\nprint(counts)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:25:05.545047Z","iopub.execute_input":"2022-07-11T18:25:05.546578Z","iopub.status.idle":"2022-07-11T18:25:05.573353Z","shell.execute_reply.started":"2022-07-11T18:25:05.546507Z","shell.execute_reply":"2022-07-11T18:25:05.572128Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA","metadata":{}},{"cell_type":"code","source":"fig = go.Figure(go.Bar(\n            x= counts.values,\n            y=counts.index,\n            orientation='h'))\n\nfig.update_layout(title='Data Distribution in Bars',font_size=15,title_x=0.45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:36:11.407885Z","iopub.execute_input":"2022-07-11T18:36:11.408620Z","iopub.status.idle":"2022-07-11T18:36:11.557678Z","shell.execute_reply.started":"2022-07-11T18:36:11.408560Z","shell.execute_reply":"2022-07-11T18:36:11.556815Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig=px.pie(counts.head(10),values= 'target', names=dataset['target'].unique(),hole=0.425)\nfig.update_layout(title='Data Distribution of Data',font_size=15,title_x=0.45,annotations=[dict(text='Cat vs Dog',font_size=18, showarrow=False,height=800,width=700)])\nfig.update_traces(textfont_size=15,textinfo='percent')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:37:10.797761Z","iopub.execute_input":"2022-07-11T18:37:10.798269Z","iopub.status.idle":"2022-07-11T18:37:12.320481Z","shell.execute_reply.started":"2022-07-11T18:37:10.798228Z","shell.execute_reply":"2022-07-11T18:37:12.319042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Displaying Images of Cat","metadata":{}},{"cell_type":"code","source":"rows = 4\ncols = 4\naxes = []\nfig=plt.figure(figsize=(10,10))\ni = 0\n\nfor a in range(rows*cols):\n    b = img.imread(train_cats_dir[i])\n    axes.append(fig.add_subplot(rows,cols,a+1))\n    plt.imshow(b)\n    i+=1\nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:38:52.088097Z","iopub.execute_input":"2022-07-11T18:38:52.088613Z","iopub.status.idle":"2022-07-11T18:38:54.549409Z","shell.execute_reply.started":"2022-07-11T18:38:52.088571Z","shell.execute_reply":"2022-07-11T18:38:54.548168Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Displaying Images of Dog","metadata":{}},{"cell_type":"code","source":"rows = 4\ncols = 4\naxes = []\nfig=plt.figure(figsize=(10,10))\ni = 0\n\nfor a in range(rows*cols):\n    b = img.imread(train_dogs_dir[i])\n    axes.append(fig.add_subplot(rows,cols,a+1))\n    plt.imshow(b)\n    i+=1\nfig.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:39:30.666275Z","iopub.execute_input":"2022-07-11T18:39:30.666678Z","iopub.status.idle":"2022-07-11T18:39:32.915597Z","shell.execute_reply.started":"2022-07-11T18:39:30.666648Z","shell.execute_reply":"2022-07-11T18:39:32.914298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Splitting the data into train and test data","metadata":{}},{"cell_type":"code","source":"dataset_train, dataset_test = train_test_split(dataset, test_size=0.2, random_state=seed)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:39:57.665743Z","iopub.execute_input":"2022-07-11T18:39:57.666227Z","iopub.status.idle":"2022-07-11T18:39:57.679388Z","shell.execute_reply.started":"2022-07-11T18:39:57.666187Z","shell.execute_reply":"2022-07-11T18:39:57.678241Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Distribution of Train Data ","metadata":{}},{"cell_type":"code","source":"class_id_distributionTrain = dataset_train['target'].value_counts()\nclass_id_distributionTrain","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:41:32.214486Z","iopub.execute_input":"2022-07-11T18:41:32.215014Z","iopub.status.idle":"2022-07-11T18:41:32.230113Z","shell.execute_reply.started":"2022-07-11T18:41:32.214972Z","shell.execute_reply":"2022-07-11T18:41:32.229174Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure(go.Bar(\n            x=class_id_distributionTrain.values,\n            y=class_id_distributionTrain.index,\n            orientation='h'))\n\nfig.update_layout(title='Data Distribution Of Train Data in Bars',font_size=15,title_x=0.45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:42:05.588172Z","iopub.execute_input":"2022-07-11T18:42:05.588591Z","iopub.status.idle":"2022-07-11T18:42:05.604875Z","shell.execute_reply.started":"2022-07-11T18:42:05.588557Z","shell.execute_reply":"2022-07-11T18:42:05.604051Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig=px.pie(class_id_distributionTrain.head(10),values= 'target', names=dataset_train['target'].unique(),hole=0.425)\nfig.update_layout(title='Data Distribution of Train Data in Pie Chart',font_size=15,title_x=0.45,annotations=[dict(text='Cat vs Dog',font_size=18, showarrow=False,height=800,width=700)])\nfig.update_traces(textfont_size=15,textinfo='percent')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:42:14.649193Z","iopub.execute_input":"2022-07-11T18:42:14.649661Z","iopub.status.idle":"2022-07-11T18:42:14.728256Z","shell.execute_reply.started":"2022-07-11T18:42:14.649623Z","shell.execute_reply":"2022-07-11T18:42:14.725761Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Data Distribution of Test data","metadata":{}},{"cell_type":"code","source":"class_id_distributionTest = dataset_test['target'].value_counts()\nclass_id_distributionTest","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:42:51.913635Z","iopub.execute_input":"2022-07-11T18:42:51.914046Z","iopub.status.idle":"2022-07-11T18:42:51.927583Z","shell.execute_reply.started":"2022-07-11T18:42:51.914004Z","shell.execute_reply":"2022-07-11T18:42:51.926148Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig = go.Figure(go.Bar(\n            x=class_id_distributionTest.values,\n            y=class_id_distributionTest.index,\n            orientation='h'))\n\nfig.update_layout(title='Data Distribution Of Train Data in Bars',font_size=15,title_x=0.45)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:43:07.464081Z","iopub.execute_input":"2022-07-11T18:43:07.464910Z","iopub.status.idle":"2022-07-11T18:43:07.485206Z","shell.execute_reply.started":"2022-07-11T18:43:07.464814Z","shell.execute_reply":"2022-07-11T18:43:07.483828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"fig=px.pie(class_id_distributionTrain.head(10),values= 'target', names=dataset_train['target'].unique(),hole=0.425)\nfig.update_layout(title='Data Distribution of Train Data in Pie Chart',font_size=15,title_x=0.45,annotations=[dict(text='Cat vs Dog',font_size=18, showarrow=False,height=800,width=700)])\nfig.update_traces(textfont_size=15,textinfo='percent')\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:43:16.420771Z","iopub.execute_input":"2022-07-11T18:43:16.421288Z","iopub.status.idle":"2022-07-11T18:43:16.494247Z","shell.execute_reply.started":"2022-07-11T18:43:16.421249Z","shell.execute_reply":"2022-07-11T18:43:16.493181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Now, we will use ImageDataGenerator\nWhen there is little data to train, we have to use ImageDataGenerator to increase the number of data.\n- rescale = 1./255 : change the value between 0 and 1\n- rotation_range = 15 : Random rotation within 15 degrees\n- shear_range = 0.1 : shear range 10%\n- zoom_range = 0.2 : zoom range 20%\n- horizontal_flip = True : Randomly flip horizontally\n- width_shift_range = 0.1 : Randomly move the original image horizontally within 10% of the width\n- height_shift_range=0.1 : Randomly move the original image vertically within 10% of the width","metadata":{}},{"cell_type":"code","source":"train_datagen=ImageDataGenerator(\nrotation_range=15,\nrescale=1./255,\nshear_range=0.1,\nzoom_range=0.2,\nhorizontal_flip=True,\nwidth_shift_range=0.1,\nheight_shift_range=0.1)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:46:43.259772Z","iopub.execute_input":"2022-07-11T18:46:43.260284Z","iopub.status.idle":"2022-07-11T18:46:43.267418Z","shell.execute_reply.started":"2022-07-11T18:46:43.260246Z","shell.execute_reply":"2022-07-11T18:46:43.266077Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_datagenerator=train_datagen.flow_from_dataframe(dataframe=dataset_train,\n                                                     x_col=\"image_path\",\n                                                     y_col=\"target\",\n                                                     target_size=(WIDTH, HEIGHT),\n                                                     class_mode=\"binary\",\n                                                     batch_size=150)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:46:48.395613Z","iopub.execute_input":"2022-07-11T18:46:48.396336Z","iopub.status.idle":"2022-07-11T18:46:48.729359Z","shell.execute_reply.started":"2022-07-11T18:46:48.396290Z","shell.execute_reply":"2022-07-11T18:46:48.727893Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_datagen = ImageDataGenerator(rescale=1./255)\ntest_datagenerator=test_datagen.flow_from_dataframe(dataframe=dataset_test,\n                                                   x_col=\"image_path\",\n                                                   y_col=\"target\",\n                                                   target_size=(WIDTH, HEIGHT),\n                                                   class_mode=\"binary\",\n                                                   batch_size=150)","metadata":{"execution":{"iopub.status.busy":"2022-07-11T18:46:56.917018Z","iopub.execute_input":"2022-07-11T18:46:56.917445Z","iopub.status.idle":"2022-07-11T18:46:56.996050Z","shell.execute_reply.started":"2022-07-11T18:46:56.917410Z","shell.execute_reply":"2022-07-11T18:46:56.994556Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Making CNN Model","metadata":{}},{"cell_type":"markdown","source":"![CNN](https://raw.githubusercontent.com/AayushSaxena08/Image_Classification_using_Transfer_Learning/main/Cats_Dogs_CNN.gif)","metadata":{}},{"cell_type":"markdown","source":"### The Convolutional Neural Network is made up of 3 parts:\n- Convolutional Layer\n- Pooling Layer\n- Fully Connected Layer","metadata":{}},{"cell_type":"markdown","source":"![CNN Architecture](https://github.com/AayushSaxena08/Image_Classification_using_Transfer_Learning/blob/main/cnn1.jpeg?raw=true)","metadata":{}},{"cell_type":"markdown","source":"# 1. Convolutional layer\nA convolutional layer helps to extract the information from the image with the help of filter(kernal).Please have a look the following image.\n![CNN](https://github.com/AayushSaxena08/Image_Classification_using_Transfer_Learning/blob/main/Filter.gif?raw=true)","metadata":{}},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}