{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.6.6","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"gpu","dataSources":[{"sourceId":16880,"databundleVersionId":858837,"sourceType":"competition"},{"sourceId":924245,"sourceType":"datasetVersion","datasetId":464091}],"dockerImageVersionId":29845,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":true}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Title","metadata":{"id":"Q6tasuafvT2O"}},{"cell_type":"markdown","source":"Deep Fake Image and Video Detection using CNN's and RNN's","metadata":{"id":"LvJBeHv3vfE_"}},{"cell_type":"markdown","source":"# Problem Statement","metadata":{"id":"xyukrIAev64Y"}},{"cell_type":"markdown","source":"DeepFake is composed from Deep Learning and Fake and means taking one person from an image or video and replacing with someone else likeness using technology such as Deep Artificial Neural Networks. Large companies like Google invest very much in fighting the DeepFake, this including release of large datasets to help training models to counter this threat.The phenomen invades rapidly the film industry and threatens to compromise news agencies. Large digital companies, including content providers and social platforms are in the frontrun of fighting Deep Fakes. GANs that generate DeepFakes becomes better every day and, of course, if you include in a new GAN model all the information we collected until now how to combat various existent models, we create a model that cannot be beatten by the existing ones.\n\nFirst we will work on detecting faces that were forged and we will work on developing a model to detect videos.","metadata":{"id":"MZsX2Wrrv9Op"}},{"cell_type":"markdown","source":"# DeepFake Image detection","metadata":{"id":"vGkKRV1kY8Z-"}},{"cell_type":"markdown","source":"## About the Dataset","metadata":{"id":"ZLLkoLT1wsPL"}},{"cell_type":"markdown","source":"This dataset contains faces extracted from deepfake-detection-challenge. All images were of size 224x224. \n\nDue to memory issue we will only use a sample of the entire dataset for prediction.","metadata":{"id":"y_3tT_2lwv-v"}},{"cell_type":"markdown","source":"## Importing Required libraries","metadata":{"id":"dFXIv9qNpKzt","tags":[]}},{"cell_type":"code","source":"!pip install -U --upgrade tensorflow","metadata":{"execution":{"iopub.status.busy":"2024-02-24T01:26:13.884344Z","iopub.execute_input":"2024-02-24T01:26:13.884694Z","iopub.status.idle":"2024-02-24T01:27:39.288965Z","shell.execute_reply.started":"2024-02-24T01:26:13.884635Z","shell.execute_reply":"2024-02-24T01:27:39.287972Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import sys\nimport sklearn\nimport tensorflow as tf\n\nimport cv2\nimport pandas as pd\nimport numpy as np\n\nimport plotly.graph_objs as go\nfrom plotly.offline import iplot\nfrom matplotlib import pyplot as plt","metadata":{"id":"TFSU3FCOpKzu","execution":{"iopub.status.busy":"2024-02-24T01:28:01.573047Z","iopub.execute_input":"2024-02-24T01:28:01.573455Z","iopub.status.idle":"2024-02-24T01:28:01.579101Z","shell.execute_reply.started":"2024-02-24T01:28:01.573387Z","shell.execute_reply":"2024-02-24T01:28:01.578072Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.test.is_gpu_available()","metadata":{"execution":{"iopub.status.busy":"2024-02-24T01:28:04.357558Z","iopub.execute_input":"2024-02-24T01:28:04.357863Z","iopub.status.idle":"2024-02-24T01:28:04.511093Z","shell.execute_reply.started":"2024-02-24T01:28:04.35782Z","shell.execute_reply":"2024-02-24T01:28:04.510296Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.__version__","metadata":{"execution":{"iopub.status.busy":"2024-02-24T01:28:05.975172Z","iopub.execute_input":"2024-02-24T01:28:05.975585Z","iopub.status.idle":"2024-02-24T01:28:05.980808Z","shell.execute_reply.started":"2024-02-24T01:28:05.975522Z","shell.execute_reply":"2024-02-24T01:28:05.980047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import matplotlib.pyplot as plt\n\nplt.rc('font', size=14)\nplt.rc('axes', labelsize=14, titlesize=14)\nplt.rc('legend', fontsize=14)\nplt.rc('xtick', labelsize=10)\nplt.rc('ytick', labelsize=10)","metadata":{"id":"8d4TH3NbpKzx","execution":{"iopub.status.busy":"2024-02-24T01:28:07.362738Z","iopub.execute_input":"2024-02-24T01:28:07.363026Z","iopub.status.idle":"2024-02-24T01:28:07.368782Z","shell.execute_reply.started":"2024-02-24T01:28:07.362985Z","shell.execute_reply":"2024-02-24T01:28:07.367852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Visualisation","metadata":{"id":"NL3Ht4wC9b3n"}},{"cell_type":"code","source":"import os\n\ndef get_data():\n    return pd.read_csv('../input/deepfake-faces/metadata.csv')","metadata":{"id":"jfv9PxSB4tM8","execution":{"iopub.status.busy":"2024-02-24T01:28:11.618075Z","iopub.execute_input":"2024-02-24T01:28:11.618396Z","iopub.status.idle":"2024-02-24T01:28:11.622106Z","shell.execute_reply.started":"2024-02-24T01:28:11.618349Z","shell.execute_reply":"2024-02-24T01:28:11.621328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta=get_data()\nmeta.head()","metadata":{"id":"tDW7BRph9ehF","outputId":"97de18b5-0a37-4302-8804-8a16a7d2ed2f","execution":{"iopub.status.busy":"2024-02-24T01:28:13.409413Z","iopub.execute_input":"2024-02-24T01:28:13.409757Z","iopub.status.idle":"2024-02-24T01:28:13.588637Z","shell.execute_reply.started":"2024-02-24T01:28:13.409694Z","shell.execute_reply":"2024-02-24T01:28:13.587917Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"meta.shape","metadata":{"id":"n7FSdDifbZxn","outputId":"5451a127-405a-4c0b-a197-c920b796adbb","execution":{"iopub.status.busy":"2024-02-24T01:28:13.877869Z","iopub.execute_input":"2024-02-24T01:28:13.878193Z","iopub.status.idle":"2024-02-24T01:28:13.884044Z","shell.execute_reply.started":"2024-02-24T01:28:13.878132Z","shell.execute_reply":"2024-02-24T01:28:13.883304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"len(meta[meta.label=='FAKE']),len(meta[meta.label=='REAL'])","metadata":{"id":"_FJcz2IthxVG","outputId":"274c3f65-7acb-4f99-8aa9-a5b2a23bf06a","execution":{"iopub.status.busy":"2024-02-24T01:28:15.612443Z","iopub.execute_input":"2024-02-24T01:28:15.612774Z","iopub.status.idle":"2024-02-24T01:28:15.658146Z","shell.execute_reply.started":"2024-02-24T01:28:15.612717Z","shell.execute_reply":"2024-02-24T01:28:15.657252Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"real_df = meta[meta[\"label\"] == \"REAL\"]\nfake_df = meta[meta[\"label\"] == \"FAKE\"]\nsample_size = 8000\n\nreal_df = real_df.sample(sample_size, random_state=42)\nfake_df = fake_df.sample(sample_size, random_state=42)\n\nsample_meta = pd.concat([real_df, fake_df])","metadata":{"id":"IgMfzY-PjjtH","execution":{"iopub.status.busy":"2024-02-24T01:28:16.167117Z","iopub.execute_input":"2024-02-24T01:28:16.167446Z","iopub.status.idle":"2024-02-24T01:28:16.217628Z","shell.execute_reply.started":"2024-02-24T01:28:16.1674Z","shell.execute_reply":"2024-02-24T01:28:16.216872Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As mentioned instead of using 95k images we will only use 16000 images.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nTrain_set, Test_set = train_test_split(sample_meta,test_size=0.2,random_state=42,stratify=sample_meta['label'])\nTrain_set, Val_set  = train_test_split(Train_set,test_size=0.3,random_state=42,stratify=Train_set['label'])","metadata":{"id":"5eB86S6K-T5Z","execution":{"iopub.status.busy":"2024-02-24T01:28:26.110356Z","iopub.execute_input":"2024-02-24T01:28:26.110669Z","iopub.status.idle":"2024-02-24T01:28:26.745982Z","shell.execute_reply.started":"2024-02-24T01:28:26.110625Z","shell.execute_reply":"2024-02-24T01:28:26.745228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"Train_set.shape,Val_set.shape,Test_set.shape","metadata":{"id":"8p-TONijb4qA","outputId":"56d0b529-9d81-4019-d8fa-618c8cdba90f","execution":{"iopub.status.busy":"2024-02-24T01:28:27.095288Z","iopub.execute_input":"2024-02-24T01:28:27.095587Z","iopub.status.idle":"2024-02-24T01:28:27.1102Z","shell.execute_reply.started":"2024-02-24T01:28:27.095545Z","shell.execute_reply":"2024-02-24T01:28:27.109227Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y = dict()\n\ny[0] = []\ny[1] = []\n\nfor set_name in (np.array(Train_set['label']), np.array(Val_set['label']), np.array(Test_set['label'])):\n    y[0].append(np.sum(set_name == 'REAL'))\n    y[1].append(np.sum(set_name == 'FAKE'))\n\ntrace0 = go.Bar(\n    x=['Train Set', 'Validation Set', 'Test Set'],\n    y=y[0],\n    name='REAL',\n    marker=dict(color='#33cc33'),\n    opacity=0.7\n)\ntrace1 = go.Bar(\n    x=['Train Set', 'Validation Set', 'Test Set'],\n    y=y[1],\n    name='FAKE',\n    marker=dict(color='#ff3300'),\n    opacity=0.7\n)\n\ndata = [trace0, trace1]\nlayout = go.Layout(\n    title='Count of classes in each set',\n    xaxis={'title': 'Set'},\n    yaxis={'title': 'Count'}\n)\n\nfig = go.Figure(data, layout)\niplot(fig)","metadata":{"id":"hzNGtCWd-mTk","outputId":"5178c3ed-cbba-4f99-99d0-26bde11a5dab","execution":{"iopub.status.busy":"2024-02-24T01:28:27.729791Z","iopub.execute_input":"2024-02-24T01:28:27.730092Z","iopub.status.idle":"2024-02-24T01:28:28.917035Z","shell.execute_reply.started":"2024-02-24T01:28:27.730048Z","shell.execute_reply":"2024-02-24T01:28:28.916185Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The original image dataset were biased with more fake images than real since we are taking a sample of it its better to take equal proportion of real and fake images.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(15,15))\nfor cur,i in enumerate(Train_set.index[25:50]):\n    plt.subplot(5,5,cur+1)\n    plt.xticks([])\n    plt.yticks([])\n    plt.grid(False)\n    \n    plt.imshow(cv2.imread('../input/deepfake-faces/faces_224/'+Train_set.loc[i,'videoname'][:-4]+'.jpg'))\n    \n    if(Train_set.loc[i,'label']=='FAKE'):\n        plt.xlabel('FAKE Image')\n    else:\n        plt.xlabel('REAL Image')\n        \nplt.show()","metadata":{"id":"VR7Uly2fcUYi","outputId":"c1f47a82-ef4f-4bcd-b51c-d4738142fc0f","execution":{"iopub.status.busy":"2024-02-24T01:28:31.209162Z","iopub.execute_input":"2024-02-24T01:28:31.209498Z","iopub.status.idle":"2024-02-24T01:28:33.059514Z","shell.execute_reply.started":"2024-02-24T01:28:31.209454Z","shell.execute_reply":"2024-02-24T01:28:33.058819Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Modelling","metadata":{"id":"dOvN_divkl-N"}},{"cell_type":"markdown","source":"Before jumping to use pretrained model lets develop some base line model to test how our pretrained model outperforms.","metadata":{}},{"cell_type":"markdown","source":"### Custom CNN Architecture","metadata":{"id":"oid44Xx-pKz6"}},{"cell_type":"code","source":"def retreive_dataset(set_name):\n    images,labels=[],[]\n    for (img, imclass) in zip(set_name['videoname'], set_name['label']):\n        images.append(cv2.imread('../input/deepfake-faces/faces_224/'+img[:-4]+'.jpg'))\n        if(imclass=='FAKE'):\n            labels.append(1)\n        else:\n            labels.append(0)\n    \n    return np.array(images),np.array(labels)","metadata":{"id":"Hz0ZdQ_fgHhG","execution":{"iopub.status.busy":"2024-02-24T01:28:42.836033Z","iopub.execute_input":"2024-02-24T01:28:42.836408Z","iopub.status.idle":"2024-02-24T01:28:42.843336Z","shell.execute_reply.started":"2024-02-24T01:28:42.836345Z","shell.execute_reply":"2024-02-24T01:28:42.842444Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train,y_train=retreive_dataset(Train_set)\nX_val,y_val=retreive_dataset(Val_set)\nX_test,y_test=retreive_dataset(Test_set)","metadata":{"id":"zeAGRcAbguKU","execution":{"iopub.status.busy":"2024-02-24T01:28:51.107944Z","iopub.execute_input":"2024-02-24T01:28:51.108328Z","iopub.status.idle":"2024-02-24T01:30:17.228658Z","shell.execute_reply.started":"2024-02-24T01:28:51.108262Z","shell.execute_reply":"2024-02-24T01:30:17.227707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from functools import partial\n\ntf.random.set_seed(42) \nDefaultConv2D = partial(tf.keras.layers.Conv2D, kernel_size=3, padding=\"same\",\n                        activation=\"relu\", kernel_initializer=\"he_normal\")\n\nmodel = tf.keras.Sequential([\n    DefaultConv2D(filters=64, kernel_size=7, input_shape=[224, 224, 3]),\n    tf.keras.layers.MaxPool2D(),\n    DefaultConv2D(filters=128),\n    DefaultConv2D(filters=128),\n    tf.keras.layers.MaxPool2D(),\n    tf.keras.layers.Flatten(),\n    tf.keras.layers.Dense(units=128, activation=\"relu\",\n                          kernel_initializer=\"he_normal\"),\n    tf.keras.layers.Dropout(0.5),\n    tf.keras.layers.Dense(units=64, activation=\"relu\",\n                          kernel_initializer=\"he_normal\"),\n    tf.keras.layers.Dropout(0.5),\n    tf.keras.layers.Dense(units=1, activation=\"sigmoid\")\n])","metadata":{"id":"34upiak4pKz6","execution":{"iopub.status.busy":"2024-02-24T01:31:21.183153Z","iopub.execute_input":"2024-02-24T01:31:21.183575Z","iopub.status.idle":"2024-02-24T01:31:22.095962Z","shell.execute_reply.started":"2024-02-24T01:31:21.183519Z","shell.execute_reply":"2024-02-24T01:31:22.095329Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.compile(loss=\"binary_crossentropy\", optimizer=\"nadam\",\n              metrics=[\"accuracy\"])\nmodel.summary()","metadata":{"execution":{"iopub.status.busy":"2024-02-24T01:31:24.634365Z","iopub.execute_input":"2024-02-24T01:31:24.63468Z","iopub.status.idle":"2024-02-24T01:31:24.663547Z","shell.execute_reply.started":"2024-02-24T01:31:24.634636Z","shell.execute_reply":"2024-02-24T01:31:24.662351Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"history = model.fit(X_train, y_train, epochs=5,batch_size=64,\n                    validation_data=(X_val, y_val))","metadata":{"id":"KZbWeIBYpKz6","outputId":"deb6f56a-7b93-4241-a1bd-b210c0f2d426","execution":{"iopub.status.busy":"2024-02-24T01:31:27.237631Z","iopub.execute_input":"2024-02-24T01:31:27.237936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"score = model.evaluate(X_test, y_test)","metadata":{"id":"6HDDr4uehast","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot model performance\nacc = history.history['accuracy']\nval_acc = history.history['val_accuracy']\nloss = history.history['loss']\nval_loss = history.history['val_loss']\nepochs_range = range(1, len(history.epoch) + 1)\n\nplt.figure(figsize=(15,5))\n\nplt.subplot(1, 2, 1)\nplt.plot(epochs_range, acc, label='Train Set')\nplt.plot(epochs_range, val_acc, label='Val Set')\nplt.legend(loc=\"best\")\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Model Accuracy')\n\nplt.subplot(1, 2, 2)\nplt.plot(epochs_range, loss, label='Train Set')\nplt.plot(epochs_range, val_loss, label='Val Set')\nplt.legend(loc=\"best\")\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Model Loss')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"A baseline score of 50.06% is good to go let's finetune some pretrained model","metadata":{}},{"cell_type":"markdown","source":"# Pretrained Models for Transfer Learning","metadata":{"id":"hqxnSBJ3pKz8"}},{"cell_type":"markdown","source":"Here i used Xception model for fine-tuning feel free to try the performance of other pretrained models.","metadata":{}},{"cell_type":"markdown","source":"All three datasets contain individual images. We need to batch them, but for this we first need to ensure they all have the same size, or else batching will not work. We can use a `Resizing` layer for this. We must also call the `tf.keras.applications.xception.preprocess_input()` function to preprocess the images appropriately for the Xception model. We will also add shuffling and prefetching to the training dataset.","metadata":{"id":"gXG6iv8XpKz9"}},{"cell_type":"code","source":"train_set_raw=tf.data.Dataset.from_tensor_slices((X_train,y_train))\nvalid_set_raw=tf.data.Dataset.from_tensor_slices((X_val,y_val))\ntest_set_raw=tf.data.Dataset.from_tensor_slices((X_test,y_test))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"tf.keras.backend.clear_session()  # extra code – resets layer name counter\n\nbatch_size = 32\npreprocess = tf.keras.applications.xception.preprocess_input\ntrain_set = train_set_raw.map(lambda X, y: (preprocess(tf.cast(X, tf.float32)), y))\ntrain_set = train_set.shuffle(1000, seed=42).batch(batch_size).prefetch(1)\nvalid_set = valid_set_raw.map(lambda X, y: (preprocess(tf.cast(X, tf.float32)), y)).batch(batch_size)\ntest_set = test_set_raw.map(lambda X, y: (preprocess(tf.cast(X, tf.float32)), y)).batch(batch_size)","metadata":{"id":"Bnz0n9XApKz9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's take a look again at the first 9 images from the validation set: they're all with values ranging from -1 to 1:","metadata":{"id":"ovNEMky-pKz9"}},{"cell_type":"code","source":"# extra code – displays the first 9 images in the first batch of valid_set\n\nplt.figure(figsize=(12, 12))\nfor X_batch, y_batch in valid_set.take(1):\n    for index in range(9):\n        plt.subplot(3, 3, index + 1)\n        plt.imshow((X_batch[index] + 1) / 2)  # rescale to 0–1 for imshow()\n        if(y_batch[index]==1):\n            classt='FAKE'\n        else:\n            classt='REAL'\n        plt.title(f\"Class: {classt}\")\n        plt.axis(\"off\")\n\nplt.show()","metadata":{"id":"ZL3c3i4opKz9","outputId":"38847d8d-8822-41a3-cfb2-27479aa5debe","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data_augmentation = tf.keras.Sequential([\n    tf.keras.layers.RandomFlip(mode=\"horizontal\", seed=42),\n    tf.keras.layers.RandomRotation(factor=0.05, seed=42),\n    tf.keras.layers.RandomContrast(factor=0.2, seed=42)\n])","metadata":{"id":"Ib0cA8Y1pKz9","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Try running the following cell multiple times to see different random data augmentations:","metadata":{"id":"G7GrQjsspKz-"}},{"cell_type":"code","source":"# extra code – displays the same first 9 images, after augmentation\n\nplt.figure(figsize=(12, 12))\nfor X_batch, y_batch in valid_set.take(1):\n    X_batch_augmented = data_augmentation(X_batch, training=True)\n    for index in range(9):\n        plt.subplot(3, 3, index + 1)\n        # We must rescale the images to the 0-1 range for imshow(), and also\n        # clip the result to that range, because data augmentation may\n        # make some values go out of bounds (e.g., RandomContrast in this case).\n        plt.imshow(np.clip((X_batch_augmented[index] + 1) / 2, 0, 1))\n        if(y_batch[index]==1):\n            classt='FAKE'\n        else:\n            classt='REAL'\n        plt.title(f\"Class: {classt}\")\n        plt.axis(\"off\")\n\nplt.show()","metadata":{"id":"w6GH5_vupKz-","outputId":"eeb2c924-2f4f-4aa1-bea9-951bebef4bf0","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now let's load the pretrained model, without its top layers, and replace them with our own task","metadata":{"id":"kNL9AOsDpKz-"}},{"cell_type":"code","source":"tf.random.set_seed(42)  # extra code – ensures reproducibility\nbase_model = tf.keras.applications.xception.Xception(weights=\"imagenet\",\n                                                     include_top=False)\navg = tf.keras.layers.GlobalAveragePooling2D()(base_model.output)\noutput = tf.keras.layers.Dense(1, activation=\"sigmoid\")(avg)\nmodel = tf.keras.Model(inputs=base_model.input, outputs=output)","metadata":{"id":"lRyCgvaKpKz-","outputId":"a825e173-8b1d-4217-a1c4-5491b49c3e82","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for layer in base_model.layers:\n    layer.trainable = False","metadata":{"id":"KBlyG6ElpKz-","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's train the model for a few epochs, while keeping the base model weights fixed:","metadata":{"id":"WFEFw7GKpKz-"}},{"cell_type":"code","source":"optimizer = tf.keras.optimizers.SGD(learning_rate=0.1, momentum=0.9)\nmodel.compile(loss=\"binary_crossentropy\", optimizer=optimizer,\n              metrics=[\"accuracy\"])\nhistory = model.fit(train_set, validation_data=valid_set, epochs=3)","metadata":{"id":"GGxK2yPcpKz-","outputId":"6b64214a-e104-4b6c-9b7a-3388fc9aa15f","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for indices in zip(range(33), range(33, 66), range(66, 99), range(99, 132)):\n    for idx in indices:\n        print(f\"{idx:3}: {base_model.layers[idx].name:22}\", end=\"\")\n    print()","metadata":{"id":"GvGMiJMLpKz-","outputId":"91f2c96c-c058-45e0-e428-66fa6076ad56","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.evaluate(test_set)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now with the finetuning the top layers of xception model the model performance jumps to 63.8% ","metadata":{}},{"cell_type":"markdown","source":"Now that the weights of our new top layers are not too bad, we can make the top part of the base model trainable again, and continue training, but with a lower learning rate:","metadata":{"id":"L_bEwL8KpKz_"}},{"cell_type":"code","source":"for layer in base_model.layers[56:]:\n    layer.trainable = True\n\noptimizer = tf.keras.optimizers.SGD(learning_rate=0.01, momentum=0.9)\nmodel.compile(loss=\"binary_crossentropy\", optimizer=optimizer,\n              metrics=[\"accuracy\"])\nhistory = model.fit(train_set, validation_data=valid_set, epochs=10)","metadata":{"id":"GEUNGlhvpKz_","outputId":"c622a91d-f634-4443-b87e-8d46defdb578","trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# plot model performance\nacc = history.history['accuracy']\nval_acc = history.history['val_accuracy']\nloss = history.history['loss']\nval_loss = history.history['val_loss']\nepochs_range = range(1, len(history.epoch) + 1)\n\nplt.figure(figsize=(15,5))\n\nplt.subplot(1, 2, 1)\nplt.plot(epochs_range, acc, label='Train Set')\nplt.plot(epochs_range, val_acc, label='Val Set')\nplt.legend(loc=\"best\")\nplt.xlabel('Epochs')\nplt.ylabel('Accuracy')\nplt.title('Model Accuracy')\n\nplt.subplot(1, 2, 2)\nplt.plot(epochs_range, loss, label='Train Set')\nplt.plot(epochs_range, val_loss, label='Val Set')\nplt.legend(loc=\"best\")\nplt.xlabel('Epochs')\nplt.ylabel('Loss')\nplt.title('Model Loss')\n\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"model.evaluate(test_set)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The model accuracy finally reaches to 81.9%","metadata":{}},{"cell_type":"code","source":"model.save('xception_deepfake_image.h5')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Add Explainability to the model","metadata":{}},{"cell_type":"markdown","source":"Lets try to interpret the trained model on how it finds a image FAKE","metadata":{}},{"cell_type":"code","source":"!pip install lime","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from lime import lime_image\n\nexplainer = lime_image.LimeImageExplainer()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(12, 12))\n\nfor index in range(9):\n    plt.subplot(3, 3, index + 1)\n    plt.imshow((x[index] + 1) / 2)  # rescale to 0–1 for imshow()\n    if(y[index]==1):\n        classt='FAKE'\n    else:\n        classt='REAL'\n    plt.title(f\"Class: {classt}\")\n    plt.axis(\"off\")\n\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data=x[2,:,:,:]\ntest_data.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"explanation = explainer.explain_instance(test_data.astype('double'), model.predict,  \n                                         top_labels=3, hide_color=0, num_samples=1000)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from skimage.segmentation import mark_boundaries\n\ntemp_1, mask_1 = explanation.get_image_and_mask(explanation.top_labels[0], positive_only=True, num_features=5, hide_rest=True)\ntemp_2, mask_2 = explanation.get_image_and_mask(explanation.top_labels[0], positive_only=False, num_features=10, hide_rest=False)\n\nfig, (ax1, ax2) = plt.subplots(1, 2, figsize=(15,15))\nax1.imshow(mark_boundaries(temp_1, mask_1))\nax2.imshow(mark_boundaries(temp_2, mask_2))\nax1.axis('off')\nax2.axis('off')","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Deep Fake Video Classification","metadata":{}},{"cell_type":"markdown","source":"## About the dataset","metadata":{}},{"cell_type":"markdown","source":"Files\n\n* train_sample_videos.zip - a ZIP file containing a sample set of training videos and a metadata.json with labels. the full set of training videos is available through the links provided above.\n* sample_submission.csv - a sample submission file in the correct format.\n* test_videos.zip - a zip file containing a small set of videos to be used as a public validation set. To understand the datasets available for this competition, review the Getting Started information.\n\nMetadata Columns\n\n* filename - the filename of the video\n* label - whether the video is REAL or FAKE\n* original - in the case that a train set video is FAKE, the original video is listed here\n* split - this is always equal to \"train\".","metadata":{}},{"cell_type":"markdown","source":"## Importing required libraries","metadata":{}},{"cell_type":"code","source":"!pip install -U --upgrade tensorflow","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from tensorflow_docs.vis import embed\nfrom tensorflow import keras\n#from imutils import paths\n\nimport matplotlib.pyplot as plt\nimport tensorflow as tf\nimport pandas as pd\nimport numpy as np\nimport imageio\nimport cv2\nimport os","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Data Visualisation","metadata":{}},{"cell_type":"code","source":"DATA_FOLDER = '../input/deepfake-detection-challenge'\nTRAIN_SAMPLE_FOLDER = 'train_sample_videos'\nTEST_FOLDER = 'test_videos'\n\nprint(f\"Train samples: {len(os.listdir(os.path.join(DATA_FOLDER, TRAIN_SAMPLE_FOLDER)))}\")\nprint(f\"Test samples: {len(os.listdir(os.path.join(DATA_FOLDER, TEST_FOLDER)))}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample_metadata = pd.read_json('../input/deepfake-detection-challenge/train_sample_videos/metadata.json').T\ntrain_sample_metadata.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample_metadata.groupby('label')['label'].count().plot(figsize=(15, 5), kind='bar', title='Distribution of Labels in the Training Set')\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_sample_metadata.shape","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's visualize now the data.\n\nWe select first a list of fake videos.","metadata":{}},{"cell_type":"markdown","source":"### Few fake videos","metadata":{}},{"cell_type":"code","source":"fake_train_sample_video = list(train_sample_metadata.loc[train_sample_metadata.label=='FAKE'].sample(3).index)\nfake_train_sample_video","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def display_image_from_video(video_path):\n    '''\n    input: video_path - path for video\n    process:\n    1. perform a video capture from the video\n    2. read the image\n    3. display the image\n    '''\n    capture_image = cv2.VideoCapture(video_path) \n    ret, frame = capture_image.read()\n    fig = plt.figure(figsize=(10,10))\n    ax = fig.add_subplot(111)\n    frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n    ax.imshow(frame)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for video_file in fake_train_sample_video:\n    display_image_from_video(os.path.join(DATA_FOLDER, TRAIN_SAMPLE_FOLDER, video_file))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's try now the same for few of the images that are real.","metadata":{}},{"cell_type":"markdown","source":"### Few Real Videos","metadata":{}},{"cell_type":"code","source":"real_train_sample_video = list(train_sample_metadata.loc[train_sample_metadata.label=='REAL'].sample(3).index)\nreal_train_sample_video","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for video_file in real_train_sample_video:\n    display_image_from_video(os.path.join(DATA_FOLDER, TRAIN_SAMPLE_FOLDER, video_file))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Videos with same original","metadata":{}},{"cell_type":"markdown","source":"Let's look now to set of samples with the same original.","metadata":{}},{"cell_type":"code","source":"train_sample_metadata['original'].value_counts()[0:5]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We pick one of the originals with largest number of samples.\n\nWe also modify our visualization function to work with multiple images.","metadata":{}},{"cell_type":"code","source":"def display_image_from_video_list(video_path_list, video_folder=TRAIN_SAMPLE_FOLDER):\n    '''\n    input: video_path_list - path for video\n    process:\n    0. for each video in the video path list\n        1. perform a video capture from the video\n        2. read the image\n        3. display the image\n    '''\n    plt.figure()\n    fig, ax = plt.subplots(2,3,figsize=(16,8))\n    # we only show images extracted from the first 6 videos\n    for i, video_file in enumerate(video_path_list[0:6]):\n        video_path = os.path.join(DATA_FOLDER, video_folder,video_file)\n        capture_image = cv2.VideoCapture(video_path) \n        ret, frame = capture_image.read()\n        frame = cv2.cvtColor(frame, cv2.COLOR_BGR2RGB)\n        ax[i//3, i%3].imshow(frame)\n        ax[i//3, i%3].set_title(f\"Video: {video_file}\")\n        ax[i//3, i%3].axis('on')\n","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"same_original_fake_train_sample_video = list(train_sample_metadata.loc[train_sample_metadata.original=='atvmxvwyns.mp4'].index)\ndisplay_image_from_video_list(same_original_fake_train_sample_video)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Test video files","metadata":{}},{"cell_type":"markdown","source":"Let's also look to few of the test data files.","metadata":{}},{"cell_type":"code","source":"test_videos = pd.DataFrame(list(os.listdir(os.path.join(DATA_FOLDER, TEST_FOLDER))), columns=['video'])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_videos.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Let's visualize now one of the videos.","metadata":{}},{"cell_type":"code","source":"display_image_from_video(os.path.join(DATA_FOLDER, TEST_FOLDER, test_videos.iloc[2].video))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Play video files","metadata":{}},{"cell_type":"markdown","source":"Let's look to few fake videos.","metadata":{}},{"cell_type":"code","source":"fake_videos = list(train_sample_metadata.loc[train_sample_metadata.label=='FAKE'].index)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from IPython.display import HTML\nfrom base64 import b64encode\n\ndef play_video(video_file, subset=TRAIN_SAMPLE_FOLDER):\n    '''\n    Display video\n    param: video_file - the name of the video file to display\n    param: subset - the folder where the video file is located (can be TRAIN_SAMPLE_FOLDER or TEST_Folder)\n    '''\n    video_url = open(os.path.join(DATA_FOLDER, subset,video_file),'rb').read()\n    data_url = \"data:video/mp4;base64,\" + b64encode(video_url).decode()\n    return HTML(\"\"\"<video width=500 controls><source src=\"%s\" type=\"video/mp4\"></video>\"\"\" % data_url)\n\nplay_video(fake_videos[10])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From visual inspection of these fakes videos, in some cases is very easy to spot the anomalies created when engineering the deep fake, in some cases is more difficult.","metadata":{}},{"cell_type":"markdown","source":"## Modelling","metadata":{}},{"cell_type":"markdown","source":"### A CNN-RNN Architecture","metadata":{}},{"cell_type":"code","source":"IMG_SIZE = 224\nBATCH_SIZE = 64\nEPOCHS = 10\n\nMAX_SEQ_LENGTH = 20\nNUM_FEATURES = 2048","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" In this example we will do the following:\n\n* Capture the frames of a video.\n* Extract frames from the videos until a maximum frame count is reached.\n* In the case, where a video's frame count is lesser than the maximum frame count we will pad the video with zeros.","metadata":{}},{"cell_type":"code","source":"def crop_center_square(frame):\n    y, x = frame.shape[0:2]\n    min_dim = min(y, x)\n    start_x = (x // 2) - (min_dim // 2)\n    start_y = (y // 2) - (min_dim // 2)\n    return frame[start_y : start_y + min_dim, start_x : start_x + min_dim]\n\n\ndef load_video(path, max_frames=0, resize=(IMG_SIZE, IMG_SIZE)):\n    cap = cv2.VideoCapture(path)\n    frames = []\n    try:\n        while True:\n            ret, frame = cap.read()\n            if not ret:\n                break\n            frame = crop_center_square(frame)\n            frame = cv2.resize(frame, resize)\n            frame = frame[:, :, [2, 1, 0]]\n            frames.append(frame)\n\n            if len(frames) == max_frames:\n                break\n    finally:\n        cap.release()\n    return np.array(frames)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can use a pre-trained network to extract meaningful features from the extracted frames. The Keras Applications module provides a number of state-of-the-art models pre-trained on the ImageNet-1k dataset. We will be using the InceptionV3 model for this purpose.","metadata":{}},{"cell_type":"code","source":"def build_feature_extractor():\n    feature_extractor = keras.applications.InceptionV3(\n        weights=\"imagenet\",\n        include_top=False,\n        pooling=\"avg\",\n        input_shape=(IMG_SIZE, IMG_SIZE, 3),\n    )\n    preprocess_input = keras.applications.inception_v3.preprocess_input\n\n    inputs = keras.Input((IMG_SIZE, IMG_SIZE, 3))\n    preprocessed = preprocess_input(inputs)\n\n    outputs = feature_extractor(preprocessed)\n    return keras.Model(inputs, outputs, name=\"feature_extractor\")\n\n\nfeature_extractor = build_feature_extractor()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Finally, we can put all the pieces together to create our data processing utility.","metadata":{}},{"cell_type":"code","source":"def prepare_all_videos(df, root_dir):\n    num_samples = len(df)\n    video_paths = list(df.index)\n    labels = df[\"label\"].values\n    labels = np.array(labels=='FAKE').astype(np.int)\n\n    # `frame_masks` and `frame_features` are what we will feed to our sequence model.\n    # `frame_masks` will contain a bunch of booleans denoting if a timestep is\n    # masked with padding or not.\n    frame_masks = np.zeros(shape=(num_samples, MAX_SEQ_LENGTH), dtype=\"bool\")\n    frame_features = np.zeros(\n        shape=(num_samples, MAX_SEQ_LENGTH, NUM_FEATURES), dtype=\"float32\"\n    )\n\n    # For each video.\n    for idx, path in enumerate(video_paths):\n        # Gather all its frames and add a batch dimension.\n        frames = load_video(os.path.join(root_dir, path))\n        frames = frames[None, ...]\n\n        # Initialize placeholders to store the masks and features of the current video.\n        temp_frame_mask = np.zeros(shape=(1, MAX_SEQ_LENGTH,), dtype=\"bool\")\n        temp_frame_features = np.zeros(\n            shape=(1, MAX_SEQ_LENGTH, NUM_FEATURES), dtype=\"float32\"\n        )\n\n        # Extract features from the frames of the current video.\n        for i, batch in enumerate(frames):\n            video_length = batch.shape[0]\n            length = min(MAX_SEQ_LENGTH, video_length)\n            for j in range(length):\n                temp_frame_features[i, j, :] = feature_extractor.predict(\n                    batch[None, j, :]\n                )\n            temp_frame_mask[i, :length] = 1  # 1 = not masked, 0 = masked\n\n        frame_features[idx,] = temp_frame_features.squeeze()\n        frame_masks[idx,] = temp_frame_mask.squeeze()\n\n    return (frame_features, frame_masks), labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since we don't have test labels we split the training data to find its performance in unseen data","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nTrain_set, Test_set = train_test_split(train_sample_metadata,test_size=0.1,random_state=42,stratify=train_sample_metadata['label'])\n\nprint(Train_set.shape, Test_set.shape )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data, train_labels = prepare_all_videos(Train_set, \"train\")\ntest_data, test_labels = prepare_all_videos(Test_set, \"test\")\n\nprint(f\"Frame features in train set: {train_data[0].shape}\")\nprint(f\"Frame masks in train set: {train_data[1].shape}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## The sequence model","metadata":{}},{"cell_type":"markdown","source":"Now, we can feed this data to a sequence model consisting of recurrent layers like GRU.","metadata":{}},{"cell_type":"code","source":"frame_features_input = keras.Input((MAX_SEQ_LENGTH, NUM_FEATURES))\nmask_input = keras.Input((MAX_SEQ_LENGTH,), dtype=\"bool\")\n\n# Refer to the following tutorial to understand the significance of using `mask`:\n# https://keras.io/api/layers/recurrent_layers/gru/\nx = keras.layers.GRU(16, return_sequences=True)(\n    frame_features_input, mask=mask_input\n)\nx = keras.layers.GRU(8)(x)\nx = keras.layers.Dropout(0.4)(x)\nx = keras.layers.Dense(8, activation=\"relu\")(x)\noutput = keras.layers.Dense(1, activation=\"sigmoid\")(x)\n\nmodel = keras.Model([frame_features_input, mask_input], output)\n\nmodel.compile(loss=\"binary_crossentropy\", optimizer=\"adam\", metrics=[\"accuracy\"])\nmodel.summary()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"checkpoint = keras.callbacks.ModelCheckpoint('./', save_weights_only=True, save_best_only=True)\nhistory = model.fit(\n        [train_data[0], train_data[1]],\n        train_labels,\n        validation_data=([test_data[0], test_data[1]],test_labels),\n        callbacks=[checkpoint],\n        epochs=EPOCHS,\n        batch_size=8\n    )","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Inference","metadata":{}},{"cell_type":"code","source":"def prepare_single_video(frames):\n    frames = frames[None, ...]\n    frame_mask = np.zeros(shape=(1, MAX_SEQ_LENGTH,), dtype=\"bool\")\n    frame_features = np.zeros(shape=(1, MAX_SEQ_LENGTH, NUM_FEATURES), dtype=\"float32\")\n\n    for i, batch in enumerate(frames):\n        video_length = batch.shape[0]\n        length = min(MAX_SEQ_LENGTH, video_length)\n        for j in range(length):\n            frame_features[i, j, :] = feature_extractor.predict(batch[None, j, :])\n        frame_mask[i, :length] = 1  # 1 = not masked, 0 = masked\n\n    return frame_features, frame_mask\n\ndef sequence_prediction(path):\n    frames = load_video(os.path.join(DATA_FOLDER, TEST_FOLDER,path))\n    frame_features, frame_mask = prepare_single_video(frames)\n    return model.predict([frame_features, frame_mask])[0]\n    \n# This utility is for visualization.\n# Referenced from:\n# https://www.tensorflow.org/hub/tutorials/action_recognition_with_tf_hub\ndef to_gif(images):\n    converted_images = images.astype(np.uint8)\n    imageio.mimsave(\"animation.gif\", converted_images, fps=10)\n    return embed.embed_file(\"animation.gif\")\n\n\ntest_video = np.random.choice(test_videos[\"video\"].values.tolist())\nprint(f\"Test video path: {test_video}\")\n\nif(sequence_prediction(test_video)>=0.5):\n    print(f'The predicted class of the video is FAKE')\nelse:\n    print(f'The predicted class of the video is REAL')\n\nplay_video(test_video,TEST_FOLDER)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Here we used simple RNN model feel free to try some complex Attention based and Transformer based models","metadata":{}},{"cell_type":"markdown","source":"# Conclusion","metadata":{}},{"cell_type":"markdown","source":"The deep fake image classifier and deep fake video classifier model gave a generalisation error of around 80%","metadata":{}},{"cell_type":"markdown","source":"# Reference and Thanks to ","metadata":{}},{"cell_type":"markdown","source":"https://keras.io/examples/vision/video_classification/\n\nhttps://www.kaggle.com/code/gpreda/deepfake-starter-kit\n\nhttps://www.kaggle.com/code/robikscube/kaggle-deepfake-detection-introduction\n\nhttps://www.kaggle.com/code/humananalog/binary-image-classifier-training-demo\n\nhttps://www.kaggle.com/datasets/dagnelies/deepfake-faces","metadata":{}},{"cell_type":"markdown","source":"**Please UPVOTE if you found the notebook useful**","metadata":{}}]}