{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Link to actual file: https://www.kaggle.com/code/st17954/deepfake-starter-kit-0607ce/edit\nLink to kaggle https://www.kaggle.com/code/gpreda/deepfake-starter-kit","metadata":{}},{"cell_type":"markdown","source":"<h1>DeepFake Starter Kit</h1>\n\n\n\n# <a id='0'>Content</a>\n\n- <a href='#1'>Introduction</a>  \n- <a href='#2'>Preliminary data exploration</a>  \n    * Load the packages  \n    * Load the data  \n    * Check files type  \n- <a href='#3'>Meta data exploration</a>  \n     * Missing data   \n     * Unique values  \n     * Most frequent originals  \n- <a href='#4'>Video data exploration</a>  \n     * Missing video (or meta) data  \n     * Few fake videos  \n     * Few real videos  \n     * Videos with same original  \n     * Test video files  \n     * Play video files\n- <a href='#5'>Face detection</a>  \n- <a href='#6'>Resources</a> \n- <a href='#7'>References</a>     \n\n","metadata":{"id":"HumNVMv2DwdQ"}},{"cell_type":"markdown","source":"# <a id='1'>Introduction</a>\n\n\nDeepFake is composed from Deep Learning and Fake and means taking one person  from an image or video and replacing with someone else\nlikeness using technology such as Deep Artificial Neural Networks [1]. Large companies like Google invest very much in fighting the DeepFake, this including release of large datasets to help training models to counter this threat [2].The phenomen invades rapidly the film industry and threatens to compromise news agencies. Large digital companies, including content providers and social platforms are in the frontrun of fighting Deep Fakes. GANs that generate DeepFakes becomes better every day and, of course, if you include in a new GAN model all the information we collected until now how to combat various existent models, we create a model that cannot be beatten by the existing ones.   \n\nIn the **Data Exploration** section we perform a (partial) Exploratory Data Analysis (EDA) on the training and testing data. After we are checking the files types, we are focusing first on the **metadata** files, which we are exploring in details, after we are importing in dataframes. Then, we move to explore video files, by looking first to a sample of fake videos, then to real videos. After that, we are also exploring few of the videos with the same origin. We are visualizing one frame extracted from the video, for both real and fake videos. Then we are also playing few videos.  \nThen, we move to perform face (and other `objects` from the persons in the videos) extraction. More precisely, we are using OpenCV Haar Cascade resources to identify frontal face, eyes, smile and profile face from still images in the videos.\n\n**Important note**: The data we analyze here is just a very small sample of data. The competition specifies that the train data is provided as archived chunks. Training of models should pe performed offline using the data provided by Kaggle as archives, models should be loaded (max 1GB memory) in a Kernel, where inference should be performed (submission sample file provided) and prediction should be prepared as an output file from the Kernel.\n\n\nIn the **Resources** section I provide a short list of various resources for GAN and DeepFake, with blog posts, Kaggle Kernels and Github repos.   \n\n","metadata":{"id":"9QyKL-TjDwdT"}},{"cell_type":"markdown","source":"\n","metadata":{"id":"0CmleVorDwdU"}},{"cell_type":"markdown","source":"# <a id='2'>Preliminary data exploration</a>","metadata":{"id":"KRNYCnmKDwdU"}},{"cell_type":"markdown","source":"## Load packages","metadata":{"id":"qJPJjalkDwdV"}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport matplotlib\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm_notebook\n%matplotlib inline \nimport cv2 as cv","metadata":{"_kg_hide-input":true,"id":"iluRVWErDwdV","execution":{"iopub.status.busy":"2023-11-04T13:53:23.617175Z","iopub.execute_input":"2023-11-04T13:53:23.617511Z","iopub.status.idle":"2023-11-04T13:53:23.625520Z","shell.execute_reply.started":"2023-11-04T13:53:23.617460Z","shell.execute_reply":"2023-11-04T13:53:23.624834Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code imports several Python libraries that are commonly used in data analysis, visualization, and computer vision tasks.\n\nThe first line of the code, \"import numpy as np,\" imports the NumPy library and assigns it the shorthand name \"np\" for convenience. NumPy is a library for numerical computing in Python that provides support for multi-dimensional arrays and matrices. It is widely used for scientific computing, data analysis, and machine learning tasks.\n\nThe second line, \"import pandas as pd,\" imports the Pandas library and assigns it the shorthand name \"pd\" for convenience. Pandas is a library for data manipulation and analysis. It provides a way to handle and process data in a tabular form, which is similar to spreadsheets or SQL tables.\n\nThe third line, \"import os,\" imports the built-in os module, which provides a way to interact with the operating system in Python. This can be used for tasks like reading and writing files, creating directories, and executing system commands.\n\nThe fourth line, \"import matplotlib,\" imports the Matplotlib library, which is a popular data visualization library in Python. Matplotlib provides a way to create static, animated, and interactive visualizations in Python.\n\nThe fifth line, \"import seaborn as sns,\" imports the Seaborn library, which is another data visualization library in Python that builds on top of Matplotlib. Seaborn provides a high-level interface for creating informative and attractive statistical graphics.\n\nThe sixth line, \"import matplotlib.pyplot as plt,\" imports the pyplot module from Matplotlib and assigns it the shorthand name \"plt\" for convenience. The pyplot module provides a simple interface for creating and customizing plots, which is especially useful for interactive data analysis and exploration.\n\nThe seventh line, \"from tqdm import tqdm_notebook,\" imports the tqdm_notebook function from the tqdm library, which is a tool for adding progress bars to loops in Python. This can be helpful for tracking the progress of time-consuming tasks and debugging code.\n\nThe eighth line, \"%matplotlib inline,\" is a magic command in Jupyter notebooks that tells Matplotlib to display its plots inline in the notebook. This makes it easy to explore data and visualize results within the context of a Jupyter notebook.\n\nThe last line, \"import cv2 as cv,\" imports the OpenCV library and assigns it the shorthand name \"cv\" for convenience. OpenCV is a library for computer vision tasks like image and video processing. It provides a range of functions for tasks like image filtering, feature detection, and object recognition. Overall, these libraries are essential tools for data analysis, visualization, and computer vision tasks in Python.","metadata":{"id":"hoHwinYMDyqG"}},{"cell_type":"markdown","source":"## Load data","metadata":{"id":"Gv91kRWADwdW"}},{"cell_type":"code","source":"DATA_FOLDER = '../input/deepfake-detection-challenge'\nTRAIN_SAMPLE_FOLDER = 'train_sample_videos'\nTEST_FOLDER = 'test_videos'\n\nprint(f\"Train samples: {len(os.listdir(os.path.join(DATA_FOLDER, TRAIN_SAMPLE_FOLDER)))}\")\nprint(f\"Test samples: {len(os.listdir(os.path.join(DATA_FOLDER, TEST_FOLDER)))}\")","metadata":{"_kg_hide-input":true,"id":"K0b-WCZ1DwdX","execution":{"iopub.status.busy":"2023-11-04T13:53:23.627272Z","iopub.execute_input":"2023-11-04T13:53:23.627773Z","iopub.status.idle":"2023-11-04T13:53:23.642749Z","shell.execute_reply.started":"2023-11-04T13:53:23.627726Z","shell.execute_reply":"2023-11-04T13:53:23.641703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code sets up variables for the paths to the data files for a deepfake detection challenge. The first line of code creates a variable called \"DATA_FOLDER\" and sets it to the path \"../input/deepfake-detection-challenge\". This is the main folder where the data for the challenge is stored.\n\nThe second line of code creates a variable called \"TRAIN_SAMPLE_FOLDER\" and sets it to \"train_sample_videos\". This folder contains a sample of the training data for the challenge, which is a set of videos that have been labeled as real or fake.\n\nThe third line of code creates a variable called \"TEST_FOLDER\" and sets it to \"test_videos\". This folder contains a set of videos that will be used to test the accuracy of deepfake detection models.\n\nThe next two lines of code use the \"os\" module to count the number of files in each of these folders. The \"os.listdir\" function returns a list of all the files in a directory, and the \"len\" function returns the number of items in that list.\n\nThe fourth line of code uses f-strings, a feature introduced in Python 3.6, to print the number of samples in the training folder. The f-string starts with the letter \"f\" and includes curly braces {} that contain the variable or expression to be evaluated and displayed. The output of this line of code will be something like \"Train samples: 100\".\n\nThe fifth line of code uses a similar f-string to print the number of samples in the test folder. The output of this line of code will be something like \"Test samples: 50\".\n\nOverall, this code is a simple way to check the number of data samples in the training and test folders for a deepfake detection challenge. It is a useful step in data exploration and understanding the scope of the challenge.","metadata":{"id":"qxteMx7_D2Rc"}},{"cell_type":"markdown","source":"We also added a face detection resource.","metadata":{"id":"D71XCFjADwdX"}},{"cell_type":"code","source":"FACE_DETECTION_FOLDER = '../input/haar-cascades-for-face-detection'\nprint(f\"Face detection resources: {os.listdir(FACE_DETECTION_FOLDER)}\")","metadata":{"_kg_hide-input":true,"id":"QU57LsE-DwdX","execution":{"iopub.status.busy":"2023-11-04T13:53:23.644178Z","iopub.execute_input":"2023-11-04T13:53:23.644655Z","iopub.status.idle":"2023-11-04T13:53:23.657758Z","shell.execute_reply.started":"2023-11-04T13:53:23.644610Z","shell.execute_reply":"2023-11-04T13:53:23.656859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code defines variables that contain the paths to different folders within a dataset for a deepfake detection challenge. The first line creates a variable named \"DATA_FOLDER\" and assigns it the string value \"../input/deepfake-detection-challenge\". This variable points to the main folder where the data for the challenge is stored.\n\nThe second line creates a variable named \"TRAIN_SAMPLE_FOLDER\" and assigns it the string value \"train_sample_videos\". This variable points to a folder within the \"DATA_FOLDER\" that contains a subset of the training videos. These videos have been labeled as either real or fake.\n\nThe third line creates a variable named \"TEST_FOLDER\" and assigns it the string value \"test_videos\". This variable points to a folder within the \"DATA_FOLDER\" that contains the test videos that will be used to evaluate the accuracy of deepfake detection models.\n\nThe next two lines of code use the \"os\" module to count the number of files in each of the folders. Specifically, the \"os.listdir()\" function is used to return a list of all the files in the folder specified by the path, and the \"len()\" function is used to count the number of files in the returned list.\n\nThe fourth line of code uses an f-string, a string formatting technique introduced in Python 3.6, to display the number of training samples. The f-string uses curly braces to insert the number of files returned by the \"os.listdir()\" function into the string. The \"os.path.join()\" function is used to create a path to the \"TRAIN_SAMPLE_FOLDER\" within the \"DATA_FOLDER\" directory.\n\nThe fifth line of code uses a similar f-string to display the number of test samples. It uses the same \"os.path.join()\" function to create a path to the \"TEST_FOLDER\" within the \"DATA_FOLDER\" directory.\n\nOverall, this code is a simple way to count the number of samples in the training and test folders of a deepfake detection dataset. It is a useful step in understanding the size and scope of the dataset and in preparing for further data analysis and modeling tasks.","metadata":{"id":"o46x5vbPD4qu"}},{"cell_type":"markdown","source":"## Check files type\n\nHere we check the train data files extensions. Most of the files looks to have `mp4` extension, let's check if there is other extension as well.","metadata":{"id":"l2eJugmGDwdY"}},{"cell_type":"code","source":"train_list = list(os.listdir(os.path.join(DATA_FOLDER, TRAIN_SAMPLE_FOLDER)))\next_dict = []\nfor file in train_list:\n    file_ext = file.split('.')[1]\n    if (file_ext not in ext_dict):\n        ext_dict.append(file_ext)\nprint(f\"Extensions: {ext_dict}\")      ","metadata":{"_kg_hide-input":true,"id":"7taVnLPpDwdY","execution":{"iopub.status.busy":"2023-11-04T13:53:23.659514Z","iopub.execute_input":"2023-11-04T13:53:23.659982Z","iopub.status.idle":"2023-11-04T13:53:23.669119Z","shell.execute_reply.started":"2023-11-04T13:53:23.659788Z","shell.execute_reply":"2023-11-04T13:53:23.667965Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet reads a list of files located in a specific directory path DATA_FOLDER and TRAIN_SAMPLE_FOLDER subdirectory. It creates an empty list ext_dict to store unique file extensions present in the directory.\n\nThe for loop iterates through each file in the directory and extracts the file extension by splitting the filename using the split() method and taking the second part of the resulting list. If the file extension is not already in the ext_dict list, it gets appended to the list.\n\nFinally, the code prints the unique extensions found in the directory, using an f-string to format the output as a string preceded by the word \"Extensions:\". Overall, this code is useful for finding out what types of files exist in a directory by examining their extensions.","metadata":{"id":"T3hAlWKdEKgJ"}},{"cell_type":"markdown","source":"Let's count how many files with each extensions there are.","metadata":{"id":"5FA_wVmwDwdb"}},{"cell_type":"code","source":"for file_ext in ext_dict:\n    print(f\"Files with extension `{file_ext}`: {len([file for file in train_list if  file.endswith(file_ext)])}\")","metadata":{"_kg_hide-input":true,"id":"TEJ6dX2MDwdb","execution":{"iopub.status.busy":"2023-11-04T13:53:23.671519Z","iopub.execute_input":"2023-11-04T13:53:23.671990Z","iopub.status.idle":"2023-11-04T13:53:23.679434Z","shell.execute_reply.started":"2023-11-04T13:53:23.671936Z","shell.execute_reply":"2023-11-04T13:53:23.678744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet is used to print out the count of files in a specific directory with a particular file extension. The loop iterates over each unique file extension present in the ext_dict list generated by the previous code snippet.\n\nFor each file extension, the code uses a list comprehension to create a new list of filenames in train_list that end with the extension file_ext. The endswith() method checks whether the file extension of each file matches the current extension in the loop.\n\nThe code then prints a string that shows the number of files with the current file extension. The f-string uses curly braces {} to interpolate the variable file_ext and the length of the filtered list. The output is a string that shows the file extension and the count of files in the directory that have that extension.\n\nOverall, this code is useful for obtaining information about the files in a directory based on their extensions, providing a summary of the types of files present and their relative proportions.","metadata":{"id":"VxzhCJ-4ES2O"}},{"cell_type":"markdown","source":"Let's repeat the same process for test videos folder.","metadata":{"id":"B0uQDwS8Dwdc"}},{"cell_type":"code","source":"test_list = list(os.listdir(os.path.join(DATA_FOLDER, TEST_FOLDER)))\next_dict = []\nfor file in test_list:\n    file_ext = file.split('.')[1]\n    if (file_ext not in ext_dict):\n        ext_dict.append(file_ext)\nprint(f\"Extensions: {ext_dict}\")\nfor file_ext in ext_dict:\n    print(f\"Files with extension `{file_ext}`: {len([file for file in train_list if  file.endswith(file_ext)])}\")","metadata":{"_kg_hide-input":true,"id":"ZWQZGpalDwdc","execution":{"iopub.status.busy":"2023-11-04T13:53:23.680869Z","iopub.execute_input":"2023-11-04T13:53:23.681272Z","iopub.status.idle":"2023-11-04T13:53:23.692620Z","shell.execute_reply.started":"2023-11-04T13:53:23.681230Z","shell.execute_reply":"2023-11-04T13:53:23.691935Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet is used to list the file extensions and the count of files with each extension in a specific directory.\n\nThe first part of the code reads a list of files located in the TEST_FOLDER directory within the DATA_FOLDER path. It creates an empty list ext_dict to store unique file extensions present in the directory.\n\nThe second part of the code iterates over each file in the test_list and extracts the file extension by splitting the filename using the split() method and taking the second part of the resulting list. If the file extension is not already in the ext_dict list, it gets appended to the list.\n\nThe third part of the code prints the unique extensions found in the test_list directory, using an f-string to format the output as a string preceded by the word \"Extensions:\".\n\nThe fourth part of the code loops through each unique file extension present in the ext_dict list generated by the previous loop.\n\nFor each file extension, the code uses a list comprehension to create a new list of filenames in the train_list that end with the extension file_ext. The endswith() method checks whether the file extension of each file matches the current extension in the loop.\n\nFinally, the code prints a string that shows the number of files with the current file extension in the train_list directory. The f-string uses curly braces {} to interpolate the variable file_ext and the length of the filtered list. The output is a string that shows the file extension and the count of files in the train_list directory that have that extension.\n\nOverall, this code provides information about the file extensions and the count of files with each extension in both train_list and test_list directories, allowing for comparison between the two sets of files.\n\n\n\n","metadata":{"id":"mFQDyvxPEXgs"}},{"cell_type":"markdown","source":"Let's check the `json` file first.","metadata":{"id":"gcQV7niMDwdd"}},{"cell_type":"code","source":"json_file = [file for file in train_list if  file.endswith('json')][0]\nprint(f\"JSON file: {json_file}\")","metadata":{"_kg_hide-input":true,"id":"4h0mOSosDwdd","execution":{"iopub.status.busy":"2023-11-04T13:53:23.693913Z","iopub.execute_input":"2023-11-04T13:53:23.694320Z","iopub.status.idle":"2023-11-04T13:53:23.711557Z","shell.execute_reply.started":"2023-11-04T13:53:23.694276Z","shell.execute_reply":"2023-11-04T13:53:23.710841Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet searches for a file in the train_list that ends with the extension .json and assigns it to the variable json_file.\n\nThe first part of the code uses a list comprehension to filter out all the filenames in the train_list that do not end with the .json extension. It creates a new list that contains only filenames that end with .json.\n\nThe [0] at the end of the list comprehension is used to extract the first filename that matches the condition. Since we are expecting only one file with the .json extension in the train_list directory, this operation ensures that json_file is a string containing the name of the JSON file.\n\nThe second part of the code prints a string that shows the name of the JSON file that was found, using an f-string to format the output. The curly braces {} are used to interpolate the variable json_file into the string preceded by the word \"JSON file:\".\n\nOverall, this code is useful for finding a specific file with a particular extension in a directory and storing the filename in a variable for later use. In this case, it is used to identify the JSON file among other files in the train_list directory.","metadata":{"id":"-8qjfgcYEb0T"}},{"cell_type":"markdown","source":"Aparently here is a metadata file. Let's explore this JSON file.","metadata":{"id":"SqdFBiYRDwde"}},{"cell_type":"code","source":"def get_meta_from_json(path):\n    df = pd.read_json(os.path.join(DATA_FOLDER, path, json_file))\n    df = df.T\n    return df\n\nmeta_train_df = get_meta_from_json(TRAIN_SAMPLE_FOLDER)\nmeta_train_df.head()","metadata":{"_kg_hide-input":true,"id":"b5l1RSf2Dwde","execution":{"iopub.status.busy":"2023-11-04T13:53:23.712702Z","iopub.execute_input":"2023-11-04T13:53:23.713019Z","iopub.status.idle":"2023-11-04T13:53:24.438774Z","shell.execute_reply.started":"2023-11-04T13:53:23.712962Z","shell.execute_reply":"2023-11-04T13:53:24.438066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code snippet defines a function get_meta_from_json(path) that reads a JSON file located at a specific directory path and returns a pandas DataFrame object containing the data from the JSON file.\n\nThe function takes one argument path, which is a string that specifies the subdirectory path where the JSON file is located. The JSON file name is specified by the json_file variable that was previously defined.\n\nThe function first reads the JSON file into a DataFrame object df using the read_json() method from the pandas library. The os.path.join() method is used to join the DATA_FOLDER path, the path argument, and the json_file variable to obtain the full path to the JSON file.\n\nThe df.T method is used to transpose the DataFrame object so that the rows become columns and vice versa. This is useful because the original JSON data is usually structured as a list of dictionaries, where each dictionary corresponds to a row of data. By transposing the DataFrame, we can make the dictionaries become the columns of the DataFrame.\n\nFinally, the function returns the DataFrame object df containing the JSON data.\n\nThe last part of the code uses the get_meta_from_json() function to read the JSON file located in the TRAIN_SAMPLE_FOLDER subdirectory and assign the resulting DataFrame object to the variable meta_train_df. The head() method is then used to print the first five rows of the meta_train_df DataFrame to verify that the data has been read correctly.\n\nOverall, this code is useful for reading structured data stored in JSON format into a pandas DataFrame object, which can then be used for data analysis and manipulation.","metadata":{"id":"CTaRL0Y8ElZU"}},{"cell_type":"markdown","source":"# <a id=\"3\">Meta data exploration</a>\n\nLet's explore now the meta data in train sample. \n\n## Missing data\n\nWe start by checking for any missing values.  ","metadata":{"id":"Lp8olkxIDwdg"}},{"cell_type":"code","source":"def missing_data(data):\n    total = data.isnull().sum()\n    percent = (data.isnull().sum()/data.isnull().count()*100)\n    tt = pd.concat([total, percent], axis=1, keys=['Total', 'Percent'])\n    types = []\n    for col in data.columns:\n        dtype = str(data[col].dtype)\n        types.append(dtype)\n    tt['Types'] = types\n    return(np.transpose(tt))","metadata":{"_kg_hide-input":true,"id":"uPywIgSkDwdg","execution":{"iopub.status.busy":"2023-11-04T13:53:24.439962Z","iopub.execute_input":"2023-11-04T13:53:24.440346Z","iopub.status.idle":"2023-11-04T13:53:24.448630Z","shell.execute_reply.started":"2023-11-04T13:53:24.440303Z","shell.execute_reply":"2023-11-04T13:53:24.447566Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code defines a function missing_data(data) that takes a pandas DataFrame object data as input and returns a summary of the missing data in the DataFrame.\n\nThe function first uses the isnull() method to create a boolean mask of the same shape as the DataFrame, with True values where the original DataFrame has missing values, and False elsewhere. Then, the sum() method is used to sum up the number of True values (i.e., the number of missing values) along each column of the DataFrame. The resulting counts are stored in a new DataFrame called total.\n\nNext, the code computes the percentage of missing data in each column by dividing the number of missing values by the total number of values in the column and multiplying by 100. The resulting percentages are stored in a new DataFrame called percent.\n\nThe concat() method is used to concatenate the total and percent DataFrames along the axis=1 (i.e., horizontally) into a new DataFrame called tt. This new DataFrame has two columns, Total and Percent, representing the total count and percentage of missing values for each column in the original DataFrame.\n\nThe code then iterates over each column of the data DataFrame and extracts its data type using the dtype attribute. The data types are stored in a list called types.\n\nFinally, the tt DataFrame is augmented with a new column called Types, which contains the data types of the columns in the original DataFrame. The function then transposes the resulting DataFrame tt using the transpose() method and returns it.\n\nOverall, this code is useful for quickly identifying missing values in a pandas DataFrame, providing a summary of the number and percentage of missing values in each column of the DataFrame, along with the data types of the columns.","metadata":{"id":"3U_Up2cCEv1o"}},{"cell_type":"code","source":"missing_data(meta_train_df)","metadata":{"_kg_hide-input":true,"id":"IiqNJUtFDwdg","execution":{"iopub.status.busy":"2023-11-04T13:53:24.449968Z","iopub.execute_input":"2023-11-04T13:53:24.450267Z","iopub.status.idle":"2023-11-04T13:53:24.579888Z","shell.execute_reply.started":"2023-11-04T13:53:24.450209Z","shell.execute_reply":"2023-11-04T13:53:24.578851Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code is calling the missing_data() function and passing the meta_train_df DataFrame as an argument.\n\nThe missing_data() function takes a pandas DataFrame as input and returns a summary of the missing data in the DataFrame, including the number and percentage of missing values in each column, as well as the data types of the columns.\n\nBy calling the missing_data() function on the meta_train_df DataFrame, this code produces a summary of the missing data in the meta_train_df DataFrame. The output is a transposed DataFrame object that displays the number and percentage of missing values for each column in meta_train_df, along with the data types of the columns.\n\nOverall, this code is useful for identifying and understanding the extent of missing data in a pandas DataFrame, which is a crucial step in data cleaning and preparation.","metadata":{"id":"gg9PdrkVE545"}},{"cell_type":"markdown","source":"There are missing data 19.25% of the samples (or 77). We suspect that actually the real data has missing original (if we generalize from the data we glimpsed). Let's check this hypothesis.","metadata":{"id":"CvfjT6foDwdh"}},{"cell_type":"code","source":"missing_data(meta_train_df.loc[meta_train_df.label=='REAL'])","metadata":{"_kg_hide-input":true,"id":"E24Y6k6LDwdh","execution":{"iopub.status.busy":"2023-11-04T13:53:24.584921Z","iopub.execute_input":"2023-11-04T13:53:24.585207Z","iopub.status.idle":"2023-11-04T13:53:24.607367Z","shell.execute_reply.started":"2023-11-04T13:53:24.585161Z","shell.execute_reply":"2023-11-04T13:53:24.606283Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code is calling the missing_data() function on a subset of the meta_train_df DataFrame that meets a specific condition, using the .loc method to select rows based on the value of the label column.\n\nThe condition meta_train_df.label=='REAL' selects only the rows in meta_train_df where the value of the label column is equal to the string 'REAL'.\n\nThe resulting DataFrame subset is then passed as an argument to the missing_data() function, which returns a summary of the missing data in the subset. The output is a transposed DataFrame object that displays the number and percentage of missing values for each column in the subset, along with the data types of the columns.\n\nOverall, this code is useful for obtaining a summary of the missing data in a subset of a larger DataFrame that meets a specific condition, which can be useful for investigating the data quality and completeness for specific categories or groups. In this case, the code is useful for examining the missing data for the subset of meta_train_df corresponding to the 'REAL' label.\n","metadata":{"id":"BsdhipHmE-lc"}},{"cell_type":"markdown","source":"Indeed, all missing `original` data are the one associated with `REAL` label.  \n\n## Unique values\n\nLet's check into more details the unique values.","metadata":{"id":"sJ9DjpokDwdh"}},{"cell_type":"code","source":"def unique_values(data):\n    total = data.count()\n    tt = pd.DataFrame(total)\n    tt.columns = ['Total']\n    uniques = []\n    for col in data.columns:\n        unique = data[col].nunique()\n        uniques.append(unique)\n    tt['Uniques'] = uniques\n    return(np.transpose(tt))","metadata":{"_kg_hide-input":true,"id":"GfCs83XpDwdi","execution":{"iopub.status.busy":"2023-11-04T13:53:24.611498Z","iopub.execute_input":"2023-11-04T13:53:24.611959Z","iopub.status.idle":"2023-11-04T13:53:24.619274Z","shell.execute_reply.started":"2023-11-04T13:53:24.611781Z","shell.execute_reply":"2023-11-04T13:53:24.617744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code defines a function unique_values(data) that takes a pandas DataFrame object data as input and returns a summary of the unique values in the DataFrame.\n\nThe function first uses the count() method to count the number of non-missing values for each column in the DataFrame. The resulting counts are stored in a new DataFrame called total.\n\nThe DataFrame() method is used to create a new DataFrame object called tt that has the same shape as total. The columns attribute is used to set the column names of tt to ['Total'].\n\nThe code then iterates over each column of the data DataFrame and uses the nunique() method to count the number of unique values in the column. The resulting counts are stored in a list called uniques.\n\nFinally, the tt DataFrame is augmented with a new column called Uniques, which contains the number of unique values for each column in the original DataFrame. The function then transposes the resulting DataFrame tt using the transpose() method and returns it.\n\nOverall, this code is useful for quickly identifying the number of unique values in a pandas DataFrame, providing a summary of the number of unique values for each column in the DataFrame.","metadata":{"id":"sMxEorCIFIPQ"}},{"cell_type":"code","source":"unique_values(meta_train_df)","metadata":{"_kg_hide-input":true,"id":"BLtquuf8Dwdi","execution":{"iopub.status.busy":"2023-11-04T13:53:24.620944Z","iopub.execute_input":"2023-11-04T13:53:24.621447Z","iopub.status.idle":"2023-11-04T13:53:24.642418Z","shell.execute_reply.started":"2023-11-04T13:53:24.621398Z","shell.execute_reply":"2023-11-04T13:53:24.641560Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code is calling the unique_values() function and passing the meta_train_df DataFrame as an argument.\n\nThe unique_values() function takes a pandas DataFrame as input and returns a summary of the unique values in the DataFrame, including the total count of non-missing values and the number of unique values in each column.\n\nBy calling the unique_values() function on the meta_train_df DataFrame, this code produces a summary of the unique values in the meta_train_df DataFrame. The output is a transposed DataFrame object that displays the total count of non-missing values and the number of unique values for each column in meta_train_df.\n\nOverall, this code is useful for quickly identifying the number of unique values in each column of a pandas DataFrame, which is a useful step in data exploration and understanding.","metadata":{"id":"SvNm3mIuFNNe"}},{"cell_type":"markdown","source":"* We observe that `original` label has the same pattern for uniques values. We know that we have 77 missing data (that's why total is only 323) and we observe that we do have 209 unique examples.  \n\n## Most frequent originals\n\nLet's look now to the most frequent originals uniques in train sample data.  ","metadata":{"id":"uxKI0AfBDwdi"}},{"cell_type":"code","source":"def most_frequent_values(data):\n    total = data.count()\n    tt = pd.DataFrame(total)\n    tt.columns = ['Total']\n    items = []\n    vals = []\n    for col in data.columns:\n        itm = data[col].value_counts().index[0]\n        val = data[col].value_counts().values[0]\n        items.append(itm)\n        vals.append(val)\n    tt['Most frequent item'] = items\n    tt['Frequence'] = vals\n    tt['Percent from total'] = np.round(vals / total * 100, 3)\n    return(np.transpose(tt))","metadata":{"_kg_hide-input":true,"id":"MV0kS0wKDwdj","execution":{"iopub.status.busy":"2023-11-04T13:53:24.644022Z","iopub.execute_input":"2023-11-04T13:53:24.644370Z","iopub.status.idle":"2023-11-04T13:53:24.655699Z","shell.execute_reply.started":"2023-11-04T13:53:24.644296Z","shell.execute_reply":"2023-11-04T13:53:24.654746Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code defines a Python function named \"most_frequent_values\" that takes a pandas DataFrame as input and computes the most frequently occurring value(s) for each column in the DataFrame, along with some additional information.\n\nThe function first calculates the total number of non-missing values in each column of the input DataFrame using the \"count()\" method of pandas. It then creates a new DataFrame named \"tt\" using the \"total\" variable. The \"tt\" DataFrame contains a single column called \"Total\", which shows the number of non-missing values for each column of the input DataFrame.\n\nThe function then iterates over each column of the input DataFrame using a for loop. For each column, it computes the most frequently occurring value(s) using the \"value_counts()\" method of pandas. It then appends the most frequent value and its frequency to two lists named \"items\" and \"vals\", respectively.\n\nAfter computing the most frequent value(s) and their frequency for each column, the function adds three new columns to the \"tt\" DataFrame. The first column is named \"Most frequent item\" and contains the most frequent value for each column. The second column is named \"Frequency\" and contains the frequency of the most frequent value for each column. The third column is named \"Percent from total\" and contains the percentage of the total number of non-missing values that the most frequent value represents for each column.\n\nFinally, the function returns the transposed \"tt\" DataFrame, which has rows as columns and columns as rows. The resulting DataFrame shows the most frequent value(s), frequency, and percentage of the total for each column of the input DataFrame.","metadata":{"id":"Lau2i2IBFgW2"}},{"cell_type":"code","source":"most_frequent_values(meta_train_df)","metadata":{"_kg_hide-input":true,"id":"DhLjLCyXDwdj","execution":{"iopub.status.busy":"2023-11-04T13:53:24.656883Z","iopub.execute_input":"2023-11-04T13:53:24.657260Z","iopub.status.idle":"2023-11-04T13:53:24.690608Z","shell.execute_reply.started":"2023-11-04T13:53:24.657220Z","shell.execute_reply":"2023-11-04T13:53:24.689838Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code \"most_frequent_values(meta_train_df)\" is calling the \"most_frequent_values\" function with an argument named \"meta_train_df\". This suggests that \"meta_train_df\" is a pandas DataFrame, and the function is being used to calculate the most frequent value(s) and additional information for each column in this DataFrame.\n\nWhen the \"most_frequent_values\" function is called with \"meta_train_df\" as an argument, it computes the most frequently occurring value(s) for each column in the DataFrame along with some additional information, as described in the previous answer. The function then returns a transposed DataFrame that shows the most frequent value(s), frequency, and percentage of the total for each column of the input DataFrame.\n\nSo, by calling \"most_frequent_values(meta_train_df)\", the user is able to quickly and easily calculate the most frequent value(s) and additional information for each column of the \"meta_train_df\" DataFrame, which can be useful for exploring the data and identifying potential issues, such as missing or outlier values, that may need to be addressed before further analysis.","metadata":{"id":"gZD3mX8pFlXs"}},{"cell_type":"markdown","source":"We see that most frequent **label** is `FAKE` (80.75%), `meawmsgiti.mp4` is the most frequent **original** (6 samples).","metadata":{"id":"GCMRegb3Dwdj"}},{"cell_type":"markdown","source":"Let's do now some data distribution visualizations.","metadata":{"id":"qXwt6UjFDwdj"}},{"cell_type":"code","source":"def plot_count(feature, title, df, size=1):\n    '''\n    Plot count of classes / feature\n    param: feature - the feature to analyze\n    param: title - title to add to the graph\n    param: df - dataframe from which we plot feature's classes distribution \n    param: size - default 1.\n    '''\n    f, ax = plt.subplots(1,1, figsize=(4*size,4))\n    total = float(len(df))\n    g = sns.countplot(df[feature], order = df[feature].value_counts().index[:20], palette='Set3')\n    g.set_title(\"Number and percentage of {}\".format(title))\n    if(size > 2):\n        plt.xticks(rotation=90, size=8)\n    for p in ax.patches:\n        height = p.get_height()\n        ax.text(p.get_x()+p.get_width()/2.,\n                height + 3,\n                '{:1.2f}%'.format(100*height/total),\n                ha=\"center\") \n    plt.show()    ","metadata":{"_kg_hide-input":true,"id":"2nSyyV-HDwdk","execution":{"iopub.status.busy":"2023-11-04T13:53:24.691856Z","iopub.execute_input":"2023-11-04T13:53:24.692282Z","iopub.status.idle":"2023-11-04T13:53:24.703309Z","shell.execute_reply.started":"2023-11-04T13:53:24.692216Z","shell.execute_reply":"2023-11-04T13:53:24.702582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code defines a Python function named \"plot_count\" that takes four parameters: \"feature\", \"title\", \"df\", and \"size\". The purpose of the function is to plot the count of classes for a given feature in a pandas DataFrame (\"df\") and display the number and percentage of each class in the plot.\n\nThe function first creates a new figure using the \"subplots\" method of matplotlib and sets the figure size based on the \"size\" parameter. It then calculates the total number of instances in the input DataFrame using the \"len()\" function and assigns the result to the \"total\" variable.\n\nThe function then creates a countplot using seaborn's \"countplot\" function. The countplot shows the number of instances for each class of the specified \"feature\" in the input DataFrame. The classes are ordered based on their frequency using the \"value_counts()\" method of pandas. The countplot is displayed with a title that includes the specified \"title\" parameter.\n\nThe function then checks the \"size\" parameter to determine whether to rotate the x-axis labels and decrease their font size. For each bar in the countplot, the function adds a text label above the bar that shows the percentage of the total instances that the class represents.\n\nFinally, the function displays the countplot using matplotlib's \"show\" function.\n\nOverall, the \"plot_count\" function is a useful tool for visualizing the distribution of a categorical feature in a pandas DataFrame and quickly identifying the most common classes and their relative frequencies.","metadata":{"id":"OF4Vp2oNFwLg"}},{"cell_type":"code","source":"plot_count('split', 'split (train)', meta_train_df)","metadata":{"_kg_hide-input":true,"id":"sfo8GwcCDwdk","execution":{"iopub.status.busy":"2023-11-04T13:53:24.704470Z","iopub.execute_input":"2023-11-04T13:53:24.704872Z","iopub.status.idle":"2023-11-04T13:53:24.976033Z","shell.execute_reply.started":"2023-11-04T13:53:24.704822Z","shell.execute_reply":"2023-11-04T13:53:24.974734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code \"plot_count('split', 'split (train)', meta_train_df)\" is calling the \"plot_count\" function with three arguments. The first argument is a string \"split\", which represents the name of a feature/column in the \"meta_train_df\" DataFrame that contains information about whether a data point belongs to the training, validation, or testing set. The second argument is a string \"split (train)\", which is the title of the plot that will be generated by the function. The third argument is the \"meta_train_df\" DataFrame itself.\n\nBy calling the \"plot_count\" function with these arguments, the code is generating a countplot of the \"split\" feature in the \"meta_train_df\" DataFrame. The countplot shows the number of instances for each class of the \"split\" feature, which in this case represents whether a data point belongs to the training set or not. The classes are ordered based on their frequency using the \"value_counts()\" method of pandas.\n\nThe title of the plot is set to \"split (train)\", which suggests that the plot is specifically showing the distribution of the training set among all the data points in the \"meta_train_df\" DataFrame. The function will display the countplot with the number and percentage of instances in each class, which can be useful for quickly assessing the balance of the training set relative to the other sets, such as the validation and testing sets.","metadata":{"id":"enbKqPGUF2ZZ"}},{"cell_type":"code","source":"plot_count('label', 'label (train)', meta_train_df)","metadata":{"_kg_hide-input":true,"id":"q5YSPB7XDwdk","execution":{"iopub.status.busy":"2023-11-04T13:53:24.978281Z","iopub.execute_input":"2023-11-04T13:53:24.979043Z","iopub.status.idle":"2023-11-04T13:53:25.240292Z","shell.execute_reply.started":"2023-11-04T13:53:24.978961Z","shell.execute_reply":"2023-11-04T13:53:25.239006Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The code \"plot_count('label', 'label (train)', meta_train_df)\" is calling the \"plot_count\" function with three arguments. The first argument is a string \"label\", which represents the name of a feature/column in the \"meta_train_df\" DataFrame that contains the class labels of the data points. The second argument is a string \"label (train)\", which is the title of the plot that will be generated by the function. The third argument is the \"meta_train_df\" DataFrame itself.\n\nBy calling the \"plot_count\" function with these arguments, the code is generating a countplot of the \"label\" feature in the \"meta_train_df\" DataFrame. The countplot shows the number of instances for each class of the \"label\" feature, which represents the class labels of the data points. The classes are ordered based on their frequency using the \"value_counts()\" method of pandas.\n\nThe title of the plot is set to \"label (train)\", which suggests that the plot is specifically showing the distribution of the class labels among all the data points in the training set of the \"meta_train_df\" DataFrame. The function will display the countplot with the number and percentage of instances in each class, which can be useful for quickly assessing the balance of the classes in the training set and identifying any potential issues, such as class imbalance or skewed distributions, that may need to be addressed before building a machine learning model.","metadata":{"id":"MOHJxEa0GA09"}},{"cell_type":"markdown","source":"As we can see, the `REAL` are only 19.25% in train sample videos, with the `FAKE`s acounting for 80.75% of the samples. \n\n\n# <a id=\"4\">Video data exploration</a>\n\n\nIn the following we will explore some of the video data. \n\n\n## Missing video (or meta) data\n\nWe check first if the list of files in the meta info and the list from the folder are the same.\n\n","metadata":{"id":"OslkRUq_Dwdl"}},{"cell_type":"code","source":"meta = np.array(list(meta_train_df.index))\nstorage = np.array([file for file in train_list if  file.endswith('mp4')])\nprint(f\"Metadata: {meta.shape[0]}, Folder: {storage.shape[0]}\")\nprint(f\"Files in metadata and not in folder: {np.setdiff1d(meta,storage,assume_unique=False).shape[0]}\")\nprint(f\"Files in folder and not in metadata: {np.setdiff1d(storage,meta,assume_unique=False).shape[0]}\")","metadata":{"id":"pYHGB9kxDwdl","execution":{"iopub.status.busy":"2023-11-04T13:53:25.242456Z","iopub.execute_input":"2023-11-04T13:53:25.243165Z","iopub.status.idle":"2023-11-04T13:53:25.261237Z","shell.execute_reply.started":"2023-11-04T13:53:25.243082Z","shell.execute_reply":"2023-11-04T13:53:25.259716Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code performs some comparisons between the filenames of videos in a folder and the filenames stored in a metadata DataFrame, named \"meta_train_df\".\n\nThe first line of the code creates a numpy array \"meta\" with the index of the \"meta_train_df\" DataFrame. The index contains the filenames of videos, which have been previously loaded in the DataFrame.\n\nThe second line creates another numpy array \"storage\" by scanning a directory containing the video files. The list of files is filtered by selecting only those which end with the extension 'mp4'.\n\nThe third line prints the total number of files in the metadata and storage numpy arrays, which represents the total number of videos in the metadata and storage folders, respectively.\n\nThe fourth line uses the numpy method \"setdiff1d\" to find the set difference between the \"meta\" and \"storage\" numpy arrays. It prints the number of files that are present in the metadata but not in the storage folder. These files may be missing or have been deleted from the storage folder.\n\nThe fifth line uses the \"setdiff1d\" method again to find the set difference between the \"storage\" and \"meta\" numpy arrays. It prints the number of files that are present in the storage folder but not in the metadata. These files may be new or have been added to the storage folder.\n\nOverall, the code is useful for identifying any discrepancies or inconsistencies between the filenames in the metadata and storage folders, which can help ensure that all necessary videos are available for analysis.","metadata":{"id":"hgLiHpFHGKot"}},{"cell_type":"markdown","source":"Let's visualize now the data.  \n\nWe select first a list of fake videos.\n\n## Few fake videos","metadata":{"id":"JVna_AFODwdl"}},{"cell_type":"code","source":"fake_train_sample_video = list(meta_train_df.loc[meta_train_df.label=='FAKE'].sample(3).index)\nfake_train_sample_video","metadata":{"_kg_hide-input":true,"id":"e2ec8VeeDwdl","execution":{"iopub.status.busy":"2023-11-04T13:53:25.263865Z","iopub.execute_input":"2023-11-04T13:53:25.264587Z","iopub.status.idle":"2023-11-04T13:53:25.280603Z","shell.execute_reply.started":"2023-11-04T13:53:25.264510Z","shell.execute_reply":"2023-11-04T13:53:25.279205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code is creating a list of three randomly selected video filenames from a metadata DataFrame, named \"meta_train_df\", that have been labeled as fake (i.e., the \"label\" column is equal to \"FAKE\").\n\nThe code first uses the \"loc\" method of pandas to select rows from the \"meta_train_df\" DataFrame where the \"label\" column is equal to \"FAKE\". It then uses the \"sample\" method of pandas to randomly select three rows from the resulting DataFrame. The \"sample\" method returns a new DataFrame containing the selected rows.\n\nFinally, the \"index\" attribute of the resulting DataFrame is accessed to get the index labels of the selected rows, which are the video filenames in this case. These index labels are converted to a list using the \"list\" function and assigned to the variable \"fake_train_sample_video\".\n\nTherefore, the \"fake_train_sample_video\" list contains the filenames of three randomly selected videos from the training set of the metadata DataFrame that have been labeled as fake. This list can be used to load and visualize the corresponding videos and evaluate the performance of a machine learning model trained on the training set.","metadata":{"id":"X-HCUlruGUFl"}},{"cell_type":"markdown","source":"From [4] ([Basic EDA Face Detection, split video, ROI](https://www.kaggle.com/marcovasquez/basic-eda-face-detection-split-video-roi)) we modified a function for displaying a selected image from a video.","metadata":{"id":"NK9mKwGBDwdm"}},{"cell_type":"code","source":"def display_image_from_video(video_path):\n    '''\n    input: video_path - path for video\n    process:\n    1. perform a video capture from the video\n    2. read the image\n    3. display the image\n    '''\n    capture_image = cv.VideoCapture(video_path) \n    ret, frame = capture_image.read()\n    fig = plt.figure(figsize=(10,10))\n    ax = fig.add_subplot(111)\n    frame = cv.cvtColor(frame, cv.COLOR_BGR2RGB)\n    ax.imshow(frame)","metadata":{"_kg_hide-input":true,"id":"KgGeexCZDwdm","execution":{"iopub.status.busy":"2023-11-04T13:53:25.283162Z","iopub.execute_input":"2023-11-04T13:53:25.283913Z","iopub.status.idle":"2023-11-04T13:53:25.295417Z","shell.execute_reply.started":"2023-11-04T13:53:25.283831Z","shell.execute_reply":"2023-11-04T13:53:25.294183Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code defines a Python function named \"display_image_from_video\" that takes a video file path as an argument. The purpose of the function is to read and display the first frame of the video file as an image.\n\nThe function first creates a new \"VideoCapture\" object using OpenCV's \"cv.VideoCapture\" method, which takes the video file path as an argument. This object is used to read the frames of the video.\n\nThe function then reads the first frame of the video file using the \"read\" method of the \"VideoCapture\" object. The resulting image data is stored in a variable named \"frame\".\n\nNext, the function creates a new matplotlib figure with a size of 10x10 using the \"figure\" method of matplotlib. It then adds a subplot to the figure using the \"add_subplot\" method and assigns it to the variable \"ax\".\n\nThe function then converts the color space of the image from BGR to RGB using the \"cvtColor\" method of OpenCV. This is necessary because OpenCV reads images in BGR format by default, while matplotlib expects images in RGB format.\n\nFinally, the function displays the image using the \"imshow\" method of the \"ax\" subplot. The resulting image will be displayed with the RGB color space and can be used to visualize the content of the video file.\n\nOverall, the \"display_image_from_video\" function is a useful tool for quickly viewing the content of a video file without having to play the entire video.","metadata":{"id":"svsXZKVaGbCO"}},{"cell_type":"code","source":"for video_file in fake_train_sample_video:\n    display_image_from_video(os.path.join(DATA_FOLDER, TRAIN_SAMPLE_FOLDER, video_file))","metadata":{"_kg_hide-input":true,"id":"eaKey1nTDwdn","execution":{"iopub.status.busy":"2023-11-04T13:53:25.297503Z","iopub.execute_input":"2023-11-04T13:53:25.298429Z","iopub.status.idle":"2023-11-04T13:53:27.526885Z","shell.execute_reply.started":"2023-11-04T13:53:25.298345Z","shell.execute_reply":"2023-11-04T13:53:27.525806Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code is using a \"for\" loop to iterate over each video filename in a list called \"fake_train_sample_video\". The purpose of the loop is to display the first frame of each selected video file as an image using the \"display_image_from_video\" function.\n\nFor each video file in the \"fake_train_sample_video\" list, the code first creates a file path by joining the \"DATA_FOLDER\" variable, \"TRAIN_SAMPLE_FOLDER\" constant, and the video filename using the \"os.path.join\" method. This creates a full path to the selected video file in the data directory.\n\nThe \"display_image_from_video\" function is then called with this file path as an argument, which reads and displays the first frame of the video file as an image.\n\nTherefore, the overall purpose of the code is to display the first frame of each randomly selected fake video file in the training set of the metadata DataFrame, allowing the user to visually inspect the content of the videos and assess the quality of the data.","metadata":{"id":"gP9luElFGg-Z"}},{"cell_type":"markdown","source":"Let's try now the same for few of the images that are real.  \n\n\n## Few real videos","metadata":{"id":"AtVAuiCwDwdo"}},{"cell_type":"code","source":"real_train_sample_video = list(meta_train_df.loc[meta_train_df.label=='REAL'].sample(3).index)\nreal_train_sample_video","metadata":{"_kg_hide-input":true,"id":"T1J8Q3w_Dwdo","execution":{"iopub.status.busy":"2023-11-04T13:53:27.528679Z","iopub.execute_input":"2023-11-04T13:53:27.529008Z","iopub.status.idle":"2023-11-04T13:53:27.537667Z","shell.execute_reply.started":"2023-11-04T13:53:27.528948Z","shell.execute_reply":"2023-11-04T13:53:27.536704Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code is creating a list of three randomly selected video filenames from a metadata DataFrame, named \"meta_train_df\", that have been labeled as real (i.e., the \"label\" column is equal to \"REAL\").\n\nThe code first uses the \"loc\" method of pandas to select rows from the \"meta_train_df\" DataFrame where the \"label\" column is equal to \"REAL\". It then uses the \"sample\" method of pandas to randomly select three rows from the resulting DataFrame. The \"sample\" method returns a new DataFrame containing the selected rows.\n\nFinally, the \"index\" attribute of the resulting DataFrame is accessed to get the index labels of the selected rows, which are the video filenames in this case. These index labels are converted to a list using the \"list\" function and assigned to the variable \"real_train_sample_video\".\n\nTherefore, the \"real_train_sample_video\" list contains the filenames of three randomly selected videos from the training set of the metadata DataFrame that have been labeled as real. This list can be used to load and visualize the corresponding videos and evaluate the performance of a machine learning model trained on the training set.","metadata":{"id":"c_btYgeSGm_a"}},{"cell_type":"code","source":"for video_file in real_train_sample_video:\n    display_image_from_video(os.path.join(DATA_FOLDER, TRAIN_SAMPLE_FOLDER, video_file))","metadata":{"_kg_hide-input":true,"id":"wC8WQB9pDwdo","execution":{"iopub.status.busy":"2023-11-04T13:53:27.539476Z","iopub.execute_input":"2023-11-04T13:53:27.539794Z","iopub.status.idle":"2023-11-04T13:53:29.472424Z","shell.execute_reply.started":"2023-11-04T13:53:27.539736Z","shell.execute_reply":"2023-11-04T13:53:29.471404Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code is using a \"for\" loop to iterate over each video filename in a list called \"real_train_sample_video\". The purpose of the loop is to display the first frame of each selected video file as an image using the \"display_image_from_video\" function.\n\nFor each video file in the \"real_train_sample_video\" list, the code first creates a file path by joining the \"DATA_FOLDER\" variable, \"TRAIN_SAMPLE_FOLDER\" constant, and the video filename using the \"os.path.join\" method. This creates a full path to the selected video file in the data directory.\n\nThe \"display_image_from_video\" function is then called with this file path as an argument, which reads and displays the first frame of the video file as an image.\n\nTherefore, the overall purpose of the code is to display the first frame of each randomly selected real video file in the training set of the metadata DataFrame, allowing the user to visually inspect the content of the videos and assess the quality of the data.","metadata":{"id":"7S5PkbeeGv5b"}},{"cell_type":"markdown","source":"## Videos with same original\n\nLet's look now to set of samples with the same original.","metadata":{"id":"IKayBwauDwdo"}},{"cell_type":"code","source":"meta_train_df['original'].value_counts()[0:5]","metadata":{"_kg_hide-input":true,"id":"iWjagyYoDwdo","execution":{"iopub.status.busy":"2023-11-04T13:53:29.474057Z","iopub.execute_input":"2023-11-04T13:53:29.474550Z","iopub.status.idle":"2023-11-04T13:53:29.486927Z","shell.execute_reply.started":"2023-11-04T13:53:29.474496Z","shell.execute_reply":"2023-11-04T13:53:29.485839Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code is using pandas to perform value counts on a specific column called \"original\" in the \"meta_train_df\" DataFrame. The code is returning the count of unique values in the \"original\" column, which represents the name of the original video file from which the corresponding fake video was generated.\n\nThe \"[0:5]\" at the end of the code is a slicing operation that limits the output to the top 5 most frequent values in the \"original\" column, sorted in descending order.\n\nOverall, the code is useful for quickly identifying the original videos that have been most frequently used to generate the fake videos in the training set of the metadata DataFrame.","metadata":{"id":"Qzhr2xmCG3T8"}},{"cell_type":"markdown","source":"We pick one of the originals with largest number of samples.   \n\nWe also modify our visualization function to work with multiple images.","metadata":{"id":"jRCrJNFEDwdp"}},{"cell_type":"code","source":"def display_image_from_video_list(video_path_list, video_folder=TRAIN_SAMPLE_FOLDER):\n    '''\n    input: video_path_list - path for video\n    process:\n    0. for each video in the video path list\n        1. perform a video capture from the video\n        2. read the image\n        3. display the image\n    '''\n    plt.figure()\n    fig, ax = plt.subplots(2,3,figsize=(16,8))\n    # we only show images extracted from the first 6 videos\n    for i, video_file in enumerate(video_path_list[0:6]):\n        video_path = os.path.join(DATA_FOLDER, video_folder,video_file)\n        capture_image = cv.VideoCapture(video_path) \n        ret, frame = capture_image.read()\n        frame = cv.cvtColor(frame, cv.COLOR_BGR2RGB)\n        ax[i//3, i%3].imshow(frame)\n        ax[i//3, i%3].set_title(f\"Video: {video_file}\")\n        ax[i//3, i%3].axis('on')","metadata":{"_kg_hide-input":true,"id":"GzscoMnMDwdp","execution":{"iopub.status.busy":"2023-11-04T13:53:29.488656Z","iopub.execute_input":"2023-11-04T13:53:29.489068Z","iopub.status.idle":"2023-11-04T13:53:29.500883Z","shell.execute_reply.started":"2023-11-04T13:53:29.488998Z","shell.execute_reply":"2023-11-04T13:53:29.499934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code defines a Python function named \"display_image_from_video_list\" that takes a list of video file paths as an argument. The purpose of the function is to read and display the first frame of each video file in the list as an image.\n\nThe function first creates a new matplotlib figure and subplot with 2 rows and 3 columns using the \"subplots\" method of matplotlib. The figure size is set to 16x8.\n\nThe function then iterates over each video file in the provided list. For each video file, it first creates a file path by joining the \"DATA_FOLDER\" variable, \"TRAIN_SAMPLE_FOLDER\" constant, and the video filename using the \"os.path.join\" method. This creates a full path to the video file in the data directory.\n\nThe function then creates a new \"VideoCapture\" object using OpenCV's \"cv.VideoCapture\" method, which takes the video file path as an argument. This object is used to read the frames of the video.\n\nThe function then reads the first frame of the video file using the \"read\" method of the \"VideoCapture\" object. The resulting image data is stored in a variable named \"frame\".\n\nNext, the function converts the color space of the image from BGR to RGB using the \"cvtColor\" method of OpenCV. This is necessary because OpenCV reads images in BGR format by default, while matplotlib expects images in RGB format.\n\nFinally, the function displays the image using the \"imshow\" method of the corresponding subplot, sets the title of the subplot to the video filename, and turns on the axis for the subplot.\n\nTherefore, the overall purpose of the code is to display the first frame of each video file in a list of video file paths, allowing the user to quickly assess the content and quality of the videos. By default, the function will display images for the first 6 videos in the list, arranged in a 2x3 grid.","metadata":{"id":"0I6knaGjG-Yb"}},{"cell_type":"code","source":"same_original_fake_train_sample_video = list(meta_train_df.loc[meta_train_df.original=='meawmsgiti.mp4'].index)\ndisplay_image_from_video_list(same_original_fake_train_sample_video)","metadata":{"_kg_hide-input":true,"id":"8l4iLKOsDwdp","execution":{"iopub.status.busy":"2023-11-04T13:53:29.502460Z","iopub.execute_input":"2023-11-04T13:53:29.503067Z","iopub.status.idle":"2023-11-04T13:53:32.764280Z","shell.execute_reply.started":"2023-11-04T13:53:29.502999Z","shell.execute_reply":"2023-11-04T13:53:32.763221Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code first creates a new list named \"same_original_fake_train_sample_video\" by using pandas to select rows from the \"meta_train_df\" DataFrame where the \"original\" column is equal to \"meawmsgiti.mp4\". In this case, the \"original\" column represents the name of the original video file from which the corresponding fake video was generated.\n\nThe code then calls the \"display_image_from_video_list\" function, passing the \"same_original_fake_train_sample_video\" list as an argument. The purpose of the function is to display the first frame of each video file in the list as an image.\n\nTherefore, the overall purpose of the code is to display the first frame of each fake video file in the training set of the metadata DataFrame that was generated from the original video file named \"meawmsgiti.mp4\". This can be useful for analyzing the quality and characteristics of the fake videos generated from a specific original video.\n\n\n\n","metadata":{"id":"3bUNlN74HC7V"}},{"cell_type":"markdown","source":"Let's look now to a different selection of videos with the same original. ","metadata":{"id":"1J9-56rKDwdp"}},{"cell_type":"code","source":"same_original_fake_train_sample_video = list(meta_train_df.loc[meta_train_df.original=='atvmxvwyns.mp4'].index)\ndisplay_image_from_video_list(same_original_fake_train_sample_video)","metadata":{"_kg_hide-input":true,"id":"0ZEcrMTqDwdq","execution":{"iopub.status.busy":"2023-11-04T13:53:32.766119Z","iopub.execute_input":"2023-11-04T13:53:32.767006Z","iopub.status.idle":"2023-11-04T13:53:35.800700Z","shell.execute_reply.started":"2023-11-04T13:53:32.766928Z","shell.execute_reply":"2023-11-04T13:53:35.799933Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code first creates a new list named \"same_original_fake_train_sample_video\" by using pandas to select rows from the \"meta_train_df\" DataFrame where the \"original\" column is equal to \"atvmxvwyns.mp4\". In this case, the \"original\" column represents the name of the original video file from which the corresponding fake video was generated.\n\nThe code then calls the \"display_image_from_video_list\" function, passing the \"same_original_fake_train_sample_video\" list as an argument. The purpose of the function is to display the first frame of each video file in the list as an image.\n\nTherefore, the overall purpose of the code is to display the first frame of each fake video file in the training set of the metadata DataFrame that was generated from the original video file named \"atvmxvwyns.mp4\". This can be useful for analyzing the quality and characteristics of the fake videos generated from a specific original video.","metadata":{"id":"90gMNF5vHKI9"}},{"cell_type":"code","source":"same_original_fake_train_sample_video = list(meta_train_df.loc[meta_train_df.original=='qeumxirsme.mp4'].index)\ndisplay_image_from_video_list(same_original_fake_train_sample_video)","metadata":{"_kg_hide-input":true,"id":"Iy8ofbgtDwdq","execution":{"iopub.status.busy":"2023-11-04T13:53:35.801951Z","iopub.execute_input":"2023-11-04T13:53:35.802229Z","iopub.status.idle":"2023-11-04T13:53:38.278511Z","shell.execute_reply.started":"2023-11-04T13:53:35.802182Z","shell.execute_reply":"2023-11-04T13:53:38.277402Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code first creates a new list named \"same_original_fake_train_sample_video\" by using pandas to select rows from the \"meta_train_df\" DataFrame where the \"original\" column is equal to \"qeumxirsme.mp4\". In this case, the \"original\" column represents the name of the original video file from which the corresponding fake video was generated.\n\nThe code then calls the \"display_image_from_video_list\" function, passing the \"same_original_fake_train_sample_video\" list as an argument. The purpose of the function is to display the first frame of each video file in the list as an image.\n\nTherefore, the overall purpose of the code is to display the first frame of each fake video file in the training set of the metadata DataFrame that was generated from the original video file named \"qeumxirsme.mp4\". This can be useful for analyzing the quality and characteristics of the fake videos generated from a specific original video.","metadata":{"id":"UP5B1d4MHQEG"}},{"cell_type":"code","source":"same_original_fake_train_sample_video = list(meta_train_df.loc[meta_train_df.original=='kgbkktcjxf.mp4'].index)\ndisplay_image_from_video_list(same_original_fake_train_sample_video)","metadata":{"_kg_hide-input":true,"id":"Wi7Vu43KDwdr","execution":{"iopub.status.busy":"2023-11-04T13:53:38.279943Z","iopub.execute_input":"2023-11-04T13:53:38.280217Z","iopub.status.idle":"2023-11-04T13:53:41.192584Z","shell.execute_reply.started":"2023-11-04T13:53:38.280171Z","shell.execute_reply":"2023-11-04T13:53:41.191679Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code first creates a new list named \"same_original_fake_train_sample_video\" by using pandas to select rows from the \"meta_train_df\" DataFrame where the \"original\" column is equal to \"kgbkktcjxf.mp4\". In this case, the \"original\" column represents the name of the original video file from which the corresponding fake video was generated.\n\nThe code then calls the \"display_image_from_video_list\" function, passing the \"same_original_fake_train_sample_video\" list as an argument. The purpose of the function is to display the first frame of each video file in the list as an image.\n\nTherefore, the overall purpose of the code is to display the first frame of each fake video file in the training set of the metadata DataFrame that was generated from the original video file named \"kgbkktcjxf.mp4\". This can be useful for analyzing the quality and characteristics of the fake videos generated from a specific original video.","metadata":{"id":"32IvgqOeHWJ0"}},{"cell_type":"markdown","source":"## Test video files\n\nLet's also look to few of the test data files.","metadata":{"id":"mRoDV2gtDwdr"}},{"cell_type":"code","source":"test_videos = pd.DataFrame(list(os.listdir(os.path.join(DATA_FOLDER, TEST_FOLDER))), columns=['video'])","metadata":{"_kg_hide-input":true,"id":"UzgOP73PDwdr","execution":{"iopub.status.busy":"2023-11-04T13:53:41.194471Z","iopub.execute_input":"2023-11-04T13:53:41.195073Z","iopub.status.idle":"2023-11-04T13:53:41.202602Z","shell.execute_reply.started":"2023-11-04T13:53:41.195015Z","shell.execute_reply":"2023-11-04T13:53:41.201707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code creates a new pandas DataFrame called \"test_videos\". The DataFrame is constructed by first using the \"os.path.join\" method to create a file path to the \"test\" folder in the data directory, and then using the \"os.listdir\" method to list all files in that folder.\n\nThe list of files is passed to the pandas \"DataFrame\" constructor as an argument, which creates a new DataFrame with a single column named \"video\". Each row in the DataFrame represents a single video file in the \"test\" folder.\n\nThe code also uses the \"columns\" parameter to explicitly name the column in the DataFrame as \"video\".\n\nTherefore, the overall purpose of the code is to create a pandas DataFrame containing a list of all video files in the \"test\" folder of the data directory, allowing for easy processing and analysis of the test data.","metadata":{"id":"VH8oLa-GHb5k"}},{"cell_type":"code","source":"test_videos.head()","metadata":{"_kg_hide-input":true,"id":"AoppDEL6Dwdr","execution":{"iopub.status.busy":"2023-11-04T13:53:41.204183Z","iopub.execute_input":"2023-11-04T13:53:41.204501Z","iopub.status.idle":"2023-11-04T13:53:41.222198Z","shell.execute_reply.started":"2023-11-04T13:53:41.204446Z","shell.execute_reply":"2023-11-04T13:53:41.221193Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code calls the \"head\" method on the \"test_videos\" DataFrame, which returns the first 5 rows of the DataFrame.\n\nThis code is useful for quickly inspecting the contents of the \"test_videos\" DataFrame to ensure that the video files have been loaded correctly and that the DataFrame has the expected structure.\n\nThe output of the \"head\" method will display the first 5 rows of the \"test_videos\" DataFrame, showing the filenames of the first 5 video files in the \"test\" folder of the data directory.","metadata":{"id":"B8Rs9MPvHiuj"}},{"cell_type":"markdown","source":"Let's visualize now one of the videos.","metadata":{"id":"UeZN4Pk7Dwdr"}},{"cell_type":"code","source":"display_image_from_video(os.path.join(DATA_FOLDER, TEST_FOLDER, test_videos.iloc[0].video))","metadata":{"_kg_hide-input":true,"id":"k6BObOcGDwds","execution":{"iopub.status.busy":"2023-11-04T13:53:41.223719Z","iopub.execute_input":"2023-11-04T13:53:41.224022Z","iopub.status.idle":"2023-11-04T13:53:41.744720Z","shell.execute_reply.started":"2023-11-04T13:53:41.223974Z","shell.execute_reply":"2023-11-04T13:53:41.743363Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code calls the \"display_image_from_video\" function with a single argument, which is a string representing the full file path of a video file. The file path is created by joining together the \"DATA_FOLDER\", \"TEST_FOLDER\", and the name of the first video file in the \"test_videos\" DataFrame using the \"os.path.join\" method.\n\nThe \"test_videos.iloc[0]\" code is used to select the first row of the \"test_videos\" DataFrame, and the \".video\" attribute is used to extract the filename of the video file from that row. The resulting filename is then included in the file path to the video file.\n\nThe purpose of the \"display_image_from_video\" function is to display the first frame of the specified video file as an image. Therefore, the overall purpose of the code is to display the first frame of the first video file in the \"test\" folder of the data directory, allowing for easy inspection of the content and quality of the video.","metadata":{"id":"z3v3AlgEHypJ"}},{"cell_type":"markdown","source":"Let's look to some more videos from test set.","metadata":{"id":"ZGlDJDo2Dwds"}},{"cell_type":"code","source":"display_image_from_video_list(test_videos.sample(6).video, TEST_FOLDER)","metadata":{"_kg_hide-input":true,"id":"bkv0bvOXDwds","execution":{"iopub.status.busy":"2023-11-04T13:53:41.746678Z","iopub.execute_input":"2023-11-04T13:53:41.747319Z","iopub.status.idle":"2023-11-04T13:53:44.363424Z","shell.execute_reply.started":"2023-11-04T13:53:41.746989Z","shell.execute_reply":"2023-11-04T13:53:44.362493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code calls the \"display_image_from_video_list\" function with two arguments. The first argument is created by calling the \"sample\" method on the \"test_videos\" DataFrame with an argument of 6, which randomly selects 6 rows from the DataFrame. The \".video\" attribute is then used to extract the filenames of the video files from those 6 rows.\n\nThe resulting list of video filenames is passed as the first argument to the \"display_image_from_video_list\" function. The second argument to the function is the name of the \"test\" folder in the data directory.\n\nThe purpose of the \"display_image_from_video_list\" function is to display the first frame of each specified video file as an image. Therefore, the overall purpose of the code is to display the first frames of 6 randomly selected video files from the \"test\" folder of the data directory, allowing for easy inspection of the content and quality of the videos.","metadata":{"id":"kwBzrw8AH6n2"}},{"cell_type":"markdown","source":"# <a id='5'>Face detection</a>  \n\nFrom [5] ([Face Detection using OpenCV](https://www.kaggle.com/serkanpeldek/face-detection-with-opencv)) by [@serkanpeldek](https://www.kaggle.com/serkanpeldek) we got and slightly modified the functions to extract face, profile face, eyes and smile.  \n\nThe class ObjectDetector initialize the cascade classifier (using the imported resource). The function **detect** uses a method of the CascadeClassifier to detect objects into images - in this case the face, eye, smile or profile face.","metadata":{"id":"KaIDF2oRDwds"}},{"cell_type":"code","source":"class ObjectDetector():\n    '''\n    Class for Object Detection\n    '''\n    def __init__(self,object_cascade_path):\n        '''\n        param: object_cascade_path - path for the *.xml defining the parameters for {face, eye, smile, profile}\n        detection algorithm\n        source of the haarcascade resource is: https://github.com/opencv/opencv/tree/master/data/haarcascades\n        '''\n\n        self.objectCascade=cv.CascadeClassifier(object_cascade_path)\n\n\n    def detect(self, image, scale_factor=1.3,\n               min_neighbors=5,\n               min_size=(20,20)):\n        '''\n        Function return rectangle coordinates of object for given image\n        param: image - image to process\n        param: scale_factor - scale factor used for object detection\n        param: min_neighbors - minimum number of parameters considered during object detection\n        param: min_size - minimum size of bounding box for object detected\n        '''\n        rects=self.objectCascade.detectMultiScale(image,\n                                                scaleFactor=scale_factor,\n                                                minNeighbors=min_neighbors,\n                                                minSize=min_size)\n        return rects","metadata":{"_kg_hide-input":true,"id":"729bEvkVDwds","execution":{"iopub.status.busy":"2023-11-04T13:53:44.365132Z","iopub.execute_input":"2023-11-04T13:53:44.365906Z","iopub.status.idle":"2023-11-04T13:53:44.376427Z","shell.execute_reply.started":"2023-11-04T13:53:44.365799Z","shell.execute_reply":"2023-11-04T13:53:44.375370Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code defines a Python class called \"ObjectDetector\". The class has a constructor method called \"init\", which takes a single argument \"object_cascade_path\", which is the path to an XML file that defines the parameters for object detection algorithm.\n\nThe class also has a method called \"detect\", which takes an image and three optional arguments: \"scale_factor\", \"min_neighbors\", and \"min_size\". The method applies object detection on the given image using the object detection algorithm defined by the XML file specified in the constructor. The method then returns the rectangle coordinates of the detected object in the form of a list of tuples. The optional arguments are used to control the detection parameters such as the minimum size of the bounding box and the minimum number of neighbors required to identify an object.\n\nTherefore, the overall purpose of the \"ObjectDetector\" class is to provide an interface for detecting objects in images using the specified object detection algorithm. The class can be instantiated with a specific XML file and then used to apply object detection on images using the \"detect\" method with customizable detection parameters.","metadata":{"id":"vWqEA-kPIIXn"}},{"cell_type":"markdown","source":"We load the resources for frontal face, eye, smile and profile face detection.  \n\nThen we initialize the `ObjectDetector` objects defined above with the respective resources, to use CascadeClassfier for each specific task.","metadata":{"id":"uWrZLiDcDwdt"}},{"cell_type":"code","source":"#Frontal face, profile, eye and smile  haar cascade loaded\nfrontal_cascade_path= os.path.join(FACE_DETECTION_FOLDER,'haarcascade_frontalface_default.xml')\neye_cascade_path= os.path.join(FACE_DETECTION_FOLDER,'haarcascade_eye.xml')\nprofile_cascade_path= os.path.join(FACE_DETECTION_FOLDER,'haarcascade_profileface.xml')\nsmile_cascade_path= os.path.join(FACE_DETECTION_FOLDER,'haarcascade_smile.xml')\n\n#Detector object created\n# frontal face\nfd=ObjectDetector(frontal_cascade_path)\n# eye\ned=ObjectDetector(eye_cascade_path)\n# profile face\npd=ObjectDetector(profile_cascade_path)\n# smile\nsd=ObjectDetector(smile_cascade_path)","metadata":{"_kg_hide-input":true,"id":"q0T521CnDwdt","execution":{"iopub.status.busy":"2023-11-04T13:53:44.378131Z","iopub.execute_input":"2023-11-04T13:53:44.378689Z","iopub.status.idle":"2023-11-04T13:53:44.507045Z","shell.execute_reply.started":"2023-11-04T13:53:44.378632Z","shell.execute_reply":"2023-11-04T13:53:44.506173Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code loads four XML files that define the parameters for object detection algorithms: \"haarcascade_frontalface_default.xml\", \"haarcascade_eye.xml\", \"haarcascade_profileface.xml\", and \"haarcascade_smile.xml\". These XML files contain information such as the size and shape of the objects to be detected and the features that the algorithm should look for to identify those objects.\n\nThe code then creates four instances of the \"ObjectDetector\" class, one for each type of object to be detected: frontal faces, eyes, profile faces, and smiles. Each instance is created with the path to the corresponding XML file as an argument, allowing the object detection algorithm to be customized for each type of object.\n\nTherefore, the overall purpose of the code is to load the XML files and create instances of the \"ObjectDetector\" class for each type of object to be detected, allowing for object detection on images using these specific detection algorithms.","metadata":{"id":"vVz9YRZCISXA"}},{"cell_type":"markdown","source":"We also define a function for detection and display of all these specific objects.  \n\nThe function call the **detect** method of the **ObjectDetector** object. For each object we are using a different shape and color, as following:\n* Frontal face: green rectangle;  \n* Eye: red circle;  \n* Smile: red rectangle;  \n* Profile face: blue rectangle.  \n\nNote: due to a huge amount of false positive, we deactivate for now the smile detector.","metadata":{"id":"OIldwawDDwdt"}},{"cell_type":"code","source":"def detect_objects(image, scale_factor, min_neighbors, min_size):\n    '''\n    Objects detection function\n    Identify frontal face, eyes, smile and profile face and display the detected objects over the image\n    param: image - the image extracted from the video\n    param: scale_factor - scale factor parameter for `detect` function of ObjectDetector object\n    param: min_neighbors - min neighbors parameter for `detect` function of ObjectDetector object\n    param: min_size - minimum size parameter for f`detect` function of ObjectDetector object\n    '''\n    \n    image_gray=cv.cvtColor(image, cv.COLOR_BGR2GRAY)\n\n\n    eyes=ed.detect(image_gray,\n                   scale_factor=scale_factor,\n                   min_neighbors=min_neighbors,\n                   min_size=(int(min_size[0]/2), int(min_size[1]/2)))\n\n    for x, y, w, h in eyes:\n        #detected eyes shown in color image\n        cv.circle(image,(int(x+w/2),int(y+h/2)),(int((w + h)/4)),(0, 0,255),3)\n\n    profiles=pd.detect(image_gray,\n                   scale_factor=scale_factor,\n                   min_neighbors=min_neighbors,\n                   min_size=min_size)\n\n    for x, y, w, h in profiles:\n        #detected profiles shown in color image\n        cv.rectangle(image,(x,y),(x+w, y+h),(255, 0,0),3)\n\n    faces=fd.detect(image_gray,\n                   scale_factor=scale_factor,\n                   min_neighbors=min_neighbors,\n                   min_size=min_size)\n\n    for x, y, w, h in faces:\n        #detected faces shown in color image\n        cv.rectangle(image,(x,y),(x+w, y+h),(0, 255,0),3)\n\n    # image\n    fig = plt.figure(figsize=(10,10))\n    ax = fig.add_subplot(111)\n    image = cv.cvtColor(image, cv.COLOR_BGR2RGB)\n    ax.imshow(image)","metadata":{"_kg_hide-input":true,"id":"6llAE2ciDwdu","execution":{"iopub.status.busy":"2023-11-04T13:53:44.508650Z","iopub.execute_input":"2023-11-04T13:53:44.508994Z","iopub.status.idle":"2023-11-04T13:53:44.525547Z","shell.execute_reply.started":"2023-11-04T13:53:44.508913Z","shell.execute_reply":"2023-11-04T13:53:44.524846Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code defines a Python function called \"detect_objects\". The function takes an image as an argument, along with three optional arguments: \"scale_factor\", \"min_neighbors\", and \"min_size\". The function applies object detection algorithms to the image using four \"ObjectDetector\" instances that were created in the previous code block. The algorithms are used to detect frontal faces, eyes, smiles, and profile faces in the image.\n\nThe detected objects are then highlighted on the original image with different colored rectangles or circles, depending on the type of object detected. Specifically, detected eyes are shown in red, detected profiles in blue, and detected faces in green.\n\nFinally, the function displays the modified image with detected objects overlaid using matplotlib.\n\nTherefore, the overall purpose of the \"detect_objects\" function is to apply object detection algorithms to an image and visually display the results by highlighting the detected objects with different colored rectangles or circles.","metadata":{"id":"RCADC5z1Id_6"}},{"cell_type":"markdown","source":"The following function extracts an image from a video and then call the function that extracts the face rectangle from the image and display the rectangle above the image.","metadata":{"id":"fqOVenzxDwdu"}},{"cell_type":"code","source":"def extract_image_objects(video_file, video_set_folder=TRAIN_SAMPLE_FOLDER):\n    '''\n    Extract one image from the video and then perform face/eyes/smile/profile detection on the image\n    param: video_file - the video from which to extract the image from which we extract the face\n    '''\n    video_path = os.path.join(DATA_FOLDER, video_set_folder,video_file)\n    capture_image = cv.VideoCapture(video_path) \n    ret, frame = capture_image.read()\n    #frame = cv.cvtColor(frame, cv.COLOR_BGR2RGB)\n    detect_objects(image=frame, \n            scale_factor=1.3, \n            min_neighbors=5, \n            min_size=(50, 50))  \n  ","metadata":{"_kg_hide-input":true,"id":"UG7TJ2_LDwdv","execution":{"iopub.status.busy":"2023-11-04T13:53:44.530403Z","iopub.execute_input":"2023-11-04T13:53:44.530846Z","iopub.status.idle":"2023-11-04T13:53:44.543549Z","shell.execute_reply.started":"2023-11-04T13:53:44.530789Z","shell.execute_reply":"2023-11-04T13:53:44.542616Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code defines a Python function called \"extract_image_objects\". The function takes two arguments: \"video_file\", which is the name of the video file from which to extract the image, and \"video_set_folder\", which specifies the folder containing the video file (by default, it is set to TRAIN_SAMPLE_FOLDER).\n\nThe function reads the video file specified by \"video_file\" using the OpenCV library and extracts a single image from the video. This image is then passed to the \"detect_objects\" function defined earlier, which applies object detection algorithms to the image to detect faces, eyes, smiles, and profile faces. The detected objects are then highlighted on the image, and the modified image is displayed using matplotlib.\n\nTherefore, the overall purpose of the \"extract_image_objects\" function is to extract an image from a specified video file, apply object detection algorithms to the image to detect faces, eyes, smiles, and profile faces, and display the modified image with detected objects overlaid.","metadata":{"id":"RS-O8Kq_Ij1n"}},{"cell_type":"markdown","source":"We apply the function for face detection for a selection of images from train sample videos.","metadata":{"id":"DKPnCInRDwdv"}},{"cell_type":"code","source":"same_original_fake_train_sample_video = list(meta_train_df.loc[meta_train_df.original=='kgbkktcjxf.mp4'].index)\nfor video_file in same_original_fake_train_sample_video[1:4]:\n    print(video_file)\n    extract_image_objects(video_file)","metadata":{"_kg_hide-input":true,"id":"YvtKwo9zDwdv","execution":{"iopub.status.busy":"2023-11-04T13:53:44.544896Z","iopub.execute_input":"2023-11-04T13:53:44.545289Z","iopub.status.idle":"2023-11-04T13:53:46.905116Z","shell.execute_reply.started":"2023-11-04T13:53:44.545247Z","shell.execute_reply":"2023-11-04T13:53:46.904376Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code extracts a list of video files with the same original video file (i.e., \"kgbkktcjxf.mp4\") from the training set, using the \"meta_train_df\" dataframe. It then selects a subset of three video files from the list using list slicing.\n\nThe function \"extract_image_objects\" is then called three times in a loop, with each iteration corresponding to one of the selected video files. For each iteration, the function is called with the current video file as an argument. The function reads the video file, extracts a single image from it, applies object detection algorithms to the image to detect faces, eyes, smiles, and profile faces, highlights the detected objects on the image, and displays the modified image using matplotlib.\n\nOverall, the code loops through a subset of three video files with the same original video and applies object detection algorithms to the images extracted from each of these videos to detect faces, eyes, smiles, and profile faces.","metadata":{"id":"vW5RWec2Iw7R"}},{"cell_type":"code","source":"train_subsample_video = list(meta_train_df.sample(3).index)\nfor video_file in train_subsample_video:\n    print(video_file)\n    extract_image_objects(video_file)","metadata":{"id":"vU8EE6zxDwdv","execution":{"iopub.status.busy":"2023-11-04T13:53:46.906292Z","iopub.execute_input":"2023-11-04T13:53:46.906694Z","iopub.status.idle":"2023-11-04T13:53:49.739802Z","shell.execute_reply.started":"2023-11-04T13:53:46.906650Z","shell.execute_reply":"2023-11-04T13:53:49.738870Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code first selects a random sample of 3 video files from the training set using the \"meta_train_df\" dataframe.\n\nThen, the code loops through each of the selected video files and prints its name. For each iteration, it calls the \"extract_image_objects\" function with the current video file as an argument. The function reads the video file, extracts a single image from it, applies object detection algorithms to the image to detect faces, eyes, smiles, and profile faces, highlights the detected objects on the image, and displays the modified image using matplotlib.\n\nTherefore, the code is basically a demonstration of applying object detection algorithms to random videos from the training set to detect faces, eyes, smiles, and profile faces.","metadata":{"id":"DJa4LYe1I4oe"}},{"cell_type":"markdown","source":"Let's look to a small collection of samples from test videos.","metadata":{"id":"rceizGXEDwdw"}},{"cell_type":"code","source":"subsample_test_videos = list(test_videos.sample(3).video)\nfor video_file in subsample_test_videos:\n    print(video_file)\n    extract_image_objects(video_file, TEST_FOLDER)","metadata":{"_kg_hide-input":true,"id":"Owupb5oeDwdw","execution":{"iopub.status.busy":"2023-11-04T13:53:49.741285Z","iopub.execute_input":"2023-11-04T13:53:49.741566Z","iopub.status.idle":"2023-11-04T13:53:52.264153Z","shell.execute_reply.started":"2023-11-04T13:53:49.741518Z","shell.execute_reply":"2023-11-04T13:53:52.263075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code first selects a random sample of 3 video files from the test set using the \"test_videos\" dataframe.\n\nThen, the code loops through each of the selected video files and prints its name. For each iteration, it calls the \"extract_image_objects\" function with the current video file and \"TEST_FOLDER\" as arguments. The function reads the video file, extracts a single image from it, applies object detection algorithms to the image to detect faces, eyes, smiles, and profile faces, highlights the detected objects on the image, and displays the modified image using matplotlib.\n\nTherefore, the code is basically a demonstration of applying object detection algorithms to random videos from the test set to detect faces, eyes, smiles, and profile faces. ","metadata":{"id":"MSM67zpgI-4k"}},{"cell_type":"markdown","source":"We can observe that in some cases, when the subject is not looking frontaly or when the luminosity is low, the algorithm for face detection is not detecting the face or eyes correctly. Due to a large amount of false positive, we deactivated for now the smile detector.","metadata":{"id":"Pu-b77yTDwdw"}},{"cell_type":"markdown","source":"## Play video files  \n\nFrom [Play video and processing](https://www.kaggle.com/hamditarek/play-video-and-processing) Kernel by [@hamditarek](https://www.kaggle.com/hamditarek) we learned how to play video files in a Kaggle Kernel.  \nLet's look to few fake videos.","metadata":{"id":"XK-0LDWhDwdw"}},{"cell_type":"code","source":"fake_videos = list(meta_train_df.loc[meta_train_df.label=='FAKE'].index)","metadata":{"_kg_hide-input":true,"id":"7elkHp2qDwdw","execution":{"iopub.status.busy":"2023-11-04T13:53:52.265609Z","iopub.execute_input":"2023-11-04T13:53:52.265961Z","iopub.status.idle":"2023-11-04T13:53:52.272157Z","shell.execute_reply.started":"2023-11-04T13:53:52.265896Z","shell.execute_reply":"2023-11-04T13:53:52.271239Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The given code selects all the video files from the training set which have a label of \"FAKE\". It does this by using the \"loc\" method of the \"meta_train_df\" dataframe to select all rows where the \"label\" column has the value \"FAKE\". Then it extracts the index of these selected rows using the \"index\" attribute and converts them to a list using the \"list\" function.\n\nTherefore, the resulting \"fake_videos\" list contains the names of all the video files from the training set that are labeled as \"FAKE\".","metadata":{"id":"r7qKB4ZFJHO7"}},{"cell_type":"code","source":"from IPython.display import HTML\nfrom base64 import b64encode\n\ndef play_video(video_file, subset=TRAIN_SAMPLE_FOLDER):\n    '''\n    Display video\n    param: video_file - the name of the video file to display\n    param: subset - the folder where the video file is located (can be TRAIN_SAMPLE_FOLDER or TEST_Folder)\n    '''\n    video_url = open(os.path.join(DATA_FOLDER, subset,video_file),'rb').read()\n    data_url = \"data:video/mp4;base64,\" + b64encode(video_url).decode()\n    return HTML(\"\"\"<video width=500 controls><source src=\"%s\" type=\"video/mp4\"></video>\"\"\" % data_url)","metadata":{"_kg_hide-input":true,"id":"Bh4km-AwDwdw","execution":{"iopub.status.busy":"2023-11-04T13:53:52.273604Z","iopub.execute_input":"2023-11-04T13:53:52.273949Z","iopub.status.idle":"2023-11-04T13:53:52.286964Z","shell.execute_reply.started":"2023-11-04T13:53:52.273890Z","shell.execute_reply":"2023-11-04T13:53:52.285975Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The play_video function takes in a video_file parameter and a subset parameter. The video_file parameter specifies the name of the video file to display, while the subset parameter specifies the folder where the video file is located. The function opens the video file in binary mode, encodes the contents of the file in Base64, and constructs a data URL that represents the encoded video. The function then returns an HTML video element that plays the video represented by the data URL when it is rendered in the notebook. When the controls attribute is set to True, the video element displays the video playback controls (such as the play/pause, volume, and seek controls).","metadata":{"id":"PBP-UwAxJM7P"}},{"cell_type":"code","source":"play_video(fake_videos[0])","metadata":{"_kg_hide-input":true,"id":"2EZyVdT-Dwdx","execution":{"iopub.status.busy":"2023-11-04T13:53:52.288347Z","iopub.execute_input":"2023-11-04T13:53:52.288631Z","iopub.status.idle":"2023-11-04T13:53:52.788540Z","shell.execute_reply.started":"2023-11-04T13:53:52.288581Z","shell.execute_reply":"2023-11-04T13:53:52.787030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This code displays a video file named fake_videos[0] from the TRAIN_SAMPLE_FOLDER using HTML5 video tag and base64 encoding. The function play_video() takes the video_file as input, and reads the video file as bytes from the path using open(). Then, it encodes the bytes using base64 encoding and creates a data URL of the video with MIME type video/mp4. Finally, it returns an HTML video tag with the source as the data URL of the encoded video, which is then displayed as an embedded video player in the output.","metadata":{"id":"sgWT5G0OJWTH"}},{"cell_type":"code","source":"play_video(fake_videos[1])","metadata":{"_kg_hide-input":true,"id":"RfF2DBEeDwdx","execution":{"iopub.status.busy":"2023-11-04T13:53:52.790009Z","iopub.execute_input":"2023-11-04T13:53:52.790316Z","iopub.status.idle":"2023-11-04T13:53:53.013723Z","shell.execute_reply.started":"2023-11-04T13:53:52.790266Z","shell.execute_reply":"2023-11-04T13:53:53.012325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[2])","metadata":{"_kg_hide-input":true,"id":"C2Uji4x3Dwdx","execution":{"iopub.status.busy":"2023-11-04T13:53:53.015799Z","iopub.execute_input":"2023-11-04T13:53:53.016492Z","iopub.status.idle":"2023-11-04T13:53:53.100649Z","shell.execute_reply.started":"2023-11-04T13:53:53.016417Z","shell.execute_reply":"2023-11-04T13:53:53.096003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[3])","metadata":{"id":"dCOmLm7sDwdx","execution":{"iopub.status.busy":"2023-11-04T13:53:53.102531Z","iopub.execute_input":"2023-11-04T13:53:53.103279Z","iopub.status.idle":"2023-11-04T13:53:53.232225Z","shell.execute_reply.started":"2023-11-04T13:53:53.103072Z","shell.execute_reply":"2023-11-04T13:53:53.231431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[4])","metadata":{"id":"6Bx9u6NjDwdx","execution":{"iopub.status.busy":"2023-11-04T13:53:53.233485Z","iopub.execute_input":"2023-11-04T13:53:53.233958Z","iopub.status.idle":"2023-11-04T13:53:53.452581Z","shell.execute_reply.started":"2023-11-04T13:53:53.233909Z","shell.execute_reply":"2023-11-04T13:53:53.450887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[5])","metadata":{"id":"LcWGuYK4Dwdx","execution":{"iopub.status.busy":"2023-11-04T13:53:53.454295Z","iopub.execute_input":"2023-11-04T13:53:53.454755Z","iopub.status.idle":"2023-11-04T13:53:53.560756Z","shell.execute_reply.started":"2023-11-04T13:53:53.454700Z","shell.execute_reply":"2023-11-04T13:53:53.559766Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[10])","metadata":{"id":"c8l69e_NDwdy","execution":{"iopub.status.busy":"2023-11-04T13:53:53.562069Z","iopub.execute_input":"2023-11-04T13:53:53.562558Z","iopub.status.idle":"2023-11-04T13:53:53.895792Z","shell.execute_reply.started":"2023-11-04T13:53:53.562507Z","shell.execute_reply":"2023-11-04T13:53:53.892236Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[12])","metadata":{"id":"k971YjhEDwdy","execution":{"iopub.status.busy":"2023-11-04T13:53:53.897352Z","iopub.execute_input":"2023-11-04T13:53:53.897759Z","iopub.status.idle":"2023-11-04T13:53:54.041328Z","shell.execute_reply.started":"2023-11-04T13:53:53.897603Z","shell.execute_reply":"2023-11-04T13:53:54.038657Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[15])","metadata":{"id":"btCYWtH1Dwdy","execution":{"iopub.status.busy":"2023-11-04T13:53:54.042581Z","iopub.execute_input":"2023-11-04T13:53:54.042890Z","iopub.status.idle":"2023-11-04T13:53:54.242044Z","shell.execute_reply.started":"2023-11-04T13:53:54.042841Z","shell.execute_reply":"2023-11-04T13:53:54.240389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"play_video(fake_videos[18])","metadata":{"id":"xdv7wVBlDwdy","execution":{"iopub.status.busy":"2023-11-04T13:53:54.243853Z","iopub.execute_input":"2023-11-04T13:53:54.244350Z","iopub.status.idle":"2023-11-04T13:53:54.384085Z","shell.execute_reply.started":"2023-11-04T13:53:54.244296Z","shell.execute_reply":"2023-11-04T13:53:54.382489Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"From visual inspection of these fakes videos, in some cases is very easy to spot the anomalies created when engineering the deep fake, in some cases is more difficult.","metadata":{"id":"D7psN3ALDwdy"}},{"cell_type":"markdown","source":"# <a id=\"6\">Resources</a>  \n\nThis resources list is not exhaustive, it provides just a starting point for Kagglers that would like to join this competition in order to learn, like myself.    \n\nI provide here technical articles links, small blog articles, links to Github projects and some inspirational Kaggle Kernels about DeepFake, DCGANs and GANs.   \n\n**Blogs, Technical Articles**\n\n* Towards Data Science DeepFakes subject selection: [DeepFakes](https://towardsdatascience.com/tagged/deepfakes)   \n\n* An introduction to DeepFakes from Towards Data Science: [Deepfakes: The Ugly, and The Good](https://towardsdatascience.com/deepfakes-the-ugly-and-the-good-49115643d8dd)  \n\n* Introduction to GAN from Towards Data Science (with code): [GANs from Scratch 1: A deep introduction. With code in PyTorch and TensorFlow](https://medium.com/ai-society/gans-from-scratch-1-a-deep-introduction-with-code-in-pytorch-and-tensorflow-cb03cdcdba0f)  \n\n* A cGAN introduction from Mastering Data Science (with code): [How to Develop a Conditional GAN (cGAN) From Scratch](https://machinelearningmastery.com/how-to-develop-a-conditional-generative-adversarial-network-from-scratch/)   \n\n* An introduction to GANs on Analytics Vidhya,  [GANs — A Brief Introduction to Generative Adversarial Networks](https://medium.com/analytics-vidhya/gans-a-brief-introduction-to-generative-adversarial-networks-f06216c7200e?)\n\n**Tutorials**\n\n* A nice tutorial for GAN using Pytorch: [DCGAN Faces Tutorial](https://pytorch.org/tutorials/beginner/dcgan_faces_tutorial.html)  \n\n* A tutorial for GANs: [A Beginner's Guide to Generative Adversarial Networks (GANs)](https://pathmind.com/wiki/generative-adversarial-network-gan)  \n\n* Deep Convolutional Generative Adversial Networks, Tensorflow, [DCGAN](https://www.tensorflow.org/tutorials/generative/dcgan)    \n\n* A tutorial for video and image colorization and resolution improvement using fast.ai and PyTorch recommended by [@init27](https://www.kaggle.com/init27), [Decrappification, DeOldification, and Super Resolution](https://www.fast.ai/2019/05/03/decrappify/)  \n\n* A tutorial from OpenCV for face detection using Cascade Classifiers, [Cascade Classifier](https://docs.opencv.org/3.4/db/d28/tutorial_cascade_classifier.html)  \n\n**Kaggle Kernels**\n\n* A very good intro to GAN, by [@nanashi](http://kaggle/nanashi): [GAN Introduction](https://www.kaggle.com/jesucristo/gan-introduction)    \n\n* A GAN developed for Dog face generation competition by [@cdeotte](http://kaggle/cdeotte): [Dog Memorizer GAN](https://www.kaggle.com/cdeotte/dog-memorizer-gan)    \n\n* A Kernel for Face detection using OpenCV with Haarcascade, by [@serkanpeldek](https://www.kaggle.com/serkanpeldek), [Face Detection with OpenCV](https://www.kaggle.com/serkanpeldek/face-detection-with-opencv)   \n\n* Play video and processing, by [@hamditarek](https://www.kaggle.com/hamditarek),  https://www.kaggle.com/hamditarek/play-video-and-processing   \n\n\n**Github repos**\n\n* Github topic for DeeFakes: [deepfakes](https://github.com/topics/)  \n\n* A Github project using Pytorch: [Faceswap-Deepfake-Pytorch](https://github.com/Oldpan/Faceswap-Deepfake-Pytorch)   \n\n* A Github project for GAN with PyTorch: [PyTorch-GAN](https://github.com/eriklindernoren/PyTorch-GAN)  \n\n* A Github recommended by [@shwetagoyal4](https://www.kaggle.com/shwetagoyal4), [Generative-model-using-PyTorch](https://github.com/Shwetago/Generative-model-using-PyTorch)","metadata":{"id":"ODcKsJM2Dwdy"}},{"cell_type":"markdown","source":"# <a id=\"7\">References</a>\n\n[1] Deepfake, Wikipedia, https://en.wikipedia.org/wiki/Deepfake  \n[2] Google DeepFake Database, Endgadget, https://www.engadget.com/2019/09/25/google-deepfake-database/  \n[3] A quick look at the first frame of each video,  https://www.kaggle.com/brassmonkey381/a-quick-look-at-the-first-frame-of-each-video  \n[4] Basic EDA Face Detection, split video, ROI, https://www.kaggle.com/marcovasquez/basic-eda-face-detection-split-video-roi  \n[5] Face Detection with OpenCV, https://www.kaggle.com/serkanpeldek/face-detection-with-opencv   \n[6] Play video and processing, https://www.kaggle.com/hamditarek/play-video-and-processing/\n","metadata":{"_cell_guid":"79c7e3d0-c299-4dcb-8224-4455121ee9b0","_uuid":"d629ff2d2480ee46fbb7e2d37f6b5fab8052498a","id":"x26iIaqVDwdz"}}]}