{
  "id": 128954,
  "title": "Other useful datasets",
  "url": "/competitions/deepfake-detection-challenge/discussion/128954",
  "author_name": "Hieu Phung",
  "post_date": "2020-02-04T14:51:51.957000",
  "votes": 71,
  "comment_count": 80,
  "views": 0,
  "content": "<p>I've created some datasets that include all detectable faces of all videos in each part of the full dataset. Kaggle and the host expected and encouraged us to train our models outside of Kaggle’s notebooks environment; however, for someone who prefers to stick to Kaggle's kernels, these preprocessed datasets would help a lot 😄.</p>\n\n<p>The whole process to create and upload these datasets is time-consuming, so I'll gradually upload the rest for the next days; here are some completed ones:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-0\">Deepfake Detection - Faces - Part 0_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-1\">Deepfake Detection - Faces - Part 0_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-0\">Deepfake Detection - Faces - Part 1_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-1\">Deepfake Detection - Faces - Part 1_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-0\">Deepfake Detection - Faces - Part 2_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-1\">Deepfake Detection - Faces - Part 2_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-2\">Deepfake Detection - Faces - Part 2_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-0\">Deepfake Detection - Faces - Part 3_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-1\">Deepfake Detection - Faces - Part 3_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-0\">Deepfake Detection - Faces - Part 4_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-1\">Deepfake Detection - Faces - Part 4_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-2\">Deepfake Detection - Faces - Part 4_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-0\">Deepfake Detection - Faces - Part 5_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-1\">Deepfake Detection - Faces - Part 5_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-2\">Deepfake Detection - Faces - Part 5_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-3\">Deepfake Detection - Faces - Part 5_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-0\">Deepfake Detection - Faces - Part 6_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-1\">Deepfake Detection - Faces - Part 6_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-2\">Deepfake Detection - Faces - Part 6_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-3\">Deepfake Detection - Faces - Part 6_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-4\">Deepfake Detection - Faces - Part 6_4</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-0\">Deepfake Detection - Faces - Part 7_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-1\">Deepfake Detection - Faces - Part 7_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-2\">Deepfake Detection - Faces - Part 7_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-3\">Deepfake Detection - Faces - Part 7_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-0\">Deepfake Detection - Faces - Part 8_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-1\">Deepfake Detection - Faces - Part 8_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-2\">Deepfake Detection - Faces - Part 8_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-0\">Deepfake Detection - Faces - Part 9_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-1\">Deepfake Detection - Faces - Part 9_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-2\">Deepfake Detection - Faces - Part 9_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-0\">Deepfake Detection - Faces - Part 10_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-1\">Deepfake Detection - Faces - Part 10_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-2\">Deepfake Detection - Faces - Part 10_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-3\">Deepfake Detection - Faces - Part 10_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-4\">Deepfake Detection - Faces - Part 10_4</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-0\">Deepfake Detection - Faces - Part 11_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-1\">Deepfake Detection - Faces - Part 11_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-2\">Deepfake Detection - Faces - Part 11_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-0\">Deepfake Detection - Faces - Part 12_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-1\">Deepfake Detection - Faces - Part 12_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-2\">Deepfake Detection - Faces - Part 12_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-0\">Deepfake Detection - Faces - Part 13_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-1\">Deepfake Detection - Faces - Part 13_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-2\">Deepfake Detection - Faces - Part 13_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-3\">Deepfake Detection - Faces - Part 13_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-4\">Deepfake Detection - Faces - Part 13_4</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-0\">Deepfake Detection - Faces - Part 14_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-1\">Deepfake Detection - Faces - Part 14_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-2\">Deepfake Detection - Faces - Part 14_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-3\">Deepfake Detection - Faces - Part 14_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-0\">Deepfake Detection - Faces - Part 15_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-1\">Deepfake Detection - Faces - Part 15_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-2\">Deepfake Detection - Faces - Part 15_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-sample\">Deepfake Detection - Faces - Sample</a>\n<em>Updating...</em></li>\n</ul>\n\n<p>How I create these datasets? Let's check out this demo kernel 👉 <a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor\">\n<em>Deepfake Detection - Face Extractor</em></a>.</p>\n\n<hr>\n\n<p>Because of the big size of each part in the original full dataset, I have to break all extracted faces in each part into 2 or 3 smaller parts before zipping them, e.g. the name <code>deepfake-detection-faces-part-1-0</code> means it created from <code>part 1</code> of the full dataset and is the <code>first split</code>.</p>\n\n<p>The name of each face image corresponds to the index of the frame that this face appears, plus a suffix _2, or _3, etc. if the number of faces in a frame is greater than 1.</p>\n\n<p>In each dataset, I also attach a metadata.csv file which stores all information needed. The format of each file will look like this:\n|     |    filename    | split |    original    | label |\n|:---:|:--------------:|:-----:|:--------------:|:-----:|\n|  0  | aagfhgtpmv.mp4 | train | vudstovrck.mp4 | FAKE  |\n|  1  | aapnvogymq.mp4 | train | jdubbvfswz.mp4 | FAKE  |\n|  2  | abarnvbtwb.mp4 | train |                | REAL  |\n| ... | ...            | ...   | ...            | ...   |</p>\n\n<hr>\n\n<p>Want something to get started using these datasets, let see <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-training\"><em>DFDC-Multiface-Training</em></a> &amp; <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-inference\"><em>DFDC-Multiface-Inference</em></a>.</p>\n\n<p>Don't know how to load and merge multiple Kaggle datasets at once in a kernel, let check this <a href=\"https://www.kaggle.com/phunghieu/loading-merging-multiple-kaggle-datasets-demo\"><em>demo</em></a>.</p>\n\n<p>Here are a few lines of code to demonstrate how to load and prepare these data for the training process:</p>\n\n<p><code># Get path of metadata.csv</code>\n<code>metadata_path = os.path.join(TRAIN_DIR, 'metadata.csv')</code></p>\n\n<p><code># Create DataFrame from metadata.csv</code>\n<code>train_df = pd.read_csv(metadata_path)</code>\n<code>train_df['label'].replace({'FAKE': 1, 'REAL': 0}, inplace=True)</code></p>\n\n<p><code>X = train_df['filename'].to_numpy()</code>\n<code>y = train_df['label'].to_numpy()</code></p>\n\n<p><code># Split the dataset</code>\n<code>X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.3, random_state=123, stratify=y)</code></p>\n\n<hr>\n\n<p>Consider to upvote these datasets if you think they are worth using 😉, this will motivate me much to continue this dull work 💪, thanks!!!</p>",
  "messages": [
    {
      "id": 736780,
      "postDate": "2020-02-04T14:51:51.957Z",
      "content": "<p>I've created some datasets that include all detectable faces of all videos in each part of the full dataset. Kaggle and the host expected and encouraged us to train our models outside of Kaggle’s notebooks environment; however, for someone who prefers to stick to Kaggle's kernels, these preprocessed datasets would help a lot 😄.</p>\n\n<p>The whole process to create and upload these datasets is time-consuming, so I'll gradually upload the rest for the next days; here are some completed ones:</p>\n\n<ul>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-0\">Deepfake Detection - Faces - Part 0_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-1\">Deepfake Detection - Faces - Part 0_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-0\">Deepfake Detection - Faces - Part 1_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-1\">Deepfake Detection - Faces - Part 1_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-0\">Deepfake Detection - Faces - Part 2_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-1\">Deepfake Detection - Faces - Part 2_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-2\">Deepfake Detection - Faces - Part 2_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-0\">Deepfake Detection - Faces - Part 3_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-1\">Deepfake Detection - Faces - Part 3_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-0\">Deepfake Detection - Faces - Part 4_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-1\">Deepfake Detection - Faces - Part 4_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-2\">Deepfake Detection - Faces - Part 4_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-0\">Deepfake Detection - Faces - Part 5_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-1\">Deepfake Detection - Faces - Part 5_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-2\">Deepfake Detection - Faces - Part 5_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-3\">Deepfake Detection - Faces - Part 5_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-0\">Deepfake Detection - Faces - Part 6_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-1\">Deepfake Detection - Faces - Part 6_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-2\">Deepfake Detection - Faces - Part 6_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-3\">Deepfake Detection - Faces - Part 6_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-4\">Deepfake Detection - Faces - Part 6_4</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-0\">Deepfake Detection - Faces - Part 7_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-1\">Deepfake Detection - Faces - Part 7_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-2\">Deepfake Detection - Faces - Part 7_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-3\">Deepfake Detection - Faces - Part 7_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-0\">Deepfake Detection - Faces - Part 8_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-1\">Deepfake Detection - Faces - Part 8_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-2\">Deepfake Detection - Faces - Part 8_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-0\">Deepfake Detection - Faces - Part 9_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-1\">Deepfake Detection - Faces - Part 9_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-2\">Deepfake Detection - Faces - Part 9_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-0\">Deepfake Detection - Faces - Part 10_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-1\">Deepfake Detection - Faces - Part 10_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-2\">Deepfake Detection - Faces - Part 10_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-3\">Deepfake Detection - Faces - Part 10_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-4\">Deepfake Detection - Faces - Part 10_4</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-0\">Deepfake Detection - Faces - Part 11_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-1\">Deepfake Detection - Faces - Part 11_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-2\">Deepfake Detection - Faces - Part 11_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-0\">Deepfake Detection - Faces - Part 12_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-1\">Deepfake Detection - Faces - Part 12_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-2\">Deepfake Detection - Faces - Part 12_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-0\">Deepfake Detection - Faces - Part 13_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-1\">Deepfake Detection - Faces - Part 13_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-2\">Deepfake Detection - Faces - Part 13_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-3\">Deepfake Detection - Faces - Part 13_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-4\">Deepfake Detection - Faces - Part 13_4</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-0\">Deepfake Detection - Faces - Part 14_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-1\">Deepfake Detection - Faces - Part 14_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-2\">Deepfake Detection - Faces - Part 14_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-3\">Deepfake Detection - Faces - Part 14_3</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-0\">Deepfake Detection - Faces - Part 15_0</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-1\">Deepfake Detection - Faces - Part 15_1</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-2\">Deepfake Detection - Faces - Part 15_2</a></li>\n<li><a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-faces-sample\">Deepfake Detection - Faces - Sample</a>\n<em>Updating...</em></li>\n</ul>\n\n<p>How I create these datasets? Let's check out this demo kernel 👉 <a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor\">\n<em>Deepfake Detection - Face Extractor</em></a>.</p>\n\n<hr>\n\n<p>Because of the big size of each part in the original full dataset, I have to break all extracted faces in each part into 2 or 3 smaller parts before zipping them, e.g. the name <code>deepfake-detection-faces-part-1-0</code> means it created from <code>part 1</code> of the full dataset and is the <code>first split</code>.</p>\n\n<p>The name of each face image corresponds to the index of the frame that this face appears, plus a suffix _2, or _3, etc. if the number of faces in a frame is greater than 1.</p>\n\n<p>In each dataset, I also attach a metadata.csv file which stores all information needed. The format of each file will look like this:\n|     |    filename    | split |    original    | label |\n|:---:|:--------------:|:-----:|:--------------:|:-----:|\n|  0  | aagfhgtpmv.mp4 | train | vudstovrck.mp4 | FAKE  |\n|  1  | aapnvogymq.mp4 | train | jdubbvfswz.mp4 | FAKE  |\n|  2  | abarnvbtwb.mp4 | train |                | REAL  |\n| ... | ...            | ...   | ...            | ...   |</p>\n\n<hr>\n\n<p>Want something to get started using these datasets, let see <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-training\"><em>DFDC-Multiface-Training</em></a> &amp; <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-inference\"><em>DFDC-Multiface-Inference</em></a>.</p>\n\n<p>Don't know how to load and merge multiple Kaggle datasets at once in a kernel, let check this <a href=\"https://www.kaggle.com/phunghieu/loading-merging-multiple-kaggle-datasets-demo\"><em>demo</em></a>.</p>\n\n<p>Here are a few lines of code to demonstrate how to load and prepare these data for the training process:</p>\n\n<p><code># Get path of metadata.csv</code>\n<code>metadata_path = os.path.join(TRAIN_DIR, 'metadata.csv')</code></p>\n\n<p><code># Create DataFrame from metadata.csv</code>\n<code>train_df = pd.read_csv(metadata_path)</code>\n<code>train_df['label'].replace({'FAKE': 1, 'REAL': 0}, inplace=True)</code></p>\n\n<p><code>X = train_df['filename'].to_numpy()</code>\n<code>y = train_df['label'].to_numpy()</code></p>\n\n<p><code># Split the dataset</code>\n<code>X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.3, random_state=123, stratify=y)</code></p>\n\n<hr>\n\n<p>Consider to upvote these datasets if you think they are worth using 😉, this will motivate me much to continue this dull work 💪, thanks!!!</p>",
      "rawMarkdown": "I've created some datasets that include all detectable faces of all videos in each part of the full dataset. Kaggle and the host expected and encouraged us to train our models outside of Kaggle’s notebooks environment; however, for someone who prefers to stick to Kaggle's kernels, these preprocessed datasets would help a lot 😄.\n\nThe whole process to create and upload these datasets is time-consuming, so I'll gradually upload the rest for the next days; here are some completed ones:\n\n* [Deepfake Detection - Faces - Part 0_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-0)\n* [Deepfake Detection - Faces - Part 0_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-1)\n* [Deepfake Detection - Faces - Part 1_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-0)\n* [Deepfake Detection - Faces - Part 1_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-1)\n* [Deepfake Detection - Faces - Part 2_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-0)\n* [Deepfake Detection - Faces - Part 2_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-1)\n* [Deepfake Detection - Faces - Part 2_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-2)\n* [Deepfake Detection - Faces - Part 3_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-0)\n* [Deepfake Detection - Faces - Part 3_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-1)\n* [Deepfake Detection - Faces - Part 4_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-0)\n* [Deepfake Detection - Faces - Part 4_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-1)\n* [Deepfake Detection - Faces - Part 4_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-2)\n* [Deepfake Detection - Faces - Part 5_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-0)\n* [Deepfake Detection - Faces - Part 5_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-1)\n* [Deepfake Detection - Faces - Part 5_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-2)\n* [Deepfake Detection - Faces - Part 5_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-3)\n* [Deepfake Detection - Faces - Part 6_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-0)\n* [Deepfake Detection - Faces - Part 6_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-1)\n* [Deepfake Detection - Faces - Part 6_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-2)\n* [Deepfake Detection - Faces - Part 6_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-3)\n* [Deepfake Detection - Faces - Part 6_4](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-4)\n* [Deepfake Detection - Faces - Part 7_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-0)\n* [Deepfake Detection - Faces - Part 7_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-1)\n* [Deepfake Detection - Faces - Part 7_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-2)\n* [Deepfake Detection - Faces - Part 7_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-3)\n* [Deepfake Detection - Faces - Part 8_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-0)\n* [Deepfake Detection - Faces - Part 8_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-1)\n* [Deepfake Detection - Faces - Part 8_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-2)\n* [Deepfake Detection - Faces - Part 9_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-0)\n* [Deepfake Detection - Faces - Part 9_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-1)\n* [Deepfake Detection - Faces - Part 9_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-2)\n* [Deepfake Detection - Faces - Part 10_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-0)\n* [Deepfake Detection - Faces - Part 10_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-1)\n* [Deepfake Detection - Faces - Part 10_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-2)\n* [Deepfake Detection - Faces - Part 10_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-3)\n* [Deepfake Detection - Faces - Part 10_4](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-4)\n* [Deepfake Detection - Faces - Part 11_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-0)\n* [Deepfake Detection - Faces - Part 11_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-1)\n* [Deepfake Detection - Faces - Part 11_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-2)\n* [Deepfake Detection - Faces - Part 12_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-0)\n* [Deepfake Detection - Faces - Part 12_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-1)\n* [Deepfake Detection - Faces - Part 12_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-2)\n* [Deepfake Detection - Faces - Part 13_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-0)\n* [Deepfake Detection - Faces - Part 13_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-1)\n* [Deepfake Detection - Faces - Part 13_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-2)\n* [Deepfake Detection - Faces - Part 13_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-3)\n* [Deepfake Detection - Faces - Part 13_4](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-4)\n* [Deepfake Detection - Faces - Part 14_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-0)\n* [Deepfake Detection - Faces - Part 14_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-1)\n* [Deepfake Detection - Faces - Part 14_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-2)\n* [Deepfake Detection - Faces - Part 14_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-3)\n* [Deepfake Detection - Faces - Part 15_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-0)\n* [Deepfake Detection - Faces - Part 15_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-1)\n* [Deepfake Detection - Faces - Part 15_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-2)\n* [Deepfake Detection - Faces - Sample](https://www.kaggle.com/phunghieu/deepfake-detection-faces-sample)\n*Updating...*\n\nHow I create these datasets? Let's check out this demo kernel 👉 [\n*Deepfake Detection - Face Extractor*](https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor).\n\n---\n\nBecause of the big size of each part in the original full dataset, I have to break all extracted faces in each part into 2 or 3 smaller parts before zipping them, e.g. the name `deepfake-detection-faces-part-1-0` means it created from `part 1` of the full dataset and is the `first split`.\n\nThe name of each face image corresponds to the index of the frame that this face appears, plus a suffix _2, or _3, etc. if the number of faces in a frame is greater than 1.\n\nIn each dataset, I also attach a metadata.csv file which stores all information needed. The format of each file will look like this:\n|     |    filename    | split |    original    | label |\n|:---:|:--------------:|:-----:|:--------------:|:-----:|\n|  0  | aagfhgtpmv.mp4 | train | vudstovrck.mp4 | FAKE  |\n|  1  | aapnvogymq.mp4 | train | jdubbvfswz.mp4 | FAKE  |\n|  2  | abarnvbtwb.mp4 | train |                | REAL  |\n| ... | ...            | ...   | ...            | ...   |\n\n---\n\nWant something to get started using these datasets, let see [*DFDC-Multiface-Training*](https://www.kaggle.com/phunghieu/dfdc-multiface-training) &amp; [*DFDC-Multiface-Inference*](https://www.kaggle.com/phunghieu/dfdc-multiface-inference).\n\nDon't know how to load and merge multiple Kaggle datasets at once in a kernel, let check this [*demo*](https://www.kaggle.com/phunghieu/loading-merging-multiple-kaggle-datasets-demo).\n\nHere are a few lines of code to demonstrate how to load and prepare these data for the training process:\n\n`# Get path of metadata.csv`\n`metadata_path = os.path.join(TRAIN_DIR, 'metadata.csv')`\n\n` # Create DataFrame from metadata.csv`\n`train_df = pd.read_csv(metadata_path)`\n`train_df['label'].replace({'FAKE': 1, 'REAL': 0}, inplace=True)`\n\n`X = train_df['filename'].to_numpy()`\n`y = train_df['label'].to_numpy()`\n\n`# Split the dataset`\n`X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.3, random_state=123, stratify=y)`\n\n---\n\nConsider to upvote these datasets if you think they are worth using 😉, this will motivate me much to continue this dull work 💪, thanks!!!",
      "votes": 70
    },
    {
      "id": 760532,
      "postDate": "2020-03-01T12:26:16.110Z",
      "content": "<p>Thanks for sharing! Can you also share the full list(metadata.json) of labels of the whole dataset? </p>",
      "rawMarkdown": "Thanks for sharing! Can you also share the full list(metadata.json) of labels of the whole dataset? ",
      "votes": 1,
      "replies": [
        {
          "id": 766344,
          "postDate": "2020-03-08T03:13:49.813Z",
          "content": "<p>Hi <a href=\"/yuanzhezhou\">@yuanzhezhou</a>,</p>\n\n<p>Someone has already done this job for the whole community, you can find this <code>metadata.json</code> in this dataset -&gt; <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">Train set metadata for DFDC</a> created by <a href=\"/zaharch\">@zaharch</a>.</p>\n\n<p>Sorry for the late response, happy kaggling!</p>",
          "rawMarkdown": "Hi @yuanzhezhou,\n\nSomeone has already done this job for the whole community, you can find this `metadata.json` in this dataset -&gt; [Train set metadata for DFDC](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc) created by @zaharch.\n\nSorry for the late response, happy kaggling!",
          "votes": 2
        }
      ]
    },
    {
      "id": 747916,
      "postDate": "2020-02-17T02:27:27.683Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>15_0</code>, <code>15_1</code>, <code>15_2</code></p>",
      "rawMarkdown": "### Update\nPart `15_0`, `15_1`, `15_2`",
      "votes": 1
    },
    {
      "id": 747582,
      "postDate": "2020-02-16T15:47:14.620Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>14_0</code>, <code>14_1</code>, <code>14_2</code>, <code>14_3</code></p>",
      "rawMarkdown": "### Update\nPart `14_0`, `14_1`, `14_2`, `14_3`",
      "votes": 1
    },
    {
      "id": 747310,
      "postDate": "2020-02-16T08:33:35.350Z",
      "content": "<h3>Add examples</h3>\n\n<p><a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-training\"><em>DFDC-Multiface-Training</em></a> &amp; <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-inference\"><em>DFDC-Multiface-Inference</em></a></p>",
      "rawMarkdown": "### Add examples\n[*DFDC-Multiface-Training*](https://www.kaggle.com/phunghieu/dfdc-multiface-training) &amp; [*DFDC-Multiface-Inference*](https://www.kaggle.com/phunghieu/dfdc-multiface-inference)",
      "votes": 1
    },
    {
      "id": 747193,
      "postDate": "2020-02-16T05:12:21.877Z",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>14_0</code>, <code>14_1</code>, <code>14_2</code>, <code>14_3</code>, <code>15_0</code>, <code>15_1</code>, <code>15_2</code></p>",
      "rawMarkdown": "### Coming soon\nPart `14_0`, `14_1`, `14_2`, `14_3`, `15_0`, `15_1`, `15_2`",
      "votes": 1
    },
    {
      "id": 746502,
      "postDate": "2020-02-15T05:40:10.580Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>13_3</code>, <code>13_4</code></p>",
      "rawMarkdown": "### Update\nPart `13_3`, `13_4`",
      "votes": 1
    },
    {
      "id": 746153,
      "postDate": "2020-02-14T17:21:39.440Z",
      "content": "<p>Thanks for the  share</p>",
      "rawMarkdown": "Thanks for the  share",
      "votes": 1,
      "replies": [
        {
          "id": 746390,
          "postDate": "2020-02-14T23:47:54.037Z",
          "content": "<p>You're welcome! <a href=\"/psywarrior\">@psywarrior</a> 😊</p>",
          "rawMarkdown": "You're welcome! @psywarrior 😊",
          "votes": 2
        }
      ]
    },
    {
      "id": 746129,
      "postDate": "2020-02-14T16:56:42.593Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>13_0</code>, <code>13_1</code>, <code>13_2</code></p>\n\n<h3>Coming soon</h3>\n\n<p>Part <code>13_3</code>, <code>13_4</code></p>",
      "rawMarkdown": "### Update\nPart `13_0`, `13_1`, `13_2`\n\n### Coming soon\nPart `13_3`, `13_4`",
      "votes": 1
    },
    {
      "id": 745751,
      "postDate": "2020-02-14T06:45:09.633Z",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>13_0</code>, <code>13_1</code>, <code>13_2</code></p>",
      "rawMarkdown": "### Coming soon\nPart `13_0`, `13_1`, `13_2`",
      "votes": 1
    },
    {
      "id": 745057,
      "postDate": "2020-02-13T12:38:24.113Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>12_0</code>, <code>12_1</code>, <code>12_2</code></p>",
      "rawMarkdown": "### Update\nPart `12_0`, `12_1`, `12_2`",
      "votes": 1
    },
    {
      "id": 744681,
      "postDate": "2020-02-13T03:10:18.563Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>10_4</code></p>",
      "rawMarkdown": "### Update\nPart `10_4`",
      "votes": 1
    },
    {
      "id": 744617,
      "postDate": "2020-02-13T01:31:58.890Z",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>12_0</code>, <code>12_1</code>, <code>12_2</code></p>",
      "rawMarkdown": "### Coming soon\nPart `12_0`, `12_1`, `12_2`",
      "votes": 1
    },
    {
      "id": 744592,
      "postDate": "2020-02-13T00:40:37.663Z",
      "content": "<p>Awesome work! I just want to confirm whether the different 'parts' (i.e. Part 0_1 and Part 8_2) are pre-processed in the exact same way? I'm assuming the numbering is basically denoting the pre-processing being done in distinct batches.</p>",
      "rawMarkdown": "Awesome work! I just want to confirm whether the different 'parts' (i.e. Part 0_1 and Part 8_2) are pre-processed in the exact same way? I'm assuming the numbering is basically denoting the pre-processing being done in distinct batches.",
      "votes": 1,
      "replies": [
        {
          "id": 744594,
          "postDate": "2020-02-13T00:43:23.867Z",
          "content": "<p><a href=\"/stephendlee94\">@stephendlee94</a> yes, they are pre-processed in the exact same way. In addition, all batches have the same FAKE/REAL ratio as well.</p>",
          "rawMarkdown": "@stephendlee94 yes, they are pre-processed in the exact same way. In addition, all batches have the same FAKE/REAL ratio as well.",
          "votes": 1
        }
      ]
    },
    {
      "id": 743682,
      "postDate": "2020-02-12T07:32:14.723Z",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>10_0</code>, <code>10_1</code>, <code>10_2</code>, <code>10_3</code>, <code>10_4</code>, <code>11_0</code>, <code>11_1</code>, <code>11_2</code></p>",
      "rawMarkdown": "### Coming soon\nPart `10_0`, `10_1`, `10_2`, `10_3`, `10_4`, `11_0`, `11_1`, `11_2`",
      "votes": 1
    },
    {
      "id": 741402,
      "postDate": "2020-02-10T15:35:21.150Z",
      "content": "<p>Nice work man! I'm struggling to process this huge dataset. Yours will help a lot!</p>",
      "rawMarkdown": "Nice work man! I'm struggling to process this huge dataset. Yours will help a lot!",
      "votes": 1,
      "replies": [
        {
          "id": 741406,
          "postDate": "2020-02-10T15:38:50.763Z",
          "content": "<p>Thanks, <a href=\"/ronaldokun\">@ronaldokun</a> 😄</p>",
          "rawMarkdown": "Thanks, @ronaldokun 😄"
        }
      ]
    },
    {
      "id": 740172,
      "postDate": "2020-02-09T01:22:01.327Z",
      "content": "<p>thanks for sharing!!! 😃 </p>",
      "rawMarkdown": "thanks for sharing!!! 😃 ",
      "votes": 1,
      "replies": [
        {
          "id": 740187,
          "postDate": "2020-02-09T02:12:19.637Z",
          "content": "<p>You're welcome! <a href=\"/victormenuzzo\">@victormenuzzo</a> 😁</p>",
          "rawMarkdown": "You're welcome! @victormenuzzo 😁"
        }
      ]
    },
    {
      "id": 740003,
      "postDate": "2020-02-08T17:33:00.313Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>6_0</code>, <code>6_1</code>, <code>6_2</code>, <code>7_3</code></p>",
      "rawMarkdown": "### Update\nPart `6_0`, `6_1`, `6_2`, `7_3`",
      "votes": 1
    },
    {
      "id": 739996,
      "postDate": "2020-02-08T17:22:23.533Z",
      "content": "<p>Good job,dude\nThanks.. 👍 </p>",
      "rawMarkdown": "Good job,dude\nThanks.. 👍 ",
      "votes": 1,
      "replies": [
        {
          "id": 739998,
          "postDate": "2020-02-08T17:24:06.743Z",
          "content": "<p>You're welcome! <a href=\"/anubhav1302\">@anubhav1302</a> 😊</p>",
          "rawMarkdown": "You're welcome! @anubhav1302 😊"
        }
      ]
    },
    {
      "id": 739952,
      "postDate": "2020-02-08T15:55:14.603Z",
      "content": "<p>Very useful great work thanks for sharing </p>",
      "rawMarkdown": "Very useful great work thanks for sharing ",
      "votes": 1,
      "replies": [
        {
          "id": 739956,
          "postDate": "2020-02-08T16:08:32.657Z",
          "content": "<p>You're welcome! <a href=\"/modoucair\">@modoucair</a> 😁</p>",
          "rawMarkdown": "You're welcome! @modoucair 😁"
        }
      ]
    },
    {
      "id": 739826,
      "postDate": "2020-02-08T13:15:05.760Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>7_2</code></p>",
      "rawMarkdown": "### Update\nPart `7_2`",
      "votes": 1
    },
    {
      "id": 739685,
      "postDate": "2020-02-08T07:36:20.337Z",
      "content": "<h2>Risk of disqualification</h2>\n\n<p>There are doubts about the effectiveness of your manual face search approach for this competition. If you teach the model in this way, then the test data should be processed in the same way, because your model cannot find the face itself. But if you just touch manually the test data, then immediately violate the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/rules\">rules of the competition</a>:</p>\n\n<p><em>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</em></p>",
      "rawMarkdown": "## Risk of disqualification\n\nThere are doubts about the effectiveness of your manual face search approach for this competition. If you teach the model in this way, then the test data should be processed in the same way, because your model cannot find the face itself. But if you just touch manually the test data, then immediately violate the [rules of the competition](https://www.kaggle.com/c/deepfake-detection-challenge/rules):\n\n*Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.*",
      "votes": 1,
      "replies": [
        {
          "id": 739690,
          "postDate": "2020-02-08T07:49:40.697Z",
          "content": "<p>Don't worry, I'll never touch the test dataset manually; instead, I'll use at least one or even some pre-trained models to find faces in each test video. Thanks for your warning, <a href=\"/vbmokin\">@vbmokin</a> 👍</p>",
          "rawMarkdown": "Don't worry, I'll never touch the test dataset manually; instead, I'll use at least one or even some pre-trained models to find faces in each test video. Thanks for your warning, @vbmokin 👍",
          "votes": 1
        }
      ]
    },
    {
      "id": 739670,
      "postDate": "2020-02-08T06:54:47.847Z",
      "content": "<p><a href=\"/phunghieu\">@phunghieu</a>  You are doing a great job. One suggestion is to save them as a video instead of images. It'll reduce size for each dataset to less then 2GB. Or even if you want to save as images try with .jpg extension. Ask for any code help and keep up the great work.</p>",
      "rawMarkdown": "@phunghieu  You are doing a great job. One suggestion is to save them as a video instead of images. It'll reduce size for each dataset to less then 2GB. Or even if you want to save as images try with .jpg extension. Ask for any code help and keep up the great work.",
      "votes": 1,
      "replies": [
        {
          "id": 739688,
          "postDate": "2020-02-08T07:46:11.690Z",
          "content": "<p>Thanks for your suggestion <a href=\"/ankitsainiankit\">@ankitsainiankit</a>, the only reason for why I saved the results in .png extension is to preserve as much information as possible, we can take further preprocessing if needed without worrying about missing important information. BTW, this is a great idea, maybe, I can create other datasets or it would help if someone did it instead of me 😊.</p>",
          "rawMarkdown": "Thanks for your suggestion @ankitsainiankit, the only reason for why I saved the results in .png extension is to preserve as much information as possible, we can take further preprocessing if needed without worrying about missing important information. BTW, this is a great idea, maybe, I can create other datasets or it would help if someone did it instead of me 😊."
        },
        {
          "id": 739703,
          "postDate": "2020-02-08T08:22:02.703Z",
          "content": "<p>I've done it for a few parts. Due to hardware limitations, it is hard for me. And one more problem is that I am using opencv face detection which is not so good (1 in 20 videos is not a valid face). And other algorithms are time-consuming. My Current score uses the face dataset by <a href=\"/unkownhihi\">@unkownhihi</a> and face containing frames from my dataset.  </p>",
          "rawMarkdown": "I've done it for a few parts. Due to hardware limitations, it is hard for me. And one more problem is that I am using opencv face detection which is not so good (1 in 20 videos is not a valid face). And other algorithms are time-consuming. My Current score uses the face dataset by @unkownhihi and face containing frames from my dataset.  "
        }
      ]
    },
    {
      "id": 739584,
      "postDate": "2020-02-08T02:39:49.550Z",
      "content": "<p>Great job! Thanks for sharing <a href=\"/phunghieu\">@phunghieu</a> !!😄 👍 </p>",
      "rawMarkdown": "Great job! Thanks for sharing @phunghieu !!😄 👍 ",
      "votes": 1,
      "replies": [
        {
          "id": 739660,
          "postDate": "2020-02-08T06:08:55.383Z",
          "content": "<p>You're welcome! <a href=\"/mashlyn\">@mashlyn</a> 😄</p>",
          "rawMarkdown": "You're welcome! @mashlyn 😄"
        }
      ]
    },
    {
      "id": 739493,
      "postDate": "2020-02-07T22:02:07.833Z",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": 1,
      "replies": [
        {
          "id": 739537,
          "postDate": "2020-02-07T23:11:42.453Z",
          "content": "<p>You're welcome, <a href=\"/azelis\">@azelis</a> 😆</p>",
          "rawMarkdown": "You're welcome, @azelis 😆"
        }
      ]
    },
    {
      "id": 739238,
      "postDate": "2020-02-07T15:38:56.450Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>7_0</code>, <code>7_1</code></p>",
      "rawMarkdown": "### Update\nPart `7_0`, `7_1`",
      "votes": 1
    },
    {
      "id": 739002,
      "postDate": "2020-02-07T09:18:54.190Z",
      "content": "<p>Great datasets</p>",
      "rawMarkdown": "Great datasets",
      "votes": 1,
      "replies": [
        {
          "id": 739016,
          "postDate": "2020-02-07T09:40:49.380Z",
          "content": "<p>Thanks, <a href=\"/madhubswl\">@madhubswl</a> 😊</p>",
          "rawMarkdown": "Thanks, @madhubswl 😊"
        }
      ]
    },
    {
      "id": 738916,
      "postDate": "2020-02-07T07:20:45.960Z",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>7_0</code>, <code>7_1</code>, <code>7_2</code>, <code>7_3</code></p>",
      "rawMarkdown": "### Coming soon\nPart `7_0`, `7_1`, `7_2`, `7_3`",
      "votes": 1
    },
    {
      "id": 738364,
      "postDate": "2020-02-06T12:41:38.917Z",
      "content": "<p>Deepfake Detection - Faces - Part 0_1 data set does not have download link. Do you maybe know why <a href=\"/phunghieu\">@phunghieu</a> ?</p>",
      "rawMarkdown": "Deepfake Detection - Faces - Part 0_1 data set does not have download link. Do you maybe know why @phunghieu ?",
      "votes": 1,
      "replies": [
        {
          "id": 738365,
          "postDate": "2020-02-06T12:48:12.307Z",
          "content": "<p>Let me check. Maybe, there are some problems when I upload it to Kaggle. Thanks for your reply!</p>",
          "rawMarkdown": "Let me check. Maybe, there are some problems when I upload it to Kaggle. Thanks for your reply!"
        },
        {
          "id": 738476,
          "postDate": "2020-02-06T15:18:49.077Z",
          "content": "<p><a href=\"/dandrocec\">@dandrocec</a> now we have the download link for Part 0_1, let check it out 😉</p>",
          "rawMarkdown": "@dandrocec now we have the download link for Part 0_1, let check it out 😉",
          "votes": 1
        },
        {
          "id": 738722,
          "postDate": "2020-02-06T22:58:02.917Z",
          "content": "<p>Thanks.</p>",
          "rawMarkdown": "Thanks."
        }
      ]
    },
    {
      "id": 738190,
      "postDate": "2020-02-06T08:47:57.450Z",
      "content": "<p>Thanks a lot for sharing this.</p>",
      "rawMarkdown": "Thanks a lot for sharing this.",
      "votes": 1,
      "replies": [
        {
          "id": 738207,
          "postDate": "2020-02-06T09:07:33.503Z",
          "content": "<p>You're welcome! <a href=\"/senolcomert\">@senolcomert</a> 😁</p>",
          "rawMarkdown": "You're welcome! @senolcomert 😁"
        }
      ]
    },
    {
      "id": 738133,
      "postDate": "2020-02-06T07:19:05.757Z",
      "content": "<h3>Update</h3>\n\n<p><a href=\"https://www.kaggle.com/phunghieu/loading-merging-multiple-kaggle-datasets-demo\"><em>Loading &amp; Merging Multiple Kaggle Datasets (Demo)</em></a></p>",
      "rawMarkdown": "### Update\n[*Loading &amp; Merging Multiple Kaggle Datasets (Demo)*](https://www.kaggle.com/phunghieu/loading-merging-multiple-kaggle-datasets-demo)",
      "votes": 1
    },
    {
      "id": 737232,
      "postDate": "2020-02-05T04:32:05.093Z",
      "content": "<p>Great work! thanks for sharing this out.</p>",
      "rawMarkdown": "Great work! thanks for sharing this out.",
      "votes": 1,
      "replies": [
        {
          "id": 737237,
          "postDate": "2020-02-05T04:38:49.937Z",
          "content": "<p>You're welcome! <a href=\"/jmelodyxiao\">@jmelodyxiao</a> 😁</p>",
          "rawMarkdown": "You're welcome! @jmelodyxiao 😁"
        },
        {
          "id": 737249,
          "postDate": "2020-02-05T05:11:56.557Z",
          "content": "<p>Have one question though: Why did you choose png format to save these faces? The output datasets is even bigger than the original datasets for every face's size is 75.24kB, which means for a single video, there are roughly 75.24kB*300 =  22MB storage needed.  I might be wrong, but IMHO png might be unnecessary large for a single face and jpg is enough for training cause the original dataset is not in that high quality though.</p>",
          "rawMarkdown": "Have one question though: Why did you choose png format to save these faces? The output datasets is even bigger than the original datasets for every face's size is 75.24kB, which means for a single video, there are roughly 75.24kB*300 =  22MB storage needed.  I might be wrong, but IMHO png might be unnecessary large for a single face and jpg is enough for training cause the original dataset is not in that high quality though.",
          "votes": 1
        },
        {
          "id": 737256,
          "postDate": "2020-02-05T05:25:20.527Z",
          "content": "<p><a href=\"/jmelodyxiao\">@jmelodyxiao</a> Since each image is pretty small, 160x160, I think it's better to preserve all the information we can get rather than using a lossy compression method 🤓, then we can have further preprocessing if needed 😉.</p>",
          "rawMarkdown": "@jmelodyxiao Since each image is pretty small, 160x160, I think it's better to preserve all the information we can get rather than using a lossy compression method 🤓, then we can have further preprocessing if needed 😉."
        },
        {
          "id": 737264,
          "postDate": "2020-02-05T05:35:19.780Z",
          "content": "<p>understood~Just thought it would be too heavy for training (need to download/upload all these datasets), which makes me still stuck in the same stereotype just like the original datasets gave.😂</p>",
          "rawMarkdown": "understood~Just thought it would be too heavy for training (need to download/upload all these datasets), which makes me still stuck in the same stereotype just like the original datasets gave.😂\n "
        },
        {
          "id": 737272,
          "postDate": "2020-02-05T05:43:24.187Z",
          "content": "<p>This is the problem that competitors who don't have enough storage n computing power face 😂. Currently, I'm using the data from 3 parts + sample part to train my classifier, it seems to work just fine; I'll public this kernel soon after confirming that it truly works.</p>",
          "rawMarkdown": "This is the problem that competitors who don't have enough storage n computing power face 😂. Currently, I'm using the data from 3 parts + sample part to train my classifier, it seems to work just fine; I'll public this kernel soon after confirming that it truly works.",
          "votes": 1
        },
        {
          "id": 737345,
          "postDate": "2020-02-05T08:05:00.403Z",
          "content": "<p>Great! Looking forward to that~</p>",
          "rawMarkdown": "Great! Looking forward to that~"
        }
      ]
    },
    {
      "id": 736979,
      "postDate": "2020-02-04T19:05:57.537Z",
      "content": "<p>Hi, I think it would be helpful to share the face extraction code, as someone using this dataset would like to use the exact face extraction process during inference! Thanks for the hard work, upvoted.</p>",
      "rawMarkdown": "Hi, I think it would be helpful to share the face extraction code, as someone using this dataset would like to use the exact face extraction process during inference! Thanks for the hard work, upvoted.",
      "votes": 1,
      "replies": [
        {
          "id": 737063,
          "postDate": "2020-02-04T21:38:04.350Z",
          "content": "<p>Yes, u r right. Im reorganizing and beautifying the Face Extractor, it'll available soon ^^. Besides, Im working on a simple pipeline of using these datasets to train a basic classifier, I'll try my best to finish it as soon as possible.\nThanks for your suggestion, <a href=\"/debanga\">@debanga</a> 👍</p>",
          "rawMarkdown": "Yes, u r right. Im reorganizing and beautifying the Face Extractor, it'll available soon ^^. Besides, Im working on a simple pipeline of using these datasets to train a basic classifier, I'll try my best to finish it as soon as possible.\nThanks for your suggestion, @debanga 👍",
          "votes": 1
        },
        {
          "id": 737203,
          "postDate": "2020-02-05T03:38:42.163Z",
          "content": "<p>So, I've finished the job with the face-extractor, let have a quick look <a href=\"/debanga\">@debanga</a> 😊.\n<a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor\"><em>Deepfake Detection - Face Extractor</em></a></p>",
          "rawMarkdown": "So, I've finished the job with the face-extractor, let have a quick look @debanga 😊.\n[*Deepfake Detection - Face Extractor*](https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor)"
        },
        {
          "id": 737217,
          "postDate": "2020-02-05T04:01:05.913Z",
          "content": "<p>Great! Will look, thanks :)))</p>",
          "rawMarkdown": "Great! Will look, thanks :)))",
          "votes": 1
        }
      ]
    },
    {
      "id": 736922,
      "postDate": "2020-02-04T17:59:15.390Z",
      "content": "<p>Thanks for your job, man! It`s great</p>",
      "rawMarkdown": "Thanks for your job, man! It`s great\n",
      "votes": 1,
      "replies": [
        {
          "id": 737056,
          "postDate": "2020-02-04T21:29:02.283Z",
          "content": "<p>You're welcome! <a href=\"/slavapasedko\">@slavapasedko</a> 😊</p>",
          "rawMarkdown": "You're welcome! @slavapasedko 😊"
        }
      ]
    },
    {
      "id": 747382,
      "postDate": "2020-02-16T11:10:26.190Z",
      "content": "<p>You are a new great asset to Kaggle by your wholeheartedly contribution.</p>",
      "rawMarkdown": "You are a new great asset to Kaggle by your wholeheartedly contribution.",
      "votes": 2,
      "replies": [
        {
          "id": 747442,
          "postDate": "2020-02-16T12:52:45.653Z",
          "content": "<p>Many thanks, <a href=\"/khahuras\">@khahuras</a>! I'm really happy that many people can, somewhat, benefit from my work 💛 </p>",
          "rawMarkdown": "Many thanks, @khahuras! I'm really happy that many people can, somewhat, benefit from my work 💛 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 746563,
      "postDate": "2020-02-15T07:49:14.683Z",
      "content": "<p>Thanku <a href=\"/phunghieu\">@phunghieu</a> \n Thanku <a href=\"/phunghieu\">@phunghieu</a> \nIs this based on the kaggle dataset ?\n1) Which model u used used in for face extraction MTCNN/dlib/Blaze..\n2) Have u taken multiple frames for face extraction,if yes how many ?\n3) How we can label this ,can we assumed image id be used for getting the label\nI hope you keeping safe if from china :)</p>",
      "rawMarkdown": "Thanku @phunghieu \n Thanku @phunghieu \nIs this based on the kaggle dataset ?\n1) Which model u used used in for face extraction MTCNN/dlib/Blaze..\n2) Have u taken multiple frames for face extraction,if yes how many ?\n3) How we can label this ,can we assumed image id be used for getting the label\nI hope you keeping safe if from china :)",
      "votes": 2,
      "replies": [
        {
          "id": 746687,
          "postDate": "2020-02-15T12:01:14.940Z",
          "content": "<p>Hi <a href=\"/jaideepvalani\">@jaideepvalani</a>,</p>\n\n<ul>\n<li>For your first question, the answer is yes; all extracted faces come from the datasets provided by the host. In addition, I use MTCNN to obtain them, the whole process can be done through my <a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor\"><em>Deepfake Detection - Face Extractor</em></a> kernel.</li>\n<li>Second, I get as many as faces possible (I run the detector through all frames of each video), the name of each image also denote the index of the frame I got it. Plus, if there is more than one face appear in the frame, the results will be appended suffix, e.g. <code>_2</code>, <code>_3</code>, and so on.</li>\n<li>Finally, I've organized the results into the separated folder for each video, so the users can easily get the corresponding label for each face image.</li>\n</ul>\n\n<p>And thanks, I still safe in Vietnam, hope that China and the whole world can overcome this epidemic as quickly as possible 💛!</p>\n\n<p>Cheers!</p>",
          "rawMarkdown": "Hi @jaideepvalani,\n\n* For your first question, the answer is yes; all extracted faces come from the datasets provided by the host. In addition, I use MTCNN to obtain them, the whole process can be done through my [*Deepfake Detection - Face Extractor*](https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor) kernel.\n* Second, I get as many as faces possible (I run the detector through all frames of each video), the name of each image also denote the index of the frame I got it. Plus, if there is more than one face appear in the frame, the results will be appended suffix, e.g. `_2`, `_3`, and so on.\n* Finally, I've organized the results into the separated folder for each video, so the users can easily get the corresponding label for each face image.\n\nAnd thanks, I still safe in Vietnam, hope that China and the whole world can overcome this epidemic as quickly as possible 💛!\n\nCheers!",
          "votes": 1
        }
      ]
    },
    {
      "id": 740080,
      "postDate": "2020-02-08T20:15:07.650Z",
      "content": "<p>Thank you for sharing!!! Great job!</p>",
      "rawMarkdown": "Thank you for sharing!!! Great job!",
      "votes": 2,
      "replies": [
        {
          "id": 740122,
          "postDate": "2020-02-08T23:13:13.977Z",
          "content": "<p>You're welcome! <a href=\"/lihyalan\">@lihyalan</a> 😄</p>",
          "rawMarkdown": "You're welcome! @lihyalan 😄"
        }
      ]
    },
    {
      "id": 744060,
      "postDate": "2020-02-12T14:25:48.330Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>10_0</code>, <code>10_1</code>, <code>10_2</code>, <code>10_3</code></p>",
      "rawMarkdown": "### Update\nPart `10_0`, `10_1`, `10_2`, `10_3`"
    },
    {
      "id": 744014,
      "postDate": "2020-02-12T13:30:46.787Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>11_0</code>, <code>11_1</code>, <code>11_2</code></p>",
      "rawMarkdown": "### Update\nPart `11_0`, `11_1`, `11_2`"
    },
    {
      "id": 740456,
      "postDate": "2020-02-09T12:49:28.313Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>9_0</code>, <code>9_1</code>, <code>9_2</code></p>",
      "rawMarkdown": "### Update\nPart `9_0`, `9_1`, `9_2`"
    },
    {
      "id": 740438,
      "postDate": "2020-02-09T12:33:51.817Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>8_0</code>, <code>8_1</code>, <code>8_2</code></p>",
      "rawMarkdown": "### Update\nPart `8_0`, `8_1`, `8_2`"
    },
    {
      "id": 740236,
      "postDate": "2020-02-09T04:11:07.933Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>6_3</code>, <code>6_4</code></p>\n\n<h3>Coming soon</h3>\n\n<p>Part <code>8_0</code>, <code>8_1</code>, <code>8_2</code>, <code>9_0</code>, <code>9_1</code>, <code>9_2</code></p>",
      "rawMarkdown": "### Update\nPart `6_3`, `6_4`\n\n### Coming soon\nPart `8_0`, `8_1`, `8_2`, `9_0`, `9_1`, `9_2`"
    },
    {
      "id": 738458,
      "postDate": "2020-02-06T14:48:20.893Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>5_2</code>, <code>5_3</code></p>",
      "rawMarkdown": "### Update\nPart `5_2`, `5_3`"
    },
    {
      "id": 738217,
      "postDate": "2020-02-06T09:20:04.837Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>5_0</code>, <code>5_1</code></p>",
      "rawMarkdown": "### Update\nPart `5_0`, `5_1`"
    },
    {
      "id": 737517,
      "postDate": "2020-02-05T13:10:18.730Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>4_0</code>, <code>4_1</code>, <code>4_2</code></p>",
      "rawMarkdown": "### Update\nPart `4_0`, `4_1`, `4_2`"
    },
    {
      "id": 737405,
      "postDate": "2020-02-05T10:07:26.407Z",
      "content": "<h3>Update</h3>\n\n<p>Part <code>2_0</code>, <code>2_1</code>, <code>2_2</code></p>\n\n<h3>Coming soon</h3>\n\n<p>Part <code>4_0</code>, <code>4_1</code>, <code>4_2</code></p>",
      "rawMarkdown": "### Update\nPart `2_0`, `2_1`, `2_2`\n\n### Coming soon\nPart `4_0`, `4_1`, `4_2`",
      "replies": [
        {
          "id": 785929,
          "postDate": "2020-03-25T14:17:17.307Z",
          "content": "<p>Hieu Phung - thanks so much for providing these! Do the parts you've uploaded to the publicly available datasets also include the testing data?</p>",
          "rawMarkdown": "Hieu Phung - thanks so much for providing these! Do the parts you've uploaded to the publicly available datasets also include the testing data?",
          "votes": 1
        },
        {
          "id": 786450,
          "postDate": "2020-03-25T22:38:42.963Z",
          "content": "<p>Hi <a href=\"/mrwynx\">@mrwynx</a>,\nAll the parts I've prepared and uploaded here only include training data.</p>",
          "rawMarkdown": "Hi @mrwynx,\nAll the parts I've prepared and uploaded here only include training data."
        }
      ]
    },
    {
      "id": 749466,
      "postDate": "2020-02-18T18:13:12.153Z",
      "content": "<p>Thanks for sharing !! </p>",
      "rawMarkdown": "Thanks for sharing !! ",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 749752,
          "postDate": "2020-02-18T22:14:20.787Z",
          "content": "<p>You're welcome! <a href=\"/muerbingsha\">@muerbingsha</a> 😁</p>",
          "rawMarkdown": "You're welcome! @muerbingsha 😁"
        }
      ]
    },
    {
      "id": 741189,
      "postDate": "2020-02-10T10:00:17.970Z",
      "content": "<p>Thanks for sharing this!</p>",
      "rawMarkdown": "Thanks for sharing this!",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 760532,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2020-03-01T12:26:16.110000",
      "content": "<p>Thanks for sharing! Can you also share the full list(metadata.json) of labels of the whole dataset? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 766344,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-03-08T03:13:49.813000",
          "content": "<p>Hi <a href=\"/yuanzhezhou\">@yuanzhezhou</a>,</p>\n\n<p>Someone has already done this job for the whole community, you can find this <code>metadata.json</code> in this dataset -&gt; <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">Train set metadata for DFDC</a> created by <a href=\"/zaharch\">@zaharch</a>.</p>\n\n<p>Sorry for the late response, happy kaggling!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 747916,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-17T02:27:27.683000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>15_0</code>, <code>15_1</code>, <code>15_2</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 747582,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-16T15:47:14.620000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>14_0</code>, <code>14_1</code>, <code>14_2</code>, <code>14_3</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 747310,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-16T08:33:35.350000",
      "content": "<h3>Add examples</h3>\n\n<p><a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-training\"><em>DFDC-Multiface-Training</em></a> &amp; <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-inference\"><em>DFDC-Multiface-Inference</em></a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 747193,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-16T05:12:21.877000",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>14_0</code>, <code>14_1</code>, <code>14_2</code>, <code>14_3</code>, <code>15_0</code>, <code>15_1</code>, <code>15_2</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 746502,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-15T05:40:10.580000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>13_3</code>, <code>13_4</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 746153,
      "author_name": "Sanyam Sanjay Sharma",
      "author_url": "",
      "post_date": "2020-02-14T17:21:39.440000",
      "content": "<p>Thanks for the  share</p>",
      "votes": 1,
      "replies": [
        {
          "id": 746390,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-14T23:47:54.037000",
          "content": "<p>You're welcome! <a href=\"/psywarrior\">@psywarrior</a> 😊</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 746129,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-14T16:56:42.593000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>13_0</code>, <code>13_1</code>, <code>13_2</code></p>\n\n<h3>Coming soon</h3>\n\n<p>Part <code>13_3</code>, <code>13_4</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 745751,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-14T06:45:09.633000",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>13_0</code>, <code>13_1</code>, <code>13_2</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 745057,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-13T12:38:24.113000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>12_0</code>, <code>12_1</code>, <code>12_2</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 744681,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-13T03:10:18.563000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>10_4</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 744617,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-13T01:31:58.890000",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>12_0</code>, <code>12_1</code>, <code>12_2</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 744592,
      "author_name": "Stephen Lee",
      "author_url": "",
      "post_date": "2020-02-13T00:40:37.663000",
      "content": "<p>Awesome work! I just want to confirm whether the different 'parts' (i.e. Part 0_1 and Part 8_2) are pre-processed in the exact same way? I'm assuming the numbering is basically denoting the pre-processing being done in distinct batches.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 744594,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-13T00:43:23.867000",
          "content": "<p><a href=\"/stephendlee94\">@stephendlee94</a> yes, they are pre-processed in the exact same way. In addition, all batches have the same FAKE/REAL ratio as well.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 743682,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-12T07:32:14.723000",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>10_0</code>, <code>10_1</code>, <code>10_2</code>, <code>10_3</code>, <code>10_4</code>, <code>11_0</code>, <code>11_1</code>, <code>11_2</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 741402,
      "author_name": "Ronaldo S.A. Batista",
      "author_url": "",
      "post_date": "2020-02-10T15:35:21.150000",
      "content": "<p>Nice work man! I'm struggling to process this huge dataset. Yours will help a lot!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 741406,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-10T15:38:50.763000",
          "content": "<p>Thanks, <a href=\"/ronaldokun\">@ronaldokun</a> 😄</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 740172,
      "author_name": "Victor Menuzzo",
      "author_url": "",
      "post_date": "2020-02-09T01:22:01.327000",
      "content": "<p>thanks for sharing!!! 😃 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 740187,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-09T02:12:19.637000",
          "content": "<p>You're welcome! <a href=\"/victormenuzzo\">@victormenuzzo</a> 😁</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 740003,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-08T17:33:00.313000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>6_0</code>, <code>6_1</code>, <code>6_2</code>, <code>7_3</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 739996,
      "author_name": "Anubhav Singh",
      "author_url": "",
      "post_date": "2020-02-08T17:22:23.533000",
      "content": "<p>Good job,dude\nThanks.. 👍 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 739998,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-08T17:24:06.743000",
          "content": "<p>You're welcome! <a href=\"/anubhav1302\">@anubhav1302</a> 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 739952,
      "author_name": "Ndiaye",
      "author_url": "",
      "post_date": "2020-02-08T15:55:14.603000",
      "content": "<p>Very useful great work thanks for sharing </p>",
      "votes": 1,
      "replies": [
        {
          "id": 739956,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-08T16:08:32.657000",
          "content": "<p>You're welcome! <a href=\"/modoucair\">@modoucair</a> 😁</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 739826,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-08T13:15:05.760000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>7_2</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 739685,
      "author_name": "Vitalii Mokin",
      "author_url": "",
      "post_date": "2020-02-08T07:36:20.337000",
      "content": "<h2>Risk of disqualification</h2>\n\n<p>There are doubts about the effectiveness of your manual face search approach for this competition. If you teach the model in this way, then the test data should be processed in the same way, because your model cannot find the face itself. But if you just touch manually the test data, then immediately violate the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/rules\">rules of the competition</a>:</p>\n\n<p><em>Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</em></p>",
      "votes": 1,
      "replies": [
        {
          "id": 739690,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-08T07:49:40.697000",
          "content": "<p>Don't worry, I'll never touch the test dataset manually; instead, I'll use at least one or even some pre-trained models to find faces in each test video. Thanks for your warning, <a href=\"/vbmokin\">@vbmokin</a> 👍</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 739670,
      "author_name": "Ankit Saini",
      "author_url": "",
      "post_date": "2020-02-08T06:54:47.847000",
      "content": "<p><a href=\"/phunghieu\">@phunghieu</a>  You are doing a great job. One suggestion is to save them as a video instead of images. It'll reduce size for each dataset to less then 2GB. Or even if you want to save as images try with .jpg extension. Ask for any code help and keep up the great work.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 739688,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-08T07:46:11.690000",
          "content": "<p>Thanks for your suggestion <a href=\"/ankitsainiankit\">@ankitsainiankit</a>, the only reason for why I saved the results in .png extension is to preserve as much information as possible, we can take further preprocessing if needed without worrying about missing important information. BTW, this is a great idea, maybe, I can create other datasets or it would help if someone did it instead of me 😊.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 739703,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-02-08T08:22:02.703000",
          "content": "<p>I've done it for a few parts. Due to hardware limitations, it is hard for me. And one more problem is that I am using opencv face detection which is not so good (1 in 20 videos is not a valid face). And other algorithms are time-consuming. My Current score uses the face dataset by <a href=\"/unkownhihi\">@unkownhihi</a> and face containing frames from my dataset.  </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 739584,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2020-02-08T02:39:49.550000",
      "content": "<p>Great job! Thanks for sharing <a href=\"/phunghieu\">@phunghieu</a> !!😄 👍 </p>",
      "votes": 1,
      "replies": [
        {
          "id": 739660,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-08T06:08:55.383000",
          "content": "<p>You're welcome! <a href=\"/mashlyn\">@mashlyn</a> 😄</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 739493,
      "author_name": "Azelis",
      "author_url": "",
      "post_date": "2020-02-07T22:02:07.833000",
      "content": "<p>Thanks!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 739537,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-07T23:11:42.453000",
          "content": "<p>You're welcome, <a href=\"/azelis\">@azelis</a> 😆</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 739238,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-07T15:38:56.450000",
      "content": "<h3>Update</h3>\n\n<p>Part <code>7_0</code>, <code>7_1</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 739002,
      "author_name": "Madhusmita Biswal",
      "author_url": "",
      "post_date": "2020-02-07T09:18:54.190000",
      "content": "<p>Great datasets</p>",
      "votes": 1,
      "replies": [
        {
          "id": 739016,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-07T09:40:49.380000",
          "content": "<p>Thanks, <a href=\"/madhubswl\">@madhubswl</a> 😊</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 738916,
      "author_name": "Hieu Phung",
      "author_url": "",
      "post_date": "2020-02-07T07:20:45.960000",
      "content": "<h3>Coming soon</h3>\n\n<p>Part <code>7_0</code>, <code>7_1</code>, <code>7_2</code>, <code>7_3</code></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 738364,
      "author_name": "Darko Androcec",
      "author_url": "",
      "post_date": "2020-02-06T12:41:38.917000",
      "content": "<p>Deepfake Detection - Faces - Part 0_1 data set does not have download link. Do you maybe know why <a href=\"/phunghieu\">@phunghieu</a> ?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 738365,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-06T12:48:12.307000",
          "content": "<p>Let me check. Maybe, there are some problems when I upload it to Kaggle. Thanks for your reply!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 738476,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-06T15:18:49.077000",
          "content": "<p><a href=\"/dandrocec\">@dandrocec</a> now we have the download link for Part 0_1, let check it out 😉</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 738722,
          "author_name": "Darko Androcec",
          "author_url": "",
          "post_date": "2020-02-06T22:58:02.917000",
          "content": "<p>Thanks.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 738190,
      "author_name": "Senol Comert",
      "author_url": "",
      "post_date": "2020-02-06T08:47:57.450000",
      "content": "<p>Thanks a lot for sharing this.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 738207,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-06T09:07:33.503000",
          "content": "<p>You're welcome! <a href=\"/senolcomert\">@senolcomert</a> 😁</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 738133,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-06T07:19:05.757000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 737232,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-05T04:32:05.093000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 737237,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T04:38:49.937000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737249,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T05:11:56.557000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737256,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T05:25:20.527000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737264,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T05:35:19.780000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737272,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T05:43:24.187000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737345,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T08:05:00.403000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 736979,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-04T19:05:57.537000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 737063,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-04T21:38:04.350000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 737203,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T03:38:42.163000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 737217,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-05T04:01:05.913000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 736922,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-04T17:59:15.390000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 737056,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-04T21:29:02.283000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 747382,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-16T11:10:26.190000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 747442,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-16T12:52:45.653000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 746563,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-15T07:49:14.683000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 746687,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-15T12:01:14.940000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 740080,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-08T20:15:07.650000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 740122,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-08T23:13:13.977000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 744060,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-12T14:25:48.330000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 744014,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-12T13:30:46.787000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 740456,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-09T12:49:28.313000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 740438,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-09T12:33:51.817000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 740236,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-09T04:11:07.933000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 738458,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-06T14:48:20.893000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 738217,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-06T09:20:04.837000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 737517,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-05T13:10:18.730000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 737405,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-05T10:07:26.407000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 785929,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-25T14:17:17.307000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 786450,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-25T22:38:42.963000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 749466,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-18T18:13:12.153000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 749752,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-18T22:14:20.787000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 741189,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-10T10:00:17.970000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "736780": "I've created some datasets that include all detectable faces of all videos in each part of the full dataset. Kaggle and the host expected and encouraged us to train our models outside of Kaggle’s notebooks environment; however, for someone who prefers to stick to Kaggle's kernels, these preprocessed datasets would help a lot 😄.\n\nThe whole process to create and upload these datasets is time-consuming, so I'll gradually upload the rest for the next days; here are some completed ones:\n\n* [Deepfake Detection - Faces - Part 0_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-0)\n* [Deepfake Detection - Faces - Part 0_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-0-1)\n* [Deepfake Detection - Faces - Part 1_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-0)\n* [Deepfake Detection - Faces - Part 1_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-1-1)\n* [Deepfake Detection - Faces - Part 2_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-0)\n* [Deepfake Detection - Faces - Part 2_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-1)\n* [Deepfake Detection - Faces - Part 2_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-2-2)\n* [Deepfake Detection - Faces - Part 3_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-0)\n* [Deepfake Detection - Faces - Part 3_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-3-1)\n* [Deepfake Detection - Faces - Part 4_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-0)\n* [Deepfake Detection - Faces - Part 4_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-1)\n* [Deepfake Detection - Faces - Part 4_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-4-2)\n* [Deepfake Detection - Faces - Part 5_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-0)\n* [Deepfake Detection - Faces - Part 5_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-1)\n* [Deepfake Detection - Faces - Part 5_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-2)\n* [Deepfake Detection - Faces - Part 5_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-5-3)\n* [Deepfake Detection - Faces - Part 6_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-0)\n* [Deepfake Detection - Faces - Part 6_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-1)\n* [Deepfake Detection - Faces - Part 6_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-2)\n* [Deepfake Detection - Faces - Part 6_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-3)\n* [Deepfake Detection - Faces - Part 6_4](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-6-4)\n* [Deepfake Detection - Faces - Part 7_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-0)\n* [Deepfake Detection - Faces - Part 7_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-1)\n* [Deepfake Detection - Faces - Part 7_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-2)\n* [Deepfake Detection - Faces - Part 7_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-7-3)\n* [Deepfake Detection - Faces - Part 8_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-0)\n* [Deepfake Detection - Faces - Part 8_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-1)\n* [Deepfake Detection - Faces - Part 8_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-8-2)\n* [Deepfake Detection - Faces - Part 9_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-0)\n* [Deepfake Detection - Faces - Part 9_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-1)\n* [Deepfake Detection - Faces - Part 9_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-9-2)\n* [Deepfake Detection - Faces - Part 10_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-0)\n* [Deepfake Detection - Faces - Part 10_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-1)\n* [Deepfake Detection - Faces - Part 10_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-2)\n* [Deepfake Detection - Faces - Part 10_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-3)\n* [Deepfake Detection - Faces - Part 10_4](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-10-4)\n* [Deepfake Detection - Faces - Part 11_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-0)\n* [Deepfake Detection - Faces - Part 11_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-1)\n* [Deepfake Detection - Faces - Part 11_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-11-2)\n* [Deepfake Detection - Faces - Part 12_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-0)\n* [Deepfake Detection - Faces - Part 12_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-1)\n* [Deepfake Detection - Faces - Part 12_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-12-2)\n* [Deepfake Detection - Faces - Part 13_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-0)\n* [Deepfake Detection - Faces - Part 13_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-1)\n* [Deepfake Detection - Faces - Part 13_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-2)\n* [Deepfake Detection - Faces - Part 13_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-3)\n* [Deepfake Detection - Faces - Part 13_4](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-13-4)\n* [Deepfake Detection - Faces - Part 14_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-0)\n* [Deepfake Detection - Faces - Part 14_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-1)\n* [Deepfake Detection - Faces - Part 14_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-2)\n* [Deepfake Detection - Faces - Part 14_3](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-14-3)\n* [Deepfake Detection - Faces - Part 15_0](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-0)\n* [Deepfake Detection - Faces - Part 15_1](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-1)\n* [Deepfake Detection - Faces - Part 15_2](https://www.kaggle.com/phunghieu/deepfake-detection-faces-part-15-2)\n* [Deepfake Detection - Faces - Sample](https://www.kaggle.com/phunghieu/deepfake-detection-faces-sample)\n*Updating...*\n\nHow I create these datasets? Let's check out this demo kernel 👉 [\n*Deepfake Detection - Face Extractor*](https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor).\n\n---\n\nBecause of the big size of each part in the original full dataset, I have to break all extracted faces in each part into 2 or 3 smaller parts before zipping them, e.g. the name `deepfake-detection-faces-part-1-0` means it created from `part 1` of the full dataset and is the `first split`.\n\nThe name of each face image corresponds to the index of the frame that this face appears, plus a suffix _2, or _3, etc. if the number of faces in a frame is greater than 1.\n\nIn each dataset, I also attach a metadata.csv file which stores all information needed. The format of each file will look like this:\n|     |    filename    | split |    original    | label |\n|:---:|:--------------:|:-----:|:--------------:|:-----:|\n|  0  | aagfhgtpmv.mp4 | train | vudstovrck.mp4 | FAKE  |\n|  1  | aapnvogymq.mp4 | train | jdubbvfswz.mp4 | FAKE  |\n|  2  | abarnvbtwb.mp4 | train |                | REAL  |\n| ... | ...            | ...   | ...            | ...   |\n\n---\n\nWant something to get started using these datasets, let see [*DFDC-Multiface-Training*](https://www.kaggle.com/phunghieu/dfdc-multiface-training) &amp; [*DFDC-Multiface-Inference*](https://www.kaggle.com/phunghieu/dfdc-multiface-inference).\n\nDon't know how to load and merge multiple Kaggle datasets at once in a kernel, let check this [*demo*](https://www.kaggle.com/phunghieu/loading-merging-multiple-kaggle-datasets-demo).\n\nHere are a few lines of code to demonstrate how to load and prepare these data for the training process:\n\n`# Get path of metadata.csv`\n`metadata_path = os.path.join(TRAIN_DIR, 'metadata.csv')`\n\n` # Create DataFrame from metadata.csv`\n`train_df = pd.read_csv(metadata_path)`\n`train_df['label'].replace({'FAKE': 1, 'REAL': 0}, inplace=True)`\n\n`X = train_df['filename'].to_numpy()`\n`y = train_df['label'].to_numpy()`\n\n`# Split the dataset`\n`X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.3, random_state=123, stratify=y)`\n\n---\n\nConsider to upvote these datasets if you think they are worth using 😉, this will motivate me much to continue this dull work 💪, thanks!!!",
    "760532": "Thanks for sharing! Can you also share the full list(metadata.json) of labels of the whole dataset? ",
    "747916": "### Update\nPart `15_0`, `15_1`, `15_2`",
    "747582": "### Update\nPart `14_0`, `14_1`, `14_2`, `14_3`",
    "747310": "### Add examples\n[*DFDC-Multiface-Training*](https://www.kaggle.com/phunghieu/dfdc-multiface-training) &amp; [*DFDC-Multiface-Inference*](https://www.kaggle.com/phunghieu/dfdc-multiface-inference)",
    "747193": "### Coming soon\nPart `14_0`, `14_1`, `14_2`, `14_3`, `15_0`, `15_1`, `15_2`",
    "746502": "### Update\nPart `13_3`, `13_4`",
    "746153": "Thanks for the  share",
    "746129": "### Update\nPart `13_0`, `13_1`, `13_2`\n\n### Coming soon\nPart `13_3`, `13_4`",
    "745751": "### Coming soon\nPart `13_0`, `13_1`, `13_2`",
    "745057": "### Update\nPart `12_0`, `12_1`, `12_2`",
    "744681": "### Update\nPart `10_4`",
    "744617": "### Coming soon\nPart `12_0`, `12_1`, `12_2`",
    "744592": "Awesome work! I just want to confirm whether the different 'parts' (i.e. Part 0_1 and Part 8_2) are pre-processed in the exact same way? I'm assuming the numbering is basically denoting the pre-processing being done in distinct batches.",
    "743682": "### Coming soon\nPart `10_0`, `10_1`, `10_2`, `10_3`, `10_4`, `11_0`, `11_1`, `11_2`",
    "741402": "Nice work man! I'm struggling to process this huge dataset. Yours will help a lot!",
    "740172": "thanks for sharing!!! 😃 ",
    "740003": "### Update\nPart `6_0`, `6_1`, `6_2`, `7_3`",
    "739996": "Good job,dude\nThanks.. 👍 ",
    "739952": "Very useful great work thanks for sharing ",
    "739826": "### Update\nPart `7_2`",
    "739685": "## Risk of disqualification\n\nThere are doubts about the effectiveness of your manual face search approach for this competition. If you teach the model in this way, then the test data should be processed in the same way, because your model cannot find the face itself. But if you just touch manually the test data, then immediately violate the [rules of the competition](https://www.kaggle.com/c/deepfake-detection-challenge/rules):\n\n*Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.*",
    "739670": "@phunghieu  You are doing a great job. One suggestion is to save them as a video instead of images. It'll reduce size for each dataset to less then 2GB. Or even if you want to save as images try with .jpg extension. Ask for any code help and keep up the great work.",
    "739584": "Great job! Thanks for sharing @phunghieu !!😄 👍 ",
    "739493": "Thanks!",
    "739238": "### Update\nPart `7_0`, `7_1`",
    "739002": "Great datasets",
    "738916": "### Coming soon\nPart `7_0`, `7_1`, `7_2`, `7_3`",
    "738364": "Deepfake Detection - Faces - Part 0_1 data set does not have download link. Do you maybe know why @phunghieu ?",
    "738190": "Thanks a lot for sharing this.",
    "738133": "### Update\n[*Loading &amp; Merging Multiple Kaggle Datasets (Demo)*](https://www.kaggle.com/phunghieu/loading-merging-multiple-kaggle-datasets-demo)",
    "737232": "Great work! thanks for sharing this out.",
    "736979": "Hi, I think it would be helpful to share the face extraction code, as someone using this dataset would like to use the exact face extraction process during inference! Thanks for the hard work, upvoted.",
    "736922": "Thanks for your job, man! It`s great\n",
    "747382": "You are a new great asset to Kaggle by your wholeheartedly contribution.",
    "746563": "Thanku @phunghieu \n Thanku @phunghieu \nIs this based on the kaggle dataset ?\n1) Which model u used used in for face extraction MTCNN/dlib/Blaze..\n2) Have u taken multiple frames for face extraction,if yes how many ?\n3) How we can label this ,can we assumed image id be used for getting the label\nI hope you keeping safe if from china :)",
    "740080": "Thank you for sharing!!! Great job!",
    "744060": "### Update\nPart `10_0`, `10_1`, `10_2`, `10_3`",
    "744014": "### Update\nPart `11_0`, `11_1`, `11_2`",
    "740456": "### Update\nPart `9_0`, `9_1`, `9_2`",
    "740438": "### Update\nPart `8_0`, `8_1`, `8_2`",
    "740236": "### Update\nPart `6_3`, `6_4`\n\n### Coming soon\nPart `8_0`, `8_1`, `8_2`, `9_0`, `9_1`, `9_2`",
    "738458": "### Update\nPart `5_2`, `5_3`",
    "738217": "### Update\nPart `5_0`, `5_1`",
    "737517": "### Update\nPart `4_0`, `4_1`, `4_2`",
    "737405": "### Update\nPart `2_0`, `2_1`, `2_2`\n\n### Coming soon\nPart `4_0`, `4_1`, `4_2`",
    "749466": "Thanks for sharing !! ",
    "741189": "Thanks for sharing this!"
  }
}