{
  "id": 121441,
  "title": "DFDC:  Insights about the Dataset creation, baselines...",
  "url": "/competitions/deepfake-detection-challenge/discussion/121441",
  "author_name": "Nanashi",
  "post_date": "2019-12-13T10:43:36.243000",
  "votes": 52,
  "comment_count": 4,
  "views": 0,
  "content": "<p>All the information comes from the paper: <strong><a href=\"https://arxiv.org/pdf/1910.08854.pdf\">The Deepfake Detection Challenge (DFDC) Preview Dataset</a></strong>\n<em>Authors:  Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, Cristian Canton Ferrer (23 Oct 2019)</em></p>\n\n<p>I think is really important to read this paper in order to understand the challenge and the datset, and I guess you have time until you get all the data downloaded ;)</p>\n\n<p><code>This post is a summary with the key points</code></p>\n\n<hr>\n\n<p><strong>Abstract:</strong> <em>In this paper, we introduce a preview of the Deepfakes Detection Challenge (DFDC) dataset consisting of <strong>5K videos</strong> featuring <strong>two facial modification algorithms.</strong> A\ndata collection campaign has been carried out where participating actors have entered into an agreement to the use and manipulation of their likenesses in our creation of the dataset. Diversity in several axes (gender, skin-tone, age, etc.) has been considered and actors recorded videos with arbitrary backgrounds thus bringing visual variability. Finally, a set of specific metrics to evaluate the performance have been defined and two existing models for detecting deepfakes have been tested to provide a reference performance baseline.</em></p>\n\n<h1>Dataset Construction</h1>\n\n<ul>\n<li><p><strong>Actors</strong> were crowdsourced ensuring a variability in gender, skin tone and age. </p></li>\n<li><p>The videos include varied lighting conditions and head poses, and participants were able to record their videos with any background they desired, which yielded visually diverse backgrounds.</p></li>\n<li><p>The rough approximation of the <strong>general distribution of gender and race</strong> across this preview dataset is 74% female and 26% male; and 68% Caucasian, 20% African-American, 9% east-Asian, and 3% south-Asian.</p></li>\n<li><p>For this first version of the DFDC dataset, a small set of 66 individuals where chosen from the pool of crowdsourced actors, and split into a training and a testing set. This was done to avoid cross-set face swaps.</p></li>\n<li><p><strong>Two methods were selected to generate face swaps (noted as methods A and B in the dataset).</strong> A number of state-of-the-art methods will be applied to generate these videos, exploring the whole gamut of manipulation techniques available to generate such tamperings.</p></li>\n<li><p>A number of face swaps were computed across subjects with similar appearances, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.). After a given pairwise model was trained on two identities, we swapped each identity onto the other’s videos.</p></li>\n</ul>\n\n<p><strong>Algorithms</strong>\n- The initial algorithm in the dataset used to produce this dataset <code>method A</code> does not produce sharp or believable face swaps if the subject’s face is too close to the camera, so selfie videos or other close-up recordings resulted in easyto-spot fakes. \n- All swaps were performed on videos where the average face size ratio was less than 0.25\n- Second method (method B) that generally produces lower-quality swaps, but is similar to other\noff-the-shelf face swap algorithms.</p>\n\n<p><strong>Filtering</strong>\n- For each original or swapped video, we removed the first five seconds of the video as subjects\nwere often seen setting up their cameras during this time. From the remaining length of the video, we extracted multiple <strong>15 second clips</strong> - generally three clips per video if\nthe length was over 50 seconds. All of the clips that comprised the training set were left at their <strong>original resolution and quality</strong>, so deriving appropriate augmentations of the training set is left as an exercise to the researcher. </p>\n\n<p><strong>Augmentations</strong>\nIn this dataset, <strong>no video was subjected to more than one augmentation.</strong> \n1. reduce the FPS of the video to 15\n2. reduce the resolution of the video to 1/4 of its original size\n3. reduce the overall encoding quality. </p>\n\n<p><strong>This is important, a possible leak?</strong>\n- All information regarding the swapped and target identities for each video, along with the train or test set assignment and any augmentations applied to a video are listed in the file dataset.json, located in the dataset root directory. <strong>Swapped video filenames contain two IDs - the first is the swapped ID, and the second is the target ID.</strong> The final identifiers in the video refer to which target video the swap was produced from, and the clip within the target video that\nwas used.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F78fb12eedb98db83f264f841aa6e5bff%2FScreenshot%20from%202019-12-13%2011-25-39.png?generation=1576232855868296&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3602d0ea652ad5ef778d15b32ccd6d69%2FScreenshot%20from%202019-12-13%2011-26-04.png?generation=1576232855798180&amp;alt=media\" alt=\"\"></p>\n\n<h1>Baseline</h1>\n\n<p>3 simple detection models. </p>\n\n<ul>\n<li><p>The first model was a frame-based model which we denote as <strong>TamperNet.</strong>\nTamperNet is a small DNN (6 convolutional layers plus a 1 fully connected layer) trained to detect low-level image manipulations, such as cut-and-pasted objects or the addition of digital text to an image, and although it was not trained only on deepfake images, it performs well in identifying digitally-altered images in general (including face swaps).</p></li>\n<li><p>The other two models are the <strong>XceptionNet</strong> face detection and full-image models, trained on the FaceForensics data set. For these models, <strong>one frame was sampled per second of video.</strong> </p></li>\n<li><p>When <strong>using frame-based models for detection</strong>, there are two thresholds to tune:</p>\n\n<ol><li>the per-frame detection threshold</li>\n<li>a threshold that specifies how many frames must exceed the per-frame threshold in order to identify a video as fake (or the frames-per-video threshold). </li></ol></li>\n</ul>\n\n<p>These <strong>thresholds must be tuned</strong> in tandem - for good performance, a low per-frame threshold will probably result in a high framesper-video threshold, and vice-versa. </p>\n\n<ul>\n<li><p>To <strong>normalize for video length</strong>, we only evaluated the frames-per-video threshold over frames that contained a detectable face. </p></li>\n<li><p>During <strong>crossvalidation</strong> on the train set, we found the optimal frame- and video-thresholds that maximized the log-WP over each fold, while still maintaining the desired level of recall. </p></li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F0225c7226e8b529ce551620ae5a40f4c%2FScreenshot%20from%202019-12-13%2011-41-39.png?generation=1576233772170941&amp;alt=media\" alt=\"\"></p>\n\n<p>(More information about the metric on the paper)</p>\n\n<p><br></p>\n\n<p><code>I hope it helps you, I want to know your opinion on the dataset :)</code></p>",
  "messages": [
    {
      "id": 694246,
      "postDate": "2019-12-13T10:43:36.243Z",
      "content": "<p>All the information comes from the paper: <strong><a href=\"https://arxiv.org/pdf/1910.08854.pdf\">The Deepfake Detection Challenge (DFDC) Preview Dataset</a></strong>\n<em>Authors:  Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, Cristian Canton Ferrer (23 Oct 2019)</em></p>\n\n<p>I think is really important to read this paper in order to understand the challenge and the datset, and I guess you have time until you get all the data downloaded ;)</p>\n\n<p><code>This post is a summary with the key points</code></p>\n\n<hr>\n\n<p><strong>Abstract:</strong> <em>In this paper, we introduce a preview of the Deepfakes Detection Challenge (DFDC) dataset consisting of <strong>5K videos</strong> featuring <strong>two facial modification algorithms.</strong> A\ndata collection campaign has been carried out where participating actors have entered into an agreement to the use and manipulation of their likenesses in our creation of the dataset. Diversity in several axes (gender, skin-tone, age, etc.) has been considered and actors recorded videos with arbitrary backgrounds thus bringing visual variability. Finally, a set of specific metrics to evaluate the performance have been defined and two existing models for detecting deepfakes have been tested to provide a reference performance baseline.</em></p>\n\n<h1>Dataset Construction</h1>\n\n<ul>\n<li><p><strong>Actors</strong> were crowdsourced ensuring a variability in gender, skin tone and age. </p></li>\n<li><p>The videos include varied lighting conditions and head poses, and participants were able to record their videos with any background they desired, which yielded visually diverse backgrounds.</p></li>\n<li><p>The rough approximation of the <strong>general distribution of gender and race</strong> across this preview dataset is 74% female and 26% male; and 68% Caucasian, 20% African-American, 9% east-Asian, and 3% south-Asian.</p></li>\n<li><p>For this first version of the DFDC dataset, a small set of 66 individuals where chosen from the pool of crowdsourced actors, and split into a training and a testing set. This was done to avoid cross-set face swaps.</p></li>\n<li><p><strong>Two methods were selected to generate face swaps (noted as methods A and B in the dataset).</strong> A number of state-of-the-art methods will be applied to generate these videos, exploring the whole gamut of manipulation techniques available to generate such tamperings.</p></li>\n<li><p>A number of face swaps were computed across subjects with similar appearances, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.). After a given pairwise model was trained on two identities, we swapped each identity onto the other’s videos.</p></li>\n</ul>\n\n<p><strong>Algorithms</strong>\n- The initial algorithm in the dataset used to produce this dataset <code>method A</code> does not produce sharp or believable face swaps if the subject’s face is too close to the camera, so selfie videos or other close-up recordings resulted in easyto-spot fakes. \n- All swaps were performed on videos where the average face size ratio was less than 0.25\n- Second method (method B) that generally produces lower-quality swaps, but is similar to other\noff-the-shelf face swap algorithms.</p>\n\n<p><strong>Filtering</strong>\n- For each original or swapped video, we removed the first five seconds of the video as subjects\nwere often seen setting up their cameras during this time. From the remaining length of the video, we extracted multiple <strong>15 second clips</strong> - generally three clips per video if\nthe length was over 50 seconds. All of the clips that comprised the training set were left at their <strong>original resolution and quality</strong>, so deriving appropriate augmentations of the training set is left as an exercise to the researcher. </p>\n\n<p><strong>Augmentations</strong>\nIn this dataset, <strong>no video was subjected to more than one augmentation.</strong> \n1. reduce the FPS of the video to 15\n2. reduce the resolution of the video to 1/4 of its original size\n3. reduce the overall encoding quality. </p>\n\n<p><strong>This is important, a possible leak?</strong>\n- All information regarding the swapped and target identities for each video, along with the train or test set assignment and any augmentations applied to a video are listed in the file dataset.json, located in the dataset root directory. <strong>Swapped video filenames contain two IDs - the first is the swapped ID, and the second is the target ID.</strong> The final identifiers in the video refer to which target video the swap was produced from, and the clip within the target video that\nwas used.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F78fb12eedb98db83f264f841aa6e5bff%2FScreenshot%20from%202019-12-13%2011-25-39.png?generation=1576232855868296&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3602d0ea652ad5ef778d15b32ccd6d69%2FScreenshot%20from%202019-12-13%2011-26-04.png?generation=1576232855798180&amp;alt=media\" alt=\"\"></p>\n\n<h1>Baseline</h1>\n\n<p>3 simple detection models. </p>\n\n<ul>\n<li><p>The first model was a frame-based model which we denote as <strong>TamperNet.</strong>\nTamperNet is a small DNN (6 convolutional layers plus a 1 fully connected layer) trained to detect low-level image manipulations, such as cut-and-pasted objects or the addition of digital text to an image, and although it was not trained only on deepfake images, it performs well in identifying digitally-altered images in general (including face swaps).</p></li>\n<li><p>The other two models are the <strong>XceptionNet</strong> face detection and full-image models, trained on the FaceForensics data set. For these models, <strong>one frame was sampled per second of video.</strong> </p></li>\n<li><p>When <strong>using frame-based models for detection</strong>, there are two thresholds to tune:</p>\n\n<ol><li>the per-frame detection threshold</li>\n<li>a threshold that specifies how many frames must exceed the per-frame threshold in order to identify a video as fake (or the frames-per-video threshold). </li></ol></li>\n</ul>\n\n<p>These <strong>thresholds must be tuned</strong> in tandem - for good performance, a low per-frame threshold will probably result in a high framesper-video threshold, and vice-versa. </p>\n\n<ul>\n<li><p>To <strong>normalize for video length</strong>, we only evaluated the frames-per-video threshold over frames that contained a detectable face. </p></li>\n<li><p>During <strong>crossvalidation</strong> on the train set, we found the optimal frame- and video-thresholds that maximized the log-WP over each fold, while still maintaining the desired level of recall. </p></li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F0225c7226e8b529ce551620ae5a40f4c%2FScreenshot%20from%202019-12-13%2011-41-39.png?generation=1576233772170941&amp;alt=media\" alt=\"\"></p>\n\n<p>(More information about the metric on the paper)</p>\n\n<p><br></p>\n\n<p><code>I hope it helps you, I want to know your opinion on the dataset :)</code></p>",
      "rawMarkdown": "All the information comes from the paper: **[The Deepfake Detection Challenge (DFDC) Preview Dataset](https://arxiv.org/pdf/1910.08854.pdf)**\n*Authors:  Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, Cristian Canton Ferrer (23 Oct 2019)*\n\nI think is really important to read this paper in order to understand the challenge and the datset, and I guess you have time until you get all the data downloaded ;)\n\n```This post is a summary with the key points```\n\n---\n\n**Abstract:** *In this paper, we introduce a preview of the Deepfakes Detection Challenge (DFDC) dataset consisting of **5K videos** featuring **two facial modification algorithms.** A\ndata collection campaign has been carried out where participating actors have entered into an agreement to the use and manipulation of their likenesses in our creation of the dataset. Diversity in several axes (gender, skin-tone, age, etc.) has been considered and actors recorded videos with arbitrary backgrounds thus bringing visual variability. Finally, a set of specific metrics to evaluate the performance have been defined and two existing models for detecting deepfakes have been tested to provide a reference performance baseline.*\n\n# Dataset Construction\n\n- **Actors** were crowdsourced ensuring a variability in gender, skin tone and age. \n\n- The videos include varied lighting conditions and head poses, and participants were able to record their videos with any background they desired, which yielded visually diverse backgrounds.\n\n- The rough approximation of the **general distribution of gender and race** across this preview dataset is 74% female and 26% male; and 68% Caucasian, 20% African-American, 9% east-Asian, and 3% south-Asian.\n\n- For this first version of the DFDC dataset, a small set of 66 individuals where chosen from the pool of crowdsourced actors, and split into a training and a testing set. This was done to avoid cross-set face swaps.\n\n- **Two methods were selected to generate face swaps (noted as methods A and B in the dataset).** A number of state-of-the-art methods will be applied to generate these videos, exploring the whole gamut of manipulation techniques available to generate such tamperings.\n\n- A number of face swaps were computed across subjects with similar appearances, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.). After a given pairwise model was trained on two identities, we swapped each identity onto the other’s videos.\n\n**Algorithms**\n- The initial algorithm in the dataset used to produce this dataset ```method A``` does not produce sharp or believable face swaps if the subject’s face is too close to the camera, so selfie videos or other close-up recordings resulted in easyto-spot fakes. \n- All swaps were performed on videos where the average face size ratio was less than 0.25\n- Second method (method B) that generally produces lower-quality swaps, but is similar to other\noff-the-shelf face swap algorithms.\n\n**Filtering**\n- For each original or swapped video, we removed the first five seconds of the video as subjects\nwere often seen setting up their cameras during this time. From the remaining length of the video, we extracted multiple **15 second clips** - generally three clips per video if\nthe length was over 50 seconds. All of the clips that comprised the training set were left at their **original resolution and quality**, so deriving appropriate augmentations of the training set is left as an exercise to the researcher. \n\n**Augmentations**\nIn this dataset, **no video was subjected to more than one augmentation.** \n1. reduce the FPS of the video to 15\n2. reduce the resolution of the video to 1/4 of its original size\n3. reduce the overall encoding quality. \n\n\n**This is important, a possible leak?**\n- All information regarding the swapped and target identities for each video, along with the train or test set assignment and any augmentations applied to a video are listed in the file dataset.json, located in the dataset root directory. **Swapped video filenames contain two IDs - the first is the swapped ID, and the second is the target ID.** The final identifiers in the video refer to which target video the swap was produced from, and the clip within the target video that\nwas used.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F78fb12eedb98db83f264f841aa6e5bff%2FScreenshot%20from%202019-12-13%2011-25-39.png?generation=1576232855868296&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3602d0ea652ad5ef778d15b32ccd6d69%2FScreenshot%20from%202019-12-13%2011-26-04.png?generation=1576232855798180&amp;alt=media)\n\n# Baseline\n\n3 simple detection models. \n\n- The first model was a frame-based model which we denote as **TamperNet.**\nTamperNet is a small DNN (6 convolutional layers plus a 1 fully connected layer) trained to detect low-level image manipulations, such as cut-and-pasted objects or the addition of digital text to an image, and although it was not trained only on deepfake images, it performs well in identifying digitally-altered images in general (including face swaps).\n\n- The other two models are the **XceptionNet** face detection and full-image models, trained on the FaceForensics data set. For these models, **one frame was sampled per second of video.** \n\n- When **using frame-based models for detection**, there are two thresholds to tune:\n1. the per-frame detection threshold\n2. a threshold that specifies how many frames must exceed the per-frame threshold in order to identify a video as fake (or the frames-per-video threshold). \n\nThese **thresholds must be tuned** in tandem - for good performance, a low per-frame threshold will probably result in a high framesper-video threshold, and vice-versa. \n\n- To **normalize for video length**, we only evaluated the frames-per-video threshold over frames that contained a detectable face. \n\n- During **crossvalidation** on the train set, we found the optimal frame- and video-thresholds that maximized the log-WP over each fold, while still maintaining the desired level of recall. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F0225c7226e8b529ce551620ae5a40f4c%2FScreenshot%20from%202019-12-13%2011-41-39.png?generation=1576233772170941&amp;alt=media)\n\n(More information about the metric on the paper)\n\n<br>\n\n```I hope it helps you, I want to know your opinion on the dataset :)```",
      "votes": 52
    },
    {
      "id": 694494,
      "postDate": "2019-12-13T17:42:24.037Z",
      "content": "<p>Thanks for sharing! I don't really understand the thresholds here. The pipeline of the baseline is:\nTamperNet(First detection of bad swaps) &gt; XceptionNet(Face) &gt; XceptionNet Full image\nLet's focus on the frame-based models:one frame sampled per seconds of the video. so we don't have all the data on the video but sample of images of this video. The frame is the sample.  Am I right? What is the threshold here ? is it the number of images in each sample?</p>",
      "rawMarkdown": "Thanks for sharing! I don't really understand the thresholds here. The pipeline of the baseline is:\nTamperNet(First detection of bad swaps) &gt; XceptionNet(Face) &gt; XceptionNet Full image\nLet's focus on the frame-based models:one frame sampled per seconds of the video. so we don't have all the data on the video but sample of images of this video. The frame is the sample.  Am I right? What is the threshold here ? is it the number of images in each sample?",
      "votes": 1
    },
    {
      "id": 760316,
      "postDate": "2020-03-01T05:34:18.417Z",
      "content": "<p>The methods A and B used for the creation of Deepfakes are not disclosed? And when I check FPS using OpenCV, I am getting approximately 30 FPS for each video, but in the paper you mentioned, it's written 15 FPS. Why?</p>",
      "rawMarkdown": "The methods A and B used for the creation of Deepfakes are not disclosed? And when I check FPS using OpenCV, I am getting approximately 30 FPS for each video, but in the paper you mentioned, it's written 15 FPS. Why?"
    },
    {
      "id": 757842,
      "postDate": "2020-02-27T06:54:36.800Z",
      "content": "<p>Thanks a lot!</p>",
      "rawMarkdown": "Thanks a lot!"
    },
    {
      "id": 754085,
      "postDate": "2020-02-23T03:14:54.257Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 694494,
      "author_name": "AmelNozieres",
      "author_url": "",
      "post_date": "2019-12-13T17:42:24.037000",
      "content": "<p>Thanks for sharing! I don't really understand the thresholds here. The pipeline of the baseline is:\nTamperNet(First detection of bad swaps) &gt; XceptionNet(Face) &gt; XceptionNet Full image\nLet's focus on the frame-based models:one frame sampled per seconds of the video. so we don't have all the data on the video but sample of images of this video. The frame is the sample.  Am I right? What is the threshold here ? is it the number of images in each sample?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 760316,
      "author_name": "Saanika Gupta",
      "author_url": "",
      "post_date": "2020-03-01T05:34:18.417000",
      "content": "<p>The methods A and B used for the creation of Deepfakes are not disclosed? And when I check FPS using OpenCV, I am getting approximately 30 FPS for each video, but in the paper you mentioned, it's written 15 FPS. Why?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 757842,
      "author_name": "YiningLiu",
      "author_url": "",
      "post_date": "2020-02-27T06:54:36.800000",
      "content": "<p>Thanks a lot!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 754085,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-23T03:14:54.257000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "694246": "All the information comes from the paper: **[The Deepfake Detection Challenge (DFDC) Preview Dataset](https://arxiv.org/pdf/1910.08854.pdf)**\n*Authors:  Brian Dolhansky, Russ Howes, Ben Pflaum, Nicole Baram, Cristian Canton Ferrer (23 Oct 2019)*\n\nI think is really important to read this paper in order to understand the challenge and the datset, and I guess you have time until you get all the data downloaded ;)\n\n```This post is a summary with the key points```\n\n---\n\n**Abstract:** *In this paper, we introduce a preview of the Deepfakes Detection Challenge (DFDC) dataset consisting of **5K videos** featuring **two facial modification algorithms.** A\ndata collection campaign has been carried out where participating actors have entered into an agreement to the use and manipulation of their likenesses in our creation of the dataset. Diversity in several axes (gender, skin-tone, age, etc.) has been considered and actors recorded videos with arbitrary backgrounds thus bringing visual variability. Finally, a set of specific metrics to evaluate the performance have been defined and two existing models for detecting deepfakes have been tested to provide a reference performance baseline.*\n\n# Dataset Construction\n\n- **Actors** were crowdsourced ensuring a variability in gender, skin tone and age. \n\n- The videos include varied lighting conditions and head poses, and participants were able to record their videos with any background they desired, which yielded visually diverse backgrounds.\n\n- The rough approximation of the **general distribution of gender and race** across this preview dataset is 74% female and 26% male; and 68% Caucasian, 20% African-American, 9% east-Asian, and 3% south-Asian.\n\n- For this first version of the DFDC dataset, a small set of 66 individuals where chosen from the pool of crowdsourced actors, and split into a training and a testing set. This was done to avoid cross-set face swaps.\n\n- **Two methods were selected to generate face swaps (noted as methods A and B in the dataset).** A number of state-of-the-art methods will be applied to generate these videos, exploring the whole gamut of manipulation techniques available to generate such tamperings.\n\n- A number of face swaps were computed across subjects with similar appearances, where each appearance was inferred from facial attributes (skin tone, facial hair, glasses, etc.). After a given pairwise model was trained on two identities, we swapped each identity onto the other’s videos.\n\n**Algorithms**\n- The initial algorithm in the dataset used to produce this dataset ```method A``` does not produce sharp or believable face swaps if the subject’s face is too close to the camera, so selfie videos or other close-up recordings resulted in easyto-spot fakes. \n- All swaps were performed on videos where the average face size ratio was less than 0.25\n- Second method (method B) that generally produces lower-quality swaps, but is similar to other\noff-the-shelf face swap algorithms.\n\n**Filtering**\n- For each original or swapped video, we removed the first five seconds of the video as subjects\nwere often seen setting up their cameras during this time. From the remaining length of the video, we extracted multiple **15 second clips** - generally three clips per video if\nthe length was over 50 seconds. All of the clips that comprised the training set were left at their **original resolution and quality**, so deriving appropriate augmentations of the training set is left as an exercise to the researcher. \n\n**Augmentations**\nIn this dataset, **no video was subjected to more than one augmentation.** \n1. reduce the FPS of the video to 15\n2. reduce the resolution of the video to 1/4 of its original size\n3. reduce the overall encoding quality. \n\n\n**This is important, a possible leak?**\n- All information regarding the swapped and target identities for each video, along with the train or test set assignment and any augmentations applied to a video are listed in the file dataset.json, located in the dataset root directory. **Swapped video filenames contain two IDs - the first is the swapped ID, and the second is the target ID.** The final identifiers in the video refer to which target video the swap was produced from, and the clip within the target video that\nwas used.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F78fb12eedb98db83f264f841aa6e5bff%2FScreenshot%20from%202019-12-13%2011-25-39.png?generation=1576232855868296&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F3602d0ea652ad5ef778d15b32ccd6d69%2FScreenshot%20from%202019-12-13%2011-26-04.png?generation=1576232855798180&amp;alt=media)\n\n# Baseline\n\n3 simple detection models. \n\n- The first model was a frame-based model which we denote as **TamperNet.**\nTamperNet is a small DNN (6 convolutional layers plus a 1 fully connected layer) trained to detect low-level image manipulations, such as cut-and-pasted objects or the addition of digital text to an image, and although it was not trained only on deepfake images, it performs well in identifying digitally-altered images in general (including face swaps).\n\n- The other two models are the **XceptionNet** face detection and full-image models, trained on the FaceForensics data set. For these models, **one frame was sampled per second of video.** \n\n- When **using frame-based models for detection**, there are two thresholds to tune:\n1. the per-frame detection threshold\n2. a threshold that specifies how many frames must exceed the per-frame threshold in order to identify a video as fake (or the frames-per-video threshold). \n\nThese **thresholds must be tuned** in tandem - for good performance, a low per-frame threshold will probably result in a high framesper-video threshold, and vice-versa. \n\n- To **normalize for video length**, we only evaluated the frames-per-video threshold over frames that contained a detectable face. \n\n- During **crossvalidation** on the train set, we found the optimal frame- and video-thresholds that maximized the log-WP over each fold, while still maintaining the desired level of recall. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2779868%2F0225c7226e8b529ce551620ae5a40f4c%2FScreenshot%20from%202019-12-13%2011-41-39.png?generation=1576233772170941&amp;alt=media)\n\n(More information about the metric on the paper)\n\n<br>\n\n```I hope it helps you, I want to know your opinion on the dataset :)```",
    "694494": "Thanks for sharing! I don't really understand the thresholds here. The pipeline of the baseline is:\nTamperNet(First detection of bad swaps) &gt; XceptionNet(Face) &gt; XceptionNet Full image\nLet's focus on the frame-based models:one frame sampled per seconds of the video. so we don't have all the data on the video but sample of images of this video. The frame is the sample.  Am I right? What is the threshold here ? is it the number of images in each sample?",
    "760316": "The methods A and B used for the creation of Deepfakes are not disclosed? And when I check FPS using OpenCV, I am getting approximately 30 FPS for each video, but in the paper you mentioned, it's written 15 FPS. Why?",
    "757842": "Thanks a lot!",
    "754085": "Thanks for sharing"
  }
}