{
  "id": 121179,
  "title": "Welcome!",
  "url": "/competitions/deepfake-detection-challenge/discussion/121179",
  "author_name": "Phil Culliton",
  "post_date": "2019-12-11T17:58:46.586000",
  "votes": 19,
  "comment_count": 28,
  "views": 0,
  "content": "<p>Welcome to the Deepfake Detection Challenge.</p>\n\n<p>In this competition, we're identifying deepfake videos. The training set is very large - please read over our Data page for instructions on how to access it! We have provided the training data in smaller chunks for those who can't access the entire dataset. The metric is simple log loss.</p>\n\n<p>Note that the timeline for this competition is slightly different from others. For example, the entry / merger deadline is earlier than usual. Please make sure to check out the timeline so you'll be aware of important dates!</p>\n\n<p>Please let us know if you have any questions. Best of luck!</p>",
  "messages": [
    {
      "id": 692775,
      "postDate": "2019-12-11T17:58:46.587Z",
      "content": "<p>Welcome to the Deepfake Detection Challenge.</p>\n\n<p>In this competition, we're identifying deepfake videos. The training set is very large - please read over our Data page for instructions on how to access it! We have provided the training data in smaller chunks for those who can't access the entire dataset. The metric is simple log loss.</p>\n\n<p>Note that the timeline for this competition is slightly different from others. For example, the entry / merger deadline is earlier than usual. Please make sure to check out the timeline so you'll be aware of important dates!</p>\n\n<p>Please let us know if you have any questions. Best of luck!</p>",
      "rawMarkdown": "Welcome to the Deepfake Detection Challenge.\n\nIn this competition, we're identifying deepfake videos. The training set is very large - please read over our Data page for instructions on how to access it! We have provided the training data in smaller chunks for those who can't access the entire dataset. The metric is simple log loss.\n\nNote that the timeline for this competition is slightly different from others. For example, the entry / merger deadline is earlier than usual. Please make sure to check out the timeline so you'll be aware of important dates!\n\nPlease let us know if you have any questions. Best of luck!",
      "votes": 19
    },
    {
      "id": 702926,
      "postDate": "2019-12-25T11:13:20.030Z",
      "content": "<p>I'd like to double check that <code>test_videos.zip</code> in the Overview and Data sections actually means <code>test_videos/</code> directory. </p>\n\n<p>For example:</p>\n\n<blockquote>\n  <p>Your code must output a submission.csv that predicts on any set of test_videos.zip</p>\n</blockquote>\n\n<p>We never need to handle zip files to make predictions through notebooks, incuding in the private test phase, right? Thanks.</p>",
      "rawMarkdown": "I'd like to double check that `test_videos.zip` in the Overview and Data sections actually means `test_videos/` directory. \n\nFor example:\n\n&gt; Your code must output a submission.csv that predicts on any set of test_videos.zip\n\nWe never need to handle zip files to make predictions through notebooks, incuding in the private test phase, right? Thanks.",
      "votes": 3
    },
    {
      "id": 700247,
      "postDate": "2019-12-21T17:30:34.117Z",
      "content": "<p>Hi <a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> I had a question about the Private Test Set. The description says that the Private Test Set contains: \"videos with a similar format and nature... <strong>but are real, organic videos</strong> with and without deepfakes\".\nCould you elaborate more about this if possible? What does it mean to have a similar format &amp; nature? And what is meant by <strong>real and organic</strong>, e.g. not of paid actors like the DFDC training dataset?</p>",
      "rawMarkdown": "Hi @juliaelliott @philculliton I had a question about the Private Test Set. The description says that the Private Test Set contains: \"videos with a similar format and nature... **but are real, organic videos** with and without deepfakes\".\nCould you elaborate more about this if possible? What does it mean to have a similar format &amp; nature? And what is meant by **real and organic**, e.g. not of paid actors like the DFDC training dataset?",
      "votes": 3
    },
    {
      "id": 694753,
      "postDate": "2019-12-14T03:14:46.780Z",
      "content": "<p>Hi, Culliton\nThank you for sparing a time to answer my questions.\n1、I notice that Kaggle/DFDC Dataset contains not only face manipulation data but also voice manipulation data. For now, no matter the data are manipulated in face, voice or both, all the data are labeled as FAKE. Are there any plans on giving face or voice-manipulation-label?\n2、What are the manipulation range of the datasets? In detail, as for the face manipulation data, are all the data face replacement just like the DFDC Preview Dataset OR some replacement and others reenactment? As for the voice manipulation data, whether the source actor giving the same speech as target actor or the source actor giving random speech different from the original one?\n3、In the Getting Started session, the 400 videos named as test_videos make up the Public Validation Set. And the leaderboard-score should be log-loss of prediction on the Public Test Set. However, I find out that kagglers are submitting csv of the Public Validation Set after observing open notebooks. The contradiction make me confused.\n4、The validation set should provide both data and label in most situations. For now, the Public Validation Set doesn’t provide label which makes it looks like a test set. In that case, we still need a validation set to evaluate algorithm during offline training.</p>",
      "rawMarkdown": "Hi, Culliton\nThank you for sparing a time to answer my questions.\n1、I notice that Kaggle/DFDC Dataset contains not only face manipulation data but also voice manipulation data. For now, no matter the data are manipulated in face, voice or both, all the data are labeled as FAKE. Are there any plans on giving face or voice-manipulation-label?\n2、What are the manipulation range of the datasets? In detail, as for the face manipulation data, are all the data face replacement just like the DFDC Preview Dataset OR some replacement and others reenactment? As for the voice manipulation data, whether the source actor giving the same speech as target actor or the source actor giving random speech different from the original one?\n3、In the Getting Started session, the 400 videos named as test_videos make up the Public Validation Set. And the leaderboard-score should be log-loss of prediction on the Public Test Set. However, I find out that kagglers are submitting csv of the Public Validation Set after observing open notebooks. The contradiction make me confused.\n4、The validation set should provide both data and label in most situations. For now, the Public Validation Set doesn’t provide label which makes it looks like a test set. In that case, we still need a validation set to evaluate algorithm during offline training.\n",
      "votes": 3,
      "replies": [
        {
          "id": 738529,
          "postDate": "2020-02-06T16:46:25.693Z",
          "content": "<p>I have the same doubts , a lot people are only exploring  face detection algorithms, but for  what I understood the videos could be manipulated in more different ways  not only the face</p>",
          "rawMarkdown": "I have the same doubts , a lot people are only exploring  face detection algorithms, but for  what I understood the videos could be manipulated in more different ways  not only the face"
        },
        {
          "id": 747077,
          "postDate": "2020-02-15T23:39:15.337Z",
          "content": "<p><a href=\"/eduardohd\">@eduardohd</a> Please refer to the subtitle of this competition.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F84c3194e3d49769cce0f9ae8dd8cc4e2%2FScreen%20Shot%202020-02-15%20at%203.37.58%20PM.png?generation=1581809915436241&amp;alt=media\" alt=\"\"></p>\n\n<p>This is only \"FACIAL\" or voice manipulation. </p>",
          "rawMarkdown": "@eduardohd Please refer to the subtitle of this competition.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F84c3194e3d49769cce0f9ae8dd8cc4e2%2FScreen%20Shot%202020-02-15%20at%203.37.58%20PM.png?generation=1581809915436241&amp;alt=media)\n\nThis is only \"FACIAL\" or voice manipulation. "
        }
      ]
    },
    {
      "id": 696091,
      "postDate": "2019-12-16T05:29:05.067Z",
      "content": "<p>Hi Phil,</p>\n\n<p>Could you please let me know if the all the frames in the FAKE videos are deep faked?</p>\n\n<p>Thanks</p>\n\n<p>Michael</p>",
      "rawMarkdown": "Hi Phil,\n\nCould you please let me know if the all the frames in the FAKE videos are deep faked?\n\nThanks\n\nMichael",
      "votes": 1,
      "replies": [
        {
          "id": 696314,
          "postDate": "2019-12-16T13:01:36.100Z",
          "content": "<p>I guess that's your job to find out</p>",
          "rawMarkdown": "I guess that's your job to find out"
        },
        {
          "id": 778027,
          "postDate": "2020-03-18T05:01:43.090Z",
          "content": "<p>I don't think that is possible.</p>",
          "rawMarkdown": "I don't think that is possible."
        }
      ]
    },
    {
      "id": 694130,
      "postDate": "2019-12-13T08:17:14.070Z",
      "content": "<blockquote>\n  <p>External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models</p>\n</blockquote>\n\n<p>We are obliged to load our trained models as external data, right?\nShould we share them publicly according to this rule too?</p>\n\n<p>Sorry for a probably naive question. This rule confused me a bit.</p>",
      "rawMarkdown": "&gt; External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models\n\nWe are obliged to load our trained models as external data, right?\nShould we share them publicly according to this rule too?\n\nSorry for a probably naive question. This rule confused me a bit.",
      "votes": 1,
      "replies": [
        {
          "id": 694465,
          "postDate": "2019-12-13T16:43:40.847Z",
          "content": "<p><a href=\"/hokmund\">@hokmund</a> Not, naive at all! Your trained model <strong>does need</strong> to be uploaded as external data into your submission notebook and subject to the 1GB constraint. However no, your trained model <strong>does not need</strong> to be declared publicly on the forum. What needs to be publicly declared are any datasets you’re using from external sources - like pretrained models or other datasets that you use for your training, for example.</p>",
          "rawMarkdown": "@hokmund Not, naive at all! Your trained model **does need** to be uploaded as external data into your submission notebook and subject to the 1GB constraint. However no, your trained model **does not need** to be declared publicly on the forum. What needs to be publicly declared are any datasets you’re using from external sources - like pretrained models or other datasets that you use for your training, for example.",
          "votes": 3
        }
      ]
    },
    {
      "id": 693972,
      "postDate": "2019-12-13T02:13:34.633Z",
      "content": "<p>Hi, Culliton</p>\n\n<p>I am a participant of DFDC challenge. I still have some questions about some rule details.</p>\n\n<ol>\n<li><p>First question is about this rule \"External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models\" \nIt means 1GB size files can be submitted (containing code &amp; models)? \nIs it permissible that using external and private training data (generate additional fake videos by ourself) during training the detection model offline?</p></li>\n<li><p>Second question is about \"A deepfake could be either a face or voice swap (or both)\" \nIt means if the voice of the video are tampered (while face is REAL), the video is still seen as a FAKE one ?\nAnd how to define a voice swap? Is the video voice changed to another different person's voice which matches the mouth shape, or another kind of sound which is totally different from the frame and the mouth shape.</p></li>\n</ol>",
      "rawMarkdown": "Hi, Culliton\n\nI am a participant of DFDC challenge. I still have some questions about some rule details.\n\n1. First question is about this rule \"External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models\" \nIt means 1GB size files can be submitted (containing code &amp; models)? \nIs it permissible that using external and private training data (generate additional fake videos by ourself) during training the detection model offline?\n\n2. Second question is about \"A deepfake could be either a face or voice swap (or both)\" \nIt means if the voice of the video are tampered (while face is REAL), the video is still seen as a FAKE one ?\nAnd how to define a voice swap? Is the video voice changed to another different person's voice which matches the mouth shape, or another kind of sound which is totally different from the frame and the mouth shape.",
      "votes": 1,
      "replies": [
        {
          "id": 694112,
          "postDate": "2019-12-13T07:53:37.360Z",
          "content": "<p>Hi <a href=\"/xiaofengmao\">@xiaofengmao</a>:\n1) All external data loaded into your Kaggle submission notebook must be limited to 1 GB or less. If your submission contains greater than than this size of external data, it will be invalidated. Any external data used, in offline training or otherwise, must be shared on the external data forum thread prior to the entry deadline. So if you use any datasets of videos (including privately generated) as external data, it must be posted there.\n2) A voice swap is exactly as it sounds - the voice has been altered on the deepfake video.</p>",
          "rawMarkdown": "Hi @xiaofengmao:\n1) All external data loaded into your Kaggle submission notebook must be limited to 1 GB or less. If your submission contains greater than than this size of external data, it will be invalidated. Any external data used, in offline training or otherwise, must be shared on the external data forum thread prior to the entry deadline. So if you use any datasets of videos (including privately generated) as external data, it must be posted there.\n2) A voice swap is exactly as it sounds - the voice has been altered on the deepfake video.",
          "votes": 3
        },
        {
          "id": 694133,
          "postDate": "2019-12-13T08:19:29.297Z",
          "content": "<blockquote>\n  <p>So if you use any datasets of videos (including privately generated) as external data, it must be posted there.</p>\n</blockquote>\n\n<p>Let's assume I generated some deepfakes for the sake of data augmentation. Do I have to post my generated deepfakes? Or do I have to post the data which I used to train my GAN?</p>",
          "rawMarkdown": "&gt; So if you use any datasets of videos (including privately generated) as external data, it must be posted there.\n\nLet's assume I generated some deepfakes for the sake of data augmentation. Do I have to post my generated deepfakes? Or do I have to post the data which I used to train my GAN?",
          "votes": 2
        },
        {
          "id": 698627,
          "postDate": "2019-12-19T13:45:10.990Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> How do you measure the size of the dataset? If I have a kaggle dataset with a gzip in it, does this count the gzipped size of the dataset (the one when including the dataset) or the unzipped size - given that kaggle datasets automatically unzips your files.</p>",
          "rawMarkdown": "@juliaelliott How do you measure the size of the dataset? If I have a kaggle dataset with a gzip in it, does this count the gzipped size of the dataset (the one when including the dataset) or the unzipped size - given that kaggle datasets automatically unzips your files."
        }
      ]
    },
    {
      "id": 3107220,
      "postDate": "2025-01-26T06:44:30.797Z",
      "content": "<p>Can anyone help me downloading this dataset locally on my machine in 2025? <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> </p>",
      "rawMarkdown": "Can anyone help me downloading this dataset locally on my machine in 2025? @philculliton "
    },
    {
      "id": 2957950,
      "postDate": "2024-08-13T15:01:09.690Z",
      "content": "<p>Have a good time, I want to download the Dataset of your competition and I used the Kaggle API, CLI and did everything to download this Dataset and made every setting, but it is not possible for me to download and it gives me the error 403 Forbidden to download this Dataset according to the same command as you To download the data set in the data section, please answer me as soon as possible and guide me.</p>",
      "rawMarkdown": "Have a good time, I want to download the Dataset of your competition and I used the Kaggle API, CLI and did everything to download this Dataset and made every setting, but it is not possible for me to download and it gives me the error 403 Forbidden to download this Dataset according to the same command as you To download the data set in the data section, please answer me as soon as possible and guide me."
    },
    {
      "id": 785462,
      "postDate": "2020-03-25T04:48:57.167Z",
      "content": "<p>I‘d like to know if there are videos in the testset with much larger dimensions, something larger than 1920x1080, or some videos with much smaller dimensions, like 480P or 320P. </p>\n\n<p><a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> .</p>",
      "rawMarkdown": "I‘d like to know if there are videos in the testset with much larger dimensions, something larger than 1920x1080, or some videos with much smaller dimensions, like 480P or 320P. \n\n@juliaelliott @philculliton ."
    },
    {
      "id": 776320,
      "postDate": "2020-03-17T09:56:08.190Z",
      "content": "<p>Hi <a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> </p>\n\n<p>I have a question about Kernel. I have 2 kernels. They have different version of Cuda. When it runs for private test set, can I suppose that they still keep using different versions?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Ff48d69b771b388a27db273c3f14277fe%2F202003175456.png?generation=1584438929841908&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Fc27a9b13365da816d9a40ad09e0f59c0%2F202003175338.png?generation=1584438932311347&amp;alt=media\" alt=\"\"></p>\n\n<p>Thanks in advance,</p>",
      "rawMarkdown": "Hi @juliaelliott @philculliton \n\nI have a question about Kernel. I have 2 kernels. They have different version of Cuda. When it runs for private test set, can I suppose that they still keep using different versions?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Ff48d69b771b388a27db273c3f14277fe%2F202003175456.png?generation=1584438929841908&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Fc27a9b13365da816d9a40ad09e0f59c0%2F202003175338.png?generation=1584438932311347&amp;alt=media)\n\nThanks in advance,\n"
    },
    {
      "id": 765383,
      "postDate": "2020-03-06T14:59:48.923Z",
      "content": "<p>Are any other method to deal with train data(size=470GB) without download?</p>",
      "rawMarkdown": "Are any other method to deal with train data(size=470GB) without download?"
    },
    {
      "id": 700868,
      "postDate": "2019-12-22T18:51:07.543Z",
      "content": "<p>By training from outside server other than Kaggle Notebook it seems that you are giving advantage to those with more computing power. Shouldn't you make this available to people who do not have access to servers with powerful GPUs? Is it possible to train the model in google colab, for example? </p>",
      "rawMarkdown": "By training from outside server other than Kaggle Notebook it seems that you are giving advantage to those with more computing power. Shouldn't you make this available to people who do not have access to servers with powerful GPUs? Is it possible to train the model in google colab, for example? "
    },
    {
      "id": 760492,
      "postDate": "2020-03-01T11:17:14.050Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 760707,
          "postDate": "2020-03-01T16:29:34.580Z",
          "content": "<p>Does it mean, that more people (even from 2 organizations) work behind the scene for \"your\" results? Isn't it prohibited?</p>",
          "rawMarkdown": "Does it mean, that more people (even from 2 organizations) work behind the scene for \"your\" results? Isn't it prohibited?"
        }
      ]
    },
    {
      "id": 694761,
      "postDate": "2019-12-14T03:23:01.937Z",
      "rawMarkdown": "",
      "votes": 3,
      "isDeleted": true,
      "replies": [
        {
          "id": 697332,
          "postDate": "2019-12-17T18:26:40.660Z",
          "content": "<p>Hi <a href=\"/yzfyzf\">@yzfyzf</a> - thanks for your questions.</p>\n\n<p>1-3. Good questions! We don't have plans to reveal any of that information. Models should be based on what's provided.</p>\n\n<ol>\n<li>We've structured the challenge so that the resources provided would be adequate. If people are running into difficulty loading the videos that may be a separate issue - I'll check it out.</li>\n</ol>\n\n<p>Thanks!</p>",
          "rawMarkdown": "Hi @yzfyzf - thanks for your questions.\n\n1-3. Good questions! We don't have plans to reveal any of that information. Models should be based on what's provided.\n\n4. We've structured the challenge so that the resources provided would be adequate. If people are running into difficulty loading the videos that may be a separate issue - I'll check it out.\n\nThanks!"
        },
        {
          "id": 698180,
          "postDate": "2019-12-18T22:39:36.253Z",
          "content": "<p>+1 for this - it takes me about 4 seconds to load each video, which corresponds to about four and half hours to load all the test data (out of a total 9 hours in kernels). This leaves just 4 seconds per video - meaning on top of this, our model has to be 2x realtime, not an easy feat! Especially in the constrained kaggle compute environment which makes optimisation difficult.</p>",
          "rawMarkdown": "+1 for this - it takes me about 4 seconds to load each video, which corresponds to about four and half hours to load all the test data (out of a total 9 hours in kernels). This leaves just 4 seconds per video - meaning on top of this, our model has to be 2x realtime, not an easy feat! Especially in the constrained kaggle compute environment which makes optimisation difficult.",
          "votes": 1
        }
      ]
    },
    {
      "id": 694665,
      "postDate": "2019-12-14T00:23:11.880Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 697323,
          "postDate": "2019-12-17T18:08:56.427Z",
          "content": "<p>Hi <a href=\"/algohunt\">@algohunt</a> - thanks for your feedback. I very much understand where you're coming from, but we have to lean towards safety on this one. We don't plan to give exemptions.</p>",
          "rawMarkdown": "Hi @algohunt - thanks for your feedback. I very much understand where you're coming from, but we have to lean towards safety on this one. We don't plan to give exemptions.",
          "votes": 1
        }
      ]
    },
    {
      "id": 802475,
      "postDate": "2020-04-09T14:43:19.180Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    }
  ],
  "comments": [
    {
      "id": 702926,
      "author_name": "🐢 Jun Koda",
      "author_url": "",
      "post_date": "2019-12-25T11:13:20.030000",
      "content": "<p>I'd like to double check that <code>test_videos.zip</code> in the Overview and Data sections actually means <code>test_videos/</code> directory. </p>\n\n<p>For example:</p>\n\n<blockquote>\n  <p>Your code must output a submission.csv that predicts on any set of test_videos.zip</p>\n</blockquote>\n\n<p>We never need to handle zip files to make predictions through notebooks, incuding in the private test phase, right? Thanks.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 700247,
      "author_name": "Rafi Hai",
      "author_url": "",
      "post_date": "2019-12-21T17:30:34.117000",
      "content": "<p>Hi <a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> I had a question about the Private Test Set. The description says that the Private Test Set contains: \"videos with a similar format and nature... <strong>but are real, organic videos</strong> with and without deepfakes\".\nCould you elaborate more about this if possible? What does it mean to have a similar format &amp; nature? And what is meant by <strong>real and organic</strong>, e.g. not of paid actors like the DFDC training dataset?</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 694753,
      "author_name": "Jasmine",
      "author_url": "",
      "post_date": "2019-12-14T03:14:46.780000",
      "content": "<p>Hi, Culliton\nThank you for sparing a time to answer my questions.\n1、I notice that Kaggle/DFDC Dataset contains not only face manipulation data but also voice manipulation data. For now, no matter the data are manipulated in face, voice or both, all the data are labeled as FAKE. Are there any plans on giving face or voice-manipulation-label?\n2、What are the manipulation range of the datasets? In detail, as for the face manipulation data, are all the data face replacement just like the DFDC Preview Dataset OR some replacement and others reenactment? As for the voice manipulation data, whether the source actor giving the same speech as target actor or the source actor giving random speech different from the original one?\n3、In the Getting Started session, the 400 videos named as test_videos make up the Public Validation Set. And the leaderboard-score should be log-loss of prediction on the Public Test Set. However, I find out that kagglers are submitting csv of the Public Validation Set after observing open notebooks. The contradiction make me confused.\n4、The validation set should provide both data and label in most situations. For now, the Public Validation Set doesn’t provide label which makes it looks like a test set. In that case, we still need a validation set to evaluate algorithm during offline training.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 738529,
          "author_name": "EduardoHD",
          "author_url": "",
          "post_date": "2020-02-06T16:46:25.693000",
          "content": "<p>I have the same doubts , a lot people are only exploring  face detection algorithms, but for  what I understood the videos could be manipulated in more different ways  not only the face</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747077,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-15T23:39:15.337000",
          "content": "<p><a href=\"/eduardohd\">@eduardohd</a> Please refer to the subtitle of this competition.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F84c3194e3d49769cce0f9ae8dd8cc4e2%2FScreen%20Shot%202020-02-15%20at%203.37.58%20PM.png?generation=1581809915436241&amp;alt=media\" alt=\"\"></p>\n\n<p>This is only \"FACIAL\" or voice manipulation. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 696091,
      "author_name": "Schroter",
      "author_url": "",
      "post_date": "2019-12-16T05:29:05.067000",
      "content": "<p>Hi Phil,</p>\n\n<p>Could you please let me know if the all the frames in the FAKE videos are deep faked?</p>\n\n<p>Thanks</p>\n\n<p>Michael</p>",
      "votes": 1,
      "replies": [
        {
          "id": 696314,
          "author_name": "Behrouz",
          "author_url": "",
          "post_date": "2019-12-16T13:01:36.100000",
          "content": "<p>I guess that's your job to find out</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 778027,
          "author_name": "Schroter",
          "author_url": "",
          "post_date": "2020-03-18T05:01:43.090000",
          "content": "<p>I don't think that is possible.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 694130,
      "author_name": "Dmytro Panchenko",
      "author_url": "",
      "post_date": "2019-12-13T08:17:14.070000",
      "content": "<blockquote>\n  <p>External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models</p>\n</blockquote>\n\n<p>We are obliged to load our trained models as external data, right?\nShould we share them publicly according to this rule too?</p>\n\n<p>Sorry for a probably naive question. This rule confused me a bit.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 694465,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-12-13T16:43:40.847000",
          "content": "<p><a href=\"/hokmund\">@hokmund</a> Not, naive at all! Your trained model <strong>does need</strong> to be uploaded as external data into your submission notebook and subject to the 1GB constraint. However no, your trained model <strong>does not need</strong> to be declared publicly on the forum. What needs to be publicly declared are any datasets you’re using from external sources - like pretrained models or other datasets that you use for your training, for example.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 693972,
      "author_name": "vtddggg",
      "author_url": "",
      "post_date": "2019-12-13T02:13:34.633000",
      "content": "<p>Hi, Culliton</p>\n\n<p>I am a participant of DFDC challenge. I still have some questions about some rule details.</p>\n\n<ol>\n<li><p>First question is about this rule \"External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models\" \nIt means 1GB size files can be submitted (containing code &amp; models)? \nIs it permissible that using external and private training data (generate additional fake videos by ourself) during training the detection model offline?</p></li>\n<li><p>Second question is about \"A deepfake could be either a face or voice swap (or both)\" \nIt means if the voice of the video are tampered (while face is REAL), the video is still seen as a FAKE one ?\nAnd how to define a voice swap? Is the video voice changed to another different person's voice which matches the mouth shape, or another kind of sound which is totally different from the frame and the mouth shape.</p></li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 694112,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2019-12-13T07:53:37.360000",
          "content": "<p>Hi <a href=\"/xiaofengmao\">@xiaofengmao</a>:\n1) All external data loaded into your Kaggle submission notebook must be limited to 1 GB or less. If your submission contains greater than than this size of external data, it will be invalidated. Any external data used, in offline training or otherwise, must be shared on the external data forum thread prior to the entry deadline. So if you use any datasets of videos (including privately generated) as external data, it must be posted there.\n2) A voice swap is exactly as it sounds - the voice has been altered on the deepfake video.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 694133,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2019-12-13T08:19:29.297000",
          "content": "<blockquote>\n  <p>So if you use any datasets of videos (including privately generated) as external data, it must be posted there.</p>\n</blockquote>\n\n<p>Let's assume I generated some deepfakes for the sake of data augmentation. Do I have to post my generated deepfakes? Or do I have to post the data which I used to train my GAN?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 698627,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2019-12-19T13:45:10.990000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> How do you measure the size of the dataset? If I have a kaggle dataset with a gzip in it, does this count the gzipped size of the dataset (the one when including the dataset) or the unzipped size - given that kaggle datasets automatically unzips your files.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3107220,
      "author_name": "Meet Rajpal",
      "author_url": "",
      "post_date": "2025-01-26T06:44:30.797000",
      "content": "<p>Can anyone help me downloading this dataset locally on my machine in 2025? <a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2957950,
      "author_name": "Sofia*Amouei",
      "author_url": "",
      "post_date": "2024-08-13T15:01:09.690000",
      "content": "<p>Have a good time, I want to download the Dataset of your competition and I used the Kaggle API, CLI and did everything to download this Dataset and made every setting, but it is not possible for me to download and it gives me the error 403 Forbidden to download this Dataset according to the same command as you To download the data set in the data section, please answer me as soon as possible and guide me.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 785462,
      "author_name": "ascu",
      "author_url": "",
      "post_date": "2020-03-25T04:48:57.167000",
      "content": "<p>I‘d like to know if there are videos in the testset with much larger dimensions, something larger than 1920x1080, or some videos with much smaller dimensions, like 480P or 320P. </p>\n\n<p><a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 776320,
      "author_name": "AkiraSosa",
      "author_url": "",
      "post_date": "2020-03-17T09:56:08.190000",
      "content": "<p>Hi <a href=\"/juliaelliott\">@juliaelliott</a> <a href=\"/philculliton\">@philculliton</a> </p>\n\n<p>I have a question about Kernel. I have 2 kernels. They have different version of Cuda. When it runs for private test set, can I suppose that they still keep using different versions?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Ff48d69b771b388a27db273c3f14277fe%2F202003175456.png?generation=1584438929841908&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Fc27a9b13365da816d9a40ad09e0f59c0%2F202003175338.png?generation=1584438932311347&amp;alt=media\" alt=\"\"></p>\n\n<p>Thanks in advance,</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 765383,
      "author_name": "Shivam",
      "author_url": "",
      "post_date": "2020-03-06T14:59:48.923000",
      "content": "<p>Are any other method to deal with train data(size=470GB) without download?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 700868,
      "author_name": "friendlyjun",
      "author_url": "",
      "post_date": "2019-12-22T18:51:07.543000",
      "content": "<p>By training from outside server other than Kaggle Notebook it seems that you are giving advantage to those with more computing power. Shouldn't you make this available to people who do not have access to servers with powerful GPUs? Is it possible to train the model in google colab, for example? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 760492,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-01T11:17:14.050000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 760707,
          "author_name": "Allie K.",
          "author_url": "",
          "post_date": "2020-03-01T16:29:34.580000",
          "content": "<p>Does it mean, that more people (even from 2 organizations) work behind the scene for \"your\" results? Isn't it prohibited?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 694761,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-14T03:23:01.937000",
      "content": "",
      "votes": 3,
      "replies": [
        {
          "id": 697332,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-12-17T18:26:40.660000",
          "content": "<p>Hi <a href=\"/yzfyzf\">@yzfyzf</a> - thanks for your questions.</p>\n\n<p>1-3. Good questions! We don't have plans to reveal any of that information. Models should be based on what's provided.</p>\n\n<ol>\n<li>We've structured the challenge so that the resources provided would be adequate. If people are running into difficulty loading the videos that may be a separate issue - I'll check it out.</li>\n</ol>\n\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 698180,
          "author_name": "anokas",
          "author_url": "",
          "post_date": "2019-12-18T22:39:36.253000",
          "content": "<p>+1 for this - it takes me about 4 seconds to load each video, which corresponds to about four and half hours to load all the test data (out of a total 9 hours in kernels). This leaves just 4 seconds per video - meaning on top of this, our model has to be 2x realtime, not an easy feat! Especially in the constrained kaggle compute environment which makes optimisation difficult.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 694665,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-14T00:23:11.880000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 697323,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-12-17T18:08:56.427000",
          "content": "<p>Hi <a href=\"/algohunt\">@algohunt</a> - thanks for your feedback. I very much understand where you're coming from, but we have to lean towards safety on this one. We don't plan to give exemptions.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 802475,
      "author_name": "Pranav M R",
      "author_url": "",
      "post_date": "2020-04-09T14:43:19.180000",
      "content": "<p>Thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "692775": "Welcome to the Deepfake Detection Challenge.\n\nIn this competition, we're identifying deepfake videos. The training set is very large - please read over our Data page for instructions on how to access it! We have provided the training data in smaller chunks for those who can't access the entire dataset. The metric is simple log loss.\n\nNote that the timeline for this competition is slightly different from others. For example, the entry / merger deadline is earlier than usual. Please make sure to check out the timeline so you'll be aware of important dates!\n\nPlease let us know if you have any questions. Best of luck!",
    "702926": "I'd like to double check that `test_videos.zip` in the Overview and Data sections actually means `test_videos/` directory. \n\nFor example:\n\n&gt; Your code must output a submission.csv that predicts on any set of test_videos.zip\n\nWe never need to handle zip files to make predictions through notebooks, incuding in the private test phase, right? Thanks.",
    "700247": "Hi @juliaelliott @philculliton I had a question about the Private Test Set. The description says that the Private Test Set contains: \"videos with a similar format and nature... **but are real, organic videos** with and without deepfakes\".\nCould you elaborate more about this if possible? What does it mean to have a similar format &amp; nature? And what is meant by **real and organic**, e.g. not of paid actors like the DFDC training dataset?",
    "694753": "Hi, Culliton\nThank you for sparing a time to answer my questions.\n1、I notice that Kaggle/DFDC Dataset contains not only face manipulation data but also voice manipulation data. For now, no matter the data are manipulated in face, voice or both, all the data are labeled as FAKE. Are there any plans on giving face or voice-manipulation-label?\n2、What are the manipulation range of the datasets? In detail, as for the face manipulation data, are all the data face replacement just like the DFDC Preview Dataset OR some replacement and others reenactment? As for the voice manipulation data, whether the source actor giving the same speech as target actor or the source actor giving random speech different from the original one?\n3、In the Getting Started session, the 400 videos named as test_videos make up the Public Validation Set. And the leaderboard-score should be log-loss of prediction on the Public Test Set. However, I find out that kagglers are submitting csv of the Public Validation Set after observing open notebooks. The contradiction make me confused.\n4、The validation set should provide both data and label in most situations. For now, the Public Validation Set doesn’t provide label which makes it looks like a test set. In that case, we still need a validation set to evaluate algorithm during offline training.\n",
    "696091": "Hi Phil,\n\nCould you please let me know if the all the frames in the FAKE videos are deep faked?\n\nThanks\n\nMichael",
    "694130": "&gt; External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models\n\nWe are obliged to load our trained models as external data, right?\nShould we share them publicly according to this rule too?\n\nSorry for a probably naive question. This rule confused me a bit.",
    "693972": "Hi, Culliton\n\nI am a participant of DFDC challenge. I still have some questions about some rule details.\n\n1. First question is about this rule \"External data is allowed up to 1 GB in size. External data must be freely &amp; publicly available, including pre-trained models\" \nIt means 1GB size files can be submitted (containing code &amp; models)? \nIs it permissible that using external and private training data (generate additional fake videos by ourself) during training the detection model offline?\n\n2. Second question is about \"A deepfake could be either a face or voice swap (or both)\" \nIt means if the voice of the video are tampered (while face is REAL), the video is still seen as a FAKE one ?\nAnd how to define a voice swap? Is the video voice changed to another different person's voice which matches the mouth shape, or another kind of sound which is totally different from the frame and the mouth shape.",
    "3107220": "Can anyone help me downloading this dataset locally on my machine in 2025? @philculliton ",
    "2957950": "Have a good time, I want to download the Dataset of your competition and I used the Kaggle API, CLI and did everything to download this Dataset and made every setting, but it is not possible for me to download and it gives me the error 403 Forbidden to download this Dataset according to the same command as you To download the data set in the data section, please answer me as soon as possible and guide me.",
    "785462": "I‘d like to know if there are videos in the testset with much larger dimensions, something larger than 1920x1080, or some videos with much smaller dimensions, like 480P or 320P. \n\n@juliaelliott @philculliton .",
    "776320": "Hi @juliaelliott @philculliton \n\nI have a question about Kernel. I have 2 kernels. They have different version of Cuda. When it runs for private test set, can I suppose that they still keep using different versions?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Ff48d69b771b388a27db273c3f14277fe%2F202003175456.png?generation=1584438929841908&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1333985%2Fc27a9b13365da816d9a40ad09e0f59c0%2F202003175338.png?generation=1584438932311347&amp;alt=media)\n\nThanks in advance,\n",
    "765383": "Are any other method to deal with train data(size=470GB) without download?",
    "700868": "By training from outside server other than Kaggle Notebook it seems that you are giving advantage to those with more computing power. Shouldn't you make this available to people who do not have access to servers with powerful GPUs? Is it possible to train the model in google colab, for example? ",
    "760492": "",
    "694761": "",
    "694665": "",
    "802475": "Thanks"
  }
}