{
  "id": 140417,
  "title": "Next steps in the DFDC competition",
  "url": "/competitions/deepfake-detection-challenge/discussion/140417",
  "author_name": "Mozaic",
  "post_date": "2020-04-01T18:10:06.827000",
  "votes": 27,
  "comment_count": 29,
  "views": 0,
  "content": "<p>Submissions for the Deepfake Detection Challenge closed on March 31st at 11:59pm UTC. First of all, thank you for your participation in building innovative new technologies to help detect deepfakes. As detailed on the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started page</a> here's what you can expect for the next few weeks:\n- Facebook is working with Kaggle to export the code for your 2 selected submissions to our local servers, where we'll re-run your code submissions against a privately held test set. This private set has been created by a reduced closed set of engineers from Facebook who did not participate in the competition.\n- The private test set includes:\n  - Synthetic videos with a similar format and nature as the training and public validation/test sets (generated with the same methods used to create the training dataset)\n  - Organic videos with and without deepfakes\n- With the predictions generated by this private re-run, Kaggle will calculate your scores on the private leaderboard.\n- We'll reveal those scores on the private leaderboard in ~3 weeks time.\n- Those in the winners circle will have their code inspected by Facebook to validate that no violations have been committed. As a reminder, to be eligible for the challenge prizes, winners must agree to open-source their work so others in the research community can benefit. Entrants will retain rights to their models trained on the training dataset.\n- We'll officially announce the final winners in May!</p>\n\n<p>Keep tuned!</p>",
  "messages": [
    {
      "id": 794298,
      "postDate": "2020-04-01T18:10:06.827Z",
      "content": "<p>Submissions for the Deepfake Detection Challenge closed on March 31st at 11:59pm UTC. First of all, thank you for your participation in building innovative new technologies to help detect deepfakes. As detailed on the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started page</a> here's what you can expect for the next few weeks:\n- Facebook is working with Kaggle to export the code for your 2 selected submissions to our local servers, where we'll re-run your code submissions against a privately held test set. This private set has been created by a reduced closed set of engineers from Facebook who did not participate in the competition.\n- The private test set includes:\n  - Synthetic videos with a similar format and nature as the training and public validation/test sets (generated with the same methods used to create the training dataset)\n  - Organic videos with and without deepfakes\n- With the predictions generated by this private re-run, Kaggle will calculate your scores on the private leaderboard.\n- We'll reveal those scores on the private leaderboard in ~3 weeks time.\n- Those in the winners circle will have their code inspected by Facebook to validate that no violations have been committed. As a reminder, to be eligible for the challenge prizes, winners must agree to open-source their work so others in the research community can benefit. Entrants will retain rights to their models trained on the training dataset.\n- We'll officially announce the final winners in May!</p>\n\n<p>Keep tuned!</p>",
      "rawMarkdown": "Submissions for the Deepfake Detection Challenge closed on March 31st at 11:59pm UTC. First of all, thank you for your participation in building innovative new technologies to help detect deepfakes. As detailed on the [Getting Started page](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started) here's what you can expect for the next few weeks:\n- Facebook is working with Kaggle to export the code for your 2 selected submissions to our local servers, where we'll re-run your code submissions against a privately held test set. This private set has been created by a reduced closed set of engineers from Facebook who did not participate in the competition.\n- The private test set includes:\n  - Synthetic videos with a similar format and nature as the training and public validation/test sets (generated with the same methods used to create the training dataset)\n  - Organic videos with and without deepfakes\n- With the predictions generated by this private re-run, Kaggle will calculate your scores on the private leaderboard.\n- We'll reveal those scores on the private leaderboard in ~3 weeks time.\n- Those in the winners circle will have their code inspected by Facebook to validate that no violations have been committed. As a reminder, to be eligible for the challenge prizes, winners must agree to open-source their work so others in the research community can benefit. Entrants will retain rights to their models trained on the training dataset.\n- We'll officially announce the final winners in May!\n\nKeep tuned!\n",
      "votes": 27
    },
    {
      "id": 794385,
      "postDate": "2020-04-01T19:11:50.983Z",
      "content": "<p>Exciting! Out of interest, can you tell us the relative proportion of synthetic vs organic videos?</p>",
      "rawMarkdown": "Exciting! Out of interest, can you tell us the relative proportion of synthetic vs organic videos?",
      "votes": 8
    },
    {
      "id": 815834,
      "postDate": "2020-04-21T21:47:33.507Z",
      "content": "<p><strong>Update</strong>: We've pushed the expected private code re-run completion date to April 24th, as we're needing a few extra days to complete the re-runs. We'll post an update when the tentative private leaderboard (pending full code &amp; documentation validation by the host) is revealed.</p>",
      "rawMarkdown": "**Update**: We've pushed the expected private code re-run completion date to April 24th, as we're needing a few extra days to complete the re-runs. We'll post an update when the tentative private leaderboard (pending full code &amp; documentation validation by the host) is revealed.",
      "votes": 6,
      "replies": [
        {
          "id": 816406,
          "postDate": "2020-04-22T10:52:38.697Z",
          "content": "<p>Hi, have you adjusted the number of teams recently, since yesterday? Saw that I droped in the medalzone but thats maybe because of less teams and than a tighter medalzone. Is the PL medelzone calculated based on the original total of teams? Now my result will certainly differ against LB but good to know.</p>",
          "rawMarkdown": "Hi, have you adjusted the number of teams recently, since yesterday? Saw that I droped in the medalzone but thats maybe because of less teams and than a tighter medalzone. Is the PL medelzone calculated based on the original total of teams? Now my result will certainly differ against LB but good to know."
        },
        {
          "id": 816893,
          "postDate": "2020-04-22T17:26:47.440Z",
          "content": "<p>Private scoring is actively happening. You may expect to see some shuffling/changes in the leaderboard appearance during this process. It's not to be considered final until we announce that it is. There may be some continuous changes in total numbers of teams since the competition close, as we work through rules violators' eliminations.</p>",
          "rawMarkdown": "Private scoring is actively happening. You may expect to see some shuffling/changes in the leaderboard appearance during this process. It's not to be considered final until we announce that it is. There may be some continuous changes in total numbers of teams since the competition close, as we work through rules violators' eliminations.",
          "votes": 1
        },
        {
          "id": 816900,
          "postDate": "2020-04-22T17:35:42.193Z",
          "content": "<p>So Public LB shows both the public and private or a combo during the update process? Or is the shuffling/changes in the Public leaderboard happening due to rules violators' eliminations only?</p>",
          "rawMarkdown": "So Public LB shows both the public and private or a combo during the update process? Or is the shuffling/changes in the Public leaderboard happening due to rules violators' eliminations only?",
          "votes": 1
        },
        {
          "id": 818908,
          "postDate": "2020-04-24T07:45:33.003Z",
          "content": "<p>I would recommend organizers put their data on kaggle and re-run the submitted notebooks again. The results are really unbelievable.</p>",
          "rawMarkdown": "I would recommend organizers put their data on kaggle and re-run the submitted notebooks again. The results are really unbelievable.",
          "votes": 2
        }
      ]
    },
    {
      "id": 852654,
      "postDate": "2020-05-18T15:31:54.093Z",
      "content": "<p>Nearly a month already since deadline but this comp still cannot be concluded.</p>",
      "rawMarkdown": "Nearly a month already since deadline but this comp still cannot be concluded.",
      "votes": 3,
      "replies": [
        {
          "id": 852841,
          "postDate": "2020-05-18T18:01:03.777Z",
          "content": "<p>I guess they are checking the code thoroughly? 🤔 </p>",
          "rawMarkdown": "I guess they are checking the code thoroughly? 🤔 ",
          "votes": 1
        }
      ]
    },
    {
      "id": 823325,
      "postDate": "2020-04-27T15:14:56.507Z",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> I would hope that you will be able keep even the private testset on a benchmark server or something or release it as part of the full dataset for the broader community to work with ... I'm have a feeling this would be the imagenet of forensics ;)</p>",
      "rawMarkdown": "@cristiancanton I would hope that you will be able keep even the private testset on a benchmark server or something or release it as part of the full dataset for the broader community to work with ... I'm have a feeling this would be the imagenet of forensics ;)",
      "votes": 1
    },
    {
      "id": 797153,
      "postDate": "2020-04-04T08:44:42.587Z",
      "content": "<p>I’m asking a basic question, please somebody explain what is “organic real/fake video”?</p>",
      "rawMarkdown": "I’m asking a basic question, please somebody explain what is “organic real/fake video”?",
      "votes": 1,
      "replies": [
        {
          "id": 797749,
          "postDate": "2020-04-04T21:31:59.920Z",
          "content": "<p>The videos in the competition training set were synthetic, meaning they were composed and recorded for this competition. Organic means \"real\" videos from YouTube /Facebook (probably ones that were removed) made by many different people using different methods. Organic is the opposite of synthetic. </p>",
          "rawMarkdown": "The videos in the competition training set were synthetic, meaning they were composed and recorded for this competition. Organic means \"real\" videos from YouTube /Facebook (probably ones that were removed) made by many different people using different methods. Organic is the opposite of synthetic. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 794516,
      "postDate": "2020-04-01T21:54:19.200Z",
      "content": "<p>Great work 👍 </p>",
      "rawMarkdown": "Great work 👍 ",
      "votes": 1
    },
    {
      "id": 800278,
      "postDate": "2020-04-07T08:35:59.330Z",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> </p>\n\n<p>I have a question regarding the inclusion of synthetic videos in the private test set:</p>\n\n<p><strong>Will the recently announced synthetic videos in the private test set be inspected/chosen in a way to avoid inclusion of synthetic videos whose properties have virtually no resemblance to “real” organic deepfakes?</strong></p>\n\n<hr>\n\n<p><strong>Background</strong></p>\n\n<p>To my understanding the information that the private test set will include synthetic videos (and not exclusively real “organic” videos) was disclosed entirely after the submission deadline. </p>\n\n<p>During the competition the information given was that the evaluation would be made entirely on \"real, organic videos with and without deepfakes\".</p>\n\n<p>In the instructions on the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started page</a>, it says:</p>\n\n<blockquote>\n  <p>“<strong>Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard.</strong> It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.”</p>\n</blockquote>\n\n<p>Also, towards the end of the competition the exact above information was quoted in your post <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421\">Reminder from the organizers</a>.</p>\n\n<p>Aside the apparent possible difference between described evaluation in instructions during the competition and actual evaluation afterwards, I see another problem with an evaluation in part based on synthetic videos (given that the selection of synthetic videos would directly resemble the overall DFDC training set):</p>\n\n<p>The DFDC training set includes a fair proportion of videos labeled as fake but whose synthetic manipulation obviously has little, if any, resemblance to what could be considered the properties of a real “organic” audio- and/or video-deepfake. In relation to the competition objective to identify real organic deepfakes this could be viewed as problematic. </p>\n\n<p>Such examples in the DFDC training set include videos labeled as “fake” where the manipulation has been done, not on the actors themselves but on “thin air” beside the actors' bodies (see picture example below). Other examples include \"fake\" videos in the training set with close to zero difference (both audio and video) to their corresponding original videos. A third example is videos so dark that virtually nothing can be seen without ridiculously heavy brightness adjustment.</p>\n\n<p>For example, if included as basis for evaluation, should a video similiar to \"dlpaibcwsy.mp4\" be considered a deepfake? The tilted “fake” square to the left of the actor never gets even near the actors face region throughout the video. A strange special effect, sure. But how good does it serve as basis for deepfake detection evaluation? </p>\n\n<p>Fake (first frame of dlpaibcwsy.mp4 from training part 0)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2Fd59c60112a9f0dc14a48d346f9c0fc89%2Fdlpaibcwsy.mp4.png?generation=1586212515405190&amp;alt=media\" alt=\"\"></p>\n\n<p>Original (first frame of fopjiyxiqd.mp4 from training part 0)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2F0e273c093bb8d6b8e9ac47f9032d26be%2Ffopjiyxiqd.mp4.png?generation=1586212516017440&amp;alt=media\" alt=\"\"></p>\n\n<p>I am of course in no position to object to the organizers decision to include synthetic videos in the private test set. In fact I can see good motives for doing so. However, my opinion is that studying the winning submissions will be a much more interesting subject if the evaluation, as far as possible – with or without inclusion of synthetic videos, is done in a way relevant for the overall objective, i.e. to identify real world “organic” deepfakes.</p>\n\n<p>A broad inclusion of synthetic videos, without any kind of selection, could obviously have the undesirable effect of benefitting submissions over-adjusted towards non representative properties of the training set.</p>\n\n<p>Thus, if possible, I believe synthetic videos in the private test set should be inspected/chosen in a way that avoids inclusion of synthetic videos without any resemblance to “real” organic deepfakes.</p>\n\n<p>Best regards</p>",
      "rawMarkdown": "@cristiancanton \n\nI have a question regarding the inclusion of synthetic videos in the private test set:\n\n**Will the recently announced synthetic videos in the private test set be inspected/chosen in a way to avoid inclusion of synthetic videos whose properties have virtually no resemblance to “real” organic deepfakes?**\n\n______\n**Background**\n\nTo my understanding the information that the private test set will include synthetic videos (and not exclusively real “organic” videos) was disclosed entirely after the submission deadline. \n\nDuring the competition the information given was that the evaluation would be made entirely on \"real, organic videos with and without deepfakes\".\n\nIn the instructions on the [Getting Started page](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started), it says:\n&gt; “**Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard.** It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.”\n\nAlso, towards the end of the competition the exact above information was quoted in your post [Reminder from the organizers](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421).\n\nAside the apparent possible difference between described evaluation in instructions during the competition and actual evaluation afterwards, I see another problem with an evaluation in part based on synthetic videos (given that the selection of synthetic videos would directly resemble the overall DFDC training set):\n\nThe DFDC training set includes a fair proportion of videos labeled as fake but whose synthetic manipulation obviously has little, if any, resemblance to what could be considered the properties of a real “organic” audio- and/or video-deepfake. In relation to the competition objective to identify real organic deepfakes this could be viewed as problematic. \n\nSuch examples in the DFDC training set include videos labeled as “fake” where the manipulation has been done, not on the actors themselves but on “thin air” beside the actors' bodies (see picture example below). Other examples include \"fake\" videos in the training set with close to zero difference (both audio and video) to their corresponding original videos. A third example is videos so dark that virtually nothing can be seen without ridiculously heavy brightness adjustment.\n\nFor example, if included as basis for evaluation, should a video similiar to \"dlpaibcwsy.mp4\" be considered a deepfake? The tilted “fake” square to the left of the actor never gets even near the actors face region throughout the video. A strange special effect, sure. But how good does it serve as basis for deepfake detection evaluation? \n\nFake (first frame of dlpaibcwsy.mp4 from training part 0)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2Fd59c60112a9f0dc14a48d346f9c0fc89%2Fdlpaibcwsy.mp4.png?generation=1586212515405190&amp;alt=media)\n\nOriginal (first frame of fopjiyxiqd.mp4 from training part 0)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2F0e273c093bb8d6b8e9ac47f9032d26be%2Ffopjiyxiqd.mp4.png?generation=1586212516017440&amp;alt=media)\n\n\n\n\nI am of course in no position to object to the organizers decision to include synthetic videos in the private test set. In fact I can see good motives for doing so. However, my opinion is that studying the winning submissions will be a much more interesting subject if the evaluation, as far as possible – with or without inclusion of synthetic videos, is done in a way relevant for the overall objective, i.e. to identify real world “organic” deepfakes.\n\nA broad inclusion of synthetic videos, without any kind of selection, could obviously have the undesirable effect of benefitting submissions over-adjusted towards non representative properties of the training set.\n\nThus, if possible, I believe synthetic videos in the private test set should be inspected/chosen in a way that avoids inclusion of synthetic videos without any resemblance to “real” organic deepfakes.\n\nBest regards\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 800402,
          "postDate": "2020-04-07T11:32:58.810Z",
          "content": "<p>Couldn't agree more.. No Organic video out there in the real deepfake world has data like the ones above with things moving around</p>",
          "rawMarkdown": "Couldn't agree more.. No Organic video out there in the real deepfake world has data like the ones above with things moving around",
          "votes": 1
        }
      ]
    },
    {
      "id": 794504,
      "postDate": "2020-04-01T21:37:29.527Z",
      "content": "<p>For those asking: Sorry, late submissions will not be possible until the winners have been locked in and the private leaderboard is finalized. Check back in late April/May!</p>\n\n<p>The \"end date\" for the competition will continue to be a date into the future (currently April 22nd is our projected date), until late submissions (which don't count towards the leaderboard) open back up.</p>",
      "rawMarkdown": "For those asking: Sorry, late submissions will not be possible until the winners have been locked in and the private leaderboard is finalized. Check back in late April/May!\n\nThe \"end date\" for the competition will continue to be a date into the future (currently April 22nd is our projected date), until late submissions (which don't count towards the leaderboard) open back up.",
      "replies": [
        {
          "id": 794691,
          "postDate": "2020-04-02T02:36:43.657Z",
          "content": "<p>Could you open the test videos of public leaderboard, so we can evaluate models locally</p>",
          "rawMarkdown": "Could you open the test videos of public leaderboard, so we can evaluate models locally",
          "votes": 1
        },
        {
          "id": 795588,
          "postDate": "2020-04-02T21:05:20.597Z",
          "content": "<p><a href=\"/gaohaibin\">@gaohaibin</a> The host will be releasing the dataset and notify participants when that's available for continued exploration/work.</p>",
          "rawMarkdown": "@gaohaibin The host will be releasing the dataset and notify participants when that's available for continued exploration/work.",
          "votes": 1
        },
        {
          "id": 796974,
          "postDate": "2020-04-04T06:23:29.100Z",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a>  post declaration of winners.. if some one is able to produce new state of Art results  is any thing there for such solutions ?</p>",
          "rawMarkdown": "@juliaelliott  post declaration of winners.. if some one is able to produce new state of Art results  is any thing there for such solutions ?"
        },
        {
          "id": 809837,
          "postDate": "2020-04-16T14:09:44.240Z",
          "content": "<p>Is the private test set going to be available at some point later?</p>",
          "rawMarkdown": "Is the private test set going to be available at some point later?"
        },
        {
          "id": 830790,
          "postDate": "2020-05-02T22:42:43.507Z",
          "content": "<p>Is there any update as to when late submissions will open up? Thanks. </p>",
          "rawMarkdown": "Is there any update as to when late submissions will open up? Thanks. "
        }
      ]
    },
    {
      "id": 814208,
      "postDate": "2020-04-20T13:00:01.163Z",
      "content": "<p>Hi, hope the re-runs went well. Regarding the approximal PL releaseday, is Wednesday still in the loop you think?</p>",
      "rawMarkdown": "Hi, hope the re-runs went well. Regarding the approximal PL releaseday, is Wednesday still in the loop you think?"
    },
    {
      "id": 794430,
      "postDate": "2020-04-01T20:09:26.880Z",
      "content": "<p>Great.\nJust wondering if the late submissions will open earlier ..</p>",
      "rawMarkdown": "Great.\nJust wondering if the late submissions will open earlier .."
    },
    {
      "id": 794410,
      "postDate": "2020-04-01T19:40:34.843Z",
      "content": "<p>Awesome. Will there be a paper about the testing dataset generation? </p>",
      "rawMarkdown": "Awesome. Will there be a paper about the testing dataset generation? ",
      "replies": [
        {
          "id": 794769,
          "postDate": "2020-04-02T04:34:26.050Z",
          "content": "<p>Everything at due time but yes, there will be a paper.</p>",
          "rawMarkdown": "Everything at due time but yes, there will be a paper.",
          "votes": 3
        }
      ]
    },
    {
      "id": 853092,
      "postDate": "2020-05-19T00:45:20.173Z",
      "content": "<p>Just wonder, If there is no late submission, what are we gonna to do with the benchmark?</p>",
      "rawMarkdown": "Just wonder, If there is no late submission, what are we gonna to do with the benchmark?",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 853149,
          "postDate": "2020-05-19T02:11:23.310Z",
          "content": "<p>Nothing. :D </p>",
          "rawMarkdown": "Nothing. :D "
        },
        {
          "id": 853164,
          "postDate": "2020-05-19T02:18:57.030Z",
          "content": "<p>I don't think the results on private board could be regarded as the final result since nearly 6% competitors got no result caused by env differences between public and private LB. It's said that kaggles' employees are dealing with this. But nearly one month has gone....We got no updates.</p>",
          "rawMarkdown": "I don't think the results on private board could be regarded as the final result since nearly 6% competitors got no result caused by env differences between public and private LB. It's said that kaggles' employees are dealing with this. But nearly one month has gone....We got no updates.",
          "votes": 3
        }
      ]
    },
    {
      "id": 799304,
      "postDate": "2020-04-06T11:00:45.590Z",
      "content": "<p>Hi, can the learderboard keep open when the competion is over?</p>",
      "rawMarkdown": "Hi, can the learderboard keep open when the competion is over?",
      "isDeleted": true
    },
    {
      "id": 795456,
      "postDate": "2020-04-02T18:17:23.137Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 794385,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-04-01T19:11:50.983000",
      "content": "<p>Exciting! Out of interest, can you tell us the relative proportion of synthetic vs organic videos?</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 815834,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2020-04-21T21:47:33.507000",
      "content": "<p><strong>Update</strong>: We've pushed the expected private code re-run completion date to April 24th, as we're needing a few extra days to complete the re-runs. We'll post an update when the tentative private leaderboard (pending full code &amp; documentation validation by the host) is revealed.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 816406,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-04-22T10:52:38.697000",
          "content": "<p>Hi, have you adjusted the number of teams recently, since yesterday? Saw that I droped in the medalzone but thats maybe because of less teams and than a tighter medalzone. Is the PL medelzone calculated based on the original total of teams? Now my result will certainly differ against LB but good to know.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 816893,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-04-22T17:26:47.440000",
          "content": "<p>Private scoring is actively happening. You may expect to see some shuffling/changes in the leaderboard appearance during this process. It's not to be considered final until we announce that it is. There may be some continuous changes in total numbers of teams since the competition close, as we work through rules violators' eliminations.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 816900,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-04-22T17:35:42.193000",
          "content": "<p>So Public LB shows both the public and private or a combo during the update process? Or is the shuffling/changes in the Public leaderboard happening due to rules violators' eliminations only?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 818908,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-04-24T07:45:33.003000",
          "content": "<p>I would recommend organizers put their data on kaggle and re-run the submitted notebooks again. The results are really unbelievable.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 852654,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-05-18T15:31:54.093000",
      "content": "<p>Nearly a month already since deadline but this comp still cannot be concluded.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 852841,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-05-18T18:01:03.777000",
          "content": "<p>I guess they are checking the code thoroughly? 🤔 </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 823325,
      "author_name": "mjsML",
      "author_url": "",
      "post_date": "2020-04-27T15:14:56.507000",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> I would hope that you will be able keep even the private testset on a benchmark server or something or release it as part of the full dataset for the broader community to work with ... I'm have a feeling this would be the imagenet of forensics ;)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 797153,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-04-04T08:44:42.587000",
      "content": "<p>I’m asking a basic question, please somebody explain what is “organic real/fake video”?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 797749,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-04-04T21:31:59.920000",
          "content": "<p>The videos in the competition training set were synthetic, meaning they were composed and recorded for this competition. Organic means \"real\" videos from YouTube /Facebook (probably ones that were removed) made by many different people using different methods. Organic is the opposite of synthetic. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 794516,
      "author_name": "podsyp",
      "author_url": "",
      "post_date": "2020-04-01T21:54:19.200000",
      "content": "<p>Great work 👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 800278,
      "author_name": "Richard_U",
      "author_url": "",
      "post_date": "2020-04-07T08:35:59.330000",
      "content": "<p><a href=\"/cristiancanton\">@cristiancanton</a> </p>\n\n<p>I have a question regarding the inclusion of synthetic videos in the private test set:</p>\n\n<p><strong>Will the recently announced synthetic videos in the private test set be inspected/chosen in a way to avoid inclusion of synthetic videos whose properties have virtually no resemblance to “real” organic deepfakes?</strong></p>\n\n<hr>\n\n<p><strong>Background</strong></p>\n\n<p>To my understanding the information that the private test set will include synthetic videos (and not exclusively real “organic” videos) was disclosed entirely after the submission deadline. </p>\n\n<p>During the competition the information given was that the evaluation would be made entirely on \"real, organic videos with and without deepfakes\".</p>\n\n<p>In the instructions on the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started page</a>, it says:</p>\n\n<blockquote>\n  <p>“<strong>Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard.</strong> It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.”</p>\n</blockquote>\n\n<p>Also, towards the end of the competition the exact above information was quoted in your post <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421\">Reminder from the organizers</a>.</p>\n\n<p>Aside the apparent possible difference between described evaluation in instructions during the competition and actual evaluation afterwards, I see another problem with an evaluation in part based on synthetic videos (given that the selection of synthetic videos would directly resemble the overall DFDC training set):</p>\n\n<p>The DFDC training set includes a fair proportion of videos labeled as fake but whose synthetic manipulation obviously has little, if any, resemblance to what could be considered the properties of a real “organic” audio- and/or video-deepfake. In relation to the competition objective to identify real organic deepfakes this could be viewed as problematic. </p>\n\n<p>Such examples in the DFDC training set include videos labeled as “fake” where the manipulation has been done, not on the actors themselves but on “thin air” beside the actors' bodies (see picture example below). Other examples include \"fake\" videos in the training set with close to zero difference (both audio and video) to their corresponding original videos. A third example is videos so dark that virtually nothing can be seen without ridiculously heavy brightness adjustment.</p>\n\n<p>For example, if included as basis for evaluation, should a video similiar to \"dlpaibcwsy.mp4\" be considered a deepfake? The tilted “fake” square to the left of the actor never gets even near the actors face region throughout the video. A strange special effect, sure. But how good does it serve as basis for deepfake detection evaluation? </p>\n\n<p>Fake (first frame of dlpaibcwsy.mp4 from training part 0)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2Fd59c60112a9f0dc14a48d346f9c0fc89%2Fdlpaibcwsy.mp4.png?generation=1586212515405190&amp;alt=media\" alt=\"\"></p>\n\n<p>Original (first frame of fopjiyxiqd.mp4 from training part 0)\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2F0e273c093bb8d6b8e9ac47f9032d26be%2Ffopjiyxiqd.mp4.png?generation=1586212516017440&amp;alt=media\" alt=\"\"></p>\n\n<p>I am of course in no position to object to the organizers decision to include synthetic videos in the private test set. In fact I can see good motives for doing so. However, my opinion is that studying the winning submissions will be a much more interesting subject if the evaluation, as far as possible – with or without inclusion of synthetic videos, is done in a way relevant for the overall objective, i.e. to identify real world “organic” deepfakes.</p>\n\n<p>A broad inclusion of synthetic videos, without any kind of selection, could obviously have the undesirable effect of benefitting submissions over-adjusted towards non representative properties of the training set.</p>\n\n<p>Thus, if possible, I believe synthetic videos in the private test set should be inspected/chosen in a way that avoids inclusion of synthetic videos without any resemblance to “real” organic deepfakes.</p>\n\n<p>Best regards</p>",
      "votes": 2,
      "replies": [
        {
          "id": 800402,
          "author_name": "WiseLearner",
          "author_url": "",
          "post_date": "2020-04-07T11:32:58.810000",
          "content": "<p>Couldn't agree more.. No Organic video out there in the real deepfake world has data like the ones above with things moving around</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 794504,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2020-04-01T21:37:29.527000",
      "content": "<p>For those asking: Sorry, late submissions will not be possible until the winners have been locked in and the private leaderboard is finalized. Check back in late April/May!</p>\n\n<p>The \"end date\" for the competition will continue to be a date into the future (currently April 22nd is our projected date), until late submissions (which don't count towards the leaderboard) open back up.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 794691,
          "author_name": "Gao Haibin",
          "author_url": "",
          "post_date": "2020-04-02T02:36:43.657000",
          "content": "<p>Could you open the test videos of public leaderboard, so we can evaluate models locally</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 795588,
          "author_name": "Julia Elliott",
          "author_url": "",
          "post_date": "2020-04-02T21:05:20.597000",
          "content": "<p><a href=\"/gaohaibin\">@gaohaibin</a> The host will be releasing the dataset and notify participants when that's available for continued exploration/work.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 796974,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-04-04T06:23:29.100000",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a>  post declaration of winners.. if some one is able to produce new state of Art results  is any thing there for such solutions ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 809837,
          "author_name": "MohamedMaher",
          "author_url": "",
          "post_date": "2020-04-16T14:09:44.240000",
          "content": "<p>Is the private test set going to be available at some point later?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 830790,
          "author_name": "qwerty",
          "author_url": "",
          "post_date": "2020-05-02T22:42:43.507000",
          "content": "<p>Is there any update as to when late submissions will open up? Thanks. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 814208,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-04-20T13:00:01.163000",
      "content": "<p>Hi, hope the re-runs went well. Regarding the approximal PL releaseday, is Wednesday still in the loop you think?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 794430,
      "author_name": "Saurabh7",
      "author_url": "",
      "post_date": "2020-04-01T20:09:26.880000",
      "content": "<p>Great.\nJust wondering if the late submissions will open earlier ..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 794410,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-04-01T19:40:34.843000",
      "content": "<p>Awesome. Will there be a paper about the testing dataset generation? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 794769,
          "author_name": "Mozaic",
          "author_url": "",
          "post_date": "2020-04-02T04:34:26.050000",
          "content": "<p>Everything at due time but yes, there will be a paper.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 853092,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-19T00:45:20.173000",
      "content": "<p>Just wonder, If there is no late submission, what are we gonna to do with the benchmark?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 853149,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-05-19T02:11:23.310000",
          "content": "<p>Nothing. :D </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 853164,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-05-19T02:18:57.030000",
          "content": "<p>I don't think the results on private board could be regarded as the final result since nearly 6% competitors got no result caused by env differences between public and private LB. It's said that kaggles' employees are dealing with this. But nearly one month has gone....We got no updates.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 799304,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-06T11:00:45.590000",
      "content": "<p>Hi, can the learderboard keep open when the competion is over?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 795456,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-02T18:17:23.137000",
      "content": "",
      "votes": -1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "794298": "Submissions for the Deepfake Detection Challenge closed on March 31st at 11:59pm UTC. First of all, thank you for your participation in building innovative new technologies to help detect deepfakes. As detailed on the [Getting Started page](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started) here's what you can expect for the next few weeks:\n- Facebook is working with Kaggle to export the code for your 2 selected submissions to our local servers, where we'll re-run your code submissions against a privately held test set. This private set has been created by a reduced closed set of engineers from Facebook who did not participate in the competition.\n- The private test set includes:\n  - Synthetic videos with a similar format and nature as the training and public validation/test sets (generated with the same methods used to create the training dataset)\n  - Organic videos with and without deepfakes\n- With the predictions generated by this private re-run, Kaggle will calculate your scores on the private leaderboard.\n- We'll reveal those scores on the private leaderboard in ~3 weeks time.\n- Those in the winners circle will have their code inspected by Facebook to validate that no violations have been committed. As a reminder, to be eligible for the challenge prizes, winners must agree to open-source their work so others in the research community can benefit. Entrants will retain rights to their models trained on the training dataset.\n- We'll officially announce the final winners in May!\n\nKeep tuned!\n",
    "794385": "Exciting! Out of interest, can you tell us the relative proportion of synthetic vs organic videos?",
    "815834": "**Update**: We've pushed the expected private code re-run completion date to April 24th, as we're needing a few extra days to complete the re-runs. We'll post an update when the tentative private leaderboard (pending full code &amp; documentation validation by the host) is revealed.",
    "852654": "Nearly a month already since deadline but this comp still cannot be concluded.",
    "823325": "@cristiancanton I would hope that you will be able keep even the private testset on a benchmark server or something or release it as part of the full dataset for the broader community to work with ... I'm have a feeling this would be the imagenet of forensics ;)",
    "797153": "I’m asking a basic question, please somebody explain what is “organic real/fake video”?",
    "794516": "Great work 👍 ",
    "800278": "@cristiancanton \n\nI have a question regarding the inclusion of synthetic videos in the private test set:\n\n**Will the recently announced synthetic videos in the private test set be inspected/chosen in a way to avoid inclusion of synthetic videos whose properties have virtually no resemblance to “real” organic deepfakes?**\n\n______\n**Background**\n\nTo my understanding the information that the private test set will include synthetic videos (and not exclusively real “organic” videos) was disclosed entirely after the submission deadline. \n\nDuring the competition the information given was that the evaluation would be made entirely on \"real, organic videos with and without deepfakes\".\n\nIn the instructions on the [Getting Started page](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started), it says:\n&gt; “**Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard.** It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.”\n\nAlso, towards the end of the competition the exact above information was quoted in your post [Reminder from the organizers](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421).\n\nAside the apparent possible difference between described evaluation in instructions during the competition and actual evaluation afterwards, I see another problem with an evaluation in part based on synthetic videos (given that the selection of synthetic videos would directly resemble the overall DFDC training set):\n\nThe DFDC training set includes a fair proportion of videos labeled as fake but whose synthetic manipulation obviously has little, if any, resemblance to what could be considered the properties of a real “organic” audio- and/or video-deepfake. In relation to the competition objective to identify real organic deepfakes this could be viewed as problematic. \n\nSuch examples in the DFDC training set include videos labeled as “fake” where the manipulation has been done, not on the actors themselves but on “thin air” beside the actors' bodies (see picture example below). Other examples include \"fake\" videos in the training set with close to zero difference (both audio and video) to their corresponding original videos. A third example is videos so dark that virtually nothing can be seen without ridiculously heavy brightness adjustment.\n\nFor example, if included as basis for evaluation, should a video similiar to \"dlpaibcwsy.mp4\" be considered a deepfake? The tilted “fake” square to the left of the actor never gets even near the actors face region throughout the video. A strange special effect, sure. But how good does it serve as basis for deepfake detection evaluation? \n\nFake (first frame of dlpaibcwsy.mp4 from training part 0)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2Fd59c60112a9f0dc14a48d346f9c0fc89%2Fdlpaibcwsy.mp4.png?generation=1586212515405190&amp;alt=media)\n\nOriginal (first frame of fopjiyxiqd.mp4 from training part 0)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4218673%2F0e273c093bb8d6b8e9ac47f9032d26be%2Ffopjiyxiqd.mp4.png?generation=1586212516017440&amp;alt=media)\n\n\n\n\nI am of course in no position to object to the organizers decision to include synthetic videos in the private test set. In fact I can see good motives for doing so. However, my opinion is that studying the winning submissions will be a much more interesting subject if the evaluation, as far as possible – with or without inclusion of synthetic videos, is done in a way relevant for the overall objective, i.e. to identify real world “organic” deepfakes.\n\nA broad inclusion of synthetic videos, without any kind of selection, could obviously have the undesirable effect of benefitting submissions over-adjusted towards non representative properties of the training set.\n\nThus, if possible, I believe synthetic videos in the private test set should be inspected/chosen in a way that avoids inclusion of synthetic videos without any resemblance to “real” organic deepfakes.\n\nBest regards\n\n",
    "794504": "For those asking: Sorry, late submissions will not be possible until the winners have been locked in and the private leaderboard is finalized. Check back in late April/May!\n\nThe \"end date\" for the competition will continue to be a date into the future (currently April 22nd is our projected date), until late submissions (which don't count towards the leaderboard) open back up.",
    "814208": "Hi, hope the re-runs went well. Regarding the approximal PL releaseday, is Wednesday still in the loop you think?",
    "794430": "Great.\nJust wondering if the late submissions will open earlier ..",
    "794410": "Awesome. Will there be a paper about the testing dataset generation? ",
    "853092": "Just wonder, If there is no late submission, what are we gonna to do with the benchmark?",
    "799304": "Hi, can the learderboard keep open when the competion is over?",
    "795456": ""
  }
}