{
  "id": 121524,
  "title": "Some stats",
  "url": "/competitions/deepfake-detection-challenge/discussion/121524",
  "author_name": "",
  "post_date": "2019-12-13T21:06:38.492105200Z",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>The training dataset contains 99,992 fake videos. (The metadata says 100,000 but 8 of those files seem to be missing.)</p>\n\n<p>The training dataset also contains 19154 real videos.</p>\n\n<p>This means many of the real videos have been used to create more than one fake video (5 on average). In fact, some real videos have been used to create up to 40 fake videos.</p>\n\n<p>I compared the file sizes of the fake videos that belong to any given real video, and in 60% of the case the fake videos have a larger file size than the real ones. (In 20% of the cases, all the fake videos were smaller than the original, and in the remaining 20% they sometimes were smaller, sometimes larger.) </p>\n\n<p>So <em>if</em> (and it's a big if) you could determine for any set of videos that they are all the fake videos that are derived from a given real video -- but you don't know which is which -- then you can classify the smallest file as real and the rest as fake and you'd be 60% accurate. </p>\n\n<p>(Of course, this assumes that the test set also contains the real videos that were used to create the fake videos, which may not be the case. So it's probably not a very useful statistic to know but there you have it in any case.)</p>",
  "messages": [
    {
      "id": "694616",
      "postDate": "12/13/2019 21:06:38",
      "content": "<p>The training dataset contains 99,992 fake videos. (The metadata says 100,000 but 8 of those files seem to be missing.)</p>\n\n<p>The training dataset also contains 19154 real videos.</p>\n\n<p>This means many of the real videos have been used to create more than one fake video (5 on average). In fact, some real videos have been used to create up to 40 fake videos.</p>\n\n<p>I compared the file sizes of the fake videos that belong to any given real video, and in 60% of the case the fake videos have a larger file size than the real ones. (In 20% of the cases, all the fake videos were smaller than the original, and in the remaining 20% they sometimes were smaller, sometimes larger.) </p>\n\n<p>So <em>if</em> (and it's a big if) you could determine for any set of videos that they are all the fake videos that are derived from a given real video -- but you don't know which is which -- then you can classify the smallest file as real and the rest as fake and you'd be 60% accurate. </p>\n\n<p>(Of course, this assumes that the test set also contains the real videos that were used to create the fake videos, which may not be the case. So it's probably not a very useful statistic to know but there you have it in any case.)</p>",
      "rawMarkdown": "The training dataset contains 99,992 fake videos. (The metadata says 100,000 but 8 of those files seem to be missing.)\n\nThe training dataset also contains 19154 real videos.\n\nThis means many of the real videos have been used to create more than one fake video (5 on average). In fact, some real videos have been used to create up to 40 fake videos.\n\nI compared the file sizes of the fake videos that belong to any given real video, and in 60% of the case the fake videos have a larger file size than the real ones. (In 20% of the cases, all the fake videos were smaller than the original, and in the remaining 20% they sometimes were smaller, sometimes larger.) \n\nSo _if_ (and it's a big if) you could determine for any set of videos that they are all the fake videos that are derived from a given real video -- but you don't know which is which -- then you can classify the smallest file as real and the rest as fake and you'd be 60% accurate. \n\n(Of course, this assumes that the test set also contains the real videos that were used to create the fake videos, which may not be the case. So it's probably not a very useful statistic to know but there you have it in any case.)",
      "votes": null
    },
    {
      "id": "694652",
      "postDate": "12/13/2019 22:47:04",
      "content": "<p>I've determined in another thread that public test set has 50-50 fake and real videos, which would fortunately make it harder to apply that technique.</p>",
      "rawMarkdown": "I've determined in another thread that public test set has 50-50 fake and real videos, which would fortunately make it harder to apply that technique.",
      "votes": null
    },
    {
      "id": "694656",
      "postDate": "12/13/2019 23:06:17",
      "content": "<p>How is the variance of +- 20% technically possible? I'm wondering since all videos do have the same duration and pixels. Can the bitrate vary locally on an image?</p>\n\n<p>I'm only wondering about the technical explanation behind this behaviour. </p>",
      "rawMarkdown": "How is the variance of +- 20% technically possible? I'm wondering since all videos do have the same duration and pixels. Can the bitrate vary locally on an image?\n\nI'm only wondering about the technical explanation behind this behaviour.",
      "votes": null
    },
    {
      "id": "694657",
      "postDate": "12/13/2019 23:20:13",
      "content": "<p>It would still work if there's a guarantee that each (real, fake) pair of videos is present in that test set, but that's probably not the case.</p>",
      "rawMarkdown": "It would still work if there's a guarantee that each (real, fake) pair of videos is present in that test set, but that's probably not the case.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 694652,
      "author_name": "sdoria",
      "author_url": "",
      "post_date": "12/13/2019 22:47:04",
      "content": "<p>I've determined in another thread that public test set has 50-50 fake and real videos, which would fortunately make it harder to apply that technique.</p>",
      "votes": null,
      "replies": [
        {
          "id": 694657,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/13/2019 23:20:13",
          "content": "<p>It would still work if there's a guarantee that each (real, fake) pair of videos is present in that test set, but that's probably not the case.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 694656,
      "author_name": "obrinkmann",
      "author_url": "",
      "post_date": "12/13/2019 23:06:17",
      "content": "<p>How is the variance of +- 20% technically possible? I'm wondering since all videos do have the same duration and pixels. Can the bitrate vary locally on an image?</p>\n\n<p>I'm only wondering about the technical explanation behind this behaviour. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "694616": "The training dataset contains 99,992 fake videos. (The metadata says 100,000 but 8 of those files seem to be missing.)\n\nThe training dataset also contains 19154 real videos.\n\nThis means many of the real videos have been used to create more than one fake video (5 on average). In fact, some real videos have been used to create up to 40 fake videos.\n\nI compared the file sizes of the fake videos that belong to any given real video, and in 60% of the case the fake videos have a larger file size than the real ones. (In 20% of the cases, all the fake videos were smaller than the original, and in the remaining 20% they sometimes were smaller, sometimes larger.) \n\nSo _if_ (and it's a big if) you could determine for any set of videos that they are all the fake videos that are derived from a given real video -- but you don't know which is which -- then you can classify the smallest file as real and the rest as fake and you'd be 60% accurate. \n\n(Of course, this assumes that the test set also contains the real videos that were used to create the fake videos, which may not be the case. So it's probably not a very useful statistic to know but there you have it in any case.)",
    "694652": "I've determined in another thread that public test set has 50-50 fake and real videos, which would fortunately make it harder to apply that technique.",
    "694656": "How is the variance of +- 20% technically possible? I'm wondering since all videos do have the same duration and pixels. Can the bitrate vary locally on an image?\n\nI'm only wondering about the technical explanation behind this behaviour.",
    "694657": "It would still work if there's a guarantee that each (real, fake) pair of videos is present in that test set, but that's probably not the case."
  },
  "source": "meta"
}