{
  "id": 145749,
  "title": "Is deepfake detection a lost cause?",
  "url": "/competitions/deepfake-detection-challenge/discussion/145749",
  "author_name": "Andrés Miguel Torrubia Sáez",
  "post_date": "2020-04-24T11:30:18.665000",
  "votes": 25,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Naïve translation of logloss of top solutions 0.42 to accuracy corresponds to ~66% accuracy. Two observations:</p>\n\n<p>1) This figure(s) are in stark contrast to most academic papers that report accuracies of 90% and above. Clearly academic studies need a more realistic validation set. </p>\n\n<p>2) I wonder if after $1m spent and top solutions are 0.42 logloss one should consider current deepfake detection as a lost cause to deploy commercially (likely top solutions would tank in hard deep fakes, which are the troublesome ones).</p>",
  "messages": [
    {
      "id": 2195509,
      "postDate": "2023-03-24T17:34:25.780Z",
      "content": "<p>Any thoughts on this contest 3 years later with deep fakes running amok? Would be curious to hear thoughts</p>",
      "rawMarkdown": "Any thoughts on this contest 3 years later with deep fakes running amok? Would be curious to hear thoughts",
      "votes": 1
    },
    {
      "id": 819132,
      "postDate": "2020-04-24T11:30:18.667Z",
      "content": "<p>Naïve translation of logloss of top solutions 0.42 to accuracy corresponds to ~66% accuracy. Two observations:</p>\n\n<p>1) This figure(s) are in stark contrast to most academic papers that report accuracies of 90% and above. Clearly academic studies need a more realistic validation set. </p>\n\n<p>2) I wonder if after $1m spent and top solutions are 0.42 logloss one should consider current deepfake detection as a lost cause to deploy commercially (likely top solutions would tank in hard deep fakes, which are the troublesome ones).</p>",
      "rawMarkdown": "Naïve translation of logloss of top solutions 0.42 to accuracy corresponds to ~66% accuracy. Two observations:\n\n1) This figure(s) are in stark contrast to most academic papers that report accuracies of 90% and above. Clearly academic studies need a more realistic validation set. \n\n2) I wonder if after $1m spent and top solutions are 0.42 logloss one should consider current deepfake detection as a lost cause to deploy commercially (likely top solutions would tank in hard deep fakes, which are the troublesome ones).",
      "votes": 25
    },
    {
      "id": 819154,
      "postDate": "2020-04-24T11:46:46.810Z",
      "content": "<ol>\n<li>training dataset has really low quality deepfakes - generalization from that set is really a big problem</li>\n<li>private set has very high quality deepfakes - I'm pretty sure that good models will go below 0.35 if they add some part of organic data to  the training set.</li>\n<li>some competitors may have used youtube organic data as I predicted in February, in that case that's totally a waste of money for hosts as the leaderboard doesn't show a fair comparison of different methods. </li>\n</ol>",
      "rawMarkdown": "1. training dataset has really low quality deepfakes - generalization from that set is really a big problem\n2. private set has very high quality deepfakes - I'm pretty sure that good models will go below 0.35 if they add some part of organic data to  the training set.\n3. some competitors may have used youtube organic data as I predicted in February, in that case that's totally a waste of money for hosts as the leaderboard doesn't show a fair comparison of different methods. ",
      "votes": 12
    },
    {
      "id": 819349,
      "postDate": "2020-04-24T14:30:24.933Z",
      "content": "<p>Agree with many of the points here, but we played with what we are given.</p>\n\n<p>As you guys point out, knowing the logloss score without knowing at least the underlying distribution of the hidden dataset, makes it difficult to guess how would our score translated into more business meaningful metrics precision/recall or F1.</p>\n\n<p>If I am the organiser, and if I am serious to tackle ongoing challenge of deepfake, then I would organise some exchange session to invite top teams from both public and private LB to have a full review on what worked and what hasn't worked.  </p>",
      "rawMarkdown": "Agree with many of the points here, but we played with what we are given.\n\nAs you guys point out, knowing the logloss score without knowing at least the underlying distribution of the hidden dataset, makes it difficult to guess how would our score translated into more business meaningful metrics precision/recall or F1.\n\nIf I am the organiser, and if I am serious to tackle ongoing challenge of deepfake, then I would organise some exchange session to invite top teams from both public and private LB to have a full review on what worked and what hasn't worked.  \n\n",
      "votes": 6,
      "replies": [
        {
          "id": 821811,
          "postDate": "2020-04-26T12:45:22.663Z",
          "content": "<p>agreed!</p>",
          "rawMarkdown": "agreed!",
          "votes": 2
        }
      ]
    },
    {
      "id": 819141,
      "postDate": "2020-04-24T11:37:20.433Z",
      "content": "<p>I think it's very difficult to know how to translate that logloss into a real accuracy due to the exponential punishment nature at extremes. Would be very interesting to find this data from the organisers, though.</p>\n\n<p>If not, did any of the high scorers use very strict clamping? If we find someone who clamped at 0.9 / 0.1, for example, where the punishment curve is relatively flat, we might be able to get a reasonable estimate of accuracy...</p>\n\n<p>However, I do agree with you, and I hope the organisers aren't disappointed.</p>",
      "rawMarkdown": "I think it's very difficult to know how to translate that logloss into a real accuracy due to the exponential punishment nature at extremes. Would be very interesting to find this data from the organisers, though.\n\nIf not, did any of the high scorers use very strict clamping? If we find someone who clamped at 0.9 / 0.1, for example, where the punishment curve is relatively flat, we might be able to get a reasonable estimate of accuracy...\n\nHowever, I do agree with you, and I hope the organisers aren't disappointed.",
      "votes": 6,
      "replies": [
        {
          "id": 819280,
          "postDate": "2020-04-24T13:26:20.267Z",
          "content": "<p>I have used the way that 0.9-fake  and 0.1-real  to caculate the accuracy in public LB. My acc in public LB(0.42) is about 85%.  </p>",
          "rawMarkdown": "I have used the way that 0.9-fake  and 0.1-real  to caculate the accuracy in public LB. My acc in public LB(0.42) is about 85%.  ",
          "votes": 1
        },
        {
          "id": 819320,
          "postDate": "2020-04-24T14:06:21.657Z",
          "content": "<p>Agreed, it's 85% assuming that clamp.\nWe clamped at 0.001 and 0.999. If we assume the network always predicted with these (it didn't, but it illustrates my point), then our loss of 0.437 corresponds with an accuracy of over 93%.</p>",
          "rawMarkdown": "Agreed, it's 85% assuming that clamp.\nWe clamped at 0.001 and 0.999. If we assume the network always predicted with these (it didn't, but it illustrates my point), then our loss of 0.437 corresponds with an accuracy of over 93%.",
          "votes": 2
        },
        {
          "id": 819642,
          "postDate": "2020-04-24T18:36:34.957Z",
          "content": "<p>We clamped at 0.95/0.05 and got the rank that we got if that helps. It roughly corresponds to ~60% accuracy.</p>",
          "rawMarkdown": "We clamped at 0.95/0.05 and got the rank that we got if that helps. It roughly corresponds to ~60% accuracy."
        }
      ]
    },
    {
      "id": 819139,
      "postDate": "2020-04-24T11:36:50.783Z",
      "content": "<p>useless dataset in reality and dont allow external dataset\nThis was a farce competition</p>",
      "rawMarkdown": "useless dataset in reality and dont allow external dataset\nThis was a farce competition\n",
      "votes": 3
    },
    {
      "id": 819917,
      "postDate": "2020-04-25T02:18:40.590Z",
      "content": "<p>Even a small false-positive rate will be overwhelming to human censors.</p>\n\n<p>Moreover, I suspect that people who look \"strange\" (birthmarks, unlucky tats, vitiligo) are liable to trigger the automatic deepfake censors all the time, leading to bad publicity.</p>\n\n<p><img src=\"https://www.favoriteplus.com/blog/wp-content/uploads/2019/06/Screen-Shot-2018-05-10-at-10.13.42-AM-300x255.png\" alt=\"\"></p>",
      "rawMarkdown": "Even a small false-positive rate will be overwhelming to human censors.\n\nMoreover, I suspect that people who look \"strange\" (birthmarks, unlucky tats, vitiligo) are liable to trigger the automatic deepfake censors all the time, leading to bad publicity.\n\n![](https://www.favoriteplus.com/blog/wp-content/uploads/2019/06/Screen-Shot-2018-05-10-at-10.13.42-AM-300x255.png)\n",
      "votes": 3
    },
    {
      "id": 819163,
      "postDate": "2020-04-24T11:56:56.483Z",
      "content": "<p>1) Or perhaps DFDC should have offered rather a better solution?\n2) Although a decent score, I personally don't think FB will deploy a 0.42 solution without further extensive tweaks</p>",
      "rawMarkdown": "1) Or perhaps DFDC should have offered rather a better solution?\n2) Although a decent score, I personally don't think FB will deploy a 0.42 solution without further extensive tweaks",
      "votes": 1
    },
    {
      "id": 819250,
      "postDate": "2020-04-24T13:09:13.607Z",
      "content": "<p>The world has changed folks!! for now detecting deep fakes seems to be a lost cause. AI is the only hope to detecting fakes and if the accuracy of the detection models is not extremely high, the only option is Judge, Jury and the expert witness that states \" based on my experience and analysis and obscure set of facts! that video is a FAKE.\"</p>",
      "rawMarkdown": "The world has changed folks!! for now detecting deep fakes seems to be a lost cause. AI is the only hope to detecting fakes and if the accuracy of the detection models is not extremely high, the only option is Judge, Jury and the expert witness that states \" based on my experience and analysis and obscure set of facts! that video is a FAKE.\""
    },
    {
      "id": 820240,
      "postDate": "2020-04-25T09:11:51.220Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2195509,
      "author_name": "Stephen Keller",
      "author_url": "",
      "post_date": "2023-03-24T17:34:25.780000",
      "content": "<p>Any thoughts on this contest 3 years later with deep fakes running amok? Would be curious to hear thoughts</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 819154,
      "author_name": "Selim Seferbekov",
      "author_url": "",
      "post_date": "2020-04-24T11:46:46.810000",
      "content": "<ol>\n<li>training dataset has really low quality deepfakes - generalization from that set is really a big problem</li>\n<li>private set has very high quality deepfakes - I'm pretty sure that good models will go below 0.35 if they add some part of organic data to  the training set.</li>\n<li>some competitors may have used youtube organic data as I predicted in February, in that case that's totally a waste of money for hosts as the leaderboard doesn't show a fair comparison of different methods. </li>\n</ol>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 819349,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-04-24T14:30:24.933000",
      "content": "<p>Agree with many of the points here, but we played with what we are given.</p>\n\n<p>As you guys point out, knowing the logloss score without knowing at least the underlying distribution of the hidden dataset, makes it difficult to guess how would our score translated into more business meaningful metrics precision/recall or F1.</p>\n\n<p>If I am the organiser, and if I am serious to tackle ongoing challenge of deepfake, then I would organise some exchange session to invite top teams from both public and private LB to have a full review on what worked and what hasn't worked.  </p>",
      "votes": 6,
      "replies": [
        {
          "id": 821811,
          "author_name": "WestLake",
          "author_url": "",
          "post_date": "2020-04-26T12:45:22.663000",
          "content": "<p>agreed!</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 819141,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-04-24T11:37:20.433000",
      "content": "<p>I think it's very difficult to know how to translate that logloss into a real accuracy due to the exponential punishment nature at extremes. Would be very interesting to find this data from the organisers, though.</p>\n\n<p>If not, did any of the high scorers use very strict clamping? If we find someone who clamped at 0.9 / 0.1, for example, where the punishment curve is relatively flat, we might be able to get a reasonable estimate of accuracy...</p>\n\n<p>However, I do agree with you, and I hope the organisers aren't disappointed.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 819280,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-04-24T13:26:20.267000",
          "content": "<p>I have used the way that 0.9-fake  and 0.1-real  to caculate the accuracy in public LB. My acc in public LB(0.42) is about 85%.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 819320,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-04-24T14:06:21.657000",
          "content": "<p>Agreed, it's 85% assuming that clamp.\nWe clamped at 0.001 and 0.999. If we assume the network always predicted with these (it didn't, but it illustrates my point), then our loss of 0.437 corresponds with an accuracy of over 93%.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 819642,
          "author_name": "BayesianKitten",
          "author_url": "",
          "post_date": "2020-04-24T18:36:34.957000",
          "content": "<p>We clamped at 0.95/0.05 and got the rank that we got if that helps. It roughly corresponds to ~60% accuracy.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 819139,
      "author_name": "あほ",
      "author_url": "",
      "post_date": "2020-04-24T11:36:50.783000",
      "content": "<p>useless dataset in reality and dont allow external dataset\nThis was a farce competition</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 819917,
      "author_name": "Oleg Trott",
      "author_url": "",
      "post_date": "2020-04-25T02:18:40.590000",
      "content": "<p>Even a small false-positive rate will be overwhelming to human censors.</p>\n\n<p>Moreover, I suspect that people who look \"strange\" (birthmarks, unlucky tats, vitiligo) are liable to trigger the automatic deepfake censors all the time, leading to bad publicity.</p>\n\n<p><img src=\"https://www.favoriteplus.com/blog/wp-content/uploads/2019/06/Screen-Shot-2018-05-10-at-10.13.42-AM-300x255.png\" alt=\"\"></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 819163,
      "author_name": "WiseLearner",
      "author_url": "",
      "post_date": "2020-04-24T11:56:56.483000",
      "content": "<p>1) Or perhaps DFDC should have offered rather a better solution?\n2) Although a decent score, I personally don't think FB will deploy a 0.42 solution without further extensive tweaks</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 819250,
      "author_name": "DavidGbodiOdaibo",
      "author_url": "",
      "post_date": "2020-04-24T13:09:13.607000",
      "content": "<p>The world has changed folks!! for now detecting deep fakes seems to be a lost cause. AI is the only hope to detecting fakes and if the accuracy of the detection models is not extremely high, the only option is Judge, Jury and the expert witness that states \" based on my experience and analysis and obscure set of facts! that video is a FAKE.\"</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820240,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T09:11:51.220000",
      "content": "",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2195509": "Any thoughts on this contest 3 years later with deep fakes running amok? Would be curious to hear thoughts",
    "819132": "Naïve translation of logloss of top solutions 0.42 to accuracy corresponds to ~66% accuracy. Two observations:\n\n1) This figure(s) are in stark contrast to most academic papers that report accuracies of 90% and above. Clearly academic studies need a more realistic validation set. \n\n2) I wonder if after $1m spent and top solutions are 0.42 logloss one should consider current deepfake detection as a lost cause to deploy commercially (likely top solutions would tank in hard deep fakes, which are the troublesome ones).",
    "819154": "1. training dataset has really low quality deepfakes - generalization from that set is really a big problem\n2. private set has very high quality deepfakes - I'm pretty sure that good models will go below 0.35 if they add some part of organic data to  the training set.\n3. some competitors may have used youtube organic data as I predicted in February, in that case that's totally a waste of money for hosts as the leaderboard doesn't show a fair comparison of different methods. ",
    "819349": "Agree with many of the points here, but we played with what we are given.\n\nAs you guys point out, knowing the logloss score without knowing at least the underlying distribution of the hidden dataset, makes it difficult to guess how would our score translated into more business meaningful metrics precision/recall or F1.\n\nIf I am the organiser, and if I am serious to tackle ongoing challenge of deepfake, then I would organise some exchange session to invite top teams from both public and private LB to have a full review on what worked and what hasn't worked.  \n\n",
    "819141": "I think it's very difficult to know how to translate that logloss into a real accuracy due to the exponential punishment nature at extremes. Would be very interesting to find this data from the organisers, though.\n\nIf not, did any of the high scorers use very strict clamping? If we find someone who clamped at 0.9 / 0.1, for example, where the punishment curve is relatively flat, we might be able to get a reasonable estimate of accuracy...\n\nHowever, I do agree with you, and I hope the organisers aren't disappointed.",
    "819139": "useless dataset in reality and dont allow external dataset\nThis was a farce competition\n",
    "819917": "Even a small false-positive rate will be overwhelming to human censors.\n\nMoreover, I suspect that people who look \"strange\" (birthmarks, unlucky tats, vitiligo) are liable to trigger the automatic deepfake censors all the time, leading to bad publicity.\n\n![](https://www.favoriteplus.com/blog/wp-content/uploads/2019/06/Screen-Shot-2018-05-10-at-10.13.42-AM-300x255.png)\n",
    "819163": "1) Or perhaps DFDC should have offered rather a better solution?\n2) Although a decent score, I personally don't think FB will deploy a 0.42 solution without further extensive tweaks",
    "819250": "The world has changed folks!! for now detecting deep fakes seems to be a lost cause. AI is the only hope to detecting fakes and if the accuracy of the detection models is not extremely high, the only option is Judge, Jury and the expert witness that states \" based on my experience and analysis and obscure set of facts! that video is a FAKE.\"",
    "820240": ""
  }
}