{
  "id": 142223,
  "title": "Evaluation on synthetic videos",
  "url": "/competitions/deepfake-detection-challenge/discussion/142223",
  "author_name": "Richard_U",
  "post_date": "2020-04-09T14:04:59.060000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>@cristiancanton @juliaelliott </p>\n\n<p>It has after the end of the competition been announced that the final evaluation on Private Test Set will be calculated using synthetic videos <strong>and</strong> real world organic videos, <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417\">Next steps in the DFDC competition</a>.</p>\n\n<p>However, during the competition on the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started page</a>, and in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421\">Reminder from the organizers</a> close to the end of competition, the impression was given that final evaluation would be made solely on organic videos. They both said (in original and as quote) the following regarding the Private Test Set:</p>\n\n<blockquote>\n  <p>\"It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.\"</p>\n</blockquote>\n\n<p>Since classifying synthetic videos, generated with the same methods as the provided extensive Training Set, obviously is a much easier task than classification of real world deepfakes, the decision to include synthetic videos may result in the following consequences:</p>\n\n<ol>\n<li><p>Submissions over-adjusted towards properties of the training set are likely to be among top positions on the final Private Leaderboard. The simple reason is that few, if any, submissions could be expected to have an organic score in level with their synthetic score (given of course that the competition rules were followed). My guess is that many legit submissions will generate a score not too far below 0.69 if evaluated only on organic videos. Therefore, if synthetic videos are included, differences in score on the synthetic part of evaluation could be expected to be the main source of overall score difference on the Private Leaderboard.</p></li>\n<li><p>The external validity of the results from the competition is threatened if final top scoring Private Leaderboard submissions would generalize poorly to the overall objective of identifying real world organic deepfakes.</p></li>\n</ol>\n\n<p>Maybe I should declare that point number 1 is really not a problem for me personally as I do not expect a high scoring position regardless evaluation. But I do pity teams that may have put a bigger effort than I did in the competition and who choose to follow the instructions and encouragements to not over-adjust towards the Training Set and Public Leaderboard.</p>\n\n<p>However, the biggest concern from my perspective is point number 2. What made this competition, and its anticipated results, interesting to me – and I am sure many others – was an assumed evaluation solely based on real organic audio- and video-deepfakes. Even if this, despite the substantial price money, would have resulted in a possible close to null result it would definitely have been a valuable result. This as it would emphasize the difficulties of real world deepfake identification (and hence the urgency of improved methods). I do however believe that the competition may have generated some really good submissions in respect to real world deepfake identification. Submissions that may not get the acknowledgement they deserve if they not also happen to perform very well on synthetic evaluation.</p>\n\n<p>The organizers decision to include synthetic videos could become especially problematic if the constitution of synthetic videos would directly resemble the Training Set. This as many of the synthetic deepfake videos in the Training Set clearly have little, if any, resemblance to real world deepfakes. I give a concrete example in this post <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278</a>, but there are a lot of videos in the Training Set that really does not resemble anything that even remotely could be found in the real world of deepfakes. Roughly this could be the case for up to as many as 50% of the “fake” videos in the DFDC Training Set.</p>\n\n<p>In summary, to ensure external validity of the competition I encourage the organizers to select a subset of synthetic videos (if at all included for evaluation) in a way that excludes synthetic videos with obvious little to no resemblance to real world deepfakes. Of course the methods of evaluation is a decision solely for the organizers. Also, an exclusion of clearly non representative synthetic videos may very well already be the case. But since we do not know I definately think this is a subject worth rising.</p>\n\n<p>Best regards and happy Easter to all of you!</p>",
  "messages": [
    {
      "id": 802435,
      "postDate": "2020-04-09T14:04:59.060Z",
      "content": "<p>@cristiancanton @juliaelliott </p>\n\n<p>It has after the end of the competition been announced that the final evaluation on Private Test Set will be calculated using synthetic videos <strong>and</strong> real world organic videos, <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417\">Next steps in the DFDC competition</a>.</p>\n\n<p>However, during the competition on the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started page</a>, and in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421\">Reminder from the organizers</a> close to the end of competition, the impression was given that final evaluation would be made solely on organic videos. They both said (in original and as quote) the following regarding the Private Test Set:</p>\n\n<blockquote>\n  <p>\"It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.\"</p>\n</blockquote>\n\n<p>Since classifying synthetic videos, generated with the same methods as the provided extensive Training Set, obviously is a much easier task than classification of real world deepfakes, the decision to include synthetic videos may result in the following consequences:</p>\n\n<ol>\n<li><p>Submissions over-adjusted towards properties of the training set are likely to be among top positions on the final Private Leaderboard. The simple reason is that few, if any, submissions could be expected to have an organic score in level with their synthetic score (given of course that the competition rules were followed). My guess is that many legit submissions will generate a score not too far below 0.69 if evaluated only on organic videos. Therefore, if synthetic videos are included, differences in score on the synthetic part of evaluation could be expected to be the main source of overall score difference on the Private Leaderboard.</p></li>\n<li><p>The external validity of the results from the competition is threatened if final top scoring Private Leaderboard submissions would generalize poorly to the overall objective of identifying real world organic deepfakes.</p></li>\n</ol>\n\n<p>Maybe I should declare that point number 1 is really not a problem for me personally as I do not expect a high scoring position regardless evaluation. But I do pity teams that may have put a bigger effort than I did in the competition and who choose to follow the instructions and encouragements to not over-adjust towards the Training Set and Public Leaderboard.</p>\n\n<p>However, the biggest concern from my perspective is point number 2. What made this competition, and its anticipated results, interesting to me – and I am sure many others – was an assumed evaluation solely based on real organic audio- and video-deepfakes. Even if this, despite the substantial price money, would have resulted in a possible close to null result it would definitely have been a valuable result. This as it would emphasize the difficulties of real world deepfake identification (and hence the urgency of improved methods). I do however believe that the competition may have generated some really good submissions in respect to real world deepfake identification. Submissions that may not get the acknowledgement they deserve if they not also happen to perform very well on synthetic evaluation.</p>\n\n<p>The organizers decision to include synthetic videos could become especially problematic if the constitution of synthetic videos would directly resemble the Training Set. This as many of the synthetic deepfake videos in the Training Set clearly have little, if any, resemblance to real world deepfakes. I give a concrete example in this post <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278</a>, but there are a lot of videos in the Training Set that really does not resemble anything that even remotely could be found in the real world of deepfakes. Roughly this could be the case for up to as many as 50% of the “fake” videos in the DFDC Training Set.</p>\n\n<p>In summary, to ensure external validity of the competition I encourage the organizers to select a subset of synthetic videos (if at all included for evaluation) in a way that excludes synthetic videos with obvious little to no resemblance to real world deepfakes. Of course the methods of evaluation is a decision solely for the organizers. Also, an exclusion of clearly non representative synthetic videos may very well already be the case. But since we do not know I definately think this is a subject worth rising.</p>\n\n<p>Best regards and happy Easter to all of you!</p>",
      "rawMarkdown": "@cristiancanton @juliaelliott \n\nIt has after the end of the competition been announced that the final evaluation on Private Test Set will be calculated using synthetic videos **and** real world organic videos, [Next steps in the DFDC competition](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417).\n\nHowever, during the competition on the [Getting Started page](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started), and in the [Reminder from the organizers](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421) close to the end of competition, the impression was given that final evaluation would be made solely on organic videos. They both said (in original and as quote) the following regarding the Private Test Set:\n&gt; \"It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.\"\n\nSince classifying synthetic videos, generated with the same methods as the provided extensive Training Set, obviously is a much easier task than classification of real world deepfakes, the decision to include synthetic videos may result in the following consequences:\n\n1. Submissions over-adjusted towards properties of the training set are likely to be among top positions on the final Private Leaderboard. The simple reason is that few, if any, submissions could be expected to have an organic score in level with their synthetic score (given of course that the competition rules were followed). My guess is that many legit submissions will generate a score not too far below 0.69 if evaluated only on organic videos. Therefore, if synthetic videos are included, differences in score on the synthetic part of evaluation could be expected to be the main source of overall score difference on the Private Leaderboard.\n\n2. The external validity of the results from the competition is threatened if final top scoring Private Leaderboard submissions would generalize poorly to the overall objective of identifying real world organic deepfakes.\n\nMaybe I should declare that point number 1 is really not a problem for me personally as I do not expect a high scoring position regardless evaluation. But I do pity teams that may have put a bigger effort than I did in the competition and who choose to follow the instructions and encouragements to not over-adjust towards the Training Set and Public Leaderboard.\n\nHowever, the biggest concern from my perspective is point number 2. What made this competition, and its anticipated results, interesting to me – and I am sure many others – was an assumed evaluation solely based on real organic audio- and video-deepfakes. Even if this, despite the substantial price money, would have resulted in a possible close to null result it would definitely have been a valuable result. This as it would emphasize the difficulties of real world deepfake identification (and hence the urgency of improved methods). I do however believe that the competition may have generated some really good submissions in respect to real world deepfake identification. Submissions that may not get the acknowledgement they deserve if they not also happen to perform very well on synthetic evaluation.\n\nThe organizers decision to include synthetic videos could become especially problematic if the constitution of synthetic videos would directly resemble the Training Set. This as many of the synthetic deepfake videos in the Training Set clearly have little, if any, resemblance to real world deepfakes. I give a concrete example in this post [https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278), but there are a lot of videos in the Training Set that really does not resemble anything that even remotely could be found in the real world of deepfakes. Roughly this could be the case for up to as many as 50% of the “fake” videos in the DFDC Training Set.\n\nIn summary, to ensure external validity of the competition I encourage the organizers to select a subset of synthetic videos (if at all included for evaluation) in a way that excludes synthetic videos with obvious little to no resemblance to real world deepfakes. Of course the methods of evaluation is a decision solely for the organizers. Also, an exclusion of clearly non representative synthetic videos may very well already be the case. But since we do not know I definately think this is a subject worth rising.\n\nBest regards and happy Easter to all of you!",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "802435": "@cristiancanton @juliaelliott \n\nIt has after the end of the competition been announced that the final evaluation on Private Test Set will be calculated using synthetic videos **and** real world organic videos, [Next steps in the DFDC competition](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417).\n\nHowever, during the competition on the [Getting Started page](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started), and in the [Reminder from the organizers](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133421) close to the end of competition, the impression was given that final evaluation would be made solely on organic videos. They both said (in original and as quote) the following regarding the Private Test Set:\n&gt; \"It contains videos with a similar format and nature as the Training and Public Validation/Test Sets, but are real, organic videos with and without deepfakes.\"\n\nSince classifying synthetic videos, generated with the same methods as the provided extensive Training Set, obviously is a much easier task than classification of real world deepfakes, the decision to include synthetic videos may result in the following consequences:\n\n1. Submissions over-adjusted towards properties of the training set are likely to be among top positions on the final Private Leaderboard. The simple reason is that few, if any, submissions could be expected to have an organic score in level with their synthetic score (given of course that the competition rules were followed). My guess is that many legit submissions will generate a score not too far below 0.69 if evaluated only on organic videos. Therefore, if synthetic videos are included, differences in score on the synthetic part of evaluation could be expected to be the main source of overall score difference on the Private Leaderboard.\n\n2. The external validity of the results from the competition is threatened if final top scoring Private Leaderboard submissions would generalize poorly to the overall objective of identifying real world organic deepfakes.\n\nMaybe I should declare that point number 1 is really not a problem for me personally as I do not expect a high scoring position regardless evaluation. But I do pity teams that may have put a bigger effort than I did in the competition and who choose to follow the instructions and encouragements to not over-adjust towards the Training Set and Public Leaderboard.\n\nHowever, the biggest concern from my perspective is point number 2. What made this competition, and its anticipated results, interesting to me – and I am sure many others – was an assumed evaluation solely based on real organic audio- and video-deepfakes. Even if this, despite the substantial price money, would have resulted in a possible close to null result it would definitely have been a valuable result. This as it would emphasize the difficulties of real world deepfake identification (and hence the urgency of improved methods). I do however believe that the competition may have generated some really good submissions in respect to real world deepfake identification. Submissions that may not get the acknowledgement they deserve if they not also happen to perform very well on synthetic evaluation.\n\nThe organizers decision to include synthetic videos could become especially problematic if the constitution of synthetic videos would directly resemble the Training Set. This as many of the synthetic deepfake videos in the Training Set clearly have little, if any, resemblance to real world deepfakes. I give a concrete example in this post [https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/140417#800278), but there are a lot of videos in the Training Set that really does not resemble anything that even remotely could be found in the real world of deepfakes. Roughly this could be the case for up to as many as 50% of the “fake” videos in the DFDC Training Set.\n\nIn summary, to ensure external validity of the competition I encourage the organizers to select a subset of synthetic videos (if at all included for evaluation) in a way that excludes synthetic videos with obvious little to no resemblance to real world deepfakes. Of course the methods of evaluation is a decision solely for the organizers. Also, an exclusion of clearly non representative synthetic videos may very well already be the case. But since we do not know I definately think this is a subject worth rising.\n\nBest regards and happy Easter to all of you!"
  }
}