{
  "id": 121694,
  "title": "REAL/FAKE is not enough, we need to know if its video/audio/both",
  "url": "/competitions/deepfake-detection-challenge/discussion/121694",
  "author_name": "",
  "post_date": "2019-12-14T21:30:20.385244400Z",
  "votes": 14,
  "comment_count": 13,
  "views": 0,
  "content": "<p>@juliaelliott , @addisonhoward the description is \"detect video OR voice manipulation\". However, the metadata only includes \"REAL/FAKE\" with no indication if video/audio/both were manipulated. Would it be possible to extend the metadata to include this information? This is just making things hard for us. Training based on the assumption that FAKE means video manipulation will hurt the model performance if only audio was manipulated.</p>",
  "messages": [
    {
      "id": "695261",
      "postDate": "12/14/2019 21:30:20",
      "content": "<p>@juliaelliott , @addisonhoward the description is \"detect video OR voice manipulation\". However, the metadata only includes \"REAL/FAKE\" with no indication if video/audio/both were manipulated. Would it be possible to extend the metadata to include this information? This is just making things hard for us. Training based on the assumption that FAKE means video manipulation will hurt the model performance if only audio was manipulated.</p>",
      "rawMarkdown": "juliaelliott , @addisonhoward the description is \"detect video OR voice manipulation\". However, the metadata only includes \"REAL/FAKE\" with no indication if video/audio/both were manipulated. Would it be possible to extend the metadata to include this information? This is just making things hard for us. Training based on the assumption that FAKE means video manipulation will hurt the model performance if only audio was manipulated.",
      "votes": null
    },
    {
      "id": "695265",
      "postDate": "12/14/2019 21:41:56",
      "content": "<p>is this the case then?\nFAKE is when FAKE video and/or FAKE audio\nREAL is when REAL video and REAL audio</p>",
      "rawMarkdown": "is this the case then?\nFAKE is when FAKE video and/or FAKE audio\nREAL is when REAL video and REAL audio",
      "votes": null
    },
    {
      "id": "695291",
      "postDate": "12/14/2019 22:47:20",
      "content": "<p>the harder, the better! that's the point, isn't it? ;)</p>",
      "rawMarkdown": "the harder, the better! that's the point, isn't it? ;)",
      "votes": null
    },
    {
      "id": "695293",
      "postDate": "12/14/2019 22:52:24",
      "content": "<p>harder in training, easier in battle? not sure....</p>",
      "rawMarkdown": "harder in training, easier in battle? not sure....",
      "votes": null
    },
    {
      "id": "696172",
      "postDate": "12/16/2019 07:57:27",
      "content": "<p>In some of the REAL training the audio is NOT the person being filmed - its can be a female giving directions in some of the first 200 videos I looked at.</p>\n\n<p>So REAL apparently also can mean another NOT SEEN person, who happens to be real.</p>\n\n<p>Going to make for a real difficult time when the train data is crap.  Just finishing the ASHRAE were the train was also a large pile of something - looks like dirty training data is the future of Kaggle.</p>",
      "rawMarkdown": "In some of the REAL training the audio is NOT the person being filmed - its can be a female giving directions in some of the first 200 videos I looked at.\n\nSo REAL apparently also can mean another NOT SEEN person, who happens to be real.\n\nGoing to make for a real difficult time when the train data is crap.  Just finishing the ASHRAE were the train was also a large pile of something - looks like dirty training data is the future of Kaggle.",
      "votes": null
    },
    {
      "id": "698848",
      "postDate": "12/19/2019 19:04:47",
      "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Hm, then it seems to be a mistake in labelling... Could you give an example of such REAL item?</p>",
      "rawMarkdown": "pcjimmmy Hm, then it seems to be a mistake in labelling... Could you give an example of such REAL item?",
      "votes": null
    },
    {
      "id": "698943",
      "postDate": "12/19/2019 23:34:40",
      "content": "<p>Will take me a day or two - I need to rebuild my data file.  During the last 24 hours I have been building a Freenas server for this challenge, cleaning out all my ASHRAE files, upgrading my Ubuntu builds to 19.10, getting all my WIndows machines updated (Linux had been the OS of choice for the ASHRAE competition), and buying Christmas gifts for the kids.  One or more of those activities lead me to delete the file I built while watching that first 200.  Since the kids getting coal for Xmas that was not the root cause - but it was doing the Ubuntu - without much thought I did a fresh install rather than an upgrade on the machine I had watched the videos on.</p>\n\n<p>There were a couple of videos where off camera person was giving directions - and it seems like most were REAL.  It did lead me to the conclusion that I was NOT going to have to do lip reading to determine fake audio.  So I am thinking that the REAL designation was not a labeling mistake.  Also in that first 200 is a video were the audio is out of sync with the video - also a REAL.  Gave me further confidence that lip reading was not going to be a needed skill to program into a model.</p>\n\n<p>I did write down a decent number of design requirements while watching those 200 - so re-watching and adding a few more hundred to my Mark One Eyeball Model is on my early to do list.  </p>\n\n<p>Would suggest to all that binge watch a bunch of these videos is a good idea.</p>",
      "rawMarkdown": "Will take me a day or two - I need to rebuild my data file.  During the last 24 hours I have been building a Freenas server for this challenge, cleaning out all my ASHRAE files, upgrading my Ubuntu builds to 19.10, getting all my WIndows machines updated (Linux had been the OS of choice for the ASHRAE competition), and buying Christmas gifts for the kids.  One or more of those activities lead me to delete the file I built while watching that first 200.  Since the kids getting coal for Xmas that was not the root cause - but it was doing the Ubuntu - without much thought I did a fresh install rather than an upgrade on the machine I had watched the videos on.\n\nThere were a couple of videos where off camera person was giving directions - and it seems like most were REAL.  It did lead me to the conclusion that I was NOT going to have to do lip reading to determine fake audio.  So I am thinking that the REAL designation was not a labeling mistake.  Also in that first 200 is a video were the audio is out of sync with the video - also a REAL.  Gave me further confidence that lip reading was not going to be a needed skill to program into a model.\n\nI did write down a decent number of design requirements while watching those 200 - so re-watching and adding a few more hundred to my Mark One Eyeball Model is on my early to do list.  \n\nWould suggest to all that binge watch a bunch of these videos is a good idea.",
      "votes": null
    },
    {
      "id": "699148",
      "postDate": "12/20/2019 05:11:42",
      "content": "<p>I am not sure these are \"proofs\" of anything. For example, you might be required to detect the slightly out of sync audio as false if it is false... But I get your drift. </p>",
      "rawMarkdown": "I am not sure these are \"proofs\" of anything. For example, you might be required to detect the slightly out of sync audio as false if it is false... But I get your drift.",
      "votes": null
    },
    {
      "id": "708235",
      "postDate": "01/02/2020 05:19:54",
      "content": "<p>Yes - I am suffering from a lack of a nice definition of what is fake audio.  For example, if the audio is down sampled from 48000 to 16000 - does that make it fake?</p>",
      "rawMarkdown": "Yes - I am suffering from a lack of a nice definition of what is fake audio.  For example, if the audio is down sampled from 48000 to 16000 - does that make it fake?",
      "votes": null
    },
    {
      "id": "708946",
      "postDate": "01/02/2020 21:14:53",
      "content": "<p>Let me provide some comments here: \n- If audio and/or video are modified, you should predict the video as FAKE, otherwise as REAL.\n- Audio modifications are intended to change the voice of the actor (e.g. pitch), not alter what they were saying.\n- If audio has been downsampled, this is not considered a FAKE.\nHope this helps!</p>",
      "rawMarkdown": "Let me provide some comments here: \n- If audio and/or video are modified, you should predict the video as FAKE, otherwise as REAL.\n- Audio modifications are intended to change the voice of the actor (e.g. pitch), not alter what they were saying.\n- If audio has been downsampled, this is not considered a FAKE.\nHope this helps!",
      "votes": null
    },
    {
      "id": "709031",
      "postDate": "01/03/2020 00:41:49",
      "content": "<p>Thanks for the added clarification.  The method I used did result in considering that down sampling was a fake.  That process was strengthed by the fact that no 16 khz audio was real.</p>\n\n<p>I do strongly agree with Moshel - it seems a real shame that so much time was invested in creating the data set and than an apparent decision was made to cripple it, to me, almost to the point of making me wonder if this is just some intern project gone crazy.  Would love to see the logic that lead to that decision.</p>\n\n<p>I am referring to the fact that for FAKE we don't know if the audio or the video or both are FAKE.  </p>",
      "rawMarkdown": "Thanks for the added clarification.  The method I used did result in considering that down sampling was a fake.  That process was strengthed by the fact that no 16 khz audio was real.\n\nI do strongly agree with Moshel - it seems a real shame that so much time was invested in creating the data set and than an apparent decision was made to cripple it, to me, almost to the point of making me wonder if this is just some intern project gone crazy.  Would love to see the logic that lead to that decision.\n\nI am referring to the fact that for FAKE we don't know if the audio or the video or both are FAKE.",
      "votes": null
    },
    {
      "id": "709099",
      "postDate": "01/03/2020 03:29:04",
      "content": "<p>+1, knowing whether a fake is an audio fake or a video fake is important for model building.</p>",
      "rawMarkdown": "1, knowing whether a fake is an audio fake or a video fake is important for model building.",
      "votes": null
    },
    {
      "id": "711867",
      "postDate": "01/06/2020 15:40:14",
      "content": "<p>What I've got by going through comments is that the fake file could contain as swapped faces, so swapped voices, maybe one of them, or both of them.\nMy approach is to get some voice swap data from external resources, and make a classifier for their mel-spectograms.</p>\n\n<p>One of the resources: <a href=\"https://github.com/dessa-public/fake-voice-detection\">https://github.com/dessa-public/fake-voice-detection</a></p>\n\n<p>Will write Medium post later on :) </p>",
      "rawMarkdown": "What I've got by going through comments is that the fake file could contain as swapped faces, so swapped voices, maybe one of them, or both of them.\nMy approach is to get some voice swap data from external resources, and make a classifier for their mel-spectograms.\n\nOne of the resources: https://github.com/dessa-public/fake-voice-detection\n\nWill write Medium post later on :)",
      "votes": null
    },
    {
      "id": "735894",
      "postDate": "02/03/2020 15:16:34",
      "content": "<p>Do we have metadata for training set which can be used to identify the cases where only audio modified ?</p>",
      "rawMarkdown": "Do we have metadata for training set which can be used to identify the cases where only audio modified ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 695265,
      "author_name": "miwojc",
      "author_url": "",
      "post_date": "12/14/2019 21:41:56",
      "content": "<p>is this the case then?\nFAKE is when FAKE video and/or FAKE audio\nREAL is when REAL video and REAL audio</p>",
      "votes": null,
      "replies": [
        {
          "id": 696172,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/16/2019 07:57:27",
          "content": "<p>In some of the REAL training the audio is NOT the person being filmed - its can be a female giving directions in some of the first 200 videos I looked at.</p>\n\n<p>So REAL apparently also can mean another NOT SEEN person, who happens to be real.</p>\n\n<p>Going to make for a real difficult time when the train data is crap.  Just finishing the ASHRAE were the train was also a large pile of something - looks like dirty training data is the future of Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 698848,
          "author_name": "olafmat",
          "author_url": "",
          "post_date": "12/19/2019 19:04:47",
          "content": "<p><a href=\"/pcjimmmy\">@pcjimmmy</a> Hm, then it seems to be a mistake in labelling... Could you give an example of such REAL item?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 698943,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "12/19/2019 23:34:40",
          "content": "<p>Will take me a day or two - I need to rebuild my data file.  During the last 24 hours I have been building a Freenas server for this challenge, cleaning out all my ASHRAE files, upgrading my Ubuntu builds to 19.10, getting all my WIndows machines updated (Linux had been the OS of choice for the ASHRAE competition), and buying Christmas gifts for the kids.  One or more of those activities lead me to delete the file I built while watching that first 200.  Since the kids getting coal for Xmas that was not the root cause - but it was doing the Ubuntu - without much thought I did a fresh install rather than an upgrade on the machine I had watched the videos on.</p>\n\n<p>There were a couple of videos where off camera person was giving directions - and it seems like most were REAL.  It did lead me to the conclusion that I was NOT going to have to do lip reading to determine fake audio.  So I am thinking that the REAL designation was not a labeling mistake.  Also in that first 200 is a video were the audio is out of sync with the video - also a REAL.  Gave me further confidence that lip reading was not going to be a needed skill to program into a model.</p>\n\n<p>I did write down a decent number of design requirements while watching those 200 - so re-watching and adding a few more hundred to my Mark One Eyeball Model is on my early to do list.  </p>\n\n<p>Would suggest to all that binge watch a bunch of these videos is a good idea.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 699148,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/20/2019 05:11:42",
          "content": "<p>I am not sure these are \"proofs\" of anything. For example, you might be required to detect the slightly out of sync audio as false if it is false... But I get your drift. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708235,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/02/2020 05:19:54",
          "content": "<p>Yes - I am suffering from a lack of a nice definition of what is fake audio.  For example, if the audio is down sampled from 48000 to 16000 - does that make it fake?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 695291,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "12/14/2019 22:47:20",
      "content": "<p>the harder, the better! that's the point, isn't it? ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 695293,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/14/2019 22:52:24",
          "content": "<p>harder in training, easier in battle? not sure....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 708946,
      "author_name": "cristiancanton",
      "author_url": "",
      "post_date": "01/02/2020 21:14:53",
      "content": "<p>Let me provide some comments here: \n- If audio and/or video are modified, you should predict the video as FAKE, otherwise as REAL.\n- Audio modifications are intended to change the voice of the actor (e.g. pitch), not alter what they were saying.\n- If audio has been downsampled, this is not considered a FAKE.\nHope this helps!</p>",
      "votes": null,
      "replies": [
        {
          "id": 709031,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "01/03/2020 00:41:49",
          "content": "<p>Thanks for the added clarification.  The method I used did result in considering that down sampling was a fake.  That process was strengthed by the fact that no 16 khz audio was real.</p>\n\n<p>I do strongly agree with Moshel - it seems a real shame that so much time was invested in creating the data set and than an apparent decision was made to cripple it, to me, almost to the point of making me wonder if this is just some intern project gone crazy.  Would love to see the logic that lead to that decision.</p>\n\n<p>I am referring to the fact that for FAKE we don't know if the audio or the video or both are FAKE.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709099,
          "author_name": "ravivadapalli",
          "author_url": "",
          "post_date": "01/03/2020 03:29:04",
          "content": "<p>+1, knowing whether a fake is an audio fake or a video fake is important for model building.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 735894,
          "author_name": "reachkishore",
          "author_url": "",
          "post_date": "02/03/2020 15:16:34",
          "content": "<p>Do we have metadata for training set which can be used to identify the cases where only audio modified ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 711867,
      "author_name": "lchkhetiani",
      "author_url": "",
      "post_date": "01/06/2020 15:40:14",
      "content": "<p>What I've got by going through comments is that the fake file could contain as swapped faces, so swapped voices, maybe one of them, or both of them.\nMy approach is to get some voice swap data from external resources, and make a classifier for their mel-spectograms.</p>\n\n<p>One of the resources: <a href=\"https://github.com/dessa-public/fake-voice-detection\">https://github.com/dessa-public/fake-voice-detection</a></p>\n\n<p>Will write Medium post later on :) </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "695261": "juliaelliott , @addisonhoward the description is \"detect video OR voice manipulation\". However, the metadata only includes \"REAL/FAKE\" with no indication if video/audio/both were manipulated. Would it be possible to extend the metadata to include this information? This is just making things hard for us. Training based on the assumption that FAKE means video manipulation will hurt the model performance if only audio was manipulated.",
    "695265": "is this the case then?\nFAKE is when FAKE video and/or FAKE audio\nREAL is when REAL video and REAL audio",
    "695291": "the harder, the better! that's the point, isn't it? ;)",
    "695293": "harder in training, easier in battle? not sure....",
    "696172": "In some of the REAL training the audio is NOT the person being filmed - its can be a female giving directions in some of the first 200 videos I looked at.\n\nSo REAL apparently also can mean another NOT SEEN person, who happens to be real.\n\nGoing to make for a real difficult time when the train data is crap.  Just finishing the ASHRAE were the train was also a large pile of something - looks like dirty training data is the future of Kaggle.",
    "698848": "pcjimmmy Hm, then it seems to be a mistake in labelling... Could you give an example of such REAL item?",
    "698943": "Will take me a day or two - I need to rebuild my data file.  During the last 24 hours I have been building a Freenas server for this challenge, cleaning out all my ASHRAE files, upgrading my Ubuntu builds to 19.10, getting all my WIndows machines updated (Linux had been the OS of choice for the ASHRAE competition), and buying Christmas gifts for the kids.  One or more of those activities lead me to delete the file I built while watching that first 200.  Since the kids getting coal for Xmas that was not the root cause - but it was doing the Ubuntu - without much thought I did a fresh install rather than an upgrade on the machine I had watched the videos on.\n\nThere were a couple of videos where off camera person was giving directions - and it seems like most were REAL.  It did lead me to the conclusion that I was NOT going to have to do lip reading to determine fake audio.  So I am thinking that the REAL designation was not a labeling mistake.  Also in that first 200 is a video were the audio is out of sync with the video - also a REAL.  Gave me further confidence that lip reading was not going to be a needed skill to program into a model.\n\nI did write down a decent number of design requirements while watching those 200 - so re-watching and adding a few more hundred to my Mark One Eyeball Model is on my early to do list.  \n\nWould suggest to all that binge watch a bunch of these videos is a good idea.",
    "699148": "I am not sure these are \"proofs\" of anything. For example, you might be required to detect the slightly out of sync audio as false if it is false... But I get your drift.",
    "708235": "Yes - I am suffering from a lack of a nice definition of what is fake audio.  For example, if the audio is down sampled from 48000 to 16000 - does that make it fake?",
    "708946": "Let me provide some comments here: \n- If audio and/or video are modified, you should predict the video as FAKE, otherwise as REAL.\n- Audio modifications are intended to change the voice of the actor (e.g. pitch), not alter what they were saying.\n- If audio has been downsampled, this is not considered a FAKE.\nHope this helps!",
    "709031": "Thanks for the added clarification.  The method I used did result in considering that down sampling was a fake.  That process was strengthed by the fact that no 16 khz audio was real.\n\nI do strongly agree with Moshel - it seems a real shame that so much time was invested in creating the data set and than an apparent decision was made to cripple it, to me, almost to the point of making me wonder if this is just some intern project gone crazy.  Would love to see the logic that lead to that decision.\n\nI am referring to the fact that for FAKE we don't know if the audio or the video or both are FAKE.",
    "709099": "1, knowing whether a fake is an audio fake or a video fake is important for model building.",
    "711867": "What I've got by going through comments is that the fake file could contain as swapped faces, so swapped voices, maybe one of them, or both of them.\nMy approach is to get some voice swap data from external resources, and make a classifier for their mel-spectograms.\n\nOne of the resources: https://github.com/dessa-public/fake-voice-detection\n\nWill write Medium post later on :)",
    "735894": "Do we have metadata for training set which can be used to identify the cases where only audio modified ?"
  },
  "source": "meta"
}