{
  "id": 147265,
  "title": "Manually Labelled Actor Database",
  "url": "/competitions/deepfake-detection-challenge/discussion/147265",
  "author_name": "Ces Bertino",
  "post_date": "2020-04-30T04:12:04.183000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>In case anyone is planning on doing any more research on this dataset, I would like to share a manually labelled actor database I built for the real videos.\nEach actor (386 of them) has a unique ID consisting of a pair of numbers. The first is the folder number from training set, and the second is a relative index within that folder of the actor.\nI only labelled the main actor of each video. There are some actors which appear in multiple folders. As far as I can tell, these actors that appear in multiple folders, only appear in multi-person  videos of the main actor. So, as far as I know, the actor IDs I labelled only appear in a single folder (in single person videos). This is one reason why I ignored multi-person videos during training. The other reason why I ignored multi-person videos during training was because I never found a single multi-person video where more than one person appeared to be fake in it.</p>\n\n<p>The notebook is in <a href=\"https://www.kaggle.com/cesb45/deepfake-detection-actors\">https://www.kaggle.com/cesb45/deepfake-detection-actors</a></p>\n\n<p>It contains the dataset, a sample photo of each actor and at the end of the notebook I also have 6 samples of face swaps which if I found the source actor correctly, show how mixed the source and target actors are in the training dataset, such that it seems impossible to find a good validation set of folders whose actors (source and target) are not in the training set. In my opinion, this was a significant weakness of the dataset.</p>",
  "messages": [
    {
      "id": 827079,
      "postDate": "2020-04-30T04:12:04.183Z",
      "content": "<p>In case anyone is planning on doing any more research on this dataset, I would like to share a manually labelled actor database I built for the real videos.\nEach actor (386 of them) has a unique ID consisting of a pair of numbers. The first is the folder number from training set, and the second is a relative index within that folder of the actor.\nI only labelled the main actor of each video. There are some actors which appear in multiple folders. As far as I can tell, these actors that appear in multiple folders, only appear in multi-person  videos of the main actor. So, as far as I know, the actor IDs I labelled only appear in a single folder (in single person videos). This is one reason why I ignored multi-person videos during training. The other reason why I ignored multi-person videos during training was because I never found a single multi-person video where more than one person appeared to be fake in it.</p>\n\n<p>The notebook is in <a href=\"https://www.kaggle.com/cesb45/deepfake-detection-actors\">https://www.kaggle.com/cesb45/deepfake-detection-actors</a></p>\n\n<p>It contains the dataset, a sample photo of each actor and at the end of the notebook I also have 6 samples of face swaps which if I found the source actor correctly, show how mixed the source and target actors are in the training dataset, such that it seems impossible to find a good validation set of folders whose actors (source and target) are not in the training set. In my opinion, this was a significant weakness of the dataset.</p>",
      "rawMarkdown": "In case anyone is planning on doing any more research on this dataset, I would like to share a manually labelled actor database I built for the real videos.\nEach actor (386 of them) has a unique ID consisting of a pair of numbers. The first is the folder number from training set, and the second is a relative index within that folder of the actor.\nI only labelled the main actor of each video. There are some actors which appear in multiple folders. As far as I can tell, these actors that appear in multiple folders, only appear in multi-person  videos of the main actor. So, as far as I know, the actor IDs I labelled only appear in a single folder (in single person videos). This is one reason why I ignored multi-person videos during training. The other reason why I ignored multi-person videos during training was because I never found a single multi-person video where more than one person appeared to be fake in it.\n\nThe notebook is in https://www.kaggle.com/cesb45/deepfake-detection-actors\n\nIt contains the dataset, a sample photo of each actor and at the end of the notebook I also have 6 samples of face swaps which if I found the source actor correctly, show how mixed the source and target actors are in the training dataset, such that it seems impossible to find a good validation set of folders whose actors (source and target) are not in the training set. In my opinion, this was a significant weakness of the dataset.",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "827079": "In case anyone is planning on doing any more research on this dataset, I would like to share a manually labelled actor database I built for the real videos.\nEach actor (386 of them) has a unique ID consisting of a pair of numbers. The first is the folder number from training set, and the second is a relative index within that folder of the actor.\nI only labelled the main actor of each video. There are some actors which appear in multiple folders. As far as I can tell, these actors that appear in multiple folders, only appear in multi-person  videos of the main actor. So, as far as I know, the actor IDs I labelled only appear in a single folder (in single person videos). This is one reason why I ignored multi-person videos during training. The other reason why I ignored multi-person videos during training was because I never found a single multi-person video where more than one person appeared to be fake in it.\n\nThe notebook is in https://www.kaggle.com/cesb45/deepfake-detection-actors\n\nIt contains the dataset, a sample photo of each actor and at the end of the notebook I also have 6 samples of face swaps which if I found the source actor correctly, show how mixed the source and target actors are in the training dataset, such that it seems impossible to find a good validation set of folders whose actors (source and target) are not in the training set. In my opinion, this was a significant weakness of the dataset."
  }
}