{
  "id": 136869,
  "title": "Anyone's tried Bojan's dataset?",
  "url": "/competitions/deepfake-detection-challenge/discussion/136869",
  "author_name": "Trigram",
  "post_date": "2020-03-18T04:45:48.909000",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Has anyone tried @tunguz 's dataset of a million fake faces? I'm currently trying to make it work along with the celebrity faces dataset.</p>",
  "messages": [
    {
      "id": 778409,
      "postDate": "2020-03-18T12:17:37.447Z",
      "content": "<p>I would say: very unlikely. Note that the methods used for generating StyleGAN fake faces and deepfake creation have significant differences that make generalization difficult (eg. StyleGAN is a mapping from the latent code, deepfake techniques map from the distribution of real faces,  StyleGAN generates entire faces and backgrounds , while deepfakes only modify the face with pre-processing steps such as face alignment and 3d modelling etc. ) Furthermore, the datasets have widely differing characteristics as well; whereas StyleGAN faces are always forward-facing and in high quality, video data is typically in lower resolution with actors in variable poses. I have also attempted using the StyleGAN dataset for this competition and while the resulting model performed well on a holdout validation set, it performed poorly on the actual videos. </p>",
      "rawMarkdown": "I would say: very unlikely. Note that the methods used for generating StyleGAN fake faces and deepfake creation have significant differences that make generalization difficult (eg. StyleGAN is a mapping from the latent code, deepfake techniques map from the distribution of real faces,  StyleGAN generates entire faces and backgrounds , while deepfakes only modify the face with pre-processing steps such as face alignment and 3d modelling etc. ) Furthermore, the datasets have widely differing characteristics as well; whereas StyleGAN faces are always forward-facing and in high quality, video data is typically in lower resolution with actors in variable poses. I have also attempted using the StyleGAN dataset for this competition and while the resulting model performed well on a holdout validation set, it performed poorly on the actual videos. ",
      "votes": 2,
      "replies": [
        {
          "id": 778566,
          "postDate": "2020-03-18T15:06:15.550Z",
          "content": "<p>So did you try to blur the image or make them shaper?</p>",
          "rawMarkdown": "So did you try to blur the image or make them shaper?",
          "isDeleted": true
        },
        {
          "id": 778586,
          "postDate": "2020-03-18T15:20:01.747Z",
          "content": "<p>You could try, but I'm doubtful that you would see any results. The main problem is that the underlying distributions of the two datasets are just too dissimilar. This makes sense, considering that they have different objectives. In particular, this specific StyleGAN dataset is known to have distinctive artefacts  that make it trivial for any modern CNN to classify correctly.   </p>",
          "rawMarkdown": "You could try, but I'm doubtful that you would see any results. The main problem is that the underlying distributions of the two datasets are just too dissimilar. This makes sense, considering that they have different objectives. In particular, this specific StyleGAN dataset is known to have distinctive artefacts  that make it trivial for any modern CNN to classify correctly.   ",
          "votes": 1
        }
      ]
    },
    {
      "id": 778017,
      "postDate": "2020-03-18T04:45:48.910Z",
      "content": "<p>Has anyone tried @tunguz 's dataset of a million fake faces? I'm currently trying to make it work along with the celebrity faces dataset.</p>",
      "rawMarkdown": "Has anyone tried @tunguz 's dataset of a million fake faces? I'm currently trying to make it work along with the celebrity faces dataset.",
      "votes": 2
    },
    {
      "id": 778324,
      "postDate": "2020-03-18T10:49:25.727Z",
      "content": "<p>I have tried to predict fake/real faces using NASNet and got the results pretty close to 0 log loss on hold out data set. I believe it is much easier to identify these fakes because of the method. GANs perform well on generating the face, but do pretty bad on background and edges between face and background. I am not sure these faces can help in the competition since the methods for creating the dataset is different. Maybe the real faces can help.</p>",
      "rawMarkdown": "I have tried to predict fake/real faces using NASNet and got the results pretty close to 0 log loss on hold out data set. I believe it is much easier to identify these fakes because of the method. GANs perform well on generating the face, but do pretty bad on background and edges between face and background. I am not sure these faces can help in the competition since the methods for creating the dataset is different. Maybe the real faces can help.",
      "replies": [
        {
          "id": 778414,
          "postDate": "2020-03-18T12:20:36.743Z",
          "content": "<p>I'm using the Celeba dataset for real faces.</p>",
          "rawMarkdown": "I'm using the Celeba dataset for real faces."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 778409,
      "author_name": "seraphimstreets",
      "author_url": "",
      "post_date": "2020-03-18T12:17:37.447000",
      "content": "<p>I would say: very unlikely. Note that the methods used for generating StyleGAN fake faces and deepfake creation have significant differences that make generalization difficult (eg. StyleGAN is a mapping from the latent code, deepfake techniques map from the distribution of real faces,  StyleGAN generates entire faces and backgrounds , while deepfakes only modify the face with pre-processing steps such as face alignment and 3d modelling etc. ) Furthermore, the datasets have widely differing characteristics as well; whereas StyleGAN faces are always forward-facing and in high quality, video data is typically in lower resolution with actors in variable poses. I have also attempted using the StyleGAN dataset for this competition and while the resulting model performed well on a holdout validation set, it performed poorly on the actual videos. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 778566,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-18T15:06:15.550000",
          "content": "<p>So did you try to blur the image or make them shaper?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 778586,
          "author_name": "seraphimstreets",
          "author_url": "",
          "post_date": "2020-03-18T15:20:01.747000",
          "content": "<p>You could try, but I'm doubtful that you would see any results. The main problem is that the underlying distributions of the two datasets are just too dissimilar. This makes sense, considering that they have different objectives. In particular, this specific StyleGAN dataset is known to have distinctive artefacts  that make it trivial for any modern CNN to classify correctly.   </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 778324,
      "author_name": "Victor Paslay",
      "author_url": "",
      "post_date": "2020-03-18T10:49:25.727000",
      "content": "<p>I have tried to predict fake/real faces using NASNet and got the results pretty close to 0 log loss on hold out data set. I believe it is much easier to identify these fakes because of the method. GANs perform well on generating the face, but do pretty bad on background and edges between face and background. I am not sure these faces can help in the competition since the methods for creating the dataset is different. Maybe the real faces can help.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 778414,
          "author_name": "Trigram",
          "author_url": "",
          "post_date": "2020-03-18T12:20:36.743000",
          "content": "<p>I'm using the Celeba dataset for real faces.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "778409": "I would say: very unlikely. Note that the methods used for generating StyleGAN fake faces and deepfake creation have significant differences that make generalization difficult (eg. StyleGAN is a mapping from the latent code, deepfake techniques map from the distribution of real faces,  StyleGAN generates entire faces and backgrounds , while deepfakes only modify the face with pre-processing steps such as face alignment and 3d modelling etc. ) Furthermore, the datasets have widely differing characteristics as well; whereas StyleGAN faces are always forward-facing and in high quality, video data is typically in lower resolution with actors in variable poses. I have also attempted using the StyleGAN dataset for this competition and while the resulting model performed well on a holdout validation set, it performed poorly on the actual videos. ",
    "778017": "Has anyone tried @tunguz 's dataset of a million fake faces? I'm currently trying to make it work along with the celebrity faces dataset.",
    "778324": "I have tried to predict fake/real faces using NASNet and got the results pretty close to 0 log loss on hold out data set. I believe it is much easier to identify these fakes because of the method. GANs perform well on generating the face, but do pretty bad on background and edges between face and background. I am not sure these faces can help in the competition since the methods for creating the dataset is different. Maybe the real faces can help."
  }
}