{
  "id": 121235,
  "title": "Huge amount of fakes?",
  "url": "/competitions/deepfake-detection-challenge/discussion/121235",
  "author_name": "",
  "post_date": "2019-12-12T00:14:48.769928900Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I downloaded one of the files and peeked at one of the json files and found 1591 fakes and 108 real. That seems like a massive amount of fakes. Not sure if that file was just an outlier. Will download the rest and explore some more.</p>\n\n<p>I also feel like something is amiss with the scoring, but we'll see what happens with that too. I entered a naive submission with all .49's and got the same score as the public notebook guessing all .5's. Not sure how that is possible. </p>",
  "messages": [
    {
      "id": "693005",
      "postDate": "12/12/2019 00:14:48",
      "content": "<p>I downloaded one of the files and peeked at one of the json files and found 1591 fakes and 108 real. That seems like a massive amount of fakes. Not sure if that file was just an outlier. Will download the rest and explore some more.</p>\n\n<p>I also feel like something is amiss with the scoring, but we'll see what happens with that too. I entered a naive submission with all .49's and got the same score as the public notebook guessing all .5's. Not sure how that is possible. </p>",
      "rawMarkdown": "I downloaded one of the files and peeked at one of the json files and found 1591 fakes and 108 real. That seems like a massive amount of fakes. Not sure if that file was just an outlier. Will download the rest and explore some more.\n\nI also feel like something is amiss with the scoring, but we'll see what happens with that too. I entered a naive submission with all .49's and got the same score as the public notebook guessing all .5's. Not sure how that is possible.",
      "votes": null
    },
    {
      "id": "693835",
      "postDate": "12/12/2019 20:44:49",
      "content": "<p>Yes there is definitely some skewness in the dataset. The sample dataset provided by Kaggle is also skewed. \nOne reason behind this might be that there are multiple fake videos generated from a single real video.</p>",
      "rawMarkdown": "Yes there is definitely some skewness in the dataset. The sample dataset provided by Kaggle is also skewed. \nOne reason behind this might be that there are multiple fake videos generated from a single real video.",
      "votes": null
    },
    {
      "id": "694255",
      "postDate": "12/13/2019 10:56:58",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a>, on the preview dataset the ratio <code>tempered:original</code> was <code>1:0.28</code> and 5214 videos.\nThey took long videos and created 15s clips, then augmentate and finally face swap. More information about the datset and baselines at: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121441#latest-694246\">DFDC: Insights about the Dataset creation, baselines...</a></p>",
      "rawMarkdown": "Hi @ryches, on the preview dataset the ratio ```tempered:original``` was ```1:0.28``` and 5214 videos.\nThey took long videos and created 15s clips, then augmentate and finally face swap. More information about the datset and baselines at: [DFDC: Insights about the Dataset creation, baselines...](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121441#latest-694246)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 693835,
      "author_name": "techytushar",
      "author_url": "",
      "post_date": "12/12/2019 20:44:49",
      "content": "<p>Yes there is definitely some skewness in the dataset. The sample dataset provided by Kaggle is also skewed. \nOne reason behind this might be that there are multiple fake videos generated from a single real video.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 694255,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "12/13/2019 10:56:58",
      "content": "<p>Hi <a href=\"/ryches\">@ryches</a>, on the preview dataset the ratio <code>tempered:original</code> was <code>1:0.28</code> and 5214 videos.\nThey took long videos and created 15s clips, then augmentate and finally face swap. More information about the datset and baselines at: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121441#latest-694246\">DFDC: Insights about the Dataset creation, baselines...</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "693005": "I downloaded one of the files and peeked at one of the json files and found 1591 fakes and 108 real. That seems like a massive amount of fakes. Not sure if that file was just an outlier. Will download the rest and explore some more.\n\nI also feel like something is amiss with the scoring, but we'll see what happens with that too. I entered a naive submission with all .49's and got the same score as the public notebook guessing all .5's. Not sure how that is possible.",
    "693835": "Yes there is definitely some skewness in the dataset. The sample dataset provided by Kaggle is also skewed. \nOne reason behind this might be that there are multiple fake videos generated from a single real video.",
    "694255": "Hi @ryches, on the preview dataset the ratio ```tempered:original``` was ```1:0.28``` and 5214 videos.\nThey took long videos and created 15s clips, then augmentate and finally face swap. More information about the datset and baselines at: [DFDC: Insights about the Dataset creation, baselines...](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121441#latest-694246)"
  },
  "source": "meta"
}