{
  "id": 126524,
  "title": "class balance on public data set",
  "url": "/competitions/deepfake-detection-challenge/discussion/126524",
  "author_name": "pete",
  "post_date": "2020-01-18T04:16:41.750000",
  "votes": 8,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Many probably already know this but it is possible to show that there roughly equal numbers of FAKE and REAL videos in the public data set. To show this, make two runs, one with all 0.5 and one with all 0.1. You get the usual 0.694 for the 0.5 case and around 1.2 for the 0.1 case. Matching against expectations for logloss shows that the labels have a 50-50 distribution.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F986843%2F471bdc70568a87c523460f0eb46e125d%2Fmeasured_ll.png?generation=1579320939622421&amp;alt=media\" alt=\"\"></p>\n\n<p>Note that this is not useful for the private leaderboard as we cannot make runs on it prior to the end of the contest.</p>",
  "messages": [
    {
      "id": 722080,
      "postDate": "2020-01-18T04:16:41.750Z",
      "content": "<p>Many probably already know this but it is possible to show that there roughly equal numbers of FAKE and REAL videos in the public data set. To show this, make two runs, one with all 0.5 and one with all 0.1. You get the usual 0.694 for the 0.5 case and around 1.2 for the 0.1 case. Matching against expectations for logloss shows that the labels have a 50-50 distribution.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F986843%2F471bdc70568a87c523460f0eb46e125d%2Fmeasured_ll.png?generation=1579320939622421&amp;alt=media\" alt=\"\"></p>\n\n<p>Note that this is not useful for the private leaderboard as we cannot make runs on it prior to the end of the contest.</p>",
      "rawMarkdown": "Many probably already know this but it is possible to show that there roughly equal numbers of FAKE and REAL videos in the public data set. To show this, make two runs, one with all 0.5 and one with all 0.1. You get the usual 0.694 for the 0.5 case and around 1.2 for the 0.1 case. Matching against expectations for logloss shows that the labels have a 50-50 distribution.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F986843%2F471bdc70568a87c523460f0eb46e125d%2Fmeasured_ll.png?generation=1579320939622421&amp;alt=media)\n\nNote that this is not useful for the private leaderboard as we cannot make runs on it prior to the end of the contest.\n",
      "votes": 8
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "722080": "Many probably already know this but it is possible to show that there roughly equal numbers of FAKE and REAL videos in the public data set. To show this, make two runs, one with all 0.5 and one with all 0.1. You get the usual 0.694 for the 0.5 case and around 1.2 for the 0.1 case. Matching against expectations for logloss shows that the labels have a 50-50 distribution.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F986843%2F471bdc70568a87c523460f0eb46e125d%2Fmeasured_ll.png?generation=1579320939622421&amp;alt=media)\n\nNote that this is not useful for the private leaderboard as we cannot make runs on it prior to the end of the contest.\n"
  }
}