{
  "id": 140450,
  "title": "What are people’s thoughts on fake/real split bias?",
  "url": "/competitions/deepfake-detection-challenge/discussion/140450",
  "author_name": "fergusoci",
  "post_date": "2020-04-01T20:44:44.058000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>One of the items that I was trying to get my head around was the potential for a large drop in private LB owing to the model’s bias in predicting the ratio of fake to real videos.</p>\n\n<p>If we’ve trained models correctly and if the test set is reasonably similar to the training set, then the model <em>should</em> predict the ratio of fake:real correctly. However, if the test set is <em>very</em> different to the train set and we haven’t adequately catered for that, then the models can easily generate very biased results: i.e. predict many more fake than real, or vica versa. </p>\n\n<p>I tried to cater for this by fitting one submission to a beta distribution (guaranteeing a 50:50 split between real and fake), leaving the other sub without this adjustment in case the 50:50 assumption was incorrect. </p>\n\n<p>What do we know about that assumption? Do we know for sure that the private test set is a 50:50 split between fake and real?</p>\n\n<p>Did anyone else come up with ways of addressing this?</p>",
  "messages": [
    {
      "id": 794462,
      "postDate": "2020-04-01T20:44:44.060Z",
      "content": "<p>One of the items that I was trying to get my head around was the potential for a large drop in private LB owing to the model’s bias in predicting the ratio of fake to real videos.</p>\n\n<p>If we’ve trained models correctly and if the test set is reasonably similar to the training set, then the model <em>should</em> predict the ratio of fake:real correctly. However, if the test set is <em>very</em> different to the train set and we haven’t adequately catered for that, then the models can easily generate very biased results: i.e. predict many more fake than real, or vica versa. </p>\n\n<p>I tried to cater for this by fitting one submission to a beta distribution (guaranteeing a 50:50 split between real and fake), leaving the other sub without this adjustment in case the 50:50 assumption was incorrect. </p>\n\n<p>What do we know about that assumption? Do we know for sure that the private test set is a 50:50 split between fake and real?</p>\n\n<p>Did anyone else come up with ways of addressing this?</p>",
      "rawMarkdown": "One of the items that I was trying to get my head around was the potential for a large drop in private LB owing to the model’s bias in predicting the ratio of fake to real videos.\n\nIf we’ve trained models correctly and if the test set is reasonably similar to the training set, then the model *should* predict the ratio of fake:real correctly. However, if the test set is *very* different to the train set and we haven’t adequately catered for that, then the models can easily generate very biased results: i.e. predict many more fake than real, or vica versa. \n\nI tried to cater for this by fitting one submission to a beta distribution (guaranteeing a 50:50 split between real and fake), leaving the other sub without this adjustment in case the 50:50 assumption was incorrect. \n\nWhat do we know about that assumption? Do we know for sure that the private test set is a 50:50 split between fake and real?\n\nDid anyone else come up with ways of addressing this?\n",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "794462": "One of the items that I was trying to get my head around was the potential for a large drop in private LB owing to the model’s bias in predicting the ratio of fake to real videos.\n\nIf we’ve trained models correctly and if the test set is reasonably similar to the training set, then the model *should* predict the ratio of fake:real correctly. However, if the test set is *very* different to the train set and we haven’t adequately catered for that, then the models can easily generate very biased results: i.e. predict many more fake than real, or vica versa. \n\nI tried to cater for this by fitting one submission to a beta distribution (guaranteeing a 50:50 split between real and fake), leaving the other sub without this adjustment in case the 50:50 assumption was incorrect. \n\nWhat do we know about that assumption? Do we know for sure that the private test set is a 50:50 split between fake and real?\n\nDid anyone else come up with ways of addressing this?\n"
  }
}