{
  "id": 121534,
  "title": "Unskew data",
  "url": "/competitions/deepfake-detection-challenge/discussion/121534",
  "author_name": "",
  "post_date": "2019-12-13T22:19:28.920502100Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Hey there,</p>\n\n<p>I have the following rookie question. Can we unskew the data by training on more screenshots from REAL videos?</p>\n\n<p>The result would be a training set of video screenshots with a close to 50:50 FAKE/REAL label distribution opposed to the given 80/20 distribution.</p>\n\n<p>Cheers, NH :)</p>",
  "messages": [
    {
      "id": "694649",
      "postDate": "12/13/2019 22:19:28",
      "content": "<p>Hey there,</p>\n\n<p>I have the following rookie question. Can we unskew the data by training on more screenshots from REAL videos?</p>\n\n<p>The result would be a training set of video screenshots with a close to 50:50 FAKE/REAL label distribution opposed to the given 80/20 distribution.</p>\n\n<p>Cheers, NH :)</p>",
      "rawMarkdown": "Hey there,\n\nI have the following rookie question. Can we unskew the data by training on more screenshots from REAL videos?\n\nThe result would be a training set of video screenshots with a close to 50:50 FAKE/REAL label distribution opposed to the given 80/20 distribution.\n\nCheers, NH :)",
      "votes": null
    },
    {
      "id": "694651",
      "postDate": "12/13/2019 22:42:25",
      "content": "<p>You could provide supplemental datasets but I think they're limited to 1GB (check the rules on that). A better approach is to use some kind of balancing algorithm. For example, if you are using a neural network, your dataloader could over-sample REAL examples and under-sample FAKE examples to train using a 1:1 balance. Look here for some more tips:\n<a href=\"https://www.kdnuggets.com/2017/06/7-techniques-handle-imbalanced-data.html\">https://www.kdnuggets.com/2017/06/7-techniques-handle-imbalanced-data.html</a></p>",
      "rawMarkdown": "You could provide supplemental datasets but I think they're limited to 1GB (check the rules on that). A better approach is to use some kind of balancing algorithm. For example, if you are using a neural network, your dataloader could over-sample REAL examples and under-sample FAKE examples to train using a 1:1 balance. Look here for some more tips:\nhttps://www.kdnuggets.com/2017/06/7-techniques-handle-imbalanced-data.html",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 694651,
      "author_name": "matthewmasters",
      "author_url": "",
      "post_date": "12/13/2019 22:42:25",
      "content": "<p>You could provide supplemental datasets but I think they're limited to 1GB (check the rules on that). A better approach is to use some kind of balancing algorithm. For example, if you are using a neural network, your dataloader could over-sample REAL examples and under-sample FAKE examples to train using a 1:1 balance. Look here for some more tips:\n<a href=\"https://www.kdnuggets.com/2017/06/7-techniques-handle-imbalanced-data.html\">https://www.kdnuggets.com/2017/06/7-techniques-handle-imbalanced-data.html</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "694649": "Hey there,\n\nI have the following rookie question. Can we unskew the data by training on more screenshots from REAL videos?\n\nThe result would be a training set of video screenshots with a close to 50:50 FAKE/REAL label distribution opposed to the given 80/20 distribution.\n\nCheers, NH :)",
    "694651": "You could provide supplemental datasets but I think they're limited to 1GB (check the rules on that). A better approach is to use some kind of balancing algorithm. For example, if you are using a neural network, your dataloader could over-sample REAL examples and under-sample FAKE examples to train using a 1:1 balance. Look here for some more tips:\nhttps://www.kdnuggets.com/2017/06/7-techniques-handle-imbalanced-data.html"
  },
  "source": "meta"
}