{
  "id": 132932,
  "title": "A better way to sort out training examples?",
  "url": "/competitions/deepfake-detection-challenge/discussion/132932",
  "author_name": "",
  "post_date": "2020-02-28T19:46:23.816649400Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I used BlazeFace to pull out faces from the videos and store them as jpg files, but some of the \"faces\" are random objects. I'm trying to find a better method to separate out the images that do not contain faces. For now, I have been using Matplotlib to show 25 images in a row with the index number above them, then I keep a notepad open, and write down the indexes of the jpg files that are not faces, and remove them later. It is the most tedious process. </p>\n\n<p>I'm kicking around the idea of using perl to display all the images as a webpage with a form that I can submit to mark the images, but I was curious if there is a better way? Surely someone else has already tackled this?</p>\n\n<p>Note: I'm using Jupyter on Paperspace, and I'd like to keep all of my data on paperspace.com </p>",
  "messages": [
    {
      "id": "759255",
      "postDate": "02/28/2020 19:46:23",
      "content": "<p>I used BlazeFace to pull out faces from the videos and store them as jpg files, but some of the \"faces\" are random objects. I'm trying to find a better method to separate out the images that do not contain faces. For now, I have been using Matplotlib to show 25 images in a row with the index number above them, then I keep a notepad open, and write down the indexes of the jpg files that are not faces, and remove them later. It is the most tedious process. </p>\n\n<p>I'm kicking around the idea of using perl to display all the images as a webpage with a form that I can submit to mark the images, but I was curious if there is a better way? Surely someone else has already tackled this?</p>\n\n<p>Note: I'm using Jupyter on Paperspace, and I'd like to keep all of my data on paperspace.com </p>",
      "rawMarkdown": "I used BlazeFace to pull out faces from the videos and store them as jpg files, but some of the \"faces\" are random objects. I'm trying to find a better method to separate out the images that do not contain faces. For now, I have been using Matplotlib to show 25 images in a row with the index number above them, then I keep a notepad open, and write down the indexes of the jpg files that are not faces, and remove them later. It is the most tedious process. \n\nI'm kicking around the idea of using perl to display all the images as a webpage with a form that I can submit to mark the images, but I was curious if there is a better way? Surely someone else has already tackled this?\n\nNote: I'm using Jupyter on Paperspace, and I'd like to keep all of my data on paperspace.com",
      "votes": null
    },
    {
      "id": "759262",
      "postDate": "02/28/2020 19:58:00",
      "content": "<p>Use a better face detector. You can find good alternatives in the Notebooks. Even the slowest one will be a lot faster than you manually labeling the data ;)</p>",
      "rawMarkdown": "Use a better face detector. You can find good alternatives in the Notebooks. Even the slowest one will be a lot faster than you manually labeling the data ;)",
      "votes": null
    },
    {
      "id": "759270",
      "postDate": "02/28/2020 20:23:42",
      "content": "<p>I guess I should have mentioned that I plan to work through all the fake faces next and pull out frames that do not appear to be fake. I saw in another thread that some of the frames from the fake videos are not fake. I wonder if anyone else has taken this approach?  </p>",
      "rawMarkdown": "I guess I should have mentioned that I plan to work through all the fake faces next and pull out frames that do not appear to be fake. I saw in another thread that some of the frames from the fake videos are not fake. I wonder if anyone else has taken this approach?",
      "votes": null
    },
    {
      "id": "759372",
      "postDate": "02/29/2020 00:16:34",
      "content": "<p>Those videos, occurs not very commonly, besides that, I had created a program, to know that if the faces looks to be real or they are real, it turns out in a fake video, the face is ALWAYS changed, does not look any problem to us, but the pixels are different, and there, having this as a noise in the data will not cause any effect, even if you do remove it, the chances of improvement and how much improvement will not be even a any deal, also, I would suggest this notebook for the face extractor as it is out, <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\">https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison</a></p>",
      "rawMarkdown": "Those videos, occurs not very commonly, besides that, I had created a program, to know that if the faces looks to be real or they are real, it turns out in a fake video, the face is ALWAYS changed, does not look any problem to us, but the pixels are different, and there, having this as a noise in the data will not cause any effect, even if you do remove it, the chances of improvement and how much improvement will not be even a any deal, also, I would suggest this notebook for the face extractor as it is out, https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison",
      "votes": null
    },
    {
      "id": "759426",
      "postDate": "02/29/2020 02:40:46",
      "content": "<p>Thanks Harshit Sheoran ! </p>",
      "rawMarkdown": "Thanks Harshit Sheoran !",
      "votes": null
    },
    {
      "id": "760515",
      "postDate": "03/01/2020 11:59:04",
      "content": "<p>I was wondering the same. But manually labeling sounds like very cumbersome. Also, it does not reflect the fact that you will have the same problem during evaluation.</p>\n\n<p>In my case, I was rather leaning towards using two different face detectors, and two neural nets, and take their average. That way, one producing nonsense can be better compensated by the other. Just a concept though at this point.</p>",
      "rawMarkdown": "I was wondering the same. But manually labeling sounds like very cumbersome. Also, it does not reflect the fact that you will have the same problem during evaluation.\n\nIn my case, I was rather leaning towards using two different face detectors, and two neural nets, and take their average. That way, one producing nonsense can be better compensated by the other. Just a concept though at this point.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 759262,
      "author_name": "seesee",
      "author_url": "",
      "post_date": "02/28/2020 19:58:00",
      "content": "<p>Use a better face detector. You can find good alternatives in the Notebooks. Even the slowest one will be a lot faster than you manually labeling the data ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 759270,
          "author_name": "jonmacpherson",
          "author_url": "",
          "post_date": "02/28/2020 20:23:42",
          "content": "<p>I guess I should have mentioned that I plan to work through all the fake faces next and pull out frames that do not appear to be fake. I saw in another thread that some of the frames from the fake videos are not fake. I wonder if anyone else has taken this approach?  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759372,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/29/2020 00:16:34",
          "content": "<p>Those videos, occurs not very commonly, besides that, I had created a program, to know that if the faces looks to be real or they are real, it turns out in a fake video, the face is ALWAYS changed, does not look any problem to us, but the pixels are different, and there, having this as a noise in the data will not cause any effect, even if you do remove it, the chances of improvement and how much improvement will not be even a any deal, also, I would suggest this notebook for the face extractor as it is out, <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison\">https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759426,
          "author_name": "jonmacpherson",
          "author_url": "",
          "post_date": "02/29/2020 02:40:46",
          "content": "<p>Thanks Harshit Sheoran ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 760515,
      "author_name": "dagnelies",
      "author_url": "",
      "post_date": "03/01/2020 11:59:04",
      "content": "<p>I was wondering the same. But manually labeling sounds like very cumbersome. Also, it does not reflect the fact that you will have the same problem during evaluation.</p>\n\n<p>In my case, I was rather leaning towards using two different face detectors, and two neural nets, and take their average. That way, one producing nonsense can be better compensated by the other. Just a concept though at this point.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "759255": "I used BlazeFace to pull out faces from the videos and store them as jpg files, but some of the \"faces\" are random objects. I'm trying to find a better method to separate out the images that do not contain faces. For now, I have been using Matplotlib to show 25 images in a row with the index number above them, then I keep a notepad open, and write down the indexes of the jpg files that are not faces, and remove them later. It is the most tedious process. \n\nI'm kicking around the idea of using perl to display all the images as a webpage with a form that I can submit to mark the images, but I was curious if there is a better way? Surely someone else has already tackled this?\n\nNote: I'm using Jupyter on Paperspace, and I'd like to keep all of my data on paperspace.com",
    "759262": "Use a better face detector. You can find good alternatives in the Notebooks. Even the slowest one will be a lot faster than you manually labeling the data ;)",
    "759270": "I guess I should have mentioned that I plan to work through all the fake faces next and pull out frames that do not appear to be fake. I saw in another thread that some of the frames from the fake videos are not fake. I wonder if anyone else has taken this approach?",
    "759372": "Those videos, occurs not very commonly, besides that, I had created a program, to know that if the faces looks to be real or they are real, it turns out in a fake video, the face is ALWAYS changed, does not look any problem to us, but the pixels are different, and there, having this as a noise in the data will not cause any effect, even if you do remove it, the chances of improvement and how much improvement will not be even a any deal, also, I would suggest this notebook for the face extractor as it is out, https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-comparison",
    "759426": "Thanks Harshit Sheoran !",
    "760515": "I was wondering the same. But manually labeling sounds like very cumbersome. Also, it does not reflect the fact that you will have the same problem during evaluation.\n\nIn my case, I was rather leaning towards using two different face detectors, and two neural nets, and take their average. That way, one producing nonsense can be better compensated by the other. Just a concept though at this point."
  },
  "source": "meta"
}