{
  "id": 133654,
  "title": "Cropping Faces",
  "url": "/competitions/deepfake-detection-challenge/discussion/133654",
  "author_name": "",
  "post_date": "2020-03-03T17:38:33.781940100Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi all,\nI've been playing around with cropping the faces out of the image before feeding it into a CNN.  I am currently using dlib to do this but it would seem like this would be prohibitively computationally expensive to actually use.  On the model without gpu support (quicker but not as good) it takes about 1.5 seconds per frame and on the model with gpu support, but not using gpu, it takes about 24 seconds.  It has taken me several days (still running) to preprocess the small dataset and there is no way it will be able to be run on the entire dataset.  Additionally, when doing inference on test videos it will take forever to run since each video is around 300 frames, and even if I get 1 second of preprocessing it will still be about 5 minutes a video.\nHas anyone figured out a way to make this a practical part of the preprocessing pipeline?  Maybe some kind of sampling or a quicker model?</p>",
  "messages": [
    {
      "id": "762660",
      "postDate": "03/03/2020 17:38:33",
      "content": "<p>Hi all,\nI've been playing around with cropping the faces out of the image before feeding it into a CNN.  I am currently using dlib to do this but it would seem like this would be prohibitively computationally expensive to actually use.  On the model without gpu support (quicker but not as good) it takes about 1.5 seconds per frame and on the model with gpu support, but not using gpu, it takes about 24 seconds.  It has taken me several days (still running) to preprocess the small dataset and there is no way it will be able to be run on the entire dataset.  Additionally, when doing inference on test videos it will take forever to run since each video is around 300 frames, and even if I get 1 second of preprocessing it will still be about 5 minutes a video.\nHas anyone figured out a way to make this a practical part of the preprocessing pipeline?  Maybe some kind of sampling or a quicker model?</p>",
      "rawMarkdown": "Hi all,\nI've been playing around with cropping the faces out of the image before feeding it into a CNN.  I am currently using dlib to do this but it would seem like this would be prohibitively computationally expensive to actually use.  On the model without gpu support (quicker but not as good) it takes about 1.5 seconds per frame and on the model with gpu support, but not using gpu, it takes about 24 seconds.  It has taken me several days (still running) to preprocess the small dataset and there is no way it will be able to be run on the entire dataset.  Additionally, when doing inference on test videos it will take forever to run since each video is around 300 frames, and even if I get 1 second of preprocessing it will still be about 5 minutes a video.\nHas anyone figured out a way to make this a practical part of the preprocessing pipeline?  Maybe some kind of sampling or a quicker model?",
      "votes": null
    },
    {
      "id": "762716",
      "postDate": "03/03/2020 18:49:34",
      "content": "<p>check the top scoring notebooks <a href=\"/baruchg\">@baruchg</a> they certainly solved that same problem ;)</p>",
      "rawMarkdown": "check the top scoring notebooks @baruchg they certainly solved that same problem ;)",
      "votes": null
    },
    {
      "id": "764826",
      "postDate": "03/06/2020 00:24:43",
      "content": "<p>Thanks!  Yeah I've been experimenting with mtcnn per the notebook and the increase in speed is incredible.</p>",
      "rawMarkdown": "Thanks!  Yeah I've been experimenting with mtcnn per the notebook and the increase in speed is incredible.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 762716,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "03/03/2020 18:49:34",
      "content": "<p>check the top scoring notebooks <a href=\"/baruchg\">@baruchg</a> they certainly solved that same problem ;)</p>",
      "votes": null,
      "replies": [
        {
          "id": 764826,
          "author_name": "baruchg",
          "author_url": "",
          "post_date": "03/06/2020 00:24:43",
          "content": "<p>Thanks!  Yeah I've been experimenting with mtcnn per the notebook and the increase in speed is incredible.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "762660": "Hi all,\nI've been playing around with cropping the faces out of the image before feeding it into a CNN.  I am currently using dlib to do this but it would seem like this would be prohibitively computationally expensive to actually use.  On the model without gpu support (quicker but not as good) it takes about 1.5 seconds per frame and on the model with gpu support, but not using gpu, it takes about 24 seconds.  It has taken me several days (still running) to preprocess the small dataset and there is no way it will be able to be run on the entire dataset.  Additionally, when doing inference on test videos it will take forever to run since each video is around 300 frames, and even if I get 1 second of preprocessing it will still be about 5 minutes a video.\nHas anyone figured out a way to make this a practical part of the preprocessing pipeline?  Maybe some kind of sampling or a quicker model?",
    "762716": "check the top scoring notebooks @baruchg they certainly solved that same problem ;)",
    "764826": "Thanks!  Yeah I've been experimenting with mtcnn per the notebook and the increase in speed is incredible."
  },
  "source": "meta"
}