{
  "id": 127492,
  "title": "Dealing with multiple faces detection",
  "url": "/competitions/deepfake-detection-challenge/discussion/127492",
  "author_name": "",
  "post_date": "2020-01-24T08:01:18.050245200Z",
  "votes": 4,
  "comment_count": 8,
  "views": 0,
  "content": "<p>For now I tried only MTCNN for face detection. It works quite well, even if there are some false positives due to faces detected on paintings or t-shirts (see video <code>apatcsqejh.mp4</code> and image below), or simply errors from the model (giving most of the time very small false faces from what I've seen).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1609504%2F985bcab26dcbffc5c3a8bcbad0dbd8a5%2Fapatcsqejh_frame_14_face_1.jpg?generation=1579851710716986&amp;alt=media\" alt=\"t-shirts\"></p>\n\n<p>The issue I'd like to discuss here is detecting faces always in the same order. If two people (let's name them A and B) are in the video, when executing MTCNN on each frame, it does not always detect A first and then B, even if A and B are not moving from where they stand or sit.</p>\n\n<p>It becomes even harder to detect A first and then B if false positives occur and MTCNN doesn't detect the same number of faces in all frames.</p>\n\n<p>We could of course launch a clustering algorithm after extracting faces to check the order of detection of faces and fix it if necessary, but this is clearly computationally expensive. Moreover, if MTCNN detects a face on frames N and N+2 but doesn't detect it on frame N + 1, I'd like to infer the position without some computationally expensive post processing.</p>\n\n<p>Thus here are my questions to start this discussion on dealing with multiple faces detection:\n- Is it possible to detect faces always in the same order with MTCNN or another face detection model?\n- Is there a way to easily extract faces that haven't been detected on only very few frames, to do some kind of inference with MTCNN or another face detection model?</p>",
  "messages": [
    {
      "id": "727930",
      "postDate": "01/24/2020 08:01:18",
      "content": "<p>For now I tried only MTCNN for face detection. It works quite well, even if there are some false positives due to faces detected on paintings or t-shirts (see video <code>apatcsqejh.mp4</code> and image below), or simply errors from the model (giving most of the time very small false faces from what I've seen).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1609504%2F985bcab26dcbffc5c3a8bcbad0dbd8a5%2Fapatcsqejh_frame_14_face_1.jpg?generation=1579851710716986&amp;alt=media\" alt=\"t-shirts\"></p>\n\n<p>The issue I'd like to discuss here is detecting faces always in the same order. If two people (let's name them A and B) are in the video, when executing MTCNN on each frame, it does not always detect A first and then B, even if A and B are not moving from where they stand or sit.</p>\n\n<p>It becomes even harder to detect A first and then B if false positives occur and MTCNN doesn't detect the same number of faces in all frames.</p>\n\n<p>We could of course launch a clustering algorithm after extracting faces to check the order of detection of faces and fix it if necessary, but this is clearly computationally expensive. Moreover, if MTCNN detects a face on frames N and N+2 but doesn't detect it on frame N + 1, I'd like to infer the position without some computationally expensive post processing.</p>\n\n<p>Thus here are my questions to start this discussion on dealing with multiple faces detection:\n- Is it possible to detect faces always in the same order with MTCNN or another face detection model?\n- Is there a way to easily extract faces that haven't been detected on only very few frames, to do some kind of inference with MTCNN or another face detection model?</p>",
      "rawMarkdown": "For now I tried only MTCNN for face detection. It works quite well, even if there are some false positives due to faces detected on paintings or t-shirts (see video `apatcsqejh.mp4` and image below), or simply errors from the model (giving most of the time very small false faces from what I've seen).\n\n![t-shirts](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1609504%2F985bcab26dcbffc5c3a8bcbad0dbd8a5%2Fapatcsqejh_frame_14_face_1.jpg?generation=1579851710716986&amp;alt=media)\n\nThe issue I'd like to discuss here is detecting faces always in the same order. If two people (let's name them A and B) are in the video, when executing MTCNN on each frame, it does not always detect A first and then B, even if A and B are not moving from where they stand or sit.\n\nIt becomes even harder to detect A first and then B if false positives occur and MTCNN doesn't detect the same number of faces in all frames.\n\nWe could of course launch a clustering algorithm after extracting faces to check the order of detection of faces and fix it if necessary, but this is clearly computationally expensive. Moreover, if MTCNN detects a face on frames N and N+2 but doesn't detect it on frame N + 1, I'd like to infer the position without some computationally expensive post processing.\n\nThus here are my questions to start this discussion on dealing with multiple faces detection:\n- Is it possible to detect faces always in the same order with MTCNN or another face detection model?\n- Is there a way to easily extract faces that haven't been detected on only very few frames, to do some kind of inference with MTCNN or another face detection model?",
      "votes": null
    },
    {
      "id": "728018",
      "postDate": "01/24/2020 10:39:44",
      "content": "<p>interesting</p>",
      "rawMarkdown": "interesting",
      "votes": null
    },
    {
      "id": "728105",
      "postDate": "01/24/2020 12:05:09",
      "content": "<p>I don't know any video face detection toolbox for multiple faces (perhaps some will come out of this comp!)\nYet you can use MTCNN to return the boxes and match them to the previous boxes by calculating the overlaps, or the L2 distances.\n<code>boxes, probs = mtcnn.detect(frames)</code></p>",
      "rawMarkdown": "I don't know any video face detection toolbox for multiple faces (perhaps some will come out of this comp!)\nYet you can use MTCNN to return the boxes and match them to the previous boxes by calculating the overlaps, or the L2 distances.\n`boxes, probs = mtcnn.detect(frames)`",
      "votes": null
    },
    {
      "id": "728405",
      "postDate": "01/24/2020 17:28:03",
      "content": "<p>Yes, I would use IOU to calculate the overlap.</p>",
      "rawMarkdown": "Yes, I would use IOU to calculate the overlap.",
      "votes": null
    },
    {
      "id": "728450",
      "postDate": "01/24/2020 18:22:27",
      "content": "<p>By what you wrote it seems you are describing a face tracker, which there are a lot of ways to do it. Now, about inferring the location for future frames you could use a kalman filter</p>",
      "rawMarkdown": "By what you wrote it seems you are describing a face tracker, which there are a lot of ways to do it. Now, about inferring the location for future frames you could use a kalman filter",
      "votes": null
    },
    {
      "id": "728492",
      "postDate": "01/24/2020 19:51:28",
      "content": "<p>Indeed, it seems a face tracker would help to detect faces in the correct order. I looked for a package doing face tracking but I only found pieces of code based on MTCNN or Facenet.</p>",
      "rawMarkdown": "Indeed, it seems a face tracker would help to detect faces in the correct order. I looked for a package doing face tracking but I only found pieces of code based on MTCNN or Facenet.",
      "votes": null
    },
    {
      "id": "728496",
      "postDate": "01/24/2020 19:55:53",
      "content": "<p>I think these are indeed good ways to match faces, this would probably be less expensive than some clustering algorithm afterwards based on distances between images.</p>",
      "rawMarkdown": "I think these are indeed good ways to match faces, this would probably be less expensive than some clustering algorithm afterwards based on distances between images.",
      "votes": null
    },
    {
      "id": "728673",
      "postDate": "01/25/2020 04:13:57",
      "content": "<p>You can store the centers of faces for each frame. Then you could assume all the faces which cluster around a particular center are the same face.</p>",
      "rawMarkdown": "You can store the centers of faces for each frame. Then you could assume all the faces which cluster around a particular center are the same face.",
      "votes": null
    },
    {
      "id": "758714",
      "postDate": "02/28/2020 04:34:08",
      "content": "<p>Just sort by x coordinate. Simple and fast.</p>",
      "rawMarkdown": "Just sort by x coordinate. Simple and fast.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 728018,
      "author_name": "viggo1999",
      "author_url": "",
      "post_date": "01/24/2020 10:39:44",
      "content": "<p>interesting</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 728105,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "01/24/2020 12:05:09",
      "content": "<p>I don't know any video face detection toolbox for multiple faces (perhaps some will come out of this comp!)\nYet you can use MTCNN to return the boxes and match them to the previous boxes by calculating the overlaps, or the L2 distances.\n<code>boxes, probs = mtcnn.detect(frames)</code></p>",
      "votes": null,
      "replies": [
        {
          "id": 728405,
          "author_name": "ynirkin",
          "author_url": "",
          "post_date": "01/24/2020 17:28:03",
          "content": "<p>Yes, I would use IOU to calculate the overlap.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 728496,
          "author_name": "mika30",
          "author_url": "",
          "post_date": "01/24/2020 19:55:53",
          "content": "<p>I think these are indeed good ways to match faces, this would probably be less expensive than some clustering algorithm afterwards based on distances between images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 728450,
      "author_name": "arc144",
      "author_url": "",
      "post_date": "01/24/2020 18:22:27",
      "content": "<p>By what you wrote it seems you are describing a face tracker, which there are a lot of ways to do it. Now, about inferring the location for future frames you could use a kalman filter</p>",
      "votes": null,
      "replies": [
        {
          "id": 728492,
          "author_name": "mika30",
          "author_url": "",
          "post_date": "01/24/2020 19:51:28",
          "content": "<p>Indeed, it seems a face tracker would help to detect faces in the correct order. I looked for a package doing face tracking but I only found pieces of code based on MTCNN or Facenet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 728673,
      "author_name": "teknas",
      "author_url": "",
      "post_date": "01/25/2020 04:13:57",
      "content": "<p>You can store the centers of faces for each frame. Then you could assume all the faces which cluster around a particular center are the same face.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 758714,
      "author_name": "harald",
      "author_url": "",
      "post_date": "02/28/2020 04:34:08",
      "content": "<p>Just sort by x coordinate. Simple and fast.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "727930": "For now I tried only MTCNN for face detection. It works quite well, even if there are some false positives due to faces detected on paintings or t-shirts (see video `apatcsqejh.mp4` and image below), or simply errors from the model (giving most of the time very small false faces from what I've seen).\n\n![t-shirts](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1609504%2F985bcab26dcbffc5c3a8bcbad0dbd8a5%2Fapatcsqejh_frame_14_face_1.jpg?generation=1579851710716986&amp;alt=media)\n\nThe issue I'd like to discuss here is detecting faces always in the same order. If two people (let's name them A and B) are in the video, when executing MTCNN on each frame, it does not always detect A first and then B, even if A and B are not moving from where they stand or sit.\n\nIt becomes even harder to detect A first and then B if false positives occur and MTCNN doesn't detect the same number of faces in all frames.\n\nWe could of course launch a clustering algorithm after extracting faces to check the order of detection of faces and fix it if necessary, but this is clearly computationally expensive. Moreover, if MTCNN detects a face on frames N and N+2 but doesn't detect it on frame N + 1, I'd like to infer the position without some computationally expensive post processing.\n\nThus here are my questions to start this discussion on dealing with multiple faces detection:\n- Is it possible to detect faces always in the same order with MTCNN or another face detection model?\n- Is there a way to easily extract faces that haven't been detected on only very few frames, to do some kind of inference with MTCNN or another face detection model?",
    "728018": "interesting",
    "728105": "I don't know any video face detection toolbox for multiple faces (perhaps some will come out of this comp!)\nYet you can use MTCNN to return the boxes and match them to the previous boxes by calculating the overlaps, or the L2 distances.\n`boxes, probs = mtcnn.detect(frames)`",
    "728405": "Yes, I would use IOU to calculate the overlap.",
    "728450": "By what you wrote it seems you are describing a face tracker, which there are a lot of ways to do it. Now, about inferring the location for future frames you could use a kalman filter",
    "728492": "Indeed, it seems a face tracker would help to detect faces in the correct order. I looked for a package doing face tracking but I only found pieces of code based on MTCNN or Facenet.",
    "728496": "I think these are indeed good ways to match faces, this would probably be less expensive than some clustering algorithm afterwards based on distances between images.",
    "728673": "You can store the centers of faces for each frame. Then you could assume all the faces which cluster around a particular center are the same face.",
    "758714": "Just sort by x coordinate. Simple and fast."
  },
  "source": "meta"
}