{
  "id": 127338,
  "title": "Public test face recognition videos fundamentally different from training/public val? Issues with face_recognition",
  "url": "/competitions/deepfake-detection-challenge/discussion/127338",
  "author_name": "James Howard",
  "post_date": "2020-01-23T11:29:17.302000",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p>EDIT: The below findings were due to Kaggle moving on to the next code block when an exception was encountered; wrapping every function is a try/except is strongly advised.</p>\n\n<p>I seem to have fairly strong evidence that the \"face_recognition\" library from pip/conda performs much much worse on the public test set, compared with the training and public validation sets.</p>\n\n<p>I've got a working notebook which allows me to submit scores, which early on in the pipeline uses the standard \"face_recognition\" library. I do the detection at 50% of the native resolution, in the interest of speed.</p>\n\n<p>In the notebook I have an assert statement which bails the notebook if I'm not able to detect a face in at least 10 consecutive frames for at least 50% of videos (before putting them through any networks etc. to save on GPU time).</p>\n\n<p>On the \"/kaggle/input/deepfake-detection-challenge/test_videos\" (public validation set) this is easy - I can detect &gt;= 10 frames of faces in 98% of videos, much more than the required 50%.</p>\n\n<p>An identical submission (therefore using the public test set), however, fails. Without this assert statement, the code runs and submits a score, though, so I know it's the assert statement rather than a crash. That score is almost identical to predicting 0.5 for every video (which I've told it to do for every video it can't find a face), and it runs very quickly (a similar time to the 400 cases, despite it being 4000 cases), so I strongly suspect it's a failure of face_recognition to find faces in the public test set, and so the more time-intensive tasks are skipped.</p>\n\n<p>Has anyone else had such an experience?</p>",
  "messages": [
    {
      "id": 727011,
      "postDate": "2020-01-23T11:29:17.303Z",
      "content": "<p>EDIT: The below findings were due to Kaggle moving on to the next code block when an exception was encountered; wrapping every function is a try/except is strongly advised.</p>\n\n<p>I seem to have fairly strong evidence that the \"face_recognition\" library from pip/conda performs much much worse on the public test set, compared with the training and public validation sets.</p>\n\n<p>I've got a working notebook which allows me to submit scores, which early on in the pipeline uses the standard \"face_recognition\" library. I do the detection at 50% of the native resolution, in the interest of speed.</p>\n\n<p>In the notebook I have an assert statement which bails the notebook if I'm not able to detect a face in at least 10 consecutive frames for at least 50% of videos (before putting them through any networks etc. to save on GPU time).</p>\n\n<p>On the \"/kaggle/input/deepfake-detection-challenge/test_videos\" (public validation set) this is easy - I can detect &gt;= 10 frames of faces in 98% of videos, much more than the required 50%.</p>\n\n<p>An identical submission (therefore using the public test set), however, fails. Without this assert statement, the code runs and submits a score, though, so I know it's the assert statement rather than a crash. That score is almost identical to predicting 0.5 for every video (which I've told it to do for every video it can't find a face), and it runs very quickly (a similar time to the 400 cases, despite it being 4000 cases), so I strongly suspect it's a failure of face_recognition to find faces in the public test set, and so the more time-intensive tasks are skipped.</p>\n\n<p>Has anyone else had such an experience?</p>",
      "rawMarkdown": "EDIT: The below findings were due to Kaggle moving on to the next code block when an exception was encountered; wrapping every function is a try/except is strongly advised.\n\nI seem to have fairly strong evidence that the \"face_recognition\" library from pip/conda performs much much worse on the public test set, compared with the training and public validation sets.\n\nI've got a working notebook which allows me to submit scores, which early on in the pipeline uses the standard \"face_recognition\" library. I do the detection at 50% of the native resolution, in the interest of speed.\n\nIn the notebook I have an assert statement which bails the notebook if I'm not able to detect a face in at least 10 consecutive frames for at least 50% of videos (before putting them through any networks etc. to save on GPU time).\n\nOn the \"/kaggle/input/deepfake-detection-challenge/test_videos\" (public validation set) this is easy - I can detect &gt;= 10 frames of faces in 98% of videos, much more than the required 50%.\n\nAn identical submission (therefore using the public test set), however, fails. Without this assert statement, the code runs and submits a score, though, so I know it's the assert statement rather than a crash. That score is almost identical to predicting 0.5 for every video (which I've told it to do for every video it can't find a face), and it runs very quickly (a similar time to the 400 cases, despite it being 4000 cases), so I strongly suspect it's a failure of face_recognition to find faces in the public test set, and so the more time-intensive tasks are skipped.\n\nHas anyone else had such an experience?",
      "votes": 6
    },
    {
      "id": 747273,
      "postDate": "2020-02-16T07:40:41.203Z",
      "content": "<p>maybe that explains why my scoring on LB was so bad despite good validation scores. 😭 Thanks for the tip. I will try to make a public notebook to confirm that.</p>",
      "rawMarkdown": "maybe that explains why my scoring on LB was so bad despite good validation scores. 😭 Thanks for the tip. I will try to make a public notebook to confirm that."
    },
    {
      "id": 727314,
      "postDate": "2020-01-23T16:15:37.953Z",
      "content": "<p>The face_rocognition is based on dlib and it runs slower than dlib. I recommend you use mtcnn, unless you want to use the 68 facial landmarks that dlib can detect.</p>",
      "rawMarkdown": "The face_rocognition is based on dlib and it runs slower than dlib. I recommend you use mtcnn, unless you want to use the 68 facial landmarks that dlib can detect.",
      "replies": [
        {
          "id": 727325,
          "postDate": "2020-01-23T16:21:33.383Z",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> Interesting. Is MTCNN more accurate than the optional CUDA CNN model that facerecognition has? By using dlib compiled with CUDNN support and doing facerecognition's <code>batch_face_locations</code> I could analyse a video per second, which is very reasonable.</p>\n\n<p>I was very impressed by face_recognition's accuracy this way, but as I said it seems to perform completely differently in the public test set.</p>",
          "rawMarkdown": "@feifeizaici Interesting. Is MTCNN more accurate than the optional CUDA CNN model that facerecognition has? By using dlib compiled with CUDNN support and doing facerecognition's `batch_face_locations` I could analyse a video per second, which is very reasonable.\n\nI was very impressed by face_recognition's accuracy this way, but as I said it seems to perform completely differently in the public test set."
        }
      ]
    },
    {
      "id": 727322,
      "postDate": "2020-01-23T16:21:05.460Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 747273,
      "author_name": "dagnelies",
      "author_url": "",
      "post_date": "2020-02-16T07:40:41.203000",
      "content": "<p>maybe that explains why my scoring on LB was so bad despite good validation scores. 😭 Thanks for the tip. I will try to make a public notebook to confirm that.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727314,
      "author_name": "FrazierLei",
      "author_url": "",
      "post_date": "2020-01-23T16:15:37.953000",
      "content": "<p>The face_rocognition is based on dlib and it runs slower than dlib. I recommend you use mtcnn, unless you want to use the 68 facial landmarks that dlib can detect.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 727325,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-01-23T16:21:33.383000",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a> Interesting. Is MTCNN more accurate than the optional CUDA CNN model that facerecognition has? By using dlib compiled with CUDNN support and doing facerecognition's <code>batch_face_locations</code> I could analyse a video per second, which is very reasonable.</p>\n\n<p>I was very impressed by face_recognition's accuracy this way, but as I said it seems to perform completely differently in the public test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727322,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T16:21:05.460000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "727011": "EDIT: The below findings were due to Kaggle moving on to the next code block when an exception was encountered; wrapping every function is a try/except is strongly advised.\n\nI seem to have fairly strong evidence that the \"face_recognition\" library from pip/conda performs much much worse on the public test set, compared with the training and public validation sets.\n\nI've got a working notebook which allows me to submit scores, which early on in the pipeline uses the standard \"face_recognition\" library. I do the detection at 50% of the native resolution, in the interest of speed.\n\nIn the notebook I have an assert statement which bails the notebook if I'm not able to detect a face in at least 10 consecutive frames for at least 50% of videos (before putting them through any networks etc. to save on GPU time).\n\nOn the \"/kaggle/input/deepfake-detection-challenge/test_videos\" (public validation set) this is easy - I can detect &gt;= 10 frames of faces in 98% of videos, much more than the required 50%.\n\nAn identical submission (therefore using the public test set), however, fails. Without this assert statement, the code runs and submits a score, though, so I know it's the assert statement rather than a crash. That score is almost identical to predicting 0.5 for every video (which I've told it to do for every video it can't find a face), and it runs very quickly (a similar time to the 400 cases, despite it being 4000 cases), so I strongly suspect it's a failure of face_recognition to find faces in the public test set, and so the more time-intensive tasks are skipped.\n\nHas anyone else had such an experience?",
    "747273": "maybe that explains why my scoring on LB was so bad despite good validation scores. 😭 Thanks for the tip. I will try to make a public notebook to confirm that.",
    "727314": "The face_rocognition is based on dlib and it runs slower than dlib. I recommend you use mtcnn, unless you want to use the 68 facial landmarks that dlib can detect.",
    "727322": ""
  }
}