{
  "id": 135134,
  "title": "Analyzing Errors (the images)",
  "url": "/competitions/deepfake-detection-challenge/discussion/135134",
  "author_name": "",
  "post_date": "2020-03-12T07:22:07.872348900Z",
  "votes": 18,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Building off of my previous error analysis I decided to narrow in on the videos that had large losses. This was only a fairly small set so it should be possible to analyze these both with automatic summary statistics and also qualitatively viewing the videos and making observations. </p>\n\n<p>The first thing is to inspect the pipeline and see if there is anywhere that is obviously introducing errors. Like if the facial recognition is finding the incorrect region altogether, making it impossible for the model to correctly predict the alterations. </p>\n\n<p>To start off I will look at the false negatives (model deemed real when the correct label was fake)</p>\n\n<p>Here is the table with the false negatives, the number of frames that are in my dataset, the id of the video, the amount of error they contribute, my models predictions and the label. First thing I notice is there are some videos where I only successfully grabbed 6 frames. I am intentionally sampling from the frames, but that is less than I would expect. Probably worth checking those videos out, but there does not seem to be any obvious correlation between those and extremely large error rates.  </p>\n\n<p>| frame_count | image_id       | error    | prediction | label |\n|-------------|----------------|----------|------------|-------|\n| 30          | kgsszrmscq.mp4 | 2.985236 | 0.050528   | 1.0   |\n| 36          | zmtiukzllb.mp4 | 2.792829 | 0.061248   | 1.0   |\n| 90          | lqpitpmzmp.mp4 | 2.720629 | 0.065833   | 1.0   |\n| 90          | ndhmanzwwd.mp4 | 2.694459 | 0.067579   | 1.0   |\n| 48          | nfzbslpenp.mp4 | 2.563123 | 0.077064   | 1.0   |\n| 90          | wqdyfdihyu.mp4 | 2.552531 | 0.077884   | 1.0   |\n| 81          | zhoszukpks.mp4 | 2.540373 | 0.078837   | 1.0   |\n| 90          | favklanlaa.mp4 | 2.528824 | 0.079753   | 1.0   |\n| 30          | radbccigbq.mp4 | 2.452371 | 0.086089   | 1.0   |\n| 30          | jiptqojggg.mp4 | 2.403550 | 0.090396   | 1.0   |\n| ...         | ...            | ...      | ...        | ...   |\n| 6           | gfwgmmwylo.mp4 | 0.760708 | 0.467335   | 1.0   |\n| 30          | rgmyjcyejd.mp4 | 0.756483 | 0.469314   | 1.0   |\n| 30          | zwzlsfkjqv.mp4 | 0.755911 | 0.469583   | 1.0   |\n| 30          | fcujmrcwbl.mp4 | 0.753807 | 0.470572   | 1.0   |\n| 30          | zzdqvspjwv.mp4 | 0.741823 | 0.476245   | 1.0   |\n| 90          | lkkzrnbwtq.mp4 | 0.734829 | 0.479587   | 1.0   |\n| 36          | btlqqvfuck.mp4 | 0.712069 | 0.490628   | 1.0   |\n| 30          | rqsjnyjukt.mp4 | 0.710752 | 0.491275   | 1.0   |\n| 90          | aygsanilyf.mp4 | 0.710258 | 0.491517   | 1.0   |\n| 48          | pceaundfcd.mp4 | 0.707198 | 0.493024   | 1.0   |</p>\n\n<p>Doing a quick correlation between frame count and error reveals a correlation of .218. Obviously not ideal, but only a handful of these videos occur and so might not be worth pursuing. </p>\n\n<p>After looking at the videos where low frame counts occur it appears over half of them are from a specific video with a side-on view of the face and the others are very low contrast. Not surprising that either of these cause issues. It might be possible to remedy this, but does not seem to be a major systemic issue in the pipeline. </p>\n\n<p>| gfwgmmwylo.mp4 |  hrupkssvds.mp4 |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F92fd1af0d45268e5928c6b36d7795f95%2Fdownload%20(4\" alt=\"\">.png?generation=1583995795634408&amp;alt=media) | <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fedf7451f83efa37a77b7db0df6b9fb77%2Fdownload%20(5\" alt=\"\">.png?generation=1583995826501815&amp;alt=media) |</p>\n\n<p>Interestingly we can even see that from the side-on video that had very low frames sometimes the facial recognition correctly found the fake face floating in the sky. That is actually a positive in my book. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff57a02e195b350622dd5785864edb06d%2Fdownload%20(6\" alt=\"\">.png?generation=1583995916342625&amp;alt=media)</p>\n\n<p>Looking at one of the original, non-cropped problematic side-on low frame count videos we can see that there is actually a little distortion ghost that goes from the man's belt, out to the sky and then back and seems to become larger and smaller as it moves from one position to the other. </p>\n\n<p>|aoigsecafe.mp4 aberration in sky| aoigsecafe.mp4 aberration near belt|\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff06c00c1efec3996f86ac1429023a72f%2FScreenshot%20from%202020-03-11%2023-54-35.png?generation=1583996228127765&amp;alt=media\" alt=\"\"> | <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fea4b271b44e64071082bc2643f95d654%2FScreenshot%20from%202020-03-11%2023-56-12.png?generation=1583996241120593&amp;alt=media\" alt=\"\"> |</p>\n\n<p>Alright, well the frame count doesn't appear to be a major driver in these images since it was only a small fraction and didnt seem to be a guaranteed terrible loss. Seems like mostly just an edge case rather than a pattern. </p>\n\n<p>Now we will look at the false negatives to see if there are any obvious trends we can see from the data. Interestingly looking at similar stats for the false positives there are about a dozen videos that are showing up with only 1 frame and it is blank. Likely just an error from my data prep process, but we can look at the source videos to see what the error may have been. The correlation between frames and error is only .09, but if there is something obvious preventing the model from even seeing any frames it is probably worth fixing. </p>\n\n<p>| mfnowqfdwl.mp4     | iqzoqolccf.mp4 |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F3904f14b2dc22382482c9ec7d4979133%2FScreenshot%20from%202020-03-12%2000-10-55.png?generation=1583997070000896&amp;alt=media\" alt=\"\">|<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F178d1f1034e5d59aa581a4d0bd4be33d%2FScreenshot%20from%202020-03-12%2000-11-45.png?generation=1583997125034111&amp;alt=media\" alt=\"\">\n|</p>\n\n<p>The first single frame video shows a very dark, low contrast video, that likely explains the the failed facial recognition and the second video is the same culprit as before. Probably not worth focusing on. Curious that the facial recognition found less faces in the real video than the fake one though. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F0aaa61d7abff60a9b6eaea86ccd3b228%2Fdownload%20(7\" alt=\"\">.png?generation=1583997538119123&amp;alt=media)</p>\n\n<p>Looking at some of the others, I can see that their file size is a magnitude smaller and they look significantly compressed. That's not great, not much we can about those, but it's likely important for the final pipeline to find these. Maybe lowering the threshold for face detection would help or maybe there is some super-resolution method that could be applied here. Still seems like a very small handful of videos this effects though. It is also important to make sure the fake detection model is also robust to compression. This is fairly easy to do with albumentations and other methods as a training augmentation. </p>\n\n<p>From this, I would conclude that the facial detection pipeline largely seems to be working.</p>\n\n<p>Looking at the frame by frame difference between images we can see where and how exactly the videos are being altered. A simple side by side between a real and fake of a certain video is a good simple step to see if the changes are easily perceivable. </p>\n\n<p>| Real | Fake |\n| --- | --- |\n|<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F40ad4bc5e1e6f68de4cdafa9bba63ebb%2Freal.jpg?generation=1584070079461883&amp;alt=media\" alt=\"\">| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8e8853ab19377a98cb1b74132c63d9ee%2Ffake.jpg?generation=1584070097361394&amp;alt=media\" alt=\"\">\n |</p>\n\n<p>The difference is somewhat subtle, especially once kaggle makes the image a bit smaller, but the alteration is definitely apparent on the left man's face. The right man's face appears to be unaltered at least in this frame. </p>\n\n<p>Inspecting the false negative with the highest error ('kgsszrmscq.mp4') it looks like a real video to me. Maybe there is something imperceptible to the human eye that I am missing. Playing the videos side by side and seeing the byte size is within the thousands makes me think there might just be very little different between the real and fake. </p>\n\n<p>The first graph probably worth looking at is just the frame level difference. Not expecting a whole ton out of this as the compression likely introduces pixel-level noise. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8fd604cb0d2e15414459fd793365304d%2Fdownload%20(8\" alt=\"\">.png?generation=1584071463782065&amp;alt=media)</p>\n\n<p>Looking at this graph of the mean absolute difference between pixels over the frames we see the pixel level difference is quite large peaking at 40. Let's inspect that frame. </p>\n\n<p>| Real | Fake |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb0c1b2e3904c4569238520ce988787c9%2Fdownload%20(9\" alt=\"\">.png?generation=1584072287346421&amp;alt=media)| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F81402569a9b6bdbe387e6b9acca955a8%2Fdownload%20(10\" alt=\"\">.png?generation=1584072309888310&amp;alt=media)\n |</p>\n\n<p>To me, both of these look real, but possibly just frame shifted. It may be the case that a frame was dropped in one of the videos or their indexes just werent aligned. I wouldnt put too much weight in the results of this video. </p>\n\n<p>Looking at the next highest error false-negative (zmtiukzllb.mp4) we see a very different profile\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F9b3687d7d286e9f905e37898094d4c67%2Fdownload%20(11\" alt=\"\">.png?generation=1584072490014540&amp;alt=media)</p>\n\n<p>| Real | Fake |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fc13f55bf8cb060c633e0c45561eafa95%2Fdownload%20(12\" alt=\"\">.png?generation=1584072537488215&amp;alt=media)| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F109760847e8bca452ef56ea2810439d2%2Fdownload%20(13\" alt=\"\">.png?generation=1584072558813406&amp;alt=media)\n |</p>\n\n<p>This plot and these images make much more sense now. The frames look to largely be matched, up to a reasonable noise level that would be seen from mp4 compression and the fake and real videos look aligned, albeit the fake face has been destroyed by some blurring. </p>\n\n<p>Looking back through the inputs the model saw I can't really say why the model viewed these as possibly real. The image is obviously blurred that is reaching the model even after face detection. It may have to do with the model seeing many regular looking inputs, the ones that had small differences and only seeing a small set that appeared fake. In the difference plot we can see it seems like the fake maybe only locks on in some frames and not all. </p>\n\n<h2>Meta Analysis</h2>\n\n<p>Next thing we can look at is view the small subset of videos the model is getting wrong and look for commonalities by hand labeling them with various characteristics we might deem relevant. </p>\n\n<p>I chose some simple distinctions I thought might be useful to look into. Specifically I was checking for multiple faces, if there were side-on faces, if the face in the video was moving, if the camera was moving, if the subject was black, if it was shot in low-light and if the face is flickering on and off and if they had what I deemed \"ghosts\" typically a swapped in face that would move around either between faces or just rest in open space. </p>\n\n<p>Looking at the first 50 of the false positives I labeled them with the following characteristics</p>\n\n<p>| Side-on | Moving subject | Moving camera | Black | Low light | flickering | Ghost |\n|---------|----------------|---------------|-------|-----------|------------|-------|\n| 44      | 7              | 7             | 7     | 9         | 29         | 11    |</p>\n\n<p>Looks like there are some more common occurrences than others. These numbers arent necessarily significant until they are compared against the false negative and baseline reals and fakes to determine if the issue is they are coming from a slightly different distribution. </p>\n\n<p>That will come at a later date. </p>",
  "messages": [
    {
      "id": "769722",
      "postDate": "03/12/2020 07:22:07",
      "content": "<p>Building off of my previous error analysis I decided to narrow in on the videos that had large losses. This was only a fairly small set so it should be possible to analyze these both with automatic summary statistics and also qualitatively viewing the videos and making observations. </p>\n\n<p>The first thing is to inspect the pipeline and see if there is anywhere that is obviously introducing errors. Like if the facial recognition is finding the incorrect region altogether, making it impossible for the model to correctly predict the alterations. </p>\n\n<p>To start off I will look at the false negatives (model deemed real when the correct label was fake)</p>\n\n<p>Here is the table with the false negatives, the number of frames that are in my dataset, the id of the video, the amount of error they contribute, my models predictions and the label. First thing I notice is there are some videos where I only successfully grabbed 6 frames. I am intentionally sampling from the frames, but that is less than I would expect. Probably worth checking those videos out, but there does not seem to be any obvious correlation between those and extremely large error rates.  </p>\n\n<p>| frame_count | image_id       | error    | prediction | label |\n|-------------|----------------|----------|------------|-------|\n| 30          | kgsszrmscq.mp4 | 2.985236 | 0.050528   | 1.0   |\n| 36          | zmtiukzllb.mp4 | 2.792829 | 0.061248   | 1.0   |\n| 90          | lqpitpmzmp.mp4 | 2.720629 | 0.065833   | 1.0   |\n| 90          | ndhmanzwwd.mp4 | 2.694459 | 0.067579   | 1.0   |\n| 48          | nfzbslpenp.mp4 | 2.563123 | 0.077064   | 1.0   |\n| 90          | wqdyfdihyu.mp4 | 2.552531 | 0.077884   | 1.0   |\n| 81          | zhoszukpks.mp4 | 2.540373 | 0.078837   | 1.0   |\n| 90          | favklanlaa.mp4 | 2.528824 | 0.079753   | 1.0   |\n| 30          | radbccigbq.mp4 | 2.452371 | 0.086089   | 1.0   |\n| 30          | jiptqojggg.mp4 | 2.403550 | 0.090396   | 1.0   |\n| ...         | ...            | ...      | ...        | ...   |\n| 6           | gfwgmmwylo.mp4 | 0.760708 | 0.467335   | 1.0   |\n| 30          | rgmyjcyejd.mp4 | 0.756483 | 0.469314   | 1.0   |\n| 30          | zwzlsfkjqv.mp4 | 0.755911 | 0.469583   | 1.0   |\n| 30          | fcujmrcwbl.mp4 | 0.753807 | 0.470572   | 1.0   |\n| 30          | zzdqvspjwv.mp4 | 0.741823 | 0.476245   | 1.0   |\n| 90          | lkkzrnbwtq.mp4 | 0.734829 | 0.479587   | 1.0   |\n| 36          | btlqqvfuck.mp4 | 0.712069 | 0.490628   | 1.0   |\n| 30          | rqsjnyjukt.mp4 | 0.710752 | 0.491275   | 1.0   |\n| 90          | aygsanilyf.mp4 | 0.710258 | 0.491517   | 1.0   |\n| 48          | pceaundfcd.mp4 | 0.707198 | 0.493024   | 1.0   |</p>\n\n<p>Doing a quick correlation between frame count and error reveals a correlation of .218. Obviously not ideal, but only a handful of these videos occur and so might not be worth pursuing. </p>\n\n<p>After looking at the videos where low frame counts occur it appears over half of them are from a specific video with a side-on view of the face and the others are very low contrast. Not surprising that either of these cause issues. It might be possible to remedy this, but does not seem to be a major systemic issue in the pipeline. </p>\n\n<p>| gfwgmmwylo.mp4 |  hrupkssvds.mp4 |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F92fd1af0d45268e5928c6b36d7795f95%2Fdownload%20(4\" alt=\"\">.png?generation=1583995795634408&amp;alt=media) | <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fedf7451f83efa37a77b7db0df6b9fb77%2Fdownload%20(5\" alt=\"\">.png?generation=1583995826501815&amp;alt=media) |</p>\n\n<p>Interestingly we can even see that from the side-on video that had very low frames sometimes the facial recognition correctly found the fake face floating in the sky. That is actually a positive in my book. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff57a02e195b350622dd5785864edb06d%2Fdownload%20(6\" alt=\"\">.png?generation=1583995916342625&amp;alt=media)</p>\n\n<p>Looking at one of the original, non-cropped problematic side-on low frame count videos we can see that there is actually a little distortion ghost that goes from the man's belt, out to the sky and then back and seems to become larger and smaller as it moves from one position to the other. </p>\n\n<p>|aoigsecafe.mp4 aberration in sky| aoigsecafe.mp4 aberration near belt|\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff06c00c1efec3996f86ac1429023a72f%2FScreenshot%20from%202020-03-11%2023-54-35.png?generation=1583996228127765&amp;alt=media\" alt=\"\"> | <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fea4b271b44e64071082bc2643f95d654%2FScreenshot%20from%202020-03-11%2023-56-12.png?generation=1583996241120593&amp;alt=media\" alt=\"\"> |</p>\n\n<p>Alright, well the frame count doesn't appear to be a major driver in these images since it was only a small fraction and didnt seem to be a guaranteed terrible loss. Seems like mostly just an edge case rather than a pattern. </p>\n\n<p>Now we will look at the false negatives to see if there are any obvious trends we can see from the data. Interestingly looking at similar stats for the false positives there are about a dozen videos that are showing up with only 1 frame and it is blank. Likely just an error from my data prep process, but we can look at the source videos to see what the error may have been. The correlation between frames and error is only .09, but if there is something obvious preventing the model from even seeing any frames it is probably worth fixing. </p>\n\n<p>| mfnowqfdwl.mp4     | iqzoqolccf.mp4 |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F3904f14b2dc22382482c9ec7d4979133%2FScreenshot%20from%202020-03-12%2000-10-55.png?generation=1583997070000896&amp;alt=media\" alt=\"\">|<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F178d1f1034e5d59aa581a4d0bd4be33d%2FScreenshot%20from%202020-03-12%2000-11-45.png?generation=1583997125034111&amp;alt=media\" alt=\"\">\n|</p>\n\n<p>The first single frame video shows a very dark, low contrast video, that likely explains the the failed facial recognition and the second video is the same culprit as before. Probably not worth focusing on. Curious that the facial recognition found less faces in the real video than the fake one though. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F0aaa61d7abff60a9b6eaea86ccd3b228%2Fdownload%20(7\" alt=\"\">.png?generation=1583997538119123&amp;alt=media)</p>\n\n<p>Looking at some of the others, I can see that their file size is a magnitude smaller and they look significantly compressed. That's not great, not much we can about those, but it's likely important for the final pipeline to find these. Maybe lowering the threshold for face detection would help or maybe there is some super-resolution method that could be applied here. Still seems like a very small handful of videos this effects though. It is also important to make sure the fake detection model is also robust to compression. This is fairly easy to do with albumentations and other methods as a training augmentation. </p>\n\n<p>From this, I would conclude that the facial detection pipeline largely seems to be working.</p>\n\n<p>Looking at the frame by frame difference between images we can see where and how exactly the videos are being altered. A simple side by side between a real and fake of a certain video is a good simple step to see if the changes are easily perceivable. </p>\n\n<p>| Real | Fake |\n| --- | --- |\n|<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F40ad4bc5e1e6f68de4cdafa9bba63ebb%2Freal.jpg?generation=1584070079461883&amp;alt=media\" alt=\"\">| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8e8853ab19377a98cb1b74132c63d9ee%2Ffake.jpg?generation=1584070097361394&amp;alt=media\" alt=\"\">\n |</p>\n\n<p>The difference is somewhat subtle, especially once kaggle makes the image a bit smaller, but the alteration is definitely apparent on the left man's face. The right man's face appears to be unaltered at least in this frame. </p>\n\n<p>Inspecting the false negative with the highest error ('kgsszrmscq.mp4') it looks like a real video to me. Maybe there is something imperceptible to the human eye that I am missing. Playing the videos side by side and seeing the byte size is within the thousands makes me think there might just be very little different between the real and fake. </p>\n\n<p>The first graph probably worth looking at is just the frame level difference. Not expecting a whole ton out of this as the compression likely introduces pixel-level noise. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8fd604cb0d2e15414459fd793365304d%2Fdownload%20(8\" alt=\"\">.png?generation=1584071463782065&amp;alt=media)</p>\n\n<p>Looking at this graph of the mean absolute difference between pixels over the frames we see the pixel level difference is quite large peaking at 40. Let's inspect that frame. </p>\n\n<p>| Real | Fake |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb0c1b2e3904c4569238520ce988787c9%2Fdownload%20(9\" alt=\"\">.png?generation=1584072287346421&amp;alt=media)| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F81402569a9b6bdbe387e6b9acca955a8%2Fdownload%20(10\" alt=\"\">.png?generation=1584072309888310&amp;alt=media)\n |</p>\n\n<p>To me, both of these look real, but possibly just frame shifted. It may be the case that a frame was dropped in one of the videos or their indexes just werent aligned. I wouldnt put too much weight in the results of this video. </p>\n\n<p>Looking at the next highest error false-negative (zmtiukzllb.mp4) we see a very different profile\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F9b3687d7d286e9f905e37898094d4c67%2Fdownload%20(11\" alt=\"\">.png?generation=1584072490014540&amp;alt=media)</p>\n\n<p>| Real | Fake |\n| --- | --- |\n| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fc13f55bf8cb060c633e0c45561eafa95%2Fdownload%20(12\" alt=\"\">.png?generation=1584072537488215&amp;alt=media)| <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F109760847e8bca452ef56ea2810439d2%2Fdownload%20(13\" alt=\"\">.png?generation=1584072558813406&amp;alt=media)\n |</p>\n\n<p>This plot and these images make much more sense now. The frames look to largely be matched, up to a reasonable noise level that would be seen from mp4 compression and the fake and real videos look aligned, albeit the fake face has been destroyed by some blurring. </p>\n\n<p>Looking back through the inputs the model saw I can't really say why the model viewed these as possibly real. The image is obviously blurred that is reaching the model even after face detection. It may have to do with the model seeing many regular looking inputs, the ones that had small differences and only seeing a small set that appeared fake. In the difference plot we can see it seems like the fake maybe only locks on in some frames and not all. </p>\n\n<h2>Meta Analysis</h2>\n\n<p>Next thing we can look at is view the small subset of videos the model is getting wrong and look for commonalities by hand labeling them with various characteristics we might deem relevant. </p>\n\n<p>I chose some simple distinctions I thought might be useful to look into. Specifically I was checking for multiple faces, if there were side-on faces, if the face in the video was moving, if the camera was moving, if the subject was black, if it was shot in low-light and if the face is flickering on and off and if they had what I deemed \"ghosts\" typically a swapped in face that would move around either between faces or just rest in open space. </p>\n\n<p>Looking at the first 50 of the false positives I labeled them with the following characteristics</p>\n\n<p>| Side-on | Moving subject | Moving camera | Black | Low light | flickering | Ghost |\n|---------|----------------|---------------|-------|-----------|------------|-------|\n| 44      | 7              | 7             | 7     | 9         | 29         | 11    |</p>\n\n<p>Looks like there are some more common occurrences than others. These numbers arent necessarily significant until they are compared against the false negative and baseline reals and fakes to determine if the issue is they are coming from a slightly different distribution. </p>\n\n<p>That will come at a later date. </p>",
      "rawMarkdown": "Building off of my previous error analysis I decided to narrow in on the videos that had large losses. This was only a fairly small set so it should be possible to analyze these both with automatic summary statistics and also qualitatively viewing the videos and making observations. \n\nThe first thing is to inspect the pipeline and see if there is anywhere that is obviously introducing errors. Like if the facial recognition is finding the incorrect region altogether, making it impossible for the model to correctly predict the alterations. \n\nTo start off I will look at the false negatives (model deemed real when the correct label was fake)\n\nHere is the table with the false negatives, the number of frames that are in my dataset, the id of the video, the amount of error they contribute, my models predictions and the label. First thing I notice is there are some videos where I only successfully grabbed 6 frames. I am intentionally sampling from the frames, but that is less than I would expect. Probably worth checking those videos out, but there does not seem to be any obvious correlation between those and extremely large error rates.  \n\n| frame_count | image_id       | error    | prediction | label |\n|-------------|----------------|----------|------------|-------|\n| 30          | kgsszrmscq.mp4 | 2.985236 | 0.050528   | 1.0   |\n| 36          | zmtiukzllb.mp4 | 2.792829 | 0.061248   | 1.0   |\n| 90          | lqpitpmzmp.mp4 | 2.720629 | 0.065833   | 1.0   |\n| 90          | ndhmanzwwd.mp4 | 2.694459 | 0.067579   | 1.0   |\n| 48          | nfzbslpenp.mp4 | 2.563123 | 0.077064   | 1.0   |\n| 90          | wqdyfdihyu.mp4 | 2.552531 | 0.077884   | 1.0   |\n| 81          | zhoszukpks.mp4 | 2.540373 | 0.078837   | 1.0   |\n| 90          | favklanlaa.mp4 | 2.528824 | 0.079753   | 1.0   |\n| 30          | radbccigbq.mp4 | 2.452371 | 0.086089   | 1.0   |\n| 30          | jiptqojggg.mp4 | 2.403550 | 0.090396   | 1.0   |\n| ...         | ...            | ...      | ...        | ...   |\n| 6           | gfwgmmwylo.mp4 | 0.760708 | 0.467335   | 1.0   |\n| 30          | rgmyjcyejd.mp4 | 0.756483 | 0.469314   | 1.0   |\n| 30          | zwzlsfkjqv.mp4 | 0.755911 | 0.469583   | 1.0   |\n| 30          | fcujmrcwbl.mp4 | 0.753807 | 0.470572   | 1.0   |\n| 30          | zzdqvspjwv.mp4 | 0.741823 | 0.476245   | 1.0   |\n| 90          | lkkzrnbwtq.mp4 | 0.734829 | 0.479587   | 1.0   |\n| 36          | btlqqvfuck.mp4 | 0.712069 | 0.490628   | 1.0   |\n| 30          | rqsjnyjukt.mp4 | 0.710752 | 0.491275   | 1.0   |\n| 90          | aygsanilyf.mp4 | 0.710258 | 0.491517   | 1.0   |\n| 48          | pceaundfcd.mp4 | 0.707198 | 0.493024   | 1.0   |\n\nDoing a quick correlation between frame count and error reveals a correlation of .218. Obviously not ideal, but only a handful of these videos occur and so might not be worth pursuing. \n\nAfter looking at the videos where low frame counts occur it appears over half of them are from a specific video with a side-on view of the face and the others are very low contrast. Not surprising that either of these cause issues. It might be possible to remedy this, but does not seem to be a major systemic issue in the pipeline. \n\n| gfwgmmwylo.mp4 |  hrupkssvds.mp4 |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F92fd1af0d45268e5928c6b36d7795f95%2Fdownload%20(4).png?generation=1583995795634408&amp;alt=media) | ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fedf7451f83efa37a77b7db0df6b9fb77%2Fdownload%20(5).png?generation=1583995826501815&amp;alt=media) |\n\nInterestingly we can even see that from the side-on video that had very low frames sometimes the facial recognition correctly found the fake face floating in the sky. That is actually a positive in my book. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff57a02e195b350622dd5785864edb06d%2Fdownload%20(6).png?generation=1583995916342625&amp;alt=media)\n\nLooking at one of the original, non-cropped problematic side-on low frame count videos we can see that there is actually a little distortion ghost that goes from the man's belt, out to the sky and then back and seems to become larger and smaller as it moves from one position to the other. \n\n|aoigsecafe.mp4 aberration in sky| aoigsecafe.mp4 aberration near belt|\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff06c00c1efec3996f86ac1429023a72f%2FScreenshot%20from%202020-03-11%2023-54-35.png?generation=1583996228127765&amp;alt=media) | ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fea4b271b44e64071082bc2643f95d654%2FScreenshot%20from%202020-03-11%2023-56-12.png?generation=1583996241120593&amp;alt=media) |\n\nAlright, well the frame count doesn't appear to be a major driver in these images since it was only a small fraction and didnt seem to be a guaranteed terrible loss. Seems like mostly just an edge case rather than a pattern. \n\nNow we will look at the false negatives to see if there are any obvious trends we can see from the data. Interestingly looking at similar stats for the false positives there are about a dozen videos that are showing up with only 1 frame and it is blank. Likely just an error from my data prep process, but we can look at the source videos to see what the error may have been. The correlation between frames and error is only .09, but if there is something obvious preventing the model from even seeing any frames it is probably worth fixing. \n\n| mfnowqfdwl.mp4\t | iqzoqolccf.mp4 |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F3904f14b2dc22382482c9ec7d4979133%2FScreenshot%20from%202020-03-12%2000-10-55.png?generation=1583997070000896&amp;alt=media)|![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F178d1f1034e5d59aa581a4d0bd4be33d%2FScreenshot%20from%202020-03-12%2000-11-45.png?generation=1583997125034111&amp;alt=media)\n|\n\nThe first single frame video shows a very dark, low contrast video, that likely explains the the failed facial recognition and the second video is the same culprit as before. Probably not worth focusing on. Curious that the facial recognition found less faces in the real video than the fake one though. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F0aaa61d7abff60a9b6eaea86ccd3b228%2Fdownload%20(7).png?generation=1583997538119123&amp;alt=media)\n\n\nLooking at some of the others, I can see that their file size is a magnitude smaller and they look significantly compressed. That's not great, not much we can about those, but it's likely important for the final pipeline to find these. Maybe lowering the threshold for face detection would help or maybe there is some super-resolution method that could be applied here. Still seems like a very small handful of videos this effects though. It is also important to make sure the fake detection model is also robust to compression. This is fairly easy to do with albumentations and other methods as a training augmentation. \n\nFrom this, I would conclude that the facial detection pipeline largely seems to be working.\n\nLooking at the frame by frame difference between images we can see where and how exactly the videos are being altered. A simple side by side between a real and fake of a certain video is a good simple step to see if the changes are easily perceivable. \n\n| Real | Fake |\n| --- | --- |\n|![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F40ad4bc5e1e6f68de4cdafa9bba63ebb%2Freal.jpg?generation=1584070079461883&amp;alt=media)| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8e8853ab19377a98cb1b74132c63d9ee%2Ffake.jpg?generation=1584070097361394&amp;alt=media)\n |\n\nThe difference is somewhat subtle, especially once kaggle makes the image a bit smaller, but the alteration is definitely apparent on the left man's face. The right man's face appears to be unaltered at least in this frame. \n\nInspecting the false negative with the highest error ('kgsszrmscq.mp4') it looks like a real video to me. Maybe there is something imperceptible to the human eye that I am missing. Playing the videos side by side and seeing the byte size is within the thousands makes me think there might just be very little different between the real and fake. \n\nThe first graph probably worth looking at is just the frame level difference. Not expecting a whole ton out of this as the compression likely introduces pixel-level noise. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8fd604cb0d2e15414459fd793365304d%2Fdownload%20(8).png?generation=1584071463782065&amp;alt=media)\n\nLooking at this graph of the mean absolute difference between pixels over the frames we see the pixel level difference is quite large peaking at 40. Let's inspect that frame. \n\n| Real | Fake |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb0c1b2e3904c4569238520ce988787c9%2Fdownload%20(9).png?generation=1584072287346421&amp;alt=media)| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F81402569a9b6bdbe387e6b9acca955a8%2Fdownload%20(10).png?generation=1584072309888310&amp;alt=media)\n |\n\nTo me, both of these look real, but possibly just frame shifted. It may be the case that a frame was dropped in one of the videos or their indexes just werent aligned. I wouldnt put too much weight in the results of this video. \n\nLooking at the next highest error false-negative (zmtiukzllb.mp4) we see a very different profile\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F9b3687d7d286e9f905e37898094d4c67%2Fdownload%20(11).png?generation=1584072490014540&amp;alt=media)\n\n| Real | Fake |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fc13f55bf8cb060c633e0c45561eafa95%2Fdownload%20(12).png?generation=1584072537488215&amp;alt=media)| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F109760847e8bca452ef56ea2810439d2%2Fdownload%20(13).png?generation=1584072558813406&amp;alt=media)\n |\n\nThis plot and these images make much more sense now. The frames look to largely be matched, up to a reasonable noise level that would be seen from mp4 compression and the fake and real videos look aligned, albeit the fake face has been destroyed by some blurring. \n\nLooking back through the inputs the model saw I can't really say why the model viewed these as possibly real. The image is obviously blurred that is reaching the model even after face detection. It may have to do with the model seeing many regular looking inputs, the ones that had small differences and only seeing a small set that appeared fake. In the difference plot we can see it seems like the fake maybe only locks on in some frames and not all. \n\n##Meta Analysis\nNext thing we can look at is view the small subset of videos the model is getting wrong and look for commonalities by hand labeling them with various characteristics we might deem relevant. \n\nI chose some simple distinctions I thought might be useful to look into. Specifically I was checking for multiple faces, if there were side-on faces, if the face in the video was moving, if the camera was moving, if the subject was black, if it was shot in low-light and if the face is flickering on and off and if they had what I deemed \"ghosts\" typically a swapped in face that would move around either between faces or just rest in open space. \n\nLooking at the first 50 of the false positives I labeled them with the following characteristics\n\n| Side-on | Moving subject | Moving camera | Black | Low light | flickering | Ghost |\n|---------|----------------|---------------|-------|-----------|------------|-------|\n| 44      | 7              | 7             | 7     | 9         | 29         | 11    |\n\nLooks like there are some more common occurrences than others. These numbers arent necessarily significant until they are compared against the false negative and baseline reals and fakes to determine if the issue is they are coming from a slightly different distribution. \n\nThat will come at a later date.",
      "votes": null
    },
    {
      "id": "770807",
      "postDate": "03/13/2020 12:40:03",
      "content": "<p>Thanks! Very good analysis :)</p>",
      "rawMarkdown": "Thanks! Very good analysis :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 770807,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "03/13/2020 12:40:03",
      "content": "<p>Thanks! Very good analysis :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "769722": "Building off of my previous error analysis I decided to narrow in on the videos that had large losses. This was only a fairly small set so it should be possible to analyze these both with automatic summary statistics and also qualitatively viewing the videos and making observations. \n\nThe first thing is to inspect the pipeline and see if there is anywhere that is obviously introducing errors. Like if the facial recognition is finding the incorrect region altogether, making it impossible for the model to correctly predict the alterations. \n\nTo start off I will look at the false negatives (model deemed real when the correct label was fake)\n\nHere is the table with the false negatives, the number of frames that are in my dataset, the id of the video, the amount of error they contribute, my models predictions and the label. First thing I notice is there are some videos where I only successfully grabbed 6 frames. I am intentionally sampling from the frames, but that is less than I would expect. Probably worth checking those videos out, but there does not seem to be any obvious correlation between those and extremely large error rates.  \n\n| frame_count | image_id       | error    | prediction | label |\n|-------------|----------------|----------|------------|-------|\n| 30          | kgsszrmscq.mp4 | 2.985236 | 0.050528   | 1.0   |\n| 36          | zmtiukzllb.mp4 | 2.792829 | 0.061248   | 1.0   |\n| 90          | lqpitpmzmp.mp4 | 2.720629 | 0.065833   | 1.0   |\n| 90          | ndhmanzwwd.mp4 | 2.694459 | 0.067579   | 1.0   |\n| 48          | nfzbslpenp.mp4 | 2.563123 | 0.077064   | 1.0   |\n| 90          | wqdyfdihyu.mp4 | 2.552531 | 0.077884   | 1.0   |\n| 81          | zhoszukpks.mp4 | 2.540373 | 0.078837   | 1.0   |\n| 90          | favklanlaa.mp4 | 2.528824 | 0.079753   | 1.0   |\n| 30          | radbccigbq.mp4 | 2.452371 | 0.086089   | 1.0   |\n| 30          | jiptqojggg.mp4 | 2.403550 | 0.090396   | 1.0   |\n| ...         | ...            | ...      | ...        | ...   |\n| 6           | gfwgmmwylo.mp4 | 0.760708 | 0.467335   | 1.0   |\n| 30          | rgmyjcyejd.mp4 | 0.756483 | 0.469314   | 1.0   |\n| 30          | zwzlsfkjqv.mp4 | 0.755911 | 0.469583   | 1.0   |\n| 30          | fcujmrcwbl.mp4 | 0.753807 | 0.470572   | 1.0   |\n| 30          | zzdqvspjwv.mp4 | 0.741823 | 0.476245   | 1.0   |\n| 90          | lkkzrnbwtq.mp4 | 0.734829 | 0.479587   | 1.0   |\n| 36          | btlqqvfuck.mp4 | 0.712069 | 0.490628   | 1.0   |\n| 30          | rqsjnyjukt.mp4 | 0.710752 | 0.491275   | 1.0   |\n| 90          | aygsanilyf.mp4 | 0.710258 | 0.491517   | 1.0   |\n| 48          | pceaundfcd.mp4 | 0.707198 | 0.493024   | 1.0   |\n\nDoing a quick correlation between frame count and error reveals a correlation of .218. Obviously not ideal, but only a handful of these videos occur and so might not be worth pursuing. \n\nAfter looking at the videos where low frame counts occur it appears over half of them are from a specific video with a side-on view of the face and the others are very low contrast. Not surprising that either of these cause issues. It might be possible to remedy this, but does not seem to be a major systemic issue in the pipeline. \n\n| gfwgmmwylo.mp4 |  hrupkssvds.mp4 |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F92fd1af0d45268e5928c6b36d7795f95%2Fdownload%20(4).png?generation=1583995795634408&amp;alt=media) | ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fedf7451f83efa37a77b7db0df6b9fb77%2Fdownload%20(5).png?generation=1583995826501815&amp;alt=media) |\n\nInterestingly we can even see that from the side-on video that had very low frames sometimes the facial recognition correctly found the fake face floating in the sky. That is actually a positive in my book. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff57a02e195b350622dd5785864edb06d%2Fdownload%20(6).png?generation=1583995916342625&amp;alt=media)\n\nLooking at one of the original, non-cropped problematic side-on low frame count videos we can see that there is actually a little distortion ghost that goes from the man's belt, out to the sky and then back and seems to become larger and smaller as it moves from one position to the other. \n\n|aoigsecafe.mp4 aberration in sky| aoigsecafe.mp4 aberration near belt|\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff06c00c1efec3996f86ac1429023a72f%2FScreenshot%20from%202020-03-11%2023-54-35.png?generation=1583996228127765&amp;alt=media) | ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fea4b271b44e64071082bc2643f95d654%2FScreenshot%20from%202020-03-11%2023-56-12.png?generation=1583996241120593&amp;alt=media) |\n\nAlright, well the frame count doesn't appear to be a major driver in these images since it was only a small fraction and didnt seem to be a guaranteed terrible loss. Seems like mostly just an edge case rather than a pattern. \n\nNow we will look at the false negatives to see if there are any obvious trends we can see from the data. Interestingly looking at similar stats for the false positives there are about a dozen videos that are showing up with only 1 frame and it is blank. Likely just an error from my data prep process, but we can look at the source videos to see what the error may have been. The correlation between frames and error is only .09, but if there is something obvious preventing the model from even seeing any frames it is probably worth fixing. \n\n| mfnowqfdwl.mp4\t | iqzoqolccf.mp4 |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F3904f14b2dc22382482c9ec7d4979133%2FScreenshot%20from%202020-03-12%2000-10-55.png?generation=1583997070000896&amp;alt=media)|![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F178d1f1034e5d59aa581a4d0bd4be33d%2FScreenshot%20from%202020-03-12%2000-11-45.png?generation=1583997125034111&amp;alt=media)\n|\n\nThe first single frame video shows a very dark, low contrast video, that likely explains the the failed facial recognition and the second video is the same culprit as before. Probably not worth focusing on. Curious that the facial recognition found less faces in the real video than the fake one though. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F0aaa61d7abff60a9b6eaea86ccd3b228%2Fdownload%20(7).png?generation=1583997538119123&amp;alt=media)\n\n\nLooking at some of the others, I can see that their file size is a magnitude smaller and they look significantly compressed. That's not great, not much we can about those, but it's likely important for the final pipeline to find these. Maybe lowering the threshold for face detection would help or maybe there is some super-resolution method that could be applied here. Still seems like a very small handful of videos this effects though. It is also important to make sure the fake detection model is also robust to compression. This is fairly easy to do with albumentations and other methods as a training augmentation. \n\nFrom this, I would conclude that the facial detection pipeline largely seems to be working.\n\nLooking at the frame by frame difference between images we can see where and how exactly the videos are being altered. A simple side by side between a real and fake of a certain video is a good simple step to see if the changes are easily perceivable. \n\n| Real | Fake |\n| --- | --- |\n|![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F40ad4bc5e1e6f68de4cdafa9bba63ebb%2Freal.jpg?generation=1584070079461883&amp;alt=media)| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8e8853ab19377a98cb1b74132c63d9ee%2Ffake.jpg?generation=1584070097361394&amp;alt=media)\n |\n\nThe difference is somewhat subtle, especially once kaggle makes the image a bit smaller, but the alteration is definitely apparent on the left man's face. The right man's face appears to be unaltered at least in this frame. \n\nInspecting the false negative with the highest error ('kgsszrmscq.mp4') it looks like a real video to me. Maybe there is something imperceptible to the human eye that I am missing. Playing the videos side by side and seeing the byte size is within the thousands makes me think there might just be very little different between the real and fake. \n\nThe first graph probably worth looking at is just the frame level difference. Not expecting a whole ton out of this as the compression likely introduces pixel-level noise. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F8fd604cb0d2e15414459fd793365304d%2Fdownload%20(8).png?generation=1584071463782065&amp;alt=media)\n\nLooking at this graph of the mean absolute difference between pixels over the frames we see the pixel level difference is quite large peaking at 40. Let's inspect that frame. \n\n| Real | Fake |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fb0c1b2e3904c4569238520ce988787c9%2Fdownload%20(9).png?generation=1584072287346421&amp;alt=media)| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F81402569a9b6bdbe387e6b9acca955a8%2Fdownload%20(10).png?generation=1584072309888310&amp;alt=media)\n |\n\nTo me, both of these look real, but possibly just frame shifted. It may be the case that a frame was dropped in one of the videos or their indexes just werent aligned. I wouldnt put too much weight in the results of this video. \n\nLooking at the next highest error false-negative (zmtiukzllb.mp4) we see a very different profile\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F9b3687d7d286e9f905e37898094d4c67%2Fdownload%20(11).png?generation=1584072490014540&amp;alt=media)\n\n| Real | Fake |\n| --- | --- |\n| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Fc13f55bf8cb060c633e0c45561eafa95%2Fdownload%20(12).png?generation=1584072537488215&amp;alt=media)| ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F109760847e8bca452ef56ea2810439d2%2Fdownload%20(13).png?generation=1584072558813406&amp;alt=media)\n |\n\nThis plot and these images make much more sense now. The frames look to largely be matched, up to a reasonable noise level that would be seen from mp4 compression and the fake and real videos look aligned, albeit the fake face has been destroyed by some blurring. \n\nLooking back through the inputs the model saw I can't really say why the model viewed these as possibly real. The image is obviously blurred that is reaching the model even after face detection. It may have to do with the model seeing many regular looking inputs, the ones that had small differences and only seeing a small set that appeared fake. In the difference plot we can see it seems like the fake maybe only locks on in some frames and not all. \n\n##Meta Analysis\nNext thing we can look at is view the small subset of videos the model is getting wrong and look for commonalities by hand labeling them with various characteristics we might deem relevant. \n\nI chose some simple distinctions I thought might be useful to look into. Specifically I was checking for multiple faces, if there were side-on faces, if the face in the video was moving, if the camera was moving, if the subject was black, if it was shot in low-light and if the face is flickering on and off and if they had what I deemed \"ghosts\" typically a swapped in face that would move around either between faces or just rest in open space. \n\nLooking at the first 50 of the false positives I labeled them with the following characteristics\n\n| Side-on | Moving subject | Moving camera | Black | Low light | flickering | Ghost |\n|---------|----------------|---------------|-------|-----------|------------|-------|\n| 44      | 7              | 7             | 7     | 9         | 29         | 11    |\n\nLooks like there are some more common occurrences than others. These numbers arent necessarily significant until they are compared against the false negative and baseline reals and fakes to determine if the issue is they are coming from a slightly different distribution. \n\nThat will come at a later date.",
    "770807": "Thanks! Very good analysis :)"
  },
  "source": "meta"
}