{
  "id": 140303,
  "title": "Public 34th/Private 47th place insights - Part1 (Inference)",
  "url": "/competitions/deepfake-detection-challenge/writeups/tom-mpware-public-34th-private-47th-place-insights",
  "author_name": "",
  "post_date": "2020-04-27T19:27:30.940Z",
  "votes": 41,
  "comment_count": 15,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>There is a lot to say about this competition. First I would like to share about our inference pipeline here and then I will make another post about the training part. Our solution uses only videos/frames (no audio at all) and total frames used in inference was quite important for us, the more the better.  We tried to move frames into GPU early in the pipeline to benefit from GPU fast computation in the next stages. We've fighted with NVidia DALI and Decord as they can load/decode videos through GPU, unfortunately it never worked, Decord had a memory leak and DALI finished with timeout during submission (impossible to troubleshoot). So we used OpenCV and moved decoded frames in GPU just after. </p>\n\n<p>Then, for <strong>faces extraction</strong>, we've used MTCNN with facenet-pytorch with small changes to make it support GPU input tensors. Prior to MTCNN, we added a GAMMA correction (still in GPU) on frames to help MTCNN (some dark videos or bad contrast). Outputs are faces resized to 256x256, aspect ratio preserved and large margin (30 pixels on each border).</p>\n\n<p>Next steps is <strong>face tracking</strong> to identify how many faces in a videos, cleanup artifacts (face detected but not face) with different rules (based on tracked confidences and maximum faces). It is based on centroid boxes tracked across even spaced frames.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F682f224bba7f3b69371c8903ac717158%2Fpipeline_inference.png?generation=1585728553095780&amp;alt=media\" alt=\"\"></p>\n\n<p>Next is <strong>models inference</strong>. We've 3 to 4 EfficientNet CV5 models with different input shapes (256x256, 240x240, 224x224) cropped by the normalizers. Each model was trained with a different validation strategy. Some with full data, some with partial data to make hold-out validation.</p>\n\n<p>Final step is <strong>post-processing</strong>. One may notice that some videos have blinking fake faces (a few frames real and a few frames fake within the same video). Probabilities of our models was moving up and down which means it worked but using simple average for final probablity would not work:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F58a2b09dc3e876009533a523b95eb571%2Ffake_frames.png?generation=1585729302156617&amp;alt=media\" alt=\"\"></p>\n\n<p>So we decided to evaluate what could be the thresholds to detect such behavior and use maximum probability instead of average in this case only. The question was how many frames with high probability should we have to consider it's a fake:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fae7596e7d77a9f827e375d39a0ebedcb%2Ftheshold.png?generation=1585729646472824&amp;alt=media\" alt=\"\"></p>\n\n<p>And the answer was around 20%. So if 20% of frames have probability higher than around 0.85 then we prefer selecting maximum probability instead of average.</p>\n\n<p>There is an additional/optional step with ridge regression (not classification) built on hold-out data and applied to models' output.</p>\n\n<p>This pipeline works only if we have enough frames and we were able to run it up to 100 frames per video. Each single EFB model got around LB=0.32. Ridge regression on ensemble got LB=0.29 and thresholds provided boost to LB=0.27.</p>",
  "messages": [
    {
      "id": "793743",
      "postDate": "04/01/2020 08:40:54",
      "content": "<p>Hi all,</p>\n\n<p>There is a lot to say about this competition. First I would like to share about our inference pipeline here and then I will make another post about the training part. Our solution uses only videos/frames (no audio at all) and total frames used in inference was quite important for us, the more the better.  We tried to move frames into GPU early in the pipeline to benefit from GPU fast computation in the next stages. We've fighted with NVidia DALI and Decord as they can load/decode videos through GPU, unfortunately it never worked, Decord had a memory leak and DALI finished with timeout during submission (impossible to troubleshoot). So we used OpenCV and moved decoded frames in GPU just after. </p>\n\n<p>Then, for <strong>faces extraction</strong>, we've used MTCNN with facenet-pytorch with small changes to make it support GPU input tensors. Prior to MTCNN, we added a GAMMA correction (still in GPU) on frames to help MTCNN (some dark videos or bad contrast). Outputs are faces resized to 256x256, aspect ratio preserved and large margin (30 pixels on each border).</p>\n\n<p>Next steps is <strong>face tracking</strong> to identify how many faces in a videos, cleanup artifacts (face detected but not face) with different rules (based on tracked confidences and maximum faces). It is based on centroid boxes tracked across even spaced frames.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F682f224bba7f3b69371c8903ac717158%2Fpipeline_inference.png?generation=1585728553095780&amp;alt=media\" alt=\"\"></p>\n\n<p>Next is <strong>models inference</strong>. We've 3 to 4 EfficientNet CV5 models with different input shapes (256x256, 240x240, 224x224) cropped by the normalizers. Each model was trained with a different validation strategy. Some with full data, some with partial data to make hold-out validation.</p>\n\n<p>Final step is <strong>post-processing</strong>. One may notice that some videos have blinking fake faces (a few frames real and a few frames fake within the same video). Probabilities of our models was moving up and down which means it worked but using simple average for final probablity would not work:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F58a2b09dc3e876009533a523b95eb571%2Ffake_frames.png?generation=1585729302156617&amp;alt=media\" alt=\"\"></p>\n\n<p>So we decided to evaluate what could be the thresholds to detect such behavior and use maximum probability instead of average in this case only. The question was how many frames with high probability should we have to consider it's a fake:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fae7596e7d77a9f827e375d39a0ebedcb%2Ftheshold.png?generation=1585729646472824&amp;alt=media\" alt=\"\"></p>\n\n<p>And the answer was around 20%. So if 20% of frames have probability higher than around 0.85 then we prefer selecting maximum probability instead of average.</p>\n\n<p>There is an additional/optional step with ridge regression (not classification) built on hold-out data and applied to models' output.</p>\n\n<p>This pipeline works only if we have enough frames and we were able to run it up to 100 frames per video. Each single EFB model got around LB=0.32. Ridge regression on ensemble got LB=0.29 and thresholds provided boost to LB=0.27.</p>",
      "rawMarkdown": "Hi all,\n\nThere is a lot to say about this competition. First I would like to share about our inference pipeline here and then I will make another post about the training part. Our solution uses only videos/frames (no audio at all) and total frames used in inference was quite important for us, the more the better.  We tried to move frames into GPU early in the pipeline to benefit from GPU fast computation in the next stages. We've fighted with NVidia DALI and Decord as they can load/decode videos through GPU, unfortunately it never worked, Decord had a memory leak and DALI finished with timeout during submission (impossible to troubleshoot). So we used OpenCV and moved decoded frames in GPU just after. \n\nThen, for **faces extraction**, we've used MTCNN with facenet-pytorch with small changes to make it support GPU input tensors. Prior to MTCNN, we added a GAMMA correction (still in GPU) on frames to help MTCNN (some dark videos or bad contrast). Outputs are faces resized to 256x256, aspect ratio preserved and large margin (30 pixels on each border).\n\nNext steps is **face tracking** to identify how many faces in a videos, cleanup artifacts (face detected but not face) with different rules (based on tracked confidences and maximum faces). It is based on centroid boxes tracked across even spaced frames.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F682f224bba7f3b69371c8903ac717158%2Fpipeline_inference.png?generation=1585728553095780&amp;alt=media)\n\nNext is **models inference**. We've 3 to 4 EfficientNet CV5 models with different input shapes (256x256, 240x240, 224x224) cropped by the normalizers. Each model was trained with a different validation strategy. Some with full data, some with partial data to make hold-out validation.\n\nFinal step is **post-processing**. One may notice that some videos have blinking fake faces (a few frames real and a few frames fake within the same video). Probabilities of our models was moving up and down which means it worked but using simple average for final probablity would not work:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F58a2b09dc3e876009533a523b95eb571%2Ffake_frames.png?generation=1585729302156617&amp;alt=media)\n\nSo we decided to evaluate what could be the thresholds to detect such behavior and use maximum probability instead of average in this case only. The question was how many frames with high probability should we have to consider it's a fake:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fae7596e7d77a9f827e375d39a0ebedcb%2Ftheshold.png?generation=1585729646472824&amp;alt=media)\n\nAnd the answer was around 20%. So if 20% of frames have probability higher than around 0.85 then we prefer selecting maximum probability instead of average.\n\nThere is an additional/optional step with ridge regression (not classification) built on hold-out data and applied to models' output.\n\nThis pipeline works only if we have enough frames and we were able to run it up to 100 frames per video. Each single EFB model got around LB=0.32. Ridge regression on ensemble got LB=0.29 and thresholds provided boost to LB=0.27.",
      "votes": null
    },
    {
      "id": "793808",
      "postDate": "04/01/2020 09:53:34",
      "content": "<p>Interesting approach. Congrats!</p>",
      "rawMarkdown": "Interesting approach. Congrats!",
      "votes": null
    },
    {
      "id": "793849",
      "postDate": "04/01/2020 10:49:57",
      "content": "<p>I would like to add that:\n- TTA was not used as it was more important to have various frames than same frames with TTA. Score with TTA was just around +0.001. \n- Floating Point 16 was used to speed up inference. It did not modify probabilities (minor).</p>",
      "rawMarkdown": "I would like to add that:\n- TTA was not used as it was more important to have various frames than same frames with TTA. Score with TTA was just around +0.001. \n- Floating Point 16 was used to speed up inference. It did not modify probabilities (minor).",
      "votes": null
    },
    {
      "id": "793897",
      "postDate": "04/01/2020 11:38:54",
      "content": "<p>I used the same methodology with the 20% but instead of maximum i averaged only the predictions that were in this direction (ignoring the opposite predictions) . I wonder if this would work better or worse than maximum. </p>",
      "rawMarkdown": "I used the same methodology with the 20% but instead of maximum i averaged only the predictions that were in this direction (ignoring the opposite predictions) . I wonder if this would work better or worse than maximum.",
      "votes": null
    },
    {
      "id": "793898",
      "postDate": "04/01/2020 11:43:08",
      "content": "<p>It's a good idea to use all frames and threshold out to determine fake videos.\nI tried in similar way with 25 frames per video just to improve 0.001 of loss.\nI'm curious this method will still work for private dataset with same theshold .85\nThank you for sharing!</p>",
      "rawMarkdown": "It's a good idea to use all frames and threshold out to determine fake videos.\nI tried in similar way with 25 frames per video just to improve 0.001 of loss.\nI'm curious this method will still work for private dataset with same theshold .85\nThank you for sharing!",
      "votes": null
    },
    {
      "id": "793900",
      "postDate": "04/01/2020 11:44:55",
      "content": "<p><a href=\"/moshel\">@moshel</a> Average of probabilities &gt; 0.85 if more than 20%, correct? It should give similar results.</p>",
      "rawMarkdown": "moshel Average of probabilities &gt; 0.85 if more than 20%, correct? It should give similar results.",
      "votes": null
    },
    {
      "id": "793905",
      "postDate": "04/01/2020 11:51:37",
      "content": "<p>We've tried 0.7 to 0.9 range and variation was around 0.002. But true if private set is totally different then it could fallback to average of probabilities and then -0.02</p>",
      "rawMarkdown": "We've tried 0.7 to 0.9 range and variation was around 0.002. But true if private set is totally different then it could fallback to average of probabilities and then -0.02",
      "votes": null
    },
    {
      "id": "793906",
      "postDate": "04/01/2020 11:52:08",
      "content": "<p>Thank you for sharing! Nice postprocessing technique!</p>",
      "rawMarkdown": "Thank you for sharing! Nice postprocessing technique!",
      "votes": null
    },
    {
      "id": "793907",
      "postDate": "04/01/2020 11:53:04",
      "content": "<p>Nice way to calibrate threshold. I noticed the blinking as well but didn't account for it. The challenge with calibrating for a particular scenario is will we see a similar distribution in the private set. We will find out in 22 days :)</p>",
      "rawMarkdown": "Nice way to calibrate threshold. I noticed the blinking as well but didn't account for it. The challenge with calibrating for a particular scenario is will we see a similar distribution in the private set. We will find out in 22 days :)",
      "votes": null
    },
    {
      "id": "794946",
      "postDate": "04/02/2020 08:39:15",
      "content": "<p>34th after LB cleanup 😏 . </p>",
      "rawMarkdown": "34th after LB cleanup 😏 .",
      "votes": null
    },
    {
      "id": "794982",
      "postDate": "04/02/2020 09:08:22",
      "content": "<p>Nice approach <a href=\"/mpware\">@mpware</a> Thanks for sharing!\nI'm struggling to understand your last graph, tho. \nCould you please tell us what are the 2 axis and T lines?\nThank you again</p>",
      "rawMarkdown": "Nice approach @mpware Thanks for sharing!\nI'm struggling to understand your last graph, tho. \nCould you please tell us what are the 2 axis and T lines?\nThank you again",
      "votes": null
    },
    {
      "id": "795058",
      "postDate": "04/02/2020 10:58:36",
      "content": "<p>X = frames ratio threshold: 10%, 20%, 30% of frames per video with high probability.\nY = log loss.\nT = different values for high probability (0.7, .0.8, 0.9 ...)</p>\n\n<p>Dataset to build this graph is our hold-out data.</p>",
      "rawMarkdown": "X = frames ratio threshold: 10%, 20%, 30% of frames per video with high probability.\nY = log loss.\nT = different values for high probability (0.7, .0.8, 0.9 ...)\n\nDataset to build this graph is our hold-out data.",
      "votes": null
    },
    {
      "id": "795192",
      "postDate": "04/02/2020 13:53:31",
      "content": "<p>ah nice, if I got it right it's something like below?\n<code>\nif preds.quantile(0.8) &amp;gt; 0.85:  # top20%\n    return preds.max()\nelse:\n    return preds.mean()\n</code>\nHas that also balanced the fake/real ratio in your hold out?\nFrom your graph it looks like the top5% / <code>quantile(0.95) &amp;gt; 0.85</code> would give you a better log-loss? Did you get the top20% from the LB?</p>\n\n<p>I tried to use preds.quantile(0.68) directly but your idea is probably better!\nAnother idea was to fit a generalized mean, but we run out of time and submissions... </p>\n\n<p>Thanks for the insights <a href=\"/mpware\">@mpware</a> !</p>",
      "rawMarkdown": "ah nice, if I got it right it's something like below?\n```\nif preds.quantile(0.8) &gt; 0.85:  # top20%\n    return preds.max()\nelse:\n    return preds.mean()\n```\nHas that also balanced the fake/real ratio in your hold out?\nFrom your graph it looks like the top5% / `quantile(0.95) &gt; 0.85` would give you a better log-loss? Did you get the top20% from the LB?\n\nI tried to use preds.quantile(0.68) directly but your idea is probably better!\nAnother idea was to fit a generalized mean, but we run out of time and submissions... \n\nThanks for the insights @mpware !",
      "votes": null
    },
    {
      "id": "795250",
      "postDate": "04/02/2020 14:54:39",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a>  Our implementation is:\n```</p>\n\n<h1>A few frames detected as fake then it's fake video</h1>\n\n<h1>Probability of each frame (mean of each model), around 65 frames</h1>\n\n<p>filename_prob = np.nanmean(clf_predict_probas, axis=1)  # (65 rows)</p>\n\n<h1>Number for frames above prob 0.85</h1>\n\n<p>T = 0.85\nTH = np.sum(filename_prob &gt; T)/len(filename_prob) <br>\nif TH &gt; 0.20:\n    filename_prob = np.nanmax(filename_prob)\nelse:\n    filename_prob = np.nanmean(filename_prob)\n```\nYes, our HO data is balanced 1:1. In the graph you can notice that any values between 0.05 and 0.25 should work. On LB we've tried TH in [0.20, 0.22] with T in [0.7, 0.8, 0.85] and it provided around +0.02 boost whatever the ensemble. We've also tried TH=0.02 (blind test prior working on HO data) and it was not good -0.03.</p>\n\n<p>We also have another threshold that I've not described because benefit was very low: If 100% of frames are below 0.5 and 90% of frames are below 0.2 than use min(probabilities).</p>\n\n<p>Finally, if video has 2 faces detected then we use the highest face probability because one (or more) face could be fake and the other not. </p>",
      "rawMarkdown": "hmendonca  Our implementation is:\n```\n# A few frames detected as fake then it's fake video\n# Probability of each frame (mean of each model), around 65 frames\nfilename_prob = np.nanmean(clf_predict_probas, axis=1)  # (65 rows)\n\n# Number for frames above prob 0.85\nT = 0.85\nTH = np.sum(filename_prob &gt; T)/len(filename_prob)       \nif TH &gt; 0.20:\n\tfilename_prob = np.nanmax(filename_prob)\nelse:\n\tfilename_prob = np.nanmean(filename_prob)\n```\nYes, our HO data is balanced 1:1. In the graph you can notice that any values between 0.05 and 0.25 should work. On LB we've tried TH in [0.20, 0.22] with T in [0.7, 0.8, 0.85] and it provided around +0.02 boost whatever the ensemble. We've also tried TH=0.02 (blind test prior working on HO data) and it was not good -0.03.\n\nWe also have another threshold that I've not described because benefit was very low: If 100% of frames are below 0.5 and 90% of frames are below 0.2 than use min(probabilities).\n\nFinally, if video has 2 faces detected then we use the highest face probability because one (or more) face could be fake and the other not.",
      "votes": null
    },
    {
      "id": "819140",
      "postDate": "04/24/2020 11:37:14",
      "content": "<p>49th on private. Not so bad, we survived the shake up!</p>\n\n<p><a href=\"/philculliton\">@philculliton</a>\nBut only one of our submissions had been evaluated. The other failed and we don't know why (we didn't remove any dataset and it was running fine for public LB). It would be great if Kaggle/Organizers would allow us (and all other teams with the same issue - around 6.5%) to make it evaluated.</p>\n\n<p>First they should provide the underlying reason/exception that lead to not have the submission.csv.\nSecond, we could instruct them to try minor changes in the code (e.g. number of workers or total frames) to have it run successfully. And as soon as it works, it's considered as a result, no fine tune allowed.</p>\n\n<p>I'm sure each evaluation failure is due to a minor issue such one video bigger than expected that causes Out Of Memory (RAM or GPU).</p>",
      "rawMarkdown": "49th on private. Not so bad, we survived the shake up!\n\n@philculliton\nBut only one of our submissions had been evaluated. The other failed and we don't know why (we didn't remove any dataset and it was running fine for public LB). It would be great if Kaggle/Organizers would allow us (and all other teams with the same issue - around 6.5%) to make it evaluated.\n\nFirst they should provide the underlying reason/exception that lead to not have the submission.csv.\nSecond, we could instruct them to try minor changes in the code (e.g. number of workers or total frames) to have it run successfully. And as soon as it works, it's considered as a result, no fine tune allowed.\n\nI'm sure each evaluation failure is due to a minor issue such one video bigger than expected that causes Out Of Memory (RAM or GPU).",
      "votes": null
    },
    {
      "id": "822315",
      "postDate": "04/26/2020 20:24:06",
      "content": "<p>Thanks for sharing and congrats on the medal +1</p>",
      "rawMarkdown": "Thanks for sharing and congrats on the medal +1",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 793808,
      "author_name": "xro7nis",
      "author_url": "",
      "post_date": "04/01/2020 09:53:34",
      "content": "<p>Interesting approach. Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 793849,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "04/01/2020 10:49:57",
      "content": "<p>I would like to add that:\n- TTA was not used as it was more important to have various frames than same frames with TTA. Score with TTA was just around +0.001. \n- Floating Point 16 was used to speed up inference. It did not modify probabilities (minor).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 793897,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "04/01/2020 11:38:54",
      "content": "<p>I used the same methodology with the 20% but instead of maximum i averaged only the predictions that were in this direction (ignoring the opposite predictions) . I wonder if this would work better or worse than maximum. </p>",
      "votes": null,
      "replies": [
        {
          "id": 793900,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "04/01/2020 11:44:55",
          "content": "<p><a href=\"/moshel\">@moshel</a> Average of probabilities &gt; 0.85 if more than 20%, correct? It should give similar results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 793898,
      "author_name": "gwsong",
      "author_url": "",
      "post_date": "04/01/2020 11:43:08",
      "content": "<p>It's a good idea to use all frames and threshold out to determine fake videos.\nI tried in similar way with 25 frames per video just to improve 0.001 of loss.\nI'm curious this method will still work for private dataset with same theshold .85\nThank you for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 793905,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "04/01/2020 11:51:37",
          "content": "<p>We've tried 0.7 to 0.9 range and variation was around 0.002. But true if private set is totally different then it could fallback to average of probabilities and then -0.02</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 793906,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "04/01/2020 11:52:08",
      "content": "<p>Thank you for sharing! Nice postprocessing technique!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 793907,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "04/01/2020 11:53:04",
      "content": "<p>Nice way to calibrate threshold. I noticed the blinking as well but didn't account for it. The challenge with calibrating for a particular scenario is will we see a similar distribution in the private set. We will find out in 22 days :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 794946,
      "author_name": "mpware",
      "author_url": "",
      "post_date": "04/02/2020 08:39:15",
      "content": "<p>34th after LB cleanup 😏 . </p>",
      "votes": null,
      "replies": [
        {
          "id": 819140,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "04/24/2020 11:37:14",
          "content": "<p>49th on private. Not so bad, we survived the shake up!</p>\n\n<p><a href=\"/philculliton\">@philculliton</a>\nBut only one of our submissions had been evaluated. The other failed and we don't know why (we didn't remove any dataset and it was running fine for public LB). It would be great if Kaggle/Organizers would allow us (and all other teams with the same issue - around 6.5%) to make it evaluated.</p>\n\n<p>First they should provide the underlying reason/exception that lead to not have the submission.csv.\nSecond, we could instruct them to try minor changes in the code (e.g. number of workers or total frames) to have it run successfully. And as soon as it works, it's considered as a result, no fine tune allowed.</p>\n\n<p>I'm sure each evaluation failure is due to a minor issue such one video bigger than expected that causes Out Of Memory (RAM or GPU).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 794982,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "04/02/2020 09:08:22",
      "content": "<p>Nice approach <a href=\"/mpware\">@mpware</a> Thanks for sharing!\nI'm struggling to understand your last graph, tho. \nCould you please tell us what are the 2 axis and T lines?\nThank you again</p>",
      "votes": null,
      "replies": [
        {
          "id": 795058,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "04/02/2020 10:58:36",
          "content": "<p>X = frames ratio threshold: 10%, 20%, 30% of frames per video with high probability.\nY = log loss.\nT = different values for high probability (0.7, .0.8, 0.9 ...)</p>\n\n<p>Dataset to build this graph is our hold-out data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 795192,
          "author_name": "hmendonca",
          "author_url": "",
          "post_date": "04/02/2020 13:53:31",
          "content": "<p>ah nice, if I got it right it's something like below?\n<code>\nif preds.quantile(0.8) &amp;gt; 0.85:  # top20%\n    return preds.max()\nelse:\n    return preds.mean()\n</code>\nHas that also balanced the fake/real ratio in your hold out?\nFrom your graph it looks like the top5% / <code>quantile(0.95) &amp;gt; 0.85</code> would give you a better log-loss? Did you get the top20% from the LB?</p>\n\n<p>I tried to use preds.quantile(0.68) directly but your idea is probably better!\nAnother idea was to fit a generalized mean, but we run out of time and submissions... </p>\n\n<p>Thanks for the insights <a href=\"/mpware\">@mpware</a> !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 795250,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "04/02/2020 14:54:39",
          "content": "<p><a href=\"/hmendonca\">@hmendonca</a>  Our implementation is:\n```</p>\n\n<h1>A few frames detected as fake then it's fake video</h1>\n\n<h1>Probability of each frame (mean of each model), around 65 frames</h1>\n\n<p>filename_prob = np.nanmean(clf_predict_probas, axis=1)  # (65 rows)</p>\n\n<h1>Number for frames above prob 0.85</h1>\n\n<p>T = 0.85\nTH = np.sum(filename_prob &gt; T)/len(filename_prob) <br>\nif TH &gt; 0.20:\n    filename_prob = np.nanmax(filename_prob)\nelse:\n    filename_prob = np.nanmean(filename_prob)\n```\nYes, our HO data is balanced 1:1. In the graph you can notice that any values between 0.05 and 0.25 should work. On LB we've tried TH in [0.20, 0.22] with T in [0.7, 0.8, 0.85] and it provided around +0.02 boost whatever the ensemble. We've also tried TH=0.02 (blind test prior working on HO data) and it was not good -0.03.</p>\n\n<p>We also have another threshold that I've not described because benefit was very low: If 100% of frames are below 0.5 and 90% of frames are below 0.2 than use min(probabilities).</p>\n\n<p>Finally, if video has 2 faces detected then we use the highest face probability because one (or more) face could be fake and the other not. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 822315,
      "author_name": "",
      "author_url": "",
      "post_date": "04/26/2020 20:24:06",
      "content": "<p>Thanks for sharing and congrats on the medal +1</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "793743": "Hi all,\n\nThere is a lot to say about this competition. First I would like to share about our inference pipeline here and then I will make another post about the training part. Our solution uses only videos/frames (no audio at all) and total frames used in inference was quite important for us, the more the better.  We tried to move frames into GPU early in the pipeline to benefit from GPU fast computation in the next stages. We've fighted with NVidia DALI and Decord as they can load/decode videos through GPU, unfortunately it never worked, Decord had a memory leak and DALI finished with timeout during submission (impossible to troubleshoot). So we used OpenCV and moved decoded frames in GPU just after. \n\nThen, for **faces extraction**, we've used MTCNN with facenet-pytorch with small changes to make it support GPU input tensors. Prior to MTCNN, we added a GAMMA correction (still in GPU) on frames to help MTCNN (some dark videos or bad contrast). Outputs are faces resized to 256x256, aspect ratio preserved and large margin (30 pixels on each border).\n\nNext steps is **face tracking** to identify how many faces in a videos, cleanup artifacts (face detected but not face) with different rules (based on tracked confidences and maximum faces). It is based on centroid boxes tracked across even spaced frames.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F682f224bba7f3b69371c8903ac717158%2Fpipeline_inference.png?generation=1585728553095780&amp;alt=media)\n\nNext is **models inference**. We've 3 to 4 EfficientNet CV5 models with different input shapes (256x256, 240x240, 224x224) cropped by the normalizers. Each model was trained with a different validation strategy. Some with full data, some with partial data to make hold-out validation.\n\nFinal step is **post-processing**. One may notice that some videos have blinking fake faces (a few frames real and a few frames fake within the same video). Probabilities of our models was moving up and down which means it worked but using simple average for final probablity would not work:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2F58a2b09dc3e876009533a523b95eb571%2Ffake_frames.png?generation=1585729302156617&amp;alt=media)\n\nSo we decided to evaluate what could be the thresholds to detect such behavior and use maximum probability instead of average in this case only. The question was how many frames with high probability should we have to consider it's a fake:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F698363%2Fae7596e7d77a9f827e375d39a0ebedcb%2Ftheshold.png?generation=1585729646472824&amp;alt=media)\n\nAnd the answer was around 20%. So if 20% of frames have probability higher than around 0.85 then we prefer selecting maximum probability instead of average.\n\nThere is an additional/optional step with ridge regression (not classification) built on hold-out data and applied to models' output.\n\nThis pipeline works only if we have enough frames and we were able to run it up to 100 frames per video. Each single EFB model got around LB=0.32. Ridge regression on ensemble got LB=0.29 and thresholds provided boost to LB=0.27.",
    "793808": "Interesting approach. Congrats!",
    "793849": "I would like to add that:\n- TTA was not used as it was more important to have various frames than same frames with TTA. Score with TTA was just around +0.001. \n- Floating Point 16 was used to speed up inference. It did not modify probabilities (minor).",
    "793897": "I used the same methodology with the 20% but instead of maximum i averaged only the predictions that were in this direction (ignoring the opposite predictions) . I wonder if this would work better or worse than maximum.",
    "793898": "It's a good idea to use all frames and threshold out to determine fake videos.\nI tried in similar way with 25 frames per video just to improve 0.001 of loss.\nI'm curious this method will still work for private dataset with same theshold .85\nThank you for sharing!",
    "793900": "moshel Average of probabilities &gt; 0.85 if more than 20%, correct? It should give similar results.",
    "793905": "We've tried 0.7 to 0.9 range and variation was around 0.002. But true if private set is totally different then it could fallback to average of probabilities and then -0.02",
    "793906": "Thank you for sharing! Nice postprocessing technique!",
    "793907": "Nice way to calibrate threshold. I noticed the blinking as well but didn't account for it. The challenge with calibrating for a particular scenario is will we see a similar distribution in the private set. We will find out in 22 days :)",
    "794946": "34th after LB cleanup 😏 .",
    "794982": "Nice approach @mpware Thanks for sharing!\nI'm struggling to understand your last graph, tho. \nCould you please tell us what are the 2 axis and T lines?\nThank you again",
    "795058": "X = frames ratio threshold: 10%, 20%, 30% of frames per video with high probability.\nY = log loss.\nT = different values for high probability (0.7, .0.8, 0.9 ...)\n\nDataset to build this graph is our hold-out data.",
    "795192": "ah nice, if I got it right it's something like below?\n```\nif preds.quantile(0.8) &gt; 0.85:  # top20%\n    return preds.max()\nelse:\n    return preds.mean()\n```\nHas that also balanced the fake/real ratio in your hold out?\nFrom your graph it looks like the top5% / `quantile(0.95) &gt; 0.85` would give you a better log-loss? Did you get the top20% from the LB?\n\nI tried to use preds.quantile(0.68) directly but your idea is probably better!\nAnother idea was to fit a generalized mean, but we run out of time and submissions... \n\nThanks for the insights @mpware !",
    "795250": "hmendonca  Our implementation is:\n```\n# A few frames detected as fake then it's fake video\n# Probability of each frame (mean of each model), around 65 frames\nfilename_prob = np.nanmean(clf_predict_probas, axis=1)  # (65 rows)\n\n# Number for frames above prob 0.85\nT = 0.85\nTH = np.sum(filename_prob &gt; T)/len(filename_prob)       \nif TH &gt; 0.20:\n\tfilename_prob = np.nanmax(filename_prob)\nelse:\n\tfilename_prob = np.nanmean(filename_prob)\n```\nYes, our HO data is balanced 1:1. In the graph you can notice that any values between 0.05 and 0.25 should work. On LB we've tried TH in [0.20, 0.22] with T in [0.7, 0.8, 0.85] and it provided around +0.02 boost whatever the ensemble. We've also tried TH=0.02 (blind test prior working on HO data) and it was not good -0.03.\n\nWe also have another threshold that I've not described because benefit was very low: If 100% of frames are below 0.5 and 90% of frames are below 0.2 than use min(probabilities).\n\nFinally, if video has 2 faces detected then we use the highest face probability because one (or more) face could be fake and the other not.",
    "819140": "49th on private. Not so bad, we survived the shake up!\n\n@philculliton\nBut only one of our submissions had been evaluated. The other failed and we don't know why (we didn't remove any dataset and it was running fine for public LB). It would be great if Kaggle/Organizers would allow us (and all other teams with the same issue - around 6.5%) to make it evaluated.\n\nFirst they should provide the underlying reason/exception that lead to not have the submission.csv.\nSecond, we could instruct them to try minor changes in the code (e.g. number of workers or total frames) to have it run successfully. And as soon as it works, it's considered as a result, no fine tune allowed.\n\nI'm sure each evaluation failure is due to a minor issue such one video bigger than expected that causes Out Of Memory (RAM or GPU).",
    "822315": "Thanks for sharing and congrats on the medal +1"
  },
  "source": "meta"
}