{
  "id": 125044,
  "title": "Let's talk about Preprocessing",
  "url": "/competitions/deepfake-detection-challenge/discussion/125044",
  "author_name": "Bibek",
  "post_date": "2020-01-08T08:59:27.864000",
  "votes": 14,
  "comment_count": 46,
  "views": 0,
  "content": "<p>I'm thinking about implementing this pipeline from <a href=\"https://arxiv.org/pdf/1910.12467.pdf\">Capsule Network</a> paper for this competition. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F79f2d5d9866d9b72a3a2fa93e053978c%2Fpipeline.png?generation=1578473585651924&amp;alt=media\" alt=\"\"></p>\n\n<p>According to paper, preprocessing(in case of videos) includes separating the frames and extracting the face; this is what I plan to do but there are many variables to consider like \n1. number of frames\n2. should preprocessing be done beforehand or on-the-fly?\n3. how to tackle face-extractor(eg: mtcnn) overload?</p>\n\n<p>I wonder how other Kagglers are approaching these problems?</p>",
  "messages": [
    {
      "id": 713405,
      "postDate": "2020-01-08T08:59:27.863Z",
      "content": "<p>I'm thinking about implementing this pipeline from <a href=\"https://arxiv.org/pdf/1910.12467.pdf\">Capsule Network</a> paper for this competition. \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F79f2d5d9866d9b72a3a2fa93e053978c%2Fpipeline.png?generation=1578473585651924&amp;alt=media\" alt=\"\"></p>\n\n<p>According to paper, preprocessing(in case of videos) includes separating the frames and extracting the face; this is what I plan to do but there are many variables to consider like \n1. number of frames\n2. should preprocessing be done beforehand or on-the-fly?\n3. how to tackle face-extractor(eg: mtcnn) overload?</p>\n\n<p>I wonder how other Kagglers are approaching these problems?</p>",
      "rawMarkdown": "I'm thinking about implementing this pipeline from [Capsule Network](https://arxiv.org/pdf/1910.12467.pdf) paper for this competition. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F79f2d5d9866d9b72a3a2fa93e053978c%2Fpipeline.png?generation=1578473585651924&amp;alt=media)\n\nAccording to paper, preprocessing(in case of videos) includes separating the frames and extracting the face; this is what I plan to do but there are many variables to consider like \n1. number of frames\n2. should preprocessing be done beforehand or on-the-fly?\n3. how to tackle face-extractor(eg: mtcnn) overload?\n\nI wonder how other Kagglers are approaching these problems?",
      "votes": 13
    },
    {
      "id": 714228,
      "postDate": "2020-01-09T07:54:28.510Z",
      "content": "<p>As the fake videos are much more than real ones, I extracted 2 frames from each FAKE and 10 frames from each REAL. And of course beforehand. It took me about two days.\nUsing other data augmentation methods(eg. HorizontalFlip), I believe faces from a series of single frames are enough.\nI think we could use other types of data and classifier, like time series and facial landmarks. But I don't have any time and GPU this month thanks to the Spring Festival.</p>\n\n<p>PS: I think the extracted faces should be processed again, for some of them are not really a face.</p>",
      "rawMarkdown": "As the fake videos are much more than real ones, I extracted 2 frames from each FAKE and 10 frames from each REAL. And of course beforehand. It took me about two days.\nUsing other data augmentation methods(eg. HorizontalFlip), I believe faces from a series of single frames are enough.\nI think we could use other types of data and classifier, like time series and facial landmarks. But I don't have any time and GPU this month thanks to the Spring Festival.\n\nPS: I think the extracted faces should be processed again, for some of them are not really a face.",
      "votes": 6,
      "replies": [
        {
          "id": 714458,
          "postDate": "2020-01-09T13:09:39.957Z",
          "content": "<p>Good idea to fix the class imbalance by extracting different # of frames depending on whether the video is fake or real... What do you mean by re-preprocess extracted faces? Isn't it the same as only considering valid faces the ones with probability higher than a threshold (i.e. &gt;90%)?</p>",
          "rawMarkdown": "Good idea to fix the class imbalance by extracting different # of frames depending on whether the video is fake or real... What do you mean by re-preprocess extracted faces? Isn't it the same as only considering valid faces the ones with probability higher than a threshold (i.e. &gt;90%)?"
        },
        {
          "id": 714495,
          "postDate": "2020-01-09T13:52:47.773Z",
          "content": "<p>I don't have any good idea by now. I use facenet-pytorch to do face crop, but it doesn't give a confidence for the output face(or perhaps I just haven't figure out how). </p>",
          "rawMarkdown": "I don't have any good idea by now. I use facenet-pytorch to do face crop, but it doesn't give a confidence for the output face(or perhaps I just haven't figure out how). \n"
        },
        {
          "id": 714517,
          "postDate": "2020-01-09T14:06:01.973Z",
          "content": "<p>I'm also using facenet-pytorch, and yes it returns probabilities! Just make sure to use <code>select_largest=False</code> when instantiating the face detector. Something like:\n<code>\nfrom facenet_pytorch import MTCNN\nmtcnn = MTCNN(keep_all=False, select_largest=False, device=device)\nboxes, probs = mtcnn.detect(pil_image)\n</code></p>",
          "rawMarkdown": "I'm also using facenet-pytorch, and yes it returns probabilities! Just make sure to use `select_largest=False` when instantiating the face detector. Something like:\n```\nfrom facenet_pytorch import MTCNN\nmtcnn = MTCNN(keep_all=False, select_largest=False, device=device)\nboxes, probs = mtcnn.detect(pil_image)\n```\n",
          "votes": 5
        },
        {
          "id": 714616,
          "postDate": "2020-01-09T15:43:38.120Z",
          "content": "<p>Thanks a lot, I would try it!</p>",
          "rawMarkdown": "Thanks a lot, I would try it!"
        },
        {
          "id": 720410,
          "postDate": "2020-01-16T11:54:16.203Z",
          "content": "<p>Hi <a href=\"/feifeizaici\">@feifeizaici</a> ,\nFor the past week I'm trying to implement your idea: use faces from all videos and fix class imbalance by extracting faces from real/fake videos in 5:1 proportion. However, I'm finding a hard time to avoid overfitting: my models end up with an unrealistic amazing performance, generalizing very poorly in the LB/test set. Taking a closer look on the internals, I see that the model ends up with a high bias towards classifying videos as fakes (2:1), a problem that I saw in my first models when I was not treating class imbalance. The dataset I generated (5 faces from every real video, to 1 face from every fake video, using all 119K videos) is balanced; however, the problem still occurs.\nSo far, my best models have been trained in an undersampled dataset; i.e., for every real video, I sample one and only one fake video. This approach solves class imbalance, however it ends up using only ~32% of the data. That's why I found your approach so great and decided to explore it: it solves class imbalance and enables us to use the whole dataset.\nDid you see this problem while developing your solution?\nAlso, I imagine the problem can be treated by using random augmentations (flips, crops, etc) - is that critical to avoid overfitting in your approach?\nI'd love to hear your perspective :)\nThanks!\nCarlos</p>",
          "rawMarkdown": "Hi @feifeizaici ,\nFor the past week I'm trying to implement your idea: use faces from all videos and fix class imbalance by extracting faces from real/fake videos in 5:1 proportion. However, I'm finding a hard time to avoid overfitting: my models end up with an unrealistic amazing performance, generalizing very poorly in the LB/test set. Taking a closer look on the internals, I see that the model ends up with a high bias towards classifying videos as fakes (2:1), a problem that I saw in my first models when I was not treating class imbalance. The dataset I generated (5 faces from every real video, to 1 face from every fake video, using all 119K videos) is balanced; however, the problem still occurs.\nSo far, my best models have been trained in an undersampled dataset; i.e., for every real video, I sample one and only one fake video. This approach solves class imbalance, however it ends up using only ~32% of the data. That's why I found your approach so great and decided to explore it: it solves class imbalance and enables us to use the whole dataset.\nDid you see this problem while developing your solution?\nAlso, I imagine the problem can be treated by using random augmentations (flips, crops, etc) - is that critical to avoid overfitting in your approach?\nI'd love to hear your perspective :)\nThanks!\nCarlos",
          "votes": 1
        },
        {
          "id": 720549,
          "postDate": "2020-01-16T14:16:35.023Z",
          "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a>\nNice to hear you are implementing my idea. My model is also a bit overfitting. When testing on my test set (including 2000 REAL and 2000 FAKE), I found the number of FN is much more than FP(it tends to predict more fakes). I have tried some methods but none of them worked well. And I'm still working on it. As for data augmentation, I did try some methods. Random-horizontal-flip improves the accuracy a bit, but random-clip doesn't. I‘m also trying other methods to solve this problem, such as modifying the loss function or the weights of classes.\nGood luck!</p>",
          "rawMarkdown": "Hi @carlossouza\nNice to hear you are implementing my idea. My model is also a bit overfitting. When testing on my test set (including 2000 REAL and 2000 FAKE), I found the number of FN is much more than FP(it tends to predict more fakes). I have tried some methods but none of them worked well. And I'm still working on it. As for data augmentation, I did try some methods. Random-horizontal-flip improves the accuracy a bit, but random-clip doesn't. I‘m also trying other methods to solve this problem, such as modifying the loss function or the weights of classes.\nGood luck!",
          "votes": 1
        },
        {
          "id": 721401,
          "postDate": "2020-01-17T10:47:00.650Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> first of all, thanks for sharing your thought and experiment process so far. </p>\n\n<p><a href=\"/feifeizaici\">@feifeizaici</a> also thanks for sharing your experience.</p>\n\n<p>I have benefited from this thread, so I feel obligated to share some of my ad-hoc experience with this dataset. </p>\n\n<ul>\n<li>I didn't do the careful sampling as <a href=\"/carlossouza\">@carlossouza</a>  did, neither did I do the pre-process according to the DFDC paper. </li>\n<li>I did follow the FaceForensic paper's approach of using running two-stages fine-tuning, i.e. 3 epochs with only top layers unfrozen, and then several epochs of unfrozen layers for all parameters, this was done on a dataset of extracted faces of about 40k images. This gave some good improvement compared to my previous training methods. Not as nice as the LB0.53 score you shared, mine was on about LB0.57</li>\n<li>However, the above approch doesn't work on my side, when trying to scale to larger datasets. I failed to generate any better model than LB0.57 with this approach for datasets with more than 300k images, and I have experimented with various well-known architectures including Xception. </li>\n<li>My current best single model was trained on a roughly balanced dataset of 300k images, with images from 3 groups of videos used as hold-out</li>\n<li>I have been using some run-of-the-mill image augmentation - horizontal flip, gaussian blur,  changing brigthness, etc. These are things that I have carried our from a previous project, so I might have to revisit their contribution to the modelling one by one in future experiments - in the unlikely event that I find time to do it :)</li>\n<li>Regarding overfitting, so far I don't have issue generalising from hold-out score to LB, hold-out score and LB are more or less 0.04 from each other either way.</li>\n</ul>\n\n<p>As far as modelling goes, so far the majority of my considering has been on choosing between base-architecture, add-on architecture configuration, fine-tuning methods, and learning rate related issue.  It is my first time to try hard on Video/Image competitions so I am flying blind in very long training cycle. </p>\n\n<p>I hope the above practical insights are useful to some of us. </p>",
          "rawMarkdown": "@carlossouza first of all, thanks for sharing your thought and experiment process so far. \n\n@feifeizaici also thanks for sharing your experience.\n\nI have benefited from this thread, so I feel obligated to share some of my ad-hoc experience with this dataset. \n\n- I didn't do the careful sampling as @carlossouza  did, neither did I do the pre-process according to the DFDC paper. \n- I did follow the FaceForensic paper's approach of using running two-stages fine-tuning, i.e. 3 epochs with only top layers unfrozen, and then several epochs of unfrozen layers for all parameters, this was done on a dataset of extracted faces of about 40k images. This gave some good improvement compared to my previous training methods. Not as nice as the LB0.53 score you shared, mine was on about LB0.57\n- However, the above approch doesn't work on my side, when trying to scale to larger datasets. I failed to generate any better model than LB0.57 with this approach for datasets with more than 300k images, and I have experimented with various well-known architectures including Xception. \n- My current best single model was trained on a roughly balanced dataset of 300k images, with images from 3 groups of videos used as hold-out\n- I have been using some run-of-the-mill image augmentation - horizontal flip, gaussian blur,  changing brigthness, etc. These are things that I have carried our from a previous project, so I might have to revisit their contribution to the modelling one by one in future experiments - in the unlikely event that I find time to do it :)\n- Regarding overfitting, so far I don't have issue generalising from hold-out score to LB, hold-out score and LB are more or less 0.04 from each other either way.\n\nAs far as modelling goes, so far the majority of my considering has been on choosing between base-architecture, add-on architecture configuration, fine-tuning methods, and learning rate related issue.  It is my first time to try hard on Video/Image competitions so I am flying blind in very long training cycle. \n \nI hope the above practical insights are useful to some of us. ",
          "votes": 9
        },
        {
          "id": 731391,
          "postDate": "2020-01-28T15:19:50.660Z",
          "content": "<p>Hi, <a href=\"/feifeizaici\">@feifeizaici</a>, thanks for your sharing,</p>\n\n<p>I only download the files two days ago, maybe im not familiar to dataset, sorry if im wrong.\nThe FAKE video also point to a REAL video, there only 20% REAL video which don't have FAKE video.\nIf we extract image from all videos, It means that ratio of FAKE : REAL = 4:5, not 4:1</p>\n\n<p>Don't know if its better to use all video, or only use video list in metadata.json?\nThanks,</p>",
          "rawMarkdown": "Hi, @feifeizaici, thanks for your sharing,\n\nI only download the files two days ago, maybe im not familiar to dataset, sorry if im wrong.\nThe FAKE video also point to a REAL video, there only 20% REAL video which don't have FAKE video.\nIf we extract image from all videos, It means that ratio of FAKE : REAL = 4:5, not 4:1\n\nDon't know if its better to use all video, or only use video list in metadata.json?\nThanks,"
        },
        {
          "id": 731415,
          "postDate": "2020-01-28T15:57:38.793Z",
          "content": "<p>Hi, <a href=\"/gody7334\">@gody7334</a> \nSome REAL videos was processed by different deepfake algorithms, and generated not only one FAKE videos. The FAKE/REAL ratio in the whole dataset is about 6:1, which is unbalaced.\nBest wishes!</p>",
          "rawMarkdown": "Hi, @gody7334 \nSome REAL videos was processed by different deepfake algorithms, and generated not only one FAKE videos. The FAKE/REAL ratio in the whole dataset is about 6:1, which is unbalaced.\nBest wishes!"
        },
        {
          "id": 732322,
          "postDate": "2020-01-29T17:22:00.210Z",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a>  I see, thanks for your reply, cheers.</p>",
          "rawMarkdown": "@feifeizaici  I see, thanks for your reply, cheers."
        }
      ]
    },
    {
      "id": 713640,
      "postDate": "2020-01-08T13:52:10.587Z",
      "content": "<p>On (2), I'd suggest doing it beforehand: it this way, you can reuse the frames to train new models, saving quite some time. Assuming your model can detect faces at 30 FPS, and that you will extract 30 faces/video, if you use 15,000 videos (~10% of the data), that would take 15,000 * 30 / 30 / 3,600 = 4 hours 10 min! And that's only to pre-process faces! Imagine spending that time/money every time you train a new classifier... :)</p>",
      "rawMarkdown": "On (2), I'd suggest doing it beforehand: it this way, you can reuse the frames to train new models, saving quite some time. Assuming your model can detect faces at 30 FPS, and that you will extract 30 faces/video, if you use 15,000 videos (~10% of the data), that would take 15,000 * 30 / 30 / 3,600 = 4 hours 10 min! And that's only to pre-process faces! Imagine spending that time/money every time you train a new classifier... :)",
      "votes": 1
    },
    {
      "id": 713562,
      "postDate": "2020-01-08T12:34:07.337Z",
      "content": "<ol>\n<li>Num Frames - Keep it simple lets say 30-50 frames are enough from 1 video.</li>\n<li>Preprocess after you extract face.</li>\n<li>What do you mean by overload, explain it in detail, if it is what I think it is, use multiple images for one time input say 1000 images each time you ask mtcnn for boxes to optimize it, I really do not recommend mtcnn though, it is slow and not the most accurate, still not gonna reveal what I am using yet.</li>\n</ol>",
      "rawMarkdown": "1. Num Frames - Keep it simple lets say 30-50 frames are enough from 1 video.\n2. Preprocess after you extract face.\n3. What do you mean by overload, explain it in detail, if it is what I think it is, use multiple images for one time input say 1000 images each time you ask mtcnn for boxes to optimize it, I really do not recommend mtcnn though, it is slow and not the most accurate, still not gonna reveal what I am using yet.",
      "votes": 2
    },
    {
      "id": 727074,
      "postDate": "2020-01-23T12:45:54.827Z",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> good to find you ..\nwe are new entrant to competition. COuld u tell what sort of preprocessing is needed.. is it extraction of faces using standard Face reco models like MTCNN and then using those images ?</p>",
      "rawMarkdown": "@bibek777 good to find you ..\nwe are new entrant to competition. COuld u tell what sort of preprocessing is needed.. is it extraction of faces using standard Face reco models like MTCNN and then using those images ?"
    },
    {
      "id": 722912,
      "postDate": "2020-01-19T09:46:53.557Z",
      "content": "<p>Excuse me. I'm confused about how to use facenet-pytorch in kaggle notebook. \"No custom packages enabled in your submission notebook\".</p>",
      "rawMarkdown": "Excuse me. I'm confused about how to use facenet-pytorch in kaggle notebook. \"No custom packages enabled in your submission notebook\"."
    },
    {
      "id": 718332,
      "postDate": "2020-01-14T09:50:42.040Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice"
    },
    {
      "id": 714653,
      "postDate": "2020-01-09T16:11:40.783Z",
      "content": "<p><a href=\"/carlossouza\">@carlossouza</a> and <a href=\"/feifeizaici\">@feifeizaici</a> might I ask what is the size of image dataset you extracted from the video to genreate  your model?  My models so far are trained on about 200k face-cropped images with different sampling proportion. Just trying to get an indication what is the best way to go forward, i.e. more images v.s. smarter way of building models - or both...  </p>",
      "rawMarkdown": "@carlossouza and @feifeizaici might I ask what is the size of image dataset you extracted from the video to genreate  your model?  My models so far are trained on about 200k face-cropped images with different sampling proportion. Just trying to get an indication what is the best way to go forward, i.e. more images v.s. smarter way of building models - or both...  ",
      "replies": [
        {
          "id": 714702,
          "postDate": "2020-01-09T17:03:50.270Z",
          "content": "<p>That's actually a very good question, one I'm trying to figure out right now.\nAll my experiments so far used 30 frames/video, and:\n- 1K videos =&gt; 30K faces\n- 5K videos =&gt; 150K faces\n- 15K videos =&gt; 450K faces</p>\n\n<p>However, my trainings usually last few epochs only (my early stopping criteria may be too easy to hit).\nAfter learning from others' experiences, will try less frames/video, more videos, and train for longer (harder early stopping criteria).\nMy next experiment will be 1 frame/video only, approx. 40K videos =&gt; 40K faces, but with a harder early stopping criteria, to train for 20+ epochs. Let's see... :)</p>",
          "rawMarkdown": "That's actually a very good question, one I'm trying to figure out right now.\nAll my experiments so far used 30 frames/video, and:\n- 1K videos =&gt; 30K faces\n- 5K videos =&gt; 150K faces\n- 15K videos =&gt; 450K faces\n\nHowever, my trainings usually last few epochs only (my early stopping criteria may be too easy to hit).\nAfter learning from others' experiences, will try less frames/video, more videos, and train for longer (harder early stopping criteria).\nMy next experiment will be 1 frame/video only, approx. 40K videos =&gt; 40K faces, but with a harder early stopping criteria, to train for 20+ epochs. Let's see... :)",
          "votes": 2
        },
        {
          "id": 714717,
          "postDate": "2020-01-09T17:22:40.243Z",
          "content": "<p>500K faces in total by now.\nI split some of them(about 50K) to do little experiment</p>",
          "rawMarkdown": "500K faces in total by now.\nI split some of them(about 50K) to do little experiment",
          "votes": 2
        },
        {
          "id": 714769,
          "postDate": "2020-01-09T18:19:54.853Z",
          "content": "<p>Adding to this, I would suggest trying less frames per video for training, as this avoids overfitting.\nFor validation it does not make much difference, maybe more frames and doing an average of prediction for frame is better than using one frame only (I haven't tried yet).</p>",
          "rawMarkdown": "Adding to this, I would suggest trying less frames per video for training, as this avoids overfitting.\nFor validation it does not make much difference, maybe more frames and doing an average of prediction for frame is better than using one frame only (I haven't tried yet).",
          "votes": 2
        },
        {
          "id": 714863,
          "postDate": "2020-01-09T21:29:33.617Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 714864,
          "postDate": "2020-01-09T21:30:00.710Z",
          "content": "<blockquote>\n  <p><strong>Carlos Souza wrote:</strong></p>\n  \n  <p>That's actually a very good question, one I'm trying to figure out right now.\n  All my experiments so far used 30 frames/video, and:\n  - 1K videos =&gt; 30K faces\n  - 5K videos =&gt; 150K faces\n  - 15K videos =&gt; 450K faces</p>\n  \n  <p>However, my trainings usually last few epochs only (my early stopping criteria may be too easy to hit).\n  After learning from others' experiences, will try less frames/video, more videos, and train for longer (harder early stopping criteria).\n  My next experiment will be 1 frame/video only, approx. 40K videos =&gt; 40K faces, but with a harder early stopping criteria, to train for 20+ epochs. Let's see... :)</p>\n</blockquote>\n\n<p>I must say your experiment planning is a lot more structured than mine. I am just doing random 1:1 sampling between images from true and fake videos, plus some group stratification 😹</p>\n\n<p>I have tried training earlier with 40k~50k faces, and had some overfitting validation score - it didn't generalise so well on the private test set, perhaps my validation set was too small</p>",
          "rawMarkdown": "&gt; **Carlos Souza wrote:**\n&gt; \n&gt; That's actually a very good question, one I'm trying to figure out right now.\n&gt; All my experiments so far used 30 frames/video, and:\n&gt; - 1K videos =&gt; 30K faces\n&gt; - 5K videos =&gt; 150K faces\n&gt; - 15K videos =&gt; 450K faces\n&gt; \n&gt; However, my trainings usually last few epochs only (my early stopping criteria may be too easy to hit).\n&gt; After learning from others' experiences, will try less frames/video, more videos, and train for longer (harder early stopping criteria).\n&gt; My next experiment will be 1 frame/video only, approx. 40K videos =&gt; 40K faces, but with a harder early stopping criteria, to train for 20+ epochs. Let's see... :)\n\nI must say your experiment planning is a lot more structured than mine. I am just doing random 1:1 sampling between images from true and fake videos, plus some group stratification 😹\n\nI have tried training earlier with 40k~50k faces, and had some overfitting validation score - it didn't generalise so well on the private test set, perhaps my validation set was too small"
        },
        {
          "id": 714866,
          "postDate": "2020-01-09T21:32:14.963Z",
          "content": "<blockquote>\n  <p><strong>FrazierLei wrote:</strong></p>\n  \n  <p>500K faces in total by now.\n  I split some of them(about 50K) to do little experiment</p>\n</blockquote>\n\n<p>500k faces sounds a lot, in my set-up that would take close to 2 hours for one epoch with one of the lighter architecture.</p>\n\n<p>I wish I can find smarter ways to build models so that I don't have to set my GPUs on fire :) </p>",
          "rawMarkdown": "&gt; **FrazierLei wrote:**\n&gt; \n&gt; 500K faces in total by now.\n&gt; I split some of them(about 50K) to do little experiment\n\n\n500k faces sounds a lot, in my set-up that would take close to 2 hours for one epoch with one of the lighter architecture.\n\nI wish I can find smarter ways to build models so that I don't have to set my GPUs on fire :) "
        },
        {
          "id": 714867,
          "postDate": "2020-01-09T21:35:14.227Z",
          "content": "<blockquote>\n  <p><strong>Pedro Bernardo wrote:</strong></p>\n  \n  <p>Adding to this, I would suggest trying less frames per video for training, as this avoids overfitting.\n  For validation it does not make much difference, maybe more frames and doing an average of prediction for frame is better than using one frame only (I haven't tried yet).</p>\n</blockquote>\n\n<p>I have tried a bit at using different numbers of frames during inference time, yes, larger amount of frames for averaging would yield slightly better scores. but at least in my pipeline that would take more memory overhead, I have run over the computing resource limit twice, and both time I was able to fixed it reducing the number of frames per video used during prediction </p>",
          "rawMarkdown": "&gt; **Pedro Bernardo wrote:**\n&gt; \n&gt; Adding to this, I would suggest trying less frames per video for training, as this avoids overfitting.\n&gt; For validation it does not make much difference, maybe more frames and doing an average of prediction for frame is better than using one frame only (I haven't tried yet).\n\nI have tried a bit at using different numbers of frames during inference time, yes, larger amount of frames for averaging would yield slightly better scores. but at least in my pipeline that would take more memory overhead, I have run over the computing resource limit twice, and both time I was able to fixed it reducing the number of frames per video used during prediction "
        },
        {
          "id": 714881,
          "postDate": "2020-01-09T22:04:54.453Z",
          "content": "<p>You can try overcome this issue predicting one video at a time, it will take longer but since you will only have one video at a time in memory you won't run into memory issues, you just need to be sure to not run out of time (time limit is nine hours), one submission that I tried using 50 frames per video took 5 hours to score, so you should be fine depending on your model.</p>",
          "rawMarkdown": "You can try overcome this issue predicting one video at a time, it will take longer but since you will only have one video at a time in memory you won't run into memory issues, you just need to be sure to not run out of time (time limit is nine hours), one submission that I tried using 50 frames per video took 5 hours to score, so you should be fine depending on your model."
        },
        {
          "id": 714909,
          "postDate": "2020-01-09T23:05:12.107Z",
          "content": "<p>I do predict one video at a time - and as you said, in theory, it should limit the amount of RAM used. But at the same time I am also applying different models, so that might be the case. will need to do a more detail study to profile the RAM usage.   </p>",
          "rawMarkdown": "I do predict one video at a time - and as you said, in theory, it should limit the amount of RAM used. But at the same time I am also applying different models, so that might be the case. will need to do a more detail study to profile the RAM usage.   "
        },
        {
          "id": 715006,
          "postDate": "2020-01-10T02:49:21.623Z",
          "content": "<p>HI，I'm a senior in China who received a condition offer from Bristol. I noticed that you were also a student(?) from Bristol. So may I get your email and  ask your some questions about the preprocessing and life in Bristol?(I'm new to kaggle,so I cant contact you by kaggle)</p>",
          "rawMarkdown": "HI，I'm a senior in China who received a condition offer from Bristol. I noticed that you were also a student(?) from Bristol. So may I get your email and  ask your some questions about the preprocessing and life in Bristol?(I'm new to kaggle,so I cant contact you by kaggle)"
        },
        {
          "id": 715126,
          "postDate": "2020-01-10T06:36:57.777Z",
          "content": "<blockquote>\n  <p><strong>Emmettj wrote:</strong></p>\n  \n  <p>HI，I'm a senior in China who received a condition offer from Bristol. I noticed that you were also a student(?) from Bristol. So may I get your email and  ask your some questions about the preprocessing and life in Bristol?(I'm new to kaggle,so I cant contact you by kaggle)</p>\n</blockquote>\n\n<p>Hello, congrat for your offer, and welcome to the kaggle community :)</p>\n\n<p>For specific kaggle competition-related topic(i.e. preprocessing), we should ONLY talk about it in the competition forum - we have a rule against private sharing, and as community members, we should help to reinforce that. </p>\n\n<p>I am not willing to share my email publicly, so I suggest you \"upgrade\" yourself to Contributor level so that you can start engaging with fellow kagglers via direct message. See the criteria <a href=\"https://www.kaggle.com/progression\">here</a>.</p>\n\n<p>An alternative way is to join the community-run <a href=\"https://www.kaggle.com/getting-started/20577\">kagglenoobs</a> slack chat, where you can direct message to others including myself.  </p>",
          "rawMarkdown": "&gt; **Emmettj wrote:**\n&gt; \n&gt; HI，I'm a senior in China who received a condition offer from Bristol. I noticed that you were also a student(?) from Bristol. So may I get your email and  ask your some questions about the preprocessing and life in Bristol?(I'm new to kaggle,so I cant contact you by kaggle)\n\nHello, congrat for your offer, and welcome to the kaggle community :)\n\nFor specific kaggle competition-related topic(i.e. preprocessing), we should ONLY talk about it in the competition forum - we have a rule against private sharing, and as community members, we should help to reinforce that. \n\nI am not willing to share my email publicly, so I suggest you \"upgrade\" yourself to Contributor level so that you can start engaging with fellow kagglers via direct message. See the criteria [here](https://www.kaggle.com/progression).\n\nAn alternative way is to join the community-run [kagglenoobs](https://www.kaggle.com/getting-started/20577) slack chat, where you can direct message to others including myself.  \n\n",
          "votes": 1
        },
        {
          "id": 715145,
          "postDate": "2020-01-10T07:13:52.223Z",
          "content": "<p>Thx for your patience.  I'm in  trouble with preprocessing the videos and training models.I saw what Sandeep Attree(11st) said and tried to do what he has done. But unfortunately, I only got a lower score, about 1.5. So can I ask some questions?\n1.what model do you use(if you are unwilling to disclose, it's ok to refuse\n2.have you considered about fake audio? how did you cope with them \n3.Did you notice that there are some noisy videos(fail to swap faces, for example,aeqpxjlbwu.mp4 in train0)</p>",
          "rawMarkdown": "Thx for your patience.  I'm in  trouble with preprocessing the videos and training models.I saw what Sandeep Attree(11st) said and tried to do what he has done. But unfortunately, I only got a lower score, about 1.5. So can I ask some questions?\n1.what model do you use(if you are unwilling to disclose, it's ok to refuse\n2.have you considered about fake audio? how did you cope with them \n3.Did you notice that there are some noisy videos(fail to swap faces, for example,aeqpxjlbwu.mp4 in train0)\n"
        },
        {
          "id": 715256,
          "postDate": "2020-01-10T10:04:39.277Z",
          "content": "<blockquote>\n  <p><strong>Emmettj wrote:</strong></p>\n  \n  <p>Thx for your patience.  I'm in  trouble with preprocessing the videos and training models.I saw what Sandeep Attree(11st) said and tried to do what he has done. But unfortunately, I only got a lower score, about 1.5. So can I ask some questions?\n  1.what model do you use(if you are unwilling to disclose, it's ok to refuse\n  2.have you considered about fake audio? how did you cope with them \n  3.Did you notice that there are some noisy videos(fail to swap faces, for example,aeqpxjlbwu.mp4 in train0)</p>\n</blockquote>\n\n<p>I am also not able to replicate Sandeep's score - especially he mentioned he was only using 1 frame per video.  Now regarding your points:\n1. I am only using the well-published architectures - you would know what they are, if not a quick search will find out.\n2. No, I haven't considered fake audio yet \n3. There had been some discussion about this on this forum, I think it is one of the admins who mentioned that with this much data, some noisy data, and even mislabelling is expected, and shouldn't hamper the modelling effort too much. I tend to agree with this </p>\n\n<p>Regarding your score, I would start by double-checking the basic things in your pipeline, is your face recognition module working? are you making the right classification? - For a long time, I was classifying False as 0 and Real as 1 - yeah I know... </p>",
          "rawMarkdown": "&gt; **Emmettj wrote:**\n&gt; \n&gt; Thx for your patience.  I'm in  trouble with preprocessing the videos and training models.I saw what Sandeep Attree(11st) said and tried to do what he has done. But unfortunately, I only got a lower score, about 1.5. So can I ask some questions?\n&gt; 1.what model do you use(if you are unwilling to disclose, it's ok to refuse\n&gt; 2.have you considered about fake audio? how did you cope with them \n&gt; 3.Did you notice that there are some noisy videos(fail to swap faces, for example,aeqpxjlbwu.mp4 in train0)\n&gt; \n\nI am also not able to replicate Sandeep's score - especially he mentioned he was only using 1 frame per video.  Now regarding your points:\n1. I am only using the well-published architectures - you would know what they are, if not a quick search will find out.\n2. No, I haven't considered fake audio yet \n3. There had been some discussion about this on this forum, I think it is one of the admins who mentioned that with this much data, some noisy data, and even mislabelling is expected, and shouldn't hamper the modelling effort too much. I tend to agree with this \n\nRegarding your score, I would start by double-checking the basic things in your pipeline, is your face recognition module working? are you making the right classification? - For a long time, I was classifying False as 0 and Real as 1 - yeah I know... "
        },
        {
          "id": 715310,
          "postDate": "2020-01-10T11:37:31.107Z",
          "content": "<p>Just finished the experiment with 1 frame/video and 40K videos (40K faces) instead of 30 frames/video and 15K videos (450K faces). Was able to get a slight improvement: from 0.63 to 0.62. However, considering that it was using way less images, I consider it a good improvement.</p>\n\n<p>So far I've only implemented one paper, FaceForensis++, and wasn't able to replicate their 81-99% accuracy: the best I got was 74%. This accuracy is, btw, exactly in the range they report as what they achieved using full images instead of face images (70-82%), which suggests my face detection &amp; image folder preparation pipeline can be improved.</p>\n\n<p>Also, I saw Sandeep's great score using only 1 frame/video. My guess is that he achieved that because he used a simpler model. My current hypothesis is that transfer learning using a complex base model, like Xception, with well over 20 million training parameters, might be overfitting to actors' faces, and not generalizing well. Anyway, I think the way forward is to test this hypothesis by trying simpler models from other papers. And if no improvement, replace my face detector algorithm.</p>",
          "rawMarkdown": "Just finished the experiment with 1 frame/video and 40K videos (40K faces) instead of 30 frames/video and 15K videos (450K faces). Was able to get a slight improvement: from 0.63 to 0.62. However, considering that it was using way less images, I consider it a good improvement.\n\nSo far I've only implemented one paper, FaceForensis++, and wasn't able to replicate their 81-99% accuracy: the best I got was 74%. This accuracy is, btw, exactly in the range they report as what they achieved using full images instead of face images (70-82%), which suggests my face detection &amp; image folder preparation pipeline can be improved.\n\nAlso, I saw Sandeep's great score using only 1 frame/video. My guess is that he achieved that because he used a simpler model. My current hypothesis is that transfer learning using a complex base model, like Xception, with well over 20 million training parameters, might be overfitting to actors' faces, and not generalizing well. Anyway, I think the way forward is to test this hypothesis by trying simpler models from other papers. And if no improvement, replace my face detector algorithm.",
          "votes": 2
        },
        {
          "id": 715316,
          "postDate": "2020-01-10T11:48:03.827Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>   thank you very much for following up, and also for your analysis. </p>\n\n<p>I did have a model with about 50k images, but I wasn't carefully sampling image from each video - it was just a random sampling and was at 0.66 - my first model that beat the 0.5 naive prediction. </p>\n\n<p>I will review this myself - and good luck with your ongoing experiment.</p>",
          "rawMarkdown": "@carlossouza   thank you very much for following up, and also for your analysis. \n\nI did have a model with about 50k images, but I wasn't carefully sampling image from each video - it was just a random sampling and was at 0.66 - my first model that beat the 0.5 naive prediction. \n\nI will review this myself - and good luck with your ongoing experiment.\n"
        },
        {
          "id": 715331,
          "postDate": "2020-01-10T12:03:00.177Z",
          "content": "<p>one more thing to add, so far on my experiments with more than 100k images, I have observed that more complex network architecture are yielding better validation &amp; LB score - so based on this I would be surprised that a simpler network would work better - although this doesn't take into the account that your image sampling (i.e. 1 - n frames per video) approach is more carefully implemented than mine. </p>",
          "rawMarkdown": "one more thing to add, so far on my experiments with more than 100k images, I have observed that more complex network architecture are yielding better validation &amp; LB score - so based on this I would be surprised that a simpler network would work better - although this doesn't take into the account that your image sampling (i.e. 1 - n frames per video) approach is more carefully implemented than mine. \n",
          "votes": 1
        },
        {
          "id": 715849,
          "postDate": "2020-01-10T22:13:55.930Z",
          "content": "<p><a href=\"/yifanxie\">@yifanxie</a> , I found a monster bug in code. Just fixed it, here's the result: LB log loss from 0.62 to 0.52 (CV 0.34)! And this was indeed achieved with only 1 frame/video with 40K videos (40K faces). :)\nAccuracy now is 85%, inline with paper's findings:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F74b5d3e43229f05e3857fcb43e6b88bf%2FScreen%20Shot%202020-01-10%20at%2019.12.55.png?generation=1578694414787332&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@yifanxie , I found a monster bug in code. Just fixed it, here's the result: LB log loss from 0.62 to 0.52 (CV 0.34)! And this was indeed achieved with only 1 frame/video with 40K videos (40K faces). :)\nAccuracy now is 85%, inline with paper's findings:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F74b5d3e43229f05e3857fcb43e6b88bf%2FScreen%20Shot%202020-01-10%20at%2019.12.55.png?generation=1578694414787332&amp;alt=media)\n",
          "votes": 2
        },
        {
          "id": 715880,
          "postDate": "2020-01-10T23:14:59.067Z",
          "content": "<p>wow, awesome, congratulation!!!!</p>",
          "rawMarkdown": "wow, awesome, congratulation!!!!"
        },
        {
          "id": 715966,
          "postDate": "2020-01-11T03:55:14.880Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> <a href=\"/yifanxie\">@yifanxie</a>\nWhat is the margin of MTCNN during your training and testing? I guessed that there might be some noise around the face in the FAKE video, so I tried to increase the margin of the detector, but found that it didn't help.\nIn addition, I use 5 frames/video to train, but the model still has serious overfitting. Do you have any data augmentation techniques? Thank you very much!</p>",
          "rawMarkdown": "@carlossouza @yifanxie\nWhat is the margin of MTCNN during your training and testing? I guessed that there might be some noise around the face in the FAKE video, so I tried to increase the margin of the detector, but found that it didn't help.\nIn addition, I use 5 frames/video to train, but the model still has serious overfitting. Do you have any data augmentation techniques? Thank you very much!\n",
          "votes": 1
        },
        {
          "id": 716185,
          "postDate": "2020-01-11T11:23:43.410Z",
          "content": "<blockquote>\n  <p>What is the margin of MTCNN during your training and testing? I guessed that there might be some noise around the face in the FAKE video, so I tried to increase the margin of the detector, but found that it didn't help.\n  In addition, I use 5 frames/video to train, but the model still has serious overfitting. Do you have any data augmentation techniques? Thank you very much!</p>\n</blockquote>\n\n<p>I don't use MTCNN, the implementation in Keras that I tried was too slow and not accurate enough for me. I use some augmentation techniques, but really nothing out of ordinary.</p>",
          "rawMarkdown": "&gt; What is the margin of MTCNN during your training and testing? I guessed that there might be some noise around the face in the FAKE video, so I tried to increase the margin of the detector, but found that it didn't help.\n&gt; In addition, I use 5 frames/video to train, but the model still has serious overfitting. Do you have any data augmentation techniques? Thank you very much!\n\nI don't use MTCNN, the implementation in Keras that I tried was too slow and not accurate enough for me. I use some augmentation techniques, but really nothing out of ordinary."
        },
        {
          "id": 717660,
          "postDate": "2020-01-13T12:21:55.883Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>  What data augmentation did you use? </p>",
          "rawMarkdown": "@carlossouza  What data augmentation did you use? "
        },
        {
          "id": 717668,
          "postDate": "2020-01-13T12:33:49.410Z",
          "content": "<p><a href=\"/fionalxd\">@fionalxd</a> , here's what I used so far:\n- 1/3 of the videos I kept unchanged\n- 2/9 of the videos I resized to 1/4 of their sizes\n- 2/9 of the videos I reduced FPS to 15\n- 2/9 of the videos I applied a hard compression</p>\n\n<p>This is exactly what was described in DFDC paper. \nFrom that point on, I extract faces from the videos, and (at least for now) I only do the transformations required from the pre-trained model I'm using (scale/normalize/center-crop). I intend to add later random crop/h&amp;v flips, and see how that might help.</p>",
          "rawMarkdown": "@fionalxd , here's what I used so far:\n- 1/3 of the videos I kept unchanged\n- 2/9 of the videos I resized to 1/4 of their sizes\n- 2/9 of the videos I reduced FPS to 15\n- 2/9 of the videos I applied a hard compression\n\nThis is exactly what was described in DFDC paper. \nFrom that point on, I extract faces from the videos, and (at least for now) I only do the transformations required from the pre-trained model I'm using (scale/normalize/center-crop). I intend to add later random crop/h&amp;v flips, and see how that might help.",
          "votes": 4
        },
        {
          "id": 717671,
          "postDate": "2020-01-13T12:47:28.390Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Thank you for sharing. According to your description, did you process all the videos, including training and validation videos?</p>",
          "rawMarkdown": "@carlossouza Thank you for sharing. According to your description, did you process all the videos, including training and validation videos?"
        },
        {
          "id": 717674,
          "postDate": "2020-01-13T12:58:01.690Z",
          "content": "<p>That's correct, I pre-processed all 119K videos. And made sure that, while splitting in train &amp; validation sets, all the above proportions are respected.</p>\n\n<p>If you want to do that, I learned 2 very important things that can significantly reduce the processing time:\n1. It is better to use a high # of cores instead a high memory size; FFMPEG is very fast in multi-processing\n2. The bottleneck in processing time is how fast the system can read the video files from the disk. If you are using AWS, my suggestion is using the fastest EBS SSD type (io1) with at least 30 IOPS/GB. The default 3 IOPS/GB is too slow. (I learned that 30 IOPS/GB is a great number because that's the default fast setting in Google Cloud)</p>\n\n<p>If you do that with the same settings I used, you should end up with approx. 260 GB of videos (instead of the original 470 GB).</p>\n\n<p>Hope it helps! Cheers!</p>",
          "rawMarkdown": "That's correct, I pre-processed all 119K videos. And made sure that, while splitting in train &amp; validation sets, all the above proportions are respected.\n\nIf you want to do that, I learned 2 very important things that can significantly reduce the processing time:\n1. It is better to use a high # of cores instead a high memory size; FFMPEG is very fast in multi-processing\n2. The bottleneck in processing time is how fast the system can read the video files from the disk. If you are using AWS, my suggestion is using the fastest EBS SSD type (io1) with at least 30 IOPS/GB. The default 3 IOPS/GB is too slow. (I learned that 30 IOPS/GB is a great number because that's the default fast setting in Google Cloud)\n\nIf you do that with the same settings I used, you should end up with approx. 260 GB of videos (instead of the original 470 GB).\n\nHope it helps! Cheers!",
          "votes": 3
        },
        {
          "id": 717680,
          "postDate": "2020-01-13T13:14:43.857Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>  Thank you very much for your suggestion. One last question is about the compress rate.  I have studied ffmpeg compression. But it seems there're several kinds of compress parameters. What command did you use for hard compression? Sorry for so much questions.....</p>",
          "rawMarkdown": "@carlossouza  Thank you very much for your suggestion. One last question is about the compress rate.  I have studied ffmpeg compression. But it seems there're several kinds of compress parameters. What command did you use for hard compression? Sorry for so much questions....."
        },
        {
          "id": 717689,
          "postDate": "2020-01-13T13:33:23.333Z",
          "content": "<p>No problem :), here's the line to compress I used:\n<code>\nffmpeg -i input.mp4 -c:v libx264 -crf 23 output.mp4\n</code></p>\n\n<p>Replace 23 with any number in the 0-51 range. 23 is a light compression. Several papers report hard compression with CRF 40. I played with both.\nMore about CRF here: <a href=\"https://trac.ffmpeg.org/wiki/Encode/H.264\">https://trac.ffmpeg.org/wiki/Encode/H.264</a></p>",
          "rawMarkdown": "No problem :), here's the line to compress I used:\n```\nffmpeg -i input.mp4 -c:v libx264 -crf 23 output.mp4\n```\n\nReplace 23 with any number in the 0-51 range. 23 is a light compression. Several papers report hard compression with CRF 40. I played with both.\nMore about CRF here: https://trac.ffmpeg.org/wiki/Encode/H.264",
          "votes": 8
        },
        {
          "id": 717704,
          "postDate": "2020-01-13T13:57:07.403Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>  I see, thank you very much for your help.</p>",
          "rawMarkdown": "@carlossouza  I see, thank you very much for your help."
        },
        {
          "id": 727909,
          "postDate": "2020-01-24T07:16:32.040Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> thanks for all your trips. Great info</p>",
          "rawMarkdown": "@carlossouza thanks for all your trips. Great info",
          "votes": 1
        }
      ]
    },
    {
      "id": 731390,
      "postDate": "2020-01-28T15:19:27.673Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 714228,
      "author_name": "FrazierLei",
      "author_url": "",
      "post_date": "2020-01-09T07:54:28.510000",
      "content": "<p>As the fake videos are much more than real ones, I extracted 2 frames from each FAKE and 10 frames from each REAL. And of course beforehand. It took me about two days.\nUsing other data augmentation methods(eg. HorizontalFlip), I believe faces from a series of single frames are enough.\nI think we could use other types of data and classifier, like time series and facial landmarks. But I don't have any time and GPU this month thanks to the Spring Festival.</p>\n\n<p>PS: I think the extracted faces should be processed again, for some of them are not really a face.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 714458,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-09T13:09:39.957000",
          "content": "<p>Good idea to fix the class imbalance by extracting different # of frames depending on whether the video is fake or real... What do you mean by re-preprocess extracted faces? Isn't it the same as only considering valid faces the ones with probability higher than a threshold (i.e. &gt;90%)?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714495,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-09T13:52:47.773000",
          "content": "<p>I don't have any good idea by now. I use facenet-pytorch to do face crop, but it doesn't give a confidence for the output face(or perhaps I just haven't figure out how). </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714517,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-09T14:06:01.973000",
          "content": "<p>I'm also using facenet-pytorch, and yes it returns probabilities! Just make sure to use <code>select_largest=False</code> when instantiating the face detector. Something like:\n<code>\nfrom facenet_pytorch import MTCNN\nmtcnn = MTCNN(keep_all=False, select_largest=False, device=device)\nboxes, probs = mtcnn.detect(pil_image)\n</code></p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 714616,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-09T15:43:38.120000",
          "content": "<p>Thanks a lot, I would try it!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 720410,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-16T11:54:16.203000",
          "content": "<p>Hi <a href=\"/feifeizaici\">@feifeizaici</a> ,\nFor the past week I'm trying to implement your idea: use faces from all videos and fix class imbalance by extracting faces from real/fake videos in 5:1 proportion. However, I'm finding a hard time to avoid overfitting: my models end up with an unrealistic amazing performance, generalizing very poorly in the LB/test set. Taking a closer look on the internals, I see that the model ends up with a high bias towards classifying videos as fakes (2:1), a problem that I saw in my first models when I was not treating class imbalance. The dataset I generated (5 faces from every real video, to 1 face from every fake video, using all 119K videos) is balanced; however, the problem still occurs.\nSo far, my best models have been trained in an undersampled dataset; i.e., for every real video, I sample one and only one fake video. This approach solves class imbalance, however it ends up using only ~32% of the data. That's why I found your approach so great and decided to explore it: it solves class imbalance and enables us to use the whole dataset.\nDid you see this problem while developing your solution?\nAlso, I imagine the problem can be treated by using random augmentations (flips, crops, etc) - is that critical to avoid overfitting in your approach?\nI'd love to hear your perspective :)\nThanks!\nCarlos</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 720549,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-16T14:16:35.023000",
          "content": "<p>Hi <a href=\"/carlossouza\">@carlossouza</a>\nNice to hear you are implementing my idea. My model is also a bit overfitting. When testing on my test set (including 2000 REAL and 2000 FAKE), I found the number of FN is much more than FP(it tends to predict more fakes). I have tried some methods but none of them worked well. And I'm still working on it. As for data augmentation, I did try some methods. Random-horizontal-flip improves the accuracy a bit, but random-clip doesn't. I‘m also trying other methods to solve this problem, such as modifying the loss function or the weights of classes.\nGood luck!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 721401,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-17T10:47:00.650000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> first of all, thanks for sharing your thought and experiment process so far. </p>\n\n<p><a href=\"/feifeizaici\">@feifeizaici</a> also thanks for sharing your experience.</p>\n\n<p>I have benefited from this thread, so I feel obligated to share some of my ad-hoc experience with this dataset. </p>\n\n<ul>\n<li>I didn't do the careful sampling as <a href=\"/carlossouza\">@carlossouza</a>  did, neither did I do the pre-process according to the DFDC paper. </li>\n<li>I did follow the FaceForensic paper's approach of using running two-stages fine-tuning, i.e. 3 epochs with only top layers unfrozen, and then several epochs of unfrozen layers for all parameters, this was done on a dataset of extracted faces of about 40k images. This gave some good improvement compared to my previous training methods. Not as nice as the LB0.53 score you shared, mine was on about LB0.57</li>\n<li>However, the above approch doesn't work on my side, when trying to scale to larger datasets. I failed to generate any better model than LB0.57 with this approach for datasets with more than 300k images, and I have experimented with various well-known architectures including Xception. </li>\n<li>My current best single model was trained on a roughly balanced dataset of 300k images, with images from 3 groups of videos used as hold-out</li>\n<li>I have been using some run-of-the-mill image augmentation - horizontal flip, gaussian blur,  changing brigthness, etc. These are things that I have carried our from a previous project, so I might have to revisit their contribution to the modelling one by one in future experiments - in the unlikely event that I find time to do it :)</li>\n<li>Regarding overfitting, so far I don't have issue generalising from hold-out score to LB, hold-out score and LB are more or less 0.04 from each other either way.</li>\n</ul>\n\n<p>As far as modelling goes, so far the majority of my considering has been on choosing between base-architecture, add-on architecture configuration, fine-tuning methods, and learning rate related issue.  It is my first time to try hard on Video/Image competitions so I am flying blind in very long training cycle. </p>\n\n<p>I hope the above practical insights are useful to some of us. </p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 731391,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2020-01-28T15:19:50.660000",
          "content": "<p>Hi, <a href=\"/feifeizaici\">@feifeizaici</a>, thanks for your sharing,</p>\n\n<p>I only download the files two days ago, maybe im not familiar to dataset, sorry if im wrong.\nThe FAKE video also point to a REAL video, there only 20% REAL video which don't have FAKE video.\nIf we extract image from all videos, It means that ratio of FAKE : REAL = 4:5, not 4:1</p>\n\n<p>Don't know if its better to use all video, or only use video list in metadata.json?\nThanks,</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 731415,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-28T15:57:38.793000",
          "content": "<p>Hi, <a href=\"/gody7334\">@gody7334</a> \nSome REAL videos was processed by different deepfake algorithms, and generated not only one FAKE videos. The FAKE/REAL ratio in the whole dataset is about 6:1, which is unbalaced.\nBest wishes!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 732322,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "2020-01-29T17:22:00.210000",
          "content": "<p><a href=\"/feifeizaici\">@feifeizaici</a>  I see, thanks for your reply, cheers.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 713640,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-01-08T13:52:10.587000",
      "content": "<p>On (2), I'd suggest doing it beforehand: it this way, you can reuse the frames to train new models, saving quite some time. Assuming your model can detect faces at 30 FPS, and that you will extract 30 faces/video, if you use 15,000 videos (~10% of the data), that would take 15,000 * 30 / 30 / 3,600 = 4 hours 10 min! And that's only to pre-process faces! Imagine spending that time/money every time you train a new classifier... :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 713562,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2020-01-08T12:34:07.337000",
      "content": "<ol>\n<li>Num Frames - Keep it simple lets say 30-50 frames are enough from 1 video.</li>\n<li>Preprocess after you extract face.</li>\n<li>What do you mean by overload, explain it in detail, if it is what I think it is, use multiple images for one time input say 1000 images each time you ask mtcnn for boxes to optimize it, I really do not recommend mtcnn though, it is slow and not the most accurate, still not gonna reveal what I am using yet.</li>\n</ol>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 727074,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-01-23T12:45:54.827000",
      "content": "<p><a href=\"/bibek777\">@bibek777</a> good to find you ..\nwe are new entrant to competition. COuld u tell what sort of preprocessing is needed.. is it extraction of faces using standard Face reco models like MTCNN and then using those images ?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 722912,
      "author_name": "alazycoder",
      "author_url": "",
      "post_date": "2020-01-19T09:46:53.557000",
      "content": "<p>Excuse me. I'm confused about how to use facenet-pytorch in kaggle notebook. \"No custom packages enabled in your submission notebook\".</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 718332,
      "author_name": "Data-RanDan",
      "author_url": "",
      "post_date": "2020-01-14T09:50:42.040000",
      "content": "<p>nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 714653,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-01-09T16:11:40.783000",
      "content": "<p><a href=\"/carlossouza\">@carlossouza</a> and <a href=\"/feifeizaici\">@feifeizaici</a> might I ask what is the size of image dataset you extracted from the video to genreate  your model?  My models so far are trained on about 200k face-cropped images with different sampling proportion. Just trying to get an indication what is the best way to go forward, i.e. more images v.s. smarter way of building models - or both...  </p>",
      "votes": 0,
      "replies": [
        {
          "id": 714702,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-09T17:03:50.270000",
          "content": "<p>That's actually a very good question, one I'm trying to figure out right now.\nAll my experiments so far used 30 frames/video, and:\n- 1K videos =&gt; 30K faces\n- 5K videos =&gt; 150K faces\n- 15K videos =&gt; 450K faces</p>\n\n<p>However, my trainings usually last few epochs only (my early stopping criteria may be too easy to hit).\nAfter learning from others' experiences, will try less frames/video, more videos, and train for longer (harder early stopping criteria).\nMy next experiment will be 1 frame/video only, approx. 40K videos =&gt; 40K faces, but with a harder early stopping criteria, to train for 20+ epochs. Let's see... :)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 714717,
          "author_name": "FrazierLei",
          "author_url": "",
          "post_date": "2020-01-09T17:22:40.243000",
          "content": "<p>500K faces in total by now.\nI split some of them(about 50K) to do little experiment</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 714769,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-09T18:19:54.853000",
          "content": "<p>Adding to this, I would suggest trying less frames per video for training, as this avoids overfitting.\nFor validation it does not make much difference, maybe more frames and doing an average of prediction for frame is better than using one frame only (I haven't tried yet).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 714863,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-09T21:29:33.617000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714864,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-09T21:30:00.710000",
          "content": "<blockquote>\n  <p><strong>Carlos Souza wrote:</strong></p>\n  \n  <p>That's actually a very good question, one I'm trying to figure out right now.\n  All my experiments so far used 30 frames/video, and:\n  - 1K videos =&gt; 30K faces\n  - 5K videos =&gt; 150K faces\n  - 15K videos =&gt; 450K faces</p>\n  \n  <p>However, my trainings usually last few epochs only (my early stopping criteria may be too easy to hit).\n  After learning from others' experiences, will try less frames/video, more videos, and train for longer (harder early stopping criteria).\n  My next experiment will be 1 frame/video only, approx. 40K videos =&gt; 40K faces, but with a harder early stopping criteria, to train for 20+ epochs. Let's see... :)</p>\n</blockquote>\n\n<p>I must say your experiment planning is a lot more structured than mine. I am just doing random 1:1 sampling between images from true and fake videos, plus some group stratification 😹</p>\n\n<p>I have tried training earlier with 40k~50k faces, and had some overfitting validation score - it didn't generalise so well on the private test set, perhaps my validation set was too small</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714866,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-09T21:32:14.963000",
          "content": "<blockquote>\n  <p><strong>FrazierLei wrote:</strong></p>\n  \n  <p>500K faces in total by now.\n  I split some of them(about 50K) to do little experiment</p>\n</blockquote>\n\n<p>500k faces sounds a lot, in my set-up that would take close to 2 hours for one epoch with one of the lighter architecture.</p>\n\n<p>I wish I can find smarter ways to build models so that I don't have to set my GPUs on fire :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714867,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-09T21:35:14.227000",
          "content": "<blockquote>\n  <p><strong>Pedro Bernardo wrote:</strong></p>\n  \n  <p>Adding to this, I would suggest trying less frames per video for training, as this avoids overfitting.\n  For validation it does not make much difference, maybe more frames and doing an average of prediction for frame is better than using one frame only (I haven't tried yet).</p>\n</blockquote>\n\n<p>I have tried a bit at using different numbers of frames during inference time, yes, larger amount of frames for averaging would yield slightly better scores. but at least in my pipeline that would take more memory overhead, I have run over the computing resource limit twice, and both time I was able to fixed it reducing the number of frames per video used during prediction </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714881,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-09T22:04:54.453000",
          "content": "<p>You can try overcome this issue predicting one video at a time, it will take longer but since you will only have one video at a time in memory you won't run into memory issues, you just need to be sure to not run out of time (time limit is nine hours), one submission that I tried using 50 frames per video took 5 hours to score, so you should be fine depending on your model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714909,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-09T23:05:12.107000",
          "content": "<p>I do predict one video at a time - and as you said, in theory, it should limit the amount of RAM used. But at the same time I am also applying different models, so that might be the case. will need to do a more detail study to profile the RAM usage.   </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715006,
          "author_name": "Emmettj",
          "author_url": "",
          "post_date": "2020-01-10T02:49:21.623000",
          "content": "<p>HI，I'm a senior in China who received a condition offer from Bristol. I noticed that you were also a student(?) from Bristol. So may I get your email and  ask your some questions about the preprocessing and life in Bristol?(I'm new to kaggle,so I cant contact you by kaggle)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715126,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-10T06:36:57.777000",
          "content": "<blockquote>\n  <p><strong>Emmettj wrote:</strong></p>\n  \n  <p>HI，I'm a senior in China who received a condition offer from Bristol. I noticed that you were also a student(?) from Bristol. So may I get your email and  ask your some questions about the preprocessing and life in Bristol?(I'm new to kaggle,so I cant contact you by kaggle)</p>\n</blockquote>\n\n<p>Hello, congrat for your offer, and welcome to the kaggle community :)</p>\n\n<p>For specific kaggle competition-related topic(i.e. preprocessing), we should ONLY talk about it in the competition forum - we have a rule against private sharing, and as community members, we should help to reinforce that. </p>\n\n<p>I am not willing to share my email publicly, so I suggest you \"upgrade\" yourself to Contributor level so that you can start engaging with fellow kagglers via direct message. See the criteria <a href=\"https://www.kaggle.com/progression\">here</a>.</p>\n\n<p>An alternative way is to join the community-run <a href=\"https://www.kaggle.com/getting-started/20577\">kagglenoobs</a> slack chat, where you can direct message to others including myself.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 715145,
          "author_name": "Emmettj",
          "author_url": "",
          "post_date": "2020-01-10T07:13:52.223000",
          "content": "<p>Thx for your patience.  I'm in  trouble with preprocessing the videos and training models.I saw what Sandeep Attree(11st) said and tried to do what he has done. But unfortunately, I only got a lower score, about 1.5. So can I ask some questions?\n1.what model do you use(if you are unwilling to disclose, it's ok to refuse\n2.have you considered about fake audio? how did you cope with them \n3.Did you notice that there are some noisy videos(fail to swap faces, for example,aeqpxjlbwu.mp4 in train0)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715256,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-10T10:04:39.277000",
          "content": "<blockquote>\n  <p><strong>Emmettj wrote:</strong></p>\n  \n  <p>Thx for your patience.  I'm in  trouble with preprocessing the videos and training models.I saw what Sandeep Attree(11st) said and tried to do what he has done. But unfortunately, I only got a lower score, about 1.5. So can I ask some questions?\n  1.what model do you use(if you are unwilling to disclose, it's ok to refuse\n  2.have you considered about fake audio? how did you cope with them \n  3.Did you notice that there are some noisy videos(fail to swap faces, for example,aeqpxjlbwu.mp4 in train0)</p>\n</blockquote>\n\n<p>I am also not able to replicate Sandeep's score - especially he mentioned he was only using 1 frame per video.  Now regarding your points:\n1. I am only using the well-published architectures - you would know what they are, if not a quick search will find out.\n2. No, I haven't considered fake audio yet \n3. There had been some discussion about this on this forum, I think it is one of the admins who mentioned that with this much data, some noisy data, and even mislabelling is expected, and shouldn't hamper the modelling effort too much. I tend to agree with this </p>\n\n<p>Regarding your score, I would start by double-checking the basic things in your pipeline, is your face recognition module working? are you making the right classification? - For a long time, I was classifying False as 0 and Real as 1 - yeah I know... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715310,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-10T11:37:31.107000",
          "content": "<p>Just finished the experiment with 1 frame/video and 40K videos (40K faces) instead of 30 frames/video and 15K videos (450K faces). Was able to get a slight improvement: from 0.63 to 0.62. However, considering that it was using way less images, I consider it a good improvement.</p>\n\n<p>So far I've only implemented one paper, FaceForensis++, and wasn't able to replicate their 81-99% accuracy: the best I got was 74%. This accuracy is, btw, exactly in the range they report as what they achieved using full images instead of face images (70-82%), which suggests my face detection &amp; image folder preparation pipeline can be improved.</p>\n\n<p>Also, I saw Sandeep's great score using only 1 frame/video. My guess is that he achieved that because he used a simpler model. My current hypothesis is that transfer learning using a complex base model, like Xception, with well over 20 million training parameters, might be overfitting to actors' faces, and not generalizing well. Anyway, I think the way forward is to test this hypothesis by trying simpler models from other papers. And if no improvement, replace my face detector algorithm.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 715316,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-10T11:48:03.827000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>   thank you very much for following up, and also for your analysis. </p>\n\n<p>I did have a model with about 50k images, but I wasn't carefully sampling image from each video - it was just a random sampling and was at 0.66 - my first model that beat the 0.5 naive prediction. </p>\n\n<p>I will review this myself - and good luck with your ongoing experiment.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715331,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-10T12:03:00.177000",
          "content": "<p>one more thing to add, so far on my experiments with more than 100k images, I have observed that more complex network architecture are yielding better validation &amp; LB score - so based on this I would be surprised that a simpler network would work better - although this doesn't take into the account that your image sampling (i.e. 1 - n frames per video) approach is more carefully implemented than mine. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 715849,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-10T22:13:55.930000",
          "content": "<p><a href=\"/yifanxie\">@yifanxie</a> , I found a monster bug in code. Just fixed it, here's the result: LB log loss from 0.62 to 0.52 (CV 0.34)! And this was indeed achieved with only 1 frame/video with 40K videos (40K faces). :)\nAccuracy now is 85%, inline with paper's findings:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F74b5d3e43229f05e3857fcb43e6b88bf%2FScreen%20Shot%202020-01-10%20at%2019.12.55.png?generation=1578694414787332&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 715880,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-10T23:14:59.067000",
          "content": "<p>wow, awesome, congratulation!!!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 715966,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-01-11T03:55:14.880000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> <a href=\"/yifanxie\">@yifanxie</a>\nWhat is the margin of MTCNN during your training and testing? I guessed that there might be some noise around the face in the FAKE video, so I tried to increase the margin of the detector, but found that it didn't help.\nIn addition, I use 5 frames/video to train, but the model still has serious overfitting. Do you have any data augmentation techniques? Thank you very much!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 716185,
          "author_name": "Yifan Xie",
          "author_url": "",
          "post_date": "2020-01-11T11:23:43.410000",
          "content": "<blockquote>\n  <p>What is the margin of MTCNN during your training and testing? I guessed that there might be some noise around the face in the FAKE video, so I tried to increase the margin of the detector, but found that it didn't help.\n  In addition, I use 5 frames/video to train, but the model still has serious overfitting. Do you have any data augmentation techniques? Thank you very much!</p>\n</blockquote>\n\n<p>I don't use MTCNN, the implementation in Keras that I tried was too slow and not accurate enough for me. I use some augmentation techniques, but really nothing out of ordinary.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717660,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-01-13T12:21:55.883000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>  What data augmentation did you use? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717668,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-13T12:33:49.410000",
          "content": "<p><a href=\"/fionalxd\">@fionalxd</a> , here's what I used so far:\n- 1/3 of the videos I kept unchanged\n- 2/9 of the videos I resized to 1/4 of their sizes\n- 2/9 of the videos I reduced FPS to 15\n- 2/9 of the videos I applied a hard compression</p>\n\n<p>This is exactly what was described in DFDC paper. \nFrom that point on, I extract faces from the videos, and (at least for now) I only do the transformations required from the pre-trained model I'm using (scale/normalize/center-crop). I intend to add later random crop/h&amp;v flips, and see how that might help.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 717671,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-01-13T12:47:28.390000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Thank you for sharing. According to your description, did you process all the videos, including training and validation videos?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717674,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-13T12:58:01.690000",
          "content": "<p>That's correct, I pre-processed all 119K videos. And made sure that, while splitting in train &amp; validation sets, all the above proportions are respected.</p>\n\n<p>If you want to do that, I learned 2 very important things that can significantly reduce the processing time:\n1. It is better to use a high # of cores instead a high memory size; FFMPEG is very fast in multi-processing\n2. The bottleneck in processing time is how fast the system can read the video files from the disk. If you are using AWS, my suggestion is using the fastest EBS SSD type (io1) with at least 30 IOPS/GB. The default 3 IOPS/GB is too slow. (I learned that 30 IOPS/GB is a great number because that's the default fast setting in Google Cloud)</p>\n\n<p>If you do that with the same settings I used, you should end up with approx. 260 GB of videos (instead of the original 470 GB).</p>\n\n<p>Hope it helps! Cheers!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 717680,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-01-13T13:14:43.857000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>  Thank you very much for your suggestion. One last question is about the compress rate.  I have studied ffmpeg compression. But it seems there're several kinds of compress parameters. What command did you use for hard compression? Sorry for so much questions.....</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 717689,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-13T13:33:23.333000",
          "content": "<p>No problem :), here's the line to compress I used:\n<code>\nffmpeg -i input.mp4 -c:v libx264 -crf 23 output.mp4\n</code></p>\n\n<p>Replace 23 with any number in the 0-51 range. 23 is a light compression. Several papers report hard compression with CRF 40. I played with both.\nMore about CRF here: <a href=\"https://trac.ffmpeg.org/wiki/Encode/H.264\">https://trac.ffmpeg.org/wiki/Encode/H.264</a></p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 717704,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2020-01-13T13:57:07.403000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a>  I see, thank you very much for your help.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 727909,
          "author_name": "Nick Sciarrilli",
          "author_url": "",
          "post_date": "2020-01-24T07:16:32.040000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> thanks for all your trips. Great info</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 731390,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-28T15:19:27.673000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "713405": "I'm thinking about implementing this pipeline from [Capsule Network](https://arxiv.org/pdf/1910.12467.pdf) paper for this competition. \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1528571%2F79f2d5d9866d9b72a3a2fa93e053978c%2Fpipeline.png?generation=1578473585651924&amp;alt=media)\n\nAccording to paper, preprocessing(in case of videos) includes separating the frames and extracting the face; this is what I plan to do but there are many variables to consider like \n1. number of frames\n2. should preprocessing be done beforehand or on-the-fly?\n3. how to tackle face-extractor(eg: mtcnn) overload?\n\nI wonder how other Kagglers are approaching these problems?",
    "714228": "As the fake videos are much more than real ones, I extracted 2 frames from each FAKE and 10 frames from each REAL. And of course beforehand. It took me about two days.\nUsing other data augmentation methods(eg. HorizontalFlip), I believe faces from a series of single frames are enough.\nI think we could use other types of data and classifier, like time series and facial landmarks. But I don't have any time and GPU this month thanks to the Spring Festival.\n\nPS: I think the extracted faces should be processed again, for some of them are not really a face.",
    "713640": "On (2), I'd suggest doing it beforehand: it this way, you can reuse the frames to train new models, saving quite some time. Assuming your model can detect faces at 30 FPS, and that you will extract 30 faces/video, if you use 15,000 videos (~10% of the data), that would take 15,000 * 30 / 30 / 3,600 = 4 hours 10 min! And that's only to pre-process faces! Imagine spending that time/money every time you train a new classifier... :)",
    "713562": "1. Num Frames - Keep it simple lets say 30-50 frames are enough from 1 video.\n2. Preprocess after you extract face.\n3. What do you mean by overload, explain it in detail, if it is what I think it is, use multiple images for one time input say 1000 images each time you ask mtcnn for boxes to optimize it, I really do not recommend mtcnn though, it is slow and not the most accurate, still not gonna reveal what I am using yet.",
    "727074": "@bibek777 good to find you ..\nwe are new entrant to competition. COuld u tell what sort of preprocessing is needed.. is it extraction of faces using standard Face reco models like MTCNN and then using those images ?",
    "722912": "Excuse me. I'm confused about how to use facenet-pytorch in kaggle notebook. \"No custom packages enabled in your submission notebook\".",
    "718332": "nice",
    "714653": "@carlossouza and @feifeizaici might I ask what is the size of image dataset you extracted from the video to genreate  your model?  My models so far are trained on about 200k face-cropped images with different sampling proportion. Just trying to get an indication what is the best way to go forward, i.e. more images v.s. smarter way of building models - or both...  ",
    "731390": ""
  }
}