{
  "id": 130759,
  "title": "A simple approach to classifying videos",
  "url": "/competitions/deepfake-detection-challenge/discussion/130759",
  "author_name": "Hieu Phung",
  "post_date": "2020-02-16T06:52:07.338000",
  "votes": 17,
  "comment_count": 22,
  "views": 0,
  "content": "<p>It seems that many teams are doing well in this competition, yet some still struggle to get started. For learning and research purposes, I propose a simple solution named <code>multiface classifier</code> to deal with the problem in this competition.</p>\n\n<p>Note:\n* This solution only used for learning and research purposes, and I do not guarantee that it will produce good results.\n* All the datasets I used to train the classifier are well prepared beforehand, and you can find the full list of them in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954\"><em>Other useful datasets</em></a> discussion.</p>\n\n<hr>\n\n<h2>Idea</h2>\n\n<p>Sometimes, it is hard to determine a video if it is <code>real</code> or <code>fake</code> by only using a single face appears in this video, or separately classify each face then combine the results in some ways, for example averaging, to predict the label of the input video. I think we can give the classifier more meaningful information by feeding it with multiple face images at once (the strategy to sample these images will be discussed later.) Then, based on a sequence of faces, the classifier can give a better judgment on the video it gets.</p>\n\n<hr>\n\n<h2>Multiface's general diagram</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1853329%2Ffa31c1ff391ce738e8da7273a9620153%2Fmultiface.png?generation=1581835338165826&amp;alt=media\" alt=\"diagram\"></p>\n\n<hr>\n\n<h2>Implementation</h2>\n\n<p>I have prepared a <code>kernel</code> named <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-training\"><em>DFDC-Multiface-Training</em></a> dedicated to realizing this idea. Along with it, I have also created the <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-inference\"><em>DFDC-Multiface-Inference</em></a> kernel for the inference process.</p>\n\n<p>To start with, I've tried ResNet18 as the classifier and chose to use 5 sampled faces in each video to classify it.\n* Each chosen face image will be preprocessed separately then stacked together depth-wise (along the third axis) into a single tensor before fed into the model.\n* In the training process, I will generate a uniform random sample of size 5 (the number of faces.) Note: I am not sure random sampling can help prevent overfitting so further experiments must be conducted.\n* In the validation process, I will not choose input faces randomly like before; instead, I use a different strategy to obtain these images. I will try to get enough faces throughout the video by evenly spaced sampling; if I cannot get enough in the first run, I will continue this strategy but with a little shift (or stride) in the interval [<em>start, stop</em>] to get different faces if possible, and continue this process until the fifth try (just a hyper-parameter to prevent infinite loop.)\n* The model will be led by the Focal Loss and optimized by the Adam algorithm.\n* In the inference process, I will loop through all test videos and try to get face images by using the same strategy as I have applied to the validation process above. The only difference is instead of having well-prepared data, I must run a face-detector, the same as I used to prepare the training dataset in the <a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor\"><em>Data Preparation</em></a> kernel, to directly extract faces from each frame of one input video. If I fail to get enough faces from a video, I will mark it as <code>invalid</code> and assign a <code>default predicted value</code> (probability) to this video.</p>\n\n<hr>\n\n<p>If you have any questions or suggestions, please let me know!</p>\n\n<p>Thanks for reading!</p>",
  "messages": [
    {
      "id": 747249,
      "postDate": "2020-02-16T06:52:07.337Z",
      "content": "<p>It seems that many teams are doing well in this competition, yet some still struggle to get started. For learning and research purposes, I propose a simple solution named <code>multiface classifier</code> to deal with the problem in this competition.</p>\n\n<p>Note:\n* This solution only used for learning and research purposes, and I do not guarantee that it will produce good results.\n* All the datasets I used to train the classifier are well prepared beforehand, and you can find the full list of them in the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954\"><em>Other useful datasets</em></a> discussion.</p>\n\n<hr>\n\n<h2>Idea</h2>\n\n<p>Sometimes, it is hard to determine a video if it is <code>real</code> or <code>fake</code> by only using a single face appears in this video, or separately classify each face then combine the results in some ways, for example averaging, to predict the label of the input video. I think we can give the classifier more meaningful information by feeding it with multiple face images at once (the strategy to sample these images will be discussed later.) Then, based on a sequence of faces, the classifier can give a better judgment on the video it gets.</p>\n\n<hr>\n\n<h2>Multiface's general diagram</h2>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1853329%2Ffa31c1ff391ce738e8da7273a9620153%2Fmultiface.png?generation=1581835338165826&amp;alt=media\" alt=\"diagram\"></p>\n\n<hr>\n\n<h2>Implementation</h2>\n\n<p>I have prepared a <code>kernel</code> named <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-training\"><em>DFDC-Multiface-Training</em></a> dedicated to realizing this idea. Along with it, I have also created the <a href=\"https://www.kaggle.com/phunghieu/dfdc-multiface-inference\"><em>DFDC-Multiface-Inference</em></a> kernel for the inference process.</p>\n\n<p>To start with, I've tried ResNet18 as the classifier and chose to use 5 sampled faces in each video to classify it.\n* Each chosen face image will be preprocessed separately then stacked together depth-wise (along the third axis) into a single tensor before fed into the model.\n* In the training process, I will generate a uniform random sample of size 5 (the number of faces.) Note: I am not sure random sampling can help prevent overfitting so further experiments must be conducted.\n* In the validation process, I will not choose input faces randomly like before; instead, I use a different strategy to obtain these images. I will try to get enough faces throughout the video by evenly spaced sampling; if I cannot get enough in the first run, I will continue this strategy but with a little shift (or stride) in the interval [<em>start, stop</em>] to get different faces if possible, and continue this process until the fifth try (just a hyper-parameter to prevent infinite loop.)\n* The model will be led by the Focal Loss and optimized by the Adam algorithm.\n* In the inference process, I will loop through all test videos and try to get face images by using the same strategy as I have applied to the validation process above. The only difference is instead of having well-prepared data, I must run a face-detector, the same as I used to prepare the training dataset in the <a href=\"https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor\"><em>Data Preparation</em></a> kernel, to directly extract faces from each frame of one input video. If I fail to get enough faces from a video, I will mark it as <code>invalid</code> and assign a <code>default predicted value</code> (probability) to this video.</p>\n\n<hr>\n\n<p>If you have any questions or suggestions, please let me know!</p>\n\n<p>Thanks for reading!</p>",
      "rawMarkdown": "It seems that many teams are doing well in this competition, yet some still struggle to get started. For learning and research purposes, I propose a simple solution named `multiface classifier` to deal with the problem in this competition.\n\nNote:\n* This solution only used for learning and research purposes, and I do not guarantee that it will produce good results.\n* All the datasets I used to train the classifier are well prepared beforehand, and you can find the full list of them in the [*Other useful datasets*](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954) discussion.\n\n---\n## Idea\nSometimes, it is hard to determine a video if it is `real` or `fake` by only using a single face appears in this video, or separately classify each face then combine the results in some ways, for example averaging, to predict the label of the input video. I think we can give the classifier more meaningful information by feeding it with multiple face images at once (the strategy to sample these images will be discussed later.) Then, based on a sequence of faces, the classifier can give a better judgment on the video it gets.\n\n---\n## Multiface's general diagram\n![diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1853329%2Ffa31c1ff391ce738e8da7273a9620153%2Fmultiface.png?generation=1581835338165826&amp;alt=media)\n\n---\n## Implementation\nI have prepared a `kernel` named [*DFDC-Multiface-Training*](https://www.kaggle.com/phunghieu/dfdc-multiface-training) dedicated to realizing this idea. Along with it, I have also created the [*DFDC-Multiface-Inference*](https://www.kaggle.com/phunghieu/dfdc-multiface-inference) kernel for the inference process.\n\nTo start with, I've tried ResNet18 as the classifier and chose to use 5 sampled faces in each video to classify it.\n* Each chosen face image will be preprocessed separately then stacked together depth-wise (along the third axis) into a single tensor before fed into the model.\n* In the training process, I will generate a uniform random sample of size 5 (the number of faces.) Note: I am not sure random sampling can help prevent overfitting so further experiments must be conducted.\n* In the validation process, I will not choose input faces randomly like before; instead, I use a different strategy to obtain these images. I will try to get enough faces throughout the video by evenly spaced sampling; if I cannot get enough in the first run, I will continue this strategy but with a little shift (or stride) in the interval [*start, stop*] to get different faces if possible, and continue this process until the fifth try (just a hyper-parameter to prevent infinite loop.)\n* The model will be led by the Focal Loss and optimized by the Adam algorithm.\n* In the inference process, I will loop through all test videos and try to get face images by using the same strategy as I have applied to the validation process above. The only difference is instead of having well-prepared data, I must run a face-detector, the same as I used to prepare the training dataset in the [*Data Preparation*](https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor) kernel, to directly extract faces from each frame of one input video. If I fail to get enough faces from a video, I will mark it as `invalid` and assign a `default predicted value` (probability) to this video.\n\n---\nIf you have any questions or suggestions, please let me know!\n\nThanks for reading!",
      "votes": 17
    },
    {
      "id": 747326,
      "postDate": "2020-02-16T09:18:17.340Z",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice",
      "votes": 3,
      "replies": [
        {
          "id": 747328,
          "postDate": "2020-02-16T09:22:09.853Z",
          "content": "<p>Thanks! <a href=\"/dataraj\">@dataraj</a> 😊 </p>",
          "rawMarkdown": "Thanks! @dataraj 😊 "
        }
      ]
    },
    {
      "id": 747994,
      "postDate": "2020-02-17T05:09:10.497Z",
      "content": "<p>I am new to deep learning; trying to learn. And luckily during my learning process i saw this competition. When you have a goal it is faster to learn. So it is helping me a lot. \n(also the fastai course series encouraged me and helped me to understand some concepts). So my comments might be just noob comments. Read it accordingly:)</p>\n\n<p>The first thing I have tried in this competition was a very similar approach; I stacked the images and used 3d version of the functions. But I have faced some problems.\nThe first problem was; for the 3d version; I was not able to find a good pre-trained model to use. I had to train from scratch. I think this is a big disadvantage.\nThe second problem I faced was choosing the conv kernel size; when it is 3d; it is getting complicated because on the plane the neighbor pixels have a different meaning but accross the planes it has a different meaning(probably the network would figure out the different meanings but still) so wasn't sure about choosing a cubic kernel size.\nThe third issue was I wasn't still persuaded myself if the relationship between frames are really that meaningful to say fake or not? Does the extra contribution worth the effort?</p>\n\n<p>Anyways later I decided to go with 2d network; I got better results this way(during predictions I took the mean of the per frame results).</p>\n\n<p>But I still did not throw away the stacked faces idea; if score from my 2d approach stops improving i might re-visit the idea of the stacked faces.</p>",
      "rawMarkdown": "I am new to deep learning; trying to learn. And luckily during my learning process i saw this competition. When you have a goal it is faster to learn. So it is helping me a lot. \n(also the fastai course series encouraged me and helped me to understand some concepts). So my comments might be just noob comments. Read it accordingly:)\n\nThe first thing I have tried in this competition was a very similar approach; I stacked the images and used 3d version of the functions. But I have faced some problems.\nThe first problem was; for the 3d version; I was not able to find a good pre-trained model to use. I had to train from scratch. I think this is a big disadvantage.\nThe second problem I faced was choosing the conv kernel size; when it is 3d; it is getting complicated because on the plane the neighbor pixels have a different meaning but accross the planes it has a different meaning(probably the network would figure out the different meanings but still) so wasn't sure about choosing a cubic kernel size.\nThe third issue was I wasn't still persuaded myself if the relationship between frames are really that meaningful to say fake or not? Does the extra contribution worth the effort?\n\nAnyways later I decided to go with 2d network; I got better results this way(during predictions I took the mean of the per frame results).\n\nBut I still did not throw away the stacked faces idea; if score from my 2d approach stops improving i might re-visit the idea of the stacked faces.",
      "votes": 1,
      "replies": [
        {
          "id": 748383,
          "postDate": "2020-02-17T13:20:14.050Z",
          "content": "<p>Hi <a href=\"/emrebayram\">@emrebayram</a>,</p>\n\n<p>Awesome! Good to hear that someone has tried 3D classification approach. I agree with you, training the model from scratch is really a big disadvantage, except you have tons of labeled data, lots of computing power, and time of course. Yet, I think it's good to do so, you can learn new interesting things, gain more insights, etc. many many benefits from this job. I think you should publish this solution at the end of the competition, this will definitely help the community a lot.</p>\n\n<p>BTW, good luck 😉.</p>",
          "rawMarkdown": "Hi @emrebayram,\n\nAwesome! Good to hear that someone has tried 3D classification approach. I agree with you, training the model from scratch is really a big disadvantage, except you have tons of labeled data, lots of computing power, and time of course. Yet, I think it's good to do so, you can learn new interesting things, gain more insights, etc. many many benefits from this job. I think you should publish this solution at the end of the competition, this will definitely help the community a lot.\n\nBTW, good luck 😉."
        }
      ]
    },
    {
      "id": 747626,
      "postDate": "2020-02-16T16:54:52.140Z",
      "content": "<p>Well here's a problem. Some videos have 1 face, some videos have 2 faces. You cannot determine the exact input shape. Besides, sometimes videos have 1 fake face, the other one is real.</p>",
      "rawMarkdown": "Well here's a problem. Some videos have 1 face, some videos have 2 faces. You cannot determine the exact input shape. Besides, sometimes videos have 1 fake face, the other one is real.",
      "votes": 1,
      "replies": [
        {
          "id": 747942,
          "postDate": "2020-02-17T03:08:31.783Z",
          "content": "<p>Hi <a href=\"/unkownhihi\">@unkownhihi</a>,</p>\n\n<p>Yes, I know about this problem, so I have designed a simple strategy to sample face images from each video and if I cannot get enough faces the model needs, I'll assign a default value to this video in the final submission.</p>",
          "rawMarkdown": "Hi @unkownhihi,\n\nYes, I know about this problem, so I have designed a simple strategy to sample face images from each video and if I cannot get enough faces the model needs, I'll assign a default value to this video in the final submission."
        },
        {
          "id": 747945,
          "postDate": "2020-02-17T03:12:10.167Z",
          "content": "<p>And about \"sometimes videos have 1 fake face, the other one is real,\"  I also aware of this problem; however, I've not come up with any ideas yet. Do you have any suggestions 😃?</p>",
          "rawMarkdown": "And about \"sometimes videos have 1 fake face, the other one is real,\"  I also aware of this problem; however, I've not come up with any ideas yet. Do you have any suggestions 😃?"
        },
        {
          "id": 747946,
          "postDate": "2020-02-17T03:12:49.383Z",
          "content": "<p>but from my observation, most videos have 1 face, instead of 2 faces.</p>",
          "rawMarkdown": "but from my observation, most videos have 1 face, instead of 2 faces."
        },
        {
          "id": 747954,
          "postDate": "2020-02-17T03:25:16.870Z",
          "content": "<p>I mean, I will sample face images throughout the whole video's frames.</p>",
          "rawMarkdown": "I mean, I will sample face images throughout the whole video's frames."
        },
        {
          "id": 747955,
          "postDate": "2020-02-17T03:27:11.847Z",
          "content": "<p>Nice! What's the LB?</p>",
          "rawMarkdown": "Nice! What's the LB?"
        },
        {
          "id": 747959,
          "postDate": "2020-02-17T03:32:54.793Z",
          "content": "<p>It hasn't worked well yet, the best I can get is 0.79725 at early epochs. When trying to fit the model further with the dataset, it seems to have good improvements in local logloss value; however, LB is skyrocketed due to overfitting.</p>",
          "rawMarkdown": "It hasn't worked well yet, the best I can get is 0.79725 at early epochs. When trying to fit the model further with the dataset, it seems to have good improvements in local logloss value; however, LB is skyrocketed due to overfitting."
        },
        {
          "id": 747960,
          "postDate": "2020-02-17T03:38:48.013Z",
          "content": "<p>well, I fell sad for you. I want to help but since \"no private sharing outside of team\", and my teammate will be mad at me if I merge with you(I'm not the leader either so). :-(</p>",
          "rawMarkdown": "well, I fell sad for you. I want to help but since \"no private sharing outside of team\", and my teammate will be mad at me if I merge with you(I'm not the leader either so). :-("
        },
        {
          "id": 747962,
          "postDate": "2020-02-17T03:43:05.290Z",
          "content": "<p>Ok, it's fine, I know the rules. BTW, thanks!</p>",
          "rawMarkdown": "Ok, it's fine, I know the rules. BTW, thanks!"
        },
        {
          "id": 747966,
          "postDate": "2020-02-17T03:57:13.817Z",
          "content": "<p>btw, good idea. One thing you can do to boost is to change your model architecture, data preparation. A good model should be able to get under 0.69314(in training loss) after around 100 samples.\nIf your training/validation loss were all good but LB is really bad, you probably want to look into the inference part.</p>",
          "rawMarkdown": "btw, good idea. One thing you can do to boost is to change your model architecture, data preparation. A good model should be able to get under 0.69314(in training loss) after around 100 samples.\nIf your training/validation loss were all good but LB is really bad, you probably want to look into the inference part.",
          "votes": 2
        }
      ]
    },
    {
      "id": 747574,
      "postDate": "2020-02-16T15:28:56.357Z",
      "content": "<p>Good approach</p>",
      "rawMarkdown": "Good approach",
      "votes": 1,
      "replies": [
        {
          "id": 747576,
          "postDate": "2020-02-16T15:31:42.263Z",
          "content": "<p>Thanks! <a href=\"/prashant4m\">@prashant4m</a> 😁 </p>",
          "rawMarkdown": "Thanks! @prashant4m 😁 "
        }
      ]
    },
    {
      "id": 747445,
      "postDate": "2020-02-16T12:56:42.937Z",
      "content": "<p>Well done for continuing to think outside the box. I like the idea of stacking a sequence of images. The challenge will be order then matters and we cannot presume the same order will appear in the test set. One approach you might want to consider is feeding the network pairs or triplets of images so it becomes order invariant. For example, if you have 10 images in a set feed it all combinations. The other consideration is not all faces are fake. A clip may be marked fake but it can contain multiple faces with one or more real and/or fake faces. You have to control for this scenario which is my current dilemma.</p>",
      "rawMarkdown": "Well done for continuing to think outside the box. I like the idea of stacking a sequence of images. The challenge will be order then matters and we cannot presume the same order will appear in the test set. One approach you might want to consider is feeding the network pairs or triplets of images so it becomes order invariant. For example, if you have 10 images in a set feed it all combinations. The other consideration is not all faces are fake. A clip may be marked fake but it can contain multiple faces with one or more real and/or fake faces. You have to control for this scenario which is my current dilemma.",
      "votes": 1,
      "replies": [
        {
          "id": 747556,
          "postDate": "2020-02-16T15:06:21.013Z",
          "content": "<p>Good idea <a href=\"/maralski\">@maralski</a>, so I should change my current architecture a little bit to see if it can learn better from the training dataset. And I also aware of the problems you and, maybe, many other competitors face when dealing with this kind of data, I will be careful and try to think some simple solutions to overcome them, at least temporarily.</p>",
          "rawMarkdown": "Good idea @maralski, so I should change my current architecture a little bit to see if it can learn better from the training dataset. And I also aware of the problems you and, maybe, many other competitors face when dealing with this kind of data, I will be careful and try to think some simple solutions to overcome them, at least temporarily.",
          "votes": 1
        }
      ]
    },
    {
      "id": 748164,
      "postDate": "2020-02-17T07:55:37.697Z",
      "content": "<p>I think many (including myself) are overfitting because we are using too many images that are very similar. Need to look into identifying common images (near duplicates) and excluding them, particularly from the majority class.</p>",
      "rawMarkdown": "I think many (including myself) are overfitting because we are using too many images that are very similar. Need to look into identifying common images (near duplicates) and excluding them, particularly from the majority class.",
      "votes": 2,
      "replies": [
        {
          "id": 748390,
          "postDate": "2020-02-17T13:24:14.463Z",
          "content": "<p>Yep, you're right, <a href=\"/maralski\">@maralski</a>! That what I've seen when examining the preprocessed data. We should find better ways to sample frames or faces, from the original videos, to prevent this problem.</p>",
          "rawMarkdown": "Yep, you're right, @maralski! That what I've seen when examining the preprocessed data. We should find better ways to sample frames or faces, from the original videos, to prevent this problem."
        }
      ]
    },
    {
      "id": 767986,
      "postDate": "2020-03-10T10:44:20.677Z",
      "content": "<p>Thank you for sharing. </p>",
      "rawMarkdown": "Thank you for sharing. ",
      "votes": 1,
      "isDeleted": true,
      "replies": [
        {
          "id": 768102,
          "postDate": "2020-03-10T13:23:53.353Z",
          "content": "<p>You're welcome!</p>",
          "rawMarkdown": "You're welcome!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 747326,
      "author_name": "Raju Kumar Mishra",
      "author_url": "",
      "post_date": "2020-02-16T09:18:17.340000",
      "content": "<p>Nice</p>",
      "votes": 3,
      "replies": [
        {
          "id": 747328,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-16T09:22:09.853000",
          "content": "<p>Thanks! <a href=\"/dataraj\">@dataraj</a> 😊 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 747994,
      "author_name": "Emre Bayram",
      "author_url": "",
      "post_date": "2020-02-17T05:09:10.497000",
      "content": "<p>I am new to deep learning; trying to learn. And luckily during my learning process i saw this competition. When you have a goal it is faster to learn. So it is helping me a lot. \n(also the fastai course series encouraged me and helped me to understand some concepts). So my comments might be just noob comments. Read it accordingly:)</p>\n\n<p>The first thing I have tried in this competition was a very similar approach; I stacked the images and used 3d version of the functions. But I have faced some problems.\nThe first problem was; for the 3d version; I was not able to find a good pre-trained model to use. I had to train from scratch. I think this is a big disadvantage.\nThe second problem I faced was choosing the conv kernel size; when it is 3d; it is getting complicated because on the plane the neighbor pixels have a different meaning but accross the planes it has a different meaning(probably the network would figure out the different meanings but still) so wasn't sure about choosing a cubic kernel size.\nThe third issue was I wasn't still persuaded myself if the relationship between frames are really that meaningful to say fake or not? Does the extra contribution worth the effort?</p>\n\n<p>Anyways later I decided to go with 2d network; I got better results this way(during predictions I took the mean of the per frame results).</p>\n\n<p>But I still did not throw away the stacked faces idea; if score from my 2d approach stops improving i might re-visit the idea of the stacked faces.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 748383,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-17T13:20:14.050000",
          "content": "<p>Hi <a href=\"/emrebayram\">@emrebayram</a>,</p>\n\n<p>Awesome! Good to hear that someone has tried 3D classification approach. I agree with you, training the model from scratch is really a big disadvantage, except you have tons of labeled data, lots of computing power, and time of course. Yet, I think it's good to do so, you can learn new interesting things, gain more insights, etc. many many benefits from this job. I think you should publish this solution at the end of the competition, this will definitely help the community a lot.</p>\n\n<p>BTW, good luck 😉.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 747626,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-02-16T16:54:52.140000",
      "content": "<p>Well here's a problem. Some videos have 1 face, some videos have 2 faces. You cannot determine the exact input shape. Besides, sometimes videos have 1 fake face, the other one is real.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 747942,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-17T03:08:31.783000",
          "content": "<p>Hi <a href=\"/unkownhihi\">@unkownhihi</a>,</p>\n\n<p>Yes, I know about this problem, so I have designed a simple strategy to sample face images from each video and if I cannot get enough faces the model needs, I'll assign a default value to this video in the final submission.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747945,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-17T03:12:10.167000",
          "content": "<p>And about \"sometimes videos have 1 fake face, the other one is real,\"  I also aware of this problem; however, I've not come up with any ideas yet. Do you have any suggestions 😃?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747946,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-17T03:12:49.383000",
          "content": "<p>but from my observation, most videos have 1 face, instead of 2 faces.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747954,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-17T03:25:16.870000",
          "content": "<p>I mean, I will sample face images throughout the whole video's frames.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747955,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-17T03:27:11.847000",
          "content": "<p>Nice! What's the LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747959,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-17T03:32:54.793000",
          "content": "<p>It hasn't worked well yet, the best I can get is 0.79725 at early epochs. When trying to fit the model further with the dataset, it seems to have good improvements in local logloss value; however, LB is skyrocketed due to overfitting.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747960,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-17T03:38:48.013000",
          "content": "<p>well, I fell sad for you. I want to help but since \"no private sharing outside of team\", and my teammate will be mad at me if I merge with you(I'm not the leader either so). :-(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747962,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-17T03:43:05.290000",
          "content": "<p>Ok, it's fine, I know the rules. BTW, thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747966,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-17T03:57:13.817000",
          "content": "<p>btw, good idea. One thing you can do to boost is to change your model architecture, data preparation. A good model should be able to get under 0.69314(in training loss) after around 100 samples.\nIf your training/validation loss were all good but LB is really bad, you probably want to look into the inference part.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 747574,
      "author_name": "Prashant Kumar Maurya",
      "author_url": "",
      "post_date": "2020-02-16T15:28:56.357000",
      "content": "<p>Good approach</p>",
      "votes": 1,
      "replies": [
        {
          "id": 747576,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-16T15:31:42.263000",
          "content": "<p>Thanks! <a href=\"/prashant4m\">@prashant4m</a> 😁 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 747445,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-02-16T12:56:42.937000",
      "content": "<p>Well done for continuing to think outside the box. I like the idea of stacking a sequence of images. The challenge will be order then matters and we cannot presume the same order will appear in the test set. One approach you might want to consider is feeding the network pairs or triplets of images so it becomes order invariant. For example, if you have 10 images in a set feed it all combinations. The other consideration is not all faces are fake. A clip may be marked fake but it can contain multiple faces with one or more real and/or fake faces. You have to control for this scenario which is my current dilemma.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 747556,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-16T15:06:21.013000",
          "content": "<p>Good idea <a href=\"/maralski\">@maralski</a>, so I should change my current architecture a little bit to see if it can learn better from the training dataset. And I also aware of the problems you and, maybe, many other competitors face when dealing with this kind of data, I will be careful and try to think some simple solutions to overcome them, at least temporarily.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 748164,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-02-17T07:55:37.697000",
      "content": "<p>I think many (including myself) are overfitting because we are using too many images that are very similar. Need to look into identifying common images (near duplicates) and excluding them, particularly from the majority class.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 748390,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-02-17T13:24:14.463000",
          "content": "<p>Yep, you're right, <a href=\"/maralski\">@maralski</a>! That what I've seen when examining the preprocessed data. We should find better ways to sample frames or faces, from the original videos, to prevent this problem.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 767986,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-10T10:44:20.677000",
      "content": "<p>Thank you for sharing. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 768102,
          "author_name": "Hieu Phung",
          "author_url": "",
          "post_date": "2020-03-10T13:23:53.353000",
          "content": "<p>You're welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "747249": "It seems that many teams are doing well in this competition, yet some still struggle to get started. For learning and research purposes, I propose a simple solution named `multiface classifier` to deal with the problem in this competition.\n\nNote:\n* This solution only used for learning and research purposes, and I do not guarantee that it will produce good results.\n* All the datasets I used to train the classifier are well prepared beforehand, and you can find the full list of them in the [*Other useful datasets*](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954) discussion.\n\n---\n## Idea\nSometimes, it is hard to determine a video if it is `real` or `fake` by only using a single face appears in this video, or separately classify each face then combine the results in some ways, for example averaging, to predict the label of the input video. I think we can give the classifier more meaningful information by feeding it with multiple face images at once (the strategy to sample these images will be discussed later.) Then, based on a sequence of faces, the classifier can give a better judgment on the video it gets.\n\n---\n## Multiface's general diagram\n![diagram](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1853329%2Ffa31c1ff391ce738e8da7273a9620153%2Fmultiface.png?generation=1581835338165826&amp;alt=media)\n\n---\n## Implementation\nI have prepared a `kernel` named [*DFDC-Multiface-Training*](https://www.kaggle.com/phunghieu/dfdc-multiface-training) dedicated to realizing this idea. Along with it, I have also created the [*DFDC-Multiface-Inference*](https://www.kaggle.com/phunghieu/dfdc-multiface-inference) kernel for the inference process.\n\nTo start with, I've tried ResNet18 as the classifier and chose to use 5 sampled faces in each video to classify it.\n* Each chosen face image will be preprocessed separately then stacked together depth-wise (along the third axis) into a single tensor before fed into the model.\n* In the training process, I will generate a uniform random sample of size 5 (the number of faces.) Note: I am not sure random sampling can help prevent overfitting so further experiments must be conducted.\n* In the validation process, I will not choose input faces randomly like before; instead, I use a different strategy to obtain these images. I will try to get enough faces throughout the video by evenly spaced sampling; if I cannot get enough in the first run, I will continue this strategy but with a little shift (or stride) in the interval [*start, stop*] to get different faces if possible, and continue this process until the fifth try (just a hyper-parameter to prevent infinite loop.)\n* The model will be led by the Focal Loss and optimized by the Adam algorithm.\n* In the inference process, I will loop through all test videos and try to get face images by using the same strategy as I have applied to the validation process above. The only difference is instead of having well-prepared data, I must run a face-detector, the same as I used to prepare the training dataset in the [*Data Preparation*](https://www.kaggle.com/phunghieu/deepfake-detection-face-extractor) kernel, to directly extract faces from each frame of one input video. If I fail to get enough faces from a video, I will mark it as `invalid` and assign a `default predicted value` (probability) to this video.\n\n---\nIf you have any questions or suggestions, please let me know!\n\nThanks for reading!",
    "747326": "Nice",
    "747994": "I am new to deep learning; trying to learn. And luckily during my learning process i saw this competition. When you have a goal it is faster to learn. So it is helping me a lot. \n(also the fastai course series encouraged me and helped me to understand some concepts). So my comments might be just noob comments. Read it accordingly:)\n\nThe first thing I have tried in this competition was a very similar approach; I stacked the images and used 3d version of the functions. But I have faced some problems.\nThe first problem was; for the 3d version; I was not able to find a good pre-trained model to use. I had to train from scratch. I think this is a big disadvantage.\nThe second problem I faced was choosing the conv kernel size; when it is 3d; it is getting complicated because on the plane the neighbor pixels have a different meaning but accross the planes it has a different meaning(probably the network would figure out the different meanings but still) so wasn't sure about choosing a cubic kernel size.\nThe third issue was I wasn't still persuaded myself if the relationship between frames are really that meaningful to say fake or not? Does the extra contribution worth the effort?\n\nAnyways later I decided to go with 2d network; I got better results this way(during predictions I took the mean of the per frame results).\n\nBut I still did not throw away the stacked faces idea; if score from my 2d approach stops improving i might re-visit the idea of the stacked faces.",
    "747626": "Well here's a problem. Some videos have 1 face, some videos have 2 faces. You cannot determine the exact input shape. Besides, sometimes videos have 1 fake face, the other one is real.",
    "747574": "Good approach",
    "747445": "Well done for continuing to think outside the box. I like the idea of stacking a sequence of images. The challenge will be order then matters and we cannot presume the same order will appear in the test set. One approach you might want to consider is feeding the network pairs or triplets of images so it becomes order invariant. For example, if you have 10 images in a set feed it all combinations. The other consideration is not all faces are fake. A clip may be marked fake but it can contain multiple faces with one or more real and/or fake faces. You have to control for this scenario which is my current dilemma.",
    "748164": "I think many (including myself) are overfitting because we are using too many images that are very similar. Need to look into identifying common images (near duplicates) and excluding them, particularly from the majority class.",
    "767986": "Thank you for sharing. "
  }
}