{
  "id": 138953,
  "title": "Tips. How to get top 10%(for now)",
  "url": "/competitions/deepfake-detection-challenge/discussion/138953",
  "author_name": "Victor Paslay",
  "post_date": "2020-03-26T21:52:01.162000",
  "votes": 36,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Hi! I joined competition relatively late - 3 weeks ago and made only 15 submissions for now. Since I did not have much computer vision experience I am quite happy that I was able to climb to 0.36 zone without ensembling - with just one model. Let me share with you some things which I needed to know to get this position. You still have time. \nWhat was useful:\n- Proper validation and splitting strategy: I do folder-wise validation but also I added some proper data generation connected with the origin of fake video. \n- Clipping. Dramatically changed my public results.\n- Augmentation - not only flipping but some augmentations like/or similar to cutmix, for instance.\n- I do not know if it helped because I use it from the beginning: I interpolated face positions in the frames where the detector was not able to detect the face. For example, if detector did not find a face on frame 0 and 1, I just take the position of face from frame 2.</p>\n\n<p>What did not work for me well:\n- Using additional real faces dataset(unfortunately! :( .\n- LRCNN. Moreover - things like concatenating results for each frame in a video, then passing this through fully connected layers also did not work.\n- Excluding faces which have low confidence(like 0.03), I use Single Shot Detector.</p>\n\n<p>I use EfficientNet B1 and it takes nearly 1 hour to train it on my dataset and get the score I have now.</p>",
  "messages": [
    {
      "id": 787547,
      "postDate": "2020-03-26T21:52:01.163Z",
      "content": "<p>Hi! I joined competition relatively late - 3 weeks ago and made only 15 submissions for now. Since I did not have much computer vision experience I am quite happy that I was able to climb to 0.36 zone without ensembling - with just one model. Let me share with you some things which I needed to know to get this position. You still have time. \nWhat was useful:\n- Proper validation and splitting strategy: I do folder-wise validation but also I added some proper data generation connected with the origin of fake video. \n- Clipping. Dramatically changed my public results.\n- Augmentation - not only flipping but some augmentations like/or similar to cutmix, for instance.\n- I do not know if it helped because I use it from the beginning: I interpolated face positions in the frames where the detector was not able to detect the face. For example, if detector did not find a face on frame 0 and 1, I just take the position of face from frame 2.</p>\n\n<p>What did not work for me well:\n- Using additional real faces dataset(unfortunately! :( .\n- LRCNN. Moreover - things like concatenating results for each frame in a video, then passing this through fully connected layers also did not work.\n- Excluding faces which have low confidence(like 0.03), I use Single Shot Detector.</p>\n\n<p>I use EfficientNet B1 and it takes nearly 1 hour to train it on my dataset and get the score I have now.</p>",
      "rawMarkdown": "Hi! I joined competition relatively late - 3 weeks ago and made only 15 submissions for now. Since I did not have much computer vision experience I am quite happy that I was able to climb to 0.36 zone without ensembling - with just one model. Let me share with you some things which I needed to know to get this position. You still have time. \nWhat was useful:\n- Proper validation and splitting strategy: I do folder-wise validation but also I added some proper data generation connected with the origin of fake video. \n- Clipping. Dramatically changed my public results.\n- Augmentation - not only flipping but some augmentations like/or similar to cutmix, for instance.\n- I do not know if it helped because I use it from the beginning: I interpolated face positions in the frames where the detector was not able to detect the face. For example, if detector did not find a face on frame 0 and 1, I just take the position of face from frame 2.\n\nWhat did not work for me well:\n- Using additional real faces dataset(unfortunately! :( .\n- LRCNN. Moreover - things like concatenating results for each frame in a video, then passing this through fully connected layers also did not work.\n- Excluding faces which have low confidence(like 0.03), I use Single Shot Detector.\n\nI use EfficientNet B1 and it takes nearly 1 hour to train it on my dataset and get the score I have now.",
      "votes": 35
    },
    {
      "id": 787638,
      "postDate": "2020-03-27T00:25:30.370Z",
      "content": "<p>Thanks for sharing! I still have not gotten a single model down to .36, you should try ensembling it gave me a large boost!\nSome questions:\nHow much has clipping given you? Any clipping for me always hurts the score.\nHow many frames do you use for inference?\nI haven't tried any complicated augmentations, does your hard augmentation give a good boost over simple augmentations?\nHave you tried cleaning the dataset like removing false positives or removing bad videos (multiple actors, bad fakes)?\nDoes your model overfit? Do you trust your validation scores? How do you choose a model to submit?</p>",
      "rawMarkdown": "Thanks for sharing! I still have not gotten a single model down to .36, you should try ensembling it gave me a large boost!\nSome questions:\nHow much has clipping given you? Any clipping for me always hurts the score.\nHow many frames do you use for inference?\nI haven't tried any complicated augmentations, does your hard augmentation give a good boost over simple augmentations?\nHave you tried cleaning the dataset like removing false positives or removing bad videos (multiple actors, bad fakes)?\nDoes your model overfit? Do you trust your validation scores? How do you choose a model to submit?",
      "votes": 5,
      "replies": [
        {
          "id": 787949,
          "postDate": "2020-03-27T08:47:08.960Z",
          "content": "<p>1) I am not sure about this result - clipping may worse it. Previously I had a model which got 0.47 before clipping and 0.41 after. I have not tried this submission without clipping.\n2) 20. The more the better but after 20 I did not find significant improvement on validation score.\n3) Yes, it gives.\n4) Nope, and what is funny - picking only good detections does not make any difference or even make it worse.\n5) My val score 0.24. In 3-5 submissions I found that my val score correlates with public score result. I pick the models with the best validation scores. </p>",
          "rawMarkdown": "1) I am not sure about this result - clipping may worse it. Previously I had a model which got 0.47 before clipping and 0.41 after. I have not tried this submission without clipping.\n2) 20. The more the better but after 20 I did not find significant improvement on validation score.\n3) Yes, it gives.\n4) Nope, and what is funny - picking only good detections does not make any difference or even make it worse.\n5) My val score 0.24. In 3-5 submissions I found that my val score correlates with public score result. I pick the models with the best validation scores. \n",
          "votes": 4
        }
      ]
    },
    {
      "id": 787581,
      "postDate": "2020-03-26T22:33:59.173Z",
      "content": "<p>Good work! Hope you get into the silver zone.</p>\n\n<p>what do you mean by this?</p>\n\n<blockquote>\n  <p>I added some proper data generation connected with the origin of fake video</p>\n</blockquote>",
      "rawMarkdown": "Good work! Hope you get into the silver zone.\n\nwhat do you mean by this?\n&gt; I added some proper data generation connected with the origin of fake video",
      "votes": 3,
      "replies": [
        {
          "id": 787943,
          "postDate": "2020-03-27T08:37:18.933Z",
          "content": "<p>Thanks. Can not go into details, but the fake videos are chosen carefully.</p>",
          "rawMarkdown": "Thanks. Can not go into details, but the fake videos are chosen carefully.",
          "votes": 1
        }
      ]
    },
    {
      "id": 787718,
      "postDate": "2020-03-27T02:22:34.590Z",
      "content": "<p>So, if your GPU is 1080TI and your batch size is 32,  5 minutes an epoch means you just use about 40k images to train.  what's more, you said the fake video is undersample, the data has 20k real videos. So I want to know if you just take one frame pre video?</p>",
      "rawMarkdown": "So, if your GPU is 1080TI and your batch size is 32,  5 minutes an epoch means you just use about 40k images to train.  what's more, you said the fake video is undersample, the data has 20k real videos. So I want to know if you just take one frame pre video?",
      "votes": 1,
      "replies": [
        {
          "id": 788334,
          "postDate": "2020-03-27T15:23:07.320Z",
          "content": "<p>You think in the right direction :) The process is a bit more complicated but the idea is what you have said.</p>",
          "rawMarkdown": "You think in the right direction :) The process is a bit more complicated but the idea is what you have said."
        },
        {
          "id": 788800,
          "postDate": "2020-03-28T02:55:19.457Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        }
      ]
    },
    {
      "id": 787556,
      "postDate": "2020-03-26T21:59:37.537Z",
      "content": "<p>what gpu do you use? 1 hour for one epoch? </p>",
      "rawMarkdown": "what gpu do you use? 1 hour for one epoch? ",
      "votes": 1,
      "replies": [
        {
          "id": 787557,
          "postDate": "2020-03-26T22:01:53.880Z",
          "content": "<p>GTX 1080 Ti. No, 1 epoch takes like 5 min. I have all face images on my SSD.\nBut most probably we have different epochs :)</p>",
          "rawMarkdown": "GTX 1080 Ti. No, 1 epoch takes like 5 min. I have all face images on my SSD.\nBut most probably we have different epochs :)",
          "votes": 1
        },
        {
          "id": 787565,
          "postDate": "2020-03-26T22:10:44.303Z",
          "content": "<p>5min per epoch is pretty sweet! i guess you froze the pretrain weights right? </p>",
          "rawMarkdown": "5min per epoch is pretty sweet! i guess you froze the pretrain weights right? "
        },
        {
          "id": 787574,
          "postDate": "2020-03-26T22:22:18.167Z",
          "content": "<p>No, batch size=32, images are 224x224x3, but the fakes are undersampled.</p>",
          "rawMarkdown": "No, batch size=32, images are 224x224x3, but the fakes are undersampled.",
          "votes": 2
        },
        {
          "id": 787587,
          "postDate": "2020-03-26T22:41:35.053Z",
          "content": "<p>emmm. I must do something wrong... </p>",
          "rawMarkdown": "emmm. I must do something wrong... "
        },
        {
          "id": 790100,
          "postDate": "2020-03-29T09:44:13.507Z",
          "content": "<p>I think it also depends on how many frames you cut from a video.</p>",
          "rawMarkdown": "I think it also depends on how many frames you cut from a video."
        }
      ]
    },
    {
      "id": 791213,
      "postDate": "2020-03-30T07:37:52.240Z",
      "content": "<p>Thanks a lot, mate! I've tampered with all of the techniques you've mentioned, but never gave them a proper number of attempts, I guess. I've used clipping at the very beginning but dropped it as it gave no boost to my first (mediocre) models, detectors and datasets. \nSo when I became stuck,  I haven't thought of using those again as I'd discarded them long ago. Looks like running through those tweaks once again got me some decent points here. Cheers!</p>",
      "rawMarkdown": "Thanks a lot, mate! I've tampered with all of the techniques you've mentioned, but never gave them a proper number of attempts, I guess. I've used clipping at the very beginning but dropped it as it gave no boost to my first (mediocre) models, detectors and datasets. \nSo when I became stuck,  I haven't thought of using those again as I'd discarded them long ago. Looks like running through those tweaks once again got me some decent points here. Cheers!",
      "votes": 2
    },
    {
      "id": 790496,
      "postDate": "2020-03-29T16:17:57.313Z",
      "content": "<p>Thanks for sharing! My approach is almost the same except for non-traditional augmentation. I predict what we do in terms of original-fake sampling is very similar :) I just implemented an augmentation idea which I think might help - will see kernel is still running :)  </p>",
      "rawMarkdown": "Thanks for sharing! My approach is almost the same except for non-traditional augmentation. I predict what we do in terms of original-fake sampling is very similar :) I just implemented an augmentation idea which I think might help - will see kernel is still running :)  "
    },
    {
      "id": 789049,
      "postDate": "2020-03-28T10:02:03.273Z",
      "content": "<p>thanks for sharing such a wonderful strategy</p>",
      "rawMarkdown": "thanks for sharing such a wonderful strategy"
    },
    {
      "id": 788171,
      "postDate": "2020-03-27T12:47:16.267Z",
      "content": "<p>Gr8 tips</p>",
      "rawMarkdown": "Gr8 tips"
    },
    {
      "id": 787746,
      "postDate": "2020-03-27T03:10:46.650Z",
      "content": "<p>Thanks for sharing!! Can you explain more about your clipping technique?</p>",
      "rawMarkdown": "Thanks for sharing!! Can you explain more about your clipping technique?"
    },
    {
      "id": 791041,
      "postDate": "2020-03-30T04:04:38.127Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 788322,
      "postDate": "2020-03-27T15:13:41.780Z",
      "content": "<p>Thanks for sharing!!</p>",
      "rawMarkdown": "Thanks for sharing!!",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 787638,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-03-27T00:25:30.370000",
      "content": "<p>Thanks for sharing! I still have not gotten a single model down to .36, you should try ensembling it gave me a large boost!\nSome questions:\nHow much has clipping given you? Any clipping for me always hurts the score.\nHow many frames do you use for inference?\nI haven't tried any complicated augmentations, does your hard augmentation give a good boost over simple augmentations?\nHave you tried cleaning the dataset like removing false positives or removing bad videos (multiple actors, bad fakes)?\nDoes your model overfit? Do you trust your validation scores? How do you choose a model to submit?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 787949,
          "author_name": "Victor Paslay",
          "author_url": "",
          "post_date": "2020-03-27T08:47:08.960000",
          "content": "<p>1) I am not sure about this result - clipping may worse it. Previously I had a model which got 0.47 before clipping and 0.41 after. I have not tried this submission without clipping.\n2) 20. The more the better but after 20 I did not find significant improvement on validation score.\n3) Yes, it gives.\n4) Nope, and what is funny - picking only good detections does not make any difference or even make it worse.\n5) My val score 0.24. In 3-5 submissions I found that my val score correlates with public score result. I pick the models with the best validation scores. </p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 787581,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2020-03-26T22:33:59.173000",
      "content": "<p>Good work! Hope you get into the silver zone.</p>\n\n<p>what do you mean by this?</p>\n\n<blockquote>\n  <p>I added some proper data generation connected with the origin of fake video</p>\n</blockquote>",
      "votes": 3,
      "replies": [
        {
          "id": 787943,
          "author_name": "Victor Paslay",
          "author_url": "",
          "post_date": "2020-03-27T08:37:18.933000",
          "content": "<p>Thanks. Can not go into details, but the fake videos are chosen carefully.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 787718,
      "author_name": "ZhiHao Wang",
      "author_url": "",
      "post_date": "2020-03-27T02:22:34.590000",
      "content": "<p>So, if your GPU is 1080TI and your batch size is 32,  5 minutes an epoch means you just use about 40k images to train.  what's more, you said the fake video is undersample, the data has 20k real videos. So I want to know if you just take one frame pre video?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 788334,
          "author_name": "Victor Paslay",
          "author_url": "",
          "post_date": "2020-03-27T15:23:07.320000",
          "content": "<p>You think in the right direction :) The process is a bit more complicated but the idea is what you have said.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 788800,
          "author_name": "ZhiHao Wang",
          "author_url": "",
          "post_date": "2020-03-28T02:55:19.457000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 787556,
      "author_name": "yimacs",
      "author_url": "",
      "post_date": "2020-03-26T21:59:37.537000",
      "content": "<p>what gpu do you use? 1 hour for one epoch? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 787557,
          "author_name": "Victor Paslay",
          "author_url": "",
          "post_date": "2020-03-26T22:01:53.880000",
          "content": "<p>GTX 1080 Ti. No, 1 epoch takes like 5 min. I have all face images on my SSD.\nBut most probably we have different epochs :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 787565,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2020-03-26T22:10:44.303000",
          "content": "<p>5min per epoch is pretty sweet! i guess you froze the pretrain weights right? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 787574,
          "author_name": "Victor Paslay",
          "author_url": "",
          "post_date": "2020-03-26T22:22:18.167000",
          "content": "<p>No, batch size=32, images are 224x224x3, but the fakes are undersampled.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 787587,
          "author_name": "yimacs",
          "author_url": "",
          "post_date": "2020-03-26T22:41:35.053000",
          "content": "<p>emmm. I must do something wrong... </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790100,
          "author_name": "Ilia Zaitsev",
          "author_url": "",
          "post_date": "2020-03-29T09:44:13.507000",
          "content": "<p>I think it also depends on how many frames you cut from a video.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 791213,
      "author_name": "petya",
      "author_url": "",
      "post_date": "2020-03-30T07:37:52.240000",
      "content": "<p>Thanks a lot, mate! I've tampered with all of the techniques you've mentioned, but never gave them a proper number of attempts, I guess. I've used clipping at the very beginning but dropped it as it gave no boost to my first (mediocre) models, detectors and datasets. \nSo when I became stuck,  I haven't thought of using those again as I'd discarded them long ago. Looks like running through those tweaks once again got me some decent points here. Cheers!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 790496,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-03-29T16:17:57.313000",
      "content": "<p>Thanks for sharing! My approach is almost the same except for non-traditional augmentation. I predict what we do in terms of original-fake sampling is very similar :) I just implemented an augmentation idea which I think might help - will see kernel is still running :)  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 789049,
      "author_name": "Aviral Gupta",
      "author_url": "",
      "post_date": "2020-03-28T10:02:03.273000",
      "content": "<p>thanks for sharing such a wonderful strategy</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 788171,
      "author_name": "ritik dubey",
      "author_url": "",
      "post_date": "2020-03-27T12:47:16.267000",
      "content": "<p>Gr8 tips</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 787746,
      "author_name": "Seungki Kim",
      "author_url": "",
      "post_date": "2020-03-27T03:10:46.650000",
      "content": "<p>Thanks for sharing!! Can you explain more about your clipping technique?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 791041,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-30T04:04:38.127000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 788322,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-27T15:13:41.780000",
      "content": "<p>Thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "787547": "Hi! I joined competition relatively late - 3 weeks ago and made only 15 submissions for now. Since I did not have much computer vision experience I am quite happy that I was able to climb to 0.36 zone without ensembling - with just one model. Let me share with you some things which I needed to know to get this position. You still have time. \nWhat was useful:\n- Proper validation and splitting strategy: I do folder-wise validation but also I added some proper data generation connected with the origin of fake video. \n- Clipping. Dramatically changed my public results.\n- Augmentation - not only flipping but some augmentations like/or similar to cutmix, for instance.\n- I do not know if it helped because I use it from the beginning: I interpolated face positions in the frames where the detector was not able to detect the face. For example, if detector did not find a face on frame 0 and 1, I just take the position of face from frame 2.\n\nWhat did not work for me well:\n- Using additional real faces dataset(unfortunately! :( .\n- LRCNN. Moreover - things like concatenating results for each frame in a video, then passing this through fully connected layers also did not work.\n- Excluding faces which have low confidence(like 0.03), I use Single Shot Detector.\n\nI use EfficientNet B1 and it takes nearly 1 hour to train it on my dataset and get the score I have now.",
    "787638": "Thanks for sharing! I still have not gotten a single model down to .36, you should try ensembling it gave me a large boost!\nSome questions:\nHow much has clipping given you? Any clipping for me always hurts the score.\nHow many frames do you use for inference?\nI haven't tried any complicated augmentations, does your hard augmentation give a good boost over simple augmentations?\nHave you tried cleaning the dataset like removing false positives or removing bad videos (multiple actors, bad fakes)?\nDoes your model overfit? Do you trust your validation scores? How do you choose a model to submit?",
    "787581": "Good work! Hope you get into the silver zone.\n\nwhat do you mean by this?\n&gt; I added some proper data generation connected with the origin of fake video",
    "787718": "So, if your GPU is 1080TI and your batch size is 32,  5 minutes an epoch means you just use about 40k images to train.  what's more, you said the fake video is undersample, the data has 20k real videos. So I want to know if you just take one frame pre video?",
    "787556": "what gpu do you use? 1 hour for one epoch? ",
    "791213": "Thanks a lot, mate! I've tampered with all of the techniques you've mentioned, but never gave them a proper number of attempts, I guess. I've used clipping at the very beginning but dropped it as it gave no boost to my first (mediocre) models, detectors and datasets. \nSo when I became stuck,  I haven't thought of using those again as I'd discarded them long ago. Looks like running through those tweaks once again got me some decent points here. Cheers!",
    "790496": "Thanks for sharing! My approach is almost the same except for non-traditional augmentation. I predict what we do in terms of original-fake sampling is very similar :) I just implemented an augmentation idea which I think might help - will see kernel is still running :)  ",
    "789049": "thanks for sharing such a wonderful strategy",
    "788171": "Gr8 tips",
    "787746": "Thanks for sharing!! Can you explain more about your clipping technique?",
    "791041": "",
    "788322": "Thanks for sharing!!"
  }
}