{
  "id": 140191,
  "title": "Model Thread - What model/s did you use?",
  "url": "/competitions/deepfake-detection-challenge/discussion/140191",
  "author_name": "Thomas Sherk",
  "post_date": "2020-03-31T20:19:17.995000",
  "votes": 22,
  "comment_count": 66,
  "views": 0,
  "content": "<p>I'm interested to see what others have done. We created a CNN-LSTM, an audio-model, and other models but didn't have great success. </p>",
  "messages": [
    {
      "id": 793121,
      "postDate": "2020-03-31T20:19:17.997Z",
      "content": "<p>I'm interested to see what others have done. We created a CNN-LSTM, an audio-model, and other models but didn't have great success. </p>",
      "rawMarkdown": "I'm interested to see what others have done. We created a CNN-LSTM, an audio-model, and other models but didn't have great success. ",
      "votes": 22
    },
    {
      "id": 793323,
      "postDate": "2020-04-01T00:23:38.763Z",
      "content": "<p>A TON honestly. </p>\n\n<p>Start off with our preprocessing method with multiplier system, take the same amount of frames in real video(equally distributed) as the number of the corresponding fake that video have, making data balance without data loss. (note its single frame model)</p>\n\n<p>Our best single model is effb6 with linear 256, relu 128, and sigmoid 1 head. We got 0.33LB for that model. We ensembled mixnet, efficientnetb1, b3, b6 with all different paddings and resolution, seresnet 101(couldn't fit seresnet 154 in that 1GB constraint) and got our current score. Random Hflip, random brightness for argumentation. </p>\n\n<p>What we tried but didn't work:\n- simple LRCN: since it couldn't fit in our multiplier system, it have data loss, which reached its limit before single frame models\n- LRCN with CNN froze with our best single single frame model weight preloaded(using our old data without multiplier system): It had the best CV, but unfortunately, it had a pretty bad LB(0.38)\n- filter out frames with face that have too much yaw: Its very likely due to low quality input data that have compression noise\n- Ensemble more model: This is not a really valid reason, but our aws credits ran out.............\n- cutout</p>\n\n<p>What we didn't try, but could've helped:</p>\n\n<ul>\n<li>SWA</li>\n<li>More LR schedule(<em>here I want to note that 1e-4 is our magic lr that other lr and several lr schedules couldn't surpass</em></li>\n<li>dual shot face detector(its very very slow, so we didn't try, and we possibly can't fit all those models in 9 hour constraint)</li>\n<li>cutmix, gridmask</li>\n<li>face embedding based validation set</li>\n<li>proper kfolds(this is really a shame to say we didn't even get a chance to do proper kfolds, still, our quota ran out before we could actually make the data)</li>\n<li>audio model</li>\n</ul>\n\n<p>I'm really excited about the topper's solution.... Good luck for the private LB! </p>",
      "rawMarkdown": "A TON honestly. \n\nStart off with our preprocessing method with multiplier system, take the same amount of frames in real video(equally distributed) as the number of the corresponding fake that video have, making data balance without data loss. (note its single frame model)\n\nOur best single model is effb6 with linear 256, relu 128, and sigmoid 1 head. We got 0.33LB for that model. We ensembled mixnet, efficientnetb1, b3, b6 with all different paddings and resolution, seresnet 101(couldn't fit seresnet 154 in that 1GB constraint) and got our current score. Random Hflip, random brightness for argumentation. \n\nWhat we tried but didn't work:\n- simple LRCN: since it couldn't fit in our multiplier system, it have data loss, which reached its limit before single frame models\n- LRCN with CNN froze with our best single single frame model weight preloaded(using our old data without multiplier system): It had the best CV, but unfortunately, it had a pretty bad LB(0.38)\n- filter out frames with face that have too much yaw: Its very likely due to low quality input data that have compression noise\n- Ensemble more model: This is not a really valid reason, but our aws credits ran out.............\n- cutout\n\nWhat we didn't try, but could've helped:\n\n- SWA\n- More LR schedule(*here I want to note that 1e-4 is our magic lr that other lr and several lr schedules couldn't surpass*\n- dual shot face detector(its very very slow, so we didn't try, and we possibly can't fit all those models in 9 hour constraint)\n- cutmix, gridmask\n- face embedding based validation set\n- proper kfolds(this is really a shame to say we didn't even get a chance to do proper kfolds, still, our quota ran out before we could actually make the data)\n- audio model\n\nI'm really excited about the topper's solution.... Good luck for the private LB! ",
      "votes": 5,
      "replies": [
        {
          "id": 793328,
          "postDate": "2020-04-01T00:28:48.040Z",
          "content": "<p>Not sure if I understand what you mean in the first part. Are you saying that if for example you pull 90 frames from a real video then you would pull 90 frames across the ~5 fake videos instead of 90 frames from each fake video. That way you have the same number of unique frames for both real and fake for a given video?</p>\n\n<p>Is that what you did? </p>",
          "rawMarkdown": "Not sure if I understand what you mean in the first part. Are you saying that if for example you pull 90 frames from a real video then you would pull 90 frames across the ~5 fake videos instead of 90 frames from each fake video. That way you have the same number of unique frames for both real and fake for a given video?\n\nIs that what you did? "
        },
        {
          "id": 793329,
          "postDate": "2020-04-01T00:29:53.543Z",
          "content": "<p>say a real video have 5 fakes, you will pull 5 frames from the real video and 1 frame from each fake video. Sorry for consufion.</p>",
          "rawMarkdown": "say a real video have 5 fakes, you will pull 5 frames from the real video and 1 frame from each fake video. Sorry for consufion."
        }
      ]
    },
    {
      "id": 793194,
      "postDate": "2020-03-31T21:48:46.227Z",
      "content": "<p>I actually used a very sophisticated solution, despite my poor showing (still not sure what went wrong).</p>\n\n<ol>\n<li><p>Input: developed a RAM-efficient multi-threaded cv2-based reader: one thread is servicing the GPU, while 1 or more others are reading the next video. This showed processing speed-up on my workstation.</p></li>\n<li><p>Face detection. Very fast and very accurate based on a modified version of Faceboxes. Did all the resizing on the GPU and kept the selected faces on the GPU to save time. Also, did face detection in batches (small enough to not cause OOM errors on a P100)</p></li>\n<li><p>Face tracking: interpolation based, to get rid of the jitter in the NN-based detections and to fill in the rare gaps.</p></li>\n<li><p>Actual deepfake detection. Tried single-face ones and short sequences. Ended up with a 2-frame sequences. Trained only on one such sequence for each video. Chose WideResnet-50-2 as the backbone. ResNext was marginally better but much slower.</p></li>\n<li><p>Augmentation: flips during training and I also re-encoded 2/3 of the videos as the DPDC Preview paper suggested.</p></li>\n<li><p>Inference: 256 frames per video (but I had to give up the multithreaded (1) reader, which turned out to be slower on Kaggle's machines than a single-threaded one)</p></li>\n<li><p>Ensemble: 2 models, trained on different frames and augmentations.</p></li>\n</ol>\n\n<p>Two ideas I wish I had time to try:</p>\n\n<ol>\n<li><p>Use sudden changes in face embeddings. This was actually the first idea I had, but went on to actually implement other things. Now I think it was promising, at least as part of an ensemble.</p></li>\n<li><p>Highlight pixelation / noise. You can do that via a high-pass filter or using NNs. Also, use the original frame resolution, but only a patch of the image. This approach does not care about the content, but only the \"noise\".</p></li>\n</ol>\n\n<p>I obviously made a mistake of actually getting started on this contest very late, and on top of that, certain events and the pandemic were a distraction.</p>",
      "rawMarkdown": "I actually used a very sophisticated solution, despite my poor showing (still not sure what went wrong).\n\n1. Input: developed a RAM-efficient multi-threaded cv2-based reader: one thread is servicing the GPU, while 1 or more others are reading the next video. This showed processing speed-up on my workstation.\n\n2. Face detection. Very fast and very accurate based on a modified version of Faceboxes. Did all the resizing on the GPU and kept the selected faces on the GPU to save time. Also, did face detection in batches (small enough to not cause OOM errors on a P100)\n\n3. Face tracking: interpolation based, to get rid of the jitter in the NN-based detections and to fill in the rare gaps.\n\n4. Actual deepfake detection. Tried single-face ones and short sequences. Ended up with a 2-frame sequences. Trained only on one such sequence for each video. Chose WideResnet-50-2 as the backbone. ResNext was marginally better but much slower.\n\n5. Augmentation: flips during training and I also re-encoded 2/3 of the videos as the DPDC Preview paper suggested.\n\n6. Inference: 256 frames per video (but I had to give up the multithreaded (1) reader, which turned out to be slower on Kaggle's machines than a single-threaded one)\n\n7. Ensemble: 2 models, trained on different frames and augmentations.\n\nTwo ideas I wish I had time to try:\n\n1. Use sudden changes in face embeddings. This was actually the first idea I had, but went on to actually implement other things. Now I think it was promising, at least as part of an ensemble.\n\n2. Highlight pixelation / noise. You can do that via a high-pass filter or using NNs. Also, use the original frame resolution, but only a patch of the image. This approach does not care about the content, but only the \"noise\".\n\nI obviously made a mistake of actually getting started on this contest very late, and on top of that, certain events and the pandemic were a distraction.",
      "votes": 5,
      "replies": [
        {
          "id": 793203,
          "postDate": "2020-03-31T22:01:33.163Z",
          "content": "<p>I tried the change in face embedding. It wasn't very good approach, if it makes you feel better about it. Face embedding is designed to actually be relaxed about minor changes. Also differences in embedding while the actor was turning were more than the difference between real and fake.  Could have been just my bad. </p>",
          "rawMarkdown": "I tried the change in face embedding. It wasn't very good approach, if it makes you feel better about it. Face embedding is designed to actually be relaxed about minor changes. Also differences in embedding while the actor was turning were more than the difference between real and fake.  Could have been just my bad. ",
          "votes": 1
        },
        {
          "id": 793206,
          "postDate": "2020-03-31T22:07:47.937Z",
          "content": "<p>An approch i tried a few times and failed miserably was to train on difference between consecutive frames. Visual examination shows that the difference is very different with fake and real. However, the nn did not pick this up and accuracy and loss were near random. </p>",
          "rawMarkdown": "An approch i tried a few times and failed miserably was to train on difference between consecutive frames. Visual examination shows that the difference is very different with fake and real. However, the nn did not pick this up and accuracy and loss were near random. ",
          "votes": 2
        },
        {
          "id": 793221,
          "postDate": "2020-03-31T22:20:03.717Z",
          "content": "<p><a href=\"/moshel\">@moshel</a> </p>\n\n<p>Face embeddings are designed to be insensitive to changes in pose and expression, yes. However, the deepfakes exhibited sudden changes in \"identity\". </p>\n\n<p>During training, one can use the difference between the fake video and the real one to find only those frames where this difference changes suddenly (A sudden difference difference, if you will)</p>\n\n<p>This happens when the forgery algorithm fails at face detection (Side note: They should have used a good detector + interpolation, as I did) </p>\n\n<p>Now, if you trained the model on just those frames as \"fakes\", I think it would learn to detect the unnatural \"facial ticks\" many deepfakes exhibit in some parts of the video.</p>\n\n<p>Well, I haven't actually implemented this, so this is a guess.</p>\n\n<p>The 2-frame model I used was also intended to be sensitive to these, but it apparently wasn't very good. However, it was pre-trained on ImageNet, not the facial triplet-loss, and I chose the frames randomly, not as described above.</p>",
          "rawMarkdown": "@moshel \n\nFace embeddings are designed to be insensitive to changes in pose and expression, yes. However, the deepfakes exhibited sudden changes in \"identity\". \n\nDuring training, one can use the difference between the fake video and the real one to find only those frames where this difference changes suddenly (A sudden difference difference, if you will)\n\nThis happens when the forgery algorithm fails at face detection (Side note: They should have used a good detector + interpolation, as I did) \n\nNow, if you trained the model on just those frames as \"fakes\", I think it would learn to detect the unnatural \"facial ticks\" many deepfakes exhibit in some parts of the video.\n\nWell, I haven't actually implemented this, so this is a guess.\n\nThe 2-frame model I used was also intended to be sensitive to these, but it apparently wasn't very good. However, it was pre-trained on ImageNet, not the facial triplet-loss, and I chose the frames randomly, not as described above."
        },
        {
          "id": 793234,
          "postDate": "2020-03-31T22:34:46.293Z",
          "content": "<p><a href=\"/moshel\">@moshel</a> I tried similar things. Lstm over the embeddings generated from the face net models. Never able to get anything out of it. I also tried frame differencing and optical flow. Surprising it could not extract signal from those representations but that's what I found in my experiments. </p>",
          "rawMarkdown": "@moshel I tried similar things. Lstm over the embeddings generated from the face net models. Never able to get anything out of it. I also tried frame differencing and optical flow. Surprising it could not extract signal from those representations but that's what I found in my experiments. "
        }
      ]
    },
    {
      "id": 793338,
      "postDate": "2020-04-01T00:41:28.727Z",
      "content": "<p>The 2 models we submitted:\n1) Ensemble of 9 EffB3 + Flip TTA + 64 Frame inference\n2) Ensemble of 19 models including EffB1, EffB2, EffB3, Xception, ResNext + Flip TTA + 64 frame inference.</p>\n\n<p>Not gonna think about it for a few days, God knows what the private set looks like 😂 </p>",
      "rawMarkdown": "The 2 models we submitted:\n1) Ensemble of 9 EffB3 + Flip TTA + 64 Frame inference\n2) Ensemble of 19 models including EffB1, EffB2, EffB3, Xception, ResNext + Flip TTA + 64 frame inference.\n\nNot gonna think about it for a few days, God knows what the private set looks like 😂 ",
      "votes": 3,
      "replies": [
        {
          "id": 793339,
          "postDate": "2020-04-01T00:43:02.730Z",
          "content": "<p>what is the resolution you used? What is your balancing technique?</p>",
          "rawMarkdown": "what is the resolution you used? What is your balancing technique?"
        },
        {
          "id": 793344,
          "postDate": "2020-04-01T00:50:08.350Z",
          "content": "<p>Resolution is 224. I tried others but did not work for us. Some components models are ensembles done by using median instead of mean. Balancing: change fake corresponding to real frame every epoch.</p>",
          "rawMarkdown": "Resolution is 224. I tried others but did not work for us. Some components models are ensembles done by using median instead of mean. Balancing: change fake corresponding to real frame every epoch.",
          "votes": 1
        },
        {
          "id": 793792,
          "postDate": "2020-04-01T09:32:18.970Z",
          "content": "<p>How did you ensemble? Just averaging, fixed hold-out stacking, CV-stacking?\nAnd what is  Flip TTA + 64 Frame inference?\nThank you!</p>",
          "rawMarkdown": "How did you ensemble? Just averaging, fixed hold-out stacking, CV-stacking?\nAnd what is  Flip TTA + 64 Frame inference?\nThank you!"
        },
        {
          "id": 794255,
          "postDate": "2020-04-01T17:25:47.223Z",
          "content": "<p>Ensemble: Combination of weighted mean and median of the model predictions\nFlip TTA and 64 frames: During inference, I used 64 frames + 64 horizontally flipped frames</p>",
          "rawMarkdown": "Ensemble: Combination of weighted mean and median of the model predictions\nFlip TTA and 64 frames: During inference, I used 64 frames + 64 horizontally flipped frames\n"
        },
        {
          "id": 794285,
          "postDate": "2020-04-01T17:57:47.107Z",
          "content": "<p>Thank you! And how exactly us the weight calculated?</p>",
          "rawMarkdown": "Thank you! And how exactly us the weight calculated?"
        },
        {
          "id": 794358,
          "postDate": "2020-04-01T18:54:48.393Z",
          "content": "<p>Higher weights to models performing better in LB. Exact numbers are hit and trial :D Regarding median, they don't have weights, so less headache!</p>",
          "rawMarkdown": "Higher weights to models performing better in LB. Exact numbers are hit and trial :D Regarding median, they don't have weights, so less headache!",
          "votes": 1
        }
      ]
    },
    {
      "id": 793307,
      "postDate": "2020-03-31T23:57:18.370Z",
      "content": "<p><strong>Network:</strong>\n- 5 x Efficientnet b3 ns w/ pre-trained weights (<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a>)\n- 300 x 300 face cutout\n- Inference on 60 frames (np.linspace)\n- Lots off augmentation (<a href=\"https://github.com/albumentations-team/albumentations\">https://github.com/albumentations-team/albumentations</a>)\n- brightness, contrast, gamma, jpeg compression, blur (inc motion), gaussian noise, scaling down\n- Dropout (0.2)</p>\n\n<p><strong>Training:</strong>\n- Pytorch\n- Batch size 8\n- 2-3 epochs per model\n- SGD\n- Momentum 0.9\n- Weight Decay 1e-4\n- LR 0.001</p>\n\n<p><strong>Training set:</strong>\n- 5 folds\n- 0-9, 10-19, 20-29, 30-39, 40-49 validation sets\n- 5 frames from real video and 5 frames across set of fake videos where frame position align between real and fake\n- ~134k training and ~37k validation images per fold</p>\n\n<p><strong>Face detector:</strong>\n- RetinaFace (<a href=\"https://github.com/biubug6/Pytorch_Retinaface\">https://github.com/biubug6/Pytorch_Retinaface</a>)\n- Resized largest axis down to 960 before detection\n- Center crop\n- Add 10% margin\n- Extract 5 frames all faces\n- Nms 0.4</p>\n\n<p><strong>Training rig:</strong>\n- Dual socket Xeon\n- 192 GB RAM\n- 1 x RTX 2080Ti\n- 1 x GTX 1080\n- 2 TB SSD</p>\n\n<p><strong>What helped:</strong>\n- Augment validation set as well (downscale, compression)\n- Center the face cutout and add margin\n- Focus on high quality faces (detection threshold &gt;= 0.99)\n- Remove fake faces that had very low and very high structural similarity to equivalent real face frame (skimage)\n- Balance real/fake training set (described in previous post)\n- Avoid just averaging faces when video has multiple faces</p>\n\n<p><strong>What didn't work</strong>\n- Label smoothing\n- Weighted average during inference weighted on # real/fake faces per video</p>\n\n<p><strong>What I didn't try (I suspect these would have helped)</strong>\n- LSTM\n- ffmpeg over cv2\n- Custom CNN trained from scratch</p>",
      "rawMarkdown": "**Network:**\n- 5 x Efficientnet b3 ns w/ pre-trained weights (https://github.com/rwightman/pytorch-image-models)\n- 300 x 300 face cutout\n- Inference on 60 frames (np.linspace)\n- Lots off augmentation (https://github.com/albumentations-team/albumentations)\n- brightness, contrast, gamma, jpeg compression, blur (inc motion), gaussian noise, scaling down\n- Dropout (0.2)\n\n**Training:**\n- Pytorch\n- Batch size 8\n- 2-3 epochs per model\n- SGD\n- Momentum 0.9\n- Weight Decay 1e-4\n- LR 0.001\n\n**Training set:**\n- 5 folds\n- 0-9, 10-19, 20-29, 30-39, 40-49 validation sets\n- 5 frames from real video and 5 frames across set of fake videos where frame position align between real and fake\n- ~134k training and ~37k validation images per fold\n\n**Face detector:**\n- RetinaFace (https://github.com/biubug6/Pytorch_Retinaface)\n- Resized largest axis down to 960 before detection\n- Center crop\n- Add 10% margin\n- Extract 5 frames all faces\n- Nms 0.4\n\n**Training rig:**\n- Dual socket Xeon\n- 192 GB RAM\n- 1 x RTX 2080Ti\n- 1 x GTX 1080\n- 2 TB SSD\n\n**What helped:**\n- Augment validation set as well (downscale, compression)\n- Center the face cutout and add margin\n- Focus on high quality faces (detection threshold &gt;= 0.99)\n- Remove fake faces that had very low and very high structural similarity to equivalent real face frame (skimage)\n- Balance real/fake training set (described in previous post)\n- Avoid just averaging faces when video has multiple faces\n\n**What didn't work**\n- Label smoothing\n- Weighted average during inference weighted on # real/fake faces per video\n\n**What I didn't try (I suspect these would have helped)**\n- LSTM\n- ffmpeg over cv2\n- Custom CNN trained from scratch",
      "votes": 3,
      "replies": [
        {
          "id": 796533,
          "postDate": "2020-04-03T17:03:57.193Z",
          "content": "<p><a href=\"/maralski\">@maralski</a> 192MB RAM?!?!?! how were you even able to get a proper OS running?</p>",
          "rawMarkdown": "@maralski 192MB RAM?!?!?! how were you even able to get a proper OS running?"
        },
        {
          "id": 796551,
          "postDate": "2020-04-03T17:23:24.777Z",
          "content": "<p>He is joking come on, thats 192 gb.</p>",
          "rawMarkdown": "He is joking come on, thats 192 gb."
        },
        {
          "id": 797148,
          "postDate": "2020-04-04T08:36:27.143Z",
          "content": "<p>Can I ask what's the differences between your 5 models?</p>",
          "rawMarkdown": "Can I ask what's the differences between your 5 models?"
        },
        {
          "id": 797157,
          "postDate": "2020-04-04T08:51:17.190Z",
          "content": "<p>Hehe my boo boo 192 GB</p>",
          "rawMarkdown": "Hehe my boo boo 192 GB"
        },
        {
          "id": 797159,
          "postDate": "2020-04-04T08:52:38.380Z",
          "content": "<p><a href=\"/suthidasukhonn\">@suthidasukhonn</a> each was trained on a different fold</p>",
          "rawMarkdown": "@suthidasukhonn each was trained on a different fold"
        }
      ]
    },
    {
      "id": 793129,
      "postDate": "2020-03-31T20:25:00.353Z",
      "content": "<p>I couldn't get lstm/gru to work. It overfitted no matter what I did.\nMy model is a single model, on full frame (not faces) of seresnext50. </p>",
      "rawMarkdown": "I couldn't get lstm/gru to work. It overfitted no matter what I did.\nMy model is a single model, on full frame (not faces) of seresnext50. ",
      "votes": 3,
      "replies": [
        {
          "id": 793135,
          "postDate": "2020-03-31T20:29:34.353Z",
          "content": "<p>Do you rescale the frame? How many frames did you take per video during training? How many during inference?</p>",
          "rawMarkdown": "Do you rescale the frame? How many frames did you take per video during training? How many during inference?"
        },
        {
          "id": 793147,
          "postDate": "2020-03-31T20:44:53.583Z",
          "content": "<p>We had an issue of over-fitting, I added regularization to the LSTM layer but that did not seem to help. If I had more time, I probably would be able to figure out what was wrong but it is what it is.</p>",
          "rawMarkdown": "We had an issue of over-fitting, I added regularization to the LSTM layer but that did not seem to help. If I had more time, I probably would be able to figure out what was wrong but it is what it is.",
          "votes": 1
        },
        {
          "id": 793167,
          "postDate": "2020-03-31T21:03:34.020Z",
          "content": "<p>Yes, rescaled to 512. 1 frame per movie during training, 30 frames inference.</p>",
          "rawMarkdown": "Yes, rescaled to 512. 1 frame per movie during training, 30 frames inference."
        },
        {
          "id": 793172,
          "postDate": "2020-03-31T21:09:30.277Z",
          "content": "<p>I had a feeling gru was the way to go, but got distracted and didn't have time to figure out why it is over fitting so quickly. I am sure it is something stupid... Basically gru should be able to \"see\" that frames are changing differently in fake and real. However, it focused on something else and the process of training these is very time consuming. Would be interesting to see what top places did. I bet they used lstm/gru. </p>",
          "rawMarkdown": "I had a feeling gru was the way to go, but got distracted and didn't have time to figure out why it is over fitting so quickly. I am sure it is something stupid... Basically gru should be able to \"see\" that frames are changing differently in fake and real. However, it focused on something else and the process of training these is very time consuming. Would be interesting to see what top places did. I bet they used lstm/gru. ",
          "votes": 1
        },
        {
          "id": 793192,
          "postDate": "2020-03-31T21:46:03.630Z",
          "content": "<p><a href=\"/moshel\">@moshel</a> And your full frame (no face) single model (frames resized to 512) scored 0.342?</p>",
          "rawMarkdown": "@moshel And your full frame (no face) single model (frames resized to 512) scored 0.342?"
        },
        {
          "id": 793196,
          "postDate": "2020-03-31T21:49:32.353Z",
          "content": "<p>Yes! </p>",
          "rawMarkdown": "Yes! ",
          "votes": 1
        },
        {
          "id": 793197,
          "postDate": "2020-03-31T21:51:41.180Z",
          "content": "<p>Naturally the devil is in the details. Selecting the cv, training generator that balance real /fake while preventing over fitting and augmentation. </p>",
          "rawMarkdown": "Naturally the devil is in the details. Selecting the cv, training generator that balance real /fake while preventing over fitting and augmentation. ",
          "votes": 3
        },
        {
          "id": 793198,
          "postDate": "2020-03-31T21:52:39.600Z",
          "content": "<p>Also the score calculation function was very important. </p>",
          "rawMarkdown": "Also the score calculation function was very important. ",
          "votes": 1
        },
        {
          "id": 793205,
          "postDate": "2020-03-31T22:07:00.173Z",
          "content": "<p>It's a good result. We've tried such model but with smaller resize. During CV we've noticed that it learnt interesting features but we stopped as CV was not so good for us.</p>",
          "rawMarkdown": "It's a good result. We've tried such model but with smaller resize. During CV we've noticed that it learnt interesting features but we stopped as CV was not so good for us."
        },
        {
          "id": 793213,
          "postDate": "2020-03-31T22:11:45.740Z",
          "content": "<p>I specialize in squeezing the most out of my models ;)\nYour score is pretty amazing. Care to share a bit about your approach? It's too late to submit anyways. </p>",
          "rawMarkdown": "I specialize in squeezing the most out of my models ;)\nYour score is pretty amazing. Care to share a bit about your approach? It's too late to submit anyways. "
        },
        {
          "id": 793227,
          "postDate": "2020-03-31T22:27:18.880Z",
          "content": "<p>Yes, will do it tomorrow (time to sleep here).</p>",
          "rawMarkdown": "Yes, will do it tomorrow (time to sleep here)."
        }
      ]
    },
    {
      "id": 795132,
      "postDate": "2020-04-02T12:57:08.177Z",
      "content": "<h1>Train</h1>\n\n<ul>\n<li>Face Detector: margin(0.05) based DSFD/BlazeFace</li>\n<li>Dataloader: under sample、5 frame/video</li>\n<li>Augmentation: ImageCompression、RandomBrightness、GaussianNoise、RandomCrop, etc</li>\n<li>Backbone: NoisyStudent pretrained EfficientNet</li>\n<li>Attention: similar as C-FAN, which adaptively aggregates deep feature vectors into a single vector</li>\n<li>Two stage training:\nStage-1: train a base CNN on real/fake images;\nStage-2: freeze the pre-trained base CNN to extract the embeddings of faces, and train attention module on videos (5 frame / video). </li>\n</ul>\n\n<h1>Val</h1>\n\n<ul>\n<li>DFDC: split by folder-wise and original videos (CV: 0.14、LB: 0.30)</li>\n<li>FaceForensics++: a generalized validation set, which only include deepfake, faceswap related videos (CV: 0.22、LB: 0.27) \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F3ba58aefd964b742a098d933ed1ef42a%2F1.png?generation=1585831881937088&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<h1>Tried but not work</h1>\n\n<ul>\n<li>Audio model</li>\n<li>LSTM, 3D Cov, TSM</li>\n<li>Spatial attention、Multi-scale attention、Multi-head attention、Prob attentioin</li>\n<li>Adversarial attack</li>\n<li>Multi-patch fusion</li>\n<li>YCbCr color space transformation</li>\n<li>use GMM to fuse n frames probabilities</li>\n<li>ArcFace、SphereFace、Contrastive Loss、Focal Loss、Label Smooth, etc</li>\n<li>decrease resolution</li>\n<li>remove multi-face videos</li>\n<li>Face Alignment</li>\n<li>Initialize efficientnet with pretrained weights on the face dataset</li>\n<li>Face Xray</li>\n<li>Frequency Learning</li>\n</ul>",
      "rawMarkdown": "# Train\n- Face Detector: margin(0.05) based DSFD/BlazeFace\n- Dataloader: under sample、5 frame/video\n- Augmentation: ImageCompression、RandomBrightness、GaussianNoise、RandomCrop, etc\n- Backbone: NoisyStudent pretrained EfficientNet\n- Attention: similar as C-FAN, which adaptively aggregates deep feature vectors into a single vector\n- Two stage training:\nStage-1: train a base CNN on real/fake images;\nStage-2: freeze the pre-trained base CNN to extract the embeddings of faces, and train attention module on videos (5 frame / video). \n\n# Val\n- DFDC: split by folder-wise and original videos (CV: 0.14、LB: 0.30)\n- FaceForensics++: a generalized validation set, which only include deepfake, faceswap related videos (CV: 0.22、LB: 0.27) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F3ba58aefd964b742a098d933ed1ef42a%2F1.png?generation=1585831881937088&amp;alt=media)\n\n# Tried but not work\n- Audio model\n- LSTM, 3D Cov, TSM\n- Spatial attention、Multi-scale attention、Multi-head attention、Prob attentioin\n- Adversarial attack\n- Multi-patch fusion\n- YCbCr color space transformation\n- use GMM to fuse n frames probabilities\n- ArcFace、SphereFace、Contrastive Loss、Focal Loss、Label Smooth, etc\n- decrease resolution\n- remove multi-face videos\n- Face Alignment\n- Initialize efficientnet with pretrained weights on the face dataset\n- Face Xray\n- Frequency Learning",
      "votes": 4,
      "replies": [
        {
          "id": 795184,
          "postDate": "2020-04-02T13:45:09.993Z",
          "content": "<p>Great! Really want to see how attention is implemented</p>",
          "rawMarkdown": "Great! Really want to see how attention is implemented"
        },
        {
          "id": 795246,
          "postDate": "2020-04-02T14:52:58.200Z",
          "content": "<p>Cool, what are <strong>3D cov</strong> and <strong>TSM</strong>? I would also like to read about the attention solution. Thank you!</p>",
          "rawMarkdown": "Cool, what are **3D cov** and **TSM**? I would also like to read about the attention solution. Thank you!"
        },
        {
          "id": 795768,
          "postDate": "2020-04-03T02:59:46.467Z",
          "content": "<p>Hard Luck, we did try that exact stage 2 approach, and validation loss did go much better than ever!, but did not perform good in LB.</p>",
          "rawMarkdown": "Hard Luck, we did try that exact stage 2 approach, and validation loss did go much better than ever!, but did not perform good in LB."
        },
        {
          "id": 795788,
          "postDate": "2020-04-03T03:46:55.587Z",
          "content": "<p><a href=\"/azamatk\">@azamatk</a> <a href=\"/leonidboytsov\">@leonidboytsov</a> The attention module is similar to SE-block, and I will share the detailed solution after the busy days.</p>",
          "rawMarkdown": "@azamatk @leonidboytsov The attention module is similar to SE-block, and I will share the detailed solution after the busy days."
        },
        {
          "id": 795789,
          "postDate": "2020-04-03T03:49:45.253Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> We encountered the same problem and later found out that it was mainly due to overfitting. Therefore, we use the warm-up strategy and train only one epoch, which greatly improves the performance in both CV and LB.</p>",
          "rawMarkdown": "@harshitsheoran We encountered the same problem and later found out that it was mainly due to overfitting. Therefore, we use the warm-up strategy and train only one epoch, which greatly improves the performance in both CV and LB."
        },
        {
          "id": 795790,
          "postDate": "2020-04-03T03:51:05.027Z",
          "content": "<p>Now I am starting to hate myself, also in my case, there was not  a lot improvement like 0.12 to 0.11 from first epoch to best.</p>",
          "rawMarkdown": "Now I am starting to hate myself, also in my case, there was not  a lot improvement like 0.12 to 0.11 from first epoch to best."
        }
      ]
    },
    {
      "id": 793267,
      "postDate": "2020-03-31T23:16:14.107Z",
      "content": "<h2>Simple Solution</h2>\n\n<p>I joined the competition relatively late, last month but thanks to everyone contributing with kernels and discussion helped me get up to speed relatively easy. For me, the biggest challenge was to work with video data since it was the first time. But it was a great experience for learning efficient video batching approaches. Spent ~3 weeks on data and ~1 week on modeling.</p>\n\n<p>All my work can be obtained <a href=\"https://github.com/KeremTurgutlu/dfdc\">here</a>. I used <a href=\"https://github.com/fastai/nbdev\">nbdev</a> to speed up my code base development process using notebooks. I highly recommend for other competitions if you are more productive with jupyter notebooks but still need the power of modularizing your code.</p>\n\n<h3>Video Loading/Batching</h3>\n\n<p>I used decoder cpu version (thanks to <a href=\"https://www.kaggle.com/leighplt/decord-videoreader/notebook\">this</a>). Tried out DALI and decord gpu versions but they had issues in terms of memory leakage and I didn't have enough time to debug and also cpu version was fast enough for my pipeline. As for opencv video loading I tried both vanilla and multithreaded options but in aws p3 instance decoder cpu was a little bit faster (eventhough it wasn't the case in Kaggle kernel).</p>\n\n<h3>Detection</h3>\n\n<p>For detection I used and adopted mobilenet detector in <a href=\"https://github.com/biubug6/Pytorch_Retinaface\">here</a>. It's a very lightweight model in terms of memory, space and it's pretty fast during inference. I didn't have time to prepare a better dataset by dealing with potential false negatives/positives. That may have allowed a gain in overall model performance. I extracted ~30 frames (1/10 of each frame with equal intervals) for each video and stored only faces. I used the <a href=\"https://github.com/KeremTurgutlu/dfdc/blob/master/dfdc/face_detection/download_detect_crop_save.py\">script</a> which downloads a video chunk (1 of 50) does detections, saves cropped faces and finally deletes the processed chunk. This is repeated sequentially for all 50 chunks. This allowed me to save a lot of disk space. </p>\n\n<p>Each face crop enlarged by x1.3 after detection.</p>\n\n<h3>Validation</h3>\n\n<p>Discussions were really helpful in terms of deciding a proper validation set. Thanks to everyone who shared what worked and what didn't. This allowed me to save a lot of time and not to spend too much time on methods like grouping based on face embeddings. I finally used chunks 1-40 as training, 41-45 as validation and 46-50 as test. Once I decided on modeling I merged training and test for final training.</p>\n\n<h3>Data Sampling</h3>\n\n<p>Once I started training a baseline model I quickly noticed that a vanilla random batch sampling won't do good. So I decided on a custom batch sampler. In a batch, this sampler basically randomly picks n number of original videos and a random fake video for each of these n original videos. Then for each real-fake pair video same frame crop is used. This strategy helped to get a decent public LB score but still, the model was overfitting pretty bad and fast.</p>\n\n<p><strong>What didn't work:</strong> I also had another sampler that utilized real video and it's all corresponding fakes again using the same frame crop for each. This was significantly worse compared to balanced sampling (1 real - 1 random fake).</p>\n\n<h3>Regularization</h3>\n\n<p>I think the biggest challenge was avoiding overfitting and the potential risk of memorizing individual faces. At first I tried trivial data augmentation strategies.</p>\n\n<ul>\n<li>Crappification: downsampling image and upsample again to mimic low resolution/compression scenarios.</li>\n<li>Left right flipping</li>\n<li>Brightness change - darker or lighter</li>\n<li>Random zoom - x1.15 - x1.35</li>\n</ul>\n\n<p>Just 3 days ago I added a custom augmentation strategy similar to cutmix which is inspired by another <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection\">compeition</a> from <a href=\"/godaibo\">@godaibo</a> from his TWIML podcast interview.</p>\n\n<ul>\n<li>For a given face crop, I merge it vertically (50% - 50%) with another face crop from the same class. This gave a boost of 0.02 in pubic LB. The idea was to not memorize individual faces. The reason for 50% and vertical merge was assumptions based on the symmetry of fake alterations, e.g. either change of eyes, mouth ...etc.</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F558069%2F1244fc86ea6a89ecd9001c1844d48292%2FScreen%20Shot%202020-04-01%20at%205.29.36%20PM.png?generation=1585751421952949&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Also tried mixup but didn't improve the LB.</li>\n</ul>\n\n<h3>Models</h3>\n\n<ul>\n<li>Resnet34, EfficientNet b5, Efficient b7 - Efficient alone is enough to get my best score. I did simple averaging. Tried both with and without TTA but didn't get much of a boost. All models are finetuned using gradual unfreezing and 1-cycle policy. FocalLoss is used at a later stage of training to update model with harder samples. Also used simple clipping (0.01-0.99) just to be safe about very high false positives and very low false negatives.</li>\n</ul>\n\n<h3>What didn't work</h3>\n\n<ul>\n<li>LRCN based modeling</li>\n<li>1 real all fakes sampling</li>\n</ul>\n\n<p>Thanks to everyone for their great contributions! </p>\n\n<p>Thanks to organizers for contributing with data and the challenge, Kaggle for hosting and AWS for the compute.</p>\n\n<p>It was a great learning experience and I wish everyone luck with private LB :)</p>",
      "rawMarkdown": "## Simple Solution\n\nI joined the competition relatively late, last month but thanks to everyone contributing with kernels and discussion helped me get up to speed relatively easy. For me, the biggest challenge was to work with video data since it was the first time. But it was a great experience for learning efficient video batching approaches. Spent ~3 weeks on data and ~1 week on modeling.\n\nAll my work can be obtained [here](https://github.com/KeremTurgutlu/dfdc). I used [nbdev](https://github.com/fastai/nbdev) to speed up my code base development process using notebooks. I highly recommend for other competitions if you are more productive with jupyter notebooks but still need the power of modularizing your code.\n\n### Video Loading/Batching\n\nI used decoder cpu version (thanks to [this](https://www.kaggle.com/leighplt/decord-videoreader/notebook)). Tried out DALI and decord gpu versions but they had issues in terms of memory leakage and I didn't have enough time to debug and also cpu version was fast enough for my pipeline. As for opencv video loading I tried both vanilla and multithreaded options but in aws p3 instance decoder cpu was a little bit faster (eventhough it wasn't the case in Kaggle kernel).\n\n### Detection\n\nFor detection I used and adopted mobilenet detector in [here](https://github.com/biubug6/Pytorch_Retinaface). It's a very lightweight model in terms of memory, space and it's pretty fast during inference. I didn't have time to prepare a better dataset by dealing with potential false negatives/positives. That may have allowed a gain in overall model performance. I extracted ~30 frames (1/10 of each frame with equal intervals) for each video and stored only faces. I used the [script](https://github.com/KeremTurgutlu/dfdc/blob/master/dfdc/face_detection/download_detect_crop_save.py) which downloads a video chunk (1 of 50) does detections, saves cropped faces and finally deletes the processed chunk. This is repeated sequentially for all 50 chunks. This allowed me to save a lot of disk space. \n\nEach face crop enlarged by x1.3 after detection.\n\n### Validation\n\nDiscussions were really helpful in terms of deciding a proper validation set. Thanks to everyone who shared what worked and what didn't. This allowed me to save a lot of time and not to spend too much time on methods like grouping based on face embeddings. I finally used chunks 1-40 as training, 41-45 as validation and 46-50 as test. Once I decided on modeling I merged training and test for final training.\n\n### Data Sampling\n\nOnce I started training a baseline model I quickly noticed that a vanilla random batch sampling won't do good. So I decided on a custom batch sampler. In a batch, this sampler basically randomly picks n number of original videos and a random fake video for each of these n original videos. Then for each real-fake pair video same frame crop is used. This strategy helped to get a decent public LB score but still, the model was overfitting pretty bad and fast.\n\n**What didn't work:** I also had another sampler that utilized real video and it's all corresponding fakes again using the same frame crop for each. This was significantly worse compared to balanced sampling (1 real - 1 random fake).\n\n### Regularization\n\nI think the biggest challenge was avoiding overfitting and the potential risk of memorizing individual faces. At first I tried trivial data augmentation strategies.\n\n- Crappification: downsampling image and upsample again to mimic low resolution/compression scenarios.\n- Left right flipping\n- Brightness change - darker or lighter\n- Random zoom - x1.15 - x1.35\n\nJust 3 days ago I added a custom augmentation strategy similar to cutmix which is inspired by another [compeition](https://www.kaggle.com/c/state-farm-distracted-driver-detection) from @godaibo from his TWIML podcast interview.\n\n- For a given face crop, I merge it vertically (50% - 50%) with another face crop from the same class. This gave a boost of 0.02 in pubic LB. The idea was to not memorize individual faces. The reason for 50% and vertical merge was assumptions based on the symmetry of fake alterations, e.g. either change of eyes, mouth ...etc.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F558069%2F1244fc86ea6a89ecd9001c1844d48292%2FScreen%20Shot%202020-04-01%20at%205.29.36%20PM.png?generation=1585751421952949&amp;alt=media)\n\n- Also tried mixup but didn't improve the LB.\n\n### Models\n\n- Resnet34, EfficientNet b5, Efficient b7 - Efficient alone is enough to get my best score. I did simple averaging. Tried both with and without TTA but didn't get much of a boost. All models are finetuned using gradual unfreezing and 1-cycle policy. FocalLoss is used at a later stage of training to update model with harder samples. Also used simple clipping (0.01-0.99) just to be safe about very high false positives and very low false negatives.\n\n### What didn't work\n\n- LRCN based modeling\n- 1 real all fakes sampling\n\n\nThanks to everyone for their great contributions! \n\nThanks to organizers for contributing with data and the challenge, Kaggle for hosting and AWS for the compute.\n\nIt was a great learning experience and I wish everyone luck with private LB :)",
      "votes": 4
    },
    {
      "id": 793450,
      "postDate": "2020-04-01T03:18:25.927Z",
      "content": "<p>During this competition, I have been struggling with the gap between CV and LB. Finally, I cannot resolve this problem...</p>\n\n<p>Validation</p>\n\n<ul>\n<li>5 fold</li>\n<li>Stratified group k fold</li>\n<li>Stratified by the number of fake and set the group to folder.</li>\n<li>2/9 Downscale and 2/9 JpegCompression</li>\n<li>logloss weighted to have the same FAKE and ORIGINAL influences</li>\n<li>CV score is 0.16 and Public is 0.33 by this approach.</li>\n</ul>\n\n<p>Network</p>\n\n<ul>\n<li>Sample original video(A) and fake video (generated from A) in the same minibatch.\n<ul><li>randomly select fake video (from A) every epoch </li>\n<li>randomly select frame every epoch, but the frame index in the pair (A, fakeA) is same.</li></ul></li>\n<li>Face crop by SSFD (<a href=\"https://github.com/cs-giung/face-detection-pytorch\">https://github.com/cs-giung/face-detection-pytorch</a>)</li>\n<li>Stacking by predictions of 20 frames per video from efficientnet-b4 and se-resnext 50</li>\n<li>Train LightGBM by extract features of video from above predictions (mean, median , min, max, std, etc..)</li>\n</ul>\n\n<p>PostProcess</p>\n\n<ul>\n<li>There are large gap train dataset and Private dataset and model’s output tend to be low confidence. So, I applied postprocess as below.\n<ul><li>sigmoid (lgbm output) -&gt; logit -&gt; multiply s -&gt; sigmoid</li></ul></li>\n</ul>",
      "rawMarkdown": "During this competition, I have been struggling with the gap between CV and LB. Finally, I cannot resolve this problem...\n\nValidation\n\n- 5 fold\n- Stratified group k fold\n- Stratified by the number of fake and set the group to folder.\n- 2/9 Downscale and 2/9 JpegCompression\n- logloss weighted to have the same FAKE and ORIGINAL influences\n- CV score is 0.16 and Public is 0.33 by this approach.\n\nNetwork\n\n- Sample original video(A) and fake video (generated from A) in the same minibatch.\n  - randomly select fake video (from A) every epoch \n  - randomly select frame every epoch, but the frame index in the pair (A, fakeA) is same.\n- Face crop by SSFD (https://github.com/cs-giung/face-detection-pytorch)\n- Stacking by predictions of 20 frames per video from efficientnet-b4 and se-resnext 50\n- Train LightGBM by extract features of video from above predictions (mean, median , min, max, std, etc..)\n\nPostProcess\n\n- There are large gap train dataset and Private dataset and model’s output tend to be low confidence. So, I applied postprocess as below.\n  - sigmoid (lgbm output) -&gt; logit -&gt; multiply s -&gt; sigmoid",
      "votes": 1
    },
    {
      "id": 793320,
      "postDate": "2020-04-01T00:15:56.187Z",
      "content": "<p>EfficientNet with Transformer which I just implemented this week\n<a href=\"https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986\">https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986</a></p>\n\n<p>Failed to submit by 5 min. I was just hoping it would improve my terrible score :)</p>",
      "rawMarkdown": "EfficientNet with Transformer which I just implemented this week\nhttps://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986\n\nFailed to submit by 5 min. I was just hoping it would improve my terrible score :)",
      "votes": 1
    },
    {
      "id": 793302,
      "postDate": "2020-03-31T23:51:57.010Z",
      "content": "<p>Xception and efficientnetb1, the LB score is always 0.2 higher than my local validation score ... I didn't use augmentations(only horizontal flip).</p>",
      "rawMarkdown": "Xception and efficientnetb1, the LB score is always 0.2 higher than my local validation score ... I didn't use augmentations(only horizontal flip).",
      "votes": 1
    },
    {
      "id": 793910,
      "postDate": "2020-04-01T11:56:34.760Z",
      "content": "<p>We started working on this a bit late. Our solution was rather simple:</p>\n\n<ul>\n<li>Extracting faces with Retinaface detector</li>\n<li>Training an ensemble of 3 EfficientNet-B5s  (training time was ~3hrs on 1 RTX6000)</li>\n<li>Data Augmentation with JPEGCompression, ISONoise, Blur and DownScale</li>\n<li>Simple weighted crossentropy loss with [0.8,0.2] weights.</li>\n<li>No validation set used (Creating this seemed painful, since folders had mixed actors).</li>\n<li>Cap to [0.05,0.95] for predictions (to avoid big penalties)</li>\n</ul>\n\n<p>Our pipeline was probably very standard- faces extraction was rescaled to 224x224, our pipeline similar to the kernel posted by Human Analog. Our training code was simply applying the above tweaks applied to the pytorch/examples/imagenet/main.py with augmentations from Albumentations. </p>\n\n<p>Our rank increase was mostly in the last 5 days of the contest. I guess we were just relatively lucky? If so, hope that our luck is not too bad the private test set.. ~~</p>",
      "rawMarkdown": "We started working on this a bit late. Our solution was rather simple:\n\n- Extracting faces with Retinaface detector\n- Training an ensemble of 3 EfficientNet-B5s  (training time was ~3hrs on 1 RTX6000)\n- Data Augmentation with JPEGCompression, ISONoise, Blur and DownScale\n- Simple weighted crossentropy loss with [0.8,0.2] weights.\n- No validation set used (Creating this seemed painful, since folders had mixed actors).\n- Cap to [0.05,0.95] for predictions (to avoid big penalties)\n\nOur pipeline was probably very standard- faces extraction was rescaled to 224x224, our pipeline similar to the kernel posted by Human Analog. Our training code was simply applying the above tweaks applied to the pytorch/examples/imagenet/main.py with augmentations from Albumentations. \n\nOur rank increase was mostly in the last 5 days of the contest. I guess we were just relatively lucky? If so, hope that our luck is not too bad the private test set.. ~~",
      "votes": 2,
      "replies": [
        {
          "id": 793922,
          "postDate": "2020-04-01T12:04:53.893Z",
          "content": "<p>How did you decide when to stop training?</p>",
          "rawMarkdown": "How did you decide when to stop training?"
        },
        {
          "id": 793932,
          "postDate": "2020-04-01T12:18:58.477Z",
          "content": "<p>Since we never did a full pass over the dataset, the training accuracy was weakly indicative of the performance, and it seemed to stay at 94%ish, which seemed reasonably accurate but not overfitting. So, we stopped it at those iterations. </p>\n\n<p>Also, one more factor was that we could make it in time to try submit 2 submissions everyday (so 6 models train with 3 GPUs). Training finishing in 3-4 hours was perfect in timing too, and I could train and upload two batches of models everyday.</p>",
          "rawMarkdown": "Since we never did a full pass over the dataset, the training accuracy was weakly indicative of the performance, and it seemed to stay at 94%ish, which seemed reasonably accurate but not overfitting. So, we stopped it at those iterations. \n\nAlso, one more factor was that we could make it in time to try submit 2 submissions everyday (so 6 models train with 3 GPUs). Training finishing in 3-4 hours was perfect in timing too, and I could train and upload two batches of models everyday.",
          "votes": 1
        },
        {
          "id": 793949,
          "postDate": "2020-04-01T12:32:08.357Z",
          "content": "<p>Using leaderboard for early stopping, wow, I should have thought about it. 👍 </p>",
          "rawMarkdown": "Using leaderboard for early stopping, wow, I should have thought about it. 👍 ",
          "votes": 1
        },
        {
          "id": 794069,
          "postDate": "2020-04-01T14:31:17.030Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 793371,
      "postDate": "2020-04-01T01:27:37.563Z",
      "content": "<p>Ensemble of 9 EfB1  LRCNN, courtesy of <a href=\"/harshitsheoran\">@harshitsheoran</a> and <a href=\"/unkownhihi\">@unkownhihi</a> , with modified face extraction (preventing it from switching face), and all trained on different parts of the data (some random, some pair of true-fake). I lost track which one is which, I chose those with CV score ranging 0.1 - 0.2 roughly (validation on folder 45-49)\nTried very hard to get MobileNet to work, but it was futile (loss doesn't go below 0.35). Tried EfB0 as well, but doesn't work out. Xcetion and Densenet both too big (a.k.a. slow) and I'm scared of burning my AWS credit. </p>\n\n<p>I'm not sure if training using pair of true-fake is really useful , i wonder if anyone can shed light on this</p>",
      "rawMarkdown": "Ensemble of 9 EfB1  LRCNN, courtesy of @harshitsheoran and @unkownhihi , with modified face extraction (preventing it from switching face), and all trained on different parts of the data (some random, some pair of true-fake). I lost track which one is which, I chose those with CV score ranging 0.1 - 0.2 roughly (validation on folder 45-49)\nTried very hard to get MobileNet to work, but it was futile (loss doesn't go below 0.35). Tried EfB0 as well, but doesn't work out. Xcetion and Densenet both too big (a.k.a. slow) and I'm scared of burning my AWS credit. \n\nI'm not sure if training using pair of true-fake is really useful , i wonder if anyone can shed light on this\n",
      "votes": 2
    },
    {
      "id": 793461,
      "postDate": "2020-04-01T03:38:37.307Z",
      "content": "<p>After most of times struggling to get a good classifier model. I followed EfB1 + LSTM model by <a href=\"/unkownhihi\">@unkownhihi</a>. And I spent times creating a sequence dataset to feed the model. I use custom Keras Sequence, balanced pair real-fake with different fake for each epoch. I use 160x160 image size of face crop\nBut this also another struggle, could not get the model to get a good validation. My validation set is folder 0,1,2,43-49\n- with frozen backbone, training not able to converge\n- with retrain backbone, my model can converge well but badly overfit. then found bug in training data, since I used padding, some fake data just blank because could not find the images/faces\n- after fix the bug, still could not get good validation loss, somehow my model very confident to guess real as fake or the other way\n- changed  image size to 120x120, center crop from original 160x160, slightly better\n- tried RMSProp, Adam and SGD with different learning rates: RMSProp and Adam can converge fast, but validation loss is erratic. SGD is more stable and not able to converge well. My validation loss is 1.x and tried to submit one get 0.8 on LB\n- last night I change the model to:\n    y = TimeDistributed(cnn)(input_layer) \n    y = LSTM(256, return_sequences=False)(y) <br>\n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\nI used SGD with momentum and cyclic learning rate. I was not sure this going to be different and just leave it run. And this morning surprisingly, I got:\npoch 7/20\n1291/1291 [==============================] - 476s 369ms/step - loss: 0.1620 - acc: 0.9330 - val_loss: 0.0334 - val_acc: 0.8122\nI realize that this is not great and may not give me super LB increase, I am just happy I finally breaking my own many attemps hard to reach good validation. But have to wait for 22 more days to do late submission :( to know for sure</p>\n\n<p>what a journey... congrats to you all. thanks for sharing</p>",
      "rawMarkdown": "After most of times struggling to get a good classifier model. I followed EfB1 + LSTM model by @unkownhihi. And I spent times creating a sequence dataset to feed the model. I use custom Keras Sequence, balanced pair real-fake with different fake for each epoch. I use 160x160 image size of face crop\nBut this also another struggle, could not get the model to get a good validation. My validation set is folder 0,1,2,43-49\n- with frozen backbone, training not able to converge\n- with retrain backbone, my model can converge well but badly overfit. then found bug in training data, since I used padding, some fake data just blank because could not find the images/faces\n- after fix the bug, still could not get good validation loss, somehow my model very confident to guess real as fake or the other way\n- changed  image size to 120x120, center crop from original 160x160, slightly better\n- tried RMSProp, Adam and SGD with different learning rates: RMSProp and Adam can converge fast, but validation loss is erratic. SGD is more stable and not able to converge well. My validation loss is 1.x and tried to submit one get 0.8 on LB\n- last night I change the model to:\n    y = TimeDistributed(cnn)(input_layer) \n    y = LSTM(256, return_sequences=False)(y)    \n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\nI used SGD with momentum and cyclic learning rate. I was not sure this going to be different and just leave it run. And this morning surprisingly, I got:\npoch 7/20\n1291/1291 [==============================] - 476s 369ms/step - loss: 0.1620 - acc: 0.9330 - val_loss: 0.0334 - val_acc: 0.8122\nI realize that this is not great and may not give me super LB increase, I am just happy I finally breaking my own many attemps hard to reach good validation. But have to wait for 22 more days to do late submission :( to know for sure\n\nwhat a journey... congrats to you all. thanks for sharing\n",
      "replies": [
        {
          "id": 794000,
          "postDate": "2020-04-01T13:24:22.650Z",
          "content": "<p>for me SGD worked alone with momentum and weight_decay and lower the lr with factor 0.97\nadamw and 1-cycle worked fine together\nalways started with lr(batch*5e-04, eg 0.064 for batch 128)</p>",
          "rawMarkdown": "for me SGD worked alone with momentum and weight_decay and lower the lr with factor 0.97\nadamw and 1-cycle worked fine together\nalways started with lr(batch*5e-04, eg 0.064 for batch 128)"
        }
      ]
    },
    {
      "id": 793252,
      "postDate": "2020-03-31T22:55:43.090Z",
      "content": "<p>efficientnet b3 pretrained model (transfer learning)\n + dropout &amp; dense layers\n + softmax output.</p>\n\n<p>it showed 88%+ acc &amp; 0.4 loss in 10 mins of GPU run (batchsize = 160)</p>",
      "rawMarkdown": "efficientnet b3 pretrained model (transfer learning)\n + dropout &amp; dense layers\n + softmax output.\n\nit showed 88%+ acc &amp; 0.4 loss in 10 mins of GPU run (batchsize = 160)"
    },
    {
      "id": 793202,
      "postDate": "2020-03-31T22:01:25.153Z",
      "content": "<p>BlazeFace for face extraction\nfirst sub- stacked efficientnet b0ns+b1ns+b2ns+b3ns\nsecond sub - stacked(b0ns+b1ns+xception)+ensemble with best model of b1ns\nmixup, \nrandom cutout,\naugmentation, \n1-cycle with more runs,\nweighted average checkpoints\nNo CV but I hope that the large public testset was not too misleading</p>",
      "rawMarkdown": "BlazeFace for face extraction\nfirst sub- stacked efficientnet b0ns+b1ns+b2ns+b3ns\nsecond sub - stacked(b0ns+b1ns+xception)+ensemble with best model of b1ns\nmixup, \nrandom cutout,\naugmentation, \n1-cycle with more runs,\nweighted average checkpoints\nNo CV but I hope that the large public testset was not too misleading",
      "replies": [
        {
          "id": 793237,
          "postDate": "2020-03-31T22:36:47.737Z",
          "content": "<p>What do you mean by mixup, cutout, and 1-cycle?</p>",
          "rawMarkdown": "What do you mean by mixup, cutout, and 1-cycle?"
        },
        {
          "id": 793245,
          "postDate": "2020-03-31T22:42:58.530Z",
          "content": "<p>mixup and cutout are augmentation methods (a bit old-fashioned, apparently bested by \"cutmix\")</p>",
          "rawMarkdown": "mixup and cutout are augmentation methods (a bit old-fashioned, apparently bested by \"cutmix\")"
        },
        {
          "id": 793263,
          "postDate": "2020-03-31T23:09:58.370Z",
          "content": "<p>Those papers are both from 2017. Are we calling 2017 old fashioned already? haha</p>",
          "rawMarkdown": "Those papers are both from 2017. Are we calling 2017 old fashioned already? haha",
          "votes": 2
        },
        {
          "id": 793292,
          "postDate": "2020-03-31T23:35:53.623Z",
          "content": "<p>Tried them all, dropblock, shakedrop, cutmix etc, this combination worked fine now :) but cutmix is best on paper, thats right.</p>",
          "rawMarkdown": "Tried them all, dropblock, shakedrop, cutmix etc, this combination worked fine now :) but cutmix is best on paper, thats right."
        },
        {
          "id": 793296,
          "postDate": "2020-03-31T23:39:26.827Z",
          "content": "<p>1-cycle scheduler, but i used it with re-runs and lower rates every time</p>",
          "rawMarkdown": "1-cycle scheduler, but i used it with re-runs and lower rates every time"
        },
        {
          "id": 793390,
          "postDate": "2020-04-01T02:02:06.830Z",
          "content": "<p><a href=\"/ryches\">@ryches</a> Deep learning's been popular for 8 years. 3 years is ages in dog-years.</p>",
          "rawMarkdown": "@ryches Deep learning's been popular for 8 years. 3 years is ages in dog-years.",
          "votes": 1
        },
        {
          "id": 793962,
          "postDate": "2020-04-01T12:47:32.973Z",
          "content": "<p><a href=\"/ryches\">@ryches</a> Everything before EfficientNet is outdated, haha :D</p>",
          "rawMarkdown": "@ryches Everything before EfficientNet is outdated, haha :D",
          "votes": 2
        }
      ]
    },
    {
      "id": 793153,
      "postDate": "2020-03-31T20:50:31.333Z",
      "content": "<p>GCloud, ensemble</p>",
      "rawMarkdown": "GCloud, ensemble"
    },
    {
      "id": 793139,
      "postDate": "2020-03-31T20:31:57.943Z",
      "content": "<p>Step1: Trained model on Fake/Real faces - efficientnetb1 ( other deepfake faces )\nStep2: Used step1 model with LSTM time distribution with Augmentation ( Jpeg compression, Horizontal flip )</p>",
      "rawMarkdown": "Step1: Trained model on Fake/Real faces - efficientnetb1 ( other deepfake faces )\nStep2: Used step1 model with LSTM time distribution with Augmentation ( Jpeg compression, Horizontal flip )"
    },
    {
      "id": 793131,
      "postDate": "2020-03-31T20:25:35.243Z",
      "content": "<p>Simple binary classifiers: Xception and Effi B1</p>",
      "rawMarkdown": "Simple binary classifiers: Xception and Effi B1"
    },
    {
      "id": 794621,
      "postDate": "2020-04-02T00:16:06.013Z",
      "rawMarkdown": "",
      "votes": 7,
      "isDeleted": true,
      "replies": [
        {
          "id": 794636,
          "postDate": "2020-04-02T01:01:01.513Z",
          "content": "<p>Oh wow never thought of using dense layer instead of gru/lstm to figure out the changes! Well done. </p>",
          "rawMarkdown": "Oh wow never thought of using dense layer instead of gru/lstm to figure out the changes! Well done. "
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 793323,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-04-01T00:23:38.763000",
      "content": "<p>A TON honestly. </p>\n\n<p>Start off with our preprocessing method with multiplier system, take the same amount of frames in real video(equally distributed) as the number of the corresponding fake that video have, making data balance without data loss. (note its single frame model)</p>\n\n<p>Our best single model is effb6 with linear 256, relu 128, and sigmoid 1 head. We got 0.33LB for that model. We ensembled mixnet, efficientnetb1, b3, b6 with all different paddings and resolution, seresnet 101(couldn't fit seresnet 154 in that 1GB constraint) and got our current score. Random Hflip, random brightness for argumentation. </p>\n\n<p>What we tried but didn't work:\n- simple LRCN: since it couldn't fit in our multiplier system, it have data loss, which reached its limit before single frame models\n- LRCN with CNN froze with our best single single frame model weight preloaded(using our old data without multiplier system): It had the best CV, but unfortunately, it had a pretty bad LB(0.38)\n- filter out frames with face that have too much yaw: Its very likely due to low quality input data that have compression noise\n- Ensemble more model: This is not a really valid reason, but our aws credits ran out.............\n- cutout</p>\n\n<p>What we didn't try, but could've helped:</p>\n\n<ul>\n<li>SWA</li>\n<li>More LR schedule(<em>here I want to note that 1e-4 is our magic lr that other lr and several lr schedules couldn't surpass</em></li>\n<li>dual shot face detector(its very very slow, so we didn't try, and we possibly can't fit all those models in 9 hour constraint)</li>\n<li>cutmix, gridmask</li>\n<li>face embedding based validation set</li>\n<li>proper kfolds(this is really a shame to say we didn't even get a chance to do proper kfolds, still, our quota ran out before we could actually make the data)</li>\n<li>audio model</li>\n</ul>\n\n<p>I'm really excited about the topper's solution.... Good luck for the private LB! </p>",
      "votes": 5,
      "replies": [
        {
          "id": 793328,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-04-01T00:28:48.040000",
          "content": "<p>Not sure if I understand what you mean in the first part. Are you saying that if for example you pull 90 frames from a real video then you would pull 90 frames across the ~5 fake videos instead of 90 frames from each fake video. That way you have the same number of unique frames for both real and fake for a given video?</p>\n\n<p>Is that what you did? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793329,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-01T00:29:53.543000",
          "content": "<p>say a real video have 5 fakes, you will pull 5 frames from the real video and 1 frame from each fake video. Sorry for consufion.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793194,
      "author_name": "Oleg Trott",
      "author_url": "",
      "post_date": "2020-03-31T21:48:46.227000",
      "content": "<p>I actually used a very sophisticated solution, despite my poor showing (still not sure what went wrong).</p>\n\n<ol>\n<li><p>Input: developed a RAM-efficient multi-threaded cv2-based reader: one thread is servicing the GPU, while 1 or more others are reading the next video. This showed processing speed-up on my workstation.</p></li>\n<li><p>Face detection. Very fast and very accurate based on a modified version of Faceboxes. Did all the resizing on the GPU and kept the selected faces on the GPU to save time. Also, did face detection in batches (small enough to not cause OOM errors on a P100)</p></li>\n<li><p>Face tracking: interpolation based, to get rid of the jitter in the NN-based detections and to fill in the rare gaps.</p></li>\n<li><p>Actual deepfake detection. Tried single-face ones and short sequences. Ended up with a 2-frame sequences. Trained only on one such sequence for each video. Chose WideResnet-50-2 as the backbone. ResNext was marginally better but much slower.</p></li>\n<li><p>Augmentation: flips during training and I also re-encoded 2/3 of the videos as the DPDC Preview paper suggested.</p></li>\n<li><p>Inference: 256 frames per video (but I had to give up the multithreaded (1) reader, which turned out to be slower on Kaggle's machines than a single-threaded one)</p></li>\n<li><p>Ensemble: 2 models, trained on different frames and augmentations.</p></li>\n</ol>\n\n<p>Two ideas I wish I had time to try:</p>\n\n<ol>\n<li><p>Use sudden changes in face embeddings. This was actually the first idea I had, but went on to actually implement other things. Now I think it was promising, at least as part of an ensemble.</p></li>\n<li><p>Highlight pixelation / noise. You can do that via a high-pass filter or using NNs. Also, use the original frame resolution, but only a patch of the image. This approach does not care about the content, but only the \"noise\".</p></li>\n</ol>\n\n<p>I obviously made a mistake of actually getting started on this contest very late, and on top of that, certain events and the pandemic were a distraction.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 793203,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T22:01:33.163000",
          "content": "<p>I tried the change in face embedding. It wasn't very good approach, if it makes you feel better about it. Face embedding is designed to actually be relaxed about minor changes. Also differences in embedding while the actor was turning were more than the difference between real and fake.  Could have been just my bad. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793206,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T22:07:47.937000",
          "content": "<p>An approch i tried a few times and failed miserably was to train on difference between consecutive frames. Visual examination shows that the difference is very different with fake and real. However, the nn did not pick this up and accuracy and loss were near random. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 793221,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-31T22:20:03.717000",
          "content": "<p><a href=\"/moshel\">@moshel</a> </p>\n\n<p>Face embeddings are designed to be insensitive to changes in pose and expression, yes. However, the deepfakes exhibited sudden changes in \"identity\". </p>\n\n<p>During training, one can use the difference between the fake video and the real one to find only those frames where this difference changes suddenly (A sudden difference difference, if you will)</p>\n\n<p>This happens when the forgery algorithm fails at face detection (Side note: They should have used a good detector + interpolation, as I did) </p>\n\n<p>Now, if you trained the model on just those frames as \"fakes\", I think it would learn to detect the unnatural \"facial ticks\" many deepfakes exhibit in some parts of the video.</p>\n\n<p>Well, I haven't actually implemented this, so this is a guess.</p>\n\n<p>The 2-frame model I used was also intended to be sensitive to these, but it apparently wasn't very good. However, it was pre-trained on ImageNet, not the facial triplet-loss, and I chose the frames randomly, not as described above.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793234,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-31T22:34:46.293000",
          "content": "<p><a href=\"/moshel\">@moshel</a> I tried similar things. Lstm over the embeddings generated from the face net models. Never able to get anything out of it. I also tried frame differencing and optical flow. Surprising it could not extract signal from those representations but that's what I found in my experiments. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793338,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-04-01T00:41:28.727000",
      "content": "<p>The 2 models we submitted:\n1) Ensemble of 9 EffB3 + Flip TTA + 64 Frame inference\n2) Ensemble of 19 models including EffB1, EffB2, EffB3, Xception, ResNext + Flip TTA + 64 frame inference.</p>\n\n<p>Not gonna think about it for a few days, God knows what the private set looks like 😂 </p>",
      "votes": 3,
      "replies": [
        {
          "id": 793339,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-01T00:43:02.730000",
          "content": "<p>what is the resolution you used? What is your balancing technique?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793344,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-04-01T00:50:08.350000",
          "content": "<p>Resolution is 224. I tried others but did not work for us. Some components models are ensembles done by using median instead of mean. Balancing: change fake corresponding to real frame every epoch.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793792,
          "author_name": "Leonid Boytsov",
          "author_url": "",
          "post_date": "2020-04-01T09:32:18.970000",
          "content": "<p>How did you ensemble? Just averaging, fixed hold-out stacking, CV-stacking?\nAnd what is  Flip TTA + 64 Frame inference?\nThank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794255,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-04-01T17:25:47.223000",
          "content": "<p>Ensemble: Combination of weighted mean and median of the model predictions\nFlip TTA and 64 frames: During inference, I used 64 frames + 64 horizontally flipped frames</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794285,
          "author_name": "Leonid Boytsov",
          "author_url": "",
          "post_date": "2020-04-01T17:57:47.107000",
          "content": "<p>Thank you! And how exactly us the weight calculated?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 794358,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-04-01T18:54:48.393000",
          "content": "<p>Higher weights to models performing better in LB. Exact numbers are hit and trial :D Regarding median, they don't have weights, so less headache!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 793307,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-03-31T23:57:18.370000",
      "content": "<p><strong>Network:</strong>\n- 5 x Efficientnet b3 ns w/ pre-trained weights (<a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a>)\n- 300 x 300 face cutout\n- Inference on 60 frames (np.linspace)\n- Lots off augmentation (<a href=\"https://github.com/albumentations-team/albumentations\">https://github.com/albumentations-team/albumentations</a>)\n- brightness, contrast, gamma, jpeg compression, blur (inc motion), gaussian noise, scaling down\n- Dropout (0.2)</p>\n\n<p><strong>Training:</strong>\n- Pytorch\n- Batch size 8\n- 2-3 epochs per model\n- SGD\n- Momentum 0.9\n- Weight Decay 1e-4\n- LR 0.001</p>\n\n<p><strong>Training set:</strong>\n- 5 folds\n- 0-9, 10-19, 20-29, 30-39, 40-49 validation sets\n- 5 frames from real video and 5 frames across set of fake videos where frame position align between real and fake\n- ~134k training and ~37k validation images per fold</p>\n\n<p><strong>Face detector:</strong>\n- RetinaFace (<a href=\"https://github.com/biubug6/Pytorch_Retinaface\">https://github.com/biubug6/Pytorch_Retinaface</a>)\n- Resized largest axis down to 960 before detection\n- Center crop\n- Add 10% margin\n- Extract 5 frames all faces\n- Nms 0.4</p>\n\n<p><strong>Training rig:</strong>\n- Dual socket Xeon\n- 192 GB RAM\n- 1 x RTX 2080Ti\n- 1 x GTX 1080\n- 2 TB SSD</p>\n\n<p><strong>What helped:</strong>\n- Augment validation set as well (downscale, compression)\n- Center the face cutout and add margin\n- Focus on high quality faces (detection threshold &gt;= 0.99)\n- Remove fake faces that had very low and very high structural similarity to equivalent real face frame (skimage)\n- Balance real/fake training set (described in previous post)\n- Avoid just averaging faces when video has multiple faces</p>\n\n<p><strong>What didn't work</strong>\n- Label smoothing\n- Weighted average during inference weighted on # real/fake faces per video</p>\n\n<p><strong>What I didn't try (I suspect these would have helped)</strong>\n- LSTM\n- ffmpeg over cv2\n- Custom CNN trained from scratch</p>",
      "votes": 3,
      "replies": [
        {
          "id": 796533,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-04-03T17:03:57.193000",
          "content": "<p><a href=\"/maralski\">@maralski</a> 192MB RAM?!?!?! how were you even able to get a proper OS running?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 796551,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-04-03T17:23:24.777000",
          "content": "<p>He is joking come on, thats 192 gb.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 797148,
          "author_name": "ss",
          "author_url": "",
          "post_date": "2020-04-04T08:36:27.143000",
          "content": "<p>Can I ask what's the differences between your 5 models?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 797157,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-04-04T08:51:17.190000",
          "content": "<p>Hehe my boo boo 192 GB</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 797159,
          "author_name": "maralski",
          "author_url": "",
          "post_date": "2020-04-04T08:52:38.380000",
          "content": "<p><a href=\"/suthidasukhonn\">@suthidasukhonn</a> each was trained on a different fold</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793129,
      "author_name": "Moshel",
      "author_url": "",
      "post_date": "2020-03-31T20:25:00.353000",
      "content": "<p>I couldn't get lstm/gru to work. It overfitted no matter what I did.\nMy model is a single model, on full frame (not faces) of seresnext50. </p>",
      "votes": 3,
      "replies": [
        {
          "id": 793135,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-31T20:29:34.353000",
          "content": "<p>Do you rescale the frame? How many frames did you take per video during training? How many during inference?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793147,
          "author_name": "Thomas Sherk",
          "author_url": "",
          "post_date": "2020-03-31T20:44:53.583000",
          "content": "<p>We had an issue of over-fitting, I added regularization to the LSTM layer but that did not seem to help. If I had more time, I probably would be able to figure out what was wrong but it is what it is.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793167,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T21:03:34.020000",
          "content": "<p>Yes, rescaled to 512. 1 frame per movie during training, 30 frames inference.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793172,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T21:09:30.277000",
          "content": "<p>I had a feeling gru was the way to go, but got distracted and didn't have time to figure out why it is over fitting so quickly. I am sure it is something stupid... Basically gru should be able to \"see\" that frames are changing differently in fake and real. However, it focused on something else and the process of training these is very time consuming. Would be interesting to see what top places did. I bet they used lstm/gru. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793192,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-03-31T21:46:03.630000",
          "content": "<p><a href=\"/moshel\">@moshel</a> And your full frame (no face) single model (frames resized to 512) scored 0.342?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793196,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T21:49:32.353000",
          "content": "<p>Yes! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793197,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T21:51:41.180000",
          "content": "<p>Naturally the devil is in the details. Selecting the cv, training generator that balance real /fake while preventing over fitting and augmentation. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 793198,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T21:52:39.600000",
          "content": "<p>Also the score calculation function was very important. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793205,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-03-31T22:07:00.173000",
          "content": "<p>It's a good result. We've tried such model but with smaller resize. During CV we've noticed that it learnt interesting features but we stopped as CV was not so good for us.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793213,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-03-31T22:11:45.740000",
          "content": "<p>I specialize in squeezing the most out of my models ;)\nYour score is pretty amazing. Care to share a bit about your approach? It's too late to submit anyways. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793227,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-03-31T22:27:18.880000",
          "content": "<p>Yes, will do it tomorrow (time to sleep here).</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 795132,
      "author_name": "Chason",
      "author_url": "",
      "post_date": "2020-04-02T12:57:08.177000",
      "content": "<h1>Train</h1>\n\n<ul>\n<li>Face Detector: margin(0.05) based DSFD/BlazeFace</li>\n<li>Dataloader: under sample、5 frame/video</li>\n<li>Augmentation: ImageCompression、RandomBrightness、GaussianNoise、RandomCrop, etc</li>\n<li>Backbone: NoisyStudent pretrained EfficientNet</li>\n<li>Attention: similar as C-FAN, which adaptively aggregates deep feature vectors into a single vector</li>\n<li>Two stage training:\nStage-1: train a base CNN on real/fake images;\nStage-2: freeze the pre-trained base CNN to extract the embeddings of faces, and train attention module on videos (5 frame / video). </li>\n</ul>\n\n<h1>Val</h1>\n\n<ul>\n<li>DFDC: split by folder-wise and original videos (CV: 0.14、LB: 0.30)</li>\n<li>FaceForensics++: a generalized validation set, which only include deepfake, faceswap related videos (CV: 0.22、LB: 0.27) \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F3ba58aefd964b742a098d933ed1ef42a%2F1.png?generation=1585831881937088&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<h1>Tried but not work</h1>\n\n<ul>\n<li>Audio model</li>\n<li>LSTM, 3D Cov, TSM</li>\n<li>Spatial attention、Multi-scale attention、Multi-head attention、Prob attentioin</li>\n<li>Adversarial attack</li>\n<li>Multi-patch fusion</li>\n<li>YCbCr color space transformation</li>\n<li>use GMM to fuse n frames probabilities</li>\n<li>ArcFace、SphereFace、Contrastive Loss、Focal Loss、Label Smooth, etc</li>\n<li>decrease resolution</li>\n<li>remove multi-face videos</li>\n<li>Face Alignment</li>\n<li>Initialize efficientnet with pretrained weights on the face dataset</li>\n<li>Face Xray</li>\n<li>Frequency Learning</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 795184,
          "author_name": "apppa",
          "author_url": "",
          "post_date": "2020-04-02T13:45:09.993000",
          "content": "<p>Great! Really want to see how attention is implemented</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795246,
          "author_name": "Leonid Boytsov",
          "author_url": "",
          "post_date": "2020-04-02T14:52:58.200000",
          "content": "<p>Cool, what are <strong>3D cov</strong> and <strong>TSM</strong>? I would also like to read about the attention solution. Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795768,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-04-03T02:59:46.467000",
          "content": "<p>Hard Luck, we did try that exact stage 2 approach, and validation loss did go much better than ever!, but did not perform good in LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795788,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-04-03T03:46:55.587000",
          "content": "<p><a href=\"/azamatk\">@azamatk</a> <a href=\"/leonidboytsov\">@leonidboytsov</a> The attention module is similar to SE-block, and I will share the detailed solution after the busy days.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795789,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-04-03T03:49:45.253000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> We encountered the same problem and later found out that it was mainly due to overfitting. Therefore, we use the warm-up strategy and train only one epoch, which greatly improves the performance in both CV and LB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 795790,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-04-03T03:51:05.027000",
          "content": "<p>Now I am starting to hate myself, also in my case, there was not  a lot improvement like 0.12 to 0.11 from first epoch to best.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793267,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2020-03-31T23:16:14.107000",
      "content": "<h2>Simple Solution</h2>\n\n<p>I joined the competition relatively late, last month but thanks to everyone contributing with kernels and discussion helped me get up to speed relatively easy. For me, the biggest challenge was to work with video data since it was the first time. But it was a great experience for learning efficient video batching approaches. Spent ~3 weeks on data and ~1 week on modeling.</p>\n\n<p>All my work can be obtained <a href=\"https://github.com/KeremTurgutlu/dfdc\">here</a>. I used <a href=\"https://github.com/fastai/nbdev\">nbdev</a> to speed up my code base development process using notebooks. I highly recommend for other competitions if you are more productive with jupyter notebooks but still need the power of modularizing your code.</p>\n\n<h3>Video Loading/Batching</h3>\n\n<p>I used decoder cpu version (thanks to <a href=\"https://www.kaggle.com/leighplt/decord-videoreader/notebook\">this</a>). Tried out DALI and decord gpu versions but they had issues in terms of memory leakage and I didn't have enough time to debug and also cpu version was fast enough for my pipeline. As for opencv video loading I tried both vanilla and multithreaded options but in aws p3 instance decoder cpu was a little bit faster (eventhough it wasn't the case in Kaggle kernel).</p>\n\n<h3>Detection</h3>\n\n<p>For detection I used and adopted mobilenet detector in <a href=\"https://github.com/biubug6/Pytorch_Retinaface\">here</a>. It's a very lightweight model in terms of memory, space and it's pretty fast during inference. I didn't have time to prepare a better dataset by dealing with potential false negatives/positives. That may have allowed a gain in overall model performance. I extracted ~30 frames (1/10 of each frame with equal intervals) for each video and stored only faces. I used the <a href=\"https://github.com/KeremTurgutlu/dfdc/blob/master/dfdc/face_detection/download_detect_crop_save.py\">script</a> which downloads a video chunk (1 of 50) does detections, saves cropped faces and finally deletes the processed chunk. This is repeated sequentially for all 50 chunks. This allowed me to save a lot of disk space. </p>\n\n<p>Each face crop enlarged by x1.3 after detection.</p>\n\n<h3>Validation</h3>\n\n<p>Discussions were really helpful in terms of deciding a proper validation set. Thanks to everyone who shared what worked and what didn't. This allowed me to save a lot of time and not to spend too much time on methods like grouping based on face embeddings. I finally used chunks 1-40 as training, 41-45 as validation and 46-50 as test. Once I decided on modeling I merged training and test for final training.</p>\n\n<h3>Data Sampling</h3>\n\n<p>Once I started training a baseline model I quickly noticed that a vanilla random batch sampling won't do good. So I decided on a custom batch sampler. In a batch, this sampler basically randomly picks n number of original videos and a random fake video for each of these n original videos. Then for each real-fake pair video same frame crop is used. This strategy helped to get a decent public LB score but still, the model was overfitting pretty bad and fast.</p>\n\n<p><strong>What didn't work:</strong> I also had another sampler that utilized real video and it's all corresponding fakes again using the same frame crop for each. This was significantly worse compared to balanced sampling (1 real - 1 random fake).</p>\n\n<h3>Regularization</h3>\n\n<p>I think the biggest challenge was avoiding overfitting and the potential risk of memorizing individual faces. At first I tried trivial data augmentation strategies.</p>\n\n<ul>\n<li>Crappification: downsampling image and upsample again to mimic low resolution/compression scenarios.</li>\n<li>Left right flipping</li>\n<li>Brightness change - darker or lighter</li>\n<li>Random zoom - x1.15 - x1.35</li>\n</ul>\n\n<p>Just 3 days ago I added a custom augmentation strategy similar to cutmix which is inspired by another <a href=\"https://www.kaggle.com/c/state-farm-distracted-driver-detection\">compeition</a> from <a href=\"/godaibo\">@godaibo</a> from his TWIML podcast interview.</p>\n\n<ul>\n<li>For a given face crop, I merge it vertically (50% - 50%) with another face crop from the same class. This gave a boost of 0.02 in pubic LB. The idea was to not memorize individual faces. The reason for 50% and vertical merge was assumptions based on the symmetry of fake alterations, e.g. either change of eyes, mouth ...etc.</li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F558069%2F1244fc86ea6a89ecd9001c1844d48292%2FScreen%20Shot%202020-04-01%20at%205.29.36%20PM.png?generation=1585751421952949&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Also tried mixup but didn't improve the LB.</li>\n</ul>\n\n<h3>Models</h3>\n\n<ul>\n<li>Resnet34, EfficientNet b5, Efficient b7 - Efficient alone is enough to get my best score. I did simple averaging. Tried both with and without TTA but didn't get much of a boost. All models are finetuned using gradual unfreezing and 1-cycle policy. FocalLoss is used at a later stage of training to update model with harder samples. Also used simple clipping (0.01-0.99) just to be safe about very high false positives and very low false negatives.</li>\n</ul>\n\n<h3>What didn't work</h3>\n\n<ul>\n<li>LRCN based modeling</li>\n<li>1 real all fakes sampling</li>\n</ul>\n\n<p>Thanks to everyone for their great contributions! </p>\n\n<p>Thanks to organizers for contributing with data and the challenge, Kaggle for hosting and AWS for the compute.</p>\n\n<p>It was a great learning experience and I wish everyone luck with private LB :)</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 793450,
      "author_name": "shimacos",
      "author_url": "",
      "post_date": "2020-04-01T03:18:25.927000",
      "content": "<p>During this competition, I have been struggling with the gap between CV and LB. Finally, I cannot resolve this problem...</p>\n\n<p>Validation</p>\n\n<ul>\n<li>5 fold</li>\n<li>Stratified group k fold</li>\n<li>Stratified by the number of fake and set the group to folder.</li>\n<li>2/9 Downscale and 2/9 JpegCompression</li>\n<li>logloss weighted to have the same FAKE and ORIGINAL influences</li>\n<li>CV score is 0.16 and Public is 0.33 by this approach.</li>\n</ul>\n\n<p>Network</p>\n\n<ul>\n<li>Sample original video(A) and fake video (generated from A) in the same minibatch.\n<ul><li>randomly select fake video (from A) every epoch </li>\n<li>randomly select frame every epoch, but the frame index in the pair (A, fakeA) is same.</li></ul></li>\n<li>Face crop by SSFD (<a href=\"https://github.com/cs-giung/face-detection-pytorch\">https://github.com/cs-giung/face-detection-pytorch</a>)</li>\n<li>Stacking by predictions of 20 frames per video from efficientnet-b4 and se-resnext 50</li>\n<li>Train LightGBM by extract features of video from above predictions (mean, median , min, max, std, etc..)</li>\n</ul>\n\n<p>PostProcess</p>\n\n<ul>\n<li>There are large gap train dataset and Private dataset and model’s output tend to be low confidence. So, I applied postprocess as below.\n<ul><li>sigmoid (lgbm output) -&gt; logit -&gt; multiply s -&gt; sigmoid</li></ul></li>\n</ul>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 793320,
      "author_name": "Mont3z Claro5",
      "author_url": "",
      "post_date": "2020-04-01T00:15:56.187000",
      "content": "<p>EfficientNet with Transformer which I just implemented this week\n<a href=\"https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986\">https://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986</a></p>\n\n<p>Failed to submit by 5 min. I was just hoping it would improve my terrible score :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 793302,
      "author_name": "yuanzhe zhou",
      "author_url": "",
      "post_date": "2020-03-31T23:51:57.010000",
      "content": "<p>Xception and efficientnetb1, the LB score is always 0.2 higher than my local validation score ... I didn't use augmentations(only horizontal flip).</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 793910,
      "author_name": "BayesianKitten",
      "author_url": "",
      "post_date": "2020-04-01T11:56:34.760000",
      "content": "<p>We started working on this a bit late. Our solution was rather simple:</p>\n\n<ul>\n<li>Extracting faces with Retinaface detector</li>\n<li>Training an ensemble of 3 EfficientNet-B5s  (training time was ~3hrs on 1 RTX6000)</li>\n<li>Data Augmentation with JPEGCompression, ISONoise, Blur and DownScale</li>\n<li>Simple weighted crossentropy loss with [0.8,0.2] weights.</li>\n<li>No validation set used (Creating this seemed painful, since folders had mixed actors).</li>\n<li>Cap to [0.05,0.95] for predictions (to avoid big penalties)</li>\n</ul>\n\n<p>Our pipeline was probably very standard- faces extraction was rescaled to 224x224, our pipeline similar to the kernel posted by Human Analog. Our training code was simply applying the above tweaks applied to the pytorch/examples/imagenet/main.py with augmentations from Albumentations. </p>\n\n<p>Our rank increase was mostly in the last 5 days of the contest. I guess we were just relatively lucky? If so, hope that our luck is not too bad the private test set.. ~~</p>",
      "votes": 2,
      "replies": [
        {
          "id": 793922,
          "author_name": "Leonid Boytsov",
          "author_url": "",
          "post_date": "2020-04-01T12:04:53.893000",
          "content": "<p>How did you decide when to stop training?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793932,
          "author_name": "BayesianKitten",
          "author_url": "",
          "post_date": "2020-04-01T12:18:58.477000",
          "content": "<p>Since we never did a full pass over the dataset, the training accuracy was weakly indicative of the performance, and it seemed to stay at 94%ish, which seemed reasonably accurate but not overfitting. So, we stopped it at those iterations. </p>\n\n<p>Also, one more factor was that we could make it in time to try submit 2 submissions everyday (so 6 models train with 3 GPUs). Training finishing in 3-4 hours was perfect in timing too, and I could train and upload two batches of models everyday.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793949,
          "author_name": "Leonid Boytsov",
          "author_url": "",
          "post_date": "2020-04-01T12:32:08.357000",
          "content": "<p>Using leaderboard for early stopping, wow, I should have thought about it. 👍 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 794069,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-01T14:31:17.030000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793371,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2020-04-01T01:27:37.563000",
      "content": "<p>Ensemble of 9 EfB1  LRCNN, courtesy of <a href=\"/harshitsheoran\">@harshitsheoran</a> and <a href=\"/unkownhihi\">@unkownhihi</a> , with modified face extraction (preventing it from switching face), and all trained on different parts of the data (some random, some pair of true-fake). I lost track which one is which, I chose those with CV score ranging 0.1 - 0.2 roughly (validation on folder 45-49)\nTried very hard to get MobileNet to work, but it was futile (loss doesn't go below 0.35). Tried EfB0 as well, but doesn't work out. Xcetion and Densenet both too big (a.k.a. slow) and I'm scared of burning my AWS credit. </p>\n\n<p>I'm not sure if training using pair of true-fake is really useful , i wonder if anyone can shed light on this</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 793461,
      "author_name": "Zungmann",
      "author_url": "",
      "post_date": "2020-04-01T03:38:37.307000",
      "content": "<p>After most of times struggling to get a good classifier model. I followed EfB1 + LSTM model by <a href=\"/unkownhihi\">@unkownhihi</a>. And I spent times creating a sequence dataset to feed the model. I use custom Keras Sequence, balanced pair real-fake with different fake for each epoch. I use 160x160 image size of face crop\nBut this also another struggle, could not get the model to get a good validation. My validation set is folder 0,1,2,43-49\n- with frozen backbone, training not able to converge\n- with retrain backbone, my model can converge well but badly overfit. then found bug in training data, since I used padding, some fake data just blank because could not find the images/faces\n- after fix the bug, still could not get good validation loss, somehow my model very confident to guess real as fake or the other way\n- changed  image size to 120x120, center crop from original 160x160, slightly better\n- tried RMSProp, Adam and SGD with different learning rates: RMSProp and Adam can converge fast, but validation loss is erratic. SGD is more stable and not able to converge well. My validation loss is 1.x and tried to submit one get 0.8 on LB\n- last night I change the model to:\n    y = TimeDistributed(cnn)(input_layer) \n    y = LSTM(256, return_sequences=False)(y) <br>\n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\nI used SGD with momentum and cyclic learning rate. I was not sure this going to be different and just leave it run. And this morning surprisingly, I got:\npoch 7/20\n1291/1291 [==============================] - 476s 369ms/step - loss: 0.1620 - acc: 0.9330 - val_loss: 0.0334 - val_acc: 0.8122\nI realize that this is not great and may not give me super LB increase, I am just happy I finally breaking my own many attemps hard to reach good validation. But have to wait for 22 more days to do late submission :( to know for sure</p>\n\n<p>what a journey... congrats to you all. thanks for sharing</p>",
      "votes": 0,
      "replies": [
        {
          "id": 794000,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-04-01T13:24:22.650000",
          "content": "<p>for me SGD worked alone with momentum and weight_decay and lower the lr with factor 0.97\nadamw and 1-cycle worked fine together\nalways started with lr(batch*5e-04, eg 0.064 for batch 128)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 793252,
      "author_name": "Mats",
      "author_url": "",
      "post_date": "2020-03-31T22:55:43.090000",
      "content": "<p>efficientnet b3 pretrained model (transfer learning)\n + dropout &amp; dense layers\n + softmax output.</p>\n\n<p>it showed 88%+ acc &amp; 0.4 loss in 10 mins of GPU run (batchsize = 160)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 793202,
      "author_name": "Kirderf",
      "author_url": "",
      "post_date": "2020-03-31T22:01:25.153000",
      "content": "<p>BlazeFace for face extraction\nfirst sub- stacked efficientnet b0ns+b1ns+b2ns+b3ns\nsecond sub - stacked(b0ns+b1ns+xception)+ensemble with best model of b1ns\nmixup, \nrandom cutout,\naugmentation, \n1-cycle with more runs,\nweighted average checkpoints\nNo CV but I hope that the large public testset was not too misleading</p>",
      "votes": 0,
      "replies": [
        {
          "id": 793237,
          "author_name": "Thomas Sherk",
          "author_url": "",
          "post_date": "2020-03-31T22:36:47.737000",
          "content": "<p>What do you mean by mixup, cutout, and 1-cycle?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793245,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-03-31T22:42:58.530000",
          "content": "<p>mixup and cutout are augmentation methods (a bit old-fashioned, apparently bested by \"cutmix\")</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793263,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-31T23:09:58.370000",
          "content": "<p>Those papers are both from 2017. Are we calling 2017 old fashioned already? haha</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 793292,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-03-31T23:35:53.623000",
          "content": "<p>Tried them all, dropblock, shakedrop, cutmix etc, this combination worked fine now :) but cutmix is best on paper, thats right.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793296,
          "author_name": "Kirderf",
          "author_url": "",
          "post_date": "2020-03-31T23:39:26.827000",
          "content": "<p>1-cycle scheduler, but i used it with re-runs and lower rates every time</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 793390,
          "author_name": "Oleg Trott",
          "author_url": "",
          "post_date": "2020-04-01T02:02:06.830000",
          "content": "<p><a href=\"/ryches\">@ryches</a> Deep learning's been popular for 8 years. 3 years is ages in dog-years.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 793962,
          "author_name": "Victor Paslay",
          "author_url": "",
          "post_date": "2020-04-01T12:47:32.973000",
          "content": "<p><a href=\"/ryches\">@ryches</a> Everything before EfficientNet is outdated, haha :D</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 793153,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-31T20:50:31.333000",
      "content": "<p>GCloud, ensemble</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 793139,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2020-03-31T20:31:57.943000",
      "content": "<p>Step1: Trained model on Fake/Real faces - efficientnetb1 ( other deepfake faces )\nStep2: Used step1 model with LSTM time distribution with Augmentation ( Jpeg compression, Horizontal flip )</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 793131,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-03-31T20:25:35.243000",
      "content": "<p>Simple binary classifiers: Xception and Effi B1</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 794621,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-02T00:16:06.013000",
      "content": "",
      "votes": 7,
      "replies": [
        {
          "id": 794636,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-04-02T01:01:01.513000",
          "content": "<p>Oh wow never thought of using dense layer instead of gru/lstm to figure out the changes! Well done. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "793121": "I'm interested to see what others have done. We created a CNN-LSTM, an audio-model, and other models but didn't have great success. ",
    "793323": "A TON honestly. \n\nStart off with our preprocessing method with multiplier system, take the same amount of frames in real video(equally distributed) as the number of the corresponding fake that video have, making data balance without data loss. (note its single frame model)\n\nOur best single model is effb6 with linear 256, relu 128, and sigmoid 1 head. We got 0.33LB for that model. We ensembled mixnet, efficientnetb1, b3, b6 with all different paddings and resolution, seresnet 101(couldn't fit seresnet 154 in that 1GB constraint) and got our current score. Random Hflip, random brightness for argumentation. \n\nWhat we tried but didn't work:\n- simple LRCN: since it couldn't fit in our multiplier system, it have data loss, which reached its limit before single frame models\n- LRCN with CNN froze with our best single single frame model weight preloaded(using our old data without multiplier system): It had the best CV, but unfortunately, it had a pretty bad LB(0.38)\n- filter out frames with face that have too much yaw: Its very likely due to low quality input data that have compression noise\n- Ensemble more model: This is not a really valid reason, but our aws credits ran out.............\n- cutout\n\nWhat we didn't try, but could've helped:\n\n- SWA\n- More LR schedule(*here I want to note that 1e-4 is our magic lr that other lr and several lr schedules couldn't surpass*\n- dual shot face detector(its very very slow, so we didn't try, and we possibly can't fit all those models in 9 hour constraint)\n- cutmix, gridmask\n- face embedding based validation set\n- proper kfolds(this is really a shame to say we didn't even get a chance to do proper kfolds, still, our quota ran out before we could actually make the data)\n- audio model\n\nI'm really excited about the topper's solution.... Good luck for the private LB! ",
    "793194": "I actually used a very sophisticated solution, despite my poor showing (still not sure what went wrong).\n\n1. Input: developed a RAM-efficient multi-threaded cv2-based reader: one thread is servicing the GPU, while 1 or more others are reading the next video. This showed processing speed-up on my workstation.\n\n2. Face detection. Very fast and very accurate based on a modified version of Faceboxes. Did all the resizing on the GPU and kept the selected faces on the GPU to save time. Also, did face detection in batches (small enough to not cause OOM errors on a P100)\n\n3. Face tracking: interpolation based, to get rid of the jitter in the NN-based detections and to fill in the rare gaps.\n\n4. Actual deepfake detection. Tried single-face ones and short sequences. Ended up with a 2-frame sequences. Trained only on one such sequence for each video. Chose WideResnet-50-2 as the backbone. ResNext was marginally better but much slower.\n\n5. Augmentation: flips during training and I also re-encoded 2/3 of the videos as the DPDC Preview paper suggested.\n\n6. Inference: 256 frames per video (but I had to give up the multithreaded (1) reader, which turned out to be slower on Kaggle's machines than a single-threaded one)\n\n7. Ensemble: 2 models, trained on different frames and augmentations.\n\nTwo ideas I wish I had time to try:\n\n1. Use sudden changes in face embeddings. This was actually the first idea I had, but went on to actually implement other things. Now I think it was promising, at least as part of an ensemble.\n\n2. Highlight pixelation / noise. You can do that via a high-pass filter or using NNs. Also, use the original frame resolution, but only a patch of the image. This approach does not care about the content, but only the \"noise\".\n\nI obviously made a mistake of actually getting started on this contest very late, and on top of that, certain events and the pandemic were a distraction.",
    "793338": "The 2 models we submitted:\n1) Ensemble of 9 EffB3 + Flip TTA + 64 Frame inference\n2) Ensemble of 19 models including EffB1, EffB2, EffB3, Xception, ResNext + Flip TTA + 64 frame inference.\n\nNot gonna think about it for a few days, God knows what the private set looks like 😂 ",
    "793307": "**Network:**\n- 5 x Efficientnet b3 ns w/ pre-trained weights (https://github.com/rwightman/pytorch-image-models)\n- 300 x 300 face cutout\n- Inference on 60 frames (np.linspace)\n- Lots off augmentation (https://github.com/albumentations-team/albumentations)\n- brightness, contrast, gamma, jpeg compression, blur (inc motion), gaussian noise, scaling down\n- Dropout (0.2)\n\n**Training:**\n- Pytorch\n- Batch size 8\n- 2-3 epochs per model\n- SGD\n- Momentum 0.9\n- Weight Decay 1e-4\n- LR 0.001\n\n**Training set:**\n- 5 folds\n- 0-9, 10-19, 20-29, 30-39, 40-49 validation sets\n- 5 frames from real video and 5 frames across set of fake videos where frame position align between real and fake\n- ~134k training and ~37k validation images per fold\n\n**Face detector:**\n- RetinaFace (https://github.com/biubug6/Pytorch_Retinaface)\n- Resized largest axis down to 960 before detection\n- Center crop\n- Add 10% margin\n- Extract 5 frames all faces\n- Nms 0.4\n\n**Training rig:**\n- Dual socket Xeon\n- 192 GB RAM\n- 1 x RTX 2080Ti\n- 1 x GTX 1080\n- 2 TB SSD\n\n**What helped:**\n- Augment validation set as well (downscale, compression)\n- Center the face cutout and add margin\n- Focus on high quality faces (detection threshold &gt;= 0.99)\n- Remove fake faces that had very low and very high structural similarity to equivalent real face frame (skimage)\n- Balance real/fake training set (described in previous post)\n- Avoid just averaging faces when video has multiple faces\n\n**What didn't work**\n- Label smoothing\n- Weighted average during inference weighted on # real/fake faces per video\n\n**What I didn't try (I suspect these would have helped)**\n- LSTM\n- ffmpeg over cv2\n- Custom CNN trained from scratch",
    "793129": "I couldn't get lstm/gru to work. It overfitted no matter what I did.\nMy model is a single model, on full frame (not faces) of seresnext50. ",
    "795132": "# Train\n- Face Detector: margin(0.05) based DSFD/BlazeFace\n- Dataloader: under sample、5 frame/video\n- Augmentation: ImageCompression、RandomBrightness、GaussianNoise、RandomCrop, etc\n- Backbone: NoisyStudent pretrained EfficientNet\n- Attention: similar as C-FAN, which adaptively aggregates deep feature vectors into a single vector\n- Two stage training:\nStage-1: train a base CNN on real/fake images;\nStage-2: freeze the pre-trained base CNN to extract the embeddings of faces, and train attention module on videos (5 frame / video). \n\n# Val\n- DFDC: split by folder-wise and original videos (CV: 0.14、LB: 0.30)\n- FaceForensics++: a generalized validation set, which only include deepfake, faceswap related videos (CV: 0.22、LB: 0.27) \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3117946%2F3ba58aefd964b742a098d933ed1ef42a%2F1.png?generation=1585831881937088&amp;alt=media)\n\n# Tried but not work\n- Audio model\n- LSTM, 3D Cov, TSM\n- Spatial attention、Multi-scale attention、Multi-head attention、Prob attentioin\n- Adversarial attack\n- Multi-patch fusion\n- YCbCr color space transformation\n- use GMM to fuse n frames probabilities\n- ArcFace、SphereFace、Contrastive Loss、Focal Loss、Label Smooth, etc\n- decrease resolution\n- remove multi-face videos\n- Face Alignment\n- Initialize efficientnet with pretrained weights on the face dataset\n- Face Xray\n- Frequency Learning",
    "793267": "## Simple Solution\n\nI joined the competition relatively late, last month but thanks to everyone contributing with kernels and discussion helped me get up to speed relatively easy. For me, the biggest challenge was to work with video data since it was the first time. But it was a great experience for learning efficient video batching approaches. Spent ~3 weeks on data and ~1 week on modeling.\n\nAll my work can be obtained [here](https://github.com/KeremTurgutlu/dfdc). I used [nbdev](https://github.com/fastai/nbdev) to speed up my code base development process using notebooks. I highly recommend for other competitions if you are more productive with jupyter notebooks but still need the power of modularizing your code.\n\n### Video Loading/Batching\n\nI used decoder cpu version (thanks to [this](https://www.kaggle.com/leighplt/decord-videoreader/notebook)). Tried out DALI and decord gpu versions but they had issues in terms of memory leakage and I didn't have enough time to debug and also cpu version was fast enough for my pipeline. As for opencv video loading I tried both vanilla and multithreaded options but in aws p3 instance decoder cpu was a little bit faster (eventhough it wasn't the case in Kaggle kernel).\n\n### Detection\n\nFor detection I used and adopted mobilenet detector in [here](https://github.com/biubug6/Pytorch_Retinaface). It's a very lightweight model in terms of memory, space and it's pretty fast during inference. I didn't have time to prepare a better dataset by dealing with potential false negatives/positives. That may have allowed a gain in overall model performance. I extracted ~30 frames (1/10 of each frame with equal intervals) for each video and stored only faces. I used the [script](https://github.com/KeremTurgutlu/dfdc/blob/master/dfdc/face_detection/download_detect_crop_save.py) which downloads a video chunk (1 of 50) does detections, saves cropped faces and finally deletes the processed chunk. This is repeated sequentially for all 50 chunks. This allowed me to save a lot of disk space. \n\nEach face crop enlarged by x1.3 after detection.\n\n### Validation\n\nDiscussions were really helpful in terms of deciding a proper validation set. Thanks to everyone who shared what worked and what didn't. This allowed me to save a lot of time and not to spend too much time on methods like grouping based on face embeddings. I finally used chunks 1-40 as training, 41-45 as validation and 46-50 as test. Once I decided on modeling I merged training and test for final training.\n\n### Data Sampling\n\nOnce I started training a baseline model I quickly noticed that a vanilla random batch sampling won't do good. So I decided on a custom batch sampler. In a batch, this sampler basically randomly picks n number of original videos and a random fake video for each of these n original videos. Then for each real-fake pair video same frame crop is used. This strategy helped to get a decent public LB score but still, the model was overfitting pretty bad and fast.\n\n**What didn't work:** I also had another sampler that utilized real video and it's all corresponding fakes again using the same frame crop for each. This was significantly worse compared to balanced sampling (1 real - 1 random fake).\n\n### Regularization\n\nI think the biggest challenge was avoiding overfitting and the potential risk of memorizing individual faces. At first I tried trivial data augmentation strategies.\n\n- Crappification: downsampling image and upsample again to mimic low resolution/compression scenarios.\n- Left right flipping\n- Brightness change - darker or lighter\n- Random zoom - x1.15 - x1.35\n\nJust 3 days ago I added a custom augmentation strategy similar to cutmix which is inspired by another [compeition](https://www.kaggle.com/c/state-farm-distracted-driver-detection) from @godaibo from his TWIML podcast interview.\n\n- For a given face crop, I merge it vertically (50% - 50%) with another face crop from the same class. This gave a boost of 0.02 in pubic LB. The idea was to not memorize individual faces. The reason for 50% and vertical merge was assumptions based on the symmetry of fake alterations, e.g. either change of eyes, mouth ...etc.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F558069%2F1244fc86ea6a89ecd9001c1844d48292%2FScreen%20Shot%202020-04-01%20at%205.29.36%20PM.png?generation=1585751421952949&amp;alt=media)\n\n- Also tried mixup but didn't improve the LB.\n\n### Models\n\n- Resnet34, EfficientNet b5, Efficient b7 - Efficient alone is enough to get my best score. I did simple averaging. Tried both with and without TTA but didn't get much of a boost. All models are finetuned using gradual unfreezing and 1-cycle policy. FocalLoss is used at a later stage of training to update model with harder samples. Also used simple clipping (0.01-0.99) just to be safe about very high false positives and very low false negatives.\n\n### What didn't work\n\n- LRCN based modeling\n- 1 real all fakes sampling\n\n\nThanks to everyone for their great contributions! \n\nThanks to organizers for contributing with data and the challenge, Kaggle for hosting and AWS for the compute.\n\nIt was a great learning experience and I wish everyone luck with private LB :)",
    "793450": "During this competition, I have been struggling with the gap between CV and LB. Finally, I cannot resolve this problem...\n\nValidation\n\n- 5 fold\n- Stratified group k fold\n- Stratified by the number of fake and set the group to folder.\n- 2/9 Downscale and 2/9 JpegCompression\n- logloss weighted to have the same FAKE and ORIGINAL influences\n- CV score is 0.16 and Public is 0.33 by this approach.\n\nNetwork\n\n- Sample original video(A) and fake video (generated from A) in the same minibatch.\n  - randomly select fake video (from A) every epoch \n  - randomly select frame every epoch, but the frame index in the pair (A, fakeA) is same.\n- Face crop by SSFD (https://github.com/cs-giung/face-detection-pytorch)\n- Stacking by predictions of 20 frames per video from efficientnet-b4 and se-resnext 50\n- Train LightGBM by extract features of video from above predictions (mean, median , min, max, std, etc..)\n\nPostProcess\n\n- There are large gap train dataset and Private dataset and model’s output tend to be low confidence. So, I applied postprocess as below.\n  - sigmoid (lgbm output) -&gt; logit -&gt; multiply s -&gt; sigmoid",
    "793320": "EfficientNet with Transformer which I just implemented this week\nhttps://www.kaggle.com/mont3z/fork-of-deepfake-transformer-971986\n\nFailed to submit by 5 min. I was just hoping it would improve my terrible score :)",
    "793302": "Xception and efficientnetb1, the LB score is always 0.2 higher than my local validation score ... I didn't use augmentations(only horizontal flip).",
    "793910": "We started working on this a bit late. Our solution was rather simple:\n\n- Extracting faces with Retinaface detector\n- Training an ensemble of 3 EfficientNet-B5s  (training time was ~3hrs on 1 RTX6000)\n- Data Augmentation with JPEGCompression, ISONoise, Blur and DownScale\n- Simple weighted crossentropy loss with [0.8,0.2] weights.\n- No validation set used (Creating this seemed painful, since folders had mixed actors).\n- Cap to [0.05,0.95] for predictions (to avoid big penalties)\n\nOur pipeline was probably very standard- faces extraction was rescaled to 224x224, our pipeline similar to the kernel posted by Human Analog. Our training code was simply applying the above tweaks applied to the pytorch/examples/imagenet/main.py with augmentations from Albumentations. \n\nOur rank increase was mostly in the last 5 days of the contest. I guess we were just relatively lucky? If so, hope that our luck is not too bad the private test set.. ~~",
    "793371": "Ensemble of 9 EfB1  LRCNN, courtesy of @harshitsheoran and @unkownhihi , with modified face extraction (preventing it from switching face), and all trained on different parts of the data (some random, some pair of true-fake). I lost track which one is which, I chose those with CV score ranging 0.1 - 0.2 roughly (validation on folder 45-49)\nTried very hard to get MobileNet to work, but it was futile (loss doesn't go below 0.35). Tried EfB0 as well, but doesn't work out. Xcetion and Densenet both too big (a.k.a. slow) and I'm scared of burning my AWS credit. \n\nI'm not sure if training using pair of true-fake is really useful , i wonder if anyone can shed light on this\n",
    "793461": "After most of times struggling to get a good classifier model. I followed EfB1 + LSTM model by @unkownhihi. And I spent times creating a sequence dataset to feed the model. I use custom Keras Sequence, balanced pair real-fake with different fake for each epoch. I use 160x160 image size of face crop\nBut this also another struggle, could not get the model to get a good validation. My validation set is folder 0,1,2,43-49\n- with frozen backbone, training not able to converge\n- with retrain backbone, my model can converge well but badly overfit. then found bug in training data, since I used padding, some fake data just blank because could not find the images/faces\n- after fix the bug, still could not get good validation loss, somehow my model very confident to guess real as fake or the other way\n- changed  image size to 120x120, center crop from original 160x160, slightly better\n- tried RMSProp, Adam and SGD with different learning rates: RMSProp and Adam can converge fast, but validation loss is erratic. SGD is more stable and not able to converge well. My validation loss is 1.x and tried to submit one get 0.8 on LB\n- last night I change the model to:\n    y = TimeDistributed(cnn)(input_layer) \n    y = LSTM(256, return_sequences=False)(y)    \n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\n    y = Dense(512, activation='relu')(y)\n    y = Dropout(dropout)(y)\nI used SGD with momentum and cyclic learning rate. I was not sure this going to be different and just leave it run. And this morning surprisingly, I got:\npoch 7/20\n1291/1291 [==============================] - 476s 369ms/step - loss: 0.1620 - acc: 0.9330 - val_loss: 0.0334 - val_acc: 0.8122\nI realize that this is not great and may not give me super LB increase, I am just happy I finally breaking my own many attemps hard to reach good validation. But have to wait for 22 more days to do late submission :( to know for sure\n\nwhat a journey... congrats to you all. thanks for sharing\n",
    "793252": "efficientnet b3 pretrained model (transfer learning)\n + dropout &amp; dense layers\n + softmax output.\n\nit showed 88%+ acc &amp; 0.4 loss in 10 mins of GPU run (batchsize = 160)",
    "793202": "BlazeFace for face extraction\nfirst sub- stacked efficientnet b0ns+b1ns+b2ns+b3ns\nsecond sub - stacked(b0ns+b1ns+xception)+ensemble with best model of b1ns\nmixup, \nrandom cutout,\naugmentation, \n1-cycle with more runs,\nweighted average checkpoints\nNo CV but I hope that the large public testset was not too misleading",
    "793153": "GCloud, ensemble",
    "793139": "Step1: Trained model on Fake/Real faces - efficientnetb1 ( other deepfake faces )\nStep2: Used step1 model with LSTM time distribution with Augmentation ( Jpeg compression, Horizontal flip )",
    "793131": "Simple binary classifiers: Xception and Effi B1",
    "794621": ""
  }
}