{
  "id": 145721,
  "title": "1st place solution",
  "url": "/competitions/deepfake-detection-challenge/discussion/145721",
  "author_name": "Selim Seferbekov",
  "post_date": "2020-04-24T08:55:14.612000",
  "votes": 339,
  "comment_count": 102,
  "views": 0,
  "content": "<h1>Keep it simple</h1>\n\n<p>I used a frame-by-frame classification approach as many other competitors did. \nTried a lot of other complex things but in the end it was better to just use a classifier.</p>\n\n<h3>Data preparation</h3>\n\n<ul>\n<li>extracted boxes and landmarks with MTCNN and saved them as json</li>\n<li>extracted crops in original size and saved them as png</li>\n<li>extracted SSIM masks with difference between real and fake and saved them as png</li>\n</ul>\n\n<h3>Face-Detector</h3>\n\n<p>I used simple MTCNN detector.\nInput size for face detector was caluclated for each video depending on video resolution. \n- 2x resize for videos with less than 300 pixels wider side\n- no resize for videos with wider side between 300 and 1000\n- 0.5x resize for videos with wider side &gt; 1000 pixels\n- 0.33x resize for videos with wider side &gt; 1900 pixels</p>\n\n<h3>Input size</h3>\n\n<p>As soon as I discovered that EfficientNets significantly outperform other encoders I used only them in my solution. \nAs I started with B4 I decided to use \"native\" size for that network (380x380). \nDue to memory costraints I did not increase input size even for B7 encoder.</p>\n\n<h3>Margin</h3>\n\n<p>When I generated crops for training I added 30% of face crop size from each side and used only this setting during the competition.</p>\n\n<h3>Encoders</h3>\n\n<p>I tried multiple variants of efficient nets at the beginning:</p>\n\n<ul>\n<li>solo B3 (300x300) - 0.29 public</li>\n<li>solo B4 (380x380) - 0.27 public</li>\n<li>solo B5 (380x380) - 0.25 public</li>\n<li>solo B6 (380x380) - 0.27 public (surprisingly it was worse than B5 and I have not tried B7 until competition last week)</li>\n<li>solo B7 (380x380) - 0.24 public </li>\n</ul>\n\n<p>In the end I used two submits with:\n- 15xB5 (different seeds) with heursitic overfitted for Public LB which was trained with standard augmentations - 10th place on private\n- 7xB7 (different seeds) with more conservative avergaing heursitic and trained with hardcore augmentations - 3rd place on private </p>\n\n<h3>Averaging predictions:</h3>\n\n<p>I used 32 frames for each video.\nFor each model output instead of simple averaging I used the following heuristic  which worked quite well on public leaderbord (0.25 -&gt; 0.22 solo B5).</p>\n\n<p><code>\ndef confident_strategy(pred, t=0.87):\n    pred = np.array(pred)\n    size = len(pred)\n    fakes = np.count_nonzero(pred &amp;gt; t)\n    if fakes &amp;gt; size // 3 and fakes &amp;gt; 11:\n        return np.mean(pred[pred &amp;gt; t])\n    elif np.count_nonzero(pred &amp;lt; 0.2) &amp;gt; 0.6 * size:\n        return np.mean(pred[pred &amp;lt; 0.2])\n    else:\n        return np.mean(pred)\n</code></p>\n\n<p>I.e. I used only confident predictions for averaging if they passed some thresholds. \nThough it worked well on public leaderbord to be safe for the second submit (3rd private) I used more conservative thersholds for real videos. </p>\n\n<h3>Validation strategy</h3>\n\n<p>I used 0-2 folders as holdout at the beginning. \nBut 400 videos from the public test had more correlation with Public LB and I switched to this approach.</p>\n\n<p>I tracked two log loss metrics (on averaged probabilities per video)\n- logloss on real videos \n- logloss on fake videos </p>\n\n<p>From the validation I can say logloss on FAKE videos was much lower than on REAL. I.e. fakes were too easy to spot.\nWhich was not encouraging. I guess public leaderboard also has this characteristic as my validation had strong correlation with public LB. \nIt is not the case on private set as it would contain new/different FaceSwap methods. </p>\n\n<h3>Augmentations</h3>\n\n<p>I used image compression, noise, blur, resize with different interpolations, color jittering, scaling and rotations\n<code>\ndef create_train_transforms(size=380):\n    return Compose([\n        ImageCompression(quality_lower=60, quality_upper=100, p=0.5),\n        GaussNoise(p=0.1),\n        GaussianBlur(blur_limit=3, p=0.05),\n        HorizontalFlip(),\n        IsotropicResize(max_side=size)\n        PadIfNeeded(min_height=size, min_width=size, border_mode=cv2.BORDER_CONSTANT),\n        OneOf([RandomBrightnessContrast(), FancyPCA(), HueSaturationValue()], p=0.7),\n        ToGray(p=0.2),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5),\n    ]\n    )\n</code></p>\n\n<h3>Generalization approach</h3>\n\n<p>I expected it won't be enough for generalization and spent a few weeks working on domain specific augmentations. </p>\n\n<p>I wanted to push models to learn the following properties:\n- visual artifacts (models learn that easily even without any augmentations)\n- different encoding of face from other part of the image. Big margin helps with that. \n- face warping artifacts. Big marging helps with that as well. Related article <a href=\"https://arxiv.org/abs/1811.00656\">https://arxiv.org/abs/1811.00656</a> to catch FWA\n- blending artifacts - here we need either to predict blending mask (<a href=\"https://arxiv.org/abs/1912.13458\">https://arxiv.org/abs/1912.13458</a>) but it's not possible to obtain ground truth mask as there is no big difference in pixels/SSIM on the edge due to blurring and other techniques used to reduce blending artifacts or come up with augmentations that <strong>destroy</strong> visual artifacts.</p>\n\n<p>To catch face blending artefacts:\n1. removed half face horisontally or vertically. Used dlib face convex hulls. \n2. blacked out landmarks (eyes, nose or mouth). Used MTCNN landmarks for that.\n3. blacked out half of the image. To be safe I checked that it will not delete highly confident difference from masks generated with SSIM. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F534152%2Fcd745d9fb666f2316ac699165d629c4f%2Faugmentations.jpg?generation=1587717153025245&amp;alt=media\" alt=\"\"></p>\n\n<p>I don't really know if it helped on private leaderboard though.</p>\n\n<h3>Training schedule</h3>\n\n<ul>\n<li>sampled fake crops based  on number of real crops i.e. <code>fakes.sample(n=num_real, replace=False, random_state=seed)</code></li>\n<li>2500 iterations per epoch</li>\n<li>SGD, momentum=0.9, weight decay=1e-4</li>\n<li>PolyLR with 0.01 starting LR</li>\n<li>75k iterations</li>\n<li>used Apex with mixed precision</li>\n<li>trained on 4 GPUs with SyncBN and DDP. Batch size 16x4 for B5, 12x4 for B7.</li>\n<li>label smoothing - for me label smoothing with 0.01 eps was optimal on public leaderboard. Also it allowed not to use clipping at all. </li>\n</ul>\n\n<h3>Hardware</h3>\n\n<p>Unfortunately I did not ask aws credits from organizers, hoped that my home workstations would be enough -  I was wrong!\nI have two devboxes: one with 2xTitan Vs, the other with 4xTitan Vs</p>\n\n<p>Huge thanks to hostkey provider (<a href=\"https://www.hostkey.com/\">https://www.hostkey.com/</a>) that gave me a grant <a href=\"http://landing.hostkey.com/grants?_ga=2.3307699.2051560741.1587714719-1038061670.1587714719\">http://landing.hostkey.com/grants?_ga=2.3307699.2051560741.1587714719-1038061670.1587714719</a> !\nI got a devbox with 4x2080Ti for two months! Additional 4 GPUs helped me a lot to iterate faster as my models required all of them for training a single model. </p>\n\n<h3>Things that I tried but did not work well enough</h3>\n\n<ul>\n<li>Metric learning - was worse than a simple classificator</li>\n<li>UNet with SSIM difference prediction - got the same LB score.</li>\n<li>Self supervision with blending real faces. Extracted convex hull for real faces warped/resized/compressed face and then blended it back and changed label to fake. Models learned that quickly but that did not give any boost on validation/public leaderboard.</li>\n<li>Self supervision with rotations like here <a href=\"https://stanford-cs221.github.io/autumn2019-extra/posters/110.pdf\">https://stanford-cs221.github.io/autumn2019-extra/posters/110.pdf</a> - did not improve my score</li>\n<li>Temperature scaling - risky, no big difference. Decided not to use it. </li>\n<li>Using steganalysis. Did not imporve validation score.</li>\n<li>RNN - I guess it would work, but due to CPU bottleneck it would be hard to extract multiple chunks in kernel. My first model with RNN + Resnet34 scored around 0.4 on public and a simple classifier was better. </li>\n</ul>\n\n<h3>No External data</h3>\n\n<p>That was a huge concern for me because nothing is basically allowed.  I decided to be safe and did not use any external data.</p>\n\n<h3>Github</h3>\n\n<p><a href=\"https://github.com/selimsef/dfdc_deepfake_challenge\">https://github.com/selimsef/dfdc_deepfake_challenge</a></p>",
  "messages": [
    {
      "id": 818975,
      "postDate": "2020-04-24T08:55:14.613Z",
      "content": "<h1>Keep it simple</h1>\n\n<p>I used a frame-by-frame classification approach as many other competitors did. \nTried a lot of other complex things but in the end it was better to just use a classifier.</p>\n\n<h3>Data preparation</h3>\n\n<ul>\n<li>extracted boxes and landmarks with MTCNN and saved them as json</li>\n<li>extracted crops in original size and saved them as png</li>\n<li>extracted SSIM masks with difference between real and fake and saved them as png</li>\n</ul>\n\n<h3>Face-Detector</h3>\n\n<p>I used simple MTCNN detector.\nInput size for face detector was caluclated for each video depending on video resolution. \n- 2x resize for videos with less than 300 pixels wider side\n- no resize for videos with wider side between 300 and 1000\n- 0.5x resize for videos with wider side &gt; 1000 pixels\n- 0.33x resize for videos with wider side &gt; 1900 pixels</p>\n\n<h3>Input size</h3>\n\n<p>As soon as I discovered that EfficientNets significantly outperform other encoders I used only them in my solution. \nAs I started with B4 I decided to use \"native\" size for that network (380x380). \nDue to memory costraints I did not increase input size even for B7 encoder.</p>\n\n<h3>Margin</h3>\n\n<p>When I generated crops for training I added 30% of face crop size from each side and used only this setting during the competition.</p>\n\n<h3>Encoders</h3>\n\n<p>I tried multiple variants of efficient nets at the beginning:</p>\n\n<ul>\n<li>solo B3 (300x300) - 0.29 public</li>\n<li>solo B4 (380x380) - 0.27 public</li>\n<li>solo B5 (380x380) - 0.25 public</li>\n<li>solo B6 (380x380) - 0.27 public (surprisingly it was worse than B5 and I have not tried B7 until competition last week)</li>\n<li>solo B7 (380x380) - 0.24 public </li>\n</ul>\n\n<p>In the end I used two submits with:\n- 15xB5 (different seeds) with heursitic overfitted for Public LB which was trained with standard augmentations - 10th place on private\n- 7xB7 (different seeds) with more conservative avergaing heursitic and trained with hardcore augmentations - 3rd place on private </p>\n\n<h3>Averaging predictions:</h3>\n\n<p>I used 32 frames for each video.\nFor each model output instead of simple averaging I used the following heuristic  which worked quite well on public leaderbord (0.25 -&gt; 0.22 solo B5).</p>\n\n<p><code>\ndef confident_strategy(pred, t=0.87):\n    pred = np.array(pred)\n    size = len(pred)\n    fakes = np.count_nonzero(pred &amp;gt; t)\n    if fakes &amp;gt; size // 3 and fakes &amp;gt; 11:\n        return np.mean(pred[pred &amp;gt; t])\n    elif np.count_nonzero(pred &amp;lt; 0.2) &amp;gt; 0.6 * size:\n        return np.mean(pred[pred &amp;lt; 0.2])\n    else:\n        return np.mean(pred)\n</code></p>\n\n<p>I.e. I used only confident predictions for averaging if they passed some thresholds. \nThough it worked well on public leaderbord to be safe for the second submit (3rd private) I used more conservative thersholds for real videos. </p>\n\n<h3>Validation strategy</h3>\n\n<p>I used 0-2 folders as holdout at the beginning. \nBut 400 videos from the public test had more correlation with Public LB and I switched to this approach.</p>\n\n<p>I tracked two log loss metrics (on averaged probabilities per video)\n- logloss on real videos \n- logloss on fake videos </p>\n\n<p>From the validation I can say logloss on FAKE videos was much lower than on REAL. I.e. fakes were too easy to spot.\nWhich was not encouraging. I guess public leaderboard also has this characteristic as my validation had strong correlation with public LB. \nIt is not the case on private set as it would contain new/different FaceSwap methods. </p>\n\n<h3>Augmentations</h3>\n\n<p>I used image compression, noise, blur, resize with different interpolations, color jittering, scaling and rotations\n<code>\ndef create_train_transforms(size=380):\n    return Compose([\n        ImageCompression(quality_lower=60, quality_upper=100, p=0.5),\n        GaussNoise(p=0.1),\n        GaussianBlur(blur_limit=3, p=0.05),\n        HorizontalFlip(),\n        IsotropicResize(max_side=size)\n        PadIfNeeded(min_height=size, min_width=size, border_mode=cv2.BORDER_CONSTANT),\n        OneOf([RandomBrightnessContrast(), FancyPCA(), HueSaturationValue()], p=0.7),\n        ToGray(p=0.2),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5),\n    ]\n    )\n</code></p>\n\n<h3>Generalization approach</h3>\n\n<p>I expected it won't be enough for generalization and spent a few weeks working on domain specific augmentations. </p>\n\n<p>I wanted to push models to learn the following properties:\n- visual artifacts (models learn that easily even without any augmentations)\n- different encoding of face from other part of the image. Big margin helps with that. \n- face warping artifacts. Big marging helps with that as well. Related article <a href=\"https://arxiv.org/abs/1811.00656\">https://arxiv.org/abs/1811.00656</a> to catch FWA\n- blending artifacts - here we need either to predict blending mask (<a href=\"https://arxiv.org/abs/1912.13458\">https://arxiv.org/abs/1912.13458</a>) but it's not possible to obtain ground truth mask as there is no big difference in pixels/SSIM on the edge due to blurring and other techniques used to reduce blending artifacts or come up with augmentations that <strong>destroy</strong> visual artifacts.</p>\n\n<p>To catch face blending artefacts:\n1. removed half face horisontally or vertically. Used dlib face convex hulls. \n2. blacked out landmarks (eyes, nose or mouth). Used MTCNN landmarks for that.\n3. blacked out half of the image. To be safe I checked that it will not delete highly confident difference from masks generated with SSIM. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F534152%2Fcd745d9fb666f2316ac699165d629c4f%2Faugmentations.jpg?generation=1587717153025245&amp;alt=media\" alt=\"\"></p>\n\n<p>I don't really know if it helped on private leaderboard though.</p>\n\n<h3>Training schedule</h3>\n\n<ul>\n<li>sampled fake crops based  on number of real crops i.e. <code>fakes.sample(n=num_real, replace=False, random_state=seed)</code></li>\n<li>2500 iterations per epoch</li>\n<li>SGD, momentum=0.9, weight decay=1e-4</li>\n<li>PolyLR with 0.01 starting LR</li>\n<li>75k iterations</li>\n<li>used Apex with mixed precision</li>\n<li>trained on 4 GPUs with SyncBN and DDP. Batch size 16x4 for B5, 12x4 for B7.</li>\n<li>label smoothing - for me label smoothing with 0.01 eps was optimal on public leaderboard. Also it allowed not to use clipping at all. </li>\n</ul>\n\n<h3>Hardware</h3>\n\n<p>Unfortunately I did not ask aws credits from organizers, hoped that my home workstations would be enough -  I was wrong!\nI have two devboxes: one with 2xTitan Vs, the other with 4xTitan Vs</p>\n\n<p>Huge thanks to hostkey provider (<a href=\"https://www.hostkey.com/\">https://www.hostkey.com/</a>) that gave me a grant <a href=\"http://landing.hostkey.com/grants?_ga=2.3307699.2051560741.1587714719-1038061670.1587714719\">http://landing.hostkey.com/grants?_ga=2.3307699.2051560741.1587714719-1038061670.1587714719</a> !\nI got a devbox with 4x2080Ti for two months! Additional 4 GPUs helped me a lot to iterate faster as my models required all of them for training a single model. </p>\n\n<h3>Things that I tried but did not work well enough</h3>\n\n<ul>\n<li>Metric learning - was worse than a simple classificator</li>\n<li>UNet with SSIM difference prediction - got the same LB score.</li>\n<li>Self supervision with blending real faces. Extracted convex hull for real faces warped/resized/compressed face and then blended it back and changed label to fake. Models learned that quickly but that did not give any boost on validation/public leaderboard.</li>\n<li>Self supervision with rotations like here <a href=\"https://stanford-cs221.github.io/autumn2019-extra/posters/110.pdf\">https://stanford-cs221.github.io/autumn2019-extra/posters/110.pdf</a> - did not improve my score</li>\n<li>Temperature scaling - risky, no big difference. Decided not to use it. </li>\n<li>Using steganalysis. Did not imporve validation score.</li>\n<li>RNN - I guess it would work, but due to CPU bottleneck it would be hard to extract multiple chunks in kernel. My first model with RNN + Resnet34 scored around 0.4 on public and a simple classifier was better. </li>\n</ul>\n\n<h3>No External data</h3>\n\n<p>That was a huge concern for me because nothing is basically allowed.  I decided to be safe and did not use any external data.</p>\n\n<h3>Github</h3>\n\n<p><a href=\"https://github.com/selimsef/dfdc_deepfake_challenge\">https://github.com/selimsef/dfdc_deepfake_challenge</a></p>",
      "rawMarkdown": "\n# Keep it simple\nI used a frame-by-frame classification approach as many other competitors did. \nTried a lot of other complex things but in the end it was better to just use a classifier.\n\n### Data preparation\n- extracted boxes and landmarks with MTCNN and saved them as json\n- extracted crops in original size and saved them as png\n- extracted SSIM masks with difference between real and fake and saved them as png\n\n### Face-Detector\nI used simple MTCNN detector.\nInput size for face detector was caluclated for each video depending on video resolution. \n- 2x resize for videos with less than 300 pixels wider side\n- no resize for videos with wider side between 300 and 1000\n- 0.5x resize for videos with wider side &gt; 1000 pixels\n- 0.33x resize for videos with wider side &gt; 1900 pixels\n\n### Input size\nAs soon as I discovered that EfficientNets significantly outperform other encoders I used only them in my solution. \nAs I started with B4 I decided to use \"native\" size for that network (380x380). \nDue to memory costraints I did not increase input size even for B7 encoder.\n\n### Margin\nWhen I generated crops for training I added 30% of face crop size from each side and used only this setting during the competition.\n\n\n### Encoders\nI tried multiple variants of efficient nets at the beginning:\n\n- solo B3 (300x300) - 0.29 public\n- solo B4 (380x380) - 0.27 public\n- solo B5 (380x380) - 0.25 public\n- solo B6 (380x380) - 0.27 public (surprisingly it was worse than B5 and I have not tried B7 until competition last week)\n- solo B7 (380x380) - 0.24 public \n\nIn the end I used two submits with:\n- 15xB5 (different seeds) with heursitic overfitted for Public LB which was trained with standard augmentations - 10th place on private\n- 7xB7 (different seeds) with more conservative avergaing heursitic and trained with hardcore augmentations - 3rd place on private \n\n### Averaging predictions:\nI used 32 frames for each video.\nFor each model output instead of simple averaging I used the following heuristic  which worked quite well on public leaderbord (0.25 -&gt; 0.22 solo B5).\n\n```\ndef confident_strategy(pred, t=0.87):\n    pred = np.array(pred)\n    size = len(pred)\n    fakes = np.count_nonzero(pred &gt; t)\n    if fakes &gt; size // 3 and fakes &gt; 11:\n        return np.mean(pred[pred &gt; t])\n    elif np.count_nonzero(pred &lt; 0.2) &gt; 0.6 * size:\n        return np.mean(pred[pred &lt; 0.2])\n    else:\n        return np.mean(pred)\n```\n\nI.e. I used only confident predictions for averaging if they passed some thresholds. \nThough it worked well on public leaderbord to be safe for the second submit (3rd private) I used more conservative thersholds for real videos. \n\n### Validation strategy\n\nI used 0-2 folders as holdout at the beginning. \nBut 400 videos from the public test had more correlation with Public LB and I switched to this approach.\n\nI tracked two log loss metrics (on averaged probabilities per video)\n- logloss on real videos \n- logloss on fake videos \n\nFrom the validation I can say logloss on FAKE videos was much lower than on REAL. I.e. fakes were too easy to spot.\nWhich was not encouraging. I guess public leaderboard also has this characteristic as my validation had strong correlation with public LB. \nIt is not the case on private set as it would contain new/different FaceSwap methods. \n\n### Augmentations\n\nI used image compression, noise, blur, resize with different interpolations, color jittering, scaling and rotations\n```\ndef create_train_transforms(size=380):\n    return Compose([\n        ImageCompression(quality_lower=60, quality_upper=100, p=0.5),\n        GaussNoise(p=0.1),\n        GaussianBlur(blur_limit=3, p=0.05),\n        HorizontalFlip(),\n        IsotropicResize(max_side=size)\n        PadIfNeeded(min_height=size, min_width=size, border_mode=cv2.BORDER_CONSTANT),\n        OneOf([RandomBrightnessContrast(), FancyPCA(), HueSaturationValue()], p=0.7),\n        ToGray(p=0.2),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5),\n    ]\n    )\n```\n### Generalization approach\nI expected it won't be enough for generalization and spent a few weeks working on domain specific augmentations. \n\nI wanted to push models to learn the following properties:\n- visual artifacts (models learn that easily even without any augmentations)\n- different encoding of face from other part of the image. Big margin helps with that. \n- face warping artifacts. Big marging helps with that as well. Related article https://arxiv.org/abs/1811.00656 to catch FWA\n- blending artifacts - here we need either to predict blending mask (https://arxiv.org/abs/1912.13458) but it's not possible to obtain ground truth mask as there is no big difference in pixels/SSIM on the edge due to blurring and other techniques used to reduce blending artifacts or come up with augmentations that **destroy** visual artifacts.\n\nTo catch face blending artefacts:\n1. removed half face horisontally or vertically. Used dlib face convex hulls. \n2. blacked out landmarks (eyes, nose or mouth). Used MTCNN landmarks for that.\n3. blacked out half of the image. To be safe I checked that it will not delete highly confident difference from masks generated with SSIM. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F534152%2Fcd745d9fb666f2316ac699165d629c4f%2Faugmentations.jpg?generation=1587717153025245&amp;alt=media)\n\nI don't really know if it helped on private leaderboard though.\n\n### Training schedule\n- sampled fake crops based  on number of real crops i.e. `fakes.sample(n=num_real, replace=False, random_state=seed)`\n- 2500 iterations per epoch\n- SGD, momentum=0.9, weight decay=1e-4\n- PolyLR with 0.01 starting LR\n- 75k iterations\n- used Apex with mixed precision\n- trained on 4 GPUs with SyncBN and DDP. Batch size 16x4 for B5, 12x4 for B7.\n- label smoothing - for me label smoothing with 0.01 eps was optimal on public leaderboard. Also it allowed not to use clipping at all. \n\n \n\n### Hardware\nUnfortunately I did not ask aws credits from organizers, hoped that my home workstations would be enough -  I was wrong!\nI have two devboxes: one with 2xTitan Vs, the other with 4xTitan Vs\n\nHuge thanks to hostkey provider (https://www.hostkey.com/) that gave me a grant http://landing.hostkey.com/grants?_ga=2.3307699.2051560741.1587714719-1038061670.1587714719 !\nI got a devbox with 4x2080Ti for two months! Additional 4 GPUs helped me a lot to iterate faster as my models required all of them for training a single model. \n\n\n### Things that I tried but did not work well enough\n- Metric learning - was worse than a simple classificator\n- UNet with SSIM difference prediction - got the same LB score.\n- Self supervision with blending real faces. Extracted convex hull for real faces warped/resized/compressed face and then blended it back and changed label to fake. Models learned that quickly but that did not give any boost on validation/public leaderboard.\n- Self supervision with rotations like here https://stanford-cs221.github.io/autumn2019-extra/posters/110.pdf - did not improve my score\n- Temperature scaling - risky, no big difference. Decided not to use it. \n- Using steganalysis. Did not imporve validation score.\n- RNN - I guess it would work, but due to CPU bottleneck it would be hard to extract multiple chunks in kernel. My first model with RNN + Resnet34 scored around 0.4 on public and a simple classifier was better. \n\n\n### No External data\nThat was a huge concern for me because nothing is basically allowed.  I decided to be safe and did not use any external data.\n\n\n### Github\nhttps://github.com/selimsef/dfdc_deepfake_challenge\n",
      "votes": 338
    },
    {
      "id": 876545,
      "postDate": "2020-06-06T19:00:59.370Z",
      "content": "<p>Opensourced my solution <a href=\"https://github.com/selimsef/dfdc_deepfake_challenge\">https://github.com/selimsef/dfdc_deepfake_challenge</a></p>",
      "rawMarkdown": "Opensourced my solution https://github.com/selimsef/dfdc_deepfake_challenge",
      "votes": 27,
      "replies": [
        {
          "id": 876557,
          "postDate": "2020-06-06T19:20:31.117Z",
          "content": "<p>Awesome! </p>",
          "rawMarkdown": "Awesome! ",
          "votes": 1
        },
        {
          "id": 883568,
          "postDate": "2020-06-12T18:30:18.823Z",
          "content": "<p>Bravo Selim! From top#3 to top#1 after final review. Weird, some teams have been re-scored.</p>",
          "rawMarkdown": "Bravo Selim! From top#3 to top#1 after final review. Weird, some teams have been re-scored.",
          "votes": 4
        }
      ]
    },
    {
      "id": 883763,
      "postDate": "2020-06-12T22:35:25.113Z",
      "content": "<p>Just want to say huge congratulation, I really enjoyed reading your solution, and would learn a lot from further studying your technique :)</p>",
      "rawMarkdown": "Just want to say huge congratulation, I really enjoyed reading your solution, and would learn a lot from further studying your technique :)",
      "votes": 12,
      "replies": [
        {
          "id": 883773,
          "postDate": "2020-06-12T22:55:58.267Z",
          "content": "<p>Thank you! </p>\n\n<p>I was really surprised that your submission that used external data with CC-BY license led to its disqualification. \nAt the beginning I wanted to collect CC-BY videos as well as they seemed to be legit. Asked about that a lot of times. Beacuse I have not received a confirmation from organizers I considered this data in the gray area (neither organizers nor host could provide definitive answers) and decided to be on the safe side. <br>\nHopefully this situation will be a lesson for Kaggle to make rules very clear for the future competitions.</p>",
          "rawMarkdown": "Thank you! \n\nI was really surprised that your submission that used external data with CC-BY license led to its disqualification. \nAt the beginning I wanted to collect CC-BY videos as well as they seemed to be legit. Asked about that a lot of times. Beacuse I have not received a confirmation from organizers I considered this data in the gray area (neither organizers nor host could provide definitive answers) and decided to be on the safe side.  \nHopefully this situation will be a lesson for Kaggle to make rules very clear for the future competitions.",
          "votes": 22
        }
      ]
    },
    {
      "id": 1229858,
      "postDate": "2021-03-07T16:29:39.273Z",
      "content": "<p>Great work! Congratulations!</p>\n<ol>\n<li>MTCNN frequently has false detections. How did you get rid of them?</li>\n<li>What did you do with multiple actor videos?</li>\n</ol>",
      "rawMarkdown": "Great work! Congratulations!\n\n1. MTCNN frequently has false detections. How did you get rid of them?\n2. What did you do with multiple actor videos?",
      "votes": 1
    },
    {
      "id": 819009,
      "postDate": "2020-04-24T09:20:43.720Z",
      "content": "<p>Really excellent work. I like the domain-specific augmentations and I'm sure it's things like that that helped you hold on to your excellent public LB position on the 'wild' private LB data, even if it didn't manifest in a big public LB change.</p>",
      "rawMarkdown": "Really excellent work. I like the domain-specific augmentations and I'm sure it's things like that that helped you hold on to your excellent public LB position on the 'wild' private LB data, even if it didn't manifest in a big public LB change.",
      "votes": 4
    },
    {
      "id": 887934,
      "postDate": "2020-06-16T02:40:48.100Z",
      "content": "<p>Thanks a lot for sharing and congratulations for winning, this is amazing really to have the first in such competition firstly saying: \"<strong>Keep it simple</strong>\".\nI do have some questions if you wouldn't mind: 1- how many frames did you use per video, only one frame? \n2- you trained the model with the extracted faces or also these SSIM masks?\n3- when you tried only B5 efficientNet model; for about how many epochs did you train the model to yield this loss?\nlastly, actually I don't understand what polyLR you mean, do you mean that you tried several learning rates? if so, then what ranges of values did you use like from 0.01 to what with how many steps?</p>",
      "rawMarkdown": "Thanks a lot for sharing and congratulations for winning, this is amazing really to have the first in such competition firstly saying: \"**Keep it simple**\".\nI do have some questions if you wouldn't mind: 1- how many frames did you use per video, only one frame? \n2- you trained the model with the extracted faces or also these SSIM masks?\n3- when you tried only B5 efficientNet model; for about how many epochs did you train the model to yield this loss?\nlastly, actually I don't understand what polyLR you mean, do you mean that you tried several learning rates? if so, then what ranges of values did you use like from 0.01 to what with how many steps?",
      "votes": 1,
      "replies": [
        {
          "id": 889149,
          "postDate": "2020-06-16T19:50:35.173Z",
          "content": "<ol>\n<li>32 frames</li>\n<li>only faces, used SSIM masks during my UNet experiments</li>\n<li>for B5 it was 30 epochs with 2500 steps each, batch size=80</li>\n</ol>\n\n<p>PolyLR - polynomial learning rate decay. It's usually used for semantic segmentation models like DeepLabV3. Cosine decay would be better I guess. </p>",
          "rawMarkdown": "1. 32 frames\n2. only faces, used SSIM masks during my UNet experiments\n3. for B5 it was 30 epochs with 2500 steps each, batch size=80\n\nPolyLR - polynomial learning rate decay. It's usually used for semantic segmentation models like DeepLabV3. Cosine decay would be better I guess. ",
          "votes": 4
        },
        {
          "id": 889264,
          "postDate": "2020-06-16T21:29:45.770Z",
          "content": "<p>very good detailed answer, thanks for your help ✌</p>",
          "rawMarkdown": "very good detailed answer, thanks for your help ✌"
        }
      ]
    },
    {
      "id": 826090,
      "postDate": "2020-04-29T12:57:40.743Z",
      "content": "<p>I'm wondering about the less glamorous side of this (the things that didn't work), can you please share what other complex things you tried that didn't work :) </p>",
      "rawMarkdown": "I'm wondering about the less glamorous side of this (the things that didn't work), can you please share what other complex things you tried that didn't work :) ",
      "votes": 1,
      "replies": [
        {
          "id": 826186,
          "postDate": "2020-04-29T13:51:10.257Z",
          "content": "<p>First of all I read overviews of existing methods <a href=\"https://arxiv.org/abs/2001.06564\">https://arxiv.org/abs/2001.06564</a> and <a href=\"https://arxiv.org/abs/2001.00179\">https://arxiv.org/abs/2001.00179</a> . Then investigated some methods if they seemed promising or came up with my own ideas.</p>\n\n<p><strong>Multitask-Learning</strong> (did not affect public leaderboard)\n- segmentation with soft probabilities obtained by SSIM masks. It worked and highlighed artefacts correctly.\n- segmentation with thresholded SSIM diffs masks\n- segmentation with face convex hull masks</p>\n\n<p><strong>Using low level CNN features</strong> (did not affect public leaderboard) - <a href=\"https://arxiv.org/abs/1909.06122\">https://arxiv.org/abs/1909.06122</a></p>\n\n<p><strong>Using contrastive loss and pairwise learning</strong></p>\n\n<p><strong>CNN + RNN on features - was overfitting badly</strong></p>\n\n<p><strong>Different pooling layers</strong>, GWAP worked well enough but again no improvements, paper <a href=\"https://arxiv.org/pdf/1809.08264.pdf\">https://arxiv.org/pdf/1809.08264.pdf</a>, <a href=\"https://github.com/BloodAxe/pytorch-toolbelt/blob/c830e02eac3162f6066ccd9fd5b02c3c2cb19018/pytorch_toolbelt/modules/pooling.py#L45\">https://github.com/BloodAxe/pytorch-toolbelt/blob/c830e02eac3162f6066ccd9fd5b02c3c2cb19018/pytorch_toolbelt/modules/pooling.py#L45</a></p>\n\n<p><strong>Using deep and shallow networks simultaneusly</strong></p>\n\n<p><strong>Multicrop network</strong> - i.e. 5 crops from image: (topleft, topright, bottomleft, bottomright, center crop), for 256x256  got five 128x128 =&gt; concatenated CNN features + FC + classifier. </p>\n\n<p><strong>Self-supervision</strong>\n- real face warping + blending and changing label to fake\n- rotate image 0, 90, 180, 120 and predicting angle as well</p>\n\n<p><strong>Data cleaning</strong>\n- used SSIM masks and removed failed fakes, or not fakes from multiface videos\n- the loss on validation dropped dramatically, but it gave worse results on public LB</p>\n\n<p>I regret that I have not tried NeXtVlad to achieve better video classification quality. </p>",
          "rawMarkdown": "First of all I read overviews of existing methods https://arxiv.org/abs/2001.06564 and https://arxiv.org/abs/2001.00179 . Then investigated some methods if they seemed promising or came up with my own ideas.\n\n**Multitask-Learning** (did not affect public leaderboard)\n- segmentation with soft probabilities obtained by SSIM masks. It worked and highlighed artefacts correctly.\n- segmentation with thresholded SSIM diffs masks\n- segmentation with face convex hull masks\n\n**Using low level CNN features** (did not affect public leaderboard) - https://arxiv.org/abs/1909.06122\n\n**Using contrastive loss and pairwise learning**\n\n**CNN + RNN on features - was overfitting badly**\n\n**Different pooling layers**, GWAP worked well enough but again no improvements, paper https://arxiv.org/pdf/1809.08264.pdf, https://github.com/BloodAxe/pytorch-toolbelt/blob/c830e02eac3162f6066ccd9fd5b02c3c2cb19018/pytorch_toolbelt/modules/pooling.py#L45\n\n**Using deep and shallow networks simultaneusly**\n\n**Multicrop network** - i.e. 5 crops from image: (topleft, topright, bottomleft, bottomright, center crop), for 256x256  got five 128x128 =&gt; concatenated CNN features + FC + classifier. \n\n**Self-supervision**\n- real face warping + blending and changing label to fake\n- rotate image 0, 90, 180, 120 and predicting angle as well\n\n**Data cleaning**\n- used SSIM masks and removed failed fakes, or not fakes from multiface videos\n- the loss on validation dropped dramatically, but it gave worse results on public LB\n\nI regret that I have not tried NeXtVlad to achieve better video classification quality. \n\n\n\n",
          "votes": 8
        }
      ]
    },
    {
      "id": 820046,
      "postDate": "2020-04-25T05:45:34.273Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": 1
    },
    {
      "id": 819994,
      "postDate": "2020-04-25T04:38:53.123Z",
      "content": "<p>Well done Selim. Combining public+private you have the best model by far 👍  Nice idea to track log loss for each class separately. Did you perform augmentation on local validation set?</p>",
      "rawMarkdown": "Well done Selim. Combining public+private you have the best model by far 👍  Nice idea to track log loss for each class separately. Did you perform augmentation on local validation set?",
      "votes": 1,
      "replies": [
        {
          "id": 820212,
          "postDate": "2020-04-25T08:37:30.253Z",
          "content": "<p>Actually no. If we consider that private set has 50% of synthetic data then other top-5 competitiors' models performed much better on organic videos than my models. One of the way to achieve that (no brainer) - to add organic videos the to trainining set. </p>\n\n<p>I did not perform any augmentations on my local validation set. </p>",
          "rawMarkdown": "Actually no. If we consider that private set has 50% of synthetic data then other top-5 competitiors' models performed much better on organic videos than my models. One of the way to achieve that (no brainer) - to add organic videos the to trainining set. \n\nI did not perform any augmentations on my local validation set. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 888228,
      "postDate": "2020-06-16T07:54:50.760Z",
      "content": "<p>Congrats on the solution.  Good writeup too, it feels like if this was easy.  I'm sure it wasn't.  Great work.</p>\n\n<blockquote>\n  <p>decided to be safe and did not use any external data.</p>\n</blockquote>\n\n<p>Only paranoids survive.</p>\n\n<p>Unfortunately.</p>",
      "rawMarkdown": "Congrats on the solution.  Good writeup too, it feels like if this was easy.  I'm sure it wasn't.  Great work.\n\n&gt;  decided to be safe and did not use any external data.\n\nOnly paranoids survive.\n\nUnfortunately.\n",
      "votes": 2
    },
    {
      "id": 828913,
      "postDate": "2020-05-01T11:36:20.857Z",
      "content": "<p>Quite a robust solution for sure, Are there any sources you'd suggest to improve one's understanding of image recognition?</p>",
      "rawMarkdown": "Quite a robust solution for sure, Are there any sources you'd suggest to improve one's understanding of image recognition?",
      "votes": 2
    },
    {
      "id": 822500,
      "postDate": "2020-04-27T00:40:01.830Z",
      "content": "<p>Congratulations and thank you for sharing! Your solution is very similar to mine for the most part.  One difference is that I only used 5 B5 models in my final ensemble (due to time and hardware constraints). I also included audio which I think severely hurt me on the private LB but marginally helped on CV and public LB.</p>\n\n<p>1)  What EfficientNet library did you use?\n2)  What was your experience with audio analysis and why did you decide not to use it?</p>\n\n<p>I too did not use any datasets besides the provided one or scrape any videos from YouTube as I was afraid it might not be allowed. It is unfortunate that the hosts did not clarify. It is unfair to people who only used the provided dataset.</p>",
      "rawMarkdown": "Congratulations and thank you for sharing! Your solution is very similar to mine for the most part.  One difference is that I only used 5 B5 models in my final ensemble (due to time and hardware constraints). I also included audio which I think severely hurt me on the private LB but marginally helped on CV and public LB.\n\n1)\tWhat EfficientNet library did you use?\n2)\tWhat was your experience with audio analysis and why did you decide not to use it?\n\nI too did not use any datasets besides the provided one or scrape any videos from YouTube as I was afraid it might not be allowed. It is unfortunate that the hosts did not clarify. It is unfair to people who only used the provided dataset.",
      "votes": 2,
      "replies": [
        {
          "id": 822901,
          "postDate": "2020-04-27T08:40:39.870Z",
          "content": "<p>I used timm library <a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a> and fine-tuned noisy student weights. But ordinary imagenet weights gave the same results.\nI have not even tried audio fake detection as it seemed that the main purpose of the challenge is to detect faceswap related deepfakes.  That was a right choice I guess. Though maybe other winning solutions use audio deepfakes, hard to say as they did not post descriptions yet.</p>",
          "rawMarkdown": "I used timm library https://github.com/rwightman/pytorch-image-models and fine-tuned noisy student weights. But ordinary imagenet weights gave the same results.\nI have not even tried audio fake detection as it seemed that the main purpose of the challenge is to detect faceswap related deepfakes.  That was a right choice I guess. Though maybe other winning solutions use audio deepfakes, hard to say as they did not post descriptions yet.",
          "votes": 3
        }
      ]
    },
    {
      "id": 821739,
      "postDate": "2020-04-26T11:41:26.910Z",
      "content": "<p>This looks great</p>",
      "rawMarkdown": "This looks great",
      "votes": 2
    },
    {
      "id": 819306,
      "postDate": "2020-04-24T14:01:09.293Z",
      "content": "<p><a href=\"/selimsef\">@selimsef</a>  many congrats and thanks</p>\n\n<p>could u help understand this portion\n\"extracted SSIM masks with difference between real and fake and saved them as png\"</p>\n\n<p>1) what ssim stands for \n2) how to perform that</p>",
      "rawMarkdown": "@selimsef  many congrats and thanks\n\ncould u help understand this portion\n\"extracted SSIM masks with difference between real and fake and saved them as png\"\n\n1) what ssim stands for \n2) how to perform that",
      "votes": 2,
      "replies": [
        {
          "id": 819473,
          "postDate": "2020-04-24T16:05:44.830Z",
          "content": "<p>answered before <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145721#819137\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145721#819137</a></p>",
          "rawMarkdown": "answered before https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145721#819137",
          "votes": 2
        },
        {
          "id": 822111,
          "postDate": "2020-04-26T17:12:16.887Z",
          "content": "<p><a href=\"/selimsef\">@selimsef</a>  thanks i got the ans below while u were explaining bibek..\n1) how this masks  were used in training \n2) How did u decide the augmentations to use</p>",
          "rawMarkdown": "@selimsef  thanks i got the ans below while u were explaining bibek..\n1) how this masks  were used in training \n2) How did u decide the augmentations to use"
        },
        {
          "id": 822262,
          "postDate": "2020-04-26T19:38:06.707Z",
          "content": "<ol>\n<li>used for U-Net experiments and partially for augmentations</li>\n<li>My intuition was that visual artefacts (strange eyes, mouth) are quite easy to spot. Therefore I needed to destroy them. In addition a lot of works use something like grid dropout (<a href=\"https://arxiv.org/abs/2001.04086\">https://arxiv.org/abs/2001.04086</a>) etc. for better generalization. Here it applies as well, just needs to be done carefully. </li>\n</ol>",
          "rawMarkdown": "1. used for U-Net experiments and partially for augmentations\n2. My intuition was that visual artefacts (strange eyes, mouth) are quite easy to spot. Therefore I needed to destroy them. In addition a lot of works use something like grid dropout (https://arxiv.org/abs/2001.04086) etc. for better generalization. Here it applies as well, just needs to be done carefully. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 819302,
      "postDate": "2020-04-24T13:55:01.463Z",
      "content": "<p>I think what kept me from getting a gold medal in this competition was the margin around the faces. My approach was very similar to yours in the sense of keeping it simple and doing heavy augmentations, although I added histogram equalization on the pipeline (this got me a huge boost on the public leaderboard at the time). </p>\n\n<p>But I kept a very tiny margin at the end, and looking back at it, I believe this was what kept me from reaching a better position (to be honest I never thought about this during the competition). Congratz on the results!</p>",
      "rawMarkdown": "I think what kept me from getting a gold medal in this competition was the margin around the faces. My approach was very similar to yours in the sense of keeping it simple and doing heavy augmentations, although I added histogram equalization on the pipeline (this got me a huge boost on the public leaderboard at the time). \n\nBut I kept a very tiny margin at the end, and looking back at it, I believe this was what kept me from reaching a better position (to be honest I never thought about this during the competition). Congratz on the results!",
      "votes": 2
    },
    {
      "id": 825391,
      "postDate": "2020-04-29T01:31:50.703Z",
      "content": "<p>Congratulations and thanks for your sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for your sharing!"
    },
    {
      "id": 2597081,
      "postDate": "2024-01-11T13:54:41.937Z",
      "content": "<p>Hi from 2024))) Very cool solution. Thanks.</p>",
      "rawMarkdown": "Hi from 2024))) Very cool solution. Thanks."
    },
    {
      "id": 1064740,
      "postDate": "2020-10-30T13:26:14.793Z",
      "content": "<p>Hi Selim, congratulations for your first place!<br>\nMay I ask one simple question for your solution? Did you ensemble models which are all trained with same image size? Do you think ensembling different image size models will harm prediction score?<br>\nI'm just curious of your ensemble was based on just different seed and same image size :)<br>\nthanks!</p>",
      "rawMarkdown": "Hi Selim, congratulations for your first place!\nMay I ask one simple question for your solution? Did you ensemble models which are all trained with same image size? Do you think ensembling different image size models will harm prediction score?\nI'm just curious of your ensemble was based on just different seed and same image size :)\nthanks!",
      "replies": [
        {
          "id": 1066921,
          "postDate": "2020-11-02T08:03:25.200Z",
          "content": "<p>Yes, this ensemble only helps to calibrate probabilities. Using different resolutions would be better. I trained B7 in the last days of the challenge and did not have time to test different resolutions, that's why I decided to just ensemble multiple checkpoints trained on differently sampled data. </p>",
          "rawMarkdown": "Yes, this ensemble only helps to calibrate probabilities. Using different resolutions would be better. I trained B7 in the last days of the challenge and did not have time to test different resolutions, that's why I decided to just ensemble multiple checkpoints trained on differently sampled data. ",
          "votes": 1
        },
        {
          "id": 1067294,
          "postDate": "2020-11-02T13:11:19.403Z",
          "content": "<p>thanks for the comment :) have a nice day!</p>",
          "rawMarkdown": "thanks for the comment :) have a nice day!"
        }
      ]
    },
    {
      "id": 904221,
      "postDate": "2020-06-27T12:53:14.757Z",
      "content": "<p>In your code uploaded on github; thanks for sharing, I've to concerns:\n1- Is there a reason for choosing AvgPooling over MaxPooling, same for GlobalAvgP over GlbalMaxP? or you think that it didn't differ much.\n2-  Did you freeze the layers of  efficientNet before usage?</p>",
      "rawMarkdown": "In your code uploaded on github; thanks for sharing, I've to concerns:\n1- Is there a reason for choosing AvgPooling over MaxPooling, same for GlobalAvgP over GlbalMaxP? or you think that it didn't differ much.\n2-  Did you freeze the layers of  efficientNet before usage?",
      "replies": [
        {
          "id": 904588,
          "postDate": "2020-06-27T18:34:47.910Z",
          "content": "<p>1 - GlobalAvgP is a default choice for classification. Actually I tried GlobalMaxP and GWAP - they were not better.\n2 - no, did not freeze anything</p>",
          "rawMarkdown": "1 - GlobalAvgP is a default choice for classification. Actually I tried GlobalMaxP and GWAP - they were not better.\n2 - no, did not freeze anything",
          "votes": 1
        }
      ]
    },
    {
      "id": 891668,
      "postDate": "2020-06-18T11:17:22.133Z",
      "content": "<p>Great solution!\nI have one question about frame preprocessing technique: you used <code>IsotropicResize</code>. I read the source code and don't quite understand, what is the difference between your <code>IsotropicResize</code> and <code>A.LongestMaxSize</code>?</p>",
      "rawMarkdown": "Great solution!\nI have one question about frame preprocessing technique: you used `IsotropicResize`. I read the source code and don't quite understand, what is the difference between your `IsotropicResize` and `A.LongestMaxSize`?",
      "replies": [
        {
          "id": 892382,
          "postDate": "2020-06-18T20:51:44.790Z",
          "content": "<p>I did not know it existed, otherwise I would reuse it. The only difference is that I used different interpolations for scaling up/down. </p>",
          "rawMarkdown": "I did not know it existed, otherwise I would reuse it. The only difference is that I used different interpolations for scaling up/down. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 885751,
      "postDate": "2020-06-14T13:05:58.777Z",
      "content": "<p><strong>Easy does it!!</strong></p>\n\n<blockquote>\n  <p>Great work. Inspirational for me. 🙌 </p>\n</blockquote>",
      "rawMarkdown": "**Easy does it!!**\n&gt; Great work. Inspirational for me. 🙌 "
    },
    {
      "id": 883945,
      "postDate": "2020-06-13T06:25:35.697Z",
      "content": "<p>The result is so SWEET to you.</p>",
      "rawMarkdown": "The result is so SWEET to you."
    },
    {
      "id": 883561,
      "postDate": "2020-06-12T18:23:39.173Z",
      "content": "<p>Congratulations on the upgrade <a href=\"/selimsef\">@selimsef</a>. Does anyone know what happened here? Is it a reproducibility issue with the other competitors?</p>",
      "rawMarkdown": "Congratulations on the upgrade @selimsef. Does anyone know what happened here? Is it a reproducibility issue with the other competitors?"
    },
    {
      "id": 883533,
      "postDate": "2020-06-12T18:04:04.150Z",
      "content": "<p>Congrats <a href=\"/selimsef\">@selimsef</a> for 1st place!!!!!</p>",
      "rawMarkdown": "Congrats @selimsef for 1st place!!!!!"
    },
    {
      "id": 828514,
      "postDate": "2020-05-01T05:38:57.440Z",
      "content": "<p>thnks</p>",
      "rawMarkdown": "thnks\n"
    },
    {
      "id": 828505,
      "postDate": "2020-05-01T05:21:21.567Z",
      "content": "<p>I like your term: \"<strong>hardcore augmentations</strong>\"</p>",
      "rawMarkdown": "I like your term: \"**hardcore augmentations**\""
    },
    {
      "id": 828096,
      "postDate": "2020-04-30T19:02:28.683Z",
      "content": "<p>Awesome!</p>",
      "rawMarkdown": "Awesome!"
    },
    {
      "id": 827768,
      "postDate": "2020-04-30T14:28:23.550Z",
      "content": "<p>Thank you so much, very detailed explanation</p>",
      "rawMarkdown": "Thank you so much, very detailed explanation"
    },
    {
      "id": 826795,
      "postDate": "2020-04-29T21:38:10.233Z",
      "content": "<p>Good job!</p>",
      "rawMarkdown": "Good job!"
    },
    {
      "id": 825870,
      "postDate": "2020-04-29T09:46:03.687Z",
      "content": "<p>Great job! Will you make a github repo for the training code ?  thanks for sharing !</p>",
      "rawMarkdown": "Great job! Will you make a github repo for the training code ?  thanks for sharing !",
      "replies": [
        {
          "id": 825894,
          "postDate": "2020-04-29T10:04:00.240Z",
          "content": "<p>Yes, after the official results are announced. </p>",
          "rawMarkdown": "Yes, after the official results are announced. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 825551,
      "postDate": "2020-04-29T05:01:25.673Z",
      "content": "<p>Congratulations and thanks for your sharing :)\nI have a question. Did you use early-stopping via validation set? If you don't, what is the reason for that?</p>",
      "rawMarkdown": "Congratulations and thanks for your sharing :)\nI have a question. Did you use early-stopping via validation set? If you don't, what is the reason for that?",
      "replies": [
        {
          "id": 825637,
          "postDate": "2020-04-29T06:26:07.760Z",
          "content": "<p>I saved checkpoint for every epoch, then selected best ones looking at validation loss graphs (fake loss, real loss, loss)</p>",
          "rawMarkdown": " I saved checkpoint for every epoch, then selected best ones looking at validation loss graphs (fake loss, real loss, loss)",
          "votes": 2
        }
      ]
    },
    {
      "id": 824076,
      "postDate": "2020-04-28T06:23:38.240Z",
      "content": "<p>Hearty Congratulations Selim ! Thank you for sharing the solution. Hopefully we can learn from this.</p>",
      "rawMarkdown": "Hearty Congratulations Selim ! Thank you for sharing the solution. Hopefully we can learn from this."
    },
    {
      "id": 823672,
      "postDate": "2020-04-27T19:52:16.917Z",
      "content": "<p>Great work!</p>",
      "rawMarkdown": "Great work!"
    },
    {
      "id": 822516,
      "postDate": "2020-04-27T01:00:59.590Z",
      "content": "<p>Congrats, great work!</p>",
      "rawMarkdown": "Congrats, great work!"
    },
    {
      "id": 822434,
      "postDate": "2020-04-26T22:28:59.660Z",
      "content": "<p>Good work.</p>",
      "rawMarkdown": "Good work."
    },
    {
      "id": 822350,
      "postDate": "2020-04-26T20:43:17.877Z",
      "content": "<p>This seems good to me</p>",
      "rawMarkdown": "This seems good to me"
    },
    {
      "id": 822199,
      "postDate": "2020-04-26T18:41:03.537Z",
      "content": "<p>Good job👍 </p>",
      "rawMarkdown": "Good job👍 "
    },
    {
      "id": 822190,
      "postDate": "2020-04-26T18:31:13.473Z",
      "content": "<p>Congratulations! Awesome work👍 </p>",
      "rawMarkdown": "Congratulations! Awesome work👍 "
    },
    {
      "id": 822127,
      "postDate": "2020-04-26T17:26:12.660Z",
      "content": "<p>this is great</p>",
      "rawMarkdown": "this is great"
    },
    {
      "id": 821988,
      "postDate": "2020-04-26T15:10:17.273Z",
      "content": "<p><a href=\"/selimsef\">@selimsef</a> thanks for your good sharing. I saw your backbone model is efficient. and the baseline B3 model is 0.29. so good score in LB. Did you finetune imageNet base pre-trained model?  we also tried efficient, the score was not good and gave up. </p>",
      "rawMarkdown": "@selimsef thanks for your good sharing. I saw your backbone model is efficient. and the baseline B3 model is 0.29. so good score in LB. Did you finetune imageNet base pre-trained model?  we also tried efficient, the score was not good and gave up. ",
      "replies": [
        {
          "id": 822146,
          "postDate": "2020-04-26T17:46:19.773Z",
          "content": "<p>Yes, I fine-tuned ImageNet pretrained weights. Though it would be interesting to compare results with training from scratch, the task is quite different, maybe pretrained weights don't even matter here. </p>",
          "rawMarkdown": "Yes, I fine-tuned ImageNet pretrained weights. Though it would be interesting to compare results with training from scratch, the task is quite different, maybe pretrained weights don't even matter here. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 821783,
      "postDate": "2020-04-26T12:22:45.383Z",
      "content": "<p>I guess blazeface wud have given better results instead of using MTCNN to detect faces...but good job indeed</p>",
      "rawMarkdown": "I guess blazeface wud have given better results instead of using MTCNN to detect faces...but good job indeed"
    },
    {
      "id": 821231,
      "postDate": "2020-04-26T03:11:10.527Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice"
    },
    {
      "id": 821165,
      "postDate": "2020-04-26T00:48:46.717Z",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice"
    },
    {
      "id": 820149,
      "postDate": "2020-04-25T07:26:20.863Z",
      "content": "<p>Thanks for your insights and congratz on keeping the place after the shakedown!</p>\n\n<p>Have you used face landmarks and SSIM masks you've extracted?\nIf I get it right, you've used landmarks to crop face segments and SSIM masks to check that the significant difference is maintained?</p>\n\n<p>How many frames per video have you extracted for yourt training dataset? (and how many have you actually used - because <code>75k iterations * 16*4 = 4'800'000</code> given you use different frames each epoch)</p>",
      "rawMarkdown": "Thanks for your insights and congratz on keeping the place after the shakedown!\n\nHave you used face landmarks and SSIM masks you've extracted?\nIf I get it right, you've used landmarks to crop face segments and SSIM masks to check that the significant difference is maintained?\n\nHow many frames per video have you extracted for yourt training dataset? (and how many have you actually used - because `75k iterations * 16*4 = 4'800'000` given you use different frames each epoch)",
      "replies": [
        {
          "id": 820207,
          "postDate": "2020-04-25T08:31:33.460Z",
          "content": "<p>For SSIM masks and landmarks\n1. used precomputed MTCNN landmarks to dropout eyes, nose, mouth from images\n2. used precomputed SSIM masks  to check that  half image/grid dropout augmentation doesn't remove all fake data\n4. calculated convex hull landmarks with dlib on the fly to dropout half face\n3. used precomputed SSIM masks for UNet experiments</p>\n\n<p>In my training set I extracted 30 frames per video. Used random sampling for each epoch.</p>",
          "rawMarkdown": "For SSIM masks and landmarks\n1. used precomputed MTCNN landmarks to dropout eyes, nose, mouth from images\n2. used precomputed SSIM masks  to check that  half image/grid dropout augmentation doesn't remove all fake data\n4. calculated convex hull landmarks with dlib on the fly to dropout half face\n3. used precomputed SSIM masks for UNet experiments\n\nIn my training set I extracted 30 frames per video. Used random sampling for each epoch.\n\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 820132,
      "postDate": "2020-04-25T07:02:24.500Z",
      "content": "<p>Well done</p>",
      "rawMarkdown": "Well done"
    },
    {
      "id": 820103,
      "postDate": "2020-04-25T06:23:54.483Z",
      "content": "<p>Good One</p>",
      "rawMarkdown": "Good One"
    },
    {
      "id": 820015,
      "postDate": "2020-04-25T05:03:07.583Z",
      "content": "<p>Awesome work Selim</p>",
      "rawMarkdown": "Awesome work Selim"
    },
    {
      "id": 819285,
      "postDate": "2020-04-24T13:28:06.853Z",
      "content": "<p>Congratulations! Very helpful!</p>",
      "rawMarkdown": "Congratulations! Very helpful!"
    },
    {
      "id": 819234,
      "postDate": "2020-04-24T12:42:30.087Z",
      "content": "<p>super! thanks for sharing, very to the point</p>",
      "rawMarkdown": "super! thanks for sharing, very to the point"
    },
    {
      "id": 819164,
      "postDate": "2020-04-24T11:56:58.347Z",
      "content": "<p>I tried multi-task learning including semantic segmentation,\nBut I had been put into trouble to get exact label though I tried various of adaptive-thresholding.</p>\n\n<p>SSIM-difference is so excellent methods.\nThank you for sharing!</p>",
      "rawMarkdown": "I tried multi-task learning including semantic segmentation,\nBut I had been put into trouble to get exact label though I tried various of adaptive-thresholding.\n\nSSIM-difference is so excellent methods.\nThank you for sharing!"
    },
    {
      "id": 819157,
      "postDate": "2020-04-24T11:47:55.467Z",
      "content": "<p>Interesting to see if your <code>confident_strategy</code> heuristic applies to other solutions. I only used straight averaging; if it could deliver a -0.03 to that solution it would mean that efficientnet-b0 with 180-224 input size could be reasonably competitive (0.44x). It's increasingly difficult to generalise this stuff well!</p>",
      "rawMarkdown": "Interesting to see if your ``confident_strategy`` heuristic applies to other solutions. I only used straight averaging; if it could deliver a -0.03 to that solution it would mean that efficientnet-b0 with 180-224 input size could be reasonably competitive (0.44x). It's increasingly difficult to generalise this stuff well!"
    },
    {
      "id": 819131,
      "postDate": "2020-04-24T11:28:31.587Z",
      "content": "<p>Thanks for sharing! You said that face blending did not help on private LB. What about public?</p>",
      "rawMarkdown": "Thanks for sharing! You said that face blending did not help on private LB. What about public?",
      "replies": [
        {
          "id": 819144,
          "postDate": "2020-04-24T11:39:07.043Z",
          "content": "<p>Relied on local validation and public leaderboard. Regarding private set I don't know as I have not used it. It could work though... </p>",
          "rawMarkdown": "Relied on local validation and public leaderboard. Regarding private set I don't know as I have not used it. It could work though... ",
          "votes": 2
        }
      ]
    },
    {
      "id": 819127,
      "postDate": "2020-04-24T11:22:09.607Z",
      "content": "<p>Bravo! Your place is well deserved!</p>",
      "rawMarkdown": "Bravo! Your place is well deserved!"
    },
    {
      "id": 819126,
      "postDate": "2020-04-24T11:21:43.597Z",
      "content": "<p>Congratulations Serif and Thanks for sharing!! I've a question though:</p>\n\n<blockquote>\n  <p>extracted SSIM masks with difference between real and fake and saved them as png</p>\n</blockquote>\n\n<p>What is SSIM? How does it look like in code?</p>",
      "rawMarkdown": "Congratulations Serif and Thanks for sharing!! I've a question though:\n&gt; extracted SSIM masks with difference between real and fake and saved them as png\n\n\nWhat is SSIM? How does it look like in code?",
      "replies": [
        {
          "id": 819137,
          "postDate": "2020-04-24T11:36:42.593Z",
          "content": "<p><code>\nfrom skimage.measure import compare_ssim\nscore, score_mask = compare_ssim(real_img, fake_img, multichannel=True, full=True)\ndiff = ((1 - score_mask) * 255).astype(np.uint8)\ndiff = cv2.cvtColor(diff, cv2.COLOR_BGR2GRAY)\n</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F534152%2F05c52a369beab0a56db3f044c81a8073%2F34_0_diff.png?generation=1587728193334275&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "```\nfrom skimage.measure import compare_ssim\nscore, score_mask = compare_ssim(real_img, fake_img, multichannel=True, full=True)\ndiff = ((1 - score_mask) * 255).astype(np.uint8)\ndiff = cv2.cvtColor(diff, cv2.COLOR_BGR2GRAY)\n```\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F534152%2F05c52a369beab0a56db3f044c81a8073%2F34_0_diff.png?generation=1587728193334275&amp;alt=media)\n",
          "votes": 7
        },
        {
          "id": 823228,
          "postDate": "2020-04-27T13:57:14.707Z",
          "content": "<p>Awesome!!! Bravo 👏 . Two clarifications pls\n1. Did you feed this mask as the input to the  training model?\n2. At the time of inference did you calculate a similar ssim mask of consecutive frames and use that for inference?\nThank you </p>",
          "rawMarkdown": "Awesome!!! Bravo 👏 . Two clarifications pls\n1. Did you feed this mask as the input to the  training model?\n2. At the time of inference did you calculate a similar ssim mask of consecutive frames and use that for inference?\nThank you "
        },
        {
          "id": 823524,
          "postDate": "2020-04-27T17:52:00.150Z",
          "content": "<p>it is calculated between real video and fake video. Obviously this info is not availble during inference. The only way is to predict this mask with U-Net, but I did not get any improvement from segmentation. </p>",
          "rawMarkdown": "it is calculated between real video and fake video. Obviously this info is not availble during inference. The only way is to predict this mask with U-Net, but I did not get any improvement from segmentation. "
        },
        {
          "id": 823702,
          "postDate": "2020-04-27T20:21:08.340Z",
          "content": "<p>Thank you for the reply selim. In that case, how did you use this mask?  In the grand scheme of things, where does this mask fit in?  was it here \" To be safe I checked that it will not delete highly confident difference from masks generated with SSIM\"</p>",
          "rawMarkdown": "Thank you for the reply selim. In that case, how did you use this mask?  In the grand scheme of things, where does this mask fit in?  was it here \" To be safe I checked that it will not delete highly confident difference from masks generated with SSIM\"",
          "votes": 1
        },
        {
          "id": 823706,
          "postDate": "2020-04-27T20:26:17.327Z",
          "content": "<p>When performed image dropout augmentations checked that it doesn't delete all the pixels that make fake face really fake. I.e. left some part of artefacts on the image, otherwise this augmentation will hurt model performance. </p>",
          "rawMarkdown": "When performed image dropout augmentations checked that it doesn't delete all the pixels that make fake face really fake. I.e. left some part of artefacts on the image, otherwise this augmentation will hurt model performance. \n",
          "votes": 1
        }
      ]
    },
    {
      "id": 819077,
      "postDate": "2020-04-24T10:43:34.400Z",
      "content": "<p>Thanks for sharing your solution. That's a lot of GPUs you got (compared to my single 1080Ti) :D \nWell done!</p>",
      "rawMarkdown": "Thanks for sharing your solution. That's a lot of GPUs you got (compared to my single 1080Ti) :D \nWell done!"
    },
    {
      "id": 819016,
      "postDate": "2020-04-24T09:25:48.820Z",
      "content": "<p>Amazing work <a href=\"/selimsef\">@selimsef</a>! Thanks for sharing. Do you think choosing png over jpg makes any difference?</p>",
      "rawMarkdown": "Amazing work @selimsef! Thanks for sharing. Do you think choosing png over jpg makes any difference?",
      "replies": [
        {
          "id": 819018,
          "postDate": "2020-04-24T09:28:03.820Z",
          "content": "<p>yes, because I applied different level image compression augmentation after that.  </p>",
          "rawMarkdown": "yes, because I applied different level image compression augmentation after that.  ",
          "votes": 3
        }
      ]
    },
    {
      "id": 818986,
      "postDate": "2020-04-24T09:09:06.257Z",
      "content": "<p>Thanks for the writeup, Selim. Did you do anything special to handle the class imbalance?</p>",
      "rawMarkdown": "Thanks for the writeup, Selim. Did you do anything special to handle the class imbalance?",
      "replies": [
        {
          "id": 818997,
          "postDate": "2020-04-24T09:12:38.307Z",
          "content": "<p>used undersampling as described in the post <code>fakes.sample(n=num_real, replace=False, random_state=seed)</code>. Class weights did not work for me.</p>",
          "rawMarkdown": "used undersampling as described in the post `fakes.sample(n=num_real, replace=False, random_state=seed)`. Class weights did not work for me.\n",
          "votes": 1
        },
        {
          "id": 819050,
          "postDate": "2020-04-24T10:09:22.223Z",
          "content": "<p>btw <a href=\"/humananalog\">@humananalog</a>  I modified your kernel as a starter. Thanks for nice work!</p>",
          "rawMarkdown": "btw @humananalog  I modified your kernel as a starter. Thanks for nice work!",
          "votes": 4
        }
      ]
    },
    {
      "id": 818982,
      "postDate": "2020-04-24T09:05:23.953Z",
      "content": "<p>Awesome work Selim. Thank you for the write-up. Were you able to recreate anything from the face warping artifact paper? I tried out their pretrained model and it did exceptionally poorly and I tried to simulate the deepfake process using the faceswap github repo and extracting out the process the deepfakes are made with, but was never able to get that approach to generalize. </p>\n\n<p>Wishing I had done more with the face margin and trying larger architectures. I stuck with just 10 pixels static padding on all sides and resnet18 because initial results with resnet34 and 50 did not look particularly promising. </p>\n\n<p>Very interesting masking technique. I extracted keypoints and tried various things with them, but did not think to use them like you did. </p>",
      "rawMarkdown": "Awesome work Selim. Thank you for the write-up. Were you able to recreate anything from the face warping artifact paper? I tried out their pretrained model and it did exceptionally poorly and I tried to simulate the deepfake process using the faceswap github repo and extracting out the process the deepfakes are made with, but was never able to get that approach to generalize. \n\nWishing I had done more with the face margin and trying larger architectures. I stuck with just 10 pixels static padding on all sides and resnet18 because initial results with resnet34 and 50 did not look particularly promising. \n\nVery interesting masking technique. I extracted keypoints and tried various things with them, but did not think to use them like you did. ",
      "replies": [
        {
          "id": 818992,
          "postDate": "2020-04-24T09:11:12.123Z",
          "content": "<p>No, actually I think that models learn to detect FWA by default provided that the margin is large enough and there is lot of training data. Agressive augmentations help with that as well.</p>",
          "rawMarkdown": "No, actually I think that models learn to detect FWA by default provided that the margin is large enough and there is lot of training data. Agressive augmentations help with that as well.",
          "votes": 2,
          "replies": [
            {
              "id": 2633731,
              "postDate": "2024-02-03T08:52:25.003Z",
              "content": "<p>How can I see the training code? I need a sample for my university thesis</p>",
              "rawMarkdown": "How can I see the training code? I need a sample for my university thesis"
            }
          ]
        }
      ]
    },
    {
      "id": 967559,
      "postDate": "2020-08-12T10:30:49.113Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 842013,
      "postDate": "2020-05-11T06:05:27.563Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 827527,
      "postDate": "2020-04-30T11:10:58.260Z",
      "content": "<p>Thanks for sharing. It is very detailed. You are very generous!</p>",
      "rawMarkdown": "Thanks for sharing. It is very detailed. You are very generous!",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 826720,
      "postDate": "2020-04-29T20:12:45.873Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 826427,
      "postDate": "2020-04-29T16:37:23.320Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 821765,
      "postDate": "2020-04-26T12:04:17.587Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 821236,
      "postDate": "2020-04-26T03:36:55.447Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 820745,
      "postDate": "2020-04-25T17:17:05.130Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2633744,
          "postDate": "2024-02-03T08:58:05.780Z",
          "content": "<p>Where can I see the training code? I need a sample for my university thesis</p>",
          "rawMarkdown": "Where can I see the training code? I need a sample for my university thesis"
        }
      ]
    },
    {
      "id": 2602687,
      "postDate": "2024-01-15T09:28:15.063Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!\n"
    },
    {
      "id": 848179,
      "postDate": "2020-05-14T19:34:09.647Z",
      "content": "<p>Thank you for your contribution!!</p>",
      "rawMarkdown": "Thank you for your contribution!!"
    },
    {
      "id": 828599,
      "postDate": "2020-05-01T07:15:30.603Z",
      "content": "<p>Thank you for your contribution!!</p>",
      "rawMarkdown": "Thank you for your contribution!!"
    },
    {
      "id": 825632,
      "postDate": "2020-04-29T06:21:00.477Z",
      "content": "<p>Thanks for sharing! Very helpful!</p>",
      "rawMarkdown": "Thanks for sharing! Very helpful!"
    },
    {
      "id": 824629,
      "postDate": "2020-04-28T13:54:01.823Z",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing! "
    },
    {
      "id": 820668,
      "postDate": "2020-04-25T16:07:23.023Z",
      "content": "<p>Thank you.</p>",
      "rawMarkdown": "Thank you."
    }
  ],
  "comments": [
    {
      "id": 876545,
      "author_name": "Selim Seferbekov",
      "author_url": "",
      "post_date": "2020-06-06T19:00:59.370000",
      "content": "<p>Opensourced my solution <a href=\"https://github.com/selimsef/dfdc_deepfake_challenge\">https://github.com/selimsef/dfdc_deepfake_challenge</a></p>",
      "votes": 27,
      "replies": [
        {
          "id": 876557,
          "author_name": "Yassine Alouini",
          "author_url": "",
          "post_date": "2020-06-06T19:20:31.117000",
          "content": "<p>Awesome! </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 883568,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-06-12T18:30:18.823000",
          "content": "<p>Bravo Selim! From top#3 to top#1 after final review. Weird, some teams have been re-scored.</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 883763,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-06-12T22:35:25.113000",
      "content": "<p>Just want to say huge congratulation, I really enjoyed reading your solution, and would learn a lot from further studying your technique :)</p>",
      "votes": 12,
      "replies": [
        {
          "id": 883773,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-06-12T22:55:58.267000",
          "content": "<p>Thank you! </p>\n\n<p>I was really surprised that your submission that used external data with CC-BY license led to its disqualification. \nAt the beginning I wanted to collect CC-BY videos as well as they seemed to be legit. Asked about that a lot of times. Beacuse I have not received a confirmation from organizers I considered this data in the gray area (neither organizers nor host could provide definitive answers) and decided to be on the safe side. <br>\nHopefully this situation will be a lesson for Kaggle to make rules very clear for the future competitions.</p>",
          "votes": 22,
          "replies": []
        }
      ]
    },
    {
      "id": 1229858,
      "author_name": "Ilkin Huseynli",
      "author_url": "",
      "post_date": "2021-03-07T16:29:39.273000",
      "content": "<p>Great work! Congratulations!</p>\n<ol>\n<li>MTCNN frequently has false detections. How did you get rid of them?</li>\n<li>What did you do with multiple actor videos?</li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 819009,
      "author_name": "James Howard",
      "author_url": "",
      "post_date": "2020-04-24T09:20:43.720000",
      "content": "<p>Really excellent work. I like the domain-specific augmentations and I'm sure it's things like that that helped you hold on to your excellent public LB position on the 'wild' private LB data, even if it didn't manifest in a big public LB change.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 887934,
      "author_name": "hosna",
      "author_url": "",
      "post_date": "2020-06-16T02:40:48.100000",
      "content": "<p>Thanks a lot for sharing and congratulations for winning, this is amazing really to have the first in such competition firstly saying: \"<strong>Keep it simple</strong>\".\nI do have some questions if you wouldn't mind: 1- how many frames did you use per video, only one frame? \n2- you trained the model with the extracted faces or also these SSIM masks?\n3- when you tried only B5 efficientNet model; for about how many epochs did you train the model to yield this loss?\nlastly, actually I don't understand what polyLR you mean, do you mean that you tried several learning rates? if so, then what ranges of values did you use like from 0.01 to what with how many steps?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 889149,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-06-16T19:50:35.173000",
          "content": "<ol>\n<li>32 frames</li>\n<li>only faces, used SSIM masks during my UNet experiments</li>\n<li>for B5 it was 30 epochs with 2500 steps each, batch size=80</li>\n</ol>\n\n<p>PolyLR - polynomial learning rate decay. It's usually used for semantic segmentation models like DeepLabV3. Cosine decay would be better I guess. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 889264,
          "author_name": "hosna",
          "author_url": "",
          "post_date": "2020-06-16T21:29:45.770000",
          "content": "<p>very good detailed answer, thanks for your help ✌</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 826090,
      "author_name": "mjsML",
      "author_url": "",
      "post_date": "2020-04-29T12:57:40.743000",
      "content": "<p>I'm wondering about the less glamorous side of this (the things that didn't work), can you please share what other complex things you tried that didn't work :) </p>",
      "votes": 1,
      "replies": [
        {
          "id": 826186,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-04-29T13:51:10.257000",
          "content": "<p>First of all I read overviews of existing methods <a href=\"https://arxiv.org/abs/2001.06564\">https://arxiv.org/abs/2001.06564</a> and <a href=\"https://arxiv.org/abs/2001.00179\">https://arxiv.org/abs/2001.00179</a> . Then investigated some methods if they seemed promising or came up with my own ideas.</p>\n\n<p><strong>Multitask-Learning</strong> (did not affect public leaderboard)\n- segmentation with soft probabilities obtained by SSIM masks. It worked and highlighed artefacts correctly.\n- segmentation with thresholded SSIM diffs masks\n- segmentation with face convex hull masks</p>\n\n<p><strong>Using low level CNN features</strong> (did not affect public leaderboard) - <a href=\"https://arxiv.org/abs/1909.06122\">https://arxiv.org/abs/1909.06122</a></p>\n\n<p><strong>Using contrastive loss and pairwise learning</strong></p>\n\n<p><strong>CNN + RNN on features - was overfitting badly</strong></p>\n\n<p><strong>Different pooling layers</strong>, GWAP worked well enough but again no improvements, paper <a href=\"https://arxiv.org/pdf/1809.08264.pdf\">https://arxiv.org/pdf/1809.08264.pdf</a>, <a href=\"https://github.com/BloodAxe/pytorch-toolbelt/blob/c830e02eac3162f6066ccd9fd5b02c3c2cb19018/pytorch_toolbelt/modules/pooling.py#L45\">https://github.com/BloodAxe/pytorch-toolbelt/blob/c830e02eac3162f6066ccd9fd5b02c3c2cb19018/pytorch_toolbelt/modules/pooling.py#L45</a></p>\n\n<p><strong>Using deep and shallow networks simultaneusly</strong></p>\n\n<p><strong>Multicrop network</strong> - i.e. 5 crops from image: (topleft, topright, bottomleft, bottomright, center crop), for 256x256  got five 128x128 =&gt; concatenated CNN features + FC + classifier. </p>\n\n<p><strong>Self-supervision</strong>\n- real face warping + blending and changing label to fake\n- rotate image 0, 90, 180, 120 and predicting angle as well</p>\n\n<p><strong>Data cleaning</strong>\n- used SSIM masks and removed failed fakes, or not fakes from multiface videos\n- the loss on validation dropped dramatically, but it gave worse results on public LB</p>\n\n<p>I regret that I have not tried NeXtVlad to achieve better video classification quality. </p>",
          "votes": 8,
          "replies": []
        }
      ]
    },
    {
      "id": 820046,
      "author_name": "Beast",
      "author_url": "",
      "post_date": "2020-04-25T05:45:34.273000",
      "content": "<p>nice</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 819994,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "2020-04-25T04:38:53.123000",
      "content": "<p>Well done Selim. Combining public+private you have the best model by far 👍  Nice idea to track log loss for each class separately. Did you perform augmentation on local validation set?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 820212,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-04-25T08:37:30.253000",
          "content": "<p>Actually no. If we consider that private set has 50% of synthetic data then other top-5 competitiors' models performed much better on organic videos than my models. One of the way to achieve that (no brainer) - to add organic videos the to trainining set. </p>\n\n<p>I did not perform any augmentations on my local validation set. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 888228,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-06-16T07:54:50.760000",
      "content": "<p>Congrats on the solution.  Good writeup too, it feels like if this was easy.  I'm sure it wasn't.  Great work.</p>\n\n<blockquote>\n  <p>decided to be safe and did not use any external data.</p>\n</blockquote>\n\n<p>Only paranoids survive.</p>\n\n<p>Unfortunately.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 828913,
      "author_name": "Naman Sood",
      "author_url": "",
      "post_date": "2020-05-01T11:36:20.857000",
      "content": "<p>Quite a robust solution for sure, Are there any sources you'd suggest to improve one's understanding of image recognition?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 822500,
      "author_name": "David",
      "author_url": "",
      "post_date": "2020-04-27T00:40:01.830000",
      "content": "<p>Congratulations and thank you for sharing! Your solution is very similar to mine for the most part.  One difference is that I only used 5 B5 models in my final ensemble (due to time and hardware constraints). I also included audio which I think severely hurt me on the private LB but marginally helped on CV and public LB.</p>\n\n<p>1)  What EfficientNet library did you use?\n2)  What was your experience with audio analysis and why did you decide not to use it?</p>\n\n<p>I too did not use any datasets besides the provided one or scrape any videos from YouTube as I was afraid it might not be allowed. It is unfortunate that the hosts did not clarify. It is unfair to people who only used the provided dataset.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 822901,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-04-27T08:40:39.870000",
          "content": "<p>I used timm library <a href=\"https://github.com/rwightman/pytorch-image-models\">https://github.com/rwightman/pytorch-image-models</a> and fine-tuned noisy student weights. But ordinary imagenet weights gave the same results.\nI have not even tried audio fake detection as it seemed that the main purpose of the challenge is to detect faceswap related deepfakes.  That was a right choice I guess. Though maybe other winning solutions use audio deepfakes, hard to say as they did not post descriptions yet.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 821739,
      "author_name": "Sanchit Tanwar",
      "author_url": "",
      "post_date": "2020-04-26T11:41:26.910000",
      "content": "<p>This looks great</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 819306,
      "author_name": "Jaideep",
      "author_url": "",
      "post_date": "2020-04-24T14:01:09.293000",
      "content": "<p><a href=\"/selimsef\">@selimsef</a>  many congrats and thanks</p>\n\n<p>could u help understand this portion\n\"extracted SSIM masks with difference between real and fake and saved them as png\"</p>\n\n<p>1) what ssim stands for \n2) how to perform that</p>",
      "votes": 2,
      "replies": [
        {
          "id": 819473,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-04-24T16:05:44.830000",
          "content": "<p>answered before <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145721#819137\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/145721#819137</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 822111,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-04-26T17:12:16.887000",
          "content": "<p><a href=\"/selimsef\">@selimsef</a>  thanks i got the ans below while u were explaining bibek..\n1) how this masks  were used in training \n2) How did u decide the augmentations to use</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 822262,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-04-26T19:38:06.707000",
          "content": "<ol>\n<li>used for U-Net experiments and partially for augmentations</li>\n<li>My intuition was that visual artefacts (strange eyes, mouth) are quite easy to spot. Therefore I needed to destroy them. In addition a lot of works use something like grid dropout (<a href=\"https://arxiv.org/abs/2001.04086\">https://arxiv.org/abs/2001.04086</a>) etc. for better generalization. Here it applies as well, just needs to be done carefully. </li>\n</ol>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 819302,
      "author_name": "Pedro Bernardo",
      "author_url": "",
      "post_date": "2020-04-24T13:55:01.463000",
      "content": "<p>I think what kept me from getting a gold medal in this competition was the margin around the faces. My approach was very similar to yours in the sense of keeping it simple and doing heavy augmentations, although I added histogram equalization on the pipeline (this got me a huge boost on the public leaderboard at the time). </p>\n\n<p>But I kept a very tiny margin at the end, and looking back at it, I believe this was what kept me from reaching a better position (to be honest I never thought about this during the competition). Congratz on the results!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 825391,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T01:31:50.703000",
      "content": "<p>Congratulations and thanks for your sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2597081,
      "author_name": "Oleg Shpagin",
      "author_url": "",
      "post_date": "2024-01-11T13:54:41.937000",
      "content": "<p>Hi from 2024))) Very cool solution. Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1064740,
      "author_name": "saewonYang",
      "author_url": "",
      "post_date": "2020-10-30T13:26:14.793000",
      "content": "<p>Hi Selim, congratulations for your first place!<br>\nMay I ask one simple question for your solution? Did you ensemble models which are all trained with same image size? Do you think ensembling different image size models will harm prediction score?<br>\nI'm just curious of your ensemble was based on just different seed and same image size :)<br>\nthanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1066921,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-11-02T08:03:25.200000",
          "content": "<p>Yes, this ensemble only helps to calibrate probabilities. Using different resolutions would be better. I trained B7 in the last days of the challenge and did not have time to test different resolutions, that's why I decided to just ensemble multiple checkpoints trained on differently sampled data. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1067294,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2020-11-02T13:11:19.403000",
          "content": "<p>thanks for the comment :) have a nice day!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 904221,
      "author_name": "hosna",
      "author_url": "",
      "post_date": "2020-06-27T12:53:14.757000",
      "content": "<p>In your code uploaded on github; thanks for sharing, I've to concerns:\n1- Is there a reason for choosing AvgPooling over MaxPooling, same for GlobalAvgP over GlbalMaxP? or you think that it didn't differ much.\n2-  Did you freeze the layers of  efficientNet before usage?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 904588,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-06-27T18:34:47.910000",
          "content": "<p>1 - GlobalAvgP is a default choice for classification. Actually I tried GlobalMaxP and GWAP - they were not better.\n2 - no, did not freeze anything</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 891668,
      "author_name": "Veselyev Aleksandr",
      "author_url": "",
      "post_date": "2020-06-18T11:17:22.133000",
      "content": "<p>Great solution!\nI have one question about frame preprocessing technique: you used <code>IsotropicResize</code>. I read the source code and don't quite understand, what is the difference between your <code>IsotropicResize</code> and <code>A.LongestMaxSize</code>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 892382,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-06-18T20:51:44.790000",
          "content": "<p>I did not know it existed, otherwise I would reuse it. The only difference is that I used different interpolations for scaling up/down. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 885751,
      "author_name": "Shubham Singh",
      "author_url": "",
      "post_date": "2020-06-14T13:05:58.777000",
      "content": "<p><strong>Easy does it!!</strong></p>\n\n<blockquote>\n  <p>Great work. Inspirational for me. 🙌 </p>\n</blockquote>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 883945,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-06-13T06:25:35.697000",
      "content": "<p>The result is so SWEET to you.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 883561,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-06-12T18:23:39.173000",
      "content": "<p>Congratulations on the upgrade <a href=\"/selimsef\">@selimsef</a>. Does anyone know what happened here? Is it a reproducibility issue with the other competitors?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 883533,
      "author_name": "Raghawendra Singh",
      "author_url": "",
      "post_date": "2020-06-12T18:04:04.150000",
      "content": "<p>Congrats <a href=\"/selimsef\">@selimsef</a> for 1st place!!!!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 828514,
      "author_name": "saumyadeepta sen",
      "author_url": "",
      "post_date": "2020-05-01T05:38:57.440000",
      "content": "<p>thnks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 828505,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-05-01T05:21:21.567000",
      "content": "<p>I like your term: \"<strong>hardcore augmentations</strong>\"</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 828096,
      "author_name": "VVRudenko",
      "author_url": "",
      "post_date": "2020-04-30T19:02:28.683000",
      "content": "<p>Awesome!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 827768,
      "author_name": "Rohith Pudari",
      "author_url": "",
      "post_date": "2020-04-30T14:28:23.550000",
      "content": "<p>Thank you so much, very detailed explanation</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 826795,
      "author_name": "Anastasia Rizzo",
      "author_url": "",
      "post_date": "2020-04-29T21:38:10.233000",
      "content": "<p>Good job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 825870,
      "author_name": "apache2046",
      "author_url": "",
      "post_date": "2020-04-29T09:46:03.687000",
      "content": "<p>Great job! Will you make a github repo for the training code ?  thanks for sharing !</p>",
      "votes": 0,
      "replies": [
        {
          "id": 825894,
          "author_name": "Selim Seferbekov",
          "author_url": "",
          "post_date": "2020-04-29T10:04:00.240000",
          "content": "<p>Yes, after the official results are announced. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 825551,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T05:01:25.673000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 825637,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-29T06:26:07.760000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 824076,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-28T06:23:38.240000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 823672,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-27T19:52:16.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 822516,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-27T01:00:59.590000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 822434,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T22:28:59.660000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 822350,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T20:43:17.877000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 822199,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T18:41:03.537000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 822190,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T18:31:13.473000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 822127,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T17:26:12.660000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 821988,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T15:10:17.273000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 822146,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-26T17:46:19.773000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 821783,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T12:22:45.383000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 821231,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T03:11:10.527000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 821165,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T00:48:46.717000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820149,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T07:26:20.863000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 820207,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-25T08:31:33.460000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 820132,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T07:02:24.500000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820103,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T06:23:54.483000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820015,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T05:03:07.583000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819285,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T13:28:06.853000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819234,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T12:42:30.087000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819164,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T11:56:58.347000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819157,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T11:47:55.467000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819131,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T11:28:31.587000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 819144,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-24T11:39:07.043000",
          "content": "",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 819127,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T11:22:09.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819126,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T11:21:43.597000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 819137,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-24T11:36:42.593000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 823228,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-27T13:57:14.707000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 823524,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-27T17:52:00.150000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 823702,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-27T20:21:08.340000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 823706,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-27T20:26:17.327000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 819077,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T10:43:34.400000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 819016,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T09:25:48.820000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 819018,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-24T09:28:03.820000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 818986,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T09:09:06.257000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 818997,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-24T09:12:38.307000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 819050,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-24T10:09:22.223000",
          "content": "",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 818982,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-24T09:05:23.953000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 818992,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-04-24T09:11:12.123000",
          "content": "",
          "votes": 2,
          "replies": [
            {
              "id": 2633731,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-02-03T08:52:25.003000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 967559,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-12T10:30:49.113000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 842013,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-11T06:05:27.563000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 827527,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-30T11:10:58.260000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 826720,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T20:12:45.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 826427,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T16:37:23.320000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 821765,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T12:04:17.587000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 821236,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-26T03:36:55.447000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820745,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T17:17:05.130000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2633744,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-02-03T08:58:05.780000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2602687,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-01-15T09:28:15.063000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 848179,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-14T19:34:09.647000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 828599,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-01T07:15:30.603000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 825632,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T06:21:00.477000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 824629,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-28T13:54:01.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820668,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-25T16:07:23.023000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "818975": "\n# Keep it simple\nI used a frame-by-frame classification approach as many other competitors did. \nTried a lot of other complex things but in the end it was better to just use a classifier.\n\n### Data preparation\n- extracted boxes and landmarks with MTCNN and saved them as json\n- extracted crops in original size and saved them as png\n- extracted SSIM masks with difference between real and fake and saved them as png\n\n### Face-Detector\nI used simple MTCNN detector.\nInput size for face detector was caluclated for each video depending on video resolution. \n- 2x resize for videos with less than 300 pixels wider side\n- no resize for videos with wider side between 300 and 1000\n- 0.5x resize for videos with wider side &gt; 1000 pixels\n- 0.33x resize for videos with wider side &gt; 1900 pixels\n\n### Input size\nAs soon as I discovered that EfficientNets significantly outperform other encoders I used only them in my solution. \nAs I started with B4 I decided to use \"native\" size for that network (380x380). \nDue to memory costraints I did not increase input size even for B7 encoder.\n\n### Margin\nWhen I generated crops for training I added 30% of face crop size from each side and used only this setting during the competition.\n\n\n### Encoders\nI tried multiple variants of efficient nets at the beginning:\n\n- solo B3 (300x300) - 0.29 public\n- solo B4 (380x380) - 0.27 public\n- solo B5 (380x380) - 0.25 public\n- solo B6 (380x380) - 0.27 public (surprisingly it was worse than B5 and I have not tried B7 until competition last week)\n- solo B7 (380x380) - 0.24 public \n\nIn the end I used two submits with:\n- 15xB5 (different seeds) with heursitic overfitted for Public LB which was trained with standard augmentations - 10th place on private\n- 7xB7 (different seeds) with more conservative avergaing heursitic and trained with hardcore augmentations - 3rd place on private \n\n### Averaging predictions:\nI used 32 frames for each video.\nFor each model output instead of simple averaging I used the following heuristic  which worked quite well on public leaderbord (0.25 -&gt; 0.22 solo B5).\n\n```\ndef confident_strategy(pred, t=0.87):\n    pred = np.array(pred)\n    size = len(pred)\n    fakes = np.count_nonzero(pred &gt; t)\n    if fakes &gt; size // 3 and fakes &gt; 11:\n        return np.mean(pred[pred &gt; t])\n    elif np.count_nonzero(pred &lt; 0.2) &gt; 0.6 * size:\n        return np.mean(pred[pred &lt; 0.2])\n    else:\n        return np.mean(pred)\n```\n\nI.e. I used only confident predictions for averaging if they passed some thresholds. \nThough it worked well on public leaderbord to be safe for the second submit (3rd private) I used more conservative thersholds for real videos. \n\n### Validation strategy\n\nI used 0-2 folders as holdout at the beginning. \nBut 400 videos from the public test had more correlation with Public LB and I switched to this approach.\n\nI tracked two log loss metrics (on averaged probabilities per video)\n- logloss on real videos \n- logloss on fake videos \n\nFrom the validation I can say logloss on FAKE videos was much lower than on REAL. I.e. fakes were too easy to spot.\nWhich was not encouraging. I guess public leaderboard also has this characteristic as my validation had strong correlation with public LB. \nIt is not the case on private set as it would contain new/different FaceSwap methods. \n\n### Augmentations\n\nI used image compression, noise, blur, resize with different interpolations, color jittering, scaling and rotations\n```\ndef create_train_transforms(size=380):\n    return Compose([\n        ImageCompression(quality_lower=60, quality_upper=100, p=0.5),\n        GaussNoise(p=0.1),\n        GaussianBlur(blur_limit=3, p=0.05),\n        HorizontalFlip(),\n        IsotropicResize(max_side=size)\n        PadIfNeeded(min_height=size, min_width=size, border_mode=cv2.BORDER_CONSTANT),\n        OneOf([RandomBrightnessContrast(), FancyPCA(), HueSaturationValue()], p=0.7),\n        ToGray(p=0.2),\n        ShiftScaleRotate(shift_limit=0.1, scale_limit=0.2, rotate_limit=10, border_mode=cv2.BORDER_CONSTANT, p=0.5),\n    ]\n    )\n```\n### Generalization approach\nI expected it won't be enough for generalization and spent a few weeks working on domain specific augmentations. \n\nI wanted to push models to learn the following properties:\n- visual artifacts (models learn that easily even without any augmentations)\n- different encoding of face from other part of the image. Big margin helps with that. \n- face warping artifacts. Big marging helps with that as well. Related article https://arxiv.org/abs/1811.00656 to catch FWA\n- blending artifacts - here we need either to predict blending mask (https://arxiv.org/abs/1912.13458) but it's not possible to obtain ground truth mask as there is no big difference in pixels/SSIM on the edge due to blurring and other techniques used to reduce blending artifacts or come up with augmentations that **destroy** visual artifacts.\n\nTo catch face blending artefacts:\n1. removed half face horisontally or vertically. Used dlib face convex hulls. \n2. blacked out landmarks (eyes, nose or mouth). Used MTCNN landmarks for that.\n3. blacked out half of the image. To be safe I checked that it will not delete highly confident difference from masks generated with SSIM. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F534152%2Fcd745d9fb666f2316ac699165d629c4f%2Faugmentations.jpg?generation=1587717153025245&amp;alt=media)\n\nI don't really know if it helped on private leaderboard though.\n\n### Training schedule\n- sampled fake crops based  on number of real crops i.e. `fakes.sample(n=num_real, replace=False, random_state=seed)`\n- 2500 iterations per epoch\n- SGD, momentum=0.9, weight decay=1e-4\n- PolyLR with 0.01 starting LR\n- 75k iterations\n- used Apex with mixed precision\n- trained on 4 GPUs with SyncBN and DDP. Batch size 16x4 for B5, 12x4 for B7.\n- label smoothing - for me label smoothing with 0.01 eps was optimal on public leaderboard. Also it allowed not to use clipping at all. \n\n \n\n### Hardware\nUnfortunately I did not ask aws credits from organizers, hoped that my home workstations would be enough -  I was wrong!\nI have two devboxes: one with 2xTitan Vs, the other with 4xTitan Vs\n\nHuge thanks to hostkey provider (https://www.hostkey.com/) that gave me a grant http://landing.hostkey.com/grants?_ga=2.3307699.2051560741.1587714719-1038061670.1587714719 !\nI got a devbox with 4x2080Ti for two months! Additional 4 GPUs helped me a lot to iterate faster as my models required all of them for training a single model. \n\n\n### Things that I tried but did not work well enough\n- Metric learning - was worse than a simple classificator\n- UNet with SSIM difference prediction - got the same LB score.\n- Self supervision with blending real faces. Extracted convex hull for real faces warped/resized/compressed face and then blended it back and changed label to fake. Models learned that quickly but that did not give any boost on validation/public leaderboard.\n- Self supervision with rotations like here https://stanford-cs221.github.io/autumn2019-extra/posters/110.pdf - did not improve my score\n- Temperature scaling - risky, no big difference. Decided not to use it. \n- Using steganalysis. Did not imporve validation score.\n- RNN - I guess it would work, but due to CPU bottleneck it would be hard to extract multiple chunks in kernel. My first model with RNN + Resnet34 scored around 0.4 on public and a simple classifier was better. \n\n\n### No External data\nThat was a huge concern for me because nothing is basically allowed.  I decided to be safe and did not use any external data.\n\n\n### Github\nhttps://github.com/selimsef/dfdc_deepfake_challenge\n",
    "876545": "Opensourced my solution https://github.com/selimsef/dfdc_deepfake_challenge",
    "883763": "Just want to say huge congratulation, I really enjoyed reading your solution, and would learn a lot from further studying your technique :)",
    "1229858": "Great work! Congratulations!\n\n1. MTCNN frequently has false detections. How did you get rid of them?\n2. What did you do with multiple actor videos?",
    "819009": "Really excellent work. I like the domain-specific augmentations and I'm sure it's things like that that helped you hold on to your excellent public LB position on the 'wild' private LB data, even if it didn't manifest in a big public LB change.",
    "887934": "Thanks a lot for sharing and congratulations for winning, this is amazing really to have the first in such competition firstly saying: \"**Keep it simple**\".\nI do have some questions if you wouldn't mind: 1- how many frames did you use per video, only one frame? \n2- you trained the model with the extracted faces or also these SSIM masks?\n3- when you tried only B5 efficientNet model; for about how many epochs did you train the model to yield this loss?\nlastly, actually I don't understand what polyLR you mean, do you mean that you tried several learning rates? if so, then what ranges of values did you use like from 0.01 to what with how many steps?",
    "826090": "I'm wondering about the less glamorous side of this (the things that didn't work), can you please share what other complex things you tried that didn't work :) ",
    "820046": "nice",
    "819994": "Well done Selim. Combining public+private you have the best model by far 👍  Nice idea to track log loss for each class separately. Did you perform augmentation on local validation set?",
    "888228": "Congrats on the solution.  Good writeup too, it feels like if this was easy.  I'm sure it wasn't.  Great work.\n\n&gt;  decided to be safe and did not use any external data.\n\nOnly paranoids survive.\n\nUnfortunately.\n",
    "828913": "Quite a robust solution for sure, Are there any sources you'd suggest to improve one's understanding of image recognition?",
    "822500": "Congratulations and thank you for sharing! Your solution is very similar to mine for the most part.  One difference is that I only used 5 B5 models in my final ensemble (due to time and hardware constraints). I also included audio which I think severely hurt me on the private LB but marginally helped on CV and public LB.\n\n1)\tWhat EfficientNet library did you use?\n2)\tWhat was your experience with audio analysis and why did you decide not to use it?\n\nI too did not use any datasets besides the provided one or scrape any videos from YouTube as I was afraid it might not be allowed. It is unfortunate that the hosts did not clarify. It is unfair to people who only used the provided dataset.",
    "821739": "This looks great",
    "819306": "@selimsef  many congrats and thanks\n\ncould u help understand this portion\n\"extracted SSIM masks with difference between real and fake and saved them as png\"\n\n1) what ssim stands for \n2) how to perform that",
    "819302": "I think what kept me from getting a gold medal in this competition was the margin around the faces. My approach was very similar to yours in the sense of keeping it simple and doing heavy augmentations, although I added histogram equalization on the pipeline (this got me a huge boost on the public leaderboard at the time). \n\nBut I kept a very tiny margin at the end, and looking back at it, I believe this was what kept me from reaching a better position (to be honest I never thought about this during the competition). Congratz on the results!",
    "825391": "Congratulations and thanks for your sharing!",
    "2597081": "Hi from 2024))) Very cool solution. Thanks.",
    "1064740": "Hi Selim, congratulations for your first place!\nMay I ask one simple question for your solution? Did you ensemble models which are all trained with same image size? Do you think ensembling different image size models will harm prediction score?\nI'm just curious of your ensemble was based on just different seed and same image size :)\nthanks!",
    "904221": "In your code uploaded on github; thanks for sharing, I've to concerns:\n1- Is there a reason for choosing AvgPooling over MaxPooling, same for GlobalAvgP over GlbalMaxP? or you think that it didn't differ much.\n2-  Did you freeze the layers of  efficientNet before usage?",
    "891668": "Great solution!\nI have one question about frame preprocessing technique: you used `IsotropicResize`. I read the source code and don't quite understand, what is the difference between your `IsotropicResize` and `A.LongestMaxSize`?",
    "885751": "**Easy does it!!**\n&gt; Great work. Inspirational for me. 🙌 ",
    "883945": "The result is so SWEET to you.",
    "883561": "Congratulations on the upgrade @selimsef. Does anyone know what happened here? Is it a reproducibility issue with the other competitors?",
    "883533": "Congrats @selimsef for 1st place!!!!!",
    "828514": "thnks\n",
    "828505": "I like your term: \"**hardcore augmentations**\"",
    "828096": "Awesome!",
    "827768": "Thank you so much, very detailed explanation",
    "826795": "Good job!",
    "825870": "Great job! Will you make a github repo for the training code ?  thanks for sharing !",
    "825551": "Congratulations and thanks for your sharing :)\nI have a question. Did you use early-stopping via validation set? If you don't, what is the reason for that?",
    "824076": "Hearty Congratulations Selim ! Thank you for sharing the solution. Hopefully we can learn from this.",
    "823672": "Great work!",
    "822516": "Congrats, great work!",
    "822434": "Good work.",
    "822350": "This seems good to me",
    "822199": "Good job👍 ",
    "822190": "Congratulations! Awesome work👍 ",
    "822127": "this is great",
    "821988": "@selimsef thanks for your good sharing. I saw your backbone model is efficient. and the baseline B3 model is 0.29. so good score in LB. Did you finetune imageNet base pre-trained model?  we also tried efficient, the score was not good and gave up. ",
    "821783": "I guess blazeface wud have given better results instead of using MTCNN to detect faces...but good job indeed",
    "821231": "nice",
    "821165": "nice",
    "820149": "Thanks for your insights and congratz on keeping the place after the shakedown!\n\nHave you used face landmarks and SSIM masks you've extracted?\nIf I get it right, you've used landmarks to crop face segments and SSIM masks to check that the significant difference is maintained?\n\nHow many frames per video have you extracted for yourt training dataset? (and how many have you actually used - because `75k iterations * 16*4 = 4'800'000` given you use different frames each epoch)",
    "820132": "Well done",
    "820103": "Good One",
    "820015": "Awesome work Selim",
    "819285": "Congratulations! Very helpful!",
    "819234": "super! thanks for sharing, very to the point",
    "819164": "I tried multi-task learning including semantic segmentation,\nBut I had been put into trouble to get exact label though I tried various of adaptive-thresholding.\n\nSSIM-difference is so excellent methods.\nThank you for sharing!",
    "819157": "Interesting to see if your ``confident_strategy`` heuristic applies to other solutions. I only used straight averaging; if it could deliver a -0.03 to that solution it would mean that efficientnet-b0 with 180-224 input size could be reasonably competitive (0.44x). It's increasingly difficult to generalise this stuff well!",
    "819131": "Thanks for sharing! You said that face blending did not help on private LB. What about public?",
    "819127": "Bravo! Your place is well deserved!",
    "819126": "Congratulations Serif and Thanks for sharing!! I've a question though:\n&gt; extracted SSIM masks with difference between real and fake and saved them as png\n\n\nWhat is SSIM? How does it look like in code?",
    "819077": "Thanks for sharing your solution. That's a lot of GPUs you got (compared to my single 1080Ti) :D \nWell done!",
    "819016": "Amazing work @selimsef! Thanks for sharing. Do you think choosing png over jpg makes any difference?",
    "818986": "Thanks for the writeup, Selim. Did you do anything special to handle the class imbalance?",
    "818982": "Awesome work Selim. Thank you for the write-up. Were you able to recreate anything from the face warping artifact paper? I tried out their pretrained model and it did exceptionally poorly and I tried to simulate the deepfake process using the faceswap github repo and extracting out the process the deepfakes are made with, but was never able to get that approach to generalize. \n\nWishing I had done more with the face margin and trying larger architectures. I stuck with just 10 pixels static padding on all sides and resnet18 because initial results with resnet34 and 50 did not look particularly promising. \n\nVery interesting masking technique. I extracted keypoints and tried various things with them, but did not think to use them like you did. ",
    "967559": "",
    "842013": "",
    "827527": "Thanks for sharing. It is very detailed. You are very generous!",
    "826720": "",
    "826427": "",
    "821765": "",
    "821236": "",
    "820745": "",
    "2602687": "Thanks for sharing!\n",
    "848179": "Thank you for your contribution!!",
    "828599": "Thank you for your contribution!!",
    "825632": "Thanks for sharing! Very helpful!",
    "824629": "Thanks for sharing! ",
    "820668": "Thank you."
  }
}