{
  "id": 145822,
  "title": "12th place in a nutshell",
  "url": "/competitions/deepfake-detection-challenge/writeups/atsamaz-gatsoev-12th-place-in-a-nutshell",
  "author_name": "",
  "post_date": "2020-06-17T19:38:26.470Z",
  "votes": 19,
  "comment_count": 8,
  "views": 0,
  "content": "<p><em>Disclaimer: It is my first experience in CV competition, so I apologize if my code or approach insult someone.</em> </p>\n\n<h2>Data preparation</h2>\n\n<ul>\n<li>1st dataset: extracted single random frame from each video</li>\n<li>2nd dataset: extracted 10 random frames from each video (used for stacking)</li>\n<li><strong>Blazeface</strong> for face extraction from <a href=\"/humananalog\">@humananalog</a> kernel</li>\n<li>output images resized to 256x256</li>\n<li>main augmentations are jpeg compression and downscale with a chance of 0.5</li>\n<li>secondary augmentations (not all were used at the same time): blur, gaussian noise, random brightness, horizontal flip</li>\n<li>classes were balanced by oversampling for single model training and by undersampling for stacked model training</li>\n</ul>\n\n<h2>Models</h2>\n\n<ul>\n<li><strong>EfficientNet B0-B6</strong> (imagenet)</li>\n<li><strong>EfficientNet B1-B6</strong> (noisy student)</li>\n</ul>\n\n<p>Efficientnet gave me a huge boost on LB, but more importantly was to choose augmentations and hyperparameters right for this kind of models.</p>\n\n<h2>Training and Validation</h2>\n\n<p>I took #40-49 data chunks as validation data and it correlated with LB pretty well (about 10% difference), input size is 256x256 for every model\n- single model training: input size depends on model (from 224 to 260), adam 1e-4, StepLR with step 2 and gamma 0.1\n- stacked model training: input size 256x256, adam 1e-3</p>\n\n<p>For every model it took about 3-5 epochs to train and start overfitting.</p>\n\n<h2>Stacking &amp; Ensembling</h2>\n\n<p>I trained a bunch of decent single models with public LB 0.38-0.4 on the first dataset and then combined different models into a stacking using second dataset. </p>\n\n<p>Then I took best 5 stacked models from this experiment, their validation score and made frame-wise weighted ensemble</p>\n\n<h2>What didn't work for me</h2>\n\n<ul>\n<li><strong>MesoNet, RNN</strong>  - underfits on train data</li>\n<li><strong>Strong augmentations</strong>  (e.g. cutout, huesaturation, 90 rotate)</li>\n<li><strong>Stacking of stacked models</strong> - seemed like it overfits, but not too much</li>\n<li><strong>FaceForensics++</strong> - the most I managed to get from this thing is 0.68 public LB</li>\n</ul>\n\n<h2>Links</h2>\n\n<ul>\n<li><strong>Refactored inference</strong> - <a href=\"https://www.kaggle.com/ims0rry/dfdc-inference-15-lb-private\">https://www.kaggle.com/ims0rry/dfdc-inference-15-lb-private</a></li>\n<li><strong>Original inference</strong> - <a href=\"https://www.kaggle.com/ims0rry/inference-demo\">https://www.kaggle.com/ims0rry/inference-demo</a></li>\n<li><strong>The whole project on github</strong> - <a href=\"https://github.com/1M50RRY/dfdc-kaggle-solution\">https://github.com/1M50RRY/dfdc-kaggle-solution</a></li>\n</ul>",
  "messages": [
    {
      "id": "819509",
      "postDate": "04/24/2020 16:33:28",
      "content": "<p><em>Disclaimer: It is my first experience in CV competition, so I apologize if my code or approach insult someone.</em> </p>\n\n<h2>Data preparation</h2>\n\n<ul>\n<li>1st dataset: extracted single random frame from each video</li>\n<li>2nd dataset: extracted 10 random frames from each video (used for stacking)</li>\n<li><strong>Blazeface</strong> for face extraction from <a href=\"/humananalog\">@humananalog</a> kernel</li>\n<li>output images resized to 256x256</li>\n<li>main augmentations are jpeg compression and downscale with a chance of 0.5</li>\n<li>secondary augmentations (not all were used at the same time): blur, gaussian noise, random brightness, horizontal flip</li>\n<li>classes were balanced by oversampling for single model training and by undersampling for stacked model training</li>\n</ul>\n\n<h2>Models</h2>\n\n<ul>\n<li><strong>EfficientNet B0-B6</strong> (imagenet)</li>\n<li><strong>EfficientNet B1-B6</strong> (noisy student)</li>\n</ul>\n\n<p>Efficientnet gave me a huge boost on LB, but more importantly was to choose augmentations and hyperparameters right for this kind of models.</p>\n\n<h2>Training and Validation</h2>\n\n<p>I took #40-49 data chunks as validation data and it correlated with LB pretty well (about 10% difference), input size is 256x256 for every model\n- single model training: input size depends on model (from 224 to 260), adam 1e-4, StepLR with step 2 and gamma 0.1\n- stacked model training: input size 256x256, adam 1e-3</p>\n\n<p>For every model it took about 3-5 epochs to train and start overfitting.</p>\n\n<h2>Stacking &amp; Ensembling</h2>\n\n<p>I trained a bunch of decent single models with public LB 0.38-0.4 on the first dataset and then combined different models into a stacking using second dataset. </p>\n\n<p>Then I took best 5 stacked models from this experiment, their validation score and made frame-wise weighted ensemble</p>\n\n<h2>What didn't work for me</h2>\n\n<ul>\n<li><strong>MesoNet, RNN</strong>  - underfits on train data</li>\n<li><strong>Strong augmentations</strong>  (e.g. cutout, huesaturation, 90 rotate)</li>\n<li><strong>Stacking of stacked models</strong> - seemed like it overfits, but not too much</li>\n<li><strong>FaceForensics++</strong> - the most I managed to get from this thing is 0.68 public LB</li>\n</ul>\n\n<h2>Links</h2>\n\n<ul>\n<li><strong>Refactored inference</strong> - <a href=\"https://www.kaggle.com/ims0rry/dfdc-inference-15-lb-private\">https://www.kaggle.com/ims0rry/dfdc-inference-15-lb-private</a></li>\n<li><strong>Original inference</strong> - <a href=\"https://www.kaggle.com/ims0rry/inference-demo\">https://www.kaggle.com/ims0rry/inference-demo</a></li>\n<li><strong>The whole project on github</strong> - <a href=\"https://github.com/1M50RRY/dfdc-kaggle-solution\">https://github.com/1M50RRY/dfdc-kaggle-solution</a></li>\n</ul>",
      "rawMarkdown": "*Disclaimer: It is my first experience in CV competition, so I apologize if my code or approach insult someone.* \n\n## Data preparation\n- 1st dataset: extracted single random frame from each video\n- 2nd dataset: extracted 10 random frames from each video (used for stacking)\n- **Blazeface** for face extraction from @humananalog kernel\n- output images resized to 256x256\n- main augmentations are jpeg compression and downscale with a chance of 0.5\n- secondary augmentations (not all were used at the same time): blur, gaussian noise, random brightness, horizontal flip\n- classes were balanced by oversampling for single model training and by undersampling for stacked model training\n\n## Models\n- **EfficientNet B0-B6** (imagenet)\n- **EfficientNet B1-B6** (noisy student)\n\nEfficientnet gave me a huge boost on LB, but more importantly was to choose augmentations and hyperparameters right for this kind of models.\n\n## Training and Validation\nI took #40-49 data chunks as validation data and it correlated with LB pretty well (about 10% difference), input size is 256x256 for every model\n- single model training: input size depends on model (from 224 to 260), adam 1e-4, StepLR with step 2 and gamma 0.1\n- stacked model training: input size 256x256, adam 1e-3\n\nFor every model it took about 3-5 epochs to train and start overfitting.\n\n## Stacking &amp; Ensembling\nI trained a bunch of decent single models with public LB 0.38-0.4 on the first dataset and then combined different models into a stacking using second dataset. \n\nThen I took best 5 stacked models from this experiment, their validation score and made frame-wise weighted ensemble\n\n## What didn't work for me\n- **MesoNet, RNN**  - underfits on train data\n- **Strong augmentations**  (e.g. cutout, huesaturation, 90 rotate)\n- **Stacking of stacked models** - seemed like it overfits, but not too much\n- **FaceForensics++** - the most I managed to get from this thing is 0.68 public LB\n\n## Links\n- **Refactored inference** - https://www.kaggle.com/ims0rry/dfdc-inference-15-lb-private\n- **Original inference** - https://www.kaggle.com/ims0rry/inference-demo\n- **The whole project on github** - https://github.com/1M50RRY/dfdc-kaggle-solution",
      "votes": null
    },
    {
      "id": "819543",
      "postDate": "04/24/2020 16:53:22",
      "content": "<p>Thanks for sharing :)</p>",
      "rawMarkdown": "Thanks for sharing :)",
      "votes": null
    },
    {
      "id": "819666",
      "postDate": "04/24/2020 18:57:14",
      "content": "<p>Cool! How were you able to use noisy-student weights? Were they already implemented in the libraries you used or did you use the Tensorflow weights from the original Github repository?</p>",
      "rawMarkdown": "Cool! How were you able to use noisy-student weights? Were they already implemented in the libraries you used or did you use the Tensorflow weights from the original Github repository?",
      "votes": null
    },
    {
      "id": "820147",
      "postDate": "04/25/2020 07:23:37",
      "content": "<p>NS weigths are ported for pytorch in <strong>timm</strong> library which I used for some of the models</p>",
      "rawMarkdown": "NS weigths are ported for pytorch in **timm** library which I used for some of the models",
      "votes": null
    },
    {
      "id": "821649",
      "postDate": "04/26/2020 10:24:13",
      "content": "<p>UPD: uploaded all pretrained models to <a href=\"https://github.com/1M50RRY/dfdc-kaggle-solution/tree/master/inference/pretrained\">github</a></p>",
      "rawMarkdown": "UPD: uploaded all pretrained models to [github](https://github.com/1M50RRY/dfdc-kaggle-solution/tree/master/inference/pretrained)",
      "votes": null
    },
    {
      "id": "824610",
      "postDate": "04/28/2020 13:43:33",
      "content": "<p>The first medal and immediately gold, congratulations!</p>",
      "rawMarkdown": "The first medal and immediately gold, congratulations!",
      "votes": null
    },
    {
      "id": "824677",
      "postDate": "04/28/2020 14:27:25",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "867234",
      "postDate": "05/30/2020 05:36:39",
      "content": "<p>Hey @1M50RRY, thanks for sharing your amazing work with us.\nHowever, I have some serialization error while loading the models using torch.load, then I noticed that the size of the models is in bytes. Not sure if there was a upload issue or if I am missing something here.\nThanks.</p>",
      "rawMarkdown": "Hey @1M50RRY, thanks for sharing your amazing work with us.\nHowever, I have some serialization error while loading the models using torch.load, then I noticed that the size of the models is in bytes. Not sure if there was a upload issue or if I am missing something here.\nThanks.",
      "votes": null
    },
    {
      "id": "867427",
      "postDate": "05/30/2020 09:23:23",
      "content": "<p>Models which size is in bytes are meta-models, they take a bunch of pretrained effnets outputs as input and generalize them. Look at the second cell of the train_stacked_model.ipynb, there is its structure.</p>",
      "rawMarkdown": "Models which size is in bytes are meta-models, they take a bunch of pretrained effnets outputs as input and generalize them. Look at the second cell of the train_stacked_model.ipynb, there is its structure.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 819543,
      "author_name": "albeffe",
      "author_url": "",
      "post_date": "04/24/2020 16:53:22",
      "content": "<p>Thanks for sharing :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 819666,
      "author_name": "carlolepelaars",
      "author_url": "",
      "post_date": "04/24/2020 18:57:14",
      "content": "<p>Cool! How were you able to use noisy-student weights? Were they already implemented in the libraries you used or did you use the Tensorflow weights from the original Github repository?</p>",
      "votes": null,
      "replies": [
        {
          "id": 820147,
          "author_name": "ims0rry",
          "author_url": "",
          "post_date": "04/25/2020 07:23:37",
          "content": "<p>NS weigths are ported for pytorch in <strong>timm</strong> library which I used for some of the models</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 821649,
      "author_name": "ims0rry",
      "author_url": "",
      "post_date": "04/26/2020 10:24:13",
      "content": "<p>UPD: uploaded all pretrained models to <a href=\"https://github.com/1M50RRY/dfdc-kaggle-solution/tree/master/inference/pretrained\">github</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 867234,
          "author_name": "abhikumar07",
          "author_url": "",
          "post_date": "05/30/2020 05:36:39",
          "content": "<p>Hey @1M50RRY, thanks for sharing your amazing work with us.\nHowever, I have some serialization error while loading the models using torch.load, then I noticed that the size of the models is in bytes. Not sure if there was a upload issue or if I am missing something here.\nThanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 867427,
          "author_name": "ims0rry",
          "author_url": "",
          "post_date": "05/30/2020 09:23:23",
          "content": "<p>Models which size is in bytes are meta-models, they take a bunch of pretrained effnets outputs as input and generalize them. Look at the second cell of the train_stacked_model.ipynb, there is its structure.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 824610,
      "author_name": "azamatk",
      "author_url": "",
      "post_date": "04/28/2020 13:43:33",
      "content": "<p>The first medal and immediately gold, congratulations!</p>",
      "votes": null,
      "replies": [
        {
          "id": 824677,
          "author_name": "ims0rry",
          "author_url": "",
          "post_date": "04/28/2020 14:27:25",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "819509": "*Disclaimer: It is my first experience in CV competition, so I apologize if my code or approach insult someone.* \n\n## Data preparation\n- 1st dataset: extracted single random frame from each video\n- 2nd dataset: extracted 10 random frames from each video (used for stacking)\n- **Blazeface** for face extraction from @humananalog kernel\n- output images resized to 256x256\n- main augmentations are jpeg compression and downscale with a chance of 0.5\n- secondary augmentations (not all were used at the same time): blur, gaussian noise, random brightness, horizontal flip\n- classes were balanced by oversampling for single model training and by undersampling for stacked model training\n\n## Models\n- **EfficientNet B0-B6** (imagenet)\n- **EfficientNet B1-B6** (noisy student)\n\nEfficientnet gave me a huge boost on LB, but more importantly was to choose augmentations and hyperparameters right for this kind of models.\n\n## Training and Validation\nI took #40-49 data chunks as validation data and it correlated with LB pretty well (about 10% difference), input size is 256x256 for every model\n- single model training: input size depends on model (from 224 to 260), adam 1e-4, StepLR with step 2 and gamma 0.1\n- stacked model training: input size 256x256, adam 1e-3\n\nFor every model it took about 3-5 epochs to train and start overfitting.\n\n## Stacking &amp; Ensembling\nI trained a bunch of decent single models with public LB 0.38-0.4 on the first dataset and then combined different models into a stacking using second dataset. \n\nThen I took best 5 stacked models from this experiment, their validation score and made frame-wise weighted ensemble\n\n## What didn't work for me\n- **MesoNet, RNN**  - underfits on train data\n- **Strong augmentations**  (e.g. cutout, huesaturation, 90 rotate)\n- **Stacking of stacked models** - seemed like it overfits, but not too much\n- **FaceForensics++** - the most I managed to get from this thing is 0.68 public LB\n\n## Links\n- **Refactored inference** - https://www.kaggle.com/ims0rry/dfdc-inference-15-lb-private\n- **Original inference** - https://www.kaggle.com/ims0rry/inference-demo\n- **The whole project on github** - https://github.com/1M50RRY/dfdc-kaggle-solution",
    "819543": "Thanks for sharing :)",
    "819666": "Cool! How were you able to use noisy-student weights? Were they already implemented in the libraries you used or did you use the Tensorflow weights from the original Github repository?",
    "820147": "NS weigths are ported for pytorch in **timm** library which I used for some of the models",
    "821649": "UPD: uploaded all pretrained models to [github](https://github.com/1M50RRY/dfdc-kaggle-solution/tree/master/inference/pretrained)",
    "824610": "The first medal and immediately gold, congratulations!",
    "824677": "Thanks!",
    "867234": "Hey @1M50RRY, thanks for sharing your amazing work with us.\nHowever, I have some serialization error while loading the models using torch.load, then I noticed that the size of the models is in bytes. Not sure if there was a upload issue or if I am missing something here.\nThanks.",
    "867427": "Models which size is in bytes are meta-models, they take a bunch of pretrained effnets outputs as input and generalize them. Look at the second cell of the train_stacked_model.ipynb, there is its structure."
  },
  "source": "meta"
}