{
  "id": 129663,
  "title": "Resnet + LSTM",
  "url": "/competitions/deepfake-detection-challenge/discussion/129663",
  "author_name": "",
  "post_date": "2020-02-09T21:06:34.714520400Z",
  "votes": 17,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Did someone try it? Resnet + LSTM on videos?\nTimesteps would be frames with unique face detected in.\nThis paper gives some references: <a href=\"https://arxiv.org/pdf/2001.03024.pdf\">https://arxiv.org/pdf/2001.03024.pdf</a> but no implementation.</p>\n\n<p>Results should be a bit better than Xception baseline (<a href=\"https://arxiv.org/pdf/1901.08971.pdf\">Faces Forensics++</a> paper).</p>",
  "messages": [
    {
      "id": "740824",
      "postDate": "02/09/2020 21:06:34",
      "content": "<p>Did someone try it? Resnet + LSTM on videos?\nTimesteps would be frames with unique face detected in.\nThis paper gives some references: <a href=\"https://arxiv.org/pdf/2001.03024.pdf\">https://arxiv.org/pdf/2001.03024.pdf</a> but no implementation.</p>\n\n<p>Results should be a bit better than Xception baseline (<a href=\"https://arxiv.org/pdf/1901.08971.pdf\">Faces Forensics++</a> paper).</p>",
      "rawMarkdown": "Did someone try it? Resnet + LSTM on videos?\nTimesteps would be frames with unique face detected in.\nThis paper gives some references: https://arxiv.org/pdf/2001.03024.pdf but no implementation.\n\nResults should be a bit better than Xception baseline ([Faces Forensics++](https://arxiv.org/pdf/1901.08971.pdf) paper).",
      "votes": null
    },
    {
      "id": "740836",
      "postDate": "02/09/2020 21:36:54",
      "content": "<p>Thanks for raising this.</p>\n\n<p>This question is for experts in sequence modeling: for LSTM to work, should we take consecutive frames (to encode temporal changes) rather than a few faces uniformly sampled from a video? I already have a uniformly sampled dataset, but I want to make sure if it would be any help if I go for an LSTM model.</p>",
      "rawMarkdown": "Thanks for raising this.\n\nThis question is for experts in sequence modeling: for LSTM to work, should we take consecutive frames (to encode temporal changes) rather than a few faces uniformly sampled from a video? I already have a uniformly sampled dataset, but I want to make sure if it would be any help if I go for an LSTM model.",
      "votes": null
    },
    {
      "id": "740882",
      "postDate": "02/09/2020 22:56:30",
      "content": "<p>I've only tried to implement Xception+LSTM from this paper <a href=\"https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf\">https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf</a>, but didn't succeed much, FaceForensics++ has better accuracy.</p>",
      "rawMarkdown": "I've only tried to implement Xception+LSTM from this paper https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf, but didn't succeed much, FaceForensics++ has better accuracy.",
      "votes": null
    },
    {
      "id": "740891",
      "postDate": "02/09/2020 23:25:52",
      "content": "<p>Hi <a href=\"/ims0rry\">@ims0rry</a>, Thanks for your reply! Could you let us know what was the issue? Was is overfitting or something else?</p>",
      "rawMarkdown": "Hi @ims0rry, Thanks for your reply! Could you let us know what was the issue? Was is overfitting or something else?",
      "votes": null
    },
    {
      "id": "741116",
      "postDate": "02/10/2020 07:47:26",
      "content": "<p><a href=\"/ims0rry\">@ims0rry</a> Also, did you try a one stage or two stages training? i.e pretrain resnet backbone only first, freeze resnet weights and then plug and train LSTM only?</p>",
      "rawMarkdown": "ims0rry Also, did you try a one stage or two stages training? i.e pretrain resnet backbone only first, freeze resnet weights and then plug and train LSTM only?",
      "votes": null
    },
    {
      "id": "741147",
      "postDate": "02/10/2020 08:56:17",
      "content": "<p>I'm not an expert but according to me, LSTM is a better choice only if we use consecutive frames or at max 3-4 frame jumps between frames. Simple Linear layer should perform better on faces uniformly sampled from video.</p>",
      "rawMarkdown": "I'm not an expert but according to me, LSTM is a better choice only if we use consecutive frames or at max 3-4 frame jumps between frames. Simple Linear layer should perform better on faces uniformly sampled from video.",
      "votes": null
    },
    {
      "id": "741160",
      "postDate": "02/10/2020 09:16:18",
      "content": "<p><a href=\"/debanga\">@debanga</a> It was underfitting, back then I've been training on 30% of original dataset just to see which models are doing well. </p>\n\n<p><a href=\"/mpware\">@mpware</a> No I didn't try it yet. It's a good idea though.</p>",
      "rawMarkdown": "debanga It was underfitting, back then I've been training on 30% of original dataset just to see which models are doing well. \n\n@mpware No I didn't try it yet. It's a good idea though.",
      "votes": null
    },
    {
      "id": "741162",
      "postDate": "02/10/2020 09:19:15",
      "content": "<p>I tried similar approach with no success so far, I tested two different models:\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then for each frame extact features vector using facenet, then feed these vectors to a Bidirectional LSTM ==&gt; score of 0.69\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then feed them to a multilayer ConvLSTM2D model ==&gt; score of 0.69</p>\n\n<p>I use DSFD (Dual Shot Face Detector) for face detection, i found it to be the only model that can detect faces even in the dark videos or videos where the characters are not facing the camera or are too far away. on kaggle notebooks with GPU, detecting and extracting 10 faces takes approximatly 2.5s, so I think we can go up to 20 faces without exceeding the 9h execution time.</p>\n\n<p>However I only used 50% of the training data, that I pre processed to get the same number of real and fake videos.\nBoth my models are overfitting after 15 epoch of training.</p>\n\n<p>I still have not figured out how to go below the 0.69 score</p>",
      "rawMarkdown": "I tried similar approach with no success so far, I tested two different models:\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then for each frame extact features vector using facenet, then feed these vectors to a Bidirectional LSTM ==&gt; score of 0.69\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then feed them to a multilayer ConvLSTM2D model ==&gt; score of 0.69\n\nI use DSFD (Dual Shot Face Detector) for face detection, i found it to be the only model that can detect faces even in the dark videos or videos where the characters are not facing the camera or are too far away. on kaggle notebooks with GPU, detecting and extracting 10 faces takes approximatly 2.5s, so I think we can go up to 20 faces without exceeding the 9h execution time.\n\nHowever I only used 50% of the training data, that I pre processed to get the same number of real and fake videos.\nBoth my models are overfitting after 15 epoch of training.\n\nI still have not figured out how to go below the 0.69 score",
      "votes": null
    },
    {
      "id": "741210",
      "postDate": "02/10/2020 10:19:23",
      "content": "<p>If you consistently get a score of 0.69, I would suspect something is wrong in your code. For a model that doesn't work well, you'd expect to get a higher loss than 0.69.</p>",
      "rawMarkdown": "If you consistently get a score of 0.69, I would suspect something is wrong in your code. For a model that doesn't work well, you'd expect to get a higher loss than 0.69.",
      "votes": null
    },
    {
      "id": "741213",
      "postDate": "02/10/2020 10:22:26",
      "content": "<p>96x96 is too small I guess, try to increase image size</p>",
      "rawMarkdown": "96x96 is too small I guess, try to increase image size",
      "votes": null
    },
    {
      "id": "741289",
      "postDate": "02/10/2020 12:51:52",
      "content": "<p>Does  anyone  who  want  to   try   <code>Transformer</code> ? \nI  think  <code>Transformer</code> is  better  than   <code>LSTM</code>,  I'm  gonna  to   have  a  try  ,  if   I  made  any  progress  ,I  will  publish  here.</p>",
      "rawMarkdown": "Does  anyone  who  want  to   try   `Transformer` ? \nI  think  `Transformer` is  better  than   `LSTM`,  I'm  gonna  to   have  a  try  ,  if   I  made  any  progress  ,I  will  publish  here.",
      "votes": null
    },
    {
      "id": "742039",
      "postDate": "02/11/2020 03:37:48",
      "content": "<p>I did something similar that got my current score. I changed Resnet to another model(I can't share because my team mate would be mad at me).</p>",
      "rawMarkdown": "I did something similar that got my current score. I changed Resnet to another model(I can't share because my team mate would be mad at me).",
      "votes": null
    },
    {
      "id": "742259",
      "postDate": "02/11/2020 07:32:57",
      "content": "<p>So LSTM works for you? I've just made some tests with 2 CNN + LSTM models and as for <a href=\"/ims0rry\">@ims0rry</a> <strong>it underfits</strong> (valid logloss lower than train logloss). I don't see why (yet). Same model (only CNN backbone) with same data and validation strategy reaches LB 0.42 but with LSTM it's 0.47. There could be several reasons for underfitting but none for them seems to be the root cause:\n- Regularization: High dropout, aggressive augmentation on train dataset\n- Hard samples in train, easy samples in valid\n- Valid dataset too small\n- Leak in valid dataset\n- LR too small or number epochs too small.</p>\n\n<p><a href=\"/unkownhihi\">@unkownhihi</a>  How many frames/timesteps did you use for the sequence?</p>\n\n<p>The 2 models I've tried:\n- CNN backbone + AvgPool (with 2048 features) + LSTM (2 layers, 512 hidden, BiDirectionnal) + Linear\n- CNN backbone + AvgPool + Linear/BN + LSTM (1 layer, 512 hidden, BiDirectionnal) + Attention + Linear</p>",
      "rawMarkdown": "So LSTM works for you? I've just made some tests with 2 CNN + LSTM models and as for @ims0rry **it underfits** (valid logloss lower than train logloss). I don't see why (yet). Same model (only CNN backbone) with same data and validation strategy reaches LB 0.42 but with LSTM it's 0.47. There could be several reasons for underfitting but none for them seems to be the root cause:\n- Regularization: High dropout, aggressive augmentation on train dataset\n- Hard samples in train, easy samples in valid\n- Valid dataset too small\n- Leak in valid dataset\n- LR too small or number epochs too small.\n\n@unkownhihi  How many frames/timesteps did you use for the sequence?\n\nThe 2 models I've tried:\n- CNN backbone + AvgPool (with 2048 features) + LSTM (2 layers, 512 hidden, BiDirectionnal) + Linear\n- CNN backbone + AvgPool + Linear/BN + LSTM (1 layer, 512 hidden, BiDirectionnal) + Attention + Linear",
      "votes": null
    },
    {
      "id": "742292",
      "postDate": "02/11/2020 08:08:39",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> Yep i'm suspecting something is wrong with my training data generator, i'm currently reworking that.\n<a href=\"/ims0rry\">@ims0rry</a> maybe you're right, I will try with bigger frames, and maybe add face alignment</p>\n\n<p>Also On some of the papers they mention the use of softmax as activation of the last layer, when I tried that my loss score would stay at the same value during the whole training, the validation loss fluctuate a lot however, I can't understant why...</p>",
      "rawMarkdown": "humananalog Yep i'm suspecting something is wrong with my training data generator, i'm currently reworking that.\n@ims0rry maybe you're right, I will try with bigger frames, and maybe add face alignment\n\nAlso On some of the papers they mention the use of softmax as activation of the last layer, when I tried that my loss score would stay at the same value during the whole training, the validation loss fluctuate a lot however, I can't understant why...",
      "votes": null
    },
    {
      "id": "742527",
      "postDate": "02/11/2020 10:57:55",
      "content": "<p>I tried doing CNN features (post averagepool) -&gt;BiLSTM to my current model and it seemed to perform no better than doing a CNN of multiple frames per video and taking an average. The performance is disappointing, given the validation loss is better for the CNN-&gt;LSTM; it seems the LSTM overfits when judged on the testing sets. It might be that using 1D convolutions of the averagepooling layer is better, though I haven't experimented.</p>",
      "rawMarkdown": "I tried doing CNN features (post averagepool) -&gt;BiLSTM to my current model and it seemed to perform no better than doing a CNN of multiple frames per video and taking an average. The performance is disappointing, given the validation loss is better for the CNN-&gt;LSTM; it seems the LSTM overfits when judged on the testing sets. It might be that using 1D convolutions of the averagepooling layer is better, though I haven't experimented.",
      "votes": null
    },
    {
      "id": "743235",
      "postDate": "02/11/2020 22:58:13",
      "content": "<p>I didn't do any dropout, no augment, not bidirection, LR of 1e-4. I literally wanted a baseline but turned out to be our best single model score. I didn't add pooling at the end of backbone. Instead, I did backbone(input_size=(...),pooling='avg'). And a hint, the backbone's output vectors are much less.</p>",
      "rawMarkdown": "I didn't do any dropout, no augment, not bidirection, LR of 1e-4. I literally wanted a baseline but turned out to be our best single model score. I didn't add pooling at the end of backbone. Instead, I did backbone(input_size=(...),pooling='avg'). And a hint, the backbone's output vectors are much less.",
      "votes": null
    },
    {
      "id": "743701",
      "postDate": "02/12/2020 07:58:51",
      "content": "<p>Thanks for such info. Simpler model seems to be perform better for you. I will try another attempt in this way. I will also try to add more consecutive frames.</p>",
      "rawMarkdown": "Thanks for such info. Simpler model seems to be perform better for you. I will try another attempt in this way. I will also try to add more consecutive frames.",
      "votes": null
    },
    {
      "id": "743771",
      "postDate": "02/12/2020 09:42:25",
      "content": "<p>I've tried self-attention at the penultimate layer of a backbone, instead of LSTM. It gave a better result than LSTM in CV, but LB score was worse by a large margin.\nActually, I observed similar phenomena like <a href=\"/jamesphoward\">@jamesphoward</a> said, the gap of CV v.s. LB is bigger than frame-by-frame models. So we may need to find a way to make CV correlates LB even in temporal modeling... (I've used split by chunk)</p>",
      "rawMarkdown": "I've tried self-attention at the penultimate layer of a backbone, instead of LSTM. It gave a better result than LSTM in CV, but LB score was worse by a large margin.\nActually, I observed similar phenomena like @jamesphoward said, the gap of CV v.s. LB is bigger than frame-by-frame models. So we may need to find a way to make CV correlates LB even in temporal modeling... (I've used split by chunk)",
      "votes": null
    },
    {
      "id": "745992",
      "postDate": "02/14/2020 13:03:40",
      "content": "<p>I tried it, but without following the paper you mentioned. Just feeding a number of consecutive face-cropped frames to the pretrained resnet (with final layer removed), then into a LSTM and finally into two dense layers.</p>\n\n<p>Result are ok so far, but nothing too shocking (around 0.5 LB score). Didn't experiment yet with many of the hyper parameters, so there is still some room to improve. </p>\n\n<p>The biggest challenge I'm having is how long to freeze the pre-trained resnet layers before making them trainable and what learning rate then to use.  </p>",
      "rawMarkdown": "I tried it, but without following the paper you mentioned. Just feeding a number of consecutive face-cropped frames to the pretrained resnet (with final layer removed), then into a LSTM and finally into two dense layers.\n\nResult are ok so far, but nothing too shocking (around 0.5 LB score). Didn't experiment yet with many of the hyper parameters, so there is still some room to improve. \n\nThe biggest challenge I'm having is how long to freeze the pre-trained resnet layers before making them trainable and what learning rate then to use.",
      "votes": null
    },
    {
      "id": "747997",
      "postDate": "02/17/2020 05:12:46",
      "content": "<p>EDIT: I made a model stacking of 2 identical lrcn models trained with different data balancing which got my newer score(0.34012)</p>",
      "rawMarkdown": "EDIT: I made a model stacking of 2 identical lrcn models trained with different data balancing which got my newer score(0.34012)",
      "votes": null
    },
    {
      "id": "748167",
      "postDate": "02/17/2020 08:00:54",
      "content": "<p>Good to see that there are some good scores possible with this approach. Means there is still some upside for my models :) </p>",
      "rawMarkdown": "Good to see that there are some good scores possible with this approach. Means there is still some upside for my models :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 740836,
      "author_name": "debanga",
      "author_url": "",
      "post_date": "02/09/2020 21:36:54",
      "content": "<p>Thanks for raising this.</p>\n\n<p>This question is for experts in sequence modeling: for LSTM to work, should we take consecutive frames (to encode temporal changes) rather than a few faces uniformly sampled from a video? I already have a uniformly sampled dataset, but I want to make sure if it would be any help if I go for an LSTM model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 741147,
          "author_name": "ankitsainiankit",
          "author_url": "",
          "post_date": "02/10/2020 08:56:17",
          "content": "<p>I'm not an expert but according to me, LSTM is a better choice only if we use consecutive frames or at max 3-4 frame jumps between frames. Simple Linear layer should perform better on faces uniformly sampled from video.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 740882,
      "author_name": "ims0rry",
      "author_url": "",
      "post_date": "02/09/2020 22:56:30",
      "content": "<p>I've only tried to implement Xception+LSTM from this paper <a href=\"https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf\">https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf</a>, but didn't succeed much, FaceForensics++ has better accuracy.</p>",
      "votes": null,
      "replies": [
        {
          "id": 740891,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/09/2020 23:25:52",
          "content": "<p>Hi <a href=\"/ims0rry\">@ims0rry</a>, Thanks for your reply! Could you let us know what was the issue? Was is overfitting or something else?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741116,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "02/10/2020 07:47:26",
          "content": "<p><a href=\"/ims0rry\">@ims0rry</a> Also, did you try a one stage or two stages training? i.e pretrain resnet backbone only first, freeze resnet weights and then plug and train LSTM only?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741160,
          "author_name": "ims0rry",
          "author_url": "",
          "post_date": "02/10/2020 09:16:18",
          "content": "<p><a href=\"/debanga\">@debanga</a> It was underfitting, back then I've been training on 30% of original dataset just to see which models are doing well. </p>\n\n<p><a href=\"/mpware\">@mpware</a> No I didn't try it yet. It's a good idea though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 741162,
      "author_name": "soviet",
      "author_url": "",
      "post_date": "02/10/2020 09:19:15",
      "content": "<p>I tried similar approach with no success so far, I tested two different models:\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then for each frame extact features vector using facenet, then feed these vectors to a Bidirectional LSTM ==&gt; score of 0.69\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then feed them to a multilayer ConvLSTM2D model ==&gt; score of 0.69</p>\n\n<p>I use DSFD (Dual Shot Face Detector) for face detection, i found it to be the only model that can detect faces even in the dark videos or videos where the characters are not facing the camera or are too far away. on kaggle notebooks with GPU, detecting and extracting 10 faces takes approximatly 2.5s, so I think we can go up to 20 faces without exceeding the 9h execution time.</p>\n\n<p>However I only used 50% of the training data, that I pre processed to get the same number of real and fake videos.\nBoth my models are overfitting after 15 epoch of training.</p>\n\n<p>I still have not figured out how to go below the 0.69 score</p>",
      "votes": null,
      "replies": [
        {
          "id": 741210,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/10/2020 10:19:23",
          "content": "<p>If you consistently get a score of 0.69, I would suspect something is wrong in your code. For a model that doesn't work well, you'd expect to get a higher loss than 0.69.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741213,
          "author_name": "ims0rry",
          "author_url": "",
          "post_date": "02/10/2020 10:22:26",
          "content": "<p>96x96 is too small I guess, try to increase image size</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 742292,
          "author_name": "soviet",
          "author_url": "",
          "post_date": "02/11/2020 08:08:39",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> Yep i'm suspecting something is wrong with my training data generator, i'm currently reworking that.\n<a href=\"/ims0rry\">@ims0rry</a> maybe you're right, I will try with bigger frames, and maybe add face alignment</p>\n\n<p>Also On some of the papers they mention the use of softmax as activation of the last layer, when I tried that my loss score would stay at the same value during the whole training, the validation loss fluctuate a lot however, I can't understant why...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 741289,
      "author_name": "xujingzhao",
      "author_url": "",
      "post_date": "02/10/2020 12:51:52",
      "content": "<p>Does  anyone  who  want  to   try   <code>Transformer</code> ? \nI  think  <code>Transformer</code> is  better  than   <code>LSTM</code>,  I'm  gonna  to   have  a  try  ,  if   I  made  any  progress  ,I  will  publish  here.</p>",
      "votes": null,
      "replies": [
        {
          "id": 743771,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "02/12/2020 09:42:25",
          "content": "<p>I've tried self-attention at the penultimate layer of a backbone, instead of LSTM. It gave a better result than LSTM in CV, but LB score was worse by a large margin.\nActually, I observed similar phenomena like <a href=\"/jamesphoward\">@jamesphoward</a> said, the gap of CV v.s. LB is bigger than frame-by-frame models. So we may need to find a way to make CV correlates LB even in temporal modeling... (I've used split by chunk)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 742039,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/11/2020 03:37:48",
      "content": "<p>I did something similar that got my current score. I changed Resnet to another model(I can't share because my team mate would be mad at me).</p>",
      "votes": null,
      "replies": [
        {
          "id": 742259,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "02/11/2020 07:32:57",
          "content": "<p>So LSTM works for you? I've just made some tests with 2 CNN + LSTM models and as for <a href=\"/ims0rry\">@ims0rry</a> <strong>it underfits</strong> (valid logloss lower than train logloss). I don't see why (yet). Same model (only CNN backbone) with same data and validation strategy reaches LB 0.42 but with LSTM it's 0.47. There could be several reasons for underfitting but none for them seems to be the root cause:\n- Regularization: High dropout, aggressive augmentation on train dataset\n- Hard samples in train, easy samples in valid\n- Valid dataset too small\n- Leak in valid dataset\n- LR too small or number epochs too small.</p>\n\n<p><a href=\"/unkownhihi\">@unkownhihi</a>  How many frames/timesteps did you use for the sequence?</p>\n\n<p>The 2 models I've tried:\n- CNN backbone + AvgPool (with 2048 features) + LSTM (2 layers, 512 hidden, BiDirectionnal) + Linear\n- CNN backbone + AvgPool + Linear/BN + LSTM (1 layer, 512 hidden, BiDirectionnal) + Attention + Linear</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 743235,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/11/2020 22:58:13",
          "content": "<p>I didn't do any dropout, no augment, not bidirection, LR of 1e-4. I literally wanted a baseline but turned out to be our best single model score. I didn't add pooling at the end of backbone. Instead, I did backbone(input_size=(...),pooling='avg'). And a hint, the backbone's output vectors are much less.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 743701,
          "author_name": "mpware",
          "author_url": "",
          "post_date": "02/12/2020 07:58:51",
          "content": "<p>Thanks for such info. Simpler model seems to be perform better for you. I will try another attempt in this way. I will also try to add more consecutive frames.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747997,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/17/2020 05:12:46",
          "content": "<p>EDIT: I made a model stacking of 2 identical lrcn models trained with different data balancing which got my newer score(0.34012)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748167,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "02/17/2020 08:00:54",
          "content": "<p>Good to see that there are some good scores possible with this approach. Means there is still some upside for my models :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 742527,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "02/11/2020 10:57:55",
      "content": "<p>I tried doing CNN features (post averagepool) -&gt;BiLSTM to my current model and it seemed to perform no better than doing a CNN of multiple frames per video and taking an average. The performance is disappointing, given the validation loss is better for the CNN-&gt;LSTM; it seems the LSTM overfits when judged on the testing sets. It might be that using 1D convolutions of the averagepooling layer is better, though I haven't experimented.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 745992,
      "author_name": "peterdekkers101",
      "author_url": "",
      "post_date": "02/14/2020 13:03:40",
      "content": "<p>I tried it, but without following the paper you mentioned. Just feeding a number of consecutive face-cropped frames to the pretrained resnet (with final layer removed), then into a LSTM and finally into two dense layers.</p>\n\n<p>Result are ok so far, but nothing too shocking (around 0.5 LB score). Didn't experiment yet with many of the hyper parameters, so there is still some room to improve. </p>\n\n<p>The biggest challenge I'm having is how long to freeze the pre-trained resnet layers before making them trainable and what learning rate then to use.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "740824": "Did someone try it? Resnet + LSTM on videos?\nTimesteps would be frames with unique face detected in.\nThis paper gives some references: https://arxiv.org/pdf/2001.03024.pdf but no implementation.\n\nResults should be a bit better than Xception baseline ([Faces Forensics++](https://arxiv.org/pdf/1901.08971.pdf) paper).",
    "740836": "Thanks for raising this.\n\nThis question is for experts in sequence modeling: for LSTM to work, should we take consecutive frames (to encode temporal changes) rather than a few faces uniformly sampled from a video? I already have a uniformly sampled dataset, but I want to make sure if it would be any help if I go for an LSTM model.",
    "740882": "I've only tried to implement Xception+LSTM from this paper https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf, but didn't succeed much, FaceForensics++ has better accuracy.",
    "740891": "Hi @ims0rry, Thanks for your reply! Could you let us know what was the issue? Was is overfitting or something else?",
    "741116": "ims0rry Also, did you try a one stage or two stages training? i.e pretrain resnet backbone only first, freeze resnet weights and then plug and train LSTM only?",
    "741147": "I'm not an expert but according to me, LSTM is a better choice only if we use consecutive frames or at max 3-4 frame jumps between frames. Simple Linear layer should perform better on faces uniformly sampled from video.",
    "741160": "debanga It was underfitting, back then I've been training on 30% of original dataset just to see which models are doing well. \n\n@mpware No I didn't try it yet. It's a good idea though.",
    "741162": "I tried similar approach with no success so far, I tested two different models:\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then for each frame extact features vector using facenet, then feed these vectors to a Bidirectional LSTM ==&gt; score of 0.69\n-Detect and crop to 96x96,  10 consecutives face frames from the video, then feed them to a multilayer ConvLSTM2D model ==&gt; score of 0.69\n\nI use DSFD (Dual Shot Face Detector) for face detection, i found it to be the only model that can detect faces even in the dark videos or videos where the characters are not facing the camera or are too far away. on kaggle notebooks with GPU, detecting and extracting 10 faces takes approximatly 2.5s, so I think we can go up to 20 faces without exceeding the 9h execution time.\n\nHowever I only used 50% of the training data, that I pre processed to get the same number of real and fake videos.\nBoth my models are overfitting after 15 epoch of training.\n\nI still have not figured out how to go below the 0.69 score",
    "741210": "If you consistently get a score of 0.69, I would suspect something is wrong in your code. For a model that doesn't work well, you'd expect to get a higher loss than 0.69.",
    "741213": "96x96 is too small I guess, try to increase image size",
    "741289": "Does  anyone  who  want  to   try   `Transformer` ? \nI  think  `Transformer` is  better  than   `LSTM`,  I'm  gonna  to   have  a  try  ,  if   I  made  any  progress  ,I  will  publish  here.",
    "742039": "I did something similar that got my current score. I changed Resnet to another model(I can't share because my team mate would be mad at me).",
    "742259": "So LSTM works for you? I've just made some tests with 2 CNN + LSTM models and as for @ims0rry **it underfits** (valid logloss lower than train logloss). I don't see why (yet). Same model (only CNN backbone) with same data and validation strategy reaches LB 0.42 but with LSTM it's 0.47. There could be several reasons for underfitting but none for them seems to be the root cause:\n- Regularization: High dropout, aggressive augmentation on train dataset\n- Hard samples in train, easy samples in valid\n- Valid dataset too small\n- Leak in valid dataset\n- LR too small or number epochs too small.\n\n@unkownhihi  How many frames/timesteps did you use for the sequence?\n\nThe 2 models I've tried:\n- CNN backbone + AvgPool (with 2048 features) + LSTM (2 layers, 512 hidden, BiDirectionnal) + Linear\n- CNN backbone + AvgPool + Linear/BN + LSTM (1 layer, 512 hidden, BiDirectionnal) + Attention + Linear",
    "742292": "humananalog Yep i'm suspecting something is wrong with my training data generator, i'm currently reworking that.\n@ims0rry maybe you're right, I will try with bigger frames, and maybe add face alignment\n\nAlso On some of the papers they mention the use of softmax as activation of the last layer, when I tried that my loss score would stay at the same value during the whole training, the validation loss fluctuate a lot however, I can't understant why...",
    "742527": "I tried doing CNN features (post averagepool) -&gt;BiLSTM to my current model and it seemed to perform no better than doing a CNN of multiple frames per video and taking an average. The performance is disappointing, given the validation loss is better for the CNN-&gt;LSTM; it seems the LSTM overfits when judged on the testing sets. It might be that using 1D convolutions of the averagepooling layer is better, though I haven't experimented.",
    "743235": "I didn't do any dropout, no augment, not bidirection, LR of 1e-4. I literally wanted a baseline but turned out to be our best single model score. I didn't add pooling at the end of backbone. Instead, I did backbone(input_size=(...),pooling='avg'). And a hint, the backbone's output vectors are much less.",
    "743701": "Thanks for such info. Simpler model seems to be perform better for you. I will try another attempt in this way. I will also try to add more consecutive frames.",
    "743771": "I've tried self-attention at the penultimate layer of a backbone, instead of LSTM. It gave a better result than LSTM in CV, but LB score was worse by a large margin.\nActually, I observed similar phenomena like @jamesphoward said, the gap of CV v.s. LB is bigger than frame-by-frame models. So we may need to find a way to make CV correlates LB even in temporal modeling... (I've used split by chunk)",
    "745992": "I tried it, but without following the paper you mentioned. Just feeding a number of consecutive face-cropped frames to the pretrained resnet (with final layer removed), then into a LSTM and finally into two dense layers.\n\nResult are ok so far, but nothing too shocking (around 0.5 LB score). Didn't experiment yet with many of the hyper parameters, so there is still some room to improve. \n\nThe biggest challenge I'm having is how long to freeze the pre-trained resnet layers before making them trainable and what learning rate then to use.",
    "747997": "EDIT: I made a model stacking of 2 identical lrcn models trained with different data balancing which got my newer score(0.34012)",
    "748167": "Good to see that there are some good scores possible with this approach. Means there is still some upside for my models :)"
  },
  "source": "meta"
}