{
  "id": 140835,
  "title": "Why RNN (like LSTM) didn't worked?",
  "url": "/competitions/deepfake-detection-challenge/discussion/140835",
  "author_name": "",
  "post_date": "2020-04-03T11:21:04.833318900Z",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>As I am going through the discussion, I see that LSTM didn't worked well. Any reason why it is the case? I saw many deepfake videos have certain face regions where there is a lot of flickering, I was hoping LSTM would be able to pick up on those but it didn't generalized.</p>\n\n<p>In my case it was severely over-fitting for the faces the model has seen before: If I would randomly sample my train-test set from same folders my test accuracy was good(close to top public score here), but if I choose test set as different folder then model becomes way worse than random guessing. Did anyone else encountered this issue? I tried reducing the model capacity but it didn't generalized at all.</p>",
  "messages": [
    {
      "id": "796175",
      "postDate": "04/03/2020 11:21:04",
      "content": "<p>As I am going through the discussion, I see that LSTM didn't worked well. Any reason why it is the case? I saw many deepfake videos have certain face regions where there is a lot of flickering, I was hoping LSTM would be able to pick up on those but it didn't generalized.</p>\n\n<p>In my case it was severely over-fitting for the faces the model has seen before: If I would randomly sample my train-test set from same folders my test accuracy was good(close to top public score here), but if I choose test set as different folder then model becomes way worse than random guessing. Did anyone else encountered this issue? I tried reducing the model capacity but it didn't generalized at all.</p>",
      "rawMarkdown": "As I am going through the discussion, I see that LSTM didn't worked well. Any reason why it is the case? I saw many deepfake videos have certain face regions where there is a lot of flickering, I was hoping LSTM would be able to pick up on those but it didn't generalized.\n\nIn my case it was severely over-fitting for the faces the model has seen before: If I would randomly sample my train-test set from same folders my test accuracy was good(close to top public score here), but if I choose test set as different folder then model becomes way worse than random guessing. Did anyone else encountered this issue? I tried reducing the model capacity but it didn't generalized at all.",
      "votes": null
    },
    {
      "id": "796185",
      "postDate": "04/03/2020 11:31:48",
      "content": "<p>Hi!I also try EfficientNet + LSTM structure.But my model didn't converge.In a balanced dataset,I read 20 frames,then i ut it into LSTM,at the last I use FullyConnect layer,and average the output,and its shape is squeezed as [1, 1], and I use crossEntropy to compute loss,but my model's accuracy is around 50%~60%.</p>\n\n<p>I'm so confused......Is there any problem in my process? hope you can give me some suggestions!</p>",
      "rawMarkdown": "Hi!I also try EfficientNet + LSTM structure.But my model didn't converge.In a balanced dataset,I read 20 frames,then i ut it into LSTM,at the last I use FullyConnect layer,and average the output,and its shape is squeezed as [1, 1], and I use crossEntropy to compute loss,but my model's accuracy is around 50%~60%.\n\nI'm so confused......Is there any problem in my process? hope you can give me some suggestions!",
      "votes": null
    },
    {
      "id": "796200",
      "postDate": "04/03/2020 11:48:02",
      "content": "<p>I had exactly the same experience. Lstm/gru model on extracted features overfitted very quickly (1 epoch) no matter what I did. I even tried just 2/3 frames sequence as i have noticed a clear visual difference but to no avail. I was actually hoping this worked for someone as this seems so logical and correct way of doing this... Alas, closest thing was feeding several feature vectors to dense layer. I have no explanation as to why it doesn't work. Not even a goid hypothesis. </p>",
      "rawMarkdown": "I had exactly the same experience. Lstm/gru model on extracted features overfitted very quickly (1 epoch) no matter what I did. I even tried just 2/3 frames sequence as i have noticed a clear visual difference but to no avail. I was actually hoping this worked for someone as this seems so logical and correct way of doing this... Alas, closest thing was feeding several feature vectors to dense layer. I have no explanation as to why it doesn't work. Not even a goid hypothesis.",
      "votes": null
    },
    {
      "id": "796450",
      "postDate": "04/03/2020 15:36:36",
      "content": "<p>I found an interesting article that might explain some of the stuff that was happening:\n<a href=\"https://towardsdatascience.com/sigmoid-activation-and-binary-crossentropy-a-less-than-perfect-match-b801e130e31\">https://towardsdatascience.com/sigmoid-activation-and-binary-crossentropy-a-less-than-perfect-match-b801e130e31</a></p>",
      "rawMarkdown": "I found an interesting article that might explain some of the stuff that was happening:\nhttps://towardsdatascience.com/sigmoid-activation-and-binary-crossentropy-a-less-than-perfect-match-b801e130e31",
      "votes": null
    },
    {
      "id": "796871",
      "postDate": "04/04/2020 02:41:53",
      "content": "<p>I committed to using a CNN LSTM model with Xception. My test loss was consistently less than 0.40 but the LB score would always be at 0.69. I tried different sampling techniques considering the different folders. Adding a significant amount of Dropout helped to address some of the overfitting (I found 0.75 Dropout to be optimal). I used both face detection cropping and full frames and achieved the same results.</p>",
      "rawMarkdown": "I committed to using a CNN LSTM model with Xception. My test loss was consistently less than 0.40 but the LB score would always be at 0.69. I tried different sampling techniques considering the different folders. Adding a significant amount of Dropout helped to address some of the overfitting (I found 0.75 Dropout to be optimal). I used both face detection cropping and full frames and achieved the same results.",
      "votes": null
    },
    {
      "id": "796876",
      "postDate": "04/04/2020 03:03:26",
      "content": "<p>0.69 is the magical 0.5. Perhaps your error handling was at fault? The hidden test videos had some corrupt videos. If you initialized the submission to 0.5 and exited on corrupt video you would get 0.69.</p>",
      "rawMarkdown": "0.69 is the magical 0.5. Perhaps your error handling was at fault? The hidden test videos had some corrupt videos. If you initialized the submission to 0.5 and exited on corrupt video you would get 0.69.",
      "votes": null
    },
    {
      "id": "796879",
      "postDate": "04/04/2020 03:08:10",
      "content": "<p>The lstm models are using the features vectors so its before the Sigmoid and crossentropy. It might explain why the single cnn training val loss was so not smooth but i doubt it. I think it was just the problem of getting the cv to work well. </p>",
      "rawMarkdown": "The lstm models are using the features vectors so its before the Sigmoid and crossentropy. It might explain why the single cnn training val loss was so not smooth but i doubt it. I think it was just the problem of getting the cv to work well.",
      "votes": null
    },
    {
      "id": "797989",
      "postDate": "04/05/2020 04:53:50",
      "content": "<p>I initially had this issue for quite long time: EfficientNet + LSTM. Training was converted but validation is very bad, loss: 1.x and accuracy 60%. Then I added two Dense(512) and Dropout. Trained with SGD with momentum. Finally validation loss and accuracy better, loss &lt; 0.1 and accuracy 80%.</p>",
      "rawMarkdown": "I initially had this issue for quite long time: EfficientNet + LSTM. Training was converted but validation is very bad, loss: 1.x and accuracy 60%. Then I added two Dense(512) and Dropout. Trained with SGD with momentum. Finally validation loss and accuracy better, loss &lt; 0.1 and accuracy 80%.",
      "votes": null
    },
    {
      "id": "798079",
      "postDate": "04/05/2020 06:42:23",
      "content": "<p>Did you submit? My lstm showed amazingly score on my cv. On lb, it was horrible. </p>",
      "rawMarkdown": "Did you submit? My lstm showed amazingly score on my cv. On lb, it was horrible.",
      "votes": null
    },
    {
      "id": "798935",
      "postDate": "04/06/2020 02:19:45",
      "content": "<p><a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\">https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference</a></p>\n\n<p>Isn't this considered as RNN as well ? It's time distributed LSTM with final Dense layer. </p>",
      "rawMarkdown": "https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\n\nIsn't this considered as RNN as well ? It's time distributed LSTM with final Dense layer.",
      "votes": null
    },
    {
      "id": "798941",
      "postDate": "04/06/2020 02:24:40",
      "content": "<p>No I did not. It was the last try on last day which after so many attempts very bad and I did not believe it would give a different result, so let it run and I slept :). Later I checked again with other hold out validation set, turned out not that good. </p>",
      "rawMarkdown": "No I did not. It was the last try on last day which after so many attempts very bad and I did not believe it would give a different result, so let it run and I slept :). Later I checked again with other hold out validation set, turned out not that good.",
      "votes": null
    },
    {
      "id": "798945",
      "postDate": "04/06/2020 02:26:55",
      "content": "<p>Yes it is but it didn't get a very good score iirc. Also didn't scale well to more frames or higher resolution. </p>",
      "rawMarkdown": "Yes it is but it didn't get a very good score iirc. Also didn't scale well to more frames or higher resolution.",
      "votes": null
    },
    {
      "id": "798957",
      "postDate": "04/06/2020 02:42:16",
      "content": "<p>I would guess it would show around 0.5 on the lb. a mystery. All instincts points to this being the right way to go. The lstm should have noticed the difference between natural and synthetic facial changes. But, it doesn't. I don't even know how to see activation maps on lstm so don't know what it was looking at that produced such good cv but bad lb. </p>",
      "rawMarkdown": "I would guess it would show around 0.5 on the lb. a mystery. All instincts points to this being the right way to go. The lstm should have noticed the difference between natural and synthetic facial changes. But, it doesn't. I don't even know how to see activation maps on lstm so don't know what it was looking at that produced such good cv but bad lb.",
      "votes": null
    },
    {
      "id": "799000",
      "postDate": "04/06/2020 04:16:25",
      "content": "<p>Large BCE in LB does not mean that it does not work. LSTM may also have high accuracy with large BCE in LB.</p>",
      "rawMarkdown": "Large BCE in LB does not mean that it does not work. LSTM may also have high accuracy with large BCE in LB.",
      "votes": null
    },
    {
      "id": "800156",
      "postDate": "04/07/2020 06:22:04",
      "content": "<p>I got 0.33 using lstm on LB, but used embeddings from my binary classifier, which scored better than 0.33, I think it was not the right way to do this. Maybe some inconsistency in the frame rate can lead to poor results (if you trained it on every 15th frame and made inference on every 5th frame )  </p>",
      "rawMarkdown": "I got 0.33 using lstm on LB, but used embeddings from my binary classifier, which scored better than 0.33, I think it was not the right way to do this. Maybe some inconsistency in the frame rate can lead to poor results (if you trained it on every 15th frame and made inference on every 5th frame )",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 796185,
      "author_name": "zzk7love",
      "author_url": "",
      "post_date": "04/03/2020 11:31:48",
      "content": "<p>Hi!I also try EfficientNet + LSTM structure.But my model didn't converge.In a balanced dataset,I read 20 frames,then i ut it into LSTM,at the last I use FullyConnect layer,and average the output,and its shape is squeezed as [1, 1], and I use crossEntropy to compute loss,but my model's accuracy is around 50%~60%.</p>\n\n<p>I'm so confused......Is there any problem in my process? hope you can give me some suggestions!</p>",
      "votes": null,
      "replies": [
        {
          "id": 797989,
          "author_name": "zungmann",
          "author_url": "",
          "post_date": "04/05/2020 04:53:50",
          "content": "<p>I initially had this issue for quite long time: EfficientNet + LSTM. Training was converted but validation is very bad, loss: 1.x and accuracy 60%. Then I added two Dense(512) and Dropout. Trained with SGD with momentum. Finally validation loss and accuracy better, loss &lt; 0.1 and accuracy 80%.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 798079,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "04/05/2020 06:42:23",
          "content": "<p>Did you submit? My lstm showed amazingly score on my cv. On lb, it was horrible. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 798941,
          "author_name": "zungmann",
          "author_url": "",
          "post_date": "04/06/2020 02:24:40",
          "content": "<p>No I did not. It was the last try on last day which after so many attempts very bad and I did not believe it would give a different result, so let it run and I slept :). Later I checked again with other hold out validation set, turned out not that good. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 798957,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "04/06/2020 02:42:16",
          "content": "<p>I would guess it would show around 0.5 on the lb. a mystery. All instincts points to this being the right way to go. The lstm should have noticed the difference between natural and synthetic facial changes. But, it doesn't. I don't even know how to see activation maps on lstm so don't know what it was looking at that produced such good cv but bad lb. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 800156,
          "author_name": "azamatk",
          "author_url": "",
          "post_date": "04/07/2020 06:22:04",
          "content": "<p>I got 0.33 using lstm on LB, but used embeddings from my binary classifier, which scored better than 0.33, I think it was not the right way to do this. Maybe some inconsistency in the frame rate can lead to poor results (if you trained it on every 15th frame and made inference on every 5th frame )  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 796200,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "04/03/2020 11:48:02",
      "content": "<p>I had exactly the same experience. Lstm/gru model on extracted features overfitted very quickly (1 epoch) no matter what I did. I even tried just 2/3 frames sequence as i have noticed a clear visual difference but to no avail. I was actually hoping this worked for someone as this seems so logical and correct way of doing this... Alas, closest thing was feeding several feature vectors to dense layer. I have no explanation as to why it doesn't work. Not even a goid hypothesis. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 796450,
      "author_name": "mont3z",
      "author_url": "",
      "post_date": "04/03/2020 15:36:36",
      "content": "<p>I found an interesting article that might explain some of the stuff that was happening:\n<a href=\"https://towardsdatascience.com/sigmoid-activation-and-binary-crossentropy-a-less-than-perfect-match-b801e130e31\">https://towardsdatascience.com/sigmoid-activation-and-binary-crossentropy-a-less-than-perfect-match-b801e130e31</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 796879,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "04/04/2020 03:08:10",
          "content": "<p>The lstm models are using the features vectors so its before the Sigmoid and crossentropy. It might explain why the single cnn training val loss was so not smooth but i doubt it. I think it was just the problem of getting the cv to work well. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 796871,
      "author_name": "dalekube",
      "author_url": "",
      "post_date": "04/04/2020 02:41:53",
      "content": "<p>I committed to using a CNN LSTM model with Xception. My test loss was consistently less than 0.40 but the LB score would always be at 0.69. I tried different sampling techniques considering the different folders. Adding a significant amount of Dropout helped to address some of the overfitting (I found 0.75 Dropout to be optimal). I used both face detection cropping and full frames and achieved the same results.</p>",
      "votes": null,
      "replies": [
        {
          "id": 796876,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "04/04/2020 03:03:26",
          "content": "<p>0.69 is the magical 0.5. Perhaps your error handling was at fault? The hidden test videos had some corrupt videos. If you initialized the submission to 0.5 and exited on corrupt video you would get 0.69.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 798935,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "04/06/2020 02:19:45",
      "content": "<p><a href=\"https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\">https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference</a></p>\n\n<p>Isn't this considered as RNN as well ? It's time distributed LSTM with final Dense layer. </p>",
      "votes": null,
      "replies": [
        {
          "id": 798945,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "04/06/2020 02:26:55",
          "content": "<p>Yes it is but it didn't get a very good score iirc. Also didn't scale well to more frames or higher resolution. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 799000,
      "author_name": "yuanzhezhou",
      "author_url": "",
      "post_date": "04/06/2020 04:16:25",
      "content": "<p>Large BCE in LB does not mean that it does not work. LSTM may also have high accuracy with large BCE in LB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "796175": "As I am going through the discussion, I see that LSTM didn't worked well. Any reason why it is the case? I saw many deepfake videos have certain face regions where there is a lot of flickering, I was hoping LSTM would be able to pick up on those but it didn't generalized.\n\nIn my case it was severely over-fitting for the faces the model has seen before: If I would randomly sample my train-test set from same folders my test accuracy was good(close to top public score here), but if I choose test set as different folder then model becomes way worse than random guessing. Did anyone else encountered this issue? I tried reducing the model capacity but it didn't generalized at all.",
    "796185": "Hi!I also try EfficientNet + LSTM structure.But my model didn't converge.In a balanced dataset,I read 20 frames,then i ut it into LSTM,at the last I use FullyConnect layer,and average the output,and its shape is squeezed as [1, 1], and I use crossEntropy to compute loss,but my model's accuracy is around 50%~60%.\n\nI'm so confused......Is there any problem in my process? hope you can give me some suggestions!",
    "796200": "I had exactly the same experience. Lstm/gru model on extracted features overfitted very quickly (1 epoch) no matter what I did. I even tried just 2/3 frames sequence as i have noticed a clear visual difference but to no avail. I was actually hoping this worked for someone as this seems so logical and correct way of doing this... Alas, closest thing was feeding several feature vectors to dense layer. I have no explanation as to why it doesn't work. Not even a goid hypothesis.",
    "796450": "I found an interesting article that might explain some of the stuff that was happening:\nhttps://towardsdatascience.com/sigmoid-activation-and-binary-crossentropy-a-less-than-perfect-match-b801e130e31",
    "796871": "I committed to using a CNN LSTM model with Xception. My test loss was consistently less than 0.40 but the LB score would always be at 0.69. I tried different sampling techniques considering the different folders. Adding a significant amount of Dropout helped to address some of the overfitting (I found 0.75 Dropout to be optimal). I used both face detection cropping and full frames and achieved the same results.",
    "796876": "0.69 is the magical 0.5. Perhaps your error handling was at fault? The hidden test videos had some corrupt videos. If you initialized the submission to 0.5 and exited on corrupt video you would get 0.69.",
    "796879": "The lstm models are using the features vectors so its before the Sigmoid and crossentropy. It might explain why the single cnn training val loss was so not smooth but i doubt it. I think it was just the problem of getting the cv to work well.",
    "797989": "I initially had this issue for quite long time: EfficientNet + LSTM. Training was converted but validation is very bad, loss: 1.x and accuracy 60%. Then I added two Dense(512) and Dropout. Trained with SGD with momentum. Finally validation loss and accuracy better, loss &lt; 0.1 and accuracy 80%.",
    "798079": "Did you submit? My lstm showed amazingly score on my cv. On lb, it was horrible.",
    "798935": "https://www.kaggle.com/unkownhihi/dfdc-lrcn-inference\n\nIsn't this considered as RNN as well ? It's time distributed LSTM with final Dense layer.",
    "798941": "No I did not. It was the last try on last day which after so many attempts very bad and I did not believe it would give a different result, so let it run and I slept :). Later I checked again with other hold out validation set, turned out not that good.",
    "798945": "Yes it is but it didn't get a very good score iirc. Also didn't scale well to more frames or higher resolution.",
    "798957": "I would guess it would show around 0.5 on the lb. a mystery. All instincts points to this being the right way to go. The lstm should have noticed the difference between natural and synthetic facial changes. But, it doesn't. I don't even know how to see activation maps on lstm so don't know what it was looking at that produced such good cv but bad lb.",
    "799000": "Large BCE in LB does not mean that it does not work. LSTM may also have high accuracy with large BCE in LB.",
    "800156": "I got 0.33 using lstm on LB, but used embeddings from my binary classifier, which scored better than 0.33, I think it was not the right way to do this. Maybe some inconsistency in the frame rate can lead to poor results (if you trained it on every 15th frame and made inference on every 5th frame )"
  },
  "source": "meta"
}