{
  "id": 134366,
  "title": "LRCN Implementation (PyTorch)",
  "url": "/competitions/deepfake-detection-challenge/discussion/134366",
  "author_name": "GreatGameDota",
  "post_date": "2020-03-07T18:51:19.407000",
  "votes": 33,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Based on this repo: <a href=\"https://github.com/HHTseng/video-classification\">https://github.com/HHTseng/video-classification</a>\nThis model combines the Xception model in my <a href=\"https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537\">baseline</a> with a LSTM.\nI'm not 100% sure if this is correct but I have been able to train it normally. Any advice/fixes appreciated!\n```\n!pip install pytorchcv --quiet\nfrom pytorchcv.model_provider import get_model\nbase_model = get_model(\"xception\", pretrained=True)\nbase_model = nn.Sequential(*list(model.children())[:-1]) # Remove original output layer\nbase_model[0].final_block.pool = nn.Sequential(nn.AdaptiveAvgPool2d(1))</p>\n\n<p>class Head(torch.nn.Module):\n  def <strong>init</strong>(self, in_f, out_f):\n    super(Head, self).<strong>init</strong>()</p>\n\n<pre><code>self.f = nn.Flatten()\nself.l = nn.Linear(in_f, 512)\nself.d = nn.Dropout(0.75)\nself.o = nn.Linear(512, out_f)\n# self.o = nn.Linear(in_f, out_f)\nself.b1 = nn.BatchNorm1d(in_f)\nself.b2 = nn.BatchNorm1d(512)\nself.r = nn.ReLU()\n</code></pre>\n\n<p>def forward(self, x):\n    x = self.f(x)\n    x = self.b1(x)\n    x = self.d(x)</p>\n\n<pre><code>x = self.l(x)\nx = self.r(x)\nx = self.b2(x)\nx = self.d(x)\n\nout = self.o(x)\nreturn out\n</code></pre>\n\n<p>class CNNEncoder(torch.nn.Module):\n  def <strong>init</strong>(self, base, in_f, out_f):\n    super(CNNEncoder, self).<strong>init</strong>()\n    self.base = base\n    self.h1 = Head(in_f, out_f)</p>\n\n<p>def forward(self, x_3d):\n    cnn_embed_seq = []</p>\n\n<pre><code>for i in range(x_3d.size(1)):\n  x = self.base(x_3d[:,i,:,:,:])\n  x = self.h1(x)\n  cnn_embed_seq.append(x)\n\ncnn_embed_seq = torch.stack(cnn_embed_seq, dim=0).transpose_(0, 1)\nreturn cnn_embed_seq\n</code></pre>\n\n<p>class RNNDecoder(nn.Module):\n  def <strong>init</strong>(self, in_f, out_f):\n    super(RNNDecoder, self).<strong>init</strong>()\n    self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=3,\n        batch_first=True\n    )</p>\n\n<pre><code>self.f1 = nn.Linear(256, 128)\nself.f2 = nn.Linear(128, out_f)\nself.r = nn.ReLU()\nself.d = nn.Dropout(0.3)\n</code></pre>\n\n<p>def forward(self, x):\n    self.LSTM.flatten_parameters()\n    x, (hn,hc) = self.LSTM(x)\n    x = self.d(self.r(self.f1(x[:,-1,:])))\n    x = self.f2(x)\n    return x</p>\n\n<p>class LRCN(nn.Module):\n    def <strong>init</strong>(self, base_model, cnn_features, out_f):\n        super(LRCN, self).<strong>init</strong>()\n        self.cnn = CNNEncoder(base_model, 2048, cnn_features)\n        self.rnn = RNNDecoder(cnn_features, out_f)</p>\n\n<pre><code>def forward(self, x_3d):\n    return self.rnn(self.cnn(x_3d))\n</code></pre>\n\n<p>cnn_features = 300\nmodel = LRCN(base_model, cnn_features, 1)\n```</p>",
  "messages": [
    {
      "id": 766156,
      "postDate": "2020-03-07T18:51:19.407Z",
      "content": "<p>Based on this repo: <a href=\"https://github.com/HHTseng/video-classification\">https://github.com/HHTseng/video-classification</a>\nThis model combines the Xception model in my <a href=\"https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537\">baseline</a> with a LSTM.\nI'm not 100% sure if this is correct but I have been able to train it normally. Any advice/fixes appreciated!\n```\n!pip install pytorchcv --quiet\nfrom pytorchcv.model_provider import get_model\nbase_model = get_model(\"xception\", pretrained=True)\nbase_model = nn.Sequential(*list(model.children())[:-1]) # Remove original output layer\nbase_model[0].final_block.pool = nn.Sequential(nn.AdaptiveAvgPool2d(1))</p>\n\n<p>class Head(torch.nn.Module):\n  def <strong>init</strong>(self, in_f, out_f):\n    super(Head, self).<strong>init</strong>()</p>\n\n<pre><code>self.f = nn.Flatten()\nself.l = nn.Linear(in_f, 512)\nself.d = nn.Dropout(0.75)\nself.o = nn.Linear(512, out_f)\n# self.o = nn.Linear(in_f, out_f)\nself.b1 = nn.BatchNorm1d(in_f)\nself.b2 = nn.BatchNorm1d(512)\nself.r = nn.ReLU()\n</code></pre>\n\n<p>def forward(self, x):\n    x = self.f(x)\n    x = self.b1(x)\n    x = self.d(x)</p>\n\n<pre><code>x = self.l(x)\nx = self.r(x)\nx = self.b2(x)\nx = self.d(x)\n\nout = self.o(x)\nreturn out\n</code></pre>\n\n<p>class CNNEncoder(torch.nn.Module):\n  def <strong>init</strong>(self, base, in_f, out_f):\n    super(CNNEncoder, self).<strong>init</strong>()\n    self.base = base\n    self.h1 = Head(in_f, out_f)</p>\n\n<p>def forward(self, x_3d):\n    cnn_embed_seq = []</p>\n\n<pre><code>for i in range(x_3d.size(1)):\n  x = self.base(x_3d[:,i,:,:,:])\n  x = self.h1(x)\n  cnn_embed_seq.append(x)\n\ncnn_embed_seq = torch.stack(cnn_embed_seq, dim=0).transpose_(0, 1)\nreturn cnn_embed_seq\n</code></pre>\n\n<p>class RNNDecoder(nn.Module):\n  def <strong>init</strong>(self, in_f, out_f):\n    super(RNNDecoder, self).<strong>init</strong>()\n    self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=3,\n        batch_first=True\n    )</p>\n\n<pre><code>self.f1 = nn.Linear(256, 128)\nself.f2 = nn.Linear(128, out_f)\nself.r = nn.ReLU()\nself.d = nn.Dropout(0.3)\n</code></pre>\n\n<p>def forward(self, x):\n    self.LSTM.flatten_parameters()\n    x, (hn,hc) = self.LSTM(x)\n    x = self.d(self.r(self.f1(x[:,-1,:])))\n    x = self.f2(x)\n    return x</p>\n\n<p>class LRCN(nn.Module):\n    def <strong>init</strong>(self, base_model, cnn_features, out_f):\n        super(LRCN, self).<strong>init</strong>()\n        self.cnn = CNNEncoder(base_model, 2048, cnn_features)\n        self.rnn = RNNDecoder(cnn_features, out_f)</p>\n\n<pre><code>def forward(self, x_3d):\n    return self.rnn(self.cnn(x_3d))\n</code></pre>\n\n<p>cnn_features = 300\nmodel = LRCN(base_model, cnn_features, 1)\n```</p>",
      "rawMarkdown": "Based on this repo: https://github.com/HHTseng/video-classification\nThis model combines the Xception model in my [baseline](https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537) with a LSTM.\nI'm not 100% sure if this is correct but I have been able to train it normally. Any advice/fixes appreciated!\n```\n!pip install pytorchcv --quiet\nfrom pytorchcv.model_provider import get_model\nbase_model = get_model(\"xception\", pretrained=True)\nbase_model = nn.Sequential(*list(model.children())[:-1]) # Remove original output layer\nbase_model[0].final_block.pool = nn.Sequential(nn.AdaptiveAvgPool2d(1))\n\nclass Head(torch.nn.Module):\n  def __init__(self, in_f, out_f):\n    super(Head, self).__init__()\n    \n    self.f = nn.Flatten()\n    self.l = nn.Linear(in_f, 512)\n    self.d = nn.Dropout(0.75)\n    self.o = nn.Linear(512, out_f)\n    # self.o = nn.Linear(in_f, out_f)\n    self.b1 = nn.BatchNorm1d(in_f)\n    self.b2 = nn.BatchNorm1d(512)\n    self.r = nn.ReLU()\n\n  def forward(self, x):\n    x = self.f(x)\n    x = self.b1(x)\n    x = self.d(x)\n\n    x = self.l(x)\n    x = self.r(x)\n    x = self.b2(x)\n    x = self.d(x)\n\n    out = self.o(x)\n    return out\n\nclass CNNEncoder(torch.nn.Module):\n  def __init__(self, base, in_f, out_f):\n    super(CNNEncoder, self).__init__()\n    self.base = base\n    self.h1 = Head(in_f, out_f)\n  \n  def forward(self, x_3d):\n    cnn_embed_seq = []\n    \n    for i in range(x_3d.size(1)):\n      x = self.base(x_3d[:,i,:,:,:])\n      x = self.h1(x)\n      cnn_embed_seq.append(x)\n    \n    cnn_embed_seq = torch.stack(cnn_embed_seq, dim=0).transpose_(0, 1)\n    return cnn_embed_seq\n\n\nclass RNNDecoder(nn.Module):\n  def __init__(self, in_f, out_f):\n    super(RNNDecoder, self).__init__()\n    self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=3,\n        batch_first=True\n    )\n\n    self.f1 = nn.Linear(256, 128)\n    self.f2 = nn.Linear(128, out_f)\n    self.r = nn.ReLU()\n    self.d = nn.Dropout(0.3)\n  \n  def forward(self, x):\n    self.LSTM.flatten_parameters()\n    x, (hn,hc) = self.LSTM(x)\n    x = self.d(self.r(self.f1(x[:,-1,:])))\n    x = self.f2(x)\n    return x\n\n\nclass LRCN(nn.Module):\n    def __init__(self, base_model, cnn_features, out_f):\n        super(LRCN, self).__init__()\n        self.cnn = CNNEncoder(base_model, 2048, cnn_features)\n        self.rnn = RNNDecoder(cnn_features, out_f)\n    \n    def forward(self, x_3d):\n        return self.rnn(self.cnn(x_3d))\n\ncnn_features = 300\nmodel = LRCN(base_model, cnn_features, 1)\n```",
      "votes": 32
    },
    {
      "id": 766887,
      "postDate": "2020-03-08T22:36:03.603Z",
      "content": "<p>I've been experimenting with a similar type of model. The LCRN paper mentions that they needed to use a lot of dropout (p=0.9) or the model would overfit. </p>\n\n<p>That's also been my problem: I can get the model to overfit quite easily but haven't quite figured out how to stop that and also get a good validation score.</p>",
      "rawMarkdown": "I've been experimenting with a similar type of model. The LCRN paper mentions that they needed to use a lot of dropout (p=0.9) or the model would overfit. \n\nThat's also been my problem: I can get the model to overfit quite easily but haven't quite figured out how to stop that and also get a good validation score.",
      "votes": 3
    },
    {
      "id": 766563,
      "postDate": "2020-03-08T11:34:59.077Z",
      "content": "<p>I also use it, but didn't get a good score。 </p>",
      "rawMarkdown": "I also use it, but didn't get a good score。 ",
      "votes": 2,
      "replies": [
        {
          "id": 766852,
          "postDate": "2020-03-08T21:02:12.800Z",
          "content": "<p>Same here, simple CNN performs better. Maybe CNN does better with more data and avging over a lot of frames.</p>",
          "rawMarkdown": "Same here, simple CNN performs better. Maybe CNN does better with more data and avging over a lot of frames.",
          "votes": 1
        },
        {
          "id": 768455,
          "postDate": "2020-03-10T20:24:37.227Z",
          "content": "<p>I've tried something similar a few weeks ago. How many frames are you using in the sequence? And are you using consecutive frames in your sequence? I've tried with 4 frames and 10 frames (spaced at regular interval) and it did not work.</p>",
          "rawMarkdown": "I've tried something similar a few weeks ago. How many frames are you using in the sequence? And are you using consecutive frames in your sequence? I've tried with 4 frames and 10 frames (spaced at regular interval) and it did not work."
        },
        {
          "id": 768461,
          "postDate": "2020-03-10T20:36:08.513Z",
          "content": "<p>5 frames, (I used the exact model here) val loss doesn't spike randomly but it still overfit badly.</p>",
          "rawMarkdown": "5 frames, (I used the exact model here) val loss doesn't spike randomly but it still overfit badly.",
          "votes": 1
        },
        {
          "id": 768946,
          "postDate": "2020-03-11T11:30:15.370Z",
          "content": "<p>I use 30 frames(skip frame) in training, and use 10 frames in kaggle, moreover, I ensamble two model results(pred = model_A * 0.5 + model_B * 0.5) ， finally I clipping the preds(preds = preds.clip(0.050.95)) to get a good score (0.36846) in LB</p>",
          "rawMarkdown": "I use 30 frames(skip frame) in training, and use 10 frames in kaggle, moreover, I ensamble two model results(pred = model_A * 0.5 + model_B * 0.5) ， finally I clipping the preds(preds = preds.clip(0.050.95)) to get a good score (0.36846) in LB",
          "votes": 4
        },
        {
          "id": 769073,
          "postDate": "2020-03-11T14:03:27.103Z",
          "content": "<p>How much of a boost does clipping give you?</p>",
          "rawMarkdown": "How much of a boost does clipping give you?"
        },
        {
          "id": 769090,
          "postDate": "2020-03-11T14:20:59.723Z",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  I didn't test the ensamble without clipping, but single CRNN(model_B) just LB=0.43 with the same clipping and without it are 0.67932 (model_A)and 0.66504(model_B), model_A and model_B are two different imagenet CNN model. However the single CNN model's LB is 0.35</p>",
          "rawMarkdown": "@greatgamedota  I didn't test the ensamble without clipping, but single CRNN(model_B) just LB=0.43 with the same clipping and without it are 0.67932 (model_A)and 0.66504(model_B), model_A and model_B are two different imagenet CNN model. However the single CNN model's LB is 0.35",
          "votes": 2
        },
        {
          "id": 776636,
          "postDate": "2020-03-17T14:23:16.797Z",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> Have you tried this lrcn model to get the good score 0.36?</p>",
          "rawMarkdown": "@chenbaoying Have you tried this lrcn model to get the good score 0.36?"
        },
        {
          "id": 776840,
          "postDate": "2020-03-17T17:09:21.797Z",
          "content": "<p><a href=\"/zhuolin\">@zhuolin</a> yes,but I ensamble two model ,now my LCRN score is 0.35</p>",
          "rawMarkdown": "@zhuolin yes,but I ensamble two model ,now my LCRN score is 0.35"
        },
        {
          "id": 776853,
          "postDate": "2020-03-17T17:18:06.057Z",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a>  Do you mean two LRCN models? Thanks.</p>",
          "rawMarkdown": "@chenbaoying  Do you mean two LRCN models? Thanks."
        },
        {
          "id": 778223,
          "postDate": "2020-03-18T08:44:15.883Z",
          "content": "<p>yes, use two LCRN to predict the videos, and the ensamble by final_pred = 0.5*model1_pred+0.5*model2*0.5</p>",
          "rawMarkdown": "yes, use two LCRN to predict the videos, and the ensamble by final_pred = 0.5*model1_pred+0.5*model2*0.5"
        },
        {
          "id": 785589,
          "postDate": "2020-03-25T07:23:43.493Z",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> \nHello, what is your single LRCN model's LB score?  Have you tested?</p>",
          "rawMarkdown": "@chenbaoying \nHello, what is your single LRCN model's LB score?  Have you tested?"
        }
      ]
    },
    {
      "id": 778408,
      "postDate": "2020-03-18T12:16:37.300Z",
      "content": "<p>I have a question <a href=\"/greatgamedota\">@greatgamedota</a> , since I am using your baseline Kernel . We already have the dataset created from multiple videos . How are we going to pass the information or group the sequence of 10-15 images from the videos ? Is it like process images sequentially and group it ? \nI am using the same dataset as your public kernel without ffhq . Like I have shown in my public kernel of densenet169 .</p>",
      "rawMarkdown": "I have a question @greatgamedota , since I am using your baseline Kernel . We already have the dataset created from multiple videos . How are we going to pass the information or group the sequence of 10-15 images from the videos ? Is it like process images sequentially and group it ? \nI am using the same dataset as your public kernel without ffhq . Like I have shown in my public kernel of densenet169 .",
      "replies": [
        {
          "id": 778491,
          "postDate": "2020-03-18T13:49:09.223Z",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> That dataset won't work with this model since it is only 1 image per video. If you want a bigger dataset I have shared one here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134420\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134420</a>\nThat dataset has ~10 images per video, some have more or less.\nTo pass multiple images into the model you just pass a (batch,frame,channel,width,height) tensor.\nI have not gotten this model to perform very well, current score is still using my public model.</p>",
          "rawMarkdown": "@phoenix9032 That dataset won't work with this model since it is only 1 image per video. If you want a bigger dataset I have shared one here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134420\nThat dataset has ~10 images per video, some have more or less.\nTo pass multiple images into the model you just pass a (batch,frame,channel,width,height) tensor.\nI have not gotten this model to perform very well, current score is still using my public model.",
          "votes": 1
        },
        {
          "id": 780156,
          "postDate": "2020-03-20T02:46:13.097Z",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> it could work..... I'm already revealing stuff that can get you high in LB like crazy.</p>",
          "rawMarkdown": "@greatgamedota it could work..... I'm already revealing stuff that can get you high in LB like crazy."
        }
      ]
    },
    {
      "id": 776600,
      "postDate": "2020-03-17T13:54:46.583Z",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Thanks for sharing. Why the <code>num_frames</code> is the depth parameter in the LSTM module \n<code>self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=depth,\n        batch_first=True\n    )</code>?</p>",
      "rawMarkdown": "@greatgamedota Thanks for sharing. Why the `num_frames` is the depth parameter in the LSTM module \n`self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=depth,\n        batch_first=True\n    )`?",
      "replies": [
        {
          "id": 776683,
          "postDate": "2020-03-17T14:53:01.953Z",
          "content": "<p>Looked into it. I was under the impression you had to tell the LSTM how long the time distributed data was but it turns out you don't. I will rename it to just num hidden RNN layers. Thanks!</p>",
          "rawMarkdown": "Looked into it. I was under the impression you had to tell the LSTM how long the time distributed data was but it turns out you don't. I will rename it to just num hidden RNN layers. Thanks!"
        },
        {
          "id": 776848,
          "postDate": "2020-03-17T17:15:15.023Z",
          "content": "<p>Yes. When I read your code, it is strange. Maybe you will get good score after you fix this :)</p>",
          "rawMarkdown": "Yes. When I read your code, it is strange. Maybe you will get good score after you fix this :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 766996,
      "postDate": "2020-03-09T02:33:40.790Z",
      "content": "<p>Do we need to crop for face when using RNN model ?  Or should we fixed the area where embedding is extracted ? </p>",
      "rawMarkdown": "Do we need to crop for face when using RNN model ?  Or should we fixed the area where embedding is extracted ? \n\n",
      "replies": [
        {
          "id": 767001,
          "postDate": "2020-03-09T02:42:14.930Z",
          "content": "<p>You could run it with an entire frame but I'd run it with just the faces cropped.</p>",
          "rawMarkdown": "You could run it with an entire frame but I'd run it with just the faces cropped."
        }
      ]
    },
    {
      "id": 769619,
      "postDate": "2020-03-12T05:02:40.193Z",
      "content": "<p>Thanks for sharing bro</p>",
      "rawMarkdown": "Thanks for sharing bro",
      "votes": 1
    },
    {
      "id": 769027,
      "postDate": "2020-03-11T12:51:47.157Z",
      "content": "<p>thanks bro</p>",
      "rawMarkdown": "thanks bro"
    },
    {
      "id": 768228,
      "postDate": "2020-03-10T15:04:33.003Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing"
    },
    {
      "id": 767243,
      "postDate": "2020-03-09T11:09:50.933Z",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks"
    },
    {
      "id": 769655,
      "postDate": "2020-03-12T05:54:07.243Z",
      "content": "<p>thanks a lot</p>",
      "rawMarkdown": "thanks a lot",
      "isDeleted": true
    },
    {
      "id": 767617,
      "postDate": "2020-03-09T22:33:35.110Z",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 766887,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2020-03-08T22:36:03.603000",
      "content": "<p>I've been experimenting with a similar type of model. The LCRN paper mentions that they needed to use a lot of dropout (p=0.9) or the model would overfit. </p>\n\n<p>That's also been my problem: I can get the model to overfit quite easily but haven't quite figured out how to stop that and also get a good validation score.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 766563,
      "author_name": "BokingChen",
      "author_url": "",
      "post_date": "2020-03-08T11:34:59.077000",
      "content": "<p>I also use it, but didn't get a good score。 </p>",
      "votes": 2,
      "replies": [
        {
          "id": 766852,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-08T21:02:12.800000",
          "content": "<p>Same here, simple CNN performs better. Maybe CNN does better with more data and avging over a lot of frames.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 768455,
          "author_name": "MPWARE",
          "author_url": "",
          "post_date": "2020-03-10T20:24:37.227000",
          "content": "<p>I've tried something similar a few weeks ago. How many frames are you using in the sequence? And are you using consecutive frames in your sequence? I've tried with 4 frames and 10 frames (spaced at regular interval) and it did not work.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768461,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-10T20:36:08.513000",
          "content": "<p>5 frames, (I used the exact model here) val loss doesn't spike randomly but it still overfit badly.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 768946,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-03-11T11:30:15.370000",
          "content": "<p>I use 30 frames(skip frame) in training, and use 10 frames in kaggle, moreover, I ensamble two model results(pred = model_A * 0.5 + model_B * 0.5) ， finally I clipping the preds(preds = preds.clip(0.050.95)) to get a good score (0.36846) in LB</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 769073,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-11T14:03:27.103000",
          "content": "<p>How much of a boost does clipping give you?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769090,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-03-11T14:20:59.723000",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  I didn't test the ensamble without clipping, but single CRNN(model_B) just LB=0.43 with the same clipping and without it are 0.67932 (model_A)and 0.66504(model_B), model_A and model_B are two different imagenet CNN model. However the single CNN model's LB is 0.35</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 776636,
          "author_name": "zjiang",
          "author_url": "",
          "post_date": "2020-03-17T14:23:16.797000",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> Have you tried this lrcn model to get the good score 0.36?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 776840,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-03-17T17:09:21.797000",
          "content": "<p><a href=\"/zhuolin\">@zhuolin</a> yes,but I ensamble two model ,now my LCRN score is 0.35</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 776853,
          "author_name": "zjiang",
          "author_url": "",
          "post_date": "2020-03-17T17:18:06.057000",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a>  Do you mean two LRCN models? Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 778223,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-03-18T08:44:15.883000",
          "content": "<p>yes, use two LCRN to predict the videos, and the ensamble by final_pred = 0.5*model1_pred+0.5*model2*0.5</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 785589,
          "author_name": "xmaocai",
          "author_url": "",
          "post_date": "2020-03-25T07:23:43.493000",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> \nHello, what is your single LRCN model's LB score?  Have you tested?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 778408,
      "author_name": "Nirjhar Roy",
      "author_url": "",
      "post_date": "2020-03-18T12:16:37.300000",
      "content": "<p>I have a question <a href=\"/greatgamedota\">@greatgamedota</a> , since I am using your baseline Kernel . We already have the dataset created from multiple videos . How are we going to pass the information or group the sequence of 10-15 images from the videos ? Is it like process images sequentially and group it ? \nI am using the same dataset as your public kernel without ffhq . Like I have shown in my public kernel of densenet169 .</p>",
      "votes": 0,
      "replies": [
        {
          "id": 778491,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-18T13:49:09.223000",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> That dataset won't work with this model since it is only 1 image per video. If you want a bigger dataset I have shared one here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134420\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/134420</a>\nThat dataset has ~10 images per video, some have more or less.\nTo pass multiple images into the model you just pass a (batch,frame,channel,width,height) tensor.\nI have not gotten this model to perform very well, current score is still using my public model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 780156,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-03-20T02:46:13.097000",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> it could work..... I'm already revealing stuff that can get you high in LB like crazy.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 776600,
      "author_name": "zjiang",
      "author_url": "",
      "post_date": "2020-03-17T13:54:46.583000",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Thanks for sharing. Why the <code>num_frames</code> is the depth parameter in the LSTM module \n<code>self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=depth,\n        batch_first=True\n    )</code>?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 776683,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-17T14:53:01.953000",
          "content": "<p>Looked into it. I was under the impression you had to tell the LSTM how long the time distributed data was but it turns out you don't. I will rename it to just num hidden RNN layers. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 776848,
          "author_name": "zjiang",
          "author_url": "",
          "post_date": "2020-03-17T17:15:15.023000",
          "content": "<p>Yes. When I read your code, it is strange. Maybe you will get good score after you fix this :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 766996,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2020-03-09T02:33:40.790000",
      "content": "<p>Do we need to crop for face when using RNN model ?  Or should we fixed the area where embedding is extracted ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 767001,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-09T02:42:14.930000",
          "content": "<p>You could run it with an entire frame but I'd run it with just the faces cropped.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 769619,
      "author_name": "Abhishek Kumar Pandit",
      "author_url": "",
      "post_date": "2020-03-12T05:02:40.193000",
      "content": "<p>Thanks for sharing bro</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 769027,
      "author_name": "Asd_123",
      "author_url": "",
      "post_date": "2020-03-11T12:51:47.157000",
      "content": "<p>thanks bro</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 768228,
      "author_name": "Kranti Kumar",
      "author_url": "",
      "post_date": "2020-03-10T15:04:33.003000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767243,
      "author_name": "mustafa mahmutoglu",
      "author_url": "",
      "post_date": "2020-03-09T11:09:50.933000",
      "content": "<p>thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 769655,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-12T05:54:07.243000",
      "content": "<p>thanks a lot</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 767617,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-09T22:33:35.110000",
      "content": "<p>Thanks for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "766156": "Based on this repo: https://github.com/HHTseng/video-classification\nThis model combines the Xception model in my [baseline](https://www.kaggle.com/greatgamedota/xception-classifier-w-ffhq-training-lb-537) with a LSTM.\nI'm not 100% sure if this is correct but I have been able to train it normally. Any advice/fixes appreciated!\n```\n!pip install pytorchcv --quiet\nfrom pytorchcv.model_provider import get_model\nbase_model = get_model(\"xception\", pretrained=True)\nbase_model = nn.Sequential(*list(model.children())[:-1]) # Remove original output layer\nbase_model[0].final_block.pool = nn.Sequential(nn.AdaptiveAvgPool2d(1))\n\nclass Head(torch.nn.Module):\n  def __init__(self, in_f, out_f):\n    super(Head, self).__init__()\n    \n    self.f = nn.Flatten()\n    self.l = nn.Linear(in_f, 512)\n    self.d = nn.Dropout(0.75)\n    self.o = nn.Linear(512, out_f)\n    # self.o = nn.Linear(in_f, out_f)\n    self.b1 = nn.BatchNorm1d(in_f)\n    self.b2 = nn.BatchNorm1d(512)\n    self.r = nn.ReLU()\n\n  def forward(self, x):\n    x = self.f(x)\n    x = self.b1(x)\n    x = self.d(x)\n\n    x = self.l(x)\n    x = self.r(x)\n    x = self.b2(x)\n    x = self.d(x)\n\n    out = self.o(x)\n    return out\n\nclass CNNEncoder(torch.nn.Module):\n  def __init__(self, base, in_f, out_f):\n    super(CNNEncoder, self).__init__()\n    self.base = base\n    self.h1 = Head(in_f, out_f)\n  \n  def forward(self, x_3d):\n    cnn_embed_seq = []\n    \n    for i in range(x_3d.size(1)):\n      x = self.base(x_3d[:,i,:,:,:])\n      x = self.h1(x)\n      cnn_embed_seq.append(x)\n    \n    cnn_embed_seq = torch.stack(cnn_embed_seq, dim=0).transpose_(0, 1)\n    return cnn_embed_seq\n\n\nclass RNNDecoder(nn.Module):\n  def __init__(self, in_f, out_f):\n    super(RNNDecoder, self).__init__()\n    self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=3,\n        batch_first=True\n    )\n\n    self.f1 = nn.Linear(256, 128)\n    self.f2 = nn.Linear(128, out_f)\n    self.r = nn.ReLU()\n    self.d = nn.Dropout(0.3)\n  \n  def forward(self, x):\n    self.LSTM.flatten_parameters()\n    x, (hn,hc) = self.LSTM(x)\n    x = self.d(self.r(self.f1(x[:,-1,:])))\n    x = self.f2(x)\n    return x\n\n\nclass LRCN(nn.Module):\n    def __init__(self, base_model, cnn_features, out_f):\n        super(LRCN, self).__init__()\n        self.cnn = CNNEncoder(base_model, 2048, cnn_features)\n        self.rnn = RNNDecoder(cnn_features, out_f)\n    \n    def forward(self, x_3d):\n        return self.rnn(self.cnn(x_3d))\n\ncnn_features = 300\nmodel = LRCN(base_model, cnn_features, 1)\n```",
    "766887": "I've been experimenting with a similar type of model. The LCRN paper mentions that they needed to use a lot of dropout (p=0.9) or the model would overfit. \n\nThat's also been my problem: I can get the model to overfit quite easily but haven't quite figured out how to stop that and also get a good validation score.",
    "766563": "I also use it, but didn't get a good score。 ",
    "778408": "I have a question @greatgamedota , since I am using your baseline Kernel . We already have the dataset created from multiple videos . How are we going to pass the information or group the sequence of 10-15 images from the videos ? Is it like process images sequentially and group it ? \nI am using the same dataset as your public kernel without ffhq . Like I have shown in my public kernel of densenet169 .",
    "776600": "@greatgamedota Thanks for sharing. Why the `num_frames` is the depth parameter in the LSTM module \n`self.LSTM = nn.LSTM(\n        input_size=in_f,\n        hidden_size=256,\n        num_layers=depth,\n        batch_first=True\n    )`?",
    "766996": "Do we need to crop for face when using RNN model ?  Or should we fixed the area where embedding is extracted ? \n\n",
    "769619": "Thanks for sharing bro",
    "769027": "thanks bro",
    "768228": "Thanks for sharing",
    "767243": "thanks",
    "769655": "thanks a lot",
    "767617": "Thanks for sharing"
  }
}