{
  "id": 131274,
  "title": "Best single model score",
  "url": "/competitions/deepfake-detection-challenge/discussion/131274",
  "author_name": "Shangqiu Li",
  "post_date": "2020-02-19T04:33:21.054000",
  "votes": 28,
  "comment_count": 139,
  "views": 0,
  "content": "<p>Please share your best single model LB(not ensemble).\nTo start with ours: 0.38504\nEDIT: 0.35040\nEDIT2: 0.34</p>",
  "messages": [
    {
      "id": 750030,
      "postDate": "2020-02-19T04:33:21.053Z",
      "content": "<p>Please share your best single model LB(not ensemble).\nTo start with ours: 0.38504\nEDIT: 0.35040\nEDIT2: 0.34</p>",
      "rawMarkdown": "Please share your best single model LB(not ensemble).\nTo start with ours: 0.38504\nEDIT: 0.35040\nEDIT2: 0.34",
      "votes": 27
    },
    {
      "id": 751313,
      "postDate": "2020-02-20T05:11:48.030Z",
      "content": "<p>Frame-by-frame model (only image): 0.33966 </p>",
      "rawMarkdown": "Frame-by-frame model (only image): 0.33966 ",
      "votes": 9,
      "replies": [
        {
          "id": 751314,
          "postDate": "2020-02-20T05:16:51.860Z",
          "content": "<p>what do you mean by frame-by-frame model?</p>",
          "rawMarkdown": "what do you mean by frame-by-frame model?"
        },
        {
          "id": 751352,
          "postDate": "2020-02-20T06:35:50.630Z",
          "content": "<p>It means a model makes a prediction to each frame (and then aggregates over frames).</p>",
          "rawMarkdown": "It means a model makes a prediction to each frame (and then aggregates over frames).",
          "votes": 1
        },
        {
          "id": 751516,
          "postDate": "2020-02-20T08:49:28.887Z",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a> Great score for a frame only model. How did you choose your validation set? And are you using any external data during training? </p>",
          "rawMarkdown": "@lyakaap Great score for a frame only model. How did you choose your validation set? And are you using any external data during training? "
        },
        {
          "id": 751646,
          "postDate": "2020-02-20T11:14:06.047Z",
          "content": "<p>Just selecting 20% of folder as validation set. Used no external data.</p>",
          "rawMarkdown": "Just selecting 20% of folder as validation set. Used no external data.",
          "votes": 1
        },
        {
          "id": 751652,
          "postDate": "2020-02-20T11:17:56.207Z",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a> You mean 20% from each folder, was your validation able to track your leaderboard?</p>",
          "rawMarkdown": "@lyakaap You mean 20% from each folder, was your validation able to track your leaderboard?"
        },
        {
          "id": 751964,
          "postDate": "2020-02-20T17:23:00.673Z",
          "content": "<p>No, split by folder index. This strategy can track well.</p>",
          "rawMarkdown": "No, split by folder index. This strategy can track well."
        },
        {
          "id": 752019,
          "postDate": "2020-02-20T17:50:02.343Z",
          "content": "<p>what is folder index?</p>",
          "rawMarkdown": "what is folder index?"
        },
        {
          "id": 752101,
          "postDate": "2020-02-20T18:39:46.010Z",
          "content": "<p>sorry for confusing, what I wanted to say is I split val/train based on the index of input zip files (dfdc_train_part_XX).</p>",
          "rawMarkdown": "sorry for confusing, what I wanted to say is I split val/train based on the index of input zip files (dfdc_train_part_XX).",
          "votes": 1
        },
        {
          "id": 752150,
          "postDate": "2020-02-20T19:12:24.757Z",
          "content": "<p>I am also doing the same but in particular last 10 dfdc_train_part_XX where XX is range(40,50), is our split same or you are randomly selecting those? </p>",
          "rawMarkdown": "I am also doing the same but in particular last 10 dfdc_train_part_XX where XX is range(40,50), is our split same or you are randomly selecting those? ",
          "votes": 2
        },
        {
          "id": 752474,
          "postDate": "2020-02-21T04:03:33.333Z",
          "content": "<p>Linear division may works.</p>",
          "rawMarkdown": "Linear division may works."
        },
        {
          "id": 752492,
          "postDate": "2020-02-21T04:33:18.960Z",
          "content": "<p>Same as yours.</p>",
          "rawMarkdown": "Same as yours."
        },
        {
          "id": 752507,
          "postDate": "2020-02-21T05:00:03.577Z",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a> That is a very good score for a frame by frame model(our's best is just 0.4418). Do you consider teaming with us? A model stacking will be really helpful.</p>",
          "rawMarkdown": "@lyakaap That is a very good score for a frame by frame model(our's best is just 0.4418). Do you consider teaming with us? A model stacking will be really helpful.",
          "votes": 1
        },
        {
          "id": 752516,
          "postDate": "2020-02-21T05:24:26.253Z",
          "content": "<p>Thank you for your proposal, but I want to merge with a team having a similar score to me with a single model.</p>",
          "rawMarkdown": "Thank you for your proposal, but I want to merge with a team having a similar score to me with a single model.",
          "votes": 1
        },
        {
          "id": 752523,
          "postDate": "2020-02-21T05:43:42.007Z",
          "content": "<p>We are working on an approach. Will ask again after a better score...</p>",
          "rawMarkdown": "We are working on an approach. Will ask again after a better score...",
          "votes": 1
        },
        {
          "id": 753617,
          "postDate": "2020-02-22T13:26:38.787Z",
          "content": "<p>how many images used for training when you mean 'frame by frame'?</p>",
          "rawMarkdown": "how many images used for training when you mean 'frame by frame'?"
        },
        {
          "id": 753695,
          "postDate": "2020-02-22T15:06:34.210Z",
          "content": "<p>We used somewhere between 1m-1.5m images <a href=\"/yangsaewon\">@yangsaewon</a> in 'frame by frame'</p>",
          "rawMarkdown": "We used somewhere between 1m-1.5m images @yangsaewon in 'frame by frame'",
          "votes": 3
        },
        {
          "id": 753723,
          "postDate": "2020-02-22T15:46:36.263Z",
          "content": "<p>So are you guys over-sampled the Real frames?</p>",
          "rawMarkdown": "So are you guys over-sampled the Real frames?"
        },
        {
          "id": 753724,
          "postDate": "2020-02-22T15:47:56.027Z",
          "content": "<p>No., The trick we are using is pretty simple, I am not sharing just yet, I am pretty sure many people are using that, saving it for sometime later in the competition.</p>",
          "rawMarkdown": "No., The trick we are using is pretty simple, I am not sharing just yet, I am pretty sure many people are using that, saving it for sometime later in the competition."
        },
        {
          "id": 758366,
          "postDate": "2020-02-27T17:22:32.420Z",
          "content": "<p><a href=\"/harshit\">@harshit</a> are you using external data published ?</p>",
          "rawMarkdown": "@harshit are you using external data published ?"
        },
        {
          "id": 758384,
          "postDate": "2020-02-27T17:34:00.880Z",
          "content": "<p>Did you compile the data yourself, if so how long did it take?</p>",
          "rawMarkdown": "Did you compile the data yourself, if so how long did it take?"
        },
        {
          "id": 763074,
          "postDate": "2020-03-04T05:34:27.270Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> As far as I understand, if you use the competition data then somehow preprocess it, you don't need to make public and declare.\n<a href=\"/greatgamedota\">@greatgamedota</a> Took us 1-2 days running 24 hours on 1080 TI.</p>",
          "rawMarkdown": "@jaideepvalani As far as I understand, if you use the competition data then somehow preprocess it, you don't need to make public and declare.\n@greatgamedota Took us 1-2 days running 24 hours on 1080 TI."
        },
        {
          "id": 763237,
          "postDate": "2020-03-04T09:23:33.987Z",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a>  how do you deal with data imbalance？</p>",
          "rawMarkdown": "@lyakaap  how do you deal with data imbalance？"
        }
      ]
    },
    {
      "id": 750624,
      "postDate": "2020-02-19T14:50:07.253Z",
      "content": "<p>0.32</p>",
      "rawMarkdown": "0.32",
      "votes": 7,
      "replies": [
        {
          "id": 750821,
          "postDate": "2020-02-19T18:02:19.010Z",
          "content": "<p>Amazing!, Any insights you can share, please?</p>",
          "rawMarkdown": "Amazing!, Any insights you can share, please?"
        },
        {
          "id": 750897,
          "postDate": "2020-02-19T19:46:36.403Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> Are you only using images, or also audio?</p>",
          "rawMarkdown": "@chenshen03 Are you only using images, or also audio?"
        },
        {
          "id": 751065,
          "postDate": "2020-02-20T01:06:07.783Z",
          "content": "<p>😃 </p>",
          "rawMarkdown": "😃 "
        },
        {
          "id": 751136,
          "postDate": "2020-02-20T02:42:04.557Z",
          "content": "<p>image only</p>",
          "rawMarkdown": "image only",
          "votes": 3
        },
        {
          "id": 751310,
          "postDate": "2020-02-20T05:10:17.513Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> Any insights?</p>",
          "rawMarkdown": "@chenshen03 Any insights?"
        },
        {
          "id": 752477,
          "postDate": "2020-02-21T04:04:27.207Z",
          "content": "<p>You can really dance!</p>",
          "rawMarkdown": "You can really dance!"
        },
        {
          "id": 754113,
          "postDate": "2020-02-23T04:27:35.967Z",
          "content": "<p>what's your face extractor?</p>",
          "rawMarkdown": "what's your face extractor?"
        },
        {
          "id": 754382,
          "postDate": "2020-02-23T13:38:44.677Z",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> blazeface</p>",
          "rawMarkdown": "@chenbaoying blazeface",
          "votes": 1
        },
        {
          "id": 754386,
          "postDate": "2020-02-23T13:43:59.480Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> There is a lot of noise in the train dataset. For example, some FAKE videos have multiple faces, but the faces may all be REAL. Removing these noises may improve performance.</p>",
          "rawMarkdown": "@harshitsheoran There is a lot of noise in the train dataset. For example, some FAKE videos have multiple faces, but the faces may all be REAL. Removing these noises may improve performance.",
          "votes": 8
        },
        {
          "id": 754400,
          "postDate": "2020-02-23T13:57:31.287Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> so you are removing noises in both training set and test set?</p>",
          "rawMarkdown": "@chenshen03 so you are removing noises in both training set and test set?",
          "votes": 1
        },
        {
          "id": 757177,
          "postDate": "2020-02-26T13:55:32.450Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> Are you using any external data?</p>",
          "rawMarkdown": "@chenshen03 Are you using any external data?",
          "votes": 1
        },
        {
          "id": 757243,
          "postDate": "2020-02-26T15:03:19.513Z",
          "content": "<p><a href=\"/dsfhe49854\">@dsfhe49854</a> No</p>",
          "rawMarkdown": "@dsfhe49854 No",
          "votes": 2
        },
        {
          "id": 759206,
          "postDate": "2020-02-28T18:12:16.847Z",
          "content": "<p>@shen chen .. how do we detect those noises in the face ?</p>",
          "rawMarkdown": "@shen chen .. how do we detect those noises in the face ?"
        },
        {
          "id": 759861,
          "postDate": "2020-02-29T14:54:43.967Z",
          "content": "<p>Does your current score  still use the single model?CNN or CNN+RNN</p>",
          "rawMarkdown": "Does your current score  still use the single model?CNN or CNN+RNN"
        },
        {
          "id": 759906,
          "postDate": "2020-02-29T15:59:35.060Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> the area of the face and the detection score</p>",
          "rawMarkdown": "@jaideepvalani the area of the face and the detection score\n\n",
          "votes": 2
        },
        {
          "id": 759908,
          "postDate": "2020-02-29T16:00:19.633Z",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> CNN + sequential model</p>",
          "rawMarkdown": "@chenbaoying CNN + sequential model"
        },
        {
          "id": 760042,
          "postDate": "2020-02-29T19:10:16.510Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> What is your input size?</p>",
          "rawMarkdown": "@chenshen03 What is your input size?"
        },
        {
          "id": 760190,
          "postDate": "2020-03-01T00:19:03.860Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> 224*224</p>",
          "rawMarkdown": "@unkownhihi 224*224"
        },
        {
          "id": 762753,
          "postDate": "2020-03-03T19:28:50.323Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> \nplease help understand below queries...\n1) There are multiple fake videos for a given original video ,can we assume that all the fake videos should have same  person or set of persons speaking as is in original video</p>\n\n<p>2) Can we say all Fake videos will have fake audios also  ?</p>\n\n<p>3)  is it possible that in pair of fake and original video both will have real faces but only different voices ? eg. in both the fakes i am the only person but just that in fake its some one elses' voice  </p>",
          "rawMarkdown": "@chenshen03 \nplease help understand below queries...\n1) There are multiple fake videos for a given original video ,can we assume that all the fake videos should have same  person or set of persons speaking as is in original video\n\n2) Can we say all Fake videos will have fake audios also  ?\n\n3)  is it possible that in pair of fake and original video both will have real faces but only different voices ? eg. in both the fakes i am the only person but just that in fake its some one elses' voice  \n"
        },
        {
          "id": 762755,
          "postDate": "2020-03-03T19:34:43.637Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \n1. It is true.\n2. No\n3. It is possible. </p>",
          "rawMarkdown": "@jaideepvalani \n1. It is true.\n2. No\n3. It is possible. "
        },
        {
          "id": 763024,
          "postDate": "2020-03-04T03:22:44.917Z",
          "content": "<p>Thanks <a href=\"/harshitsheoran\">@harshitsheoran</a> <br>\nSo 3rd case if there are good number then image based model could making an error in prediction ,any way to separate such videos ?\nAlso would u mind giving  brief  on the this model framework..  do we need only images..  here or videos ?</p>",
          "rawMarkdown": "Thanks @harshitsheoran  \nSo 3rd case if there are good number then image based model could making an error in prediction ,any way to separate such videos ?\nAlso would u mind giving  brief  on the this model framework..  do we need only images..  here or videos ?"
        },
        {
          "id": 763068,
          "postDate": "2020-03-04T05:27:51.687Z",
          "content": "<p>There is, if you train on it, but it does not worth the time at this stage.</p>",
          "rawMarkdown": "There is, if you train on it, but it does not worth the time at this stage."
        },
        {
          "id": 763258,
          "postDate": "2020-03-04T09:54:16.913Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> So far, we haven't found any benefits from the audio model,  but it's still worth trying.</p>",
          "rawMarkdown": "@jaideepvalani So far, we haven't found any benefits from the audio model,  but it's still worth trying."
        },
        {
          "id": 765758,
          "postDate": "2020-03-07T03:48:49.427Z",
          "content": "<p><a href=\"/chason\">@chason</a>\nm using detection score  only if score is more than this add it in list,\nhow about the area how to  find the relavant one based on this </p>",
          "rawMarkdown": "@chason\nm using detection score  only if score is more than this add it in list,\nhow about the area how to  find the relavant one based on this "
        },
        {
          "id": 766128,
          "postDate": "2020-03-07T17:37:30.077Z",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> What's your local validation score and how did you choose your validation set?</p>",
          "rawMarkdown": "@chenshen03 What's your local validation score and how did you choose your validation set?"
        }
      ]
    },
    {
      "id": 750044,
      "postDate": "2020-02-19T04:52:36.813Z",
      "content": "<p>“Need to think of a good model” :)</p>",
      "rawMarkdown": "“Need to think of a good model” :)",
      "votes": 3,
      "replies": [
        {
          "id": 750047,
          "postDate": "2020-02-19T04:57:29.837Z",
          "content": "<p>LOL I am trying to think of a better model too! There's a ton of good kernels and discussions of good models. I recommend you to look into some of them(will get you to top 10%).</p>",
          "rawMarkdown": "LOL I am trying to think of a better model too! There's a ton of good kernels and discussions of good models. I recommend you to look into some of them(will get you to top 10%).",
          "votes": 1
        },
        {
          "id": 750462,
          "postDate": "2020-02-19T12:03:59.650Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> could u point me to one such  please which u took as base.. thanks in advance </p>",
          "rawMarkdown": "@unkownhihi could u point me to one such  please which u took as base.. thanks in advance "
        },
        {
          "id": 750605,
          "postDate": "2020-02-19T14:41:16.980Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> ResNext/Xception is a good one to start with.</p>",
          "rawMarkdown": "@jaideepvalani ResNext/Xception is a good one to start with."
        },
        {
          "id": 750660,
          "postDate": "2020-02-19T15:27:37.560Z",
          "content": "<p>Yes, these models get me 0.372</p>",
          "rawMarkdown": "Yes, these models get me 0.372",
          "votes": 1
        },
        {
          "id": 750724,
          "postDate": "2020-02-19T16:29:31.373Z",
          "content": "<p>I actually use resnext and get a very good score on local CV ,balanced loss bt still very poor score at LB. \nWhat mistake i i could be making .. \nI use MTCNN in pre process pipeline.</p>",
          "rawMarkdown": "I actually use resnext and get a very good score on local CV ,balanced loss bt still very poor score at LB. \nWhat mistake i i could be making .. \nI use MTCNN in pre process pipeline."
        },
        {
          "id": 750822,
          "postDate": "2020-02-19T18:03:49.103Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Take your time in finding a better face detector, If ResNext does not work the best, use it with lstm, you should be able to figure rest of it yourself. ;-)</p>",
          "rawMarkdown": "@jaideepvalani Take your time in finding a better face detector, If ResNext does not work the best, use it with lstm, you should be able to figure rest of it yourself. ;-)"
        },
        {
          "id": 750842,
          "postDate": "2020-02-19T18:31:33.937Z",
          "content": "<p>Thanku Harshit for your advice... I heard about LSTM used as decoder to encoder models. But havent seen its implementation if you can point out to some kernel or link.. would be useful :) </p>",
          "rawMarkdown": "Thanku Harshit for your advice... I heard about LSTM used as decoder to encoder models. But havent seen its implementation if you can point out to some kernel or link.. would be useful :) "
        },
        {
          "id": 750875,
          "postDate": "2020-02-19T19:07:49.680Z",
          "content": "<p>In my opinion, try the normal model to get started. If you did things right, it should get you to the top 10%. We will publish lrcn(the LSTM model <a href=\"/harshitsheoran\">@harshitsheoran</a> was talking about) starter kernel soon(along with the better face extractor). </p>",
          "rawMarkdown": "In my opinion, try the normal model to get started. If you did things right, it should get you to the top 10%. We will publish lrcn(the LSTM model @harshitsheoran was talking about) starter kernel soon(along with the better face extractor). ",
          "votes": 5
        },
        {
          "id": 751565,
          "postDate": "2020-02-20T09:31:44.457Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> great work. I'm sort of stuck at this point. Maybe your kernel will help me.</p>",
          "rawMarkdown": "@unkownhihi great work. I'm sort of stuck at this point. Maybe your kernel will help me."
        },
        {
          "id": 752455,
          "postDate": "2020-02-21T03:15:37.507Z",
          "content": "<p>I went past two bench Mark's 467 and 451 ... at 404 now ,will try interesting things now shift from rank 590 to 80 ...long way to go  </p>",
          "rawMarkdown": "I went past two bench Mark's 467 and 451 ... at 404 now ,will try interesting things now shift from rank 590 to 80 ...long way to go  ",
          "votes": 1
        },
        {
          "id": 752804,
          "postDate": "2020-02-21T12:30:24.603Z",
          "content": "<p>Hi, I want to know did you get a score of 0.38 through lrcn?\nIs the lstm an effective approach?</p>",
          "rawMarkdown": "Hi, I want to know did you get a score of 0.38 through lrcn?\nIs the lstm an effective approach?"
        },
        {
          "id": 753016,
          "postDate": "2020-02-21T16:05:54.487Z",
          "content": "<p><a href=\"/yzcwansui\">@yzcwansui</a> yes it was an effective approach for us.</p>",
          "rawMarkdown": "@yzcwansui yes it was an effective approach for us."
        },
        {
          "id": 753032,
          "postDate": "2020-02-21T16:37:40.693Z",
          "content": "<p><a href=\"/yzcwansui\">@yzcwansui</a> although for some folks, single frame model is better. But the best score we can get for single frame model is just 0.44</p>",
          "rawMarkdown": "@yzcwansui although for some folks, single frame model is better. But the best score we can get for single frame model is just 0.44",
          "votes": 1
        },
        {
          "id": 768934,
          "postDate": "2020-03-11T11:14:57.350Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> <br>\nI try the resnet+lstm model, it performs bad in the test dataset while gets good score in the train dataset,  I use dropout=0.9 to prevent overfiting but it doesn't work, \ndo u have any advice? appreciate it first.\nwe train/test the model using 5 frames per video. and the num of real and fake is balanced.</p>",
          "rawMarkdown": "@harshitsheoran  \nI try the resnet+lstm model, it performs bad in the test dataset while gets good score in the train dataset,  I use dropout=0.9 to prevent overfiting but it doesn't work, \ndo u have any advice? appreciate it first.\nwe train/test the model using 5 frames per video. and the num of real and fake is balanced."
        }
      ]
    },
    {
      "id": 762742,
      "postDate": "2020-03-03T19:19:17.933Z",
      "content": "<p>New Single Frame by frame model Score 0.34000 used 160k images.</p>",
      "rawMarkdown": "New Single Frame by frame model Score 0.34000 used 160k images.",
      "votes": 4,
      "replies": [
        {
          "id": 763253,
          "postDate": "2020-03-04T09:46:37.970Z",
          "content": "<p>Wow, how do you deal with data imbalance？and u are using RNN model ?</p>",
          "rawMarkdown": "Wow, how do you deal with data imbalance？and u are using RNN model ?"
        },
        {
          "id": 763490,
          "postDate": "2020-03-04T14:35:53.107Z",
          "content": "<p>amazing score, \nwhat's your input size?</p>",
          "rawMarkdown": "amazing score, \nwhat's your input size?"
        },
        {
          "id": 763519,
          "postDate": "2020-03-04T15:11:31.457Z",
          "content": "<p>input size 200x200, no rnn, but working on that, yes dealt with data imbalance, but doing the same thing in an LRCN has some drawbacks. Ensemble of 3 models between 0.34-0.35 (all single model), was 0.318, So, I guess, ensembling right now is much less effective than it was back then.</p>",
          "rawMarkdown": "input size 200x200, no rnn, but working on that, yes dealt with data imbalance, but doing the same thing in an LRCN has some drawbacks. Ensemble of 3 models between 0.34-0.35 (all single model), was 0.318, So, I guess, ensembling right now is much less effective than it was back then.",
          "votes": 4
        },
        {
          "id": 763641,
          "postDate": "2020-03-04T17:27:59.217Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a>  could u point me to  any good public lrcn model available in git hub ?\nwant to try this stuff .. will be newest of its kind for me :)\nALso if you could brief me about the overall framework of this combo. \nDo we need videos to train this model or images sufficient ?</p>",
          "rawMarkdown": "@harshitsheoran  could u point me to  any good public lrcn model available in git hub ?\nwant to try this stuff .. will be newest of its kind for me :)\nALso if you could brief me about the overall framework of this combo. \nDo we need videos to train this model or images sufficient ?"
        },
        {
          "id": 763667,
          "postDate": "2020-03-04T18:04:45.680Z",
          "content": "<p>an LRCN, have a input in this format (batch_size, N_samples, size, size, n_channels), here N_samples can be your continuous or equally gapped number of images of a same video or different, that is up to you, I dont know how to do timedistributed layers in pytorch but in keras you can search timedistributed on google, to know more about it, after that you can wrap up your pretrained model in a timedistributed layer and then use that cnn model with a rnn.</p>",
          "rawMarkdown": "an LRCN, have a input in this format (batch_size, N_samples, size, size, n_channels), here N_samples can be your continuous or equally gapped number of images of a same video or different, that is up to you, I dont know how to do timedistributed layers in pytorch but in keras you can search timedistributed on google, to know more about it, after that you can wrap up your pretrained model in a timedistributed layer and then use that cnn model with a rnn.",
          "votes": 2
        },
        {
          "id": 763676,
          "postDate": "2020-03-04T18:16:50.173Z",
          "content": "<p>It's actually fairly easy, you can switch from frames-by-video to a linear series of frames (for the CNN) and then back to videos, like this:</p>\n\n<p><code>\n    def forward(self, x):\n        batch_size, n_channels, n_frames, height, width = x.size()  # e.g. 32 x 3 x 64 x 224 x 224\n        assert n_channels == 3, f\"Expecting 3 channels but got {n_channels}\"\n        x = x.permute(0,2,1,3,4)  # 32 x 64 x 3 x 224 x 224\n        x = x.reshape(batch_size*n_frames, n_channels, height, width)  # (32*64) x 3 x 224 x 224\n        x = self.cnn_model(x)  # (32*64) x 2048\n        x = x.view(batch_size, n_frames, self.n_cnn_features)\n        x, (h_n, h_c) = self.lstm(x)   # x = seq_len, batch, num_directions*hidden_size\n        x = x[:,-1]  # Take last time step, which will be BATCH_SIZE * (2*HIDDEN) -&amp;gt; e.g. 12 * 128\n        x = self.dropout(x)\n        x = self.linear(x)\n        return x\n</code></p>",
          "rawMarkdown": "It's actually fairly easy, you can switch from frames-by-video to a linear series of frames (for the CNN) and then back to videos, like this:\n\n```\n    def forward(self, x):\n        batch_size, n_channels, n_frames, height, width = x.size()  # e.g. 32 x 3 x 64 x 224 x 224\n        assert n_channels == 3, f\"Expecting 3 channels but got {n_channels}\"\n        x = x.permute(0,2,1,3,4)  # 32 x 64 x 3 x 224 x 224\n        x = x.reshape(batch_size*n_frames, n_channels, height, width)  # (32*64) x 3 x 224 x 224\n        x = self.cnn_model(x)  # (32*64) x 2048\n        x = x.view(batch_size, n_frames, self.n_cnn_features)\n        x, (h_n, h_c) = self.lstm(x)   # x = seq_len, batch, num_directions*hidden_size\n        x = x[:,-1]  # Take last time step, which will be BATCH_SIZE * (2*HIDDEN) -&gt; e.g. 12 * 128\n        x = self.dropout(x)\n        x = self.linear(x)\n        return x\n```",
          "votes": 17
        },
        {
          "id": 763710,
          "postDate": "2020-03-04T19:08:38.173Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 764388,
          "postDate": "2020-03-05T12:11:47.770Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Dose \"Frame by frame model\" mean that you take each frame as input rather than the cropped face?</p>",
          "rawMarkdown": "@harshitsheoran Dose \"Frame by frame model\" mean that you take each frame as input rather than the cropped face?"
        },
        {
          "id": 764563,
          "postDate": "2020-03-05T15:55:13.590Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> \nThanku Dr. james \nIf i understand correctly this how it would be \n1 ) CNN first layer  nn.Conv2d(3,n ....) \n2) Input  size ( 32*64=2048,3,224,224)..i wonder how we can fit such high batch size  ?\n3) CNN Output \n4) resize  output to  32 ,64 frames * 2048 for LSTM </p>\n\n<p>5) output of LSTM passed to linear layer\n6) sigmoid  bce loss</p>",
          "rawMarkdown": "@jamesphoward \nThanku Dr. james \nIf i understand correctly this how it would be \n1 ) CNN first layer  nn.Conv2d(3,n ....) \n2) Input  size ( 32*64=2048,3,224,224)..i wonder how we can fit such high batch size  ?\n3) CNN Output \n4) resize  output to  32 ,64 frames * 2048 for LSTM \n\n5) output of LSTM passed to linear layer\n6) sigmoid  bce loss\n"
        },
        {
          "id": 764657,
          "postDate": "2020-03-05T17:46:29.493Z",
          "content": "<p>Frame by frame have data input like this, (batch_size, face_size, face_size, n_channels), the face are cropped from the frame but they are not fed into an lstm or any transformer, it makes it a simple binary classfication problem, but here also we find that a simple LRCN performs much better than our frame-by-frame model, it looks like we are not even close to done improving. Good Luck!</p>",
          "rawMarkdown": "Frame by frame have data input like this, (batch_size, face_size, face_size, n_channels), the face are cropped from the frame but they are not fed into an lstm or any transformer, it makes it a simple binary classfication problem, but here also we find that a simple LRCN performs much better than our frame-by-frame model, it looks like we are not even close to done improving. Good Luck!",
          "votes": 2
        },
        {
          "id": 764902,
          "postDate": "2020-03-06T03:54:34.483Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> I did this few weeks ago, but face memory problem. How did you fit 32x64x224x224x3 for each batch? It's quite a lot. Using accumulated gradient seems to be faulty as BatchNorm won't work if we fit 1x64x224x224x3 each time.</p>",
          "rawMarkdown": "@jamesphoward I did this few weeks ago, but face memory problem. How did you fit 32x64x224x224x3 for each batch? It's quite a lot. Using accumulated gradient seems to be faulty as BatchNorm won't work if we fit 1x64x224x224x3 each time."
        },
        {
          "id": 764980,
          "postDate": "2020-03-06T06:43:50.570Z",
          "content": "<p>Can you clarify what you mean by 160k images? <a href=\"/harshitsheoran\">@harshitsheoran</a> Given that we have &gt;100k videos does that mean youre only looking at a couple frames per video? Or are you using some method to choose to train on a subset of videos?</p>",
          "rawMarkdown": "Can you clarify what you mean by 160k images? @harshitsheoran Given that we have &gt;100k videos does that mean youre only looking at a couple frames per video? Or are you using some method to choose to train on a subset of videos?"
        },
        {
          "id": 764982,
          "postDate": "2020-03-06T06:49:03.763Z",
          "content": "<p><a href=\"/ryches\">@ryches</a> good question, first of all I am only using first 40 folders, yes, that does mean, that on an average we do have only a few frames from each video but honestly they are selected so perfectly anything more would be nothing but an overkill and the chances were that it would 1. not letting the model any advantage 2. making the model starting overfitting faster than its normal epoch.</p>",
          "rawMarkdown": "@ryches good question, first of all I am only using first 40 folders, yes, that does mean, that on an average we do have only a few frames from each video but honestly they are selected so perfectly anything more would be nothing but an overkill and the chances were that it would 1. not letting the model any advantage 2. making the model starting overfitting faster than its normal epoch.",
          "votes": 1
        },
        {
          "id": 765129,
          "postDate": "2020-03-06T09:29:01.957Z",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> It was just an example, what batch size of videos and how many frames per video you use depend on how complex the CNN is, how many parameters you use for the RNN, your GPU, and whether you freeze the CNN before embedding it. I advise the latter, at least for earlier stages of training.</p>",
          "rawMarkdown": "@khahuras It was just an example, what batch size of videos and how many frames per video you use depend on how complex the CNN is, how many parameters you use for the RNN, your GPU, and whether you freeze the CNN before embedding it. I advise the latter, at least for earlier stages of training.",
          "votes": 2
        },
        {
          "id": 765135,
          "postDate": "2020-03-06T09:34:20.720Z",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> \nwhy did we take last time step in your sample code above ?</p>",
          "rawMarkdown": "@jamesphoward \nwhy did we take last time step in your sample code above ?"
        },
        {
          "id": 765138,
          "postDate": "2020-03-06T09:39:36.473Z",
          "content": "<p>Because of the way RNNs work in pytorch you get an output (prediction) after every time step (i.e. every frame of the video). However, you probably just want the RNN's prediction once it's seen the whole video, so I take the last time step. However, maybe that's not the best method, who knows?</p>",
          "rawMarkdown": "Because of the way RNNs work in pytorch you get an output (prediction) after every time step (i.e. every frame of the video). However, you probably just want the RNN's prediction once it's seen the whole video, so I take the last time step. However, maybe that's not the best method, who knows?"
        },
        {
          "id": 765269,
          "postDate": "2020-03-06T12:41:42.873Z",
          "content": "<p>What is the size of your self.linear ?\nsince output of lstm would be for eg. 12 * 128 </p>",
          "rawMarkdown": "What is the size of your self.linear ?\nsince output of lstm would be for eg. 12 * 128 "
        },
        {
          "id": 765298,
          "postDate": "2020-03-06T13:21:51.317Z",
          "content": "<p>The input size would be the the number of hidden units, or double this if your LSTM/GRU was bidirectional.</p>",
          "rawMarkdown": "The input size would be the the number of hidden units, or double this if your LSTM/GRU was bidirectional."
        },
        {
          "id": 765320,
          "postDate": "2020-03-06T13:42:15.960Z",
          "content": "<p>Sorry to bother you more..\nThis is my input \n<code>\nself.LSTM=nn.LSTM(1024,512,2) \nself.linear=nn.linear(512,1)\nx torch.Size([192, 1024])-&amp;gt; reshaped to  16,12,1024  for nn.LSTM\n</code>\nif i do x=x[:,-1]  my original batch size of 192 gets reduced to 12 .\nIs there an additional layer needed before feeding to Loss function as  my  output of linear wont meet requirement of loss function which is 192*1</p>",
          "rawMarkdown": "Sorry to bother you more..\nThis is my input \n```\nself.LSTM=nn.LSTM(1024,512,2) \nself.linear=nn.linear(512,1)\nx torch.Size([192, 1024])-&gt; reshaped to  16,12,1024  for nn.LSTM\n```\nif i do x=x[:,-1]  my original batch size of 192 gets reduced to 12 .\nIs there an additional layer needed before feeding to Loss function as  my  output of linear wont meet requirement of loss function which is 192*1\n \n"
        },
        {
          "id": 765354,
          "postDate": "2020-03-06T14:21:19.873Z",
          "content": "<p>Have you set <code>batch_first=True</code> in your LSTM if you are supplying the batch first?</p>",
          "rawMarkdown": "Have you set `batch_first=True` in your LSTM if you are supplying the batch first?"
        },
        {
          "id": 765368,
          "postDate": "2020-03-06T14:39:16.267Z",
          "content": "<p>Yup i did as first part.. \n```\nclass lstm(nn.Module):\n    def <strong>init</strong>(self,bs,num,  lstm1_input=1024,lstm_hidden=512):\n        super().<strong>init</strong>()\n        self.batch=bs #16\n        self.num=num #12</p>\n\n<pre><code>    self.lstm = nn.LSTM(1024, 512 ,2,batch_first=True)# input elements,hidden state size\n    self.linear=nn.Linear(512,1)\n    self.drop=nn.Dropout(0.5)\ndef forward(self,input):\n    #print(input.size())\n    x=input.view(self.batch,self.num,input.size(1)) # resized from 192 to 16,12\n    x,y=self.lstm(x)\n    print(x.size()) #see below\n    x = x[:,-1,:]\n    print(x.size()) #see below\n    x=self.drop(x)\n    x=self.linear(x)\n    return x\n</code></pre>\n\n<p>```</p>\n\n<p>torch.Size([16, 12, 512])\ntorch.Size([16, 512])</p>",
          "rawMarkdown": "Yup i did as first part.. \n```\nclass lstm(nn.Module):\n    def __init__(self,bs,num,  lstm1_input=1024,lstm_hidden=512):\n        super().__init__()\n        self.batch=bs #16\n        self.num=num #12\n        \n        \n        self.lstm = nn.LSTM(1024, 512 ,2,batch_first=True)# input elements,hidden state size\n        self.linear=nn.Linear(512,1)\n        self.drop=nn.Dropout(0.5)\n    def forward(self,input):\n        #print(input.size())\n        x=input.view(self.batch,self.num,input.size(1)) # resized from 192 to 16,12\n        x,y=self.lstm(x)\n        print(x.size()) #see below\n        x = x[:,-1,:]\n        print(x.size()) #see below\n        x=self.drop(x)\n        x=self.linear(x)\n        return x\n```\n \ntorch.Size([16, 12, 512])\ntorch.Size([16, 512])"
        },
        {
          "id": 765378,
          "postDate": "2020-03-06T14:47:03.843Z",
          "content": "<p>Isn't that correct? You have fed a batch of 16 into your network and your linear layer is receiving 16 sets of 512 hidden units' outputs? I don't see the problem...</p>",
          "rawMarkdown": "Isn't that correct? You have fed a batch of 16 into your network and your linear layer is receiving 16 sets of 512 hidden units' outputs? I don't see the problem..."
        },
        {
          "id": 765386,
          "postDate": "2020-03-06T15:04:11.990Z",
          "content": "<p>yes i dont have issue here but in  the loss calculation.. \n data loader will give labels with bs =batch * num of frames * num of classes\nlooks like i should generate 12 steps using some Time Distribute layer like in keras and pack them. while keeping the original bs as 16 only for data loader...\nis this wha</p>",
          "rawMarkdown": "yes i dont have issue here but in  the loss calculation.. \n data loader will give labels with bs =batch * num of frames * num of classes\nlooks like i should generate 12 steps using some Time Distribute layer like in keras and pack them. while keeping the original bs as 16 only for data loader...\nis this wha\n"
        },
        {
          "id": 765402,
          "postDate": "2020-03-06T15:26:44.247Z",
          "content": "<p>Hi, I think that's really a problem you are going to need to solve as this isn't really appropriate for the forums, and we're now getting on to fundamental issues of network training, sorry!</p>",
          "rawMarkdown": "Hi, I think that's really a problem you are going to need to solve as this isn't really appropriate for the forums, and we're now getting on to fundamental issues of network training, sorry!"
        },
        {
          "id": 768364,
          "postDate": "2020-03-10T17:37:36.477Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> \nI overcame all the bs related issue after chosing right way of reshaping and transposing the dims. \nNow i face another issue which is overfitting. Model done best than ever models i tried on CV but miserably overfits on test data,especially fails to predict reals correctly. \nDid u faced this issue with LSTM.. if yes what can be f help. i purposefully dint pay attn on split strategy.</p>\n\n<p><code>\nepoch   train_loss  valid_loss  accuracy    time\n0   0.450687    0.451665    0.834688    09:08\n1   0.375398    0.342386    0.859259    09:06\n2   0.261341           0.245293        0.896838         09:02\n3   0.199818             0.183108    0.928907   09:06 \n4   0.169864    0.163689    0.938302    09:07\n5   0.140049    0.210106           0.915447     09:10\n6   0.126233            0.136548    0.947064    09:11\n</code> </p>",
          "rawMarkdown": "@harshitsheoran \nI overcame all the bs related issue after chosing right way of reshaping and transposing the dims. \nNow i face another issue which is overfitting. Model done best than ever models i tried on CV but miserably overfits on test data,especially fails to predict reals correctly. \nDid u faced this issue with LSTM.. if yes what can be f help. i purposefully dint pay attn on split strategy.\n\n```\nepoch\ttrain_loss\tvalid_loss\taccuracy\ttime\n0\t0.450687\t0.451665\t0.834688\t09:08\n1\t0.375398\t0.342386\t0.859259\t09:06\n2\t0.261341\t       0.245293\t       0.896838\t        09:02\n3\t0.199818\t         0.183108\t 0.928907\t09:06 \n4\t0.169864\t0.163689\t0.938302\t09:07\n5\t0.140049\t0.210106\t       0.915447  \t09:10\n6\t0.126233\t        0.136548\t0.947064\t09:11\n``` "
        },
        {
          "id": 768522,
          "postDate": "2020-03-10T23:13:40.947Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> well, that loss is kinda close to ours, with accuracy also close to ours every epoch, (our 1 epoch takes 3 hours so idk how you are overfitting and I am not, I also have a much much much larger dataset which I train on compared to yours) I never had overfitting issues and please if you are using keras dont focus on valid_loss it lies, I guess you should have a benchmark, of last 10 folders, after every epoch check your model on that, and see which epoch is starts overfitting, in our case</p>\n\n<p>model__e1 -&gt; 0.212 / 0.934\nmodel_e2 -&gt; 0.183 / 0.946\nmodel_e3 -&gt; 0.175 / 0.952\nmodel_e4 -&gt; 0.186 / 0.950</p>\n\n<p>even something like this happens where you think overfitting is starting, still train it might get a better improvement in next 3 epochs.</p>\n\n<p>try dropout, and batch norm (however I dont use any of them) maybe to overcome overfitting, also no noise function in the augmentations, check the labels and your generator again manually. We had some problem and that was cuz our generator was messed up once.</p>",
          "rawMarkdown": "@jaideepvalani well, that loss is kinda close to ours, with accuracy also close to ours every epoch, (our 1 epoch takes 3 hours so idk how you are overfitting and I am not, I also have a much much much larger dataset which I train on compared to yours) I never had overfitting issues and please if you are using keras dont focus on valid_loss it lies, I guess you should have a benchmark, of last 10 folders, after every epoch check your model on that, and see which epoch is starts overfitting, in our case\n\nmodel__e1 -&gt; 0.212 / 0.934\nmodel_e2 -&gt; 0.183 / 0.946\nmodel_e3 -&gt; 0.175 / 0.952\nmodel_e4 -&gt; 0.186 / 0.950\n\neven something like this happens where you think overfitting is starting, still train it might get a better improvement in next 3 epochs.\n\ntry dropout, and batch norm (however I dont use any of them) maybe to overcome overfitting, also no noise function in the augmentations, check the labels and your generator again manually. We had some problem and that was cuz our generator was messed up once.",
          "votes": 3
        },
        {
          "id": 768526,
          "postDate": "2020-03-10T23:32:20.117Z",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> what is your bce val score each epoch?</p>",
          "rawMarkdown": "@harshitsheoran what is your bce val score each epoch?"
        },
        {
          "id": 768565,
          "postDate": "2020-03-11T01:16:58.600Z",
          "content": "<p><a href=\"/khahuras\">@khahuras</a>  that is the validation loss I shared and not the training loss</p>\n\n<p>0.160 on that bench means 0.318 on leaderboard, always almost double.</p>",
          "rawMarkdown": "@khahuras  that is the validation loss I shared and not the training loss\n\n0.160 on that bench means 0.318 on leaderboard, always almost double.",
          "votes": 2
        },
        {
          "id": 768642,
          "postDate": "2020-03-11T03:43:04.843Z",
          "content": "<p>1) how does your hist gram of 400 sub videos look like  assuming fake =0.99 and real =0.... \nleft and right height bars ,which one more ?\n2) bi direction LSTM or uni ?\n3) what are you lstm dimensions..   bs,num_frames ,no of elements  does keeping higher bs helps more ?\n4) how is your batch distribution  real:fake ratio in batch</p>",
          "rawMarkdown": "1) how does your hist gram of 400 sub videos look like  assuming fake =0.99 and real =0.... \nleft and right height bars ,which one more ?\n2) bi direction LSTM or uni ?\n3) what are you lstm dimensions..   bs,num_frames ,no of elements  does keeping higher bs helps more ?\n4) how is your batch distribution  real:fake ratio in batch\n\n"
        },
        {
          "id": 768661,
          "postDate": "2020-03-11T04:24:13.177Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> here's my response(I'm basically in charge of model training right now):\n1. We used a post-processing technique. Here's the histogram<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fd8ce84bed0e9d697476308f24ae354a7%2FScreen%20Shot%202020-03-10%20at%209.20.51%20PM.png?generation=1583900473994383&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>We used an uni.</li>\n<li>batch size is 6, the rest, I can't tell you. Higher batch size does sometimes help a little bit in experiment, but we consider it is due to random, because it is higher like 0.001, and sometimes, its actually worse.</li>\n</ol>",
          "rawMarkdown": "@jaideepvalani here's my response(I'm basically in charge of model training right now):\n1. We used a post-processing technique. Here's the histogram![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fd8ce84bed0e9d697476308f24ae354a7%2FScreen%20Shot%202020-03-10%20at%209.20.51%20PM.png?generation=1583900473994383&amp;alt=media)\n\n2. We used an uni.\n3. batch size is 6, the rest, I can't tell you. Higher batch size does sometimes help a little bit in experiment, but we consider it is due to random, because it is higher like 0.001, and sometimes, its actually worse.",
          "votes": 3
        },
        {
          "id": 769048,
          "postDate": "2020-03-11T13:23:26.830Z",
          "content": "<p>thanku.. during inference how many video frames used...\nis it necessary to keep same size as in training ?</p>",
          "rawMarkdown": "thanku.. during inference how many video frames used...\nis it necessary to keep same size as in training ?"
        },
        {
          "id": 769132,
          "postDate": "2020-03-11T15:02:04.417Z",
          "content": "<p>The answer might surprise you. We only used not a constant amount of frame per video.... I don't get the second question, sry.</p>",
          "rawMarkdown": "The answer might surprise you. We only used not a constant amount of frame per video.... I don't get the second question, sry."
        },
        {
          "id": 769193,
          "postDate": "2020-03-11T16:15:01.213Z",
          "content": "<p>i meant the batch size and num of frames  during inference should they be  same as in training for lstm to give expected output as in training</p>",
          "rawMarkdown": "i meant the batch size and num of frames  during inference should they be  same as in training for lstm to give expected output as in training"
        },
        {
          "id": 769201,
          "postDate": "2020-03-11T16:23:32.310Z",
          "content": "<p>ok. I got you now. the num of frames during inference is 10(they don't really affect much). (secret hint: we don't pass that 10 frames to lrcn at once, I can't really go into details.). for batch size during inference, its 1.</p>",
          "rawMarkdown": "ok. I got you now. the num of frames during inference is 10(they don't really affect much). (secret hint: we don't pass that 10 frames to lrcn at once, I can't really go into details.). for batch size during inference, its 1.",
          "votes": 1
        },
        {
          "id": 770631,
          "postDate": "2020-03-13T07:14:12.937Z",
          "content": "<p>Thanks.. is public score calculated based on the 400 test videos or it could  be some other videos which are subset of private one.</p>",
          "rawMarkdown": "Thanks.. is public score calculated based on the 400 test videos or it could  be some other videos which are subset of private one."
        },
        {
          "id": 770917,
          "postDate": "2020-03-13T15:09:45.687Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> I think the public score is calculated based on the 4000 videos that will be tested in a black box environment.</p>",
          "rawMarkdown": "@jaideepvalani I think the public score is calculated based on the 4000 videos that will be tested in a black box environment."
        }
      ]
    },
    {
      "id": 756584,
      "postDate": "2020-02-25T21:33:50.670Z",
      "content": "<p>My latest model scores: 0.3422</p>",
      "rawMarkdown": "My latest model scores: 0.3422",
      "votes": 4,
      "replies": [
        {
          "id": 758365,
          "postDate": "2020-02-27T17:20:11.153Z",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a>  amazing score\nwhich model did u use ?\nHave u used ur own dataset,if yes which face detector and no of frames used ?</p>",
          "rawMarkdown": "@ngcferreira  amazing score\nwhich model did u use ?\nHave u used ur own dataset,if yes which face detector and no of frames used ?"
        },
        {
          "id": 758416,
          "postDate": "2020-02-27T18:19:41.723Z",
          "content": "<p>I've use a CNN based model, using only one frame.\nI've used no external data and used yolov3 for detecting faces and create the training set and blazeface for  inference. \nThe prediction is just an average of 40 frames.</p>",
          "rawMarkdown": "I've use a CNN based model, using only one frame.\nI've used no external data and used yolov3 for detecting faces and create the training set and blazeface for  inference. \nThe prediction is just an average of 40 frames.",
          "votes": 4
        },
        {
          "id": 758419,
          "postDate": "2020-02-27T18:22:00.390Z",
          "content": "<p>what is your input face size?</p>",
          "rawMarkdown": "what is your input face size?"
        },
        {
          "id": 759069,
          "postDate": "2020-02-28T14:27:16.680Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> I'm using face images scaled to 224x224</p>",
          "rawMarkdown": "@unkownhihi I'm using face images scaled to 224x224",
          "votes": 1
        },
        {
          "id": 763234,
          "postDate": "2020-03-04T09:21:03.777Z",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a>  how do you deal with data imbalance？</p>",
          "rawMarkdown": "@ngcferreira  how do you deal with data imbalance？"
        },
        {
          "id": 766231,
          "postDate": "2020-03-07T22:00:21.667Z",
          "content": "<p><a href=\"/xmaocai\">@xmaocai</a>  So first I split the data into 3 sets (train, val and test set). In every epoch, I randomly choose the same number of fake images as the number of real images. Basically I'm under sampling the data.</p>",
          "rawMarkdown": "@xmaocai  So first I split the data into 3 sets (train, val and test set). In every epoch, I randomly choose the same number of fake images as the number of real images. Basically I'm under sampling the data.",
          "votes": 1
        }
      ]
    },
    {
      "id": 765464,
      "postDate": "2020-03-06T16:48:22.063Z",
      "content": "<p>CNN+LSTM ,You can refer to <a href=\"https://github.com/HHTseng/video-classification\">https://github.com/HHTseng/video-classification</a></p>",
      "rawMarkdown": "CNN+LSTM ,You can refer to https://github.com/HHTseng/video-classification",
      "votes": 1,
      "replies": [
        {
          "id": 768635,
          "postDate": "2020-03-11T03:32:03.570Z",
          "content": "<p>Hello, thanks for u advice. I try it but the model doesn't perform well in the val dataset while gets good acc in the train dataset. I use dropout=0.9 to prevent overfiting but it doesn't work, \ndo u have any advice? \nbtw, 5 frames per video.</p>",
          "rawMarkdown": "Hello, thanks for u advice. I try it but the model doesn't perform well in the val dataset while gets good acc in the train dataset. I use dropout=0.9 to prevent overfiting but it doesn't work, \ndo u have any advice? \nbtw, 5 frames per video."
        },
        {
          "id": 769109,
          "postDate": "2020-03-11T14:38:15.880Z",
          "content": "<p>@xmacai what is the dim of LSTM and input X to CNN.. \nmay u could be mistaking  at any  place</p>",
          "rawMarkdown": "@xmacai what is the dim of LSTM and input X to CNN.. \nmay u could be mistaking  at any  place"
        },
        {
          "id": 770798,
          "postDate": "2020-03-13T12:27:33.170Z",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \ninput X dim: torch.Size([64, 1, 3, 224, 224])\nthe dim of input to LSTM: torch.Size([8, 1, 512])\nLSTM dim: \n<code>self.LSTM =nn.LSTM(\n            input_size=300,\n            hidden_size=256,\n            num_layers=3,\n            batch_first=True,)</code>\nHere are the input data handle:\n<code>for index in data:</code>\n            <code>image = Image.open(index).convert('RGB')</code>\n            <code>if self.data_transforms is not None:</code>\n                <code>try:</code>\n                    <code>image = self.data_transforms(image)</code>\n                <code>except:</code>\n                    <code>print(\"Cannot transform image: {}\".format(index))</code>\n            <code>X.append(image)</code>\n<code>X = torch.stack(X, dim=0)</code>\n<code>Y = torch.LongTensor([label])</code></p>",
          "rawMarkdown": "@jaideepvalani \ninput X dim: torch.Size([64, 1, 3, 224, 224])\nthe dim of input to LSTM: torch.Size([8, 1, 512])\nLSTM dim: \n`self.LSTM =nn.LSTM(\n            input_size=300,\n            hidden_size=256,\n            num_layers=3,\n            batch_first=True,)`\nHere are the input data handle:\n`for index in data: `\n            `image = Image.open(index).convert('RGB')`\n            `if self.data_transforms is not None:`\n                `try:`\n                    `image = self.data_transforms(image)`\n                `except:`\n                    `print(\"Cannot transform image: {}\".format(index))`\n            `X.append(image)`\n`X = torch.stack(X, dim=0)`\n`Y = torch.LongTensor([label])`\n"
        }
      ]
    },
    {
      "id": 757639,
      "postDate": "2020-02-27T01:22:17.800Z",
      "content": "<p>current best score is 0.40621 after trying for weeks :)</p>",
      "rawMarkdown": "current best score is 0.40621 after trying for weeks :)",
      "votes": 1
    },
    {
      "id": 753355,
      "postDate": "2020-02-22T04:17:41.253Z",
      "content": "<p>Well, I have my train split as 0-29 and 40 to 49 and val split as 30-39, the same as <a href=\"/jamesphoward\">@jamesphoward</a> . Currently, I'm getting a validation score of 0.295 for \"Image Only\" model. But I have 0 trust in it! Or maybe I have a bug in my inference pipeline. Who knows!\nAlso, the fact that the model starts overfitting from 3rd epoch onwards gives me the chills.</p>",
      "rawMarkdown": "Well, I have my train split as 0-29 and 40 to 49 and val split as 30-39, the same as @jamesphoward . Currently, I'm getting a validation score of 0.295 for \"Image Only\" model. But I have 0 trust in it! Or maybe I have a bug in my inference pipeline. Who knows!\nAlso, the fact that the model starts overfitting from 3rd epoch onwards gives me the chills.",
      "votes": 1
    },
    {
      "id": 750291,
      "postDate": "2020-02-19T09:07:17.483Z",
      "content": "<p>My best model (images only) scores: 0.40457</p>",
      "rawMarkdown": "My best model (images only) scores: 0.40457\n",
      "votes": 1
    },
    {
      "id": 750589,
      "postDate": "2020-02-19T14:34:03.210Z",
      "content": "<p>0.489 :P</p>\n\n<p>Update 3: .402\nUpdate 4: .392</p>",
      "rawMarkdown": "0.489 :P\n\nUpdate 3: .402\nUpdate 4: .392",
      "votes": 2,
      "replies": [
        {
          "id": 765336,
          "postDate": "2020-03-06T14:07:27.887Z",
          "content": "<p>0.43* All I did was make a different inference kernal (different face detector) and pick a specific model trained on the data from my baseline kernal (one I trained over 2 weeks ago).</p>",
          "rawMarkdown": "0.43* All I did was make a different inference kernal (different face detector) and pick a specific model trained on the data from my baseline kernal (one I trained over 2 weeks ago).",
          "votes": 1
        }
      ]
    },
    {
      "id": 785367,
      "postDate": "2020-03-25T02:14:25.803Z",
      "content": "<p>what kind of scheduler do you guys use for training? fixed lr or cosine scheduling, etc..</p>",
      "rawMarkdown": "what kind of scheduler do you guys use for training? fixed lr or cosine scheduling, etc.."
    },
    {
      "id": 769950,
      "postDate": "2020-03-12T12:39:41.487Z",
      "content": "<p>0.44202.</p>",
      "rawMarkdown": "0.44202.",
      "replies": [
        {
          "id": 776564,
          "postDate": "2020-03-17T13:28:24.857Z",
          "content": "<p>0.42352  .. Just near :) </p>",
          "rawMarkdown": "0.42352  .. Just near :) ",
          "votes": 1
        }
      ]
    },
    {
      "id": 755733,
      "postDate": "2020-02-25T04:23:27.997Z",
      "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a>  hi，when using lrcn，how many frames of  one video do you use？are they sequential frames  or not。 appreciate for your advises</p>",
      "rawMarkdown": "@harshitsheoran  hi，when using lrcn，how many frames of  one video do you use？are they sequential frames  or not。 appreciate for your advises",
      "replies": [
        {
          "id": 757654,
          "postDate": "2020-02-27T01:58:20.820Z",
          "content": "<p>10 frames per video, you can increase or decrease does not really matter, they are not sequential, they are equally gapped.</p>",
          "rawMarkdown": "10 frames per video, you can increase or decrease does not really matter, they are not sequential, they are equally gapped.",
          "votes": 2
        }
      ]
    },
    {
      "id": 752910,
      "postDate": "2020-02-21T14:37:16.007Z",
      "content": "<p>0.41136 with a single CNN model. I wonder, how are you dealing with inference on multi-faced videos? I mean, one face  is usually okay in each frame, another is distorted, rarely both of them. Should I pick the max prediction per frame or something? \nupd: 0.40463\nupd: 0.40128</p>",
      "rawMarkdown": "0.41136 with a single CNN model. I wonder, how are you dealing with inference on multi-faced videos? I mean, one face  is usually okay in each frame, another is distorted, rarely both of them. Should I pick the max prediction per frame or something? \nupd: 0.40463\nupd: 0.40128",
      "replies": [
        {
          "id": 752923,
          "postDate": "2020-02-21T14:46:02.813Z",
          "content": "<p>Excellent question. I've experimented with this and I've found the mean to be very slightly better than either the max or min.</p>\n\n<p>I used to think the maximum would be best, but that assumes you don't have false positives for face detection; false positives are likely to be unnatural and therefore fakes, and so it might be the mean is a little more resilient to this.</p>\n\n<p>I would highlight, though, that this is very likely to depend on your processing pipeline.</p>",
          "rawMarkdown": "Excellent question. I've experimented with this and I've found the mean to be very slightly better than either the max or min.\n\nI used to think the maximum would be best, but that assumes you don't have false positives for face detection; false positives are likely to be unnatural and therefore fakes, and so it might be the mean is a little more resilient to this.\n\nI would highlight, though, that this is very likely to depend on your processing pipeline.",
          "votes": 8
        },
        {
          "id": 752964,
          "postDate": "2020-02-21T15:25:10.303Z",
          "content": "<p>Distorted face found during detection are likely to be predicted as Fake. So if there is frame with distorted  in list of frames  then that would most likely contribute to false positive and hence with max prob compared to all. </p>",
          "rawMarkdown": "Distorted face found during detection are likely to be predicted as Fake. So if there is frame with distorted  in list of frames  then that would most likely contribute to false positive and hence with max prob compared to all. ",
          "votes": 1
        },
        {
          "id": 753088,
          "postDate": "2020-02-21T17:39:24.780Z",
          "content": "<p>Thanks for the insight. I guess I'll play with inference detector a little bit more. </p>",
          "rawMarkdown": "Thanks for the insight. I guess I'll play with inference detector a little bit more. "
        },
        {
          "id": 754121,
          "postDate": "2020-02-23T04:41:19.190Z",
          "content": "<p>what's your face extractor? I use MTCNN but it can't get a good LB</p>",
          "rawMarkdown": "what's your face extractor? I use MTCNN but it can't get a good LB"
        },
        {
          "id": 758870,
          "postDate": "2020-02-28T09:22:12.697Z",
          "content": "<p>I've played with blazeface, yolov2, yolov3, mtcnn, and yolov3 got me the best score of 0.40463 for a single model</p>",
          "rawMarkdown": "I've played with blazeface, yolov2, yolov3, mtcnn, and yolov3 got me the best score of 0.40463 for a single model",
          "votes": 1
        },
        {
          "id": 761607,
          "postDate": "2020-03-02T18:32:21.367Z",
          "content": "<p><a href=\"/defileroff\">@defileroff</a> did you find that there was much difference between the face extractors in terms of overall score? I haven't played around with this yet, wondering what kind of priority it should be relative to everything else :)</p>",
          "rawMarkdown": "@defileroff did you find that there was much difference between the face extractors in terms of overall score? I haven't played around with this yet, wondering what kind of priority it should be relative to everything else :)"
        },
        {
          "id": 764394,
          "postDate": "2020-03-05T12:19:56.800Z",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a>  Playing with different frames, faces, extractors for train/valid sets gave me a boost of 0.05, playing with different extractors at inference time gave me a boost of 0.01 </p>\n\n<p>I guess if your classifier is weak in terms of false-positive numbers, you'd need a better detector to avoid false-positive face detections. But if your classifier is good enough, you can take a less picky detector, that will detect everything that looks like a face, lol, and let the classifier deal with it.</p>",
          "rawMarkdown": "@fergusoci  Playing with different frames, faces, extractors for train/valid sets gave me a boost of 0.05, playing with different extractors at inference time gave me a boost of 0.01 \n\nI guess if your classifier is weak in terms of false-positive numbers, you'd need a better detector to avoid false-positive face detections. But if your classifier is good enough, you can take a less picky detector, that will detect everything that looks like a face, lol, and let the classifier deal with it.",
          "votes": 2
        },
        {
          "id": 774959,
          "postDate": "2020-03-16T05:33:29.230Z",
          "content": "<p><a href=\"/defileroff\">@defileroff</a> pre-trained yolov3 face detector or trained with dfdc dataset? </p>",
          "rawMarkdown": "@defileroff pre-trained yolov3 face detector or trained with dfdc dataset? "
        },
        {
          "id": 790715,
          "postDate": "2020-03-29T19:53:32.240Z",
          "content": "<p>Pretrained one, I don't even use multi-frames like others do, works fine to some extent.</p>",
          "rawMarkdown": "Pretrained one, I don't even use multi-frames like others do, works fine to some extent."
        }
      ]
    },
    {
      "id": 752227,
      "postDate": "2020-02-20T19:52:05.620Z",
      "content": "<p>good</p>",
      "rawMarkdown": "good"
    },
    {
      "id": 765463,
      "postDate": "2020-03-06T16:46:12.143Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 765277,
      "postDate": "2020-03-06T12:56:25.547Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 759856,
      "postDate": "2020-02-29T14:52:08.033Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 759946,
          "postDate": "2020-02-29T16:39:44.357Z",
          "content": "<p>We are using a ensemble.</p>",
          "rawMarkdown": "We are using a ensemble."
        },
        {
          "id": 763356,
          "postDate": "2020-03-04T11:57:23.233Z",
          "content": "<p>thanks, I ensamble two model also get a better score</p>",
          "rawMarkdown": "thanks, I ensamble two model also get a better score"
        }
      ]
    },
    {
      "id": 757172,
      "postDate": "2020-02-26T13:52:53.370Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 751313,
      "author_name": "lyakaap",
      "author_url": "",
      "post_date": "2020-02-20T05:11:48.030000",
      "content": "<p>Frame-by-frame model (only image): 0.33966 </p>",
      "votes": 9,
      "replies": [
        {
          "id": 751314,
          "author_name": "Bibek",
          "author_url": "",
          "post_date": "2020-02-20T05:16:51.860000",
          "content": "<p>what do you mean by frame-by-frame model?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751352,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2020-02-20T06:35:50.630000",
          "content": "<p>It means a model makes a prediction to each frame (and then aggregates over frames).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 751516,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-02-20T08:49:28.887000",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a> Great score for a frame only model. How did you choose your validation set? And are you using any external data during training? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751646,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2020-02-20T11:14:06.047000",
          "content": "<p>Just selecting 20% of folder as validation set. Used no external data.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 751652,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-20T11:17:56.207000",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a> You mean 20% from each folder, was your validation able to track your leaderboard?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751964,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2020-02-20T17:23:00.673000",
          "content": "<p>No, split by folder index. This strategy can track well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752019,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-20T17:50:02.343000",
          "content": "<p>what is folder index?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752101,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2020-02-20T18:39:46.010000",
          "content": "<p>sorry for confusing, what I wanted to say is I split val/train based on the index of input zip files (dfdc_train_part_XX).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 752150,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-20T19:12:24.757000",
          "content": "<p>I am also doing the same but in particular last 10 dfdc_train_part_XX where XX is range(40,50), is our split same or you are randomly selecting those? </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 752474,
          "author_name": "Daic",
          "author_url": "",
          "post_date": "2020-02-21T04:03:33.333000",
          "content": "<p>Linear division may works.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752492,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2020-02-21T04:33:18.960000",
          "content": "<p>Same as yours.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752507,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-21T05:00:03.577000",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a> That is a very good score for a frame by frame model(our's best is just 0.4418). Do you consider teaming with us? A model stacking will be really helpful.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 752516,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2020-02-21T05:24:26.253000",
          "content": "<p>Thank you for your proposal, but I want to merge with a team having a similar score to me with a single model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 752523,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-21T05:43:42.007000",
          "content": "<p>We are working on an approach. Will ask again after a better score...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 753617,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2020-02-22T13:26:38.787000",
          "content": "<p>how many images used for training when you mean 'frame by frame'?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753695,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-22T15:06:34.210000",
          "content": "<p>We used somewhere between 1m-1.5m images <a href=\"/yangsaewon\">@yangsaewon</a> in 'frame by frame'</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 753723,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-22T15:46:36.263000",
          "content": "<p>So are you guys over-sampled the Real frames?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753724,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-22T15:47:56.027000",
          "content": "<p>No., The trick we are using is pretty simple, I am not sharing just yet, I am pretty sure many people are using that, saving it for sometime later in the competition.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758366,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-27T17:22:32.420000",
          "content": "<p><a href=\"/harshit\">@harshit</a> are you using external data published ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758384,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-27T17:34:00.880000",
          "content": "<p>Did you compile the data yourself, if so how long did it take?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763074,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-03-04T05:34:27.270000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> As far as I understand, if you use the competition data then somehow preprocess it, you don't need to make public and declare.\n<a href=\"/greatgamedota\">@greatgamedota</a> Took us 1-2 days running 24 hours on 1080 TI.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763237,
          "author_name": "xmaocai",
          "author_url": "",
          "post_date": "2020-03-04T09:23:33.987000",
          "content": "<p><a href=\"/lyakaap\">@lyakaap</a>  how do you deal with data imbalance？</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 750624,
      "author_name": "Chason",
      "author_url": "",
      "post_date": "2020-02-19T14:50:07.253000",
      "content": "<p>0.32</p>",
      "votes": 7,
      "replies": [
        {
          "id": 750821,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-19T18:02:19.010000",
          "content": "<p>Amazing!, Any insights you can share, please?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750897,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-02-19T19:46:36.403000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> Are you only using images, or also audio?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751065,
          "author_name": "xinshuaidong",
          "author_url": "",
          "post_date": "2020-02-20T01:06:07.783000",
          "content": "<p>😃 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 751136,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-20T02:42:04.557000",
          "content": "<p>image only</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 751310,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-20T05:10:17.513000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> Any insights?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752477,
          "author_name": "Daic",
          "author_url": "",
          "post_date": "2020-02-21T04:04:27.207000",
          "content": "<p>You can really dance!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754113,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-02-23T04:27:35.967000",
          "content": "<p>what's your face extractor?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754382,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-23T13:38:44.677000",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> blazeface</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 754386,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-23T13:43:59.480000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> There is a lot of noise in the train dataset. For example, some FAKE videos have multiple faces, but the faces may all be REAL. Removing these noises may improve performance.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 754400,
          "author_name": "saewonYang",
          "author_url": "",
          "post_date": "2020-02-23T13:57:31.287000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> so you are removing noises in both training set and test set?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 757177,
          "author_name": "RZ",
          "author_url": "",
          "post_date": "2020-02-26T13:55:32.450000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> Are you using any external data?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 757243,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-26T15:03:19.513000",
          "content": "<p><a href=\"/dsfhe49854\">@dsfhe49854</a> No</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 759206,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-28T18:12:16.847000",
          "content": "<p>@shen chen .. how do we detect those noises in the face ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759861,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-02-29T14:54:43.967000",
          "content": "<p>Does your current score  still use the single model?CNN or CNN+RNN</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759906,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-29T15:59:35.060000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> the area of the face and the detection score</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 759908,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-02-29T16:00:19.633000",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> CNN + sequential model</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 760042,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-29T19:10:16.510000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> What is your input size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 760190,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-03-01T00:19:03.860000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> 224*224</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762753,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-03T19:28:50.323000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> \nplease help understand below queries...\n1) There are multiple fake videos for a given original video ,can we assume that all the fake videos should have same  person or set of persons speaking as is in original video</p>\n\n<p>2) Can we say all Fake videos will have fake audios also  ?</p>\n\n<p>3)  is it possible that in pair of fake and original video both will have real faces but only different voices ? eg. in both the fakes i am the only person but just that in fake its some one elses' voice  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 762755,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-03T19:34:43.637000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \n1. It is true.\n2. No\n3. It is possible. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763024,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-04T03:22:44.917000",
          "content": "<p>Thanks <a href=\"/harshitsheoran\">@harshitsheoran</a> <br>\nSo 3rd case if there are good number then image based model could making an error in prediction ,any way to separate such videos ?\nAlso would u mind giving  brief  on the this model framework..  do we need only images..  here or videos ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763068,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-04T05:27:51.687000",
          "content": "<p>There is, if you train on it, but it does not worth the time at this stage.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763258,
          "author_name": "Chason",
          "author_url": "",
          "post_date": "2020-03-04T09:54:16.913000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> So far, we haven't found any benefits from the audio model,  but it's still worth trying.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765758,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-07T03:48:49.427000",
          "content": "<p><a href=\"/chason\">@chason</a>\nm using detection score  only if score is more than this add it in list,\nhow about the area how to  find the relavant one based on this </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766128,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-03-07T17:37:30.077000",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> What's your local validation score and how did you choose your validation set?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 750044,
      "author_name": "Kha Vo",
      "author_url": "",
      "post_date": "2020-02-19T04:52:36.813000",
      "content": "<p>“Need to think of a good model” :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 750047,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-19T04:57:29.837000",
          "content": "<p>LOL I am trying to think of a better model too! There's a ton of good kernels and discussions of good models. I recommend you to look into some of them(will get you to top 10%).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 750462,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-19T12:03:59.650000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> could u point me to one such  please which u took as base.. thanks in advance </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750605,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-19T14:41:16.980000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> ResNext/Xception is a good one to start with.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750660,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-19T15:27:37.560000",
          "content": "<p>Yes, these models get me 0.372</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 750724,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-19T16:29:31.373000",
          "content": "<p>I actually use resnext and get a very good score on local CV ,balanced loss bt still very poor score at LB. \nWhat mistake i i could be making .. \nI use MTCNN in pre process pipeline.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750822,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-19T18:03:49.103000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> Take your time in finding a better face detector, If ResNext does not work the best, use it with lstm, you should be able to figure rest of it yourself. ;-)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750842,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-19T18:31:33.937000",
          "content": "<p>Thanku Harshit for your advice... I heard about LSTM used as decoder to encoder models. But havent seen its implementation if you can point out to some kernel or link.. would be useful :) </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 750875,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-19T19:07:49.680000",
          "content": "<p>In my opinion, try the normal model to get started. If you did things right, it should get you to the top 10%. We will publish lrcn(the LSTM model <a href=\"/harshitsheoran\">@harshitsheoran</a> was talking about) starter kernel soon(along with the better face extractor). </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 751565,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-02-20T09:31:44.457000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> great work. I'm sort of stuck at this point. Maybe your kernel will help me.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 752455,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-21T03:15:37.507000",
          "content": "<p>I went past two bench Mark's 467 and 451 ... at 404 now ,will try interesting things now shift from rank 590 to 80 ...long way to go  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 752804,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-21T12:30:24.603000",
          "content": "<p>Hi, I want to know did you get a score of 0.38 through lrcn?\nIs the lstm an effective approach?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753016,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-21T16:05:54.487000",
          "content": "<p><a href=\"/yzcwansui\">@yzcwansui</a> yes it was an effective approach for us.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753032,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-21T16:37:40.693000",
          "content": "<p><a href=\"/yzcwansui\">@yzcwansui</a> although for some folks, single frame model is better. But the best score we can get for single frame model is just 0.44</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 768934,
          "author_name": "xmaocai",
          "author_url": "",
          "post_date": "2020-03-11T11:14:57.350000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> <br>\nI try the resnet+lstm model, it performs bad in the test dataset while gets good score in the train dataset,  I use dropout=0.9 to prevent overfiting but it doesn't work, \ndo u have any advice? appreciate it first.\nwe train/test the model using 5 frames per video. and the num of real and fake is balanced.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 762742,
      "author_name": "Harshit Sheoran",
      "author_url": "",
      "post_date": "2020-03-03T19:19:17.933000",
      "content": "<p>New Single Frame by frame model Score 0.34000 used 160k images.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 763253,
          "author_name": "xmaocai",
          "author_url": "",
          "post_date": "2020-03-04T09:46:37.970000",
          "content": "<p>Wow, how do you deal with data imbalance？and u are using RNN model ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763490,
          "author_name": "YangHang978",
          "author_url": "",
          "post_date": "2020-03-04T14:35:53.107000",
          "content": "<p>amazing score, \nwhat's your input size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763519,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-04T15:11:31.457000",
          "content": "<p>input size 200x200, no rnn, but working on that, yes dealt with data imbalance, but doing the same thing in an LRCN has some drawbacks. Ensemble of 3 models between 0.34-0.35 (all single model), was 0.318, So, I guess, ensembling right now is much less effective than it was back then.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 763641,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-04T17:27:59.217000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a>  could u point me to  any good public lrcn model available in git hub ?\nwant to try this stuff .. will be newest of its kind for me :)\nALso if you could brief me about the overall framework of this combo. \nDo we need videos to train this model or images sufficient ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763667,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-04T18:04:45.680000",
          "content": "<p>an LRCN, have a input in this format (batch_size, N_samples, size, size, n_channels), here N_samples can be your continuous or equally gapped number of images of a same video or different, that is up to you, I dont know how to do timedistributed layers in pytorch but in keras you can search timedistributed on google, to know more about it, after that you can wrap up your pretrained model in a timedistributed layer and then use that cnn model with a rnn.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 763676,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-03-04T18:16:50.173000",
          "content": "<p>It's actually fairly easy, you can switch from frames-by-video to a linear series of frames (for the CNN) and then back to videos, like this:</p>\n\n<p><code>\n    def forward(self, x):\n        batch_size, n_channels, n_frames, height, width = x.size()  # e.g. 32 x 3 x 64 x 224 x 224\n        assert n_channels == 3, f\"Expecting 3 channels but got {n_channels}\"\n        x = x.permute(0,2,1,3,4)  # 32 x 64 x 3 x 224 x 224\n        x = x.reshape(batch_size*n_frames, n_channels, height, width)  # (32*64) x 3 x 224 x 224\n        x = self.cnn_model(x)  # (32*64) x 2048\n        x = x.view(batch_size, n_frames, self.n_cnn_features)\n        x, (h_n, h_c) = self.lstm(x)   # x = seq_len, batch, num_directions*hidden_size\n        x = x[:,-1]  # Take last time step, which will be BATCH_SIZE * (2*HIDDEN) -&amp;gt; e.g. 12 * 128\n        x = self.dropout(x)\n        x = self.linear(x)\n        return x\n</code></p>",
          "votes": 17,
          "replies": []
        },
        {
          "id": 763710,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-03-04T19:08:38.173000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764388,
          "author_name": "GJLin",
          "author_url": "",
          "post_date": "2020-03-05T12:11:47.770000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> Dose \"Frame by frame model\" mean that you take each frame as input rather than the cropped face?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764563,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-05T15:55:13.590000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> \nThanku Dr. james \nIf i understand correctly this how it would be \n1 ) CNN first layer  nn.Conv2d(3,n ....) \n2) Input  size ( 32*64=2048,3,224,224)..i wonder how we can fit such high batch size  ?\n3) CNN Output \n4) resize  output to  32 ,64 frames * 2048 for LSTM </p>\n\n<p>5) output of LSTM passed to linear layer\n6) sigmoid  bce loss</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764657,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-05T17:46:29.493000",
          "content": "<p>Frame by frame have data input like this, (batch_size, face_size, face_size, n_channels), the face are cropped from the frame but they are not fed into an lstm or any transformer, it makes it a simple binary classfication problem, but here also we find that a simple LRCN performs much better than our frame-by-frame model, it looks like we are not even close to done improving. Good Luck!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 764902,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2020-03-06T03:54:34.483000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> I did this few weeks ago, but face memory problem. How did you fit 32x64x224x224x3 for each batch? It's quite a lot. Using accumulated gradient seems to be faulty as BatchNorm won't work if we fit 1x64x224x224x3 each time.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764980,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-03-06T06:43:50.570000",
          "content": "<p>Can you clarify what you mean by 160k images? <a href=\"/harshitsheoran\">@harshitsheoran</a> Given that we have &gt;100k videos does that mean youre only looking at a couple frames per video? Or are you using some method to choose to train on a subset of videos?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764982,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-06T06:49:03.763000",
          "content": "<p><a href=\"/ryches\">@ryches</a> good question, first of all I am only using first 40 folders, yes, that does mean, that on an average we do have only a few frames from each video but honestly they are selected so perfectly anything more would be nothing but an overkill and the chances were that it would 1. not letting the model any advantage 2. making the model starting overfitting faster than its normal epoch.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 765129,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-03-06T09:29:01.957000",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> It was just an example, what batch size of videos and how many frames per video you use depend on how complex the CNN is, how many parameters you use for the RNN, your GPU, and whether you freeze the CNN before embedding it. I advise the latter, at least for earlier stages of training.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 765135,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-06T09:34:20.720000",
          "content": "<p><a href=\"/jamesphoward\">@jamesphoward</a> \nwhy did we take last time step in your sample code above ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765138,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-03-06T09:39:36.473000",
          "content": "<p>Because of the way RNNs work in pytorch you get an output (prediction) after every time step (i.e. every frame of the video). However, you probably just want the RNN's prediction once it's seen the whole video, so I take the last time step. However, maybe that's not the best method, who knows?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765269,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-06T12:41:42.873000",
          "content": "<p>What is the size of your self.linear ?\nsince output of lstm would be for eg. 12 * 128 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765298,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-03-06T13:21:51.317000",
          "content": "<p>The input size would be the the number of hidden units, or double this if your LSTM/GRU was bidirectional.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765320,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-06T13:42:15.960000",
          "content": "<p>Sorry to bother you more..\nThis is my input \n<code>\nself.LSTM=nn.LSTM(1024,512,2) \nself.linear=nn.linear(512,1)\nx torch.Size([192, 1024])-&amp;gt; reshaped to  16,12,1024  for nn.LSTM\n</code>\nif i do x=x[:,-1]  my original batch size of 192 gets reduced to 12 .\nIs there an additional layer needed before feeding to Loss function as  my  output of linear wont meet requirement of loss function which is 192*1</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765354,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-03-06T14:21:19.873000",
          "content": "<p>Have you set <code>batch_first=True</code> in your LSTM if you are supplying the batch first?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765368,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-06T14:39:16.267000",
          "content": "<p>Yup i did as first part.. \n```\nclass lstm(nn.Module):\n    def <strong>init</strong>(self,bs,num,  lstm1_input=1024,lstm_hidden=512):\n        super().<strong>init</strong>()\n        self.batch=bs #16\n        self.num=num #12</p>\n\n<pre><code>    self.lstm = nn.LSTM(1024, 512 ,2,batch_first=True)# input elements,hidden state size\n    self.linear=nn.Linear(512,1)\n    self.drop=nn.Dropout(0.5)\ndef forward(self,input):\n    #print(input.size())\n    x=input.view(self.batch,self.num,input.size(1)) # resized from 192 to 16,12\n    x,y=self.lstm(x)\n    print(x.size()) #see below\n    x = x[:,-1,:]\n    print(x.size()) #see below\n    x=self.drop(x)\n    x=self.linear(x)\n    return x\n</code></pre>\n\n<p>```</p>\n\n<p>torch.Size([16, 12, 512])\ntorch.Size([16, 512])</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765378,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-03-06T14:47:03.843000",
          "content": "<p>Isn't that correct? You have fed a batch of 16 into your network and your linear layer is receiving 16 sets of 512 hidden units' outputs? I don't see the problem...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765386,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-06T15:04:11.990000",
          "content": "<p>yes i dont have issue here but in  the loss calculation.. \n data loader will give labels with bs =batch * num of frames * num of classes\nlooks like i should generate 12 steps using some Time Distribute layer like in keras and pack them. while keeping the original bs as 16 only for data loader...\nis this wha</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 765402,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-03-06T15:26:44.247000",
          "content": "<p>Hi, I think that's really a problem you are going to need to solve as this isn't really appropriate for the forums, and we're now getting on to fundamental issues of network training, sorry!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768364,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-10T17:37:36.477000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> \nI overcame all the bs related issue after chosing right way of reshaping and transposing the dims. \nNow i face another issue which is overfitting. Model done best than ever models i tried on CV but miserably overfits on test data,especially fails to predict reals correctly. \nDid u faced this issue with LSTM.. if yes what can be f help. i purposefully dint pay attn on split strategy.</p>\n\n<p><code>\nepoch   train_loss  valid_loss  accuracy    time\n0   0.450687    0.451665    0.834688    09:08\n1   0.375398    0.342386    0.859259    09:06\n2   0.261341           0.245293        0.896838         09:02\n3   0.199818             0.183108    0.928907   09:06 \n4   0.169864    0.163689    0.938302    09:07\n5   0.140049    0.210106           0.915447     09:10\n6   0.126233            0.136548    0.947064    09:11\n</code> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768522,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-10T23:13:40.947000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> well, that loss is kinda close to ours, with accuracy also close to ours every epoch, (our 1 epoch takes 3 hours so idk how you are overfitting and I am not, I also have a much much much larger dataset which I train on compared to yours) I never had overfitting issues and please if you are using keras dont focus on valid_loss it lies, I guess you should have a benchmark, of last 10 folders, after every epoch check your model on that, and see which epoch is starts overfitting, in our case</p>\n\n<p>model__e1 -&gt; 0.212 / 0.934\nmodel_e2 -&gt; 0.183 / 0.946\nmodel_e3 -&gt; 0.175 / 0.952\nmodel_e4 -&gt; 0.186 / 0.950</p>\n\n<p>even something like this happens where you think overfitting is starting, still train it might get a better improvement in next 3 epochs.</p>\n\n<p>try dropout, and batch norm (however I dont use any of them) maybe to overcome overfitting, also no noise function in the augmentations, check the labels and your generator again manually. We had some problem and that was cuz our generator was messed up once.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 768526,
          "author_name": "Kha Vo",
          "author_url": "",
          "post_date": "2020-03-10T23:32:20.117000",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> what is your bce val score each epoch?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768565,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-03-11T01:16:58.600000",
          "content": "<p><a href=\"/khahuras\">@khahuras</a>  that is the validation loss I shared and not the training loss</p>\n\n<p>0.160 on that bench means 0.318 on leaderboard, always almost double.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 768642,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-11T03:43:04.843000",
          "content": "<p>1) how does your hist gram of 400 sub videos look like  assuming fake =0.99 and real =0.... \nleft and right height bars ,which one more ?\n2) bi direction LSTM or uni ?\n3) what are you lstm dimensions..   bs,num_frames ,no of elements  does keeping higher bs helps more ?\n4) how is your batch distribution  real:fake ratio in batch</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 768661,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-03-11T04:24:13.177000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> here's my response(I'm basically in charge of model training right now):\n1. We used a post-processing technique. Here's the histogram<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fd8ce84bed0e9d697476308f24ae354a7%2FScreen%20Shot%202020-03-10%20at%209.20.51%20PM.png?generation=1583900473994383&amp;alt=media\" alt=\"\"></p>\n\n<ol>\n<li>We used an uni.</li>\n<li>batch size is 6, the rest, I can't tell you. Higher batch size does sometimes help a little bit in experiment, but we consider it is due to random, because it is higher like 0.001, and sometimes, its actually worse.</li>\n</ol>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 769048,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-11T13:23:26.830000",
          "content": "<p>thanku.. during inference how many video frames used...\nis it necessary to keep same size as in training ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769132,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-03-11T15:02:04.417000",
          "content": "<p>The answer might surprise you. We only used not a constant amount of frame per video.... I don't get the second question, sry.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769193,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-11T16:15:01.213000",
          "content": "<p>i meant the batch size and num of frames  during inference should they be  same as in training for lstm to give expected output as in training</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769201,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-03-11T16:23:32.310000",
          "content": "<p>ok. I got you now. the num of frames during inference is 10(they don't really affect much). (secret hint: we don't pass that 10 frames to lrcn at once, I can't really go into details.). for batch size during inference, its 1.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 770631,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-13T07:14:12.937000",
          "content": "<p>Thanks.. is public score calculated based on the 400 test videos or it could  be some other videos which are subset of private one.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770917,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-03-13T15:09:45.687000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> I think the public score is calculated based on the 4000 videos that will be tested in a black box environment.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 756584,
      "author_name": "Nuno Ferreira",
      "author_url": "",
      "post_date": "2020-02-25T21:33:50.670000",
      "content": "<p>My latest model scores: 0.3422</p>",
      "votes": 4,
      "replies": [
        {
          "id": 758365,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-27T17:20:11.153000",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a>  amazing score\nwhich model did u use ?\nHave u used ur own dataset,if yes which face detector and no of frames used ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758416,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-02-27T18:19:41.723000",
          "content": "<p>I've use a CNN based model, using only one frame.\nI've used no external data and used yolov3 for detecting faces and create the training set and blazeface for  inference. \nThe prediction is just an average of 40 frames.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 758419,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-27T18:22:00.390000",
          "content": "<p>what is your input face size?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759069,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-02-28T14:27:16.680000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> I'm using face images scaled to 224x224</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 763234,
          "author_name": "xmaocai",
          "author_url": "",
          "post_date": "2020-03-04T09:21:03.777000",
          "content": "<p><a href=\"/ngcferreira\">@ngcferreira</a>  how do you deal with data imbalance？</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 766231,
          "author_name": "Nuno Ferreira",
          "author_url": "",
          "post_date": "2020-03-07T22:00:21.667000",
          "content": "<p><a href=\"/xmaocai\">@xmaocai</a>  So first I split the data into 3 sets (train, val and test set). In every epoch, I randomly choose the same number of fake images as the number of real images. Basically I'm under sampling the data.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 765464,
      "author_name": "BokingChen",
      "author_url": "",
      "post_date": "2020-03-06T16:48:22.063000",
      "content": "<p>CNN+LSTM ,You can refer to <a href=\"https://github.com/HHTseng/video-classification\">https://github.com/HHTseng/video-classification</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 768635,
          "author_name": "xmaocai",
          "author_url": "",
          "post_date": "2020-03-11T03:32:03.570000",
          "content": "<p>Hello, thanks for u advice. I try it but the model doesn't perform well in the val dataset while gets good acc in the train dataset. I use dropout=0.9 to prevent overfiting but it doesn't work, \ndo u have any advice? \nbtw, 5 frames per video.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 769109,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-03-11T14:38:15.880000",
          "content": "<p>@xmacai what is the dim of LSTM and input X to CNN.. \nmay u could be mistaking  at any  place</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 770798,
          "author_name": "xmaocai",
          "author_url": "",
          "post_date": "2020-03-13T12:27:33.170000",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \ninput X dim: torch.Size([64, 1, 3, 224, 224])\nthe dim of input to LSTM: torch.Size([8, 1, 512])\nLSTM dim: \n<code>self.LSTM =nn.LSTM(\n            input_size=300,\n            hidden_size=256,\n            num_layers=3,\n            batch_first=True,)</code>\nHere are the input data handle:\n<code>for index in data:</code>\n            <code>image = Image.open(index).convert('RGB')</code>\n            <code>if self.data_transforms is not None:</code>\n                <code>try:</code>\n                    <code>image = self.data_transforms(image)</code>\n                <code>except:</code>\n                    <code>print(\"Cannot transform image: {}\".format(index))</code>\n            <code>X.append(image)</code>\n<code>X = torch.stack(X, dim=0)</code>\n<code>Y = torch.LongTensor([label])</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757639,
      "author_name": "A*DF",
      "author_url": "",
      "post_date": "2020-02-27T01:22:17.800000",
      "content": "<p>current best score is 0.40621 after trying for weeks :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 753355,
      "author_name": "Akash",
      "author_url": "",
      "post_date": "2020-02-22T04:17:41.253000",
      "content": "<p>Well, I have my train split as 0-29 and 40 to 49 and val split as 30-39, the same as <a href=\"/jamesphoward\">@jamesphoward</a> . Currently, I'm getting a validation score of 0.295 for \"Image Only\" model. But I have 0 trust in it! Or maybe I have a bug in my inference pipeline. Who knows!\nAlso, the fact that the model starts overfitting from 3rd epoch onwards gives me the chills.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 750291,
      "author_name": "Nuno Ferreira",
      "author_url": "",
      "post_date": "2020-02-19T09:07:17.483000",
      "content": "<p>My best model (images only) scores: 0.40457</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 750589,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-02-19T14:34:03.210000",
      "content": "<p>0.489 :P</p>\n\n<p>Update 3: .402\nUpdate 4: .392</p>",
      "votes": 2,
      "replies": [
        {
          "id": 765336,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-03-06T14:07:27.887000",
          "content": "<p>0.43* All I did was make a different inference kernal (different face detector) and pick a specific model trained on the data from my baseline kernal (one I trained over 2 weeks ago).</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 785367,
      "author_name": "saewonYang",
      "author_url": "",
      "post_date": "2020-03-25T02:14:25.803000",
      "content": "<p>what kind of scheduler do you guys use for training? fixed lr or cosine scheduling, etc..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 769950,
      "author_name": "Trigram",
      "author_url": "",
      "post_date": "2020-03-12T12:39:41.487000",
      "content": "<p>0.44202.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 776564,
          "author_name": "Nirjhar Roy",
          "author_url": "",
          "post_date": "2020-03-17T13:28:24.857000",
          "content": "<p>0.42352  .. Just near :) </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 755733,
      "author_name": "Phil",
      "author_url": "",
      "post_date": "2020-02-25T04:23:27.997000",
      "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a>  hi，when using lrcn，how many frames of  one video do you use？are they sequential frames  or not。 appreciate for your advises</p>",
      "votes": 0,
      "replies": [
        {
          "id": 757654,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-27T01:58:20.820000",
          "content": "<p>10 frames per video, you can increase or decrease does not really matter, they are not sequential, they are equally gapped.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 752910,
      "author_name": "petya",
      "author_url": "",
      "post_date": "2020-02-21T14:37:16.007000",
      "content": "<p>0.41136 with a single CNN model. I wonder, how are you dealing with inference on multi-faced videos? I mean, one face  is usually okay in each frame, another is distorted, rarely both of them. Should I pick the max prediction per frame or something? \nupd: 0.40463\nupd: 0.40128</p>",
      "votes": 0,
      "replies": [
        {
          "id": 752923,
          "author_name": "James Howard",
          "author_url": "",
          "post_date": "2020-02-21T14:46:02.813000",
          "content": "<p>Excellent question. I've experimented with this and I've found the mean to be very slightly better than either the max or min.</p>\n\n<p>I used to think the maximum would be best, but that assumes you don't have false positives for face detection; false positives are likely to be unnatural and therefore fakes, and so it might be the mean is a little more resilient to this.</p>\n\n<p>I would highlight, though, that this is very likely to depend on your processing pipeline.</p>",
          "votes": 8,
          "replies": []
        },
        {
          "id": 752964,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-21T15:25:10.303000",
          "content": "<p>Distorted face found during detection are likely to be predicted as Fake. So if there is frame with distorted  in list of frames  then that would most likely contribute to false positive and hence with max prob compared to all. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 753088,
          "author_name": "petya",
          "author_url": "",
          "post_date": "2020-02-21T17:39:24.780000",
          "content": "<p>Thanks for the insight. I guess I'll play with inference detector a little bit more. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 754121,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-02-23T04:41:19.190000",
          "content": "<p>what's your face extractor? I use MTCNN but it can't get a good LB</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 758870,
          "author_name": "petya",
          "author_url": "",
          "post_date": "2020-02-28T09:22:12.697000",
          "content": "<p>I've played with blazeface, yolov2, yolov3, mtcnn, and yolov3 got me the best score of 0.40463 for a single model</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 761607,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "2020-03-02T18:32:21.367000",
          "content": "<p><a href=\"/defileroff\">@defileroff</a> did you find that there was much difference between the face extractors in terms of overall score? I haven't played around with this yet, wondering what kind of priority it should be relative to everything else :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 764394,
          "author_name": "petya",
          "author_url": "",
          "post_date": "2020-03-05T12:19:56.800000",
          "content": "<p><a href=\"/fergusoci\">@fergusoci</a>  Playing with different frames, faces, extractors for train/valid sets gave me a boost of 0.05, playing with different extractors at inference time gave me a boost of 0.01 </p>\n\n<p>I guess if your classifier is weak in terms of false-positive numbers, you'd need a better detector to avoid false-positive face detections. But if your classifier is good enough, you can take a less picky detector, that will detect everything that looks like a face, lol, and let the classifier deal with it.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 774959,
          "author_name": "kkk man",
          "author_url": "",
          "post_date": "2020-03-16T05:33:29.230000",
          "content": "<p><a href=\"/defileroff\">@defileroff</a> pre-trained yolov3 face detector or trained with dfdc dataset? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 790715,
          "author_name": "petya",
          "author_url": "",
          "post_date": "2020-03-29T19:53:32.240000",
          "content": "<p>Pretrained one, I don't even use multi-frames like others do, works fine to some extent.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 752227,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-20T19:52:05.620000",
      "content": "<p>good</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 765463,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-06T16:46:12.143000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 765277,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-06T12:56:25.547000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 759856,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-29T14:52:08.033000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 759946,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-29T16:39:44.357000",
          "content": "<p>We are using a ensemble.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 763356,
          "author_name": "BokingChen",
          "author_url": "",
          "post_date": "2020-03-04T11:57:23.233000",
          "content": "<p>thanks, I ensamble two model also get a better score</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 757172,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-26T13:52:53.370000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "750030": "Please share your best single model LB(not ensemble).\nTo start with ours: 0.38504\nEDIT: 0.35040\nEDIT2: 0.34",
    "751313": "Frame-by-frame model (only image): 0.33966 ",
    "750624": "0.32",
    "750044": "“Need to think of a good model” :)",
    "762742": "New Single Frame by frame model Score 0.34000 used 160k images.",
    "756584": "My latest model scores: 0.3422",
    "765464": "CNN+LSTM ,You can refer to https://github.com/HHTseng/video-classification",
    "757639": "current best score is 0.40621 after trying for weeks :)",
    "753355": "Well, I have my train split as 0-29 and 40 to 49 and val split as 30-39, the same as @jamesphoward . Currently, I'm getting a validation score of 0.295 for \"Image Only\" model. But I have 0 trust in it! Or maybe I have a bug in my inference pipeline. Who knows!\nAlso, the fact that the model starts overfitting from 3rd epoch onwards gives me the chills.",
    "750291": "My best model (images only) scores: 0.40457\n",
    "750589": "0.489 :P\n\nUpdate 3: .402\nUpdate 4: .392",
    "785367": "what kind of scheduler do you guys use for training? fixed lr or cosine scheduling, etc..",
    "769950": "0.44202.",
    "755733": "@harshitsheoran  hi，when using lrcn，how many frames of  one video do you use？are they sequential frames  or not。 appreciate for your advises",
    "752910": "0.41136 with a single CNN model. I wonder, how are you dealing with inference on multi-faced videos? I mean, one face  is usually okay in each frame, another is distorted, rarely both of them. Should I pick the max prediction per frame or something? \nupd: 0.40463\nupd: 0.40128",
    "752227": "good",
    "765463": "",
    "765277": "",
    "759856": "",
    "757172": ""
  }
}