{
  "id": 128702,
  "title": "[placeholder] pytorch starter kit",
  "url": "/competitions/deepfake-detection-challenge/discussion/128702",
  "author_name": "",
  "post_date": "2020-02-02T16:31:07.516425500Z",
  "votes": 36,
  "comment_count": 36,
  "views": 0,
  "content": "<p>objective is to implement a baseline system based on the paper:\n<a href=\"https://arxiv.org/pdf/2001.07444.pdf\">https://arxiv.org/pdf/2001.07444.pdf</a>\nDetecting Face2Face Facial Reenactment in Videos</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe6f0e559bacd9ebb937436ddd7909b68%2FSelection_051.png?generation=1580661047964292&amp;alt=media\" alt=\"\"></p>\n\n<p>...\n<a href=\"https://drive.google.com/open?id=1El6zsA0_k5IcW2XLwQRd6pAXshURrFo6\">https://drive.google.com/open?id=1El6zsA0_k5IcW2XLwQRd6pAXshURrFo6</a></p>\n\n<p>(to be updated)</p>",
  "messages": [
    {
      "id": "735137",
      "postDate": "02/02/2020 16:31:07",
      "content": "<p>objective is to implement a baseline system based on the paper:\n<a href=\"https://arxiv.org/pdf/2001.07444.pdf\">https://arxiv.org/pdf/2001.07444.pdf</a>\nDetecting Face2Face Facial Reenactment in Videos</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe6f0e559bacd9ebb937436ddd7909b68%2FSelection_051.png?generation=1580661047964292&amp;alt=media\" alt=\"\"></p>\n\n<p>...\n<a href=\"https://drive.google.com/open?id=1El6zsA0_k5IcW2XLwQRd6pAXshURrFo6\">https://drive.google.com/open?id=1El6zsA0_k5IcW2XLwQRd6pAXshURrFo6</a></p>\n\n<p>(to be updated)</p>",
      "rawMarkdown": "objective is to implement a baseline system based on the paper:\nhttps://arxiv.org/pdf/2001.07444.pdf\nDetecting Face2Face Facial Reenactment in Videos\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe6f0e559bacd9ebb937436ddd7909b68%2FSelection_051.png?generation=1580661047964292&amp;alt=media)\n\n\n...\nhttps://drive.google.com/open?id=1El6zsA0_k5IcW2XLwQRd6pAXshURrFo6\n\n(to be updated)",
      "votes": null
    },
    {
      "id": "735141",
      "postDate": "02/02/2020 16:39:31",
      "content": "<p>a few notes:</p>\n\n<ol>\n<li><p>you are given train videos as :  a true video (identity A) and fake videos (identity A is swap to B, unfortunately video of B is not given). during training, you should sample frame in the pair, i.e both the same time frame of real and fake video must be present in a single batch </p></li>\n<li><p>you can \"force attention\" to the face region, i.e. the classifier should make decision based on the weird/unnatural face pixels and avoid background pixels (unless they contains information about inherit image noise signature, etc)</p></li>\n<li><p>there are several labels you can get from training set. e.g.:</p>\n\n<ul><li>image label : is it from real or fake video</li>\n<li>pixel label : subtract fake frame from real frame and see if  the pixel is altered / manipulated</li>\n<li>block label : divide image into say 8x8 block.  subtract fake block from real block </li>\n<li>etc ... same idea apply in the temporal/frame difference/ time axis ...</li></ul></li>\n</ol>",
      "rawMarkdown": "a few notes:\n\n1. you are given train videos as :  a true video (identity A) and fake videos (identity A is swap to B, unfortunately video of B is not given). during training, you should sample frame in the pair, i.e both the same time frame of real and fake video must be present in a single batch \n\n2. you can \"force attention\" to the face region, i.e. the classifier should make decision based on the weird/unnatural face pixels and avoid background pixels (unless they contains information about inherit image noise signature, etc)\n\n3. there are several labels you can get from training set. e.g.:\n   - image label : is it from real or fake video\n   - pixel label : subtract fake frame from real frame and see if  the pixel is altered / manipulated\n   - block label : divide image into say 8x8 block.  subtract fake block from real block \n   - etc ... same idea apply in the temporal/frame difference/ time axis ...",
      "votes": null
    },
    {
      "id": "735240",
      "postDate": "02/02/2020 19:01:57",
      "content": "<p>Placeholder comment</p>",
      "rawMarkdown": "Placeholder comment",
      "votes": null
    },
    {
      "id": "735247",
      "postDate": "02/02/2020 19:22:40",
      "content": "<p>[Placeholder] comment😄 😄 </p>",
      "rawMarkdown": "[Placeholder] comment😄 😄",
      "votes": null
    },
    {
      "id": "735373",
      "postDate": "02/02/2020 23:54:25",
      "content": "<p>placeholder fork - kidding, nice pointers for ideas already, thanks for sharing <a href=\"/hengck23\">@hengck23</a> </p>",
      "rawMarkdown": "placeholder fork - kidding, nice pointers for ideas already, thanks for sharing @hengck23",
      "votes": null
    },
    {
      "id": "735729",
      "postDate": "02/03/2020 11:48:38",
      "content": "<p>StyleGAN2: Near-Perfect Human Face Synthesis...and More\n<a href=\"https://www.youtube.com/watch?v=SWoravHhsUU\">https://www.youtube.com/watch?v=SWoravHhsUU</a></p>\n\n<p>teeth and eye are the telltale signs</p>",
      "rawMarkdown": "StyleGAN2: Near-Perfect Human Face Synthesis...and More\nhttps://www.youtube.com/watch?v=SWoravHhsUU\n\nteeth and eye are the telltale signs",
      "votes": null
    },
    {
      "id": "735807",
      "postDate": "02/03/2020 13:30:57",
      "content": "<p>Hey, <a href=\"/hengck23\">@hengck23</a>, I had tried that method in the past, did not work out very well, However that was a few weeks ago, I guess, I am gonna retry this thing, One mistake you would want to avoid is even trying to go for resnet-18, But thinking about it, you would have already got your mind which net would be replacing that. Also, I tried your note number 1 (to keep same videos/same frames/ different label in 1 batch) few days before, It was not performing any good, Right now, Can you answer me, Why these approaches should work, what is in your mind which makes you think it will work, Thank You!.</p>",
      "rawMarkdown": "Hey, @hengck23, I had tried that method in the past, did not work out very well, However that was a few weeks ago, I guess, I am gonna retry this thing, One mistake you would want to avoid is even trying to go for resnet-18, But thinking about it, you would have already got your mind which net would be replacing that. Also, I tried your note number 1 (to keep same videos/same frames/ different label in 1 batch) few days before, It was not performing any good, Right now, Can you answer me, Why these approaches should work, what is in your mind which makes you think it will work, Thank You!.",
      "votes": null
    },
    {
      "id": "735895",
      "postDate": "02/03/2020 15:18:33",
      "content": "<p>Paper -&gt; <a href=\"https://arxiv.org/pdf/1912.04958.pdf\">https://arxiv.org/pdf/1912.04958.pdf</a></p>",
      "rawMarkdown": "Paper -&gt; https://arxiv.org/pdf/1912.04958.pdf",
      "votes": null
    },
    {
      "id": "736706",
      "postDate": "02/04/2020 13:34:56",
      "content": "<p>this is an example of inference code for beginner kaggler. it is based on @ Human Analog 's work: <a href=\"https://www.kaggle.com/humananalog/inference-demo\">https://www.kaggle.com/humananalog/inference-demo</a></p>\n\n<p>I refactor it for my own clarity. refer to the readme.ppt for details.</p>",
      "rawMarkdown": "this is an example of inference code for beginner kaggler. it is based on @ Human Analog 's work: https://www.kaggle.com/humananalog/inference-demo\n\n\nI refactor it for my own clarity. refer to the readme.ppt for details.",
      "votes": null
    },
    {
      "id": "737282",
      "postDate": "02/05/2020 06:09:06",
      "content": "<p>Loophole?</p>\n\n<p>I am not sure if both real and fake video will be present in the test or not. If so, one can cluster the same videos together. We know that only one of the video is real and the rest are manipulated. Or with many similar videos, there exist at most one or zero real version</p>",
      "rawMarkdown": "Loophole?\n\nI am not sure if both real and fake video will be present in the test or not. If so, one can cluster the same videos together. We know that only one of the video is real and the rest are manipulated. Or with many similar videos, there exist at most one or zero real version",
      "votes": null
    },
    {
      "id": "737290",
      "postDate": "02/05/2020 06:26:52",
      "content": "<p>Bascially not all frames  are “fake” ( or more correctly, looks “fake”)There needs to be temporal localisation method, just like action recognition in a video.</p>\n\n<p>Probably the static image classifier needs to pool over the whole video, i.e over the time domain as well for training a image feature extraction</p>",
      "rawMarkdown": "Bascially not all frames  are “fake” ( or more correctly, looks “fake”)There needs to be temporal localisation method, just like action recognition in a video.\n\nProbably the static image classifier needs to pool over the whole video, i.e over the time domain as well for training a image feature extraction",
      "votes": null
    },
    {
      "id": "737346",
      "postDate": "02/05/2020 08:06:32",
      "content": "<p><a href=\"/hengck23\">@hengck23</a> Please check your email and reply.\nThank You!.</p>",
      "rawMarkdown": "hengck23 Please check your email and reply.\nThank You!.",
      "votes": null
    },
    {
      "id": "738327",
      "postDate": "02/06/2020 11:47:45",
      "content": "<p>okay</p>",
      "rawMarkdown": "okay",
      "votes": null
    },
    {
      "id": "738413",
      "postDate": "02/06/2020 13:43:13",
      "content": "<p>Well, <a href=\"/hengck23\">@hengck23</a>, I would take this as \"Most images looks to be real in a fake video\" but are fake, using this believe is much better.</p>",
      "rawMarkdown": "Well, @hengck23, I would take this as \"Most images looks to be real in a fake video\" but are fake, using this believe is much better.",
      "votes": null
    },
    {
      "id": "740799",
      "postDate": "02/09/2020 20:10:43",
      "content": "<p>Just wondering when will the kernel will be made public?</p>",
      "rawMarkdown": "Just wondering when will the kernel will be made public?",
      "votes": null
    },
    {
      "id": "741297",
      "postDate": "02/10/2020 13:00:13",
      "content": "<p>Yes ,I  think  we  should  train  a  model  to    filter   some   import   frame , but  not  input  all  the  video  frame  or  just   random  select .</p>",
      "rawMarkdown": "Yes ,I  think  we  should  train  a  model  to    filter   some   import   frame , but  not  input  all  the  video  frame  or  just   random  select .",
      "votes": null
    },
    {
      "id": "745748",
      "postDate": "02/14/2020 06:36:29",
      "content": "<p>thank you for the share. Always nice to have you in a comp 👍 </p>",
      "rawMarkdown": "thank you for the share. Always nice to have you in a comp 👍",
      "votes": null
    },
    {
      "id": "746081",
      "postDate": "02/14/2020 15:44:54",
      "content": "<p>For what it's worth, I trained the model from this paper on my dataset and it scores 0.42574 on the LB.</p>\n\n<p>The only changes I made are:</p>\n\n<ul>\n<li>Sigmoid instead of softmax, so the combined tensor is 5x1 instead of 10x1.</li>\n<li>Using 1-cycle training for 10 epochs with a maximum LR of 1e-3.</li>\n<li>Each batch consists of (real, fake) image pairs.</li>\n<li>My own data augmentation, etc.</li>\n<li>I used pretrained ResNet-18 from torchvision, and froze the model up to block 4.</li>\n</ul>\n\n<p>This took about 10 hrs to train on my 1080 Ti using 1M images (batch size 128). Final train loss is 0.2041, val loss is 0.1923 (but the val loss is fairly meaningless as my validation set is not very good; I only use it to make sure the model isn't overfitting).</p>",
      "rawMarkdown": "For what it's worth, I trained the model from this paper on my dataset and it scores 0.42574 on the LB.\n\nThe only changes I made are:\n\n- Sigmoid instead of softmax, so the combined tensor is 5x1 instead of 10x1.\n- Using 1-cycle training for 10 epochs with a maximum LR of 1e-3.\n- Each batch consists of (real, fake) image pairs.\n- My own data augmentation, etc.\n- I used pretrained ResNet-18 from torchvision, and froze the model up to block 4.\n\nThis took about 10 hrs to train on my 1080 Ti using 1M images (batch size 128). Final train loss is 0.2041, val loss is 0.1923 (but the val loss is fairly meaningless as my validation set is not very good; I only use it to make sure the model isn't overfitting).",
      "votes": null
    },
    {
      "id": "746224",
      "postDate": "02/14/2020 19:02:30",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> Thanks for the experiment!</p>",
      "rawMarkdown": "humananalog Thanks for the experiment!",
      "votes": null
    },
    {
      "id": "746260",
      "postDate": "02/14/2020 19:47:14",
      "content": "<p>I think it might be useful to do this sort of thing anyway, as it's basically an ensemble of 5 different classifiers. The linear layer that you add to the end is for taking the weighted mean of their results. So it should improve existing models, although it will be 5x slower to train obviously.</p>",
      "rawMarkdown": "I think it might be useful to do this sort of thing anyway, as it's basically an ensemble of 5 different classifiers. The linear layer that you add to the end is for taking the weighted mean of their results. So it should improve existing models, although it will be 5x slower to train obviously.",
      "votes": null
    },
    {
      "id": "746272",
      "postDate": "02/14/2020 19:54:33",
      "content": "<p>Do you have any log of how the Resnet18_1 performs (as it's a full face classifier) in your local val loss? Just curious to know how the ensemble helped in boosting!</p>",
      "rawMarkdown": "Do you have any log of how the Resnet18_1 performs (as it's a full face classifier) in your local val loss? Just curious to know how the ensemble helped in boosting!",
      "votes": null
    },
    {
      "id": "746287",
      "postDate": "02/14/2020 20:02:30",
      "content": "<p>No, unfortunately I don't have any results for a single ResNet-18 model.</p>",
      "rawMarkdown": "No, unfortunately I don't have any results for a single ResNet-18 model.",
      "votes": null
    },
    {
      "id": "747053",
      "postDate": "02/15/2020 22:22:32",
      "content": "<p>Since I wasn't running any other experiments, I decided to \"upgrade\" this model to use ResNet-34 instead of ResNet-18 and train it in exactly the same way. It only gave a minor improvement in the score on the LB test set: 0.41932 vs 0.42574. (On my local validation set the score was pretty much the same as before.)</p>",
      "rawMarkdown": "Since I wasn't running any other experiments, I decided to \"upgrade\" this model to use ResNet-34 instead of ResNet-18 and train it in exactly the same way. It only gave a minor improvement in the score on the LB test set: 0.41932 vs 0.42574. (On my local validation set the score was pretty much the same as before.)",
      "votes": null
    },
    {
      "id": "747084",
      "postDate": "02/16/2020 00:05:57",
      "content": "<p>Well, That was expected, larger model will not be doing any good, you know.</p>",
      "rawMarkdown": "Well, That was expected, larger model will not be doing any good, you know.",
      "votes": null
    },
    {
      "id": "747358",
      "postDate": "02/16/2020 10:30:34",
      "content": "<p>But when is a model too small and when is it too large? You don't know it until you try it. ;-)</p>",
      "rawMarkdown": "But when is a model too small and when is it too large? You don't know it until you try it. ;-)",
      "votes": null
    },
    {
      "id": "747494",
      "postDate": "02/16/2020 14:19:50",
      "content": "<p>Well, you can know when model is too small or too large by just seeing how the model reacts (without actually training.), EZ ;-)...</p>",
      "rawMarkdown": "Well, you can know when model is too small or too large by just seeing how the model reacts (without actually training.), EZ ;-)...",
      "votes": null
    },
    {
      "id": "747812",
      "postDate": "02/16/2020 22:18:48",
      "content": "<p>Again, I did haven't an experiment planned for today so I trained just a single ResNet-18 on my data and it scores 0.42936 on the LB. This is almost identical to the 5x ResNet-18 model. So it turns out the \"ensembling\" doesn't really help very much.</p>",
      "rawMarkdown": "Again, I did haven't an experiment planned for today so I trained just a single ResNet-18 on my data and it scores 0.42936 on the LB. This is almost identical to the 5x ResNet-18 model. So it turns out the \"ensembling\" doesn't really help very much.",
      "votes": null
    },
    {
      "id": "747824",
      "postDate": "02/16/2020 22:42:21",
      "content": "<p>Sounds like a really good job with data if even a single resnet18 model gets such score on the LB</p>",
      "rawMarkdown": "Sounds like a really good job with data if even a single resnet18 model gets such score on the LB",
      "votes": null
    },
    {
      "id": "747858",
      "postDate": "02/16/2020 23:56:40",
      "content": "<p><a href=\"/ims0rry\">@ims0rry</a> totally agree!</p>",
      "rawMarkdown": "ims0rry totally agree!",
      "votes": null
    },
    {
      "id": "747877",
      "postDate": "02/17/2020 01:01:31",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> How were you able to score 0.42 on LB with single frame? Our best single frame was just 0.44. BTW, do consider teaming with us. A model stacking might be pretty helpful. </p>",
      "rawMarkdown": "humananalog How were you able to score 0.42 on LB with single frame? Our best single frame was just 0.44. BTW, do consider teaming with us. A model stacking might be pretty helpful.",
      "votes": null
    },
    {
      "id": "748040",
      "postDate": "02/17/2020 06:13:33",
      "content": "<p>[Placeholder [Placeholder] funny placeholder comment.] ;)</p>",
      "rawMarkdown": "[Placeholder [Placeholder] funny placeholder comment.] ;)",
      "votes": null
    },
    {
      "id": "748302",
      "postDate": "02/17/2020 11:09:28",
      "content": "<blockquote>\n  <p>How were you able to score 0.42 on LB with single frame</p>\n</blockquote>\n\n<p>Ah, sorry if this wasn't clear. This score (and the others I posted in this thread) is done by averaging the predictions of 17 frames from each video. Using only a single frame would make the score worse, no doubt. (For fun, I'll try that now.)</p>",
      "rawMarkdown": "&gt; How were you able to score 0.42 on LB with single frame\n\nAh, sorry if this wasn't clear. This score (and the others I posted in this thread) is done by averaging the predictions of 17 frames from each video. Using only a single frame would make the score worse, no doubt. (For fun, I'll try that now.)",
      "votes": null
    },
    {
      "id": "748333",
      "postDate": "02/17/2020 11:52:30",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> any reply for \"do consider teaming with us\", Can we team up?</p>",
      "rawMarkdown": "humananalog any reply for \"do consider teaming with us\", Can we team up?",
      "votes": null
    },
    {
      "id": "748352",
      "postDate": "02/17/2020 12:26:44",
      "content": "<blockquote>\n  <p>Can we team up?</p>\n</blockquote>\n\n<p>Thanks for asking but I don't really know how much more time I can put into this competition, so it's probably not a good idea to have me in your team. 😉 </p>",
      "rawMarkdown": "&gt; Can we team up?\n\nThanks for asking but I don't really know how much more time I can put into this competition, so it's probably not a good idea to have me in your team. 😉",
      "votes": null
    },
    {
      "id": "748513",
      "postDate": "02/17/2020 15:59:55",
      "content": "<blockquote>\n  <p>Using only a single frame would make the score worse, no doubt</p>\n</blockquote>\n\n<p>Exact same model, using just the middle frame from each video scores 0.60600 on the LB. So that's a massive difference.</p>",
      "rawMarkdown": "&gt; Using only a single frame would make the score worse, no doubt\n\nExact same model, using just the middle frame from each video scores 0.60600 on the LB. So that's a massive difference.",
      "votes": null
    },
    {
      "id": "750782",
      "postDate": "02/19/2020 17:24:11",
      "content": "<p>Get some idea: \n1. extract same frame(image) from REAL and FAKE video, \n2. apply a t-test to this pair images to decide if the image is significant different to each other.\n3. set a threshold to keep the most different pair image(REAL and FAKE) as dataset</p>",
      "rawMarkdown": "Get some idea: \n1. extract same frame(image) from REAL and FAKE video, \n2. apply a t-test to this pair images to decide if the image is significant different to each other.\n3. set a threshold to keep the most different pair image(REAL and FAKE) as dataset",
      "votes": null
    },
    {
      "id": "750818",
      "postDate": "02/19/2020 17:59:19",
      "content": "<p>Good Idea, I might try it in near future. Unfortunately there is no fast way to do it, preprocessing to classify which image is the most different is a mess to do, but I think I have a new idea because of you so Thank You!</p>",
      "rawMarkdown": "Good Idea, I might try it in near future. Unfortunately there is no fast way to do it, preprocessing to classify which image is the most different is a mess to do, but I think I have a new idea because of you so Thank You!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 735141,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/02/2020 16:39:31",
      "content": "<p>a few notes:</p>\n\n<ol>\n<li><p>you are given train videos as :  a true video (identity A) and fake videos (identity A is swap to B, unfortunately video of B is not given). during training, you should sample frame in the pair, i.e both the same time frame of real and fake video must be present in a single batch </p></li>\n<li><p>you can \"force attention\" to the face region, i.e. the classifier should make decision based on the weird/unnatural face pixels and avoid background pixels (unless they contains information about inherit image noise signature, etc)</p></li>\n<li><p>there are several labels you can get from training set. e.g.:</p>\n\n<ul><li>image label : is it from real or fake video</li>\n<li>pixel label : subtract fake frame from real frame and see if  the pixel is altered / manipulated</li>\n<li>block label : divide image into say 8x8 block.  subtract fake block from real block </li>\n<li>etc ... same idea apply in the temporal/frame difference/ time axis ...</li></ul></li>\n</ol>",
      "votes": null,
      "replies": [
        {
          "id": 737290,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "02/05/2020 06:26:52",
          "content": "<p>Bascially not all frames  are “fake” ( or more correctly, looks “fake”)There needs to be temporal localisation method, just like action recognition in a video.</p>\n\n<p>Probably the static image classifier needs to pool over the whole video, i.e over the time domain as well for training a image feature extraction</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 738413,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/06/2020 13:43:13",
          "content": "<p>Well, <a href=\"/hengck23\">@hengck23</a>, I would take this as \"Most images looks to be real in a fake video\" but are fake, using this believe is much better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 741297,
          "author_name": "xujingzhao",
          "author_url": "",
          "post_date": "02/10/2020 13:00:13",
          "content": "<p>Yes ,I  think  we  should  train  a  model  to    filter   some   import   frame , but  not  input  all  the  video  frame  or  just   random  select .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750782,
          "author_name": "gody7334",
          "author_url": "",
          "post_date": "02/19/2020 17:24:11",
          "content": "<p>Get some idea: \n1. extract same frame(image) from REAL and FAKE video, \n2. apply a t-test to this pair images to decide if the image is significant different to each other.\n3. set a threshold to keep the most different pair image(REAL and FAKE) as dataset</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750818,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/19/2020 17:59:19",
          "content": "<p>Good Idea, I might try it in near future. Unfortunately there is no fast way to do it, preprocessing to classify which image is the most different is a mess to do, but I think I have a new idea because of you so Thank You!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 735240,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "02/02/2020 19:01:57",
      "content": "<p>Placeholder comment</p>",
      "votes": null,
      "replies": [
        {
          "id": 735247,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "02/02/2020 19:22:40",
          "content": "<p>[Placeholder] comment😄 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748040,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "02/17/2020 06:13:33",
          "content": "<p>[Placeholder [Placeholder] funny placeholder comment.] ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 735373,
      "author_name": "yifanxie",
      "author_url": "",
      "post_date": "02/02/2020 23:54:25",
      "content": "<p>placeholder fork - kidding, nice pointers for ideas already, thanks for sharing <a href=\"/hengck23\">@hengck23</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 735729,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/03/2020 11:48:38",
      "content": "<p>StyleGAN2: Near-Perfect Human Face Synthesis...and More\n<a href=\"https://www.youtube.com/watch?v=SWoravHhsUU\">https://www.youtube.com/watch?v=SWoravHhsUU</a></p>\n\n<p>teeth and eye are the telltale signs</p>",
      "votes": null,
      "replies": [
        {
          "id": 735895,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/03/2020 15:18:33",
          "content": "<p>Paper -&gt; <a href=\"https://arxiv.org/pdf/1912.04958.pdf\">https://arxiv.org/pdf/1912.04958.pdf</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 735807,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/03/2020 13:30:57",
      "content": "<p>Hey, <a href=\"/hengck23\">@hengck23</a>, I had tried that method in the past, did not work out very well, However that was a few weeks ago, I guess, I am gonna retry this thing, One mistake you would want to avoid is even trying to go for resnet-18, But thinking about it, you would have already got your mind which net would be replacing that. Also, I tried your note number 1 (to keep same videos/same frames/ different label in 1 batch) few days before, It was not performing any good, Right now, Can you answer me, Why these approaches should work, what is in your mind which makes you think it will work, Thank You!.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 736706,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/04/2020 13:34:56",
      "content": "<p>this is an example of inference code for beginner kaggler. it is based on @ Human Analog 's work: <a href=\"https://www.kaggle.com/humananalog/inference-demo\">https://www.kaggle.com/humananalog/inference-demo</a></p>\n\n<p>I refactor it for my own clarity. refer to the readme.ppt for details.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 737282,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/05/2020 06:09:06",
      "content": "<p>Loophole?</p>\n\n<p>I am not sure if both real and fake video will be present in the test or not. If so, one can cluster the same videos together. We know that only one of the video is real and the rest are manipulated. Or with many similar videos, there exist at most one or zero real version</p>",
      "votes": null,
      "replies": [
        {
          "id": 737346,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/05/2020 08:06:32",
          "content": "<p><a href=\"/hengck23\">@hengck23</a> Please check your email and reply.\nThank You!.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 738327,
      "author_name": "vineeth1999",
      "author_url": "",
      "post_date": "02/06/2020 11:47:45",
      "content": "<p>okay</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 740799,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "02/09/2020 20:10:43",
      "content": "<p>Just wondering when will the kernel will be made public?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 745748,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "02/14/2020 06:36:29",
      "content": "<p>thank you for the share. Always nice to have you in a comp 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 746081,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "02/14/2020 15:44:54",
      "content": "<p>For what it's worth, I trained the model from this paper on my dataset and it scores 0.42574 on the LB.</p>\n\n<p>The only changes I made are:</p>\n\n<ul>\n<li>Sigmoid instead of softmax, so the combined tensor is 5x1 instead of 10x1.</li>\n<li>Using 1-cycle training for 10 epochs with a maximum LR of 1e-3.</li>\n<li>Each batch consists of (real, fake) image pairs.</li>\n<li>My own data augmentation, etc.</li>\n<li>I used pretrained ResNet-18 from torchvision, and froze the model up to block 4.</li>\n</ul>\n\n<p>This took about 10 hrs to train on my 1080 Ti using 1M images (batch size 128). Final train loss is 0.2041, val loss is 0.1923 (but the val loss is fairly meaningless as my validation set is not very good; I only use it to make sure the model isn't overfitting).</p>",
      "votes": null,
      "replies": [
        {
          "id": 746224,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/14/2020 19:02:30",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> Thanks for the experiment!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746260,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/14/2020 19:47:14",
          "content": "<p>I think it might be useful to do this sort of thing anyway, as it's basically an ensemble of 5 different classifiers. The linear layer that you add to the end is for taking the weighted mean of their results. So it should improve existing models, although it will be 5x slower to train obviously.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746272,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/14/2020 19:54:33",
          "content": "<p>Do you have any log of how the Resnet18_1 performs (as it's a full face classifier) in your local val loss? Just curious to know how the ensemble helped in boosting!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746287,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/14/2020 20:02:30",
          "content": "<p>No, unfortunately I don't have any results for a single ResNet-18 model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747812,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/16/2020 22:18:48",
          "content": "<p>Again, I did haven't an experiment planned for today so I trained just a single ResNet-18 on my data and it scores 0.42936 on the LB. This is almost identical to the 5x ResNet-18 model. So it turns out the \"ensembling\" doesn't really help very much.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747824,
          "author_name": "ims0rry",
          "author_url": "",
          "post_date": "02/16/2020 22:42:21",
          "content": "<p>Sounds like a really good job with data if even a single resnet18 model gets such score on the LB</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747858,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/16/2020 23:56:40",
          "content": "<p><a href=\"/ims0rry\">@ims0rry</a> totally agree!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747877,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/17/2020 01:01:31",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> How were you able to score 0.42 on LB with single frame? Our best single frame was just 0.44. BTW, do consider teaming with us. A model stacking might be pretty helpful. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748302,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/17/2020 11:09:28",
          "content": "<blockquote>\n  <p>How were you able to score 0.42 on LB with single frame</p>\n</blockquote>\n\n<p>Ah, sorry if this wasn't clear. This score (and the others I posted in this thread) is done by averaging the predictions of 17 frames from each video. Using only a single frame would make the score worse, no doubt. (For fun, I'll try that now.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748333,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/17/2020 11:52:30",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> any reply for \"do consider teaming with us\", Can we team up?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748352,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/17/2020 12:26:44",
          "content": "<blockquote>\n  <p>Can we team up?</p>\n</blockquote>\n\n<p>Thanks for asking but I don't really know how much more time I can put into this competition, so it's probably not a good idea to have me in your team. 😉 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748513,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/17/2020 15:59:55",
          "content": "<blockquote>\n  <p>Using only a single frame would make the score worse, no doubt</p>\n</blockquote>\n\n<p>Exact same model, using just the middle frame from each video scores 0.60600 on the LB. So that's a massive difference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 747053,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "02/15/2020 22:22:32",
      "content": "<p>Since I wasn't running any other experiments, I decided to \"upgrade\" this model to use ResNet-34 instead of ResNet-18 and train it in exactly the same way. It only gave a minor improvement in the score on the LB test set: 0.41932 vs 0.42574. (On my local validation set the score was pretty much the same as before.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 747084,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/16/2020 00:05:57",
          "content": "<p>Well, That was expected, larger model will not be doing any good, you know.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747358,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/16/2020 10:30:34",
          "content": "<p>But when is a model too small and when is it too large? You don't know it until you try it. ;-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747494,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/16/2020 14:19:50",
          "content": "<p>Well, you can know when model is too small or too large by just seeing how the model reacts (without actually training.), EZ ;-)...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "735137": "objective is to implement a baseline system based on the paper:\nhttps://arxiv.org/pdf/2001.07444.pdf\nDetecting Face2Face Facial Reenactment in Videos\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F113660%2Fe6f0e559bacd9ebb937436ddd7909b68%2FSelection_051.png?generation=1580661047964292&amp;alt=media)\n\n\n...\nhttps://drive.google.com/open?id=1El6zsA0_k5IcW2XLwQRd6pAXshURrFo6\n\n(to be updated)",
    "735141": "a few notes:\n\n1. you are given train videos as :  a true video (identity A) and fake videos (identity A is swap to B, unfortunately video of B is not given). during training, you should sample frame in the pair, i.e both the same time frame of real and fake video must be present in a single batch \n\n2. you can \"force attention\" to the face region, i.e. the classifier should make decision based on the weird/unnatural face pixels and avoid background pixels (unless they contains information about inherit image noise signature, etc)\n\n3. there are several labels you can get from training set. e.g.:\n   - image label : is it from real or fake video\n   - pixel label : subtract fake frame from real frame and see if  the pixel is altered / manipulated\n   - block label : divide image into say 8x8 block.  subtract fake block from real block \n   - etc ... same idea apply in the temporal/frame difference/ time axis ...",
    "735240": "Placeholder comment",
    "735247": "[Placeholder] comment😄 😄",
    "735373": "placeholder fork - kidding, nice pointers for ideas already, thanks for sharing @hengck23",
    "735729": "StyleGAN2: Near-Perfect Human Face Synthesis...and More\nhttps://www.youtube.com/watch?v=SWoravHhsUU\n\nteeth and eye are the telltale signs",
    "735807": "Hey, @hengck23, I had tried that method in the past, did not work out very well, However that was a few weeks ago, I guess, I am gonna retry this thing, One mistake you would want to avoid is even trying to go for resnet-18, But thinking about it, you would have already got your mind which net would be replacing that. Also, I tried your note number 1 (to keep same videos/same frames/ different label in 1 batch) few days before, It was not performing any good, Right now, Can you answer me, Why these approaches should work, what is in your mind which makes you think it will work, Thank You!.",
    "735895": "Paper -&gt; https://arxiv.org/pdf/1912.04958.pdf",
    "736706": "this is an example of inference code for beginner kaggler. it is based on @ Human Analog 's work: https://www.kaggle.com/humananalog/inference-demo\n\n\nI refactor it for my own clarity. refer to the readme.ppt for details.",
    "737282": "Loophole?\n\nI am not sure if both real and fake video will be present in the test or not. If so, one can cluster the same videos together. We know that only one of the video is real and the rest are manipulated. Or with many similar videos, there exist at most one or zero real version",
    "737290": "Bascially not all frames  are “fake” ( or more correctly, looks “fake”)There needs to be temporal localisation method, just like action recognition in a video.\n\nProbably the static image classifier needs to pool over the whole video, i.e over the time domain as well for training a image feature extraction",
    "737346": "hengck23 Please check your email and reply.\nThank You!.",
    "738327": "okay",
    "738413": "Well, @hengck23, I would take this as \"Most images looks to be real in a fake video\" but are fake, using this believe is much better.",
    "740799": "Just wondering when will the kernel will be made public?",
    "741297": "Yes ,I  think  we  should  train  a  model  to    filter   some   import   frame , but  not  input  all  the  video  frame  or  just   random  select .",
    "745748": "thank you for the share. Always nice to have you in a comp 👍",
    "746081": "For what it's worth, I trained the model from this paper on my dataset and it scores 0.42574 on the LB.\n\nThe only changes I made are:\n\n- Sigmoid instead of softmax, so the combined tensor is 5x1 instead of 10x1.\n- Using 1-cycle training for 10 epochs with a maximum LR of 1e-3.\n- Each batch consists of (real, fake) image pairs.\n- My own data augmentation, etc.\n- I used pretrained ResNet-18 from torchvision, and froze the model up to block 4.\n\nThis took about 10 hrs to train on my 1080 Ti using 1M images (batch size 128). Final train loss is 0.2041, val loss is 0.1923 (but the val loss is fairly meaningless as my validation set is not very good; I only use it to make sure the model isn't overfitting).",
    "746224": "humananalog Thanks for the experiment!",
    "746260": "I think it might be useful to do this sort of thing anyway, as it's basically an ensemble of 5 different classifiers. The linear layer that you add to the end is for taking the weighted mean of their results. So it should improve existing models, although it will be 5x slower to train obviously.",
    "746272": "Do you have any log of how the Resnet18_1 performs (as it's a full face classifier) in your local val loss? Just curious to know how the ensemble helped in boosting!",
    "746287": "No, unfortunately I don't have any results for a single ResNet-18 model.",
    "747053": "Since I wasn't running any other experiments, I decided to \"upgrade\" this model to use ResNet-34 instead of ResNet-18 and train it in exactly the same way. It only gave a minor improvement in the score on the LB test set: 0.41932 vs 0.42574. (On my local validation set the score was pretty much the same as before.)",
    "747084": "Well, That was expected, larger model will not be doing any good, you know.",
    "747358": "But when is a model too small and when is it too large? You don't know it until you try it. ;-)",
    "747494": "Well, you can know when model is too small or too large by just seeing how the model reacts (without actually training.), EZ ;-)...",
    "747812": "Again, I did haven't an experiment planned for today so I trained just a single ResNet-18 on my data and it scores 0.42936 on the LB. This is almost identical to the 5x ResNet-18 model. So it turns out the \"ensembling\" doesn't really help very much.",
    "747824": "Sounds like a really good job with data if even a single resnet18 model gets such score on the LB",
    "747858": "ims0rry totally agree!",
    "747877": "humananalog How were you able to score 0.42 on LB with single frame? Our best single frame was just 0.44. BTW, do consider teaming with us. A model stacking might be pretty helpful.",
    "748040": "[Placeholder [Placeholder] funny placeholder comment.] ;)",
    "748302": "&gt; How were you able to score 0.42 on LB with single frame\n\nAh, sorry if this wasn't clear. This score (and the others I posted in this thread) is done by averaging the predictions of 17 frames from each video. Using only a single frame would make the score worse, no doubt. (For fun, I'll try that now.)",
    "748333": "humananalog any reply for \"do consider teaming with us\", Can we team up?",
    "748352": "&gt; Can we team up?\n\nThanks for asking but I don't really know how much more time I can put into this competition, so it's probably not a good idea to have me in your team. 😉",
    "748513": "&gt; Using only a single frame would make the score worse, no doubt\n\nExact same model, using just the middle frame from each video scores 0.60600 on the LB. So that's a massive difference.",
    "750782": "Get some idea: \n1. extract same frame(image) from REAL and FAKE video, \n2. apply a t-test to this pair images to decide if the image is significant different to each other.\n3. set a threshold to keep the most different pair image(REAL and FAKE) as dataset",
    "750818": "Good Idea, I might try it in near future. Unfortunately there is no fast way to do it, preprocessing to classify which image is the most different is a mess to do, but I think I have a new idea because of you so Thank You!"
  },
  "source": "meta"
}