{
  "id": 124583,
  "title": "Useful Model Structures",
  "url": "/competitions/deepfake-detection-challenge/discussion/124583",
  "author_name": "Shangqiu Li",
  "post_date": "2020-01-05T04:17:42.038000",
  "votes": 56,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Just sharing a list of model structures I discovered while looking along the papers(and code for some)\n- convolutional LSTM network according to this paper: <a href=\"https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf\">link</a>\nThey didn't provide code. According to their description,I made up this fake code:\nInput(n_frames,width,height,channels)-&gt;TimeDistributed(Conv2D(16,(5,5))<em>-&gt;TimeDistributed(Conv2D(32,(5,5))</em>-&gt;LSTM(2048)-&gt;Dropout(0.5)-&gt;Dense(512)-&gt;Dropout(0.5)-&gt;Dense(1,activation='softmax')</p>\n\n<p>*(parameters not specified in paper, I made it up)\nPS I saw lots of paper using similar structure too.\n- MesoNet: <a href=\"https://arxiv.org/abs/1809.00888\">link</a>\nThey provided a github repo: <a href=\"https://github.com/DariusAf/MesoNet\">link</a>\nI made a kernel too using pretrained weights and same structure: <a href=\"https://www.kaggle.com/unkownhihi/starter-kernel-with-cnn-ll-lb-0-69306-no-leak\">link</a></p>\n\n<ul>\n<li><p>Simple Xception, tested as the best performing network(for single image, not video) in this paper: <a href=\"https://arxiv.org/pdf/1901.08971.pdf\">link</a>\n&gt; We transfer it to our task by replacing the final\nfully connected layer with two outputs. The other layers are\ninitialized with the ImageNet weights. To set up the newly\ninserted fully connected layer, we fix all weights up to the final layers and pre-train the network for 3 epochs. After this\nstep, we train the network for 15 more epochs and choose\nthe best performing model based on validation accuracy.</p></li>\n<li><p>SVM+preprocessing: <a href=\"https://arxiv.org/pdf/1911.00686.pdf\">link</a></p></li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ffb40ebd86419c8495a00d799e412e380%2FScreen%20Shot%202020-01-04%20at%208.35.27%20PM.png?generation=1578198953525556&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Capsule Network: <a href=\"https://arxiv.org/pdf/1910.12467.pdf\">link</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F39952c3af57566a575e1f36ff706c8d7%2FScreen%20Shot%202020-01-05%20at%2010.02.24%20AM.png?generation=1578247420455320&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>Here are some GitHub projects: \nKeras Implementation <a href=\"https://github.com/XifengGuo/CapsNet-Keras\">link</a>\nPytorch Implementation <a href=\"https://github.com/nii-yamagishilab/Capsule-Forensics-v2\">link</a> mentioned by <a href=\"/bibek777\">@bibek777</a> \nTensorflow Implementation <a href=\"https://github.com/zhanpenghe/Capsule-Network\">link</a></p>\n\n<ul>\n<li>Blink Detection: <a href=\"https://arxiv.org/pdf/1806.02877.pdf\">link</a>\nSummary:\n&gt;Our current method only uses the lack of\nblinking as a cue for detection. However, the dynamic pattern\nof blinking should also be considered – too fast or frequent\nblinking that is deemed physiologically unlikely could also be\na sign of tampering. Finally, eye blinking is a relatively easy\ncue in detecting fake face videos, and sophisticated forgers can\nstill create realistic blinking effects with post-processing and\nmore advanced models and more training data. So in the long\nrun, we are interested in exploring other types of physiological\nsignals to detect fake videos.</li>\n</ul>\n\n<p>Model(LRCN):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fd6d97959668a57b2b352d4d9111a9ee0%2FScreen%20Shot%202020-01-06%20at%204.35.30%20PM.png?generation=1578357386995711&amp;alt=media\" alt=\"\"></p>\n\n<p>Blink Detection Dataset available at <a href=\"http://www.cs.albany.edu/~lsw/downloads.html\">link</a></p>\n\n<p>I will update as I discover more.</p>",
  "messages": [
    {
      "id": 710674,
      "postDate": "2020-01-05T04:17:42.040Z",
      "content": "<p>Just sharing a list of model structures I discovered while looking along the papers(and code for some)\n- convolutional LSTM network according to this paper: <a href=\"https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf\">link</a>\nThey didn't provide code. According to their description,I made up this fake code:\nInput(n_frames,width,height,channels)-&gt;TimeDistributed(Conv2D(16,(5,5))<em>-&gt;TimeDistributed(Conv2D(32,(5,5))</em>-&gt;LSTM(2048)-&gt;Dropout(0.5)-&gt;Dense(512)-&gt;Dropout(0.5)-&gt;Dense(1,activation='softmax')</p>\n\n<p>*(parameters not specified in paper, I made it up)\nPS I saw lots of paper using similar structure too.\n- MesoNet: <a href=\"https://arxiv.org/abs/1809.00888\">link</a>\nThey provided a github repo: <a href=\"https://github.com/DariusAf/MesoNet\">link</a>\nI made a kernel too using pretrained weights and same structure: <a href=\"https://www.kaggle.com/unkownhihi/starter-kernel-with-cnn-ll-lb-0-69306-no-leak\">link</a></p>\n\n<ul>\n<li><p>Simple Xception, tested as the best performing network(for single image, not video) in this paper: <a href=\"https://arxiv.org/pdf/1901.08971.pdf\">link</a>\n&gt; We transfer it to our task by replacing the final\nfully connected layer with two outputs. The other layers are\ninitialized with the ImageNet weights. To set up the newly\ninserted fully connected layer, we fix all weights up to the final layers and pre-train the network for 3 epochs. After this\nstep, we train the network for 15 more epochs and choose\nthe best performing model based on validation accuracy.</p></li>\n<li><p>SVM+preprocessing: <a href=\"https://arxiv.org/pdf/1911.00686.pdf\">link</a></p></li>\n</ul>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ffb40ebd86419c8495a00d799e412e380%2FScreen%20Shot%202020-01-04%20at%208.35.27%20PM.png?generation=1578198953525556&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>Capsule Network: <a href=\"https://arxiv.org/pdf/1910.12467.pdf\">link</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F39952c3af57566a575e1f36ff706c8d7%2FScreen%20Shot%202020-01-05%20at%2010.02.24%20AM.png?generation=1578247420455320&amp;alt=media\" alt=\"\"></li>\n</ul>\n\n<p>Here are some GitHub projects: \nKeras Implementation <a href=\"https://github.com/XifengGuo/CapsNet-Keras\">link</a>\nPytorch Implementation <a href=\"https://github.com/nii-yamagishilab/Capsule-Forensics-v2\">link</a> mentioned by <a href=\"/bibek777\">@bibek777</a> \nTensorflow Implementation <a href=\"https://github.com/zhanpenghe/Capsule-Network\">link</a></p>\n\n<ul>\n<li>Blink Detection: <a href=\"https://arxiv.org/pdf/1806.02877.pdf\">link</a>\nSummary:\n&gt;Our current method only uses the lack of\nblinking as a cue for detection. However, the dynamic pattern\nof blinking should also be considered – too fast or frequent\nblinking that is deemed physiologically unlikely could also be\na sign of tampering. Finally, eye blinking is a relatively easy\ncue in detecting fake face videos, and sophisticated forgers can\nstill create realistic blinking effects with post-processing and\nmore advanced models and more training data. So in the long\nrun, we are interested in exploring other types of physiological\nsignals to detect fake videos.</li>\n</ul>\n\n<p>Model(LRCN):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fd6d97959668a57b2b352d4d9111a9ee0%2FScreen%20Shot%202020-01-06%20at%204.35.30%20PM.png?generation=1578357386995711&amp;alt=media\" alt=\"\"></p>\n\n<p>Blink Detection Dataset available at <a href=\"http://www.cs.albany.edu/~lsw/downloads.html\">link</a></p>\n\n<p>I will update as I discover more.</p>",
      "rawMarkdown": "Just sharing a list of model structures I discovered while looking along the papers(and code for some)\n- convolutional LSTM network according to this paper: [link](https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf)\nThey didn't provide code. According to their description,I made up this fake code:\nInput(n_frames,width,height,channels)-&gt;TimeDistributed(Conv2D(16,(5,5))*-&gt;TimeDistributed(Conv2D(32,(5,5))*-&gt;LSTM(2048)-&gt;Dropout(0.5)-&gt;Dense(512)-&gt;Dropout(0.5)-&gt;Dense(1,activation='softmax')\n\n*(parameters not specified in paper, I made it up)\nPS I saw lots of paper using similar structure too.\n- MesoNet: [link](https://arxiv.org/abs/1809.00888)\nThey provided a github repo: [link](https://github.com/DariusAf/MesoNet)\nI made a kernel too using pretrained weights and same structure: [link](https://www.kaggle.com/unkownhihi/starter-kernel-with-cnn-ll-lb-0-69306-no-leak)\n\n- Simple Xception, tested as the best performing network(for single image, not video) in this paper: [link](https://arxiv.org/pdf/1901.08971.pdf)\n&gt; We transfer it to our task by replacing the final\nfully connected layer with two outputs. The other layers are\ninitialized with the ImageNet weights. To set up the newly\ninserted fully connected layer, we fix all weights up to the final layers and pre-train the network for 3 epochs. After this\nstep, we train the network for 15 more epochs and choose\nthe best performing model based on validation accuracy.\n\n- SVM+preprocessing: [link](https://arxiv.org/pdf/1911.00686.pdf)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ffb40ebd86419c8495a00d799e412e380%2FScreen%20Shot%202020-01-04%20at%208.35.27%20PM.png?generation=1578198953525556&amp;alt=media)\n\n- Capsule Network: [link](https://arxiv.org/pdf/1910.12467.pdf)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F39952c3af57566a575e1f36ff706c8d7%2FScreen%20Shot%202020-01-05%20at%2010.02.24%20AM.png?generation=1578247420455320&amp;alt=media)\n\nHere are some GitHub projects: \nKeras Implementation [link](https://github.com/XifengGuo/CapsNet-Keras)\nPytorch Implementation [link](https://github.com/nii-yamagishilab/Capsule-Forensics-v2) mentioned by @bibek777 \nTensorflow Implementation [link](https://github.com/zhanpenghe/Capsule-Network)\n\n- Blink Detection: [link](https://arxiv.org/pdf/1806.02877.pdf)\nSummary:\n&gt;Our current method only uses the lack of\nblinking as a cue for detection. However, the dynamic pattern\nof blinking should also be considered – too fast or frequent\nblinking that is deemed physiologically unlikely could also be\na sign of tampering. Finally, eye blinking is a relatively easy\ncue in detecting fake face videos, and sophisticated forgers can\nstill create realistic blinking effects with post-processing and\nmore advanced models and more training data. So in the long\nrun, we are interested in exploring other types of physiological\nsignals to detect fake videos.\n\nModel(LRCN):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fd6d97959668a57b2b352d4d9111a9ee0%2FScreen%20Shot%202020-01-06%20at%204.35.30%20PM.png?generation=1578357386995711&amp;alt=media)\n\nBlink Detection Dataset available at [link](http://www.cs.albany.edu/~lsw/downloads.html)\n\nI will update as I discover more.",
      "votes": 54
    },
    {
      "id": 717452,
      "postDate": "2020-01-13T06:03:46.807Z",
      "content": "<p>here is a pytorch implementation of <a href=\"https://github.com/nii-yamagishilab/Capsule-Forensics-v2\">Capsule Network</a></p>",
      "rawMarkdown": "here is a pytorch implementation of [Capsule Network](https://github.com/nii-yamagishilab/Capsule-Forensics-v2)",
      "votes": 3,
      "replies": [
        {
          "id": 718027,
          "postDate": "2020-01-13T22:46:59.910Z",
          "content": "<p>Thanks. Added to the post.</p>",
          "rawMarkdown": "Thanks. Added to the post."
        }
      ]
    },
    {
      "id": 710735,
      "postDate": "2020-01-05T06:28:16.857Z",
      "content": "<p>Nice </p>",
      "rawMarkdown": "Nice ",
      "votes": 3
    },
    {
      "id": 756746,
      "postDate": "2020-02-26T03:08:31.567Z",
      "content": "<p>Nice list, the idea is very interesting. I would like to add more in terms of emotional intelligence which would help overall.</p>",
      "rawMarkdown": "Nice list, the idea is very interesting. I would like to add more in terms of emotional intelligence which would help overall.",
      "votes": 1
    },
    {
      "id": 711573,
      "postDate": "2020-01-06T09:07:53.270Z",
      "content": "<p>Nice </p>",
      "rawMarkdown": "Nice ",
      "votes": 1
    },
    {
      "id": 711371,
      "postDate": "2020-01-06T01:51:57.523Z",
      "content": "<p>I just discoverd that <a href=\"/robikscube\">@robikscube</a> made a great kernel using xception(weights: FaceForensics++)<a href=\"https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet#Use-the-classification-package-per-the-github-instructions\">link</a></p>",
      "rawMarkdown": "I just discoverd that @robikscube made a great kernel using xception(weights: FaceForensics++)[link](https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet#Use-the-classification-package-per-the-github-instructions)",
      "votes": 1,
      "replies": [
        {
          "id": 752856,
          "postDate": "2020-02-21T13:30:21.430Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a>  how did u initialize the LRCN ? and which git hub repo did u use ?</p>",
          "rawMarkdown": "@unkownhihi  how did u initialize the LRCN ? and which git hub repo did u use ?"
        },
        {
          "id": 753008,
          "postDate": "2020-02-21T16:01:43.803Z",
          "content": "<p>Sample code:\n<code>\ninp=Input((10,244,244,3))\nx=TimeDistributed(backbone)(inp)\nx=LSTM(256)(x)\nx=Dense(128,activation='relu')\nx=Dense(1,activation='sigmoid')\n</code>\nbackbone is any keras.application model or your own CNN model.</p>",
          "rawMarkdown": "Sample code:\n```\ninp=Input((10,244,244,3))\nx=TimeDistributed(backbone)(inp)\nx=LSTM(256)(x)\nx=Dense(128,activation='relu')\nx=Dense(1,activation='sigmoid')\n```\nbackbone is any keras.application model or your own CNN model."
        }
      ]
    },
    {
      "id": 710720,
      "postDate": "2020-01-05T05:57:46.653Z",
      "content": "<p>Thanks a lot! A great Kaggle conversation starter :)</p>",
      "rawMarkdown": "Thanks a lot! A great Kaggle conversation starter :)",
      "votes": 1
    },
    {
      "id": 730718,
      "postDate": "2020-01-27T20:23:34.597Z",
      "content": "<p><a href=\"https://arxiv.org/pdf/2001.06232.pdf\">https://arxiv.org/pdf/2001.06232.pdf</a>\nSideways: Depth-Parallel Training of Video Models</p>\n\n<p><a href=\"https://twitter.com/i/videos/1219305774570180615?embed_source=facebook\">https://twitter.com/i/videos/1219305774570180615?embed_source=facebook</a></p>",
      "rawMarkdown": "https://arxiv.org/pdf/2001.06232.pdf\nSideways: Depth-Parallel Training of Video Models\n\n\nhttps://twitter.com/i/videos/1219305774570180615?embed_source=facebook",
      "votes": 2
    },
    {
      "id": 763952,
      "postDate": "2020-03-05T02:39:25.140Z",
      "content": "<p>Thanks a lot. Very helpful reference! :)</p>",
      "rawMarkdown": "Thanks a lot. Very helpful reference! :)"
    },
    {
      "id": 748029,
      "postDate": "2020-02-17T05:58:59.783Z",
      "content": "<p>Awesome list of models and approaches. The blink detection approach is very interesting, I think it can be combined with other physiological traits as mentioned: mouth movements, face expressions, and so on. </p>",
      "rawMarkdown": "Awesome list of models and approaches. The blink detection approach is very interesting, I think it can be combined with other physiological traits as mentioned: mouth movements, face expressions, and so on. "
    },
    {
      "id": 717426,
      "postDate": "2020-01-13T05:04:56.087Z",
      "content": "<p>Nice</p>",
      "rawMarkdown": "Nice"
    },
    {
      "id": 745053,
      "postDate": "2020-02-13T12:37:18.033Z",
      "content": "<p>Thank you for sharing !!!\n I wonder what image size do you use in CNNLSTM. I tried (None, 17, 299, 299, 3) in which 17 is frame number and 299 is image size, but in prediction Kaggle memory is not enough for the parameters. </p>",
      "rawMarkdown": "Thank you for sharing !!!\n I wonder what image size do you use in CNNLSTM. I tried (None, 17, 299, 299, 3) in which 17 is frame number and 299 is image size, but in prediction Kaggle memory is not enough for the parameters. ",
      "isDeleted": true,
      "replies": [
        {
          "id": 745318,
          "postDate": "2020-02-13T17:52:35.027Z",
          "content": "<p>I did (10,240,240,3). batch size of 4.</p>",
          "rawMarkdown": "I did (10,240,240,3). batch size of 4.",
          "votes": 1
        },
        {
          "id": 745332,
          "postDate": "2020-02-13T18:08:37.573Z",
          "content": "<p>I will try. Thank you !!!</p>",
          "rawMarkdown": "I will try. Thank you !!!",
          "isDeleted": true
        },
        {
          "id": 759008,
          "postDate": "2020-02-28T12:53:33.877Z",
          "content": "<p>Hi Shark, would you like to team up to continue working on this? I'm from China as well.</p>",
          "rawMarkdown": "Hi Shark, would you like to team up to continue working on this? I'm from China as well."
        }
      ]
    },
    {
      "id": 712211,
      "postDate": "2020-01-07T00:17:23.963Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 724183,
      "postDate": "2020-01-20T22:00:21.097Z",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!"
    },
    {
      "id": 717764,
      "postDate": "2020-01-13T15:39:01.280Z",
      "content": "<p>Thanks for the useful information</p>",
      "rawMarkdown": "Thanks for the useful information"
    }
  ],
  "comments": [
    {
      "id": 717452,
      "author_name": "Bibek",
      "author_url": "",
      "post_date": "2020-01-13T06:03:46.807000",
      "content": "<p>here is a pytorch implementation of <a href=\"https://github.com/nii-yamagishilab/Capsule-Forensics-v2\">Capsule Network</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 718027,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-13T22:46:59.910000",
          "content": "<p>Thanks. Added to the post.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 710735,
      "author_name": "Raju Kumar Mishra",
      "author_url": "",
      "post_date": "2020-01-05T06:28:16.857000",
      "content": "<p>Nice </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 756746,
      "author_name": "Pritesh Raj",
      "author_url": "",
      "post_date": "2020-02-26T03:08:31.567000",
      "content": "<p>Nice list, the idea is very interesting. I would like to add more in terms of emotional intelligence which would help overall.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 711573,
      "author_name": "AtulVerma",
      "author_url": "",
      "post_date": "2020-01-06T09:07:53.270000",
      "content": "<p>Nice </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 711371,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-01-06T01:51:57.523000",
      "content": "<p>I just discoverd that <a href=\"/robikscube\">@robikscube</a> made a great kernel using xception(weights: FaceForensics++)<a href=\"https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet#Use-the-classification-package-per-the-github-instructions\">link</a></p>",
      "votes": 1,
      "replies": [
        {
          "id": 752856,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-02-21T13:30:21.430000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a>  how did u initialize the LRCN ? and which git hub repo did u use ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 753008,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-21T16:01:43.803000",
          "content": "<p>Sample code:\n<code>\ninp=Input((10,244,244,3))\nx=TimeDistributed(backbone)(inp)\nx=LSTM(256)(x)\nx=Dense(128,activation='relu')\nx=Dense(1,activation='sigmoid')\n</code>\nbackbone is any keras.application model or your own CNN model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 710720,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-01-05T05:57:46.653000",
      "content": "<p>Thanks a lot! A great Kaggle conversation starter :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 730718,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-01-27T20:23:34.597000",
      "content": "<p><a href=\"https://arxiv.org/pdf/2001.06232.pdf\">https://arxiv.org/pdf/2001.06232.pdf</a>\nSideways: Depth-Parallel Training of Video Models</p>\n\n<p><a href=\"https://twitter.com/i/videos/1219305774570180615?embed_source=facebook\">https://twitter.com/i/videos/1219305774570180615?embed_source=facebook</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 763952,
      "author_name": "nourange",
      "author_url": "",
      "post_date": "2020-03-05T02:39:25.140000",
      "content": "<p>Thanks a lot. Very helpful reference! :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 748029,
      "author_name": "Yassine Alouini",
      "author_url": "",
      "post_date": "2020-02-17T05:58:59.783000",
      "content": "<p>Awesome list of models and approaches. The blink detection approach is very interesting, I think it can be combined with other physiological traits as mentioned: mouth movements, face expressions, and so on. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 717426,
      "author_name": "Fasato",
      "author_url": "",
      "post_date": "2020-01-13T05:04:56.087000",
      "content": "<p>Nice</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 745053,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-13T12:37:18.033000",
      "content": "<p>Thank you for sharing !!!\n I wonder what image size do you use in CNNLSTM. I tried (None, 17, 299, 299, 3) in which 17 is frame number and 299 is image size, but in prediction Kaggle memory is not enough for the parameters. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 745318,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-13T17:52:35.027000",
          "content": "<p>I did (10,240,240,3). batch size of 4.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 745332,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-13T18:08:37.573000",
          "content": "<p>I will try. Thank you !!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 759008,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-28T12:53:33.877000",
          "content": "<p>Hi Shark, would you like to team up to continue working on this? I'm from China as well.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 712211,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-07T00:17:23.963000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 724183,
      "author_name": "Nikita Karaev",
      "author_url": "",
      "post_date": "2020-01-20T22:00:21.097000",
      "content": "<p>Thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 717764,
      "author_name": "Sai Srinivas Reddy",
      "author_url": "",
      "post_date": "2020-01-13T15:39:01.280000",
      "content": "<p>Thanks for the useful information</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "710674": "Just sharing a list of model structures I discovered while looking along the papers(and code for some)\n- convolutional LSTM network according to this paper: [link](https://engineering.purdue.edu/~dgueraco/content/deepfake.pdf)\nThey didn't provide code. According to their description,I made up this fake code:\nInput(n_frames,width,height,channels)-&gt;TimeDistributed(Conv2D(16,(5,5))*-&gt;TimeDistributed(Conv2D(32,(5,5))*-&gt;LSTM(2048)-&gt;Dropout(0.5)-&gt;Dense(512)-&gt;Dropout(0.5)-&gt;Dense(1,activation='softmax')\n\n*(parameters not specified in paper, I made it up)\nPS I saw lots of paper using similar structure too.\n- MesoNet: [link](https://arxiv.org/abs/1809.00888)\nThey provided a github repo: [link](https://github.com/DariusAf/MesoNet)\nI made a kernel too using pretrained weights and same structure: [link](https://www.kaggle.com/unkownhihi/starter-kernel-with-cnn-ll-lb-0-69306-no-leak)\n\n- Simple Xception, tested as the best performing network(for single image, not video) in this paper: [link](https://arxiv.org/pdf/1901.08971.pdf)\n&gt; We transfer it to our task by replacing the final\nfully connected layer with two outputs. The other layers are\ninitialized with the ImageNet weights. To set up the newly\ninserted fully connected layer, we fix all weights up to the final layers and pre-train the network for 3 epochs. After this\nstep, we train the network for 15 more epochs and choose\nthe best performing model based on validation accuracy.\n\n- SVM+preprocessing: [link](https://arxiv.org/pdf/1911.00686.pdf)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Ffb40ebd86419c8495a00d799e412e380%2FScreen%20Shot%202020-01-04%20at%208.35.27%20PM.png?generation=1578198953525556&amp;alt=media)\n\n- Capsule Network: [link](https://arxiv.org/pdf/1910.12467.pdf)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2F39952c3af57566a575e1f36ff706c8d7%2FScreen%20Shot%202020-01-05%20at%2010.02.24%20AM.png?generation=1578247420455320&amp;alt=media)\n\nHere are some GitHub projects: \nKeras Implementation [link](https://github.com/XifengGuo/CapsNet-Keras)\nPytorch Implementation [link](https://github.com/nii-yamagishilab/Capsule-Forensics-v2) mentioned by @bibek777 \nTensorflow Implementation [link](https://github.com/zhanpenghe/Capsule-Network)\n\n- Blink Detection: [link](https://arxiv.org/pdf/1806.02877.pdf)\nSummary:\n&gt;Our current method only uses the lack of\nblinking as a cue for detection. However, the dynamic pattern\nof blinking should also be considered – too fast or frequent\nblinking that is deemed physiologically unlikely could also be\na sign of tampering. Finally, eye blinking is a relatively easy\ncue in detecting fake face videos, and sophisticated forgers can\nstill create realistic blinking effects with post-processing and\nmore advanced models and more training data. So in the long\nrun, we are interested in exploring other types of physiological\nsignals to detect fake videos.\n\nModel(LRCN):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3319985%2Fd6d97959668a57b2b352d4d9111a9ee0%2FScreen%20Shot%202020-01-06%20at%204.35.30%20PM.png?generation=1578357386995711&amp;alt=media)\n\nBlink Detection Dataset available at [link](http://www.cs.albany.edu/~lsw/downloads.html)\n\nI will update as I discover more.",
    "717452": "here is a pytorch implementation of [Capsule Network](https://github.com/nii-yamagishilab/Capsule-Forensics-v2)",
    "710735": "Nice ",
    "756746": "Nice list, the idea is very interesting. I would like to add more in terms of emotional intelligence which would help overall.",
    "711573": "Nice ",
    "711371": "I just discoverd that @robikscube made a great kernel using xception(weights: FaceForensics++)[link](https://www.kaggle.com/robikscube/faceforensics-baseline-dlib-no-internet#Use-the-classification-package-per-the-github-instructions)",
    "710720": "Thanks a lot! A great Kaggle conversation starter :)",
    "730718": "https://arxiv.org/pdf/2001.06232.pdf\nSideways: Depth-Parallel Training of Video Models\n\n\nhttps://twitter.com/i/videos/1219305774570180615?embed_source=facebook",
    "763952": "Thanks a lot. Very helpful reference! :)",
    "748029": "Awesome list of models and approaches. The blink detection approach is very interesting, I think it can be combined with other physiological traits as mentioned: mouth movements, face expressions, and so on. ",
    "717426": "Nice",
    "745053": "Thank you for sharing !!!\n I wonder what image size do you use in CNNLSTM. I tried (None, 17, 299, 299, 3) in which 17 is frame number and 299 is image size, but in prediction Kaggle memory is not enough for the parameters. ",
    "712211": "",
    "724183": "Thank you for sharing!",
    "717764": "Thanks for the useful information"
  }
}