{
  "id": 135441,
  "title": "Class Activations on ResNext50",
  "url": "/competitions/deepfake-detection-challenge/discussion/135441",
  "author_name": "",
  "post_date": "2020-03-13T21:48:18.452905800Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The basic intuition of this algorithm is that the model must have used some pixels to identify the image class. This is based on the following paper - <a href=\"https://arxiv.org/pdf/1610.02391.pdf\">Grad-CAM visualizations</a>  </p>\n\n<p>You can browse through the notebook over here - <a href=\"https://www.kaggle.com/skylord/grad-cam-on-resnext\">Link to kernel</a> </p>\n\n<p>After going over a few samples, it looks like the model is using facial features like nose, ears, eyes etc. to classify the faces.  So is the model identifying the deepfake artifacts or just noise? </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff1cb5a298387f7c826ee0009446285c0%2F__results___33_0.png?generation=1584135632083106&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8354c97a9fe69f3b785508d23e8531bb%2Fdownload.png?generation=1584135719614810&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb972b658ac15ec0f1977668c36bc844%2Fdownload%20(1\" alt=\"\">.png?generation=1584135760984220&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff5a113755a82a3b61603b65d3df3f7dd%2Fdownload%20(2\" alt=\"\">.png?generation=1584135851598595&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fc24b5b9a9e05d566da70a9ac93b4023d%2Fdownload%20(3\" alt=\"\">.png?generation=1584135889179734&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F9a84e9512c8200b01c80fd1509f774ea%2Fdownload%20(4\" alt=\"\">.png?generation=1584135929242126&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fd34c414124ff3366f5920aeb44f64234%2Fdownload%20(5\" alt=\"\">.png?generation=1584135958664344&amp;alt=media)</p>",
  "messages": [
    {
      "id": "771213",
      "postDate": "03/13/2020 21:48:18",
      "content": "<p>The basic intuition of this algorithm is that the model must have used some pixels to identify the image class. This is based on the following paper - <a href=\"https://arxiv.org/pdf/1610.02391.pdf\">Grad-CAM visualizations</a>  </p>\n\n<p>You can browse through the notebook over here - <a href=\"https://www.kaggle.com/skylord/grad-cam-on-resnext\">Link to kernel</a> </p>\n\n<p>After going over a few samples, it looks like the model is using facial features like nose, ears, eyes etc. to classify the faces.  So is the model identifying the deepfake artifacts or just noise? </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff1cb5a298387f7c826ee0009446285c0%2F__results___33_0.png?generation=1584135632083106&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8354c97a9fe69f3b785508d23e8531bb%2Fdownload.png?generation=1584135719614810&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb972b658ac15ec0f1977668c36bc844%2Fdownload%20(1\" alt=\"\">.png?generation=1584135760984220&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff5a113755a82a3b61603b65d3df3f7dd%2Fdownload%20(2\" alt=\"\">.png?generation=1584135851598595&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fc24b5b9a9e05d566da70a9ac93b4023d%2Fdownload%20(3\" alt=\"\">.png?generation=1584135889179734&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F9a84e9512c8200b01c80fd1509f774ea%2Fdownload%20(4\" alt=\"\">.png?generation=1584135929242126&amp;alt=media)</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fd34c414124ff3366f5920aeb44f64234%2Fdownload%20(5\" alt=\"\">.png?generation=1584135958664344&amp;alt=media)</p>",
      "rawMarkdown": "The basic intuition of this algorithm is that the model must have used some pixels to identify the image class. This is based on the following paper - [Grad-CAM visualizations](https://arxiv.org/pdf/1610.02391.pdf)  \n\nYou can browse through the notebook over here - [Link to kernel](https://www.kaggle.com/skylord/grad-cam-on-resnext) \n\nAfter going over a few samples, it looks like the model is using facial features like nose, ears, eyes etc. to classify the faces.  So is the model identifying the deepfake artifacts or just noise? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff1cb5a298387f7c826ee0009446285c0%2F__results___33_0.png?generation=1584135632083106&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8354c97a9fe69f3b785508d23e8531bb%2Fdownload.png?generation=1584135719614810&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb972b658ac15ec0f1977668c36bc844%2Fdownload%20(1).png?generation=1584135760984220&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff5a113755a82a3b61603b65d3df3f7dd%2Fdownload%20(2).png?generation=1584135851598595&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fc24b5b9a9e05d566da70a9ac93b4023d%2Fdownload%20(3).png?generation=1584135889179734&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F9a84e9512c8200b01c80fd1509f774ea%2Fdownload%20(4).png?generation=1584135929242126&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fd34c414124ff3366f5920aeb44f64234%2Fdownload%20(5).png?generation=1584135958664344&amp;alt=media)",
      "votes": null
    },
    {
      "id": "771481",
      "postDate": "03/14/2020 08:11:59",
      "content": "<p>Did a few more runs on the videos where the loss is high. The class activations for the top k  mis-classified videos is as follows (k=10): </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F30b5d4b25e5c90da8efb0a48834d0ca4%2FHIghDiff_04.png?generation=1584173469546915&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fb6b39cfa5283365cb06e221136c95083%2FHIghDiff_03.png?generation=1584173469367546&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb90df81a971e16bb354d38bba301d50%2FHighReal_0.png?generation=1584173215785085&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F4da7996109a4fdebc9edd2c261a6f7ff%2FHIghDiff_00.png?generation=1584173265337593&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fe0a13c7fd52b9f7891837718d93f0b05%2FHIghDiff_02.png?generation=1584173307222628&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F368499f96929a96212712239b0e2bde5%2FHighReal_12.png?generation=1584173340636721&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F60a23357b01a1b3ccd9795d2182ea164%2FHighReal_10.png?generation=1584173376716114&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F651dffdd5bdcf393c7b010c641fdd43b%2FHIghDiff_01.png?generation=1584173412740079&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8235f52380e849c4d1464983c43327a1%2FHighReal_09.png?generation=1584173512865681&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F57407660c162d40f2a59f4ebc15728e4%2FHighReal_16.png?generation=1584173515119630&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Did a few more runs on the videos where the loss is high. The class activations for the top k  mis-classified videos is as follows (k=10): \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F30b5d4b25e5c90da8efb0a48834d0ca4%2FHIghDiff_04.png?generation=1584173469546915&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fb6b39cfa5283365cb06e221136c95083%2FHIghDiff_03.png?generation=1584173469367546&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb90df81a971e16bb354d38bba301d50%2FHighReal_0.png?generation=1584173215785085&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F4da7996109a4fdebc9edd2c261a6f7ff%2FHIghDiff_00.png?generation=1584173265337593&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fe0a13c7fd52b9f7891837718d93f0b05%2FHIghDiff_02.png?generation=1584173307222628&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F368499f96929a96212712239b0e2bde5%2FHighReal_12.png?generation=1584173340636721&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F60a23357b01a1b3ccd9795d2182ea164%2FHighReal_10.png?generation=1584173376716114&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F651dffdd5bdcf393c7b010c641fdd43b%2FHIghDiff_01.png?generation=1584173412740079&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8235f52380e849c4d1464983c43327a1%2FHighReal_09.png?generation=1584173512865681&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F57407660c162d40f2a59f4ebc15728e4%2FHighReal_16.png?generation=1584173515119630&amp;alt=media)",
      "votes": null
    },
    {
      "id": "771910",
      "postDate": "03/14/2020 19:32:49",
      "content": "<p>Nice. I am thinking of passing the heat map output to another CNN. It works similar to attention.</p>",
      "rawMarkdown": "Nice. I am thinking of passing the heat map output to another CNN. It works similar to attention.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 771481,
      "author_name": "skylord",
      "author_url": "",
      "post_date": "03/14/2020 08:11:59",
      "content": "<p>Did a few more runs on the videos where the loss is high. The class activations for the top k  mis-classified videos is as follows (k=10): </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F30b5d4b25e5c90da8efb0a48834d0ca4%2FHIghDiff_04.png?generation=1584173469546915&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fb6b39cfa5283365cb06e221136c95083%2FHIghDiff_03.png?generation=1584173469367546&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb90df81a971e16bb354d38bba301d50%2FHighReal_0.png?generation=1584173215785085&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F4da7996109a4fdebc9edd2c261a6f7ff%2FHIghDiff_00.png?generation=1584173265337593&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fe0a13c7fd52b9f7891837718d93f0b05%2FHIghDiff_02.png?generation=1584173307222628&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F368499f96929a96212712239b0e2bde5%2FHighReal_12.png?generation=1584173340636721&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F60a23357b01a1b3ccd9795d2182ea164%2FHighReal_10.png?generation=1584173376716114&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F651dffdd5bdcf393c7b010c641fdd43b%2FHIghDiff_01.png?generation=1584173412740079&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8235f52380e849c4d1464983c43327a1%2FHighReal_09.png?generation=1584173512865681&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F57407660c162d40f2a59f4ebc15728e4%2FHighReal_16.png?generation=1584173515119630&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 771910,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "03/14/2020 19:32:49",
      "content": "<p>Nice. I am thinking of passing the heat map output to another CNN. It works similar to attention.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "771213": "The basic intuition of this algorithm is that the model must have used some pixels to identify the image class. This is based on the following paper - [Grad-CAM visualizations](https://arxiv.org/pdf/1610.02391.pdf)  \n\nYou can browse through the notebook over here - [Link to kernel](https://www.kaggle.com/skylord/grad-cam-on-resnext) \n\nAfter going over a few samples, it looks like the model is using facial features like nose, ears, eyes etc. to classify the faces.  So is the model identifying the deepfake artifacts or just noise? \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff1cb5a298387f7c826ee0009446285c0%2F__results___33_0.png?generation=1584135632083106&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8354c97a9fe69f3b785508d23e8531bb%2Fdownload.png?generation=1584135719614810&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb972b658ac15ec0f1977668c36bc844%2Fdownload%20(1).png?generation=1584135760984220&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Ff5a113755a82a3b61603b65d3df3f7dd%2Fdownload%20(2).png?generation=1584135851598595&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fc24b5b9a9e05d566da70a9ac93b4023d%2Fdownload%20(3).png?generation=1584135889179734&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F9a84e9512c8200b01c80fd1509f774ea%2Fdownload%20(4).png?generation=1584135929242126&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fd34c414124ff3366f5920aeb44f64234%2Fdownload%20(5).png?generation=1584135958664344&amp;alt=media)",
    "771481": "Did a few more runs on the videos where the loss is high. The class activations for the top k  mis-classified videos is as follows (k=10): \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F30b5d4b25e5c90da8efb0a48834d0ca4%2FHIghDiff_04.png?generation=1584173469546915&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fb6b39cfa5283365cb06e221136c95083%2FHIghDiff_03.png?generation=1584173469367546&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fdb90df81a971e16bb354d38bba301d50%2FHighReal_0.png?generation=1584173215785085&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F4da7996109a4fdebc9edd2c261a6f7ff%2FHIghDiff_00.png?generation=1584173265337593&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2Fe0a13c7fd52b9f7891837718d93f0b05%2FHIghDiff_02.png?generation=1584173307222628&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F368499f96929a96212712239b0e2bde5%2FHighReal_12.png?generation=1584173340636721&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F60a23357b01a1b3ccd9795d2182ea164%2FHighReal_10.png?generation=1584173376716114&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F651dffdd5bdcf393c7b010c641fdd43b%2FHIghDiff_01.png?generation=1584173412740079&amp;alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F8235f52380e849c4d1464983c43327a1%2FHighReal_09.png?generation=1584173512865681&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F252874%2F57407660c162d40f2a59f4ebc15728e4%2FHighReal_16.png?generation=1584173515119630&amp;alt=media)",
    "771910": "Nice. I am thinking of passing the heat map output to another CNN. It works similar to attention."
  },
  "source": "meta"
}