{
  "id": 132412,
  "title": "Transformer-like models?",
  "url": "/competitions/deepfake-detection-challenge/discussion/132412",
  "author_name": "",
  "post_date": "2020-02-25T21:35:08.819793600Z",
  "votes": 6,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Since transformers architectures are all the rage nowadays and have been applied many times to win a diverse set of competitions (<a href=\"https://www.kaggle.com/c/champs-scalar-coupling/discussion/106575\">here</a>, <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400\">here</a>, and even <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891\">here</a>), I was wondering if anyone can point to any resources for applying such techniques to DeepFake detection?</p>\n\n<p>Something like a temporal transformer on top of a CNN detecting face frame by frame?\nDoes that make sense to anyone? Any thoughts are appreciated!</p>",
  "messages": [
    {
      "id": "756587",
      "postDate": "02/25/2020 21:35:08",
      "content": "<p>Since transformers architectures are all the rage nowadays and have been applied many times to win a diverse set of competitions (<a href=\"https://www.kaggle.com/c/champs-scalar-coupling/discussion/106575\">here</a>, <a href=\"https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400\">here</a>, and even <a href=\"https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891\">here</a>), I was wondering if anyone can point to any resources for applying such techniques to DeepFake detection?</p>\n\n<p>Something like a temporal transformer on top of a CNN detecting face frame by frame?\nDoes that make sense to anyone? Any thoughts are appreciated!</p>",
      "rawMarkdown": "Since transformers architectures are all the rage nowadays and have been applied many times to win a diverse set of competitions ([here](https://www.kaggle.com/c/champs-scalar-coupling/discussion/106575), [here](https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400), and even [here](https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891)), I was wondering if anyone can point to any resources for applying such techniques to DeepFake detection?\n\nSomething like a temporal transformer on top of a CNN detecting face frame by frame?\nDoes that make sense to anyone? Any thoughts are appreciated!",
      "votes": null
    },
    {
      "id": "762699",
      "postDate": "03/03/2020 18:21:34",
      "content": "<p>You might find the Action Transformer interesting: <a href=\"https://arxiv.org/pdf/1812.02707.pdf\">https://arxiv.org/pdf/1812.02707.pdf</a></p>\n\n<p>The authors describe how training over person-specific queries of the Atomic Visual Actions dataset, their transformer learned to attend to faces and hands for the task of classifying human activities.</p>\n\n<p>Paraphrasing from the intro, they use I3D for base features along with a region proposal network as the input query to a transformer aggregating spatio-temporal context from the video clip.</p>\n\n<p>However, this challenge might not benefit as much from attending to spatial regions not containing faces. </p>",
      "rawMarkdown": "You might find the Action Transformer interesting: https://arxiv.org/pdf/1812.02707.pdf\n\nThe authors describe how training over person-specific queries of the Atomic Visual Actions dataset, their transformer learned to attend to faces and hands for the task of classifying human activities.\n\nParaphrasing from the intro, they use I3D for base features along with a region proposal network as the input query to a transformer aggregating spatio-temporal context from the video clip.\n\nHowever, this challenge might not benefit as much from attending to spatial regions not containing faces.",
      "votes": null
    },
    {
      "id": "762871",
      "postDate": "03/03/2020 22:05:36",
      "content": "<p>Awesome, thanks for the link! \nIndeed, the face is probably the most important part in this challenge. \nThat being said, I have noticed some \"artifacts\" outside the face (check this <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133684\">thread</a> for more details). </p>",
      "rawMarkdown": "Awesome, thanks for the link! \nIndeed, the face is probably the most important part in this challenge. \nThat being said, I have noticed some \"artifacts\" outside the face (check this [thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133684) for more details).",
      "votes": null
    },
    {
      "id": "769601",
      "postDate": "03/12/2020 03:56:18",
      "content": "<p>Hi , have you tried transformer ? How's it ? </p>",
      "rawMarkdown": "Hi , have you tried transformer ? How's it ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 762699,
      "author_name": "itsmellslikeml",
      "author_url": "",
      "post_date": "03/03/2020 18:21:34",
      "content": "<p>You might find the Action Transformer interesting: <a href=\"https://arxiv.org/pdf/1812.02707.pdf\">https://arxiv.org/pdf/1812.02707.pdf</a></p>\n\n<p>The authors describe how training over person-specific queries of the Atomic Visual Actions dataset, their transformer learned to attend to faces and hands for the task of classifying human activities.</p>\n\n<p>Paraphrasing from the intro, they use I3D for base features along with a region proposal network as the input query to a transformer aggregating spatio-temporal context from the video clip.</p>\n\n<p>However, this challenge might not benefit as much from attending to spatial regions not containing faces. </p>",
      "votes": null,
      "replies": [
        {
          "id": 762871,
          "author_name": "yassinealouini",
          "author_url": "",
          "post_date": "03/03/2020 22:05:36",
          "content": "<p>Awesome, thanks for the link! \nIndeed, the face is probably the most important part in this challenge. \nThat being said, I have noticed some \"artifacts\" outside the face (check this <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133684\">thread</a> for more details). </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 769601,
      "author_name": "fionalxd",
      "author_url": "",
      "post_date": "03/12/2020 03:56:18",
      "content": "<p>Hi , have you tried transformer ? How's it ? </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "756587": "Since transformers architectures are all the rage nowadays and have been applied many times to win a diverse set of competitions ([here](https://www.kaggle.com/c/champs-scalar-coupling/discussion/106575), [here](https://www.kaggle.com/c/nfl-big-data-bowl-2020/discussion/119400), and even [here](https://www.kaggle.com/c/data-science-bowl-2019/discussion/127891)), I was wondering if anyone can point to any resources for applying such techniques to DeepFake detection?\n\nSomething like a temporal transformer on top of a CNN detecting face frame by frame?\nDoes that make sense to anyone? Any thoughts are appreciated!",
    "762699": "You might find the Action Transformer interesting: https://arxiv.org/pdf/1812.02707.pdf\n\nThe authors describe how training over person-specific queries of the Atomic Visual Actions dataset, their transformer learned to attend to faces and hands for the task of classifying human activities.\n\nParaphrasing from the intro, they use I3D for base features along with a region proposal network as the input query to a transformer aggregating spatio-temporal context from the video clip.\n\nHowever, this challenge might not benefit as much from attending to spatial regions not containing faces.",
    "762871": "Awesome, thanks for the link! \nIndeed, the face is probably the most important part in this challenge. \nThat being said, I have noticed some \"artifacts\" outside the face (check this [thread](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/133684) for more details).",
    "769601": "Hi , have you tried transformer ? How's it ?"
  },
  "source": "meta"
}