{
  "id": 132647,
  "title": "Compressed Video Action Recognition",
  "url": "/competitions/deepfake-detection-challenge/discussion/132647",
  "author_name": "",
  "post_date": "2020-02-27T04:09:43.349789400Z",
  "votes": 22,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I've been doing quite a bit of reading and exploring various different ideas and either don't have the time or expertise or resources to fully implement them so I am going to offer up various different ideas to see if anyone else can take a crack at them. </p>\n\n<p>The first one I want to highlight is <a href=\"https://www.cs.utexas.edu/~cywu/projects/coviar/\">CoViAR</a> (Compressed Video Action Recognition). One of the major challenges people have been dealing with in this challenge has been that just decoding 100k+ mp4's takes quite a bit of time. This paper looks at the case where we skip over that step and just operate directly on the encoded files. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F6d96407d53aa28a3e12325a224877eda%2Fcoviar.png?generation=1582775224217207&amp;alt=media\" alt=\"\"></p>\n\n<p>This has the theoretical advantage of not needing to spend time decoding the videos and also potentially operate on a denser representation of the data. I only have a weak understanding of mp4's compression, but my vague knowledge of it is that there are I-frames which are basically a normal image that the other frames are based around and then P-frames and residuals that represent the motion and corrections required on subsequent frames so instead of having redundant information representing the raw pixel values for all 1920x1080x3 frames they simply represent the changes of information between frames. </p>\n\n<p>For those familiar with optical flow this sounds like it will also be a very useful approximation of that. I think this technique likely has various advantages for this competition in particular because operating in a different space than the images are generated may make certain forms of editing much more obvious, particularly since it is focusing on the discontinuities between frames. </p>\n\n<p>The disadvantage of this technique is it wont b be as straightforward to crop in on faces and the images won't be in a format that lends itself to transfer learning with imagenet networks (may or may not be useful anyway)</p>\n\n<p>For anyone interested in this approach the researchers published the repo here <a href=\"https://github.com/chaoyuaw/pytorch-coviar\">https://github.com/chaoyuaw/pytorch-coviar</a></p>\n\n<p>I spent a little bit of time looking into this and could get it working locally(albeit slowly) but working in kaggle kernels was a bit of a struggle because it assumes ffmpeg install in a specific location. I spent a bit of time trying to correct the project to work in the kernels but ended up abandoning it. </p>",
  "messages": [
    {
      "id": "757744",
      "postDate": "02/27/2020 04:09:43",
      "content": "<p>I've been doing quite a bit of reading and exploring various different ideas and either don't have the time or expertise or resources to fully implement them so I am going to offer up various different ideas to see if anyone else can take a crack at them. </p>\n\n<p>The first one I want to highlight is <a href=\"https://www.cs.utexas.edu/~cywu/projects/coviar/\">CoViAR</a> (Compressed Video Action Recognition). One of the major challenges people have been dealing with in this challenge has been that just decoding 100k+ mp4's takes quite a bit of time. This paper looks at the case where we skip over that step and just operate directly on the encoded files. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F6d96407d53aa28a3e12325a224877eda%2Fcoviar.png?generation=1582775224217207&amp;alt=media\" alt=\"\"></p>\n\n<p>This has the theoretical advantage of not needing to spend time decoding the videos and also potentially operate on a denser representation of the data. I only have a weak understanding of mp4's compression, but my vague knowledge of it is that there are I-frames which are basically a normal image that the other frames are based around and then P-frames and residuals that represent the motion and corrections required on subsequent frames so instead of having redundant information representing the raw pixel values for all 1920x1080x3 frames they simply represent the changes of information between frames. </p>\n\n<p>For those familiar with optical flow this sounds like it will also be a very useful approximation of that. I think this technique likely has various advantages for this competition in particular because operating in a different space than the images are generated may make certain forms of editing much more obvious, particularly since it is focusing on the discontinuities between frames. </p>\n\n<p>The disadvantage of this technique is it wont b be as straightforward to crop in on faces and the images won't be in a format that lends itself to transfer learning with imagenet networks (may or may not be useful anyway)</p>\n\n<p>For anyone interested in this approach the researchers published the repo here <a href=\"https://github.com/chaoyuaw/pytorch-coviar\">https://github.com/chaoyuaw/pytorch-coviar</a></p>\n\n<p>I spent a little bit of time looking into this and could get it working locally(albeit slowly) but working in kaggle kernels was a bit of a struggle because it assumes ffmpeg install in a specific location. I spent a bit of time trying to correct the project to work in the kernels but ended up abandoning it. </p>",
      "rawMarkdown": "I've been doing quite a bit of reading and exploring various different ideas and either don't have the time or expertise or resources to fully implement them so I am going to offer up various different ideas to see if anyone else can take a crack at them. \n\nThe first one I want to highlight is [CoViAR](https://www.cs.utexas.edu/~cywu/projects/coviar/) (Compressed Video Action Recognition). One of the major challenges people have been dealing with in this challenge has been that just decoding 100k+ mp4's takes quite a bit of time. This paper looks at the case where we skip over that step and just operate directly on the encoded files. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F6d96407d53aa28a3e12325a224877eda%2Fcoviar.png?generation=1582775224217207&amp;alt=media)\n\nThis has the theoretical advantage of not needing to spend time decoding the videos and also potentially operate on a denser representation of the data. I only have a weak understanding of mp4's compression, but my vague knowledge of it is that there are I-frames which are basically a normal image that the other frames are based around and then P-frames and residuals that represent the motion and corrections required on subsequent frames so instead of having redundant information representing the raw pixel values for all 1920x1080x3 frames they simply represent the changes of information between frames. \n\nFor those familiar with optical flow this sounds like it will also be a very useful approximation of that. I think this technique likely has various advantages for this competition in particular because operating in a different space than the images are generated may make certain forms of editing much more obvious, particularly since it is focusing on the discontinuities between frames. \n\nThe disadvantage of this technique is it wont b be as straightforward to crop in on faces and the images won't be in a format that lends itself to transfer learning with imagenet networks (may or may not be useful anyway)\n\nFor anyone interested in this approach the researchers published the repo here https://github.com/chaoyuaw/pytorch-coviar\n\nI spent a little bit of time looking into this and could get it working locally(albeit slowly) but working in kaggle kernels was a bit of a struggle because it assumes ffmpeg install in a specific location. I spent a bit of time trying to correct the project to work in the kernels but ended up abandoning it.",
      "votes": null
    },
    {
      "id": "779663",
      "postDate": "03/19/2020 15:21:09",
      "content": "<p>Very interesting ryches. \nAs I see you didn't include this software in the External Data Disclosure Thread, right?</p>",
      "rawMarkdown": "Very interesting ryches. \nAs I see you didn't include this software in the External Data Disclosure Thread, right?",
      "votes": null
    },
    {
      "id": "2189381",
      "postDate": "03/20/2023 12:57:22",
      "content": "<p>AI E2E codes has developed in recent years, maybe there is something new for this task.</p>",
      "rawMarkdown": "AI E2E codes has developed in recent years, maybe there is something new for this task.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2189381,
      "author_name": "yuyuxu5207",
      "author_url": "",
      "post_date": "03/20/2023 12:57:22",
      "content": "<p>AI E2E codes has developed in recent years, maybe there is something new for this task.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 779663,
      "author_name": "francescomarra",
      "author_url": "",
      "post_date": "03/19/2020 15:21:09",
      "content": "<p>Very interesting ryches. \nAs I see you didn't include this software in the External Data Disclosure Thread, right?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "757744": "I've been doing quite a bit of reading and exploring various different ideas and either don't have the time or expertise or resources to fully implement them so I am going to offer up various different ideas to see if anyone else can take a crack at them. \n\nThe first one I want to highlight is [CoViAR](https://www.cs.utexas.edu/~cywu/projects/coviar/) (Compressed Video Action Recognition). One of the major challenges people have been dealing with in this challenge has been that just decoding 100k+ mp4's takes quite a bit of time. This paper looks at the case where we skip over that step and just operate directly on the encoded files. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2F6d96407d53aa28a3e12325a224877eda%2Fcoviar.png?generation=1582775224217207&amp;alt=media)\n\nThis has the theoretical advantage of not needing to spend time decoding the videos and also potentially operate on a denser representation of the data. I only have a weak understanding of mp4's compression, but my vague knowledge of it is that there are I-frames which are basically a normal image that the other frames are based around and then P-frames and residuals that represent the motion and corrections required on subsequent frames so instead of having redundant information representing the raw pixel values for all 1920x1080x3 frames they simply represent the changes of information between frames. \n\nFor those familiar with optical flow this sounds like it will also be a very useful approximation of that. I think this technique likely has various advantages for this competition in particular because operating in a different space than the images are generated may make certain forms of editing much more obvious, particularly since it is focusing on the discontinuities between frames. \n\nThe disadvantage of this technique is it wont b be as straightforward to crop in on faces and the images won't be in a format that lends itself to transfer learning with imagenet networks (may or may not be useful anyway)\n\nFor anyone interested in this approach the researchers published the repo here https://github.com/chaoyuaw/pytorch-coviar\n\nI spent a little bit of time looking into this and could get it working locally(albeit slowly) but working in kaggle kernels was a bit of a struggle because it assumes ffmpeg install in a specific location. I spent a bit of time trying to correct the project to work in the kernels but ended up abandoning it.",
    "779663": "Very interesting ryches. \nAs I see you didn't include this software in the External Data Disclosure Thread, right?",
    "2189381": "AI E2E codes has developed in recent years, maybe there is something new for this task."
  },
  "source": "meta"
}