{
  "id": 131315,
  "title": "Today on Arxiv...",
  "url": "/competitions/deepfake-detection-challenge/discussion/131315",
  "author_name": "",
  "post_date": "2020-02-19T08:48:02.410467400Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I have seen this paper <em><a href=\"https://arxiv.org/pdf/2002.07442.pdf\">V4D:4D CONVOLUTIONAL NEURAL NETWORKS FOR\nVIDEO-LEVEL REPRESENTATION LEARNING</a></em> published today on Arxiv.com. </p>\n\n<p>It sems prety interesting, however, I have never worked with 4D convolutions. Can anyone can shed any light?</p>",
  "messages": [
    {
      "id": "750277",
      "postDate": "02/19/2020 08:48:02",
      "content": "<p>I have seen this paper <em><a href=\"https://arxiv.org/pdf/2002.07442.pdf\">V4D:4D CONVOLUTIONAL NEURAL NETWORKS FOR\nVIDEO-LEVEL REPRESENTATION LEARNING</a></em> published today on Arxiv.com. </p>\n\n<p>It sems prety interesting, however, I have never worked with 4D convolutions. Can anyone can shed any light?</p>",
      "rawMarkdown": "I have seen this paper *[V4D:4D CONVOLUTIONAL NEURAL NETWORKS FOR\nVIDEO-LEVEL REPRESENTATION LEARNING](https://arxiv.org/pdf/2002.07442.pdf)* published today on Arxiv.com. \n\nIt sems prety interesting, however, I have never worked with 4D convolutions. Can anyone can shed any light?",
      "votes": null
    },
    {
      "id": "751383",
      "postDate": "02/20/2020 07:14:08",
      "content": "<p>The 4D CNN paper is talking about using the \"entire\" video which requires very high computation and time requirements (concern is whether your submission takes less than 8 seconds per video). Then you will need to try tricks such as using every 10th frame instead of consecutive frames, etc.</p>\n\n<p>Why not consider 3D CNNs instead, which may be sufficient for spatio-temporal feature extraction.\nYou can also try Facebook approach <a href=\"https://github.com/facebookresearch/SparseConvNet\">https://github.com/facebookresearch/SparseConvNet</a>.</p>\n\n<p>In any case, this competition is really about identifying fakes i.e. inconsistent face poses, face coherence, etc.\nSee the list of papers at <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121223\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121223</a></p>",
      "rawMarkdown": "The 4D CNN paper is talking about using the \"entire\" video which requires very high computation and time requirements (concern is whether your submission takes less than 8 seconds per video). Then you will need to try tricks such as using every 10th frame instead of consecutive frames, etc.\n\nWhy not consider 3D CNNs instead, which may be sufficient for spatio-temporal feature extraction.\nYou can also try Facebook approach https://github.com/facebookresearch/SparseConvNet.\n\nIn any case, this competition is really about identifying fakes i.e. inconsistent face poses, face coherence, etc.\nSee the list of papers at https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121223",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 751383,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "02/20/2020 07:14:08",
      "content": "<p>The 4D CNN paper is talking about using the \"entire\" video which requires very high computation and time requirements (concern is whether your submission takes less than 8 seconds per video). Then you will need to try tricks such as using every 10th frame instead of consecutive frames, etc.</p>\n\n<p>Why not consider 3D CNNs instead, which may be sufficient for spatio-temporal feature extraction.\nYou can also try Facebook approach <a href=\"https://github.com/facebookresearch/SparseConvNet\">https://github.com/facebookresearch/SparseConvNet</a>.</p>\n\n<p>In any case, this competition is really about identifying fakes i.e. inconsistent face poses, face coherence, etc.\nSee the list of papers at <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121223\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121223</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "750277": "I have seen this paper *[V4D:4D CONVOLUTIONAL NEURAL NETWORKS FOR\nVIDEO-LEVEL REPRESENTATION LEARNING](https://arxiv.org/pdf/2002.07442.pdf)* published today on Arxiv.com. \n\nIt sems prety interesting, however, I have never worked with 4D convolutions. Can anyone can shed any light?",
    "751383": "The 4D CNN paper is talking about using the \"entire\" video which requires very high computation and time requirements (concern is whether your submission takes less than 8 seconds per video). Then you will need to try tricks such as using every 10th frame instead of consecutive frames, etc.\n\nWhy not consider 3D CNNs instead, which may be sufficient for spatio-temporal feature extraction.\nYou can also try Facebook approach https://github.com/facebookresearch/SparseConvNet.\n\nIn any case, this competition is really about identifying fakes i.e. inconsistent face poses, face coherence, etc.\nSee the list of papers at https://www.kaggle.com/c/deepfake-detection-challenge/discussion/121223"
  },
  "source": "meta"
}