{
  "id": 133317,
  "title": "Is this a useful exercise?",
  "url": "/competitions/deepfake-detection-challenge/discussion/133317",
  "author_name": "",
  "post_date": "2020-03-02T03:00:06.941007900Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>This dataset is a bit unwieldy if you are making a video-based model. I attempted to extract the faces and compress the cropped videos down to a manageable size. See <a href=\"https://www.kaggle.com/anlthms/dfdc-video-faces\">https://www.kaggle.com/anlthms/dfdc-video-faces</a> for examples.</p>\n\n<p>I know it is kinda late in the timeline, but does this sound like a reasonable approach to start with? The underlying assumption is that faces are the only relevant part in the videos. One could compress the data down to less than 3% of its size by excluding everything else.</p>\n\n<p>EDIT: published a notebook to show the face video clips - <a href=\"https://www.kaggle.com/anlthms/example-video-clips\">https://www.kaggle.com/anlthms/example-video-clips</a></p>",
  "messages": [
    {
      "id": "761010",
      "postDate": "03/02/2020 03:00:06",
      "content": "<p>This dataset is a bit unwieldy if you are making a video-based model. I attempted to extract the faces and compress the cropped videos down to a manageable size. See <a href=\"https://www.kaggle.com/anlthms/dfdc-video-faces\">https://www.kaggle.com/anlthms/dfdc-video-faces</a> for examples.</p>\n\n<p>I know it is kinda late in the timeline, but does this sound like a reasonable approach to start with? The underlying assumption is that faces are the only relevant part in the videos. One could compress the data down to less than 3% of its size by excluding everything else.</p>\n\n<p>EDIT: published a notebook to show the face video clips - <a href=\"https://www.kaggle.com/anlthms/example-video-clips\">https://www.kaggle.com/anlthms/example-video-clips</a></p>",
      "rawMarkdown": "This dataset is a bit unwieldy if you are making a video-based model. I attempted to extract the faces and compress the cropped videos down to a manageable size. See [https://www.kaggle.com/anlthms/dfdc-video-faces](https://www.kaggle.com/anlthms/dfdc-video-faces) for examples.\n\nI know it is kinda late in the timeline, but does this sound like a reasonable approach to start with? The underlying assumption is that faces are the only relevant part in the videos. One could compress the data down to less than 3% of its size by excluding everything else.\n\nEDIT: published a notebook to show the face video clips - [https://www.kaggle.com/anlthms/example-video-clips](https://www.kaggle.com/anlthms/example-video-clips)",
      "votes": null
    },
    {
      "id": "761036",
      "postDate": "03/02/2020 04:04:26",
      "content": "<p>Yes. And this is the approach currently being used by everyone. Welcome to the competition.</p>",
      "rawMarkdown": "Yes. And this is the approach currently being used by everyone. Welcome to the competition.",
      "votes": null
    },
    {
      "id": "761074",
      "postDate": "03/02/2020 05:10:10",
      "content": "<p>Thanks, <a href=\"/khahuras\">@khahuras</a>. If there are example notebooks using video (and not images), I missed them. I am trying to keep the data in video form maintaining the original resolution. The tricky part was motion stabilizing the videos. Has someone already done this?</p>\n\n<p>Published a kernel to demo my dataset - <a href=\"https://www.kaggle.com/anlthms/example-video-clips\">https://www.kaggle.com/anlthms/example-video-clips</a></p>",
      "rawMarkdown": "Thanks, @khahuras. If there are example notebooks using video (and not images), I missed them. I am trying to keep the data in video form maintaining the original resolution. The tricky part was motion stabilizing the videos. Has someone already done this?\n\nPublished a kernel to demo my dataset - [https://www.kaggle.com/anlthms/example-video-clips](https://www.kaggle.com/anlthms/example-video-clips)",
      "votes": null
    },
    {
      "id": "761102",
      "postDate": "03/02/2020 06:18:54",
      "content": "<p>No, at least no public thing about that to my knowledge. </p>",
      "rawMarkdown": "No, at least no public thing about that to my knowledge.",
      "votes": null
    },
    {
      "id": "761124",
      "postDate": "03/02/2020 06:59:10",
      "content": "<p><a href=\"/anlthms\">@anlthms</a> Do you mean this one. I've almost done it.\n<a href=\"https://www.kaggle.com/wuliaokaola/dfdc-face-videos-part01\">https://www.kaggle.com/wuliaokaola/dfdc-face-videos-part01</a></p>",
      "rawMarkdown": "anlthms Do you mean this one. I've almost done it.\nhttps://www.kaggle.com/wuliaokaola/dfdc-face-videos-part01",
      "votes": null
    },
    {
      "id": "761626",
      "postDate": "03/02/2020 19:18:26",
      "content": "<p>Yes, your dataset looks similar. Did it help to keep the input data as video?</p>\n\n<p>My intuition is that image-based models will miss out on features in the time dimension. Things like flicker artifacts would be easier to detect with temporal data.</p>",
      "rawMarkdown": "Yes, your dataset looks similar. Did it help to keep the input data as video?\n\nMy intuition is that image-based models will miss out on features in the time dimension. Things like flicker artifacts would be easier to detect with temporal data.",
      "votes": null
    },
    {
      "id": "761803",
      "postDate": "03/03/2020 00:23:01",
      "content": "<p>I plan to implement the time series model. But now I am working on image-based model. The current LB score (0.39692) is based on a single image-based model without ensemble.</p>\n\n<p>If anyone want to team up, please PM me. I have both AWS credit and GCP TPU quota.</p>",
      "rawMarkdown": "I plan to implement the time series model. But now I am working on image-based model. The current LB score (0.39692) is based on a single image-based model without ensemble.\n\nIf anyone want to team up, please PM me. I have both AWS credit and GCP TPU quota.",
      "votes": null
    },
    {
      "id": "764747",
      "postDate": "03/05/2020 20:39:17",
      "content": "<p>I put together some minimalistic code to train and validate using these cropped videos. </p>\n\n<p><a href=\"https://www.kaggle.com/anlthms/video-classification-with-3d-resnet\">https://www.kaggle.com/anlthms/video-classification-with-3d-resnet</a></p>\n\n<p>Using 3 folders to train and 2 to validate:</p>\n\n<blockquote>\n  <p>validation accuracy 56%\n  epoch 6 training loss 0.63 validation loss 0.68</p>\n</blockquote>\n\n<p>Was hoping for better results. Oh well...</p>",
      "rawMarkdown": "I put together some minimalistic code to train and validate using these cropped videos. \n\n[https://www.kaggle.com/anlthms/video-classification-with-3d-resnet](https://www.kaggle.com/anlthms/video-classification-with-3d-resnet)\n\nUsing 3 folders to train and 2 to validate:\n\n&gt; validation accuracy 56%\n&gt; epoch 6 training loss 0.63 validation loss 0.68\n\nWas hoping for better results. Oh well...",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 761036,
      "author_name": "khahuras",
      "author_url": "",
      "post_date": "03/02/2020 04:04:26",
      "content": "<p>Yes. And this is the approach currently being used by everyone. Welcome to the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 761074,
          "author_name": "anlthms",
          "author_url": "",
          "post_date": "03/02/2020 05:10:10",
          "content": "<p>Thanks, <a href=\"/khahuras\">@khahuras</a>. If there are example notebooks using video (and not images), I missed them. I am trying to keep the data in video form maintaining the original resolution. The tricky part was motion stabilizing the videos. Has someone already done this?</p>\n\n<p>Published a kernel to demo my dataset - <a href=\"https://www.kaggle.com/anlthms/example-video-clips\">https://www.kaggle.com/anlthms/example-video-clips</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761102,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/02/2020 06:18:54",
          "content": "<p>No, at least no public thing about that to my knowledge. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761124,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "03/02/2020 06:59:10",
          "content": "<p><a href=\"/anlthms\">@anlthms</a> Do you mean this one. I've almost done it.\n<a href=\"https://www.kaggle.com/wuliaokaola/dfdc-face-videos-part01\">https://www.kaggle.com/wuliaokaola/dfdc-face-videos-part01</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761626,
          "author_name": "anlthms",
          "author_url": "",
          "post_date": "03/02/2020 19:18:26",
          "content": "<p>Yes, your dataset looks similar. Did it help to keep the input data as video?</p>\n\n<p>My intuition is that image-based models will miss out on features in the time dimension. Things like flicker artifacts would be easier to detect with temporal data.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761803,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "03/03/2020 00:23:01",
          "content": "<p>I plan to implement the time series model. But now I am working on image-based model. The current LB score (0.39692) is based on a single image-based model without ensemble.</p>\n\n<p>If anyone want to team up, please PM me. I have both AWS credit and GCP TPU quota.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 764747,
      "author_name": "anlthms",
      "author_url": "",
      "post_date": "03/05/2020 20:39:17",
      "content": "<p>I put together some minimalistic code to train and validate using these cropped videos. </p>\n\n<p><a href=\"https://www.kaggle.com/anlthms/video-classification-with-3d-resnet\">https://www.kaggle.com/anlthms/video-classification-with-3d-resnet</a></p>\n\n<p>Using 3 folders to train and 2 to validate:</p>\n\n<blockquote>\n  <p>validation accuracy 56%\n  epoch 6 training loss 0.63 validation loss 0.68</p>\n</blockquote>\n\n<p>Was hoping for better results. Oh well...</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "761010": "This dataset is a bit unwieldy if you are making a video-based model. I attempted to extract the faces and compress the cropped videos down to a manageable size. See [https://www.kaggle.com/anlthms/dfdc-video-faces](https://www.kaggle.com/anlthms/dfdc-video-faces) for examples.\n\nI know it is kinda late in the timeline, but does this sound like a reasonable approach to start with? The underlying assumption is that faces are the only relevant part in the videos. One could compress the data down to less than 3% of its size by excluding everything else.\n\nEDIT: published a notebook to show the face video clips - [https://www.kaggle.com/anlthms/example-video-clips](https://www.kaggle.com/anlthms/example-video-clips)",
    "761036": "Yes. And this is the approach currently being used by everyone. Welcome to the competition.",
    "761074": "Thanks, @khahuras. If there are example notebooks using video (and not images), I missed them. I am trying to keep the data in video form maintaining the original resolution. The tricky part was motion stabilizing the videos. Has someone already done this?\n\nPublished a kernel to demo my dataset - [https://www.kaggle.com/anlthms/example-video-clips](https://www.kaggle.com/anlthms/example-video-clips)",
    "761102": "No, at least no public thing about that to my knowledge.",
    "761124": "anlthms Do you mean this one. I've almost done it.\nhttps://www.kaggle.com/wuliaokaola/dfdc-face-videos-part01",
    "761626": "Yes, your dataset looks similar. Did it help to keep the input data as video?\n\nMy intuition is that image-based models will miss out on features in the time dimension. Things like flicker artifacts would be easier to detect with temporal data.",
    "761803": "I plan to implement the time series model. But now I am working on image-based model. The current LB score (0.39692) is based on a single image-based model without ensemble.\n\nIf anyone want to team up, please PM me. I have both AWS credit and GCP TPU quota.",
    "764747": "I put together some minimalistic code to train and validate using these cropped videos. \n\n[https://www.kaggle.com/anlthms/video-classification-with-3d-resnet](https://www.kaggle.com/anlthms/video-classification-with-3d-resnet)\n\nUsing 3 folders to train and 2 to validate:\n\n&gt; validation accuracy 56%\n&gt; epoch 6 training loss 0.63 validation loss 0.68\n\nWas hoping for better results. Oh well..."
  },
  "source": "meta"
}