{
  "id": 121403,
  "title": "Taking requests (and a funny picture)",
  "url": "/competitions/deepfake-detection-challenge/discussion/121403",
  "author_name": "",
  "post_date": "2019-12-13T02:54:19.235656100Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff86f3bee56bb39e83569adafad0b15ea%2Fafrgmowivl.jpg?generation=1576205762267875&amp;alt=media\" alt=\"\"></p>\n\n<p>I've downloaded and unzipped the entire dataset and have started playing with it a little. My first plan is to pull each first frame and resize to a more manageable size and upload as a dataset for everyone to use. A single still looks like enough to be able to detect quite a bit. Also planning on stripping just the audio and uploading that as a complementary dataset. </p>\n\n<p>If anyone has any specific requests in terms of data prep or reductions I am taking orders as long as it doesnt take way too much time to run. I'd like to make this more accessible to more people. </p>\n\n<p>On first aggregation of the training set metadata files it looks like 0.83925 are fake and only 0.16075 are real. It appears this is not the case on the test set though.</p>",
  "messages": [
    {
      "id": "694000",
      "postDate": "12/13/2019 02:54:19",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff86f3bee56bb39e83569adafad0b15ea%2Fafrgmowivl.jpg?generation=1576205762267875&amp;alt=media\" alt=\"\"></p>\n\n<p>I've downloaded and unzipped the entire dataset and have started playing with it a little. My first plan is to pull each first frame and resize to a more manageable size and upload as a dataset for everyone to use. A single still looks like enough to be able to detect quite a bit. Also planning on stripping just the audio and uploading that as a complementary dataset. </p>\n\n<p>If anyone has any specific requests in terms of data prep or reductions I am taking orders as long as it doesnt take way too much time to run. I'd like to make this more accessible to more people. </p>\n\n<p>On first aggregation of the training set metadata files it looks like 0.83925 are fake and only 0.16075 are real. It appears this is not the case on the test set though.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff86f3bee56bb39e83569adafad0b15ea%2Fafrgmowivl.jpg?generation=1576205762267875&amp;alt=media)\n\nI've downloaded and unzipped the entire dataset and have started playing with it a little. My first plan is to pull each first frame and resize to a more manageable size and upload as a dataset for everyone to use. A single still looks like enough to be able to detect quite a bit. Also planning on stripping just the audio and uploading that as a complementary dataset. \n\nIf anyone has any specific requests in terms of data prep or reductions I am taking orders as long as it doesnt take way too much time to run. I'd like to make this more accessible to more people. \n\nOn first aggregation of the training set metadata files it looks like 0.83925 are fake and only 0.16075 are real. It appears this is not the case on the test set though.",
      "votes": null
    },
    {
      "id": "694016",
      "postDate": "12/13/2019 04:10:05",
      "content": "<p>Maybe you should make a video of just the face. I notice the difference between real and fake is not that much in the first frame. Thank you so much for sharing the dataset. It is really helpful for those of us that don't have a powerful hardware.</p>",
      "rawMarkdown": "Maybe you should make a video of just the face. I notice the difference between real and fake is not that much in the first frame. Thank you so much for sharing the dataset. It is really helpful for those of us that don't have a powerful hardware.",
      "votes": null
    },
    {
      "id": "694051",
      "postDate": "12/13/2019 05:40:35",
      "content": "<p>I will look into doing that. I can fairly quickly apply a haar cascade with opencv for narrowing down on the faces, but there are certain edge cases I would need to figure out like what to do when there are multiples faces in frame. One thing I considered doing was just zeroing out everything outside of regions that ever had faces in them. In theory that should reduce quite a bit of space. </p>\n\n<p>One alternate approach I thought of was finding the furthest most left,right, top,bottom points which features a face and then crop into that region. In theory that should be much narrower and just have the region of interest. </p>\n\n<p>Regardless, video is going to be quite large, doing something like the featurizing that was done in the youtube challenge or sampling individual frames will probably be much more information dense per gb</p>",
      "rawMarkdown": "I will look into doing that. I can fairly quickly apply a haar cascade with opencv for narrowing down on the faces, but there are certain edge cases I would need to figure out like what to do when there are multiples faces in frame. One thing I considered doing was just zeroing out everything outside of regions that ever had faces in them. In theory that should reduce quite a bit of space. \n\nOne alternate approach I thought of was finding the furthest most left,right, top,bottom points which features a face and then crop into that region. In theory that should be much narrower and just have the region of interest. \n\nRegardless, video is going to be quite large, doing something like the featurizing that was done in the youtube challenge or sampling individual frames will probably be much more information dense per gb",
      "votes": null
    },
    {
      "id": "694686",
      "postDate": "12/14/2019 01:20:50",
      "content": "<p>Yes. You are right. How about just crop the face of just the first frame of each video. The size will significantly be smaller and won't lost a lot of data.</p>",
      "rawMarkdown": "Yes. You are right. How about just crop the face of just the first frame of each video. The size will significantly be smaller and won't lost a lot of data.",
      "votes": null
    },
    {
      "id": "694746",
      "postDate": "12/14/2019 02:56:56",
      "content": "<p>Thanks!\nIf I had the time/compute/skills right now I would extract a reasonable number of frames from each video and figure out a way to crop the face(s).</p>",
      "rawMarkdown": "Thanks!\nIf I had the time/compute/skills right now I would extract a reasonable number of frames from each video and figure out a way to crop the face(s).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 694016,
      "author_name": "unkownhihi",
      "author_url": "",
      "post_date": "12/13/2019 04:10:05",
      "content": "<p>Maybe you should make a video of just the face. I notice the difference between real and fake is not that much in the first frame. Thank you so much for sharing the dataset. It is really helpful for those of us that don't have a powerful hardware.</p>",
      "votes": null,
      "replies": [
        {
          "id": 694051,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "12/13/2019 05:40:35",
          "content": "<p>I will look into doing that. I can fairly quickly apply a haar cascade with opencv for narrowing down on the faces, but there are certain edge cases I would need to figure out like what to do when there are multiples faces in frame. One thing I considered doing was just zeroing out everything outside of regions that ever had faces in them. In theory that should reduce quite a bit of space. </p>\n\n<p>One alternate approach I thought of was finding the furthest most left,right, top,bottom points which features a face and then crop into that region. In theory that should be much narrower and just have the region of interest. </p>\n\n<p>Regardless, video is going to be quite large, doing something like the featurizing that was done in the youtube challenge or sampling individual frames will probably be much more information dense per gb</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 694686,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "12/14/2019 01:20:50",
          "content": "<p>Yes. You are right. How about just crop the face of just the first frame of each video. The size will significantly be smaller and won't lost a lot of data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 694746,
      "author_name": "sdoria",
      "author_url": "",
      "post_date": "12/14/2019 02:56:56",
      "content": "<p>Thanks!\nIf I had the time/compute/skills right now I would extract a reasonable number of frames from each video and figure out a way to crop the face(s).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "694000": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1035002%2Ff86f3bee56bb39e83569adafad0b15ea%2Fafrgmowivl.jpg?generation=1576205762267875&amp;alt=media)\n\nI've downloaded and unzipped the entire dataset and have started playing with it a little. My first plan is to pull each first frame and resize to a more manageable size and upload as a dataset for everyone to use. A single still looks like enough to be able to detect quite a bit. Also planning on stripping just the audio and uploading that as a complementary dataset. \n\nIf anyone has any specific requests in terms of data prep or reductions I am taking orders as long as it doesnt take way too much time to run. I'd like to make this more accessible to more people. \n\nOn first aggregation of the training set metadata files it looks like 0.83925 are fake and only 0.16075 are real. It appears this is not the case on the test set though.",
    "694016": "Maybe you should make a video of just the face. I notice the difference between real and fake is not that much in the first frame. Thank you so much for sharing the dataset. It is really helpful for those of us that don't have a powerful hardware.",
    "694051": "I will look into doing that. I can fairly quickly apply a haar cascade with opencv for narrowing down on the faces, but there are certain edge cases I would need to figure out like what to do when there are multiples faces in frame. One thing I considered doing was just zeroing out everything outside of regions that ever had faces in them. In theory that should reduce quite a bit of space. \n\nOne alternate approach I thought of was finding the furthest most left,right, top,bottom points which features a face and then crop into that region. In theory that should be much narrower and just have the region of interest. \n\nRegardless, video is going to be quite large, doing something like the featurizing that was done in the youtube challenge or sampling individual frames will probably be much more information dense per gb",
    "694686": "Yes. You are right. How about just crop the face of just the first frame of each video. The size will significantly be smaller and won't lost a lot of data.",
    "694746": "Thanks!\nIf I had the time/compute/skills right now I would extract a reasonable number of frames from each video and figure out a way to crop the face(s)."
  },
  "source": "meta"
}