{
  "id": 123707,
  "title": "Resized first frames and labels dataset",
  "url": "/competitions/deepfake-detection-challenge/discussion/123707",
  "author_name": "ryches",
  "post_date": "2019-12-29T19:40:43.505000",
  "votes": 23,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Made a simple dataset that captures the first frame of every video resized to 512x512 along with the labels. Along with this you should be able to train a full model within the kernel in 9 hours as well as make predictions since loading first frame is quite fast. Will make a sample showing a simple resnet18 approach to this that can achieve a respectable accuracy once I get some time\n<a href=\"https://www.kaggle.com/ryches/deepfake-first-frames-and-labels\">https://www.kaggle.com/ryches/deepfake-first-frames-and-labels</a></p>",
  "messages": [
    {
      "id": 705993,
      "postDate": "2019-12-29T19:40:43.507Z",
      "content": "<p>Made a simple dataset that captures the first frame of every video resized to 512x512 along with the labels. Along with this you should be able to train a full model within the kernel in 9 hours as well as make predictions since loading first frame is quite fast. Will make a sample showing a simple resnet18 approach to this that can achieve a respectable accuracy once I get some time\n<a href=\"https://www.kaggle.com/ryches/deepfake-first-frames-and-labels\">https://www.kaggle.com/ryches/deepfake-first-frames-and-labels</a></p>",
      "rawMarkdown": "Made a simple dataset that captures the first frame of every video resized to 512x512 along with the labels. Along with this you should be able to train a full model within the kernel in 9 hours as well as make predictions since loading first frame is quite fast. Will make a sample showing a simple resnet18 approach to this that can achieve a respectable accuracy once I get some time\nhttps://www.kaggle.com/ryches/deepfake-first-frames-and-labels",
      "votes": 22
    },
    {
      "id": 706848,
      "postDate": "2019-12-30T23:12:24.017Z",
      "content": "<p>Looking forward to see your starter!</p>",
      "rawMarkdown": "Looking forward to see your starter!",
      "votes": 1
    },
    {
      "id": 708447,
      "postDate": "2020-01-02T10:01:56.713Z",
      "content": "<p><a href=\"/ryches\">@ryches</a> Many thanks. Could you kindly share your code to generate these frames as well? I think you did some good pre-processing steps.</p>",
      "rawMarkdown": "@ryches Many thanks. Could you kindly share your code to generate these frames as well? I think you did some good pre-processing steps.",
      "replies": [
        {
          "id": 708931,
          "postDate": "2020-01-02T20:56:36.957Z",
          "content": "<p>It should be linked to the dataset. It's in the description. I just extracted the first frame with opencv videocapture and then resized using the function I linked and then cast to np.uint8. </p>",
          "rawMarkdown": "It should be linked to the dataset. It's in the description. I just extracted the first frame with opencv videocapture and then resized using the function I linked and then cast to np.uint8. "
        }
      ]
    },
    {
      "id": 707090,
      "postDate": "2019-12-31T08:43:16.923Z",
      "content": "<p>Great, the public validation set is really to small to train anything here. The next step would be to share only faces detected as 256x256 images on more frames.</p>",
      "rawMarkdown": "Great, the public validation set is really to small to train anything here. The next step would be to share only faces detected as 256x256 images on more frames."
    },
    {
      "id": 706918,
      "postDate": "2019-12-31T02:21:34.343Z",
      "content": "<p>Have you tried this dataset with resnet18? Do you have a corresponding score?</p>",
      "rawMarkdown": "Have you tried this dataset with resnet18? Do you have a corresponding score?",
      "replies": [
        {
          "id": 706956,
          "postDate": "2019-12-31T04:26:12.733Z",
          "content": "<p>I tried it with resnet 50 and the local cv log loss was around 2.3. I don't know what I am doing wrong.</p>",
          "rawMarkdown": "I tried it with resnet 50 and the local cv log loss was around 2.3. I don't know what I am doing wrong."
        },
        {
          "id": 706963,
          "postDate": "2019-12-31T05:01:22.053Z",
          "content": "<p>How did you split your train and val dataset? I have tested split-by-part and split-by-video. I did get a good val accuracy in my extracted validation dataset, but the online log-loss is really high.</p>",
          "rawMarkdown": "How did you split your train and val dataset? I have tested split-by-part and split-by-video. I did get a good val accuracy in my extracted validation dataset, but the online log-loss is really high."
        },
        {
          "id": 707322,
          "postDate": "2019-12-31T16:32:17.963Z",
          "content": "<p>I quit training when I saw the loss not decreasing after one epoch. What is your val accuracy? I guess I need to be more patient.</p>",
          "rawMarkdown": "I quit training when I saw the loss not decreasing after one epoch. What is your val accuracy? I guess I need to be more patient."
        }
      ]
    },
    {
      "id": 706064,
      "postDate": "2019-12-29T22:15:41.213Z",
      "content": "<p>Might be me - but network failure three different times trying to get your data set.</p>\n\n<p>Must have been me - 4th try it worked fine.</p>\n\n<p>Thanks for the work to get this for us.</p>",
      "rawMarkdown": "Might be me - but network failure three different times trying to get your data set.\n\nMust have been me - 4th try it worked fine.\n\nThanks for the work to get this for us.\n"
    },
    {
      "id": 706011,
      "postDate": "2019-12-29T20:07:44.837Z",
      "content": "<p>I am not sure about the public sharing of the data with people who don't accept the challenge rules first. From their <a href=\"https://deepfakedetectionchallenge.ai/faqs\">website</a>:</p>\n\n<blockquote>\n  <p><em>How are you protecting against adversaries who will try to access the code and data?</em>\n  We will be gating access to the dataset so that only researchers accepted into the challenge can access it. Each participant will need to agree to the terms of use on how they use, store, and handle the data. There are also strict restrictions on sharing the data. </p>\n</blockquote>",
      "rawMarkdown": "I am not sure about the public sharing of the data with people who don't accept the challenge rules first. From their [website](https://deepfakedetectionchallenge.ai/faqs):\n\n&gt; \n *How are you protecting against adversaries who will try to access the code and data?*\nWe will be gating access to the dataset so that only researchers accepted into the challenge can access it. Each participant will need to agree to the terms of use on how they use, store, and handle the data. There are also strict restrictions on sharing the data. \n&gt; ",
      "replies": [
        {
          "id": 706013,
          "postDate": "2019-12-29T20:16:12.067Z",
          "content": "<p>I will remove it if it is requested, but I don't immediately see an issue with this as it is still hosted and tracked by kaggle. </p>",
          "rawMarkdown": "I will remove it if it is requested, but I don't immediately see an issue with this as it is still hosted and tracked by kaggle. "
        },
        {
          "id": 706019,
          "postDate": "2019-12-29T20:20:32.013Z",
          "content": "<blockquote>\n  <p>May licensees of the DFDC Dataset perform minor pre-processing or augmentations on the DFDC Dataset (e.g., blurring, rotations and color enhancements/saturations)?\n  Licensees of the DFDC Dataset may perform minor pre-processing or augmentations of the DFDC Dataset (or portions thereof); provided, that such minor pre-processing or augmentations are not for the purpose of creating a new dataset (or portion thereof) and are only used for the authorized “Purpose” stated within the Deepfake Detection Challenge Dataset License Agreement. Licensees of the DFDC Dataset may also modify the DFDC Dataset solely for the purpose of enabling compatibility with Licensees’ computer systems</p>\n</blockquote>",
          "rawMarkdown": "&gt; May licensees of the DFDC Dataset perform minor pre-processing or augmentations on the DFDC Dataset (e.g., blurring, rotations and color enhancements/saturations)?\nLicensees of the DFDC Dataset may perform minor pre-processing or augmentations of the DFDC Dataset (or portions thereof); provided, that such minor pre-processing or augmentations are not for the purpose of creating a new dataset (or portion thereof) and are only used for the authorized “Purpose” stated within the Deepfake Detection Challenge Dataset License Agreement. Licensees of the DFDC Dataset may also modify the DFDC Dataset solely for the purpose of enabling compatibility with Licensees’ computer systems"
        },
        {
          "id": 706028,
          "postDate": "2019-12-29T20:33:03.767Z",
          "content": "<p><a href=\"/tayorm\">@tayorm</a>  The rules statement you shared clearly mean that the Public Sharing of data is Allowed as long as the sharing did not breach restrictions stated in \"Term of Use\" (which mainly related to standard legal and operation issues that is not related to good faith sharing data for competition propose).  </p>\n\n<p><a href=\"https://deepfakedetectionchallenge.ai/terms\">https://deepfakedetectionchallenge.ai/terms</a></p>",
          "rawMarkdown": "@tayorm  The rules statement you shared clearly mean that the Public Sharing of data is Allowed as long as the sharing did not breach restrictions stated in \"Term of Use\" (which mainly related to standard legal and operation issues that is not related to good faith sharing data for competition propose).  \n\nhttps://deepfakedetectionchallenge.ai/terms"
        },
        {
          "id": 706035,
          "postDate": "2019-12-29T20:40:24.543Z",
          "content": "<p>My reasoning is that by offering the data as a kaggle dataset any novice kaggler can download it without agreeing on the challenge terms of use. I am still trying to understand that paragraph. I hope competition organizers clear up this from a legal perspective.</p>",
          "rawMarkdown": "My reasoning is that by offering the data as a kaggle dataset any novice kaggler can download it without agreeing on the challenge terms of use. I am still trying to understand that paragraph. I hope competition organizers clear up this from a legal perspective."
        },
        {
          "id": 706048,
          "postDate": "2019-12-29T20:56:43.280Z",
          "content": "<p>I agree this is unclear, there are slightly conflicting statements on the kaggle site and the deepfake website. I don't think the organizers of the comp want no one to even be able to take a crack at the problem and to just exploit metadata leaks which is what this competition 99% is right now. </p>",
          "rawMarkdown": "I agree this is unclear, there are slightly conflicting statements on the kaggle site and the deepfake website. I don't think the organizers of the comp want no one to even be able to take a crack at the problem and to just exploit metadata leaks which is what this competition 99% is right now. "
        }
      ]
    },
    {
      "id": 777985,
      "postDate": "2020-03-18T04:00:04.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 706844,
      "postDate": "2019-12-30T23:00:09.460Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 706843,
      "postDate": "2019-12-30T22:58:02.907Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 708988,
      "postDate": "2020-01-02T22:30:12.327Z",
      "content": "<p>thank you for sharing this <a href=\"/ryches\">@ryches</a> !</p>",
      "rawMarkdown": "thank you for sharing this @ryches !"
    }
  ],
  "comments": [
    {
      "id": 706848,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2019-12-30T23:12:24.017000",
      "content": "<p>Looking forward to see your starter!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 708447,
      "author_name": "tarobxl",
      "author_url": "",
      "post_date": "2020-01-02T10:01:56.713000",
      "content": "<p><a href=\"/ryches\">@ryches</a> Many thanks. Could you kindly share your code to generate these frames as well? I think you did some good pre-processing steps.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 708931,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-01-02T20:56:36.957000",
          "content": "<p>It should be linked to the dataset. It's in the description. I just extracted the first frame with opencv videocapture and then resized using the function I linked and then cast to np.uint8. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 707090,
      "author_name": "MPWARE",
      "author_url": "",
      "post_date": "2019-12-31T08:43:16.923000",
      "content": "<p>Great, the public validation set is really to small to train anything here. The next step would be to share only faces detected as 256x256 images on more frames.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706918,
      "author_name": "xxn-xx",
      "author_url": "",
      "post_date": "2019-12-31T02:21:34.343000",
      "content": "<p>Have you tried this dataset with resnet18? Do you have a corresponding score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 706956,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2019-12-31T04:26:12.733000",
          "content": "<p>I tried it with resnet 50 and the local cv log loss was around 2.3. I don't know what I am doing wrong.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706963,
          "author_name": "xxn-xx",
          "author_url": "",
          "post_date": "2019-12-31T05:01:22.053000",
          "content": "<p>How did you split your train and val dataset? I have tested split-by-part and split-by-video. I did get a good val accuracy in my extracted validation dataset, but the online log-loss is really high.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 707322,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2019-12-31T16:32:17.963000",
          "content": "<p>I quit training when I saw the loss not decreasing after one epoch. What is your val accuracy? I guess I need to be more patient.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 706064,
      "author_name": "PC Jimmmy",
      "author_url": "",
      "post_date": "2019-12-29T22:15:41.213000",
      "content": "<p>Might be me - but network failure three different times trying to get your data set.</p>\n\n<p>Must have been me - 4th try it worked fine.</p>\n\n<p>Thanks for the work to get this for us.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706011,
      "author_name": "Mohammed Tayor",
      "author_url": "",
      "post_date": "2019-12-29T20:07:44.837000",
      "content": "<p>I am not sure about the public sharing of the data with people who don't accept the challenge rules first. From their <a href=\"https://deepfakedetectionchallenge.ai/faqs\">website</a>:</p>\n\n<blockquote>\n  <p><em>How are you protecting against adversaries who will try to access the code and data?</em>\n  We will be gating access to the dataset so that only researchers accepted into the challenge can access it. Each participant will need to agree to the terms of use on how they use, store, and handle the data. There are also strict restrictions on sharing the data. </p>\n</blockquote>",
      "votes": 0,
      "replies": [
        {
          "id": 706013,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-12-29T20:16:12.067000",
          "content": "<p>I will remove it if it is requested, but I don't immediately see an issue with this as it is still hosted and tracked by kaggle. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706019,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-12-29T20:20:32.013000",
          "content": "<blockquote>\n  <p>May licensees of the DFDC Dataset perform minor pre-processing or augmentations on the DFDC Dataset (e.g., blurring, rotations and color enhancements/saturations)?\n  Licensees of the DFDC Dataset may perform minor pre-processing or augmentations of the DFDC Dataset (or portions thereof); provided, that such minor pre-processing or augmentations are not for the purpose of creating a new dataset (or portion thereof) and are only used for the authorized “Purpose” stated within the Deepfake Detection Challenge Dataset License Agreement. Licensees of the DFDC Dataset may also modify the DFDC Dataset solely for the purpose of enabling compatibility with Licensees’ computer systems</p>\n</blockquote>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706028,
          "author_name": "Chew Kok Wah",
          "author_url": "",
          "post_date": "2019-12-29T20:33:03.767000",
          "content": "<p><a href=\"/tayorm\">@tayorm</a>  The rules statement you shared clearly mean that the Public Sharing of data is Allowed as long as the sharing did not breach restrictions stated in \"Term of Use\" (which mainly related to standard legal and operation issues that is not related to good faith sharing data for competition propose).  </p>\n\n<p><a href=\"https://deepfakedetectionchallenge.ai/terms\">https://deepfakedetectionchallenge.ai/terms</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706035,
          "author_name": "Mohammed Tayor",
          "author_url": "",
          "post_date": "2019-12-29T20:40:24.543000",
          "content": "<p>My reasoning is that by offering the data as a kaggle dataset any novice kaggler can download it without agreeing on the challenge terms of use. I am still trying to understand that paragraph. I hope competition organizers clear up this from a legal perspective.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706048,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-12-29T20:56:43.280000",
          "content": "<p>I agree this is unclear, there are slightly conflicting statements on the kaggle site and the deepfake website. I don't think the organizers of the comp want no one to even be able to take a crack at the problem and to just exploit metadata leaks which is what this competition 99% is right now. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 777985,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-03-18T04:00:04.580000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706844,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-30T23:00:09.460000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 706843,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-30T22:58:02.907000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 708988,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-02T22:30:12.327000",
      "content": "<p>thank you for sharing this <a href=\"/ryches\">@ryches</a> !</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "705993": "Made a simple dataset that captures the first frame of every video resized to 512x512 along with the labels. Along with this you should be able to train a full model within the kernel in 9 hours as well as make predictions since loading first frame is quite fast. Will make a sample showing a simple resnet18 approach to this that can achieve a respectable accuracy once I get some time\nhttps://www.kaggle.com/ryches/deepfake-first-frames-and-labels",
    "706848": "Looking forward to see your starter!",
    "708447": "@ryches Many thanks. Could you kindly share your code to generate these frames as well? I think you did some good pre-processing steps.",
    "707090": "Great, the public validation set is really to small to train anything here. The next step would be to share only faces detected as 256x256 images on more frames.",
    "706918": "Have you tried this dataset with resnet18? Do you have a corresponding score?",
    "706064": "Might be me - but network failure three different times trying to get your data set.\n\nMust have been me - 4th try it worked fine.\n\nThanks for the work to get this for us.\n",
    "706011": "I am not sure about the public sharing of the data with people who don't accept the challenge rules first. From their [website](https://deepfakedetectionchallenge.ai/faqs):\n\n&gt; \n *How are you protecting against adversaries who will try to access the code and data?*\nWe will be gating access to the dataset so that only researchers accepted into the challenge can access it. Each participant will need to agree to the terms of use on how they use, store, and handle the data. There are also strict restrictions on sharing the data. \n&gt; ",
    "777985": "",
    "706844": "",
    "706843": "",
    "708988": "thank you for sharing this @ryches !"
  }
}