{
  "id": 126108,
  "title": "where are the labels for the public validation set?",
  "url": "/competitions/deepfake-detection-challenge/discussion/126108",
  "author_name": "zjiang",
  "post_date": "2020-01-15T18:19:41.190000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I want to test the performance of our method on the public validation set but I cannot find the labels for the 400 testing videos in the <code>deepfake-detection-challenge</code> directory.  Have I misunderstood something? </p>",
  "messages": [
    {
      "id": 719692,
      "postDate": "2020-01-15T18:35:13.307Z",
      "content": "<p>Hi <a href=\"/zhuolin\">@zhuolin</a> it's in the metadata.json:</p>\n\n<p><code>/kaggle/input/deepfake-detection-challenge/train_sample_videos/metadata.json</code></p>\n\n<p><code>\n{\"aagfhgtpmv.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"vudstovrck.mp4\"},\n\"aapnvogymq.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"jdubbvfswz.mp4\"},\n\"abarnvbtwb.mp4\":{\"label\":\"REAL\",\"split\":\"train\",\"original\":null},\n...\n</code></p>\n\n<p>You should use the videos under '<strong>test_videos</strong>' for your submission.</p>",
      "rawMarkdown": "Hi @zhuolin it's in the metadata.json:\n\n`/kaggle/input/deepfake-detection-challenge/train_sample_videos/metadata.json`\n\n```\n{\"aagfhgtpmv.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"vudstovrck.mp4\"},\n\"aapnvogymq.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"jdubbvfswz.mp4\"},\n\"abarnvbtwb.mp4\":{\"label\":\"REAL\",\"split\":\"train\",\"original\":null},\n...\n```\n\nYou should use the videos under '**test_videos**' for your submission.",
      "votes": 2
    },
    {
      "id": 1046562,
      "postDate": "2020-10-11T19:26:11.457Z",
      "content": "<p>Were the labels of the DFDC test dataset ever released? I wanted to test my model with this test-set for educational purposes.</p>",
      "rawMarkdown": "Were the labels of the DFDC test dataset ever released? I wanted to test my model with this test-set for educational purposes.",
      "replies": [
        {
          "id": 1046725,
          "postDate": "2020-10-12T00:11:28.943Z",
          "content": "<p>The labels for the private test (DFDC-like data) can be found here: <a href=\"https://ai.facebook.com/datasets/dfdc\" target=\"_blank\">https://ai.facebook.com/datasets/dfdc</a></p>",
          "rawMarkdown": "The labels for the private test (DFDC-like data) can be found here: https://ai.facebook.com/datasets/dfdc"
        }
      ]
    },
    {
      "id": 719702,
      "postDate": "2020-01-15T18:47:12.453Z",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a> <a href=\"/vishwasaha\">@vishwasaha</a> @nosound Thank you all. It sounds that <code>test_video</code> directory will be replaced by any dataset during the evaluation, so the labels are not provided.  If I want to test the performance locally, I can use the videos in the <code>train_sample_videos</code>.  How can I get the results on the public validation set? Thanks.</p>",
      "rawMarkdown": "@hmendonca @vishwasaha @nosound Thank you all. It sounds that `test_video` directory will be replaced by any dataset during the evaluation, so the labels are not provided.  If I want to test the performance locally, I can use the videos in the `train_sample_videos`.  How can I get the results on the public validation set? Thanks."
    },
    {
      "id": 719690,
      "postDate": "2020-01-15T18:29:52.813Z",
      "content": "<p>Validation set labels are not provided in the competition data that is hosted on Kaggle. But it is a subset of the full train set, so you can find the labels there. I maintain a dataset with the full train data labels, metadata etc., <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">it is here</a>. <code>pd.read_csv()</code> on the file and take a subset by filenames.</p>",
      "rawMarkdown": "Validation set labels are not provided in the competition data that is hosted on Kaggle. But it is a subset of the full train set, so you can find the labels there. I maintain a dataset with the full train data labels, metadata etc., [it is here](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc). `pd.read_csv()` on the file and take a subset by filenames."
    },
    {
      "id": 719681,
      "postDate": "2020-01-15T18:19:41.190Z",
      "content": "<p>I want to test the performance of our method on the public validation set but I cannot find the labels for the 400 testing videos in the <code>deepfake-detection-challenge</code> directory.  Have I misunderstood something? </p>",
      "rawMarkdown": "I want to test the performance of our method on the public validation set but I cannot find the labels for the 400 testing videos in the `deepfake-detection-challenge` directory.  Have I misunderstood something? "
    },
    {
      "id": 719693,
      "postDate": "2020-01-15T18:35:32.797Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 719692,
      "author_name": "Henrique Mendonça",
      "author_url": "",
      "post_date": "2020-01-15T18:35:13.307000",
      "content": "<p>Hi <a href=\"/zhuolin\">@zhuolin</a> it's in the metadata.json:</p>\n\n<p><code>/kaggle/input/deepfake-detection-challenge/train_sample_videos/metadata.json</code></p>\n\n<p><code>\n{\"aagfhgtpmv.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"vudstovrck.mp4\"},\n\"aapnvogymq.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"jdubbvfswz.mp4\"},\n\"abarnvbtwb.mp4\":{\"label\":\"REAL\",\"split\":\"train\",\"original\":null},\n...\n</code></p>\n\n<p>You should use the videos under '<strong>test_videos</strong>' for your submission.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1046562,
      "author_name": "pratikpv",
      "author_url": "",
      "post_date": "2020-10-11T19:26:11.457000",
      "content": "<p>Were the labels of the DFDC test dataset ever released? I wanted to test my model with this test-set for educational purposes.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1046725,
          "author_name": "Mozaic",
          "author_url": "",
          "post_date": "2020-10-12T00:11:28.943000",
          "content": "<p>The labels for the private test (DFDC-like data) can be found here: <a href=\"https://ai.facebook.com/datasets/dfdc\" target=\"_blank\">https://ai.facebook.com/datasets/dfdc</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 719702,
      "author_name": "zjiang",
      "author_url": "",
      "post_date": "2020-01-15T18:47:12.453000",
      "content": "<p><a href=\"/hmendonca\">@hmendonca</a> <a href=\"/vishwasaha\">@vishwasaha</a> @nosound Thank you all. It sounds that <code>test_video</code> directory will be replaced by any dataset during the evaluation, so the labels are not provided.  If I want to test the performance locally, I can use the videos in the <code>train_sample_videos</code>.  How can I get the results on the public validation set? Thanks.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 719690,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2020-01-15T18:29:52.813000",
      "content": "<p>Validation set labels are not provided in the competition data that is hosted on Kaggle. But it is a subset of the full train set, so you can find the labels there. I maintain a dataset with the full train data labels, metadata etc., <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">it is here</a>. <code>pd.read_csv()</code> on the file and take a subset by filenames.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 719693,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-15T18:35:32.797000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "719692": "Hi @zhuolin it's in the metadata.json:\n\n`/kaggle/input/deepfake-detection-challenge/train_sample_videos/metadata.json`\n\n```\n{\"aagfhgtpmv.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"vudstovrck.mp4\"},\n\"aapnvogymq.mp4\":{\"label\":\"FAKE\",\"split\":\"train\",\"original\":\"jdubbvfswz.mp4\"},\n\"abarnvbtwb.mp4\":{\"label\":\"REAL\",\"split\":\"train\",\"original\":null},\n...\n```\n\nYou should use the videos under '**test_videos**' for your submission.",
    "1046562": "Were the labels of the DFDC test dataset ever released? I wanted to test my model with this test-set for educational purposes.",
    "719702": "@hmendonca @vishwasaha @nosound Thank you all. It sounds that `test_video` directory will be replaced by any dataset during the evaluation, so the labels are not provided.  If I want to test the performance locally, I can use the videos in the `train_sample_videos`.  How can I get the results on the public validation set? Thanks.",
    "719690": "Validation set labels are not provided in the competition data that is hosted on Kaggle. But it is a subset of the full train set, so you can find the labels there. I maintain a dataset with the full train data labels, metadata etc., [it is here](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc). `pd.read_csv()` on the file and take a subset by filenames.",
    "719681": "I want to test the performance of our method on the public validation set but I cannot find the labels for the 400 testing videos in the `deepfake-detection-challenge` directory.  Have I misunderstood something? ",
    "719693": ""
  }
}