{
  "id": 121664,
  "title": "Test Videos present in Train  dataset with label ?",
  "url": "/competitions/deepfake-detection-challenge/discussion/121664",
  "author_name": "",
  "post_date": "2019-12-14T16:45:15.286067600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I have just downloaded the set01 of train data and the first video is  aassnaulhq.mp4 which is also present in public test as first video .  </p>\n\n<p>The JSON also says that its a FAKE video . Since the dataset is huge , how much can we ensure there is no leak ? Or Am I misunderstanding something about this competition ?</p>\n\n<p>Anyone can clarify ?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F069d738b7a7d9ede66fae61bda11ea61%2FTrain_data_Set01.PNG?generation=1576341900915374&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F782279656c959e0dc48b7ef9f7d78d82%2FTrain_Metadata.PNG?generation=1576341907637855&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F2e038e5b9017763c19025080a9fa834d%2FPublic_Test.PNG?generation=1576341910537769&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "695142",
      "postDate": "12/14/2019 16:45:15",
      "content": "<p>I have just downloaded the set01 of train data and the first video is  aassnaulhq.mp4 which is also present in public test as first video .  </p>\n\n<p>The JSON also says that its a FAKE video . Since the dataset is huge , how much can we ensure there is no leak ? Or Am I misunderstanding something about this competition ?</p>\n\n<p>Anyone can clarify ?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F069d738b7a7d9ede66fae61bda11ea61%2FTrain_data_Set01.PNG?generation=1576341900915374&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F782279656c959e0dc48b7ef9f7d78d82%2FTrain_Metadata.PNG?generation=1576341907637855&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F2e038e5b9017763c19025080a9fa834d%2FPublic_Test.PNG?generation=1576341910537769&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "I have just downloaded the set01 of train data and the first video is  aassnaulhq.mp4 which is also present in public test as first video .  \n\nThe JSON also says that its a FAKE video . Since the dataset is huge , how much can we ensure there is no leak ? Or Am I misunderstanding something about this competition ?\n\nAnyone can clarify ?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F069d738b7a7d9ede66fae61bda11ea61%2FTrain_data_Set01.PNG?generation=1576341900915374&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F782279656c959e0dc48b7ef9f7d78d82%2FTrain_Metadata.PNG?generation=1576341907637855&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F2e038e5b9017763c19025080a9fa834d%2FPublic_Test.PNG?generation=1576341910537769&amp;alt=media)",
      "votes": null
    },
    {
      "id": "695156",
      "postDate": "12/14/2019 17:14:38",
      "content": "<p>This appears to be true for the entire public validation set (i.e. the <code>test_videos</code> folder). All 400 test videos are present in the full training set.</p>\n\n<p>(The notebook you submit isn't scored on these videos, so it's not really leakage.)</p>",
      "rawMarkdown": "This appears to be true for the entire public validation set (i.e. the `test_videos` folder). All 400 test videos are present in the full training set.\n\n(The notebook you submit isn't scored on these videos, so it's not really leakage.)",
      "votes": null
    },
    {
      "id": "697301",
      "postDate": "12/17/2019 17:28:32",
      "content": "<p>Hi! Thanks for your question. <a href=\"/humananalog\">@humananalog</a> is correct - the public test set is not included in scoring. It's just there to test your code against.</p>",
      "rawMarkdown": "Hi! Thanks for your question. @humananalog is correct - the public test set is not included in scoring. It's just there to test your code against.",
      "votes": null
    },
    {
      "id": "1247550",
      "postDate": "03/21/2021 20:12:50",
      "content": "<p>Hello, I am working on this competition now for testing a few things. I am a bit lost with the structure of the data, especially the test set and training set and I am not a native english speaker so just to be sure:<br>\n-the metadata.json file contains both the public training and the public testing set (aka videos in the training and testing set) ?<br>\n-the labels for the document in the testing set are also in the metadata.json file ? So the 400 documents cover both training set and testing set in this file ?<br>\nOr Am I getting it completely wrong and:<br>\n-the metadata.json is only for the videos in the training folder<br>\n-the submission.csv file is the one that actually contains the videos in the testing set along with their labels ?</p>",
      "rawMarkdown": "Hello, I am working on this competition now for testing a few things. I am a bit lost with the structure of the data, especially the test set and training set and I am not a native english speaker so just to be sure:\n-the metadata.json file contains both the public training and the public testing set (aka videos in the training and testing set) ?\n-the labels for the document in the testing set are also in the metadata.json file ? So the 400 documents cover both training set and testing set in this file ?\nOr Am I getting it completely wrong and:\n-the metadata.json is only for the videos in the training folder\n-the submission.csv file is the one that actually contains the videos in the testing set along with their labels ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1247550,
      "author_name": "tatatatata",
      "author_url": "",
      "post_date": "03/21/2021 20:12:50",
      "content": "<p>Hello, I am working on this competition now for testing a few things. I am a bit lost with the structure of the data, especially the test set and training set and I am not a native english speaker so just to be sure:<br>\n-the metadata.json file contains both the public training and the public testing set (aka videos in the training and testing set) ?<br>\n-the labels for the document in the testing set are also in the metadata.json file ? So the 400 documents cover both training set and testing set in this file ?<br>\nOr Am I getting it completely wrong and:<br>\n-the metadata.json is only for the videos in the training folder<br>\n-the submission.csv file is the one that actually contains the videos in the testing set along with their labels ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 695156,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "12/14/2019 17:14:38",
      "content": "<p>This appears to be true for the entire public validation set (i.e. the <code>test_videos</code> folder). All 400 test videos are present in the full training set.</p>\n\n<p>(The notebook you submit isn't scored on these videos, so it's not really leakage.)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 697301,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "12/17/2019 17:28:32",
      "content": "<p>Hi! Thanks for your question. <a href=\"/humananalog\">@humananalog</a> is correct - the public test set is not included in scoring. It's just there to test your code against.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "695142": "I have just downloaded the set01 of train data and the first video is  aassnaulhq.mp4 which is also present in public test as first video .  \n\nThe JSON also says that its a FAKE video . Since the dataset is huge , how much can we ensure there is no leak ? Or Am I misunderstanding something about this competition ?\n\nAnyone can clarify ?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F069d738b7a7d9ede66fae61bda11ea61%2FTrain_data_Set01.PNG?generation=1576341900915374&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F782279656c959e0dc48b7ef9f7d78d82%2FTrain_Metadata.PNG?generation=1576341907637855&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F2e038e5b9017763c19025080a9fa834d%2FPublic_Test.PNG?generation=1576341910537769&amp;alt=media)",
    "695156": "This appears to be true for the entire public validation set (i.e. the `test_videos` folder). All 400 test videos are present in the full training set.\n\n(The notebook you submit isn't scored on these videos, so it's not really leakage.)",
    "697301": "Hi! Thanks for your question. @humananalog is correct - the public test set is not included in scoring. It's just there to test your code against.",
    "1247550": "Hello, I am working on this competition now for testing a few things. I am a bit lost with the structure of the data, especially the test set and training set and I am not a native english speaker so just to be sure:\n-the metadata.json file contains both the public training and the public testing set (aka videos in the training and testing set) ?\n-the labels for the document in the testing set are also in the metadata.json file ? So the 400 documents cover both training set and testing set in this file ?\nOr Am I getting it completely wrong and:\n-the metadata.json is only for the videos in the training folder\n-the submission.csv file is the one that actually contains the videos in the testing set along with their labels ?"
  },
  "source": "meta"
}