{
  "id": 131313,
  "title": "Are test labels wrong?",
  "url": "/competitions/deepfake-detection-challenge/discussion/131313",
  "author_name": "",
  "post_date": "2020-02-19T08:43:32.359600300Z",
  "votes": null,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Looking at ahjnxtiamx.mp4 test video it's clearly fake.</p>\n\n<p>When submitting a notebook with that label fake you would expect the LB score to be better than the 0.5 benchmark. However it's the same as the 0.5 benchmark (0.69314)! \n<a href=\"https://www.kaggle.com/dthrone/are-test-labels-wrong\">See in this notebook: Are test labels wrong?</a></p>\n\n<p>Does anyone have an explanation for this?</p>\n\n<p>Is the scoring performed on a random subset? \nIs the test set supposed to be reproducible when submitting the same result multiple times? If yes, are the labels wrong?</p>",
  "messages": [
    {
      "id": "750270",
      "postDate": "02/19/2020 08:43:32",
      "content": "<p>Looking at ahjnxtiamx.mp4 test video it's clearly fake.</p>\n\n<p>When submitting a notebook with that label fake you would expect the LB score to be better than the 0.5 benchmark. However it's the same as the 0.5 benchmark (0.69314)! \n<a href=\"https://www.kaggle.com/dthrone/are-test-labels-wrong\">See in this notebook: Are test labels wrong?</a></p>\n\n<p>Does anyone have an explanation for this?</p>\n\n<p>Is the scoring performed on a random subset? \nIs the test set supposed to be reproducible when submitting the same result multiple times? If yes, are the labels wrong?</p>",
      "rawMarkdown": "Looking at ahjnxtiamx.mp4 test video it's clearly fake.\n\nWhen submitting a notebook with that label fake you would expect the LB score to be better than the 0.5 benchmark. However it's the same as the 0.5 benchmark (0.69314)! \n[See in this notebook: Are test labels wrong?](https://www.kaggle.com/dthrone/are-test-labels-wrong)\n\nDoes anyone have an explanation for this?\n\nIs the scoring performed on a random subset? \nIs the test set supposed to be reproducible when submitting the same result multiple times? If yes, are the labels wrong?",
      "votes": null
    },
    {
      "id": "750279",
      "postDate": "02/19/2020 08:52:47",
      "content": "<p>Test videos are a subsample from the full dataset. I would encourage you to use a wider test set (~20% == 10 parts) in order to cope with overfiting. Moreover, <a href=\"/harshitsheoran\">@harshitsheoran</a> propose to use parts 40-50 to test <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/130915\">here</a></p>",
      "rawMarkdown": "Test videos are a subsample from the full dataset. I would encourage you to use a wider test set (~20% == 10 parts) in order to cope with overfiting. Moreover, @harshitsheoran propose to use parts 40-50 to test [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/130915)",
      "votes": null
    },
    {
      "id": "750289",
      "postDate": "02/19/2020 09:06:58",
      "content": "<p>In getting started.\"3.Public Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. \"\nI think the public validation set is not included in the public test set.</p>",
      "rawMarkdown": "In getting started.\"3.Public Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. \"\nI think the public validation set is not included in the public test set.",
      "votes": null
    },
    {
      "id": "750309",
      "postDate": "02/19/2020 09:20:53",
      "content": "<p>Thank you, this is not what I have expected. Reading your comments I have the same conclusion. If the test_videos is a subset of the dataset the LB is calculated on the score should be different. The only explanation I have is that the test_videos files are not included in the set that the LB is calculated on.</p>",
      "rawMarkdown": "Thank you, this is not what I have expected. Reading your comments I have the same conclusion. If the test_videos is a subset of the dataset the LB is calculated on the score should be different. The only explanation I have is that the test_videos files are not included in the set that the LB is calculated on.",
      "votes": null
    },
    {
      "id": "750322",
      "postDate": "02/19/2020 09:40:14",
      "content": "<p>As <a href=\"/kueno55\">@kueno55</a> has said in <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started</a> you have all the information regarding to the different datasets.</p>",
      "rawMarkdown": "As @kueno55 has said in [Getting Started](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started) you have all the information regarding to the different datasets.",
      "votes": null
    },
    {
      "id": "750424",
      "postDate": "02/19/2020 11:04:02",
      "content": "<p>All of the <code>test_videos</code> are in the training data, so it would be very strange if they were also included in the LB test set. ;-)</p>",
      "rawMarkdown": "All of the `test_videos` are in the training data, so it would be very strange if they were also included in the LB test set. ;-)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 750279,
      "author_name": "cayala",
      "author_url": "",
      "post_date": "02/19/2020 08:52:47",
      "content": "<p>Test videos are a subsample from the full dataset. I would encourage you to use a wider test set (~20% == 10 parts) in order to cope with overfiting. Moreover, <a href=\"/harshitsheoran\">@harshitsheoran</a> propose to use parts 40-50 to test <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/130915\">here</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 750289,
      "author_name": "kueno55",
      "author_url": "",
      "post_date": "02/19/2020 09:06:58",
      "content": "<p>In getting started.\"3.Public Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. \"\nI think the public validation set is not included in the public test set.</p>",
      "votes": null,
      "replies": [
        {
          "id": 750309,
          "author_name": "dthrone",
          "author_url": "",
          "post_date": "02/19/2020 09:20:53",
          "content": "<p>Thank you, this is not what I have expected. Reading your comments I have the same conclusion. If the test_videos is a subset of the dataset the LB is calculated on the score should be different. The only explanation I have is that the test_videos files are not included in the set that the LB is calculated on.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750322,
          "author_name": "cayala",
          "author_url": "",
          "post_date": "02/19/2020 09:40:14",
          "content": "<p>As <a href=\"/kueno55\">@kueno55</a> has said in <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started</a> you have all the information regarding to the different datasets.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 750424,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "02/19/2020 11:04:02",
      "content": "<p>All of the <code>test_videos</code> are in the training data, so it would be very strange if they were also included in the LB test set. ;-)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "750270": "Looking at ahjnxtiamx.mp4 test video it's clearly fake.\n\nWhen submitting a notebook with that label fake you would expect the LB score to be better than the 0.5 benchmark. However it's the same as the 0.5 benchmark (0.69314)! \n[See in this notebook: Are test labels wrong?](https://www.kaggle.com/dthrone/are-test-labels-wrong)\n\nDoes anyone have an explanation for this?\n\nIs the scoring performed on a random subset? \nIs the test set supposed to be reproducible when submitting the same result multiple times? If yes, are the labels wrong?",
    "750279": "Test videos are a subsample from the full dataset. I would encourage you to use a wider test set (~20% == 10 parts) in order to cope with overfiting. Moreover, @harshitsheoran propose to use parts 40-50 to test [here](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/130915)",
    "750289": "In getting started.\"3.Public Test Set: This dataset is completely withheld and is what Kaggle’s platform computes the public leaderboard against. \"\nI think the public validation set is not included in the public test set.",
    "750309": "Thank you, this is not what I have expected. Reading your comments I have the same conclusion. If the test_videos is a subset of the dataset the LB is calculated on the score should be different. The only explanation I have is that the test_videos files are not included in the set that the LB is calculated on.",
    "750322": "As @kueno55 has said in [Getting Started](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started) you have all the information regarding to the different datasets.",
    "750424": "All of the `test_videos` are in the training data, so it would be very strange if they were also included in the LB test set. ;-)"
  },
  "source": "meta"
}