{
  "id": 128650,
  "title": "Another data leak in public test set?",
  "url": "/competitions/deepfake-detection-challenge/discussion/128650",
  "author_name": "",
  "post_date": "2020-02-02T05:46:53.049186500Z",
  "votes": 3,
  "comment_count": 5,
  "views": 0,
  "content": "<p>@zaharch found that there are 27 corrupted videos in the public test set. And <a href=\"https://www.kaggle.com/wuliaokaola/public-test-set-probe\">this kernel</a> can get LB 0.68846 without any model. That means all of the corrupted videos are fake. \n<code>0.69314 * (4000 - 27) / 4000 = 0.68846</code>\nI think when they generated deepfake videos, something went wrong that made video corrupted. @philculliton I don't know how about the private test set, but I think this is a leakage of public test set.</p>",
  "messages": [
    {
      "id": "734873",
      "postDate": "02/02/2020 05:46:53",
      "content": "<p>@zaharch found that there are 27 corrupted videos in the public test set. And <a href=\"https://www.kaggle.com/wuliaokaola/public-test-set-probe\">this kernel</a> can get LB 0.68846 without any model. That means all of the corrupted videos are fake. \n<code>0.69314 * (4000 - 27) / 4000 = 0.68846</code>\nI think when they generated deepfake videos, something went wrong that made video corrupted. @philculliton I don't know how about the private test set, but I think this is a leakage of public test set.</p>",
      "rawMarkdown": "zaharch found that there are 27 corrupted videos in the public test set. And [this kernel](https://www.kaggle.com/wuliaokaola/public-test-set-probe) can get LB 0.68846 without any model. That means all of the corrupted videos are fake. \n`0.69314 * (4000 - 27) / 4000 = 0.68846`\nI think when they generated deepfake videos, something went wrong that made video corrupted. @philculliton I don't know how about the private test set, but I think this is a leakage of public test set.",
      "votes": null
    },
    {
      "id": "734904",
      "postDate": "02/02/2020 07:14:49",
      "content": "<p>In many of the shared kernels when an error occurs most folks are using 0.5 as the default probability to use.</p>\n\n<p>This suggests that a value close to 1 would make more sense when a file error occurs.  Many might have already decided that 0.5 was not the best to use for error default - I have been using 0.75 myself.</p>\n\n<p>Not sure that 27/4000 is a significant leak however.</p>",
      "rawMarkdown": "In many of the shared kernels when an error occurs most folks are using 0.5 as the default probability to use.\n\nThis suggests that a value close to 1 would make more sense when a file error occurs.  Many might have already decided that 0.5 was not the best to use for error default - I have been using 0.75 myself.\n\nNot sure that 27/4000 is a significant leak however.",
      "votes": null
    },
    {
      "id": "736122",
      "postDate": "02/03/2020 21:01:57",
      "content": "<p>This is one of those interesting situations where there's a real risk of overfitting to the public leaderboard. I currently predict 0.5 on any video that fails. I suspect my LB score could be improved by picking a higher score for failures, but with the logarithmic loss function there's a risk from it, too.\nI suspect failures are more likely to be deepfakes even in the final private test set, but it'd be better if we could get reassurance that such errors would be cleaned up prior?</p>",
      "rawMarkdown": "This is one of those interesting situations where there's a real risk of overfitting to the public leaderboard. I currently predict 0.5 on any video that fails. I suspect my LB score could be improved by picking a higher score for failures, but with the logarithmic loss function there's a risk from it, too.\nI suspect failures are more likely to be deepfakes even in the final private test set, but it'd be better if we could get reassurance that such errors would be cleaned up prior?",
      "votes": null
    },
    {
      "id": "736137",
      "postDate": "02/03/2020 21:35:09",
      "content": "<p>The final private test set would seem to me to have no failures or will a whole bunch more than 27.  It's pretty big shot in the dark.  But data from the wild tends to have more errors than a nicely developed data set.  </p>\n\n<p>If I was you - sitting in  4th place than I would absolutely use the safe position of 0.5.  I am not likely to be sitting so close to the top - as slow as my progress has been will almost certainly be out of the metals range.  So I am going to probably choose to gamble and go even higher than 0.75.</p>",
      "rawMarkdown": "The final private test set would seem to me to have no failures or will a whole bunch more than 27.  It's pretty big shot in the dark.  But data from the wild tends to have more errors than a nicely developed data set.  \n\nIf I was you - sitting in  4th place than I would absolutely use the safe position of 0.5.  I am not likely to be sitting so close to the top - as slow as my progress has been will almost certainly be out of the metals range.  So I am going to probably choose to gamble and go even higher than 0.75.",
      "votes": null
    },
    {
      "id": "737130",
      "postDate": "02/05/2020 00:12:30",
      "content": "<p>Thanks for raising this. For the private test set, we have been noting all of this feedback so that it goes smooth. My recommendation would be to try not to overfit or leverage these glitches. </p>",
      "rawMarkdown": "Thanks for raising this. For the private test set, we have been noting all of this feedback so that it goes smooth. My recommendation would be to try not to overfit or leverage these glitches.",
      "votes": null
    },
    {
      "id": "737142",
      "postDate": "02/05/2020 00:42:13",
      "content": "<p>Thank you for your reply and advice.</p>",
      "rawMarkdown": "Thank you for your reply and advice.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 734904,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "02/02/2020 07:14:49",
      "content": "<p>In many of the shared kernels when an error occurs most folks are using 0.5 as the default probability to use.</p>\n\n<p>This suggests that a value close to 1 would make more sense when a file error occurs.  Many might have already decided that 0.5 was not the best to use for error default - I have been using 0.75 myself.</p>\n\n<p>Not sure that 27/4000 is a significant leak however.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 736122,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "02/03/2020 21:01:57",
      "content": "<p>This is one of those interesting situations where there's a real risk of overfitting to the public leaderboard. I currently predict 0.5 on any video that fails. I suspect my LB score could be improved by picking a higher score for failures, but with the logarithmic loss function there's a risk from it, too.\nI suspect failures are more likely to be deepfakes even in the final private test set, but it'd be better if we could get reassurance that such errors would be cleaned up prior?</p>",
      "votes": null,
      "replies": [
        {
          "id": 736137,
          "author_name": "pcjimmmy",
          "author_url": "",
          "post_date": "02/03/2020 21:35:09",
          "content": "<p>The final private test set would seem to me to have no failures or will a whole bunch more than 27.  It's pretty big shot in the dark.  But data from the wild tends to have more errors than a nicely developed data set.  </p>\n\n<p>If I was you - sitting in  4th place than I would absolutely use the safe position of 0.5.  I am not likely to be sitting so close to the top - as slow as my progress has been will almost certainly be out of the metals range.  So I am going to probably choose to gamble and go even higher than 0.75.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 737130,
      "author_name": "cristiancanton",
      "author_url": "",
      "post_date": "02/05/2020 00:12:30",
      "content": "<p>Thanks for raising this. For the private test set, we have been noting all of this feedback so that it goes smooth. My recommendation would be to try not to overfit or leverage these glitches. </p>",
      "votes": null,
      "replies": [
        {
          "id": 737142,
          "author_name": "wuliaokaola",
          "author_url": "",
          "post_date": "02/05/2020 00:42:13",
          "content": "<p>Thank you for your reply and advice.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "734873": "zaharch found that there are 27 corrupted videos in the public test set. And [this kernel](https://www.kaggle.com/wuliaokaola/public-test-set-probe) can get LB 0.68846 without any model. That means all of the corrupted videos are fake. \n`0.69314 * (4000 - 27) / 4000 = 0.68846`\nI think when they generated deepfake videos, something went wrong that made video corrupted. @philculliton I don't know how about the private test set, but I think this is a leakage of public test set.",
    "734904": "In many of the shared kernels when an error occurs most folks are using 0.5 as the default probability to use.\n\nThis suggests that a value close to 1 would make more sense when a file error occurs.  Many might have already decided that 0.5 was not the best to use for error default - I have been using 0.75 myself.\n\nNot sure that 27/4000 is a significant leak however.",
    "736122": "This is one of those interesting situations where there's a real risk of overfitting to the public leaderboard. I currently predict 0.5 on any video that fails. I suspect my LB score could be improved by picking a higher score for failures, but with the logarithmic loss function there's a risk from it, too.\nI suspect failures are more likely to be deepfakes even in the final private test set, but it'd be better if we could get reassurance that such errors would be cleaned up prior?",
    "736137": "The final private test set would seem to me to have no failures or will a whole bunch more than 27.  It's pretty big shot in the dark.  But data from the wild tends to have more errors than a nicely developed data set.  \n\nIf I was you - sitting in  4th place than I would absolutely use the safe position of 0.5.  I am not likely to be sitting so close to the top - as slow as my progress has been will almost certainly be out of the metals range.  So I am going to probably choose to gamble and go even higher than 0.75.",
    "737130": "Thanks for raising this. For the private test set, we have been noting all of this feedback so that it goes smooth. My recommendation would be to try not to overfit or leverage these glitches.",
    "737142": "Thank you for your reply and advice."
  },
  "source": "meta"
}