{
  "id": 129000,
  "title": "Mismatch between local test and public test result",
  "url": "/competitions/deepfake-detection-challenge/discussion/129000",
  "author_name": "",
  "post_date": "2020-02-04T19:03:37.284949200Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Could anyone help to figure out this? </p>\n\n<p>I trained my model and tested it locally with 4000 random selected test videos from the full data set. It contains equal number of REAL and FAKE videos. \nHere is the result of my test: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4260855%2Ff16dce44fcd5e4714f91730b924c441b%2F0204_image_01.png?generation=1580842596992905&amp;alt=media\" alt=\"\"></p>\n\n<p>But whenever I uploaded the model on kaggle, before I committed, I predicted on the train -sample-videos, the distribution of the result looks alright. But after I committed, I always get a result close to naive prediction 0.693. </p>\n\n<p>(1) I wanna make sure whether randomly selected 4000 videos from the full data set is a robust test set to get a feeling about the result. \n(2) If there is any problem about the submission, could anyone give some clues about it? It didn't report any error during my submission. </p>\n\n<p>Thanks very much for the help! </p>",
  "messages": [
    {
      "id": "736976",
      "postDate": "02/04/2020 19:03:37",
      "content": "<p>Could anyone help to figure out this? </p>\n\n<p>I trained my model and tested it locally with 4000 random selected test videos from the full data set. It contains equal number of REAL and FAKE videos. \nHere is the result of my test: \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4260855%2Ff16dce44fcd5e4714f91730b924c441b%2F0204_image_01.png?generation=1580842596992905&amp;alt=media\" alt=\"\"></p>\n\n<p>But whenever I uploaded the model on kaggle, before I committed, I predicted on the train -sample-videos, the distribution of the result looks alright. But after I committed, I always get a result close to naive prediction 0.693. </p>\n\n<p>(1) I wanna make sure whether randomly selected 4000 videos from the full data set is a robust test set to get a feeling about the result. \n(2) If there is any problem about the submission, could anyone give some clues about it? It didn't report any error during my submission. </p>\n\n<p>Thanks very much for the help! </p>",
      "rawMarkdown": "Could anyone help to figure out this? \n\nI trained my model and tested it locally with 4000 random selected test videos from the full data set. It contains equal number of REAL and FAKE videos. \nHere is the result of my test: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4260855%2Ff16dce44fcd5e4714f91730b924c441b%2F0204_image_01.png?generation=1580842596992905&amp;alt=media)\n \nBut whenever I uploaded the model on kaggle, before I committed, I predicted on the train -sample-videos, the distribution of the result looks alright. But after I committed, I always get a result close to naive prediction 0.693. \n\n(1) I wanna make sure whether randomly selected 4000 videos from the full data set is a robust test set to get a feeling about the result. \n(2) If there is any problem about the submission, could anyone give some clues about it? It didn't report any error during my submission. \n\nThanks very much for the help!",
      "votes": null
    },
    {
      "id": "736980",
      "postDate": "02/04/2020 19:06:17",
      "content": "<p>Hi <a href=\"/jeffexu\">@jeffexu</a> \nThe same topic is covered here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128919\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128919</a></p>",
      "rawMarkdown": "Hi @jeffexu \nThe same topic is covered here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128919",
      "votes": null
    },
    {
      "id": "736998",
      "postDate": "02/04/2020 19:34:24",
      "content": "<p>Thanks! This helps a lot! </p>",
      "rawMarkdown": "Thanks! This helps a lot!",
      "votes": null
    },
    {
      "id": "737081",
      "postDate": "02/04/2020 22:25:34",
      "content": "<p>How quickly does your code run? If your submission is much less than 10 times the time needed to commit then I imagine you’re getting an exception which you don’t catch and your code is processing only a small bit of the dataset.</p>\n\n<p>Quite a few public test set videos cause exceptions not seen on the public validation set, I found.</p>",
      "rawMarkdown": "How quickly does your code run? If your submission is much less than 10 times the time needed to commit then I imagine you’re getting an exception which you don’t catch and your code is processing only a small bit of the dataset.\n\nQuite a few public test set videos cause exceptions not seen on the public validation set, I found.",
      "votes": null
    },
    {
      "id": "737090",
      "postDate": "02/04/2020 22:49:03",
      "content": "<p>The time is around 10 times of the commit time. I guess the problem lies in the validation leak. BTW, does anyone know that if I commit with GPU on, will the code also run in the background with GPU? </p>",
      "rawMarkdown": "The time is around 10 times of the commit time. I guess the problem lies in the validation leak. BTW, does anyone know that if I commit with GPU on, will the code also run in the background with GPU?",
      "votes": null
    },
    {
      "id": "737095",
      "postDate": "02/04/2020 23:06:09",
      "content": "<p>Of course, if you commit with GPU both the commit and the submission afterwards are running accordingly. And if I understand correctly, commit eats from the GPU quota and the submission does not. </p>",
      "rawMarkdown": "Of course, if you commit with GPU both the commit and the submission afterwards are running accordingly. And if I understand correctly, commit eats from the GPU quota and the submission does not.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 736980,
      "author_name": "hmendonca",
      "author_url": "",
      "post_date": "02/04/2020 19:06:17",
      "content": "<p>Hi <a href=\"/jeffexu\">@jeffexu</a> \nThe same topic is covered here: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128919\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128919</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 736998,
          "author_name": "jeffexu",
          "author_url": "",
          "post_date": "02/04/2020 19:34:24",
          "content": "<p>Thanks! This helps a lot! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 737081,
      "author_name": "jamesphoward",
      "author_url": "",
      "post_date": "02/04/2020 22:25:34",
      "content": "<p>How quickly does your code run? If your submission is much less than 10 times the time needed to commit then I imagine you’re getting an exception which you don’t catch and your code is processing only a small bit of the dataset.</p>\n\n<p>Quite a few public test set videos cause exceptions not seen on the public validation set, I found.</p>",
      "votes": null,
      "replies": [
        {
          "id": 737090,
          "author_name": "jeffexu",
          "author_url": "",
          "post_date": "02/04/2020 22:49:03",
          "content": "<p>The time is around 10 times of the commit time. I guess the problem lies in the validation leak. BTW, does anyone know that if I commit with GPU on, will the code also run in the background with GPU? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 737095,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "02/04/2020 23:06:09",
          "content": "<p>Of course, if you commit with GPU both the commit and the submission afterwards are running accordingly. And if I understand correctly, commit eats from the GPU quota and the submission does not. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "736976": "Could anyone help to figure out this? \n\nI trained my model and tested it locally with 4000 random selected test videos from the full data set. It contains equal number of REAL and FAKE videos. \nHere is the result of my test: \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4260855%2Ff16dce44fcd5e4714f91730b924c441b%2F0204_image_01.png?generation=1580842596992905&amp;alt=media)\n \nBut whenever I uploaded the model on kaggle, before I committed, I predicted on the train -sample-videos, the distribution of the result looks alright. But after I committed, I always get a result close to naive prediction 0.693. \n\n(1) I wanna make sure whether randomly selected 4000 videos from the full data set is a robust test set to get a feeling about the result. \n(2) If there is any problem about the submission, could anyone give some clues about it? It didn't report any error during my submission. \n\nThanks very much for the help!",
    "736980": "Hi @jeffexu \nThe same topic is covered here: https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128919",
    "736998": "Thanks! This helps a lot!",
    "737081": "How quickly does your code run? If your submission is much less than 10 times the time needed to commit then I imagine you’re getting an exception which you don’t catch and your code is processing only a small bit of the dataset.\n\nQuite a few public test set videos cause exceptions not seen on the public validation set, I found.",
    "737090": "The time is around 10 times of the commit time. I guess the problem lies in the validation leak. BTW, does anyone know that if I commit with GPU on, will the code also run in the background with GPU?",
    "737095": "Of course, if you commit with GPU both the commit and the submission afterwards are running accordingly. And if I understand correctly, commit eats from the GPU quota and the submission does not."
  },
  "source": "meta"
}