{
  "id": 122634,
  "title": "Public test set size is 4000",
  "url": "/competitions/deepfake-detection-challenge/discussion/122634",
  "author_name": "nosound",
  "post_date": "2019-12-21T15:56:34.283000",
  "votes": 30,
  "comment_count": 16,
  "views": 0,
  "content": "<p>All <code>0.5</code> submission scores <code>0.69314</code>, first row changed to <code>0.999</code> scores <code>0.69470</code>. It means that the possible range is <code>3977</code> to <code>4002</code>. And then submission with the following row passed successfully</p>\n\n<p><code>assert (len(df) == 400) or (len(df) == 4000)</code></p>\n\n<p>Additionally, there are exactly 2000 samples of each class.</p>\n\n<p>Interesting, what is the private test size?</p>",
  "messages": [
    {
      "id": 700199,
      "postDate": "2019-12-21T15:56:34.283Z",
      "content": "<p>All <code>0.5</code> submission scores <code>0.69314</code>, first row changed to <code>0.999</code> scores <code>0.69470</code>. It means that the possible range is <code>3977</code> to <code>4002</code>. And then submission with the following row passed successfully</p>\n\n<p><code>assert (len(df) == 400) or (len(df) == 4000)</code></p>\n\n<p>Additionally, there are exactly 2000 samples of each class.</p>\n\n<p>Interesting, what is the private test size?</p>",
      "rawMarkdown": "All `0.5` submission scores `0.69314`, first row changed to `0.999` scores `0.69470`. It means that the possible range is `3977` to `4002`. And then submission with the following row passed successfully\n\n```assert (len(df) == 400) or (len(df) == 4000)```\n\nAdditionally, there are exactly 2000 samples of each class.\n\nInteresting, what is the private test size?",
      "votes": 30
    },
    {
      "id": 701037,
      "postDate": "2019-12-23T02:35:33.007Z",
      "content": "<p>Looks like you are right. <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started</a>\n&gt; You are limited to 9 hours of GPU compute time. This constraint is also imposed on the Public Test Set re-run. You should anticipate the Public Test set to be 10 times the size of the Public Validation Set and budget accordingly. </p>",
      "rawMarkdown": "Looks like you are right. https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\n&gt; You are limited to 9 hours of GPU compute time. This constraint is also imposed on the Public Test Set re-run. You should anticipate the Public Test set to be 10 times the size of the Public Validation Set and budget accordingly. ",
      "votes": 3
    },
    {
      "id": 705418,
      "postDate": "2019-12-28T22:58:04.997Z",
      "content": "<p>The full dataset has 119146 files. All are in mp4 format. Not sure if the private test sets will also have all files of same format and files size will be in the same range as  in training and public test set.</p>",
      "rawMarkdown": "The full dataset has 119146 files. All are in mp4 format. Not sure if the private test sets will also have all files of same format and files size will be in the same range as  in training and public test set."
    },
    {
      "id": 702912,
      "postDate": "2019-12-25T10:41:29.650Z",
      "content": "<p>You're a genius👍  <a href=\"/zaharch\">@zaharch</a> \nHow do you get 'there are exactly 2000 samples of each class'?</p>",
      "rawMarkdown": "You're a genius👍  @zaharch \nHow do you get 'there are exactly 2000 samples of each class'?",
      "replies": [
        {
          "id": 702964,
          "postDate": "2019-12-25T12:12:35.050Z",
          "content": "<p>Initially I saw somewhere that <code>0.49</code> and <code>0.51</code> gave the same score, but it can also be calculated from all ones or all zeros scores, or immediately concluded from the fact that all ones and all zeros give the same score.</p>\n\n<p>Not related to the question, how come you have forked and re-published a variation of my kernel, but have never took time to upvote the original? :)</p>",
          "rawMarkdown": "Initially I saw somewhere that `0.49` and `0.51` gave the same score, but it can also be calculated from all ones or all zeros scores, or immediately concluded from the fact that all ones and all zeros give the same score.\n\nNot related to the question, how come you have forked and re-published a variation of my kernel, but have never took time to upvote the original? :)",
          "votes": 3
        },
        {
          "id": 702976,
          "postDate": "2019-12-25T12:32:27.820Z",
          "content": "<p>sorry, I've upvoted it.\nYeah, you're right, count(fake) == count(real). <a href=\"/zaharch\">@zaharch</a> </p>",
          "rawMarkdown": "sorry, I've upvoted it.\nYeah, you're right, count(fake) == count(real). @zaharch ",
          "votes": 1
        },
        {
          "id": 702982,
          "postDate": "2019-12-25T12:44:22.897Z",
          "content": "<p>Thanks, yes I agree with your formula, in case REAL and FAKE labels counts are equal. But if not it is in more general form <code>-x ln(p) - (1-x) ln(1-p)</code> where x is the ratio. If such scores are equal for <code>p</code> and <code>1-p</code>, it means either <code>p=0.5</code> or <code>x=0.5</code>.</p>",
          "rawMarkdown": "Thanks, yes I agree with your formula, in case REAL and FAKE labels counts are equal. But if not it is in more general form `-x ln(p) - (1-x) ln(1-p)` where x is the ratio. If such scores are equal for `p` and `1-p`, it means either `p=0.5` or `x=0.5`.",
          "votes": 1
        },
        {
          "id": 702992,
          "postDate": "2019-12-25T12:58:56.857Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 703000,
          "postDate": "2019-12-25T13:06:23.160Z",
          "content": "<p>For single sample, yes, it is a label, but the score is the average over <code>4000</code> samples, that is why I say ratio. Number of samples in one class divided by <code>4000</code>.</p>",
          "rawMarkdown": "For single sample, yes, it is a label, but the score is the average over `4000` samples, that is why I say ratio. Number of samples in one class divided by `4000`."
        }
      ]
    },
    {
      "id": 700210,
      "postDate": "2019-12-21T16:19:48.547Z",
      "content": "<p>Good question. As far as I can tell, there are 119948 total video clips in total. The private set can't be too huge or it would never finish within the time limit.</p>\n\n<p>I am equally interested in their description of the private set:</p>\n\n<blockquote>\n  <p>Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard. It contains videos with a <em>similar</em> format and nature as the Training and Public Validation/Test Sets, but are <strong>real, organic</strong> videos with and without deepfakes.</p>\n</blockquote>\n\n<p>Does this mean the methods used for tampering with the videos in private test might not all be the same as those we see in train and public? These private test videos are not among the 120k videos we are given?</p>",
      "rawMarkdown": "Good question. As far as I can tell, there are 119948 total video clips in total. The private set can't be too huge or it would never finish within the time limit.\n\nI am equally interested in their description of the private set:\n\n&gt; Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard. It contains videos with a *similar* format and nature as the Training and Public Validation/Test Sets, but are **real, organic** videos with and without deepfakes.\n\nDoes this mean the methods used for tampering with the videos in private test might not all be the same as those we see in train and public? These private test videos are not among the 120k videos we are given?\n\n",
      "replies": [
        {
          "id": 700222,
          "postDate": "2019-12-21T16:38:45.697Z",
          "content": "<p>I fear that it means exactly what your thinking - the final scoring may be on videos nothing like the current set.  Which I think means we need to use some videos from the wild to train our models?</p>",
          "rawMarkdown": "I fear that it means exactly what your thinking - the final scoring may be on videos nothing like the current set.  Which I think means we need to use some videos from the wild to train our models?",
          "votes": 1
        },
        {
          "id": 700939,
          "postDate": "2019-12-22T21:02:16.600Z",
          "content": "<p>Hi, I'm just curious how did you come to the number of 119948 total? (afaik 119154 train and 400 public test videos)</p>",
          "rawMarkdown": "Hi, I'm just curious how did you come to the number of 119948 total? (afaik 119154 train and 400 public test videos)"
        },
        {
          "id": 700970,
          "postDate": "2019-12-22T22:24:17.777Z",
          "content": "<p>That is the length of the list that results from doing os.walk in the folder I unzipped everything after removing everything not ending in .mp4.</p>\n\n<p>It is certainly possible I have some duplicates in there for whatever reason, I have not looked that carefully yet. </p>",
          "rawMarkdown": "That is the length of the list that results from doing os.walk in the folder I unzipped everything after removing everything not ending in .mp4.\n\nIt is certainly possible I have some duplicates in there for whatever reason, I have not looked that carefully yet. "
        },
        {
          "id": 700996,
          "postDate": "2019-12-23T00:21:19.433Z",
          "content": "<p>119146 is my count of the full data set - I get your number if I include the train and test samples - but they would be duplicates.</p>",
          "rawMarkdown": "119146 is my count of the full data set - I get your number if I include the train and test samples - but they would be duplicates.",
          "votes": 2
        },
        {
          "id": 701029,
          "postDate": "2019-12-23T02:22:39.517Z",
          "content": "<p>metadata says there are 8 more <a href=\"/pcjimmmy\">@pcjimmmy</a> <a href=\"https://www.kaggle.com/aleksandradeis/deepfake-challenge-eda#700928\">https://www.kaggle.com/aleksandradeis/deepfake-challenge-eda#700928</a> </p>",
          "rawMarkdown": "metadata says there are 8 more @pcjimmmy https://www.kaggle.com/aleksandradeis/deepfake-challenge-eda#700928 "
        },
        {
          "id": 701069,
          "postDate": "2019-12-23T03:35:38.813Z",
          "content": "<p>Excellent, thanks for checking up on that!</p>",
          "rawMarkdown": "Excellent, thanks for checking up on that!"
        }
      ]
    },
    {
      "id": 700969,
      "postDate": "2019-12-22T22:23:25.380Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 701037,
      "author_name": "hirviö",
      "author_url": "",
      "post_date": "2019-12-23T02:35:33.007000",
      "content": "<p>Looks like you are right. <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started</a>\n&gt; You are limited to 9 hours of GPU compute time. This constraint is also imposed on the Public Test Set re-run. You should anticipate the Public Test set to be 10 times the size of the Public Validation Set and budget accordingly. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 705418,
      "author_name": "Al Warrior",
      "author_url": "",
      "post_date": "2019-12-28T22:58:04.997000",
      "content": "<p>The full dataset has 119146 files. All are in mp4 format. Not sure if the private test sets will also have all files of same format and files size will be in the same range as  in training and public test set.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 702912,
      "author_name": "DiegoJohnson",
      "author_url": "",
      "post_date": "2019-12-25T10:41:29.650000",
      "content": "<p>You're a genius👍  <a href=\"/zaharch\">@zaharch</a> \nHow do you get 'there are exactly 2000 samples of each class'?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 702964,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-12-25T12:12:35.050000",
          "content": "<p>Initially I saw somewhere that <code>0.49</code> and <code>0.51</code> gave the same score, but it can also be calculated from all ones or all zeros scores, or immediately concluded from the fact that all ones and all zeros give the same score.</p>\n\n<p>Not related to the question, how come you have forked and re-published a variation of my kernel, but have never took time to upvote the original? :)</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 702976,
          "author_name": "DiegoJohnson",
          "author_url": "",
          "post_date": "2019-12-25T12:32:27.820000",
          "content": "<p>sorry, I've upvoted it.\nYeah, you're right, count(fake) == count(real). <a href=\"/zaharch\">@zaharch</a> </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 702982,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-12-25T12:44:22.897000",
          "content": "<p>Thanks, yes I agree with your formula, in case REAL and FAKE labels counts are equal. But if not it is in more general form <code>-x ln(p) - (1-x) ln(1-p)</code> where x is the ratio. If such scores are equal for <code>p</code> and <code>1-p</code>, it means either <code>p=0.5</code> or <code>x=0.5</code>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 702992,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-25T12:58:56.857000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 703000,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2019-12-25T13:06:23.160000",
          "content": "<p>For single sample, yes, it is a label, but the score is the average over <code>4000</code> samples, that is why I say ratio. Number of samples in one class divided by <code>4000</code>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 700210,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-12-21T16:19:48.547000",
      "content": "<p>Good question. As far as I can tell, there are 119948 total video clips in total. The private set can't be too huge or it would never finish within the time limit.</p>\n\n<p>I am equally interested in their description of the private set:</p>\n\n<blockquote>\n  <p>Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard. It contains videos with a <em>similar</em> format and nature as the Training and Public Validation/Test Sets, but are <strong>real, organic</strong> videos with and without deepfakes.</p>\n</blockquote>\n\n<p>Does this mean the methods used for tampering with the videos in private test might not all be the same as those we see in train and public? These private test videos are not among the 120k videos we are given?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 700222,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2019-12-21T16:38:45.697000",
          "content": "<p>I fear that it means exactly what your thinking - the final scoring may be on videos nothing like the current set.  Which I think means we need to use some videos from the wild to train our models?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 700939,
          "author_name": "hirviö",
          "author_url": "",
          "post_date": "2019-12-22T21:02:16.600000",
          "content": "<p>Hi, I'm just curious how did you come to the number of 119948 total? (afaik 119154 train and 400 public test videos)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 700970,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-12-22T22:24:17.777000",
          "content": "<p>That is the length of the list that results from doing os.walk in the folder I unzipped everything after removing everything not ending in .mp4.</p>\n\n<p>It is certainly possible I have some duplicates in there for whatever reason, I have not looked that carefully yet. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 700996,
          "author_name": "PC Jimmmy",
          "author_url": "",
          "post_date": "2019-12-23T00:21:19.433000",
          "content": "<p>119146 is my count of the full data set - I get your number if I include the train and test samples - but they would be duplicates.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 701029,
          "author_name": "hirviö",
          "author_url": "",
          "post_date": "2019-12-23T02:22:39.517000",
          "content": "<p>metadata says there are 8 more <a href=\"/pcjimmmy\">@pcjimmmy</a> <a href=\"https://www.kaggle.com/aleksandradeis/deepfake-challenge-eda#700928\">https://www.kaggle.com/aleksandradeis/deepfake-challenge-eda#700928</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 701069,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "2019-12-23T03:35:38.813000",
          "content": "<p>Excellent, thanks for checking up on that!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 700969,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-22T22:23:25.380000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "700199": "All `0.5` submission scores `0.69314`, first row changed to `0.999` scores `0.69470`. It means that the possible range is `3977` to `4002`. And then submission with the following row passed successfully\n\n```assert (len(df) == 400) or (len(df) == 4000)```\n\nAdditionally, there are exactly 2000 samples of each class.\n\nInteresting, what is the private test size?",
    "701037": "Looks like you are right. https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\n&gt; You are limited to 9 hours of GPU compute time. This constraint is also imposed on the Public Test Set re-run. You should anticipate the Public Test set to be 10 times the size of the Public Validation Set and budget accordingly. ",
    "705418": "The full dataset has 119146 files. All are in mp4 format. Not sure if the private test sets will also have all files of same format and files size will be in the same range as  in training and public test set.",
    "702912": "You're a genius👍  @zaharch \nHow do you get 'there are exactly 2000 samples of each class'?",
    "700210": "Good question. As far as I can tell, there are 119948 total video clips in total. The private set can't be too huge or it would never finish within the time limit.\n\nI am equally interested in their description of the private set:\n\n&gt; Private Test Set: This dataset is privately held outside of Kaggle’s platform, and is used to compute the private leaderboard. It contains videos with a *similar* format and nature as the Training and Public Validation/Test Sets, but are **real, organic** videos with and without deepfakes.\n\nDoes this mean the methods used for tampering with the videos in private test might not all be the same as those we see in train and public? These private test videos are not among the 120k videos we are given?\n\n",
    "700969": ""
  }
}