{
  "id": 121458,
  "title": "Small dataset",
  "url": "/competitions/deepfake-detection-challenge/discussion/121458",
  "author_name": "",
  "post_date": "2019-12-13T12:40:36.060668100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>If I'm reading this correctly, the full dataset has only about 5000 videos, based on 66 individuals. Is that correct ? </p>\n\n<p>That seems pretty low to me - low from ML perspective, I guess the cost to produce it is definitely not low.</p>",
  "messages": [
    {
      "id": "694328",
      "postDate": "12/13/2019 12:40:36",
      "content": "<p>If I'm reading this correctly, the full dataset has only about 5000 videos, based on 66 individuals. Is that correct ? </p>\n\n<p>That seems pretty low to me - low from ML perspective, I guess the cost to produce it is definitely not low.</p>",
      "rawMarkdown": "If I'm reading this correctly, the full dataset has only about 5000 videos, based on 66 individuals. Is that correct ? \n\nThat seems pretty low to me - low from ML perspective, I guess the cost to produce it is definitely not low.",
      "votes": null
    },
    {
      "id": "694425",
      "postDate": "12/13/2019 15:15:28",
      "content": "<p>From what I have extracted so far, it is more like 125K videos (50*2.5K approx). Maybe 5K is the number of real videos?</p>",
      "rawMarkdown": "From what I have extracted so far, it is more like 125K videos (50*2.5K approx). Maybe 5K is the number of real videos?",
      "votes": null
    },
    {
      "id": "696631",
      "postDate": "12/16/2019 21:20:20",
      "content": "<p>According to the paper, it's the total number of videos. Were you able to get all the dataset already ? I suspect, there are a lot of redundant videos (cut from a single longer source video). </p>",
      "rawMarkdown": "According to the paper, it's the total number of videos. Were you able to get all the dataset already ? I suspect, there are a lot of redundant videos (cut from a single longer source video).",
      "votes": null
    },
    {
      "id": "696984",
      "postDate": "12/17/2019 10:20:46",
      "content": "<p>The paper is about the <em>preview</em> dataset that was released in October. There's no guarantee that anything in that paper applies to the full dataset here on Kaggle.</p>",
      "rawMarkdown": "The paper is about the *preview* dataset that was released in October. There's no guarantee that anything in that paper applies to the full dataset here on Kaggle.",
      "votes": null
    },
    {
      "id": "697030",
      "postDate": "12/17/2019 11:47:50",
      "content": "<p>Ooooh ! Thanks for clarifying that. Makes quite a big difference</p>",
      "rawMarkdown": "Ooooh ! Thanks for clarifying that. Makes quite a big difference",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 694425,
      "author_name": "batangas",
      "author_url": "",
      "post_date": "12/13/2019 15:15:28",
      "content": "<p>From what I have extracted so far, it is more like 125K videos (50*2.5K approx). Maybe 5K is the number of real videos?</p>",
      "votes": null,
      "replies": [
        {
          "id": 696631,
          "author_name": "alchemist",
          "author_url": "",
          "post_date": "12/16/2019 21:20:20",
          "content": "<p>According to the paper, it's the total number of videos. Were you able to get all the dataset already ? I suspect, there are a lot of redundant videos (cut from a single longer source video). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 696984,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/17/2019 10:20:46",
          "content": "<p>The paper is about the <em>preview</em> dataset that was released in October. There's no guarantee that anything in that paper applies to the full dataset here on Kaggle.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697030,
          "author_name": "alchemist",
          "author_url": "",
          "post_date": "12/17/2019 11:47:50",
          "content": "<p>Ooooh ! Thanks for clarifying that. Makes quite a big difference</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "694328": "If I'm reading this correctly, the full dataset has only about 5000 videos, based on 66 individuals. Is that correct ? \n\nThat seems pretty low to me - low from ML perspective, I guess the cost to produce it is definitely not low.",
    "694425": "From what I have extracted so far, it is more like 125K videos (50*2.5K approx). Maybe 5K is the number of real videos?",
    "696631": "According to the paper, it's the total number of videos. Were you able to get all the dataset already ? I suspect, there are a lot of redundant videos (cut from a single longer source video).",
    "696984": "The paper is about the *preview* dataset that was released in October. There's no guarantee that anything in that paper applies to the full dataset here on Kaggle.",
    "697030": "Ooooh ! Thanks for clarifying that. Makes quite a big difference"
  },
  "source": "meta"
}