{
  "id": 134420,
  "title": "Other useful datasets finished",
  "url": "/competitions/deepfake-detection-challenge/discussion/134420",
  "author_name": "",
  "post_date": "2020-03-08T02:16:30.346305200Z",
  "votes": 17,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I wanted to use <a href=\"/phunghieu\">@phunghieu</a> 's datasets for this comp but he wasn't able to finish it and it is logistically to large (<strong>11.8+</strong> million images, <strong>800+</strong> GB for just the first <strong>15</strong> parts of the data set). So I went ahead and finished the data set:</p>\n\n<h3>First 15 parts: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954\">Link</a></h3>\n\n<h3>I wasn't consistent in how I made the datasets so I'll label each part with what it contains</h3>\n\n<p>All datasets contain jpg images to conserve space</p>\n\n<h3>10 frames all 150x150 images, mobilenet used for detection</h3>\n\n<p>Part 16: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-16\">https://www.kaggle.com/greatgamedota/dfdc-part-16</a> (121 MB)</p>\n\n<h3>150 frames all 150x150 images, facenet pytorch used for detection</h3>\n\n<p>Part 17: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-17\">https://www.kaggle.com/greatgamedota/dfdc-part-17</a> (1GB)</p>\n\n<h3>10 frames, 150x150 images, mobilenet detector</h3>\n\n<p>Part 18: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-18\">https://www.kaggle.com/greatgamedota/dfdc-part-18</a> (158 MB)\n<strong>...</strong>\nPart 32: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-32\">https://www.kaggle.com/greatgamedota/dfdc-part-32</a> (150 MB)</p>\n\n<h3>10 frames, 160x160, <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code\">mobilenet detector</a></h3>\n\n<p>Part 33: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-33\">https://www.kaggle.com/greatgamedota/dfdc-part-33</a> (139 MB)</p>\n\n<h3>10 frames, 150x150, mobilenet detector</h3>\n\n<p>Part 34: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-34\">https://www.kaggle.com/greatgamedota/dfdc-part-34</a> (147 MB)\n<strong>...</strong>\nPart 49: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-49\">https://www.kaggle.com/greatgamedota/dfdc-part-49</a> (189 MB)</p>\n\n<h3>Here are some code snippets to load all of these datasets into a kaggle kernal (after you spend 10 minutes just adding them lol)</h3>\n\n<p>```</p>\n\n<h1>First 15 parts:</h1>\n\n<p>meta = glob.glob('../input/deepfake-detection-faces-<em>/</em>.csv')\nmeta.sort(key=lambda f: int(re.sub('\\D', '', f)))</p>\n\n<p>dfs = []\nfor path in meta:\n    df = pd.read_csv(path)\n    df['path'] = ''\n    path = path.split(\"/\")[:-1]\n    path = path[0] + '/' + path[1] + '/' + path[2] + '/'\n    for i in range(len(df)):\n        df.loc[i]['path'] = f'{path}{df.loc[i][\"filename\"][:-4]}'\n    dfs.append(df)</p>\n\n<p>train_df = pd.concat(dfs)\ntrain_df = train_df.reset_index(drop=True)\nprint(len(train_df))\n```</p>\n\n<h3>My datasets require more cleaning:</h3>\n\n<p>```</p>\n\n<h1>@GreatGameDota's datasets</h1>\n\n<p>part = 16\nfor j in range((49-16)+1):\n    if part+j != 17:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.csv')\n    else:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.json', index_col=0)\n    meta['path'] = ''\n    print(part+j)\n    del_idxs = []\n    for i in range(len(meta)):\n        if os.path.isdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'):\n            if len(os.listdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}')) &lt; 5:\n                del_idxs.append(i)\n            else:\n                meta.loc[i]['path'] = f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'\n        else:\n            del_idxs.append(i)\n    print(del_idxs) # Print out removed video ids\n    for idx in del_idxs:\n        meta = meta.drop(idx)\n    train_df = pd.concat([train_df,meta])\n    train_df = train_df.reset_index(drop=True)\nprint(len(train_df)) # Should be ~120k if everything is successfully loaded\n```\nThe above snippet adds data that has at least 5 images for that video. (Fun fact: mobilenet &gt;&gt; facenet just based on the amount of videos facenet failed on and need to be removed)</p>\n\n<h3>Hope this helps anyone using/wants to use kaggle for this comp :)</h3>",
  "messages": [
    {
      "id": "766323",
      "postDate": "03/08/2020 02:16:30",
      "content": "<p>I wanted to use <a href=\"/phunghieu\">@phunghieu</a> 's datasets for this comp but he wasn't able to finish it and it is logistically to large (<strong>11.8+</strong> million images, <strong>800+</strong> GB for just the first <strong>15</strong> parts of the data set). So I went ahead and finished the data set:</p>\n\n<h3>First 15 parts: <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954\">Link</a></h3>\n\n<h3>I wasn't consistent in how I made the datasets so I'll label each part with what it contains</h3>\n\n<p>All datasets contain jpg images to conserve space</p>\n\n<h3>10 frames all 150x150 images, mobilenet used for detection</h3>\n\n<p>Part 16: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-16\">https://www.kaggle.com/greatgamedota/dfdc-part-16</a> (121 MB)</p>\n\n<h3>150 frames all 150x150 images, facenet pytorch used for detection</h3>\n\n<p>Part 17: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-17\">https://www.kaggle.com/greatgamedota/dfdc-part-17</a> (1GB)</p>\n\n<h3>10 frames, 150x150 images, mobilenet detector</h3>\n\n<p>Part 18: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-18\">https://www.kaggle.com/greatgamedota/dfdc-part-18</a> (158 MB)\n<strong>...</strong>\nPart 32: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-32\">https://www.kaggle.com/greatgamedota/dfdc-part-32</a> (150 MB)</p>\n\n<h3>10 frames, 160x160, <a href=\"https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code\">mobilenet detector</a></h3>\n\n<p>Part 33: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-33\">https://www.kaggle.com/greatgamedota/dfdc-part-33</a> (139 MB)</p>\n\n<h3>10 frames, 150x150, mobilenet detector</h3>\n\n<p>Part 34: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-34\">https://www.kaggle.com/greatgamedota/dfdc-part-34</a> (147 MB)\n<strong>...</strong>\nPart 49: <a href=\"https://www.kaggle.com/greatgamedota/dfdc-part-49\">https://www.kaggle.com/greatgamedota/dfdc-part-49</a> (189 MB)</p>\n\n<h3>Here are some code snippets to load all of these datasets into a kaggle kernal (after you spend 10 minutes just adding them lol)</h3>\n\n<p>```</p>\n\n<h1>First 15 parts:</h1>\n\n<p>meta = glob.glob('../input/deepfake-detection-faces-<em>/</em>.csv')\nmeta.sort(key=lambda f: int(re.sub('\\D', '', f)))</p>\n\n<p>dfs = []\nfor path in meta:\n    df = pd.read_csv(path)\n    df['path'] = ''\n    path = path.split(\"/\")[:-1]\n    path = path[0] + '/' + path[1] + '/' + path[2] + '/'\n    for i in range(len(df)):\n        df.loc[i]['path'] = f'{path}{df.loc[i][\"filename\"][:-4]}'\n    dfs.append(df)</p>\n\n<p>train_df = pd.concat(dfs)\ntrain_df = train_df.reset_index(drop=True)\nprint(len(train_df))\n```</p>\n\n<h3>My datasets require more cleaning:</h3>\n\n<p>```</p>\n\n<h1>@GreatGameDota's datasets</h1>\n\n<p>part = 16\nfor j in range((49-16)+1):\n    if part+j != 17:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.csv')\n    else:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.json', index_col=0)\n    meta['path'] = ''\n    print(part+j)\n    del_idxs = []\n    for i in range(len(meta)):\n        if os.path.isdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'):\n            if len(os.listdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}')) &lt; 5:\n                del_idxs.append(i)\n            else:\n                meta.loc[i]['path'] = f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'\n        else:\n            del_idxs.append(i)\n    print(del_idxs) # Print out removed video ids\n    for idx in del_idxs:\n        meta = meta.drop(idx)\n    train_df = pd.concat([train_df,meta])\n    train_df = train_df.reset_index(drop=True)\nprint(len(train_df)) # Should be ~120k if everything is successfully loaded\n```\nThe above snippet adds data that has at least 5 images for that video. (Fun fact: mobilenet &gt;&gt; facenet just based on the amount of videos facenet failed on and need to be removed)</p>\n\n<h3>Hope this helps anyone using/wants to use kaggle for this comp :)</h3>",
      "rawMarkdown": "I wanted to use @phunghieu 's datasets for this comp but he wasn't able to finish it and it is logistically to large (**11.8+** million images, **800+** GB for just the first **15** parts of the data set). So I went ahead and finished the data set:\n\n### First 15 parts: [Link](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954)\n\n### I wasn't consistent in how I made the datasets so I'll label each part with what it contains\n\nAll datasets contain jpg images to conserve space\n\n### 10 frames all 150x150 images, mobilenet used for detection\nPart 16: https://www.kaggle.com/greatgamedota/dfdc-part-16 (121 MB)\n\n### 150 frames all 150x150 images, facenet pytorch used for detection\nPart 17: https://www.kaggle.com/greatgamedota/dfdc-part-17 (1GB)\n\n### 10 frames, 150x150 images, mobilenet detector\nPart 18: https://www.kaggle.com/greatgamedota/dfdc-part-18 (158 MB)\n**...**\nPart 32: https://www.kaggle.com/greatgamedota/dfdc-part-32 (150 MB)\n\n### 10 frames, 160x160, [mobilenet detector](https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code)\nPart 33: https://www.kaggle.com/greatgamedota/dfdc-part-33 (139 MB)\n\n### 10 frames, 150x150, mobilenet detector\nPart 34: https://www.kaggle.com/greatgamedota/dfdc-part-34 (147 MB)\n**...**\nPart 49: https://www.kaggle.com/greatgamedota/dfdc-part-49 (189 MB)\n\n### Here are some code snippets to load all of these datasets into a kaggle kernal (after you spend 10 minutes just adding them lol)\n\n```\n# First 15 parts:\nmeta = glob.glob('../input/deepfake-detection-faces-*/*.csv')\nmeta.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\ndfs = []\nfor path in meta:\n    df = pd.read_csv(path)\n    df['path'] = ''\n    path = path.split(\"/\")[:-1]\n    path = path[0] + '/' + path[1] + '/' + path[2] + '/'\n    for i in range(len(df)):\n        df.loc[i]['path'] = f'{path}{df.loc[i][\"filename\"][:-4]}'\n    dfs.append(df)\n\ntrain_df = pd.concat(dfs)\ntrain_df = train_df.reset_index(drop=True)\nprint(len(train_df))\n```\n### My datasets require more cleaning:\n```\n# @GreatGameDota's datasets\npart = 16\nfor j in range((49-16)+1):\n    if part+j != 17:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.csv')\n    else:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.json', index_col=0)\n    meta['path'] = ''\n    print(part+j)\n    del_idxs = []\n    for i in range(len(meta)):\n        if os.path.isdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'):\n            if len(os.listdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}')) &lt; 5:\n                del_idxs.append(i)\n            else:\n                meta.loc[i]['path'] = f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'\n        else:\n            del_idxs.append(i)\n    print(del_idxs) # Print out removed video ids\n    for idx in del_idxs:\n        meta = meta.drop(idx)\n    train_df = pd.concat([train_df,meta])\n    train_df = train_df.reset_index(drop=True)\nprint(len(train_df)) # Should be ~120k if everything is successfully loaded\n```\nThe above snippet adds data that has at least 5 images for that video. (Fun fact: mobilenet &gt;&gt; facenet just based on the amount of videos facenet failed on and need to be removed)\n\n### Hope this helps anyone using/wants to use kaggle for this comp :)",
      "votes": null
    },
    {
      "id": "766341",
      "postDate": "03/08/2020 03:08:32",
      "content": "<p>Great job, <a href=\"/greatgamedota\">@greatgamedota</a> 👍. I hope that other competitors can benefit from these datasets to continue to push their models' performance in this competition.</p>",
      "rawMarkdown": "Great job, @greatgamedota 👍. I hope that other competitors can benefit from these datasets to continue to push their models' performance in this competition.",
      "votes": null
    },
    {
      "id": "766456",
      "postDate": "03/08/2020 06:52:22",
      "content": "<p>This is amazing work <a href=\"/greatgamedota\">@greatgamedota</a>. Thanks for putting all the hard work to create and upload this dataset for people who are using kaggle only for this comp. You and <a href=\"/phunghieu\">@phunghieu</a> are the chief reason, Kaggle community rocks. Appreciated.  </p>\n\n<p>+10 if you can upload a kernal with all the data preloaded :).  I know, its a lot to ask, However it will great for people to quickly try various training approaches on this dataset. </p>",
      "rawMarkdown": "This is amazing work @greatgamedota. Thanks for putting all the hard work to create and upload this dataset for people who are using kaggle only for this comp. You and @phunghieu are the chief reason, Kaggle community rocks. Appreciated.  \n\n+10 if you can upload a kernal with all the data preloaded :).  I know, its a lot to ask, However it will great for people to quickly try various training approaches on this dataset.",
      "votes": null
    },
    {
      "id": "766651",
      "postDate": "03/08/2020 14:15:29",
      "content": "<p>Here is a kaggle kernal with all the datasets preloaded: <a href=\"https://www.kaggle.com/greatgamedota/deepfake-detection-full-data\">https://www.kaggle.com/greatgamedota/deepfake-detection-full-data</a> (WARNING: Kaggle attempts to load a preview of ALL the datasets which causes Kaggle to tank)</p>",
      "rawMarkdown": "Here is a kaggle kernal with all the datasets preloaded: https://www.kaggle.com/greatgamedota/deepfake-detection-full-data (WARNING: Kaggle attempts to load a preview of ALL the datasets which causes Kaggle to tank)",
      "votes": null
    },
    {
      "id": "766652",
      "postDate": "03/08/2020 14:15:46",
      "content": "<p>Thank you! I have shared a kernal.</p>",
      "rawMarkdown": "Thank you! I have shared a kernal.",
      "votes": null
    },
    {
      "id": "780959",
      "postDate": "03/20/2020 19:22:53",
      "content": "<p>UPDATE: I updated datasets 16, and 18-32 with images using mobilenet as the detector instead of facenet_pytorch due to its increased accuracy.</p>",
      "rawMarkdown": "UPDATE: I updated datasets 16, and 18-32 with images using mobilenet as the detector instead of facenet_pytorch due to its increased accuracy.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 766341,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "03/08/2020 03:08:32",
      "content": "<p>Great job, <a href=\"/greatgamedota\">@greatgamedota</a> 👍. I hope that other competitors can benefit from these datasets to continue to push their models' performance in this competition.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 766456,
      "author_name": "pankymathur",
      "author_url": "",
      "post_date": "03/08/2020 06:52:22",
      "content": "<p>This is amazing work <a href=\"/greatgamedota\">@greatgamedota</a>. Thanks for putting all the hard work to create and upload this dataset for people who are using kaggle only for this comp. You and <a href=\"/phunghieu\">@phunghieu</a> are the chief reason, Kaggle community rocks. Appreciated.  </p>\n\n<p>+10 if you can upload a kernal with all the data preloaded :).  I know, its a lot to ask, However it will great for people to quickly try various training approaches on this dataset. </p>",
      "votes": null,
      "replies": [
        {
          "id": 766652,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "03/08/2020 14:15:46",
          "content": "<p>Thank you! I have shared a kernal.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 766651,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "03/08/2020 14:15:29",
      "content": "<p>Here is a kaggle kernal with all the datasets preloaded: <a href=\"https://www.kaggle.com/greatgamedota/deepfake-detection-full-data\">https://www.kaggle.com/greatgamedota/deepfake-detection-full-data</a> (WARNING: Kaggle attempts to load a preview of ALL the datasets which causes Kaggle to tank)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 780959,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "03/20/2020 19:22:53",
      "content": "<p>UPDATE: I updated datasets 16, and 18-32 with images using mobilenet as the detector instead of facenet_pytorch due to its increased accuracy.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "766323": "I wanted to use @phunghieu 's datasets for this comp but he wasn't able to finish it and it is logistically to large (**11.8+** million images, **800+** GB for just the first **15** parts of the data set). So I went ahead and finished the data set:\n\n### First 15 parts: [Link](https://www.kaggle.com/c/deepfake-detection-challenge/discussion/128954)\n\n### I wasn't consistent in how I made the datasets so I'll label each part with what it contains\n\nAll datasets contain jpg images to conserve space\n\n### 10 frames all 150x150 images, mobilenet used for detection\nPart 16: https://www.kaggle.com/greatgamedota/dfdc-part-16 (121 MB)\n\n### 150 frames all 150x150 images, facenet pytorch used for detection\nPart 17: https://www.kaggle.com/greatgamedota/dfdc-part-17 (1GB)\n\n### 10 frames, 150x150 images, mobilenet detector\nPart 18: https://www.kaggle.com/greatgamedota/dfdc-part-18 (158 MB)\n**...**\nPart 32: https://www.kaggle.com/greatgamedota/dfdc-part-32 (150 MB)\n\n### 10 frames, 160x160, [mobilenet detector](https://www.kaggle.com/unkownhihi/mobilenet-face-extractor-helper-code)\nPart 33: https://www.kaggle.com/greatgamedota/dfdc-part-33 (139 MB)\n\n### 10 frames, 150x150, mobilenet detector\nPart 34: https://www.kaggle.com/greatgamedota/dfdc-part-34 (147 MB)\n**...**\nPart 49: https://www.kaggle.com/greatgamedota/dfdc-part-49 (189 MB)\n\n### Here are some code snippets to load all of these datasets into a kaggle kernal (after you spend 10 minutes just adding them lol)\n\n```\n# First 15 parts:\nmeta = glob.glob('../input/deepfake-detection-faces-*/*.csv')\nmeta.sort(key=lambda f: int(re.sub('\\D', '', f)))\n\ndfs = []\nfor path in meta:\n    df = pd.read_csv(path)\n    df['path'] = ''\n    path = path.split(\"/\")[:-1]\n    path = path[0] + '/' + path[1] + '/' + path[2] + '/'\n    for i in range(len(df)):\n        df.loc[i]['path'] = f'{path}{df.loc[i][\"filename\"][:-4]}'\n    dfs.append(df)\n\ntrain_df = pd.concat(dfs)\ntrain_df = train_df.reset_index(drop=True)\nprint(len(train_df))\n```\n### My datasets require more cleaning:\n```\n# @GreatGameDota's datasets\npart = 16\nfor j in range((49-16)+1):\n    if part+j != 17:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.csv')\n    else:\n        meta = pd.read_csv(f'../input/dfdc-part-{part+j}/images/metadata{part+j}.json', index_col=0)\n    meta['path'] = ''\n    print(part+j)\n    del_idxs = []\n    for i in range(len(meta)):\n        if os.path.isdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'):\n            if len(os.listdir(f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}')) &lt; 5:\n                del_idxs.append(i)\n            else:\n                meta.loc[i]['path'] = f'../input/dfdc-part-{part+j}/images/{meta.loc[i][\"filename\"][:-4]}'\n        else:\n            del_idxs.append(i)\n    print(del_idxs) # Print out removed video ids\n    for idx in del_idxs:\n        meta = meta.drop(idx)\n    train_df = pd.concat([train_df,meta])\n    train_df = train_df.reset_index(drop=True)\nprint(len(train_df)) # Should be ~120k if everything is successfully loaded\n```\nThe above snippet adds data that has at least 5 images for that video. (Fun fact: mobilenet &gt;&gt; facenet just based on the amount of videos facenet failed on and need to be removed)\n\n### Hope this helps anyone using/wants to use kaggle for this comp :)",
    "766341": "Great job, @greatgamedota 👍. I hope that other competitors can benefit from these datasets to continue to push their models' performance in this competition.",
    "766456": "This is amazing work @greatgamedota. Thanks for putting all the hard work to create and upload this dataset for people who are using kaggle only for this comp. You and @phunghieu are the chief reason, Kaggle community rocks. Appreciated.  \n\n+10 if you can upload a kernal with all the data preloaded :).  I know, its a lot to ask, However it will great for people to quickly try various training approaches on this dataset.",
    "766651": "Here is a kaggle kernal with all the datasets preloaded: https://www.kaggle.com/greatgamedota/deepfake-detection-full-data (WARNING: Kaggle attempts to load a preview of ALL the datasets which causes Kaggle to tank)",
    "766652": "Thank you! I have shared a kernal.",
    "780959": "UPDATE: I updated datasets 16, and 18-32 with images using mobilenet as the detector instead of facenet_pytorch due to its increased accuracy."
  },
  "source": "meta"
}