{
  "id": 124348,
  "title": "Dataset with extracted faces from first frame",
  "url": "/competitions/deepfake-detection-challenge/discussion/124348",
  "author_name": "dagnelies",
  "post_date": "2020-01-03T15:32:34.864000",
  "votes": 11,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Here it is ;)</p>\n\n<p>Almost 100 thousand 155x155 faces from the videos first frame.</p>\n\n<p><a href=\"https://www.kaggle.com/dagnelies/deepfake-faces\">https://www.kaggle.com/dagnelies/deepfake-faces</a></p>",
  "messages": [
    {
      "id": 709499,
      "postDate": "2020-01-03T15:32:34.863Z",
      "content": "<p>Here it is ;)</p>\n\n<p>Almost 100 thousand 155x155 faces from the videos first frame.</p>\n\n<p><a href=\"https://www.kaggle.com/dagnelies/deepfake-faces\">https://www.kaggle.com/dagnelies/deepfake-faces</a></p>",
      "rawMarkdown": "Here it is ;)\n\nAlmost 100 thousand 155x155 faces from the videos first frame.\n\nhttps://www.kaggle.com/dagnelies/deepfake-faces",
      "votes": 11
    },
    {
      "id": 710536,
      "postDate": "2020-01-04T21:38:50.217Z",
      "content": "<p>Hi,</p>\n\n<p>upvoting is always welcome. ;) But of course you can use the dataset anyway. </p>\n\n<p>The whole procedure to build this dataset iss rather cumbersome and split across several shell and python scripts.</p>\n\n<p>The relevant part to extract the faces is simply:</p>\n\n<p>```python\nimport face_recognition\nfrom PIL import Image\nimport os</p>\n\n<p>def process_face(path):\n    image = face_recognition.load_image_file(path)</p>\n\n<pre><code>face_locations = face_recognition.face_locations(image)\nprint(\"faces:\", len(face_locations))\n\nif len(face_locations) != 1:\n    return # ignore uncertain cases\n\nfor face_location in face_locations:\n\n    # Print the location of each face in this image\n    top, right, bottom, left = face_location\n\n    # You can access the actual face itself like this:\n    face = image[top:bottom, left:right]\n\n\n    face_img = Image.fromarray(face)\n    face_img.save(path[:-4] + \"_face.jpg\")\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Hi,\n\nupvoting is always welcome. ;) But of course you can use the dataset anyway. \n\nThe whole procedure to build this dataset iss rather cumbersome and split across several shell and python scripts.\n\nThe relevant part to extract the faces is simply:\n\n```python\nimport face_recognition\nfrom PIL import Image\nimport os\n\ndef process_face(path):\n    image = face_recognition.load_image_file(path)\n\n    face_locations = face_recognition.face_locations(image)\n    print(\"faces:\", len(face_locations))\n    \n    if len(face_locations) != 1:\n        return # ignore uncertain cases\n\n    for face_location in face_locations:\n\n        # Print the location of each face in this image\n        top, right, bottom, left = face_location\n        \n        # You can access the actual face itself like this:\n        face = image[top:bottom, left:right]\n\n\n        face_img = Image.fromarray(face)\n        face_img.save(path[:-4] + \"_face.jpg\")\n```",
      "votes": 2,
      "replies": [
        {
          "id": 710539,
          "postDate": "2020-01-04T21:41:14.927Z",
          "content": "<p>Thanks.</p>",
          "rawMarkdown": "Thanks."
        },
        {
          "id": 728143,
          "postDate": "2020-01-24T12:49:50.817Z",
          "content": "<p><a href=\"/dagnelies\">@dagnelies</a> thanks..\n1) Does this dataset comprises of  full data set of 470 gb.. i was wondering it has curtailed to just 300 mb</p>\n\n<p>2)How can i control the no of frames i want to extract images from ?</p>",
          "rawMarkdown": "@dagnelies thanks..\n1) Does this dataset comprises of  full data set of 470 gb.. i was wondering it has curtailed to just 300 mb\n\n2)How can i control the no of frames i want to extract images from ?"
        },
        {
          "id": 728263,
          "postDate": "2020-01-24T14:48:29.923Z",
          "content": "<p>1) yes, it is on the whole training set on the first frame. Around 80% got a face recognized.\n2) there are plenty of examples to achieve that in the other notebooks</p>",
          "rawMarkdown": "1) yes, it is on the whole training set on the first frame. Around 80% got a face recognized.\n2) there are plenty of examples to achieve that in the other notebooks"
        },
        {
          "id": 728295,
          "postDate": "2020-01-24T15:16:18.110Z",
          "content": "<p><a href=\"/dagnelies\">@dagnelies</a>  thanku..\nNow next q is how we get label file for this. .json for full data set.\nSorry may be very silly q :)</p>",
          "rawMarkdown": "@dagnelies  thanku..\nNow next q is how we get label file for this. .json for full data set.\nSorry may be very silly q :)"
        },
        {
          "id": 728434,
          "postDate": "2020-01-24T18:07:13.507Z",
          "content": "<p>I think metadata.csv is included</p>",
          "rawMarkdown": "I think metadata.csv is included"
        }
      ]
    },
    {
      "id": 727644,
      "postDate": "2020-01-23T23:52:31.517Z",
      "content": "<p>Hi there\nNice work!\nHowever, could you describe what resizing algorithm/code you have used on the images?\nThis would assist in developing testing!\nThanks\nEd</p>",
      "rawMarkdown": "Hi there\nNice work!\nHowever, could you describe what resizing algorithm/code you have used on the images?\nThis would assist in developing testing!\nThanks\nEd\n",
      "replies": [
        {
          "id": 728275,
          "postDate": "2020-01-24T14:53:30.380Z",
          "content": "<p>Eeeh ...I don't remember :/</p>\n\n<p>I think it was just some PIL resizing function, like <code>resize</code> or so.</p>",
          "rawMarkdown": "Eeeh ...I don't remember :/\n\nI think it was just some PIL resizing function, like `resize` or so."
        },
        {
          "id": 728479,
          "postDate": "2020-01-24T19:22:25.550Z",
          "content": "<p>Many thanks for the reply, and your effort, saves me on spending $$$ (or free credits :) ) on GCP...</p>",
          "rawMarkdown": "Many thanks for the reply, and your effort, saves me on spending $$$ (or free credits :) ) on GCP..."
        }
      ]
    },
    {
      "id": 709808,
      "postDate": "2020-01-04T00:27:01.490Z",
      "content": "<p>Excuse me. Should I upvote your topic or your dataset if I want to use your dataset? \nBTW can you share the the code you used. It will be helpful when inference.\nThanks!</p>",
      "rawMarkdown": "Excuse me. Should I upvote your topic or your dataset if I want to use your dataset? \nBTW can you share the the code you used. It will be helpful when inference.\nThanks!"
    },
    {
      "id": 742866,
      "postDate": "2020-02-11T15:52:12.653Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 743008,
          "postDate": "2020-02-11T17:46:44.553Z",
          "content": "<p>The dataset is still there</p>\n\n<p>Since <a href=\"/humananalog\">@humananalog</a> 's kernel was based on 224x224 images, I re-extracted the faces directly to this size. Maybe that's why kernel produced an error. Simply replace the input directory \"faces_155\" by \"faces_224\" ;) ...and you can skip the resizing part of course</p>",
          "rawMarkdown": "The dataset is still there\n\nSince @humananalog 's kernel was based on 224x224 images, I re-extracted the faces directly to this size. Maybe that's why kernel produced an error. Simply replace the input directory \"faces_155\" by \"faces_224\" ;) ...and you can skip the resizing part of course"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 710536,
      "author_name": "dagnelies",
      "author_url": "",
      "post_date": "2020-01-04T21:38:50.217000",
      "content": "<p>Hi,</p>\n\n<p>upvoting is always welcome. ;) But of course you can use the dataset anyway. </p>\n\n<p>The whole procedure to build this dataset iss rather cumbersome and split across several shell and python scripts.</p>\n\n<p>The relevant part to extract the faces is simply:</p>\n\n<p>```python\nimport face_recognition\nfrom PIL import Image\nimport os</p>\n\n<p>def process_face(path):\n    image = face_recognition.load_image_file(path)</p>\n\n<pre><code>face_locations = face_recognition.face_locations(image)\nprint(\"faces:\", len(face_locations))\n\nif len(face_locations) != 1:\n    return # ignore uncertain cases\n\nfor face_location in face_locations:\n\n    # Print the location of each face in this image\n    top, right, bottom, left = face_location\n\n    # You can access the actual face itself like this:\n    face = image[top:bottom, left:right]\n\n\n    face_img = Image.fromarray(face)\n    face_img.save(path[:-4] + \"_face.jpg\")\n</code></pre>\n\n<p>```</p>",
      "votes": 2,
      "replies": [
        {
          "id": 710539,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-04T21:41:14.927000",
          "content": "<p>Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728143,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-01-24T12:49:50.817000",
          "content": "<p><a href=\"/dagnelies\">@dagnelies</a> thanks..\n1) Does this dataset comprises of  full data set of 470 gb.. i was wondering it has curtailed to just 300 mb</p>\n\n<p>2)How can i control the no of frames i want to extract images from ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728263,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "2020-01-24T14:48:29.923000",
          "content": "<p>1) yes, it is on the whole training set on the first frame. Around 80% got a face recognized.\n2) there are plenty of examples to achieve that in the other notebooks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728295,
          "author_name": "Jaideep",
          "author_url": "",
          "post_date": "2020-01-24T15:16:18.110000",
          "content": "<p><a href=\"/dagnelies\">@dagnelies</a>  thanku..\nNow next q is how we get label file for this. .json for full data set.\nSorry may be very silly q :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728434,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "2020-01-24T18:07:13.507000",
          "content": "<p>I think metadata.csv is included</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727644,
      "author_name": "Ed Austin",
      "author_url": "",
      "post_date": "2020-01-23T23:52:31.517000",
      "content": "<p>Hi there\nNice work!\nHowever, could you describe what resizing algorithm/code you have used on the images?\nThis would assist in developing testing!\nThanks\nEd</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728275,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "2020-01-24T14:53:30.380000",
          "content": "<p>Eeeh ...I don't remember :/</p>\n\n<p>I think it was just some PIL resizing function, like <code>resize</code> or so.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728479,
          "author_name": "Ed Austin",
          "author_url": "",
          "post_date": "2020-01-24T19:22:25.550000",
          "content": "<p>Many thanks for the reply, and your effort, saves me on spending $$$ (or free credits :) ) on GCP...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 709808,
      "author_name": "Shangqiu Li",
      "author_url": "",
      "post_date": "2020-01-04T00:27:01.490000",
      "content": "<p>Excuse me. Should I upvote your topic or your dataset if I want to use your dataset? \nBTW can you share the the code you used. It will be helpful when inference.\nThanks!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 742866,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-11T15:52:12.653000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 743008,
          "author_name": "dagnelies",
          "author_url": "",
          "post_date": "2020-02-11T17:46:44.553000",
          "content": "<p>The dataset is still there</p>\n\n<p>Since <a href=\"/humananalog\">@humananalog</a> 's kernel was based on 224x224 images, I re-extracted the faces directly to this size. Maybe that's why kernel produced an error. Simply replace the input directory \"faces_155\" by \"faces_224\" ;) ...and you can skip the resizing part of course</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "709499": "Here it is ;)\n\nAlmost 100 thousand 155x155 faces from the videos first frame.\n\nhttps://www.kaggle.com/dagnelies/deepfake-faces",
    "710536": "Hi,\n\nupvoting is always welcome. ;) But of course you can use the dataset anyway. \n\nThe whole procedure to build this dataset iss rather cumbersome and split across several shell and python scripts.\n\nThe relevant part to extract the faces is simply:\n\n```python\nimport face_recognition\nfrom PIL import Image\nimport os\n\ndef process_face(path):\n    image = face_recognition.load_image_file(path)\n\n    face_locations = face_recognition.face_locations(image)\n    print(\"faces:\", len(face_locations))\n    \n    if len(face_locations) != 1:\n        return # ignore uncertain cases\n\n    for face_location in face_locations:\n\n        # Print the location of each face in this image\n        top, right, bottom, left = face_location\n        \n        # You can access the actual face itself like this:\n        face = image[top:bottom, left:right]\n\n\n        face_img = Image.fromarray(face)\n        face_img.save(path[:-4] + \"_face.jpg\")\n```",
    "727644": "Hi there\nNice work!\nHowever, could you describe what resizing algorithm/code you have used on the images?\nThis would assist in developing testing!\nThanks\nEd\n",
    "709808": "Excuse me. Should I upvote your topic or your dataset if I want to use your dataset? \nBTW can you share the the code you used. It will be helpful when inference.\nThanks!",
    "742866": ""
  }
}