{
  "id": 121202,
  "title": "New to Machine Learning or Kaggle?",
  "url": "/competitions/deepfake-detection-challenge/discussion/121202",
  "author_name": "Addison Howard",
  "post_date": "2019-12-11T20:12:05.547000",
  "votes": 17,
  "comment_count": 20,
  "views": 0,
  "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\">how to enter a competition using Kaggle Notebooks</a>.</p>\n\n<p>This competition has a very unique submission method, so be sure to check out the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started</a> thread for more info!</p>",
  "messages": [
    {
      "id": 692877,
      "postDate": "2019-12-11T20:12:05.547Z",
      "content": "<p>New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! </p>\n\n<p>If you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!</p>\n\n<p>New to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about <a href=\"https://www.youtube.com/watch?v=aIus8si_Et0\">site etiquette</a>, <a href=\"https://www.youtube.com/watch?v=sEJHyuWKd-s\">Kaggle lingo</a>, and <a href=\"https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ\">how to enter a competition using Kaggle Notebooks</a>.</p>\n\n<p>This competition has a very unique submission method, so be sure to check out the <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started\">Getting Started</a> thread for more info!</p>",
      "rawMarkdown": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).\n\nThis competition has a very unique submission method, so be sure to check out the [Getting Started](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started) thread for more info!",
      "votes": 17
    },
    {
      "id": 700917,
      "postDate": "2019-12-22T20:14:16.160Z",
      "content": "<p>Hi,\nDo we need to run our code finally on sample training data of 400 files while submitting the code to build the model? Is the 470 GB that was provided was just for us to see and work locally? I am asking these questions as I am not quite sure if I am going on right direction.</p>\n\n<p>Can someone please clarify?</p>",
      "rawMarkdown": "Hi,\nDo we need to run our code finally on sample training data of 400 files while submitting the code to build the model? Is the 470 GB that was provided was just for us to see and work locally? I am asking these questions as I am not quite sure if I am going on right direction.\n\nCan someone please clarify?",
      "votes": 1,
      "replies": [
        {
          "id": 705803,
          "postDate": "2019-12-29T13:46:20.290Z",
          "content": "<p>In case you're still stuck. 470 gb data is for you to use and train your model. then load that model as dataset into kaggle. Then use that pretrained model to predict the test videos and make submission.</p>",
          "rawMarkdown": "In case you're still stuck. 470 gb data is for you to use and train your model. then load that model as dataset into kaggle. Then use that pretrained model to predict the test videos and make submission.",
          "votes": 1
        },
        {
          "id": 734724,
          "postDate": "2020-02-01T22:10:20.520Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 718861,
      "postDate": "2020-01-14T21:58:37.873Z",
      "content": "<p>This competition is definitely NOT good starting point for Machine Learning beginner ! :-) \nI would suggest to start with Titanic, Housing price prediction competition...</p>",
      "rawMarkdown": "This competition is definitely NOT good starting point for Machine Learning beginner ! :-) \nI would suggest to start with Titanic, Housing price prediction competition...",
      "votes": -4,
      "replies": [
        {
          "id": 718876,
          "postDate": "2020-01-14T22:26:04.490Z",
          "content": "<p>Actually to be fair if a beginner can understand the requirements and envisage an architecture everything else is immaterial. If they can't, they wouldn't enter in the first place. The same argument can be made for professionals come to think of it.</p>",
          "rawMarkdown": "Actually to be fair if a beginner can understand the requirements and envisage an architecture everything else is immaterial. If they can't, they wouldn't enter in the first place. The same argument can be made for professionals come to think of it.",
          "votes": 5
        },
        {
          "id": 747977,
          "postDate": "2020-02-17T04:06:41.163Z",
          "content": "<p>This competition can be a great starting point for a beginner. It is a great opportunity for beginners to start using cloud computing, dealing with \"BIG\" data, finding the optimal(fastest) way to preprocess data.</p>",
          "rawMarkdown": "This competition can be a great starting point for a beginner. It is a great opportunity for beginners to start using cloud computing, dealing with \"BIG\" data, finding the optimal(fastest) way to preprocess data.",
          "votes": 3
        }
      ]
    },
    {
      "id": 2352486,
      "postDate": "2023-07-21T03:18:41.733Z",
      "content": "<p>Thank you for motivating me to ask this directly from the community <br>\nI want to download the dataset and use it for my research where I will be using this data set to train multiple CNN's model to detect either the video is deepfake or not. If I go for the 470 gb of data is this the dataset I should be working with or there is something else as well <br>\nIs the dataset labeled still as When I downloaded the test videos and sample_submission.csv files there were no indication of labeled data as if the concerned video is deepfake or not Please help</p>",
      "rawMarkdown": "Thank you for motivating me to ask this directly from the community \nI want to download the dataset and use it for my research where I will be using this data set to train multiple CNN's model to detect either the video is deepfake or not. If I go for the 470 gb of data is this the dataset I should be working with or there is something else as well \nIs the dataset labeled still as When I downloaded the test videos and sample_submission.csv files there were no indication of labeled data as if the concerned video is deepfake or not Please help"
    },
    {
      "id": 747903,
      "postDate": "2020-02-17T02:10:44.223Z",
      "content": "<p>Hello, I'm trying to get 30 random frames from each video and store it into a np array for keras.  I feel like I'm doing it in a bad way. Can you please take a look at it?\n`\nimport numpy as np\nimport cv2 as cv\nimport random\nimport pandas as pd\ntrain_folder=\"/kaggle/input/deepfake-detection-challenge/train_sample_videos/\"\ntrain_label=data = pd.read_json(train_folder+\"metadata.json\",orient=\"index\")</p>\n\n<p>def getFrames(name):\n    cap = cv.VideoCapture( train_folder+ name)\n    images=[]\n    for i in range(30):\n        cap.set(1,random.random()*100)\n        success,image=cap.read()\n        image = cv.cvtColor(image, cv.COLOR_BGR2RGB)\n        image.resize(382,223)\n        image=image.reshape((382,223,1)).astype('float32')/255\n        images.append(image)\n    cap.release()\n    return images</p>\n\n<p>train_data=[]\ntrain_label_val=[]</p>\n\n<h1>get 30 frames from each video</h1>\n\n<p>for index,name in enumerate(train_label.index[:10]):\n    for i in range(1):\n        #save frame\n        img=getFrames(name)\n        train_data=train_data+img\n        #save train value\n        train_label_val.append(train_label[\"label\"][index]==\"REAL\")</p>\n\n<h1>make np array</h1>\n\n<p>train_data=np.array(train_data)\ntrain_label_val=np.array(train_label_val)`</p>",
      "rawMarkdown": "Hello, I'm trying to get 30 random frames from each video and store it into a np array for keras.  I feel like I'm doing it in a bad way. Can you please take a look at it?\n`\nimport numpy as np\nimport cv2 as cv\nimport random\nimport pandas as pd\ntrain_folder=\"/kaggle/input/deepfake-detection-challenge/train_sample_videos/\"\ntrain_label=data = pd.read_json(train_folder+\"metadata.json\",orient=\"index\")\n\ndef getFrames(name):\n    cap = cv.VideoCapture( train_folder+ name)\n    images=[]\n    for i in range(30):\n        cap.set(1,random.random()*100)\n        success,image=cap.read()\n        image = cv.cvtColor(image, cv.COLOR_BGR2RGB)\n        image.resize(382,223)\n        image=image.reshape((382,223,1)).astype('float32')/255\n        images.append(image)\n    cap.release()\n    return images\n\ntrain_data=[]\ntrain_label_val=[]\n\n#get 30 frames from each video\nfor index,name in enumerate(train_label.index[:10]):\n    for i in range(1):\n        #save frame\n        img=getFrames(name)\n        train_data=train_data+img\n        #save train value\n        train_label_val.append(train_label[\"label\"][index]==\"REAL\")\n\n#make np array\ntrain_data=np.array(train_data)\ntrain_label_val=np.array(train_label_val)`"
    },
    {
      "id": 745205,
      "postDate": "2020-02-13T16:02:45.750Z",
      "content": "<p>Hi ,both audio and vdeio should be checked or only vedio? seems tough work</p>",
      "rawMarkdown": "Hi ,both audio and vdeio should be checked or only vedio? seems tough work"
    },
    {
      "id": 739074,
      "postDate": "2020-02-07T11:37:22.403Z",
      "content": "<p>I have found out how to extract frames from the Video files and Extract faces from the image using MTCNN. What's next? How can I train model (say Inception or Resnet) as a whole?</p>",
      "rawMarkdown": "I have found out how to extract frames from the Video files and Extract faces from the image using MTCNN. What's next? How can I train model (say Inception or Resnet) as a whole?"
    },
    {
      "id": 696301,
      "postDate": "2019-12-16T12:33:19.243Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 694153,
      "postDate": "2019-12-13T08:48:34.847Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 2872175,
      "postDate": "2024-06-14T16:11:17.113Z",
      "content": "<p>Really worth it. Thanks .</p>",
      "rawMarkdown": "Really worth it. Thanks ."
    },
    {
      "id": 1397473,
      "postDate": "2021-07-23T08:30:28.363Z",
      "content": "<p>Thanks for the share.</p>",
      "rawMarkdown": "Thanks for the share."
    },
    {
      "id": 962367,
      "postDate": "2020-08-08T04:54:09.767Z",
      "content": "<p>Thank you very much</p>",
      "rawMarkdown": "Thank you very much"
    },
    {
      "id": 820203,
      "postDate": "2020-04-25T08:20:36.887Z",
      "content": "<p>Thank you for share</p>",
      "rawMarkdown": "Thank you for share"
    },
    {
      "id": 814479,
      "postDate": "2020-04-20T17:41:31.713Z",
      "content": "<p>Thank you so much</p>",
      "rawMarkdown": "Thank you so much"
    },
    {
      "id": 756956,
      "postDate": "2020-02-26T09:07:39.363Z",
      "content": "<p>Thank you for share.</p>",
      "rawMarkdown": "Thank you for share."
    },
    {
      "id": 694171,
      "postDate": "2019-12-13T09:10:54.223Z",
      "content": "<p>Thank you for the helpful post!</p>",
      "rawMarkdown": "Thank you for the helpful post!"
    },
    {
      "id": 693282,
      "postDate": "2019-12-12T07:57:44.363Z",
      "content": "<p>Thank you for the post. </p>",
      "rawMarkdown": "Thank you for the post. "
    }
  ],
  "comments": [
    {
      "id": 700917,
      "author_name": "Raja Suman C",
      "author_url": "",
      "post_date": "2019-12-22T20:14:16.160000",
      "content": "<p>Hi,\nDo we need to run our code finally on sample training data of 400 files while submitting the code to build the model? Is the 470 GB that was provided was just for us to see and work locally? I am asking these questions as I am not quite sure if I am going on right direction.</p>\n\n<p>Can someone please clarify?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 705803,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2019-12-29T13:46:20.290000",
          "content": "<p>In case you're still stuck. 470 gb data is for you to use and train your model. then load that model as dataset into kaggle. Then use that pretrained model to predict the test videos and make submission.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 734724,
          "author_name": "adamRR",
          "author_url": "",
          "post_date": "2020-02-01T22:10:20.520000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 718861,
      "author_name": "Vivek Bombatkar",
      "author_url": "",
      "post_date": "2020-01-14T21:58:37.873000",
      "content": "<p>This competition is definitely NOT good starting point for Machine Learning beginner ! :-) \nI would suggest to start with Titanic, Housing price prediction competition...</p>",
      "votes": -4,
      "replies": [
        {
          "id": 718876,
          "author_name": "Ed Austin",
          "author_url": "",
          "post_date": "2020-01-14T22:26:04.490000",
          "content": "<p>Actually to be fair if a beginner can understand the requirements and envisage an architecture everything else is immaterial. If they can't, they wouldn't enter in the first place. The same argument can be made for professionals come to think of it.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 747977,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-02-17T04:06:41.163000",
          "content": "<p>This competition can be a great starting point for a beginner. It is a great opportunity for beginners to start using cloud computing, dealing with \"BIG\" data, finding the optimal(fastest) way to preprocess data.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2352486,
      "author_name": "muhammad aleem",
      "author_url": "",
      "post_date": "2023-07-21T03:18:41.733000",
      "content": "<p>Thank you for motivating me to ask this directly from the community <br>\nI want to download the dataset and use it for my research where I will be using this data set to train multiple CNN's model to detect either the video is deepfake or not. If I go for the 470 gb of data is this the dataset I should be working with or there is something else as well <br>\nIs the dataset labeled still as When I downloaded the test videos and sample_submission.csv files there were no indication of labeled data as if the concerned video is deepfake or not Please help</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 747903,
      "author_name": "Tommy Dong",
      "author_url": "",
      "post_date": "2020-02-17T02:10:44.223000",
      "content": "<p>Hello, I'm trying to get 30 random frames from each video and store it into a np array for keras.  I feel like I'm doing it in a bad way. Can you please take a look at it?\n`\nimport numpy as np\nimport cv2 as cv\nimport random\nimport pandas as pd\ntrain_folder=\"/kaggle/input/deepfake-detection-challenge/train_sample_videos/\"\ntrain_label=data = pd.read_json(train_folder+\"metadata.json\",orient=\"index\")</p>\n\n<p>def getFrames(name):\n    cap = cv.VideoCapture( train_folder+ name)\n    images=[]\n    for i in range(30):\n        cap.set(1,random.random()*100)\n        success,image=cap.read()\n        image = cv.cvtColor(image, cv.COLOR_BGR2RGB)\n        image.resize(382,223)\n        image=image.reshape((382,223,1)).astype('float32')/255\n        images.append(image)\n    cap.release()\n    return images</p>\n\n<p>train_data=[]\ntrain_label_val=[]</p>\n\n<h1>get 30 frames from each video</h1>\n\n<p>for index,name in enumerate(train_label.index[:10]):\n    for i in range(1):\n        #save frame\n        img=getFrames(name)\n        train_data=train_data+img\n        #save train value\n        train_label_val.append(train_label[\"label\"][index]==\"REAL\")</p>\n\n<h1>make np array</h1>\n\n<p>train_data=np.array(train_data)\ntrain_label_val=np.array(train_label_val)`</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 745205,
      "author_name": "Feng",
      "author_url": "",
      "post_date": "2020-02-13T16:02:45.750000",
      "content": "<p>Hi ,both audio and vdeio should be checked or only vedio? seems tough work</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 739074,
      "author_name": "Mahesh Deshwal",
      "author_url": "",
      "post_date": "2020-02-07T11:37:22.403000",
      "content": "<p>I have found out how to extract frames from the Video files and Extract faces from the image using MTCNN. What's next? How can I train model (say Inception or Resnet) as a whole?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 696301,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-16T12:33:19.243000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 694153,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-13T08:48:34.847000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 2872175,
      "author_name": "Sumeet Rodiya",
      "author_url": "",
      "post_date": "2024-06-14T16:11:17.113000",
      "content": "<p>Really worth it. Thanks .</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1397473,
      "author_name": "Gabriel Ohowa Owino",
      "author_url": "",
      "post_date": "2021-07-23T08:30:28.363000",
      "content": "<p>Thanks for the share.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 962367,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-08T04:54:09.767000",
      "content": "<p>Thank you very much</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 820203,
      "author_name": "Larry Yang",
      "author_url": "",
      "post_date": "2020-04-25T08:20:36.887000",
      "content": "<p>Thank you for share</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 814479,
      "author_name": "sudesh mate",
      "author_url": "",
      "post_date": "2020-04-20T17:41:31.713000",
      "content": "<p>Thank you so much</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 756956,
      "author_name": "JackJiang",
      "author_url": "",
      "post_date": "2020-02-26T09:07:39.363000",
      "content": "<p>Thank you for share.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 694171,
      "author_name": "Akhil Singh",
      "author_url": "",
      "post_date": "2019-12-13T09:10:54.223000",
      "content": "<p>Thank you for the helpful post!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 693282,
      "author_name": "AtulVerma",
      "author_url": "",
      "post_date": "2019-12-12T07:57:44.363000",
      "content": "<p>Thank you for the post. </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "692877": "New to machine learning and data science? No question is too basic or too simple. Feel free to start your own thread, or use this thread as a place to post any first-timer clarifying questions for the Kaggle community to help you with! \n\nIf you would consider yourself a beginner but don't know where to get started, let other Kagglers help you take your first steps here!\n\nNew to Kaggle? Take a look at a few videos our very own Dr. Rachael Tatman has put together to learn a bit more about [site etiquette](https://www.youtube.com/watch?v=aIus8si_Et0), [Kaggle lingo](https://www.youtube.com/watch?v=sEJHyuWKd-s), and [how to enter a competition using Kaggle Notebooks](https://www.youtube.com/watch?&amp;v=GJBOMWpLpTQ).\n\nThis competition has a very unique submission method, so be sure to check out the [Getting Started](https://www.kaggle.com/c/deepfake-detection-challenge/overview/getting-started) thread for more info!",
    "700917": "Hi,\nDo we need to run our code finally on sample training data of 400 files while submitting the code to build the model? Is the 470 GB that was provided was just for us to see and work locally? I am asking these questions as I am not quite sure if I am going on right direction.\n\nCan someone please clarify?",
    "718861": "This competition is definitely NOT good starting point for Machine Learning beginner ! :-) \nI would suggest to start with Titanic, Housing price prediction competition...",
    "2352486": "Thank you for motivating me to ask this directly from the community \nI want to download the dataset and use it for my research where I will be using this data set to train multiple CNN's model to detect either the video is deepfake or not. If I go for the 470 gb of data is this the dataset I should be working with or there is something else as well \nIs the dataset labeled still as When I downloaded the test videos and sample_submission.csv files there were no indication of labeled data as if the concerned video is deepfake or not Please help",
    "747903": "Hello, I'm trying to get 30 random frames from each video and store it into a np array for keras.  I feel like I'm doing it in a bad way. Can you please take a look at it?\n`\nimport numpy as np\nimport cv2 as cv\nimport random\nimport pandas as pd\ntrain_folder=\"/kaggle/input/deepfake-detection-challenge/train_sample_videos/\"\ntrain_label=data = pd.read_json(train_folder+\"metadata.json\",orient=\"index\")\n\ndef getFrames(name):\n    cap = cv.VideoCapture( train_folder+ name)\n    images=[]\n    for i in range(30):\n        cap.set(1,random.random()*100)\n        success,image=cap.read()\n        image = cv.cvtColor(image, cv.COLOR_BGR2RGB)\n        image.resize(382,223)\n        image=image.reshape((382,223,1)).astype('float32')/255\n        images.append(image)\n    cap.release()\n    return images\n\ntrain_data=[]\ntrain_label_val=[]\n\n#get 30 frames from each video\nfor index,name in enumerate(train_label.index[:10]):\n    for i in range(1):\n        #save frame\n        img=getFrames(name)\n        train_data=train_data+img\n        #save train value\n        train_label_val.append(train_label[\"label\"][index]==\"REAL\")\n\n#make np array\ntrain_data=np.array(train_data)\ntrain_label_val=np.array(train_label_val)`",
    "745205": "Hi ,both audio and vdeio should be checked or only vedio? seems tough work",
    "739074": "I have found out how to extract frames from the Video files and Extract faces from the image using MTCNN. What's next? How can I train model (say Inception or Resnet) as a whole?",
    "696301": "",
    "694153": "",
    "2872175": "Really worth it. Thanks .",
    "1397473": "Thanks for the share.",
    "962367": "Thank you very much",
    "820203": "Thank you for share",
    "814479": "Thank you so much",
    "756956": "Thank you for share.",
    "694171": "Thank you for the helpful post!",
    "693282": "Thank you for the post. "
  }
}