{
  "id": 124110,
  "title": "Frame by Frame Or Stream?",
  "url": "/competitions/deepfake-detection-challenge/discussion/124110",
  "author_name": "",
  "post_date": "2020-01-02T03:20:06.852625600Z",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I am new to preprocessing videos for training and I have been struggling to find a \"definitive\" answer to my question. </p>\n\n<p>To train a model with videos, do you: \nTake in a video extract (for example) the first 50 frames, label them in their sequence on disk and then pass that to the model?\nor \nTake in a video, extract (for example) the first 50 frames and then pass that sequence of frames to the model directly? </p>\n\n<p>I could see advantages and disadvantages like flexibilty to try different distributions of frames but slower because you'd have to extract more images. The other disadvantage for the first option would be the the large amount of data storage needed to hold all those frames</p>\n\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "708161",
      "postDate": "01/02/2020 03:20:06",
      "content": "<p>Hello,</p>\n\n<p>I am new to preprocessing videos for training and I have been struggling to find a \"definitive\" answer to my question. </p>\n\n<p>To train a model with videos, do you: \nTake in a video extract (for example) the first 50 frames, label them in their sequence on disk and then pass that to the model?\nor \nTake in a video, extract (for example) the first 50 frames and then pass that sequence of frames to the model directly? </p>\n\n<p>I could see advantages and disadvantages like flexibilty to try different distributions of frames but slower because you'd have to extract more images. The other disadvantage for the first option would be the the large amount of data storage needed to hold all those frames</p>\n\n<p>Thank you!</p>",
      "rawMarkdown": "Hello,\n\nI am new to preprocessing videos for training and I have been struggling to find a \"definitive\" answer to my question. \n\nTo train a model with videos, do you: \nTake in a video extract (for example) the first 50 frames, label them in their sequence on disk and then pass that to the model?\nor \nTake in a video, extract (for example) the first 50 frames and then pass that sequence of frames to the model directly? \n\nI could see advantages and disadvantages like flexibilty to try different distributions of frames but slower because you'd have to extract more images. The other disadvantage for the first option would be the the large amount of data storage needed to hold all those frames\n\nThank you!",
      "votes": null
    },
    {
      "id": "708196",
      "postDate": "01/02/2020 04:05:10",
      "content": "<p>Thanks for sharing topic </p>",
      "rawMarkdown": "Thanks for sharing topic",
      "votes": null
    },
    {
      "id": "708468",
      "postDate": "01/02/2020 10:21:43",
      "content": "<p>I'm following your 1st approach. The 2nd approach sounds interesting, but it would have a very high cost. Imagine that:\n- You are training models using 10,000 videos (less than 10% of the total dataset)\n- You were able to come up with an algorithm that can detect faces at 30 FPS\n- You are using 50 frames/video</p>\n\n<p>In the 1st approach, you would detect the faces once, and reuse them. In the 2nd approach, every time you train a model, you'd have to detect. This means 10,000 x 50 / 30 / 3,600 ~= 4h40min every time you train a model, just to preprocess images! So, 2nd approach would be very inefficient in terms of cost. Does it make sense? :)</p>",
      "rawMarkdown": "I'm following your 1st approach. The 2nd approach sounds interesting, but it would have a very high cost. Imagine that:\n- You are training models using 10,000 videos (less than 10% of the total dataset)\n- You were able to come up with an algorithm that can detect faces at 30 FPS\n- You are using 50 frames/video\n\nIn the 1st approach, you would detect the faces once, and reuse them. In the 2nd approach, every time you train a model, you'd have to detect. This means 10,000 x 50 / 30 / 3,600 ~= 4h40min every time you train a model, just to preprocess images! So, 2nd approach would be very inefficient in terms of cost. Does it make sense? :)",
      "votes": null
    },
    {
      "id": "708476",
      "postDate": "01/02/2020 10:35:28",
      "content": "<p>This assumes that you can extract information from the frames (such as the faces) just once and reuse it many times. However, if you're always working on the entirety of the video frames (let's say you don't just look at faces) then it might be faster to make a data loader that works directly on the mp4 file.</p>",
      "rawMarkdown": "This assumes that you can extract information from the frames (such as the faces) just once and reuse it many times. However, if you're always working on the entirety of the video frames (let's say you don't just look at faces) then it might be faster to make a data loader that works directly on the mp4 file.",
      "votes": null
    },
    {
      "id": "708491",
      "postDate": "01/02/2020 10:54:29",
      "content": "<p>Yes, I'm working with faces. Every paper I've read shows that working with faces yields significantly better results than working with entire images... </p>",
      "rawMarkdown": "Yes, I'm working with faces. Every paper I've read shows that working with faces yields significantly better results than working with entire images...",
      "votes": null
    },
    {
      "id": "708673",
      "postDate": "01/02/2020 14:34:57",
      "content": "<p>Danny - I am working on the approach where the next 30-60 days I am using frames saved to disk - because a whole lot goofs and restarts come with my Python skill level. </p>\n\n<p>But my goal would be that 60-90 days I taking in the videos.  Not sure that I will get here, not sure you need to get here for a million dollars - but might need to get here because everything needs to happen within the 9 hour compute time.</p>",
      "rawMarkdown": "Danny - I am working on the approach where the next 30-60 days I am using frames saved to disk - because a whole lot goofs and restarts come with my Python skill level. \n\nBut my goal would be that 60-90 days I taking in the videos.  Not sure that I will get here, not sure you need to get here for a million dollars - but might need to get here because everything needs to happen within the 9 hour compute time.",
      "votes": null
    },
    {
      "id": "708701",
      "postDate": "01/02/2020 15:07:44",
      "content": "<p>Thanks for your reply Carlos! \nThat absolutely makes sense. I was thinking the same thing but I was hung up on how we'd handle the test set that they withhold. With your answer, I'm assuming now that my submission code will reuse the function to detect, extract, and label images will need to do that and predict in under the 9 hour limit. Hence why there are all the posts about optimizing the frame extraction from videos?</p>",
      "rawMarkdown": "Thanks for your reply Carlos! \nThat absolutely makes sense. I was thinking the same thing but I was hung up on how we'd handle the test set that they withhold. With your answer, I'm assuming now that my submission code will reuse the function to detect, extract, and label images will need to do that and predict in under the 9 hour limit. Hence why there are all the posts about optimizing the frame extraction from videos?",
      "votes": null
    },
    {
      "id": "708768",
      "postDate": "01/02/2020 16:57:20",
      "content": "<p>Because our algos have to predict unseen 4,000 videos (private test set). Using the same numbers (face detection at 30 FPS, using 50 frames/video), just detecting faces would take close to 2 hours from the 9 limit!</p>",
      "rawMarkdown": "Because our algos have to predict unseen 4,000 videos (private test set). Using the same numbers (face detection at 30 FPS, using 50 frames/video), just detecting faces would take close to 2 hours from the 9 limit!",
      "votes": null
    },
    {
      "id": "708874",
      "postDate": "01/02/2020 19:40:56",
      "content": "<p>Hi Carlos,\n  I don't quite get your point about the two strategies. Imagine we want to sample 50 frames from a video, what's difference between the two strategies? I think even in 1st one, we still need to detect face frame by frame (50 times) since the face of people may move?</p>",
      "rawMarkdown": "Hi Carlos,\n  I don't quite get your point about the two strategies. Imagine we want to sample 50 frames from a video, what's difference between the two strategies? I think even in 1st one, we still need to detect face frame by frame (50 times) since the face of people may move?",
      "votes": null
    },
    {
      "id": "708942",
      "postDate": "01/02/2020 21:04:01",
      "content": "<p>As I understood the question, in the 1st scenario, the model receives images as inputs, i.e. they can be cached in disk. In the 2nd scenario, the volume receives videos as inputs: either you cache frames from all videos (days of continuous preprocessing), or every time you sample videos to train your model, you will have to extract frames</p>",
      "rawMarkdown": "As I understood the question, in the 1st scenario, the model receives images as inputs, i.e. they can be cached in disk. In the 2nd scenario, the volume receives videos as inputs: either you cache frames from all videos (days of continuous preprocessing), or every time you sample videos to train your model, you will have to extract frames",
      "votes": null
    },
    {
      "id": "708974",
      "postDate": "01/02/2020 22:11:03",
      "content": "<p>Hi Carlos,\n  Great thanks for your explanation. Now I can get your point clearly. I should take the first solution as well since yesterday I found my submission kernel takes several hours to finish on Kaggle. And I believe I should sample some images for coming multiple rounds of training as well.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi Carlos,\n  Great thanks for your explanation. Now I can get your point clearly. I should take the first solution as well since yesterday I found my submission kernel takes several hours to finish on Kaggle. And I believe I should sample some images for coming multiple rounds of training as well.\n\n  Thanks!",
      "votes": null
    },
    {
      "id": "709017",
      "postDate": "01/02/2020 23:58:03",
      "content": "<p>I just wanted to say that Carlos did understand my questions right incase anyone else was confused by my post.</p>",
      "rawMarkdown": "I just wanted to say that Carlos did understand my questions right incase anyone else was confused by my post.",
      "votes": null
    },
    {
      "id": "714373",
      "postDate": "01/09/2020 11:15:50",
      "content": "<p>the first one is great because it requires less processing power  </p>",
      "rawMarkdown": "the first one is great because it requires less processing power",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 708196,
      "author_name": "ahmedsmara",
      "author_url": "",
      "post_date": "01/02/2020 04:05:10",
      "content": "<p>Thanks for sharing topic </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 708468,
      "author_name": "carlossouza",
      "author_url": "",
      "post_date": "01/02/2020 10:21:43",
      "content": "<p>I'm following your 1st approach. The 2nd approach sounds interesting, but it would have a very high cost. Imagine that:\n- You are training models using 10,000 videos (less than 10% of the total dataset)\n- You were able to come up with an algorithm that can detect faces at 30 FPS\n- You are using 50 frames/video</p>\n\n<p>In the 1st approach, you would detect the faces once, and reuse them. In the 2nd approach, every time you train a model, you'd have to detect. This means 10,000 x 50 / 30 / 3,600 ~= 4h40min every time you train a model, just to preprocess images! So, 2nd approach would be very inefficient in terms of cost. Does it make sense? :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 708476,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/02/2020 10:35:28",
          "content": "<p>This assumes that you can extract information from the frames (such as the faces) just once and reuse it many times. However, if you're always working on the entirety of the video frames (let's say you don't just look at faces) then it might be faster to make a data loader that works directly on the mp4 file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708491,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/02/2020 10:54:29",
          "content": "<p>Yes, I'm working with faces. Every paper I've read shows that working with faces yields significantly better results than working with entire images... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708701,
          "author_name": "dannyk2018",
          "author_url": "",
          "post_date": "01/02/2020 15:07:44",
          "content": "<p>Thanks for your reply Carlos! \nThat absolutely makes sense. I was thinking the same thing but I was hung up on how we'd handle the test set that they withhold. With your answer, I'm assuming now that my submission code will reuse the function to detect, extract, and label images will need to do that and predict in under the 9 hour limit. Hence why there are all the posts about optimizing the frame extraction from videos?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708768,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/02/2020 16:57:20",
          "content": "<p>Because our algos have to predict unseen 4,000 videos (private test set). Using the same numbers (face detection at 30 FPS, using 50 frames/video), just detecting faces would take close to 2 hours from the 9 limit!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708874,
          "author_name": "",
          "author_url": "",
          "post_date": "01/02/2020 19:40:56",
          "content": "<p>Hi Carlos,\n  I don't quite get your point about the two strategies. Imagine we want to sample 50 frames from a video, what's difference between the two strategies? I think even in 1st one, we still need to detect face frame by frame (50 times) since the face of people may move?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708942,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/02/2020 21:04:01",
          "content": "<p>As I understood the question, in the 1st scenario, the model receives images as inputs, i.e. they can be cached in disk. In the 2nd scenario, the volume receives videos as inputs: either you cache frames from all videos (days of continuous preprocessing), or every time you sample videos to train your model, you will have to extract frames</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 708974,
          "author_name": "",
          "author_url": "",
          "post_date": "01/02/2020 22:11:03",
          "content": "<p>Hi Carlos,\n  Great thanks for your explanation. Now I can get your point clearly. I should take the first solution as well since yesterday I found my submission kernel takes several hours to finish on Kaggle. And I believe I should sample some images for coming multiple rounds of training as well.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709017,
          "author_name": "dannyk2018",
          "author_url": "",
          "post_date": "01/02/2020 23:58:03",
          "content": "<p>I just wanted to say that Carlos did understand my questions right incase anyone else was confused by my post.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 708673,
      "author_name": "pcjimmmy",
      "author_url": "",
      "post_date": "01/02/2020 14:34:57",
      "content": "<p>Danny - I am working on the approach where the next 30-60 days I am using frames saved to disk - because a whole lot goofs and restarts come with my Python skill level. </p>\n\n<p>But my goal would be that 60-90 days I taking in the videos.  Not sure that I will get here, not sure you need to get here for a million dollars - but might need to get here because everything needs to happen within the 9 hour compute time.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 714373,
      "author_name": "venser",
      "author_url": "",
      "post_date": "01/09/2020 11:15:50",
      "content": "<p>the first one is great because it requires less processing power  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "708161": "Hello,\n\nI am new to preprocessing videos for training and I have been struggling to find a \"definitive\" answer to my question. \n\nTo train a model with videos, do you: \nTake in a video extract (for example) the first 50 frames, label them in their sequence on disk and then pass that to the model?\nor \nTake in a video, extract (for example) the first 50 frames and then pass that sequence of frames to the model directly? \n\nI could see advantages and disadvantages like flexibilty to try different distributions of frames but slower because you'd have to extract more images. The other disadvantage for the first option would be the the large amount of data storage needed to hold all those frames\n\nThank you!",
    "708196": "Thanks for sharing topic",
    "708468": "I'm following your 1st approach. The 2nd approach sounds interesting, but it would have a very high cost. Imagine that:\n- You are training models using 10,000 videos (less than 10% of the total dataset)\n- You were able to come up with an algorithm that can detect faces at 30 FPS\n- You are using 50 frames/video\n\nIn the 1st approach, you would detect the faces once, and reuse them. In the 2nd approach, every time you train a model, you'd have to detect. This means 10,000 x 50 / 30 / 3,600 ~= 4h40min every time you train a model, just to preprocess images! So, 2nd approach would be very inefficient in terms of cost. Does it make sense? :)",
    "708476": "This assumes that you can extract information from the frames (such as the faces) just once and reuse it many times. However, if you're always working on the entirety of the video frames (let's say you don't just look at faces) then it might be faster to make a data loader that works directly on the mp4 file.",
    "708491": "Yes, I'm working with faces. Every paper I've read shows that working with faces yields significantly better results than working with entire images...",
    "708673": "Danny - I am working on the approach where the next 30-60 days I am using frames saved to disk - because a whole lot goofs and restarts come with my Python skill level. \n\nBut my goal would be that 60-90 days I taking in the videos.  Not sure that I will get here, not sure you need to get here for a million dollars - but might need to get here because everything needs to happen within the 9 hour compute time.",
    "708701": "Thanks for your reply Carlos! \nThat absolutely makes sense. I was thinking the same thing but I was hung up on how we'd handle the test set that they withhold. With your answer, I'm assuming now that my submission code will reuse the function to detect, extract, and label images will need to do that and predict in under the 9 hour limit. Hence why there are all the posts about optimizing the frame extraction from videos?",
    "708768": "Because our algos have to predict unseen 4,000 videos (private test set). Using the same numbers (face detection at 30 FPS, using 50 frames/video), just detecting faces would take close to 2 hours from the 9 limit!",
    "708874": "Hi Carlos,\n  I don't quite get your point about the two strategies. Imagine we want to sample 50 frames from a video, what's difference between the two strategies? I think even in 1st one, we still need to detect face frame by frame (50 times) since the face of people may move?",
    "708942": "As I understood the question, in the 1st scenario, the model receives images as inputs, i.e. they can be cached in disk. In the 2nd scenario, the volume receives videos as inputs: either you cache frames from all videos (days of continuous preprocessing), or every time you sample videos to train your model, you will have to extract frames",
    "708974": "Hi Carlos,\n  Great thanks for your explanation. Now I can get your point clearly. I should take the first solution as well since yesterday I found my submission kernel takes several hours to finish on Kaggle. And I believe I should sample some images for coming multiple rounds of training as well.\n\n  Thanks!",
    "709017": "I just wanted to say that Carlos did understand my questions right incase anyone else was confused by my post.",
    "714373": "the first one is great because it requires less processing power"
  },
  "source": "meta"
}