{
  "id": 122936,
  "title": "Effective Train Size",
  "url": "/competitions/deepfake-detection-challenge/discussion/122936",
  "author_name": "",
  "post_date": "2019-12-23T18:48:23.373311900Z",
  "votes": 11,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I've been wondering if to do well in this competition we really need to use the full train set. Specially since it's imbalanced with only ~19k of real videos, and 100k of fake videos.</p>\n\n<p>I've been trying some models with small sample sizes of the train set, and increasing the size slowly to see the effect on the test set, I could see a huge jump from using 1k videos to 4k videos and keeping the train set balanced on my samples.</p>\n\n<p>However, my results are still worse than just setting 0.5 to all. Since I still don't have the computer power to train larger sets for now, I would like to know from you all, specially the people who have ~0.5x on the leaderboard at the moment, are you guys using the full train set to train or are you taking a sample? Anyone else trying models with samples of the train set and getting results worse than 0.5 to all?</p>\n\n<p>Merry Christmas to everyone working on this competition during this time of the year btw :) </p>",
  "messages": [
    {
      "id": "701669",
      "postDate": "12/23/2019 18:48:23",
      "content": "<p>I've been wondering if to do well in this competition we really need to use the full train set. Specially since it's imbalanced with only ~19k of real videos, and 100k of fake videos.</p>\n\n<p>I've been trying some models with small sample sizes of the train set, and increasing the size slowly to see the effect on the test set, I could see a huge jump from using 1k videos to 4k videos and keeping the train set balanced on my samples.</p>\n\n<p>However, my results are still worse than just setting 0.5 to all. Since I still don't have the computer power to train larger sets for now, I would like to know from you all, specially the people who have ~0.5x on the leaderboard at the moment, are you guys using the full train set to train or are you taking a sample? Anyone else trying models with samples of the train set and getting results worse than 0.5 to all?</p>\n\n<p>Merry Christmas to everyone working on this competition during this time of the year btw :) </p>",
      "rawMarkdown": "I've been wondering if to do well in this competition we really need to use the full train set. Specially since it's imbalanced with only ~19k of real videos, and 100k of fake videos.\n\nI've been trying some models with small sample sizes of the train set, and increasing the size slowly to see the effect on the test set, I could see a huge jump from using 1k videos to 4k videos and keeping the train set balanced on my samples.\n\nHowever, my results are still worse than just setting 0.5 to all. Since I still don't have the computer power to train larger sets for now, I would like to know from you all, specially the people who have ~0.5x on the leaderboard at the moment, are you guys using the full train set to train or are you taking a sample? Anyone else trying models with samples of the train set and getting results worse than 0.5 to all?\n\nMerry Christmas to everyone working on this competition during this time of the year btw :)",
      "votes": null
    },
    {
      "id": "704419",
      "postDate": "12/27/2019 13:02:55",
      "content": "<p>I had access to more computer power and was finally able to break the 0.5 submission. First I trained on 1k videos and got ~1.05, after increasing my train size to 4k videos I got ~0.71, and finally with ~17k videos I got to ~0.64, always using the same model and processing the data the same way. I also kept the train set balanced roughly 50/50 between fake and real images. I also used only 50 frames from each video for training and predicting. I also got similar results on my validation set, my last submission (that scored 0.64 on test) on validation scored 0.66.</p>\n\n<p>I believe increasing the train size might yield better results on test, but you don't need to use all of it to validate your models, this makes it easier to experiment with cheaper hardware. Good luck everyone!</p>",
      "rawMarkdown": "I had access to more computer power and was finally able to break the 0.5 submission. First I trained on 1k videos and got ~1.05, after increasing my train size to 4k videos I got ~0.71, and finally with ~17k videos I got to ~0.64, always using the same model and processing the data the same way. I also kept the train set balanced roughly 50/50 between fake and real images. I also used only 50 frames from each video for training and predicting. I also got similar results on my validation set, my last submission (that scored 0.64 on test) on validation scored 0.66.\n\nI believe increasing the train size might yield better results on test, but you don't need to use all of it to validate your models, this makes it easier to experiment with cheaper hardware. Good luck everyone!",
      "votes": null
    },
    {
      "id": "705735",
      "postDate": "12/29/2019 11:21:55",
      "content": "<p>Can you upload the metadata.json files for all the training data? It will be very helpful for analyzing data distribution.The training set is too large and I'm training on 400 training samples now...</p>",
      "rawMarkdown": "Can you upload the metadata.json files for all the training data? It will be very helpful for analyzing data distribution.The training set is too large and I'm training on 400 training samples now...",
      "votes": null
    },
    {
      "id": "705839",
      "postDate": "12/29/2019 15:09:02",
      "content": "<p>Hi <a href=\"/lanncx\">@lanncx</a> , your request was the last drop and I actually did it. Take a look at <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">the dataset</a>, and the <a href=\"https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata\">accompanying kernel</a>.</p>",
      "rawMarkdown": "Hi @lanncx , your request was the last drop and I actually did it. Take a look at [the dataset](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc), and the [accompanying kernel](https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata).",
      "votes": null
    },
    {
      "id": "705851",
      "postDate": "12/29/2019 15:23:04",
      "content": "<p>That's great, thanks for your sharing!</p>",
      "rawMarkdown": "That's great, thanks for your sharing!",
      "votes": null
    },
    {
      "id": "706922",
      "postDate": "12/31/2019 02:26:29",
      "content": "<p>How did you get your validation set? Extract them by random pick frames or parts?</p>",
      "rawMarkdown": "How did you get your validation set? Extract them by random pick frames or parts?",
      "votes": null
    },
    {
      "id": "707357",
      "postDate": "12/31/2019 17:31:01",
      "content": "<p>The results I talked about here the validation set was created using a group split (using sklearn GroupShufffleSplit) by original video. This means that all fakes of one original video were grouped together either on train or validation. </p>\n\n<p>However, for this experiments I used a really shallow CNN, when trying the same scheme for a deeper one my model overfitted, so I did the grouping by part (the 50 different files) this also keeps the constraint of original and fakes grouped together but makes the validation and train sets more different between them. </p>\n\n<p>The reason, I believe, is that the files also group videos by actors, this means that validation won't have the same actors as train. If you don't do this you might be in risk that your model will score high in validation because the network is learning facial features of the actors instead of features that generalize well for classifying fake/real, this will be specially true if you are using a sample of the data, this effect might diminish if you increase your sample and then it shouldn't be different then doing the validation set by grouping by original video. </p>",
      "rawMarkdown": "The results I talked about here the validation set was created using a group split (using sklearn GroupShufffleSplit) by original video. This means that all fakes of one original video were grouped together either on train or validation. \n\nHowever, for this experiments I used a really shallow CNN, when trying the same scheme for a deeper one my model overfitted, so I did the grouping by part (the 50 different files) this also keeps the constraint of original and fakes grouped together but makes the validation and train sets more different between them. \n\nThe reason, I believe, is that the files also group videos by actors, this means that validation won't have the same actors as train. If you don't do this you might be in risk that your model will score high in validation because the network is learning facial features of the actors instead of features that generalize well for classifying fake/real, this will be specially true if you are using a sample of the data, this effect might diminish if you increase your sample and then it shouldn't be different then doing the validation set by grouping by original video.",
      "votes": null
    },
    {
      "id": "709592",
      "postDate": "01/03/2020 17:44:19",
      "content": "<p>Using a deeper cnn we can get better results with less data: was able to achieve 0.65 in LB using frames from 5,000 videos, trained only for 3 epochs</p>",
      "rawMarkdown": "Using a deeper cnn we can get better results with less data: was able to achieve 0.65 in LB using frames from 5,000 videos, trained only for 3 epochs",
      "votes": null
    },
    {
      "id": "709703",
      "postDate": "01/03/2020 21:08:46",
      "content": "<p>Hi Carlos, \n  Did you fine tune a model pretrained on imagenet? Or you train a totally new model?</p>\n\n<p>Thanks,\n  Zanlang</p>",
      "rawMarkdown": "Hi Carlos, \n  Did you fine tune a model pretrained on imagenet? Or you train a totally new model?\n\n  Thanks,\n  Zanlang",
      "votes": null
    },
    {
      "id": "709711",
      "postDate": "01/03/2020 21:30:41",
      "content": "<p>So far I've only worked with fine tuning. Intuitively I think training a totally new model would be very expensive.. And you?</p>",
      "rawMarkdown": "So far I've only worked with fine tuning. Intuitively I think training a totally new model would be very expensive.. And you?",
      "votes": null
    },
    {
      "id": "709737",
      "postDate": "01/03/2020 22:09:15",
      "content": "<p>Hi Carlos,\n  Me too. I used a pretrained densenet and fine tune it. It actually converges quickly as well but the val_loss kept oscillating, which I think results from my unbalanced FAKE/REAL sample ratios (4:1). </p>",
      "rawMarkdown": "Hi Carlos,\n  Me too. I used a pretrained densenet and fine tune it. It actually converges quickly as well but the val_loss kept oscillating, which I think results from my unbalanced FAKE/REAL sample ratios (4:1).",
      "votes": null
    },
    {
      "id": "709742",
      "postDate": "01/03/2020 22:13:53",
      "content": "<p>And I struggle to extract faces from the test videos. It seems that cv2.CascadeClassifier() frequently failed to detect faces in a frame sampled from a video. For example, in my last committed kernel, it fails face detection in 157/400 videos...</p>\n\n<p>Any idea to improve the detection rate?\nThanks!</p>",
      "rawMarkdown": "And I struggle to extract faces from the test videos. It seems that cv2.CascadeClassifier() frequently failed to detect faces in a frame sampled from a video. For example, in my last committed kernel, it fails face detection in 157/400 videos...\n\nAny idea to improve the detection rate?\nThanks!",
      "votes": null
    },
    {
      "id": "709757",
      "postDate": "01/03/2020 22:34:30",
      "content": "<p>I tested few models: dlib (svm and cnn), and facenet-pytorch. Settled for the latter: it works pretty well. I did several tweaks to accelerate extraction... currently my code can detect a face, extract and classify it as fake/real at best @ 48 FPS... I'm happy with that, it hardly fails. I actually implemented progressive resizing. In the beginning, it resizes the image to a small size (160 x 160) and tries to detect a face. If it works, great: move to the next. If not, resize 2x to 320 x 320 and try again. If unsuccessful, double again. This logic is successful in the 1st attempt aprox 70% of the time. 2nd attempt is higher. 3rd attempt is close to 100% :)</p>",
      "rawMarkdown": "I tested few models: dlib (svm and cnn), and facenet-pytorch. Settled for the latter: it works pretty well. I did several tweaks to accelerate extraction... currently my code can detect a face, extract and classify it as fake/real at best @ 48 FPS... I'm happy with that, it hardly fails. I actually implemented progressive resizing. In the beginning, it resizes the image to a small size (160 x 160) and tries to detect a face. If it works, great: move to the next. If not, resize 2x to 320 x 320 and try again. If unsuccessful, double again. This logic is successful in the 1st attempt aprox 70% of the time. 2nd attempt is higher. 3rd attempt is close to 100% :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 704419,
      "author_name": "pedromb",
      "author_url": "",
      "post_date": "12/27/2019 13:02:55",
      "content": "<p>I had access to more computer power and was finally able to break the 0.5 submission. First I trained on 1k videos and got ~1.05, after increasing my train size to 4k videos I got ~0.71, and finally with ~17k videos I got to ~0.64, always using the same model and processing the data the same way. I also kept the train set balanced roughly 50/50 between fake and real images. I also used only 50 frames from each video for training and predicting. I also got similar results on my validation set, my last submission (that scored 0.64 on test) on validation scored 0.66.</p>\n\n<p>I believe increasing the train size might yield better results on test, but you don't need to use all of it to validate your models, this makes it easier to experiment with cheaper hardware. Good luck everyone!</p>",
      "votes": null,
      "replies": [
        {
          "id": 706922,
          "author_name": "fionalxd",
          "author_url": "",
          "post_date": "12/31/2019 02:26:29",
          "content": "<p>How did you get your validation set? Extract them by random pick frames or parts?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 707357,
          "author_name": "pedromb",
          "author_url": "",
          "post_date": "12/31/2019 17:31:01",
          "content": "<p>The results I talked about here the validation set was created using a group split (using sklearn GroupShufffleSplit) by original video. This means that all fakes of one original video were grouped together either on train or validation. </p>\n\n<p>However, for this experiments I used a really shallow CNN, when trying the same scheme for a deeper one my model overfitted, so I did the grouping by part (the 50 different files) this also keeps the constraint of original and fakes grouped together but makes the validation and train sets more different between them. </p>\n\n<p>The reason, I believe, is that the files also group videos by actors, this means that validation won't have the same actors as train. If you don't do this you might be in risk that your model will score high in validation because the network is learning facial features of the actors instead of features that generalize well for classifying fake/real, this will be specially true if you are using a sample of the data, this effect might diminish if you increase your sample and then it shouldn't be different then doing the validation set by grouping by original video. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709592,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/03/2020 17:44:19",
          "content": "<p>Using a deeper cnn we can get better results with less data: was able to achieve 0.65 in LB using frames from 5,000 videos, trained only for 3 epochs</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709703,
          "author_name": "",
          "author_url": "",
          "post_date": "01/03/2020 21:08:46",
          "content": "<p>Hi Carlos, \n  Did you fine tune a model pretrained on imagenet? Or you train a totally new model?</p>\n\n<p>Thanks,\n  Zanlang</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709711,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/03/2020 21:30:41",
          "content": "<p>So far I've only worked with fine tuning. Intuitively I think training a totally new model would be very expensive.. And you?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709737,
          "author_name": "",
          "author_url": "",
          "post_date": "01/03/2020 22:09:15",
          "content": "<p>Hi Carlos,\n  Me too. I used a pretrained densenet and fine tune it. It actually converges quickly as well but the val_loss kept oscillating, which I think results from my unbalanced FAKE/REAL sample ratios (4:1). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709742,
          "author_name": "",
          "author_url": "",
          "post_date": "01/03/2020 22:13:53",
          "content": "<p>And I struggle to extract faces from the test videos. It seems that cv2.CascadeClassifier() frequently failed to detect faces in a frame sampled from a video. For example, in my last committed kernel, it fails face detection in 157/400 videos...</p>\n\n<p>Any idea to improve the detection rate?\nThanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709757,
          "author_name": "carlossouza",
          "author_url": "",
          "post_date": "01/03/2020 22:34:30",
          "content": "<p>I tested few models: dlib (svm and cnn), and facenet-pytorch. Settled for the latter: it works pretty well. I did several tweaks to accelerate extraction... currently my code can detect a face, extract and classify it as fake/real at best @ 48 FPS... I'm happy with that, it hardly fails. I actually implemented progressive resizing. In the beginning, it resizes the image to a small size (160 x 160) and tries to detect a face. If it works, great: move to the next. If not, resize 2x to 320 x 320 and try again. If unsuccessful, double again. This logic is successful in the 1st attempt aprox 70% of the time. 2nd attempt is higher. 3rd attempt is close to 100% :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 705735,
      "author_name": "lanncx",
      "author_url": "",
      "post_date": "12/29/2019 11:21:55",
      "content": "<p>Can you upload the metadata.json files for all the training data? It will be very helpful for analyzing data distribution.The training set is too large and I'm training on 400 training samples now...</p>",
      "votes": null,
      "replies": [
        {
          "id": 705839,
          "author_name": "zaharch",
          "author_url": "",
          "post_date": "12/29/2019 15:09:02",
          "content": "<p>Hi <a href=\"/lanncx\">@lanncx</a> , your request was the last drop and I actually did it. Take a look at <a href=\"https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc\">the dataset</a>, and the <a href=\"https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata\">accompanying kernel</a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 705851,
          "author_name": "lanncx",
          "author_url": "",
          "post_date": "12/29/2019 15:23:04",
          "content": "<p>That's great, thanks for your sharing!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "701669": "I've been wondering if to do well in this competition we really need to use the full train set. Specially since it's imbalanced with only ~19k of real videos, and 100k of fake videos.\n\nI've been trying some models with small sample sizes of the train set, and increasing the size slowly to see the effect on the test set, I could see a huge jump from using 1k videos to 4k videos and keeping the train set balanced on my samples.\n\nHowever, my results are still worse than just setting 0.5 to all. Since I still don't have the computer power to train larger sets for now, I would like to know from you all, specially the people who have ~0.5x on the leaderboard at the moment, are you guys using the full train set to train or are you taking a sample? Anyone else trying models with samples of the train set and getting results worse than 0.5 to all?\n\nMerry Christmas to everyone working on this competition during this time of the year btw :)",
    "704419": "I had access to more computer power and was finally able to break the 0.5 submission. First I trained on 1k videos and got ~1.05, after increasing my train size to 4k videos I got ~0.71, and finally with ~17k videos I got to ~0.64, always using the same model and processing the data the same way. I also kept the train set balanced roughly 50/50 between fake and real images. I also used only 50 frames from each video for training and predicting. I also got similar results on my validation set, my last submission (that scored 0.64 on test) on validation scored 0.66.\n\nI believe increasing the train size might yield better results on test, but you don't need to use all of it to validate your models, this makes it easier to experiment with cheaper hardware. Good luck everyone!",
    "705735": "Can you upload the metadata.json files for all the training data? It will be very helpful for analyzing data distribution.The training set is too large and I'm training on 400 training samples now...",
    "705839": "Hi @lanncx , your request was the last drop and I actually did it. Take a look at [the dataset](https://www.kaggle.com/zaharch/train-set-metadata-for-dfdc), and the [accompanying kernel](https://www.kaggle.com/zaharch/looking-at-the-full-train-set-metadata).",
    "705851": "That's great, thanks for your sharing!",
    "706922": "How did you get your validation set? Extract them by random pick frames or parts?",
    "707357": "The results I talked about here the validation set was created using a group split (using sklearn GroupShufffleSplit) by original video. This means that all fakes of one original video were grouped together either on train or validation. \n\nHowever, for this experiments I used a really shallow CNN, when trying the same scheme for a deeper one my model overfitted, so I did the grouping by part (the 50 different files) this also keeps the constraint of original and fakes grouped together but makes the validation and train sets more different between them. \n\nThe reason, I believe, is that the files also group videos by actors, this means that validation won't have the same actors as train. If you don't do this you might be in risk that your model will score high in validation because the network is learning facial features of the actors instead of features that generalize well for classifying fake/real, this will be specially true if you are using a sample of the data, this effect might diminish if you increase your sample and then it shouldn't be different then doing the validation set by grouping by original video.",
    "709592": "Using a deeper cnn we can get better results with less data: was able to achieve 0.65 in LB using frames from 5,000 videos, trained only for 3 epochs",
    "709703": "Hi Carlos, \n  Did you fine tune a model pretrained on imagenet? Or you train a totally new model?\n\n  Thanks,\n  Zanlang",
    "709711": "So far I've only worked with fine tuning. Intuitively I think training a totally new model would be very expensive.. And you?",
    "709737": "Hi Carlos,\n  Me too. I used a pretrained densenet and fine tune it. It actually converges quickly as well but the val_loss kept oscillating, which I think results from my unbalanced FAKE/REAL sample ratios (4:1).",
    "709742": "And I struggle to extract faces from the test videos. It seems that cv2.CascadeClassifier() frequently failed to detect faces in a frame sampled from a video. For example, in my last committed kernel, it fails face detection in 157/400 videos...\n\nAny idea to improve the detection rate?\nThanks!",
    "709757": "I tested few models: dlib (svm and cnn), and facenet-pytorch. Settled for the latter: it works pretty well. I did several tweaks to accelerate extraction... currently my code can detect a face, extract and classify it as fake/real at best @ 48 FPS... I'm happy with that, it hardly fails. I actually implemented progressive resizing. In the beginning, it resizes the image to a small size (160 x 160) and tries to detect a face. If it works, great: move to the next. If not, resize 2x to 320 x 320 and try again. If unsuccessful, double again. This logic is successful in the 1st attempt aprox 70% of the time. 2nd attempt is higher. 3rd attempt is close to 100% :)"
  },
  "source": "meta"
}