{
  "id": 128919,
  "title": "Major difference between CV and LB",
  "url": "/competitions/deepfake-detection-challenge/discussion/128919",
  "author_name": "Akash",
  "post_date": "2020-02-04T09:51:55.908000",
  "votes": 17,
  "comment_count": 32,
  "views": 0,
  "content": "<p>Hello folks!</p>\n\n<p>I'm training my model on the entire dataset with each video sampled at 0.5 FPS leading to 5 frames per face in any video with image resolution 128x128. Also, I'm doing the entire thing using tf.keras if that helps.</p>\n\n<p>My training set is as follows:</p>\n\n<p>Train data: Randomly sample around 10k real video names. Sample 1 fake video for each of the real video sampled earlier. Build x_train and y_train by loading all the images from the selected set of real and fake videos and then shuffling the x_train, y_train (index is handled) one last time.</p>\n\n<p>Validation data: Randomly sample 2K real videos from the remaining pool of unsampled data for train. Do the entire charade for fake video similar to earlier. </p>\n\n<p>This set of data is then mean centered and is being trained on a DenseNet like network with parameter count of approx 900k.</p>\n\n<p>After training the network for around 120 epochs, I get train loss of 0.15 and validation loss of 0.25.</p>\n\n<p>For the purposes of video classification, I'm running at 1.1FPS and am taking the mean prediction of each set of faces (I have a reliable clustering algorithm to track faces). And if any face has a score greater than 0.5, I'm tagging the video with prediction score equivalent to the mean prediction of that face. Else, its the lowest prediction of any face in the video.\nHowever, due to some reason, it seems that the LB loss is 0.73 which is worse than 0.5 prediction for every video.</p>\n\n<p>I think the model seems to be overfitting to faces. Or is there anything else I'm missing here? Is there a better way to handle for train/val segregation? Should I be taking lesser number of images per video? I know I haven't done image augmentation in this case, but I felt that since I have round 160K images in my x_train alone, it would be okay to not augment.</p>\n\n<p>Is it possible that I'm messing up my submission file? The order in which video prediction is ordered in my submission file is governed by the order in test_videos as in the below code, which I took from <a href=\"/humananalog\">@humananalog</a> inference kernel.</p>\n\n<p>test_dir = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"</p>\n\n<p>test_videos = sorted([x for x in os.listdir(test_dir) if x[-4:] == \".mp4\"])</p>",
  "messages": [
    {
      "id": 736524,
      "postDate": "2020-02-04T09:51:55.910Z",
      "content": "<p>Hello folks!</p>\n\n<p>I'm training my model on the entire dataset with each video sampled at 0.5 FPS leading to 5 frames per face in any video with image resolution 128x128. Also, I'm doing the entire thing using tf.keras if that helps.</p>\n\n<p>My training set is as follows:</p>\n\n<p>Train data: Randomly sample around 10k real video names. Sample 1 fake video for each of the real video sampled earlier. Build x_train and y_train by loading all the images from the selected set of real and fake videos and then shuffling the x_train, y_train (index is handled) one last time.</p>\n\n<p>Validation data: Randomly sample 2K real videos from the remaining pool of unsampled data for train. Do the entire charade for fake video similar to earlier. </p>\n\n<p>This set of data is then mean centered and is being trained on a DenseNet like network with parameter count of approx 900k.</p>\n\n<p>After training the network for around 120 epochs, I get train loss of 0.15 and validation loss of 0.25.</p>\n\n<p>For the purposes of video classification, I'm running at 1.1FPS and am taking the mean prediction of each set of faces (I have a reliable clustering algorithm to track faces). And if any face has a score greater than 0.5, I'm tagging the video with prediction score equivalent to the mean prediction of that face. Else, its the lowest prediction of any face in the video.\nHowever, due to some reason, it seems that the LB loss is 0.73 which is worse than 0.5 prediction for every video.</p>\n\n<p>I think the model seems to be overfitting to faces. Or is there anything else I'm missing here? Is there a better way to handle for train/val segregation? Should I be taking lesser number of images per video? I know I haven't done image augmentation in this case, but I felt that since I have round 160K images in my x_train alone, it would be okay to not augment.</p>\n\n<p>Is it possible that I'm messing up my submission file? The order in which video prediction is ordered in my submission file is governed by the order in test_videos as in the below code, which I took from <a href=\"/humananalog\">@humananalog</a> inference kernel.</p>\n\n<p>test_dir = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"</p>\n\n<p>test_videos = sorted([x for x in os.listdir(test_dir) if x[-4:] == \".mp4\"])</p>",
      "rawMarkdown": "Hello folks!\n\nI'm training my model on the entire dataset with each video sampled at 0.5 FPS leading to 5 frames per face in any video with image resolution 128x128. Also, I'm doing the entire thing using tf.keras if that helps.\n\nMy training set is as follows:\n\nTrain data: Randomly sample around 10k real video names. Sample 1 fake video for each of the real video sampled earlier. Build x_train and y_train by loading all the images from the selected set of real and fake videos and then shuffling the x_train, y_train (index is handled) one last time.\n\nValidation data: Randomly sample 2K real videos from the remaining pool of unsampled data for train. Do the entire charade for fake video similar to earlier. \n\nThis set of data is then mean centered and is being trained on a DenseNet like network with parameter count of approx 900k.\n\nAfter training the network for around 120 epochs, I get train loss of 0.15 and validation loss of 0.25.\n\nFor the purposes of video classification, I'm running at 1.1FPS and am taking the mean prediction of each set of faces (I have a reliable clustering algorithm to track faces). And if any face has a score greater than 0.5, I'm tagging the video with prediction score equivalent to the mean prediction of that face. Else, its the lowest prediction of any face in the video.\nHowever, due to some reason, it seems that the LB loss is 0.73 which is worse than 0.5 prediction for every video.\n\nI think the model seems to be overfitting to faces. Or is there anything else I'm missing here? Is there a better way to handle for train/val segregation? Should I be taking lesser number of images per video? I know I haven't done image augmentation in this case, but I felt that since I have round 160K images in my x_train alone, it would be okay to not augment.\n\nIs it possible that I'm messing up my submission file? The order in which video prediction is ordered in my submission file is governed by the order in test_videos as in the below code, which I took from @humananalog inference kernel.\n\ntest_dir = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"\n\ntest_videos = sorted([x for x in os.listdir(test_dir) if x[-4:] == \".mp4\"])",
      "votes": 17
    },
    {
      "id": 743949,
      "postDate": "2020-02-12T12:40:01.613Z",
      "content": "<p>I just want to share this from my recent two good submissions that helped me to cross the bubble and enter sub 0.4!</p>\n\n<p>Submission A:  train BCE: 0.1551, val BCE: 0.3817, <code>LB: 0.42658</code>\nSubmission B: train BCE: 0.2273, val BCE: 0.3542, <code>LB: 0.38849</code> (added some new augmentations)</p>\n\n<p>1) Clearly, although my training loss is worse in the second case, I have improved my validation score =&gt; improved LB score. \nMy validation BCE is tracking LB very well. So, what's my validation set? Just data in 41-50 folders.</p>\n\n<p>2) Same single model as <a href=\"/humananalog\">@humananalog</a>'s public kernel, just with a dropout. Inference same with <a href=\"/humananalog\">@humananalog</a> public inference kernel.</p>\n\n<p>3) It's all about GOOD DATA + MORE DATA.</p>",
      "rawMarkdown": "I just want to share this from my recent two good submissions that helped me to cross the bubble and enter sub 0.4!\n\nSubmission A:  train BCE: 0.1551, val BCE: 0.3817, ```LB: 0.42658```\nSubmission B: train BCE: 0.2273, val BCE: 0.3542, ```LB: 0.38849``` (added some new augmentations)\n\n1) Clearly, although my training loss is worse in the second case, I have improved my validation score =&gt; improved LB score. \nMy validation BCE is tracking LB very well. So, what's my validation set? Just data in 41-50 folders.\n\n2) Same single model as @humananalog's public kernel, just with a dropout. Inference same with @humananalog public inference kernel.\n\n3) It's all about GOOD DATA + MORE DATA.\n\n",
      "votes": 12,
      "replies": [
        {
          "id": 744016,
          "postDate": "2020-02-12T13:32:20.370Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 744029,
          "postDate": "2020-02-12T13:46:04.787Z",
          "content": "<p>Hi <a href=\"/jpandas\">@jpandas</a>, I am using ALL the data (face crops) from <code>dfdc_train_part_40</code> to <code>dfdc_train_part_49</code>. Actually, I have 3 cv splits, and this is one of them. So far I am only using 1 fold, as training takes lot of time.</p>",
          "rawMarkdown": "Hi @jpandas, I am using ALL the data (face crops) from ```dfdc_train_part_40``` to ```dfdc_train_part_49```. Actually, I have 3 cv splits, and this is one of them. So far I am only using 1 fold, as training takes lot of time.",
          "votes": 1
        },
        {
          "id": 744031,
          "postDate": "2020-02-12T13:48:42.387Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 744034,
          "postDate": "2020-02-12T13:50:37.910Z",
          "content": "<p><a href=\"/jpandas\">@jpandas</a> welcome :))</p>",
          "rawMarkdown": "@jpandas welcome :))"
        },
        {
          "id": 744044,
          "postDate": "2020-02-12T14:05:48.410Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 744084,
          "postDate": "2020-02-12T14:45:57.017Z",
          "content": "<p><a href=\"/jpandas\">@jpandas</a> I am using ~120K real, 120K fake. So, out of these images from folder 41-50 are in my validation set. Maybe, I will increase number of images to see what happens to val BCE and LB. Use real/fake pairs from the same original video in the batches! To be honest, I am adding features so fast (and it is working), I should slow down to figure out what impacted most :D I need to be more scientific in planning the experiments, but again time is running out!!</p>",
          "rawMarkdown": "@jpandas I am using ~120K real, 120K fake. So, out of these images from folder 41-50 are in my validation set. Maybe, I will increase number of images to see what happens to val BCE and LB. Use real/fake pairs from the same original video in the batches! To be honest, I am adding features so fast (and it is working), I should slow down to figure out what impacted most :D I need to be more scientific in planning the experiments, but again time is running out!!",
          "votes": 3
        },
        {
          "id": 744094,
          "postDate": "2020-02-12T14:52:29.737Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 744105,
          "postDate": "2020-02-12T15:04:11.553Z",
          "content": "<p><a href=\"/jpandas\">@jpandas</a> You too ^^</p>",
          "rawMarkdown": "@jpandas You too ^^"
        },
        {
          "id": 744423,
          "postDate": "2020-02-12T20:30:44.593Z",
          "content": "<p>Hey! Thanks for piping in here. I too am currently using 40 random folders for train and the remaining 10 for val. In my current models, I seem to hit 0.33-0.36 val loss within the 2-4th epoch and strong overfitting right after that. This is significantly different from my previous experience with training a model which generally takes roughly 200 epochs to get to a decent score. I personally think I/we are doing something significantly wrong because of which the model is overfitting so quickly, and I'm not just referring to the probability of having same faces in both train and val sets despite using different folders.</p>",
          "rawMarkdown": "Hey! Thanks for piping in here. I too am currently using 40 random folders for train and the remaining 10 for val. In my current models, I seem to hit 0.33-0.36 val loss within the 2-4th epoch and strong overfitting right after that. This is significantly different from my previous experience with training a model which generally takes roughly 200 epochs to get to a decent score. I personally think I/we are doing something significantly wrong because of which the model is overfitting so quickly, and I'm not just referring to the probability of having same faces in both train and val sets despite using different folders."
        },
        {
          "id": 744433,
          "postDate": "2020-02-12T20:37:11.140Z",
          "content": "<p>My current results are in 3-4 epochs, don't know why :D even with all my efforts so far.</p>",
          "rawMarkdown": "My current results are in 3-4 epochs, don't know why :D even with all my efforts so far.",
          "votes": 1
        },
        {
          "id": 744453,
          "postDate": "2020-02-12T21:02:18.970Z",
          "content": "<p>Yea man. This is just strange. And does it start overfitting for you from the next epoch onwards? It does so for me. As per my understanding, it basically means that the number of parameters in the model is just too high. To counter that, I tried using models with lesser number of parameters. I even tried the normal methods like dropouts and regularization, but unfortunately, their Val scores never even came close. Those models could simply not generalize. I'm extremely keen to find out what is going on in there. Cause the moment we do, we will significantly be able to bring down the Val loss. </p>",
          "rawMarkdown": "Yea man. This is just strange. And does it start overfitting for you from the next epoch onwards? It does so for me. As per my understanding, it basically means that the number of parameters in the model is just too high. To counter that, I tried using models with lesser number of parameters. I even tried the normal methods like dropouts and regularization, but unfortunately, their Val scores never even came close. Those models could simply not generalize. I'm extremely keen to find out what is going on in there. Cause the moment we do, we will significantly be able to bring down the Val loss. "
        },
        {
          "id": 744482,
          "postDate": "2020-02-12T21:42:29.253Z",
          "content": "<blockquote>\n  <p>And does it start overfitting for you from the next epoch onwards?</p>\n</blockquote>\n\n<p>Yes :D I have to figure that out soon.</p>",
          "rawMarkdown": "&gt; And does it start overfitting for you from the next epoch onwards?\n\nYes :D I have to figure that out soon."
        }
      ]
    },
    {
      "id": 736556,
      "postDate": "2020-02-04T10:42:14.767Z",
      "content": "<p>Thanks you for the detailed description. A few thoughts:</p>\n\n<ol>\n<li><code>0.25</code> seems to be a little bit too optimistic, probably a leaking validation scheme. There are a lot of real videos per person, so sampling by the whole folder is safer in that regard (a single actor is usually in the same folder). Additionally, have you cleaned duplicates?</li>\n<li>What is the validation score for the video?</li>\n<li>Taking <code>min</code> introduces significant bias for the real videos, have you tried taking <code>mean</code> as well?</li>\n</ol>",
      "rawMarkdown": "Thanks you for the detailed description. A few thoughts:\n\n1. `0.25` seems to be a little bit too optimistic, probably a leaking validation scheme. There are a lot of real videos per person, so sampling by the whole folder is safer in that regard (a single actor is usually in the same folder). Additionally, have you cleaned duplicates?\n2. What is the validation score for the video?\n3. Taking `min` introduces significant bias for the real videos, have you tried taking `mean` as well?",
      "votes": 8,
      "replies": [
        {
          "id": 736640,
          "postDate": "2020-02-04T12:20:59.920Z",
          "content": "<p>I would second nosound, validation leak. As you take multiple images from same video, shuffle and split, you will end up with frames from the same video in both train and validation. I would recommend splitting the video names list and then generating the data frames. As nosound said, best to use images from another folder for validation to reduce the model ability to recognize the actors real face. </p>",
          "rawMarkdown": "I would second nosound, validation leak. As you take multiple images from same video, shuffle and split, you will end up with frames from the same video in both train and validation. I would recommend splitting the video names list and then generating the data frames. As nosound said, best to use images from another folder for validation to reduce the model ability to recognize the actors real face. ",
          "votes": 3
        },
        {
          "id": 736671,
          "postDate": "2020-02-04T12:53:14.273Z",
          "content": "<p>Actually, I have made sure that there are different \"videos\" in the train and val split. But yes, I have not taken care of the scenario where the same actor might actually end up in both the train and val. I shall try it out.</p>\n\n<p>Also <a href=\"/zaharch\">@zaharch</a> , I ran the model on the videos in the train sample, and I got .39 log loss with an 82% accuracy on classification of the video as a whole. With respect to your point 3, I understand the bias it creates. I haven't tried to take a mean across different faces. I guess I'll try that out.</p>",
          "rawMarkdown": "Actually, I have made sure that there are different \"videos\" in the train and val split. But yes, I have not taken care of the scenario where the same actor might actually end up in both the train and val. I shall try it out.\n\nAlso @zaharch , I ran the model on the videos in the train sample, and I got .39 log loss with an 82% accuracy on classification of the video as a whole. With respect to your point 3, I understand the bias it creates. I haven't tried to take a mean across different faces. I guess I'll try that out.",
          "votes": 2
        },
        {
          "id": 736680,
          "postDate": "2020-02-04T13:01:47.147Z",
          "content": "<p>Note that the train sample (also the public validation) is a subset of the full train, make sure you don't train on it if you validate on it.</p>",
          "rawMarkdown": "Note that the train sample (also the public validation) is a subset of the full train, make sure you don't train on it if you validate on it."
        },
        {
          "id": 736684,
          "postDate": "2020-02-04T13:06:23.653Z",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> Very good observations, but, the truth is, the model does overfits on faces, when you validate with the same actors our validation score was even &lt; 0.01 at some point, however our best lb model have the same actors validation score of 0.09..., We are not going to reveal our no-same actors scores yet...</p>",
          "rawMarkdown": "@zaharch Very good observations, but, the truth is, the model does overfits on faces, when you validate with the same actors our validation score was even &lt; 0.01 at some point, however our best lb model have the same actors validation score of 0.09..., We are not going to reveal our no-same actors scores yet..."
        },
        {
          "id": 736686,
          "postDate": "2020-02-04T13:08:48.783Z",
          "content": "<p><a href=\"/akashnandi\">@akashnandi</a> Honestly, Meaning all the faces scores whatever that is and clipping is the best method till the time you are not planning any complications...</p>",
          "rawMarkdown": "@akashnandi Honestly, Meaning all the faces scores whatever that is and clipping is the best method till the time you are not planning any complications..."
        },
        {
          "id": 736790,
          "postDate": "2020-02-04T14:56:04.380Z",
          "content": "<p>Right right. I'm going to try it out. Thanks <a href=\"/harshitsheoran\">@harshitsheoran</a> </p>",
          "rawMarkdown": "Right right. I'm going to try it out. Thanks @harshitsheoran "
        }
      ]
    },
    {
      "id": 736646,
      "postDate": "2020-02-04T12:26:29.393Z",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> gave some good advice above.</p>\n\n<p>I would add the following:\n-  try starting with a smaller amount of images, even with 1 frame per video\n-  pick a good pretrained network as base, if you are following <a href=\"/humananalog\">@humananalog</a>'s kernel, resnext sounds like a good one, if you are using keras, then I would recommend resnet\n-  start with a simple optimizer like SGD\n- if you use augmentation, start with simple operations</p>\n\n<p>basically, start with small dataset, use simple network, and then iterate </p>",
      "rawMarkdown": "@zaharch gave some good advice above.\n\nI would add the following:\n-  try starting with a smaller amount of images, even with 1 frame per video\n-  pick a good pretrained network as base, if you are following @humananalog's kernel, resnext sounds like a good one, if you are using keras, then I would recommend resnet\n-  start with a simple optimizer like SGD\n- if you use augmentation, start with simple operations\n\nbasically, start with small dataset, use simple network, and then iterate \n",
      "votes": 4,
      "replies": [
        {
          "id": 736672,
          "postDate": "2020-02-04T12:54:12.827Z",
          "content": "<p>Yes. I'll start off with 1 frame and augment it. And yes, currently, I'm using SGD with nestrov. Let me give a try with the pre trained network once.</p>",
          "rawMarkdown": "Yes. I'll start off with 1 frame and augment it. And yes, currently, I'm using SGD with nestrov. Let me give a try with the pre trained network once."
        },
        {
          "id": 736769,
          "postDate": "2020-02-04T14:33:20.800Z",
          "content": "<p>Why 128? For me, there was a big difference when I increased the face resolution to 256. </p>",
          "rawMarkdown": "Why 128? For me, there was a big difference when I increased the face resolution to 256. "
        },
        {
          "id": 736799,
          "postDate": "2020-02-04T15:19:27.100Z",
          "content": "<p>Thats because it'd take forever for me to actually train the model if the resolution is 256. I'm working with a RTX 2060 in my laptop. Though I could try FP16.</p>",
          "rawMarkdown": "Thats because it'd take forever for me to actually train the model if the resolution is 256. I'm working with a RTX 2060 in my laptop. Though I could try FP16."
        },
        {
          "id": 744471,
          "postDate": "2020-02-12T21:23:24.963Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 743661,
      "postDate": "2020-02-12T07:04:17.173Z",
      "content": "<p>Thanks for all the advice in the thread. I'm seeing really bad submission results compared to my local testing. I'm using the entire extra dataset, removing dupilcates, shuffling the results and splitting the dataframe using train/test split to ensure that fake and real videos by file name are kept separate. Then undersampling the fakes to match the reals in each split. I get absurdly good results after a long time training locally, and horrible submission scores. Locally 0.042 train, 0.232 validation, submission of 0.9+. I've also tried epochs that were much less over fit at 0.406/0.413 and the same results. Definitely taking an appropriate amount of time to run too.</p>\n\n<p>Based on the discussion here, I am using too small of a resolution for the faces. I also have not tried separating out a folder for validation instead of a random shuffle. </p>",
      "rawMarkdown": "Thanks for all the advice in the thread. I'm seeing really bad submission results compared to my local testing. I'm using the entire extra dataset, removing dupilcates, shuffling the results and splitting the dataframe using train/test split to ensure that fake and real videos by file name are kept separate. Then undersampling the fakes to match the reals in each split. I get absurdly good results after a long time training locally, and horrible submission scores. Locally 0.042 train, 0.232 validation, submission of 0.9+. I've also tried epochs that were much less over fit at 0.406/0.413 and the same results. Definitely taking an appropriate amount of time to run too.\n\nBased on the discussion here, I am using too small of a resolution for the faces. I also have not tried separating out a folder for validation instead of a random shuffle. ",
      "replies": [
        {
          "id": 743674,
          "postDate": "2020-02-12T07:21:49.877Z",
          "content": "<p>the most important thing is to keep actors from appearing both in your train and validation. this is not trivial to do. if same actors appears in both, you model will overfit to recognise their \"real\" face. However, this will not help you with the hidden test set that has a different set of actors. </p>",
          "rawMarkdown": "the most important thing is to keep actors from appearing both in your train and validation. this is not trivial to do. if same actors appears in both, you model will overfit to recognise their \"real\" face. However, this will not help you with the hidden test set that has a different set of actors. "
        },
        {
          "id": 743776,
          "postDate": "2020-02-12T09:49:39.183Z",
          "content": "<p>If you choose all <code>0.5</code> you get score <code>0.69314</code> so your submission score is way too high. It is not overfitting, it is more like you submit <code>1-p</code> instead of <code>p</code> or some other bug. At the end of your kernel, after running on 400 validation videos, try to print the metrics. If there is a bad score too, you can debug on those videos.</p>",
          "rawMarkdown": "If you choose all `0.5` you get score `0.69314` so your submission score is way too high. It is not overfitting, it is more like you submit `1-p` instead of `p` or some other bug. At the end of your kernel, after running on 400 validation videos, try to print the metrics. If there is a bad score too, you can debug on those videos.",
          "votes": 1
        },
        {
          "id": 744430,
          "postDate": "2020-02-12T20:35:16.567Z",
          "content": "<p>Try clipping the values to [0.1, 0.9] or [0.2, 0.8]. Or maybe just take a ROC-AUC and figure out the best thresholding. Ideally you want to do the ROC-AUC over multiple val splits.</p>",
          "rawMarkdown": "Try clipping the values to [0.1, 0.9] or [0.2, 0.8]. Or maybe just take a ROC-AUC and figure out the best thresholding. Ideally you want to do the ROC-AUC over multiple val splits."
        },
        {
          "id": 744485,
          "postDate": "2020-02-12T21:50:34.900Z",
          "content": "<p>Thanks, adding these things to the list of things to try. I've been only using 96x96 sized faces so that seems like it is way off from what others are using. </p>\n\n<p>I am also going over my submission face generator with a fine toothed comb to ensure it is generating the same as the training generator. I noticed an edge case where it potentially could return all zeros instead of face images. Visually checking showed that this was infrequent but it definitely could be an issue.</p>\n\n<p>The network I'm using is not like any I've seen used before so that may be the cause. I've put together several networks before though without issues like this.  In the end it is still a single node with sigmoid activation and keras bce loss. 1 for fake, 0 for real. Using MTCNN to grab faces, handling exceptions and errors properly etc.</p>",
          "rawMarkdown": "Thanks, adding these things to the list of things to try. I've been only using 96x96 sized faces so that seems like it is way off from what others are using. \n\nI am also going over my submission face generator with a fine toothed comb to ensure it is generating the same as the training generator. I noticed an edge case where it potentially could return all zeros instead of face images. Visually checking showed that this was infrequent but it definitely could be an issue.\n\nThe network I'm using is not like any I've seen used before so that may be the cause. I've put together several networks before though without issues like this.  In the end it is still a single node with sigmoid activation and keras bce loss. 1 for fake, 0 for real. Using MTCNN to grab faces, handling exceptions and errors properly etc."
        },
        {
          "id": 747772,
          "postDate": "2020-02-16T20:36:20.103Z",
          "content": "<p>The results are in. The changes I made \"improved\" things quite a bit. I am no longer seeing good validation results. I now get overfitting as others have described. I'm also only getting validation results of .45 or so at best. Working on bringing that down.</p>\n\n<p>After this I found the network not fully utilizing the GPU with 256x256 faces. It also significantly increased a memory leak issue in tensorflow 2.1. I could only get 7 epochs before running out of memory. Prior to 2.1 I would create my generators by subclassing a keras sequence. In 2.1 model.fit_generator shows a deprecation warning with reasonable performance but has a slow memory leak. Using a keras sequence with mode.fit in 2.1 also does not work, it triggers an exception that the class is not callable. Passing the <strong>iter</strong> method works as it is a callable, but leaks memory very quickly.</p>\n\n<p>To get around this I rewrote the generator to be a tf.data.dataset. Following the better performance guide, GPU performance is again the bottleneck and training is going much faster. No memory leak at all anymore. The GPU is once again heating up the room... </p>\n\n<p><a href=\"https://www.tensorflow.org/guide/data_performance\">https://www.tensorflow.org/guide/data_performance</a></p>",
          "rawMarkdown": "The results are in. The changes I made \"improved\" things quite a bit. I am no longer seeing good validation results. I now get overfitting as others have described. I'm also only getting validation results of .45 or so at best. Working on bringing that down.\n\nAfter this I found the network not fully utilizing the GPU with 256x256 faces. It also significantly increased a memory leak issue in tensorflow 2.1. I could only get 7 epochs before running out of memory. Prior to 2.1 I would create my generators by subclassing a keras sequence. In 2.1 model.fit_generator shows a deprecation warning with reasonable performance but has a slow memory leak. Using a keras sequence with mode.fit in 2.1 also does not work, it triggers an exception that the class is not callable. Passing the __iter__ method works as it is a callable, but leaks memory very quickly.\n\nTo get around this I rewrote the generator to be a tf.data.dataset. Following the better performance guide, GPU performance is again the bottleneck and training is going much faster. No memory leak at all anymore. The GPU is once again heating up the room... \n\n[https://www.tensorflow.org/guide/data_performance](https://www.tensorflow.org/guide/data_performance)",
          "votes": 3
        }
      ]
    },
    {
      "id": 744469,
      "postDate": "2020-02-12T21:22:48.603Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 743949,
      "author_name": "Debanga Raj Neog",
      "author_url": "",
      "post_date": "2020-02-12T12:40:01.613000",
      "content": "<p>I just want to share this from my recent two good submissions that helped me to cross the bubble and enter sub 0.4!</p>\n\n<p>Submission A:  train BCE: 0.1551, val BCE: 0.3817, <code>LB: 0.42658</code>\nSubmission B: train BCE: 0.2273, val BCE: 0.3542, <code>LB: 0.38849</code> (added some new augmentations)</p>\n\n<p>1) Clearly, although my training loss is worse in the second case, I have improved my validation score =&gt; improved LB score. \nMy validation BCE is tracking LB very well. So, what's my validation set? Just data in 41-50 folders.</p>\n\n<p>2) Same single model as <a href=\"/humananalog\">@humananalog</a>'s public kernel, just with a dropout. Inference same with <a href=\"/humananalog\">@humananalog</a> public inference kernel.</p>\n\n<p>3) It's all about GOOD DATA + MORE DATA.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 744016,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-12T13:32:20.370000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744029,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-12T13:46:04.787000",
          "content": "<p>Hi <a href=\"/jpandas\">@jpandas</a>, I am using ALL the data (face crops) from <code>dfdc_train_part_40</code> to <code>dfdc_train_part_49</code>. Actually, I have 3 cv splits, and this is one of them. So far I am only using 1 fold, as training takes lot of time.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744031,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-12T13:48:42.387000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744034,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-12T13:50:37.910000",
          "content": "<p><a href=\"/jpandas\">@jpandas</a> welcome :))</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744044,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-12T14:05:48.410000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744084,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-12T14:45:57.017000",
          "content": "<p><a href=\"/jpandas\">@jpandas</a> I am using ~120K real, 120K fake. So, out of these images from folder 41-50 are in my validation set. Maybe, I will increase number of images to see what happens to val BCE and LB. Use real/fake pairs from the same original video in the batches! To be honest, I am adding features so fast (and it is working), I should slow down to figure out what impacted most :D I need to be more scientific in planning the experiments, but again time is running out!!</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 744094,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-02-12T14:52:29.737000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744105,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-12T15:04:11.553000",
          "content": "<p><a href=\"/jpandas\">@jpandas</a> You too ^^</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744423,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-12T20:30:44.593000",
          "content": "<p>Hey! Thanks for piping in here. I too am currently using 40 random folders for train and the remaining 10 for val. In my current models, I seem to hit 0.33-0.36 val loss within the 2-4th epoch and strong overfitting right after that. This is significantly different from my previous experience with training a model which generally takes roughly 200 epochs to get to a decent score. I personally think I/we are doing something significantly wrong because of which the model is overfitting so quickly, and I'm not just referring to the probability of having same faces in both train and val sets despite using different folders.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744433,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-12T20:37:11.140000",
          "content": "<p>My current results are in 3-4 epochs, don't know why :D even with all my efforts so far.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744453,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-12T21:02:18.970000",
          "content": "<p>Yea man. This is just strange. And does it start overfitting for you from the next epoch onwards? It does so for me. As per my understanding, it basically means that the number of parameters in the model is just too high. To counter that, I tried using models with lesser number of parameters. I even tried the normal methods like dropouts and regularization, but unfortunately, their Val scores never even came close. Those models could simply not generalize. I'm extremely keen to find out what is going on in there. Cause the moment we do, we will significantly be able to bring down the Val loss. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744482,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-12T21:42:29.253000",
          "content": "<blockquote>\n  <p>And does it start overfitting for you from the next epoch onwards?</p>\n</blockquote>\n\n<p>Yes :D I have to figure that out soon.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 736556,
      "author_name": "nosound",
      "author_url": "",
      "post_date": "2020-02-04T10:42:14.767000",
      "content": "<p>Thanks you for the detailed description. A few thoughts:</p>\n\n<ol>\n<li><code>0.25</code> seems to be a little bit too optimistic, probably a leaking validation scheme. There are a lot of real videos per person, so sampling by the whole folder is safer in that regard (a single actor is usually in the same folder). Additionally, have you cleaned duplicates?</li>\n<li>What is the validation score for the video?</li>\n<li>Taking <code>min</code> introduces significant bias for the real videos, have you tried taking <code>mean</code> as well?</li>\n</ol>",
      "votes": 8,
      "replies": [
        {
          "id": 736640,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-02-04T12:20:59.920000",
          "content": "<p>I would second nosound, validation leak. As you take multiple images from same video, shuffle and split, you will end up with frames from the same video in both train and validation. I would recommend splitting the video names list and then generating the data frames. As nosound said, best to use images from another folder for validation to reduce the model ability to recognize the actors real face. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 736671,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-04T12:53:14.273000",
          "content": "<p>Actually, I have made sure that there are different \"videos\" in the train and val split. But yes, I have not taken care of the scenario where the same actor might actually end up in both the train and val. I shall try it out.</p>\n\n<p>Also <a href=\"/zaharch\">@zaharch</a> , I ran the model on the videos in the train sample, and I got .39 log loss with an 82% accuracy on classification of the video as a whole. With respect to your point 3, I understand the bias it creates. I haven't tried to take a mean across different faces. I guess I'll try that out.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 736680,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-04T13:01:47.147000",
          "content": "<p>Note that the train sample (also the public validation) is a subset of the full train, make sure you don't train on it if you validate on it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 736684,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-04T13:06:23.653000",
          "content": "<p><a href=\"/zaharch\">@zaharch</a> Very good observations, but, the truth is, the model does overfits on faces, when you validate with the same actors our validation score was even &lt; 0.01 at some point, however our best lb model have the same actors validation score of 0.09..., We are not going to reveal our no-same actors scores yet...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 736686,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2020-02-04T13:08:48.783000",
          "content": "<p><a href=\"/akashnandi\">@akashnandi</a> Honestly, Meaning all the faces scores whatever that is and clipping is the best method till the time you are not planning any complications...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 736790,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-04T14:56:04.380000",
          "content": "<p>Right right. I'm going to try it out. Thanks <a href=\"/harshitsheoran\">@harshitsheoran</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 736646,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-02-04T12:26:29.393000",
      "content": "<p><a href=\"/zaharch\">@zaharch</a> gave some good advice above.</p>\n\n<p>I would add the following:\n-  try starting with a smaller amount of images, even with 1 frame per video\n-  pick a good pretrained network as base, if you are following <a href=\"/humananalog\">@humananalog</a>'s kernel, resnext sounds like a good one, if you are using keras, then I would recommend resnet\n-  start with a simple optimizer like SGD\n- if you use augmentation, start with simple operations</p>\n\n<p>basically, start with small dataset, use simple network, and then iterate </p>",
      "votes": 4,
      "replies": [
        {
          "id": 736672,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-04T12:54:12.827000",
          "content": "<p>Yes. I'll start off with 1 frame and augment it. And yes, currently, I'm using SGD with nestrov. Let me give a try with the pre trained network once.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 736769,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2020-02-04T14:33:20.800000",
          "content": "<p>Why 128? For me, there was a big difference when I increased the face resolution to 256. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 736799,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-04T15:19:27.100000",
          "content": "<p>Thats because it'd take forever for me to actually train the model if the resolution is 256. I'm working with a RTX 2060 in my laptop. Though I could try FP16.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744471,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-02-12T21:23:24.963000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 743661,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2020-02-12T07:04:17.173000",
      "content": "<p>Thanks for all the advice in the thread. I'm seeing really bad submission results compared to my local testing. I'm using the entire extra dataset, removing dupilcates, shuffling the results and splitting the dataframe using train/test split to ensure that fake and real videos by file name are kept separate. Then undersampling the fakes to match the reals in each split. I get absurdly good results after a long time training locally, and horrible submission scores. Locally 0.042 train, 0.232 validation, submission of 0.9+. I've also tried epochs that were much less over fit at 0.406/0.413 and the same results. Definitely taking an appropriate amount of time to run too.</p>\n\n<p>Based on the discussion here, I am using too small of a resolution for the faces. I also have not tried separating out a folder for validation instead of a random shuffle. </p>",
      "votes": 0,
      "replies": [
        {
          "id": 743674,
          "author_name": "Moshel",
          "author_url": "",
          "post_date": "2020-02-12T07:21:49.877000",
          "content": "<p>the most important thing is to keep actors from appearing both in your train and validation. this is not trivial to do. if same actors appears in both, you model will overfit to recognise their \"real\" face. However, this will not help you with the hidden test set that has a different set of actors. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743776,
          "author_name": "nosound",
          "author_url": "",
          "post_date": "2020-02-12T09:49:39.183000",
          "content": "<p>If you choose all <code>0.5</code> you get score <code>0.69314</code> so your submission score is way too high. It is not overfitting, it is more like you submit <code>1-p</code> instead of <code>p</code> or some other bug. At the end of your kernel, after running on 400 validation videos, try to print the metrics. If there is a bad score too, you can debug on those videos.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 744430,
          "author_name": "Akash",
          "author_url": "",
          "post_date": "2020-02-12T20:35:16.567000",
          "content": "<p>Try clipping the values to [0.1, 0.9] or [0.2, 0.8]. Or maybe just take a ROC-AUC and figure out the best thresholding. Ideally you want to do the ROC-AUC over multiple val splits.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 744485,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2020-02-12T21:50:34.900000",
          "content": "<p>Thanks, adding these things to the list of things to try. I've been only using 96x96 sized faces so that seems like it is way off from what others are using. </p>\n\n<p>I am also going over my submission face generator with a fine toothed comb to ensure it is generating the same as the training generator. I noticed an edge case where it potentially could return all zeros instead of face images. Visually checking showed that this was infrequent but it definitely could be an issue.</p>\n\n<p>The network I'm using is not like any I've seen used before so that may be the cause. I've put together several networks before though without issues like this.  In the end it is still a single node with sigmoid activation and keras bce loss. 1 for fake, 0 for real. Using MTCNN to grab faces, handling exceptions and errors properly etc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 747772,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2020-02-16T20:36:20.103000",
          "content": "<p>The results are in. The changes I made \"improved\" things quite a bit. I am no longer seeing good validation results. I now get overfitting as others have described. I'm also only getting validation results of .45 or so at best. Working on bringing that down.</p>\n\n<p>After this I found the network not fully utilizing the GPU with 256x256 faces. It also significantly increased a memory leak issue in tensorflow 2.1. I could only get 7 epochs before running out of memory. Prior to 2.1 I would create my generators by subclassing a keras sequence. In 2.1 model.fit_generator shows a deprecation warning with reasonable performance but has a slow memory leak. Using a keras sequence with mode.fit in 2.1 also does not work, it triggers an exception that the class is not callable. Passing the <strong>iter</strong> method works as it is a callable, but leaks memory very quickly.</p>\n\n<p>To get around this I rewrote the generator to be a tf.data.dataset. Following the better performance guide, GPU performance is again the bottleneck and training is going much faster. No memory leak at all anymore. The GPU is once again heating up the room... </p>\n\n<p><a href=\"https://www.tensorflow.org/guide/data_performance\">https://www.tensorflow.org/guide/data_performance</a></p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 744469,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-12T21:22:48.603000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "736524": "Hello folks!\n\nI'm training my model on the entire dataset with each video sampled at 0.5 FPS leading to 5 frames per face in any video with image resolution 128x128. Also, I'm doing the entire thing using tf.keras if that helps.\n\nMy training set is as follows:\n\nTrain data: Randomly sample around 10k real video names. Sample 1 fake video for each of the real video sampled earlier. Build x_train and y_train by loading all the images from the selected set of real and fake videos and then shuffling the x_train, y_train (index is handled) one last time.\n\nValidation data: Randomly sample 2K real videos from the remaining pool of unsampled data for train. Do the entire charade for fake video similar to earlier. \n\nThis set of data is then mean centered and is being trained on a DenseNet like network with parameter count of approx 900k.\n\nAfter training the network for around 120 epochs, I get train loss of 0.15 and validation loss of 0.25.\n\nFor the purposes of video classification, I'm running at 1.1FPS and am taking the mean prediction of each set of faces (I have a reliable clustering algorithm to track faces). And if any face has a score greater than 0.5, I'm tagging the video with prediction score equivalent to the mean prediction of that face. Else, its the lowest prediction of any face in the video.\nHowever, due to some reason, it seems that the LB loss is 0.73 which is worse than 0.5 prediction for every video.\n\nI think the model seems to be overfitting to faces. Or is there anything else I'm missing here? Is there a better way to handle for train/val segregation? Should I be taking lesser number of images per video? I know I haven't done image augmentation in this case, but I felt that since I have round 160K images in my x_train alone, it would be okay to not augment.\n\nIs it possible that I'm messing up my submission file? The order in which video prediction is ordered in my submission file is governed by the order in test_videos as in the below code, which I took from @humananalog inference kernel.\n\ntest_dir = \"/kaggle/input/deepfake-detection-challenge/test_videos/\"\n\ntest_videos = sorted([x for x in os.listdir(test_dir) if x[-4:] == \".mp4\"])",
    "743949": "I just want to share this from my recent two good submissions that helped me to cross the bubble and enter sub 0.4!\n\nSubmission A:  train BCE: 0.1551, val BCE: 0.3817, ```LB: 0.42658```\nSubmission B: train BCE: 0.2273, val BCE: 0.3542, ```LB: 0.38849``` (added some new augmentations)\n\n1) Clearly, although my training loss is worse in the second case, I have improved my validation score =&gt; improved LB score. \nMy validation BCE is tracking LB very well. So, what's my validation set? Just data in 41-50 folders.\n\n2) Same single model as @humananalog's public kernel, just with a dropout. Inference same with @humananalog public inference kernel.\n\n3) It's all about GOOD DATA + MORE DATA.\n\n",
    "736556": "Thanks you for the detailed description. A few thoughts:\n\n1. `0.25` seems to be a little bit too optimistic, probably a leaking validation scheme. There are a lot of real videos per person, so sampling by the whole folder is safer in that regard (a single actor is usually in the same folder). Additionally, have you cleaned duplicates?\n2. What is the validation score for the video?\n3. Taking `min` introduces significant bias for the real videos, have you tried taking `mean` as well?",
    "736646": "@zaharch gave some good advice above.\n\nI would add the following:\n-  try starting with a smaller amount of images, even with 1 frame per video\n-  pick a good pretrained network as base, if you are following @humananalog's kernel, resnext sounds like a good one, if you are using keras, then I would recommend resnet\n-  start with a simple optimizer like SGD\n- if you use augmentation, start with simple operations\n\nbasically, start with small dataset, use simple network, and then iterate \n",
    "743661": "Thanks for all the advice in the thread. I'm seeing really bad submission results compared to my local testing. I'm using the entire extra dataset, removing dupilcates, shuffling the results and splitting the dataframe using train/test split to ensure that fake and real videos by file name are kept separate. Then undersampling the fakes to match the reals in each split. I get absurdly good results after a long time training locally, and horrible submission scores. Locally 0.042 train, 0.232 validation, submission of 0.9+. I've also tried epochs that were much less over fit at 0.406/0.413 and the same results. Definitely taking an appropriate amount of time to run too.\n\nBased on the discussion here, I am using too small of a resolution for the faces. I also have not tried separating out a folder for validation instead of a random shuffle. ",
    "744469": ""
  }
}