{
  "id": 124127,
  "title": "Validation strategies?",
  "url": "/competitions/deepfake-detection-challenge/discussion/124127",
  "author_name": "ryches",
  "post_date": "2020-01-02T04:54:43.425000",
  "votes": 27,
  "comment_count": 27,
  "views": 0,
  "content": "<p>I've seen a few people in the discussion mention a big discrepancy between local validation and LB. Some of the options that might be viable is a stratified hold out, group shuffle split that separates based on the source video so you don't have 3 fakes of a certain video in the training set and another 2 fakes of the same video in the validation set. Alternately a split can be done based on the folders provided to us. So then it would be a true holdout of an entirely unseen face and video set. My concern with the final one is that will likely be fragile based on the subject being filmed, we don't really have a ton of unique faces. </p>\n\n<p>Any thoughts on this? Are others seeing a large discrepancy between val and lb?</p>",
  "messages": [
    {
      "id": 708219,
      "postDate": "2020-01-02T04:54:43.427Z",
      "content": "<p>I've seen a few people in the discussion mention a big discrepancy between local validation and LB. Some of the options that might be viable is a stratified hold out, group shuffle split that separates based on the source video so you don't have 3 fakes of a certain video in the training set and another 2 fakes of the same video in the validation set. Alternately a split can be done based on the folders provided to us. So then it would be a true holdout of an entirely unseen face and video set. My concern with the final one is that will likely be fragile based on the subject being filmed, we don't really have a ton of unique faces. </p>\n\n<p>Any thoughts on this? Are others seeing a large discrepancy between val and lb?</p>",
      "rawMarkdown": "I've seen a few people in the discussion mention a big discrepancy between local validation and LB. Some of the options that might be viable is a stratified hold out, group shuffle split that separates based on the source video so you don't have 3 fakes of a certain video in the training set and another 2 fakes of the same video in the validation set. Alternately a split can be done based on the folders provided to us. So then it would be a true holdout of an entirely unseen face and video set. My concern with the final one is that will likely be fragile based on the subject being filmed, we don't really have a ton of unique faces. \n\nAny thoughts on this? Are others seeing a large discrepancy between val and lb?",
      "votes": 27
    },
    {
      "id": 709229,
      "postDate": "2020-01-03T07:52:58.677Z",
      "content": "<p>Different folders can contain the same actors. For example:\n<code>\ndfdc_train_part_44/abkclntljz.mp4\ndfdc_train_part_14/jvtgwixooc.mp4\ndfdc_train_part_47/mlwnlxchfy.mp4\ndfdc_train_part_35/iukwdhtsut.mp4\ndfdc_train_part_13/fhmjuyomqf.mp4\ndfdc_train_part_26/wlrgqwoqfx.mp4\n</code></p>",
      "rawMarkdown": "Different folders can contain the same actors. For example:\n```\ndfdc_train_part_44/abkclntljz.mp4\ndfdc_train_part_14/jvtgwixooc.mp4\ndfdc_train_part_47/mlwnlxchfy.mp4\ndfdc_train_part_35/iukwdhtsut.mp4\ndfdc_train_part_13/fhmjuyomqf.mp4\ndfdc_train_part_26/wlrgqwoqfx.mp4\n```",
      "votes": 14,
      "replies": [
        {
          "id": 714583,
          "postDate": "2020-01-09T15:15:08.703Z",
          "content": "<p>Hi Organizers,</p>\n\n<p>It makes hard for participants to use this dataset. Though I have tried some automated ways to identify faces, my conclusion is that it's impossible to make a reliable cross validation with this dataset. So, is there any plan to provide an information about actors for each videos?</p>",
          "rawMarkdown": "Hi Organizers,\n\nIt makes hard for participants to use this dataset. Though I have tried some automated ways to identify faces, my conclusion is that it's impossible to make a reliable cross validation with this dataset. So, is there any plan to provide an information about actors for each videos?",
          "votes": -1
        },
        {
          "id": 770419,
          "postDate": "2020-03-12T22:32:36.520Z",
          "content": "<p>May I know how do you find out video with same actor from different folder? </p>",
          "rawMarkdown": "May I know how do you find out video with same actor from different folder? "
        }
      ]
    },
    {
      "id": 709199,
      "postDate": "2020-01-03T06:45:27.570Z",
      "content": "<p>Just finished a 1st experiment with embeddings to group similar faces (hopefully from the same actor) automatically. Here's 3 cases, and below the code. The large image in the left is given; the small images in the right are selected by the algorithm as the most similar to the one in the left (universe = 800 sample videos provided). There's a lot to improve in the code, this is just a quick example to show how embeddings are cool :)</p>\n\n<p>Case 1\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fef9fa8a881d97538f9a7c68e73389023%2FScreen%20Shot%202020-01-03%20at%2003.34.37.png?generation=1578033677488156&amp;alt=media\" alt=\"\"></p>\n\n<p>Case 2\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F14d092eda70ccab8c0ac2ec2445aed92%2FScreen%20Shot%202020-01-03%20at%2003.31.01.png?generation=1578033761923474&amp;alt=media\" alt=\"\"></p>\n\n<p>Case 3\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fee9275d339905ec47b41793eccc5a6a4%2FScreen%20Shot%202020-01-03%20at%2003.31.57.png?generation=1578033783595516&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.kaggle.com/carlossouza/embeddings-grouping-similar-faces-automagically\">https://www.kaggle.com/carlossouza/embeddings-grouping-similar-faces-automagically</a></p>",
      "rawMarkdown": "Just finished a 1st experiment with embeddings to group similar faces (hopefully from the same actor) automatically. Here's 3 cases, and below the code. The large image in the left is given; the small images in the right are selected by the algorithm as the most similar to the one in the left (universe = 800 sample videos provided). There's a lot to improve in the code, this is just a quick example to show how embeddings are cool :)\n\nCase 1\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fef9fa8a881d97538f9a7c68e73389023%2FScreen%20Shot%202020-01-03%20at%2003.34.37.png?generation=1578033677488156&amp;alt=media)\n\nCase 2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F14d092eda70ccab8c0ac2ec2445aed92%2FScreen%20Shot%202020-01-03%20at%2003.31.01.png?generation=1578033761923474&amp;alt=media)\n\nCase 3\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fee9275d339905ec47b41793eccc5a6a4%2FScreen%20Shot%202020-01-03%20at%2003.31.57.png?generation=1578033783595516&amp;alt=media)\n\nhttps://www.kaggle.com/carlossouza/embeddings-grouping-similar-faces-automagically\n",
      "votes": 7
    },
    {
      "id": 710487,
      "postDate": "2020-01-04T19:47:44.590Z",
      "content": "<p>All the folders contain fakes and reals at a ratio of 5:1. But the actual public test set has approximately equal number of both. Will it be a good idea to use only videos from one folder which has this huge data imbalance? Even though the log loss is the metric, it can go very high in case of false fake (true positives) predictions with which this data imbalance will not help to validate.</p>\n\n<p>For folder <code>dfdc_train_part_33</code>,</p>\n\n<p>```python\nIn [1]: labels.count('FAKE')\nOut[1]: 1884</p>\n\n<p>In [2]: labels.count('REAL')\nOut[2]: 390\n```</p>",
      "rawMarkdown": "All the folders contain fakes and reals at a ratio of 5:1. But the actual public test set has approximately equal number of both. Will it be a good idea to use only videos from one folder which has this huge data imbalance? Even though the log loss is the metric, it can go very high in case of false fake (true positives) predictions with which this data imbalance will not help to validate.\n\nFor folder `dfdc_train_part_33`,\n\n```python\nIn [1]: labels.count('FAKE')\nOut[1]: 1884\n\nIn [2]: labels.count('REAL')\nOut[2]: 390\n```\n",
      "votes": 1,
      "replies": [
        {
          "id": 712214,
          "postDate": "2020-01-07T00:24:32.220Z",
          "content": "<p>Apply underbalancing technique would solve this problem. And you can use more than one folder(just make sure there's no leak).</p>",
          "rawMarkdown": "Apply underbalancing technique would solve this problem. And you can use more than one folder(just make sure there's no leak)."
        },
        {
          "id": 714268,
          "postDate": "2020-01-09T08:55:38.913Z",
          "content": "<p>Hi，I'm new to CV and DL. Could you share some underbalancing techniques with me? Does this  equal to 'sample'? Thx!</p>",
          "rawMarkdown": "Hi，I'm new to CV and DL. Could you share some underbalancing techniques with me? Does this  equal to 'sample'? Thx!"
        },
        {
          "id": 714283,
          "postDate": "2020-01-09T09:06:11.337Z",
          "content": "<p>I'm not an expert but there are two tricks I use and they seem to work. One is to repeat real videos and make batch with the real-fake pairs of same actor(known as overbalancing, thanks to <a href=\"/unkownhihi\">@unkownhihi</a> for pointing out ). and other is to change a little bit to loss function\nfor example -(2*y*log(ypred) + (1-y)*log(1-ypred)) so that classifying a real video as fake will have greater loss. </p>",
          "rawMarkdown": "I'm not an expert but there are two tricks I use and they seem to work. One is to repeat real videos and make batch with the real-fake pairs of same actor(known as overbalancing, thanks to @unkownhihi for pointing out ). and other is to change a little bit to loss function\nfor example -(2*y*log(ypred) + (1-y)*log(1-ypred)) so that classifying a real video as fake will have greater loss. "
        },
        {
          "id": 714420,
          "postDate": "2020-01-09T12:31:55.013Z",
          "content": "<p>thx!</p>",
          "rawMarkdown": "thx!"
        },
        {
          "id": 714974,
          "postDate": "2020-01-10T01:49:04.073Z",
          "content": "<p><a href=\"/ankitsainiankit\">@ankitsainiankit</a> The first one you said is overbalancing. The under balancing is to pick the same number of fake videos as real videos(in other words, delete some fake videos). </p>",
          "rawMarkdown": "@ankitsainiankit The first one you said is overbalancing. The under balancing is to pick the same number of fake videos as real videos(in other words, delete some fake videos). ",
          "votes": 1
        },
        {
          "id": 715099,
          "postDate": "2020-01-10T06:02:20.050Z",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> I'm sorry for misinformation. I'll edit it.</p>",
          "rawMarkdown": "@unkownhihi I'm sorry for misinformation. I'll edit it."
        },
        {
          "id": 785213,
          "postDate": "2020-03-24T22:08:37.730Z",
          "content": "<p>You can pick subfolders, e.g. 41-45. Then you can randomly pick 1 fake for each original. Do this multiple times (bootstrapping) calculate validation for each bootstrap and average them out to get a more robust estimate. </p>",
          "rawMarkdown": "You can pick subfolders, e.g. 41-45. Then you can randomly pick 1 fake for each original. Do this multiple times (bootstrapping) calculate validation for each bootstrap and average them out to get a more robust estimate. "
        }
      ]
    },
    {
      "id": 709985,
      "postDate": "2020-01-04T07:10:39.197Z",
      "content": "<p>I have done some local experiment with both validation by group, and validation by original videos.\nMy observation is that the gap between training and validation error is a lot wider when validating by original videos.\nFor me that is a clear indication of overfitting.</p>\n\n<p>As pointed out by <a href=\"/liuftvafas\">@liuftvafas</a>,  we do have the same actor in different groups.  nevertheless, base on the above evidence I would go with validation by groups until better idea is shared :)</p>",
      "rawMarkdown": "I have done some local experiment with both validation by group, and validation by original videos.\nMy observation is that the gap between training and validation error is a lot wider when validating by original videos.\nFor me that is a clear indication of overfitting.\n\nAs pointed out by @liuftvafas,  we do have the same actor in different groups.  nevertheless, base on the above evidence I would go with validation by groups until better idea is shared :)\n",
      "votes": 1,
      "replies": [
        {
          "id": 759228,
          "postDate": "2020-02-28T18:54:47.893Z",
          "content": "<p>Thanks for sharing! With validation by group do you mean by actors? Or what does it mean exactly? Thanks!</p>",
          "rawMarkdown": "Thanks for sharing! With validation by group do you mean by actors? Or what does it mean exactly? Thanks!"
        }
      ]
    },
    {
      "id": 709180,
      "postDate": "2020-01-03T06:04:10.040Z",
      "content": "<p>Yeah, I had large discrepancy between val score and lb score when I just did random split. Now I'm trying splitting with respect to source videos.\nOne thing to point out, as shown in <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013</a> , augmentations are applied to test set. Maybe we should simulate these augmentations to our valid set.</p>",
      "rawMarkdown": "Yeah, I had large discrepancy between val score and lb score when I just did random split. Now I'm trying splitting with respect to source videos.\nOne thing to point out, as shown in https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013 , augmentations are applied to test set. Maybe we should simulate these augmentations to our valid set.",
      "votes": 2
    },
    {
      "id": 708486,
      "postDate": "2020-01-02T10:49:41.793Z",
      "content": "<p>We can use a group of actors in training set, and a distinct group of actors in the validation set. This can be done automagically using embeddings. Will try to do a notebook in the next days to show how. And yes, I'm seeing a big discrepancy too (shown below). But at least for now, I'm focusing on other parts of the problem... validation will come later. Cheers!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fa0dfc470332d5200e9f0eb99a0a3e466%2Fistanbul-20200102.png?generation=1577968215664071&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "We can use a group of actors in training set, and a distinct group of actors in the validation set. This can be done automagically using embeddings. Will try to do a notebook in the next days to show how. And yes, I'm seeing a big discrepancy too (shown below). But at least for now, I'm focusing on other parts of the problem... validation will come later. Cheers!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fa0dfc470332d5200e9f0eb99a0a3e466%2Fistanbul-20200102.png?generation=1577968215664071&amp;alt=media)\n",
      "votes": 2,
      "replies": [
        {
          "id": 708712,
          "postDate": "2020-01-02T15:24:19.710Z",
          "content": "<p>Is it necessary to use embeddings? My understanding was that the folders already split the video to the different actors. Have not spent a ton of time verifying it myself but someone stated it and seemed to be true in my limited experience. </p>",
          "rawMarkdown": "Is it necessary to use embeddings? My understanding was that the folders already split the video to the different actors. Have not spent a ton of time verifying it myself but someone stated it and seemed to be true in my limited experience. ",
          "votes": 4
        },
        {
          "id": 708769,
          "postDate": "2020-01-02T16:59:07.820Z",
          "content": "<p>I've read that somewhere in the forum too... but will double check to make sure!</p>",
          "rawMarkdown": "I've read that somewhere in the forum too... but will double check to make sure!",
          "votes": 1
        },
        {
          "id": 708772,
          "postDate": "2020-01-02T17:03:11.733Z",
          "content": "<p>I'd be interested to see what you had in mind with the embeddings too. I get the concept but I'm not sure exactly how I'd formulate that in our video context with tons of videos and frames for each actor. </p>",
          "rawMarkdown": "I'd be interested to see what you had in mind with the embeddings too. I get the concept but I'm not sure exactly how I'd formulate that in our video context with tons of videos and frames for each actor. \n"
        },
        {
          "id": 708800,
          "postDate": "2020-01-02T17:44:29.300Z",
          "content": "<p>I'm almost sure this is the case, I think is unrealistic to actually check on that, but by checking a few and also considering the results of my experiments I do believe that the foIders also separate actors (they for sure separate original/fakes of a same video, this I've checked). I didn't quiet understand why you believe this strategy would be fragile, can you elaborate on that? I believe that doing a group split by folder is the best effortless way of creating a validation set. </p>",
          "rawMarkdown": "I'm almost sure this is the case, I think is unrealistic to actually check on that, but by checking a few and also considering the results of my experiments I do believe that the foIders also separate actors (they for sure separate original/fakes of a same video, this I've checked). I didn't quiet understand why you believe this strategy would be fragile, can you elaborate on that? I believe that doing a group split by folder is the best effortless way of creating a validation set. ",
          "votes": 1
        },
        {
          "id": 708843,
          "postDate": "2020-01-02T18:49:14.853Z",
          "content": "<p>Out of curiosity, how did you check the folders separate actors? \nEvery folder has over 2,000 videos! I actually watched few seconds of a few videos here and there, but I'm not sure at all that that is the case.\nOk, if you split ensuring all videos from a single folder belong only to either train or validation sets and see a big gain in LB, i.e. your model is generalizing better, you can infer that this is because you are splitting better the data - and a plausible explanation would be that all videos from a single actor belong to a single folder. But that could be explained by other things.\nMy idea is to extract 1 frame from each video (~120K frames) and use embeddings to group similar faces: that's very doable. After all, we already have the trained models. :)</p>",
          "rawMarkdown": "Out of curiosity, how did you check the folders separate actors? \nEvery folder has over 2,000 videos! I actually watched few seconds of a few videos here and there, but I'm not sure at all that that is the case.\nOk, if you split ensuring all videos from a single folder belong only to either train or validation sets and see a big gain in LB, i.e. your model is generalizing better, you can infer that this is because you are splitting better the data - and a plausible explanation would be that all videos from a single actor belong to a single folder. But that could be explained by other things.\nMy idea is to extract 1 frame from each video (~120K frames) and use embeddings to group similar faces: that's very doable. After all, we already have the trained models. :)",
          "votes": 1
        },
        {
          "id": 708855,
          "postDate": "2020-01-02T19:14:13.847Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Thanks for the useful points. Any suggestions on how to select a frame in a video or what’s the resolution of the cropped faces that would work well with embedding estimation approaches?</p>",
          "rawMarkdown": "@carlossouza Thanks for the useful points. Any suggestions on how to select a frame in a video or what’s the resolution of the cropped faces that would work well with embedding estimation approaches?"
        },
        {
          "id": 708860,
          "postDate": "2020-01-02T19:20:53.647Z",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> What makes me feel it will be a bit fragile is splitting by a folder may mean that we are looking at the performance of the model on that specific actor rather than generalized performance we can expect to see on future unseen data. </p>\n\n<p>For example, I can already see that the pre-existing face detection models are significantly worse on the black participants (either because of lack of contrast or because lack of training data on black people for the face detection algorithms). If you choose three folders and it's all white people then you will greatly overstate your performance. I think validating like that is probably better than the other two methods I stated, but still imperfect. </p>",
          "rawMarkdown": "@pedromb What makes me feel it will be a bit fragile is splitting by a folder may mean that we are looking at the performance of the model on that specific actor rather than generalized performance we can expect to see on future unseen data. \n\nFor example, I can already see that the pre-existing face detection models are significantly worse on the black participants (either because of lack of contrast or because lack of training data on black people for the face detection algorithms). If you choose three folders and it's all white people then you will greatly overstate your performance. I think validating like that is probably better than the other two methods I stated, but still imperfect. ",
          "votes": 2
        },
        {
          "id": 708892,
          "postDate": "2020-01-02T20:06:15.603Z",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> That's what I meant by being unrealist to check every folder, my assumption was based on my experiments alone and a few manual checks.</p>\n\n<p><a href=\"/ryches\">@ryches</a> I understand what you mean now and I totally agree (I also noticed the fact that the face detection methods are worse on black participants). Like I said, I believe this is the best effortless validation strategy to use. </p>",
          "rawMarkdown": "@carlossouza That's what I meant by being unrealist to check every folder, my assumption was based on my experiments alone and a few manual checks.\n\n@ryches I understand what you mean now and I totally agree (I also noticed the fact that the face detection methods are worse on black participants). Like I said, I believe this is the best effortless validation strategy to use. "
        },
        {
          "id": 708978,
          "postDate": "2020-01-02T22:15:23.693Z",
          "content": "<p>Hi <a href=\"/ryches\">@ryches</a> <a href=\"/carlossouza\">@carlossouza</a> ,\nThanks for your discussion. Can you explain the term \"embedding\"? I'm new to CV, so I'm not familiar with it. Does it like an \"autoencoder\"?</p>\n\n<p>Thanks!</p>",
          "rawMarkdown": "Hi @ryches @carlossouza ,\nThanks for your discussion. Can you explain the term \"embedding\"? I'm new to CV, so I'm not familiar with it. Does it like an \"autoencoder\"?\n\nThanks!"
        },
        {
          "id": 712216,
          "postDate": "2020-01-07T00:27:24.977Z",
          "content": "<p><a href=\"/stevenyin\">@stevenyin</a> embedding is generated by just using a pretrained model conv part to predict new tasks. You can get more details by going over <a href=\"/carlossouza\">@carlossouza</a> notebook. Hope this helps.</p>",
          "rawMarkdown": "@stevenyin embedding is generated by just using a pretrained model conv part to predict new tasks. You can get more details by going over @carlossouza notebook. Hope this helps."
        }
      ]
    },
    {
      "id": 710476,
      "postDate": "2020-01-04T19:18:12.843Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 709229,
      "author_name": "Liuftvafas",
      "author_url": "",
      "post_date": "2020-01-03T07:52:58.677000",
      "content": "<p>Different folders can contain the same actors. For example:\n<code>\ndfdc_train_part_44/abkclntljz.mp4\ndfdc_train_part_14/jvtgwixooc.mp4\ndfdc_train_part_47/mlwnlxchfy.mp4\ndfdc_train_part_35/iukwdhtsut.mp4\ndfdc_train_part_13/fhmjuyomqf.mp4\ndfdc_train_part_26/wlrgqwoqfx.mp4\n</code></p>",
      "votes": 14,
      "replies": [
        {
          "id": 714583,
          "author_name": "AkiraSosa",
          "author_url": "",
          "post_date": "2020-01-09T15:15:08.703000",
          "content": "<p>Hi Organizers,</p>\n\n<p>It makes hard for participants to use this dataset. Though I have tried some automated ways to identify faces, my conclusion is that it's impossible to make a reliable cross validation with this dataset. So, is there any plan to provide an information about actors for each videos?</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 770419,
          "author_name": "Chew Kok Wah",
          "author_url": "",
          "post_date": "2020-03-12T22:32:36.520000",
          "content": "<p>May I know how do you find out video with same actor from different folder? </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 709199,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-01-03T06:45:27.570000",
      "content": "<p>Just finished a 1st experiment with embeddings to group similar faces (hopefully from the same actor) automatically. Here's 3 cases, and below the code. The large image in the left is given; the small images in the right are selected by the algorithm as the most similar to the one in the left (universe = 800 sample videos provided). There's a lot to improve in the code, this is just a quick example to show how embeddings are cool :)</p>\n\n<p>Case 1\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fef9fa8a881d97538f9a7c68e73389023%2FScreen%20Shot%202020-01-03%20at%2003.34.37.png?generation=1578033677488156&amp;alt=media\" alt=\"\"></p>\n\n<p>Case 2\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F14d092eda70ccab8c0ac2ec2445aed92%2FScreen%20Shot%202020-01-03%20at%2003.31.01.png?generation=1578033761923474&amp;alt=media\" alt=\"\"></p>\n\n<p>Case 3\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fee9275d339905ec47b41793eccc5a6a4%2FScreen%20Shot%202020-01-03%20at%2003.31.57.png?generation=1578033783595516&amp;alt=media\" alt=\"\"></p>\n\n<p><a href=\"https://www.kaggle.com/carlossouza/embeddings-grouping-similar-faces-automagically\">https://www.kaggle.com/carlossouza/embeddings-grouping-similar-faces-automagically</a></p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 710487,
      "author_name": "Manideep",
      "author_url": "",
      "post_date": "2020-01-04T19:47:44.590000",
      "content": "<p>All the folders contain fakes and reals at a ratio of 5:1. But the actual public test set has approximately equal number of both. Will it be a good idea to use only videos from one folder which has this huge data imbalance? Even though the log loss is the metric, it can go very high in case of false fake (true positives) predictions with which this data imbalance will not help to validate.</p>\n\n<p>For folder <code>dfdc_train_part_33</code>,</p>\n\n<p>```python\nIn [1]: labels.count('FAKE')\nOut[1]: 1884</p>\n\n<p>In [2]: labels.count('REAL')\nOut[2]: 390\n```</p>",
      "votes": 1,
      "replies": [
        {
          "id": 712214,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-07T00:24:32.220000",
          "content": "<p>Apply underbalancing technique would solve this problem. And you can use more than one folder(just make sure there's no leak).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714268,
          "author_name": "Emmettj",
          "author_url": "",
          "post_date": "2020-01-09T08:55:38.913000",
          "content": "<p>Hi，I'm new to CV and DL. Could you share some underbalancing techniques with me? Does this  equal to 'sample'? Thx!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714283,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-01-09T09:06:11.337000",
          "content": "<p>I'm not an expert but there are two tricks I use and they seem to work. One is to repeat real videos and make batch with the real-fake pairs of same actor(known as overbalancing, thanks to <a href=\"/unkownhihi\">@unkownhihi</a> for pointing out ). and other is to change a little bit to loss function\nfor example -(2*y*log(ypred) + (1-y)*log(1-ypred)) so that classifying a real video as fake will have greater loss. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714420,
          "author_name": "Emmettj",
          "author_url": "",
          "post_date": "2020-01-09T12:31:55.013000",
          "content": "<p>thx!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 714974,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-10T01:49:04.073000",
          "content": "<p><a href=\"/ankitsainiankit\">@ankitsainiankit</a> The first one you said is overbalancing. The under balancing is to pick the same number of fake videos as real videos(in other words, delete some fake videos). </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 715099,
          "author_name": "Ankit Saini",
          "author_url": "",
          "post_date": "2020-01-10T06:02:20.050000",
          "content": "<p><a href=\"/unkownhihi\">@unkownhihi</a> I'm sorry for misinformation. I'll edit it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 785213,
          "author_name": "Kerem Turgutlu",
          "author_url": "",
          "post_date": "2020-03-24T22:08:37.730000",
          "content": "<p>You can pick subfolders, e.g. 41-45. Then you can randomly pick 1 fake for each original. Do this multiple times (bootstrapping) calculate validation for each bootstrap and average them out to get a more robust estimate. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 709985,
      "author_name": "Yifan Xie",
      "author_url": "",
      "post_date": "2020-01-04T07:10:39.197000",
      "content": "<p>I have done some local experiment with both validation by group, and validation by original videos.\nMy observation is that the gap between training and validation error is a lot wider when validating by original videos.\nFor me that is a clear indication of overfitting.</p>\n\n<p>As pointed out by <a href=\"/liuftvafas\">@liuftvafas</a>,  we do have the same actor in different groups.  nevertheless, base on the above evidence I would go with validation by groups until better idea is shared :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 759228,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "2020-02-28T18:54:47.893000",
          "content": "<p>Thanks for sharing! With validation by group do you mean by actors? Or what does it mean exactly? Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 709180,
      "author_name": "YoonSoo",
      "author_url": "",
      "post_date": "2020-01-03T06:04:10.040000",
      "content": "<p>Yeah, I had large discrepancy between val score and lb score when I just did random split. Now I'm trying splitting with respect to source videos.\nOne thing to point out, as shown in <a href=\"https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013\">https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013</a> , augmentations are applied to test set. Maybe we should simulate these augmentations to our valid set.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 708486,
      "author_name": "Carlos Souza",
      "author_url": "",
      "post_date": "2020-01-02T10:49:41.793000",
      "content": "<p>We can use a group of actors in training set, and a distinct group of actors in the validation set. This can be done automagically using embeddings. Will try to do a notebook in the next days to show how. And yes, I'm seeing a big discrepancy too (shown below). But at least for now, I'm focusing on other parts of the problem... validation will come later. Cheers!\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fa0dfc470332d5200e9f0eb99a0a3e466%2Fistanbul-20200102.png?generation=1577968215664071&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 708712,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-01-02T15:24:19.710000",
          "content": "<p>Is it necessary to use embeddings? My understanding was that the folders already split the video to the different actors. Have not spent a ton of time verifying it myself but someone stated it and seemed to be true in my limited experience. </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 708769,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-02T16:59:07.820000",
          "content": "<p>I've read that somewhere in the forum too... but will double check to make sure!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 708772,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-01-02T17:03:11.733000",
          "content": "<p>I'd be interested to see what you had in mind with the embeddings too. I get the concept but I'm not sure exactly how I'd formulate that in our video context with tons of videos and frames for each actor. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 708800,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-02T17:44:29.300000",
          "content": "<p>I'm almost sure this is the case, I think is unrealistic to actually check on that, but by checking a few and also considering the results of my experiments I do believe that the foIders also separate actors (they for sure separate original/fakes of a same video, this I've checked). I didn't quiet understand why you believe this strategy would be fragile, can you elaborate on that? I believe that doing a group split by folder is the best effortless way of creating a validation set. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 708843,
          "author_name": "Carlos Souza",
          "author_url": "",
          "post_date": "2020-01-02T18:49:14.853000",
          "content": "<p>Out of curiosity, how did you check the folders separate actors? \nEvery folder has over 2,000 videos! I actually watched few seconds of a few videos here and there, but I'm not sure at all that that is the case.\nOk, if you split ensuring all videos from a single folder belong only to either train or validation sets and see a big gain in LB, i.e. your model is generalizing better, you can infer that this is because you are splitting better the data - and a plausible explanation would be that all videos from a single actor belong to a single folder. But that could be explained by other things.\nMy idea is to extract 1 frame from each video (~120K frames) and use embeddings to group similar faces: that's very doable. After all, we already have the trained models. :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 708855,
          "author_name": "Debanga Raj Neog",
          "author_url": "",
          "post_date": "2020-01-02T19:14:13.847000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> Thanks for the useful points. Any suggestions on how to select a frame in a video or what’s the resolution of the cropped faces that would work well with embedding estimation approaches?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 708860,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2020-01-02T19:20:53.647000",
          "content": "<p><a href=\"/pedromb\">@pedromb</a> What makes me feel it will be a bit fragile is splitting by a folder may mean that we are looking at the performance of the model on that specific actor rather than generalized performance we can expect to see on future unseen data. </p>\n\n<p>For example, I can already see that the pre-existing face detection models are significantly worse on the black participants (either because of lack of contrast or because lack of training data on black people for the face detection algorithms). If you choose three folders and it's all white people then you will greatly overstate your performance. I think validating like that is probably better than the other two methods I stated, but still imperfect. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 708892,
          "author_name": "Pedro Bernardo",
          "author_url": "",
          "post_date": "2020-01-02T20:06:15.603000",
          "content": "<p><a href=\"/carlossouza\">@carlossouza</a> That's what I meant by being unrealist to check every folder, my assumption was based on my experiments alone and a few manual checks.</p>\n\n<p><a href=\"/ryches\">@ryches</a> I understand what you mean now and I totally agree (I also noticed the fact that the face detection methods are worse on black participants). Like I said, I believe this is the best effortless validation strategy to use. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 708978,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-01-02T22:15:23.693000",
          "content": "<p>Hi <a href=\"/ryches\">@ryches</a> <a href=\"/carlossouza\">@carlossouza</a> ,\nThanks for your discussion. Can you explain the term \"embedding\"? I'm new to CV, so I'm not familiar with it. Does it like an \"autoencoder\"?</p>\n\n<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 712216,
          "author_name": "Shangqiu Li",
          "author_url": "",
          "post_date": "2020-01-07T00:27:24.977000",
          "content": "<p><a href=\"/stevenyin\">@stevenyin</a> embedding is generated by just using a pretrained model conv part to predict new tasks. You can get more details by going over <a href=\"/carlossouza\">@carlossouza</a> notebook. Hope this helps.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 710476,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-04T19:18:12.843000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "708219": "I've seen a few people in the discussion mention a big discrepancy between local validation and LB. Some of the options that might be viable is a stratified hold out, group shuffle split that separates based on the source video so you don't have 3 fakes of a certain video in the training set and another 2 fakes of the same video in the validation set. Alternately a split can be done based on the folders provided to us. So then it would be a true holdout of an entirely unseen face and video set. My concern with the final one is that will likely be fragile based on the subject being filmed, we don't really have a ton of unique faces. \n\nAny thoughts on this? Are others seeing a large discrepancy between val and lb?",
    "709229": "Different folders can contain the same actors. For example:\n```\ndfdc_train_part_44/abkclntljz.mp4\ndfdc_train_part_14/jvtgwixooc.mp4\ndfdc_train_part_47/mlwnlxchfy.mp4\ndfdc_train_part_35/iukwdhtsut.mp4\ndfdc_train_part_13/fhmjuyomqf.mp4\ndfdc_train_part_26/wlrgqwoqfx.mp4\n```",
    "709199": "Just finished a 1st experiment with embeddings to group similar faces (hopefully from the same actor) automatically. Here's 3 cases, and below the code. The large image in the left is given; the small images in the right are selected by the algorithm as the most similar to the one in the left (universe = 800 sample videos provided). There's a lot to improve in the code, this is just a quick example to show how embeddings are cool :)\n\nCase 1\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fef9fa8a881d97538f9a7c68e73389023%2FScreen%20Shot%202020-01-03%20at%2003.34.37.png?generation=1578033677488156&amp;alt=media)\n\nCase 2\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2F14d092eda70ccab8c0ac2ec2445aed92%2FScreen%20Shot%202020-01-03%20at%2003.31.01.png?generation=1578033761923474&amp;alt=media)\n\nCase 3\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fee9275d339905ec47b41793eccc5a6a4%2FScreen%20Shot%202020-01-03%20at%2003.31.57.png?generation=1578033783595516&amp;alt=media)\n\nhttps://www.kaggle.com/carlossouza/embeddings-grouping-similar-faces-automagically\n",
    "710487": "All the folders contain fakes and reals at a ratio of 5:1. But the actual public test set has approximately equal number of both. Will it be a good idea to use only videos from one folder which has this huge data imbalance? Even though the log loss is the metric, it can go very high in case of false fake (true positives) predictions with which this data imbalance will not help to validate.\n\nFor folder `dfdc_train_part_33`,\n\n```python\nIn [1]: labels.count('FAKE')\nOut[1]: 1884\n\nIn [2]: labels.count('REAL')\nOut[2]: 390\n```\n",
    "709985": "I have done some local experiment with both validation by group, and validation by original videos.\nMy observation is that the gap between training and validation error is a lot wider when validating by original videos.\nFor me that is a clear indication of overfitting.\n\nAs pointed out by @liuftvafas,  we do have the same actor in different groups.  nevertheless, base on the above evidence I would go with validation by groups until better idea is shared :)\n",
    "709180": "Yeah, I had large discrepancy between val score and lb score when I just did random split. Now I'm trying splitting with respect to source videos.\nOne thing to point out, as shown in https://www.kaggle.com/c/deepfake-detection-challenge/discussion/122013 , augmentations are applied to test set. Maybe we should simulate these augmentations to our valid set.",
    "708486": "We can use a group of actors in training set, and a distinct group of actors in the validation set. This can be done automagically using embeddings. Will try to do a notebook in the next days to show how. And yes, I'm seeing a big discrepancy too (shown below). But at least for now, I'm focusing on other parts of the problem... validation will come later. Cheers!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F915913%2Fa0dfc470332d5200e9f0eb99a0a3e466%2Fistanbul-20200102.png?generation=1577968215664071&amp;alt=media)\n",
    "710476": ""
  }
}