{
  "id": 132700,
  "title": "How to train with all dataset?",
  "url": "/competitions/deepfake-detection-challenge/discussion/132700",
  "author_name": "",
  "post_date": "2020-02-27T12:10:37.143610Z",
  "votes": 15,
  "comment_count": 24,
  "views": 0,
  "content": "<p>The dataset is class imbalanced, that is, REAL:FAKE=1:5. I currently use under sample to make sure that the number of REAL and FAKE videos is the same, but this keeps most FAKE videos from being exploited and somewhat limits the generalization of model. I tried methods like weight loss and over sampler, but still perform worse than under sample. Does anyone know how to use the entire data set for training? Thank you very much!</p>",
  "messages": [
    {
      "id": "758071",
      "postDate": "02/27/2020 12:10:37",
      "content": "<p>The dataset is class imbalanced, that is, REAL:FAKE=1:5. I currently use under sample to make sure that the number of REAL and FAKE videos is the same, but this keeps most FAKE videos from being exploited and somewhat limits the generalization of model. I tried methods like weight loss and over sampler, but still perform worse than under sample. Does anyone know how to use the entire data set for training? Thank you very much!</p>",
      "rawMarkdown": "The dataset is class imbalanced, that is, REAL:FAKE=1:5. I currently use under sample to make sure that the number of REAL and FAKE videos is the same, but this keeps most FAKE videos from being exploited and somewhat limits the generalization of model. I tried methods like weight loss and over sampler, but still perform worse than under sample. Does anyone know how to use the entire data set for training? Thank you very much!",
      "votes": null
    },
    {
      "id": "758213",
      "postDate": "02/27/2020 14:50:20",
      "content": "<p>Man, you are in 6th place with underbalancing, that is truely fascinating, well, we do have our best score with underbalancing, however yesterday, I had a success in making a data, which I say that any data better than this is not possible for this competition, atleast theoratically, I will give you a hint \"1 frame should only appear once\", without underbalacing, think about that you might get it, Good Luck!</p>",
      "rawMarkdown": "Man, you are in 6th place with underbalancing, that is truely fascinating, well, we do have our best score with underbalancing, however yesterday, I had a success in making a data, which I say that any data better than this is not possible for this competition, atleast theoratically, I will give you a hint \"1 frame should only appear once\", without underbalacing, think about that you might get it, Good Luck!",
      "votes": null
    },
    {
      "id": "758338",
      "postDate": "02/27/2020 16:47:29",
      "content": "<p>An easy way around this, although I'm not sure how effective it is, is too just cycle through the fakes throughout the batches/epochs. You create a list with every fake face and cycles through this list throughout the epochs of training. For example, if you have 20k real faces and 100k fake faces your data per epoch would be:</p>\n\n<p>Epoch 1: 20k real + 0 to 20k from your fake list.\nEpoch 2: 20k real + 20k to 40k from your fake list.\n....\nEpoch 5: 20k real + 80k to 100k from your fake list.\nEpoch 6: 20k real + 0 to 20k from your fake list.</p>\n\n<p>And so on. I'm not sure if this has any advantage over just sampling 20k from the fake list every epoch, but theoretically it would allow you to use all data, somehow, at least.</p>",
      "rawMarkdown": "An easy way around this, although I'm not sure how effective it is, is too just cycle through the fakes throughout the batches/epochs. You create a list with every fake face and cycles through this list throughout the epochs of training. For example, if you have 20k real faces and 100k fake faces your data per epoch would be:\n\nEpoch 1: 20k real + 0 to 20k from your fake list.\nEpoch 2: 20k real + 20k to 40k from your fake list.\n....\nEpoch 5: 20k real + 80k to 100k from your fake list.\nEpoch 6: 20k real + 0 to 20k from your fake list.\n\nAnd so on. I'm not sure if this has any advantage over just sampling 20k from the fake list every epoch, but theoretically it would allow you to use all data, somehow, at least.",
      "votes": null
    },
    {
      "id": "758346",
      "postDate": "02/27/2020 16:52:00",
      "content": "<p>really not effective, its just normal oversampling, in our case it did not even give any advantage over undersampling.</p>",
      "rawMarkdown": "really not effective, its just normal oversampling, in our case it did not even give any advantage over undersampling.",
      "votes": null
    },
    {
      "id": "758790",
      "postDate": "02/28/2020 06:49:35",
      "content": "<p>Thank you very much! \nBTW, Could you please tell me whether \"1 frame\" means 1 frame per video or the order of frames in the video?</p>",
      "rawMarkdown": "Thank you very much! \nBTW, Could you please tell me whether \"1 frame\" means 1 frame per video or the order of frames in the video?",
      "votes": null
    },
    {
      "id": "758794",
      "postDate": "02/28/2020 06:53:59",
      "content": "<p>I also think this approach may not work, which inevitably making the model too focused much on the new coming FAKE images.</p>",
      "rawMarkdown": "I also think this approach may not work, which inevitably making the model too focused much on the new coming FAKE images.",
      "votes": null
    },
    {
      "id": "758922",
      "postDate": "02/28/2020 10:48:23",
      "content": "<p>1 frame means a single frame never repeats in the dataset, the dataset is the images of different frames from different videos, but super-highly balanced. (Its not a simple multiplier trick)</p>",
      "rawMarkdown": "1 frame means a single frame never repeats in the dataset, the dataset is the images of different frames from different videos, but super-highly balanced. (Its not a simple multiplier trick)",
      "votes": null
    },
    {
      "id": "759378",
      "postDate": "02/29/2020 00:32:47",
      "content": "<p>how about taking more frame samples from real videos vs. fake videos?</p>",
      "rawMarkdown": "how about taking more frame samples from real videos vs. fake videos?",
      "votes": null
    },
    {
      "id": "759482",
      "postDate": "02/29/2020 05:02:40",
      "content": "<p>The strategy I am going to try next is in the picture. It won't include every fake image but the important consideration is frames between real and fake are aligned to eliminate bias.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F170613a872b3ffa08a054cacd33063c1%2Fstrategy.jpg?generation=1582952448470919&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "The strategy I am going to try next is in the picture. It won't include every fake image but the important consideration is frames between real and fake are aligned to eliminate bias.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F170613a872b3ffa08a054cacd33063c1%2Fstrategy.jpg?generation=1582952448470919&amp;alt=media)",
      "votes": null
    },
    {
      "id": "759547",
      "postDate": "02/29/2020 07:09:35",
      "content": "<p>I've tried this method, but it doesn't work.</p>",
      "rawMarkdown": "I've tried this method, but it doesn't work.",
      "votes": null
    },
    {
      "id": "759550",
      "postDate": "02/29/2020 07:10:45",
      "content": "<p>Thank you. I'm trying and hope it works.</p>",
      "rawMarkdown": "Thank you. I'm trying and hope it works.",
      "votes": null
    },
    {
      "id": "759716",
      "postDate": "02/29/2020 11:28:39",
      "content": "<p>Already been trying this method for a while now. Hasn't really provided any gains.</p>",
      "rawMarkdown": "Already been trying this method for a while now. Hasn't really provided any gains.",
      "votes": null
    },
    {
      "id": "759760",
      "postDate": "02/29/2020 12:34:59",
      "content": "<p>This method, just has to be theoratically it, you must be doing a flaw in that, I did try it once, I am gonna do it again, as I got a bug the last time I tried.</p>",
      "rawMarkdown": "This method, just has to be theoratically it, you must be doing a flaw in that, I did try it once, I am gonna do it again, as I got a bug the last time I tried.",
      "votes": null
    },
    {
      "id": "763532",
      "postDate": "03/04/2020 15:17:07",
      "content": "<p>That's interesting.. If I may ask, what's the idea behind this? Do you believe there is a correlation between the Nth frame of a video with the frame N of the other videos in the dataset?\nBtw, did this work for you guys? <a href=\"/chenshen03\">@chenshen03</a> <a href=\"/harshitsheoran\">@harshitsheoran</a> </p>",
      "rawMarkdown": "That's interesting.. If I may ask, what's the idea behind this? Do you believe there is a correlation between the Nth frame of a video with the frame N of the other videos in the dataset?\nBtw, did this work for you guys? @chenshen03 @harshitsheoran",
      "votes": null
    },
    {
      "id": "763533",
      "postDate": "03/04/2020 15:20:35",
      "content": "<p>I mean, thats litterally the best we could do, <a href=\"/arc144\">@arc144</a> it does work out for us as it has no bottleneck even when using full data, there is a correlation but not with any video to any video, Trust me, you can get a lot of information from the metadata file.</p>",
      "rawMarkdown": "I mean, thats litterally the best we could do, @arc144 it does work out for us as it has no bottleneck even when using full data, there is a correlation but not with any video to any video, Trust me, you can get a lot of information from the metadata file.",
      "votes": null
    },
    {
      "id": "763570",
      "postDate": "03/04/2020 16:02:03",
      "content": "<p>Well, I noticed that if I used more than 500K approx train images using the above method, my model would start overfitting very quickly. So I've gone back to using lesser images, 200k approx.</p>",
      "rawMarkdown": "Well, I noticed that if I used more than 500K approx train images using the above method, my model would start overfitting very quickly. So I've gone back to using lesser images, 200k approx.",
      "votes": null
    },
    {
      "id": "763615",
      "postDate": "03/04/2020 17:01:17",
      "content": "<p>Using that technique I would suggest only 160k images.</p>",
      "rawMarkdown": "Using that technique I would suggest only 160k images.",
      "votes": null
    },
    {
      "id": "765876",
      "postDate": "03/07/2020 09:14:43",
      "content": "<p>do your model work in unbalance data?@Chason</p>",
      "rawMarkdown": "do your model work in unbalance data?@Chason",
      "votes": null
    },
    {
      "id": "766337",
      "postDate": "03/08/2020 03:01:04",
      "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> No, I'm still using balanced data, which constructed by under sample.</p>",
      "rawMarkdown": "chenbaoying No, I'm still using balanced data, which constructed by under sample.",
      "votes": null
    },
    {
      "id": "766395",
      "postDate": "03/08/2020 04:48:28",
      "content": "<p>I treat the excess fakes as \"data augmentation\" so in different epochs, while I use the same reals, the fakes are similar not identical. The histogram in this case is a useful diagnostic - it becomes a bit asymmetrical, preferring fake over real (although the the test videos have equal real/fake). This is in spite of a score that's not too bad (~ 0.37).</p>",
      "rawMarkdown": "I treat the excess fakes as \"data augmentation\" so in different epochs, while I use the same reals, the fakes are similar not identical. The histogram in this case is a useful diagnostic - it becomes a bit asymmetrical, preferring fake over real (although the the test videos have equal real/fake). This is in spite of a score that's not too bad (~ 0.37).",
      "votes": null
    },
    {
      "id": "767316",
      "postDate": "03/09/2020 13:24:09",
      "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> : May I ask do you mean this: suppose video A as 4 fakes: B, C, D, E. Then suppose I have frames from 1 to 8, equally gapped. Then I can take frame 1, 3, 5, 7 from video A; frame 2 from video B; frame 4 from video C, frame 6 from video D, and frame 8 from video E. Is that what you hint?</p>",
      "rawMarkdown": "harshitsheoran : May I ask do you mean this: suppose video A as 4 fakes: B, C, D, E. Then suppose I have frames from 1 to 8, equally gapped. Then I can take frame 1, 3, 5, 7 from video A; frame 2 from video B; frame 4 from video C, frame 6 from video D, and frame 8 from video E. Is that what you hint?",
      "votes": null
    },
    {
      "id": "767579",
      "postDate": "03/09/2020 20:40:12",
      "content": "<p><a href=\"/khahuras\">@khahuras</a> Good Job, You are really pretty close, this is one of finest data I had ever created and currently I am using a very similar data to what you are saying, however, yesterday, I found a big bottleneck in this technique, Yes, that was very close to my hint, but still you can get much more information from metadata file than you think you can.</p>\n\n<p>EDIT: The technique you defined above should not have a big performance difference than mine, Good Luck!</p>",
      "rawMarkdown": "khahuras Good Job, You are really pretty close, this is one of finest data I had ever created and currently I am using a very similar data to what you are saying, however, yesterday, I found a big bottleneck in this technique, Yes, that was very close to my hint, but still you can get much more information from metadata file than you think you can.\n\nEDIT: The technique you defined above should not have a big performance difference than mine, Good Luck!",
      "votes": null
    },
    {
      "id": "767649",
      "postDate": "03/10/2020 00:29:15",
      "content": "<p>Thanks <a href=\"/harshitsheoran\">@harshitsheoran</a> . Indeed, I have been doing something like this: for each epoch, for each shuffled index of all real videos, select 1 pair of (real, fake). The fake is randomly selected from the corresponding fakes of that real video. So, in 1 epoch, the total number of samples would be 2*nb_real_videos. The samples in each epoch are not fixed and built beforehand, but indeed are randomly sampled after every epoch (hence different). And different epochs have different fake versions because of this way. I'm still quite confused why this doesn't work. I think by randomly selecting 1 fake, it acts as a kind of augmentation. I may need to re-think why.</p>",
      "rawMarkdown": "Thanks @harshitsheoran . Indeed, I have been doing something like this: for each epoch, for each shuffled index of all real videos, select 1 pair of (real, fake). The fake is randomly selected from the corresponding fakes of that real video. So, in 1 epoch, the total number of samples would be 2*nb_real_videos. The samples in each epoch are not fixed and built beforehand, but indeed are randomly sampled after every epoch (hence different). And different epochs have different fake versions because of this way. I'm still quite confused why this doesn't work. I think by randomly selecting 1 fake, it acts as a kind of augmentation. I may need to re-think why.",
      "votes": null
    },
    {
      "id": "767650",
      "postDate": "03/10/2020 00:30:49",
      "content": "<p><a href=\"/khahuras\">@khahuras</a> Sry your technique, I can not say it is the best, there are a lot more better than this one.</p>",
      "rawMarkdown": "khahuras Sry your technique, I can not say it is the best, there are a lot more better than this one.",
      "votes": null
    },
    {
      "id": "768957",
      "postDate": "03/11/2020 11:40:36",
      "content": "<p><a href=\"/chenshen03\">@chenshen03</a> How many frams per videos in your training, and how many frames in your kaggle submition</p>",
      "rawMarkdown": "chenshen03 How many frams per videos in your training, and how many frames in your kaggle submition",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 758213,
      "author_name": "harshitsheoran",
      "author_url": "",
      "post_date": "02/27/2020 14:50:20",
      "content": "<p>Man, you are in 6th place with underbalancing, that is truely fascinating, well, we do have our best score with underbalancing, however yesterday, I had a success in making a data, which I say that any data better than this is not possible for this competition, atleast theoratically, I will give you a hint \"1 frame should only appear once\", without underbalacing, think about that you might get it, Good Luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 758790,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "02/28/2020 06:49:35",
          "content": "<p>Thank you very much! \nBTW, Could you please tell me whether \"1 frame\" means 1 frame per video or the order of frames in the video?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758922,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/28/2020 10:48:23",
          "content": "<p>1 frame means a single frame never repeats in the dataset, the dataset is the images of different frames from different videos, but super-highly balanced. (Its not a simple multiplier trick)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759550,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "02/29/2020 07:10:45",
          "content": "<p>Thank you. I'm trying and hope it works.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763532,
          "author_name": "arc144",
          "author_url": "",
          "post_date": "03/04/2020 15:17:07",
          "content": "<p>That's interesting.. If I may ask, what's the idea behind this? Do you believe there is a correlation between the Nth frame of a video with the frame N of the other videos in the dataset?\nBtw, did this work for you guys? <a href=\"/chenshen03\">@chenshen03</a> <a href=\"/harshitsheoran\">@harshitsheoran</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763533,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/04/2020 15:20:35",
          "content": "<p>I mean, thats litterally the best we could do, <a href=\"/arc144\">@arc144</a> it does work out for us as it has no bottleneck even when using full data, there is a correlation but not with any video to any video, Trust me, you can get a lot of information from the metadata file.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 765876,
          "author_name": "chenbaoying",
          "author_url": "",
          "post_date": "03/07/2020 09:14:43",
          "content": "<p>do your model work in unbalance data?@Chason</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766337,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "03/08/2020 03:01:04",
          "content": "<p><a href=\"/chenbaoying\">@chenbaoying</a> No, I'm still using balanced data, which constructed by under sample.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 766395,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "03/08/2020 04:48:28",
          "content": "<p>I treat the excess fakes as \"data augmentation\" so in different epochs, while I use the same reals, the fakes are similar not identical. The histogram in this case is a useful diagnostic - it becomes a bit asymmetrical, preferring fake over real (although the the test videos have equal real/fake). This is in spite of a score that's not too bad (~ 0.37).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767316,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/09/2020 13:24:09",
          "content": "<p><a href=\"/harshitsheoran\">@harshitsheoran</a> : May I ask do you mean this: suppose video A as 4 fakes: B, C, D, E. Then suppose I have frames from 1 to 8, equally gapped. Then I can take frame 1, 3, 5, 7 from video A; frame 2 from video B; frame 4 from video C, frame 6 from video D, and frame 8 from video E. Is that what you hint?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767579,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/09/2020 20:40:12",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> Good Job, You are really pretty close, this is one of finest data I had ever created and currently I am using a very similar data to what you are saying, however, yesterday, I found a big bottleneck in this technique, Yes, that was very close to my hint, but still you can get much more information from metadata file than you think you can.</p>\n\n<p>EDIT: The technique you defined above should not have a big performance difference than mine, Good Luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767649,
          "author_name": "khahuras",
          "author_url": "",
          "post_date": "03/10/2020 00:29:15",
          "content": "<p>Thanks <a href=\"/harshitsheoran\">@harshitsheoran</a> . Indeed, I have been doing something like this: for each epoch, for each shuffled index of all real videos, select 1 pair of (real, fake). The fake is randomly selected from the corresponding fakes of that real video. So, in 1 epoch, the total number of samples would be 2*nb_real_videos. The samples in each epoch are not fixed and built beforehand, but indeed are randomly sampled after every epoch (hence different). And different epochs have different fake versions because of this way. I'm still quite confused why this doesn't work. I think by randomly selecting 1 fake, it acts as a kind of augmentation. I may need to re-think why.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 767650,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/10/2020 00:30:49",
          "content": "<p><a href=\"/khahuras\">@khahuras</a> Sry your technique, I can not say it is the best, there are a lot more better than this one.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 758338,
      "author_name": "pedromb",
      "author_url": "",
      "post_date": "02/27/2020 16:47:29",
      "content": "<p>An easy way around this, although I'm not sure how effective it is, is too just cycle through the fakes throughout the batches/epochs. You create a list with every fake face and cycles through this list throughout the epochs of training. For example, if you have 20k real faces and 100k fake faces your data per epoch would be:</p>\n\n<p>Epoch 1: 20k real + 0 to 20k from your fake list.\nEpoch 2: 20k real + 20k to 40k from your fake list.\n....\nEpoch 5: 20k real + 80k to 100k from your fake list.\nEpoch 6: 20k real + 0 to 20k from your fake list.</p>\n\n<p>And so on. I'm not sure if this has any advantage over just sampling 20k from the fake list every epoch, but theoretically it would allow you to use all data, somehow, at least.</p>",
      "votes": null,
      "replies": [
        {
          "id": 758346,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/27/2020 16:52:00",
          "content": "<p>really not effective, its just normal oversampling, in our case it did not even give any advantage over undersampling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 758794,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "02/28/2020 06:53:59",
          "content": "<p>I also think this approach may not work, which inevitably making the model too focused much on the new coming FAKE images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759378,
      "author_name": "jmoscetti",
      "author_url": "",
      "post_date": "02/29/2020 00:32:47",
      "content": "<p>how about taking more frame samples from real videos vs. fake videos?</p>",
      "votes": null,
      "replies": [
        {
          "id": 759547,
          "author_name": "chenshen03",
          "author_url": "",
          "post_date": "02/29/2020 07:09:35",
          "content": "<p>I've tried this method, but it doesn't work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 768957,
          "author_name": "chenbaoying",
          "author_url": "",
          "post_date": "03/11/2020 11:40:36",
          "content": "<p><a href=\"/chenshen03\">@chenshen03</a> How many frams per videos in your training, and how many frames in your kaggle submition</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 759482,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "02/29/2020 05:02:40",
      "content": "<p>The strategy I am going to try next is in the picture. It won't include every fake image but the important consideration is frames between real and fake are aligned to eliminate bias.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F170613a872b3ffa08a054cacd33063c1%2Fstrategy.jpg?generation=1582952448470919&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 759716,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "02/29/2020 11:28:39",
          "content": "<p>Already been trying this method for a while now. Hasn't really provided any gains.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 759760,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "02/29/2020 12:34:59",
          "content": "<p>This method, just has to be theoratically it, you must be doing a flaw in that, I did try it once, I am gonna do it again, as I got a bug the last time I tried.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763570,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "03/04/2020 16:02:03",
          "content": "<p>Well, I noticed that if I used more than 500K approx train images using the above method, my model would start overfitting very quickly. So I've gone back to using lesser images, 200k approx.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 763615,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "03/04/2020 17:01:17",
          "content": "<p>Using that technique I would suggest only 160k images.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "758071": "The dataset is class imbalanced, that is, REAL:FAKE=1:5. I currently use under sample to make sure that the number of REAL and FAKE videos is the same, but this keeps most FAKE videos from being exploited and somewhat limits the generalization of model. I tried methods like weight loss and over sampler, but still perform worse than under sample. Does anyone know how to use the entire data set for training? Thank you very much!",
    "758213": "Man, you are in 6th place with underbalancing, that is truely fascinating, well, we do have our best score with underbalancing, however yesterday, I had a success in making a data, which I say that any data better than this is not possible for this competition, atleast theoratically, I will give you a hint \"1 frame should only appear once\", without underbalacing, think about that you might get it, Good Luck!",
    "758338": "An easy way around this, although I'm not sure how effective it is, is too just cycle through the fakes throughout the batches/epochs. You create a list with every fake face and cycles through this list throughout the epochs of training. For example, if you have 20k real faces and 100k fake faces your data per epoch would be:\n\nEpoch 1: 20k real + 0 to 20k from your fake list.\nEpoch 2: 20k real + 20k to 40k from your fake list.\n....\nEpoch 5: 20k real + 80k to 100k from your fake list.\nEpoch 6: 20k real + 0 to 20k from your fake list.\n\nAnd so on. I'm not sure if this has any advantage over just sampling 20k from the fake list every epoch, but theoretically it would allow you to use all data, somehow, at least.",
    "758346": "really not effective, its just normal oversampling, in our case it did not even give any advantage over undersampling.",
    "758790": "Thank you very much! \nBTW, Could you please tell me whether \"1 frame\" means 1 frame per video or the order of frames in the video?",
    "758794": "I also think this approach may not work, which inevitably making the model too focused much on the new coming FAKE images.",
    "758922": "1 frame means a single frame never repeats in the dataset, the dataset is the images of different frames from different videos, but super-highly balanced. (Its not a simple multiplier trick)",
    "759378": "how about taking more frame samples from real videos vs. fake videos?",
    "759482": "The strategy I am going to try next is in the picture. It won't include every fake image but the important consideration is frames between real and fake are aligned to eliminate bias.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F389345%2F170613a872b3ffa08a054cacd33063c1%2Fstrategy.jpg?generation=1582952448470919&amp;alt=media)",
    "759547": "I've tried this method, but it doesn't work.",
    "759550": "Thank you. I'm trying and hope it works.",
    "759716": "Already been trying this method for a while now. Hasn't really provided any gains.",
    "759760": "This method, just has to be theoratically it, you must be doing a flaw in that, I did try it once, I am gonna do it again, as I got a bug the last time I tried.",
    "763532": "That's interesting.. If I may ask, what's the idea behind this? Do you believe there is a correlation between the Nth frame of a video with the frame N of the other videos in the dataset?\nBtw, did this work for you guys? @chenshen03 @harshitsheoran",
    "763533": "I mean, thats litterally the best we could do, @arc144 it does work out for us as it has no bottleneck even when using full data, there is a correlation but not with any video to any video, Trust me, you can get a lot of information from the metadata file.",
    "763570": "Well, I noticed that if I used more than 500K approx train images using the above method, my model would start overfitting very quickly. So I've gone back to using lesser images, 200k approx.",
    "763615": "Using that technique I would suggest only 160k images.",
    "765876": "do your model work in unbalance data?@Chason",
    "766337": "chenbaoying No, I'm still using balanced data, which constructed by under sample.",
    "766395": "I treat the excess fakes as \"data augmentation\" so in different epochs, while I use the same reals, the fakes are similar not identical. The histogram in this case is a useful diagnostic - it becomes a bit asymmetrical, preferring fake over real (although the the test videos have equal real/fake). This is in spite of a score that's not too bad (~ 0.37).",
    "767316": "harshitsheoran : May I ask do you mean this: suppose video A as 4 fakes: B, C, D, E. Then suppose I have frames from 1 to 8, equally gapped. Then I can take frame 1, 3, 5, 7 from video A; frame 2 from video B; frame 4 from video C, frame 6 from video D, and frame 8 from video E. Is that what you hint?",
    "767579": "khahuras Good Job, You are really pretty close, this is one of finest data I had ever created and currently I am using a very similar data to what you are saying, however, yesterday, I found a big bottleneck in this technique, Yes, that was very close to my hint, but still you can get much more information from metadata file than you think you can.\n\nEDIT: The technique you defined above should not have a big performance difference than mine, Good Luck!",
    "767649": "Thanks @harshitsheoran . Indeed, I have been doing something like this: for each epoch, for each shuffled index of all real videos, select 1 pair of (real, fake). The fake is randomly selected from the corresponding fakes of that real video. So, in 1 epoch, the total number of samples would be 2*nb_real_videos. The samples in each epoch are not fixed and built beforehand, but indeed are randomly sampled after every epoch (hence different). And different epochs have different fake versions because of this way. I'm still quite confused why this doesn't work. I think by randomly selecting 1 fake, it acts as a kind of augmentation. I may need to re-think why.",
    "767650": "khahuras Sry your technique, I can not say it is the best, there are a lot more better than this one.",
    "768957": "chenshen03 How many frams per videos in your training, and how many frames in your kaggle submition"
  },
  "source": "meta"
}