{
  "id": 129866,
  "title": "Dealing with unbalanced data?",
  "url": "/competitions/deepfake-detection-challenge/discussion/129866",
  "author_name": "",
  "post_date": "2020-02-11T03:53:34.059045500Z",
  "votes": 6,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I am currently using underbalancing. VERY simple: Just simple <code>random.sample(fake,len(real))</code>\nSeveral Methods Available:\nUnderbalancing\nOverbalancing\n<a href=\"https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-14-106\">SMOTE</a>\n<a href=\"https://journal.r-project.org/archive/2014/RJ-2014-008/RJ-2014-008.pdf\">ROSE</a>\nWhich one are you going with?</p>",
  "messages": [
    {
      "id": "742052",
      "postDate": "02/11/2020 03:53:34",
      "content": "<p>I am currently using underbalancing. VERY simple: Just simple <code>random.sample(fake,len(real))</code>\nSeveral Methods Available:\nUnderbalancing\nOverbalancing\n<a href=\"https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-14-106\">SMOTE</a>\n<a href=\"https://journal.r-project.org/archive/2014/RJ-2014-008/RJ-2014-008.pdf\">ROSE</a>\nWhich one are you going with?</p>",
      "rawMarkdown": "I am currently using underbalancing. VERY simple: Just simple `random.sample(fake,len(real))`\nSeveral Methods Available:\nUnderbalancing\nOverbalancing\n[SMOTE](https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-14-106)\n[ROSE](https://journal.r-project.org/archive/2014/RJ-2014-008/RJ-2014-008.pdf)\nWhich one are you going with?",
      "votes": null
    },
    {
      "id": "742059",
      "postDate": "02/11/2020 03:59:49",
      "content": "<p>Currently, I'm trying replacement sampler provided by PyTorch.</p>",
      "rawMarkdown": "Currently, I'm trying replacement sampler provided by PyTorch.",
      "votes": null
    },
    {
      "id": "744619",
      "postDate": "02/13/2020 01:45:13",
      "content": "<p>Oversampling real class</p>",
      "rawMarkdown": "Oversampling real class",
      "votes": null
    },
    {
      "id": "744985",
      "postDate": "02/13/2020 11:03:05",
      "content": "<p>undersampling but on every epoch, I alternate between corresponding fake videos. So eventually it is close to overbalancing. </p>",
      "rawMarkdown": "undersampling but on every epoch, I alternate between corresponding fake videos. So eventually it is close to overbalancing.",
      "votes": null
    },
    {
      "id": "745001",
      "postDate": "02/13/2020 11:35:26",
      "content": "<p>In each batch, I load 50% real videos, then the other 50% is sampled from the fakes for each real video in the batch. So the batch always has (real, fake) pairs. Then the batch is shuffled so that the target labels aren't always (0, 1, 0, 1, ...). (This is done using a custom collate function in PyTorch and the standard data loaders.)</p>",
      "rawMarkdown": "In each batch, I load 50% real videos, then the other 50% is sampled from the fakes for each real video in the batch. So the batch always has (real, fake) pairs. Then the batch is shuffled so that the target labels aren't always (0, 1, 0, 1, ...). (This is done using a custom collate function in PyTorch and the standard data loaders.)",
      "votes": null
    },
    {
      "id": "745042",
      "postDate": "02/13/2020 12:28:04",
      "content": "<p>Why would it matter if the target labels are (0,1,0,1...) ?</p>",
      "rawMarkdown": "Why would it matter if the target labels are (0,1,0,1...) ?",
      "votes": null
    },
    {
      "id": "745082",
      "postDate": "02/13/2020 13:07:15",
      "content": "<p>Don't want the model to learn that pattern and then assume every even image is 0 and every odd image is 1.</p>",
      "rawMarkdown": "Don't want the model to learn that pattern and then assume every even image is 0 and every odd image is 1.",
      "votes": null
    },
    {
      "id": "745086",
      "postDate": "02/13/2020 13:14:36",
      "content": "<p>That isn't possible unless you are using LSTM or similar layers.</p>",
      "rawMarkdown": "That isn't possible unless you are using LSTM or similar layers.",
      "votes": null
    },
    {
      "id": "745190",
      "postDate": "02/13/2020 15:43:31",
      "content": "<p>Thanks for this information, may be I should try shuffling.</p>",
      "rawMarkdown": "Thanks for this information, may be I should try shuffling.",
      "votes": null
    },
    {
      "id": "745612",
      "postDate": "02/14/2020 01:59:45",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> I am doing the same thing as you are. However, I do not automatically include a fake corresponding to each real - the fakes are all randomly sampled, basically independently of the reals. Why would forcing the pairing improve things? (Note, though, that I DO mandate that I cannot have a real/fake pair on opposite sides of the cv/train boundary) to mitigate leakage, although I do still appear to have some leakage.</p>",
      "rawMarkdown": "humananalog I am doing the same thing as you are. However, I do not automatically include a fake corresponding to each real - the fakes are all randomly sampled, basically independently of the reals. Why would forcing the pairing improve things? (Note, though, that I DO mandate that I cannot have a real/fake pair on opposite sides of the cv/train boundary) to mitigate leakage, although I do still appear to have some leakage.",
      "votes": null
    },
    {
      "id": "745685",
      "postDate": "02/14/2020 04:48:23",
      "content": "<p><a href=\"/ankitsainiankit\">@ankitsainiankit</a> your approach is giving me the best scores.</p>",
      "rawMarkdown": "ankitsainiankit your approach is giving me the best scores.",
      "votes": null
    },
    {
      "id": "745692",
      "postDate": "02/14/2020 05:02:23",
      "content": "<p><a href=\"/ankitsainiankit\">@ankitsainiankit</a> Do you also do this while validation?</p>",
      "rawMarkdown": "ankitsainiankit Do you also do this while validation?",
      "votes": null
    },
    {
      "id": "745785",
      "postDate": "02/14/2020 07:44:15",
      "content": "<p>yes, as of now I'm doing with my validation too. But I don't think it makes a good difference in Validation. Simple undersampling gives the same results.</p>",
      "rawMarkdown": "yes, as of now I'm doing with my validation too. But I don't think it makes a good difference in Validation. Simple undersampling gives the same results.",
      "votes": null
    },
    {
      "id": "745955",
      "postDate": "02/14/2020 12:25:39",
      "content": "<blockquote>\n  <p>Why would forcing the pairing improve things?</p>\n</blockquote>\n\n<p>I don't really know except that it made a big improvement in my score. </p>\n\n<p>Maybe it is because random sampling might miss the videos of which there are only few fakes. Some videos have lots of fakes, some only have one. By always combining a real face and one of its fakes in the same batch, you make sure that the model sees at least one fake for every real face.</p>\n\n<p>Or maybe it's because doing it in the same batch, rather than in different batches that might be spread far apart in time, will create a nicer gradient update. By this I mean: when making a prediction for the real face and its corresponding fake face in the same batch, the parameters of the model are the same at that point. And so the gradient update takes into account that these two images are related. If you don't do it in the same batch, the model will see the real and fake face at different points in time, and their gradients won't reinforce each other. (But this is speculation on my part.)</p>\n\n<p>I even tried to enforce this even further: if the real image has a predicted probability P, then its fake image (in that batch) should have been predicted with 1 - P. And vice versa. If not, I added an extra penalty to the loss. You can see that this penalty indeed becomes smaller over time. (But I'm not sure it actually helped in the end.)</p>",
      "rawMarkdown": "&gt; Why would forcing the pairing improve things?\n\nI don't really know except that it made a big improvement in my score. \n\nMaybe it is because random sampling might miss the videos of which there are only few fakes. Some videos have lots of fakes, some only have one. By always combining a real face and one of its fakes in the same batch, you make sure that the model sees at least one fake for every real face.\n\nOr maybe it's because doing it in the same batch, rather than in different batches that might be spread far apart in time, will create a nicer gradient update. By this I mean: when making a prediction for the real face and its corresponding fake face in the same batch, the parameters of the model are the same at that point. And so the gradient update takes into account that these two images are related. If you don't do it in the same batch, the model will see the real and fake face at different points in time, and their gradients won't reinforce each other. (But this is speculation on my part.)\n\nI even tried to enforce this even further: if the real image has a predicted probability P, then its fake image (in that batch) should have been predicted with 1 - P. And vice versa. If not, I added an extra penalty to the loss. You can see that this penalty indeed becomes smaller over time. (But I'm not sure it actually helped in the end.)",
      "votes": null
    },
    {
      "id": "746402",
      "postDate": "02/15/2020 00:37:04",
      "content": "<p>Interesting - I like the \"nicer update\" explanation. I actually was doing this a while back (I had a data loader that returned both real and one of the fakes each time it was called) but an unrelated and since corrected  bug convinced me to stop doing it. Maybe I'll go back to that method (but if I do, I would add your method of removing the 101010 pattern)..</p>",
      "rawMarkdown": "Interesting - I like the \"nicer update\" explanation. I actually was doing this a while back (I had a data loader that returned both real and one of the fakes each time it was called) but an unrelated and since corrected  bug convinced me to stop doing it. Maybe I'll go back to that method (but if I do, I would add your method of removing the 101010 pattern)..",
      "votes": null
    },
    {
      "id": "761985",
      "postDate": "03/03/2020 05:20:45",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> I wonder how you implemented your batch-sampler of getting half real images and its <strong>*<em>corresponding</em>*</strong> fake images for the other half? I tried but failed to implement. Is it possible to share your sample code for doing this(via kernel)?</p>",
      "rawMarkdown": "humananalog I wonder how you implemented your batch-sampler of getting half real images and its ****corresponding**** fake images for the other half? I tried but failed to implement. Is it possible to share your sample code for doing this(via kernel)?",
      "votes": null
    },
    {
      "id": "762210",
      "postDate": "03/03/2020 10:23:27",
      "content": "<p>I created a <code>Dataset</code> subclass that returns the real image and one of its fake images, plus their labels. This dataset class has a dictionary that lets me look up a list of all the fake images for each real image. So the <code>__getitem()__</code> function in the dataset gets the real image at the current index, then looks in the dictionary to get the list of corresponding fake images, then randomly picks one of those. It returns both images.</p>\n\n<p>In addition, I created a custom collate function that takes these 2 images and their labels, and stacks them into a new tensor that is twice as big as the batch size, then shuffles it. The output of this collate function is two tensors: one with the real and fake faces in it, and the other one with their target labels.</p>",
      "rawMarkdown": "I created a `Dataset` subclass that returns the real image and one of its fake images, plus their labels. This dataset class has a dictionary that lets me look up a list of all the fake images for each real image. So the `__getitem()__` function in the dataset gets the real image at the current index, then looks in the dictionary to get the list of corresponding fake images, then randomly picks one of those. It returns both images.\n\nIn addition, I created a custom collate function that takes these 2 images and their labels, and stacks them into a new tensor that is twice as big as the batch size, then shuffles it. The output of this collate function is two tensors: one with the real and fake faces in it, and the other one with their target labels.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 742059,
      "author_name": "phunghieu",
      "author_url": "",
      "post_date": "02/11/2020 03:59:49",
      "content": "<p>Currently, I'm trying replacement sampler provided by PyTorch.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 744619,
      "author_name": "maralski",
      "author_url": "",
      "post_date": "02/13/2020 01:45:13",
      "content": "<p>Oversampling real class</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 744985,
      "author_name": "ankitsainiankit",
      "author_url": "",
      "post_date": "02/13/2020 11:03:05",
      "content": "<p>undersampling but on every epoch, I alternate between corresponding fake videos. So eventually it is close to overbalancing. </p>",
      "votes": null,
      "replies": [
        {
          "id": 745685,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/14/2020 04:48:23",
          "content": "<p><a href=\"/ankitsainiankit\">@ankitsainiankit</a> your approach is giving me the best scores.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745692,
          "author_name": "unkownhihi",
          "author_url": "",
          "post_date": "02/14/2020 05:02:23",
          "content": "<p><a href=\"/ankitsainiankit\">@ankitsainiankit</a> Do you also do this while validation?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745785,
          "author_name": "ankitsainiankit",
          "author_url": "",
          "post_date": "02/14/2020 07:44:15",
          "content": "<p>yes, as of now I'm doing with my validation too. But I don't think it makes a good difference in Validation. Simple undersampling gives the same results.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745001,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "02/13/2020 11:35:26",
      "content": "<p>In each batch, I load 50% real videos, then the other 50% is sampled from the fakes for each real video in the batch. So the batch always has (real, fake) pairs. Then the batch is shuffled so that the target labels aren't always (0, 1, 0, 1, ...). (This is done using a custom collate function in PyTorch and the standard data loaders.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 745042,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "02/13/2020 12:28:04",
          "content": "<p>Why would it matter if the target labels are (0,1,0,1...) ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745082,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/13/2020 13:07:15",
          "content": "<p>Don't want the model to learn that pattern and then assume every even image is 0 and every odd image is 1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745086,
          "author_name": "akashnandi",
          "author_url": "",
          "post_date": "02/13/2020 13:14:36",
          "content": "<p>That isn't possible unless you are using LSTM or similar layers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745190,
          "author_name": "debanga",
          "author_url": "",
          "post_date": "02/13/2020 15:43:31",
          "content": "<p>Thanks for this information, may be I should try shuffling.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 761985,
          "author_name": "bibek777",
          "author_url": "",
          "post_date": "03/03/2020 05:20:45",
          "content": "<p><a href=\"/humananalog\">@humananalog</a> I wonder how you implemented your batch-sampler of getting half real images and its <strong>*<em>corresponding</em>*</strong> fake images for the other half? I tried but failed to implement. Is it possible to share your sample code for doing this(via kernel)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 762210,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "03/03/2020 10:23:27",
          "content": "<p>I created a <code>Dataset</code> subclass that returns the real image and one of its fake images, plus their labels. This dataset class has a dictionary that lets me look up a list of all the fake images for each real image. So the <code>__getitem()__</code> function in the dataset gets the real image at the current index, then looks in the dictionary to get the list of corresponding fake images, then randomly picks one of those. It returns both images.</p>\n\n<p>In addition, I created a custom collate function that takes these 2 images and their labels, and stacks them into a new tensor that is twice as big as the batch size, then shuffles it. The output of this collate function is two tensors: one with the real and fake faces in it, and the other one with their target labels.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745612,
      "author_name": "petewills",
      "author_url": "",
      "post_date": "02/14/2020 01:59:45",
      "content": "<p><a href=\"/humananalog\">@humananalog</a> I am doing the same thing as you are. However, I do not automatically include a fake corresponding to each real - the fakes are all randomly sampled, basically independently of the reals. Why would forcing the pairing improve things? (Note, though, that I DO mandate that I cannot have a real/fake pair on opposite sides of the cv/train boundary) to mitigate leakage, although I do still appear to have some leakage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 745955,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "02/14/2020 12:25:39",
          "content": "<blockquote>\n  <p>Why would forcing the pairing improve things?</p>\n</blockquote>\n\n<p>I don't really know except that it made a big improvement in my score. </p>\n\n<p>Maybe it is because random sampling might miss the videos of which there are only few fakes. Some videos have lots of fakes, some only have one. By always combining a real face and one of its fakes in the same batch, you make sure that the model sees at least one fake for every real face.</p>\n\n<p>Or maybe it's because doing it in the same batch, rather than in different batches that might be spread far apart in time, will create a nicer gradient update. By this I mean: when making a prediction for the real face and its corresponding fake face in the same batch, the parameters of the model are the same at that point. And so the gradient update takes into account that these two images are related. If you don't do it in the same batch, the model will see the real and fake face at different points in time, and their gradients won't reinforce each other. (But this is speculation on my part.)</p>\n\n<p>I even tried to enforce this even further: if the real image has a predicted probability P, then its fake image (in that batch) should have been predicted with 1 - P. And vice versa. If not, I added an extra penalty to the loss. You can see that this penalty indeed becomes smaller over time. (But I'm not sure it actually helped in the end.)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746402,
          "author_name": "petewills",
          "author_url": "",
          "post_date": "02/15/2020 00:37:04",
          "content": "<p>Interesting - I like the \"nicer update\" explanation. I actually was doing this a while back (I had a data loader that returned both real and one of the fakes each time it was called) but an unrelated and since corrected  bug convinced me to stop doing it. Maybe I'll go back to that method (but if I do, I would add your method of removing the 101010 pattern)..</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "742052": "I am currently using underbalancing. VERY simple: Just simple `random.sample(fake,len(real))`\nSeveral Methods Available:\nUnderbalancing\nOverbalancing\n[SMOTE](https://bmcbioinformatics.biomedcentral.com/articles/10.1186/1471-2105-14-106)\n[ROSE](https://journal.r-project.org/archive/2014/RJ-2014-008/RJ-2014-008.pdf)\nWhich one are you going with?",
    "742059": "Currently, I'm trying replacement sampler provided by PyTorch.",
    "744619": "Oversampling real class",
    "744985": "undersampling but on every epoch, I alternate between corresponding fake videos. So eventually it is close to overbalancing.",
    "745001": "In each batch, I load 50% real videos, then the other 50% is sampled from the fakes for each real video in the batch. So the batch always has (real, fake) pairs. Then the batch is shuffled so that the target labels aren't always (0, 1, 0, 1, ...). (This is done using a custom collate function in PyTorch and the standard data loaders.)",
    "745042": "Why would it matter if the target labels are (0,1,0,1...) ?",
    "745082": "Don't want the model to learn that pattern and then assume every even image is 0 and every odd image is 1.",
    "745086": "That isn't possible unless you are using LSTM or similar layers.",
    "745190": "Thanks for this information, may be I should try shuffling.",
    "745612": "humananalog I am doing the same thing as you are. However, I do not automatically include a fake corresponding to each real - the fakes are all randomly sampled, basically independently of the reals. Why would forcing the pairing improve things? (Note, though, that I DO mandate that I cannot have a real/fake pair on opposite sides of the cv/train boundary) to mitigate leakage, although I do still appear to have some leakage.",
    "745685": "ankitsainiankit your approach is giving me the best scores.",
    "745692": "ankitsainiankit Do you also do this while validation?",
    "745785": "yes, as of now I'm doing with my validation too. But I don't think it makes a good difference in Validation. Simple undersampling gives the same results.",
    "745955": "&gt; Why would forcing the pairing improve things?\n\nI don't really know except that it made a big improvement in my score. \n\nMaybe it is because random sampling might miss the videos of which there are only few fakes. Some videos have lots of fakes, some only have one. By always combining a real face and one of its fakes in the same batch, you make sure that the model sees at least one fake for every real face.\n\nOr maybe it's because doing it in the same batch, rather than in different batches that might be spread far apart in time, will create a nicer gradient update. By this I mean: when making a prediction for the real face and its corresponding fake face in the same batch, the parameters of the model are the same at that point. And so the gradient update takes into account that these two images are related. If you don't do it in the same batch, the model will see the real and fake face at different points in time, and their gradients won't reinforce each other. (But this is speculation on my part.)\n\nI even tried to enforce this even further: if the real image has a predicted probability P, then its fake image (in that batch) should have been predicted with 1 - P. And vice versa. If not, I added an extra penalty to the loss. You can see that this penalty indeed becomes smaller over time. (But I'm not sure it actually helped in the end.)",
    "746402": "Interesting - I like the \"nicer update\" explanation. I actually was doing this a while back (I had a data loader that returned both real and one of the fakes each time it was called) but an unrelated and since corrected  bug convinced me to stop doing it. Maybe I'll go back to that method (but if I do, I would add your method of removing the 101010 pattern)..",
    "761985": "humananalog I wonder how you implemented your batch-sampler of getting half real images and its ****corresponding**** fake images for the other half? I tried but failed to implement. Is it possible to share your sample code for doing this(via kernel)?",
    "762210": "I created a `Dataset` subclass that returns the real image and one of its fake images, plus their labels. This dataset class has a dictionary that lets me look up a list of all the fake images for each real image. So the `__getitem()__` function in the dataset gets the real image at the current index, then looks in the dictionary to get the list of corresponding fake images, then randomly picks one of those. It returns both images.\n\nIn addition, I created a custom collate function that takes these 2 images and their labels, and stacks them into a new tensor that is twice as big as the batch size, then shuffles it. The output of this collate function is two tensors: one with the real and fake faces in it, and the other one with their target labels."
  },
  "source": "meta"
}