{
  "id": 165212,
  "title": "Imbalanced Dataset Sampler! Better Sampler! Bert balance dataset",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/165212",
  "author_name": "",
  "post_date": "2020-07-08T20:17:49.738265Z",
  "votes": 3,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In many machine learning applications, we often come across datasets where some types of data may be seen more than other types. Take identification of rare diseases for example, there are probably more normal samples than disease ones. In these cases, we need to make sure that the trained model is not biased towards the class that has more data. As an example, consider a dataset where there are 5 disease images and 20 normal images. If the model predicts all images to be normal, its accuracy is 80%, and F1-score of such a model is 0.88. Therefore, the model has high tendency to be biased toward the ‘normal’ class.</p>\n\n<p>To solve this problem, a widely adopted technique is called resampling. It consists of removing samples from the majority class (under-sampling) and / or adding more examples from the minority class (over-sampling). Despite the advantage of balancing classes, these techniques also have their weaknesses (there is no free lunch). The simplest implementation of over-sampling is to duplicate random records from the minority class, which can cause overfitting. In under-sampling, the simplest technique involves removing random records from the majority class, which can cause loss of information.\nIn this repo, we implement an easy-to-use PyTorch sampler ImbalancedDatasetSampler that is able to</p>\n\n<ul>\n<li>rebalance the class distributions when sampling from the imbalanced dataset</li>\n<li>estimate the sampling weights automatically</li>\n<li>avoid creating a new balanced dataset</li>\n<li>mitigate overfitting when it is used in conjunction with data augmentation techniques\n```\nimport torch\nimport torch.utils.data\nimport torchvision</li>\n</ul>\n\n<p>class ImbalancedDatasetSampler(torch.utils.data.sampler.Sampler):\n    \"\"\"Samples elements randomly from a given list of indices for imbalanced dataset\n    Arguments:\n        indices (list, optional): a list of indices\n        num_samples (int, optional): number of samples to draw\n        callback_get_label func: a callback-like function which takes two arguments - dataset and index\n    \"\"\"</p>\n\n<pre><code>def __init__(self, dataset, indices=None, num_samples=None, callback_get_label=None):\n\n    # if indices is not provided, \n    # all elements in the dataset will be considered\n    self.indices = list(range(len(dataset))) \\\n        if indices is None else indices\n\n    # define custom callback\n    self.callback_get_label = callback_get_label\n\n    # if num_samples is not provided, \n    # draw `len(indices)` samples in each iteration\n    self.num_samples = len(self.indices) \\\n        if num_samples is None else num_samples\n\n    # distribution of classes in the dataset \n    label_to_count = {}\n    for idx in self.indices:\n        label = self._get_label(dataset, idx)\n        if label in label_to_count:\n            label_to_count[label] += 1\n        else:\n            label_to_count[label] = 1\n\n    # weight for each sample\n    weights = [1.0 / label_to_count[self._get_label(dataset, idx)]\n               for idx in self.indices]\n    self.weights = torch.DoubleTensor(weights)\n\ndef _get_label(self, dataset, idx):  \n    return dataset.train_labels[idx].item()\n\ndef __iter__(self):\n    return (self.indices[i] for i in torch.multinomial(\n        self.weights, self.num_samples, replacement=True))\n\ndef __len__(self):\n    return self.num_samples\n</code></pre>\n\n<p>```</p>",
  "messages": [
    {
      "id": "920821",
      "postDate": "07/08/2020 20:17:49",
      "content": "<p>In many machine learning applications, we often come across datasets where some types of data may be seen more than other types. Take identification of rare diseases for example, there are probably more normal samples than disease ones. In these cases, we need to make sure that the trained model is not biased towards the class that has more data. As an example, consider a dataset where there are 5 disease images and 20 normal images. If the model predicts all images to be normal, its accuracy is 80%, and F1-score of such a model is 0.88. Therefore, the model has high tendency to be biased toward the ‘normal’ class.</p>\n\n<p>To solve this problem, a widely adopted technique is called resampling. It consists of removing samples from the majority class (under-sampling) and / or adding more examples from the minority class (over-sampling). Despite the advantage of balancing classes, these techniques also have their weaknesses (there is no free lunch). The simplest implementation of over-sampling is to duplicate random records from the minority class, which can cause overfitting. In under-sampling, the simplest technique involves removing random records from the majority class, which can cause loss of information.\nIn this repo, we implement an easy-to-use PyTorch sampler ImbalancedDatasetSampler that is able to</p>\n\n<ul>\n<li>rebalance the class distributions when sampling from the imbalanced dataset</li>\n<li>estimate the sampling weights automatically</li>\n<li>avoid creating a new balanced dataset</li>\n<li>mitigate overfitting when it is used in conjunction with data augmentation techniques\n```\nimport torch\nimport torch.utils.data\nimport torchvision</li>\n</ul>\n\n<p>class ImbalancedDatasetSampler(torch.utils.data.sampler.Sampler):\n    \"\"\"Samples elements randomly from a given list of indices for imbalanced dataset\n    Arguments:\n        indices (list, optional): a list of indices\n        num_samples (int, optional): number of samples to draw\n        callback_get_label func: a callback-like function which takes two arguments - dataset and index\n    \"\"\"</p>\n\n<pre><code>def __init__(self, dataset, indices=None, num_samples=None, callback_get_label=None):\n\n    # if indices is not provided, \n    # all elements in the dataset will be considered\n    self.indices = list(range(len(dataset))) \\\n        if indices is None else indices\n\n    # define custom callback\n    self.callback_get_label = callback_get_label\n\n    # if num_samples is not provided, \n    # draw `len(indices)` samples in each iteration\n    self.num_samples = len(self.indices) \\\n        if num_samples is None else num_samples\n\n    # distribution of classes in the dataset \n    label_to_count = {}\n    for idx in self.indices:\n        label = self._get_label(dataset, idx)\n        if label in label_to_count:\n            label_to_count[label] += 1\n        else:\n            label_to_count[label] = 1\n\n    # weight for each sample\n    weights = [1.0 / label_to_count[self._get_label(dataset, idx)]\n               for idx in self.indices]\n    self.weights = torch.DoubleTensor(weights)\n\ndef _get_label(self, dataset, idx):  \n    return dataset.train_labels[idx].item()\n\ndef __iter__(self):\n    return (self.indices[i] for i in torch.multinomial(\n        self.weights, self.num_samples, replacement=True))\n\ndef __len__(self):\n    return self.num_samples\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "In many machine learning applications, we often come across datasets where some types of data may be seen more than other types. Take identification of rare diseases for example, there are probably more normal samples than disease ones. In these cases, we need to make sure that the trained model is not biased towards the class that has more data. As an example, consider a dataset where there are 5 disease images and 20 normal images. If the model predicts all images to be normal, its accuracy is 80%, and F1-score of such a model is 0.88. Therefore, the model has high tendency to be biased toward the ‘normal’ class.\n\nTo solve this problem, a widely adopted technique is called resampling. It consists of removing samples from the majority class (under-sampling) and / or adding more examples from the minority class (over-sampling). Despite the advantage of balancing classes, these techniques also have their weaknesses (there is no free lunch). The simplest implementation of over-sampling is to duplicate random records from the minority class, which can cause overfitting. In under-sampling, the simplest technique involves removing random records from the majority class, which can cause loss of information.\nIn this repo, we implement an easy-to-use PyTorch sampler ImbalancedDatasetSampler that is able to\n\n* rebalance the class distributions when sampling from the imbalanced dataset\n* estimate the sampling weights automatically\n* avoid creating a new balanced dataset\n* mitigate overfitting when it is used in conjunction with data augmentation techniques\n```\nimport torch\nimport torch.utils.data\nimport torchvision\n\n\nclass ImbalancedDatasetSampler(torch.utils.data.sampler.Sampler):\n    \"\"\"Samples elements randomly from a given list of indices for imbalanced dataset\n    Arguments:\n        indices (list, optional): a list of indices\n        num_samples (int, optional): number of samples to draw\n        callback_get_label func: a callback-like function which takes two arguments - dataset and index\n    \"\"\"\n\n    def __init__(self, dataset, indices=None, num_samples=None, callback_get_label=None):\n                \n        # if indices is not provided, \n        # all elements in the dataset will be considered\n        self.indices = list(range(len(dataset))) \\\n            if indices is None else indices\n\n        # define custom callback\n        self.callback_get_label = callback_get_label\n\n        # if num_samples is not provided, \n        # draw `len(indices)` samples in each iteration\n        self.num_samples = len(self.indices) \\\n            if num_samples is None else num_samples\n            \n        # distribution of classes in the dataset \n        label_to_count = {}\n        for idx in self.indices:\n            label = self._get_label(dataset, idx)\n            if label in label_to_count:\n                label_to_count[label] += 1\n            else:\n                label_to_count[label] = 1\n                \n        # weight for each sample\n        weights = [1.0 / label_to_count[self._get_label(dataset, idx)]\n                   for idx in self.indices]\n        self.weights = torch.DoubleTensor(weights)\n\n    def _get_label(self, dataset, idx):  \n        return dataset.train_labels[idx].item()\n        \n    def __iter__(self):\n        return (self.indices[i] for i in torch.multinomial(\n            self.weights, self.num_samples, replacement=True))\n\n    def __len__(self):\n        return self.num_samples\n```",
      "votes": null
    },
    {
      "id": "920882",
      "postDate": "07/08/2020 21:57:25",
      "content": "<p>Hi\nNice solution. I also use unbalanced sampling, but in a simpler form. \nMy dataset has <code>classes</code> attribute which is either 0 or 1. </p>\n\n<p>```\npos_weight = 0.6\ndataset = MelanomaDataset()</p>\n\n<h1>Fix class inbalance.</h1>\n\n<p>if pos_weight is None:\n    sample = None\nelse:\n    label_to_weight = {\n        0: 1 - pos_weight,\n        1: pos_weight\n    }\n    weights = [label_to_weight[k] for k in dataset.classes]\n    sampler = torch.utils.data.sampler.WeightedRandomSampler(weights, num_samples=len(train_dataset), replacement=True)\n```</p>\n\n<p>Then this sampler is used in dataloader. </p>\n\n<p>As far as I know, it may be a bad idea to train network with completely balanced sampler (0.5 probability of positive class), cause it make network to produce more positive predictions. At the same time at first epochs balanced (or even oversampled) positive class can help in learning relevant features faster. In last year SIIM-Pneumothorax segmentation winner used different proportion of pos/neg classes with gradual decreasing towards real data distribution. \nThis method can also help here :)</p>",
      "rawMarkdown": "Hi\nNice solution. I also use unbalanced sampling, but in a simpler form. \nMy dataset has `classes` attribute which is either 0 or 1. \n\n```\npos_weight = 0.6\ndataset = MelanomaDataset()\n# Fix class inbalance.\nif pos_weight is None:\n    sample = None\nelse:\n    label_to_weight = {\n        0: 1 - pos_weight,\n        1: pos_weight\n    }\n    weights = [label_to_weight[k] for k in dataset.classes]\n    sampler = torch.utils.data.sampler.WeightedRandomSampler(weights, num_samples=len(train_dataset), replacement=True)\n```\n\nThen this sampler is used in dataloader. \n\nAs far as I know, it may be a bad idea to train network with completely balanced sampler (0.5 probability of positive class), cause it make network to produce more positive predictions. At the same time at first epochs balanced (or even oversampled) positive class can help in learning relevant features faster. In last year SIIM-Pneumothorax segmentation winner used different proportion of pos/neg classes with gradual decreasing towards real data distribution. \nThis method can also help here :)",
      "votes": null
    },
    {
      "id": "920887",
      "postDate": "07/08/2020 22:06:06",
      "content": "<p>Nice idea. I will try to deploy it soon. As i know, in this competition. We have a high imbalance dataset ( 2% vs 98%). So what i do not just 0.5 probability, it change defence the dataset.</p>",
      "rawMarkdown": "Nice idea. I will try to deploy it soon. As i know, in this competition. We have a high imbalance dataset ( 2% vs 98%). So what i do not just 0.5 probability, it change defence the dataset.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 920882,
      "author_name": "zakajd",
      "author_url": "",
      "post_date": "07/08/2020 21:57:25",
      "content": "<p>Hi\nNice solution. I also use unbalanced sampling, but in a simpler form. \nMy dataset has <code>classes</code> attribute which is either 0 or 1. </p>\n\n<p>```\npos_weight = 0.6\ndataset = MelanomaDataset()</p>\n\n<h1>Fix class inbalance.</h1>\n\n<p>if pos_weight is None:\n    sample = None\nelse:\n    label_to_weight = {\n        0: 1 - pos_weight,\n        1: pos_weight\n    }\n    weights = [label_to_weight[k] for k in dataset.classes]\n    sampler = torch.utils.data.sampler.WeightedRandomSampler(weights, num_samples=len(train_dataset), replacement=True)\n```</p>\n\n<p>Then this sampler is used in dataloader. </p>\n\n<p>As far as I know, it may be a bad idea to train network with completely balanced sampler (0.5 probability of positive class), cause it make network to produce more positive predictions. At the same time at first epochs balanced (or even oversampled) positive class can help in learning relevant features faster. In last year SIIM-Pneumothorax segmentation winner used different proportion of pos/neg classes with gradual decreasing towards real data distribution. \nThis method can also help here :)</p>",
      "votes": null,
      "replies": [
        {
          "id": 920887,
          "author_name": "doanquanvietnamca",
          "author_url": "",
          "post_date": "07/08/2020 22:06:06",
          "content": "<p>Nice idea. I will try to deploy it soon. As i know, in this competition. We have a high imbalance dataset ( 2% vs 98%). So what i do not just 0.5 probability, it change defence the dataset.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "920821": "In many machine learning applications, we often come across datasets where some types of data may be seen more than other types. Take identification of rare diseases for example, there are probably more normal samples than disease ones. In these cases, we need to make sure that the trained model is not biased towards the class that has more data. As an example, consider a dataset where there are 5 disease images and 20 normal images. If the model predicts all images to be normal, its accuracy is 80%, and F1-score of such a model is 0.88. Therefore, the model has high tendency to be biased toward the ‘normal’ class.\n\nTo solve this problem, a widely adopted technique is called resampling. It consists of removing samples from the majority class (under-sampling) and / or adding more examples from the minority class (over-sampling). Despite the advantage of balancing classes, these techniques also have their weaknesses (there is no free lunch). The simplest implementation of over-sampling is to duplicate random records from the minority class, which can cause overfitting. In under-sampling, the simplest technique involves removing random records from the majority class, which can cause loss of information.\nIn this repo, we implement an easy-to-use PyTorch sampler ImbalancedDatasetSampler that is able to\n\n* rebalance the class distributions when sampling from the imbalanced dataset\n* estimate the sampling weights automatically\n* avoid creating a new balanced dataset\n* mitigate overfitting when it is used in conjunction with data augmentation techniques\n```\nimport torch\nimport torch.utils.data\nimport torchvision\n\n\nclass ImbalancedDatasetSampler(torch.utils.data.sampler.Sampler):\n    \"\"\"Samples elements randomly from a given list of indices for imbalanced dataset\n    Arguments:\n        indices (list, optional): a list of indices\n        num_samples (int, optional): number of samples to draw\n        callback_get_label func: a callback-like function which takes two arguments - dataset and index\n    \"\"\"\n\n    def __init__(self, dataset, indices=None, num_samples=None, callback_get_label=None):\n                \n        # if indices is not provided, \n        # all elements in the dataset will be considered\n        self.indices = list(range(len(dataset))) \\\n            if indices is None else indices\n\n        # define custom callback\n        self.callback_get_label = callback_get_label\n\n        # if num_samples is not provided, \n        # draw `len(indices)` samples in each iteration\n        self.num_samples = len(self.indices) \\\n            if num_samples is None else num_samples\n            \n        # distribution of classes in the dataset \n        label_to_count = {}\n        for idx in self.indices:\n            label = self._get_label(dataset, idx)\n            if label in label_to_count:\n                label_to_count[label] += 1\n            else:\n                label_to_count[label] = 1\n                \n        # weight for each sample\n        weights = [1.0 / label_to_count[self._get_label(dataset, idx)]\n                   for idx in self.indices]\n        self.weights = torch.DoubleTensor(weights)\n\n    def _get_label(self, dataset, idx):  \n        return dataset.train_labels[idx].item()\n        \n    def __iter__(self):\n        return (self.indices[i] for i in torch.multinomial(\n            self.weights, self.num_samples, replacement=True))\n\n    def __len__(self):\n        return self.num_samples\n```",
    "920882": "Hi\nNice solution. I also use unbalanced sampling, but in a simpler form. \nMy dataset has `classes` attribute which is either 0 or 1. \n\n```\npos_weight = 0.6\ndataset = MelanomaDataset()\n# Fix class inbalance.\nif pos_weight is None:\n    sample = None\nelse:\n    label_to_weight = {\n        0: 1 - pos_weight,\n        1: pos_weight\n    }\n    weights = [label_to_weight[k] for k in dataset.classes]\n    sampler = torch.utils.data.sampler.WeightedRandomSampler(weights, num_samples=len(train_dataset), replacement=True)\n```\n\nThen this sampler is used in dataloader. \n\nAs far as I know, it may be a bad idea to train network with completely balanced sampler (0.5 probability of positive class), cause it make network to produce more positive predictions. At the same time at first epochs balanced (or even oversampled) positive class can help in learning relevant features faster. In last year SIIM-Pneumothorax segmentation winner used different proportion of pos/neg classes with gradual decreasing towards real data distribution. \nThis method can also help here :)",
    "920887": "Nice idea. I will try to deploy it soon. As i know, in this competition. We have a high imbalance dataset ( 2% vs 98%). So what i do not just 0.5 probability, it change defence the dataset."
  },
  "source": "meta"
}