{
  "id": 47134,
  "title": "Mixup aug simple drop in code ",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47134",
  "author_name": "",
  "post_date": "2018-01-09T05:12:44.082000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I have used the mixup technique introduced in the <a href=\"https://arxiv.org/abs/1710.09412\">paper</a> and explained in the <a href=\"http://www.inference.vc/mixup-data-dependent-data-augmentation/\">blog</a>. I have implemented the blog version as it is simpler and need no changes in existing arch. It leads to 0.5% improvement of validation score in a single model. Hope some of you may find it helpful. This implementation is in pytorch. It wraps an existing Dataset object like train dataset and can be used as:</p>\n\n<pre><code>trainDataset = MixupDataset(trainDataset, alpha=0.2)\n</code></pre>\n\n<p>Rest of the code remains same(no change in criterion). I have used this over a MFCC dataset with existing augmentations. Haven't tried with raw audio waveforms. Also semi supervised setting didn't help me much. If anyone wants to team up let me know.</p>\n\n<pre><code>import numpy as np\nimport random\nclass MixupDataset(Dataset):\n\n    def __init__(self, dataset, mixup_dataset=None, alpha=None):\n        self.dataset = dataset\n        self.alpha = alpha\n        self.mixup_dataset = mixup_dataset if mixup_dataset else dataset\n\n\n    def __getitem__(self, index):\n        x,y = self.dataset.__getitem__(index)\n        if self.alpha and self.alpha!=0:\n            x2,_ = self.mixup_dataset[random.randint(0,len(self.mixup_dataset)-1)]\n            lam = np.random.beta(self.alpha+1, self.alpha)\n            x = lam * x + (1. - lam) * x2   \n        return x,y\n\n    def __len__(self):\n            return self.dataset.__len__()\n</code></pre>",
  "messages": [
    {
      "id": 266587,
      "postDate": "2018-01-09T05:12:44.083Z",
      "content": "<p>I have used the mixup technique introduced in the <a href=\"https://arxiv.org/abs/1710.09412\">paper</a> and explained in the <a href=\"http://www.inference.vc/mixup-data-dependent-data-augmentation/\">blog</a>. I have implemented the blog version as it is simpler and need no changes in existing arch. It leads to 0.5% improvement of validation score in a single model. Hope some of you may find it helpful. This implementation is in pytorch. It wraps an existing Dataset object like train dataset and can be used as:</p>\n\n<pre><code>trainDataset = MixupDataset(trainDataset, alpha=0.2)\n</code></pre>\n\n<p>Rest of the code remains same(no change in criterion). I have used this over a MFCC dataset with existing augmentations. Haven't tried with raw audio waveforms. Also semi supervised setting didn't help me much. If anyone wants to team up let me know.</p>\n\n<pre><code>import numpy as np\nimport random\nclass MixupDataset(Dataset):\n\n    def __init__(self, dataset, mixup_dataset=None, alpha=None):\n        self.dataset = dataset\n        self.alpha = alpha\n        self.mixup_dataset = mixup_dataset if mixup_dataset else dataset\n\n\n    def __getitem__(self, index):\n        x,y = self.dataset.__getitem__(index)\n        if self.alpha and self.alpha!=0:\n            x2,_ = self.mixup_dataset[random.randint(0,len(self.mixup_dataset)-1)]\n            lam = np.random.beta(self.alpha+1, self.alpha)\n            x = lam * x + (1. - lam) * x2   \n        return x,y\n\n    def __len__(self):\n            return self.dataset.__len__()\n</code></pre>",
      "rawMarkdown": "I have used the mixup technique introduced in the [paper][1] and explained in the [blog][2]. I have implemented the blog version as it is simpler and need no changes in existing arch. It leads to 0.5% improvement of validation score in a single model. Hope some of you may find it helpful. This implementation is in pytorch. It wraps an existing Dataset object like train dataset and can be used as:\n\n    trainDataset = MixupDataset(trainDataset, alpha=0.2)\n\nRest of the code remains same(no change in criterion). I have used this over a MFCC dataset with existing augmentations. Haven't tried with raw audio waveforms. Also semi supervised setting didn't help me much. If anyone wants to team up let me know.\n\n    import numpy as np\n    import random\n    class MixupDataset(Dataset):\n\n        def __init__(self, dataset, mixup_dataset=None, alpha=None):\n            self.dataset = dataset\n            self.alpha = alpha\n            self.mixup_dataset = mixup_dataset if mixup_dataset else dataset\n            \n            \n        def __getitem__(self, index):\n            x,y = self.dataset.__getitem__(index)\n            if self.alpha and self.alpha!=0:\n                x2,_ = self.mixup_dataset[random.randint(0,len(self.mixup_dataset)-1)]\n                lam = np.random.beta(self.alpha+1, self.alpha)\n                x = lam * x + (1. - lam) * x2   \n            return x,y\n        \n        def __len__(self):\n                return self.dataset.__len__()\n\n \n\n\n  [1]: https://arxiv.org/abs/1710.09412\n  [2]: http://www.inference.vc/mixup-data-dependent-data-augmentation/",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "266587": "I have used the mixup technique introduced in the [paper][1] and explained in the [blog][2]. I have implemented the blog version as it is simpler and need no changes in existing arch. It leads to 0.5% improvement of validation score in a single model. Hope some of you may find it helpful. This implementation is in pytorch. It wraps an existing Dataset object like train dataset and can be used as:\n\n    trainDataset = MixupDataset(trainDataset, alpha=0.2)\n\nRest of the code remains same(no change in criterion). I have used this over a MFCC dataset with existing augmentations. Haven't tried with raw audio waveforms. Also semi supervised setting didn't help me much. If anyone wants to team up let me know.\n\n    import numpy as np\n    import random\n    class MixupDataset(Dataset):\n\n        def __init__(self, dataset, mixup_dataset=None, alpha=None):\n            self.dataset = dataset\n            self.alpha = alpha\n            self.mixup_dataset = mixup_dataset if mixup_dataset else dataset\n            \n            \n        def __getitem__(self, index):\n            x,y = self.dataset.__getitem__(index)\n            if self.alpha and self.alpha!=0:\n                x2,_ = self.mixup_dataset[random.randint(0,len(self.mixup_dataset)-1)]\n                lam = np.random.beta(self.alpha+1, self.alpha)\n                x = lam * x + (1. - lam) * x2   \n            return x,y\n        \n        def __len__(self):\n                return self.dataset.__len__()\n\n \n\n\n  [1]: https://arxiv.org/abs/1710.09412\n  [2]: http://www.inference.vc/mixup-data-dependent-data-augmentation/"
  }
}