{
  "id": 181162,
  "title": "Code for CutMix / MixUp for audio data",
  "url": "/competitions/birdsong-recognition/discussion/181162",
  "author_name": "",
  "post_date": "2020-09-07T20:42:35.686577700Z",
  "votes": 23,
  "comment_count": 6,
  "views": 0,
  "content": "<p>After a lot of fighting , this competition doesn't seem to be working for me , but not all competitions are there for winning right? I have learned a lot from this competition and before giving up just wanted to share the working implementation of </p>\n<blockquote>\n  <p>MIX-UP and CUT-MIX </p>\n</blockquote>\n<p>with everyone , I hope it helps some people in the community , then I would consider my hard work did pay off </p>\n<p>I thought it would be better to share it here as I have already shared a kernel for other audio augmentations , but cutmix and mixup needed to integrated in the dataset class itself . If my work here serves the purpose and helps people , I will share a working pipeline of EFFNET - B1 model as well </p>\n<blockquote>\n  <p>Below is the code , you can also find the colab notebook where I have used the code in a pipeline <a href=\"https://drive.google.com/file/d/1O3eBTMcqV4d7jp0DhRSFwwBQ2ySLcjga/view?usp=sharing\" target=\"_blank\">here</a></p>\n</blockquote>\n<p>`class SpectrogramDataset(data.Dataset):</p>\n<pre><code>def __init__(self,\n             df: pd.DataFrame,\n             img_size=224,\n             p = 1.0,\n             waveform_transforms=None,\n             spectrogram_transforms=None,\n             melspectrogram_parameters={}):\n    self.df = df\n    self.img_size = img_size\n    self.p=p\n    self.waveform_transforms = waveform_transforms\n    self.spectrogram_transforms = spectrogram_transforms\n    self.melspectrogram_parameters = melspectrogram_parameters\n\n\ndef load_preprocess_audio(self,idx:int):\n\n    sample = self.df.loc[idx, :]\n    file_path = sample[\"file_path\"]\n    ebird_code = sample[\"ebird_code\"]\n\n    y, sr = sf.read(file_path)\n\n    len_y = len(y)\n    effective_length = sr * PERIOD\n    if len_y &lt; effective_length:\n        new_y = np.zeros(effective_length, dtype=y.dtype)\n        start = np.random.randint(effective_length - len_y)\n        new_y[start:start + len_y] = y\n        y = new_y.astype(np.float32)\n    elif len_y &gt; effective_length:\n        start = np.random.randint(len_y - effective_length)\n        y = y[start:start + effective_length].astype(np.float32)\n    else:\n        y = y.astype(np.float32)\n\n    labels = np.zeros(len(BIRD_CODE), dtype=int)\n    labels[BIRD_CODE[ebird_code]] = 1\n\n    return y,labels,sr,ebird_code\n\ndef __len__(self):\n    return len(self.df)\n\ndef __getitem__(self, idx: int):\n\n    if np.random.uniform(0,1) &gt; self.p:\n        y,labels,sr,_ = self.load_preprocess_audio(idx)\n    else:\n        y,labels,sr = self.cutmix(idx)\n\n    melspec = librosa.feature.melspectrogram(y, sr=sr, **self.melspectrogram_parameters)\n    melspec = librosa.power_to_db(melspec).astype(np.float32)\n\n    if self.spectrogram_transforms:\n        melspec = self.spectrogram_transforms(melspec)\n    else:\n        pass\n\n    image = mono_to_color(melspec)\n    height, width, _ = image.shape\n    image = cv2.resize(image, (int(width * self.img_size / height), self.img_size))\n    image = np.moveaxis(image, 2, 0)\n    image = (image / 255.0).astype(np.float32)\n\n    return {\n        \"image\": image,\n        \"targets\": labels\n    }\n\n\ndef cutmix(self,idx:int):\n    #Load and preprocess the given audio\n    audio,label,sr,ebird_code = self.load_preprocess_audio(idx)\n    #Choose random audio from all files \n    mixing_audio_index = np.random.randint(0,len(self.df)-1)\n    #Loading the audio that we will use for mixup\n    audio_mix,label_mix,sr_mix,ebird_code_mix = self.load_preprocess_audio(mixing_audio_index)\n\n    #Determining the indexes where we will cut out the part and insert the new audio\n    start_ = np.random.randint(0,len(audio))\n    end_ = np.random.randint(start_,len(audio))\n    #Now cutting and mixing up\n    audio[start_:end_] = audio_mix[start_:end_]\n    # Adjusting Labels\n    if label.argmax() == label_mix.argmax():\n        new_label = label\n    else:\n        new_label = label + label_mix   # Since they were already one-hot-encoded\n\n    return audio,new_label,sr\n\ndef mixup(self,idx:int):\n    #Load and preprocess the given audio\n    audio,label,sr = self.load_preprocess_audio(idx)\n    #Choose random audio from all files\n    mixing_audio_index = np.random.randint(0,len(self.df))\n    #Loading the audio that we will use for mixup\n    audio_mix,label_mix,sr_mix = self.load_preprocess_audio(mixing_audio_index)\n    #creating mixup audio\n    audio = audio + audio_mix*0.005\n    # Adjusting Labels\n    if label.argmax() == label_mix.argmax():\n        new_label = label\n    else:\n        new_label = label + label_mix   # Since they were already one-hot-encoded\n\n    return audio,new_label,sr`\n</code></pre>\n<p>You can use either of cutmix or mixup and approach this as multilabel problem , you can also use different audios to add as no call data and then use cutmix and mixup to approach as multilabel columns . I hope this helps</p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": "1002108",
      "postDate": "09/07/2020 20:42:35",
      "content": "<p>After a lot of fighting , this competition doesn't seem to be working for me , but not all competitions are there for winning right? I have learned a lot from this competition and before giving up just wanted to share the working implementation of </p>\n<blockquote>\n  <p>MIX-UP and CUT-MIX </p>\n</blockquote>\n<p>with everyone , I hope it helps some people in the community , then I would consider my hard work did pay off </p>\n<p>I thought it would be better to share it here as I have already shared a kernel for other audio augmentations , but cutmix and mixup needed to integrated in the dataset class itself . If my work here serves the purpose and helps people , I will share a working pipeline of EFFNET - B1 model as well </p>\n<blockquote>\n  <p>Below is the code , you can also find the colab notebook where I have used the code in a pipeline <a href=\"https://drive.google.com/file/d/1O3eBTMcqV4d7jp0DhRSFwwBQ2ySLcjga/view?usp=sharing\" target=\"_blank\">here</a></p>\n</blockquote>\n<p>`class SpectrogramDataset(data.Dataset):</p>\n<pre><code>def __init__(self,\n             df: pd.DataFrame,\n             img_size=224,\n             p = 1.0,\n             waveform_transforms=None,\n             spectrogram_transforms=None,\n             melspectrogram_parameters={}):\n    self.df = df\n    self.img_size = img_size\n    self.p=p\n    self.waveform_transforms = waveform_transforms\n    self.spectrogram_transforms = spectrogram_transforms\n    self.melspectrogram_parameters = melspectrogram_parameters\n\n\ndef load_preprocess_audio(self,idx:int):\n\n    sample = self.df.loc[idx, :]\n    file_path = sample[\"file_path\"]\n    ebird_code = sample[\"ebird_code\"]\n\n    y, sr = sf.read(file_path)\n\n    len_y = len(y)\n    effective_length = sr * PERIOD\n    if len_y &lt; effective_length:\n        new_y = np.zeros(effective_length, dtype=y.dtype)\n        start = np.random.randint(effective_length - len_y)\n        new_y[start:start + len_y] = y\n        y = new_y.astype(np.float32)\n    elif len_y &gt; effective_length:\n        start = np.random.randint(len_y - effective_length)\n        y = y[start:start + effective_length].astype(np.float32)\n    else:\n        y = y.astype(np.float32)\n\n    labels = np.zeros(len(BIRD_CODE), dtype=int)\n    labels[BIRD_CODE[ebird_code]] = 1\n\n    return y,labels,sr,ebird_code\n\ndef __len__(self):\n    return len(self.df)\n\ndef __getitem__(self, idx: int):\n\n    if np.random.uniform(0,1) &gt; self.p:\n        y,labels,sr,_ = self.load_preprocess_audio(idx)\n    else:\n        y,labels,sr = self.cutmix(idx)\n\n    melspec = librosa.feature.melspectrogram(y, sr=sr, **self.melspectrogram_parameters)\n    melspec = librosa.power_to_db(melspec).astype(np.float32)\n\n    if self.spectrogram_transforms:\n        melspec = self.spectrogram_transforms(melspec)\n    else:\n        pass\n\n    image = mono_to_color(melspec)\n    height, width, _ = image.shape\n    image = cv2.resize(image, (int(width * self.img_size / height), self.img_size))\n    image = np.moveaxis(image, 2, 0)\n    image = (image / 255.0).astype(np.float32)\n\n    return {\n        \"image\": image,\n        \"targets\": labels\n    }\n\n\ndef cutmix(self,idx:int):\n    #Load and preprocess the given audio\n    audio,label,sr,ebird_code = self.load_preprocess_audio(idx)\n    #Choose random audio from all files \n    mixing_audio_index = np.random.randint(0,len(self.df)-1)\n    #Loading the audio that we will use for mixup\n    audio_mix,label_mix,sr_mix,ebird_code_mix = self.load_preprocess_audio(mixing_audio_index)\n\n    #Determining the indexes where we will cut out the part and insert the new audio\n    start_ = np.random.randint(0,len(audio))\n    end_ = np.random.randint(start_,len(audio))\n    #Now cutting and mixing up\n    audio[start_:end_] = audio_mix[start_:end_]\n    # Adjusting Labels\n    if label.argmax() == label_mix.argmax():\n        new_label = label\n    else:\n        new_label = label + label_mix   # Since they were already one-hot-encoded\n\n    return audio,new_label,sr\n\ndef mixup(self,idx:int):\n    #Load and preprocess the given audio\n    audio,label,sr = self.load_preprocess_audio(idx)\n    #Choose random audio from all files\n    mixing_audio_index = np.random.randint(0,len(self.df))\n    #Loading the audio that we will use for mixup\n    audio_mix,label_mix,sr_mix = self.load_preprocess_audio(mixing_audio_index)\n    #creating mixup audio\n    audio = audio + audio_mix*0.005\n    # Adjusting Labels\n    if label.argmax() == label_mix.argmax():\n        new_label = label\n    else:\n        new_label = label + label_mix   # Since they were already one-hot-encoded\n\n    return audio,new_label,sr`\n</code></pre>\n<p>You can use either of cutmix or mixup and approach this as multilabel problem , you can also use different audios to add as no call data and then use cutmix and mixup to approach as multilabel columns . I hope this helps</p>\n<p>Thanks</p>",
      "rawMarkdown": "After a lot of fighting , this competition doesn't seem to be working for me , but not all competitions are there for winning right? I have learned a lot from this competition and before giving up just wanted to share the working implementation of \n> MIX-UP and CUT-MIX \n\nwith everyone , I hope it helps some people in the community , then I would consider my hard work did pay off \n\nI thought it would be better to share it here as I have already shared a kernel for other audio augmentations , but cutmix and mixup needed to integrated in the dataset class itself . If my work here serves the purpose and helps people , I will share a working pipeline of EFFNET - B1 model as well \n\n> Below is the code , you can also find the colab notebook where I have used the code in a pipeline [here](https://drive.google.com/file/d/1O3eBTMcqV4d7jp0DhRSFwwBQ2ySLcjga/view?usp=sharing)\n\n\n`class SpectrogramDataset(data.Dataset):\n\n    def __init__(self,\n                 df: pd.DataFrame,\n                 img_size=224,\n                 p = 1.0,\n                 waveform_transforms=None,\n                 spectrogram_transforms=None,\n                 melspectrogram_parameters={}):\n        self.df = df\n        self.img_size = img_size\n        self.p=p\n        self.waveform_transforms = waveform_transforms\n        self.spectrogram_transforms = spectrogram_transforms\n        self.melspectrogram_parameters = melspectrogram_parameters\n        \n        \n    def load_preprocess_audio(self,idx:int):\n        \n        sample = self.df.loc[idx, :]\n        file_path = sample[\"file_path\"]\n        ebird_code = sample[\"ebird_code\"]\n        \n        y, sr = sf.read(file_path)\n\n        len_y = len(y)\n        effective_length = sr * PERIOD\n        if len_y < effective_length:\n            new_y = np.zeros(effective_length, dtype=y.dtype)\n            start = np.random.randint(effective_length - len_y)\n            new_y[start:start + len_y] = y\n            y = new_y.astype(np.float32)\n        elif len_y > effective_length:\n            start = np.random.randint(len_y - effective_length)\n            y = y[start:start + effective_length].astype(np.float32)\n        else:\n            y = y.astype(np.float32)\n            \n        labels = np.zeros(len(BIRD_CODE), dtype=int)\n        labels[BIRD_CODE[ebird_code]] = 1\n            \n        return y,labels,sr,ebird_code\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx: int):\n        \n        if np.random.uniform(0,1) > self.p:\n            y,labels,sr,_ = self.load_preprocess_audio(idx)\n        else:\n            y,labels,sr = self.cutmix(idx)\n                        \n        melspec = librosa.feature.melspectrogram(y, sr=sr, **self.melspectrogram_parameters)\n        melspec = librosa.power_to_db(melspec).astype(np.float32)\n\n        if self.spectrogram_transforms:\n            melspec = self.spectrogram_transforms(melspec)\n        else:\n            pass\n\n        image = mono_to_color(melspec)\n        height, width, _ = image.shape\n        image = cv2.resize(image, (int(width * self.img_size / height), self.img_size))\n        image = np.moveaxis(image, 2, 0)\n        image = (image / 255.0).astype(np.float32)\n        \n        return {\n            \"image\": image,\n            \"targets\": labels\n        }\n    \n    \n    def cutmix(self,idx:int):\n        #Load and preprocess the given audio\n        audio,label,sr,ebird_code = self.load_preprocess_audio(idx)\n        #Choose random audio from all files \n        mixing_audio_index = np.random.randint(0,len(self.df)-1)\n        #Loading the audio that we will use for mixup\n        audio_mix,label_mix,sr_mix,ebird_code_mix = self.load_preprocess_audio(mixing_audio_index)\n        \n        #Determining the indexes where we will cut out the part and insert the new audio\n        start_ = np.random.randint(0,len(audio))\n        end_ = np.random.randint(start_,len(audio))\n        #Now cutting and mixing up\n        audio[start_:end_] = audio_mix[start_:end_]\n        # Adjusting Labels\n        if label.argmax() == label_mix.argmax():\n            new_label = label\n        else:\n            new_label = label + label_mix   # Since they were already one-hot-encoded\n        \n        return audio,new_label,sr\n        \n    def mixup(self,idx:int):\n        #Load and preprocess the given audio\n        audio,label,sr = self.load_preprocess_audio(idx)\n        #Choose random audio from all files\n        mixing_audio_index = np.random.randint(0,len(self.df))\n        #Loading the audio that we will use for mixup\n        audio_mix,label_mix,sr_mix = self.load_preprocess_audio(mixing_audio_index)\n        #creating mixup audio\n        audio = audio + audio_mix*0.005\n        # Adjusting Labels\n        if label.argmax() == label_mix.argmax():\n            new_label = label\n        else:\n            new_label = label + label_mix   # Since they were already one-hot-encoded\n        \n        return audio,new_label,sr`\n\nYou can use either of cutmix or mixup and approach this as multilabel problem , you can also use different audios to add as no call data and then use cutmix and mixup to approach as multilabel columns . I hope this helps\n\nThanks",
      "votes": null
    },
    {
      "id": "1002111",
      "postDate": "09/07/2020 20:51:42",
      "content": "<p>Did cutmix/mixup help you?</p>",
      "rawMarkdown": "Did cutmix/mixup help you?",
      "votes": null
    },
    {
      "id": "1002502",
      "postDate": "09/08/2020 07:25:35",
      "content": "<p>NIce</p>",
      "rawMarkdown": "NIce",
      "votes": null
    },
    {
      "id": "1002659",
      "postDate": "09/08/2020 10:15:20",
      "content": "<p>Well I Can't say exactly , it did increase the cv score , but the lb was disappointing</p>",
      "rawMarkdown": "Well I Can't say exactly , it did increase the cv score , but the lb was disappointing",
      "votes": null
    },
    {
      "id": "1004053",
      "postDate": "09/09/2020 12:37:25",
      "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> don't give up till the end, you have 1 week from now. In one week anything can happen. Keep doing your work.</p>",
      "rawMarkdown": "tanulsingh077 don't give up till the end, you have 1 week from now. In one week anything can happen. Keep doing your work.",
      "votes": null
    },
    {
      "id": "1004840",
      "postDate": "09/10/2020 04:35:05",
      "content": "<p>Do you normalized the original audio and the mixed audio? or you just add them together?<br>\nAnd I'm confused that the label should be just two-hot or smoothing one(in your case[0 , 0, 1, 0, 0.005])<br>\nI tried some methods and increase in my cv a little bit, but lb is so bad:(<br>\nThanks~</p>",
      "rawMarkdown": "Do you normalized the original audio and the mixed audio? or you just add them together?\nAnd I'm confused that the label should be just two-hot or smoothing one(in your case[0 , 0, 1, 0, 0.005])\nI tried some methods and increase in my cv a little bit, but lb is so bad:(\nThanks~",
      "votes": null
    },
    {
      "id": "2238160",
      "postDate": "04/28/2023 08:54:44",
      "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> Thanks for sharing you notebook and this post as well. I am learning from it a lot! I can't open colab file you mentioned though. I know its been 3 years, but any chances its still there?😀</p>",
      "rawMarkdown": "tanulsingh077 Thanks for sharing you notebook and this post as well. I am learning from it a lot! I can't open colab file you mentioned though. I know its been 3 years, but any chances its still there?😀",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1002111,
      "author_name": "mightyrains",
      "author_url": "",
      "post_date": "09/07/2020 20:51:42",
      "content": "<p>Did cutmix/mixup help you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1002659,
          "author_name": "tanulsingh077",
          "author_url": "",
          "post_date": "09/08/2020 10:15:20",
          "content": "<p>Well I Can't say exactly , it did increase the cv score , but the lb was disappointing</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1004053,
          "author_name": "validmodel",
          "author_url": "",
          "post_date": "09/09/2020 12:37:25",
          "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> don't give up till the end, you have 1 week from now. In one week anything can happen. Keep doing your work.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1004840,
      "author_name": "joshspchang",
      "author_url": "",
      "post_date": "09/10/2020 04:35:05",
      "content": "<p>Do you normalized the original audio and the mixed audio? or you just add them together?<br>\nAnd I'm confused that the label should be just two-hot or smoothing one(in your case[0 , 0, 1, 0, 0.005])<br>\nI tried some methods and increase in my cv a little bit, but lb is so bad:(<br>\nThanks~</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2238160,
      "author_name": "azamat25",
      "author_url": "",
      "post_date": "04/28/2023 08:54:44",
      "content": "<p><a href=\"https://www.kaggle.com/tanulsingh077\" target=\"_blank\">@tanulsingh077</a> Thanks for sharing you notebook and this post as well. I am learning from it a lot! I can't open colab file you mentioned though. I know its been 3 years, but any chances its still there?😀</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1002502,
      "author_name": "aavinashvijay",
      "author_url": "",
      "post_date": "09/08/2020 07:25:35",
      "content": "<p>NIce</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1002108": "After a lot of fighting , this competition doesn't seem to be working for me , but not all competitions are there for winning right? I have learned a lot from this competition and before giving up just wanted to share the working implementation of \n> MIX-UP and CUT-MIX \n\nwith everyone , I hope it helps some people in the community , then I would consider my hard work did pay off \n\nI thought it would be better to share it here as I have already shared a kernel for other audio augmentations , but cutmix and mixup needed to integrated in the dataset class itself . If my work here serves the purpose and helps people , I will share a working pipeline of EFFNET - B1 model as well \n\n> Below is the code , you can also find the colab notebook where I have used the code in a pipeline [here](https://drive.google.com/file/d/1O3eBTMcqV4d7jp0DhRSFwwBQ2ySLcjga/view?usp=sharing)\n\n\n`class SpectrogramDataset(data.Dataset):\n\n    def __init__(self,\n                 df: pd.DataFrame,\n                 img_size=224,\n                 p = 1.0,\n                 waveform_transforms=None,\n                 spectrogram_transforms=None,\n                 melspectrogram_parameters={}):\n        self.df = df\n        self.img_size = img_size\n        self.p=p\n        self.waveform_transforms = waveform_transforms\n        self.spectrogram_transforms = spectrogram_transforms\n        self.melspectrogram_parameters = melspectrogram_parameters\n        \n        \n    def load_preprocess_audio(self,idx:int):\n        \n        sample = self.df.loc[idx, :]\n        file_path = sample[\"file_path\"]\n        ebird_code = sample[\"ebird_code\"]\n        \n        y, sr = sf.read(file_path)\n\n        len_y = len(y)\n        effective_length = sr * PERIOD\n        if len_y < effective_length:\n            new_y = np.zeros(effective_length, dtype=y.dtype)\n            start = np.random.randint(effective_length - len_y)\n            new_y[start:start + len_y] = y\n            y = new_y.astype(np.float32)\n        elif len_y > effective_length:\n            start = np.random.randint(len_y - effective_length)\n            y = y[start:start + effective_length].astype(np.float32)\n        else:\n            y = y.astype(np.float32)\n            \n        labels = np.zeros(len(BIRD_CODE), dtype=int)\n        labels[BIRD_CODE[ebird_code]] = 1\n            \n        return y,labels,sr,ebird_code\n\n    def __len__(self):\n        return len(self.df)\n\n    def __getitem__(self, idx: int):\n        \n        if np.random.uniform(0,1) > self.p:\n            y,labels,sr,_ = self.load_preprocess_audio(idx)\n        else:\n            y,labels,sr = self.cutmix(idx)\n                        \n        melspec = librosa.feature.melspectrogram(y, sr=sr, **self.melspectrogram_parameters)\n        melspec = librosa.power_to_db(melspec).astype(np.float32)\n\n        if self.spectrogram_transforms:\n            melspec = self.spectrogram_transforms(melspec)\n        else:\n            pass\n\n        image = mono_to_color(melspec)\n        height, width, _ = image.shape\n        image = cv2.resize(image, (int(width * self.img_size / height), self.img_size))\n        image = np.moveaxis(image, 2, 0)\n        image = (image / 255.0).astype(np.float32)\n        \n        return {\n            \"image\": image,\n            \"targets\": labels\n        }\n    \n    \n    def cutmix(self,idx:int):\n        #Load and preprocess the given audio\n        audio,label,sr,ebird_code = self.load_preprocess_audio(idx)\n        #Choose random audio from all files \n        mixing_audio_index = np.random.randint(0,len(self.df)-1)\n        #Loading the audio that we will use for mixup\n        audio_mix,label_mix,sr_mix,ebird_code_mix = self.load_preprocess_audio(mixing_audio_index)\n        \n        #Determining the indexes where we will cut out the part and insert the new audio\n        start_ = np.random.randint(0,len(audio))\n        end_ = np.random.randint(start_,len(audio))\n        #Now cutting and mixing up\n        audio[start_:end_] = audio_mix[start_:end_]\n        # Adjusting Labels\n        if label.argmax() == label_mix.argmax():\n            new_label = label\n        else:\n            new_label = label + label_mix   # Since they were already one-hot-encoded\n        \n        return audio,new_label,sr\n        \n    def mixup(self,idx:int):\n        #Load and preprocess the given audio\n        audio,label,sr = self.load_preprocess_audio(idx)\n        #Choose random audio from all files\n        mixing_audio_index = np.random.randint(0,len(self.df))\n        #Loading the audio that we will use for mixup\n        audio_mix,label_mix,sr_mix = self.load_preprocess_audio(mixing_audio_index)\n        #creating mixup audio\n        audio = audio + audio_mix*0.005\n        # Adjusting Labels\n        if label.argmax() == label_mix.argmax():\n            new_label = label\n        else:\n            new_label = label + label_mix   # Since they were already one-hot-encoded\n        \n        return audio,new_label,sr`\n\nYou can use either of cutmix or mixup and approach this as multilabel problem , you can also use different audios to add as no call data and then use cutmix and mixup to approach as multilabel columns . I hope this helps\n\nThanks",
    "1002111": "Did cutmix/mixup help you?",
    "1002502": "NIce",
    "1002659": "Well I Can't say exactly , it did increase the cv score , but the lb was disappointing",
    "1004053": "tanulsingh077 don't give up till the end, you have 1 week from now. In one week anything can happen. Keep doing your work.",
    "1004840": "Do you normalized the original audio and the mixed audio? or you just add them together?\nAnd I'm confused that the label should be just two-hot or smoothing one(in your case[0 , 0, 1, 0, 0.005])\nI tried some methods and increase in my cv a little bit, but lb is so bad:(\nThanks~",
    "2238160": "tanulsingh077 Thanks for sharing you notebook and this post as well. I am learning from it a lot! I can't open colab file you mentioned though. I know its been 3 years, but any chances its still there?😀"
  },
  "source": "meta"
}