{
  "id": 95409,
  "title": "9th place solution: smaller and faster",
  "url": "/competitions/freesound-audio-tagging-2019/writeups/4-people-9th-place-solution-smaller-and-faster",
  "author_name": "",
  "post_date": "2019-06-28T18:21:42.353Z",
  "votes": 35,
  "comment_count": 20,
  "views": 0,
  "content": "<p>First, I would like to congratulate all winers and participants of this competition and thank kaggle and the organizers for posting this interesting challenge and providing resources for participating in it.  Also, I want to express my highest gratitude and appreciation to my teammates: <a href=\"/suicaokhoailang\">@suicaokhoailang</a>, <a href=\"/theoviel\">@theoviel</a>, and <a href=\"/yanglehan\">@yanglehan</a> for working with me on this challenge.  It was the first time when I worked with audio related task, and it was a month of quite intense learning. I would like to give special thanks to <a href=\"/theoviel\">@theoviel</a> for working quite hard with me on optimization of our models and prepare the final submissions.</p>\n\n<p>The code used for training our final models is available now at <a href=\"https://www.kaggle.com/theoviel/9th-place-modeling-kernel\">this link</a> (private LB 0.72171 score, top 42, with a single model (5 CV folds) trained for 7 hours). Further details of our solution can be found in our <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/563780/13697/DCASE2019_Challenge.pdf\">DCASE2019 technical report</a>.</p>\n\n<p>The <strong>key points</strong> of our solutions are the following (see details below): (1) Use of <strong>small and fast models with optimized architecture</strong>, (2) Use of <strong>64 mels</strong> instead of 128 (sometimes less gives more) with <strong>4s duration</strong>, (3) Use <strong>data augmentation and noisy data</strong> for pretraining.</p>\n\n<p><strong>Stage 1</strong>: we started the competition, like many participants, with experimenting with common computer vision models. The input size of spectrograms was 256x512 pixels, with upscaling the input image along first dimension by a factor of 2. With this setup the best performance was demonstrated by DenseNet models: they outperformed the baseline model published in <a href=\"https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch\">this kernel</a> and successfully used in the competition a year ago, also Dnet121 was faster. With using pretraining on full noisy set, spectral augmentation, and MixUp, CV could reach ~0.85+, and public score for a single fold, 4 folds, and ensemble of 4 models are 0.67-0.68, ~0.70, 0.717, respectively. Despite these models are not used for our submission, these experiments have provided important insights for the next stage.</p>\n\n<p><strong>Stage 2</strong>: \n<strong>Use of noisy data</strong>: It is the main point of this competition that organizers wanted us to focus on (no external data, no pretrained models, no test data use policies). We have used 2 strategies: (1) pretraining on full noisy data and (2) pretraining on a mixture of the curated data with most confidently labeled noisy data. In both cases the pretraining is followed by fine tuning on curated only data. The most confident labels are identified based on a model trained on curated data, and further details can be provided by <a href=\"/theoviel\">@theoviel</a>. For our best setup we have the following values of CV (in stage 2 we use 5 fold scheme): training on curated data only - 0.858, pretraining on full noisy data - 0.866, curated data + 15k best noisy data examples - 0.865, curated data + 5k best noisy data examples - 0.872. We have utilized all 3 strategies of noisy data use to create a variety in our ensemble.</p>\n\n<p><strong>Preprocessing</strong>: According to our experiments, big models, like Dnet121, work better on 128 mels and even higher image resolution, while the default model reaches the best performance for 64 mels. This setup also decreases training time and improves convergence speed. 32 mels also could be considered, but the performence drops to 0.856 for our best setup. Use of 4s intervals instead of traditional 2s has gave also a considerable boost. The input image size for the model is 64x256x1. We tried both normalization of data based on image and global train set statistics, and the results were similar. Though, our best CV is reached for global normalization. In final models we used both strategies to crease a diversity. We also tried to experiment with the fft window size but did not see a significant difference and stayed with 1920. One thing to try we didn't have time for is using different window sizes mels as channels of the produced image. In particular, <a href=\"https://arxiv.org/abs/1706.07156\">this paper</a> shows that some classes prefer longer while other shorter window size. The preprocessing pipeline is similar to one described in <a href=\"https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\">this kernel</a>.</p>\n\n<p><strong>Model architecture</strong>: At stage 2 we used the model from <a href=\"https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch\">this kernel</a> as a starting point. The performance of this base model for our best setup is 0.855 CV. Based on our prior positive experience with DensNet, we added dense connections inside convolution blocks and concatenate pooling that boosted the performance to 0.868 in our best experiments (model M1):\n```\nclass ConvBlock(nn.Module):\n    def <strong>init</strong>(self, in_channels, out_channels, kernel_size=3, pool=True):\n        super().<strong>init</strong>()</p>\n\n<pre><code>    padding = kernel_size // 2\n    self.pool = pool\n\n    self.conv1 = nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size,\n            stride=1, padding=padding),\n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(),\n    )\n    self.conv2 = nn.Sequential(\n        nn.Conv2d(out_channels + in_channels, out_channels, \n            kernel_size=kernel_size, stride=1, padding=padding),\n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(),\n    )\n\ndef forward(self, x): # x.shape = [batch_size, in_channels, a, b]\n    x1 = self.conv1(x)\n    x = self.conv2(torch.cat([x, x1],1))\n    if(self.pool): x = F.avg_pool2d(x, 2)\n    return x   # x.shape = [batch_size, out_channels, a//2, b//2]\n</code></pre>\n\n<p>```\nThe increase of the number of convolution blocks from 4 to 5 gave only 0.865 CV. Use of a pyramidal pooling for 2,3 and 4-th conv blocks (M2) gave slightly worse result than M1. Finally, our ultimate setup (M3) consists of 5 conv blocks with pyramidal pooling reached 0.872 CV. DenseNet121 in the same pipeline reached only 0.836 (DenseNet121 requires higher image resolution, 256x512, to reach 0.85+ CV). From the experiments, it looks for audio it is important to have nonlinear operations before size reduction by pooling, though we did not check it in details. We used M1, M2, and M3 to create variability in our ensemble. Because of the submission limit we checked performance of only a few models in public LB, with the best single model score 0.715.</p>\n\n<p>This plot illustrates architecture of M3 model:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F4c15c1ec24d235c7834cc1ae94a3ca38%2FM3.png?generation=1561588228369529&amp;alt=media\" alt=\"\"></p>\n\n<p>Here, all tested models are summarized:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ffaca75bb8fb9dc910a648d1c141313be%2Fmodels.png?generation=1561746099636890&amp;alt=media\" alt=\"\"></p>\n\n<p>And here the model performance for our best setup:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ff7014dbf8de75685fff13c683384f068%2Fp.png?generation=1561588425442446&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Data augmentation</strong>:\nThe main thing for working with audio data is MixUp. In contrast to images of real objects, <a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\">sounds are transparent</a>: they do not overshadow each other. Therefore, MixUp is so efficient for audio and gives 0.01-0.015 CV boost. At the stage 1 the best results were achieved for alpha MixUp parameter equal to 1.0, while at the stage 2 we used 0.4. <a href=\"https://arxiv.org/abs/1904.08779\">Spectral augmentation</a> (with 2 masks for frequency and time with the coverage range between zero and 0.15 and 0.3, respectively) gave about 0.005 CV boost. We did not use stretching the spectrograms in time domain because it gave lower model performance. In several model we also used <a href=\"https://arxiv.org/abs/1905.09788\">Multisample Dropout</a> (other models were trained without dropout); though, it decreased CV by ~0.002. We did not apply horizontal flip since it decreased CV and also is not natural: I do not think that people would be able to recognize sounds played from the back. It is the same as training ImageNet with use of vertical flip.</p>\n\n<p><strong>Training</strong>: At the pretraining stage we used one cycle of cosine annealing with warm up. The maximum lr is 0.001, and the number of epochs is ranged between 50 and 70 for different setups. At the stage of fine tuning we applied ReduceLROnPlateau several times to alternate high and low lr. The code is implemented with Pytorch. The total time of training for one fold is 1-2 hours, so the training of entire model is 6-9 hours. Almost all our models were trained at kaggle, and the kernel is at <a href=\"https://www.kaggle.com/theoviel/9th-place-modeling-kernel\">this link</a>.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F51ade3b68a7ab201b15fb9736d618a7b%2Flr.png?generation=1561588570346398&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Submission</strong>: \nUse of 64 mels has allowed us also to shorten the inference time. In particular, each 5-fold model takes only 1-1.5 min for generation of a prediction with 8TTA. Therefore, we were able to use an ensemble of 15 models to generate the final submissions. \nThe final remark about the possible reason of the gap between CV and public LB: <a href=\"https://arxiv.org/pdf/1904.05635v1.pdf\">this paper</a> (Task 1B) reports 0.15 lwlrap drop for test performed on data recorded with different device vs the same device as used for training data. So, the public LB data may be selected for a different device, but we never know if it is true or not.</p>",
  "messages": [
    {
      "id": "550824",
      "postDate": "06/12/2019 04:22:07",
      "content": "<p>First, I would like to congratulate all winers and participants of this competition and thank kaggle and the organizers for posting this interesting challenge and providing resources for participating in it.  Also, I want to express my highest gratitude and appreciation to my teammates: <a href=\"/suicaokhoailang\">@suicaokhoailang</a>, <a href=\"/theoviel\">@theoviel</a>, and <a href=\"/yanglehan\">@yanglehan</a> for working with me on this challenge.  It was the first time when I worked with audio related task, and it was a month of quite intense learning. I would like to give special thanks to <a href=\"/theoviel\">@theoviel</a> for working quite hard with me on optimization of our models and prepare the final submissions.</p>\n\n<p>The code used for training our final models is available now at <a href=\"https://www.kaggle.com/theoviel/9th-place-modeling-kernel\">this link</a> (private LB 0.72171 score, top 42, with a single model (5 CV folds) trained for 7 hours). Further details of our solution can be found in our <a href=\"https://storage.googleapis.com/kaggle-forum-message-attachments/563780/13697/DCASE2019_Challenge.pdf\">DCASE2019 technical report</a>.</p>\n\n<p>The <strong>key points</strong> of our solutions are the following (see details below): (1) Use of <strong>small and fast models with optimized architecture</strong>, (2) Use of <strong>64 mels</strong> instead of 128 (sometimes less gives more) with <strong>4s duration</strong>, (3) Use <strong>data augmentation and noisy data</strong> for pretraining.</p>\n\n<p><strong>Stage 1</strong>: we started the competition, like many participants, with experimenting with common computer vision models. The input size of spectrograms was 256x512 pixels, with upscaling the input image along first dimension by a factor of 2. With this setup the best performance was demonstrated by DenseNet models: they outperformed the baseline model published in <a href=\"https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch\">this kernel</a> and successfully used in the competition a year ago, also Dnet121 was faster. With using pretraining on full noisy set, spectral augmentation, and MixUp, CV could reach ~0.85+, and public score for a single fold, 4 folds, and ensemble of 4 models are 0.67-0.68, ~0.70, 0.717, respectively. Despite these models are not used for our submission, these experiments have provided important insights for the next stage.</p>\n\n<p><strong>Stage 2</strong>: \n<strong>Use of noisy data</strong>: It is the main point of this competition that organizers wanted us to focus on (no external data, no pretrained models, no test data use policies). We have used 2 strategies: (1) pretraining on full noisy data and (2) pretraining on a mixture of the curated data with most confidently labeled noisy data. In both cases the pretraining is followed by fine tuning on curated only data. The most confident labels are identified based on a model trained on curated data, and further details can be provided by <a href=\"/theoviel\">@theoviel</a>. For our best setup we have the following values of CV (in stage 2 we use 5 fold scheme): training on curated data only - 0.858, pretraining on full noisy data - 0.866, curated data + 15k best noisy data examples - 0.865, curated data + 5k best noisy data examples - 0.872. We have utilized all 3 strategies of noisy data use to create a variety in our ensemble.</p>\n\n<p><strong>Preprocessing</strong>: According to our experiments, big models, like Dnet121, work better on 128 mels and even higher image resolution, while the default model reaches the best performance for 64 mels. This setup also decreases training time and improves convergence speed. 32 mels also could be considered, but the performence drops to 0.856 for our best setup. Use of 4s intervals instead of traditional 2s has gave also a considerable boost. The input image size for the model is 64x256x1. We tried both normalization of data based on image and global train set statistics, and the results were similar. Though, our best CV is reached for global normalization. In final models we used both strategies to crease a diversity. We also tried to experiment with the fft window size but did not see a significant difference and stayed with 1920. One thing to try we didn't have time for is using different window sizes mels as channels of the produced image. In particular, <a href=\"https://arxiv.org/abs/1706.07156\">this paper</a> shows that some classes prefer longer while other shorter window size. The preprocessing pipeline is similar to one described in <a href=\"https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data\">this kernel</a>.</p>\n\n<p><strong>Model architecture</strong>: At stage 2 we used the model from <a href=\"https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch\">this kernel</a> as a starting point. The performance of this base model for our best setup is 0.855 CV. Based on our prior positive experience with DensNet, we added dense connections inside convolution blocks and concatenate pooling that boosted the performance to 0.868 in our best experiments (model M1):\n```\nclass ConvBlock(nn.Module):\n    def <strong>init</strong>(self, in_channels, out_channels, kernel_size=3, pool=True):\n        super().<strong>init</strong>()</p>\n\n<pre><code>    padding = kernel_size // 2\n    self.pool = pool\n\n    self.conv1 = nn.Sequential(\n        nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size,\n            stride=1, padding=padding),\n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(),\n    )\n    self.conv2 = nn.Sequential(\n        nn.Conv2d(out_channels + in_channels, out_channels, \n            kernel_size=kernel_size, stride=1, padding=padding),\n        nn.BatchNorm2d(out_channels),\n        nn.ReLU(),\n    )\n\ndef forward(self, x): # x.shape = [batch_size, in_channels, a, b]\n    x1 = self.conv1(x)\n    x = self.conv2(torch.cat([x, x1],1))\n    if(self.pool): x = F.avg_pool2d(x, 2)\n    return x   # x.shape = [batch_size, out_channels, a//2, b//2]\n</code></pre>\n\n<p>```\nThe increase of the number of convolution blocks from 4 to 5 gave only 0.865 CV. Use of a pyramidal pooling for 2,3 and 4-th conv blocks (M2) gave slightly worse result than M1. Finally, our ultimate setup (M3) consists of 5 conv blocks with pyramidal pooling reached 0.872 CV. DenseNet121 in the same pipeline reached only 0.836 (DenseNet121 requires higher image resolution, 256x512, to reach 0.85+ CV). From the experiments, it looks for audio it is important to have nonlinear operations before size reduction by pooling, though we did not check it in details. We used M1, M2, and M3 to create variability in our ensemble. Because of the submission limit we checked performance of only a few models in public LB, with the best single model score 0.715.</p>\n\n<p>This plot illustrates architecture of M3 model:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F4c15c1ec24d235c7834cc1ae94a3ca38%2FM3.png?generation=1561588228369529&amp;alt=media\" alt=\"\"></p>\n\n<p>Here, all tested models are summarized:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ffaca75bb8fb9dc910a648d1c141313be%2Fmodels.png?generation=1561746099636890&amp;alt=media\" alt=\"\"></p>\n\n<p>And here the model performance for our best setup:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ff7014dbf8de75685fff13c683384f068%2Fp.png?generation=1561588425442446&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Data augmentation</strong>:\nThe main thing for working with audio data is MixUp. In contrast to images of real objects, <a href=\"https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd\">sounds are transparent</a>: they do not overshadow each other. Therefore, MixUp is so efficient for audio and gives 0.01-0.015 CV boost. At the stage 1 the best results were achieved for alpha MixUp parameter equal to 1.0, while at the stage 2 we used 0.4. <a href=\"https://arxiv.org/abs/1904.08779\">Spectral augmentation</a> (with 2 masks for frequency and time with the coverage range between zero and 0.15 and 0.3, respectively) gave about 0.005 CV boost. We did not use stretching the spectrograms in time domain because it gave lower model performance. In several model we also used <a href=\"https://arxiv.org/abs/1905.09788\">Multisample Dropout</a> (other models were trained without dropout); though, it decreased CV by ~0.002. We did not apply horizontal flip since it decreased CV and also is not natural: I do not think that people would be able to recognize sounds played from the back. It is the same as training ImageNet with use of vertical flip.</p>\n\n<p><strong>Training</strong>: At the pretraining stage we used one cycle of cosine annealing with warm up. The maximum lr is 0.001, and the number of epochs is ranged between 50 and 70 for different setups. At the stage of fine tuning we applied ReduceLROnPlateau several times to alternate high and low lr. The code is implemented with Pytorch. The total time of training for one fold is 1-2 hours, so the training of entire model is 6-9 hours. Almost all our models were trained at kaggle, and the kernel is at <a href=\"https://www.kaggle.com/theoviel/9th-place-modeling-kernel\">this link</a>.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F51ade3b68a7ab201b15fb9736d618a7b%2Flr.png?generation=1561588570346398&amp;alt=media\" alt=\"\"></p>\n\n<p><strong>Submission</strong>: \nUse of 64 mels has allowed us also to shorten the inference time. In particular, each 5-fold model takes only 1-1.5 min for generation of a prediction with 8TTA. Therefore, we were able to use an ensemble of 15 models to generate the final submissions. \nThe final remark about the possible reason of the gap between CV and public LB: <a href=\"https://arxiv.org/pdf/1904.05635v1.pdf\">this paper</a> (Task 1B) reports 0.15 lwlrap drop for test performed on data recorded with different device vs the same device as used for training data. So, the public LB data may be selected for a different device, but we never know if it is true or not.</p>",
      "rawMarkdown": "First, I would like to congratulate all winers and participants of this competition and thank kaggle and the organizers for posting this interesting challenge and providing resources for participating in it.  Also, I want to express my highest gratitude and appreciation to my teammates: @suicaokhoailang, @theoviel, and @yanglehan for working with me on this challenge.  It was the first time when I worked with audio related task, and it was a month of quite intense learning. I would like to give special thanks to @theoviel for working quite hard with me on optimization of our models and prepare the final submissions.\n\nThe code used for training our final models is available now at [this link](https://www.kaggle.com/theoviel/9th-place-modeling-kernel) (private LB 0.72171 score, top 42, with a single model (5 CV folds) trained for 7 hours). Further details of our solution can be found in our [DCASE2019 technical report](https://storage.googleapis.com/kaggle-forum-message-attachments/563780/13697/DCASE2019_Challenge.pdf).\n\nThe **key points** of our solutions are the following (see details below): (1) Use of **small and fast models with optimized architecture**, (2) Use of **64 mels** instead of 128 (sometimes less gives more) with **4s duration**, (3) Use **data augmentation and noisy data** for pretraining.\n\n**Stage 1**: we started the competition, like many participants, with experimenting with common computer vision models. The input size of spectrograms was 256x512 pixels, with upscaling the input image along first dimension by a factor of 2. With this setup the best performance was demonstrated by DenseNet models: they outperformed the baseline model published in [this kernel](https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch) and successfully used in the competition a year ago, also Dnet121 was faster. With using pretraining on full noisy set, spectral augmentation, and MixUp, CV could reach ~0.85+, and public score for a single fold, 4 folds, and ensemble of 4 models are 0.67-0.68, ~0.70, 0.717, respectively. Despite these models are not used for our submission, these experiments have provided important insights for the next stage.\n\n**Stage 2**: \n**Use of noisy data**: It is the main point of this competition that organizers wanted us to focus on (no external data, no pretrained models, no test data use policies). We have used 2 strategies: (1) pretraining on full noisy data and (2) pretraining on a mixture of the curated data with most confidently labeled noisy data. In both cases the pretraining is followed by fine tuning on curated only data. The most confident labels are identified based on a model trained on curated data, and further details can be provided by @theoviel. For our best setup we have the following values of CV (in stage 2 we use 5 fold scheme): training on curated data only - 0.858, pretraining on full noisy data - 0.866, curated data + 15k best noisy data examples - 0.865, curated data + 5k best noisy data examples - 0.872. We have utilized all 3 strategies of noisy data use to create a variety in our ensemble.\n\n**Preprocessing**: According to our experiments, big models, like Dnet121, work better on 128 mels and even higher image resolution, while the default model reaches the best performance for 64 mels. This setup also decreases training time and improves convergence speed. 32 mels also could be considered, but the performence drops to 0.856 for our best setup. Use of 4s intervals instead of traditional 2s has gave also a considerable boost. The input image size for the model is 64x256x1. We tried both normalization of data based on image and global train set statistics, and the results were similar. Though, our best CV is reached for global normalization. In final models we used both strategies to crease a diversity. We also tried to experiment with the fft window size but did not see a significant difference and stayed with 1920. One thing to try we didn't have time for is using different window sizes mels as channels of the produced image. In particular, [this paper](https://arxiv.org/abs/1706.07156) shows that some classes prefer longer while other shorter window size. The preprocessing pipeline is similar to one described in [this kernel](https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data).\n\n**Model architecture**: At stage 2 we used the model from [this kernel](https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch) as a starting point. The performance of this base model for our best setup is 0.855 CV. Based on our prior positive experience with DensNet, we added dense connections inside convolution blocks and concatenate pooling that boosted the performance to 0.868 in our best experiments (model M1):\n```\nclass ConvBlock(nn.Module):\n    def __init__(self, in_channels, out_channels, kernel_size=3, pool=True):\n        super().__init__()\n        \n        padding = kernel_size // 2\n        self.pool = pool\n        \n        self.conv1 = nn.Sequential(\n            nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size,\n                stride=1, padding=padding),\n            nn.BatchNorm2d(out_channels),\n            nn.ReLU(),\n        )\n        self.conv2 = nn.Sequential(\n            nn.Conv2d(out_channels + in_channels, out_channels, \n                kernel_size=kernel_size, stride=1, padding=padding),\n            nn.BatchNorm2d(out_channels),\n            nn.ReLU(),\n        )\n        \n    def forward(self, x): # x.shape = [batch_size, in_channels, a, b]\n        x1 = self.conv1(x)\n        x = self.conv2(torch.cat([x, x1],1))\n        if(self.pool): x = F.avg_pool2d(x, 2)\n        return x   # x.shape = [batch_size, out_channels, a//2, b//2]\n```\nThe increase of the number of convolution blocks from 4 to 5 gave only 0.865 CV. Use of a pyramidal pooling for 2,3 and 4-th conv blocks (M2) gave slightly worse result than M1. Finally, our ultimate setup (M3) consists of 5 conv blocks with pyramidal pooling reached 0.872 CV. DenseNet121 in the same pipeline reached only 0.836 (DenseNet121 requires higher image resolution, 256x512, to reach 0.85+ CV). From the experiments, it looks for audio it is important to have nonlinear operations before size reduction by pooling, though we did not check it in details. We used M1, M2, and M3 to create variability in our ensemble. Because of the submission limit we checked performance of only a few models in public LB, with the best single model score 0.715.\n\nThis plot illustrates architecture of M3 model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F4c15c1ec24d235c7834cc1ae94a3ca38%2FM3.png?generation=1561588228369529&amp;alt=media)\n\n\nHere, all tested models are summarized:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ffaca75bb8fb9dc910a648d1c141313be%2Fmodels.png?generation=1561746099636890&amp;alt=media)\n\n\n\nAnd here the model performance for our best setup:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ff7014dbf8de75685fff13c683384f068%2Fp.png?generation=1561588425442446&amp;alt=media)\n\n\n**Data augmentation**:\nThe main thing for working with audio data is MixUp. In contrast to images of real objects, [sounds are transparent](https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd): they do not overshadow each other. Therefore, MixUp is so efficient for audio and gives 0.01-0.015 CV boost. At the stage 1 the best results were achieved for alpha MixUp parameter equal to 1.0, while at the stage 2 we used 0.4. [Spectral augmentation](https://arxiv.org/abs/1904.08779) (with 2 masks for frequency and time with the coverage range between zero and 0.15 and 0.3, respectively) gave about 0.005 CV boost. We did not use stretching the spectrograms in time domain because it gave lower model performance. In several model we also used [Multisample Dropout](https://arxiv.org/abs/1905.09788) (other models were trained without dropout); though, it decreased CV by ~0.002. We did not apply horizontal flip since it decreased CV and also is not natural: I do not think that people would be able to recognize sounds played from the back. It is the same as training ImageNet with use of vertical flip.\n\n**Training**: At the pretraining stage we used one cycle of cosine annealing with warm up. The maximum lr is 0.001, and the number of epochs is ranged between 50 and 70 for different setups. At the stage of fine tuning we applied ReduceLROnPlateau several times to alternate high and low lr. The code is implemented with Pytorch. The total time of training for one fold is 1-2 hours, so the training of entire model is 6-9 hours. Almost all our models were trained at kaggle, and the kernel is at [this link](https://www.kaggle.com/theoviel/9th-place-modeling-kernel).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F51ade3b68a7ab201b15fb9736d618a7b%2Flr.png?generation=1561588570346398&amp;alt=media)\n\n\n**Submission**: \nUse of 64 mels has allowed us also to shorten the inference time. In particular, each 5-fold model takes only 1-1.5 min for generation of a prediction with 8TTA. Therefore, we were able to use an ensemble of 15 models to generate the final submissions. \nThe final remark about the possible reason of the gap between CV and public LB: [this paper](https://arxiv.org/pdf/1904.05635v1.pdf) (Task 1B) reports 0.15 lwlrap drop for test performed on data recorded with different device vs the same device as used for training data. So, the public LB data may be selected for a different device, but we never know if it is true or not.",
      "votes": null
    },
    {
      "id": "550875",
      "postDate": "06/12/2019 05:37:17",
      "content": "<p>Thanks for sharing, I didn't even imagine decreasing mels down to 64.\nAnd I agree with mixup is so efficient in audio task, \"sounds are transparent\" - that's it.\nI hope your team's good luck in 2nd stage.</p>",
      "rawMarkdown": "Thanks for sharing, I didn't even imagine decreasing mels down to 64.\nAnd I agree with mixup is so efficient in audio task, \"sounds are transparent\" - that's it.\nI hope your team's good luck in 2nd stage.",
      "votes": null
    },
    {
      "id": "550889",
      "postDate": "06/12/2019 05:54:52",
      "content": "<p>Thanks, u2. And thank you for your kernels, I used one as a starting point for mel preprocessing.</p>",
      "rawMarkdown": "Thanks, u2. And thank you for your kernels, I used one as a starting point for mel preprocessing.",
      "votes": null
    },
    {
      "id": "551052",
      "postDate": "06/12/2019 09:36:09",
      "content": "<p>Thanks so much <a href=\"/iafoss\">@iafoss</a> and all your team for the very detailed solution. I did try 4s (but with 128 mel) and had no improvement, perhaps I have to go back and redo my experiment :)</p>",
      "rawMarkdown": "Thanks so much @iafoss and all your team for the very detailed solution. I did try 4s (but with 128 mel) and had no improvement, perhaps I have to go back and redo my experiment :)",
      "votes": null
    },
    {
      "id": "551099",
      "postDate": "06/12/2019 10:39:27",
      "content": "<p>Thanks for sharing <a href=\"/iafoss\">@iafoss</a> </p>",
      "rawMarkdown": "Thanks for sharing @iafoss",
      "votes": null
    },
    {
      "id": "551143",
      "postDate": "06/12/2019 11:42:15",
      "content": "<p>Thanks for sharing!\nI should have tried 64 mels, it looks promising in your writing ;)\nI wish you the best for this competition</p>",
      "rawMarkdown": "Thanks for sharing!\nI should have tried 64 mels, it looks promising in your writing ;)\nI wish you the best for this competition",
      "votes": null
    },
    {
      "id": "551298",
      "postDate": "06/12/2019 15:05:44",
      "content": "<p><a href=\"/ebouteillon\">@ebouteillon</a>, I wish u also the best luck. The idea suggested by <a href=\"/theoviel\">@theoviel</a> about 64 mels is really unexpected to work. However, when I tried to switch back to 128 mels in our best setup later, I didn't see improvement. The thing could be also that for 128 mels ~100 epochs (in total) was not enough since stage1 Dnet models we trained for 300-400 epochs.</p>",
      "rawMarkdown": "ebouteillon, I wish u also the best luck. The idea suggested by @theoviel about 64 mels is really unexpected to work. However, when I tried to switch back to 128 mels in our best setup later, I didn't see improvement. The thing could be also that for 128 mels ~100 epochs (in total) was not enough since stage1 Dnet models we trained for 300-400 epochs.",
      "votes": null
    },
    {
      "id": "551301",
      "postDate": "06/12/2019 15:08:46",
      "content": "<p>perfect</p>",
      "rawMarkdown": "perfect",
      "votes": null
    },
    {
      "id": "551308",
      "postDate": "06/12/2019 15:19:02",
      "content": "<p>May I ask, have your team used any kind of minority class oversampling?  Have you applied MixUp with some probability or just augmented the whole dataset? Thanks for sharing and good luck in final shakeup!</p>",
      "rawMarkdown": "May I ask, have your team used any kind of minority class oversampling?  Have you applied MixUp with some probability or just augmented the whole dataset? Thanks for sharing and good luck in final shakeup!",
      "votes": null
    },
    {
      "id": "551334",
      "postDate": "06/12/2019 15:42:17",
      "content": "<p>Thanks for sharing the method!!\nAlways learned a lot from you , <a href=\"/iafoss\">@iafoss</a> .\nNice new paper implement ! (Spectral augmentation)\nHope you get a good result for this competition !!</p>",
      "rawMarkdown": "Thanks for sharing the method!!\nAlways learned a lot from you , @iafoss .\nNice new paper implement ! (Spectral augmentation)\nHope you get a good result for this competition !!",
      "votes": null
    },
    {
      "id": "551352",
      "postDate": "06/12/2019 15:57:44",
      "content": "<p>Thanks, and good luck at the stage 2.\nWe have applied it to all samples in a batch, and we did not use any special consideration for any classes. The implementation is borrowed from <a href=\"https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py\">here</a>. </p>",
      "rawMarkdown": "Thanks, and good luck at the stage 2.\nWe have applied it to all samples in a batch, and we did not use any special consideration for any classes. The implementation is borrowed from [here](https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py).",
      "votes": null
    },
    {
      "id": "551608",
      "postDate": "06/12/2019 23:11:29",
      "content": "<p>You are welcome and thanks, u2 best luck at the stage 2</p>",
      "rawMarkdown": "You are welcome and thanks, u2 best luck at the stage 2",
      "votes": null
    },
    {
      "id": "552162",
      "postDate": "06/13/2019 15:26:10",
      "content": "<p>Very nice writeup <a href=\"/iafoss\">@iafoss</a> \nI miss your awesome kernels..  There were great kernels in this competition especially those of <a href=\"/daisukelab\">@daisukelab</a> and <a href=\"/mhiro2\">@mhiro2</a> \nI had almost the same experiments like yours (only part of what you did).. I wished I had more time than the last 2 days for this comp and those 2 days was hopeless with only 4 submissions , which was surprising that with a good dose of luck it got into the silver range.</p>\n\n<p>One question about your model arch. How did you think about changing the arch into what you did? Is it try and see, or you had an idea that the model should look like this and not that?</p>\n\n<p>I wish you and your team all the best in the 2nd stage. I have learned a lot from your kernels and from <a href=\"/suicaokhoailang\">@suicaokhoailang</a> and <a href=\"/radek1\">@radek1</a>  sharing in the past competitions.</p>",
      "rawMarkdown": "Very nice writeup @iafoss \nI miss your awesome kernels..  There were great kernels in this competition especially those of @daisukelab and @mhiro2 \nI had almost the same experiments like yours (only part of what you did).. I wished I had more time than the last 2 days for this comp and those 2 days was hopeless with only 4 submissions , which was surprising that with a good dose of luck it got into the silver range.\n\nOne question about your model arch. How did you think about changing the arch into what you did? Is it try and see, or you had an idea that the model should look like this and not that?\n\nI wish you and your team all the best in the 2nd stage. I have learned a lot from your kernels and from @suicaokhoailang and @radek1  sharing in the past competitions.",
      "votes": null
    },
    {
      "id": "552214",
      "postDate": "06/13/2019 16:33:41",
      "content": "<p><a href=\"/hwasiti\">@hwasiti</a>, It is quite impressive that u could get that far just within 2 days, hope u also get a good score at the stage 2. Regarding your question, the main reason why I focuses mostly on dense connections (within the conv block and pyramid pooling) is that Dnet models worked much better than ResNet, ResNeXt, CBAM ResNeXt, NasNet, etc. at stage 1. Also, concat pooling gave quite noticeable boost in my early testing. Later I tried other things too, but they didn't work that well. The thing I didn't expect is that traditional 7x7 conv followed by pooling in the first layer doesn't really work here. Probably, if one just took Dnet121 and replace its first block, it could work quite well on 64 mels, but I didn't have time to check it. And thank you for your best wishes.</p>",
      "rawMarkdown": "hwasiti, It is quite impressive that u could get that far just within 2 days, hope u also get a good score at the stage 2. Regarding your question, the main reason why I focuses mostly on dense connections (within the conv block and pyramid pooling) is that Dnet models worked much better than ResNet, ResNeXt, CBAM ResNeXt, NasNet, etc. at stage 1. Also, concat pooling gave quite noticeable boost in my early testing. Later I tried other things too, but they didn't work that well. The thing I didn't expect is that traditional 7x7 conv followed by pooling in the first layer doesn't really work here. Probably, if one just took Dnet121 and replace its first block, it could work quite well on 64 mels, but I didn't have time to check it. And thank you for your best wishes.",
      "votes": null
    },
    {
      "id": "552833",
      "postDate": "06/14/2019 16:01:20",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "561889",
      "postDate": "06/26/2019 20:45:58",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> Thanks for sharing in detail. Maybe I should try to reduce mel to 64 and increase duration to 4s. Haven't thought about reducing mel would gain better CV or LB </p>",
      "rawMarkdown": "iafoss Thanks for sharing in detail. Maybe I should try to reduce mel to 64 and increase duration to 4s. Haven't thought about reducing mel would gain better CV or LB",
      "votes": null
    },
    {
      "id": "561949",
      "postDate": "06/26/2019 22:16:43",
      "content": "<p>The key point with 64 mel is using small models with optimized architecture. Regular computer vision models require 128 mels, and also may be up-scaling as we did in our first attempts, that makes the models quite slow at training and inference. Meanwhile with 64 mel setup we could reach 0.715 public LB score for a single 5-fold model trained within one kernel (~6 hours). Ensembling boosted it to 0.739.</p>\n\n<p>We will release the kernel used for training the models when private LB is available.</p>",
      "rawMarkdown": "The key point with 64 mel is using small models with optimized architecture. Regular computer vision models require 128 mels, and also may be up-scaling as we did in our first attempts, that makes the models quite slow at training and inference. Meanwhile with 64 mel setup we could reach 0.715 public LB score for a single 5-fold model trained within one kernel (~6 hours). Ensembling boosted it to 0.739.\n\nWe will release the kernel used for training the models when private LB is available.",
      "votes": null
    },
    {
      "id": "561972",
      "postDate": "06/26/2019 22:38:59",
      "content": "<p>Great, really keen to see your kernel.</p>",
      "rawMarkdown": "Great, really keen to see your kernel.",
      "votes": null
    },
    {
      "id": "563798",
      "postDate": "06/28/2019 17:36:09",
      "content": "<p>The kernel right now is available at <a href=\"https://www.kaggle.com/theoviel/9th-place-modeling-kernel\">https://www.kaggle.com/theoviel/9th-place-modeling-kernel</a></p>",
      "rawMarkdown": "The kernel right now is available at https://www.kaggle.com/theoviel/9th-place-modeling-kernel",
      "votes": null
    },
    {
      "id": "563944",
      "postDate": "06/28/2019 19:52:39",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> Great work. And congrats to be Kaggle Competition Master. Every time I can learn a lot from your kernel. Also learned  a lot from your Airbus kernels.</p>",
      "rawMarkdown": "iafoss Great work. And congrats to be Kaggle Competition Master. Every time I can learn a lot from your kernel. Also learned  a lot from your Airbus kernels.",
      "votes": null
    },
    {
      "id": "564043",
      "postDate": "06/28/2019 23:25:31",
      "content": "<p>Thanks so much, I'm happy to know that my kernels were useful.</p>",
      "rawMarkdown": "Thanks so much, I'm happy to know that my kernels were useful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 550875,
      "author_name": "daisukelab",
      "author_url": "",
      "post_date": "06/12/2019 05:37:17",
      "content": "<p>Thanks for sharing, I didn't even imagine decreasing mels down to 64.\nAnd I agree with mixup is so efficient in audio task, \"sounds are transparent\" - that's it.\nI hope your team's good luck in 2nd stage.</p>",
      "votes": null,
      "replies": [
        {
          "id": 550889,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/12/2019 05:54:52",
          "content": "<p>Thanks, u2. And thank you for your kernels, I used one as a starting point for mel preprocessing.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551052,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "06/12/2019 09:36:09",
      "content": "<p>Thanks so much <a href=\"/iafoss\">@iafoss</a> and all your team for the very detailed solution. I did try 4s (but with 128 mel) and had no improvement, perhaps I have to go back and redo my experiment :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 551099,
      "author_name": "karanjakhar",
      "author_url": "",
      "post_date": "06/12/2019 10:39:27",
      "content": "<p>Thanks for sharing <a href=\"/iafoss\">@iafoss</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 551143,
      "author_name": "ebouteillon",
      "author_url": "",
      "post_date": "06/12/2019 11:42:15",
      "content": "<p>Thanks for sharing!\nI should have tried 64 mels, it looks promising in your writing ;)\nI wish you the best for this competition</p>",
      "votes": null,
      "replies": [
        {
          "id": 551298,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/12/2019 15:05:44",
          "content": "<p><a href=\"/ebouteillon\">@ebouteillon</a>, I wish u also the best luck. The idea suggested by <a href=\"/theoviel\">@theoviel</a> about 64 mels is really unexpected to work. However, when I tried to switch back to 128 mels in our best setup later, I didn't see improvement. The thing could be also that for 128 mels ~100 epochs (in total) was not enough since stage1 Dnet models we trained for 300-400 epochs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551301,
      "author_name": "fatongyu",
      "author_url": "",
      "post_date": "06/12/2019 15:08:46",
      "content": "<p>perfect</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 551308,
      "author_name": "cateek",
      "author_url": "",
      "post_date": "06/12/2019 15:19:02",
      "content": "<p>May I ask, have your team used any kind of minority class oversampling?  Have you applied MixUp with some probability or just augmented the whole dataset? Thanks for sharing and good luck in final shakeup!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551352,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/12/2019 15:57:44",
          "content": "<p>Thanks, and good luck at the stage 2.\nWe have applied it to all samples in a batch, and we did not use any special consideration for any classes. The implementation is borrowed from <a href=\"https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py\">here</a>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 551334,
      "author_name": "super13579",
      "author_url": "",
      "post_date": "06/12/2019 15:42:17",
      "content": "<p>Thanks for sharing the method!!\nAlways learned a lot from you , <a href=\"/iafoss\">@iafoss</a> .\nNice new paper implement ! (Spectral augmentation)\nHope you get a good result for this competition !!</p>",
      "votes": null,
      "replies": [
        {
          "id": 551608,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/12/2019 23:11:29",
          "content": "<p>You are welcome and thanks, u2 best luck at the stage 2</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 552162,
      "author_name": "hwasiti",
      "author_url": "",
      "post_date": "06/13/2019 15:26:10",
      "content": "<p>Very nice writeup <a href=\"/iafoss\">@iafoss</a> \nI miss your awesome kernels..  There were great kernels in this competition especially those of <a href=\"/daisukelab\">@daisukelab</a> and <a href=\"/mhiro2\">@mhiro2</a> \nI had almost the same experiments like yours (only part of what you did).. I wished I had more time than the last 2 days for this comp and those 2 days was hopeless with only 4 submissions , which was surprising that with a good dose of luck it got into the silver range.</p>\n\n<p>One question about your model arch. How did you think about changing the arch into what you did? Is it try and see, or you had an idea that the model should look like this and not that?</p>\n\n<p>I wish you and your team all the best in the 2nd stage. I have learned a lot from your kernels and from <a href=\"/suicaokhoailang\">@suicaokhoailang</a> and <a href=\"/radek1\">@radek1</a>  sharing in the past competitions.</p>",
      "votes": null,
      "replies": [
        {
          "id": 552214,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/13/2019 16:33:41",
          "content": "<p><a href=\"/hwasiti\">@hwasiti</a>, It is quite impressive that u could get that far just within 2 days, hope u also get a good score at the stage 2. Regarding your question, the main reason why I focuses mostly on dense connections (within the conv block and pyramid pooling) is that Dnet models worked much better than ResNet, ResNeXt, CBAM ResNeXt, NasNet, etc. at stage 1. Also, concat pooling gave quite noticeable boost in my early testing. Later I tried other things too, but they didn't work that well. The thing I didn't expect is that traditional 7x7 conv followed by pooling in the first layer doesn't really work here. Probably, if one just took Dnet121 and replace its first block, it could work quite well on 64 mels, but I didn't have time to check it. And thank you for your best wishes.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 552833,
      "author_name": "rhhridoy",
      "author_url": "",
      "post_date": "06/14/2019 16:01:20",
      "content": "<p>Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 561889,
      "author_name": "lftuwujie",
      "author_url": "",
      "post_date": "06/26/2019 20:45:58",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> Thanks for sharing in detail. Maybe I should try to reduce mel to 64 and increase duration to 4s. Haven't thought about reducing mel would gain better CV or LB </p>",
      "votes": null,
      "replies": [
        {
          "id": 561949,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/26/2019 22:16:43",
          "content": "<p>The key point with 64 mel is using small models with optimized architecture. Regular computer vision models require 128 mels, and also may be up-scaling as we did in our first attempts, that makes the models quite slow at training and inference. Meanwhile with 64 mel setup we could reach 0.715 public LB score for a single 5-fold model trained within one kernel (~6 hours). Ensembling boosted it to 0.739.</p>\n\n<p>We will release the kernel used for training the models when private LB is available.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 561972,
          "author_name": "lftuwujie",
          "author_url": "",
          "post_date": "06/26/2019 22:38:59",
          "content": "<p>Great, really keen to see your kernel.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 563798,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/28/2019 17:36:09",
          "content": "<p>The kernel right now is available at <a href=\"https://www.kaggle.com/theoviel/9th-place-modeling-kernel\">https://www.kaggle.com/theoviel/9th-place-modeling-kernel</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 563944,
          "author_name": "lftuwujie",
          "author_url": "",
          "post_date": "06/28/2019 19:52:39",
          "content": "<p><a href=\"/iafoss\">@iafoss</a> Great work. And congrats to be Kaggle Competition Master. Every time I can learn a lot from your kernel. Also learned  a lot from your Airbus kernels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 564043,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "06/28/2019 23:25:31",
          "content": "<p>Thanks so much, I'm happy to know that my kernels were useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "550824": "First, I would like to congratulate all winers and participants of this competition and thank kaggle and the organizers for posting this interesting challenge and providing resources for participating in it.  Also, I want to express my highest gratitude and appreciation to my teammates: @suicaokhoailang, @theoviel, and @yanglehan for working with me on this challenge.  It was the first time when I worked with audio related task, and it was a month of quite intense learning. I would like to give special thanks to @theoviel for working quite hard with me on optimization of our models and prepare the final submissions.\n\nThe code used for training our final models is available now at [this link](https://www.kaggle.com/theoviel/9th-place-modeling-kernel) (private LB 0.72171 score, top 42, with a single model (5 CV folds) trained for 7 hours). Further details of our solution can be found in our [DCASE2019 technical report](https://storage.googleapis.com/kaggle-forum-message-attachments/563780/13697/DCASE2019_Challenge.pdf).\n\nThe **key points** of our solutions are the following (see details below): (1) Use of **small and fast models with optimized architecture**, (2) Use of **64 mels** instead of 128 (sometimes less gives more) with **4s duration**, (3) Use **data augmentation and noisy data** for pretraining.\n\n**Stage 1**: we started the competition, like many participants, with experimenting with common computer vision models. The input size of spectrograms was 256x512 pixels, with upscaling the input image along first dimension by a factor of 2. With this setup the best performance was demonstrated by DenseNet models: they outperformed the baseline model published in [this kernel](https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch) and successfully used in the competition a year ago, also Dnet121 was faster. With using pretraining on full noisy set, spectral augmentation, and MixUp, CV could reach ~0.85+, and public score for a single fold, 4 folds, and ensemble of 4 models are 0.67-0.68, ~0.70, 0.717, respectively. Despite these models are not used for our submission, these experiments have provided important insights for the next stage.\n\n**Stage 2**: \n**Use of noisy data**: It is the main point of this competition that organizers wanted us to focus on (no external data, no pretrained models, no test data use policies). We have used 2 strategies: (1) pretraining on full noisy data and (2) pretraining on a mixture of the curated data with most confidently labeled noisy data. In both cases the pretraining is followed by fine tuning on curated only data. The most confident labels are identified based on a model trained on curated data, and further details can be provided by @theoviel. For our best setup we have the following values of CV (in stage 2 we use 5 fold scheme): training on curated data only - 0.858, pretraining on full noisy data - 0.866, curated data + 15k best noisy data examples - 0.865, curated data + 5k best noisy data examples - 0.872. We have utilized all 3 strategies of noisy data use to create a variety in our ensemble.\n\n**Preprocessing**: According to our experiments, big models, like Dnet121, work better on 128 mels and even higher image resolution, while the default model reaches the best performance for 64 mels. This setup also decreases training time and improves convergence speed. 32 mels also could be considered, but the performence drops to 0.856 for our best setup. Use of 4s intervals instead of traditional 2s has gave also a considerable boost. The input image size for the model is 64x256x1. We tried both normalization of data based on image and global train set statistics, and the results were similar. Though, our best CV is reached for global normalization. In final models we used both strategies to crease a diversity. We also tried to experiment with the fft window size but did not see a significant difference and stayed with 1920. One thing to try we didn't have time for is using different window sizes mels as channels of the produced image. In particular, [this paper](https://arxiv.org/abs/1706.07156) shows that some classes prefer longer while other shorter window size. The preprocessing pipeline is similar to one described in [this kernel](https://www.kaggle.com/daisukelab/creating-fat2019-preprocessed-data).\n\n**Model architecture**: At stage 2 we used the model from [this kernel](https://www.kaggle.com/mhiro2/simple-2d-cnn-classifier-with-pytorch) as a starting point. The performance of this base model for our best setup is 0.855 CV. Based on our prior positive experience with DensNet, we added dense connections inside convolution blocks and concatenate pooling that boosted the performance to 0.868 in our best experiments (model M1):\n```\nclass ConvBlock(nn.Module):\n    def __init__(self, in_channels, out_channels, kernel_size=3, pool=True):\n        super().__init__()\n        \n        padding = kernel_size // 2\n        self.pool = pool\n        \n        self.conv1 = nn.Sequential(\n            nn.Conv2d(in_channels, out_channels, kernel_size=kernel_size,\n                stride=1, padding=padding),\n            nn.BatchNorm2d(out_channels),\n            nn.ReLU(),\n        )\n        self.conv2 = nn.Sequential(\n            nn.Conv2d(out_channels + in_channels, out_channels, \n                kernel_size=kernel_size, stride=1, padding=padding),\n            nn.BatchNorm2d(out_channels),\n            nn.ReLU(),\n        )\n        \n    def forward(self, x): # x.shape = [batch_size, in_channels, a, b]\n        x1 = self.conv1(x)\n        x = self.conv2(torch.cat([x, x1],1))\n        if(self.pool): x = F.avg_pool2d(x, 2)\n        return x   # x.shape = [batch_size, out_channels, a//2, b//2]\n```\nThe increase of the number of convolution blocks from 4 to 5 gave only 0.865 CV. Use of a pyramidal pooling for 2,3 and 4-th conv blocks (M2) gave slightly worse result than M1. Finally, our ultimate setup (M3) consists of 5 conv blocks with pyramidal pooling reached 0.872 CV. DenseNet121 in the same pipeline reached only 0.836 (DenseNet121 requires higher image resolution, 256x512, to reach 0.85+ CV). From the experiments, it looks for audio it is important to have nonlinear operations before size reduction by pooling, though we did not check it in details. We used M1, M2, and M3 to create variability in our ensemble. Because of the submission limit we checked performance of only a few models in public LB, with the best single model score 0.715.\n\nThis plot illustrates architecture of M3 model:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F4c15c1ec24d235c7834cc1ae94a3ca38%2FM3.png?generation=1561588228369529&amp;alt=media)\n\n\nHere, all tested models are summarized:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ffaca75bb8fb9dc910a648d1c141313be%2Fmodels.png?generation=1561746099636890&amp;alt=media)\n\n\n\nAnd here the model performance for our best setup:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2Ff7014dbf8de75685fff13c683384f068%2Fp.png?generation=1561588425442446&amp;alt=media)\n\n\n**Data augmentation**:\nThe main thing for working with audio data is MixUp. In contrast to images of real objects, [sounds are transparent](https://towardsdatascience.com/whats-wrong-with-spectrograms-and-cnns-for-audio-processing-311377d7ccd): they do not overshadow each other. Therefore, MixUp is so efficient for audio and gives 0.01-0.015 CV boost. At the stage 1 the best results were achieved for alpha MixUp parameter equal to 1.0, while at the stage 2 we used 0.4. [Spectral augmentation](https://arxiv.org/abs/1904.08779) (with 2 masks for frequency and time with the coverage range between zero and 0.15 and 0.3, respectively) gave about 0.005 CV boost. We did not use stretching the spectrograms in time domain because it gave lower model performance. In several model we also used [Multisample Dropout](https://arxiv.org/abs/1905.09788) (other models were trained without dropout); though, it decreased CV by ~0.002. We did not apply horizontal flip since it decreased CV and also is not natural: I do not think that people would be able to recognize sounds played from the back. It is the same as training ImageNet with use of vertical flip.\n\n**Training**: At the pretraining stage we used one cycle of cosine annealing with warm up. The maximum lr is 0.001, and the number of epochs is ranged between 50 and 70 for different setups. At the stage of fine tuning we applied ReduceLROnPlateau several times to alternate high and low lr. The code is implemented with Pytorch. The total time of training for one fold is 1-2 hours, so the training of entire model is 6-9 hours. Almost all our models were trained at kaggle, and the kernel is at [this link](https://www.kaggle.com/theoviel/9th-place-modeling-kernel).\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1212661%2F51ade3b68a7ab201b15fb9736d618a7b%2Flr.png?generation=1561588570346398&amp;alt=media)\n\n\n**Submission**: \nUse of 64 mels has allowed us also to shorten the inference time. In particular, each 5-fold model takes only 1-1.5 min for generation of a prediction with 8TTA. Therefore, we were able to use an ensemble of 15 models to generate the final submissions. \nThe final remark about the possible reason of the gap between CV and public LB: [this paper](https://arxiv.org/pdf/1904.05635v1.pdf) (Task 1B) reports 0.15 lwlrap drop for test performed on data recorded with different device vs the same device as used for training data. So, the public LB data may be selected for a different device, but we never know if it is true or not.",
    "550875": "Thanks for sharing, I didn't even imagine decreasing mels down to 64.\nAnd I agree with mixup is so efficient in audio task, \"sounds are transparent\" - that's it.\nI hope your team's good luck in 2nd stage.",
    "550889": "Thanks, u2. And thank you for your kernels, I used one as a starting point for mel preprocessing.",
    "551052": "Thanks so much @iafoss and all your team for the very detailed solution. I did try 4s (but with 128 mel) and had no improvement, perhaps I have to go back and redo my experiment :)",
    "551099": "Thanks for sharing @iafoss",
    "551143": "Thanks for sharing!\nI should have tried 64 mels, it looks promising in your writing ;)\nI wish you the best for this competition",
    "551298": "ebouteillon, I wish u also the best luck. The idea suggested by @theoviel about 64 mels is really unexpected to work. However, when I tried to switch back to 128 mels in our best setup later, I didn't see improvement. The thing could be also that for 128 mels ~100 epochs (in total) was not enough since stage1 Dnet models we trained for 300-400 epochs.",
    "551301": "perfect",
    "551308": "May I ask, have your team used any kind of minority class oversampling?  Have you applied MixUp with some probability or just augmented the whole dataset? Thanks for sharing and good luck in final shakeup!",
    "551334": "Thanks for sharing the method!!\nAlways learned a lot from you , @iafoss .\nNice new paper implement ! (Spectral augmentation)\nHope you get a good result for this competition !!",
    "551352": "Thanks, and good luck at the stage 2.\nWe have applied it to all samples in a batch, and we did not use any special consideration for any classes. The implementation is borrowed from [here](https://github.com/facebookresearch/mixup-cifar10/blob/master/train.py).",
    "551608": "You are welcome and thanks, u2 best luck at the stage 2",
    "552162": "Very nice writeup @iafoss \nI miss your awesome kernels..  There were great kernels in this competition especially those of @daisukelab and @mhiro2 \nI had almost the same experiments like yours (only part of what you did).. I wished I had more time than the last 2 days for this comp and those 2 days was hopeless with only 4 submissions , which was surprising that with a good dose of luck it got into the silver range.\n\nOne question about your model arch. How did you think about changing the arch into what you did? Is it try and see, or you had an idea that the model should look like this and not that?\n\nI wish you and your team all the best in the 2nd stage. I have learned a lot from your kernels and from @suicaokhoailang and @radek1  sharing in the past competitions.",
    "552214": "hwasiti, It is quite impressive that u could get that far just within 2 days, hope u also get a good score at the stage 2. Regarding your question, the main reason why I focuses mostly on dense connections (within the conv block and pyramid pooling) is that Dnet models worked much better than ResNet, ResNeXt, CBAM ResNeXt, NasNet, etc. at stage 1. Also, concat pooling gave quite noticeable boost in my early testing. Later I tried other things too, but they didn't work that well. The thing I didn't expect is that traditional 7x7 conv followed by pooling in the first layer doesn't really work here. Probably, if one just took Dnet121 and replace its first block, it could work quite well on 64 mels, but I didn't have time to check it. And thank you for your best wishes.",
    "552833": "Thanks for sharing.",
    "561889": "iafoss Thanks for sharing in detail. Maybe I should try to reduce mel to 64 and increase duration to 4s. Haven't thought about reducing mel would gain better CV or LB",
    "561949": "The key point with 64 mel is using small models with optimized architecture. Regular computer vision models require 128 mels, and also may be up-scaling as we did in our first attempts, that makes the models quite slow at training and inference. Meanwhile with 64 mel setup we could reach 0.715 public LB score for a single 5-fold model trained within one kernel (~6 hours). Ensembling boosted it to 0.739.\n\nWe will release the kernel used for training the models when private LB is available.",
    "561972": "Great, really keen to see your kernel.",
    "563798": "The kernel right now is available at https://www.kaggle.com/theoviel/9th-place-modeling-kernel",
    "563944": "iafoss Great work. And congrats to be Kaggle Competition Master. Every time I can learn a lot from your kernel. Also learned  a lot from your Airbus kernels.",
    "564043": "Thanks so much, I'm happy to know that my kernels were useful."
  },
  "source": "meta"
}