{
  "id": 64262,
  "title": "8th place solution",
  "url": "/competitions/freesound-audio-tagging/writeups/sainathadapa-8th-place-solution",
  "author_name": "",
  "post_date": "2019-06-05T15:58:31.840Z",
  "votes": 34,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I'm describing the approach that I took for the final solution here. The model achieved 8th position on the final leaderboard, with a score (MAP@3) of 0.943289. The code is available at <a href=\"https://github.com/sainathadapa/kaggle-freesound-audio-tagging\">https://github.com/sainathadapa/kaggle-freesound-audio-tagging</a>. (For the list of things that I tried before the final solution, <a href=\"https://github.com/sainathadapa/kaggle-freesound-audio-tagging/blob/master/approaches_all.md\">please refer to this document</a>).</p>\n\n<h1>Acknowledgements</h1>\n\n<p>Thanks to <a href=\"https://www.kaggle.com/amlanpraharaj\">Amlan Praharaj</a>, <a href=\"https://www.kaggle.com/opanichev\">Oleg Panichev</a> and <a href=\"https://www.kaggle.com/agehsbarg\">Aleksandrs Gehsbargs</a> for the kernels they have shared. The kernels have helped me get started on the competition. Special thanks to <a href=\"https://www.kaggle.com/daisukelab\"></a><a href=\"/daisukelab\">@daisukelab</a> for the sharing his observations about data preprocessing, augmentation and model architectures. His insights were crucial to my solution.</p>\n\n<h1>Data preprocessing</h1>\n\n<p>Leading/trailing silence in the audio may not contain much information and thus not useful for the model. Hence, the very first preprocessing step is to remove this silence. <code>librosa.effects.trim</code> function was used to achieve this.</p>\n\n<h2>Log Mel-Spectrograms</h2>\n\n<p>Often in speech recognition tasks, MFCC features are constructed from the raw audio data. Since the current data contains non-human sounds as well, using the Log Mel-Spectrogram data is better compared to the MFCC representation. Log Mel-Spectrogram for all train and test samples was pre-computed, so that compute time can be saved during training and prediction (Disk is cheaper when compared to GPU).</p>\n\n<h2>Additional features</h2>\n\n<p>Inspired by few Kaggle kernels, summary statistics of multiple spectral and time based features were calculated. Since many of these features were correlated, these features were transformed using the Principle Component Analysis (PCA). Top 350 features were used while modeling (which amount to ~97% of the total variance).</p>\n\n<h1>Architecture</h1>\n\n<p>The model at its core, uses the <a href=\"https://arxiv.org/abs/1801.04381\">MobileNetV2</a> architecture with few modifications. The input Log Mel-Spec data is sent to the MobileNetV2 after first passing the input through two 2D convolution layers. This is so that the single channel input can be converted into a 3 channel input. (Thanks to the <a href=\"http://forums.fast.ai/t/black-and-white-images-on-vgg16/2479/12\">FastAI forums for this tip</a>) The output from the MobileNetV2 is then concatenated with the PCA features, and a series of Dense layers are used before the final softmax activation layer for output.</p>\n\n<pre><code>inp1 = Input(shape=(64, None, 1), name='mel')\n\nx = BatchNormalization()(inp1)\nx = Conv2D(10, kernel_size=(1, 1), padding='same', activation='relu')(x)\nx = Conv2D(3, kernel_size=(1, 1), padding='same', activation='relu')(x)\n\nmn = MobileNetV2(include_top=False)\nmn.layers.pop(0)\n\nmn_out = mn(x)\nx = GlobalAveragePooling2D()(mn_out)\n\ninp2 = Input(shape=(350,), name='pca')\ny = BatchNormalization()(inp2)\n\nx = concatenate([x, y], axis=-1)\nx = Dense(1536, activation='relu')(x)\nx = BatchNormalization()(x)\nx = Dense(384, activation='relu')(x)\nx = BatchNormalization()(x)\nx = Dense(41, activation='softmax')(x)\n\nmodel = Model(inputs=[inp1, inp2], outputs=x)\n\n\n__________________________________________________________________________________________________\nLayer (type)                    Output Shape         Param #     Connected to                     \n==================================================================================================\nmel (InputLayer)                (None, 64, None, 1)  0                                            \n__________________________________________________________________________________________________\nbatch_normalization_1 (BatchNor (None, 64, None, 1)  4           mel[0][0]                        \n__________________________________________________________________________________________________\nconv2d_1 (Conv2D)               (None, 64, None, 10) 20          batch_normalization_1[0][0]      \n__________________________________________________________________________________________________\nconv2d_2 (Conv2D)               (None, 64, None, 3)  33          conv2d_1[0][0]                   \n__________________________________________________________________________________________________\nmobilenetv2_1.00_224 (Model)    multiple             2257984     conv2d_2[0][0]                   \n__________________________________________________________________________________________________\npca (InputLayer)                (None, 350)          0                                            \n__________________________________________________________________________________________________\nglobal_average_pooling2d_1 (Glo (None, 1280)         0           mobilenetv2_1.00_224[1][0]       \n__________________________________________________________________________________________________\nbatch_normalization_2 (BatchNor (None, 350)          1400        pca[0][0]                        \n__________________________________________________________________________________________________\nconcatenate_1 (Concatenate)     (None, 1630)         0           global_average_pooling2d_1[0][0]\n                                                                 batch_normalization_2[0][0]      \n__________________________________________________________________________________________________\ndense_1 (Dense)                 (None, 1536)         2505216     concatenate_1[0][0]              \n__________________________________________________________________________________________________\nbatch_normalization_3 (BatchNor (None, 1536)         6144        dense_1[0][0]                    \n__________________________________________________________________________________________________\ndense_2 (Dense)                 (None, 384)          590208      batch_normalization_3[0][0]      \n__________________________________________________________________________________________________\nbatch_normalization_4 (BatchNor (None, 384)          1536        dense_2[0][0]                    \n__________________________________________________________________________________________________\ndense_3 (Dense)                 (None, 41)           15785       batch_normalization_4[0][0]      \n==================================================================================================\nTotal params: 5,378,330\nTrainable params: 5,339,676\nNon-trainable params: 38,654\n__________________________________________________________________________________________________\n</code></pre>\n\n<h1>Train data generation with augmentation</h1>\n\n<p>Both the train and test audio files are of varied length. Model is designed to make use of this particular nature of the dataset. Use of Global average pooling before the Dense layers allows the model to accept inputs of various lengths. While training, at each batch generation, a random integer between the limits is chosen. 25th and 75th percentiles of train file lengths are used as min and max limits respectively. Shorter length samples than the chosen length are padded, while a random span of chosen length is extracted from longer length samples.</p>\n\n<p>The common augmentation practices for Image classification such as horizontal/vertical shift, horizontal flip were used. In addition to this, Random erasing was also used. Random erasing or Cutout selects a random rectangle in the image, and replaces it with adjacent or random values. For more information about this data augmentation technique, refer to the <a href=\"https://arxiv.org/abs/1708.04896\">original paper</a>.</p>\n\n<p>Mixup is the final augmentation technique used while training. Mixup essentially takes pairs of data points, chosen randomly, and mixes them (both X and y) using a proportion chosen from Beta distribution.\n&gt; One intuition behind this is that by linearly interpolating between datapoints, we incentivize the network to act smoothly and kind of interpolate nicely between datapoints - without sharp transitions. (Quote from <a href=\"https://www.inference.vc/mixup-data-dependent-data-augmentation/\">https://www.inference.vc/mixup-data-dependent-data-augmentation/</a>)</p>\n\n<p>While I haven't ran exhaustive trials to say for sure, anecdotally, each of the data augmentation have helped in improving the loss.</p>\n\n<h1>Training</h1>\n\n<p>Ten folds (stratified split as there is class imbalance) were generated. For each fold, a model of similar architecture but that uses only Log Mel-Spectrogram data is trained. The weights from this model are loaded into the whole model (that uses both mel and pca features), and training process continues. Attempt to train the model without using this two-stage approach didn't result in as good a model as before.</p>\n\n<h1>Predictions on test data</h1>\n\n<p>Six different lengths selected at equal intervals between the 25th and 75th percentile of train file lengths. To make use of the higher amount of information present in longer length samples, at each length, predictions are generated five times. Each time a random span of specified length is extracted from longer (than specified length) length samples. 10 (folds) x 6 (lengths) x 5 (tries) gives 300 sets of predictions for the test data. All of these predictions were combined using geometric mean, and top 3 predicted classes for each data point are selected for submission.</p>",
  "messages": [
    {
      "id": "376395",
      "postDate": "08/27/2018 12:01:37",
      "content": "<p>I'm describing the approach that I took for the final solution here. The model achieved 8th position on the final leaderboard, with a score (MAP@3) of 0.943289. The code is available at <a href=\"https://github.com/sainathadapa/kaggle-freesound-audio-tagging\">https://github.com/sainathadapa/kaggle-freesound-audio-tagging</a>. (For the list of things that I tried before the final solution, <a href=\"https://github.com/sainathadapa/kaggle-freesound-audio-tagging/blob/master/approaches_all.md\">please refer to this document</a>).</p>\n\n<h1>Acknowledgements</h1>\n\n<p>Thanks to <a href=\"https://www.kaggle.com/amlanpraharaj\">Amlan Praharaj</a>, <a href=\"https://www.kaggle.com/opanichev\">Oleg Panichev</a> and <a href=\"https://www.kaggle.com/agehsbarg\">Aleksandrs Gehsbargs</a> for the kernels they have shared. The kernels have helped me get started on the competition. Special thanks to <a href=\"https://www.kaggle.com/daisukelab\"></a><a href=\"/daisukelab\">@daisukelab</a> for the sharing his observations about data preprocessing, augmentation and model architectures. His insights were crucial to my solution.</p>\n\n<h1>Data preprocessing</h1>\n\n<p>Leading/trailing silence in the audio may not contain much information and thus not useful for the model. Hence, the very first preprocessing step is to remove this silence. <code>librosa.effects.trim</code> function was used to achieve this.</p>\n\n<h2>Log Mel-Spectrograms</h2>\n\n<p>Often in speech recognition tasks, MFCC features are constructed from the raw audio data. Since the current data contains non-human sounds as well, using the Log Mel-Spectrogram data is better compared to the MFCC representation. Log Mel-Spectrogram for all train and test samples was pre-computed, so that compute time can be saved during training and prediction (Disk is cheaper when compared to GPU).</p>\n\n<h2>Additional features</h2>\n\n<p>Inspired by few Kaggle kernels, summary statistics of multiple spectral and time based features were calculated. Since many of these features were correlated, these features were transformed using the Principle Component Analysis (PCA). Top 350 features were used while modeling (which amount to ~97% of the total variance).</p>\n\n<h1>Architecture</h1>\n\n<p>The model at its core, uses the <a href=\"https://arxiv.org/abs/1801.04381\">MobileNetV2</a> architecture with few modifications. The input Log Mel-Spec data is sent to the MobileNetV2 after first passing the input through two 2D convolution layers. This is so that the single channel input can be converted into a 3 channel input. (Thanks to the <a href=\"http://forums.fast.ai/t/black-and-white-images-on-vgg16/2479/12\">FastAI forums for this tip</a>) The output from the MobileNetV2 is then concatenated with the PCA features, and a series of Dense layers are used before the final softmax activation layer for output.</p>\n\n<pre><code>inp1 = Input(shape=(64, None, 1), name='mel')\n\nx = BatchNormalization()(inp1)\nx = Conv2D(10, kernel_size=(1, 1), padding='same', activation='relu')(x)\nx = Conv2D(3, kernel_size=(1, 1), padding='same', activation='relu')(x)\n\nmn = MobileNetV2(include_top=False)\nmn.layers.pop(0)\n\nmn_out = mn(x)\nx = GlobalAveragePooling2D()(mn_out)\n\ninp2 = Input(shape=(350,), name='pca')\ny = BatchNormalization()(inp2)\n\nx = concatenate([x, y], axis=-1)\nx = Dense(1536, activation='relu')(x)\nx = BatchNormalization()(x)\nx = Dense(384, activation='relu')(x)\nx = BatchNormalization()(x)\nx = Dense(41, activation='softmax')(x)\n\nmodel = Model(inputs=[inp1, inp2], outputs=x)\n\n\n__________________________________________________________________________________________________\nLayer (type)                    Output Shape         Param #     Connected to                     \n==================================================================================================\nmel (InputLayer)                (None, 64, None, 1)  0                                            \n__________________________________________________________________________________________________\nbatch_normalization_1 (BatchNor (None, 64, None, 1)  4           mel[0][0]                        \n__________________________________________________________________________________________________\nconv2d_1 (Conv2D)               (None, 64, None, 10) 20          batch_normalization_1[0][0]      \n__________________________________________________________________________________________________\nconv2d_2 (Conv2D)               (None, 64, None, 3)  33          conv2d_1[0][0]                   \n__________________________________________________________________________________________________\nmobilenetv2_1.00_224 (Model)    multiple             2257984     conv2d_2[0][0]                   \n__________________________________________________________________________________________________\npca (InputLayer)                (None, 350)          0                                            \n__________________________________________________________________________________________________\nglobal_average_pooling2d_1 (Glo (None, 1280)         0           mobilenetv2_1.00_224[1][0]       \n__________________________________________________________________________________________________\nbatch_normalization_2 (BatchNor (None, 350)          1400        pca[0][0]                        \n__________________________________________________________________________________________________\nconcatenate_1 (Concatenate)     (None, 1630)         0           global_average_pooling2d_1[0][0]\n                                                                 batch_normalization_2[0][0]      \n__________________________________________________________________________________________________\ndense_1 (Dense)                 (None, 1536)         2505216     concatenate_1[0][0]              \n__________________________________________________________________________________________________\nbatch_normalization_3 (BatchNor (None, 1536)         6144        dense_1[0][0]                    \n__________________________________________________________________________________________________\ndense_2 (Dense)                 (None, 384)          590208      batch_normalization_3[0][0]      \n__________________________________________________________________________________________________\nbatch_normalization_4 (BatchNor (None, 384)          1536        dense_2[0][0]                    \n__________________________________________________________________________________________________\ndense_3 (Dense)                 (None, 41)           15785       batch_normalization_4[0][0]      \n==================================================================================================\nTotal params: 5,378,330\nTrainable params: 5,339,676\nNon-trainable params: 38,654\n__________________________________________________________________________________________________\n</code></pre>\n\n<h1>Train data generation with augmentation</h1>\n\n<p>Both the train and test audio files are of varied length. Model is designed to make use of this particular nature of the dataset. Use of Global average pooling before the Dense layers allows the model to accept inputs of various lengths. While training, at each batch generation, a random integer between the limits is chosen. 25th and 75th percentiles of train file lengths are used as min and max limits respectively. Shorter length samples than the chosen length are padded, while a random span of chosen length is extracted from longer length samples.</p>\n\n<p>The common augmentation practices for Image classification such as horizontal/vertical shift, horizontal flip were used. In addition to this, Random erasing was also used. Random erasing or Cutout selects a random rectangle in the image, and replaces it with adjacent or random values. For more information about this data augmentation technique, refer to the <a href=\"https://arxiv.org/abs/1708.04896\">original paper</a>.</p>\n\n<p>Mixup is the final augmentation technique used while training. Mixup essentially takes pairs of data points, chosen randomly, and mixes them (both X and y) using a proportion chosen from Beta distribution.\n&gt; One intuition behind this is that by linearly interpolating between datapoints, we incentivize the network to act smoothly and kind of interpolate nicely between datapoints - without sharp transitions. (Quote from <a href=\"https://www.inference.vc/mixup-data-dependent-data-augmentation/\">https://www.inference.vc/mixup-data-dependent-data-augmentation/</a>)</p>\n\n<p>While I haven't ran exhaustive trials to say for sure, anecdotally, each of the data augmentation have helped in improving the loss.</p>\n\n<h1>Training</h1>\n\n<p>Ten folds (stratified split as there is class imbalance) were generated. For each fold, a model of similar architecture but that uses only Log Mel-Spectrogram data is trained. The weights from this model are loaded into the whole model (that uses both mel and pca features), and training process continues. Attempt to train the model without using this two-stage approach didn't result in as good a model as before.</p>\n\n<h1>Predictions on test data</h1>\n\n<p>Six different lengths selected at equal intervals between the 25th and 75th percentile of train file lengths. To make use of the higher amount of information present in longer length samples, at each length, predictions are generated five times. Each time a random span of specified length is extracted from longer (than specified length) length samples. 10 (folds) x 6 (lengths) x 5 (tries) gives 300 sets of predictions for the test data. All of these predictions were combined using geometric mean, and top 3 predicted classes for each data point are selected for submission.</p>",
      "rawMarkdown": "I'm describing the approach that I took for the final solution here. The model achieved 8th position on the final leaderboard, with a score (MAP@3) of 0.943289. The code is available at https://github.com/sainathadapa/kaggle-freesound-audio-tagging. (For the list of things that I tried before the final solution, [please refer to this document](https://github.com/sainathadapa/kaggle-freesound-audio-tagging/blob/master/approaches_all.md)).\n\n# Acknowledgements\nThanks to [Amlan Praharaj](https://www.kaggle.com/amlanpraharaj), [Oleg Panichev](https://www.kaggle.com/opanichev) and [Aleksandrs Gehsbargs](https://www.kaggle.com/agehsbarg) for the kernels they have shared. The kernels have helped me get started on the competition. Special thanks to [@daisukelab](https://www.kaggle.com/daisukelab) for the sharing his observations about data preprocessing, augmentation and model architectures. His insights were crucial to my solution.\n\n#  Data preprocessing\nLeading/trailing silence in the audio may not contain much information and thus not useful for the model. Hence, the very first preprocessing step is to remove this silence. `librosa.effects.trim` function was used to achieve this.\n\n## Log Mel-Spectrograms\nOften in speech recognition tasks, MFCC features are constructed from the raw audio data. Since the current data contains non-human sounds as well, using the Log Mel-Spectrogram data is better compared to the MFCC representation. Log Mel-Spectrogram for all train and test samples was pre-computed, so that compute time can be saved during training and prediction (Disk is cheaper when compared to GPU).\n\n## Additional features\nInspired by few Kaggle kernels, summary statistics of multiple spectral and time based features were calculated. Since many of these features were correlated, these features were transformed using the Principle Component Analysis (PCA). Top 350 features were used while modeling (which amount to ~97% of the total variance).\n\n# Architecture\nThe model at its core, uses the [MobileNetV2](https://arxiv.org/abs/1801.04381) architecture with few modifications. The input Log Mel-Spec data is sent to the MobileNetV2 after first passing the input through two 2D convolution layers. This is so that the single channel input can be converted into a 3 channel input. (Thanks to the [FastAI forums for this tip](http://forums.fast.ai/t/black-and-white-images-on-vgg16/2479/12)) The output from the MobileNetV2 is then concatenated with the PCA features, and a series of Dense layers are used before the final softmax activation layer for output.\n\n\n    inp1 = Input(shape=(64, None, 1), name='mel')\n    \n    x = BatchNormalization()(inp1)\n    x = Conv2D(10, kernel_size=(1, 1), padding='same', activation='relu')(x)\n    x = Conv2D(3, kernel_size=(1, 1), padding='same', activation='relu')(x)\n    \n    mn = MobileNetV2(include_top=False)\n    mn.layers.pop(0)\n    \n    mn_out = mn(x)\n    x = GlobalAveragePooling2D()(mn_out)\n    \n    inp2 = Input(shape=(350,), name='pca')\n    y = BatchNormalization()(inp2)\n    \n    x = concatenate([x, y], axis=-1)\n    x = Dense(1536, activation='relu')(x)\n    x = BatchNormalization()(x)\n    x = Dense(384, activation='relu')(x)\n    x = BatchNormalization()(x)\n    x = Dense(41, activation='softmax')(x)\n    \n    model = Model(inputs=[inp1, inp2], outputs=x)\n\n\n    __________________________________________________________________________________________________\n    Layer (type)                    Output Shape         Param #     Connected to                     \n    ==================================================================================================\n    mel (InputLayer)                (None, 64, None, 1)  0                                            \n    __________________________________________________________________________________________________\n    batch_normalization_1 (BatchNor (None, 64, None, 1)  4           mel[0][0]                        \n    __________________________________________________________________________________________________\n    conv2d_1 (Conv2D)               (None, 64, None, 10) 20          batch_normalization_1[0][0]      \n    __________________________________________________________________________________________________\n    conv2d_2 (Conv2D)               (None, 64, None, 3)  33          conv2d_1[0][0]                   \n    __________________________________________________________________________________________________\n    mobilenetv2_1.00_224 (Model)    multiple             2257984     conv2d_2[0][0]                   \n    __________________________________________________________________________________________________\n    pca (InputLayer)                (None, 350)          0                                            \n    __________________________________________________________________________________________________\n    global_average_pooling2d_1 (Glo (None, 1280)         0           mobilenetv2_1.00_224[1][0]       \n    __________________________________________________________________________________________________\n    batch_normalization_2 (BatchNor (None, 350)          1400        pca[0][0]                        \n    __________________________________________________________________________________________________\n    concatenate_1 (Concatenate)     (None, 1630)         0           global_average_pooling2d_1[0][0]\n                                                                     batch_normalization_2[0][0]      \n    __________________________________________________________________________________________________\n    dense_1 (Dense)                 (None, 1536)         2505216     concatenate_1[0][0]              \n    __________________________________________________________________________________________________\n    batch_normalization_3 (BatchNor (None, 1536)         6144        dense_1[0][0]                    \n    __________________________________________________________________________________________________\n    dense_2 (Dense)                 (None, 384)          590208      batch_normalization_3[0][0]      \n    __________________________________________________________________________________________________\n    batch_normalization_4 (BatchNor (None, 384)          1536        dense_2[0][0]                    \n    __________________________________________________________________________________________________\n    dense_3 (Dense)                 (None, 41)           15785       batch_normalization_4[0][0]      \n    ==================================================================================================\n    Total params: 5,378,330\n    Trainable params: 5,339,676\n    Non-trainable params: 38,654\n    __________________________________________________________________________________________________\n\n\n# Train data generation with augmentation\nBoth the train and test audio files are of varied length. Model is designed to make use of this particular nature of the dataset. Use of Global average pooling before the Dense layers allows the model to accept inputs of various lengths. While training, at each batch generation, a random integer between the limits is chosen. 25th and 75th percentiles of train file lengths are used as min and max limits respectively. Shorter length samples than the chosen length are padded, while a random span of chosen length is extracted from longer length samples.\n\nThe common augmentation practices for Image classification such as horizontal/vertical shift, horizontal flip were used. In addition to this, Random erasing was also used. Random erasing or Cutout selects a random rectangle in the image, and replaces it with adjacent or random values. For more information about this data augmentation technique, refer to the [original paper](https://arxiv.org/abs/1708.04896).\n\nMixup is the final augmentation technique used while training. Mixup essentially takes pairs of data points, chosen randomly, and mixes them (both X and y) using a proportion chosen from Beta distribution.\n&gt; One intuition behind this is that by linearly interpolating between datapoints, we incentivize the network to act smoothly and kind of interpolate nicely between datapoints - without sharp transitions. (Quote from https://www.inference.vc/mixup-data-dependent-data-augmentation/)\n\nWhile I haven't ran exhaustive trials to say for sure, anecdotally, each of the data augmentation have helped in improving the loss.\n\n# Training\nTen folds (stratified split as there is class imbalance) were generated. For each fold, a model of similar architecture but that uses only Log Mel-Spectrogram data is trained. The weights from this model are loaded into the whole model (that uses both mel and pca features), and training process continues. Attempt to train the model without using this two-stage approach didn't result in as good a model as before.\n\n# Predictions on test data\nSix different lengths selected at equal intervals between the 25th and 75th percentile of train file lengths. To make use of the higher amount of information present in longer length samples, at each length, predictions are generated five times. Each time a random span of specified length is extracted from longer (than specified length) length samples. 10 (folds) x 6 (lengths) x 5 (tries) gives 300 sets of predictions for the test data. All of these predictions were combined using geometric mean, and top 3 predicted classes for each data point are selected for submission.",
      "votes": null
    },
    {
      "id": "523914",
      "postDate": "04/27/2019 11:31:03",
      "content": "<p>Great work</p>",
      "rawMarkdown": "Great work",
      "votes": null
    },
    {
      "id": "538865",
      "postDate": "05/29/2019 07:48:28",
      "content": "<p>Excellent work, this is very helpful, Thanks.</p>",
      "rawMarkdown": "Excellent work, this is very helpful, Thanks.",
      "votes": null
    },
    {
      "id": "546026",
      "postDate": "06/06/2019 07:10:15",
      "content": "<p><a href=\"/sainathadapa\">@sainathadapa</a>  any perquisite or elementary domain knowledge is required for this task which may help in understanding data batter</p>",
      "rawMarkdown": "sainathadapa  any perquisite or elementary domain knowledge is required for this task which may help in understanding data batter",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 523914,
      "author_name": "ayushyuvraj",
      "author_url": "",
      "post_date": "04/27/2019 11:31:03",
      "content": "<p>Great work</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 538865,
      "author_name": "ackusingh",
      "author_url": "",
      "post_date": "05/29/2019 07:48:28",
      "content": "<p>Excellent work, this is very helpful, Thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 546026,
      "author_name": "econdata",
      "author_url": "",
      "post_date": "06/06/2019 07:10:15",
      "content": "<p><a href=\"/sainathadapa\">@sainathadapa</a>  any perquisite or elementary domain knowledge is required for this task which may help in understanding data batter</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "376395": "I'm describing the approach that I took for the final solution here. The model achieved 8th position on the final leaderboard, with a score (MAP@3) of 0.943289. The code is available at https://github.com/sainathadapa/kaggle-freesound-audio-tagging. (For the list of things that I tried before the final solution, [please refer to this document](https://github.com/sainathadapa/kaggle-freesound-audio-tagging/blob/master/approaches_all.md)).\n\n# Acknowledgements\nThanks to [Amlan Praharaj](https://www.kaggle.com/amlanpraharaj), [Oleg Panichev](https://www.kaggle.com/opanichev) and [Aleksandrs Gehsbargs](https://www.kaggle.com/agehsbarg) for the kernels they have shared. The kernels have helped me get started on the competition. Special thanks to [@daisukelab](https://www.kaggle.com/daisukelab) for the sharing his observations about data preprocessing, augmentation and model architectures. His insights were crucial to my solution.\n\n#  Data preprocessing\nLeading/trailing silence in the audio may not contain much information and thus not useful for the model. Hence, the very first preprocessing step is to remove this silence. `librosa.effects.trim` function was used to achieve this.\n\n## Log Mel-Spectrograms\nOften in speech recognition tasks, MFCC features are constructed from the raw audio data. Since the current data contains non-human sounds as well, using the Log Mel-Spectrogram data is better compared to the MFCC representation. Log Mel-Spectrogram for all train and test samples was pre-computed, so that compute time can be saved during training and prediction (Disk is cheaper when compared to GPU).\n\n## Additional features\nInspired by few Kaggle kernels, summary statistics of multiple spectral and time based features were calculated. Since many of these features were correlated, these features were transformed using the Principle Component Analysis (PCA). Top 350 features were used while modeling (which amount to ~97% of the total variance).\n\n# Architecture\nThe model at its core, uses the [MobileNetV2](https://arxiv.org/abs/1801.04381) architecture with few modifications. The input Log Mel-Spec data is sent to the MobileNetV2 after first passing the input through two 2D convolution layers. This is so that the single channel input can be converted into a 3 channel input. (Thanks to the [FastAI forums for this tip](http://forums.fast.ai/t/black-and-white-images-on-vgg16/2479/12)) The output from the MobileNetV2 is then concatenated with the PCA features, and a series of Dense layers are used before the final softmax activation layer for output.\n\n\n    inp1 = Input(shape=(64, None, 1), name='mel')\n    \n    x = BatchNormalization()(inp1)\n    x = Conv2D(10, kernel_size=(1, 1), padding='same', activation='relu')(x)\n    x = Conv2D(3, kernel_size=(1, 1), padding='same', activation='relu')(x)\n    \n    mn = MobileNetV2(include_top=False)\n    mn.layers.pop(0)\n    \n    mn_out = mn(x)\n    x = GlobalAveragePooling2D()(mn_out)\n    \n    inp2 = Input(shape=(350,), name='pca')\n    y = BatchNormalization()(inp2)\n    \n    x = concatenate([x, y], axis=-1)\n    x = Dense(1536, activation='relu')(x)\n    x = BatchNormalization()(x)\n    x = Dense(384, activation='relu')(x)\n    x = BatchNormalization()(x)\n    x = Dense(41, activation='softmax')(x)\n    \n    model = Model(inputs=[inp1, inp2], outputs=x)\n\n\n    __________________________________________________________________________________________________\n    Layer (type)                    Output Shape         Param #     Connected to                     \n    ==================================================================================================\n    mel (InputLayer)                (None, 64, None, 1)  0                                            \n    __________________________________________________________________________________________________\n    batch_normalization_1 (BatchNor (None, 64, None, 1)  4           mel[0][0]                        \n    __________________________________________________________________________________________________\n    conv2d_1 (Conv2D)               (None, 64, None, 10) 20          batch_normalization_1[0][0]      \n    __________________________________________________________________________________________________\n    conv2d_2 (Conv2D)               (None, 64, None, 3)  33          conv2d_1[0][0]                   \n    __________________________________________________________________________________________________\n    mobilenetv2_1.00_224 (Model)    multiple             2257984     conv2d_2[0][0]                   \n    __________________________________________________________________________________________________\n    pca (InputLayer)                (None, 350)          0                                            \n    __________________________________________________________________________________________________\n    global_average_pooling2d_1 (Glo (None, 1280)         0           mobilenetv2_1.00_224[1][0]       \n    __________________________________________________________________________________________________\n    batch_normalization_2 (BatchNor (None, 350)          1400        pca[0][0]                        \n    __________________________________________________________________________________________________\n    concatenate_1 (Concatenate)     (None, 1630)         0           global_average_pooling2d_1[0][0]\n                                                                     batch_normalization_2[0][0]      \n    __________________________________________________________________________________________________\n    dense_1 (Dense)                 (None, 1536)         2505216     concatenate_1[0][0]              \n    __________________________________________________________________________________________________\n    batch_normalization_3 (BatchNor (None, 1536)         6144        dense_1[0][0]                    \n    __________________________________________________________________________________________________\n    dense_2 (Dense)                 (None, 384)          590208      batch_normalization_3[0][0]      \n    __________________________________________________________________________________________________\n    batch_normalization_4 (BatchNor (None, 384)          1536        dense_2[0][0]                    \n    __________________________________________________________________________________________________\n    dense_3 (Dense)                 (None, 41)           15785       batch_normalization_4[0][0]      \n    ==================================================================================================\n    Total params: 5,378,330\n    Trainable params: 5,339,676\n    Non-trainable params: 38,654\n    __________________________________________________________________________________________________\n\n\n# Train data generation with augmentation\nBoth the train and test audio files are of varied length. Model is designed to make use of this particular nature of the dataset. Use of Global average pooling before the Dense layers allows the model to accept inputs of various lengths. While training, at each batch generation, a random integer between the limits is chosen. 25th and 75th percentiles of train file lengths are used as min and max limits respectively. Shorter length samples than the chosen length are padded, while a random span of chosen length is extracted from longer length samples.\n\nThe common augmentation practices for Image classification such as horizontal/vertical shift, horizontal flip were used. In addition to this, Random erasing was also used. Random erasing or Cutout selects a random rectangle in the image, and replaces it with adjacent or random values. For more information about this data augmentation technique, refer to the [original paper](https://arxiv.org/abs/1708.04896).\n\nMixup is the final augmentation technique used while training. Mixup essentially takes pairs of data points, chosen randomly, and mixes them (both X and y) using a proportion chosen from Beta distribution.\n&gt; One intuition behind this is that by linearly interpolating between datapoints, we incentivize the network to act smoothly and kind of interpolate nicely between datapoints - without sharp transitions. (Quote from https://www.inference.vc/mixup-data-dependent-data-augmentation/)\n\nWhile I haven't ran exhaustive trials to say for sure, anecdotally, each of the data augmentation have helped in improving the loss.\n\n# Training\nTen folds (stratified split as there is class imbalance) were generated. For each fold, a model of similar architecture but that uses only Log Mel-Spectrogram data is trained. The weights from this model are loaded into the whole model (that uses both mel and pca features), and training process continues. Attempt to train the model without using this two-stage approach didn't result in as good a model as before.\n\n# Predictions on test data\nSix different lengths selected at equal intervals between the 25th and 75th percentile of train file lengths. To make use of the higher amount of information present in longer length samples, at each length, predictions are generated five times. Each time a random span of specified length is extracted from longer (than specified length) length samples. 10 (folds) x 6 (lengths) x 5 (tries) gives 300 sets of predictions for the test data. All of these predictions were combined using geometric mean, and top 3 predicted classes for each data point are selected for submission.",
    "523914": "Great work",
    "538865": "Excellent work, this is very helpful, Thanks.",
    "546026": "sainathadapa  any perquisite or elementary domain knowledge is required for this task which may help in understanding data batter"
  },
  "source": "meta"
}