{
  "id": 47790,
  "title": "In search of best architecture: crelu, dilation, etc. ",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47790",
  "author_name": "harshml",
  "post_date": "2018-01-18T21:23:53.787000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p><em>I had no experience in dealing with audio files, so it was very interesting to try myself in this competition. Thanks all for</em> <code>hard coding</code> <em>in new year holidays!</em></p>\n\n<p>Here is my report on evaluation some nn building blocks for speech recognition task.</p>\n\n<p><strong>Stage 0. <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">Sample's</a> parameters tuning</strong></p>\n\n<p>Tensorflow's sample already can perform train/val split without picking same speaking person in both splits and has time shift/volume level augmentation, so it worth trying. It <code>conv</code> model with default parameters provides a nice baseline in <strong>77%</strong> on LB. I made one change, switched from SGD (converged slowly) to Adam, and training started to take <strong>~15 min</strong>. Next was reading a bunch of articles about audio processing, almost all suggested to use window size in 20 ms with 50% overlap and just 12 dct coefficients. With this setup, accuracy increased to <strong>80%</strong>. Now it's time to try more advanced models, decided to experiment with raw wav, since network usually learns features better, than handcrafted.</p>\n\n<p><strong>Stage 1. Raw wav experiments</strong></p>\n\n<p>Was very surprised, that there are not so much articles on this, most from 2014-16. Found useful this one: <a href=\"https://arxiv.org/abs/1610.00087\">\"Very Deep Convolutional Neural Networks for Raw Waveforms\"</a>, here the proposed models:\n<img src=\"https://lut.im/2p0Ky1mTOW/aUiNk1vH3T4h1G1g.png\" alt=\"Wav models\">\nI've chose M11 model as baseline for speed/accuracy reasons, obtained 75% best accuracy. Not so much. I've though, that now model has to many parameters, so increased percentage of unknown and silence words to 60% and 20% correspondingly. This improved score till 78%. After that added identity connections, thus switched to resnet-like architecture to help network learn increased parameters number and obtained <strong>80%</strong> accuracy. This model trains for 50 epoch, ~40 mins. Not so cool in comparison with <code>conv</code> model with the same result in ~15 minutes. This was the (unsuccessful) end of my experiments with raw wav input, I've decided to try this model on the mfcc input.</p>\n\n<p><strong>Stage 2. Architecture tuning on mfcc input</strong></p>\n\n<p>The first try of this model on mfcc input gives score <strong>84.4%</strong>. Next added second input channel - first order derivative of mfcc, known as delta. This improved LB on 0.9%. Next trick was to use <a href=\"https://arxiv.org/abs/1603.05201\">Concatenated Rectified Linear Unit (CReLU)</a> activation in shallow layers, which allows to improve generalization by learning half of filters, taking their responses and concatenating it with negated responses. This gives +0.6%. Last observation was that receptive field doesn't cover significant time scale, so inserted by one additional layer with 256 and 512 channels (after the layers with the same filters number) with dilation 2, obtained +0.4 improvement and <strong>86.3/87.3%</strong> accuracy on public/private LB. This was the best single model result. The final result is averaging results of 6 models, that gives additional 0.9% and <strong>88.2%</strong> at all.</p>\n\n<p>I've tried more extensive augmentation, add different pitch and generated noise, however this gives no improvement. Also not used test set during training, however this might boost the metric. Despite medalless result, I've learn a lot about audio classification, many thanks to the organizers and community!</p>",
  "messages": [
    {
      "id": 270750,
      "postDate": "2018-01-18T21:23:53.787Z",
      "content": "<p><em>I had no experience in dealing with audio files, so it was very interesting to try myself in this competition. Thanks all for</em> <code>hard coding</code> <em>in new year holidays!</em></p>\n\n<p>Here is my report on evaluation some nn building blocks for speech recognition task.</p>\n\n<p><strong>Stage 0. <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">Sample's</a> parameters tuning</strong></p>\n\n<p>Tensorflow's sample already can perform train/val split without picking same speaking person in both splits and has time shift/volume level augmentation, so it worth trying. It <code>conv</code> model with default parameters provides a nice baseline in <strong>77%</strong> on LB. I made one change, switched from SGD (converged slowly) to Adam, and training started to take <strong>~15 min</strong>. Next was reading a bunch of articles about audio processing, almost all suggested to use window size in 20 ms with 50% overlap and just 12 dct coefficients. With this setup, accuracy increased to <strong>80%</strong>. Now it's time to try more advanced models, decided to experiment with raw wav, since network usually learns features better, than handcrafted.</p>\n\n<p><strong>Stage 1. Raw wav experiments</strong></p>\n\n<p>Was very surprised, that there are not so much articles on this, most from 2014-16. Found useful this one: <a href=\"https://arxiv.org/abs/1610.00087\">\"Very Deep Convolutional Neural Networks for Raw Waveforms\"</a>, here the proposed models:\n<img src=\"https://lut.im/2p0Ky1mTOW/aUiNk1vH3T4h1G1g.png\" alt=\"Wav models\">\nI've chose M11 model as baseline for speed/accuracy reasons, obtained 75% best accuracy. Not so much. I've though, that now model has to many parameters, so increased percentage of unknown and silence words to 60% and 20% correspondingly. This improved score till 78%. After that added identity connections, thus switched to resnet-like architecture to help network learn increased parameters number and obtained <strong>80%</strong> accuracy. This model trains for 50 epoch, ~40 mins. Not so cool in comparison with <code>conv</code> model with the same result in ~15 minutes. This was the (unsuccessful) end of my experiments with raw wav input, I've decided to try this model on the mfcc input.</p>\n\n<p><strong>Stage 2. Architecture tuning on mfcc input</strong></p>\n\n<p>The first try of this model on mfcc input gives score <strong>84.4%</strong>. Next added second input channel - first order derivative of mfcc, known as delta. This improved LB on 0.9%. Next trick was to use <a href=\"https://arxiv.org/abs/1603.05201\">Concatenated Rectified Linear Unit (CReLU)</a> activation in shallow layers, which allows to improve generalization by learning half of filters, taking their responses and concatenating it with negated responses. This gives +0.6%. Last observation was that receptive field doesn't cover significant time scale, so inserted by one additional layer with 256 and 512 channels (after the layers with the same filters number) with dilation 2, obtained +0.4 improvement and <strong>86.3/87.3%</strong> accuracy on public/private LB. This was the best single model result. The final result is averaging results of 6 models, that gives additional 0.9% and <strong>88.2%</strong> at all.</p>\n\n<p>I've tried more extensive augmentation, add different pitch and generated noise, however this gives no improvement. Also not used test set during training, however this might boost the metric. Despite medalless result, I've learn a lot about audio classification, many thanks to the organizers and community!</p>",
      "rawMarkdown": "*I had no experience in dealing with audio files, so it was very interesting to try myself in this competition. Thanks all for* <code>hard coding</code> *in new year holidays!*\n\nHere is my report on evaluation some nn building blocks for speech recognition task.\n\n**Stage 0. [Sample's][1] parameters tuning**\n\nTensorflow's sample already can perform train/val split without picking same speaking person in both splits and has time shift/volume level augmentation, so it worth trying. It `conv` model with default parameters provides a nice baseline in **77%** on LB. I made one change, switched from SGD (converged slowly) to Adam, and training started to take **~15 min**. Next was reading a bunch of articles about audio processing, almost all suggested to use window size in 20 ms with 50% overlap and just 12 dct coefficients. With this setup, accuracy increased to **80%**. Now it's time to try more advanced models, decided to experiment with raw wav, since network usually learns features better, than handcrafted.\n\n**Stage 1. Raw wav experiments**\n\nWas very surprised, that there are not so much articles on this, most from 2014-16. Found useful this one: [\"Very Deep Convolutional Neural Networks for Raw Waveforms\"][2], here the proposed models:\n![Wav models][3]\nI've chose M11 model as baseline for speed/accuracy reasons, obtained 75% best accuracy. Not so much. I've though, that now model has to many parameters, so increased percentage of unknown and silence words to 60% and 20% correspondingly. This improved score till 78%. After that added identity connections, thus switched to resnet-like architecture to help network learn increased parameters number and obtained **80%** accuracy. This model trains for 50 epoch, ~40 mins. Not so cool in comparison with `conv` model with the same result in ~15 minutes. This was the (unsuccessful) end of my experiments with raw wav input, I've decided to try this model on the mfcc input.\n\n**Stage 2. Architecture tuning on mfcc input**\n\nThe first try of this model on mfcc input gives score **84.4%**. Next added second input channel - first order derivative of mfcc, known as delta. This improved LB on 0.9%. Next trick was to use [Concatenated Rectified Linear Unit (CReLU)][4] activation in shallow layers, which allows to improve generalization by learning half of filters, taking their responses and concatenating it with negated responses. This gives +0.6%. Last observation was that receptive field doesn't cover significant time scale, so inserted by one additional layer with 256 and 512 channels (after the layers with the same filters number) with dilation 2, obtained +0.4 improvement and **86.3/87.3%** accuracy on public/private LB. This was the best single model result. The final result is averaging results of 6 models, that gives additional 0.9% and **88.2%** at all.\n\nI've tried more extensive augmentation, add different pitch and generated noise, however this gives no improvement. Also not used test set during training, however this might boost the metric. Despite medalless result, I've learn a lot about audio classification, many thanks to the organizers and community!\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: https://arxiv.org/abs/1610.00087\n  [3]: https://lut.im/2p0Ky1mTOW/aUiNk1vH3T4h1G1g.png\n  [4]: https://arxiv.org/abs/1603.05201",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "270750": "*I had no experience in dealing with audio files, so it was very interesting to try myself in this competition. Thanks all for* <code>hard coding</code> *in new year holidays!*\n\nHere is my report on evaluation some nn building blocks for speech recognition task.\n\n**Stage 0. [Sample's][1] parameters tuning**\n\nTensorflow's sample already can perform train/val split without picking same speaking person in both splits and has time shift/volume level augmentation, so it worth trying. It `conv` model with default parameters provides a nice baseline in **77%** on LB. I made one change, switched from SGD (converged slowly) to Adam, and training started to take **~15 min**. Next was reading a bunch of articles about audio processing, almost all suggested to use window size in 20 ms with 50% overlap and just 12 dct coefficients. With this setup, accuracy increased to **80%**. Now it's time to try more advanced models, decided to experiment with raw wav, since network usually learns features better, than handcrafted.\n\n**Stage 1. Raw wav experiments**\n\nWas very surprised, that there are not so much articles on this, most from 2014-16. Found useful this one: [\"Very Deep Convolutional Neural Networks for Raw Waveforms\"][2], here the proposed models:\n![Wav models][3]\nI've chose M11 model as baseline for speed/accuracy reasons, obtained 75% best accuracy. Not so much. I've though, that now model has to many parameters, so increased percentage of unknown and silence words to 60% and 20% correspondingly. This improved score till 78%. After that added identity connections, thus switched to resnet-like architecture to help network learn increased parameters number and obtained **80%** accuracy. This model trains for 50 epoch, ~40 mins. Not so cool in comparison with `conv` model with the same result in ~15 minutes. This was the (unsuccessful) end of my experiments with raw wav input, I've decided to try this model on the mfcc input.\n\n**Stage 2. Architecture tuning on mfcc input**\n\nThe first try of this model on mfcc input gives score **84.4%**. Next added second input channel - first order derivative of mfcc, known as delta. This improved LB on 0.9%. Next trick was to use [Concatenated Rectified Linear Unit (CReLU)][4] activation in shallow layers, which allows to improve generalization by learning half of filters, taking their responses and concatenating it with negated responses. This gives +0.6%. Last observation was that receptive field doesn't cover significant time scale, so inserted by one additional layer with 256 and 512 channels (after the layers with the same filters number) with dilation 2, obtained +0.4 improvement and **86.3/87.3%** accuracy on public/private LB. This was the best single model result. The final result is averaging results of 6 models, that gives additional 0.9% and **88.2%** at all.\n\nI've tried more extensive augmentation, add different pitch and generated noise, however this gives no improvement. Also not used test set during training, however this might boost the metric. Despite medalless result, I've learn a lot about audio classification, many thanks to the organizers and community!\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: https://arxiv.org/abs/1610.00087\n  [3]: https://lut.im/2p0Ky1mTOW/aUiNk1vH3T4h1G1g.png\n  [4]: https://arxiv.org/abs/1603.05201"
  }
}