{
  "id": 47633,
  "title": "VGG16 neural network + openSMILE VAD (LB: 0.87)",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47633",
  "author_name": "",
  "post_date": "2018-01-17T06:07:24.446707600Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks for the interesting challenge! I've been working a lot with speech recognition models for my research, but this is my first time implementing one myself.</p>\n\n<p>In general, speech recognition is a difficult problem, but it’s much easier when the vocabulary is limited to a handful of words. We don’t need to use complicated language models to detect phonemes, and then string the phonemes into words, like <a href=\"http://kaldi-asr.org/\">Kaldi</a> does for speech recognition. Instead, a convolutional neural network works quite well.</p>\n\n<h1>First Steps</h1>\n\n<p>The dataset consists of about 64000 audio files which have already been split into training / validation / testing sets. You are then asked to make predictions on about 150000 audio files for which the labels are unknown.</p>\n\n<p>Actually, this dataset had already been published in academic literature, and people published code to solve the same problem. I started with <a href=\"https://github.com/adiyoss/GCommandsPytorch\">GCommandPytorch</a> by Yossi Adi, which implements a speech recognition CNN in Pytorch.</p>\n\n<p>The first step that it does is convert the audio file into a spectrogram, which is an image representation of sound. This is easily done using LibRosa.</p>\n\n<p><img src=\"https://i.imgur.com/bJ0FV9e.png\" alt=\"\">\n<em>Above: Sample spectrograms of “yes” and “no”</em></p>\n\n<p>Now we’ve converted the problem to an image classification problem, which is well studied. To an untrained human observer, all the spectrograms may look the same, but neural networks can learn things that humans can’t. Convolutional neural networks work very well for classifying images, for example <a href=\"https://arxiv.org/abs/1409.1556\">VGG16</a>:</p>\n\n<p><img src=\"https://luckytoilet.files.wordpress.com/2018/01/21.png\" alt=\"\">\n<em>Above: A Convolutional Neural Network (LeNet). VGG16 is similar, but has even more layers.</em></p>\n\n<p>For more details about this approach, refer to these papers:</p>\n\n<ol>\n<li><a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/43969.pdf\">Convolutional Neural Networks for Small-footprint Keyword Spotting</a></li>\n<li><a href=\"https://arxiv.org/abs/1710.06554\">Honk: A PyTorch Reimplementation of Convolutional Neural Networks for Keyword Spotting</a></li>\n</ol>\n\n<h1>Voice Activity Detection</h1>\n\n<p>You might ask: if somebody already implemented this, then what’s there left to do other than run their code? Well, the test data contains “silence” samples, which contain background noise but no human speech. It also has words outside the set we care about, which we need to label as “unknown”. The Pytorch CNN produces about 95% validation accuracy by itself, but the accuracy is much lower when we add these two additional requirements.</p>\n\n<p>For silence detection, I first tried the simplest thing I could think of: taking the maximum absolute value of the waveform and decide it’s “silence” if the value is below a threshold. When combined with VGG16, this gets accuracy <strong>0.78</strong> on the leaderboard. This is a crude metric because sufficiently loud noise would be considered speech.</p>\n\n<p>Next, I tried running <a href=\"http://audeering.com/technology/opensmile/\">openSMILE</a>, which I use in my research to extract various acoustic features from audio. It implements an <a href=\"http://ieeexplore.ieee.org/document/6637694/\">LSTM for voice activity detection</a>: every 0.05 seconds, it outputs a probability that someone is talking. Combining the openSMILE output with the VGG16 prediction gave a score of <strong>0.81</strong>.</p>\n\n<h1>More improvements</h1>\n\n<p>I tried a bunch of things to improve my score:</p>\n\n<ol>\n<li><p>Fiddled around with the neural network hyperparameters which boosted my score to <strong>0.85</strong>. Each epoch took about 10 minutes on a GPU, and the whole model takes about 2 hours to train. \nSomehow, <a href=\"https://arxiv.org/abs/1412.6980\">Adam</a> didn’t produce good results, and SGD with momentum worked better.</p></li>\n<li><p>Took 100% of the data for training and used the public LB for validation (don’t do this in real life lol). This improved my score to <strong>0.86</strong>.</p></li>\n<li><p>Trained an ensemble 3 versions of the same neural network with same hyperparameters but different randomly initialized weights and took a majority vote to do prediction. This improved the score to <strong>0.87</strong>. I would’ve liked to train more, but other people in my research group needed to use the GPUs.</p></li>\n</ol>\n\n<hr>\n\n<p>In the end, the top scoring model had a score of <strong>0.91</strong>, which beat my model by 4 percentage points. Although not enough to win a Kaggle medal, my model was in the top 15% of all submissions. Not bad!</p>\n\n<p><em>My source code for the contest is <a href=\"https://github.com/luckytoilet/kaggle-speech-recognition\">available here</a>.</em></p>\n\n<p><em>This post is also posted on <a href=\"https://luckytoilet.wordpress.com/2018/01/16/kaggle-speech-recognition-challenge/\">my blog</a>.</em></p>",
  "messages": [
    {
      "id": "269697",
      "postDate": "01/17/2018 06:07:24",
      "content": "<p>Thanks for the interesting challenge! I've been working a lot with speech recognition models for my research, but this is my first time implementing one myself.</p>\n\n<p>In general, speech recognition is a difficult problem, but it’s much easier when the vocabulary is limited to a handful of words. We don’t need to use complicated language models to detect phonemes, and then string the phonemes into words, like <a href=\"http://kaldi-asr.org/\">Kaldi</a> does for speech recognition. Instead, a convolutional neural network works quite well.</p>\n\n<h1>First Steps</h1>\n\n<p>The dataset consists of about 64000 audio files which have already been split into training / validation / testing sets. You are then asked to make predictions on about 150000 audio files for which the labels are unknown.</p>\n\n<p>Actually, this dataset had already been published in academic literature, and people published code to solve the same problem. I started with <a href=\"https://github.com/adiyoss/GCommandsPytorch\">GCommandPytorch</a> by Yossi Adi, which implements a speech recognition CNN in Pytorch.</p>\n\n<p>The first step that it does is convert the audio file into a spectrogram, which is an image representation of sound. This is easily done using LibRosa.</p>\n\n<p><img src=\"https://i.imgur.com/bJ0FV9e.png\" alt=\"\">\n<em>Above: Sample spectrograms of “yes” and “no”</em></p>\n\n<p>Now we’ve converted the problem to an image classification problem, which is well studied. To an untrained human observer, all the spectrograms may look the same, but neural networks can learn things that humans can’t. Convolutional neural networks work very well for classifying images, for example <a href=\"https://arxiv.org/abs/1409.1556\">VGG16</a>:</p>\n\n<p><img src=\"https://luckytoilet.files.wordpress.com/2018/01/21.png\" alt=\"\">\n<em>Above: A Convolutional Neural Network (LeNet). VGG16 is similar, but has even more layers.</em></p>\n\n<p>For more details about this approach, refer to these papers:</p>\n\n<ol>\n<li><a href=\"https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/43969.pdf\">Convolutional Neural Networks for Small-footprint Keyword Spotting</a></li>\n<li><a href=\"https://arxiv.org/abs/1710.06554\">Honk: A PyTorch Reimplementation of Convolutional Neural Networks for Keyword Spotting</a></li>\n</ol>\n\n<h1>Voice Activity Detection</h1>\n\n<p>You might ask: if somebody already implemented this, then what’s there left to do other than run their code? Well, the test data contains “silence” samples, which contain background noise but no human speech. It also has words outside the set we care about, which we need to label as “unknown”. The Pytorch CNN produces about 95% validation accuracy by itself, but the accuracy is much lower when we add these two additional requirements.</p>\n\n<p>For silence detection, I first tried the simplest thing I could think of: taking the maximum absolute value of the waveform and decide it’s “silence” if the value is below a threshold. When combined with VGG16, this gets accuracy <strong>0.78</strong> on the leaderboard. This is a crude metric because sufficiently loud noise would be considered speech.</p>\n\n<p>Next, I tried running <a href=\"http://audeering.com/technology/opensmile/\">openSMILE</a>, which I use in my research to extract various acoustic features from audio. It implements an <a href=\"http://ieeexplore.ieee.org/document/6637694/\">LSTM for voice activity detection</a>: every 0.05 seconds, it outputs a probability that someone is talking. Combining the openSMILE output with the VGG16 prediction gave a score of <strong>0.81</strong>.</p>\n\n<h1>More improvements</h1>\n\n<p>I tried a bunch of things to improve my score:</p>\n\n<ol>\n<li><p>Fiddled around with the neural network hyperparameters which boosted my score to <strong>0.85</strong>. Each epoch took about 10 minutes on a GPU, and the whole model takes about 2 hours to train. \nSomehow, <a href=\"https://arxiv.org/abs/1412.6980\">Adam</a> didn’t produce good results, and SGD with momentum worked better.</p></li>\n<li><p>Took 100% of the data for training and used the public LB for validation (don’t do this in real life lol). This improved my score to <strong>0.86</strong>.</p></li>\n<li><p>Trained an ensemble 3 versions of the same neural network with same hyperparameters but different randomly initialized weights and took a majority vote to do prediction. This improved the score to <strong>0.87</strong>. I would’ve liked to train more, but other people in my research group needed to use the GPUs.</p></li>\n</ol>\n\n<hr>\n\n<p>In the end, the top scoring model had a score of <strong>0.91</strong>, which beat my model by 4 percentage points. Although not enough to win a Kaggle medal, my model was in the top 15% of all submissions. Not bad!</p>\n\n<p><em>My source code for the contest is <a href=\"https://github.com/luckytoilet/kaggle-speech-recognition\">available here</a>.</em></p>\n\n<p><em>This post is also posted on <a href=\"https://luckytoilet.wordpress.com/2018/01/16/kaggle-speech-recognition-challenge/\">my blog</a>.</em></p>",
      "rawMarkdown": "Thanks for the interesting challenge! I've been working a lot with speech recognition models for my research, but this is my first time implementing one myself.\n\nIn general, speech recognition is a difficult problem, but it’s much easier when the vocabulary is limited to a handful of words. We don’t need to use complicated language models to detect phonemes, and then string the phonemes into words, like [Kaldi](http://kaldi-asr.org/) does for speech recognition. Instead, a convolutional neural network works quite well.\n\n# First Steps\n\nThe dataset consists of about 64000 audio files which have already been split into training / validation / testing sets. You are then asked to make predictions on about 150000 audio files for which the labels are unknown.\n\nActually, this dataset had already been published in academic literature, and people published code to solve the same problem. I started with [GCommandPytorch](https://github.com/adiyoss/GCommandsPytorch) by Yossi Adi, which implements a speech recognition CNN in Pytorch.\n\nThe first step that it does is convert the audio file into a spectrogram, which is an image representation of sound. This is easily done using LibRosa.\n\n![][1]\n*Above: Sample spectrograms of “yes” and “no”*\n\nNow we’ve converted the problem to an image classification problem, which is well studied. To an untrained human observer, all the spectrograms may look the same, but neural networks can learn things that humans can’t. Convolutional neural networks work very well for classifying images, for example [VGG16](https://arxiv.org/abs/1409.1556):\n\n![][2]\n*Above: A Convolutional Neural Network (LeNet). VGG16 is similar, but has even more layers.*\n\nFor more details about this approach, refer to these papers:\n\n1. [Convolutional Neural Networks for Small-footprint Keyword Spotting](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/43969.pdf)\n2. [Honk: A PyTorch Reimplementation of Convolutional Neural Networks for Keyword Spotting](https://arxiv.org/abs/1710.06554)\n\n#Voice Activity Detection\n\nYou might ask: if somebody already implemented this, then what’s there left to do other than run their code? Well, the test data contains “silence” samples, which contain background noise but no human speech. It also has words outside the set we care about, which we need to label as “unknown”. The Pytorch CNN produces about 95% validation accuracy by itself, but the accuracy is much lower when we add these two additional requirements.\n\nFor silence detection, I first tried the simplest thing I could think of: taking the maximum absolute value of the waveform and decide it’s “silence” if the value is below a threshold. When combined with VGG16, this gets accuracy **0.78** on the leaderboard. This is a crude metric because sufficiently loud noise would be considered speech.\n\nNext, I tried running [openSMILE](http://audeering.com/technology/opensmile/), which I use in my research to extract various acoustic features from audio. It implements an [LSTM for voice activity detection](http://ieeexplore.ieee.org/document/6637694/): every 0.05 seconds, it outputs a probability that someone is talking. Combining the openSMILE output with the VGG16 prediction gave a score of **0.81**.\n\n# More improvements\n\nI tried a bunch of things to improve my score:\n\n1. Fiddled around with the neural network hyperparameters which boosted my score to **0.85**. Each epoch took about 10 minutes on a GPU, and the whole model takes about 2 hours to train. \n Somehow, [Adam](https://arxiv.org/abs/1412.6980) didn’t produce good results, and SGD with momentum worked better.\n\n2. Took 100% of the data for training and used the public LB for validation (don’t do this in real life lol). This improved my score to **0.86**.\n\n3. Trained an ensemble 3 versions of the same neural network with same hyperparameters but different randomly initialized weights and took a majority vote to do prediction. This improved the score to **0.87**. I would’ve liked to train more, but other people in my research group needed to use the GPUs.\n\n---\n\nIn the end, the top scoring model had a score of **0.91**, which beat my model by 4 percentage points. Although not enough to win a Kaggle medal, my model was in the top 15% of all submissions. Not bad!\n\n*My source code for the contest is [available here](https://github.com/luckytoilet/kaggle-speech-recognition).*\n\n*This post is also posted on [my blog](https://luckytoilet.wordpress.com/2018/01/16/kaggle-speech-recognition-challenge/).*\n\n\n  [1]: https://i.imgur.com/bJ0FV9e.png\n  [2]: https://luckytoilet.files.wordpress.com/2018/01/21.png",
      "votes": null
    },
    {
      "id": "270309",
      "postDate": "01/18/2018 03:32:22",
      "content": "<p>Awsome，Thanks！</p>",
      "rawMarkdown": "Awsome，Thanks！",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 270309,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "01/18/2018 03:32:22",
      "content": "<p>Awsome，Thanks！</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "269697": "Thanks for the interesting challenge! I've been working a lot with speech recognition models for my research, but this is my first time implementing one myself.\n\nIn general, speech recognition is a difficult problem, but it’s much easier when the vocabulary is limited to a handful of words. We don’t need to use complicated language models to detect phonemes, and then string the phonemes into words, like [Kaldi](http://kaldi-asr.org/) does for speech recognition. Instead, a convolutional neural network works quite well.\n\n# First Steps\n\nThe dataset consists of about 64000 audio files which have already been split into training / validation / testing sets. You are then asked to make predictions on about 150000 audio files for which the labels are unknown.\n\nActually, this dataset had already been published in academic literature, and people published code to solve the same problem. I started with [GCommandPytorch](https://github.com/adiyoss/GCommandsPytorch) by Yossi Adi, which implements a speech recognition CNN in Pytorch.\n\nThe first step that it does is convert the audio file into a spectrogram, which is an image representation of sound. This is easily done using LibRosa.\n\n![][1]\n*Above: Sample spectrograms of “yes” and “no”*\n\nNow we’ve converted the problem to an image classification problem, which is well studied. To an untrained human observer, all the spectrograms may look the same, but neural networks can learn things that humans can’t. Convolutional neural networks work very well for classifying images, for example [VGG16](https://arxiv.org/abs/1409.1556):\n\n![][2]\n*Above: A Convolutional Neural Network (LeNet). VGG16 is similar, but has even more layers.*\n\nFor more details about this approach, refer to these papers:\n\n1. [Convolutional Neural Networks for Small-footprint Keyword Spotting](https://static.googleusercontent.com/media/research.google.com/en//pubs/archive/43969.pdf)\n2. [Honk: A PyTorch Reimplementation of Convolutional Neural Networks for Keyword Spotting](https://arxiv.org/abs/1710.06554)\n\n#Voice Activity Detection\n\nYou might ask: if somebody already implemented this, then what’s there left to do other than run their code? Well, the test data contains “silence” samples, which contain background noise but no human speech. It also has words outside the set we care about, which we need to label as “unknown”. The Pytorch CNN produces about 95% validation accuracy by itself, but the accuracy is much lower when we add these two additional requirements.\n\nFor silence detection, I first tried the simplest thing I could think of: taking the maximum absolute value of the waveform and decide it’s “silence” if the value is below a threshold. When combined with VGG16, this gets accuracy **0.78** on the leaderboard. This is a crude metric because sufficiently loud noise would be considered speech.\n\nNext, I tried running [openSMILE](http://audeering.com/technology/opensmile/), which I use in my research to extract various acoustic features from audio. It implements an [LSTM for voice activity detection](http://ieeexplore.ieee.org/document/6637694/): every 0.05 seconds, it outputs a probability that someone is talking. Combining the openSMILE output with the VGG16 prediction gave a score of **0.81**.\n\n# More improvements\n\nI tried a bunch of things to improve my score:\n\n1. Fiddled around with the neural network hyperparameters which boosted my score to **0.85**. Each epoch took about 10 minutes on a GPU, and the whole model takes about 2 hours to train. \n Somehow, [Adam](https://arxiv.org/abs/1412.6980) didn’t produce good results, and SGD with momentum worked better.\n\n2. Took 100% of the data for training and used the public LB for validation (don’t do this in real life lol). This improved my score to **0.86**.\n\n3. Trained an ensemble 3 versions of the same neural network with same hyperparameters but different randomly initialized weights and took a majority vote to do prediction. This improved the score to **0.87**. I would’ve liked to train more, but other people in my research group needed to use the GPUs.\n\n---\n\nIn the end, the top scoring model had a score of **0.91**, which beat my model by 4 percentage points. Although not enough to win a Kaggle medal, my model was in the top 15% of all submissions. Not bad!\n\n*My source code for the contest is [available here](https://github.com/luckytoilet/kaggle-speech-recognition).*\n\n*This post is also posted on [my blog](https://luckytoilet.wordpress.com/2018/01/16/kaggle-speech-recognition-challenge/).*\n\n\n  [1]: https://i.imgur.com/bJ0FV9e.png\n  [2]: https://luckytoilet.files.wordpress.com/2018/01/21.png",
    "270309": "Awsome，Thanks！"
  },
  "source": "meta"
}