{
  "id": 47674,
  "title": "Our approach [4th place]",
  "url": "/competitions/tensorflow-speech-recognition-challenge/writeups/high-five-our-approach-4th-place",
  "author_name": "",
  "post_date": "2018-01-17T16:10:05.718519300Z",
  "votes": 48,
  "comment_count": 6,
  "views": 0,
  "content": "<p>First of all, thanks to all participants and congratulations to the winners!</p>\n\n<p>I’ll describe my approach and our final solution. My colleagues will provide more details about their techniques and tricks as well.</p>\n\n<p>I started to participate in this competition three weeks ago. It was enough for training many different models and building an ensemble.</p>\n\n<p>I splitted train dataset on 10 folds by speaker_id and generated 6000 silence files using background noise files provided by organizers. Next 1,5 weeks I was training different neural networks.</p>\n\n<p>Almost all of my models have VGG-like and ResNet-like architecture. I just used different number of layers and different number of filters on convolutional layers. Also, I used different types of .wav preprocessing:</p>\n\n<ul>\n<li>FFT</li>\n<li>MFCC</li>\n<li>Mel spectogram</li>\n<li>Chroma fft</li>\n<li>Tempogram</li>\n<li>Raw 16000-dimensional vector</li>\n</ul>\n\n<p>My CNNs were trained using disbalanced dataset. It was okay because I did not want to use these models as is. I also trained several RNNs to increase diversity in an ensemble. For augmentation I used:</p>\n\n<ul>\n<li>Pitch shift</li>\n<li>Time stretch</li>\n<li>Time shift</li>\n<li>Random noise</li>\n</ul>\n\n<p>Next, for each fold I extracted predictions from each model. I did not apply softmax at the outputs to keep more information about predictions. Then, I trained xgboost on class-balanced data and got #11 place.</p>\n\n<p>It was time to merge with other participants :) We merged with Aleksey, Giba, Dmytro and feels_g00d_man. The next challenge was to merge our solutions. Averaging best submissions did not work and did not improve our place at leaderboard.</p>\n\n<p>We combined all L1 models and created one big csv file. It was passed to many different L2 models:</p>\n\n<ul>\n<li>XGBoost</li>\n<li>LGBM</li>\n<li>Catboost</li>\n<li>Simple neural networks</li>\n<li>Random forest</li>\n<li>Extra trees</li>\n<li>Adaboost</li>\n<li>kNN</li>\n<li>etc.</li>\n</ul>\n\n<p>I tried almost everything :) Special like to LGBM because this implementation of gradient boosting is very-very fast. XGBoost needs ~8 hours for training 10 folds on GPU and ~14 hours on CPU while LGBM needs only 15 minutes on CPU!</p>\n\n<p>Next, I used weighted geometric mean of L2 models (thus, it was very simple L3 model).\nTo find optimal weights I built a simple neural network with custom regularizer (all weights &gt;= 0 and the sum of weights is equal to 1).</p>\n\n<p><strong>SILENCE TRICK</strong></p>\n\n<p>It was very interesting to work with silence because we didn’t have train examples for this class. I found a very simple way for increasing LB score. Let’s look at the test predictions and their probabilities and sort them by confidence (maximal probability). Next, let’s select top K samples with the lowest confidence (I used K=10000) and compute power_level = np.max(librosa.feature.melspectrogram). Then, let’s interpret all selected samples with power_level &lt; L (I used L=1) as silence. This simple trick allows to increase LB score on ~0.005</p>\n\n<p>After some time, we have replaced this trick with a separate silence/no-silence model trained using semi-supervised approach. It slightly increased our score (+0.001 on private LB).</p>\n\n<p><strong>UNKNOWN TRICK</strong></p>\n\n<p>Since test set contains unknown unknowns we needed to find a way to work with such samples. feels_g00d_man found a very interesting method to do it. His approach described below.</p>\n\n<p><strong>OVERALL</strong></p>\n\n<p>The test dataset was very strange and we could not find a way how to validate our models. CV score did not correlate with public LB. Usually, increasing CV score was leading to decreasing LB score and it was annoying. Also, the hardest problem was how to select 3 final submissions. We expected a big shake-up and I was pleasantly surprised when I saw private LB. For final submissions we selected the best submit based on public LB, mode blend of 16 different submissions and one submit with a large number of unknowns in predictions. Our best public submit is our best private as well. </p>\n\n<p>Thanks to the organizers, it was a very interesting competition for me because I have never worked with speech data and because of challenge with silence and unknown words :)</p>\n\n<p>I also want to thank all my teammates, it was very cool to work with such great data scientists!</p>\n\n<p>See you in other competitions :)</p>\n\n<p><strong>feels_g00d_man's comment</strong></p>\n\n<p>&gt; \nI teamed up with guys right before the merger deadline. So my initial solution was to train around 10 models on all 31 classes using only spectrograms on 5 folds with TTA and stack them with L2 xgboost. After teaming up I was focusing on training L1 models in our big ensemble. Also I trained separate model for unknown unknowns by generating \"double words\" and we replaced unknown unknowns from this model in our final ensembles (that gave 0.24 boost on private LB).\nI pushed source code, described approach and \"double words\" on <a href=\"https://github.com/heyt0ny/TensorFlow-Speech-Recognition-Challenge-Solution\">https://github.com/heyt0ny/TensorFlow-Speech-Recognition-Challenge-Solution</a></p>\n\n<p><img src=\"https://i.imgur.com/yeh7Sh7.png\" alt=\"Double words\"></p>",
  "messages": [
    {
      "id": "269989",
      "postDate": "01/17/2018 16:10:05",
      "content": "<p>First of all, thanks to all participants and congratulations to the winners!</p>\n\n<p>I’ll describe my approach and our final solution. My colleagues will provide more details about their techniques and tricks as well.</p>\n\n<p>I started to participate in this competition three weeks ago. It was enough for training many different models and building an ensemble.</p>\n\n<p>I splitted train dataset on 10 folds by speaker_id and generated 6000 silence files using background noise files provided by organizers. Next 1,5 weeks I was training different neural networks.</p>\n\n<p>Almost all of my models have VGG-like and ResNet-like architecture. I just used different number of layers and different number of filters on convolutional layers. Also, I used different types of .wav preprocessing:</p>\n\n<ul>\n<li>FFT</li>\n<li>MFCC</li>\n<li>Mel spectogram</li>\n<li>Chroma fft</li>\n<li>Tempogram</li>\n<li>Raw 16000-dimensional vector</li>\n</ul>\n\n<p>My CNNs were trained using disbalanced dataset. It was okay because I did not want to use these models as is. I also trained several RNNs to increase diversity in an ensemble. For augmentation I used:</p>\n\n<ul>\n<li>Pitch shift</li>\n<li>Time stretch</li>\n<li>Time shift</li>\n<li>Random noise</li>\n</ul>\n\n<p>Next, for each fold I extracted predictions from each model. I did not apply softmax at the outputs to keep more information about predictions. Then, I trained xgboost on class-balanced data and got #11 place.</p>\n\n<p>It was time to merge with other participants :) We merged with Aleksey, Giba, Dmytro and feels_g00d_man. The next challenge was to merge our solutions. Averaging best submissions did not work and did not improve our place at leaderboard.</p>\n\n<p>We combined all L1 models and created one big csv file. It was passed to many different L2 models:</p>\n\n<ul>\n<li>XGBoost</li>\n<li>LGBM</li>\n<li>Catboost</li>\n<li>Simple neural networks</li>\n<li>Random forest</li>\n<li>Extra trees</li>\n<li>Adaboost</li>\n<li>kNN</li>\n<li>etc.</li>\n</ul>\n\n<p>I tried almost everything :) Special like to LGBM because this implementation of gradient boosting is very-very fast. XGBoost needs ~8 hours for training 10 folds on GPU and ~14 hours on CPU while LGBM needs only 15 minutes on CPU!</p>\n\n<p>Next, I used weighted geometric mean of L2 models (thus, it was very simple L3 model).\nTo find optimal weights I built a simple neural network with custom regularizer (all weights &gt;= 0 and the sum of weights is equal to 1).</p>\n\n<p><strong>SILENCE TRICK</strong></p>\n\n<p>It was very interesting to work with silence because we didn’t have train examples for this class. I found a very simple way for increasing LB score. Let’s look at the test predictions and their probabilities and sort them by confidence (maximal probability). Next, let’s select top K samples with the lowest confidence (I used K=10000) and compute power_level = np.max(librosa.feature.melspectrogram). Then, let’s interpret all selected samples with power_level &lt; L (I used L=1) as silence. This simple trick allows to increase LB score on ~0.005</p>\n\n<p>After some time, we have replaced this trick with a separate silence/no-silence model trained using semi-supervised approach. It slightly increased our score (+0.001 on private LB).</p>\n\n<p><strong>UNKNOWN TRICK</strong></p>\n\n<p>Since test set contains unknown unknowns we needed to find a way to work with such samples. feels_g00d_man found a very interesting method to do it. His approach described below.</p>\n\n<p><strong>OVERALL</strong></p>\n\n<p>The test dataset was very strange and we could not find a way how to validate our models. CV score did not correlate with public LB. Usually, increasing CV score was leading to decreasing LB score and it was annoying. Also, the hardest problem was how to select 3 final submissions. We expected a big shake-up and I was pleasantly surprised when I saw private LB. For final submissions we selected the best submit based on public LB, mode blend of 16 different submissions and one submit with a large number of unknowns in predictions. Our best public submit is our best private as well. </p>\n\n<p>Thanks to the organizers, it was a very interesting competition for me because I have never worked with speech data and because of challenge with silence and unknown words :)</p>\n\n<p>I also want to thank all my teammates, it was very cool to work with such great data scientists!</p>\n\n<p>See you in other competitions :)</p>\n\n<p><strong>feels_g00d_man's comment</strong></p>\n\n<p>&gt; \nI teamed up with guys right before the merger deadline. So my initial solution was to train around 10 models on all 31 classes using only spectrograms on 5 folds with TTA and stack them with L2 xgboost. After teaming up I was focusing on training L1 models in our big ensemble. Also I trained separate model for unknown unknowns by generating \"double words\" and we replaced unknown unknowns from this model in our final ensembles (that gave 0.24 boost on private LB).\nI pushed source code, described approach and \"double words\" on <a href=\"https://github.com/heyt0ny/TensorFlow-Speech-Recognition-Challenge-Solution\">https://github.com/heyt0ny/TensorFlow-Speech-Recognition-Challenge-Solution</a></p>\n\n<p><img src=\"https://i.imgur.com/yeh7Sh7.png\" alt=\"Double words\"></p>",
      "rawMarkdown": "First of all, thanks to all participants and congratulations to the winners!\n\nI’ll describe my approach and our final solution. My colleagues will provide more details about their techniques and tricks as well.\n\nI started to participate in this competition three weeks ago. It was enough for training many different models and building an ensemble.\n\nI splitted train dataset on 10 folds by speaker_id and generated 6000 silence files using background noise files provided by organizers. Next 1,5 weeks I was training different neural networks.\n\nAlmost all of my models have VGG-like and ResNet-like architecture. I just used different number of layers and different number of filters on convolutional layers. Also, I used different types of .wav preprocessing:\n\n - FFT\n - MFCC\n - Mel spectogram\n - Chroma fft\n - Tempogram\n - Raw 16000-dimensional vector\n\nMy CNNs were trained using disbalanced dataset. It was okay because I did not want to use these models as is. I also trained several RNNs to increase diversity in an ensemble. For augmentation I used:\n\n - Pitch shift\n - Time stretch\n - Time shift\n - Random noise\n\nNext, for each fold I extracted predictions from each model. I did not apply softmax at the outputs to keep more information about predictions. Then, I trained xgboost on class-balanced data and got #11 place.\n\nIt was time to merge with other participants :) We merged with Aleksey, Giba, Dmytro and feels_g00d_man. The next challenge was to merge our solutions. Averaging best submissions did not work and did not improve our place at leaderboard.\n\nWe combined all L1 models and created one big csv file. It was passed to many different L2 models:\n\n - XGBoost\n - LGBM\n - Catboost\n - Simple neural networks\n - Random forest\n - Extra trees\n - Adaboost\n - kNN\n - etc.\n\nI tried almost everything :) Special like to LGBM because this implementation of gradient boosting is very-very fast. XGBoost needs ~8 hours for training 10 folds on GPU and ~14 hours on CPU while LGBM needs only 15 minutes on CPU!\n\nNext, I used weighted geometric mean of L2 models (thus, it was very simple L3 model).\nTo find optimal weights I built a simple neural network with custom regularizer (all weights &gt;= 0 and the sum of weights is equal to 1).\n\n\n**SILENCE TRICK**\n\nIt was very interesting to work with silence because we didn’t have train examples for this class. I found a very simple way for increasing LB score. Let’s look at the test predictions and their probabilities and sort them by confidence (maximal probability). Next, let’s select top K samples with the lowest confidence (I used K=10000) and compute power_level = np.max(librosa.feature.melspectrogram). Then, let’s interpret all selected samples with power_level &lt; L (I used L=1) as silence. This simple trick allows to increase LB score on ~0.005\n\nAfter some time, we have replaced this trick with a separate silence/no-silence model trained using semi-supervised approach. It slightly increased our score (+0.001 on private LB).\n\n\n**UNKNOWN TRICK**\n\nSince test set contains unknown unknowns we needed to find a way to work with such samples. feels_g00d_man found a very interesting method to do it. His approach described below.\n\n**OVERALL**\n\nThe test dataset was very strange and we could not find a way how to validate our models. CV score did not correlate with public LB. Usually, increasing CV score was leading to decreasing LB score and it was annoying. Also, the hardest problem was how to select 3 final submissions. We expected a big shake-up and I was pleasantly surprised when I saw private LB. For final submissions we selected the best submit based on public LB, mode blend of 16 different submissions and one submit with a large number of unknowns in predictions. Our best public submit is our best private as well. \n\nThanks to the organizers, it was a very interesting competition for me because I have never worked with speech data and because of challenge with silence and unknown words :)\n\nI also want to thank all my teammates, it was very cool to work with such great data scientists!\n\nSee you in other competitions :)\n\n**feels_g00d_man's comment**\n\n&gt; \nI teamed up with guys right before the merger deadline. So my initial solution was to train around 10 models on all 31 classes using only spectrograms on 5 folds with TTA and stack them with L2 xgboost. After teaming up I was focusing on training L1 models in our big ensemble. Also I trained separate model for unknown unknowns by generating \"double words\" and we replaced unknown unknowns from this model in our final ensembles (that gave 0.24 boost on private LB).\nI pushed source code, described approach and \"double words\" on https://github.com/heyt0ny/TensorFlow-Speech-Recognition-Challenge-Solution\n\n![Double words][1]\n\n  [1]: https://i.imgur.com/yeh7Sh7.png",
      "votes": null
    },
    {
      "id": "270533",
      "postDate": "01/18/2018 12:49:20",
      "content": "<p>Thanks for the great writeup Pavel--very comprehensive and a lot of hard work.   Congratuations to you and team!   Our big team was trying to catch yours but did not quite get there--all in good Kaggle fun and a \"high five\" from us to you :)</p>",
      "rawMarkdown": "Thanks for the great writeup Pavel--very comprehensive and a lot of hard work.   Congratuations to you and team!   Our big team was trying to catch yours but did not quite get there--all in good Kaggle fun and a \"high five\" from us to you :)",
      "votes": null
    },
    {
      "id": "271550",
      "postDate": "01/20/2018 18:54:18",
      "content": "<p>Nice approach and thanks for sharing.</p>",
      "rawMarkdown": "Nice approach and thanks for sharing.",
      "votes": null
    },
    {
      "id": "274768",
      "postDate": "01/27/2018 08:45:22",
      "content": "<p>what are \"fegolib\" and \"cv2\"? which are referred in your source code(main.py, etc.)</p>",
      "rawMarkdown": "what are \"fegolib\" and \"cv2\"? which are referred in your source code(main.py, etc.)",
      "votes": null
    },
    {
      "id": "310200",
      "postDate": "04/06/2018 19:48:50",
      "content": "<p>I think <code>cv2</code> is OpenCV.</p>",
      "rawMarkdown": "I think `cv2` is OpenCV.",
      "votes": null
    },
    {
      "id": "418702",
      "postDate": "11/10/2018 13:10:50",
      "content": "<p>Good approach</p>",
      "rawMarkdown": "Good approach",
      "votes": null
    },
    {
      "id": "427038",
      "postDate": "11/24/2018 11:59:21",
      "content": "<p>666</p>",
      "rawMarkdown": "666",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 270533,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "01/18/2018 12:49:20",
      "content": "<p>Thanks for the great writeup Pavel--very comprehensive and a lot of hard work.   Congratuations to you and team!   Our big team was trying to catch yours but did not quite get there--all in good Kaggle fun and a \"high five\" from us to you :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 271550,
      "author_name": "porfyriosg",
      "author_url": "",
      "post_date": "01/20/2018 18:54:18",
      "content": "<p>Nice approach and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 274768,
      "author_name": "rtygbwwwerr",
      "author_url": "",
      "post_date": "01/27/2018 08:45:22",
      "content": "<p>what are \"fegolib\" and \"cv2\"? which are referred in your source code(main.py, etc.)</p>",
      "votes": null,
      "replies": [
        {
          "id": 310200,
          "author_name": "pointyointment",
          "author_url": "",
          "post_date": "04/06/2018 19:48:50",
          "content": "<p>I think <code>cv2</code> is OpenCV.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 418702,
      "author_name": "arunkumarramanan",
      "author_url": "",
      "post_date": "11/10/2018 13:10:50",
      "content": "<p>Good approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 427038,
      "author_name": "knightpeng",
      "author_url": "",
      "post_date": "11/24/2018 11:59:21",
      "content": "<p>666</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "269989": "First of all, thanks to all participants and congratulations to the winners!\n\nI’ll describe my approach and our final solution. My colleagues will provide more details about their techniques and tricks as well.\n\nI started to participate in this competition three weeks ago. It was enough for training many different models and building an ensemble.\n\nI splitted train dataset on 10 folds by speaker_id and generated 6000 silence files using background noise files provided by organizers. Next 1,5 weeks I was training different neural networks.\n\nAlmost all of my models have VGG-like and ResNet-like architecture. I just used different number of layers and different number of filters on convolutional layers. Also, I used different types of .wav preprocessing:\n\n - FFT\n - MFCC\n - Mel spectogram\n - Chroma fft\n - Tempogram\n - Raw 16000-dimensional vector\n\nMy CNNs were trained using disbalanced dataset. It was okay because I did not want to use these models as is. I also trained several RNNs to increase diversity in an ensemble. For augmentation I used:\n\n - Pitch shift\n - Time stretch\n - Time shift\n - Random noise\n\nNext, for each fold I extracted predictions from each model. I did not apply softmax at the outputs to keep more information about predictions. Then, I trained xgboost on class-balanced data and got #11 place.\n\nIt was time to merge with other participants :) We merged with Aleksey, Giba, Dmytro and feels_g00d_man. The next challenge was to merge our solutions. Averaging best submissions did not work and did not improve our place at leaderboard.\n\nWe combined all L1 models and created one big csv file. It was passed to many different L2 models:\n\n - XGBoost\n - LGBM\n - Catboost\n - Simple neural networks\n - Random forest\n - Extra trees\n - Adaboost\n - kNN\n - etc.\n\nI tried almost everything :) Special like to LGBM because this implementation of gradient boosting is very-very fast. XGBoost needs ~8 hours for training 10 folds on GPU and ~14 hours on CPU while LGBM needs only 15 minutes on CPU!\n\nNext, I used weighted geometric mean of L2 models (thus, it was very simple L3 model).\nTo find optimal weights I built a simple neural network with custom regularizer (all weights &gt;= 0 and the sum of weights is equal to 1).\n\n\n**SILENCE TRICK**\n\nIt was very interesting to work with silence because we didn’t have train examples for this class. I found a very simple way for increasing LB score. Let’s look at the test predictions and their probabilities and sort them by confidence (maximal probability). Next, let’s select top K samples with the lowest confidence (I used K=10000) and compute power_level = np.max(librosa.feature.melspectrogram). Then, let’s interpret all selected samples with power_level &lt; L (I used L=1) as silence. This simple trick allows to increase LB score on ~0.005\n\nAfter some time, we have replaced this trick with a separate silence/no-silence model trained using semi-supervised approach. It slightly increased our score (+0.001 on private LB).\n\n\n**UNKNOWN TRICK**\n\nSince test set contains unknown unknowns we needed to find a way to work with such samples. feels_g00d_man found a very interesting method to do it. His approach described below.\n\n**OVERALL**\n\nThe test dataset was very strange and we could not find a way how to validate our models. CV score did not correlate with public LB. Usually, increasing CV score was leading to decreasing LB score and it was annoying. Also, the hardest problem was how to select 3 final submissions. We expected a big shake-up and I was pleasantly surprised when I saw private LB. For final submissions we selected the best submit based on public LB, mode blend of 16 different submissions and one submit with a large number of unknowns in predictions. Our best public submit is our best private as well. \n\nThanks to the organizers, it was a very interesting competition for me because I have never worked with speech data and because of challenge with silence and unknown words :)\n\nI also want to thank all my teammates, it was very cool to work with such great data scientists!\n\nSee you in other competitions :)\n\n**feels_g00d_man's comment**\n\n&gt; \nI teamed up with guys right before the merger deadline. So my initial solution was to train around 10 models on all 31 classes using only spectrograms on 5 folds with TTA and stack them with L2 xgboost. After teaming up I was focusing on training L1 models in our big ensemble. Also I trained separate model for unknown unknowns by generating \"double words\" and we replaced unknown unknowns from this model in our final ensembles (that gave 0.24 boost on private LB).\nI pushed source code, described approach and \"double words\" on https://github.com/heyt0ny/TensorFlow-Speech-Recognition-Challenge-Solution\n\n![Double words][1]\n\n  [1]: https://i.imgur.com/yeh7Sh7.png",
    "270533": "Thanks for the great writeup Pavel--very comprehensive and a lot of hard work.   Congratuations to you and team!   Our big team was trying to catch yours but did not quite get there--all in good Kaggle fun and a \"high five\" from us to you :)",
    "271550": "Nice approach and thanks for sharing.",
    "274768": "what are \"fegolib\" and \"cv2\"? which are referred in your source code(main.py, etc.)",
    "310200": "I think `cv2` is OpenCV.",
    "418702": "Good approach",
    "427038": "666"
  },
  "source": "meta"
}