{
  "id": 47687,
  "title": "Top 5% Solution Source Code ",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47687",
  "author_name": "SubhojeetPramanik",
  "post_date": "2018-01-17T18:07:08.056000",
  "votes": 12,
  "comment_count": 0,
  "views": 0,
  "content": "<p><a href=\"https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\">https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution</a></p>\n\n<p>Congratulations to all the winners! This was a wonderful experience, also thank you Google for the free credits without which I cannot imagine taking part in the competition. Also thank you Heng for your posts throughout the competition, your tips helped to jump many places in the leaderboard. Although I did not get a good enough rank but I still think my source code can serve as a template for beginners for any upcoming Kaggle Challenges.  </p>\n\n<p>My solution was an ensemble of 13 models ensembled using weighted averaging and stacking.  No external dataset was used. The training data was augmented using randomly sampled background noise and time shifting.</p>\n\n<p>The list of the Models used:</p>\n\n<ol>\n<li><p>A variant of Convolutional LSTM (<a href=\"https://arxiv.org/pdf/1610.00277.pdf\">https://arxiv.org/pdf/1610.00277.pdf</a>) </p></li>\n<li><p>LSTM-L (<a href=\"https://arxiv.org/pdf/1711.07128.pdf\">https://arxiv.org/pdf/1711.07128.pdf</a>) </p></li>\n<li><p>C-RNN (<a href=\"https://arxiv.org/pdf/1711.07128.pdf\">https://arxiv.org/pdf/1711.07128.pdf</a>) </p></li>\n<li><p>GRU-L (<a href=\"https://arxiv.org/pdf/1711.07128.pdf\">https://arxiv.org/pdf/1711.07128.pdf</a>) </p></li>\n<li><p>Resnet</p></li>\n</ol>\n\n<p>The features used were MFCC and Audio spectrogram. After initial training on the actual dataset each model was retrained on a combination of the train set and psuedo labelled test set (only predictions with 95%+ confidence were used). </p>\n\n<p>In my case weighted averaging worked better than stacking using linear regression on private leaderboard.  </p>\n\n<p>The entire project was very well structured, modular to make the training and analysis easier. All models were implemented using Tensorflow 1.4. The entire source is available on GITHub.  </p>\n\n<p><a href=\"https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\">https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution</a></p>",
  "messages": [
    {
      "id": 270073,
      "postDate": "2018-01-17T18:07:08.057Z",
      "content": "<p><a href=\"https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\">https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution</a></p>\n\n<p>Congratulations to all the winners! This was a wonderful experience, also thank you Google for the free credits without which I cannot imagine taking part in the competition. Also thank you Heng for your posts throughout the competition, your tips helped to jump many places in the leaderboard. Although I did not get a good enough rank but I still think my source code can serve as a template for beginners for any upcoming Kaggle Challenges.  </p>\n\n<p>My solution was an ensemble of 13 models ensembled using weighted averaging and stacking.  No external dataset was used. The training data was augmented using randomly sampled background noise and time shifting.</p>\n\n<p>The list of the Models used:</p>\n\n<ol>\n<li><p>A variant of Convolutional LSTM (<a href=\"https://arxiv.org/pdf/1610.00277.pdf\">https://arxiv.org/pdf/1610.00277.pdf</a>) </p></li>\n<li><p>LSTM-L (<a href=\"https://arxiv.org/pdf/1711.07128.pdf\">https://arxiv.org/pdf/1711.07128.pdf</a>) </p></li>\n<li><p>C-RNN (<a href=\"https://arxiv.org/pdf/1711.07128.pdf\">https://arxiv.org/pdf/1711.07128.pdf</a>) </p></li>\n<li><p>GRU-L (<a href=\"https://arxiv.org/pdf/1711.07128.pdf\">https://arxiv.org/pdf/1711.07128.pdf</a>) </p></li>\n<li><p>Resnet</p></li>\n</ol>\n\n<p>The features used were MFCC and Audio spectrogram. After initial training on the actual dataset each model was retrained on a combination of the train set and psuedo labelled test set (only predictions with 95%+ confidence were used). </p>\n\n<p>In my case weighted averaging worked better than stacking using linear regression on private leaderboard.  </p>\n\n<p>The entire project was very well structured, modular to make the training and analysis easier. All models were implemented using Tensorflow 1.4. The entire source is available on GITHub.  </p>\n\n<p><a href=\"https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\">https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution</a></p>",
      "rawMarkdown": "https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\n\nCongratulations to all the winners! This was a wonderful experience, also thank you Google for the free credits without which I cannot imagine taking part in the competition. Also thank you Heng for your posts throughout the competition, your tips helped to jump many places in the leaderboard. Although I did not get a good enough rank but I still think my source code can serve as a template for beginners for any upcoming Kaggle Challenges.  \n\nMy solution was an ensemble of 13 models ensembled using weighted averaging and stacking.  No external dataset was used. The training data was augmented using randomly sampled background noise and time shifting.\n\nThe list of the Models used:\n\n 1. A variant of Convolutional LSTM (https://arxiv.org/pdf/1610.00277.pdf) \n\n 2. LSTM-L (https://arxiv.org/pdf/1711.07128.pdf) \n\n 3. C-RNN (https://arxiv.org/pdf/1711.07128.pdf) \n\n 4. GRU-L (https://arxiv.org/pdf/1711.07128.pdf) \n\n 5. Resnet\n\nThe features used were MFCC and Audio spectrogram. After initial training on the actual dataset each model was retrained on a combination of the train set and psuedo labelled test set (only predictions with 95%+ confidence were used). \n\nIn my case weighted averaging worked better than stacking using linear regression on private leaderboard.  \n\nThe entire project was very well structured, modular to make the training and analysis easier. All models were implemented using Tensorflow 1.4. The entire source is available on GITHub.  \n\nhttps://github.com/subho406/TF-Speech-Recognition-Challenge-Solution",
      "votes": 12
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "270073": "https://github.com/subho406/TF-Speech-Recognition-Challenge-Solution\n\nCongratulations to all the winners! This was a wonderful experience, also thank you Google for the free credits without which I cannot imagine taking part in the competition. Also thank you Heng for your posts throughout the competition, your tips helped to jump many places in the leaderboard. Although I did not get a good enough rank but I still think my source code can serve as a template for beginners for any upcoming Kaggle Challenges.  \n\nMy solution was an ensemble of 13 models ensembled using weighted averaging and stacking.  No external dataset was used. The training data was augmented using randomly sampled background noise and time shifting.\n\nThe list of the Models used:\n\n 1. A variant of Convolutional LSTM (https://arxiv.org/pdf/1610.00277.pdf) \n\n 2. LSTM-L (https://arxiv.org/pdf/1711.07128.pdf) \n\n 3. C-RNN (https://arxiv.org/pdf/1711.07128.pdf) \n\n 4. GRU-L (https://arxiv.org/pdf/1711.07128.pdf) \n\n 5. Resnet\n\nThe features used were MFCC and Audio spectrogram. After initial training on the actual dataset each model was retrained on a combination of the train set and psuedo labelled test set (only predictions with 95%+ confidence were used). \n\nIn my case weighted averaging worked better than stacking using linear regression on private leaderboard.  \n\nThe entire project was very well structured, modular to make the training and analysis easier. All models were implemented using Tensorflow 1.4. The entire source is available on GITHub.  \n\nhttps://github.com/subho406/TF-Speech-Recognition-Challenge-Solution"
  }
}