{
  "id": 47715,
  "title": "2nd Place Solution",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47715",
  "author_name": "Thomas O'Malley",
  "post_date": "2018-01-17T22:45:03.471000",
  "votes": 85,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Hi, an honor to compete with you all! I really learned a lot from doing this competition and from everyone here, and congrats to Heng CherKeng, Ryan Sun, and See!</p>\n\n<p>I thought I'd share my methodology, with a focus on parts I haven't seen discussed here yet.</p>\n\n<p>I was able to get a 90.9% private LB on a single model, but wasn't able to fully benefit from ensembling (I'm not very experienced in that, but learned a lot in a short time from the forums here). My ensembling technique was to vary a few parameters from the main model and average the square roots of the models' probabilities.</p>\n\n<p>Model Architecture:</p>\n\n<p>I used 120 log-mel filterbanks for my best model. Given this, I thought it was important to create a model that treated time and frequency differently. Specifically, I didn't do any downsampling in the time domain until the very end. With time as the first dimension and frequency as the second, my model architecture was:</p>\n\n<p>1) Conv2d(64, [7,3] ) </p>\n\n<p>I thought of this as a \"denoising\" and basic feature extraction step</p>\n\n<p>2) MaxPool( [1,3] )</p>\n\n<p>Getting back down to the standard 40 frequency features</p>\n\n<p>3) Conv2d(128, [1,7] ) </p>\n\n<p>Look for local patterns across frequency bands</p>\n\n<p>4) MaxPool( [1,4] ) </p>\n\n<p>Allow for speaker variation, similar to what worked here: <a href=\"https://link.springer.com/content/pdf/10.1186%2Fs13636-015-0068-3.pdf\">https://link.springer.com/content/pdf/10.1186%2Fs13636-015-0068-3.pdf</a></p>\n\n<p>5) Conv2d(256 [1,10], padding=\"VALID\") </p>\n\n<p>This allows it to treat each remaining freq band very differently, and compress the frequency dimension entirely. I think of this as detecting phoneme-level features</p>\n\n<p>6) Conv2d(512,[7,1])</p>\n\n<p>I think of this as looking for connected components of a short keyword at different points in time</p>\n\n<p>7) GlobalMaxPool in time</p>\n\n<p>Collect all the components</p>\n\n<p>8) Dropout + Fully Connected 256</p>\n\n<p>Because why not, and seemed to work well</p>\n\n<p>Data Augmentation/Standardization:</p>\n\n<p>In addition to time stretching (which gave a boost of +1% LB), there were two techniques I applied that I haven't seen mentioned yet here that I think really helped.</p>\n\n<p>1) Standardize Peak (Windowed) Volume</p>\n\n<p>Basically, I took every clip, split it into 20 to 50 chunks, and then standardized the volume of the clips so that every clip had the same max chunk volume. Why this approach? Well, standardizing by average volume would be fine, but since some keywords were longer than others, very short . keywords would have a low average volume even though the part of the WAV containing the actual word could be quite loud.</p>\n\n<p>When I implemented this, my LB score went from 82% to 84%</p>\n\n<p>My theory is that by doing this, the convnet didn't have to deal with as many issues in terms of different scales for the same feature, since the volumes of the WAVs spanned orders or magnitude. Obviously using Log Mels helped with this too.</p>\n\n<p>2) Vocal Tract Length Perturbation</p>\n\n<p>I used VTLP as described in this paper <a href=\"https://pdfs.semanticscholar.org/3de0/616eb3cd4554fdf9fd65c9c82f2605a17413.pdf\">https://pdfs.semanticscholar.org/3de0/616eb3cd4554fdf9fd65c9c82f2605a17413.pdf</a></p>\n\n<p>This perturbation could be applied when creating the weight matrix to convert a spectrogram into log mels, so it was a very fast augmentation method</p>\n\n<p>This increased my LB score +1%, and I saw the greatest benefit using the same VTLP factor within a batch, along the line of reasoning described here:</p>\n\n<p><a href=\"https://arxiv.org/abs/1707.00722\">https://arxiv.org/abs/1707.00722</a></p>",
  "messages": [
    {
      "id": 270205,
      "postDate": "2018-01-17T22:45:03.473Z",
      "content": "<p>Hi, an honor to compete with you all! I really learned a lot from doing this competition and from everyone here, and congrats to Heng CherKeng, Ryan Sun, and See!</p>\n\n<p>I thought I'd share my methodology, with a focus on parts I haven't seen discussed here yet.</p>\n\n<p>I was able to get a 90.9% private LB on a single model, but wasn't able to fully benefit from ensembling (I'm not very experienced in that, but learned a lot in a short time from the forums here). My ensembling technique was to vary a few parameters from the main model and average the square roots of the models' probabilities.</p>\n\n<p>Model Architecture:</p>\n\n<p>I used 120 log-mel filterbanks for my best model. Given this, I thought it was important to create a model that treated time and frequency differently. Specifically, I didn't do any downsampling in the time domain until the very end. With time as the first dimension and frequency as the second, my model architecture was:</p>\n\n<p>1) Conv2d(64, [7,3] ) </p>\n\n<p>I thought of this as a \"denoising\" and basic feature extraction step</p>\n\n<p>2) MaxPool( [1,3] )</p>\n\n<p>Getting back down to the standard 40 frequency features</p>\n\n<p>3) Conv2d(128, [1,7] ) </p>\n\n<p>Look for local patterns across frequency bands</p>\n\n<p>4) MaxPool( [1,4] ) </p>\n\n<p>Allow for speaker variation, similar to what worked here: <a href=\"https://link.springer.com/content/pdf/10.1186%2Fs13636-015-0068-3.pdf\">https://link.springer.com/content/pdf/10.1186%2Fs13636-015-0068-3.pdf</a></p>\n\n<p>5) Conv2d(256 [1,10], padding=\"VALID\") </p>\n\n<p>This allows it to treat each remaining freq band very differently, and compress the frequency dimension entirely. I think of this as detecting phoneme-level features</p>\n\n<p>6) Conv2d(512,[7,1])</p>\n\n<p>I think of this as looking for connected components of a short keyword at different points in time</p>\n\n<p>7) GlobalMaxPool in time</p>\n\n<p>Collect all the components</p>\n\n<p>8) Dropout + Fully Connected 256</p>\n\n<p>Because why not, and seemed to work well</p>\n\n<p>Data Augmentation/Standardization:</p>\n\n<p>In addition to time stretching (which gave a boost of +1% LB), there were two techniques I applied that I haven't seen mentioned yet here that I think really helped.</p>\n\n<p>1) Standardize Peak (Windowed) Volume</p>\n\n<p>Basically, I took every clip, split it into 20 to 50 chunks, and then standardized the volume of the clips so that every clip had the same max chunk volume. Why this approach? Well, standardizing by average volume would be fine, but since some keywords were longer than others, very short . keywords would have a low average volume even though the part of the WAV containing the actual word could be quite loud.</p>\n\n<p>When I implemented this, my LB score went from 82% to 84%</p>\n\n<p>My theory is that by doing this, the convnet didn't have to deal with as many issues in terms of different scales for the same feature, since the volumes of the WAVs spanned orders or magnitude. Obviously using Log Mels helped with this too.</p>\n\n<p>2) Vocal Tract Length Perturbation</p>\n\n<p>I used VTLP as described in this paper <a href=\"https://pdfs.semanticscholar.org/3de0/616eb3cd4554fdf9fd65c9c82f2605a17413.pdf\">https://pdfs.semanticscholar.org/3de0/616eb3cd4554fdf9fd65c9c82f2605a17413.pdf</a></p>\n\n<p>This perturbation could be applied when creating the weight matrix to convert a spectrogram into log mels, so it was a very fast augmentation method</p>\n\n<p>This increased my LB score +1%, and I saw the greatest benefit using the same VTLP factor within a batch, along the line of reasoning described here:</p>\n\n<p><a href=\"https://arxiv.org/abs/1707.00722\">https://arxiv.org/abs/1707.00722</a></p>",
      "rawMarkdown": "Hi, an honor to compete with you all! I really learned a lot from doing this competition and from everyone here, and congrats to Heng CherKeng, Ryan Sun, and See!\n\nI thought I'd share my methodology, with a focus on parts I haven't seen discussed here yet.\n\nI was able to get a 90.9% private LB on a single model, but wasn't able to fully benefit from ensembling (I'm not very experienced in that, but learned a lot in a short time from the forums here). My ensembling technique was to vary a few parameters from the main model and average the square roots of the models' probabilities.\n\nModel Architecture:\n\nI used 120 log-mel filterbanks for my best model. Given this, I thought it was important to create a model that treated time and frequency differently. Specifically, I didn't do any downsampling in the time domain until the very end. With time as the first dimension and frequency as the second, my model architecture was:\n\n1) Conv2d(64, [7,3] ) \n\nI thought of this as a \"denoising\" and basic feature extraction step\n\n2) MaxPool( [1,3] )\n\nGetting back down to the standard 40 frequency features\n\n3) Conv2d(128, [1,7] ) \n\nLook for local patterns across frequency bands\n\n4) MaxPool( [1,4] ) \n\nAllow for speaker variation, similar to what worked here: https://link.springer.com/content/pdf/10.1186%2Fs13636-015-0068-3.pdf\n\n5) Conv2d(256 [1,10], padding=\"VALID\") \n\nThis allows it to treat each remaining freq band very differently, and compress the frequency dimension entirely. I think of this as detecting phoneme-level features\n\n6) Conv2d(512,[7,1])\n\nI think of this as looking for connected components of a short keyword at different points in time\n\n7) GlobalMaxPool in time\n\nCollect all the components\n\n8) Dropout + Fully Connected 256\n\nBecause why not, and seemed to work well\n\nData Augmentation/Standardization:\n\nIn addition to time stretching (which gave a boost of +1% LB), there were two techniques I applied that I haven't seen mentioned yet here that I think really helped.\n\n1) Standardize Peak (Windowed) Volume\n\nBasically, I took every clip, split it into 20 to 50 chunks, and then standardized the volume of the clips so that every clip had the same max chunk volume. Why this approach? Well, standardizing by average volume would be fine, but since some keywords were longer than others, very short . keywords would have a low average volume even though the part of the WAV containing the actual word could be quite loud.\n\nWhen I implemented this, my LB score went from 82% to 84%\n\nMy theory is that by doing this, the convnet didn't have to deal with as many issues in terms of different scales for the same feature, since the volumes of the WAVs spanned orders or magnitude. Obviously using Log Mels helped with this too.\n\n2) Vocal Tract Length Perturbation\n\nI used VTLP as described in this paper https://pdfs.semanticscholar.org/3de0/616eb3cd4554fdf9fd65c9c82f2605a17413.pdf\n\nThis perturbation could be applied when creating the weight matrix to convert a spectrogram into log mels, so it was a very fast augmentation method\n\nThis increased my LB score +1%, and I saw the greatest benefit using the same VTLP factor within a batch, along the line of reasoning described here:\n\nhttps://arxiv.org/abs/1707.00722\n\n \n\n\n\n\n",
      "votes": 84
    },
    {
      "id": 277741,
      "postDate": "2018-02-04T03:55:22.123Z",
      "content": "<p>Thanks for your sharing, and I was wondering your detail of \"Standardize Peak (Windowed) Volume\", could you please share some pseudo code of that? I have add the same process to my code as you said, but it could not work as well as your.</p>",
      "rawMarkdown": "Thanks for your sharing, and I was wondering your detail of \"Standardize Peak (Windowed) Volume\", could you please share some pseudo code of that? I have add the same process to my code as you said, but it could not work as well as your.",
      "votes": 1
    },
    {
      "id": 271189,
      "postDate": "2018-01-19T21:22:34.933Z",
      "content": "<p>Congratulations Thomas on this amazing solution ! Is there a particular reason why you chose to take the square root of the probabilities when ensembling ? Thanks !</p>",
      "rawMarkdown": "Congratulations Thomas on this amazing solution ! Is there a particular reason why you chose to take the square root of the probabilities when ensembling ? Thanks !",
      "votes": 1
    },
    {
      "id": 270523,
      "postDate": "2018-01-18T12:20:25.140Z",
      "content": "<p>Thoughtful and elegant, really like attaching meaningfulness to each layer.  Well done Thomas and congrats on a great competition!</p>",
      "rawMarkdown": "Thoughtful and elegant, really like attaching meaningfulness to each layer.  Well done Thomas and congrats on a great competition!",
      "votes": 2
    },
    {
      "id": 270324,
      "postDate": "2018-01-18T04:55:18.977Z",
      "content": "<p>Thanks for the write up. It is a very elegant and beautiful solution!\nDo you mind posting your submission CSV file along with the raw probablity values here?</p>\n\n<p>I can make a comparsion between your results and mine. I posted mine at:\n<a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47728\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47728</a></p>",
      "rawMarkdown": "Thanks for the write up. It is a very elegant and beautiful solution!\nDo you mind posting your submission CSV file along with the raw probablity values here?\n\nI can make a comparsion between your results and mine. I posted mine at:\nhttps://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47728",
      "votes": 2
    },
    {
      "id": 270253,
      "postDate": "2018-01-18T00:43:25.427Z",
      "content": "<p>Looks awesome Thomas O'Malley! \nHow did you tune your layers? Any tricks? :)</p>",
      "rawMarkdown": "Looks awesome Thomas O'Malley! \nHow did you tune your layers? Any tricks? :)"
    },
    {
      "id": 271707,
      "postDate": "2018-01-21T07:52:45.463Z",
      "content": "<p>Congratulations @Thomas for winning the 2nd place! I like the solution because it feel like it much more focus on the special characteristics of speech.  Can you share your implementation of the Vocal Tract Length Perturbation? Which framework did you use for building the model and training it?</p>",
      "rawMarkdown": "Congratulations @Thomas for winning the 2nd place! I like the solution because it feel like it much more focus on the special characteristics of speech.  Can you share your implementation of the Vocal Tract Length Perturbation? Which framework did you use for building the model and training it?"
    },
    {
      "id": 271623,
      "postDate": "2018-01-21T00:17:33.327Z",
      "content": "<p>Congrats</p>",
      "rawMarkdown": "Congrats"
    },
    {
      "id": 271446,
      "postDate": "2018-01-20T13:10:27.793Z",
      "content": "<p>Nice. So without any augmentation your score was 0.86? Could you elaborate on the time stretching? Do you mean you adjusted the speed of the file (and if yes what bounds did you use)?</p>\n\n<p>Really like your volume normalization idea btw.</p>",
      "rawMarkdown": "Nice. So without any augmentation your score was 0.86? Could you elaborate on the time stretching? Do you mean you adjusted the speed of the file (and if yes what bounds did you use)?\n\nReally like your volume normalization idea btw."
    },
    {
      "id": 270315,
      "postDate": "2018-01-18T04:05:19.803Z",
      "content": "<p>Thanks for the nice description and paper references, and congrats on the great finish.</p>",
      "rawMarkdown": "Thanks for the nice description and paper references, and congrats on the great finish."
    },
    {
      "id": 294232,
      "postDate": "2018-03-11T15:54:04.380Z",
      "rawMarkdown": "",
      "votes": 5,
      "isDeleted": true
    },
    {
      "id": 270307,
      "postDate": "2018-01-18T03:28:29.897Z",
      "content": "<p>Awsome work， Thanks！</p>",
      "rawMarkdown": "Awsome work， Thanks！"
    }
  ],
  "comments": [
    {
      "id": 277741,
      "author_name": "rtyTyu",
      "author_url": "",
      "post_date": "2018-02-04T03:55:22.123000",
      "content": "<p>Thanks for your sharing, and I was wondering your detail of \"Standardize Peak (Windowed) Volume\", could you please share some pseudo code of that? I have add the same process to my code as you said, but it could not work as well as your.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 271189,
      "author_name": "Kerem Turgutlu",
      "author_url": "",
      "post_date": "2018-01-19T21:22:34.933000",
      "content": "<p>Congratulations Thomas on this amazing solution ! Is there a particular reason why you chose to take the square root of the probabilities when ensembling ? Thanks !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 270523,
      "author_name": "Russ Wolfinger",
      "author_url": "",
      "post_date": "2018-01-18T12:20:25.140000",
      "content": "<p>Thoughtful and elegant, really like attaching meaningfulness to each layer.  Well done Thomas and congrats on a great competition!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 270324,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2018-01-18T04:55:18.977000",
      "content": "<p>Thanks for the write up. It is a very elegant and beautiful solution!\nDo you mind posting your submission CSV file along with the raw probablity values here?</p>\n\n<p>I can make a comparsion between your results and mine. I posted mine at:\n<a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47728\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47728</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 270253,
      "author_name": "Little Boat",
      "author_url": "",
      "post_date": "2018-01-18T00:43:25.427000",
      "content": "<p>Looks awesome Thomas O'Malley! \nHow did you tune your layers? Any tricks? :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 271707,
      "author_name": "Ori Tal",
      "author_url": "",
      "post_date": "2018-01-21T07:52:45.463000",
      "content": "<p>Congratulations @Thomas for winning the 2nd place! I like the solution because it feel like it much more focus on the special characteristics of speech.  Can you share your implementation of the Vocal Tract Length Perturbation? Which framework did you use for building the model and training it?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 271623,
      "author_name": "Vishesh Sharma",
      "author_url": "",
      "post_date": "2018-01-21T00:17:33.327000",
      "content": "<p>Congrats</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 271446,
      "author_name": "RuAB",
      "author_url": "",
      "post_date": "2018-01-20T13:10:27.793000",
      "content": "<p>Nice. So without any augmentation your score was 0.86? Could you elaborate on the time stretching? Do you mean you adjusted the speed of the file (and if yes what bounds did you use)?</p>\n\n<p>Really like your volume normalization idea btw.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 270315,
      "author_name": "Brian Scherer",
      "author_url": "",
      "post_date": "2018-01-18T04:05:19.803000",
      "content": "<p>Thanks for the nice description and paper references, and congrats on the great finish.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 294232,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-03-11T15:54:04.380000",
      "content": "",
      "votes": 5,
      "replies": []
    },
    {
      "id": 270307,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "2018-01-18T03:28:29.897000",
      "content": "<p>Awsome work， Thanks！</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "270205": "Hi, an honor to compete with you all! I really learned a lot from doing this competition and from everyone here, and congrats to Heng CherKeng, Ryan Sun, and See!\n\nI thought I'd share my methodology, with a focus on parts I haven't seen discussed here yet.\n\nI was able to get a 90.9% private LB on a single model, but wasn't able to fully benefit from ensembling (I'm not very experienced in that, but learned a lot in a short time from the forums here). My ensembling technique was to vary a few parameters from the main model and average the square roots of the models' probabilities.\n\nModel Architecture:\n\nI used 120 log-mel filterbanks for my best model. Given this, I thought it was important to create a model that treated time and frequency differently. Specifically, I didn't do any downsampling in the time domain until the very end. With time as the first dimension and frequency as the second, my model architecture was:\n\n1) Conv2d(64, [7,3] ) \n\nI thought of this as a \"denoising\" and basic feature extraction step\n\n2) MaxPool( [1,3] )\n\nGetting back down to the standard 40 frequency features\n\n3) Conv2d(128, [1,7] ) \n\nLook for local patterns across frequency bands\n\n4) MaxPool( [1,4] ) \n\nAllow for speaker variation, similar to what worked here: https://link.springer.com/content/pdf/10.1186%2Fs13636-015-0068-3.pdf\n\n5) Conv2d(256 [1,10], padding=\"VALID\") \n\nThis allows it to treat each remaining freq band very differently, and compress the frequency dimension entirely. I think of this as detecting phoneme-level features\n\n6) Conv2d(512,[7,1])\n\nI think of this as looking for connected components of a short keyword at different points in time\n\n7) GlobalMaxPool in time\n\nCollect all the components\n\n8) Dropout + Fully Connected 256\n\nBecause why not, and seemed to work well\n\nData Augmentation/Standardization:\n\nIn addition to time stretching (which gave a boost of +1% LB), there were two techniques I applied that I haven't seen mentioned yet here that I think really helped.\n\n1) Standardize Peak (Windowed) Volume\n\nBasically, I took every clip, split it into 20 to 50 chunks, and then standardized the volume of the clips so that every clip had the same max chunk volume. Why this approach? Well, standardizing by average volume would be fine, but since some keywords were longer than others, very short . keywords would have a low average volume even though the part of the WAV containing the actual word could be quite loud.\n\nWhen I implemented this, my LB score went from 82% to 84%\n\nMy theory is that by doing this, the convnet didn't have to deal with as many issues in terms of different scales for the same feature, since the volumes of the WAVs spanned orders or magnitude. Obviously using Log Mels helped with this too.\n\n2) Vocal Tract Length Perturbation\n\nI used VTLP as described in this paper https://pdfs.semanticscholar.org/3de0/616eb3cd4554fdf9fd65c9c82f2605a17413.pdf\n\nThis perturbation could be applied when creating the weight matrix to convert a spectrogram into log mels, so it was a very fast augmentation method\n\nThis increased my LB score +1%, and I saw the greatest benefit using the same VTLP factor within a batch, along the line of reasoning described here:\n\nhttps://arxiv.org/abs/1707.00722\n\n \n\n\n\n\n",
    "277741": "Thanks for your sharing, and I was wondering your detail of \"Standardize Peak (Windowed) Volume\", could you please share some pseudo code of that? I have add the same process to my code as you said, but it could not work as well as your.",
    "271189": "Congratulations Thomas on this amazing solution ! Is there a particular reason why you chose to take the square root of the probabilities when ensembling ? Thanks !",
    "270523": "Thoughtful and elegant, really like attaching meaningfulness to each layer.  Well done Thomas and congrats on a great competition!",
    "270324": "Thanks for the write up. It is a very elegant and beautiful solution!\nDo you mind posting your submission CSV file along with the raw probablity values here?\n\nI can make a comparsion between your results and mine. I posted mine at:\nhttps://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47728",
    "270253": "Looks awesome Thomas O'Malley! \nHow did you tune your layers? Any tricks? :)",
    "271707": "Congratulations @Thomas for winning the 2nd place! I like the solution because it feel like it much more focus on the special characteristics of speech.  Can you share your implementation of the Vocal Tract Length Perturbation? Which framework did you use for building the model and training it?",
    "271623": "Congrats",
    "271446": "Nice. So without any augmentation your score was 0.86? Could you elaborate on the time stretching? Do you mean you adjusted the speed of the file (and if yes what bounds did you use)?\n\nReally like your volume normalization idea btw.",
    "270315": "Thanks for the nice description and paper references, and congrats on the great finish.",
    "294232": "",
    "270307": "Awsome work， Thanks！"
  }
}