{
  "id": 45037,
  "title": "Hello Edge: Keyword Spotting on Microcontrollers",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45037",
  "author_name": "Pete Warden",
  "post_date": "2017-12-05T17:07:31.544000",
  "votes": 21,
  "comment_count": 26,
  "views": 0,
  "content": "<p>I thought this recent research paper from ARM might be interesting to contestants, especially those looking at running on the Pi:\n<a href=\"https://arxiv.org/abs/1711.07128\">https://arxiv.org/abs/1711.07128</a></p>\n\n<p>It has accuracy, size, and speed results for a lot of different approaches to running on-device speech models, using the same dataset as this contest.</p>",
  "messages": [
    {
      "id": 253801,
      "postDate": "2017-12-05T17:07:31.543Z",
      "content": "<p>I thought this recent research paper from ARM might be interesting to contestants, especially those looking at running on the Pi:\n<a href=\"https://arxiv.org/abs/1711.07128\">https://arxiv.org/abs/1711.07128</a></p>\n\n<p>It has accuracy, size, and speed results for a lot of different approaches to running on-device speech models, using the same dataset as this contest.</p>",
      "rawMarkdown": "I thought this recent research paper from ARM might be interesting to contestants, especially those looking at running on the Pi:\nhttps://arxiv.org/abs/1711.07128\n\nIt has accuracy, size, and speed results for a lot of different approaches to running on-device speech models, using the same dataset as this contest.",
      "votes": 21
    },
    {
      "id": 254424,
      "postDate": "2017-12-06T23:50:59.300Z",
      "content": "<p>@Pete, thanks for posting our work here.</p>\n\n<p>@Jeff, all models in Section 4.1 use 40 MFCC features which is the same as used in DNN[5] and CNN[6] papers from Google Research. As discussed in Section 4.3, the number of MFCC features is also used as a hyperparameter for the models. Based on our experiments, we find 10 MFCC features are enough, which is what we used for the architecture exploration and described in the Appendix.</p>\n\n<p>@Human Analog, we will release all the Tensorflow models and scripts we have used for this paper on github so you can try them yourself 😊</p>\n\n<p>We have put the scripts and models here: <a href=\"https://github.com/ARM-software/ML-KWS-for-MCU\">https://github.com/ARM-software/ML-KWS-for-MCU</a></p>",
      "rawMarkdown": "@Pete, thanks for posting our work here.\n\n@Jeff, all models in Section 4.1 use 40 MFCC features which is the same as used in DNN[5] and CNN[6] papers from Google Research. As discussed in Section 4.3, the number of MFCC features is also used as a hyperparameter for the models. Based on our experiments, we find 10 MFCC features are enough, which is what we used for the architecture exploration and described in the Appendix.\n\n@Human Analog, we will release all the Tensorflow models and scripts we have used for this paper on github so you can try them yourself 😊\n\nWe have put the scripts and models here: https://github.com/ARM-software/ML-KWS-for-MCU",
      "votes": 8,
      "replies": [
        {
          "id": 254644,
          "postDate": "2017-12-07T10:07:56.337Z",
          "content": "<p>Awesome. :-) I ask because several of us have found that a high validation score doesn't necessarily give a high leaderboard score as well. Which would mean either the validation set isn't a good benchmark for real-world performance, of the Kaggle LB test set isn't (or we're doing something else wrong, of course).</p>",
          "rawMarkdown": "Awesome. :-) I ask because several of us have found that a high validation score doesn't necessarily give a high leaderboard score as well. Which would mean either the validation set isn't a good benchmark for real-world performance, of the Kaggle LB test set isn't (or we're doing something else wrong, of course)."
        },
        {
          "id": 256033,
          "postDate": "2017-12-11T00:58:55.577Z",
          "content": "<p>Thanks for the info, Liangzhen. I reproduced your experiment on the largest DSCNN model, but it resulted in 94.7% on the test set instead of 95.4%. On the Kaggle leaderboard, it resulted in a score of 0.55. If anyone else here tried reproducing the results from the paper, what were the results you observed?</p>",
          "rawMarkdown": "Thanks for the info, Liangzhen. I reproduced your experiment on the largest DSCNN model, but it resulted in 94.7% on the test set instead of 95.4%. On the Kaggle leaderboard, it resulted in a score of 0.55. If anyone else here tried reproducing the results from the paper, what were the results you observed?",
          "votes": 1
        },
        {
          "id": 256372,
          "postDate": "2017-12-11T20:10:05.357Z",
          "content": "<p><a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250</a>\nHave you read the following discussion about necessary change needed for submitting on Kaggle? LB~0.55 seems to be unreasonable. </p>",
          "rawMarkdown": "https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250\nHave you read the following discussion about necessary change needed for submitting on Kaggle? LB~0.55 seems to be unreasonable. "
        },
        {
          "id": 256374,
          "postDate": "2017-12-11T20:15:39.260Z",
          "content": "<p>Hi, Ben. Yes, I've read that discussion. With a different model architecture, I was able to get .84 on the leaderboard.</p>",
          "rawMarkdown": "Hi, Ben. Yes, I've read that discussion. With a different model architecture, I was able to get .84 on the leaderboard."
        },
        {
          "id": 256377,
          "postDate": "2017-12-11T20:42:13.963Z",
          "content": "<p>that's interesting....didn't expect the discrepancy can be so large. Can you try the smallest DSCNN model to see how it performs? I will train some models to give a shot as well</p>",
          "rawMarkdown": "that's interesting....didn't expect the discrepancy can be so large. Can you try the smallest DSCNN model to see how it performs? I will train some models to give a shot as well"
        },
        {
          "id": 256435,
          "postDate": "2017-12-11T23:27:43.787Z",
          "content": "<p>94.7% vs. 95.4% seems reasonable to me considering the different initializations and other factors (Adam vs. SGD, learning rate, step etc.). Also, the largest DS-CNN model (with ~400k weights) resulted in LB score of 0.83. We will be open-sourcing the scripts and models shortly on github.</p>",
          "rawMarkdown": "94.7% vs. 95.4% seems reasonable to me considering the different initializations and other factors (Adam vs. SGD, learning rate, step etc.). Also, the largest DS-CNN model (with ~400k weights) resulted in LB score of 0.83. We will be open-sourcing the scripts and models shortly on github.\n",
          "votes": 5
        },
        {
          "id": 256439,
          "postDate": "2017-12-11T23:38:26.957Z",
          "content": "<p>Thanks, Liangzhen. What were the training parameter values that you used to reach .83? I’m guessing you changed window size, stride, MFCC bins, etc.</p>",
          "rawMarkdown": "Thanks, Liangzhen. What were the training parameter values that you used to reach .83? I’m guessing you changed window size, stride, MFCC bins, etc."
        },
        {
          "id": 257233,
          "postDate": "2017-12-13T18:50:30.033Z",
          "content": "<p>We have put the scripts and models here: <a href=\"https://github.com/ARM-software/ML-KWS-for-MCU\">https://github.com/ARM-software/ML-KWS-for-MCU</a></p>\n\n<p>We have also included the exact training commands (with all hyperparameters) to generate the models.</p>",
          "rawMarkdown": "We have put the scripts and models here: https://github.com/ARM-software/ML-KWS-for-MCU\n\nWe have also included the exact training commands (with all hyperparameters) to generate the models.",
          "votes": 4
        },
        {
          "id": 257386,
          "postDate": "2017-12-14T06:15:30.843Z",
          "content": "<p>Thank you for making your code open sourced. I got 0.74 in public score using the pre-trained DNN_S model. This pre-trained model is worse than tutorial's default conv model that gave me 0.78. </p>",
          "rawMarkdown": "Thank you for making your code open sourced. I got 0.74 in public score using the pre-trained DNN_S model. This pre-trained model is worse than tutorial's default conv model that gave me 0.78. "
        },
        {
          "id": 257677,
          "postDate": "2017-12-14T19:10:56.700Z",
          "content": "<p>The DNN_S model gives 84.6% on the validation set compared to ~89% of the default conv model. So your observation in LB makes sense here, and DNN_S is a much (~10X) smaller model.</p>",
          "rawMarkdown": "The DNN_S model gives 84.6% on the validation set compared to ~89% of the default conv model. So your observation in LB makes sense here, and DNN_S is a much (~10X) smaller model."
        },
        {
          "id": 257947,
          "postDate": "2017-12-15T06:00:49.353Z",
          "content": "<p>I am sorry. I confirmed that your DS_CNN_L model gives 0.83 in LB.</p>",
          "rawMarkdown": "I am sorry. I confirmed that your DS_CNN_L model gives 0.83 in LB.",
          "votes": 2
        },
        {
          "id": 260525,
          "postDate": "2017-12-20T11:52:34.080Z",
          "content": "<p>@Liangzhen Lai Hi!, is this project only published the model? then what about the data pre-processing before MFCC feature? like sample rate and bit accuracy? and all the workflow before generate MFCC feature? seems that all the micro-controllers have only 12bit accuracy ADC for MIC. will this work?</p>",
          "rawMarkdown": "@Liangzhen Lai Hi!, is this project only published the model? then what about the data pre-processing before MFCC feature? like sample rate and bit accuracy? and all the workflow before generate MFCC feature? seems that all the micro-controllers have only 12bit accuracy ADC for MIC. will this work?"
        },
        {
          "id": 260779,
          "postDate": "2017-12-20T22:07:53.237Z",
          "content": "<p>We are planning to release the code as well, check this: <a href=\"https://github.com/ARM-software/ML-KWS-for-MCU/issues/2\">https://github.com/ARM-software/ML-KWS-for-MCU/issues/2</a></p>\n\n<p>You are right that most microcontrollers have 12-bit ADC. Some embedded development boards have dedicated audio codec chip to get higher resolution. For example, the demo we show in the paper uses STM32F746G-DISCO and is able to get 16-bit/16,000 sample rate, which is the same as this speech command dataset. We don't have any results for 12-bit, but I guess it shouldn't be too hard to just quantize the 16-bit .wav files into 12-bit and test it out.</p>",
          "rawMarkdown": "We are planning to release the code as well, check this: https://github.com/ARM-software/ML-KWS-for-MCU/issues/2\n\nYou are right that most microcontrollers have 12-bit ADC. Some embedded development boards have dedicated audio codec chip to get higher resolution. For example, the demo we show in the paper uses STM32F746G-DISCO and is able to get 16-bit/16,000 sample rate, which is the same as this speech command dataset. We don't have any results for 12-bit, but I guess it shouldn't be too hard to just quantize the 16-bit .wav files into 12-bit and test it out.",
          "votes": 1
        },
        {
          "id": 260849,
          "postDate": "2017-12-21T03:30:31.767Z",
          "content": "<p>so when will you release the code of the MFCC feature extracting for STM32-DISCO ?  </p>",
          "rawMarkdown": "so when will you release the code of the MFCC feature extracting for STM32-DISCO ?  "
        },
        {
          "id": 261185,
          "postDate": "2017-12-22T00:23:36.287Z",
          "content": "<p>We are targeting late Jan or early Feb with an end-to-end example. Meanwhile, you can also try to follow tensorflow mfcc implementation (<a href=\"https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/kernels/mfcc.cc\">https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/kernels/mfcc.cc</a>) with some functions (e.g. fft and cosine) from CMSIS-DSP.</p>",
          "rawMarkdown": "We are targeting late Jan or early Feb with an end-to-end example. Meanwhile, you can also try to follow tensorflow mfcc implementation (https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/kernels/mfcc.cc) with some functions (e.g. fft and cosine) from CMSIS-DSP."
        },
        {
          "id": 262350,
          "postDate": "2017-12-26T09:49:27.027Z",
          "content": "<p>It seems imposiable to do fft on a cheap MCU(below 100MHz) in realtime,  so this guy made a uSpeech (<a href=\"https://github.com/arjo129/uSpeech\">https://github.com/arjo129/uSpeech</a>) to take place fft&amp;mfcc. but it cant use the model you provided. waitting for you end-to-end solution...</p>",
          "rawMarkdown": "It seems imposiable to do fft on a cheap MCU(below 100MHz) in realtime,  so this guy made a uSpeech (https://github.com/arjo129/uSpeech) to take place fft&amp;mfcc. but it cant use the model you provided. waitting for you end-to-end solution..."
        },
        {
          "id": 264361,
          "postDate": "2018-01-02T23:01:00.177Z",
          "content": "<p>@Liangzhen, how to apply pre-trained model to LB test set?</p>",
          "rawMarkdown": "@Liangzhen, how to apply pre-trained model to LB test set?"
        },
        {
          "id": 264804,
          "postDate": "2018-01-03T22:51:22.620Z",
          "content": "<p>You can refer to the end of README.md. Use label_wav.py to generate the prediction for each file in the LB test set.</p>",
          "rawMarkdown": "You can refer to the end of README.md. Use label_wav.py to generate the prediction for each file in the LB test set."
        }
      ]
    },
    {
      "id": 254347,
      "postDate": "2017-12-06T19:16:59.110Z",
      "content": "<p>Nice paper! It's pretty amazing they're running on MCUs down to Cortex M0s. The MobileNet implementation with separable convolutions might be worth a look, I found a keras implementation on github. I used to work on semiconductor design and we used M0s and M4s. M4s have a Harvard architecture designed for FFTs IIRC, M0s only have a single AHB interface for both instruction and data. I never imagined they'd be used for dense matrix multiplies!</p>\n\n<p>Bill Dally (NVIDA / Stanford) has some cool papers about stochastic rounding of weights, network pruning, and storing weights in a LUT to avoid moving so much data around. Those would all help in fettling a larger model for an embedded target. Here's a great 2-hour talk from NIPS 2015: <a href=\"https://youtu.be/J-GOkwiwg4c\">https://youtu.be/J-GOkwiwg4c</a> </p>\n\n<p>I just need another couple of months to explore these techniques for the Raspberry Pi :-) !</p>",
      "rawMarkdown": "Nice paper! It's pretty amazing they're running on MCUs down to Cortex M0s. The MobileNet implementation with separable convolutions might be worth a look, I found a keras implementation on github. I used to work on semiconductor design and we used M0s and M4s. M4s have a Harvard architecture designed for FFTs IIRC, M0s only have a single AHB interface for both instruction and data. I never imagined they'd be used for dense matrix multiplies!\n\nBill Dally (NVIDA / Stanford) has some cool papers about stochastic rounding of weights, network pruning, and storing weights in a LUT to avoid moving so much data around. Those would all help in fettling a larger model for an embedded target. Here's a great 2-hour talk from NIPS 2015: https://youtu.be/J-GOkwiwg4c \n\nI just need another couple of months to explore these techniques for the Raspberry Pi :-) !",
      "votes": 1
    },
    {
      "id": 254156,
      "postDate": "2017-12-06T10:21:06.353Z",
      "content": "<p>Interesting! But how well do these models score on the Kaggle leaderboard? ;-)</p>",
      "rawMarkdown": "Interesting! But how well do these models score on the Kaggle leaderboard? ;-)",
      "votes": 2
    },
    {
      "id": 301438,
      "postDate": "2018-03-22T18:31:43.127Z",
      "content": "<p>Working on decision trees / random forests for use on microcontrollers here: <a href=\"https://github.com/jonnor/emtrees\">https://github.com/jonnor/emtrees</a>\nMaybe I'll test it on keyword spotting in the future.</p>",
      "rawMarkdown": "Working on decision trees / random forests for use on microcontrollers here: https://github.com/jonnor/emtrees\nMaybe I'll test it on keyword spotting in the future."
    },
    {
      "id": 260519,
      "postDate": "2017-12-20T11:41:20.427Z",
      "content": "<p>@Liangzhen Lai Hi!, is this project only published the model? then what about the data pre-processing before MFCC feature? like sample rate and bit accuracy? and all the workflow before generate MFCC feature? seems that all the micro-controllers have only 12bit accuracy ADC for MIC. will this work?</p>",
      "rawMarkdown": "@Liangzhen Lai Hi!, is this project only published the model? then what about the data pre-processing before MFCC feature? like sample rate and bit accuracy? and all the workflow before generate MFCC feature? seems that all the micro-controllers have only 12bit accuracy ADC for MIC. will this work?"
    },
    {
      "id": 254345,
      "postDate": "2017-12-06T19:02:02.513Z",
      "content": "<p>Thanks for the link to the paper, Pete. I noticed that the paragraph under appendix A states that 10 MFCC features were used, but section 4.1 \"Training Results\" states that 40 MFCC features were used.  Is that a typo in appendix A? Is 40 MFCC features typical in speech recognition systems?</p>",
      "rawMarkdown": "Thanks for the link to the paper, Pete. I noticed that the paragraph under appendix A states that 10 MFCC features were used, but section 4.1 \"Training Results\" states that 40 MFCC features were used.  Is that a typo in appendix A? Is 40 MFCC features typical in speech recognition systems?",
      "replies": [
        {
          "id": 254617,
          "postDate": "2017-12-07T09:00:39.310Z",
          "content": "<p>yes, 40 feature dim is more common in SR system.</p>",
          "rawMarkdown": "yes, 40 feature dim is more common in SR system."
        }
      ]
    },
    {
      "id": 459172,
      "postDate": "2019-01-21T10:34:53.970Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 254424,
      "author_name": "Liangzhen Lai",
      "author_url": "",
      "post_date": "2017-12-06T23:50:59.300000",
      "content": "<p>@Pete, thanks for posting our work here.</p>\n\n<p>@Jeff, all models in Section 4.1 use 40 MFCC features which is the same as used in DNN[5] and CNN[6] papers from Google Research. As discussed in Section 4.3, the number of MFCC features is also used as a hyperparameter for the models. Based on our experiments, we find 10 MFCC features are enough, which is what we used for the architecture exploration and described in the Appendix.</p>\n\n<p>@Human Analog, we will release all the Tensorflow models and scripts we have used for this paper on github so you can try them yourself 😊</p>\n\n<p>We have put the scripts and models here: <a href=\"https://github.com/ARM-software/ML-KWS-for-MCU\">https://github.com/ARM-software/ML-KWS-for-MCU</a></p>",
      "votes": 8,
      "replies": [
        {
          "id": 254644,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2017-12-07T10:07:56.337000",
          "content": "<p>Awesome. :-) I ask because several of us have found that a high validation score doesn't necessarily give a high leaderboard score as well. Which would mean either the validation set isn't a good benchmark for real-world performance, of the Kaggle LB test set isn't (or we're doing something else wrong, of course).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 256033,
          "author_name": "Jeff",
          "author_url": "",
          "post_date": "2017-12-11T00:58:55.577000",
          "content": "<p>Thanks for the info, Liangzhen. I reproduced your experiment on the largest DSCNN model, but it resulted in 94.7% on the test set instead of 95.4%. On the Kaggle leaderboard, it resulted in a score of 0.55. If anyone else here tried reproducing the results from the paper, what were the results you observed?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 256372,
          "author_name": "BenZhang",
          "author_url": "",
          "post_date": "2017-12-11T20:10:05.357000",
          "content": "<p><a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250</a>\nHave you read the following discussion about necessary change needed for submitting on Kaggle? LB~0.55 seems to be unreasonable. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 256374,
          "author_name": "Jeff",
          "author_url": "",
          "post_date": "2017-12-11T20:15:39.260000",
          "content": "<p>Hi, Ben. Yes, I've read that discussion. With a different model architecture, I was able to get .84 on the leaderboard.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 256377,
          "author_name": "BenZhang",
          "author_url": "",
          "post_date": "2017-12-11T20:42:13.963000",
          "content": "<p>that's interesting....didn't expect the discrepancy can be so large. Can you try the smallest DSCNN model to see how it performs? I will train some models to give a shot as well</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 256435,
          "author_name": "Liangzhen Lai",
          "author_url": "",
          "post_date": "2017-12-11T23:27:43.787000",
          "content": "<p>94.7% vs. 95.4% seems reasonable to me considering the different initializations and other factors (Adam vs. SGD, learning rate, step etc.). Also, the largest DS-CNN model (with ~400k weights) resulted in LB score of 0.83. We will be open-sourcing the scripts and models shortly on github.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 256439,
          "author_name": "Jeff",
          "author_url": "",
          "post_date": "2017-12-11T23:38:26.957000",
          "content": "<p>Thanks, Liangzhen. What were the training parameter values that you used to reach .83? I’m guessing you changed window size, stride, MFCC bins, etc.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 257233,
          "author_name": "Liangzhen Lai",
          "author_url": "",
          "post_date": "2017-12-13T18:50:30.033000",
          "content": "<p>We have put the scripts and models here: <a href=\"https://github.com/ARM-software/ML-KWS-for-MCU\">https://github.com/ARM-software/ML-KWS-for-MCU</a></p>\n\n<p>We have also included the exact training commands (with all hyperparameters) to generate the models.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 257386,
          "author_name": "Tony Y.",
          "author_url": "",
          "post_date": "2017-12-14T06:15:30.843000",
          "content": "<p>Thank you for making your code open sourced. I got 0.74 in public score using the pre-trained DNN_S model. This pre-trained model is worse than tutorial's default conv model that gave me 0.78. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 257677,
          "author_name": "Liangzhen Lai",
          "author_url": "",
          "post_date": "2017-12-14T19:10:56.700000",
          "content": "<p>The DNN_S model gives 84.6% on the validation set compared to ~89% of the default conv model. So your observation in LB makes sense here, and DNN_S is a much (~10X) smaller model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 257947,
          "author_name": "Tony Y.",
          "author_url": "",
          "post_date": "2017-12-15T06:00:49.353000",
          "content": "<p>I am sorry. I confirmed that your DS_CNN_L model gives 0.83 in LB.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 260525,
          "author_name": "zhouhua",
          "author_url": "",
          "post_date": "2017-12-20T11:52:34.080000",
          "content": "<p>@Liangzhen Lai Hi!, is this project only published the model? then what about the data pre-processing before MFCC feature? like sample rate and bit accuracy? and all the workflow before generate MFCC feature? seems that all the micro-controllers have only 12bit accuracy ADC for MIC. will this work?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 260779,
          "author_name": "Liangzhen Lai",
          "author_url": "",
          "post_date": "2017-12-20T22:07:53.237000",
          "content": "<p>We are planning to release the code as well, check this: <a href=\"https://github.com/ARM-software/ML-KWS-for-MCU/issues/2\">https://github.com/ARM-software/ML-KWS-for-MCU/issues/2</a></p>\n\n<p>You are right that most microcontrollers have 12-bit ADC. Some embedded development boards have dedicated audio codec chip to get higher resolution. For example, the demo we show in the paper uses STM32F746G-DISCO and is able to get 16-bit/16,000 sample rate, which is the same as this speech command dataset. We don't have any results for 12-bit, but I guess it shouldn't be too hard to just quantize the 16-bit .wav files into 12-bit and test it out.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 260849,
          "author_name": "zhouhua",
          "author_url": "",
          "post_date": "2017-12-21T03:30:31.767000",
          "content": "<p>so when will you release the code of the MFCC feature extracting for STM32-DISCO ?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 261185,
          "author_name": "Liangzhen Lai",
          "author_url": "",
          "post_date": "2017-12-22T00:23:36.287000",
          "content": "<p>We are targeting late Jan or early Feb with an end-to-end example. Meanwhile, you can also try to follow tensorflow mfcc implementation (<a href=\"https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/kernels/mfcc.cc\">https://github.com/tensorflow/tensorflow/blob/master/tensorflow/core/kernels/mfcc.cc</a>) with some functions (e.g. fft and cosine) from CMSIS-DSP.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 262350,
          "author_name": "zhouhua",
          "author_url": "",
          "post_date": "2017-12-26T09:49:27.027000",
          "content": "<p>It seems imposiable to do fft on a cheap MCU(below 100MHz) in realtime,  so this guy made a uSpeech (<a href=\"https://github.com/arjo129/uSpeech\">https://github.com/arjo129/uSpeech</a>) to take place fft&amp;mfcc. but it cant use the model you provided. waitting for you end-to-end solution...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 264361,
          "author_name": "Adam Wang",
          "author_url": "",
          "post_date": "2018-01-02T23:01:00.177000",
          "content": "<p>@Liangzhen, how to apply pre-trained model to LB test set?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 264804,
          "author_name": "Liangzhen Lai",
          "author_url": "",
          "post_date": "2018-01-03T22:51:22.620000",
          "content": "<p>You can refer to the end of README.md. Use label_wav.py to generate the prediction for each file in the LB test set.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 254347,
      "author_name": "tb303",
      "author_url": "",
      "post_date": "2017-12-06T19:16:59.110000",
      "content": "<p>Nice paper! It's pretty amazing they're running on MCUs down to Cortex M0s. The MobileNet implementation with separable convolutions might be worth a look, I found a keras implementation on github. I used to work on semiconductor design and we used M0s and M4s. M4s have a Harvard architecture designed for FFTs IIRC, M0s only have a single AHB interface for both instruction and data. I never imagined they'd be used for dense matrix multiplies!</p>\n\n<p>Bill Dally (NVIDA / Stanford) has some cool papers about stochastic rounding of weights, network pruning, and storing weights in a LUT to avoid moving so much data around. Those would all help in fettling a larger model for an embedded target. Here's a great 2-hour talk from NIPS 2015: <a href=\"https://youtu.be/J-GOkwiwg4c\">https://youtu.be/J-GOkwiwg4c</a> </p>\n\n<p>I just need another couple of months to explore these techniques for the Raspberry Pi :-) !</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 254156,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2017-12-06T10:21:06.353000",
      "content": "<p>Interesting! But how well do these models score on the Kaggle leaderboard? ;-)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 301438,
      "author_name": "Jon Nordby",
      "author_url": "",
      "post_date": "2018-03-22T18:31:43.127000",
      "content": "<p>Working on decision trees / random forests for use on microcontrollers here: <a href=\"https://github.com/jonnor/emtrees\">https://github.com/jonnor/emtrees</a>\nMaybe I'll test it on keyword spotting in the future.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 260519,
      "author_name": "zhouhua",
      "author_url": "",
      "post_date": "2017-12-20T11:41:20.427000",
      "content": "<p>@Liangzhen Lai Hi!, is this project only published the model? then what about the data pre-processing before MFCC feature? like sample rate and bit accuracy? and all the workflow before generate MFCC feature? seems that all the micro-controllers have only 12bit accuracy ADC for MIC. will this work?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 254345,
      "author_name": "Jeff",
      "author_url": "",
      "post_date": "2017-12-06T19:02:02.513000",
      "content": "<p>Thanks for the link to the paper, Pete. I noticed that the paragraph under appendix A states that 10 MFCC features were used, but section 4.1 \"Training Results\" states that 40 MFCC features were used.  Is that a typo in appendix A? Is 40 MFCC features typical in speech recognition systems?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 254617,
          "author_name": "Feiteng",
          "author_url": "",
          "post_date": "2017-12-07T09:00:39.310000",
          "content": "<p>yes, 40 feature dim is more common in SR system.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 459172,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-21T10:34:53.970000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "253801": "I thought this recent research paper from ARM might be interesting to contestants, especially those looking at running on the Pi:\nhttps://arxiv.org/abs/1711.07128\n\nIt has accuracy, size, and speed results for a lot of different approaches to running on-device speech models, using the same dataset as this contest.",
    "254424": "@Pete, thanks for posting our work here.\n\n@Jeff, all models in Section 4.1 use 40 MFCC features which is the same as used in DNN[5] and CNN[6] papers from Google Research. As discussed in Section 4.3, the number of MFCC features is also used as a hyperparameter for the models. Based on our experiments, we find 10 MFCC features are enough, which is what we used for the architecture exploration and described in the Appendix.\n\n@Human Analog, we will release all the Tensorflow models and scripts we have used for this paper on github so you can try them yourself 😊\n\nWe have put the scripts and models here: https://github.com/ARM-software/ML-KWS-for-MCU",
    "254347": "Nice paper! It's pretty amazing they're running on MCUs down to Cortex M0s. The MobileNet implementation with separable convolutions might be worth a look, I found a keras implementation on github. I used to work on semiconductor design and we used M0s and M4s. M4s have a Harvard architecture designed for FFTs IIRC, M0s only have a single AHB interface for both instruction and data. I never imagined they'd be used for dense matrix multiplies!\n\nBill Dally (NVIDA / Stanford) has some cool papers about stochastic rounding of weights, network pruning, and storing weights in a LUT to avoid moving so much data around. Those would all help in fettling a larger model for an embedded target. Here's a great 2-hour talk from NIPS 2015: https://youtu.be/J-GOkwiwg4c \n\nI just need another couple of months to explore these techniques for the Raspberry Pi :-) !",
    "254156": "Interesting! But how well do these models score on the Kaggle leaderboard? ;-)",
    "301438": "Working on decision trees / random forests for use on microcontrollers here: https://github.com/jonnor/emtrees\nMaybe I'll test it on keyword spotting in the future.",
    "260519": "@Liangzhen Lai Hi!, is this project only published the model? then what about the data pre-processing before MFCC feature? like sample rate and bit accuracy? and all the workflow before generate MFCC feature? seems that all the micro-controllers have only 12bit accuracy ADC for MIC. will this work?",
    "254345": "Thanks for the link to the paper, Pete. I noticed that the paragraph under appendix A states that 10 MFCC features were used, but section 4.1 \"Training Results\" states that 40 MFCC features were used.  Is that a typo in appendix A? Is 40 MFCC features typical in speech recognition systems?",
    "459172": ""
  }
}