{
  "id": 46945,
  "title": "My Tricks and Solution",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46945",
  "author_name": "hengck23",
  "post_date": "2018-01-05T15:28:59.211000",
  "votes": 92,
  "comment_count": 36,
  "views": 0,
  "content": "<p>(this discussion will be constantly updated)</p>\n\n<p>[dataset]</p>\n\n<ul>\n<li>At first, this challenge looks like a 12-class classification problem. But actually, it is \"not\". There are 10 class and 2 background class \"silence\" and \"unknown\". Note that there are no \"silence\" train samples and \"unknown\" train samples is not exactly the same as that of LB set.</li>\n<li>If you can generate correct \"silence\" and \"unknown\" train data, you should get good results (e.g. single model in the range of 0.87 to 0.88)</li>\n<li>How to probe the LB \"silence\" and \"unknown\" distribution? Train a classifier and use it to \"pseudo-label\" the LB test data. Sample the test set using the \"pseudo-label\" (choose different predicted probability) and listen to those LB test samples. It is not difficult to identify their characteristics (e.g. how does LB silence sounds like? what are the missing LB unknown from the train, ...)</li>\n<li>the next step is to generate these background train samples same the LB. For me I simply use  \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.</li>\n<li>there will be some label noise, but deep learning can accommodate 10 to 20% label noise  if your data is large enough</li>\n<li>if your sampling is correct, simple \"cnn_trad_pool2_net + mfcc\"  from the paper gives LB of around 0.82 and 0.83 in my experiments</li>\n</ul>\n\n<p>[input]</p>\n\n<ul>\n<li><p>you can choose either to to use: raw wave form (1d), logmelspectrum (2d), mfcc (2d), or mixture of these.</p></li>\n<li><p>for my case, logmelspectrum (2d) is better than mfcc. I am trying raw wave form (1d), but haven't got good results yet.</p></li>\n<li><p>simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). Convergence is very fast and you can get validation accuracy in the range 0.95 to 0.97, validation loss of 0.18 to 0.13. (the standard validation split from the tensorflow hash method). It doesn't matter of you use all validation samples or sub-sample equal ratio of train unknown samples. The validation accuracy and loss should be \"about the same\" for both of the two cases if your generalization is good enough.</p></li>\n</ul>\n\n<p>[ensemble]</p>\n\n<ul>\n<li>just keep making different models and ensemble them. In my experiments, i can get 0.87 by ensembling about 10 models of in the range of 0.85 to 0.86. Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting </li>\n</ul>\n\n<p>[other tricks]</p>\n\n<ul>\n<li><p>try wave + reflected wave (train separately and test separately).</p>\n\n<ul><li>because of padding, stride, etc ... the results are \"slightly\" different. Note that you can flip left-right, up-down.</li>\n<li>there are also other test-time augmentation you can use like cropping, stretch amplitude, etc</li></ul></li>\n<li><p>will  unsupervised learning using pseudo-label overfits?</p>\n\n<ul><li><p>Create two representation views for a train sample (e.g. view1=orginal and view2=reflected , or view1=wave(1d) and view2=mfcc (2d)</p></li>\n<li><p>use view1 to pedsuo-label and view-2 to train new model. </p></li></ul></li>\n</ul>\n\n<p>[how about raspberry pi special prize?]</p>\n\n<ul>\n<li>it may be easier to get  good accuracy with an ensemble of teacher models first, then transfer this knowledge to a low complexity student (i.e. Professor Hinton's \"knowledge distillation\" paper).</li>\n<li>In my experiment, my best single model  made from knowledge distillation has LB 0.88 (same as the ensemble)</li>\n</ul>",
  "messages": [
    {
      "id": 265463,
      "postDate": "2018-01-05T15:28:59.210Z",
      "content": "<p>(this discussion will be constantly updated)</p>\n\n<p>[dataset]</p>\n\n<ul>\n<li>At first, this challenge looks like a 12-class classification problem. But actually, it is \"not\". There are 10 class and 2 background class \"silence\" and \"unknown\". Note that there are no \"silence\" train samples and \"unknown\" train samples is not exactly the same as that of LB set.</li>\n<li>If you can generate correct \"silence\" and \"unknown\" train data, you should get good results (e.g. single model in the range of 0.87 to 0.88)</li>\n<li>How to probe the LB \"silence\" and \"unknown\" distribution? Train a classifier and use it to \"pseudo-label\" the LB test data. Sample the test set using the \"pseudo-label\" (choose different predicted probability) and listen to those LB test samples. It is not difficult to identify their characteristics (e.g. how does LB silence sounds like? what are the missing LB unknown from the train, ...)</li>\n<li>the next step is to generate these background train samples same the LB. For me I simply use  \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.</li>\n<li>there will be some label noise, but deep learning can accommodate 10 to 20% label noise  if your data is large enough</li>\n<li>if your sampling is correct, simple \"cnn_trad_pool2_net + mfcc\"  from the paper gives LB of around 0.82 and 0.83 in my experiments</li>\n</ul>\n\n<p>[input]</p>\n\n<ul>\n<li><p>you can choose either to to use: raw wave form (1d), logmelspectrum (2d), mfcc (2d), or mixture of these.</p></li>\n<li><p>for my case, logmelspectrum (2d) is better than mfcc. I am trying raw wave form (1d), but haven't got good results yet.</p></li>\n<li><p>simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). Convergence is very fast and you can get validation accuracy in the range 0.95 to 0.97, validation loss of 0.18 to 0.13. (the standard validation split from the tensorflow hash method). It doesn't matter of you use all validation samples or sub-sample equal ratio of train unknown samples. The validation accuracy and loss should be \"about the same\" for both of the two cases if your generalization is good enough.</p></li>\n</ul>\n\n<p>[ensemble]</p>\n\n<ul>\n<li>just keep making different models and ensemble them. In my experiments, i can get 0.87 by ensembling about 10 models of in the range of 0.85 to 0.86. Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting </li>\n</ul>\n\n<p>[other tricks]</p>\n\n<ul>\n<li><p>try wave + reflected wave (train separately and test separately).</p>\n\n<ul><li>because of padding, stride, etc ... the results are \"slightly\" different. Note that you can flip left-right, up-down.</li>\n<li>there are also other test-time augmentation you can use like cropping, stretch amplitude, etc</li></ul></li>\n<li><p>will  unsupervised learning using pseudo-label overfits?</p>\n\n<ul><li><p>Create two representation views for a train sample (e.g. view1=orginal and view2=reflected , or view1=wave(1d) and view2=mfcc (2d)</p></li>\n<li><p>use view1 to pedsuo-label and view-2 to train new model. </p></li></ul></li>\n</ul>\n\n<p>[how about raspberry pi special prize?]</p>\n\n<ul>\n<li>it may be easier to get  good accuracy with an ensemble of teacher models first, then transfer this knowledge to a low complexity student (i.e. Professor Hinton's \"knowledge distillation\" paper).</li>\n<li>In my experiment, my best single model  made from knowledge distillation has LB 0.88 (same as the ensemble)</li>\n</ul>",
      "rawMarkdown": "(this discussion will be constantly updated)\n\n[dataset]\n\n- At first, this challenge looks like a 12-class classification problem. But actually, it is \"not\". There are 10 class and 2 background class \"silence\" and \"unknown\". Note that there are no \"silence\" train samples and \"unknown\" train samples is not exactly the same as that of LB set.\n- If you can generate correct \"silence\" and \"unknown\" train data, you should get good results (e.g. single model in the range of 0.87 to 0.88)\n- How to probe the LB \"silence\" and \"unknown\" distribution? Train a classifier and use it to \"pseudo-label\" the LB test data. Sample the test set using the \"pseudo-label\" (choose different predicted probability) and listen to those LB test samples. It is not difficult to identify their characteristics (e.g. how does LB silence sounds like? what are the missing LB unknown from the train, ...)\n- the next step is to generate these background train samples same the LB. For me I simply use  \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.\n- there will be some label noise, but deep learning can accommodate 10 to 20% label noise  if your data is large enough\n-  if your sampling is correct, simple \"cnn_trad_pool2_net + mfcc\"  from the paper gives LB of around 0.82 and 0.83 in my experiments\n\n\n[input]\n\n - you can choose either to to use: raw wave form (1d), logmelspectrum (2d), mfcc (2d), or mixture of these.\n\n - for my case, logmelspectrum (2d) is better than mfcc. I am trying raw wave form (1d), but haven't got good results yet.\n\n - simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). Convergence is very fast and you can get validation accuracy in the range 0.95 to 0.97, validation loss of 0.18 to 0.13. (the standard validation split from the tensorflow hash method). It doesn't matter of you use all validation samples or sub-sample equal ratio of train unknown samples. The validation accuracy and loss should be \"about the same\" for both of the two cases if your generalization is good enough.\n\n\n[ensemble]\n\n - just keep making different models and ensemble them. In my experiments, i can get 0.87 by ensembling about 10 models of in the range of 0.85 to 0.86. Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting \n\n\n\n[other tricks]\n\n- try wave + reflected wave (train separately and test separately).\n  - because of padding, stride, etc ... the results are \"slightly\" different. Note that you can flip left-right, up-down.\n  - there are also other test-time augmentation you can use like cropping, stretch amplitude, etc\n\n- will  unsupervised learning using pseudo-label overfits?\n\n  - Create two representation views for a train sample (e.g. view1=orginal and view2=reflected , or view1=wave(1d) and view2=mfcc (2d)\n\n  - use view1 to pedsuo-label and view-2 to train new model. \n\n\n\n[how about raspberry pi special prize?]\n\n - it may be easier to get  good accuracy with an ensemble of teacher models first, then transfer this knowledge to a low complexity student (i.e. Professor Hinton's \"knowledge distillation\" paper).\n - In my experiment, my best single model  made from knowledge distillation has LB 0.88 (same as the ensemble)",
      "votes": 92
    },
    {
      "id": 265499,
      "postDate": "2018-01-05T17:12:14.433Z",
      "content": "<p>Thank you for sharing your insights. </p>\n\n<p>I can confirm that without your trick on the dataset, one can get still get a single model with .88. </p>\n\n<p>I do get better results with spectrogram as well. My best model on MFCC is .86, best model on the raw waveform is .85. </p>\n\n<p>My CNN models work better than RNN models (about 0.02 higher on same input features). I am running out of time (due to family events) to do experiments on hybrid approaches like CRNN or concatenate RNN and CNN output layers, but I got some architecture around .85 ~ .86 range with these approaches.</p>",
      "rawMarkdown": "Thank you for sharing your insights. \n\nI can confirm that without your trick on the dataset, one can get still get a single model with .88. \n\nI do get better results with spectrogram as well. My best model on MFCC is .86, best model on the raw waveform is .85. \n\nMy CNN models work better than RNN models (about 0.02 higher on same input features). I am running out of time (due to family events) to do experiments on hybrid approaches like CRNN or concatenate RNN and CNN output layers, but I got some architecture around .85 ~ .86 range with these approaches.",
      "votes": 8,
      "replies": [
        {
          "id": 265535,
          "postDate": "2018-01-05T19:36:26.070Z",
          "content": "<p>Hi Ren, were you able to get that score only by using provided training set ?</p>",
          "rawMarkdown": "Hi Ren, were you able to get that score only by using provided training set ?"
        },
        {
          "id": 265539,
          "postDate": "2018-01-05T19:47:48.257Z",
          "content": "<p>Yes, only thing I do to the training data is cut silence wav into pieces, add random noises and adjust volumes</p>",
          "rawMarkdown": "Yes, only thing I do to the training data is cut silence wav into pieces, add random noises and adjust volumes",
          "votes": 4
        },
        {
          "id": 265571,
          "postDate": "2018-01-05T22:21:53.247Z",
          "content": "<p>Thanks for the comments. From the discussion at <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239</a>, it is mentioned:</p>\n\n<p>\"Yes, in general, unsupervised learning is fine with the Test set. As long as, per the Rules, there is no hand labeling of the Test set. \" --inversion</p>\n\n<p>Hence, I think my method is not against the rule.</p>",
          "rawMarkdown": "Thanks for the comments. From the discussion at https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239, it is mentioned:\n\n\"Yes, in general, unsupervised learning is fine with the Test set. As long as, per the Rules, there is no hand labeling of the Test set. \" --inversion\n\nHence, I think my method is not against the rule.",
          "votes": 2
        },
        {
          "id": 265577,
          "postDate": "2018-01-05T22:47:34.280Z",
          "content": "<p>great,I missed that.</p>",
          "rawMarkdown": "great,I missed that.",
          "votes": 1
        }
      ]
    },
    {
      "id": 265593,
      "postDate": "2018-01-06T01:14:51.410Z",
      "content": "<p>@Heng, Could you point me to some sample code for computing the logmelspectrum.</p>",
      "rawMarkdown": "@Heng, Could you point me to some sample code for computing the logmelspectrum.",
      "votes": 3,
      "replies": [
        {
          "id": 265670,
          "postDate": "2018-01-06T08:52:26.203Z",
          "content": "<p>also refer to this thread: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982</a></p>",
          "rawMarkdown": "also refer to this thread: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982",
          "votes": 4
        },
        {
          "id": 265750,
          "postDate": "2018-01-06T15:15:36.750Z",
          "content": "<p>thanks</p>",
          "rawMarkdown": "thanks"
        }
      ]
    },
    {
      "id": 269494,
      "postDate": "2018-01-16T23:29:01.887Z",
      "content": "<p>Thanks for those points - I think the data augmentation was crucial in this task. I pitch-shifted, distorted, stretched, filtered and cropped the files randomly. I even bounced down some of the original noise tracks to super low bitrate mp3 then back into wav because I noticed by nets were getting fooled by the weird bubbly artefacts. I also found that data balancing during training was quite important - if I didn't train on a roughly even distribution the scores went way down.</p>\n\n<p>I got a great boost by applying a 2d Scattering Wavelet transform on a spectogram followed by a 10L Resnet - the best single model scored 88% on LB. \nScattering transforms are really interesting topic, they're basically convolutional filters (morlet &amp; gabor in my case) followed by a non-linearity, like convnets, however they have certain properties that make them favourable as feature descriptors, plus they come pre-defined so there's that less burden on the net to learn the low level filters, but stronger representational power - here's some papers/info: \n<a href=\"https://arxiv.org/abs/1304.6763\">https://arxiv.org/abs/1304.6763</a>\n<a href=\"https://arxiv.org/abs/1312.5940\">https://arxiv.org/abs/1312.5940</a>\n<a href=\"http://helper.ipam.ucla.edu/publications/gss2012/gss2012_10668.pdf\">http://helper.ipam.ucla.edu/publications/gss2012/gss2012_10668.pdf</a></p>\n\n<p>I also had fun playing with ConvLSTMs - I was training a sample-level model using stacks of these using different scale frames a la sampleRNN, they were promising but I sadly ran out of time as they obviously take a long time to train.</p>\n\n<p>If I were to do this again I'd definitely start sooner on examining and exploring/using the test data because I discovered a lot of the things you mentioned far too late - however, for a first time had a great laugh, thanks all! </p>",
      "rawMarkdown": "Thanks for those points - I think the data augmentation was crucial in this task. I pitch-shifted, distorted, stretched, filtered and cropped the files randomly. I even bounced down some of the original noise tracks to super low bitrate mp3 then back into wav because I noticed by nets were getting fooled by the weird bubbly artefacts. I also found that data balancing during training was quite important - if I didn't train on a roughly even distribution the scores went way down.\n\nI got a great boost by applying a 2d Scattering Wavelet transform on a spectogram followed by a 10L Resnet - the best single model scored 88% on LB. \nScattering transforms are really interesting topic, they're basically convolutional filters (morlet &amp; gabor in my case) followed by a non-linearity, like convnets, however they have certain properties that make them favourable as feature descriptors, plus they come pre-defined so there's that less burden on the net to learn the low level filters, but stronger representational power - here's some papers/info: \nhttps://arxiv.org/abs/1304.6763\nhttps://arxiv.org/abs/1312.5940\nhttp://helper.ipam.ucla.edu/publications/gss2012/gss2012_10668.pdf\n\nI also had fun playing with ConvLSTMs - I was training a sample-level model using stacks of these using different scale frames a la sampleRNN, they were promising but I sadly ran out of time as they obviously take a long time to train.\n\nIf I were to do this again I'd definitely start sooner on examining and exploring/using the test data because I discovered a lot of the things you mentioned far too late - however, for a first time had a great laugh, thanks all! \n ",
      "votes": 4,
      "replies": [
        {
          "id": 269512,
          "postDate": "2018-01-17T00:12:28.577Z",
          "content": "<p>I was able to get .87 LB with Baidu's KWS CRNN (<a href=\"https://arxiv.org/abs/1703.05390\">https://arxiv.org/abs/1703.05390</a>), 80-bin log mel spectrogram input, and GRU instead of LSTM.</p>",
          "rawMarkdown": "I was able to get .87 LB with Baidu's KWS CRNN (https://arxiv.org/abs/1703.05390), 80-bin log mel spectrogram input, and GRU instead of LSTM.",
          "votes": 1
        },
        {
          "id": 363147,
          "postDate": "2018-07-28T02:17:16.833Z",
          "content": "<p>Thanks for  sharing your idea. Scattering Wavelet transform is an interesting research. Could you point me to some sample code about Scattering Wavelet transform ? </p>",
          "rawMarkdown": "Thanks for  sharing your idea. Scattering Wavelet transform is an interesting research. Could you point me to some sample code about Scattering Wavelet transform ? "
        },
        {
          "id": 383292,
          "postDate": "2018-09-08T09:18:17.287Z",
          "content": "<p>Sure, here's one in pytorch <a href=\"https://github.com/edouardoyallon/pyscatwave\">https://github.com/edouardoyallon/pyscatwave</a> </p>",
          "rawMarkdown": "Sure, here's one in pytorch https://github.com/edouardoyallon/pyscatwave "
        }
      ]
    },
    {
      "id": 268202,
      "postDate": "2018-01-13T18:50:35.850Z",
      "content": "<p>I wonder what kind of results you could get with <a href=\"https://github.com/facebookresearch/wav2letter\">wav2letter</a> based on the <a href=\"https://arxiv.org/abs/1712.09444\">Letter-Based Speech Recognition with Gated ConvNets</a> paper.</p>",
      "rawMarkdown": "I wonder what kind of results you could get with [wav2letter][1] based on the [Letter-Based Speech Recognition with Gated ConvNets][2] paper.\n\n\n  [1]: https://github.com/facebookresearch/wav2letter\n  [2]: https://arxiv.org/abs/1712.09444",
      "votes": 1,
      "replies": [
        {
          "id": 268281,
          "postDate": "2018-01-14T01:55:42.190Z",
          "content": "<p>Thanks for the paper. I was searching for methods that detect phonemes and ctc loss. This paper seems to work.</p>",
          "rawMarkdown": "Thanks for the paper. I was searching for methods that detect phonemes and ctc loss. This paper seems to work.\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 267476,
      "postDate": "2018-01-11T13:08:24.367Z",
      "content": "<p>Thanks so much for sharing.</p>\n\n<p>How do you adjust the \"silence\" and \"unknown\" samples distribution in the training set? Wouldn't it be better to use more \"unknown\" samples? However, in my experiments, using more \"unknown\" training samples (in percentage) may hurt the overall accuracy: if I set the percentage of \"unknown\" in training set to 10% and 20%, they both get 0.88 but 20% is slightly worse. Setting it to 35% only gets 0.87.</p>\n\n<p>Another way to use more \"unknown\" samples may be to use different \"unknown\" samples in every epoch so that overall percentage of \"unknown\" in every epoch wouldn't be too large. Would it cause problems? (It seems that the percentage of \"unknown\" in the test result drops several percent this way)</p>",
      "rawMarkdown": "Thanks so much for sharing.\n\nHow do you adjust the \"silence\" and \"unknown\" samples distribution in the training set? Wouldn't it be better to use more \"unknown\" samples? However, in my experiments, using more \"unknown\" training samples (in percentage) may hurt the overall accuracy: if I set the percentage of \"unknown\" in training set to 10% and 20%, they both get 0.88 but 20% is slightly worse. Setting it to 35% only gets 0.87.\n\nAnother way to use more \"unknown\" samples may be to use different \"unknown\" samples in every epoch so that overall percentage of \"unknown\" in every epoch wouldn't be too large. Would it cause problems? (It seems that the percentage of \"unknown\" in the test result drops several percent this way)",
      "votes": 1
    },
    {
      "id": 266807,
      "postDate": "2018-01-09T20:32:38.967Z",
      "content": "<p>good</p>",
      "rawMarkdown": "good",
      "votes": 1
    },
    {
      "id": 265485,
      "postDate": "2018-01-05T16:26:06.320Z",
      "content": "<p>Thanks a lot for sharing ideas, I hope I will succeed implementing some of them )</p>\n\n<p>@simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). </p>\n\n<p>Did you use specgram as an input for these models? what do you think about the input size (i.e. window and step while making spectrogram)?</p>",
      "rawMarkdown": "Thanks a lot for sharing ideas, I hope I will succeed implementing some of them )\n\n@simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). \n\nDid you use specgram as an input for these models? what do you think about the input size (i.e. window and step while making spectrogram)?",
      "votes": 2,
      "replies": [
        {
          "id": 265572,
          "postDate": "2018-01-05T22:24:05.713Z",
          "content": "<p>I use spectrogram as input. It is the same size  (40x101) used in the tensorflow tutorial. </p>",
          "rawMarkdown": "I use spectrogram as input. It is the same size  (40x101) used in the tensorflow tutorial. ",
          "votes": 1
        },
        {
          "id": 265668,
          "postDate": "2018-01-06T08:51:44.280Z",
          "content": "<p>also refer to this thread: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982</a></p>",
          "rawMarkdown": "also refer to this thread: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982",
          "votes": 2
        },
        {
          "id": 267638,
          "postDate": "2018-01-11T23:24:50.557Z",
          "content": "<p>First time competitor here. On those vgg-like models, how many samples are you using in your training set (including augmentation) and about how long are you letting the model train (in terms of epochs)? Thanks</p>",
          "rawMarkdown": "First time competitor here. On those vgg-like models, how many samples are you using in your training set (including augmentation) and about how long are you letting the model train (in terms of epochs)? Thanks"
        }
      ]
    },
    {
      "id": 265969,
      "postDate": "2018-01-07T09:27:07.813Z",
      "content": "<p>\" test-time augmentation\" shouldn't be allowed, I think.</p>",
      "rawMarkdown": "\" test-time augmentation\" shouldn't be allowed, I think.",
      "votes": 1,
      "replies": [
        {
          "id": 265970,
          "postDate": "2018-01-07T09:28:22.920Z",
          "content": "<p>why not? is it in the rule? It is a common method used in other kaggle competitions.</p>",
          "rawMarkdown": "why not? is it in the rule? It is a common method used in other kaggle competitions.\n",
          "votes": 5
        }
      ]
    },
    {
      "id": 269428,
      "postDate": "2018-01-16T20:36:47.807Z",
      "content": "<p>Thank you Heng! your posts really helped me during the contest.</p>",
      "rawMarkdown": "Thank you Heng! your posts really helped me during the contest."
    },
    {
      "id": 269099,
      "postDate": "2018-01-16T07:47:14.457Z",
      "content": "<p>\"Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting \"</p>\n\n<p>Is the ^1 majority voting ? Thanks.</p>",
      "rawMarkdown": "\"Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting \"\n\nIs the ^1 majority voting ? Thanks.",
      "replies": [
        {
          "id": 269100,
          "postDate": "2018-01-16T07:48:31.020Z",
          "content": "<p>it is not. It is my mistake</p>",
          "rawMarkdown": "it is not. It is my mistake"
        }
      ]
    },
    {
      "id": 266267,
      "postDate": "2018-01-08T09:36:48.883Z",
      "content": "<p>For the regular prize (i.e. not the \"Special Prize\") is there any requirement to use only a single model? Any limitation of using the ensemble method or big models?</p>",
      "rawMarkdown": "For the regular prize (i.e. not the \"Special Prize\") is there any requirement to use only a single model? Any limitation of using the ensemble method or big models?",
      "replies": [
        {
          "id": 266346,
          "postDate": "2018-01-08T15:09:28.580Z",
          "content": "<p>I would think that so long as it fits on the Pi, it should be fine. </p>",
          "rawMarkdown": "I would think that so long as it fits on the Pi, it should be fine. "
        },
        {
          "id": 266347,
          "postDate": "2018-01-08T15:11:04.340Z",
          "content": "<p>So , just to get it clear, the non-special prize also must fit the Pi device?</p>",
          "rawMarkdown": "So , just to get it clear, the non-special prize also must fit the Pi device?"
        },
        {
          "id": 268595,
          "postDate": "2018-01-15T00:13:05.240Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 268643,
          "postDate": "2018-01-15T04:35:22.220Z",
          "content": "<p>Thanks</p>",
          "rawMarkdown": "Thanks"
        },
        {
          "id": 269572,
          "postDate": "2018-01-17T02:18:56.947Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 265528,
      "postDate": "2018-01-05T19:06:43.543Z",
      "content": "<p>Thanks @Heng? just a quick clarification.\nWhen you say </p>\n\n<blockquote>\n  <p>the next step is to generate these background train samples same the LB. For me I simply use \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.</p>\n</blockquote>\n\n<p>After we make the prediction on test set and use the results as new train data and their labels, do we also keep all the test set unchanged or do we discard those entries that now appear in training data?</p>",
      "rawMarkdown": "Thanks @Heng? just a quick clarification.\nWhen you say \n\n&gt; the next step is to generate these background train samples same the LB. For me I simply use \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.\n\nAfter we make the prediction on test set and use the results as new train data and their labels, do we also keep all the test set unchanged or do we discard those entries that now appear in training data?\n",
      "replies": [
        {
          "id": 265531,
          "postDate": "2018-01-05T19:24:35.720Z",
          "content": "<p>I think heng means the LB test set.\nIn essence, train your model, predict probs on the LB test set. Select LB test examples with high unknown or silence probs and include them in your train set and retrain again with this new train set</p>\n\n<p>EDIT:\nrelevant links\n1) <a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a>\n2) <a href=\"https://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4706\">https://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4706</a>\n3) <a href=\"https://www.analyticsvidhya.com/blog/2017/09/pseudo-labelling-semi-supervised-learning-technique/\">https://www.analyticsvidhya.com/blog/2017/09/pseudo-labelling-semi-supervised-learning-technique/</a></p>",
          "rawMarkdown": "I think heng means the LB test set.\nIn essence, train your model, predict probs on the LB test set. Select LB test examples with high unknown or silence probs and include them in your train set and retrain again with this new train set\n\nEDIT:\nrelevant links\n1) http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\n2) https://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4706\n3) https://www.analyticsvidhya.com/blog/2017/09/pseudo-labelling-semi-supervised-learning-technique/\n",
          "votes": 2
        },
        {
          "id": 265532,
          "postDate": "2018-01-05T19:32:44.133Z",
          "content": "<p>@RaviTejaGutta thanks ;)</p>",
          "rawMarkdown": "@RaviTejaGutta thanks ;)"
        }
      ]
    },
    {
      "id": 412359,
      "postDate": "2018-10-30T04:08:36.187Z",
      "content": "<p>Thank you for the guidance!</p>",
      "rawMarkdown": "Thank you for the guidance!"
    },
    {
      "id": 266831,
      "postDate": "2018-01-09T22:43:22.573Z",
      "content": "<p>Awesome post - thanks!</p>",
      "rawMarkdown": "Awesome post - thanks!"
    }
  ],
  "comments": [
    {
      "id": 265499,
      "author_name": "Ren",
      "author_url": "",
      "post_date": "2018-01-05T17:12:14.433000",
      "content": "<p>Thank you for sharing your insights. </p>\n\n<p>I can confirm that without your trick on the dataset, one can get still get a single model with .88. </p>\n\n<p>I do get better results with spectrogram as well. My best model on MFCC is .86, best model on the raw waveform is .85. </p>\n\n<p>My CNN models work better than RNN models (about 0.02 higher on same input features). I am running out of time (due to family events) to do experiments on hybrid approaches like CRNN or concatenate RNN and CNN output layers, but I got some architecture around .85 ~ .86 range with these approaches.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 265535,
          "author_name": "Ravi Teja Gutta",
          "author_url": "",
          "post_date": "2018-01-05T19:36:26.070000",
          "content": "<p>Hi Ren, were you able to get that score only by using provided training set ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 265539,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-01-05T19:47:48.257000",
          "content": "<p>Yes, only thing I do to the training data is cut silence wav into pieces, add random noises and adjust volumes</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 265571,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-01-05T22:21:53.247000",
          "content": "<p>Thanks for the comments. From the discussion at <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44239</a>, it is mentioned:</p>\n\n<p>\"Yes, in general, unsupervised learning is fine with the Test set. As long as, per the Rules, there is no hand labeling of the Test set. \" --inversion</p>\n\n<p>Hence, I think my method is not against the rule.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 265577,
          "author_name": "Ren",
          "author_url": "",
          "post_date": "2018-01-05T22:47:34.280000",
          "content": "<p>great,I missed that.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 265593,
      "author_name": "Robert",
      "author_url": "",
      "post_date": "2018-01-06T01:14:51.410000",
      "content": "<p>@Heng, Could you point me to some sample code for computing the logmelspectrum.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 265670,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-01-06T08:52:26.203000",
          "content": "<p>also refer to this thread: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982</a></p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 265750,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-01-06T15:15:36.750000",
          "content": "<p>thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 269494,
      "author_name": "cthom055",
      "author_url": "",
      "post_date": "2018-01-16T23:29:01.887000",
      "content": "<p>Thanks for those points - I think the data augmentation was crucial in this task. I pitch-shifted, distorted, stretched, filtered and cropped the files randomly. I even bounced down some of the original noise tracks to super low bitrate mp3 then back into wav because I noticed by nets were getting fooled by the weird bubbly artefacts. I also found that data balancing during training was quite important - if I didn't train on a roughly even distribution the scores went way down.</p>\n\n<p>I got a great boost by applying a 2d Scattering Wavelet transform on a spectogram followed by a 10L Resnet - the best single model scored 88% on LB. \nScattering transforms are really interesting topic, they're basically convolutional filters (morlet &amp; gabor in my case) followed by a non-linearity, like convnets, however they have certain properties that make them favourable as feature descriptors, plus they come pre-defined so there's that less burden on the net to learn the low level filters, but stronger representational power - here's some papers/info: \n<a href=\"https://arxiv.org/abs/1304.6763\">https://arxiv.org/abs/1304.6763</a>\n<a href=\"https://arxiv.org/abs/1312.5940\">https://arxiv.org/abs/1312.5940</a>\n<a href=\"http://helper.ipam.ucla.edu/publications/gss2012/gss2012_10668.pdf\">http://helper.ipam.ucla.edu/publications/gss2012/gss2012_10668.pdf</a></p>\n\n<p>I also had fun playing with ConvLSTMs - I was training a sample-level model using stacks of these using different scale frames a la sampleRNN, they were promising but I sadly ran out of time as they obviously take a long time to train.</p>\n\n<p>If I were to do this again I'd definitely start sooner on examining and exploring/using the test data because I discovered a lot of the things you mentioned far too late - however, for a first time had a great laugh, thanks all! </p>",
      "votes": 4,
      "replies": [
        {
          "id": 269512,
          "author_name": "Jeff",
          "author_url": "",
          "post_date": "2018-01-17T00:12:28.577000",
          "content": "<p>I was able to get .87 LB with Baidu's KWS CRNN (<a href=\"https://arxiv.org/abs/1703.05390\">https://arxiv.org/abs/1703.05390</a>), 80-bin log mel spectrogram input, and GRU instead of LSTM.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 363147,
          "author_name": "Zhangguanhua",
          "author_url": "",
          "post_date": "2018-07-28T02:17:16.833000",
          "content": "<p>Thanks for  sharing your idea. Scattering Wavelet transform is an interesting research. Could you point me to some sample code about Scattering Wavelet transform ? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 383292,
          "author_name": "cthom055",
          "author_url": "",
          "post_date": "2018-09-08T09:18:17.287000",
          "content": "<p>Sure, here's one in pytorch <a href=\"https://github.com/edouardoyallon/pyscatwave\">https://github.com/edouardoyallon/pyscatwave</a> </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 268202,
      "author_name": "jvent",
      "author_url": "",
      "post_date": "2018-01-13T18:50:35.850000",
      "content": "<p>I wonder what kind of results you could get with <a href=\"https://github.com/facebookresearch/wav2letter\">wav2letter</a> based on the <a href=\"https://arxiv.org/abs/1712.09444\">Letter-Based Speech Recognition with Gated ConvNets</a> paper.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 268281,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-01-14T01:55:42.190000",
          "content": "<p>Thanks for the paper. I was searching for methods that detect phonemes and ctc loss. This paper seems to work.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 267476,
      "author_name": "Weiwei Chen",
      "author_url": "",
      "post_date": "2018-01-11T13:08:24.367000",
      "content": "<p>Thanks so much for sharing.</p>\n\n<p>How do you adjust the \"silence\" and \"unknown\" samples distribution in the training set? Wouldn't it be better to use more \"unknown\" samples? However, in my experiments, using more \"unknown\" training samples (in percentage) may hurt the overall accuracy: if I set the percentage of \"unknown\" in training set to 10% and 20%, they both get 0.88 but 20% is slightly worse. Setting it to 35% only gets 0.87.</p>\n\n<p>Another way to use more \"unknown\" samples may be to use different \"unknown\" samples in every epoch so that overall percentage of \"unknown\" in every epoch wouldn't be too large. Would it cause problems? (It seems that the percentage of \"unknown\" in the test result drops several percent this way)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 266807,
      "author_name": "Mahesh",
      "author_url": "",
      "post_date": "2018-01-09T20:32:38.967000",
      "content": "<p>good</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 265485,
      "author_name": "Ivan Timoshilov",
      "author_url": "",
      "post_date": "2018-01-05T16:26:06.320000",
      "content": "<p>Thanks a lot for sharing ideas, I hope I will succeed implementing some of them )</p>\n\n<p>@simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). </p>\n\n<p>Did you use specgram as an input for these models? what do you think about the input size (i.e. window and step while making spectrogram)?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 265572,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-01-05T22:24:05.713000",
          "content": "<p>I use spectrogram as input. It is the same size  (40x101) used in the tensorflow tutorial. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 265668,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-01-06T08:51:44.280000",
          "content": "<p>also refer to this thread: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46982</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 267638,
          "author_name": "formigone",
          "author_url": "",
          "post_date": "2018-01-11T23:24:50.557000",
          "content": "<p>First time competitor here. On those vgg-like models, how many samples are you using in your training set (including augmentation) and about how long are you letting the model train (in terms of epochs)? Thanks</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 265969,
      "author_name": "Feiteng",
      "author_url": "",
      "post_date": "2018-01-07T09:27:07.813000",
      "content": "<p>\" test-time augmentation\" shouldn't be allowed, I think.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 265970,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-01-07T09:28:22.920000",
          "content": "<p>why not? is it in the rule? It is a common method used in other kaggle competitions.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 269428,
      "author_name": "Ori Tal",
      "author_url": "",
      "post_date": "2018-01-16T20:36:47.807000",
      "content": "<p>Thank you Heng! your posts really helped me during the contest.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 269099,
      "author_name": "Yibing Wu",
      "author_url": "",
      "post_date": "2018-01-16T07:47:14.457000",
      "content": "<p>\"Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting \"</p>\n\n<p>Is the ^1 majority voting ? Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 269100,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-01-16T07:48:31.020000",
          "content": "<p>it is not. It is my mistake</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 266267,
      "author_name": "Ori Tal",
      "author_url": "",
      "post_date": "2018-01-08T09:36:48.883000",
      "content": "<p>For the regular prize (i.e. not the \"Special Prize\") is there any requirement to use only a single model? Any limitation of using the ensemble method or big models?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 266346,
          "author_name": "pskwarko",
          "author_url": "",
          "post_date": "2018-01-08T15:09:28.580000",
          "content": "<p>I would think that so long as it fits on the Pi, it should be fine. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 266347,
          "author_name": "Ori Tal",
          "author_url": "",
          "post_date": "2018-01-08T15:11:04.340000",
          "content": "<p>So , just to get it clear, the non-special prize also must fit the Pi device?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 268595,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-01-15T00:13:05.240000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 268643,
          "author_name": "Ori Tal",
          "author_url": "",
          "post_date": "2018-01-15T04:35:22.220000",
          "content": "<p>Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 269572,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-01-17T02:18:56.947000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 265528,
      "author_name": "kirk",
      "author_url": "",
      "post_date": "2018-01-05T19:06:43.543000",
      "content": "<p>Thanks @Heng? just a quick clarification.\nWhen you say </p>\n\n<blockquote>\n  <p>the next step is to generate these background train samples same the LB. For me I simply use \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.</p>\n</blockquote>\n\n<p>After we make the prediction on test set and use the results as new train data and their labels, do we also keep all the test set unchanged or do we discard those entries that now appear in training data?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 265531,
          "author_name": "Ravi Teja Gutta",
          "author_url": "",
          "post_date": "2018-01-05T19:24:35.720000",
          "content": "<p>I think heng means the LB test set.\nIn essence, train your model, predict probs on the LB test set. Select LB test examples with high unknown or silence probs and include them in your train set and retrain again with this new train set</p>\n\n<p>EDIT:\nrelevant links\n1) <a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a>\n2) <a href=\"https://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4706\">https://www.kaggle.com/c/challenges-in-representation-learning-the-black-box-learning-challenge/discussion/4706</a>\n3) <a href=\"https://www.analyticsvidhya.com/blog/2017/09/pseudo-labelling-semi-supervised-learning-technique/\">https://www.analyticsvidhya.com/blog/2017/09/pseudo-labelling-semi-supervised-learning-technique/</a></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 265532,
          "author_name": "kirk",
          "author_url": "",
          "post_date": "2018-01-05T19:32:44.133000",
          "content": "<p>@RaviTejaGutta thanks ;)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 412359,
      "author_name": "John Tucker",
      "author_url": "",
      "post_date": "2018-10-30T04:08:36.187000",
      "content": "<p>Thank you for the guidance!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 266831,
      "author_name": "Abhinav R.",
      "author_url": "",
      "post_date": "2018-01-09T22:43:22.573000",
      "content": "<p>Awesome post - thanks!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "265463": "(this discussion will be constantly updated)\n\n[dataset]\n\n- At first, this challenge looks like a 12-class classification problem. But actually, it is \"not\". There are 10 class and 2 background class \"silence\" and \"unknown\". Note that there are no \"silence\" train samples and \"unknown\" train samples is not exactly the same as that of LB set.\n- If you can generate correct \"silence\" and \"unknown\" train data, you should get good results (e.g. single model in the range of 0.87 to 0.88)\n- How to probe the LB \"silence\" and \"unknown\" distribution? Train a classifier and use it to \"pseudo-label\" the LB test data. Sample the test set using the \"pseudo-label\" (choose different predicted probability) and listen to those LB test samples. It is not difficult to identify their characteristics (e.g. how does LB silence sounds like? what are the missing LB unknown from the train, ...)\n- the next step is to generate these background train samples same the LB. For me I simply use  \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.\n- there will be some label noise, but deep learning can accommodate 10 to 20% label noise  if your data is large enough\n-  if your sampling is correct, simple \"cnn_trad_pool2_net + mfcc\"  from the paper gives LB of around 0.82 and 0.83 in my experiments\n\n\n[input]\n\n - you can choose either to to use: raw wave form (1d), logmelspectrum (2d), mfcc (2d), or mixture of these.\n\n - for my case, logmelspectrum (2d) is better than mfcc. I am trying raw wave form (1d), but haven't got good results yet.\n\n - simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). Convergence is very fast and you can get validation accuracy in the range 0.95 to 0.97, validation loss of 0.18 to 0.13. (the standard validation split from the tensorflow hash method). It doesn't matter of you use all validation samples or sub-sample equal ratio of train unknown samples. The validation accuracy and loss should be \"about the same\" for both of the two cases if your generalization is good enough.\n\n\n[ensemble]\n\n - just keep making different models and ensemble them. In my experiments, i can get 0.87 by ensembling about 10 models of in the range of 0.85 to 0.86. Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting \n\n\n\n[other tricks]\n\n- try wave + reflected wave (train separately and test separately).\n  - because of padding, stride, etc ... the results are \"slightly\" different. Note that you can flip left-right, up-down.\n  - there are also other test-time augmentation you can use like cropping, stretch amplitude, etc\n\n- will  unsupervised learning using pseudo-label overfits?\n\n  - Create two representation views for a train sample (e.g. view1=orginal and view2=reflected , or view1=wave(1d) and view2=mfcc (2d)\n\n  - use view1 to pedsuo-label and view-2 to train new model. \n\n\n\n[how about raspberry pi special prize?]\n\n - it may be easier to get  good accuracy with an ensemble of teacher models first, then transfer this knowledge to a low complexity student (i.e. Professor Hinton's \"knowledge distillation\" paper).\n - In my experiment, my best single model  made from knowledge distillation has LB 0.88 (same as the ensemble)",
    "265499": "Thank you for sharing your insights. \n\nI can confirm that without your trick on the dataset, one can get still get a single model with .88. \n\nI do get better results with spectrogram as well. My best model on MFCC is .86, best model on the raw waveform is .85. \n\nMy CNN models work better than RNN models (about 0.02 higher on same input features). I am running out of time (due to family events) to do experiments on hybrid approaches like CRNN or concatenate RNN and CNN output layers, but I got some architecture around .85 ~ .86 range with these approaches.",
    "265593": "@Heng, Could you point me to some sample code for computing the logmelspectrum.",
    "269494": "Thanks for those points - I think the data augmentation was crucial in this task. I pitch-shifted, distorted, stretched, filtered and cropped the files randomly. I even bounced down some of the original noise tracks to super low bitrate mp3 then back into wav because I noticed by nets were getting fooled by the weird bubbly artefacts. I also found that data balancing during training was quite important - if I didn't train on a roughly even distribution the scores went way down.\n\nI got a great boost by applying a 2d Scattering Wavelet transform on a spectogram followed by a 10L Resnet - the best single model scored 88% on LB. \nScattering transforms are really interesting topic, they're basically convolutional filters (morlet &amp; gabor in my case) followed by a non-linearity, like convnets, however they have certain properties that make them favourable as feature descriptors, plus they come pre-defined so there's that less burden on the net to learn the low level filters, but stronger representational power - here's some papers/info: \nhttps://arxiv.org/abs/1304.6763\nhttps://arxiv.org/abs/1312.5940\nhttp://helper.ipam.ucla.edu/publications/gss2012/gss2012_10668.pdf\n\nI also had fun playing with ConvLSTMs - I was training a sample-level model using stacks of these using different scale frames a la sampleRNN, they were promising but I sadly ran out of time as they obviously take a long time to train.\n\nIf I were to do this again I'd definitely start sooner on examining and exploring/using the test data because I discovered a lot of the things you mentioned far too late - however, for a first time had a great laugh, thanks all! \n ",
    "268202": "I wonder what kind of results you could get with [wav2letter][1] based on the [Letter-Based Speech Recognition with Gated ConvNets][2] paper.\n\n\n  [1]: https://github.com/facebookresearch/wav2letter\n  [2]: https://arxiv.org/abs/1712.09444",
    "267476": "Thanks so much for sharing.\n\nHow do you adjust the \"silence\" and \"unknown\" samples distribution in the training set? Wouldn't it be better to use more \"unknown\" samples? However, in my experiments, using more \"unknown\" training samples (in percentage) may hurt the overall accuracy: if I set the percentage of \"unknown\" in training set to 10% and 20%, they both get 0.88 but 20% is slightly worse. Setting it to 35% only gets 0.87.\n\nAnother way to use more \"unknown\" samples may be to use different \"unknown\" samples in every epoch so that overall percentage of \"unknown\" in every epoch wouldn't be too large. Would it cause problems? (It seems that the percentage of \"unknown\" in the test result drops several percent this way)",
    "266807": "good",
    "265485": "Thanks a lot for sharing ideas, I hope I will succeed implementing some of them )\n\n@simple vgg-like, resnet-like net can get about 0.85 to 0.87 (plain single model, no test-augmentation). \n\nDid you use specgram as an input for these models? what do you think about the input size (i.e. window and step while making spectrogram)?",
    "265969": "\" test-time augmentation\" shouldn't be allowed, I think.",
    "269428": "Thank you Heng! your posts really helped me during the contest.",
    "269099": "\"Use ensemble_prob = SUM { model_prob^0.5 }. You can other fator like ^0.4, ^0.3 .... in the extreme case ^0 , i.e majority voting \"\n\nIs the ^1 majority voting ? Thanks.",
    "266267": "For the regular prize (i.e. not the \"Special Prize\") is there any requirement to use only a single model? Any limitation of using the ensemble method or big models?",
    "265528": "Thanks @Heng? just a quick clarification.\nWhen you say \n\n&gt; the next step is to generate these background train samples same the LB. For me I simply use \"pseudo-label\" LB samples for these two background class. I choose those high confidence ones (e.g. &gt;0.95) and add them to my train set.\n\nAfter we make the prediction on test set and use the results as new train data and their labels, do we also keep all the test set unchanged or do we discard those entries that now appear in training data?\n",
    "412359": "Thank you for the guidance!",
    "266831": "Awesome post - thanks!"
  }
}