{
  "id": 44283,
  "title": "Anyone using 1-d convolutions?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44283",
  "author_name": "",
  "post_date": "2017-11-26T15:08:29.645453800Z",
  "votes": 38,
  "comment_count": 69,
  "views": 0,
  "content": "<p>It looks like the approach taken in most of the published kernels is to first convert the wav files to spectrograms and then using a convolutional neural network on these, essentially turning the sound classification problem into an image classification problem. </p>\n\n<p>I was just curious if anyone (besides me) has used 1-dimensional convolutions on the wav data directly? This basically only looks at the temporal structure of the data, not so much at the frequencies (although convolution in the time domain is the same as multiplication in the frequency domain, so these two things are obviously related somehow).</p>\n\n<p>My models are maxing out at 0.9 validation score, which translates to \"only\" 0.82 on the public LB. So the approach seems to work in theory, but it also seems like it still could be doing better.</p>",
  "messages": [
    {
      "id": "248593",
      "postDate": "11/26/2017 15:08:29",
      "content": "<p>It looks like the approach taken in most of the published kernels is to first convert the wav files to spectrograms and then using a convolutional neural network on these, essentially turning the sound classification problem into an image classification problem. </p>\n\n<p>I was just curious if anyone (besides me) has used 1-dimensional convolutions on the wav data directly? This basically only looks at the temporal structure of the data, not so much at the frequencies (although convolution in the time domain is the same as multiplication in the frequency domain, so these two things are obviously related somehow).</p>\n\n<p>My models are maxing out at 0.9 validation score, which translates to \"only\" 0.82 on the public LB. So the approach seems to work in theory, but it also seems like it still could be doing better.</p>",
      "rawMarkdown": "It looks like the approach taken in most of the published kernels is to first convert the wav files to spectrograms and then using a convolutional neural network on these, essentially turning the sound classification problem into an image classification problem. \n\nI was just curious if anyone (besides me) has used 1-dimensional convolutions on the wav data directly? This basically only looks at the temporal structure of the data, not so much at the frequencies (although convolution in the time domain is the same as multiplication in the frequency domain, so these two things are obviously related somehow).\n\nMy models are maxing out at 0.9 validation score, which translates to \"only\" 0.82 on the public LB. So the approach seems to work in theory, but it also seems like it still could be doing better.",
      "votes": null
    },
    {
      "id": "248792",
      "postDate": "11/27/2017 00:53:44",
      "content": "<p>I plan to experiment with 1-d convolutions. I'm also at a similar point, stuck at 0.82 on the LB with local validation scores over 93%. I hit 88-90% validation with the baseline TF model. I built a custom model with more capacity, takes 2x longer to train but hits 93+% validation consistently. Both models score almost exactly the same on the pub LB though.</p>\n\n<p>I'm going to see if I can get 1-d convolutions into a similar range. I expect that stacked dilated convolutions a-la Wavenet (minus the generative aspects) will be required to hit reasonable receptive field size for the samples... </p>",
      "rawMarkdown": "I plan to experiment with 1-d convolutions. I'm also at a similar point, stuck at 0.82 on the LB with local validation scores over 93%. I hit 88-90% validation with the baseline TF model. I built a custom model with more capacity, takes 2x longer to train but hits 93+% validation consistently. Both models score almost exactly the same on the pub LB though.\n\nI'm going to see if I can get 1-d convolutions into a similar range. I expect that stacked dilated convolutions a-la Wavenet (minus the generative aspects) will be required to hit reasonable receptive field size for the samples...",
      "votes": null
    },
    {
      "id": "248876",
      "postDate": "11/27/2017 06:53:48",
      "content": "<p>If you are into PyTorch, then I have a very nice notebook here exemplifying the use of a 1d CNN. \n<a href=\"https://github.com/QuantScientist/Deep-Learning-Boot-Camp/blob/master/day02-PyTORCH-and-PyCUDA/PyTorch/55-PyTorch-using-CONV1D-on-one-dimensional-data-CNN.ipynb\">https://github.com/QuantScientist/Deep-Learning-Boot-Camp/blob/master/day02-PyTORCH-and-PyCUDA/PyTorch/55-PyTorch-using-CONV1D-on-one-dimensional-data-CNN.ipynb</a></p>",
      "rawMarkdown": "If you are into PyTorch, then I have a very nice notebook here exemplifying the use of a 1d CNN. \nhttps://github.com/QuantScientist/Deep-Learning-Boot-Camp/blob/master/day02-PyTORCH-and-PyCUDA/PyTorch/55-PyTorch-using-CONV1D-on-one-dimensional-data-CNN.ipynb",
      "votes": null
    },
    {
      "id": "250152",
      "postDate": "11/29/2017 21:47:11",
      "content": "<p>I started with 1D for the first half of my model. With the intent of condensing the {16000,1] to [250,64] and resize to follow with standard 2D conv image analysis but it kept getting stuck in a trivial result. I'm guessing there was issues with resizing the tensor.</p>",
      "rawMarkdown": "I started with 1D for the first half of my model. With the intent of condensing the {16000,1] to [250,64] and resize to follow with standard 2D conv image analysis but it kept getting stuck in a trivial result. I'm guessing there was issues with resizing the tensor.",
      "votes": null
    },
    {
      "id": "250566",
      "postDate": "11/30/2017 05:26:23",
      "content": "<p>From a DSP perspective as you say, a 1-D convolution is equivalent to multiplying by the FFT of the convolution in the frequency domain. I think if we train a 1-D convolution on the waveform it should automatically adjust the filter taps to bring out the most pertinent frequencies that help distinguish between the relevant classes. If we train multiple 1-D convolutions in parallel, it should pick out the best selection of frequencies.</p>\n\n<p>If we convert to a spectrogram using the STFT, we effectively fix our set of frequencies with a linear spacing. This isn't great because we perceive audio logarithmically in both frequency and amplitude. The MFCC improves on the linear frequency spacing using the mel scale, which spaces the frequencies so we perceive them as increasing in pitch linearly. It also uses a log transform of the amplitudes. This makes the resulting feature vectors closer to what we perceive in practice. Pretty cool stuff!</p>\n\n<p>One thing I forgot to mention - some letters and sounds are like white noise, containing frequencies across the spectrum.   The 's' in 'yes' is a good example, in <a href=\"https://www.kaggle.com/davids1992/speech-visualization-and-exploration\">DavidS's excellent kernel</a> you can see it covers frequencies from 3.2kHz to 8kHz. This would need a broad filter to cover.</p>",
      "rawMarkdown": "From a DSP perspective as you say, a 1-D convolution is equivalent to multiplying by the FFT of the convolution in the frequency domain. I think if we train a 1-D convolution on the waveform it should automatically adjust the filter taps to bring out the most pertinent frequencies that help distinguish between the relevant classes. If we train multiple 1-D convolutions in parallel, it should pick out the best selection of frequencies.\n\nIf we convert to a spectrogram using the STFT, we effectively fix our set of frequencies with a linear spacing. This isn't great because we perceive audio logarithmically in both frequency and amplitude. The MFCC improves on the linear frequency spacing using the mel scale, which spaces the frequencies so we perceive them as increasing in pitch linearly. It also uses a log transform of the amplitudes. This makes the resulting feature vectors closer to what we perceive in practice. Pretty cool stuff!\n\nOne thing I forgot to mention - some letters and sounds are like white noise, containing frequencies across the spectrum.   The 's' in 'yes' is a good example, in [DavidS's excellent kernel](https://www.kaggle.com/davids1992/speech-visualization-and-exploration) you can see it covers frequencies from 3.2kHz to 8kHz. This would need a broad filter to cover.",
      "votes": null
    },
    {
      "id": "250650",
      "postDate": "11/30/2017 07:07:56",
      "content": "<p>A pre-programmed/fixed-weight 1D convolution using FIR filters would probably do well for the Raspberry Pi side of the competition...</p>",
      "rawMarkdown": "A pre-programmed/fixed-weight 1D convolution using FIR filters would probably do well for the Raspberry Pi side of the competition...",
      "votes": null
    },
    {
      "id": "252785",
      "postDate": "12/03/2017 21:43:41",
      "content": "<p>In case anyone is interested, I used a DenseNet-type architecture consisting of only 1-d convolutions with kernel size 9 (which seemed to work better than smaller sizes) and a \"swish\" activation function. The model has about 2M parameters. </p>\n\n<p>The training set consisted of the 12 classes in equal proportions. Even though the \"unknown\" class has many more examples that the other classes, I only included 1800 unknown examples in the training set, as well as 1800 silence examples. </p>\n\n<p>However, when loading the examples into a batch for training, when the label is \"unknown\" it randomly samples from <em>all</em> of the unknown examples -- so over time the model does see all of the unknown examples. (When the label is \"silence\", it creates a random white noise example on-the-fly.) Training examples are also augmented on-the-fly using the wavs from the background noise folder.</p>\n\n<p>For the validation set I combined the suggested validation and test sets into one big validation set, just to have more examples. Here, the unknown class contains about 8500 examples versus ~500 for each other class -- so while the training set is balanced the validation set is very unbalanced.</p>\n\n<p>I tried training this model in various ways (and with various architectural changes). The best score I got was 0.95 on the validation set, but unfortunately this only counted as 0.83 on the leaderboard. I'm not sure why that difference is so huge.</p>\n\n<p>Here is the precision/recall computed on the validation set:</p>\n\n<pre><code>             precision    recall  f1-score   support\n\n        yes       0.97      0.96      0.96       517\n         no       0.94      0.89      0.91       522\n         up       0.94      0.89      0.92       532\n       down       0.96      0.93      0.94       517\n       left       0.94      0.95      0.94       514\n      right       0.96      0.88      0.92       515\n         on       0.96      0.89      0.93       503\n        off       0.89      0.93      0.91       518\n       stop       0.96      0.94      0.95       495\n         go       0.95      0.89      0.92       511\n    silence       0.78      0.95      0.86       600\n    unknown       0.97      0.97      0.97      8489\n\navg / total       0.95      0.95      0.95     14233\n</code></pre>\n\n<p>Interestingly, precision on \"silence\" is rather low but when you listen to the examples that are misclassified as silence, they actually do sound like silence. So the labels seem to be wrong for those examples.</p>\n\n<p>Anyway, 0.83 on the LB seems the best my model can do at this point. So using 1-d convolutions and nothing else seems to be able to learn this dataset quite well.</p>",
      "rawMarkdown": "In case anyone is interested, I used a DenseNet-type architecture consisting of only 1-d convolutions with kernel size 9 (which seemed to work better than smaller sizes) and a \"swish\" activation function. The model has about 2M parameters. \n\nThe training set consisted of the 12 classes in equal proportions. Even though the \"unknown\" class has many more examples that the other classes, I only included 1800 unknown examples in the training set, as well as 1800 silence examples. \n\nHowever, when loading the examples into a batch for training, when the label is \"unknown\" it randomly samples from *all* of the unknown examples -- so over time the model does see all of the unknown examples. (When the label is \"silence\", it creates a random white noise example on-the-fly.) Training examples are also augmented on-the-fly using the wavs from the background noise folder.\n\nFor the validation set I combined the suggested validation and test sets into one big validation set, just to have more examples. Here, the unknown class contains about 8500 examples versus ~500 for each other class -- so while the training set is balanced the validation set is very unbalanced.\n\nI tried training this model in various ways (and with various architectural changes). The best score I got was 0.95 on the validation set, but unfortunately this only counted as 0.83 on the leaderboard. I'm not sure why that difference is so huge.\n\nHere is the precision/recall computed on the validation set:\n\n                 precision    recall  f1-score   support\n\n            yes       0.97      0.96      0.96       517\n             no       0.94      0.89      0.91       522\n             up       0.94      0.89      0.92       532\n           down       0.96      0.93      0.94       517\n           left       0.94      0.95      0.94       514\n          right       0.96      0.88      0.92       515\n             on       0.96      0.89      0.93       503\n            off       0.89      0.93      0.91       518\n           stop       0.96      0.94      0.95       495\n             go       0.95      0.89      0.92       511\n        silence       0.78      0.95      0.86       600\n        unknown       0.97      0.97      0.97      8489\n    \n    avg / total       0.95      0.95      0.95     14233\n\nInterestingly, precision on \"silence\" is rather low but when you listen to the examples that are misclassified as silence, they actually do sound like silence. So the labels seem to be wrong for those examples.\n\nAnyway, 0.83 on the LB seems the best my model can do at this point. So using 1-d convolutions and nothing else seems to be able to learn this dataset quite well.",
      "votes": null
    },
    {
      "id": "252825",
      "postDate": "12/04/2017 00:04:06",
      "content": "<p>I still haven't put the time in to try 1-d yet, but have put some higher capacity spectrogram based models to the test. As you noticed, I can easily get cross validation above 95% on my validation set which does respect the speaker hashes (so speakers are not overlapping). It does not translate to better on the LB though.</p>\n\n<p>I did however notice something strange with the silence. I assumed my problem was unknown so I started manually thresholding. It made no difference, for the hell of it, I tried the same with silence, can usually get an extra .02 at least by playing with the silence. Let me know if you observe the same with your models... </p>",
      "rawMarkdown": "I still haven't put the time in to try 1-d yet, but have put some higher capacity spectrogram based models to the test. As you noticed, I can easily get cross validation above 95% on my validation set which does respect the speaker hashes (so speakers are not overlapping). It does not translate to better on the LB though.\n\nI did however notice something strange with the silence. I assumed my problem was unknown so I started manually thresholding. It made no difference, for the hell of it, I tried the same with silence, can usually get an extra .02 at least by playing with the silence. Let me know if you observe the same with your models...",
      "votes": null
    },
    {
      "id": "253003",
      "postDate": "12/04/2017 09:48:56",
      "content": "<p>I had a suspicion about silence too, so I replaced all \"silence\" with \"unknown\" in my submission and it scored 9% lower on the LB. Since the baseline score for all silence is also 9%, it doesn't seem that my silence predictions were too far off. </p>\n\n<p>But we don't have a good definition of exactly what silence is. There are very noisy examples in the test set that do not have any useful signal in it. Are these silence or unknown? What about an example with nothing but a short clicking sound, is that considered to be silence?</p>",
      "rawMarkdown": "I had a suspicion about silence too, so I replaced all \"silence\" with \"unknown\" in my submission and it scored 9% lower on the LB. Since the baseline score for all silence is also 9%, it doesn't seem that my silence predictions were too far off. \n\nBut we don't have a good definition of exactly what silence is. There are very noisy examples in the test set that do not have any useful signal in it. Are these silence or unknown? What about an example with nothing but a short clicking sound, is that considered to be silence?",
      "votes": null
    },
    {
      "id": "253044",
      "postDate": "12/04/2017 11:17:20",
      "content": "<p>I have the same confusion about silence.  In test data, there are many wav files consist of all zero, so they are silence certainly. But a sound with only noise should be unknown or silence ? </p>",
      "rawMarkdown": "I have the same confusion about silence.  In test data, there are many wav files consist of all zero, so they are silence certainly. But a sound with only noise should be unknown or silence ?",
      "votes": null
    },
    {
      "id": "253066",
      "postDate": "12/04/2017 12:11:27",
      "content": "<p>In the TensorFlow tutorial code, they treat silence as just background noise (from the background noises folder). That's also how I did it, with some random white noise mixed in.</p>",
      "rawMarkdown": "In the TensorFlow tutorial code, they treat silence as just background noise (from the background noises folder). That's also how I did it, with some random white noise mixed in.",
      "votes": null
    },
    {
      "id": "253116",
      "postDate": "12/04/2017 13:15:50",
      "content": "<p>Yes, in the TensorFlow tutorial code, \"silence\" in train data is zero add mix background while validation data is zero.</p>",
      "rawMarkdown": "Yes, in the TensorFlow tutorial code, \"silence\" in train data is zero add mix background while validation data is zero.",
      "votes": null
    },
    {
      "id": "253331",
      "postDate": "12/04/2017 19:48:08",
      "content": "<p>Do you subsample the input from 16000 to something smaller ?</p>\n\n<p>How better is densenet's performance compared to a simpler feedforward CNN ?</p>",
      "rawMarkdown": "Do you subsample the input from 16000 to something smaller ?\n\nHow better is densenet's performance compared to a simpler feedforward CNN ?",
      "votes": null
    },
    {
      "id": "253335",
      "postDate": "12/04/2017 20:00:25",
      "content": "<p>Yes, it gets subsampled over time using max-pooling. DenseNet does seem to work better than a simpler structure. I chose DenseNet because it's supposed to work well with small-ish datasets.</p>",
      "rawMarkdown": "Yes, it gets subsampled over time using max-pooling. DenseNet does seem to work better than a simpler structure. I chose DenseNet because it's supposed to work well with small-ish datasets.",
      "votes": null
    },
    {
      "id": "253339",
      "postDate": "12/04/2017 20:22:46",
      "content": "<p>Yeah I get that, I am talking about before feeding it to the network.\nThanks for the idea btw. </p>",
      "rawMarkdown": "Yeah I get that, I am talking about before feeding it to the network.\nThanks for the idea btw.",
      "votes": null
    },
    {
      "id": "253342",
      "postDate": "12/04/2017 20:25:07",
      "content": "<p>No, I feed the network the 16000 values for each wave form. I tried using random crops of size 8000 but this didn't seem to make it work any better.</p>",
      "rawMarkdown": "No, I feed the network the 16000 values for each wave form. I tried using random crops of size 8000 but this didn't seem to make it work any better.",
      "votes": null
    },
    {
      "id": "253968",
      "postDate": "12/05/2017 23:35:51",
      "content": "<p>I'm considering both 1d conv, 2d conv, and 1d-2d conv stacking.\nUp to now, 1d conv has approached LB 0.85, 2d conv LB 0.81, 1d-2d conv stack LB 0.81. The 2d conv and 1d-2d conv stacks require further processing.\nI would like to check the limit score first by 1d conv.</p>",
      "rawMarkdown": "I'm considering both 1d conv, 2d conv, and 1d-2d conv stacking.\nUp to now, 1d conv has approached LB 0.85, 2d conv LB 0.81, 1d-2d conv stack LB 0.81. The 2d conv and 1d-2d conv stacks require further processing.\nI would like to check the limit score first by 1d conv.",
      "votes": null
    },
    {
      "id": "254152",
      "postDate": "12/06/2017 10:09:45",
      "content": "<p>I'm curious, for your 1d conv with LB 0.85, what is your local validation score? And what sort of model architecture did you use?</p>",
      "rawMarkdown": "I'm curious, for your 1d conv with LB 0.85, what is your local validation score? And what sort of model architecture did you use?",
      "votes": null
    },
    {
      "id": "254182",
      "postDate": "12/06/2017 12:02:40",
      "content": "<p>My validation score is about 0.915.\nPlease see val_loss, val_acc in the figure below.</p>\n\n<p><img src=\"https://2.bp.blogspot.com/-q9Si5rYq1h8/WifcBHU_45I/AAAAAAAAaZc/f0TSMI1aT4suya2Ti9Al43EabZj8JAA_wCLcBGAs/s1600/1dcnn.PNG\" alt=\"1dcnn_val_loss\" title=\"\">\nI also had a difference of more than 12% between val_acc and LB SCORE like you did at first.\nI thought the valiance between the validation set and the LB set was large.\nSo we augmented the validation set and computed val_loss and val_acc.</p>\n\n<p>And I used a simple architecture like vgg style of the conv layer of filter size 3. I tried resnet style and xception style, but the improvement of val_acc was not great compared to the down speed of learning.</p>",
      "rawMarkdown": "My validation score is about 0.915.\nPlease see val_loss, val_acc in the figure below.\n\n![1dcnn_val_loss][1]\nI also had a difference of more than 12% between val_acc and LB SCORE like you did at first.\nI thought the valiance between the validation set and the LB set was large.\nSo we augmented the validation set and computed val_loss and val_acc.\n\n\nAnd I used a simple architecture like vgg style of the conv layer of filter size 3. I tried resnet style and xception style, but the improvement of val_acc was not great compared to the down speed of learning.\n\n\n  [1]: https://2.bp.blogspot.com/-q9Si5rYq1h8/WifcBHU_45I/AAAAAAAAaZc/f0TSMI1aT4suya2Ti9Al43EabZj8JAA_wCLcBGAs/s1600/1dcnn.PNG",
      "votes": null
    },
    {
      "id": "256312",
      "postDate": "12/11/2017 17:41:49",
      "content": "<p>So for a given input, do you apply your [1x1xF] on the [1x16000x1]? What does your architecture look like? Thanks</p>",
      "rawMarkdown": "So for a given input, do you apply your [1x1xF] on the [1x16000x1]? What does your architecture look like? Thanks",
      "votes": null
    },
    {
      "id": "256431",
      "postDate": "12/11/2017 23:08:54",
      "content": "<p>Its  [1xFx1] on the [1x16000x1] \nlook up convolution1D on Keras</p>",
      "rawMarkdown": "Its  [1xFx1] on the [1x16000x1] \nlook up convolution1D on Keras",
      "votes": null
    },
    {
      "id": "256461",
      "postDate": "12/12/2017 01:28:54",
      "content": "<p>Sure. F being the number of filters. I'm trying to get an idea of what the rest of the network might look like. Any pointers on where I can look for reference on this? How many layers are you using, and what are the data shape through each layer?</p>",
      "rawMarkdown": "Sure. F being the number of filters. I'm trying to get an idea of what the rest of the network might look like. Any pointers on where I can look for reference on this? How many layers are you using, and what are the data shape through each layer?",
      "votes": null
    },
    {
      "id": "256541",
      "postDate": "12/12/2017 07:24:04",
      "content": "<p>F is the size of the filter, I used F=9 like OP.\nIts a sequence of conv+relu+maxpooling layers\nthe input is of shape (16000, 1) (raw audio), applying 16 conv filters with border=\"same\" makes it (16000, 16) then maxpooling of size 4 makes it (4000, 16) and so on .... \nMy architecture does not do very well on the LB so you would probably have to wait for a better answer from  @Human Analog or @은주니(ttagu99)</p>",
      "rawMarkdown": "F is the size of the filter, I used F=9 like OP.\nIts a sequence of conv+relu+maxpooling layers\nthe input is of shape (16000, 1) (raw audio), applying 16 conv filters with border=\"same\" makes it (16000, 16) then maxpooling of size 4 makes it (4000, 16) and so on .... \nMy architecture does not do very well on the LB so you would probably have to wait for a better answer from  @Human Analog or @은주니(ttagu99)",
      "votes": null
    },
    {
      "id": "256644",
      "postDate": "12/12/2017 13:21:23",
      "content": "<p>My model architecture was very simple up to LB 0.83, as shown below.</p>\n\n<p>for i in range(6): <br>\n    x = Conv1D(8*(2 ** i), (3),padding = 'same')(x) <br>\n    x = BatchNormalization()(x) <br>\n    x = Activation('relu')(x) <br>\n    x = MaxPooling1D((2), padding='same')(x)</p>\n\n<p>My model is getting more complicated and deeper. This is not what I like.\nI will check only the limit LB Score(Now LB 0.87) of the 1D model, and explore a more simple model and 2D CNN.</p>",
      "rawMarkdown": "My model architecture was very simple up to LB 0.83, as shown below.\n\nfor i in range(6):    \n    x = Conv1D(8*(2 ** i), (3),padding = 'same')(x)    \n    x = BatchNormalization()(x)    \n    x = Activation('relu')(x)    \n    x = MaxPooling1D((2), padding='same')(x)\n    \n\n\nMy model is getting more complicated and deeper. This is not what I like.\nI will check only the limit LB Score(Now LB 0.87) of the 1D model, and explore a more simple model and 2D CNN.",
      "votes": null
    },
    {
      "id": "256649",
      "postDate": "12/12/2017 13:40:29",
      "content": "<p>@은주니(ttagu99)\nAre you using MFCC for process sounds?\nI only get (98, 13) features when I use MFCC. So making deeper is hard, I guess. What feature do you use?</p>",
      "rawMarkdown": "은주니(ttagu99)\nAre you using MFCC for process sounds?\nI only get (98, 13) features when I use MFCC. So making deeper is hard, I guess. What feature do you use?",
      "votes": null
    },
    {
      "id": "256653",
      "postDate": "12/12/2017 14:01:21",
      "content": "<p>I have not used features like MFCC, Spectogram yet. I would like to check my limit LB score with 1d CNN in end-to-end learning. So My input shape is (16000, 1).</p>",
      "rawMarkdown": "I have not used features like MFCC, Spectogram yet. I would like to check my limit LB score with 1d CNN in end-to-end learning. So My input shape is (16000, 1).",
      "votes": null
    },
    {
      "id": "256724",
      "postDate": "12/12/2017 15:55:45",
      "content": "<p>@은주니(ttagu99) Thanks for your kind explanation. It is very interesting that your model doesn't use any known feature extractor and achieve such a great result. </p>",
      "rawMarkdown": "은주니(ttagu99) Thanks for your kind explanation. It is very interesting that your model doesn't use any known feature extractor and achieve such a great result.",
      "votes": null
    },
    {
      "id": "257065",
      "postDate": "12/13/2017 10:04:17",
      "content": "<p>For me the reasoning was that deep learning is supposed to learn how to extract (the best) features by itself, and so a pure 1-d convolution network that works directly on the wav data should be able to extract any frequency spectrum information and other features without us having to explicitly compute the MFCC, FFT, etc. It looks like this is correct, since my own 1-d convolution network gets ~0.95 on the validation set. And it seems @ttagu99 gets similar results (and an even better LB score than me).</p>",
      "rawMarkdown": "For me the reasoning was that deep learning is supposed to learn how to extract (the best) features by itself, and so a pure 1-d convolution network that works directly on the wav data should be able to extract any frequency spectrum information and other features without us having to explicitly compute the MFCC, FFT, etc. It looks like this is correct, since my own 1-d convolution network gets ~0.95 on the validation set. And it seems @ttagu99 gets similar results (and an even better LB score than me).",
      "votes": null
    },
    {
      "id": "259071",
      "postDate": "12/17/2017 16:37:04",
      "content": "<p>May I ask how you feed learning rate to Tensorboard in Keras? </p>",
      "rawMarkdown": "May I ask how you feed learning rate to Tensorboard in Keras?",
      "votes": null
    },
    {
      "id": "259279",
      "postDate": "12/18/2017 03:53:48",
      "content": "<p>@은주니(ttagu99) have you used dense layers after 6 Conv+BatchNrm+ReLU+MaxPool? Because after 6 MaxPools the size will be reduced by 2**6 only.</p>",
      "rawMarkdown": "은주니(ttagu99) have you used dense layers after 6 Conv+BatchNrm+ReLU+MaxPool? Because after 6 MaxPools the size will be reduced by 2**6 only.",
      "votes": null
    },
    {
      "id": "259288",
      "postDate": "12/18/2017 04:20:29",
      "content": "<p>I added a dense layer to the last layer. As the code below</p>\n\n<pre><code>x_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)\nx_1d = Dropout(0.2)(x_1d)\nx_1d = Dense(12, activation = 'softmax',name='cls_1d')(x_1d)\n</code></pre>",
      "rawMarkdown": "I added a dense layer to the last layer. As the code below\n\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)\n    x_1d = Dropout(0.2)(x_1d)\n    x_1d = Dense(12, activation = 'softmax',name='cls_1d')(x_1d)",
      "votes": null
    },
    {
      "id": "259293",
      "postDate": "12/18/2017 04:29:52",
      "content": "<p>I used tensorboard in Keras like the code below.</p>\n\n<pre><code>callbacks = [EarlyStopping(monitor='val_loss',\n                               patience=7,\n                               verbose=1,\n                               min_delta=0.00001,\n                               mode='min'),\n                 ReduceLROnPlateau(monitor='val_loss',\n                                   factor=0.1,\n                                   patience=4,\n                                   verbose=1,\n                                   epsilon=0.0001,\n                                   mode='min'),\n                 ModelCheckpoint(monitor='val_loss',\n                                 filepath=root_dir + 'weights/' + weight_name,\n                                 save_best_only=True,\n                                 save_weights_only=True,\n                                 mode='min') ,\n                 TQDMCallback(),\n                 TensorBoard(log_dir=root_dir+ weight_name.split('.')[0], histogram_freq=0, write_graph=True, write_images=True) ]\n\n\nhistory = model.fit_generator(generator=train_generator(batch_size),\n                              steps_per_epoch=int((train_df.shape[0]/batch_size)),\n                              epochs=50,\n                              verbose=2,\n                              callbacks=callbacks,\n                              validation_data=valid_generator(batch_size),\n                              validation_steps=int(np.ceil(valid_df.shape[0]/batch_size)))\n</code></pre>",
      "rawMarkdown": "I used tensorboard in Keras like the code below.\n\n    \n    callbacks = [EarlyStopping(monitor='val_loss',\n                                   patience=7,\n                                   verbose=1,\n                                   min_delta=0.00001,\n                                   mode='min'),\n                     ReduceLROnPlateau(monitor='val_loss',\n                                       factor=0.1,\n                                       patience=4,\n                                       verbose=1,\n                                       epsilon=0.0001,\n                                       mode='min'),\n                     ModelCheckpoint(monitor='val_loss',\n                                     filepath=root_dir + 'weights/' + weight_name,\n                                     save_best_only=True,\n                                     save_weights_only=True,\n                                     mode='min') ,\n                     TQDMCallback(),\n                     TensorBoard(log_dir=root_dir+ weight_name.split('.')[0], histogram_freq=0, write_graph=True, write_images=True) ]\n    \n    \n    history = model.fit_generator(generator=train_generator(batch_size),\n                                  steps_per_epoch=int((train_df.shape[0]/batch_size)),\n                                  epochs=50,\n                                  verbose=2,\n                                  callbacks=callbacks,\n                                  validation_data=valid_generator(batch_size),\n                                  validation_steps=int(np.ceil(valid_df.shape[0]/batch_size)))",
      "votes": null
    },
    {
      "id": "259384",
      "postDate": "12/18/2017 09:01:34",
      "content": "<p>Thanks a lot! I compared the difference between my code and yours, I had the TensorBoard callback before the ReduceLROnPlateau callback in the list, and as a result that learning rate is not logged, switch their order and it will show up in my case. </p>",
      "rawMarkdown": "Thanks a lot! I compared the difference between my code and yours, I had the TensorBoard callback before the ReduceLROnPlateau callback in the list, and as a result that learning rate is not logged, switch their order and it will show up in my case.",
      "votes": null
    },
    {
      "id": "259475",
      "postDate": "12/18/2017 13:25:22",
      "content": "<p>How many trainable parameters your model has?  6 conv+pooling layers plus that big dense layer seems require more than 60 million params... </p>",
      "rawMarkdown": "How many trainable parameters your model has?  6 conv+pooling layers plus that big dense layer seems require more than 60 million params...",
      "votes": null
    },
    {
      "id": "259736",
      "postDate": "12/18/2017 23:20:09",
      "content": "<p>After that have you (250, 12) Tensor or Flatten Layer in some place?</p>",
      "rawMarkdown": "After that have you (250, 12) Tensor or Flatten Layer in some place?",
      "votes": null
    },
    {
      "id": "259760",
      "postDate": "12/19/2017 00:18:00",
      "content": "<p>@Ren   I am now using nine pooling layers. I have not tried many tests, so please let me know if you have better results. I use 9 pooling layers, and this model has 20M parameters. ps. I am also testing a model with 200M parameters, but the score is not getting any better.</p>",
      "rawMarkdown": "Ren   I am now using nine pooling layers. I have not tried many tests, so please let me know if you have better results. I use 9 pooling layers, and this model has 20M parameters. ps. I am also testing a model with 200M parameters, but the score is not getting any better.",
      "votes": null
    },
    {
      "id": "259763",
      "postDate": "12/19/2017 00:28:43",
      "content": "<p>I don’t have anything better than yours. My 1d CNN only gets public lb 0.76 but it already has 8 million params. </p>\n\n<p>UPDATE: I now can get public lb 0.84 with 1d CNN now, with a bit fix in the silence examples, add a bit more data augmentation and larger network. </p>",
      "rawMarkdown": "I don’t have anything better than yours. My 1d CNN only gets public lb 0.76 but it already has 8 million params. \n\nUPDATE: I now can get public lb 0.84 with 1d CNN now, with a bit fix in the silence examples, add a bit more data augmentation and larger network.",
      "votes": null
    },
    {
      "id": "259764",
      "postDate": "12/19/2017 00:29:38",
      "content": "<p>@Rafał Jankowski  I used the Global Average Pooling layer and the Global Max Pooling layer. Please see the code below.</p>\n\n<pre><code>x_1d_branch_1 = GlobalAveragePooling1D()(x_1d)\nx_1d_branch_2 = GlobalMaxPool1D()(x_1d)\nx_1d = concatenate([x_1d_branch_1, x_1d_branch_2])\nx_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)\n</code></pre>",
      "rawMarkdown": "Rafał Jankowski  I used the Global Average Pooling layer and the Global Max Pooling layer. Please see the code below.\n\n    x_1d_branch_1 = GlobalAveragePooling1D()(x_1d)\n    x_1d_branch_2 = GlobalMaxPool1D()(x_1d)\n    x_1d = concatenate([x_1d_branch_1, x_1d_branch_2])\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)",
      "votes": null
    },
    {
      "id": "259795",
      "postDate": "12/19/2017 02:01:55",
      "content": "<p>Well, my entries get only 0.6 LB, while maxing out on 0.95 validation score, so it could have been worse :D</p>",
      "rawMarkdown": "Well, my entries get only 0.6 LB, while maxing out on 0.95 validation score, so it could have been worse :D",
      "votes": null
    },
    {
      "id": "260608",
      "postDate": "12/20/2017 14:46:23",
      "content": "<p>Can you tell me how long will it take to train your model~? Thanks.</p>",
      "rawMarkdown": "Can you tell me how long will it take to train your model~? Thanks.",
      "votes": null
    },
    {
      "id": "260763",
      "postDate": "12/20/2017 21:27:46",
      "content": "<p>Thanks a lot, I got ~ 0.02 validation accuracy improvement replacing my Flatten() with GlobalMaxPooling alone.  </p>",
      "rawMarkdown": "Thanks a lot, I got ~ 0.02 validation accuracy improvement replacing my Flatten() with GlobalMaxPooling alone.",
      "votes": null
    },
    {
      "id": "260943",
      "postDate": "12/21/2017 09:13:33",
      "content": "<p>@Zuping Wu  My 1d model takes about 8 minutes per epoch. If I train about 100epoch, it takes about 13 hours. I use 1 titan x pascal(not xp)</p>",
      "rawMarkdown": "Zuping Wu  My 1d model takes about 8 minutes per epoch. If I train about 100epoch, it takes about 13 hours. I use 1 titan x pascal(not xp)",
      "votes": null
    },
    {
      "id": "261085",
      "postDate": "12/21/2017 18:52:52",
      "content": "<p>Maybe you should check this out?\n<a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250#249527\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250#249527</a></p>",
      "rawMarkdown": "Maybe you should check this out?\nhttps://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250#249527",
      "votes": null
    },
    {
      "id": "261732",
      "postDate": "12/23/2017 18:00:17",
      "content": "<p>@은주니(ttagu99) Thank you very much for tips. I got LB .87 just with conv1d. In my experiment, bit larger kernel size, and augmentation seems to help. I'm very surprised that simple architecture can go this far!</p>",
      "rawMarkdown": "은주니(ttagu99) Thank you very much for tips. I got LB .87 just with conv1d. In my experiment, bit larger kernel size, and augmentation seems to help. I'm very surprised that simple architecture can go this far!",
      "votes": null
    },
    {
      "id": "261839",
      "postDate": "12/24/2017 06:33:20",
      "content": "<p>Hi Sukjae, Can you give some hints on your validation strategy especially with unknown and silence classes</p>",
      "rawMarkdown": "Hi Sukjae, Can you give some hints on your validation strategy especially with unknown and silence classes",
      "votes": null
    },
    {
      "id": "261993",
      "postDate": "12/24/2017 20:22:56",
      "content": "<p>@Ravi, \nI just split train/val by person's id. Unknowns are under-sampled to be around 10% of knowns. And for silence, background noise is randomly sampled - I know this doesn't work at all as validation, but I have no idea how to handle silence yet.</p>",
      "rawMarkdown": "Ravi, \nI just split train/val by person's id. Unknowns are under-sampled to be around 10% of knowns. And for silence, background noise is randomly sampled - I know this doesn't work at all as validation, but I have no idea how to handle silence yet.",
      "votes": null
    },
    {
      "id": "261999",
      "postDate": "12/24/2017 20:48:39",
      "content": "<p>Thanks for sharing, I'm struggling to find to a good val split.</p>",
      "rawMarkdown": "Thanks for sharing, I'm struggling to find to a good val split.",
      "votes": null
    },
    {
      "id": "262004",
      "postDate": "12/24/2017 21:38:27",
      "content": "<p>Since the test data have unknown \"unknown\" compared to training data (not sure if they were used to evaluate or not though), if they were, it seems to me that it is not possible to have a split that gives you validation score that is aligned with leaderboard score. I can only get sort of directionally correct validation set up. </p>",
      "rawMarkdown": "Since the test data have unknown \"unknown\" compared to training data (not sure if they were used to evaluate or not though), if they were, it seems to me that it is not possible to have a split that gives you validation score that is aligned with leaderboard score. I can only get sort of directionally correct validation set up.",
      "votes": null
    },
    {
      "id": "262011",
      "postDate": "12/24/2017 22:32:12",
      "content": "<p>By \"directionally correct validation set\", you mean that LB decreases as val-score decreases and LB increases when val-score increases but they are not close right ?</p>",
      "rawMarkdown": "By \"directionally correct validation set\", you mean that LB decreases as val-score decreases and LB increases when val-score increases but they are not close right ?",
      "votes": null
    },
    {
      "id": "262012",
      "postDate": "12/24/2017 22:41:05",
      "content": "<p>yes.</p>",
      "rawMarkdown": "yes.",
      "votes": null
    },
    {
      "id": "262036",
      "postDate": "12/25/2017 02:42:01",
      "content": "<p>What is your architecture?</p>",
      "rawMarkdown": "What is your architecture?",
      "votes": null
    },
    {
      "id": "262425",
      "postDate": "12/26/2017 14:51:27",
      "content": "<p>What was your Conv1D architecture?</p>",
      "rawMarkdown": "What was your Conv1D architecture?",
      "votes": null
    },
    {
      "id": "262426",
      "postDate": "12/26/2017 15:05:35",
      "content": "<p>Not sure who you are asking... @은주니(ttagu99) gives very generous and clear code here about his architecture already. I had a similar architecture.</p>",
      "rawMarkdown": "Not sure who you are asking... @은주니(ttagu99) gives very generous and clear code here about his architecture already. I had a similar architecture.",
      "votes": null
    },
    {
      "id": "263747",
      "postDate": "12/31/2017 18:41:48",
      "content": "<p>Hi, </p>\n\n<p>I am also using mainly 1D conv following VGG architecture but I get \"bad\" LB (~0.65-0.7). ( I also try the model of @은주니(ttagu99) but same result ). With the validation set I get almost 0.9 so I am wondering where I am wrong. When I split the dataset for a word given, I do not put a same ID in the train and validation set at the same time. Does anyone have the same probleme ?</p>",
      "rawMarkdown": "Hi, \n\nI am also using mainly 1D conv following VGG architecture but I get \"bad\" LB (~0.65-0.7). ( I also try the model of @은주니(ttagu99) but same result ). With the validation set I get almost 0.9 so I am wondering where I am wrong. When I split the dataset for a word given, I do not put a same ID in the train and validation set at the same time. Does anyone have the same probleme ?",
      "votes": null
    },
    {
      "id": "263748",
      "postDate": "12/31/2017 18:46:25",
      "content": "<p>Make sure your labels are good</p>",
      "rawMarkdown": "Make sure your labels are good",
      "votes": null
    },
    {
      "id": "263752",
      "postDate": "12/31/2017 19:22:00",
      "content": "<p>That's what I thought too, but I checked the label in my train test (number associated to the word) and compare it to the submission and it does not seem wrong. Maybe I am doing something wrong before feeding the neural network ... ( I use the signal function in scipy and use a zero padding when the audio lasts less than 1 sec )</p>",
      "rawMarkdown": "That's what I thought too, but I checked the label in my train test (number associated to the word) and compare it to the submission and it does not seem wrong. Maybe I am doing something wrong before feeding the neural network ... ( I use the signal function in scipy and use a zero padding when the audio lasts less than 1 sec )",
      "votes": null
    },
    {
      "id": "263755",
      "postDate": "12/31/2017 19:33:12",
      "content": "<p>@Sukjae Cho How did you split your train/val ? Given a word you put different ID in the train or validation OR given an ID you put in the train or the validation set ? (in the first case we can have the same ID in the train and validation set but saying a different word, in the second case not)</p>",
      "rawMarkdown": "Sukjae Cho How did you split your train/val ? Given a word you put different ID in the train or validation OR given an ID you put in the train or the validation set ? (in the first case we can have the same ID in the train and validation set but saying a different word, in the second case not)",
      "votes": null
    },
    {
      "id": "264690",
      "postDate": "01/03/2018 17:32:14",
      "content": "<p>@ShiroK, I split train/val by person's ID(second option). But icybee seems not using this approach according to other post(shocker!) and doing better, so there may be better way.</p>",
      "rawMarkdown": "ShiroK, I split train/val by person's ID(second option). But icybee seems not using this approach according to other post(shocker!) and doing better, so there may be better way.",
      "votes": null
    },
    {
      "id": "264724",
      "postDate": "01/03/2018 19:00:42",
      "content": "<p>@Sukjae Cho thanks ! I tried the two methods I explained before and it does not seems to change a lot. \n@은주니(ttagu99) I tried your architecture but I obtain 'low' LB ~0.7 did you do a preprocess or your input is just the raw data (16000, 1) ? I do not really understand why I do not obtain similar result :/ </p>",
      "rawMarkdown": "Sukjae Cho thanks ! I tried the two methods I explained before and it does not seems to change a lot. \n@은주니(ttagu99) I tried your architecture but I obtain 'low' LB ~0.7 did you do a preprocess or your input is just the raw data (16000, 1) ? I do not really understand why I do not obtain similar result :/",
      "votes": null
    },
    {
      "id": "267557",
      "postDate": "01/11/2018 18:07:32",
      "content": "<p>@Sukjae Cho how big is your kernel size to achieve 87% LB? I tried 3, 9, 17 and 33 and only able to reach 86% LB (after 10 max pool).</p>",
      "rawMarkdown": "Sukjae Cho how big is your kernel size to achieve 87% LB? I tried 3, 9, 17 and 33 and only able to reach 86% LB (after 10 max pool).",
      "votes": null
    },
    {
      "id": "267715",
      "postDate": "01/12/2018 05:26:01",
      "content": "<p>I'm using 9 at first layers. There may be better size, but doesn't matter much - I think the network itself has enough capacity already, so other factors matter more.</p>",
      "rawMarkdown": "I'm using 9 at first layers. There may be better size, but doesn't matter much - I think the network itself has enough capacity already, so other factors matter more.",
      "votes": null
    },
    {
      "id": "269131",
      "postDate": "01/16/2018 09:29:05",
      "content": "<p>Thanks. It seems going deeper helps a lot. I started first with the ttagu99's model and increased the number of the max pool layers and was able to get around 84% LB. After that I made again my network wider (starting with 32 features) and deeper (10 max pools) and got 86% LB. On this model, I tried many kernel sizes like 3, 9, 17 and 33 so on but there was no improvement. </p>\n\n<p>After the reply of Sukjae Cho, I increased my network depth (12 max pools) and got more than 1.5% validation accuracy improvement which should be enough to get 87% LB. </p>",
      "rawMarkdown": "Thanks. It seems going deeper helps a lot. I started first with the ttagu99's model and increased the number of the max pool layers and was able to get around 84% LB. After that I made again my network wider (starting with 32 features) and deeper (10 max pools) and got 86% LB. On this model, I tried many kernel sizes like 3, 9, 17 and 33 so on but there was no improvement. \n\nAfter the reply of Sukjae Cho, I increased my network depth (12 max pools) and got more than 1.5% validation accuracy improvement which should be enough to get 87% LB.",
      "votes": null
    },
    {
      "id": "269148",
      "postDate": "01/16/2018 10:12:52",
      "content": "<p>Thanks for update. Now I understand one of the mysteries :)\nI had both 1D model and 2D model with similar performance(both upper 87%). I experimented both models but couldn't get any improvement. So I concluded that model capacity is enough and focused on augmentation. As I add different kind of augmentations, 2D model performance start to improve, but oddly 1D model performance dropped a lot. I didn't understand why and dropped 1D model, but it looks like my 1D model wasn't large enough to handle heavily augmented samples. Too bad that there are not much time left :(</p>",
      "rawMarkdown": "Thanks for update. Now I understand one of the mysteries :)\nI had both 1D model and 2D model with similar performance(both upper 87%). I experimented both models but couldn't get any improvement. So I concluded that model capacity is enough and focused on augmentation. As I add different kind of augmentations, 2D model performance start to improve, but oddly 1D model performance dropped a lot. I didn't understand why and dropped 1D model, but it looks like my 1D model wasn't large enough to handle heavily augmented samples. Too bad that there are not much time left :(",
      "votes": null
    },
    {
      "id": "269515",
      "postDate": "01/17/2018 00:20:47",
      "content": "<p>Did anyone try WaveNet? It was too slow for me to train before the submission deadline, so I'm curious if anyone else had tried it.</p>",
      "rawMarkdown": "Did anyone try WaveNet? It was too slow for me to train before the submission deadline, so I'm curious if anyone else had tried it.",
      "votes": null
    },
    {
      "id": "270591",
      "postDate": "01/18/2018 15:02:08",
      "content": "<p>Yes, our best WaveNet-like architecture had 2 res blocks with dilation rates from 1 till 2048. We followed the logic of speech-to-text wave-net from <a href=\"https://github.com/buriburisuri/speech-to-text-wavenet\">https://github.com/buriburisuri/speech-to-text-wavenet</a> but without gated activations and with averaging pooling in the end. It achieved 87% on private and 86% on public LB. The model used a lot of RAM due to a large input dimension and took approximately 24 hours to train, but it was the only model that worked for us, so we included it into final submission :)</p>",
      "rawMarkdown": "Yes, our best WaveNet-like architecture had 2 res blocks with dilation rates from 1 till 2048. We followed the logic of speech-to-text wave-net from https://github.com/buriburisuri/speech-to-text-wavenet but without gated activations and with averaging pooling in the end. It achieved 87% on private and 86% on public LB. The model used a lot of RAM due to a large input dimension and took approximately 24 hours to train, but it was the only model that worked for us, so we included it into final submission :)",
      "votes": null
    },
    {
      "id": "270623",
      "postDate": "01/18/2018 16:07:53",
      "content": "<p>Cool. I used dilation rates up to 8192 because that would have a receptive field of 16384, which is greater than the 16000 size of each audio sample.</p>",
      "rawMarkdown": "Cool. I used dilation rates up to 8192 because that would have a receptive field of 16384, which is greater than the 16000 size of each audio sample.",
      "votes": null
    },
    {
      "id": "270627",
      "postDate": "01/18/2018 16:22:25",
      "content": "<p>Yeah I followed the same logic and tried 8192 but it didn't bring any improvements for some reason. </p>",
      "rawMarkdown": "Yeah I followed the same logic and tried 8192 but it didn't bring any improvements for some reason.",
      "votes": null
    },
    {
      "id": "271096",
      "postDate": "01/19/2018 17:32:10",
      "content": "<p>Before moving to PyTorch very late in the competition, I had built a few 1d models in Tensorflow. They were doing 'no worse' than the 2d models with that training scheme.... </p>\n\n<p>Some code at the link below. The 'basic3' was consistently decent, think 'a' was alright too. I experimented with some wavenet style dilation + skip connections but it was a beast, did end up converging to something reasonable </p>\n\n<p><a href=\"https://github.com/rwightman/tensorflow-speech_commands/blob/master/models/conv1d.py\">https://github.com/rwightman/tensorflow-speech_commands/blob/master/models/conv1d.py</a></p>",
      "rawMarkdown": "Before moving to PyTorch very late in the competition, I had built a few 1d models in Tensorflow. They were doing 'no worse' than the 2d models with that training scheme.... \n\nSome code at the link below. The 'basic3' was consistently decent, think 'a' was alright too. I experimented with some wavenet style dilation + skip connections but it was a beast, did end up converging to something reasonable \n\nhttps://github.com/rwightman/tensorflow-speech_commands/blob/master/models/conv1d.py",
      "votes": null
    },
    {
      "id": "280840",
      "postDate": "02/11/2018 07:03:46",
      "content": "<p>I have also tried using 1-D convolutions. Mine did not cross 0.81 in LB. I have used kernels of size 32 with stride of 4.</p>",
      "rawMarkdown": "I have also tried using 1-D convolutions. Mine did not cross 0.81 in LB. I have used kernels of size 32 with stride of 4.",
      "votes": null
    },
    {
      "id": "869360",
      "postDate": "06/01/2020 00:59:30",
      "content": "<p>Dear friend\nI am trying to complete my master's work and I need to include a topic on audio classification through the 1d convolutional network directly from the audio signal. I saw that you argued a lot about the case. Could you help me with some python code that extracts the signal, writes it to a .npy file, then retrieves it and makes the classification using conv1d.\nThank you.\npmacedofilho@gmail.com</p>",
      "rawMarkdown": "Dear friend\nI am trying to complete my master's work and I need to include a topic on audio classification through the 1d convolutional network directly from the audio signal. I saw that you argued a lot about the case. Could you help me with some python code that extracts the signal, writes it to a .npy file, then retrieves it and makes the classification using conv1d.\nThank you.\npmacedofilho@gmail.com",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 248792,
      "author_name": "rwightman",
      "author_url": "",
      "post_date": "11/27/2017 00:53:44",
      "content": "<p>I plan to experiment with 1-d convolutions. I'm also at a similar point, stuck at 0.82 on the LB with local validation scores over 93%. I hit 88-90% validation with the baseline TF model. I built a custom model with more capacity, takes 2x longer to train but hits 93+% validation consistently. Both models score almost exactly the same on the pub LB though.</p>\n\n<p>I'm going to see if I can get 1-d convolutions into a similar range. I expect that stacked dilated convolutions a-la Wavenet (minus the generative aspects) will be required to hit reasonable receptive field size for the samples... </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 248876,
      "author_name": "solomonk",
      "author_url": "",
      "post_date": "11/27/2017 06:53:48",
      "content": "<p>If you are into PyTorch, then I have a very nice notebook here exemplifying the use of a 1d CNN. \n<a href=\"https://github.com/QuantScientist/Deep-Learning-Boot-Camp/blob/master/day02-PyTORCH-and-PyCUDA/PyTorch/55-PyTorch-using-CONV1D-on-one-dimensional-data-CNN.ipynb\">https://github.com/QuantScientist/Deep-Learning-Boot-Camp/blob/master/day02-PyTORCH-and-PyCUDA/PyTorch/55-PyTorch-using-CONV1D-on-one-dimensional-data-CNN.ipynb</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250152,
      "author_name": "gerti004",
      "author_url": "",
      "post_date": "11/29/2017 21:47:11",
      "content": "<p>I started with 1D for the first half of my model. With the intent of condensing the {16000,1] to [250,64] and resize to follow with standard 2D conv image analysis but it kept getting stuck in a trivial result. I'm guessing there was issues with resizing the tensor.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250566,
      "author_name": "timgasser",
      "author_url": "",
      "post_date": "11/30/2017 05:26:23",
      "content": "<p>From a DSP perspective as you say, a 1-D convolution is equivalent to multiplying by the FFT of the convolution in the frequency domain. I think if we train a 1-D convolution on the waveform it should automatically adjust the filter taps to bring out the most pertinent frequencies that help distinguish between the relevant classes. If we train multiple 1-D convolutions in parallel, it should pick out the best selection of frequencies.</p>\n\n<p>If we convert to a spectrogram using the STFT, we effectively fix our set of frequencies with a linear spacing. This isn't great because we perceive audio logarithmically in both frequency and amplitude. The MFCC improves on the linear frequency spacing using the mel scale, which spaces the frequencies so we perceive them as increasing in pitch linearly. It also uses a log transform of the amplitudes. This makes the resulting feature vectors closer to what we perceive in practice. Pretty cool stuff!</p>\n\n<p>One thing I forgot to mention - some letters and sounds are like white noise, containing frequencies across the spectrum.   The 's' in 'yes' is a good example, in <a href=\"https://www.kaggle.com/davids1992/speech-visualization-and-exploration\">DavidS's excellent kernel</a> you can see it covers frequencies from 3.2kHz to 8kHz. This would need a broad filter to cover.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 250650,
      "author_name": "happycube",
      "author_url": "",
      "post_date": "11/30/2017 07:07:56",
      "content": "<p>A pre-programmed/fixed-weight 1D convolution using FIR filters would probably do well for the Raspberry Pi side of the competition...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 252785,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "12/03/2017 21:43:41",
      "content": "<p>In case anyone is interested, I used a DenseNet-type architecture consisting of only 1-d convolutions with kernel size 9 (which seemed to work better than smaller sizes) and a \"swish\" activation function. The model has about 2M parameters. </p>\n\n<p>The training set consisted of the 12 classes in equal proportions. Even though the \"unknown\" class has many more examples that the other classes, I only included 1800 unknown examples in the training set, as well as 1800 silence examples. </p>\n\n<p>However, when loading the examples into a batch for training, when the label is \"unknown\" it randomly samples from <em>all</em> of the unknown examples -- so over time the model does see all of the unknown examples. (When the label is \"silence\", it creates a random white noise example on-the-fly.) Training examples are also augmented on-the-fly using the wavs from the background noise folder.</p>\n\n<p>For the validation set I combined the suggested validation and test sets into one big validation set, just to have more examples. Here, the unknown class contains about 8500 examples versus ~500 for each other class -- so while the training set is balanced the validation set is very unbalanced.</p>\n\n<p>I tried training this model in various ways (and with various architectural changes). The best score I got was 0.95 on the validation set, but unfortunately this only counted as 0.83 on the leaderboard. I'm not sure why that difference is so huge.</p>\n\n<p>Here is the precision/recall computed on the validation set:</p>\n\n<pre><code>             precision    recall  f1-score   support\n\n        yes       0.97      0.96      0.96       517\n         no       0.94      0.89      0.91       522\n         up       0.94      0.89      0.92       532\n       down       0.96      0.93      0.94       517\n       left       0.94      0.95      0.94       514\n      right       0.96      0.88      0.92       515\n         on       0.96      0.89      0.93       503\n        off       0.89      0.93      0.91       518\n       stop       0.96      0.94      0.95       495\n         go       0.95      0.89      0.92       511\n    silence       0.78      0.95      0.86       600\n    unknown       0.97      0.97      0.97      8489\n\navg / total       0.95      0.95      0.95     14233\n</code></pre>\n\n<p>Interestingly, precision on \"silence\" is rather low but when you listen to the examples that are misclassified as silence, they actually do sound like silence. So the labels seem to be wrong for those examples.</p>\n\n<p>Anyway, 0.83 on the LB seems the best my model can do at this point. So using 1-d convolutions and nothing else seems to be able to learn this dataset quite well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 252825,
          "author_name": "rwightman",
          "author_url": "",
          "post_date": "12/04/2017 00:04:06",
          "content": "<p>I still haven't put the time in to try 1-d yet, but have put some higher capacity spectrogram based models to the test. As you noticed, I can easily get cross validation above 95% on my validation set which does respect the speaker hashes (so speakers are not overlapping). It does not translate to better on the LB though.</p>\n\n<p>I did however notice something strange with the silence. I assumed my problem was unknown so I started manually thresholding. It made no difference, for the hell of it, I tried the same with silence, can usually get an extra .02 at least by playing with the silence. Let me know if you observe the same with your models... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253003,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/04/2017 09:48:56",
          "content": "<p>I had a suspicion about silence too, so I replaced all \"silence\" with \"unknown\" in my submission and it scored 9% lower on the LB. Since the baseline score for all silence is also 9%, it doesn't seem that my silence predictions were too far off. </p>\n\n<p>But we don't have a good definition of exactly what silence is. There are very noisy examples in the test set that do not have any useful signal in it. Are these silence or unknown? What about an example with nothing but a short clicking sound, is that considered to be silence?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253044,
          "author_name": "whiteworld",
          "author_url": "",
          "post_date": "12/04/2017 11:17:20",
          "content": "<p>I have the same confusion about silence.  In test data, there are many wav files consist of all zero, so they are silence certainly. But a sound with only noise should be unknown or silence ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253066,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/04/2017 12:11:27",
          "content": "<p>In the TensorFlow tutorial code, they treat silence as just background noise (from the background noises folder). That's also how I did it, with some random white noise mixed in.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253116,
          "author_name": "whiteworld",
          "author_url": "",
          "post_date": "12/04/2017 13:15:50",
          "content": "<p>Yes, in the TensorFlow tutorial code, \"silence\" in train data is zero add mix background while validation data is zero.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259795,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "12/19/2017 02:01:55",
          "content": "<p>Well, my entries get only 0.6 LB, while maxing out on 0.95 validation score, so it could have been worse :D</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261085,
          "author_name": "bsp2020",
          "author_url": "",
          "post_date": "12/21/2017 18:52:52",
          "content": "<p>Maybe you should check this out?\n<a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250#249527\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250#249527</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 253331,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "12/04/2017 19:48:08",
      "content": "<p>Do you subsample the input from 16000 to something smaller ?</p>\n\n<p>How better is densenet's performance compared to a simpler feedforward CNN ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 253335,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/04/2017 20:00:25",
          "content": "<p>Yes, it gets subsampled over time using max-pooling. DenseNet does seem to work better than a simpler structure. I chose DenseNet because it's supposed to work well with small-ish datasets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253339,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "12/04/2017 20:22:46",
          "content": "<p>Yeah I get that, I am talking about before feeding it to the network.\nThanks for the idea btw. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 253342,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/04/2017 20:25:07",
          "content": "<p>No, I feed the network the 16000 values for each wave form. I tried using random crops of size 8000 but this didn't seem to make it work any better.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 253968,
      "author_name": "ttagu99",
      "author_url": "",
      "post_date": "12/05/2017 23:35:51",
      "content": "<p>I'm considering both 1d conv, 2d conv, and 1d-2d conv stacking.\nUp to now, 1d conv has approached LB 0.85, 2d conv LB 0.81, 1d-2d conv stack LB 0.81. The 2d conv and 1d-2d conv stacks require further processing.\nI would like to check the limit score first by 1d conv.</p>",
      "votes": null,
      "replies": [
        {
          "id": 254152,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/06/2017 10:09:45",
          "content": "<p>I'm curious, for your 1d conv with LB 0.85, what is your local validation score? And what sort of model architecture did you use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 254182,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/06/2017 12:02:40",
          "content": "<p>My validation score is about 0.915.\nPlease see val_loss, val_acc in the figure below.</p>\n\n<p><img src=\"https://2.bp.blogspot.com/-q9Si5rYq1h8/WifcBHU_45I/AAAAAAAAaZc/f0TSMI1aT4suya2Ti9Al43EabZj8JAA_wCLcBGAs/s1600/1dcnn.PNG\" alt=\"1dcnn_val_loss\" title=\"\">\nI also had a difference of more than 12% between val_acc and LB SCORE like you did at first.\nI thought the valiance between the validation set and the LB set was large.\nSo we augmented the validation set and computed val_loss and val_acc.</p>\n\n<p>And I used a simple architecture like vgg style of the conv layer of filter size 3. I tried resnet style and xception style, but the improvement of val_acc was not great compared to the down speed of learning.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259071,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/17/2017 16:37:04",
          "content": "<p>May I ask how you feed learning rate to Tensorboard in Keras? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259293,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/18/2017 04:29:52",
          "content": "<p>I used tensorboard in Keras like the code below.</p>\n\n<pre><code>callbacks = [EarlyStopping(monitor='val_loss',\n                               patience=7,\n                               verbose=1,\n                               min_delta=0.00001,\n                               mode='min'),\n                 ReduceLROnPlateau(monitor='val_loss',\n                                   factor=0.1,\n                                   patience=4,\n                                   verbose=1,\n                                   epsilon=0.0001,\n                                   mode='min'),\n                 ModelCheckpoint(monitor='val_loss',\n                                 filepath=root_dir + 'weights/' + weight_name,\n                                 save_best_only=True,\n                                 save_weights_only=True,\n                                 mode='min') ,\n                 TQDMCallback(),\n                 TensorBoard(log_dir=root_dir+ weight_name.split('.')[0], histogram_freq=0, write_graph=True, write_images=True) ]\n\n\nhistory = model.fit_generator(generator=train_generator(batch_size),\n                              steps_per_epoch=int((train_df.shape[0]/batch_size)),\n                              epochs=50,\n                              verbose=2,\n                              callbacks=callbacks,\n                              validation_data=valid_generator(batch_size),\n                              validation_steps=int(np.ceil(valid_df.shape[0]/batch_size)))\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259384,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/18/2017 09:01:34",
          "content": "<p>Thanks a lot! I compared the difference between my code and yours, I had the TensorBoard callback before the ReduceLROnPlateau callback in the list, and as a result that learning rate is not logged, switch their order and it will show up in my case. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262425,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "12/26/2017 14:51:27",
          "content": "<p>What was your Conv1D architecture?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262426,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/26/2017 15:05:35",
          "content": "<p>Not sure who you are asking... @은주니(ttagu99) gives very generous and clear code here about his architecture already. I had a similar architecture.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 256312,
      "author_name": "formigone",
      "author_url": "",
      "post_date": "12/11/2017 17:41:49",
      "content": "<p>So for a given input, do you apply your [1x1xF] on the [1x16000x1]? What does your architecture look like? Thanks</p>",
      "votes": null,
      "replies": [
        {
          "id": 256431,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "12/11/2017 23:08:54",
          "content": "<p>Its  [1xFx1] on the [1x16000x1] \nlook up convolution1D on Keras</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256461,
          "author_name": "formigone",
          "author_url": "",
          "post_date": "12/12/2017 01:28:54",
          "content": "<p>Sure. F being the number of filters. I'm trying to get an idea of what the rest of the network might look like. Any pointers on where I can look for reference on this? How many layers are you using, and what are the data shape through each layer?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256541,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "12/12/2017 07:24:04",
          "content": "<p>F is the size of the filter, I used F=9 like OP.\nIts a sequence of conv+relu+maxpooling layers\nthe input is of shape (16000, 1) (raw audio), applying 16 conv filters with border=\"same\" makes it (16000, 16) then maxpooling of size 4 makes it (4000, 16) and so on .... \nMy architecture does not do very well on the LB so you would probably have to wait for a better answer from  @Human Analog or @은주니(ttagu99)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256644,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/12/2017 13:21:23",
          "content": "<p>My model architecture was very simple up to LB 0.83, as shown below.</p>\n\n<p>for i in range(6): <br>\n    x = Conv1D(8*(2 ** i), (3),padding = 'same')(x) <br>\n    x = BatchNormalization()(x) <br>\n    x = Activation('relu')(x) <br>\n    x = MaxPooling1D((2), padding='same')(x)</p>\n\n<p>My model is getting more complicated and deeper. This is not what I like.\nI will check only the limit LB Score(Now LB 0.87) of the 1D model, and explore a more simple model and 2D CNN.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256649,
          "author_name": "ildoonet",
          "author_url": "",
          "post_date": "12/12/2017 13:40:29",
          "content": "<p>@은주니(ttagu99)\nAre you using MFCC for process sounds?\nI only get (98, 13) features when I use MFCC. So making deeper is hard, I guess. What feature do you use?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256653,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/12/2017 14:01:21",
          "content": "<p>I have not used features like MFCC, Spectogram yet. I would like to check my limit LB score with 1d CNN in end-to-end learning. So My input shape is (16000, 1).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 256724,
          "author_name": "ildoonet",
          "author_url": "",
          "post_date": "12/12/2017 15:55:45",
          "content": "<p>@은주니(ttagu99) Thanks for your kind explanation. It is very interesting that your model doesn't use any known feature extractor and achieve such a great result. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 257065,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "12/13/2017 10:04:17",
          "content": "<p>For me the reasoning was that deep learning is supposed to learn how to extract (the best) features by itself, and so a pure 1-d convolution network that works directly on the wav data should be able to extract any frequency spectrum information and other features without us having to explicitly compute the MFCC, FFT, etc. It looks like this is correct, since my own 1-d convolution network gets ~0.95 on the validation set. And it seems @ttagu99 gets similar results (and an even better LB score than me).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259279,
          "author_name": "fizzbuzz",
          "author_url": "",
          "post_date": "12/18/2017 03:53:48",
          "content": "<p>@은주니(ttagu99) have you used dense layers after 6 Conv+BatchNrm+ReLU+MaxPool? Because after 6 MaxPools the size will be reduced by 2**6 only.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259288,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/18/2017 04:20:29",
          "content": "<p>I added a dense layer to the last layer. As the code below</p>\n\n<pre><code>x_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)\nx_1d = Dropout(0.2)(x_1d)\nx_1d = Dense(12, activation = 'softmax',name='cls_1d')(x_1d)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259475,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/18/2017 13:25:22",
          "content": "<p>How many trainable parameters your model has?  6 conv+pooling layers plus that big dense layer seems require more than 60 million params... </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259736,
          "author_name": "raalsky",
          "author_url": "",
          "post_date": "12/18/2017 23:20:09",
          "content": "<p>After that have you (250, 12) Tensor or Flatten Layer in some place?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259760,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/19/2017 00:18:00",
          "content": "<p>@Ren   I am now using nine pooling layers. I have not tried many tests, so please let me know if you have better results. I use 9 pooling layers, and this model has 20M parameters. ps. I am also testing a model with 200M parameters, but the score is not getting any better.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259763,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/19/2017 00:28:43",
          "content": "<p>I don’t have anything better than yours. My 1d CNN only gets public lb 0.76 but it already has 8 million params. </p>\n\n<p>UPDATE: I now can get public lb 0.84 with 1d CNN now, with a bit fix in the silence examples, add a bit more data augmentation and larger network. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 259764,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/19/2017 00:29:38",
          "content": "<p>@Rafał Jankowski  I used the Global Average Pooling layer and the Global Max Pooling layer. Please see the code below.</p>\n\n<pre><code>x_1d_branch_1 = GlobalAveragePooling1D()(x_1d)\nx_1d_branch_2 = GlobalMaxPool1D()(x_1d)\nx_1d = concatenate([x_1d_branch_1, x_1d_branch_2])\nx_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260608,
          "author_name": "wuzuping",
          "author_url": "",
          "post_date": "12/20/2017 14:46:23",
          "content": "<p>Can you tell me how long will it take to train your model~? Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260763,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/20/2017 21:27:46",
          "content": "<p>Thanks a lot, I got ~ 0.02 validation accuracy improvement replacing my Flatten() with GlobalMaxPooling alone.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 260943,
          "author_name": "ttagu99",
          "author_url": "",
          "post_date": "12/21/2017 09:13:33",
          "content": "<p>@Zuping Wu  My 1d model takes about 8 minutes per epoch. If I train about 100epoch, it takes about 13 hours. I use 1 titan x pascal(not xp)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261732,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "12/23/2017 18:00:17",
          "content": "<p>@은주니(ttagu99) Thank you very much for tips. I got LB .87 just with conv1d. In my experiment, bit larger kernel size, and augmentation seems to help. I'm very surprised that simple architecture can go this far!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261839,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "12/24/2017 06:33:20",
          "content": "<p>Hi Sukjae, Can you give some hints on your validation strategy especially with unknown and silence classes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261993,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "12/24/2017 20:22:56",
          "content": "<p>@Ravi, \nI just split train/val by person's id. Unknowns are under-sampled to be around 10% of knowns. And for silence, background noise is randomly sampled - I know this doesn't work at all as validation, but I have no idea how to handle silence yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261999,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "12/24/2017 20:48:39",
          "content": "<p>Thanks for sharing, I'm struggling to find to a good val split.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262004,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/24/2017 21:38:27",
          "content": "<p>Since the test data have unknown \"unknown\" compared to training data (not sure if they were used to evaluate or not though), if they were, it seems to me that it is not possible to have a split that gives you validation score that is aligned with leaderboard score. I can only get sort of directionally correct validation set up. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262011,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "12/24/2017 22:32:12",
          "content": "<p>By \"directionally correct validation set\", you mean that LB decreases as val-score decreases and LB increases when val-score increases but they are not close right ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262012,
          "author_name": "ryanzhang",
          "author_url": "",
          "post_date": "12/24/2017 22:41:05",
          "content": "<p>yes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 262036,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "12/25/2017 02:42:01",
          "content": "<p>What is your architecture?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263755,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "12/31/2017 19:33:12",
          "content": "<p>@Sukjae Cho How did you split your train/val ? Given a word you put different ID in the train or validation OR given an ID you put in the train or the validation set ? (in the first case we can have the same ID in the train and validation set but saying a different word, in the second case not)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264690,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "01/03/2018 17:32:14",
          "content": "<p>@ShiroK, I split train/val by person's ID(second option). But icybee seems not using this approach according to other post(shocker!) and doing better, so there may be better way.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264724,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "01/03/2018 19:00:42",
          "content": "<p>@Sukjae Cho thanks ! I tried the two methods I explained before and it does not seems to change a lot. \n@은주니(ttagu99) I tried your architecture but I obtain 'low' LB ~0.7 did you do a preprocess or your input is just the raw data (16000, 1) ? I do not really understand why I do not obtain similar result :/ </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 263747,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "12/31/2017 18:41:48",
      "content": "<p>Hi, </p>\n\n<p>I am also using mainly 1D conv following VGG architecture but I get \"bad\" LB (~0.65-0.7). ( I also try the model of @은주니(ttagu99) but same result ). With the validation set I get almost 0.9 so I am wondering where I am wrong. When I split the dataset for a word given, I do not put a same ID in the train and validation set at the same time. Does anyone have the same probleme ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 263748,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "12/31/2017 18:46:25",
          "content": "<p>Make sure your labels are good</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263752,
          "author_name": "ludovick",
          "author_url": "",
          "post_date": "12/31/2017 19:22:00",
          "content": "<p>That's what I thought too, but I checked the label in my train test (number associated to the word) and compare it to the submission and it does not seem wrong. Maybe I am doing something wrong before feeding the neural network ... ( I use the signal function in scipy and use a zero padding when the audio lasts less than 1 sec )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267557,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "01/11/2018 18:07:32",
      "content": "<p>@Sukjae Cho how big is your kernel size to achieve 87% LB? I tried 3, 9, 17 and 33 and only able to reach 86% LB (after 10 max pool).</p>",
      "votes": null,
      "replies": [
        {
          "id": 267715,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "01/12/2018 05:26:01",
          "content": "<p>I'm using 9 at first layers. There may be better size, but doesn't matter much - I think the network itself has enough capacity already, so other factors matter more.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269131,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "01/16/2018 09:29:05",
          "content": "<p>Thanks. It seems going deeper helps a lot. I started first with the ttagu99's model and increased the number of the max pool layers and was able to get around 84% LB. After that I made again my network wider (starting with 32 features) and deeper (10 max pools) and got 86% LB. On this model, I tried many kernel sizes like 3, 9, 17 and 33 so on but there was no improvement. </p>\n\n<p>After the reply of Sukjae Cho, I increased my network depth (12 max pools) and got more than 1.5% validation accuracy improvement which should be enough to get 87% LB. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 269148,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "01/16/2018 10:12:52",
          "content": "<p>Thanks for update. Now I understand one of the mysteries :)\nI had both 1D model and 2D model with similar performance(both upper 87%). I experimented both models but couldn't get any improvement. So I concluded that model capacity is enough and focused on augmentation. As I add different kind of augmentations, 2D model performance start to improve, but oddly 1D model performance dropped a lot. I didn't understand why and dropped 1D model, but it looks like my 1D model wasn't large enough to handle heavily augmented samples. Too bad that there are not much time left :(</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269515,
      "author_name": "agent007",
      "author_url": "",
      "post_date": "01/17/2018 00:20:47",
      "content": "<p>Did anyone try WaveNet? It was too slow for me to train before the submission deadline, so I'm curious if anyone else had tried it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 270591,
          "author_name": "isklyar",
          "author_url": "",
          "post_date": "01/18/2018 15:02:08",
          "content": "<p>Yes, our best WaveNet-like architecture had 2 res blocks with dilation rates from 1 till 2048. We followed the logic of speech-to-text wave-net from <a href=\"https://github.com/buriburisuri/speech-to-text-wavenet\">https://github.com/buriburisuri/speech-to-text-wavenet</a> but without gated activations and with averaging pooling in the end. It achieved 87% on private and 86% on public LB. The model used a lot of RAM due to a large input dimension and took approximately 24 hours to train, but it was the only model that worked for us, so we included it into final submission :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270623,
          "author_name": "agent007",
          "author_url": "",
          "post_date": "01/18/2018 16:07:53",
          "content": "<p>Cool. I used dilation rates up to 8192 because that would have a receptive field of 16384, which is greater than the 16000 size of each audio sample.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 270627,
          "author_name": "isklyar",
          "author_url": "",
          "post_date": "01/18/2018 16:22:25",
          "content": "<p>Yeah I followed the same logic and tried 8192 but it didn't bring any improvements for some reason. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 271096,
      "author_name": "rwightman",
      "author_url": "",
      "post_date": "01/19/2018 17:32:10",
      "content": "<p>Before moving to PyTorch very late in the competition, I had built a few 1d models in Tensorflow. They were doing 'no worse' than the 2d models with that training scheme.... </p>\n\n<p>Some code at the link below. The 'basic3' was consistently decent, think 'a' was alright too. I experimented with some wavenet style dilation + skip connections but it was a beast, did end up converging to something reasonable </p>\n\n<p><a href=\"https://github.com/rwightman/tensorflow-speech_commands/blob/master/models/conv1d.py\">https://github.com/rwightman/tensorflow-speech_commands/blob/master/models/conv1d.py</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 280840,
      "author_name": "hoticevijay",
      "author_url": "",
      "post_date": "02/11/2018 07:03:46",
      "content": "<p>I have also tried using 1-D convolutions. Mine did not cross 0.81 in LB. I have used kernels of size 32 with stride of 4.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 869360,
      "author_name": "pmacedofilho",
      "author_url": "",
      "post_date": "06/01/2020 00:59:30",
      "content": "<p>Dear friend\nI am trying to complete my master's work and I need to include a topic on audio classification through the 1d convolutional network directly from the audio signal. I saw that you argued a lot about the case. Could you help me with some python code that extracts the signal, writes it to a .npy file, then retrieves it and makes the classification using conv1d.\nThank you.\npmacedofilho@gmail.com</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "248593": "It looks like the approach taken in most of the published kernels is to first convert the wav files to spectrograms and then using a convolutional neural network on these, essentially turning the sound classification problem into an image classification problem. \n\nI was just curious if anyone (besides me) has used 1-dimensional convolutions on the wav data directly? This basically only looks at the temporal structure of the data, not so much at the frequencies (although convolution in the time domain is the same as multiplication in the frequency domain, so these two things are obviously related somehow).\n\nMy models are maxing out at 0.9 validation score, which translates to \"only\" 0.82 on the public LB. So the approach seems to work in theory, but it also seems like it still could be doing better.",
    "248792": "I plan to experiment with 1-d convolutions. I'm also at a similar point, stuck at 0.82 on the LB with local validation scores over 93%. I hit 88-90% validation with the baseline TF model. I built a custom model with more capacity, takes 2x longer to train but hits 93+% validation consistently. Both models score almost exactly the same on the pub LB though.\n\nI'm going to see if I can get 1-d convolutions into a similar range. I expect that stacked dilated convolutions a-la Wavenet (minus the generative aspects) will be required to hit reasonable receptive field size for the samples...",
    "248876": "If you are into PyTorch, then I have a very nice notebook here exemplifying the use of a 1d CNN. \nhttps://github.com/QuantScientist/Deep-Learning-Boot-Camp/blob/master/day02-PyTORCH-and-PyCUDA/PyTorch/55-PyTorch-using-CONV1D-on-one-dimensional-data-CNN.ipynb",
    "250152": "I started with 1D for the first half of my model. With the intent of condensing the {16000,1] to [250,64] and resize to follow with standard 2D conv image analysis but it kept getting stuck in a trivial result. I'm guessing there was issues with resizing the tensor.",
    "250566": "From a DSP perspective as you say, a 1-D convolution is equivalent to multiplying by the FFT of the convolution in the frequency domain. I think if we train a 1-D convolution on the waveform it should automatically adjust the filter taps to bring out the most pertinent frequencies that help distinguish between the relevant classes. If we train multiple 1-D convolutions in parallel, it should pick out the best selection of frequencies.\n\nIf we convert to a spectrogram using the STFT, we effectively fix our set of frequencies with a linear spacing. This isn't great because we perceive audio logarithmically in both frequency and amplitude. The MFCC improves on the linear frequency spacing using the mel scale, which spaces the frequencies so we perceive them as increasing in pitch linearly. It also uses a log transform of the amplitudes. This makes the resulting feature vectors closer to what we perceive in practice. Pretty cool stuff!\n\nOne thing I forgot to mention - some letters and sounds are like white noise, containing frequencies across the spectrum.   The 's' in 'yes' is a good example, in [DavidS's excellent kernel](https://www.kaggle.com/davids1992/speech-visualization-and-exploration) you can see it covers frequencies from 3.2kHz to 8kHz. This would need a broad filter to cover.",
    "250650": "A pre-programmed/fixed-weight 1D convolution using FIR filters would probably do well for the Raspberry Pi side of the competition...",
    "252785": "In case anyone is interested, I used a DenseNet-type architecture consisting of only 1-d convolutions with kernel size 9 (which seemed to work better than smaller sizes) and a \"swish\" activation function. The model has about 2M parameters. \n\nThe training set consisted of the 12 classes in equal proportions. Even though the \"unknown\" class has many more examples that the other classes, I only included 1800 unknown examples in the training set, as well as 1800 silence examples. \n\nHowever, when loading the examples into a batch for training, when the label is \"unknown\" it randomly samples from *all* of the unknown examples -- so over time the model does see all of the unknown examples. (When the label is \"silence\", it creates a random white noise example on-the-fly.) Training examples are also augmented on-the-fly using the wavs from the background noise folder.\n\nFor the validation set I combined the suggested validation and test sets into one big validation set, just to have more examples. Here, the unknown class contains about 8500 examples versus ~500 for each other class -- so while the training set is balanced the validation set is very unbalanced.\n\nI tried training this model in various ways (and with various architectural changes). The best score I got was 0.95 on the validation set, but unfortunately this only counted as 0.83 on the leaderboard. I'm not sure why that difference is so huge.\n\nHere is the precision/recall computed on the validation set:\n\n                 precision    recall  f1-score   support\n\n            yes       0.97      0.96      0.96       517\n             no       0.94      0.89      0.91       522\n             up       0.94      0.89      0.92       532\n           down       0.96      0.93      0.94       517\n           left       0.94      0.95      0.94       514\n          right       0.96      0.88      0.92       515\n             on       0.96      0.89      0.93       503\n            off       0.89      0.93      0.91       518\n           stop       0.96      0.94      0.95       495\n             go       0.95      0.89      0.92       511\n        silence       0.78      0.95      0.86       600\n        unknown       0.97      0.97      0.97      8489\n    \n    avg / total       0.95      0.95      0.95     14233\n\nInterestingly, precision on \"silence\" is rather low but when you listen to the examples that are misclassified as silence, they actually do sound like silence. So the labels seem to be wrong for those examples.\n\nAnyway, 0.83 on the LB seems the best my model can do at this point. So using 1-d convolutions and nothing else seems to be able to learn this dataset quite well.",
    "252825": "I still haven't put the time in to try 1-d yet, but have put some higher capacity spectrogram based models to the test. As you noticed, I can easily get cross validation above 95% on my validation set which does respect the speaker hashes (so speakers are not overlapping). It does not translate to better on the LB though.\n\nI did however notice something strange with the silence. I assumed my problem was unknown so I started manually thresholding. It made no difference, for the hell of it, I tried the same with silence, can usually get an extra .02 at least by playing with the silence. Let me know if you observe the same with your models...",
    "253003": "I had a suspicion about silence too, so I replaced all \"silence\" with \"unknown\" in my submission and it scored 9% lower on the LB. Since the baseline score for all silence is also 9%, it doesn't seem that my silence predictions were too far off. \n\nBut we don't have a good definition of exactly what silence is. There are very noisy examples in the test set that do not have any useful signal in it. Are these silence or unknown? What about an example with nothing but a short clicking sound, is that considered to be silence?",
    "253044": "I have the same confusion about silence.  In test data, there are many wav files consist of all zero, so they are silence certainly. But a sound with only noise should be unknown or silence ?",
    "253066": "In the TensorFlow tutorial code, they treat silence as just background noise (from the background noises folder). That's also how I did it, with some random white noise mixed in.",
    "253116": "Yes, in the TensorFlow tutorial code, \"silence\" in train data is zero add mix background while validation data is zero.",
    "253331": "Do you subsample the input from 16000 to something smaller ?\n\nHow better is densenet's performance compared to a simpler feedforward CNN ?",
    "253335": "Yes, it gets subsampled over time using max-pooling. DenseNet does seem to work better than a simpler structure. I chose DenseNet because it's supposed to work well with small-ish datasets.",
    "253339": "Yeah I get that, I am talking about before feeding it to the network.\nThanks for the idea btw.",
    "253342": "No, I feed the network the 16000 values for each wave form. I tried using random crops of size 8000 but this didn't seem to make it work any better.",
    "253968": "I'm considering both 1d conv, 2d conv, and 1d-2d conv stacking.\nUp to now, 1d conv has approached LB 0.85, 2d conv LB 0.81, 1d-2d conv stack LB 0.81. The 2d conv and 1d-2d conv stacks require further processing.\nI would like to check the limit score first by 1d conv.",
    "254152": "I'm curious, for your 1d conv with LB 0.85, what is your local validation score? And what sort of model architecture did you use?",
    "254182": "My validation score is about 0.915.\nPlease see val_loss, val_acc in the figure below.\n\n![1dcnn_val_loss][1]\nI also had a difference of more than 12% between val_acc and LB SCORE like you did at first.\nI thought the valiance between the validation set and the LB set was large.\nSo we augmented the validation set and computed val_loss and val_acc.\n\n\nAnd I used a simple architecture like vgg style of the conv layer of filter size 3. I tried resnet style and xception style, but the improvement of val_acc was not great compared to the down speed of learning.\n\n\n  [1]: https://2.bp.blogspot.com/-q9Si5rYq1h8/WifcBHU_45I/AAAAAAAAaZc/f0TSMI1aT4suya2Ti9Al43EabZj8JAA_wCLcBGAs/s1600/1dcnn.PNG",
    "256312": "So for a given input, do you apply your [1x1xF] on the [1x16000x1]? What does your architecture look like? Thanks",
    "256431": "Its  [1xFx1] on the [1x16000x1] \nlook up convolution1D on Keras",
    "256461": "Sure. F being the number of filters. I'm trying to get an idea of what the rest of the network might look like. Any pointers on where I can look for reference on this? How many layers are you using, and what are the data shape through each layer?",
    "256541": "F is the size of the filter, I used F=9 like OP.\nIts a sequence of conv+relu+maxpooling layers\nthe input is of shape (16000, 1) (raw audio), applying 16 conv filters with border=\"same\" makes it (16000, 16) then maxpooling of size 4 makes it (4000, 16) and so on .... \nMy architecture does not do very well on the LB so you would probably have to wait for a better answer from  @Human Analog or @은주니(ttagu99)",
    "256644": "My model architecture was very simple up to LB 0.83, as shown below.\n\nfor i in range(6):    \n    x = Conv1D(8*(2 ** i), (3),padding = 'same')(x)    \n    x = BatchNormalization()(x)    \n    x = Activation('relu')(x)    \n    x = MaxPooling1D((2), padding='same')(x)\n    \n\n\nMy model is getting more complicated and deeper. This is not what I like.\nI will check only the limit LB Score(Now LB 0.87) of the 1D model, and explore a more simple model and 2D CNN.",
    "256649": "은주니(ttagu99)\nAre you using MFCC for process sounds?\nI only get (98, 13) features when I use MFCC. So making deeper is hard, I guess. What feature do you use?",
    "256653": "I have not used features like MFCC, Spectogram yet. I would like to check my limit LB score with 1d CNN in end-to-end learning. So My input shape is (16000, 1).",
    "256724": "은주니(ttagu99) Thanks for your kind explanation. It is very interesting that your model doesn't use any known feature extractor and achieve such a great result.",
    "257065": "For me the reasoning was that deep learning is supposed to learn how to extract (the best) features by itself, and so a pure 1-d convolution network that works directly on the wav data should be able to extract any frequency spectrum information and other features without us having to explicitly compute the MFCC, FFT, etc. It looks like this is correct, since my own 1-d convolution network gets ~0.95 on the validation set. And it seems @ttagu99 gets similar results (and an even better LB score than me).",
    "259071": "May I ask how you feed learning rate to Tensorboard in Keras?",
    "259279": "은주니(ttagu99) have you used dense layers after 6 Conv+BatchNrm+ReLU+MaxPool? Because after 6 MaxPools the size will be reduced by 2**6 only.",
    "259288": "I added a dense layer to the last layer. As the code below\n\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)\n    x_1d = Dropout(0.2)(x_1d)\n    x_1d = Dense(12, activation = 'softmax',name='cls_1d')(x_1d)",
    "259293": "I used tensorboard in Keras like the code below.\n\n    \n    callbacks = [EarlyStopping(monitor='val_loss',\n                                   patience=7,\n                                   verbose=1,\n                                   min_delta=0.00001,\n                                   mode='min'),\n                     ReduceLROnPlateau(monitor='val_loss',\n                                       factor=0.1,\n                                       patience=4,\n                                       verbose=1,\n                                       epsilon=0.0001,\n                                       mode='min'),\n                     ModelCheckpoint(monitor='val_loss',\n                                     filepath=root_dir + 'weights/' + weight_name,\n                                     save_best_only=True,\n                                     save_weights_only=True,\n                                     mode='min') ,\n                     TQDMCallback(),\n                     TensorBoard(log_dir=root_dir+ weight_name.split('.')[0], histogram_freq=0, write_graph=True, write_images=True) ]\n    \n    \n    history = model.fit_generator(generator=train_generator(batch_size),\n                                  steps_per_epoch=int((train_df.shape[0]/batch_size)),\n                                  epochs=50,\n                                  verbose=2,\n                                  callbacks=callbacks,\n                                  validation_data=valid_generator(batch_size),\n                                  validation_steps=int(np.ceil(valid_df.shape[0]/batch_size)))",
    "259384": "Thanks a lot! I compared the difference between my code and yours, I had the TensorBoard callback before the ReduceLROnPlateau callback in the list, and as a result that learning rate is not logged, switch their order and it will show up in my case.",
    "259475": "How many trainable parameters your model has?  6 conv+pooling layers plus that big dense layer seems require more than 60 million params...",
    "259736": "After that have you (250, 12) Tensor or Flatten Layer in some place?",
    "259760": "Ren   I am now using nine pooling layers. I have not tried many tests, so please let me know if you have better results. I use 9 pooling layers, and this model has 20M parameters. ps. I am also testing a model with 200M parameters, but the score is not getting any better.",
    "259763": "I don’t have anything better than yours. My 1d CNN only gets public lb 0.76 but it already has 8 million params. \n\nUPDATE: I now can get public lb 0.84 with 1d CNN now, with a bit fix in the silence examples, add a bit more data augmentation and larger network.",
    "259764": "Rafał Jankowski  I used the Global Average Pooling layer and the Global Max Pooling layer. Please see the code below.\n\n    x_1d_branch_1 = GlobalAveragePooling1D()(x_1d)\n    x_1d_branch_2 = GlobalMaxPool1D()(x_1d)\n    x_1d = concatenate([x_1d_branch_1, x_1d_branch_2])\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024')(x_1d)",
    "259795": "Well, my entries get only 0.6 LB, while maxing out on 0.95 validation score, so it could have been worse :D",
    "260608": "Can you tell me how long will it take to train your model~? Thanks.",
    "260763": "Thanks a lot, I got ~ 0.02 validation accuracy improvement replacing my Flatten() with GlobalMaxPooling alone.",
    "260943": "Zuping Wu  My 1d model takes about 8 minutes per epoch. If I train about 100epoch, it takes about 13 hours. I use 1 titan x pascal(not xp)",
    "261085": "Maybe you should check this out?\nhttps://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44250#249527",
    "261732": "은주니(ttagu99) Thank you very much for tips. I got LB .87 just with conv1d. In my experiment, bit larger kernel size, and augmentation seems to help. I'm very surprised that simple architecture can go this far!",
    "261839": "Hi Sukjae, Can you give some hints on your validation strategy especially with unknown and silence classes",
    "261993": "Ravi, \nI just split train/val by person's id. Unknowns are under-sampled to be around 10% of knowns. And for silence, background noise is randomly sampled - I know this doesn't work at all as validation, but I have no idea how to handle silence yet.",
    "261999": "Thanks for sharing, I'm struggling to find to a good val split.",
    "262004": "Since the test data have unknown \"unknown\" compared to training data (not sure if they were used to evaluate or not though), if they were, it seems to me that it is not possible to have a split that gives you validation score that is aligned with leaderboard score. I can only get sort of directionally correct validation set up.",
    "262011": "By \"directionally correct validation set\", you mean that LB decreases as val-score decreases and LB increases when val-score increases but they are not close right ?",
    "262012": "yes.",
    "262036": "What is your architecture?",
    "262425": "What was your Conv1D architecture?",
    "262426": "Not sure who you are asking... @은주니(ttagu99) gives very generous and clear code here about his architecture already. I had a similar architecture.",
    "263747": "Hi, \n\nI am also using mainly 1D conv following VGG architecture but I get \"bad\" LB (~0.65-0.7). ( I also try the model of @은주니(ttagu99) but same result ). With the validation set I get almost 0.9 so I am wondering where I am wrong. When I split the dataset for a word given, I do not put a same ID in the train and validation set at the same time. Does anyone have the same probleme ?",
    "263748": "Make sure your labels are good",
    "263752": "That's what I thought too, but I checked the label in my train test (number associated to the word) and compare it to the submission and it does not seem wrong. Maybe I am doing something wrong before feeding the neural network ... ( I use the signal function in scipy and use a zero padding when the audio lasts less than 1 sec )",
    "263755": "Sukjae Cho How did you split your train/val ? Given a word you put different ID in the train or validation OR given an ID you put in the train or the validation set ? (in the first case we can have the same ID in the train and validation set but saying a different word, in the second case not)",
    "264690": "ShiroK, I split train/val by person's ID(second option). But icybee seems not using this approach according to other post(shocker!) and doing better, so there may be better way.",
    "264724": "Sukjae Cho thanks ! I tried the two methods I explained before and it does not seems to change a lot. \n@은주니(ttagu99) I tried your architecture but I obtain 'low' LB ~0.7 did you do a preprocess or your input is just the raw data (16000, 1) ? I do not really understand why I do not obtain similar result :/",
    "267557": "Sukjae Cho how big is your kernel size to achieve 87% LB? I tried 3, 9, 17 and 33 and only able to reach 86% LB (after 10 max pool).",
    "267715": "I'm using 9 at first layers. There may be better size, but doesn't matter much - I think the network itself has enough capacity already, so other factors matter more.",
    "269131": "Thanks. It seems going deeper helps a lot. I started first with the ttagu99's model and increased the number of the max pool layers and was able to get around 84% LB. After that I made again my network wider (starting with 32 features) and deeper (10 max pools) and got 86% LB. On this model, I tried many kernel sizes like 3, 9, 17 and 33 so on but there was no improvement. \n\nAfter the reply of Sukjae Cho, I increased my network depth (12 max pools) and got more than 1.5% validation accuracy improvement which should be enough to get 87% LB.",
    "269148": "Thanks for update. Now I understand one of the mysteries :)\nI had both 1D model and 2D model with similar performance(both upper 87%). I experimented both models but couldn't get any improvement. So I concluded that model capacity is enough and focused on augmentation. As I add different kind of augmentations, 2D model performance start to improve, but oddly 1D model performance dropped a lot. I didn't understand why and dropped 1D model, but it looks like my 1D model wasn't large enough to handle heavily augmented samples. Too bad that there are not much time left :(",
    "269515": "Did anyone try WaveNet? It was too slow for me to train before the submission deadline, so I'm curious if anyone else had tried it.",
    "270591": "Yes, our best WaveNet-like architecture had 2 res blocks with dilation rates from 1 till 2048. We followed the logic of speech-to-text wave-net from https://github.com/buriburisuri/speech-to-text-wavenet but without gated activations and with averaging pooling in the end. It achieved 87% on private and 86% on public LB. The model used a lot of RAM due to a large input dimension and took approximately 24 hours to train, but it was the only model that worked for us, so we included it into final submission :)",
    "270623": "Cool. I used dilation rates up to 8192 because that would have a receptive field of 16384, which is greater than the 16000 size of each audio sample.",
    "270627": "Yeah I followed the same logic and tried 8192 but it didn't bring any improvements for some reason.",
    "271096": "Before moving to PyTorch very late in the competition, I had built a few 1d models in Tensorflow. They were doing 'no worse' than the 2d models with that training scheme.... \n\nSome code at the link below. The 'basic3' was consistently decent, think 'a' was alright too. I experimented with some wavenet style dilation + skip connections but it was a beast, did end up converging to something reasonable \n\nhttps://github.com/rwightman/tensorflow-speech_commands/blob/master/models/conv1d.py",
    "280840": "I have also tried using 1-D convolutions. Mine did not cross 0.81 in LB. I have used kernels of size 32 with stride of 4.",
    "869360": "Dear friend\nI am trying to complete my master's work and I need to include a topic on audio classification through the 1d convolutional network directly from the audio signal. I saw that you argued a lot about the case. Could you help me with some python code that extracts the signal, writes it to a .npy file, then retrieves it and makes the classification using conv1d.\nThank you.\npmacedofilho@gmail.com"
  },
  "source": "meta"
}