{
  "id": 47618,
  "title": "LB Score 0.90637 Approach",
  "url": "/competitions/tensorflow-speech-recognition-challenge/writeups/gold-gazua-lb-score-0-90637-approach",
  "author_name": "",
  "post_date": "2018-01-17T02:44:02.646851500Z",
  "votes": 26,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I had no background in speech recognition at all, but thanks to many generous kernels/discussions and I could learned a lot during this competition. Especially thanks to  @ttagu99, @Heng, @vinvinvin</p>\n\n<p>At first I wasn't interested in this competition, but 1D conv approach looked really interesting so I just gave it a try. Here's my initial approach:</p>\n\n<pre><code>1. Used 1D conv net with 10 pooling layer, used kernel size 9, filter count 256 for first layer.  Almost similar to ttagu99's, but replaced GAP+GMP to GMP and just used single FC with no dropout.\n2. Split train/val by person id, and train/predict on all 10 folds. \n3. Listened mis-predicted samples from validation set(around 2000?) and noticed some mislabeled samples and samples without any voice in it. Smart guy would find another algorithm to identify these, but I just listened to them. Eventually I identified 640 silences, 121 mislabels.\n4. Concatenated all noise and identified silences in training set into single wav. And it is randomly sampled while augmentation.\n5. Augmentation: \n    a. Time-shift augmentation: many samples are clipped at start or end, so I thought it's better not to cut these out. So I just randomly padded samples front &amp; back with random noise and increased PCM sample count to 20k. I didn't expect this augmentation help much, somehow it helped somewhat. \n    b. Noise augmentation: Added up to x.5 noise and it improved LB little bit. \n    c. Tried other augmentations like pitch, volume, speed, but they didn't help much or even harmed the performance.\n</code></pre>\n\n<p>With this approach, I got LB score of .87 and couldn't increase the performance anymore with 1D conv. Tried some 2D Conv approach but didn't work well.</p>\n\n<p>Later I formed a team with @Ildoo Kim who used high resolution mel spectrograms + VGG like network and had similar score as mine. We got immediate boost after merging of my augmentation and Ildoo's model. After some more fiddling of models, we got little bit of improvements. But we're stuck around .88 with single model, .893 with 5 model ensembles for a while. </p>\n\n<p>I concluded that the model is large enough, so I worked more on data and found that adding heavy noise augmentation while keeping noise vs signal ratio doesn't exceed 2(yes, noise can be twice louder than voice) boost the score a lot. Just with this augmentation, same model (resnet-like net) got public LB score .898</p>\n\n<p>Unfortunately we found this 2 days just before deadline and didn't have much time and submissions to experiment more. So we just trained a few more models and ensembled blindly even without checking individual scores.</p>\n\n<p>One interesting thing was my original 1D model didn't work well after adding heavy noise. So I dropped it altogether. But later I found that 1D model also can get better score also once I add more capacity to the network.</p>",
  "messages": [
    {
      "id": "269596",
      "postDate": "01/17/2018 02:44:02",
      "content": "<p>I had no background in speech recognition at all, but thanks to many generous kernels/discussions and I could learned a lot during this competition. Especially thanks to  @ttagu99, @Heng, @vinvinvin</p>\n\n<p>At first I wasn't interested in this competition, but 1D conv approach looked really interesting so I just gave it a try. Here's my initial approach:</p>\n\n<pre><code>1. Used 1D conv net with 10 pooling layer, used kernel size 9, filter count 256 for first layer.  Almost similar to ttagu99's, but replaced GAP+GMP to GMP and just used single FC with no dropout.\n2. Split train/val by person id, and train/predict on all 10 folds. \n3. Listened mis-predicted samples from validation set(around 2000?) and noticed some mislabeled samples and samples without any voice in it. Smart guy would find another algorithm to identify these, but I just listened to them. Eventually I identified 640 silences, 121 mislabels.\n4. Concatenated all noise and identified silences in training set into single wav. And it is randomly sampled while augmentation.\n5. Augmentation: \n    a. Time-shift augmentation: many samples are clipped at start or end, so I thought it's better not to cut these out. So I just randomly padded samples front &amp; back with random noise and increased PCM sample count to 20k. I didn't expect this augmentation help much, somehow it helped somewhat. \n    b. Noise augmentation: Added up to x.5 noise and it improved LB little bit. \n    c. Tried other augmentations like pitch, volume, speed, but they didn't help much or even harmed the performance.\n</code></pre>\n\n<p>With this approach, I got LB score of .87 and couldn't increase the performance anymore with 1D conv. Tried some 2D Conv approach but didn't work well.</p>\n\n<p>Later I formed a team with @Ildoo Kim who used high resolution mel spectrograms + VGG like network and had similar score as mine. We got immediate boost after merging of my augmentation and Ildoo's model. After some more fiddling of models, we got little bit of improvements. But we're stuck around .88 with single model, .893 with 5 model ensembles for a while. </p>\n\n<p>I concluded that the model is large enough, so I worked more on data and found that adding heavy noise augmentation while keeping noise vs signal ratio doesn't exceed 2(yes, noise can be twice louder than voice) boost the score a lot. Just with this augmentation, same model (resnet-like net) got public LB score .898</p>\n\n<p>Unfortunately we found this 2 days just before deadline and didn't have much time and submissions to experiment more. So we just trained a few more models and ensembled blindly even without checking individual scores.</p>\n\n<p>One interesting thing was my original 1D model didn't work well after adding heavy noise. So I dropped it altogether. But later I found that 1D model also can get better score also once I add more capacity to the network.</p>",
      "rawMarkdown": "I had no background in speech recognition at all, but thanks to many generous kernels/discussions and I could learned a lot during this competition. Especially thanks to  @ttagu99, @Heng, @vinvinvin\n\nAt first I wasn't interested in this competition, but 1D conv approach looked really interesting so I just gave it a try. Here's my initial approach:\n\n\t1. Used 1D conv net with 10 pooling layer, used kernel size 9, filter count 256 for first layer.  Almost similar to ttagu99's, but replaced GAP+GMP to GMP and just used single FC with no dropout.\n\t2. Split train/val by person id, and train/predict on all 10 folds. \n\t3. Listened mis-predicted samples from validation set(around 2000?) and noticed some mislabeled samples and samples without any voice in it. Smart guy would find another algorithm to identify these, but I just listened to them. Eventually I identified 640 silences, 121 mislabels.\n\t4. Concatenated all noise and identified silences in training set into single wav. And it is randomly sampled while augmentation.\n\t5. Augmentation: \n\t\ta. Time-shift augmentation: many samples are clipped at start or end, so I thought it's better not to cut these out. So I just randomly padded samples front &amp; back with random noise and increased PCM sample count to 20k. I didn't expect this augmentation help much, somehow it helped somewhat. \n\t\tb. Noise augmentation: Added up to x.5 noise and it improved LB little bit. \n\t\tc. Tried other augmentations like pitch, volume, speed, but they didn't help much or even harmed the performance.\n\t\t\nWith this approach, I got LB score of .87 and couldn't increase the performance anymore with 1D conv. Tried some 2D Conv approach but didn't work well.\n\nLater I formed a team with @Ildoo Kim who used high resolution mel spectrograms + VGG like network and had similar score as mine. We got immediate boost after merging of my augmentation and Ildoo's model. After some more fiddling of models, we got little bit of improvements. But we're stuck around .88 with single model, .893 with 5 model ensembles for a while. \n\nI concluded that the model is large enough, so I worked more on data and found that adding heavy noise augmentation while keeping noise vs signal ratio doesn't exceed 2(yes, noise can be twice louder than voice) boost the score a lot. Just with this augmentation, same model (resnet-like net) got public LB score .898\n\nUnfortunately we found this 2 days just before deadline and didn't have much time and submissions to experiment more. So we just trained a few more models and ensembled blindly even without checking individual scores.\n\nOne interesting thing was my original 1D model didn't work well after adding heavy noise. So I dropped it altogether. But later I found that 1D model also can get better score also once I add more capacity to the network.",
      "votes": null
    },
    {
      "id": "269640",
      "postDate": "01/17/2018 03:44:49",
      "content": "<p>Awsome work! Thanks!</p>",
      "rawMarkdown": "Awsome work! Thanks!",
      "votes": null
    },
    {
      "id": "269650",
      "postDate": "01/17/2018 04:07:53",
      "content": "<p>Big thanks to you and ttagu99 for your comments on the \"Anyone using 1-d convolutions?\" discussion. We were stuck on  2D Resnets and tried a conv1D variant based on the interest generated in that thread (raw input, kernel 9, resnet conv1D) and got fantastic single model results (public LB 0.88706, private LB 0.89005)</p>\n\n<p>I feel like part of our silver is due to the great discussion you two had there.  </p>",
      "rawMarkdown": "Big thanks to you and ttagu99 for your comments on the \"Anyone using 1-d convolutions?\" discussion. We were stuck on  2D Resnets and tried a conv1D variant based on the interest generated in that thread (raw input, kernel 9, resnet conv1D) and got fantastic single model results (public LB 0.88706, private LB 0.89005)\n\nI feel like part of our silver is due to the great discussion you two had there.",
      "votes": null
    },
    {
      "id": "269682",
      "postDate": "01/17/2018 05:20:15",
      "content": "<p>Same here! I never thought unprocessed 1d wave would work. I only tried after reading @ttagu99, @Sukjae Cho, and many others in the discussion. So my solution are in fact ideas from many kagglers.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Same here! I never thought unprocessed 1d wave would work. I only tried after reading @ttagu99, @Sukjae Cho, and many others in the discussion. So my solution are in fact ideas from many kagglers.\n\nThanks!",
      "votes": null
    },
    {
      "id": "269727",
      "postDate": "01/17/2018 07:16:20",
      "content": "<p>I had no experience in speech recognition technology. I had no experience with 1d-cnn, but I was able to further develop it with the discussion of Human Analog, Sukjae Cho and Ren. Thank you.\nNow, 1d cnn submission files are deleted due to system problems, public score and private score.\nSo I do not know how much my individual approach got to score. When the system is fixed, we will update the contents again.\nMy 1d model used two architectures.</p>\n\n<p>Private LB 0.87571</p>\n\n<pre><code>init_filter_num = 8\nx_in_1d = Input(shape = (16000,1))\nx_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\nfor i in range(9):\n    name = 'step'+str(i)\n    x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n    x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n    if i !=0:\n        x_1d = layers.add([x_1d_concate, x_1d])\n    x_1d = Activation('relu')(x_1d)\n    for j in range(2,14):\n        short_cut = x_1d\n        x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_'+str(j))(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_'+str(j))(x_1d)\n        x_1d = layers.add([short_cut,x_1d])\n        x_1d = Activation('relu')(x_1d)\n    x_1d_max = MaxPooling1D((2), padding='same')(x_1d)\n    x_1d_avg = AvgPool1D((2), padding='same')(x_1d)\n    x_1d = layers.add([x_1d_max, x_1d_avg])\n    x_1d = BatchNormalization(name = 'batch'+name+'_avgmax_add')(x_1d)\n    x_1d = Activation('relu')(x_1d)\n    if i != 8:\n        x_1d_concate = layers.concatenate([x_1d_max, x_1d_avg])\n        x_1d_concate = BatchNormalization(name = 'batch'+name+'_avgmax_concate')(x_1d_concate)\nx_1d = Conv1D(1024, (1),name='last1024')(x_1d)\nx_1d = GlobalMaxPool1D()(x_1d) #only g max\nx_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\nx_1d = Dropout(0.2)(x_1d)\nx_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)\n</code></pre>\n\n<p>Private LB 0.87513</p>\n\n<pre><code>init_filter_num = 8\nx_in_1d = Input(shape = (16000,1))\nx_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\nfor i in range(9):\n    name = 'step'+str(i)\n    x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n    x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n    x_1d = Activation('relu')(x_1d)\n    x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_2')(x_1d)\n    x_1d = BatchNormalization(name = 'batch'+name+'_2')(x_1d)\n    x_1d = Activation('relu')(x_1d)\n    x_1d = MaxPooling1D((2), padding='same')(x_1d)\nx_1d = Conv1D(1024, (1),name='last1024')(x_1d)\nx_1d = GlobalMaxPool1D()(x_1d) #only g max\nx_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\nx_1d = Dropout(0.2)(x_1d)\nx_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)\n</code></pre>",
      "rawMarkdown": "I had no experience in speech recognition technology. I had no experience with 1d-cnn, but I was able to further develop it with the discussion of Human Analog, Sukjae Cho and Ren. Thank you.\nNow, 1d cnn submission files are deleted due to system problems, public score and private score.\nSo I do not know how much my individual approach got to score. When the system is fixed, we will update the contents again.\nMy 1d model used two architectures.\n\n\nPrivate LB 0.87571\n\n    \n    init_filter_num = 8\n    x_in_1d = Input(shape = (16000,1))\n    x_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\n    for i in range(9):\n        name = 'step'+str(i)\n        x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n        if i !=0:\n            x_1d = layers.add([x_1d_concate, x_1d])\n        x_1d = Activation('relu')(x_1d)\n        for j in range(2,14):\n            short_cut = x_1d\n            x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_'+str(j))(x_1d)\n            x_1d = BatchNormalization(name = 'batch'+name+'_'+str(j))(x_1d)\n            x_1d = layers.add([short_cut,x_1d])\n            x_1d = Activation('relu')(x_1d)\n        x_1d_max = MaxPooling1D((2), padding='same')(x_1d)\n        x_1d_avg = AvgPool1D((2), padding='same')(x_1d)\n        x_1d = layers.add([x_1d_max, x_1d_avg])\n        x_1d = BatchNormalization(name = 'batch'+name+'_avgmax_add')(x_1d)\n        x_1d = Activation('relu')(x_1d)\n        if i != 8:\n            x_1d_concate = layers.concatenate([x_1d_max, x_1d_avg])\n            x_1d_concate = BatchNormalization(name = 'batch'+name+'_avgmax_concate')(x_1d_concate)\n    x_1d = Conv1D(1024, (1),name='last1024')(x_1d)\n    x_1d = GlobalMaxPool1D()(x_1d) #only g max\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\n    x_1d = Dropout(0.2)(x_1d)\n    x_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)\n\n\n\nPrivate LB 0.87513\n\n    init_filter_num = 8\n    x_in_1d = Input(shape = (16000,1))\n    x_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\n    for i in range(9):\n        name = 'step'+str(i)\n        x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n        x_1d = Activation('relu')(x_1d)\n        x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_2')(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_2')(x_1d)\n        x_1d = Activation('relu')(x_1d)\n        x_1d = MaxPooling1D((2), padding='same')(x_1d)\n    x_1d = Conv1D(1024, (1),name='last1024')(x_1d)\n    x_1d = GlobalMaxPool1D()(x_1d) #only g max\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\n    x_1d = Dropout(0.2)(x_1d)\n    x_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)",
      "votes": null
    },
    {
      "id": "269792",
      "postDate": "01/17/2018 10:08:40",
      "content": "<p>Interesting about the SNR levels required. I spent way too much time on creating samples mixed with noise and it was quite disappointing barely getting any benefit from them. </p>",
      "rawMarkdown": "Interesting about the SNR levels required. I spent way too much time on creating samples mixed with noise and it was quite disappointing barely getting any benefit from them.",
      "votes": null
    },
    {
      "id": "270064",
      "postDate": "01/17/2018 17:53:31",
      "content": "<p>In my case, I got better result with BatchNormalization instead of Dropout at FC layer.</p>",
      "rawMarkdown": "In my case, I got better result with BatchNormalization instead of Dropout at FC layer.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 269640,
      "author_name": "yyll008",
      "author_url": "",
      "post_date": "01/17/2018 03:44:49",
      "content": "<p>Awsome work! Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 269650,
      "author_name": "atfisnotatf",
      "author_url": "",
      "post_date": "01/17/2018 04:07:53",
      "content": "<p>Big thanks to you and ttagu99 for your comments on the \"Anyone using 1-d convolutions?\" discussion. We were stuck on  2D Resnets and tried a conv1D variant based on the interest generated in that thread (raw input, kernel 9, resnet conv1D) and got fantastic single model results (public LB 0.88706, private LB 0.89005)</p>\n\n<p>I feel like part of our silver is due to the great discussion you two had there.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 269682,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/17/2018 05:20:15",
          "content": "<p>Same here! I never thought unprocessed 1d wave would work. I only tried after reading @ttagu99, @Sukjae Cho, and many others in the discussion. So my solution are in fact ideas from many kagglers.</p>\n\n<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269727,
      "author_name": "ttagu99",
      "author_url": "",
      "post_date": "01/17/2018 07:16:20",
      "content": "<p>I had no experience in speech recognition technology. I had no experience with 1d-cnn, but I was able to further develop it with the discussion of Human Analog, Sukjae Cho and Ren. Thank you.\nNow, 1d cnn submission files are deleted due to system problems, public score and private score.\nSo I do not know how much my individual approach got to score. When the system is fixed, we will update the contents again.\nMy 1d model used two architectures.</p>\n\n<p>Private LB 0.87571</p>\n\n<pre><code>init_filter_num = 8\nx_in_1d = Input(shape = (16000,1))\nx_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\nfor i in range(9):\n    name = 'step'+str(i)\n    x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n    x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n    if i !=0:\n        x_1d = layers.add([x_1d_concate, x_1d])\n    x_1d = Activation('relu')(x_1d)\n    for j in range(2,14):\n        short_cut = x_1d\n        x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_'+str(j))(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_'+str(j))(x_1d)\n        x_1d = layers.add([short_cut,x_1d])\n        x_1d = Activation('relu')(x_1d)\n    x_1d_max = MaxPooling1D((2), padding='same')(x_1d)\n    x_1d_avg = AvgPool1D((2), padding='same')(x_1d)\n    x_1d = layers.add([x_1d_max, x_1d_avg])\n    x_1d = BatchNormalization(name = 'batch'+name+'_avgmax_add')(x_1d)\n    x_1d = Activation('relu')(x_1d)\n    if i != 8:\n        x_1d_concate = layers.concatenate([x_1d_max, x_1d_avg])\n        x_1d_concate = BatchNormalization(name = 'batch'+name+'_avgmax_concate')(x_1d_concate)\nx_1d = Conv1D(1024, (1),name='last1024')(x_1d)\nx_1d = GlobalMaxPool1D()(x_1d) #only g max\nx_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\nx_1d = Dropout(0.2)(x_1d)\nx_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)\n</code></pre>\n\n<p>Private LB 0.87513</p>\n\n<pre><code>init_filter_num = 8\nx_in_1d = Input(shape = (16000,1))\nx_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\nfor i in range(9):\n    name = 'step'+str(i)\n    x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n    x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n    x_1d = Activation('relu')(x_1d)\n    x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_2')(x_1d)\n    x_1d = BatchNormalization(name = 'batch'+name+'_2')(x_1d)\n    x_1d = Activation('relu')(x_1d)\n    x_1d = MaxPooling1D((2), padding='same')(x_1d)\nx_1d = Conv1D(1024, (1),name='last1024')(x_1d)\nx_1d = GlobalMaxPool1D()(x_1d) #only g max\nx_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\nx_1d = Dropout(0.2)(x_1d)\nx_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 270064,
          "author_name": "jandjenter",
          "author_url": "",
          "post_date": "01/17/2018 17:53:31",
          "content": "<p>In my case, I got better result with BatchNormalization instead of Dropout at FC layer.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 269792,
      "author_name": "nimitz14",
      "author_url": "",
      "post_date": "01/17/2018 10:08:40",
      "content": "<p>Interesting about the SNR levels required. I spent way too much time on creating samples mixed with noise and it was quite disappointing barely getting any benefit from them. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "269596": "I had no background in speech recognition at all, but thanks to many generous kernels/discussions and I could learned a lot during this competition. Especially thanks to  @ttagu99, @Heng, @vinvinvin\n\nAt first I wasn't interested in this competition, but 1D conv approach looked really interesting so I just gave it a try. Here's my initial approach:\n\n\t1. Used 1D conv net with 10 pooling layer, used kernel size 9, filter count 256 for first layer.  Almost similar to ttagu99's, but replaced GAP+GMP to GMP and just used single FC with no dropout.\n\t2. Split train/val by person id, and train/predict on all 10 folds. \n\t3. Listened mis-predicted samples from validation set(around 2000?) and noticed some mislabeled samples and samples without any voice in it. Smart guy would find another algorithm to identify these, but I just listened to them. Eventually I identified 640 silences, 121 mislabels.\n\t4. Concatenated all noise and identified silences in training set into single wav. And it is randomly sampled while augmentation.\n\t5. Augmentation: \n\t\ta. Time-shift augmentation: many samples are clipped at start or end, so I thought it's better not to cut these out. So I just randomly padded samples front &amp; back with random noise and increased PCM sample count to 20k. I didn't expect this augmentation help much, somehow it helped somewhat. \n\t\tb. Noise augmentation: Added up to x.5 noise and it improved LB little bit. \n\t\tc. Tried other augmentations like pitch, volume, speed, but they didn't help much or even harmed the performance.\n\t\t\nWith this approach, I got LB score of .87 and couldn't increase the performance anymore with 1D conv. Tried some 2D Conv approach but didn't work well.\n\nLater I formed a team with @Ildoo Kim who used high resolution mel spectrograms + VGG like network and had similar score as mine. We got immediate boost after merging of my augmentation and Ildoo's model. After some more fiddling of models, we got little bit of improvements. But we're stuck around .88 with single model, .893 with 5 model ensembles for a while. \n\nI concluded that the model is large enough, so I worked more on data and found that adding heavy noise augmentation while keeping noise vs signal ratio doesn't exceed 2(yes, noise can be twice louder than voice) boost the score a lot. Just with this augmentation, same model (resnet-like net) got public LB score .898\n\nUnfortunately we found this 2 days just before deadline and didn't have much time and submissions to experiment more. So we just trained a few more models and ensembled blindly even without checking individual scores.\n\nOne interesting thing was my original 1D model didn't work well after adding heavy noise. So I dropped it altogether. But later I found that 1D model also can get better score also once I add more capacity to the network.",
    "269640": "Awsome work! Thanks!",
    "269650": "Big thanks to you and ttagu99 for your comments on the \"Anyone using 1-d convolutions?\" discussion. We were stuck on  2D Resnets and tried a conv1D variant based on the interest generated in that thread (raw input, kernel 9, resnet conv1D) and got fantastic single model results (public LB 0.88706, private LB 0.89005)\n\nI feel like part of our silver is due to the great discussion you two had there.",
    "269682": "Same here! I never thought unprocessed 1d wave would work. I only tried after reading @ttagu99, @Sukjae Cho, and many others in the discussion. So my solution are in fact ideas from many kagglers.\n\nThanks!",
    "269727": "I had no experience in speech recognition technology. I had no experience with 1d-cnn, but I was able to further develop it with the discussion of Human Analog, Sukjae Cho and Ren. Thank you.\nNow, 1d cnn submission files are deleted due to system problems, public score and private score.\nSo I do not know how much my individual approach got to score. When the system is fixed, we will update the contents again.\nMy 1d model used two architectures.\n\n\nPrivate LB 0.87571\n\n    \n    init_filter_num = 8\n    x_in_1d = Input(shape = (16000,1))\n    x_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\n    for i in range(9):\n        name = 'step'+str(i)\n        x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n        if i !=0:\n            x_1d = layers.add([x_1d_concate, x_1d])\n        x_1d = Activation('relu')(x_1d)\n        for j in range(2,14):\n            short_cut = x_1d\n            x_1d = Conv1D(init_filter_num*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_'+str(j))(x_1d)\n            x_1d = BatchNormalization(name = 'batch'+name+'_'+str(j))(x_1d)\n            x_1d = layers.add([short_cut,x_1d])\n            x_1d = Activation('relu')(x_1d)\n        x_1d_max = MaxPooling1D((2), padding='same')(x_1d)\n        x_1d_avg = AvgPool1D((2), padding='same')(x_1d)\n        x_1d = layers.add([x_1d_max, x_1d_avg])\n        x_1d = BatchNormalization(name = 'batch'+name+'_avgmax_add')(x_1d)\n        x_1d = Activation('relu')(x_1d)\n        if i != 8:\n            x_1d_concate = layers.concatenate([x_1d_max, x_1d_avg])\n            x_1d_concate = BatchNormalization(name = 'batch'+name+'_avgmax_concate')(x_1d_concate)\n    x_1d = Conv1D(1024, (1),name='last1024')(x_1d)\n    x_1d = GlobalMaxPool1D()(x_1d) #only g max\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\n    x_1d = Dropout(0.2)(x_1d)\n    x_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)\n\n\n\nPrivate LB 0.87513\n\n    init_filter_num = 8\n    x_in_1d = Input(shape = (16000,1))\n    x_1d = BatchNormalization(name = 'batchnormal_1d_in')(x_in_1d)\n    for i in range(9):\n        name = 'step'+str(i)\n        x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_1')(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_1')(x_1d)\n        x_1d = Activation('relu')(x_1d)\n        x_1d = Conv1D(8*(2 ** i), (3),padding = 'same', name = 'conv'+name+'_2')(x_1d)\n        x_1d = BatchNormalization(name = 'batch'+name+'_2')(x_1d)\n        x_1d = Activation('relu')(x_1d)\n        x_1d = MaxPooling1D((2), padding='same')(x_1d)\n    x_1d = Conv1D(1024, (1),name='last1024')(x_1d)\n    x_1d = GlobalMaxPool1D()(x_1d) #only g max\n    x_1d = Dense(1024, activation = 'relu', name= 'dense1024_onlygmax')(x_1d)\n    x_1d = Dropout(0.2)(x_1d)\n    x_1d = Dense(len(POSSIBLE_LABELS), activation = 'softmax',name='cls_1d')(x_1d)",
    "269792": "Interesting about the SNR levels required. I spent way too much time on creating samples mixed with noise and it was quite disappointing barely getting any benefit from them.",
    "270064": "In my case, I got better result with BatchNormalization instead of Dropout at FC layer."
  },
  "source": "meta"
}