{
  "id": 47079,
  "title": "Ensembling 0.86 models to 0.89",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47079",
  "author_name": "",
  "post_date": "2018-01-08T06:51:22.747119800Z",
  "votes": 36,
  "comment_count": 7,
  "views": 0,
  "content": "<p>[Things that worked]</p>\n\n<ul>\n<li>to get LB=0.89, you need about 20 models for ensemble, each is about 0.86 and they should be uncorrelated. How to make uncorrelated models?\n<ol><li>different input: waveform, spectrogram, MFCC</li>\n<li>different feature extraction: vgg, resnet, se-resnet,  inception, etc ... multi-scale features\nfor inception, try square and long, tall rectangular conv filters like 3x3,5x1,1x5 in inception block</li>\n<li>different network: CNN, RNN/LSTM</li>\n<li>if you are using CNN, normally you would use: feature_extractor --&gt;pooling --&gt;classifier\nfor pooling, try flatten, global avg pool, max avg pool, global avg+max avg pool, gated pool</li>\n<li>Apply unsupervised pesudo-label from some model to other models</li>\n<li>different augmentation (time, pitch,amplitude stretch, etc)</li>\n<li>different loss</li></ol></li>\n</ul>\n\n<p>[Things I haven't figured out]</p>\n\n<ul>\n<li>Some input works better than others. E.g. how to set the parameters for spectrogram/MFCC )e.g. frame length, num of FFT, window size ...). Because some kaggler reported best plain single model (without TTA, unsupervised learning) is 0.88. I only mange to get 0.87</li>\n<li>Normalization. Do we need to normlise spectrogram/MFCC?</li>\n<li>Denoising or noise reduction? Are those useful?</li>\n</ul>",
  "messages": [
    {
      "id": "266236",
      "postDate": "01/08/2018 06:51:22",
      "content": "<p>[Things that worked]</p>\n\n<ul>\n<li>to get LB=0.89, you need about 20 models for ensemble, each is about 0.86 and they should be uncorrelated. How to make uncorrelated models?\n<ol><li>different input: waveform, spectrogram, MFCC</li>\n<li>different feature extraction: vgg, resnet, se-resnet,  inception, etc ... multi-scale features\nfor inception, try square and long, tall rectangular conv filters like 3x3,5x1,1x5 in inception block</li>\n<li>different network: CNN, RNN/LSTM</li>\n<li>if you are using CNN, normally you would use: feature_extractor --&gt;pooling --&gt;classifier\nfor pooling, try flatten, global avg pool, max avg pool, global avg+max avg pool, gated pool</li>\n<li>Apply unsupervised pesudo-label from some model to other models</li>\n<li>different augmentation (time, pitch,amplitude stretch, etc)</li>\n<li>different loss</li></ol></li>\n</ul>\n\n<p>[Things I haven't figured out]</p>\n\n<ul>\n<li>Some input works better than others. E.g. how to set the parameters for spectrogram/MFCC )e.g. frame length, num of FFT, window size ...). Because some kaggler reported best plain single model (without TTA, unsupervised learning) is 0.88. I only mange to get 0.87</li>\n<li>Normalization. Do we need to normlise spectrogram/MFCC?</li>\n<li>Denoising or noise reduction? Are those useful?</li>\n</ul>",
      "rawMarkdown": "[Things that worked]\n\n- to get LB=0.89, you need about 20 models for ensemble, each is about 0.86 and they should be uncorrelated. How to make uncorrelated models?\n  1. different input: waveform, spectrogram, MFCC\n  2. different feature extraction: vgg, resnet, se-resnet,  inception, etc ... multi-scale features\n      for inception, try square and long, tall rectangular conv filters like 3x3,5x1,1x5 in inception block\n  3. different network: CNN, RNN/LSTM\n  4. if you are using CNN, normally you would use: feature_extractor --&gt;pooling --&gt;classifier\n      for pooling, try flatten, global avg pool, max avg pool, global avg+max avg pool, gated pool\n  5. Apply unsupervised pesudo-label from some model to other models\n  6. different augmentation (time, pitch,amplitude stretch, etc)\n  7. different loss\n\n\n[Things I haven't figured out]\n\n - Some input works better than others. E.g. how to set the parameters for spectrogram/MFCC )e.g. frame length, num of FFT, window size ...). Because some kaggler reported best plain single model (without TTA, unsupervised learning) is 0.88. I only mange to get 0.87\n - Normalization. Do we need to normlise spectrogram/MFCC?\n - Denoising or noise reduction? Are those useful?",
      "votes": null
    },
    {
      "id": "266270",
      "postDate": "01/08/2018 09:39:11",
      "content": "<p>some papers you are reference to make a single model from ensemble of complex models:</p>\n\n<p>[1] Data Distillation: Towards Omni-Supervised Learning -\nIlija Radosavovic, Piotr Dollár, Ross Girshick, Georgia Gkioxari, Kaiming He</p>\n\n<p>[2] Distilling the Knowledge in a Neural Network - Geoffrey Hinton, Oriol Vinyals, Jeff Dean</p>\n\n<p>[3] Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results- Antti Tarvainen, Harri Valpola</p>\n\n<p>[4] Semi-Supervised Learning with Ladder Networks - Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, Tapani Raiko</p>",
      "rawMarkdown": "some papers you are reference to make a single model from ensemble of complex models:\n\n[1] Data Distillation: Towards Omni-Supervised Learning -\nIlija Radosavovic, Piotr Dollár, Ross Girshick, Georgia Gkioxari, Kaiming He\n\n\n[2] Distilling the Knowledge in a Neural Network - Geoffrey Hinton, Oriol Vinyals, Jeff Dean\n\n[3] Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results- Antti Tarvainen, Harri Valpola\n\n[4] Semi-Supervised Learning with Ladder Networks - Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, Tapani Raiko",
      "votes": null
    },
    {
      "id": "266388",
      "postDate": "01/08/2018 17:18:14",
      "content": "<p>Re: Denoising or noise reduction? Are those useful?</p>\n\n<p>Denoising and noise reduction did't work for me in this case. Maybe I didn't set things right.</p>",
      "rawMarkdown": "Re: Denoising or noise reduction? Are those useful?\n\nDenoising and noise reduction did't work for me in this case. Maybe I didn't set things right.",
      "votes": null
    },
    {
      "id": "266392",
      "postDate": "01/08/2018 17:24:59",
      "content": "<p>can you give details on how your denoise or noise reduction is done? thanks!</p>",
      "rawMarkdown": "can you give details on how your denoise or noise reduction is done? thanks!",
      "votes": null
    },
    {
      "id": "266398",
      "postDate": "01/08/2018 17:35:49",
      "content": "<p>The way I did is for each wave in the train :</p>\n\n<p>sox wave. input -n noiseprof noise.prof</p>\n\n<p>sox wave.input wave.output noisered noise.prof 0.21</p>\n\n<p>The noise addition also didn't work. I mixed the each wave with a mixed noise ( white + pink + ..., each has a random weight) by a random weight on their energy terms. </p>",
      "rawMarkdown": "The way I did is for each wave in the train :\n\nsox wave. input -n noiseprof noise.prof\n\nsox wave.input wave.output noisered noise.prof 0.21\n\nThe noise addition also didn't work. I mixed the each wave with a mixed noise ( white + pink + ..., each has a random weight) by a random weight on their energy terms.",
      "votes": null
    },
    {
      "id": "266496",
      "postDate": "01/08/2018 23:27:27",
      "content": "<p>For what it's worth, I tried normalizing with MFCCs using a CNN and the results were poor. I'm pretty new to this stuff so my approach could be totally wrong. </p>\n\n<p>Using the tutorial code, I basically got all the data (by passing in a -1 to the <code>how_many</code> arg of  the <code>get_data</code> method from the audio processor) then used <code>np.mean</code> and <code>np.std</code> on this result, subtracting and dividing respectively each training sample. Did the same for the test data, but used the graph in freeze.py which does no preprocessing - it simply converts the clips straight to MFCCs. The mean and standard deviation were noticeably different on the test data.</p>\n\n<p>Anyway, my score went from a 0.85 to a 0.43 after normalizing so I scrapped that idea. But again, I could be totally off on my procedure.</p>",
      "rawMarkdown": "For what it's worth, I tried normalizing with MFCCs using a CNN and the results were poor. I'm pretty new to this stuff so my approach could be totally wrong. \n\nUsing the tutorial code, I basically got all the data (by passing in a -1 to the `how_many` arg of  the `get_data` method from the audio processor) then used `np.mean` and `np.std` on this result, subtracting and dividing respectively each training sample. Did the same for the test data, but used the graph in freeze.py which does no preprocessing - it simply converts the clips straight to MFCCs. The mean and standard deviation were noticeably different on the test data.\n\nAnyway, my score went from a 0.85 to a 0.43 after normalizing so I scrapped that idea. But again, I could be totally off on my procedure.",
      "votes": null
    },
    {
      "id": "267832",
      "postDate": "01/12/2018 13:16:05",
      "content": "<p>Thanks for sharing.</p>\n\n<p>What methods do you use for ensemble? Just average over the models or use second-level model for stacking?</p>",
      "rawMarkdown": "Thanks for sharing.\n\nWhat methods do you use for ensemble? Just average over the models or use second-level model for stacking?",
      "votes": null
    },
    {
      "id": "267838",
      "postDate": "01/12/2018 13:43:38",
      "content": "<p>weighted average of sqrt(raw_probability). currently, i am trying  level 2 stacking</p>",
      "rawMarkdown": "weighted average of sqrt(raw_probability). currently, i am trying  level 2 stacking",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 266270,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/08/2018 09:39:11",
      "content": "<p>some papers you are reference to make a single model from ensemble of complex models:</p>\n\n<p>[1] Data Distillation: Towards Omni-Supervised Learning -\nIlija Radosavovic, Piotr Dollár, Ross Girshick, Georgia Gkioxari, Kaiming He</p>\n\n<p>[2] Distilling the Knowledge in a Neural Network - Geoffrey Hinton, Oriol Vinyals, Jeff Dean</p>\n\n<p>[3] Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results- Antti Tarvainen, Harri Valpola</p>\n\n<p>[4] Semi-Supervised Learning with Ladder Networks - Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, Tapani Raiko</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 266388,
      "author_name": "ybwu01",
      "author_url": "",
      "post_date": "01/08/2018 17:18:14",
      "content": "<p>Re: Denoising or noise reduction? Are those useful?</p>\n\n<p>Denoising and noise reduction did't work for me in this case. Maybe I didn't set things right.</p>",
      "votes": null,
      "replies": [
        {
          "id": 266392,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/08/2018 17:24:59",
          "content": "<p>can you give details on how your denoise or noise reduction is done? thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 266398,
          "author_name": "ybwu01",
          "author_url": "",
          "post_date": "01/08/2018 17:35:49",
          "content": "<p>The way I did is for each wave in the train :</p>\n\n<p>sox wave. input -n noiseprof noise.prof</p>\n\n<p>sox wave.input wave.output noisered noise.prof 0.21</p>\n\n<p>The noise addition also didn't work. I mixed the each wave with a mixed noise ( white + pink + ..., each has a random weight) by a random weight on their energy terms. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 266496,
      "author_name": "jfaath",
      "author_url": "",
      "post_date": "01/08/2018 23:27:27",
      "content": "<p>For what it's worth, I tried normalizing with MFCCs using a CNN and the results were poor. I'm pretty new to this stuff so my approach could be totally wrong. </p>\n\n<p>Using the tutorial code, I basically got all the data (by passing in a -1 to the <code>how_many</code> arg of  the <code>get_data</code> method from the audio processor) then used <code>np.mean</code> and <code>np.std</code> on this result, subtracting and dividing respectively each training sample. Did the same for the test data, but used the graph in freeze.py which does no preprocessing - it simply converts the clips straight to MFCCs. The mean and standard deviation were noticeably different on the test data.</p>\n\n<p>Anyway, my score went from a 0.85 to a 0.43 after normalizing so I scrapped that idea. But again, I could be totally off on my procedure.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267832,
      "author_name": "bencww",
      "author_url": "",
      "post_date": "01/12/2018 13:16:05",
      "content": "<p>Thanks for sharing.</p>\n\n<p>What methods do you use for ensemble? Just average over the models or use second-level model for stacking?</p>",
      "votes": null,
      "replies": [
        {
          "id": 267838,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/12/2018 13:43:38",
          "content": "<p>weighted average of sqrt(raw_probability). currently, i am trying  level 2 stacking</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "266236": "[Things that worked]\n\n- to get LB=0.89, you need about 20 models for ensemble, each is about 0.86 and they should be uncorrelated. How to make uncorrelated models?\n  1. different input: waveform, spectrogram, MFCC\n  2. different feature extraction: vgg, resnet, se-resnet,  inception, etc ... multi-scale features\n      for inception, try square and long, tall rectangular conv filters like 3x3,5x1,1x5 in inception block\n  3. different network: CNN, RNN/LSTM\n  4. if you are using CNN, normally you would use: feature_extractor --&gt;pooling --&gt;classifier\n      for pooling, try flatten, global avg pool, max avg pool, global avg+max avg pool, gated pool\n  5. Apply unsupervised pesudo-label from some model to other models\n  6. different augmentation (time, pitch,amplitude stretch, etc)\n  7. different loss\n\n\n[Things I haven't figured out]\n\n - Some input works better than others. E.g. how to set the parameters for spectrogram/MFCC )e.g. frame length, num of FFT, window size ...). Because some kaggler reported best plain single model (without TTA, unsupervised learning) is 0.88. I only mange to get 0.87\n - Normalization. Do we need to normlise spectrogram/MFCC?\n - Denoising or noise reduction? Are those useful?",
    "266270": "some papers you are reference to make a single model from ensemble of complex models:\n\n[1] Data Distillation: Towards Omni-Supervised Learning -\nIlija Radosavovic, Piotr Dollár, Ross Girshick, Georgia Gkioxari, Kaiming He\n\n\n[2] Distilling the Knowledge in a Neural Network - Geoffrey Hinton, Oriol Vinyals, Jeff Dean\n\n[3] Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results- Antti Tarvainen, Harri Valpola\n\n[4] Semi-Supervised Learning with Ladder Networks - Antti Rasmus, Harri Valpola, Mikko Honkala, Mathias Berglund, Tapani Raiko",
    "266388": "Re: Denoising or noise reduction? Are those useful?\n\nDenoising and noise reduction did't work for me in this case. Maybe I didn't set things right.",
    "266392": "can you give details on how your denoise or noise reduction is done? thanks!",
    "266398": "The way I did is for each wave in the train :\n\nsox wave. input -n noiseprof noise.prof\n\nsox wave.input wave.output noisered noise.prof 0.21\n\nThe noise addition also didn't work. I mixed the each wave with a mixed noise ( white + pink + ..., each has a random weight) by a random weight on their energy terms.",
    "266496": "For what it's worth, I tried normalizing with MFCCs using a CNN and the results were poor. I'm pretty new to this stuff so my approach could be totally wrong. \n\nUsing the tutorial code, I basically got all the data (by passing in a -1 to the `how_many` arg of  the `get_data` method from the audio processor) then used `np.mean` and `np.std` on this result, subtracting and dividing respectively each training sample. Did the same for the test data, but used the graph in freeze.py which does no preprocessing - it simply converts the clips straight to MFCCs. The mean and standard deviation were noticeably different on the test data.\n\nAnyway, my score went from a 0.85 to a 0.43 after normalizing so I scrapped that idea. But again, I could be totally off on my procedure.",
    "267832": "Thanks for sharing.\n\nWhat methods do you use for ensemble? Just average over the models or use second-level model for stacking?",
    "267838": "weighted average of sqrt(raw_probability). currently, i am trying  level 2 stacking"
  },
  "source": "meta"
}