{
  "id": 47728,
  "title": "what i have learned and moving forward",
  "url": "/competitions/tensorflow-speech-recognition-challenge/writeups/heng-ryan-see-good-bug-what-i-have-learned-and-mov",
  "author_name": "",
  "post_date": "2018-01-18T17:04:14.773Z",
  "votes": 54,
  "comment_count": 15,
  "views": 0,
  "content": "<h2>HENG SOLUTION</h2>\n\n<p>my solution ]</p>\n\n<ul>\n<li><p>Most of the details are already posted. We are making document and code clean up for prize submission. Will share with you guys later. In short summary:</p>\n\n<ul><li><p>ensemble of about 30 models comprise of wave, log melspectrogram, mfccs. </p></li>\n<li><p>My models are weak in the range of 0.86. My teammate (@Ryan, @See) models are stronger,  in the range of  0.88</p></li>\n<li><p>mostly convolution networks</p></li>\n<li><p>pseudo labeling to train some of the networks (not all)</p></li></ul></li>\n</ul>\n\n<p>[ my approach to this competition ]</p>\n\n<ul>\n<li><p>Each kaggle competition is different. For some competitions, the winning factor could be:</p>\n\n<ul><li><p>feature engineering (or network design),  e.g. the carvana car segmentation</p></li>\n<li><p>dealing with noisy label, e.g. amazon satellite image classification</p></li>\n<li><p>dealing with large data+category  and efficiency, e.g cdiscount e-commerce product image classification</p></li></ul></li>\n</ul>\n\n<p>Most of the time, it is combination of the above. For this competition, i think the main challenge is dealing with data domain shift, i.e. \"train+validation data\" and \"LB data\" are different. Why?</p>\n\n<ul>\n<li><p>the gap between validation score and LB score is large (12% to 8%)</p></li>\n<li><p>the trained model is sensitive to distribution of the training class. e.g if you just train with random sampling,  simple cnn_trad_pool2_net  gives less than 0.80 on public LB. But if you use balanced class sampling, it increases to 0.82</p></li>\n<li><p>the trained model is also sensitive type of silence train samples, amplitude of noise level etc.</p></li>\n<li><p>lastly i notice the unknowns in LB data is not the same as that in train+valid set</p></li>\n</ul>\n\n<p>I am not familiar with speech recognition or audio processing, hence i think it would be difficult for me to design good network. So I focus on the data instead. I used the simplest approach, create train data from the LB set.</p>\n\n<p>Pseudo labeling may \"overfit\" and can be a dangerous approach. I need to make sure pseudo label LB samples are in correct label and correct distribution. The way to do is:</p>\n\n<ul>\n<li><p>let T = train set, V= validation set, L = LB set,  </p></li>\n<li><p>M = model trained on T</p></li>\n<li><p>P = subset of L, labelled by M  with label noise &lt; e.g  10%</p></li>\n<li><p>N = model trained on T+P</p></li>\n<li><p>Accept N if accuracy of N is better than M on both V and L </p></li>\n</ul>\n\n<p>you can relax the last acceptance test. Note if P={empty}, you have the original results. You can modify P  (e.g. by different threshold like 5%,15%,20% etc ...) until you pass the acceptance test. We actually use different thresholds for different classes (silence, unknowns, knowns) to ensure that the pseudo-labelled set is more balance. </p>\n\n<p>[ some surprises of the competition ]</p>\n\n<ul>\n<li>We see ourselves in the top 5 but did not expect to be the winner. We had a bug and mistakenly used a model twice (due to typo error in code) . And this model has much higher weights than the rest. This bug is our winning submission private LB 0.91060\n(public LB 0.90296), which is also attached below.</li>\n</ul>\n\n<p>Later, it reveal that our highest private LB is 0.91107  (public LB 0.90241) is another ensemble. It is hard to make selection given only 2 decimals are revealed at the competition.</p>\n\n<ul>\n<li><p>1d wave input actually works!  </p></li>\n<li><p>sometimes high score models does not improve when ensemble, especially for those greater than LB =0.88. I compare the confidence probability scores of weak (LB 0.86) and strong (LB 0.88) models. For strong model, the sample score are mostly very near to 1 or zero. So it is very hard to change the scores of test samples i think. </p></li>\n</ul>\n\n<p>[ how to go on ]</p>\n\n<ul>\n<li><p>I take this learning path. Be a master in training (hyper parameter tuning + data augmentation). Then be a master in network design and then be a ensemble expert.</p></li>\n<li><p>After the basics above,  i think semi/weak supervised learning is one way to go. From competition point of view, able to automatic label LB dataset and use it for training is very powerful.  (My next competition is National Science Bowl 2018, where i hope to use GAN  to generate \"LB data\" with label)</p></li>\n<li><p>I want to make a LB score predictor</p></li>\n<li><p>I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.</p></li>\n<li><p>I want to run some of the kaggle solutions that uses CRNN and LSTM. I believe that is the correct way to do speech.</p></li>\n</ul>\n\n<hr>\n\n<h2>SEE SOLUTION</h2>\n\n<h1>Overview of my approach</h1>\n\n<p>I started with the provided <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">tutorial</a> and could easily get better results by just adding momentum to the plain SGD solver (82-83% on the leaderboard). I have no prior experience with audio data and mostly used deep learning with images. For this domain you don't use features but feed the raw pixel values. My thinking was that this should work with audio data as well. Throughout the competition I ran experiments using raw waveforms, spectrograms and log mel features as input. I got similar results using log mel and raw waveform (86%-87%) and used the waveform data for most experiments as it was easier to interpret for me.</p>\n\n<p>For the special price the restrictions were: the network is smaller than 5.000.000 bytes and runs in less than 175ms per sample on a stock Raspberry Pi 3. Regarding the size, this allows you to build networks that have roughly 1.250.000 weight parameters. So by experimenting with these restrictions I came up with an architecture that uses Depthwise1D convolutions on the raw waveform. Using <a href=\"https://arxiv.org/pdf/1503.02531.pdf\">model distillation</a> this network predicts the correct class for 90.8% of the private leaderboard samples and runs in roughly 80ms.</p>\n\n<h1>What didn't work</h1>\n\n<ul>\n<li><p>Fancy augmentation methods: I tried flipping (i.e: <code>* -1.0</code>) the samples. You can check that they will sound exactly the same. I also modified <code>input_data.py</code> to change the foreground and background volume independently and created a separate volume range for the silence samples. My validation accuracy improved for some experiments but my leaderboard scores didn't.</p></li>\n<li><p>Predicting unknown unknowns: I didn't find a good way to consistently predict these words. Often, similar words were wrongly classified (e.g. one as on).</p></li>\n<li><p>Creating new words: I trained some networks with even more classes. I reversed the samples from the known unwanted words, e.g. <code>bird</code>, <code>bed</code>, <code>marvin</code>, and created new classes (<code>bird</code> -&gt; <code>drib</code> ...). The idea was to have more unknowns to prevent the network from wrongly mapping unknowns to the known words. For example the word <code>follow</code> was mostly predicted as <code>off</code>. However, neither my validation score not my leaderboard score improved.</p></li>\n<li><p>Cyclic learning rate schedules: The winning entry of the <a href=\"http://blog.kaggle.com/2017/12/22/carvana-image-masking-first-place-interview/\">Caravana Image Masking Challenge</a> used cyclic learning rates but for me the results got worse and you had additional hyperparameters. Maybe I just didn't implement it correctly.</p></li>\n</ul>\n\n<h1>What worked</h1>\n\n<ul>\n<li><p>Mixing tensorflow and Keras: Both frameworks work perfectly together and you can mix them wherever you want. For example: I wrapped the provided data AudioProcessor from <code>input_data.py</code> in a generator and used it with <code>keras.models.Model.fit_generator</code>. This way, I could implement new architectures really fast using Keras and later just extract and freeze the graph from the trained models.</p></li>\n<li><p>Pseudo labeling: I used consistent samples from the test set to train new networks. Choosing them was based on a.) my three best models agree on this submission. I used this version at early stages of the competition. b.) using a probability threshold on the predicted softmax probabilities. Typically, using <code>pseudo_threshold=0.6</code> were the samples that our ensembled model predicted correctly. I also implemented a schedule for pseudo labels. That is: For the first 5 epochs you only use pseudo labels and then gradually mix in data from the training data set. Though, I didn't have time to run these experiments, so I kept a fixed ratio of training and pseudo data.</p></li>\n<li><p>Test time augmentation: It is a simple way to get some boost. Just augment the samples, feed them multiple times and average the probabilities. I tried the following: time-shifting, increase/decrease the volume and time-stretching using <code>librosa.effects.time_stretch</code>.</p></li>\n</ul>\n\n<hr>\n\n<p>I posted out submission results here, raw probability score is normalised from [0,1] to [0,255]</p>",
  "messages": [
    {
      "id": "270321",
      "postDate": "01/18/2018 04:40:25",
      "content": "<h2>HENG SOLUTION</h2>\n\n<p>my solution ]</p>\n\n<ul>\n<li><p>Most of the details are already posted. We are making document and code clean up for prize submission. Will share with you guys later. In short summary:</p>\n\n<ul><li><p>ensemble of about 30 models comprise of wave, log melspectrogram, mfccs. </p></li>\n<li><p>My models are weak in the range of 0.86. My teammate (@Ryan, @See) models are stronger,  in the range of  0.88</p></li>\n<li><p>mostly convolution networks</p></li>\n<li><p>pseudo labeling to train some of the networks (not all)</p></li></ul></li>\n</ul>\n\n<p>[ my approach to this competition ]</p>\n\n<ul>\n<li><p>Each kaggle competition is different. For some competitions, the winning factor could be:</p>\n\n<ul><li><p>feature engineering (or network design),  e.g. the carvana car segmentation</p></li>\n<li><p>dealing with noisy label, e.g. amazon satellite image classification</p></li>\n<li><p>dealing with large data+category  and efficiency, e.g cdiscount e-commerce product image classification</p></li></ul></li>\n</ul>\n\n<p>Most of the time, it is combination of the above. For this competition, i think the main challenge is dealing with data domain shift, i.e. \"train+validation data\" and \"LB data\" are different. Why?</p>\n\n<ul>\n<li><p>the gap between validation score and LB score is large (12% to 8%)</p></li>\n<li><p>the trained model is sensitive to distribution of the training class. e.g if you just train with random sampling,  simple cnn_trad_pool2_net  gives less than 0.80 on public LB. But if you use balanced class sampling, it increases to 0.82</p></li>\n<li><p>the trained model is also sensitive type of silence train samples, amplitude of noise level etc.</p></li>\n<li><p>lastly i notice the unknowns in LB data is not the same as that in train+valid set</p></li>\n</ul>\n\n<p>I am not familiar with speech recognition or audio processing, hence i think it would be difficult for me to design good network. So I focus on the data instead. I used the simplest approach, create train data from the LB set.</p>\n\n<p>Pseudo labeling may \"overfit\" and can be a dangerous approach. I need to make sure pseudo label LB samples are in correct label and correct distribution. The way to do is:</p>\n\n<ul>\n<li><p>let T = train set, V= validation set, L = LB set,  </p></li>\n<li><p>M = model trained on T</p></li>\n<li><p>P = subset of L, labelled by M  with label noise &lt; e.g  10%</p></li>\n<li><p>N = model trained on T+P</p></li>\n<li><p>Accept N if accuracy of N is better than M on both V and L </p></li>\n</ul>\n\n<p>you can relax the last acceptance test. Note if P={empty}, you have the original results. You can modify P  (e.g. by different threshold like 5%,15%,20% etc ...) until you pass the acceptance test. We actually use different thresholds for different classes (silence, unknowns, knowns) to ensure that the pseudo-labelled set is more balance. </p>\n\n<p>[ some surprises of the competition ]</p>\n\n<ul>\n<li>We see ourselves in the top 5 but did not expect to be the winner. We had a bug and mistakenly used a model twice (due to typo error in code) . And this model has much higher weights than the rest. This bug is our winning submission private LB 0.91060\n(public LB 0.90296), which is also attached below.</li>\n</ul>\n\n<p>Later, it reveal that our highest private LB is 0.91107  (public LB 0.90241) is another ensemble. It is hard to make selection given only 2 decimals are revealed at the competition.</p>\n\n<ul>\n<li><p>1d wave input actually works!  </p></li>\n<li><p>sometimes high score models does not improve when ensemble, especially for those greater than LB =0.88. I compare the confidence probability scores of weak (LB 0.86) and strong (LB 0.88) models. For strong model, the sample score are mostly very near to 1 or zero. So it is very hard to change the scores of test samples i think. </p></li>\n</ul>\n\n<p>[ how to go on ]</p>\n\n<ul>\n<li><p>I take this learning path. Be a master in training (hyper parameter tuning + data augmentation). Then be a master in network design and then be a ensemble expert.</p></li>\n<li><p>After the basics above,  i think semi/weak supervised learning is one way to go. From competition point of view, able to automatic label LB dataset and use it for training is very powerful.  (My next competition is National Science Bowl 2018, where i hope to use GAN  to generate \"LB data\" with label)</p></li>\n<li><p>I want to make a LB score predictor</p></li>\n<li><p>I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.</p></li>\n<li><p>I want to run some of the kaggle solutions that uses CRNN and LSTM. I believe that is the correct way to do speech.</p></li>\n</ul>\n\n<hr>\n\n<h2>SEE SOLUTION</h2>\n\n<h1>Overview of my approach</h1>\n\n<p>I started with the provided <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">tutorial</a> and could easily get better results by just adding momentum to the plain SGD solver (82-83% on the leaderboard). I have no prior experience with audio data and mostly used deep learning with images. For this domain you don't use features but feed the raw pixel values. My thinking was that this should work with audio data as well. Throughout the competition I ran experiments using raw waveforms, spectrograms and log mel features as input. I got similar results using log mel and raw waveform (86%-87%) and used the waveform data for most experiments as it was easier to interpret for me.</p>\n\n<p>For the special price the restrictions were: the network is smaller than 5.000.000 bytes and runs in less than 175ms per sample on a stock Raspberry Pi 3. Regarding the size, this allows you to build networks that have roughly 1.250.000 weight parameters. So by experimenting with these restrictions I came up with an architecture that uses Depthwise1D convolutions on the raw waveform. Using <a href=\"https://arxiv.org/pdf/1503.02531.pdf\">model distillation</a> this network predicts the correct class for 90.8% of the private leaderboard samples and runs in roughly 80ms.</p>\n\n<h1>What didn't work</h1>\n\n<ul>\n<li><p>Fancy augmentation methods: I tried flipping (i.e: <code>* -1.0</code>) the samples. You can check that they will sound exactly the same. I also modified <code>input_data.py</code> to change the foreground and background volume independently and created a separate volume range for the silence samples. My validation accuracy improved for some experiments but my leaderboard scores didn't.</p></li>\n<li><p>Predicting unknown unknowns: I didn't find a good way to consistently predict these words. Often, similar words were wrongly classified (e.g. one as on).</p></li>\n<li><p>Creating new words: I trained some networks with even more classes. I reversed the samples from the known unwanted words, e.g. <code>bird</code>, <code>bed</code>, <code>marvin</code>, and created new classes (<code>bird</code> -&gt; <code>drib</code> ...). The idea was to have more unknowns to prevent the network from wrongly mapping unknowns to the known words. For example the word <code>follow</code> was mostly predicted as <code>off</code>. However, neither my validation score not my leaderboard score improved.</p></li>\n<li><p>Cyclic learning rate schedules: The winning entry of the <a href=\"http://blog.kaggle.com/2017/12/22/carvana-image-masking-first-place-interview/\">Caravana Image Masking Challenge</a> used cyclic learning rates but for me the results got worse and you had additional hyperparameters. Maybe I just didn't implement it correctly.</p></li>\n</ul>\n\n<h1>What worked</h1>\n\n<ul>\n<li><p>Mixing tensorflow and Keras: Both frameworks work perfectly together and you can mix them wherever you want. For example: I wrapped the provided data AudioProcessor from <code>input_data.py</code> in a generator and used it with <code>keras.models.Model.fit_generator</code>. This way, I could implement new architectures really fast using Keras and later just extract and freeze the graph from the trained models.</p></li>\n<li><p>Pseudo labeling: I used consistent samples from the test set to train new networks. Choosing them was based on a.) my three best models agree on this submission. I used this version at early stages of the competition. b.) using a probability threshold on the predicted softmax probabilities. Typically, using <code>pseudo_threshold=0.6</code> were the samples that our ensembled model predicted correctly. I also implemented a schedule for pseudo labels. That is: For the first 5 epochs you only use pseudo labels and then gradually mix in data from the training data set. Though, I didn't have time to run these experiments, so I kept a fixed ratio of training and pseudo data.</p></li>\n<li><p>Test time augmentation: It is a simple way to get some boost. Just augment the samples, feed them multiple times and average the probabilities. I tried the following: time-shifting, increase/decrease the volume and time-stretching using <code>librosa.effects.time_stretch</code>.</p></li>\n</ul>\n\n<hr>\n\n<p>I posted out submission results here, raw probability score is normalised from [0,1] to [0,255]</p>",
      "rawMarkdown": "## HENG SOLUTION ##\n\n my solution ]\n\n- Most of the details are already posted. We are making document and code clean up for prize submission. Will share with you guys later. In short summary:\n\n   - ensemble of about 30 models comprise of wave, log melspectrogram, mfccs. \n\n  - My models are weak in the range of 0.86. My teammate (@Ryan, @See) models are stronger,  in the range of  0.88\n\n   - mostly convolution networks\n\n   - pseudo labeling to train some of the networks (not all)\n\n\n[ my approach to this competition ]\n\n - Each kaggle competition is different. For some competitions, the winning factor could be:\n\n     -  feature engineering (or network design),  e.g. the carvana car segmentation\n\n     -  dealing with noisy label, e.g. amazon satellite image classification\n\n     -  dealing with large data+category  and efficiency, e.g cdiscount e-commerce product image classification\n\nMost of the time, it is combination of the above. For this competition, i think the main challenge is dealing with data domain shift, i.e. \"train+validation data\" and \"LB data\" are different. Why?\n\n   - the gap between validation score and LB score is large (12% to 8%)\n\n   - the trained model is sensitive to distribution of the training class. e.g if you just train with random sampling,  simple cnn_trad_pool2_net  gives less than 0.80 on public LB. But if you use balanced class sampling, it increases to 0.82\n\n   - the trained model is also sensitive type of silence train samples, amplitude of noise level etc.\n\n   - lastly i notice the unknowns in LB data is not the same as that in train+valid set\n\nI am not familiar with speech recognition or audio processing, hence i think it would be difficult for me to design good network. So I focus on the data instead. I used the simplest approach, create train data from the LB set.\n\nPseudo labeling may \"overfit\" and can be a dangerous approach. I need to make sure pseudo label LB samples are in correct label and correct distribution. The way to do is:\n\n   - let T = train set, V= validation set, L = LB set,  \n\n   - M = model trained on T\n\n   - P = subset of L, labelled by M  with label noise &lt; e.g  10%\n\n   - N = model trained on T+P\n\n   - Accept N if accuracy of N is better than M on both V and L \n\nyou can relax the last acceptance test. Note if P={empty}, you have the original results. You can modify P  (e.g. by different threshold like 5%,15%,20% etc ...) until you pass the acceptance test. We actually use different thresholds for different classes (silence, unknowns, knowns) to ensure that the pseudo-labelled set is more balance. \n\n[ some surprises of the competition ]\n\n -  We see ourselves in the top 5 but did not expect to be the winner. We had a bug and mistakenly used a model twice (due to typo error in code) . And this model has much higher weights than the rest. This bug is our winning submission private LB 0.91060\n(public LB 0.90296), which is also attached below.\n\nLater, it reveal that our highest private LB is 0.91107  (public LB 0.90241) is another ensemble. It is hard to make selection given only 2 decimals are revealed at the competition.\n \n -  1d wave input actually works!  \n\n - sometimes high score models does not improve when ensemble, especially for those greater than LB =0.88. I compare the confidence probability scores of weak (LB 0.86) and strong (LB 0.88) models. For strong model, the sample score are mostly very near to 1 or zero. So it is very hard to change the scores of test samples i think. \n \n\n\n[ how to go on ]\n\n- I take this learning path. Be a master in training (hyper parameter tuning + data augmentation). Then be a master in network design and then be a ensemble expert.\n\n- After the basics above,  i think semi/weak supervised learning is one way to go. From competition point of view, able to automatic label LB dataset and use it for training is very powerful.  (My next competition is National Science Bowl 2018, where i hope to use GAN  to generate \"LB data\" with label)\n\n- I want to make a LB score predictor\n\n- I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.\n\n- I want to run some of the kaggle solutions that uses CRNN and LSTM. I believe that is the correct way to do speech.\n\n---\n\n## SEE SOLUTION ##\n\n\n# Overview of my approach\nI started with the provided [tutorial](https://www.tensorflow.org/versions/master/tutorials/audio_recognition) and could easily get better results by just adding momentum to the plain SGD solver (82-83% on the leaderboard). I have no prior experience with audio data and mostly used deep learning with images. For this domain you don't use features but feed the raw pixel values. My thinking was that this should work with audio data as well. Throughout the competition I ran experiments using raw waveforms, spectrograms and log mel features as input. I got similar results using log mel and raw waveform (86%-87%) and used the waveform data for most experiments as it was easier to interpret for me.\n\nFor the special price the restrictions were: the network is smaller than 5.000.000 bytes and runs in less than 175ms per sample on a stock Raspberry Pi 3. Regarding the size, this allows you to build networks that have roughly 1.250.000 weight parameters. So by experimenting with these restrictions I came up with an architecture that uses Depthwise1D convolutions on the raw waveform. Using [model distillation](https://arxiv.org/pdf/1503.02531.pdf) this network predicts the correct class for 90.8% of the private leaderboard samples and runs in roughly 80ms.\n\n# What didn't work\n\n- Fancy augmentation methods: I tried flipping (i.e: ` * -1.0`) the samples. You can check that they will sound exactly the same. I also modified `input_data.py` to change the foreground and background volume independently and created a separate volume range for the silence samples. My validation accuracy improved for some experiments but my leaderboard scores didn't.\n\n- Predicting unknown unknowns: I didn't find a good way to consistently predict these words. Often, similar words were wrongly classified (e.g. one as on).\n\n- Creating new words: I trained some networks with even more classes. I reversed the samples from the known unwanted words, e.g. `bird`, `bed`, `marvin`, and created new classes (`bird` -&gt; `drib` ...). The idea was to have more unknowns to prevent the network from wrongly mapping unknowns to the known words. For example the word `follow` was mostly predicted as `off`. However, neither my validation score not my leaderboard score improved.\n\n- Cyclic learning rate schedules: The winning entry of the [Caravana Image Masking Challenge](http://blog.kaggle.com/2017/12/22/carvana-image-masking-first-place-interview/) used cyclic learning rates but for me the results got worse and you had additional hyperparameters. Maybe I just didn't implement it correctly.\n\n# What worked\n- Mixing tensorflow and Keras: Both frameworks work perfectly together and you can mix them wherever you want. For example: I wrapped the provided data AudioProcessor from `input_data.py` in a generator and used it with `keras.models.Model.fit_generator`. This way, I could implement new architectures really fast using Keras and later just extract and freeze the graph from the trained models.\n\n- Pseudo labeling: I used consistent samples from the test set to train new networks. Choosing them was based on a.) my three best models agree on this submission. I used this version at early stages of the competition. b.) using a probability threshold on the predicted softmax probabilities. Typically, using `pseudo_threshold=0.6` were the samples that our ensembled model predicted correctly. I also implemented a schedule for pseudo labels. That is: For the first 5 epochs you only use pseudo labels and then gradually mix in data from the training data set. Though, I didn't have time to run these experiments, so I kept a fixed ratio of training and pseudo data.\n\n- Test time augmentation: It is a simple way to get some boost. Just augment the samples, feed them multiple times and average the probabilities. I tried the following: time-shifting, increase/decrease the volume and time-stretching using `librosa.effects.time_stretch`.\n\n---\nI posted out submission results here, raw probability score is normalised from [0,1] to [0,255]",
      "votes": null
    },
    {
      "id": "270328",
      "postDate": "01/18/2018 04:58:27",
      "content": "<p>Congratulations Heng! Great post! Can you discuss on the emsembling techniques used  and the methods used to determine the emsembling weights ?</p>",
      "rawMarkdown": "Congratulations Heng! Great post! Can you discuss on the emsembling techniques used  and the methods used to determine the emsembling weights ?",
      "votes": null
    },
    {
      "id": "270331",
      "postDate": "01/18/2018 05:16:00",
      "content": "<p>Awesome, and thank you </p>",
      "rawMarkdown": "Awesome, and thank you",
      "votes": null
    },
    {
      "id": "270335",
      "postDate": "01/18/2018 05:24:53",
      "content": "<p>Congratulations Heng! Hope to see your GAN solution in the Data Science Bowl 2018.</p>",
      "rawMarkdown": "Congratulations Heng! Hope to see your GAN solution in the Data Science Bowl 2018.",
      "votes": null
    },
    {
      "id": "270354",
      "postDate": "01/18/2018 05:57:08",
      "content": "<p>Great job and Thanks o/</p>",
      "rawMarkdown": "Great job and Thanks o/",
      "votes": null
    },
    {
      "id": "270390",
      "postDate": "01/18/2018 07:14:00",
      "content": "<p>Heng, congratulations. You are an inspiration for a lot of participants here! Myself included :-)</p>\n\n<p>Re: </p>\n\n<blockquote>\n  <p>I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.</p>\n</blockquote>\n\n<p>Can you take a look @ <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47645\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47645</a> ? It's precisely what I tried and didn't work.</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Heng, congratulations. You are an inspiration for a lot of participants here! Myself included :-)\n\nRe: \n\n&gt; I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.\n\nCan you take a look @ https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47645 ? It's precisely what I tried and didn't work.\n\nThanks!",
      "votes": null
    },
    {
      "id": "270415",
      "postDate": "01/18/2018 07:36:30",
      "content": "<p>Congratulations Heng ! thank a lot for sharing... I would like to know how you treat the unknown class. I got around LB score 0.805 and my model predict wrong on unknown class mostly such as 'one to on' etc</p>",
      "rawMarkdown": "Congratulations Heng ! thank a lot for sharing... I would like to know how you treat the unknown class. I got around LB score 0.805 and my model predict wrong on unknown class mostly such as 'one to on' etc",
      "votes": null
    },
    {
      "id": "270437",
      "postDate": "01/18/2018 08:13:20",
      "content": "<p>Congrats! Another prize in a month! <br>\nPerhaps you can have a look at some competitions using RNN for time series,there are  many competitions of this type on kaggle recently and there are a lot of interesting solutions. <br>\nI found reading and learning beautiful solutions of finished competitions is so relaxing when we don't have so much time to enter a competition!</p>",
      "rawMarkdown": "Congrats! Another prize in a month! <br>\nPerhaps you can have a look at some competitions using RNN for time series,there are  many competitions of this type on kaggle recently and there are a lot of interesting solutions. <br>\nI found reading and learning beautiful solutions of finished competitions is so relaxing when we don't have so much time to enter a competition!",
      "votes": null
    },
    {
      "id": "270530",
      "postDate": "01/18/2018 12:36:13",
      "content": "<p>Heng, awesome work and insights!  Your openness and helpfulness is a true embodiment of the Kaggle spirit and #1 finish shows how you can do it all; very difficult to pull off.   Big congratulations to you, Ryan, and See!</p>",
      "rawMarkdown": "Heng, awesome work and insights!  Your openness and helpfulness is a true embodiment of the Kaggle spirit and #1 finish shows how you can do it all; very difficult to pull off.   Big congratulations to you, Ryan, and See!",
      "votes": null
    },
    {
      "id": "270648",
      "postDate": "01/18/2018 17:02:29",
      "content": "<p>Thanks! I am trying out time series solution too!</p>",
      "rawMarkdown": "Thanks! I am trying out time series solution too!",
      "votes": null
    },
    {
      "id": "270650",
      "postDate": "01/18/2018 17:08:04",
      "content": "<p>Congrats Heng and your team!</p>\n\n<p>You said you choose diffeeent thresholds for each class while preparing pseudo. Do You mean that instead of using simple idxmax, you scaled probs by threshold for each label and after that used idxmax? And the threshold were choosen, so that you had balenced core labels in test? </p>",
      "rawMarkdown": "Congrats Heng and your team!\n\nYou said you choose diffeeent thresholds for each class while preparing pseudo. Do You mean that instead of using simple idxmax, you scaled probs by threshold for each label and after that used idxmax? And the threshold were choosen, so that you had balenced core labels in test?",
      "votes": null
    },
    {
      "id": "270652",
      "postDate": "01/18/2018 17:12:57",
      "content": "<p>for each LB sample, we have p = [p0 p1 ... p11] probability values. We extract label=argmax(p) and confidence c = p[label]. </p>\n\n<p>Now for silence and unknown,  they are included in pseudo-label set only if  confidence c&gt;0.5</p>\n\n<p>For the known classes, we use confidence c&gt;0.8</p>\n\n<p>These values are different training new models to improve diversity.</p>",
      "rawMarkdown": "for each LB sample, we have p = [p0 p1 ... p11] probability values. We extract label=argmax(p) and confidence c = p[label]. \n\nNow for silence and unknown,  they are included in pseudo-label set only if  confidence c&gt;0.5\n\nFor the known classes, we use confidence c&gt;0.8\n\nThese values are different training new models to improve diversity.",
      "votes": null
    },
    {
      "id": "270661",
      "postDate": "01/18/2018 17:38:25",
      "content": "<p>Congrats, this is my first Kaggle competition. Learned a lot from you.</p>",
      "rawMarkdown": "Congrats, this is my first Kaggle competition. Learned a lot from you.",
      "votes": null
    },
    {
      "id": "275282",
      "postDate": "01/28/2018 16:35:45",
      "content": "<p>Thank for sharing !\nWill you upload your code with keras on github ? </p>",
      "rawMarkdown": "Thank for sharing !\nWill you upload your code with keras on github ?",
      "votes": null
    },
    {
      "id": "279938",
      "postDate": "02/09/2018 01:14:42",
      "content": "<p>the final ensemble for submission:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/279938/8533/overall.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/279938/8532/final.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "the final ensemble for submission:\n\n  ![enter image description here][1]\n\n\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/279938/8533/overall.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/279938/8532/final.png",
      "votes": null
    },
    {
      "id": "303423",
      "postDate": "03/26/2018 06:51:45",
      "content": "<p>Hi Heng, trying to set some solid foundations of pseudo labelling methods here, can I understand for the each stage of ensembling with psuedo-labels you removed the previous stage set of pseudo labels as well? first stage of predictions used 51088 data but after the first one there is 6798 data on top of the 51088 data and is used constantly at all the stages, is this pseudo-labelled test set data? Also, did you train each and every model with the new combination of pseudo-labelled+training data with a certain fixed threshold \"P\"  and after ensembling you validate the result to the LB and Validation data, if both increase then keep the predictions if not choose a new P and retrain all over again?</p>",
      "rawMarkdown": "Hi Heng, trying to set some solid foundations of pseudo labelling methods here, can I understand for the each stage of ensembling with psuedo-labels you removed the previous stage set of pseudo labels as well? first stage of predictions used 51088 data but after the first one there is 6798 data on top of the 51088 data and is used constantly at all the stages, is this pseudo-labelled test set data? Also, did you train each and every model with the new combination of pseudo-labelled+training data with a certain fixed threshold \"P\"  and after ensembling you validate the result to the LB and Validation data, if both increase then keep the predictions if not choose a new P and retrain all over again?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 270328,
      "author_name": "subho406",
      "author_url": "",
      "post_date": "01/18/2018 04:58:27",
      "content": "<p>Congratulations Heng! Great post! Can you discuss on the emsembling techniques used  and the methods used to determine the emsembling weights ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270331,
      "author_name": "hongta",
      "author_url": "",
      "post_date": "01/18/2018 05:16:00",
      "content": "<p>Awesome, and thank you </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270335,
      "author_name": "johnsondata",
      "author_url": "",
      "post_date": "01/18/2018 05:24:53",
      "content": "<p>Congratulations Heng! Hope to see your GAN solution in the Data Science Bowl 2018.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270354,
      "author_name": "edreams",
      "author_url": "",
      "post_date": "01/18/2018 05:57:08",
      "content": "<p>Great job and Thanks o/</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270390,
      "author_name": "antorsae",
      "author_url": "",
      "post_date": "01/18/2018 07:14:00",
      "content": "<p>Heng, congratulations. You are an inspiration for a lot of participants here! Myself included :-)</p>\n\n<p>Re: </p>\n\n<blockquote>\n  <p>I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.</p>\n</blockquote>\n\n<p>Can you take a look @ <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47645\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47645</a> ? It's precisely what I tried and didn't work.</p>\n\n<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270415,
      "author_name": "yeyinthtoon",
      "author_url": "",
      "post_date": "01/18/2018 07:36:30",
      "content": "<p>Congratulations Heng ! thank a lot for sharing... I would like to know how you treat the unknown class. I got around LB score 0.805 and my model predict wrong on unknown class mostly such as 'one to on' etc</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270437,
      "author_name": "bestfitting",
      "author_url": "",
      "post_date": "01/18/2018 08:13:20",
      "content": "<p>Congrats! Another prize in a month! <br>\nPerhaps you can have a look at some competitions using RNN for time series,there are  many competitions of this type on kaggle recently and there are a lot of interesting solutions. <br>\nI found reading and learning beautiful solutions of finished competitions is so relaxing when we don't have so much time to enter a competition!</p>",
      "votes": null,
      "replies": [
        {
          "id": 270648,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/18/2018 17:02:29",
          "content": "<p>Thanks! I am trying out time series solution too!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270530,
      "author_name": "sasrdw",
      "author_url": "",
      "post_date": "01/18/2018 12:36:13",
      "content": "<p>Heng, awesome work and insights!  Your openness and helpfulness is a true embodiment of the Kaggle spirit and #1 finish shows how you can do it all; very difficult to pull off.   Big congratulations to you, Ryan, and See!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 270650,
      "author_name": "heyt0ny",
      "author_url": "",
      "post_date": "01/18/2018 17:08:04",
      "content": "<p>Congrats Heng and your team!</p>\n\n<p>You said you choose diffeeent thresholds for each class while preparing pseudo. Do You mean that instead of using simple idxmax, you scaled probs by threshold for each label and after that used idxmax? And the threshold were choosen, so that you had balenced core labels in test? </p>",
      "votes": null,
      "replies": [
        {
          "id": 270652,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/18/2018 17:12:57",
          "content": "<p>for each LB sample, we have p = [p0 p1 ... p11] probability values. We extract label=argmax(p) and confidence c = p[label]. </p>\n\n<p>Now for silence and unknown,  they are included in pseudo-label set only if  confidence c&gt;0.5</p>\n\n<p>For the known classes, we use confidence c&gt;0.8</p>\n\n<p>These values are different training new models to improve diversity.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 270661,
      "author_name": "abnerchou",
      "author_url": "",
      "post_date": "01/18/2018 17:38:25",
      "content": "<p>Congrats, this is my first Kaggle competition. Learned a lot from you.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 275282,
      "author_name": "ludovick",
      "author_url": "",
      "post_date": "01/28/2018 16:35:45",
      "content": "<p>Thank for sharing !\nWill you upload your code with keras on github ? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 279938,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/09/2018 01:14:42",
      "content": "<p>the final ensemble for submission:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/279938/8533/overall.png\" alt=\"enter image description here\"></p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/279938/8532/final.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 303423,
          "author_name": "dicksonchin93",
          "author_url": "",
          "post_date": "03/26/2018 06:51:45",
          "content": "<p>Hi Heng, trying to set some solid foundations of pseudo labelling methods here, can I understand for the each stage of ensembling with psuedo-labels you removed the previous stage set of pseudo labels as well? first stage of predictions used 51088 data but after the first one there is 6798 data on top of the 51088 data and is used constantly at all the stages, is this pseudo-labelled test set data? Also, did you train each and every model with the new combination of pseudo-labelled+training data with a certain fixed threshold \"P\"  and after ensembling you validate the result to the LB and Validation data, if both increase then keep the predictions if not choose a new P and retrain all over again?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "270321": "## HENG SOLUTION ##\n\n my solution ]\n\n- Most of the details are already posted. We are making document and code clean up for prize submission. Will share with you guys later. In short summary:\n\n   - ensemble of about 30 models comprise of wave, log melspectrogram, mfccs. \n\n  - My models are weak in the range of 0.86. My teammate (@Ryan, @See) models are stronger,  in the range of  0.88\n\n   - mostly convolution networks\n\n   - pseudo labeling to train some of the networks (not all)\n\n\n[ my approach to this competition ]\n\n - Each kaggle competition is different. For some competitions, the winning factor could be:\n\n     -  feature engineering (or network design),  e.g. the carvana car segmentation\n\n     -  dealing with noisy label, e.g. amazon satellite image classification\n\n     -  dealing with large data+category  and efficiency, e.g cdiscount e-commerce product image classification\n\nMost of the time, it is combination of the above. For this competition, i think the main challenge is dealing with data domain shift, i.e. \"train+validation data\" and \"LB data\" are different. Why?\n\n   - the gap between validation score and LB score is large (12% to 8%)\n\n   - the trained model is sensitive to distribution of the training class. e.g if you just train with random sampling,  simple cnn_trad_pool2_net  gives less than 0.80 on public LB. But if you use balanced class sampling, it increases to 0.82\n\n   - the trained model is also sensitive type of silence train samples, amplitude of noise level etc.\n\n   - lastly i notice the unknowns in LB data is not the same as that in train+valid set\n\nI am not familiar with speech recognition or audio processing, hence i think it would be difficult for me to design good network. So I focus on the data instead. I used the simplest approach, create train data from the LB set.\n\nPseudo labeling may \"overfit\" and can be a dangerous approach. I need to make sure pseudo label LB samples are in correct label and correct distribution. The way to do is:\n\n   - let T = train set, V= validation set, L = LB set,  \n\n   - M = model trained on T\n\n   - P = subset of L, labelled by M  with label noise &lt; e.g  10%\n\n   - N = model trained on T+P\n\n   - Accept N if accuracy of N is better than M on both V and L \n\nyou can relax the last acceptance test. Note if P={empty}, you have the original results. You can modify P  (e.g. by different threshold like 5%,15%,20% etc ...) until you pass the acceptance test. We actually use different thresholds for different classes (silence, unknowns, knowns) to ensure that the pseudo-labelled set is more balance. \n\n[ some surprises of the competition ]\n\n -  We see ourselves in the top 5 but did not expect to be the winner. We had a bug and mistakenly used a model twice (due to typo error in code) . And this model has much higher weights than the rest. This bug is our winning submission private LB 0.91060\n(public LB 0.90296), which is also attached below.\n\nLater, it reveal that our highest private LB is 0.91107  (public LB 0.90241) is another ensemble. It is hard to make selection given only 2 decimals are revealed at the competition.\n \n -  1d wave input actually works!  \n\n - sometimes high score models does not improve when ensemble, especially for those greater than LB =0.88. I compare the confidence probability scores of weak (LB 0.86) and strong (LB 0.88) models. For strong model, the sample score are mostly very near to 1 or zero. So it is very hard to change the scores of test samples i think. \n \n\n\n[ how to go on ]\n\n- I take this learning path. Be a master in training (hyper parameter tuning + data augmentation). Then be a master in network design and then be a ensemble expert.\n\n- After the basics above,  i think semi/weak supervised learning is one way to go. From competition point of view, able to automatic label LB dataset and use it for training is very powerful.  (My next competition is National Science Bowl 2018, where i hope to use GAN  to generate \"LB data\" with label)\n\n- I want to make a LB score predictor\n\n- I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.\n\n- I want to run some of the kaggle solutions that uses CRNN and LSTM. I believe that is the correct way to do speech.\n\n---\n\n## SEE SOLUTION ##\n\n\n# Overview of my approach\nI started with the provided [tutorial](https://www.tensorflow.org/versions/master/tutorials/audio_recognition) and could easily get better results by just adding momentum to the plain SGD solver (82-83% on the leaderboard). I have no prior experience with audio data and mostly used deep learning with images. For this domain you don't use features but feed the raw pixel values. My thinking was that this should work with audio data as well. Throughout the competition I ran experiments using raw waveforms, spectrograms and log mel features as input. I got similar results using log mel and raw waveform (86%-87%) and used the waveform data for most experiments as it was easier to interpret for me.\n\nFor the special price the restrictions were: the network is smaller than 5.000.000 bytes and runs in less than 175ms per sample on a stock Raspberry Pi 3. Regarding the size, this allows you to build networks that have roughly 1.250.000 weight parameters. So by experimenting with these restrictions I came up with an architecture that uses Depthwise1D convolutions on the raw waveform. Using [model distillation](https://arxiv.org/pdf/1503.02531.pdf) this network predicts the correct class for 90.8% of the private leaderboard samples and runs in roughly 80ms.\n\n# What didn't work\n\n- Fancy augmentation methods: I tried flipping (i.e: ` * -1.0`) the samples. You can check that they will sound exactly the same. I also modified `input_data.py` to change the foreground and background volume independently and created a separate volume range for the silence samples. My validation accuracy improved for some experiments but my leaderboard scores didn't.\n\n- Predicting unknown unknowns: I didn't find a good way to consistently predict these words. Often, similar words were wrongly classified (e.g. one as on).\n\n- Creating new words: I trained some networks with even more classes. I reversed the samples from the known unwanted words, e.g. `bird`, `bed`, `marvin`, and created new classes (`bird` -&gt; `drib` ...). The idea was to have more unknowns to prevent the network from wrongly mapping unknowns to the known words. For example the word `follow` was mostly predicted as `off`. However, neither my validation score not my leaderboard score improved.\n\n- Cyclic learning rate schedules: The winning entry of the [Caravana Image Masking Challenge](http://blog.kaggle.com/2017/12/22/carvana-image-masking-first-place-interview/) used cyclic learning rates but for me the results got worse and you had additional hyperparameters. Maybe I just didn't implement it correctly.\n\n# What worked\n- Mixing tensorflow and Keras: Both frameworks work perfectly together and you can mix them wherever you want. For example: I wrapped the provided data AudioProcessor from `input_data.py` in a generator and used it with `keras.models.Model.fit_generator`. This way, I could implement new architectures really fast using Keras and later just extract and freeze the graph from the trained models.\n\n- Pseudo labeling: I used consistent samples from the test set to train new networks. Choosing them was based on a.) my three best models agree on this submission. I used this version at early stages of the competition. b.) using a probability threshold on the predicted softmax probabilities. Typically, using `pseudo_threshold=0.6` were the samples that our ensembled model predicted correctly. I also implemented a schedule for pseudo labels. That is: For the first 5 epochs you only use pseudo labels and then gradually mix in data from the training data set. Though, I didn't have time to run these experiments, so I kept a fixed ratio of training and pseudo data.\n\n- Test time augmentation: It is a simple way to get some boost. Just augment the samples, feed them multiple times and average the probabilities. I tried the following: time-shifting, increase/decrease the volume and time-stretching using `librosa.effects.time_stretch`.\n\n---\nI posted out submission results here, raw probability score is normalised from [0,1] to [0,255]",
    "270328": "Congratulations Heng! Great post! Can you discuss on the emsembling techniques used  and the methods used to determine the emsembling weights ?",
    "270331": "Awesome, and thank you",
    "270335": "Congratulations Heng! Hope to see your GAN solution in the Data Science Bowl 2018.",
    "270354": "Great job and Thanks o/",
    "270390": "Heng, congratulations. You are an inspiration for a lot of participants here! Myself included :-)\n\nRe: \n\n&gt; I want to make a better way to determine the ensemble weights, e.g. formulate the ensemble weights based on score distribution(which is a rough indication of error. err = 1-max(P_i) ). Maybe I can refer to boosting.\n\nCan you take a look @ https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/47645 ? It's precisely what I tried and didn't work.\n\nThanks!",
    "270415": "Congratulations Heng ! thank a lot for sharing... I would like to know how you treat the unknown class. I got around LB score 0.805 and my model predict wrong on unknown class mostly such as 'one to on' etc",
    "270437": "Congrats! Another prize in a month! <br>\nPerhaps you can have a look at some competitions using RNN for time series,there are  many competitions of this type on kaggle recently and there are a lot of interesting solutions. <br>\nI found reading and learning beautiful solutions of finished competitions is so relaxing when we don't have so much time to enter a competition!",
    "270530": "Heng, awesome work and insights!  Your openness and helpfulness is a true embodiment of the Kaggle spirit and #1 finish shows how you can do it all; very difficult to pull off.   Big congratulations to you, Ryan, and See!",
    "270648": "Thanks! I am trying out time series solution too!",
    "270650": "Congrats Heng and your team!\n\nYou said you choose diffeeent thresholds for each class while preparing pseudo. Do You mean that instead of using simple idxmax, you scaled probs by threshold for each label and after that used idxmax? And the threshold were choosen, so that you had balenced core labels in test?",
    "270652": "for each LB sample, we have p = [p0 p1 ... p11] probability values. We extract label=argmax(p) and confidence c = p[label]. \n\nNow for silence and unknown,  they are included in pseudo-label set only if  confidence c&gt;0.5\n\nFor the known classes, we use confidence c&gt;0.8\n\nThese values are different training new models to improve diversity.",
    "270661": "Congrats, this is my first Kaggle competition. Learned a lot from you.",
    "275282": "Thank for sharing !\nWill you upload your code with keras on github ?",
    "279938": "the final ensemble for submission:\n\n  ![enter image description here][1]\n\n\n  ![enter image description here][2]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/279938/8533/overall.png\n  [2]: https://kaggle2.blob.core.windows.net/forum-message-attachments/279938/8532/final.png",
    "303423": "Hi Heng, trying to set some solid foundations of pseudo labelling methods here, can I understand for the each stage of ensembling with psuedo-labels you removed the previous stage set of pseudo labels as well? first stage of predictions used 51088 data but after the first one there is 6798 data on top of the 51088 data and is used constantly at all the stages, is this pseudo-labelled test set data? Also, did you train each and every model with the new combination of pseudo-labelled+training data with a certain fixed threshold \"P\"  and after ensembling you validate the result to the LB and Validation data, if both increase then keep the predictions if not choose a new P and retrain all over again?"
  },
  "source": "meta"
}