{
  "id": 47194,
  "title": "moving from 0.90 and beyond - competition tricks or black magic ",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47194",
  "author_name": "",
  "post_date": "2018-01-10T03:10:46.048457700Z",
  "votes": 28,
  "comment_count": 21,
  "views": 0,
  "content": "<p>You can read about ideas like ensembling and pseudo-labeling from blogs and papers. Apart from the usual stuffs like the implementation steps and the theory of the methods, there are a few things that you can take note of for \"data science competition\".</p>\n\n<p>[pseudo-labeling]</p>\n\n<ul>\n<li><p>what it IS NOT:</p>\n\n<ul><li>just label LB samples and train</li></ul></li>\n<li><p>what it IS:</p>\n\n<ul><li><p>estimate the error of the LB pseudo samples. In the worst case, if you learned wrongly, how much error you will incur? How to estimate this. It is the old trick: either simulation with your train and validation samples or  \"probe\" the LB samples. </p></li>\n<li><p>(to be updated) ...</p></li></ul></li>\n</ul>\n\n<p>[ensemble weights]</p>\n\n<ul>\n<li><p>what it IS NOT:</p>\n\n<ul><li>just sum up results from each individual models</li></ul></li>\n<li><p>what it IS:</p>\n\n<ul><li><p>out of all LB test samples you have, which are the ones that are most likely to be improved if you ensemble. How  to identify them? e.g. by scores (maybe samples that have score in 0.5 range have better chance of being improved than those with score &gt;0.9 or &lt;0.1). e.g. by class (maybe those predicted to be xx has better chance of being improved).  How much score must we change then, e.g. +0.05 +0.1 .... this is how you can set your weights. How to estimate the improvements?It is again the same old trick: either simulation with your train and validation samples or  \"probe\" the LB samples!</p></li>\n<li><p>(to be updated) ...</p></li></ul></li>\n</ul>\n\n<p>of course what i described is\"manual tuning\". Based on the same idea,  you can design automatic algorithms to make it work better.</p>\n\n<p>lastly, there is  always a \"trend\", when you do things right. try to improve in  that direction. e.g. see <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/47062\">https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/47062</a></p>\n\n<p>Have fun!</p>",
  "messages": [
    {
      "id": "266871",
      "postDate": "01/10/2018 03:10:46",
      "content": "<p>You can read about ideas like ensembling and pseudo-labeling from blogs and papers. Apart from the usual stuffs like the implementation steps and the theory of the methods, there are a few things that you can take note of for \"data science competition\".</p>\n\n<p>[pseudo-labeling]</p>\n\n<ul>\n<li><p>what it IS NOT:</p>\n\n<ul><li>just label LB samples and train</li></ul></li>\n<li><p>what it IS:</p>\n\n<ul><li><p>estimate the error of the LB pseudo samples. In the worst case, if you learned wrongly, how much error you will incur? How to estimate this. It is the old trick: either simulation with your train and validation samples or  \"probe\" the LB samples. </p></li>\n<li><p>(to be updated) ...</p></li></ul></li>\n</ul>\n\n<p>[ensemble weights]</p>\n\n<ul>\n<li><p>what it IS NOT:</p>\n\n<ul><li>just sum up results from each individual models</li></ul></li>\n<li><p>what it IS:</p>\n\n<ul><li><p>out of all LB test samples you have, which are the ones that are most likely to be improved if you ensemble. How  to identify them? e.g. by scores (maybe samples that have score in 0.5 range have better chance of being improved than those with score &gt;0.9 or &lt;0.1). e.g. by class (maybe those predicted to be xx has better chance of being improved).  How much score must we change then, e.g. +0.05 +0.1 .... this is how you can set your weights. How to estimate the improvements?It is again the same old trick: either simulation with your train and validation samples or  \"probe\" the LB samples!</p></li>\n<li><p>(to be updated) ...</p></li></ul></li>\n</ul>\n\n<p>of course what i described is\"manual tuning\". Based on the same idea,  you can design automatic algorithms to make it work better.</p>\n\n<p>lastly, there is  always a \"trend\", when you do things right. try to improve in  that direction. e.g. see <a href=\"https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/47062\">https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/47062</a></p>\n\n<p>Have fun!</p>",
      "rawMarkdown": "You can read about ideas like ensembling and pseudo-labeling from blogs and papers. Apart from the usual stuffs like the implementation steps and the theory of the methods, there are a few things that you can take note of for \"data science competition\".\n\n \n[pseudo-labeling]\n\n - what it IS NOT:\n\n  -  just label LB samples and train\n\n\n - what it IS:\n\n  -  estimate the error of the LB pseudo samples. In the worst case, if you learned wrongly, how much error you will incur? How to estimate this. It is the old trick: either simulation with your train and validation samples or  \"probe\" the LB samples. \n\n  - (to be updated) ...\n\n[ensemble weights]\n\n - what it IS NOT:\n\n  -  just sum up results from each individual models\n\n - what it IS:\n\n   -  out of all LB test samples you have, which are the ones that are most likely to be improved if you ensemble. How  to identify them? e.g. by scores (maybe samples that have score in 0.5 range have better chance of being improved than those with score &gt;0.9 or &lt;0.1). e.g. by class (maybe those predicted to be xx has better chance of being improved).  How much score must we change then, e.g. +0.05 +0.1 .... this is how you can set your weights. How to estimate the improvements?It is again the same old trick: either simulation with your train and validation samples or  \"probe\" the LB samples!\n\n  - (to be updated) ...\n\nof course what i described is\"manual tuning\". Based on the same idea,  you can design automatic algorithms to make it work better.\n\nlastly, there is  always a \"trend\", when you do things right. try to improve in  that direction. e.g. see https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/47062\n\nHave fun!",
      "votes": null
    },
    {
      "id": "267031",
      "postDate": "01/10/2018 11:59:00",
      "content": "<p>How much time does it take to train a single model? Do you also parallelize on multiple gpus?</p>",
      "rawMarkdown": "How much time does it take to train a single model? Do you also parallelize on multiple gpus?",
      "votes": null
    },
    {
      "id": "267104",
      "postDate": "01/10/2018 15:34:47",
      "content": "<p>open to your own interpretation!</p>",
      "rawMarkdown": "open to your own interpretation!",
      "votes": null
    },
    {
      "id": "267105",
      "postDate": "01/10/2018 15:35:19",
      "content": "<blockquote>\n  <blockquote>\n    <p>How much time does it take to train a single model?\n    2 to 4 hr.</p>\n    \n    <p>Do you also parallelize on multiple gpus?\n    No</p>\n  </blockquote>\n</blockquote>",
      "rawMarkdown": "&gt;&gt;How much time does it take to train a single model?\n2 to 4 hr.\n\n&gt;&gt;Do you also parallelize on multiple gpus?\nNo",
      "votes": null
    },
    {
      "id": "267147",
      "postDate": "01/10/2018 17:16:44",
      "content": "<p>yet another observation</p>",
      "rawMarkdown": "yet another observation",
      "votes": null
    },
    {
      "id": "267171",
      "postDate": "01/10/2018 18:24:58",
      "content": "<p>I don't have time now but here is one paper that might be relevant when training with noisy labels\n<a href=\"https://arxiv.org/abs/1412.6596\">https://arxiv.org/abs/1412.6596</a></p>",
      "rawMarkdown": "I don't have time now but here is one paper that might be relevant when training with noisy labels\nhttps://arxiv.org/abs/1412.6596",
      "votes": null
    },
    {
      "id": "267328",
      "postDate": "01/11/2018 02:53:46",
      "content": "<p>Dose pseudo-test24b-0.6 mean pseudo data take 60% of the new train data set(original train data+pseudo data)?  Additionally , how do you generate the pseudo data set,  the NN models which get 0.86 LB or some transitional machine learning classifier (Decision Tree, XGBoost)? </p>",
      "rawMarkdown": "Dose pseudo-test24b-0.6 mean pseudo data take 60% of the new train data set(original train data+pseudo data)?  Additionally , how do you generate the pseudo data set,  the NN models which get 0.86 LB or some transitional machine learning classifier (Decision Tree, XGBoost)?",
      "votes": null
    },
    {
      "id": "267371",
      "postDate": "01/11/2018 06:54:55",
      "content": "<p>I think that achieving scores higher than 0.9 is possible, but a different and fundamental aproach is needed. </p>\n\n<p>It is clear that the test set is fundamentally different to the train set. It has different words and some of they are very tricky:</p>\n\n<ul>\n<li>down vs town</li>\n<li>left vs learn</li>\n<li>off, on vs oh</li>\n</ul>\n\n<p>This explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset. </p>",
      "rawMarkdown": "I think that achieving scores higher than 0.9 is possible, but a different and fundamental aproach is needed. \n\nIt is clear that the test set is fundamentally different to the train set. It has different words and some of they are very tricky:\n\n* down vs town\n* left vs learn\n* off, on vs oh\n\nThis explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset.",
      "votes": null
    },
    {
      "id": "267424",
      "postDate": "01/11/2018 10:19:03",
      "content": "<p>another observation</p>",
      "rawMarkdown": "another observation",
      "votes": null
    },
    {
      "id": "267538",
      "postDate": "01/11/2018 17:04:48",
      "content": "<p>you are correct. there are some solution:</p>\n\n<p>1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n   problem is that you need to align the wave.</p>\n\n<p>2) for my team, we try to sample from LB test set from pseudo-labeling</p>",
      "rawMarkdown": "you are correct. there are some solution:\n\n1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n   problem is that you need to align the wave.\n\n2) for my team, we try to sample from LB test set from pseudo-labeling",
      "votes": null
    },
    {
      "id": "267699",
      "postDate": "01/12/2018 04:26:21",
      "content": "<p><a href=\"http://benanne.github.io/2014/08/05/spotify-cnns.html\">http://benanne.github.io/2014/08/05/spotify-cnns.html</a></p>",
      "rawMarkdown": "http://benanne.github.io/2014/08/05/spotify-cnns.html",
      "votes": null
    },
    {
      "id": "267701",
      "postDate": "01/12/2018 04:27:29",
      "content": "<p><a href=\"https://arxiv.org/pdf/1709.01922.pdf\">https://arxiv.org/pdf/1709.01922.pdf</a></p>",
      "rawMarkdown": "https://arxiv.org/pdf/1709.01922.pdf",
      "votes": null
    },
    {
      "id": "267703",
      "postDate": "01/12/2018 04:29:19",
      "content": "<p><a href=\"http://ismir2015.uma.es/articles/264_Paper.pdf\">http://ismir2015.uma.es/articles/264_Paper.pdf</a></p>",
      "rawMarkdown": "http://ismir2015.uma.es/articles/264_Paper.pdf",
      "votes": null
    },
    {
      "id": "267743",
      "postDate": "01/12/2018 06:54:06",
      "content": "<blockquote>\n  <p><strong>Heng CherKeng wrote</strong></p>\n  \n  <blockquote>\n    <p>you are correct. there are some solution:</p>\n  </blockquote>\n  \n  <p>1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n     problem is that you need to align the wave.</p>\n</blockquote>\n\n<p>I like your idea. One option is to build a dictionary with different phonemes(having multiple audios for each phoneme) Then combine them randomly to create the desired words. <br>\nGathering the phonemes could be tedious but it could worth the effort.</p>",
      "rawMarkdown": "&gt; **Heng CherKeng wrote**\n&gt; \n&gt; &gt; you are correct. there are some solution:\n&gt; \n&gt; 1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n&gt;    problem is that you need to align the wave.\n\nI like your idea. One option is to build a dictionary with different phonemes(having multiple audios for each phoneme) Then combine them randomly to create the desired words.  \nGathering the phonemes could be tedious but it could worth the effort.",
      "votes": null
    },
    {
      "id": "267773",
      "postDate": "01/12/2018 08:52:49",
      "content": "<p>I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. </p>\n\n<p>Anyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).</p>\n\n<p>For me, pseudo labelling the test set was the only thing that helped.</p>",
      "rawMarkdown": "I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. \n\nAnyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).\n\nFor me, pseudo labelling the test set was the only thing that helped.",
      "votes": null
    },
    {
      "id": "267811",
      "postDate": "01/12/2018 11:33:59",
      "content": "<blockquote>\n  <p><strong>fergusoci wrote</strong></p>\n  \n  <blockquote>\n    <p>I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. </p>\n  </blockquote>\n  \n  <p>Anyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).</p>\n  \n  <p>For me, pseudo labelling the test set was the only thing that helped.</p>\n</blockquote>\n\n<p>Thanks for the info, how many new audios did you made? Maybe it's a problem of having a big number of them.</p>",
      "rawMarkdown": "&gt; **fergusoci wrote**\n&gt; \n&gt; &gt; I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. \n&gt; \n&gt; Anyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).\n&gt; \n&gt; For me, pseudo labelling the test set was the only thing that helped.\n\nThanks for the info, how many new audios did you made? Maybe it's a problem of having a big number of them.",
      "votes": null
    },
    {
      "id": "267834",
      "postDate": "01/12/2018 13:21:04",
      "content": "<p>Perhaps, yeah. I had generated 20k samples. I was subsampling unknowns during the model build (set to randomly pick 10% of the samples on each epoch), but maybe there was too many of the synthetic unknowns in comparison to the regular unknowns. </p>",
      "rawMarkdown": "Perhaps, yeah. I had generated 20k samples. I was subsampling unknowns during the model build (set to randomly pick 10% of the samples on each epoch), but maybe there was too many of the synthetic unknowns in comparison to the regular unknowns.",
      "votes": null
    },
    {
      "id": "267882",
      "postDate": "01/12/2018 16:21:29",
      "content": "<p>sophisticated way is to GAN to create samples that resembles those in the LB set. But this is beyond the scope of the competition.</p>\n\n<p>E.g. </p>\n\n<p>Generator (train_sample1,train_sample2) --&gt; fake_sample</p>\n\n<p>fake_sample /LB_sample --&gt; Discrminator --&gt; 0/1</p>\n\n<p>If Generator(train_sample1,train_sample2) = encode() --&gt; decode() --&gt; fake_sample,</p>\n\n<p>then we can deocde(LB_sample ) --&gt; latent components</p>",
      "rawMarkdown": "sophisticated way is to GAN to create samples that resembles those in the LB set. But this is beyond the scope of the competition.\n\nE.g. \n\nGenerator (train_sample1,train_sample2) --&gt; fake_sample\n\nfake_sample /LB_sample --&gt; Discrminator --&gt; 0/1\n\nIf Generator(train_sample1,train_sample2) = encode() --&gt; decode() --&gt; fake_sample,\n\nthen we can deocde(LB_sample ) --&gt; latent components",
      "votes": null
    },
    {
      "id": "268092",
      "postDate": "01/13/2018 09:18:30",
      "content": "<p>level-2 classifier. it achieve LB 0.88 without pusedo-labeling</p>",
      "rawMarkdown": "level-2 classifier. it achieve LB 0.88 without pusedo-labeling",
      "votes": null
    },
    {
      "id": "268150",
      "postDate": "01/13/2018 14:53:52",
      "content": "<p>the most confusing samples</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/268150/8255/confusion.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "the most confusing samples\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/268150/8255/confusion.png",
      "votes": null
    },
    {
      "id": "268171",
      "postDate": "01/13/2018 16:22:53",
      "content": "<p>you can try this:</p>\n\n<pre><code>if there_is_no_noise and \\\n           amplitude_is_loud_enough and \\\n           predicted_label_is_one_of_knowns and \\\n           max_score_is_not_high_as_it_should_be :\n\n           predicted_label = 'unknown'  #change\n</code></pre>\n\n<p>it is because given all these good conditions, the prediction score should be high. If not, it is probably one of the confusing words. Now is how to set and detect the conditions? </p>",
      "rawMarkdown": "you can try this:\n\n    if there_is_no_noise and \\\n               amplitude_is_loud_enough and \\\n               predicted_label_is_one_of_knowns and \\\n               max_score_is_not_high_as_it_should_be :\n   \n               predicted_label = 'unknown'  #change\n\n\nit is because given all these good conditions, the prediction score should be high. If not, it is probably one of the confusing words. Now is how to set and detect the conditions?",
      "votes": null
    },
    {
      "id": "268179",
      "postDate": "01/13/2018 16:35:00",
      "content": "<p>updated slides:</p>",
      "rawMarkdown": "updated slides:",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 267031,
      "author_name": "",
      "author_url": "",
      "post_date": "01/10/2018 11:59:00",
      "content": "<p>How much time does it take to train a single model? Do you also parallelize on multiple gpus?</p>",
      "votes": null,
      "replies": [
        {
          "id": 267105,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/10/2018 15:35:19",
          "content": "<blockquote>\n  <blockquote>\n    <p>How much time does it take to train a single model?\n    2 to 4 hr.</p>\n    \n    <p>Do you also parallelize on multiple gpus?\n    No</p>\n  </blockquote>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267104,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/10/2018 15:34:47",
      "content": "<p>open to your own interpretation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267147,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/10/2018 17:16:44",
      "content": "<p>yet another observation</p>",
      "votes": null,
      "replies": [
        {
          "id": 267328,
          "author_name": "johnsondata",
          "author_url": "",
          "post_date": "01/11/2018 02:53:46",
          "content": "<p>Dose pseudo-test24b-0.6 mean pseudo data take 60% of the new train data set(original train data+pseudo data)?  Additionally , how do you generate the pseudo data set,  the NN models which get 0.86 LB or some transitional machine learning classifier (Decision Tree, XGBoost)? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267171,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "01/10/2018 18:24:58",
      "content": "<p>I don't have time now but here is one paper that might be relevant when training with noisy labels\n<a href=\"https://arxiv.org/abs/1412.6596\">https://arxiv.org/abs/1412.6596</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267371,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "01/11/2018 06:54:55",
      "content": "<p>I think that achieving scores higher than 0.9 is possible, but a different and fundamental aproach is needed. </p>\n\n<p>It is clear that the test set is fundamentally different to the train set. It has different words and some of they are very tricky:</p>\n\n<ul>\n<li>down vs town</li>\n<li>left vs learn</li>\n<li>off, on vs oh</li>\n</ul>\n\n<p>This explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset. </p>",
      "votes": null,
      "replies": [
        {
          "id": 267538,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/11/2018 17:04:48",
          "content": "<p>you are correct. there are some solution:</p>\n\n<p>1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n   problem is that you need to align the wave.</p>\n\n<p>2) for my team, we try to sample from LB test set from pseudo-labeling</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267743,
          "author_name": "ironbar",
          "author_url": "",
          "post_date": "01/12/2018 06:54:06",
          "content": "<blockquote>\n  <p><strong>Heng CherKeng wrote</strong></p>\n  \n  <blockquote>\n    <p>you are correct. there are some solution:</p>\n  </blockquote>\n  \n  <p>1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n     problem is that you need to align the wave.</p>\n</blockquote>\n\n<p>I like your idea. One option is to build a dictionary with different phonemes(having multiple audios for each phoneme) Then combine them randomly to create the desired words. <br>\nGathering the phonemes could be tedious but it could worth the effort.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267773,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "01/12/2018 08:52:49",
          "content": "<p>I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. </p>\n\n<p>Anyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).</p>\n\n<p>For me, pseudo labelling the test set was the only thing that helped.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267811,
          "author_name": "ironbar",
          "author_url": "",
          "post_date": "01/12/2018 11:33:59",
          "content": "<blockquote>\n  <p><strong>fergusoci wrote</strong></p>\n  \n  <blockquote>\n    <p>I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. </p>\n  </blockquote>\n  \n  <p>Anyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).</p>\n  \n  <p>For me, pseudo labelling the test set was the only thing that helped.</p>\n</blockquote>\n\n<p>Thanks for the info, how many new audios did you made? Maybe it's a problem of having a big number of them.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267834,
          "author_name": "fergusoci",
          "author_url": "",
          "post_date": "01/12/2018 13:21:04",
          "content": "<p>Perhaps, yeah. I had generated 20k samples. I was subsampling unknowns during the model build (set to randomly pick 10% of the samples on each epoch), but maybe there was too many of the synthetic unknowns in comparison to the regular unknowns. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267882,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/12/2018 16:21:29",
          "content": "<p>sophisticated way is to GAN to create samples that resembles those in the LB set. But this is beyond the scope of the competition.</p>\n\n<p>E.g. </p>\n\n<p>Generator (train_sample1,train_sample2) --&gt; fake_sample</p>\n\n<p>fake_sample /LB_sample --&gt; Discrminator --&gt; 0/1</p>\n\n<p>If Generator(train_sample1,train_sample2) = encode() --&gt; decode() --&gt; fake_sample,</p>\n\n<p>then we can deocde(LB_sample ) --&gt; latent components</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 268171,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/13/2018 16:22:53",
          "content": "<p>you can try this:</p>\n\n<pre><code>if there_is_no_noise and \\\n           amplitude_is_loud_enough and \\\n           predicted_label_is_one_of_knowns and \\\n           max_score_is_not_high_as_it_should_be :\n\n           predicted_label = 'unknown'  #change\n</code></pre>\n\n<p>it is because given all these good conditions, the prediction score should be high. If not, it is probably one of the confusing words. Now is how to set and detect the conditions? </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267424,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/11/2018 10:19:03",
      "content": "<p>another observation</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267699,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/12/2018 04:26:21",
      "content": "<p><a href=\"http://benanne.github.io/2014/08/05/spotify-cnns.html\">http://benanne.github.io/2014/08/05/spotify-cnns.html</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267701,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/12/2018 04:27:29",
      "content": "<p><a href=\"https://arxiv.org/pdf/1709.01922.pdf\">https://arxiv.org/pdf/1709.01922.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 267703,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/12/2018 04:29:19",
      "content": "<p><a href=\"http://ismir2015.uma.es/articles/264_Paper.pdf\">http://ismir2015.uma.es/articles/264_Paper.pdf</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 268092,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/13/2018 09:18:30",
      "content": "<p>level-2 classifier. it achieve LB 0.88 without pusedo-labeling</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 268150,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/13/2018 14:53:52",
      "content": "<p>the most confusing samples</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/268150/8255/confusion.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 268179,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/13/2018 16:35:00",
      "content": "<p>updated slides:</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "266871": "You can read about ideas like ensembling and pseudo-labeling from blogs and papers. Apart from the usual stuffs like the implementation steps and the theory of the methods, there are a few things that you can take note of for \"data science competition\".\n\n \n[pseudo-labeling]\n\n - what it IS NOT:\n\n  -  just label LB samples and train\n\n\n - what it IS:\n\n  -  estimate the error of the LB pseudo samples. In the worst case, if you learned wrongly, how much error you will incur? How to estimate this. It is the old trick: either simulation with your train and validation samples or  \"probe\" the LB samples. \n\n  - (to be updated) ...\n\n[ensemble weights]\n\n - what it IS NOT:\n\n  -  just sum up results from each individual models\n\n - what it IS:\n\n   -  out of all LB test samples you have, which are the ones that are most likely to be improved if you ensemble. How  to identify them? e.g. by scores (maybe samples that have score in 0.5 range have better chance of being improved than those with score &gt;0.9 or &lt;0.1). e.g. by class (maybe those predicted to be xx has better chance of being improved).  How much score must we change then, e.g. +0.05 +0.1 .... this is how you can set your weights. How to estimate the improvements?It is again the same old trick: either simulation with your train and validation samples or  \"probe\" the LB samples!\n\n  - (to be updated) ...\n\nof course what i described is\"manual tuning\". Based on the same idea,  you can design automatic algorithms to make it work better.\n\nlastly, there is  always a \"trend\", when you do things right. try to improve in  that direction. e.g. see https://www.kaggle.com/c/statoil-iceberg-classifier-challenge/discussion/47062\n\nHave fun!",
    "267031": "How much time does it take to train a single model? Do you also parallelize on multiple gpus?",
    "267104": "open to your own interpretation!",
    "267105": "&gt;&gt;How much time does it take to train a single model?\n2 to 4 hr.\n\n&gt;&gt;Do you also parallelize on multiple gpus?\nNo",
    "267147": "yet another observation",
    "267171": "I don't have time now but here is one paper that might be relevant when training with noisy labels\nhttps://arxiv.org/abs/1412.6596",
    "267328": "Dose pseudo-test24b-0.6 mean pseudo data take 60% of the new train data set(original train data+pseudo data)?  Additionally , how do you generate the pseudo data set,  the NN models which get 0.86 LB or some transitional machine learning classifier (Decision Tree, XGBoost)?",
    "267371": "I think that achieving scores higher than 0.9 is possible, but a different and fundamental aproach is needed. \n\nIt is clear that the test set is fundamentally different to the train set. It has different words and some of they are very tricky:\n\n* down vs town\n* left vs learn\n* off, on vs oh\n\nThis explains the great differences between cross-validation score and test score. The problem that we are facing on the test set is much harder than the problem on the train set. Instead of detecting big differences we are asked to detect subtle differences, and that is quite hard in a already noisy dataset.",
    "267424": "another observation",
    "267538": "you are correct. there are some solution:\n\n1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n   problem is that you need to align the wave.\n\n2) for my team, we try to sample from LB test set from pseudo-labeling",
    "267699": "http://benanne.github.io/2014/08/05/spotify-cnns.html",
    "267701": "https://arxiv.org/pdf/1709.01922.pdf",
    "267703": "http://ismir2015.uma.es/articles/264_Paper.pdf",
    "267743": "&gt; **Heng CherKeng wrote**\n&gt; \n&gt; &gt; you are correct. there are some solution:\n&gt; \n&gt; 1) cut and paste to create synthetic data : e.g. from down cut own. from tree cut t. then join to get town.\n&gt;    problem is that you need to align the wave.\n\nI like your idea. One option is to build a dictionary with different phonemes(having multiple audios for each phoneme) Then combine them randomly to create the desired words.  \nGathering the phonemes could be tedious but it could worth the effort.",
    "267773": "I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. \n\nAnyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).\n\nFor me, pseudo labelling the test set was the only thing that helped.",
    "267811": "&gt; **fergusoci wrote**\n&gt; \n&gt; &gt; I tried something similar to 1), but slightly cruder. I picked out different words in the training set from the same individuals and merged the first half of one word with the second half of the other (after adjusting for amplitude... because they were from the same individuals they meshed together quite well). I was aiming to create a set of 'unknown' words, most of these weren't real words, but some ended up sounding like real words: 'town', 'toe', 'low', etc. \n&gt; \n&gt; Anyway, it didn't seem to work! I only had the LB to check on as I still can't get a local cross validation scheme that reflects the test dataset (splitting by person not aggressive enough, leaving out some unknown words in splits too aggressive...).\n&gt; \n&gt; For me, pseudo labelling the test set was the only thing that helped.\n\nThanks for the info, how many new audios did you made? Maybe it's a problem of having a big number of them.",
    "267834": "Perhaps, yeah. I had generated 20k samples. I was subsampling unknowns during the model build (set to randomly pick 10% of the samples on each epoch), but maybe there was too many of the synthetic unknowns in comparison to the regular unknowns.",
    "267882": "sophisticated way is to GAN to create samples that resembles those in the LB set. But this is beyond the scope of the competition.\n\nE.g. \n\nGenerator (train_sample1,train_sample2) --&gt; fake_sample\n\nfake_sample /LB_sample --&gt; Discrminator --&gt; 0/1\n\nIf Generator(train_sample1,train_sample2) = encode() --&gt; decode() --&gt; fake_sample,\n\nthen we can deocde(LB_sample ) --&gt; latent components",
    "268092": "level-2 classifier. it achieve LB 0.88 without pusedo-labeling",
    "268150": "the most confusing samples\n\n ![enter image description here][1]\n\n\n  [1]: https://kaggle2.blob.core.windows.net/forum-message-attachments/268150/8255/confusion.png",
    "268171": "you can try this:\n\n    if there_is_no_noise and \\\n               amplitude_is_loud_enough and \\\n               predicted_label_is_one_of_knowns and \\\n               max_score_is_not_high_as_it_should_be :\n   \n               predicted_label = 'unknown'  #change\n\n\nit is because given all these good conditions, the prediction score should be high. If not, it is probably one of the confusing words. Now is how to set and detect the conditions?",
    "268179": "updated slides:"
  },
  "source": "meta"
}