{
  "id": 46659,
  "title": "How do you split the train data?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46659",
  "author_name": "",
  "post_date": "2017-12-31T20:20:16.352175700Z",
  "votes": 5,
  "comment_count": 22,
  "views": 0,
  "content": "<p>At first I tried randomly splitting with no regard for speakers, and then I was getting about -20% accuracy on validation in comparison to LB. Then I read the best way to do it is to put different speakers in train and val sets. So I did this, and gotten 10% improvement straight away (model stopped overfitting to known speakers).</p>\n\n<p>But how do you actually go over 80% accuracy on LB? My models - even tho they are simple - they achieve 95%+ validation accuracy, but my best attempt is only 75% on LB, why is that? Why is validation set (even the one provided in validation_files.txt) not representative of leaderboard.</p>\n\n<p>I'm training my nets in Keras, and to combat class imbalance problem I use class weights inversely proportional to the number of samples.</p>",
  "messages": [
    {
      "id": "263770",
      "postDate": "12/31/2017 20:20:16",
      "content": "<p>At first I tried randomly splitting with no regard for speakers, and then I was getting about -20% accuracy on validation in comparison to LB. Then I read the best way to do it is to put different speakers in train and val sets. So I did this, and gotten 10% improvement straight away (model stopped overfitting to known speakers).</p>\n\n<p>But how do you actually go over 80% accuracy on LB? My models - even tho they are simple - they achieve 95%+ validation accuracy, but my best attempt is only 75% on LB, why is that? Why is validation set (even the one provided in validation_files.txt) not representative of leaderboard.</p>\n\n<p>I'm training my nets in Keras, and to combat class imbalance problem I use class weights inversely proportional to the number of samples.</p>",
      "rawMarkdown": "At first I tried randomly splitting with no regard for speakers, and then I was getting about -20% accuracy on validation in comparison to LB. Then I read the best way to do it is to put different speakers in train and val sets. So I did this, and gotten 10% improvement straight away (model stopped overfitting to known speakers).\n\nBut how do you actually go over 80% accuracy on LB? My models - even tho they are simple - they achieve 95%+ validation accuracy, but my best attempt is only 75% on LB, why is that? Why is validation set (even the one provided in validation_files.txt) not representative of leaderboard.\n\nI'm training my nets in Keras, and to combat class imbalance problem I use class weights inversely proportional to the number of samples.",
      "votes": null
    },
    {
      "id": "263810",
      "postDate": "01/01/2018 04:44:14",
      "content": "<p>Wait a seconed, </p>\n\n<blockquote>\n  <p>the best way to do it is to put different speakers in train and val sets. So I did this,</p>\n</blockquote>\n\n<p>How did you do that , I mean how do you know which wav file belongs to the same man? </p>",
      "rawMarkdown": "Wait a seconed, \n\n&gt; the best way to do it is to put different speakers in train and val sets. So I did this,\n\nHow did you do that , I mean how do you know which wav file belongs to the same man?",
      "votes": null
    },
    {
      "id": "263812",
      "postDate": "01/01/2018 04:57:58",
      "content": "<p>The speakerID is part of the filename in the train set, as well as an \"utterance count\" -- the number of times the same speaker said the word.</p>\n\n<p>For example, in the train \"bed\" directory, for the file 00f0204f_nohash_0.wav the speaker ID is 00f0204f and this was utterance \"0\" (the first) of the word.  The file 00f0204f_nohash_1.wav is by the same speaker and a 2nd utterance (\"_1\") of the same word by the speaker.</p>",
      "rawMarkdown": "The speakerID is part of the filename in the train set, as well as an \"utterance count\" -- the number of times the same speaker said the word.\n\nFor example, in the train \"bed\" directory, for the file 00f0204f_nohash_0.wav the speaker ID is 00f0204f and this was utterance \"0\" (the first) of the word.  The file 00f0204f_nohash_1.wav is by the same speaker and a 2nd utterance (\"_1\") of the same word by the speaker.",
      "votes": null
    },
    {
      "id": "263831",
      "postDate": "01/01/2018 07:10:01",
      "content": "<p>thank you so much...I didn'T see the information</p>",
      "rawMarkdown": "thank you so much...I didn'T see the information",
      "votes": null
    },
    {
      "id": "263874",
      "postDate": "01/01/2018 12:48:41",
      "content": "<p>When the validation set gives a much better score than the test set, it means your validation set (and by extension the training set) is not representative of the test set. In other words, by using the provided training set you are not actually building a classifier that is suitable for the test set. </p>\n\n<p>If this was not a competition then you could (and should) make the training set match the test set better or the other way around. However, we're not in control of the test set here, and so getting a good score involves figuring out how the test set differs from the train/validation set, and augmenting the train/validation set to account for those differences.</p>\n\n<p>But you also don't want to go too far with that. It's very well possible that the people who are currently at the top of the public LB are seriously overfitting their models and that they will score much worse on the private LB.</p>",
      "rawMarkdown": "When the validation set gives a much better score than the test set, it means your validation set (and by extension the training set) is not representative of the test set. In other words, by using the provided training set you are not actually building a classifier that is suitable for the test set. \n\nIf this was not a competition then you could (and should) make the training set match the test set better or the other way around. However, we're not in control of the test set here, and so getting a good score involves figuring out how the test set differs from the train/validation set, and augmenting the train/validation set to account for those differences.\n\nBut you also don't want to go too far with that. It's very well possible that the people who are currently at the top of the public LB are seriously overfitting their models and that they will score much worse on the private LB.",
      "votes": null
    },
    {
      "id": "263881",
      "postDate": "01/01/2018 13:26:19",
      "content": "<p>So what kind of differences between provided set and LB set have been already found? Because I recall only one: LB set has apparently 2 times as much unknown classes.</p>",
      "rawMarkdown": "So what kind of differences between provided set and LB set have been already found? Because I recall only one: LB set has apparently 2 times as much unknown classes.",
      "votes": null
    },
    {
      "id": "263882",
      "postDate": "01/01/2018 13:28:53",
      "content": "<p>Anecdotally, I find that the test set consists of lower quality recordings. People who are harder to understand, more noise in the recordings, etc. (I could be wrong. :-) )</p>",
      "rawMarkdown": "Anecdotally, I find that the test set consists of lower quality recordings. People who are harder to understand, more noise in the recordings, etc. (I could be wrong. :-) )",
      "votes": null
    },
    {
      "id": "263908",
      "postDate": "01/01/2018 15:24:58",
      "content": "<p>part of the file name is an anonymised user id. I use for example the following code to split the data (X is a list of filenames and y the labels). I believe there is also an example in the training directory: </p>\n\n<pre><code>    import random\n\n   def split_set(X, y, valid=0.2):\n        Xt, yt, Xv, yv = [], [], [], []\n        for idx, filename in enumerate(X): \n            user_id = filename[-21:-13]\n            random.seed(user_id)\n            if random.random() &lt; valid:\n                Xv.append(X[idx])\n                yv.append(y[idx])\n            else:\n                Xt.append(X[idx])\n                yt.append(y[idx])\n\n        assert len(X) == len(Xt) + len(Xv)        \n        return Xt, yt, Xv, yv\n</code></pre>",
      "rawMarkdown": "part of the file name is an anonymised user id. I use for example the following code to split the data (X is a list of filenames and y the labels). I believe there is also an example in the training directory: \n\n        import random\n      \n       def split_set(X, y, valid=0.2):\n            Xt, yt, Xv, yv = [], [], [], []\n            for idx, filename in enumerate(X): \n                user_id = filename[-21:-13]\n                random.seed(user_id)\n                if random.random() &lt; valid:\n                    Xv.append(X[idx])\n                    yv.append(y[idx])\n                else:\n                    Xt.append(X[idx])\n                    yt.append(y[idx])\n                    \n            assert len(X) == len(Xt) + len(Xv)        \n            return Xt, yt, Xv, yv",
      "votes": null
    },
    {
      "id": "263909",
      "postDate": "01/01/2018 15:28:18",
      "content": "<p>You mind telling how does your validation accuracy compare to your LB?</p>",
      "rawMarkdown": "You mind telling how does your validation accuracy compare to your LB?",
      "votes": null
    },
    {
      "id": "264007",
      "postDate": "01/02/2018 01:24:10",
      "content": "<p>@ Peter Dekkers  Thank you! Still a question, when you say </p>\n\n<blockquote>\n  <p>model stopped overfitting to known speakers</p>\n</blockquote>\n\n<p>do you mean you find a point that model performs better on the val set generated by \"split_set\" function, or what exactly did you do to prevent overfit?</p>",
      "rawMarkdown": "Peter Dekkers  Thank you! Still a question, when you say \n\n&gt; model stopped overfitting to known speakers\n\n \n\ndo you mean you find a point that model performs better on the val set generated by \"split_set\" function, or what exactly did you do to prevent overfit?",
      "votes": null
    },
    {
      "id": "264433",
      "postDate": "01/03/2018 03:04:32",
      "content": "<p>This was a really nice question. I was having the same question myself. Even using the provided validation_list.txt and testing_list.txt I could still see that there was a big gap between local validation score and LB score. I didn't had the chance to test yet what @Peter Dekkers suggested. That is in the todo list.</p>",
      "rawMarkdown": "This was a really nice question. I was having the same question myself. Even using the provided validation_list.txt and testing_list.txt I could still see that there was a big gap between local validation score and LB score. I didn't had the chance to test yet what @Peter Dekkers suggested. That is in the todo list.",
      "votes": null
    },
    {
      "id": "264496",
      "postDate": "01/03/2018 07:52:16",
      "content": "<p>Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. </p>\n\n<p>I guess one of the reasons is the overfitting the \"unknown\" words. In training it just learns to recognise the different unknown words in the dataset. But in LB there are new unknown words and these could be labelled unknown as a well as a known word.</p>",
      "rawMarkdown": "Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. \n\nI guess one of the reasons is the overfitting the \"unknown\" words. In training it just learns to recognise the different unknown words in the dataset. But in LB there are new unknown words and these could be labelled unknown as a well as a known word.",
      "votes": null
    },
    {
      "id": "264531",
      "postDate": "01/03/2018 10:26:19",
      "content": "<p>Sorry for off-topic, but wow, it's a pretty neat trick with random.seed here :). Although I think it changes the meaning of the \"valid\" parameter from \"percentage of train samples to be moved to validation set\" to \"probability of moving all of given user's samples to validation set\". This may result in different split than 80%-20% (if valid == 0.2) if some users are \"overrepresented\" in the set, but I guess that effect is neglible in large data sets.</p>",
      "rawMarkdown": "Sorry for off-topic, but wow, it's a pretty neat trick with random.seed here :). Although I think it changes the meaning of the \"valid\" parameter from \"percentage of train samples to be moved to validation set\" to \"probability of moving all of given user's samples to validation set\". This may result in different split than 80%-20% (if valid == 0.2) if some users are \"overrepresented\" in the set, but I guess that effect is neglible in large data sets.",
      "votes": null
    },
    {
      "id": "264537",
      "postDate": "01/03/2018 10:39:08",
      "content": "<blockquote>\n  <p><strong>Peter Dekkers wrote</strong></p>\n  \n  <blockquote>\n    <p>Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. </p>\n  </blockquote>\n</blockquote>\n\n<p>You mind sharing your architecture and data preparation?\nI constantly hit 95%+ on validation set, but I can't seem to be able to go over 76% on LB.</p>",
      "rawMarkdown": "&gt; **Peter Dekkers wrote**\n&gt; \n&gt; &gt; Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. \n&gt; \n\nYou mind sharing your architecture and data preparation?\nI constantly hit 95%+ on validation set, but I can't seem to be able to go over 76% on LB.",
      "votes": null
    },
    {
      "id": "264542",
      "postDate": "01/03/2018 10:53:33",
      "content": "<p>Few things I’m doing:</p>\n\n<p>1) I split data in 90% training, 10% validation. Not ideal, but since there is not much data to start with this works best for me.</p>\n\n<p>2) For training I use random sampling to get an even distribution among the 12 labels. This made a good improvement for me, since it avoids that the network has a preference for unknown (I also tried weights in the loss function, but that didn’t work as well)</p>\n\n<p>3) For validation i use the natural distribution (so I just use all the samples in the validation set once).</p>\n\n<p>Besides that I also do some extra feature engineering, like mixing noise and time shift a bit to create more samples and fight overfitting. The network is a straight forward conv net with use of dropout to fight overfitting and batch normalisation to control the gradients. I use PyTorch:</p>\n\n<pre><code>class ConvNet2(nn.Module):\n'''The actual network consisting of few conv layers followed by fully connected layers'''\n\ndef __init__(self):\n    super(ConvNet2, self).__init__()\n    self.features = nn.Sequential(\n\n        nn.BatchNorm2d(1, momentum=0.99), # too lazy to normalise input myself ;)\n\n        nn.Conv2d(1, 64, kernel_size=3),\n        nn.MaxPool2d(2),\n        nn.ReLU(True),\n\n        nn.Conv2d( 64, 128, kernel_size=3),\n        nn.MaxPool2d(2),\n        nn.ReLU(True),\n\n        nn.Conv2d(128, 256, kernel_size=3),\n        nn.MaxPool2d(4),\n        nn.BatchNorm2d(256, momentum=0.99),\n        nn.ReLU(True),\n        nn.Dropout(0.5),\n    ) \n\n    self.linear = nn.Sequential(\n        nn.Linear(7680, 256),\n        nn.BatchNorm1d(256, momentum=0.99),\n        nn.ReLU(True),\n        nn.Linear(256, len(WORDS)),\n    )\n\ndef forward(self, input):\n    x = self.features(input)\n    x = x.view(x.size(0), -1) # flatten\n    x = self.linear(x)\n    return F.log_softmax(x, dim=1)\n</code></pre>\n\n<p>Hopes this provides some ideas.</p>",
      "rawMarkdown": "Few things I’m doing:\n\n1) I split data in 90% training, 10% validation. Not ideal, but since there is not much data to start with this works best for me.\n\n2) For training I use random sampling to get an even distribution among the 12 labels. This made a good improvement for me, since it avoids that the network has a preference for unknown (I also tried weights in the loss function, but that didn’t work as well)\n\n3) For validation i use the natural distribution (so I just use all the samples in the validation set once).\n\nBesides that I also do some extra feature engineering, like mixing noise and time shift a bit to create more samples and fight overfitting. The network is a straight forward conv net with use of dropout to fight overfitting and batch normalisation to control the gradients. I use PyTorch:\n\n    class ConvNet2(nn.Module):\n    '''The actual network consisting of few conv layers followed by fully connected layers'''\n   \n    def __init__(self):\n        super(ConvNet2, self).__init__()\n        self.features = nn.Sequential(\n            \n            nn.BatchNorm2d(1, momentum=0.99), # too lazy to normalise input myself ;)\n            \n            nn.Conv2d(1, 64, kernel_size=3),\n            nn.MaxPool2d(2),\n            nn.ReLU(True),\n            \n            nn.Conv2d( 64, 128, kernel_size=3),\n            nn.MaxPool2d(2),\n            nn.ReLU(True),\n\n            nn.Conv2d(128, 256, kernel_size=3),\n            nn.MaxPool2d(4),\n            nn.BatchNorm2d(256, momentum=0.99),\n            nn.ReLU(True),\n            nn.Dropout(0.5),\n        ) \n        \n        self.linear = nn.Sequential(\n            nn.Linear(7680, 256),\n            nn.BatchNorm1d(256, momentum=0.99),\n            nn.ReLU(True),\n            nn.Linear(256, len(WORDS)),\n        )\n        \n    def forward(self, input):\n        x = self.features(input)\n        x = x.view(x.size(0), -1) # flatten\n        x = self.linear(x)\n        return F.log_softmax(x, dim=1)\n\n\n\nHopes this provides some ideas.",
      "votes": null
    },
    {
      "id": "264652",
      "postDate": "01/03/2018 16:22:07",
      "content": "<p>What are you feeding as the input?</p>",
      "rawMarkdown": "What are you feeding as the input?",
      "votes": null
    },
    {
      "id": "264685",
      "postDate": "01/03/2018 17:16:57",
      "content": "<p>Tried both log spectrogram and mel spectrogram, similar results (0.85 best LB).  I used 20ms windows with 50% overlap.</p>\n\n<p>Also one time I tried feeding directly a plain wav file (16000 samples) into a 1D convolutional network , but that wasn't a big success (although others showed that that can also work). However still like the simplicity and elegance of this.</p>",
      "rawMarkdown": "Tried both log spectrogram and mel spectrogram, similar results (0.85 best LB).  I used 20ms windows with 50% overlap.\n\nAlso one time I tried feeding directly a plain wav file (16000 samples) into a 1D convolutional network , but that wasn't a big success (although others showed that that can also work). However still like the simplicity and elegance of this.",
      "votes": null
    },
    {
      "id": "264743",
      "postDate": "01/03/2018 19:43:52",
      "content": "<p>@Lukasz: Indeed the approach is not very suitable if you require the split to be exactly 20% or close to it, but like you stated if the number are large enough, you won't be too far off. The plus side is that the algoritme is pretty reusable for other use cases as long as you can provide an identifier that is hash-able. </p>\n\n<p>P.S In fact there is always the possibility that it is mathematiccly impossible to achieve exactly 20% depending on the distribution of samples. So there is no algorithm that can 100% ensure this.</p>",
      "rawMarkdown": "Lukasz: Indeed the approach is not very suitable if you require the split to be exactly 20% or close to it, but like you stated if the number are large enough, you won't be too far off. The plus side is that the algoritme is pretty reusable for other use cases as long as you can provide an identifier that is hash-able. \n\nP.S In fact there is always the possibility that it is mathematiccly impossible to achieve exactly 20% depending on the distribution of samples. So there is no algorithm that can 100% ensure this.",
      "votes": null
    },
    {
      "id": "265093",
      "postDate": "01/04/2018 15:56:11",
      "content": "<p>Hi Lugi, I'm having similar experiences with keras. What preprocessing are you using?. I tried to make my preprocessing as close to tensorflow's official example but recently discovered that my preprocessing was different to tensorflow's. So I now switched back to official tensorflow code and making changes to it. So double check preprocessing stage.</p>\n\n<p>Also according to this thread <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44283#254182\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44283#254182</a>, it seems that adding augmentation to validation set improves a bit.</p>\n\n<p>One other thing to note is that if you are using spectrograms/MFCC, you will end up with matrices that are not square. In the official examples they used rectangular kernel shapes. So be wary of using square kernels, atleast in earlier layers</p>\n\n<p>A good place to start would be <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/45037\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/45037</a></p>",
      "rawMarkdown": "Hi Lugi, I'm having similar experiences with keras. What preprocessing are you using?. I tried to make my preprocessing as close to tensorflow's official example but recently discovered that my preprocessing was different to tensorflow's. So I now switched back to official tensorflow code and making changes to it. So double check preprocessing stage.\n\nAlso according to this thread https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44283#254182, it seems that adding augmentation to validation set improves a bit.\n\nOne other thing to note is that if you are using spectrograms/MFCC, you will end up with matrices that are not square. In the official examples they used rectangular kernel shapes. So be wary of using square kernels, atleast in earlier layers\n\nA good place to start would be https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/45037",
      "votes": null
    },
    {
      "id": "265241",
      "postDate": "01/05/2018 00:17:28",
      "content": "<p>Im feeding the waveform directly to the net without any augmentation, so this is probably the place where im losing most lb points.</p>",
      "rawMarkdown": "Im feeding the waveform directly to the net without any augmentation, so this is probably the place where im losing most lb points.",
      "votes": null
    },
    {
      "id": "265252",
      "postDate": "01/05/2018 00:51:28",
      "content": "<p>Hi Peter, just a quick question about the code snippet you posted. If we follow that logic then dim(Xt) = 12 and dim(Xv) = 11. I noticed that label <code>_background_noise_</code> is present only in yt. Did you notice that?</p>\n\n<pre><code> np.unique(np.array(ytr))\n array(['_background_noise_', 'bed', 'bird', 'cat', 'dog', 'down', 'eight',\n   'five', 'four', 'go', 'happy', 'house', 'left', 'marvin', 'nine',\n   'no', 'off', 'on', 'one', 'right', 'seven', 'sheila', 'six', 'stop',\n   'three', 'tree', 'two', 'up', 'wow', 'yes', 'zero'],\n  dtype='&lt;U18')\n\nnp.unique(yval)\narray(['bed', 'bird', 'cat', 'dog', 'down', 'eight', 'five', 'four', 'go',\n   'happy', 'house', 'left', 'marvin', 'nine', 'no', 'off', 'on',\n   'one', 'right', 'seven', 'sheila', 'six', 'stop', 'three', 'tree',\n   'two', 'up', 'wow', 'yes', 'zero'],\n  dtype='&lt;U6')\n</code></pre>",
      "rawMarkdown": "Hi Peter, just a quick question about the code snippet you posted. If we follow that logic then dim(Xt) = 12 and dim(Xv) = 11. I noticed that label `_background_noise_` is present only in yt. Did you notice that?\n\n     np.unique(np.array(ytr))\n     array(['_background_noise_', 'bed', 'bird', 'cat', 'dog', 'down', 'eight',\n       'five', 'four', 'go', 'happy', 'house', 'left', 'marvin', 'nine',\n       'no', 'off', 'on', 'one', 'right', 'seven', 'sheila', 'six', 'stop',\n       'three', 'tree', 'two', 'up', 'wow', 'yes', 'zero'],\n      dtype='",
      "votes": null
    },
    {
      "id": "265400",
      "postDate": "01/05/2018 13:31:54",
      "content": "<p>hi there!</p>\n\n<p>divide your dataset into 3 parts</p>\n\n<p>training set(60%)</p>\n\n<p>cross-validation set(20%)</p>\n\n<p>test set(20%)</p>\n\n<p>or you may change the ratio according to your model, but just make sure to keep the training set large enough and then use the test set to see if the model is performing well.</p>",
      "rawMarkdown": "hi there!\n\ndivide your dataset into 3 parts\n\n\ntraining set(60%)\n\ncross-validation set(20%)\n\ntest set(20%)\n\nor you may change the ratio according to your model, but just make sure to keep the training set large enough and then use the test set to see if the model is performing well.",
      "votes": null
    },
    {
      "id": "265408",
      "postDate": "01/05/2018 13:59:09",
      "content": "<p>This partitioning won't work in this competition. You will still get good score in partitioned test set, but LB will be bad. </p>\n\n<p>I guess one of the main reasons why there is a big gap between CV and LB is unknown unknowns (it's audio labels that are present in test and not present in train). So the CNN while trained hasn't seen these labels and that's why can't predict them correctly.</p>",
      "rawMarkdown": "This partitioning won't work in this competition. You will still get good score in partitioned test set, but LB will be bad. \n\nI guess one of the main reasons why there is a big gap between CV and LB is unknown unknowns (it's audio labels that are present in test and not present in train). So the CNN while trained hasn't seen these labels and that's why can't predict them correctly.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 263810,
      "author_name": "icybee",
      "author_url": "",
      "post_date": "01/01/2018 04:44:14",
      "content": "<p>Wait a seconed, </p>\n\n<blockquote>\n  <p>the best way to do it is to put different speakers in train and val sets. So I did this,</p>\n</blockquote>\n\n<p>How did you do that , I mean how do you know which wav file belongs to the same man? </p>",
      "votes": null,
      "replies": [
        {
          "id": 263908,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/01/2018 15:24:58",
          "content": "<p>part of the file name is an anonymised user id. I use for example the following code to split the data (X is a list of filenames and y the labels). I believe there is also an example in the training directory: </p>\n\n<pre><code>    import random\n\n   def split_set(X, y, valid=0.2):\n        Xt, yt, Xv, yv = [], [], [], []\n        for idx, filename in enumerate(X): \n            user_id = filename[-21:-13]\n            random.seed(user_id)\n            if random.random() &lt; valid:\n                Xv.append(X[idx])\n                yv.append(y[idx])\n            else:\n                Xt.append(X[idx])\n                yt.append(y[idx])\n\n        assert len(X) == len(Xt) + len(Xv)        \n        return Xt, yt, Xv, yv\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263909,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "01/01/2018 15:28:18",
          "content": "<p>You mind telling how does your validation accuracy compare to your LB?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264007,
          "author_name": "icybee",
          "author_url": "",
          "post_date": "01/02/2018 01:24:10",
          "content": "<p>@ Peter Dekkers  Thank you! Still a question, when you say </p>\n\n<blockquote>\n  <p>model stopped overfitting to known speakers</p>\n</blockquote>\n\n<p>do you mean you find a point that model performs better on the val set generated by \"split_set\" function, or what exactly did you do to prevent overfit?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264496,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/03/2018 07:52:16",
          "content": "<p>Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. </p>\n\n<p>I guess one of the reasons is the overfitting the \"unknown\" words. In training it just learns to recognise the different unknown words in the dataset. But in LB there are new unknown words and these could be labelled unknown as a well as a known word.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264531,
          "author_name": "laszel",
          "author_url": "",
          "post_date": "01/03/2018 10:26:19",
          "content": "<p>Sorry for off-topic, but wow, it's a pretty neat trick with random.seed here :). Although I think it changes the meaning of the \"valid\" parameter from \"percentage of train samples to be moved to validation set\" to \"probability of moving all of given user's samples to validation set\". This may result in different split than 80%-20% (if valid == 0.2) if some users are \"overrepresented\" in the set, but I guess that effect is neglible in large data sets.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264537,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "01/03/2018 10:39:08",
          "content": "<blockquote>\n  <p><strong>Peter Dekkers wrote</strong></p>\n  \n  <blockquote>\n    <p>Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. </p>\n  </blockquote>\n</blockquote>\n\n<p>You mind sharing your architecture and data preparation?\nI constantly hit 95%+ on validation set, but I can't seem to be able to go over 76% on LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264542,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/03/2018 10:53:33",
          "content": "<p>Few things I’m doing:</p>\n\n<p>1) I split data in 90% training, 10% validation. Not ideal, but since there is not much data to start with this works best for me.</p>\n\n<p>2) For training I use random sampling to get an even distribution among the 12 labels. This made a good improvement for me, since it avoids that the network has a preference for unknown (I also tried weights in the loss function, but that didn’t work as well)</p>\n\n<p>3) For validation i use the natural distribution (so I just use all the samples in the validation set once).</p>\n\n<p>Besides that I also do some extra feature engineering, like mixing noise and time shift a bit to create more samples and fight overfitting. The network is a straight forward conv net with use of dropout to fight overfitting and batch normalisation to control the gradients. I use PyTorch:</p>\n\n<pre><code>class ConvNet2(nn.Module):\n'''The actual network consisting of few conv layers followed by fully connected layers'''\n\ndef __init__(self):\n    super(ConvNet2, self).__init__()\n    self.features = nn.Sequential(\n\n        nn.BatchNorm2d(1, momentum=0.99), # too lazy to normalise input myself ;)\n\n        nn.Conv2d(1, 64, kernel_size=3),\n        nn.MaxPool2d(2),\n        nn.ReLU(True),\n\n        nn.Conv2d( 64, 128, kernel_size=3),\n        nn.MaxPool2d(2),\n        nn.ReLU(True),\n\n        nn.Conv2d(128, 256, kernel_size=3),\n        nn.MaxPool2d(4),\n        nn.BatchNorm2d(256, momentum=0.99),\n        nn.ReLU(True),\n        nn.Dropout(0.5),\n    ) \n\n    self.linear = nn.Sequential(\n        nn.Linear(7680, 256),\n        nn.BatchNorm1d(256, momentum=0.99),\n        nn.ReLU(True),\n        nn.Linear(256, len(WORDS)),\n    )\n\ndef forward(self, input):\n    x = self.features(input)\n    x = x.view(x.size(0), -1) # flatten\n    x = self.linear(x)\n    return F.log_softmax(x, dim=1)\n</code></pre>\n\n<p>Hopes this provides some ideas.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264652,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "01/03/2018 16:22:07",
          "content": "<p>What are you feeding as the input?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264685,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/03/2018 17:16:57",
          "content": "<p>Tried both log spectrogram and mel spectrogram, similar results (0.85 best LB).  I used 20ms windows with 50% overlap.</p>\n\n<p>Also one time I tried feeding directly a plain wav file (16000 samples) into a 1D convolutional network , but that wasn't a big success (although others showed that that can also work). However still like the simplicity and elegance of this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 264743,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/03/2018 19:43:52",
          "content": "<p>@Lukasz: Indeed the approach is not very suitable if you require the split to be exactly 20% or close to it, but like you stated if the number are large enough, you won't be too far off. The plus side is that the algoritme is pretty reusable for other use cases as long as you can provide an identifier that is hash-able. </p>\n\n<p>P.S In fact there is always the possibility that it is mathematiccly impossible to achieve exactly 20% depending on the distribution of samples. So there is no algorithm that can 100% ensure this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265252,
          "author_name": "kirk86",
          "author_url": "",
          "post_date": "01/05/2018 00:51:28",
          "content": "<p>Hi Peter, just a quick question about the code snippet you posted. If we follow that logic then dim(Xt) = 12 and dim(Xv) = 11. I noticed that label <code>_background_noise_</code> is present only in yt. Did you notice that?</p>\n\n<pre><code> np.unique(np.array(ytr))\n array(['_background_noise_', 'bed', 'bird', 'cat', 'dog', 'down', 'eight',\n   'five', 'four', 'go', 'happy', 'house', 'left', 'marvin', 'nine',\n   'no', 'off', 'on', 'one', 'right', 'seven', 'sheila', 'six', 'stop',\n   'three', 'tree', 'two', 'up', 'wow', 'yes', 'zero'],\n  dtype='&lt;U18')\n\nnp.unique(yval)\narray(['bed', 'bird', 'cat', 'dog', 'down', 'eight', 'five', 'four', 'go',\n   'happy', 'house', 'left', 'marvin', 'nine', 'no', 'off', 'on',\n   'one', 'right', 'seven', 'sheila', 'six', 'stop', 'three', 'tree',\n   'two', 'up', 'wow', 'yes', 'zero'],\n  dtype='&lt;U6')\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 263812,
      "author_name": "efglynn",
      "author_url": "",
      "post_date": "01/01/2018 04:57:58",
      "content": "<p>The speakerID is part of the filename in the train set, as well as an \"utterance count\" -- the number of times the same speaker said the word.</p>\n\n<p>For example, in the train \"bed\" directory, for the file 00f0204f_nohash_0.wav the speaker ID is 00f0204f and this was utterance \"0\" (the first) of the word.  The file 00f0204f_nohash_1.wav is by the same speaker and a 2nd utterance (\"_1\") of the same word by the speaker.</p>",
      "votes": null,
      "replies": [
        {
          "id": 263831,
          "author_name": "icybee",
          "author_url": "",
          "post_date": "01/01/2018 07:10:01",
          "content": "<p>thank you so much...I didn'T see the information</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 263874,
      "author_name": "humananalog",
      "author_url": "",
      "post_date": "01/01/2018 12:48:41",
      "content": "<p>When the validation set gives a much better score than the test set, it means your validation set (and by extension the training set) is not representative of the test set. In other words, by using the provided training set you are not actually building a classifier that is suitable for the test set. </p>\n\n<p>If this was not a competition then you could (and should) make the training set match the test set better or the other way around. However, we're not in control of the test set here, and so getting a good score involves figuring out how the test set differs from the train/validation set, and augmenting the train/validation set to account for those differences.</p>\n\n<p>But you also don't want to go too far with that. It's very well possible that the people who are currently at the top of the public LB are seriously overfitting their models and that they will score much worse on the private LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 263881,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "01/01/2018 13:26:19",
          "content": "<p>So what kind of differences between provided set and LB set have been already found? Because I recall only one: LB set has apparently 2 times as much unknown classes.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 263882,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "01/01/2018 13:28:53",
          "content": "<p>Anecdotally, I find that the test set consists of lower quality recordings. People who are harder to understand, more noise in the recordings, etc. (I could be wrong. :-) )</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 264433,
      "author_name": "kirk86",
      "author_url": "",
      "post_date": "01/03/2018 03:04:32",
      "content": "<p>This was a really nice question. I was having the same question myself. Even using the provided validation_list.txt and testing_list.txt I could still see that there was a big gap between local validation score and LB score. I didn't had the chance to test yet what @Peter Dekkers suggested. That is in the todo list.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 265093,
      "author_name": "rteja1113",
      "author_url": "",
      "post_date": "01/04/2018 15:56:11",
      "content": "<p>Hi Lugi, I'm having similar experiences with keras. What preprocessing are you using?. I tried to make my preprocessing as close to tensorflow's official example but recently discovered that my preprocessing was different to tensorflow's. So I now switched back to official tensorflow code and making changes to it. So double check preprocessing stage.</p>\n\n<p>Also according to this thread <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44283#254182\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44283#254182</a>, it seems that adding augmentation to validation set improves a bit.</p>\n\n<p>One other thing to note is that if you are using spectrograms/MFCC, you will end up with matrices that are not square. In the official examples they used rectangular kernel shapes. So be wary of using square kernels, atleast in earlier layers</p>\n\n<p>A good place to start would be <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/45037\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/45037</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 265241,
          "author_name": "lugi777",
          "author_url": "",
          "post_date": "01/05/2018 00:17:28",
          "content": "<p>Im feeding the waveform directly to the net without any augmentation, so this is probably the place where im losing most lb points.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265400,
      "author_name": "vishalyo990",
      "author_url": "",
      "post_date": "01/05/2018 13:31:54",
      "content": "<p>hi there!</p>\n\n<p>divide your dataset into 3 parts</p>\n\n<p>training set(60%)</p>\n\n<p>cross-validation set(20%)</p>\n\n<p>test set(20%)</p>\n\n<p>or you may change the ratio according to your model, but just make sure to keep the training set large enough and then use the test set to see if the model is performing well.</p>",
      "votes": null,
      "replies": [
        {
          "id": 265408,
          "author_name": "heyt0ny",
          "author_url": "",
          "post_date": "01/05/2018 13:59:09",
          "content": "<p>This partitioning won't work in this competition. You will still get good score in partitioned test set, but LB will be bad. </p>\n\n<p>I guess one of the main reasons why there is a big gap between CV and LB is unknown unknowns (it's audio labels that are present in test and not present in train). So the CNN while trained hasn't seen these labels and that's why can't predict them correctly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "263770": "At first I tried randomly splitting with no regard for speakers, and then I was getting about -20% accuracy on validation in comparison to LB. Then I read the best way to do it is to put different speakers in train and val sets. So I did this, and gotten 10% improvement straight away (model stopped overfitting to known speakers).\n\nBut how do you actually go over 80% accuracy on LB? My models - even tho they are simple - they achieve 95%+ validation accuracy, but my best attempt is only 75% on LB, why is that? Why is validation set (even the one provided in validation_files.txt) not representative of leaderboard.\n\nI'm training my nets in Keras, and to combat class imbalance problem I use class weights inversely proportional to the number of samples.",
    "263810": "Wait a seconed, \n\n&gt; the best way to do it is to put different speakers in train and val sets. So I did this,\n\nHow did you do that , I mean how do you know which wav file belongs to the same man?",
    "263812": "The speakerID is part of the filename in the train set, as well as an \"utterance count\" -- the number of times the same speaker said the word.\n\nFor example, in the train \"bed\" directory, for the file 00f0204f_nohash_0.wav the speaker ID is 00f0204f and this was utterance \"0\" (the first) of the word.  The file 00f0204f_nohash_1.wav is by the same speaker and a 2nd utterance (\"_1\") of the same word by the speaker.",
    "263831": "thank you so much...I didn'T see the information",
    "263874": "When the validation set gives a much better score than the test set, it means your validation set (and by extension the training set) is not representative of the test set. In other words, by using the provided training set you are not actually building a classifier that is suitable for the test set. \n\nIf this was not a competition then you could (and should) make the training set match the test set better or the other way around. However, we're not in control of the test set here, and so getting a good score involves figuring out how the test set differs from the train/validation set, and augmenting the train/validation set to account for those differences.\n\nBut you also don't want to go too far with that. It's very well possible that the people who are currently at the top of the public LB are seriously overfitting their models and that they will score much worse on the private LB.",
    "263881": "So what kind of differences between provided set and LB set have been already found? Because I recall only one: LB set has apparently 2 times as much unknown classes.",
    "263882": "Anecdotally, I find that the test set consists of lower quality recordings. People who are harder to understand, more noise in the recordings, etc. (I could be wrong. :-) )",
    "263908": "part of the file name is an anonymised user id. I use for example the following code to split the data (X is a list of filenames and y the labels). I believe there is also an example in the training directory: \n\n        import random\n      \n       def split_set(X, y, valid=0.2):\n            Xt, yt, Xv, yv = [], [], [], []\n            for idx, filename in enumerate(X): \n                user_id = filename[-21:-13]\n                random.seed(user_id)\n                if random.random() &lt; valid:\n                    Xv.append(X[idx])\n                    yv.append(y[idx])\n                else:\n                    Xt.append(X[idx])\n                    yt.append(y[idx])\n                    \n            assert len(X) == len(Xt) + len(Xv)        \n            return Xt, yt, Xv, yv",
    "263909": "You mind telling how does your validation accuracy compare to your LB?",
    "264007": "Peter Dekkers  Thank you! Still a question, when you say \n\n&gt; model stopped overfitting to known speakers\n\n \n\ndo you mean you find a point that model performs better on the val set generated by \"split_set\" function, or what exactly did you do to prevent overfit?",
    "264433": "This was a really nice question. I was having the same question myself. Even using the provided validation_list.txt and testing_list.txt I could still see that there was a big gap between local validation score and LB score. I didn't had the chance to test yet what @Peter Dekkers suggested. That is in the todo list.",
    "264496": "Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. \n\nI guess one of the reasons is the overfitting the \"unknown\" words. In training it just learns to recognise the different unknown words in the dataset. But in LB there are new unknown words and these could be labelled unknown as a well as a known word.",
    "264531": "Sorry for off-topic, but wow, it's a pretty neat trick with random.seed here :). Although I think it changes the meaning of the \"valid\" parameter from \"percentage of train samples to be moved to validation set\" to \"probability of moving all of given user's samples to validation set\". This may result in different split than 80%-20% (if valid == 0.2) if some users are \"overrepresented\" in the set, but I guess that effect is neglible in large data sets.",
    "264537": "&gt; **Peter Dekkers wrote**\n&gt; \n&gt; &gt; Unfortunately my model also still overfits, my validation accuracy is 93%-94% while LB is about 10% lower. \n&gt; \n\nYou mind sharing your architecture and data preparation?\nI constantly hit 95%+ on validation set, but I can't seem to be able to go over 76% on LB.",
    "264542": "Few things I’m doing:\n\n1) I split data in 90% training, 10% validation. Not ideal, but since there is not much data to start with this works best for me.\n\n2) For training I use random sampling to get an even distribution among the 12 labels. This made a good improvement for me, since it avoids that the network has a preference for unknown (I also tried weights in the loss function, but that didn’t work as well)\n\n3) For validation i use the natural distribution (so I just use all the samples in the validation set once).\n\nBesides that I also do some extra feature engineering, like mixing noise and time shift a bit to create more samples and fight overfitting. The network is a straight forward conv net with use of dropout to fight overfitting and batch normalisation to control the gradients. I use PyTorch:\n\n    class ConvNet2(nn.Module):\n    '''The actual network consisting of few conv layers followed by fully connected layers'''\n   \n    def __init__(self):\n        super(ConvNet2, self).__init__()\n        self.features = nn.Sequential(\n            \n            nn.BatchNorm2d(1, momentum=0.99), # too lazy to normalise input myself ;)\n            \n            nn.Conv2d(1, 64, kernel_size=3),\n            nn.MaxPool2d(2),\n            nn.ReLU(True),\n            \n            nn.Conv2d( 64, 128, kernel_size=3),\n            nn.MaxPool2d(2),\n            nn.ReLU(True),\n\n            nn.Conv2d(128, 256, kernel_size=3),\n            nn.MaxPool2d(4),\n            nn.BatchNorm2d(256, momentum=0.99),\n            nn.ReLU(True),\n            nn.Dropout(0.5),\n        ) \n        \n        self.linear = nn.Sequential(\n            nn.Linear(7680, 256),\n            nn.BatchNorm1d(256, momentum=0.99),\n            nn.ReLU(True),\n            nn.Linear(256, len(WORDS)),\n        )\n        \n    def forward(self, input):\n        x = self.features(input)\n        x = x.view(x.size(0), -1) # flatten\n        x = self.linear(x)\n        return F.log_softmax(x, dim=1)\n\n\n\nHopes this provides some ideas.",
    "264652": "What are you feeding as the input?",
    "264685": "Tried both log spectrogram and mel spectrogram, similar results (0.85 best LB).  I used 20ms windows with 50% overlap.\n\nAlso one time I tried feeding directly a plain wav file (16000 samples) into a 1D convolutional network , but that wasn't a big success (although others showed that that can also work). However still like the simplicity and elegance of this.",
    "264743": "Lukasz: Indeed the approach is not very suitable if you require the split to be exactly 20% or close to it, but like you stated if the number are large enough, you won't be too far off. The plus side is that the algoritme is pretty reusable for other use cases as long as you can provide an identifier that is hash-able. \n\nP.S In fact there is always the possibility that it is mathematiccly impossible to achieve exactly 20% depending on the distribution of samples. So there is no algorithm that can 100% ensure this.",
    "265093": "Hi Lugi, I'm having similar experiences with keras. What preprocessing are you using?. I tried to make my preprocessing as close to tensorflow's official example but recently discovered that my preprocessing was different to tensorflow's. So I now switched back to official tensorflow code and making changes to it. So double check preprocessing stage.\n\nAlso according to this thread https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/44283#254182, it seems that adding augmentation to validation set improves a bit.\n\nOne other thing to note is that if you are using spectrograms/MFCC, you will end up with matrices that are not square. In the official examples they used rectangular kernel shapes. So be wary of using square kernels, atleast in earlier layers\n\nA good place to start would be https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/45037",
    "265241": "Im feeding the waveform directly to the net without any augmentation, so this is probably the place where im losing most lb points.",
    "265252": "Hi Peter, just a quick question about the code snippet you posted. If we follow that logic then dim(Xt) = 12 and dim(Xv) = 11. I noticed that label `_background_noise_` is present only in yt. Did you notice that?\n\n     np.unique(np.array(ytr))\n     array(['_background_noise_', 'bed', 'bird', 'cat', 'dog', 'down', 'eight',\n       'five', 'four', 'go', 'happy', 'house', 'left', 'marvin', 'nine',\n       'no', 'off', 'on', 'one', 'right', 'seven', 'sheila', 'six', 'stop',\n       'three', 'tree', 'two', 'up', 'wow', 'yes', 'zero'],\n      dtype='",
    "265400": "hi there!\n\ndivide your dataset into 3 parts\n\n\ntraining set(60%)\n\ncross-validation set(20%)\n\ntest set(20%)\n\nor you may change the ratio according to your model, but just make sure to keep the training set large enough and then use the test set to see if the model is performing well.",
    "265408": "This partitioning won't work in this competition. You will still get good score in partitioned test set, but LB will be bad. \n\nI guess one of the main reasons why there is a big gap between CV and LB is unknown unknowns (it's audio labels that are present in test and not present in train). So the CNN while trained hasn't seen these labels and that's why can't predict them correctly."
  },
  "source": "meta"
}