{
  "id": 44250,
  "title": "Mismatch between validation and test accuracy",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44250",
  "author_name": "Shujian Liu",
  "post_date": "2017-11-26T01:33:38.957000",
  "votes": 13,
  "comment_count": 34,
  "views": 0,
  "content": "<p>I just started this competition yesterday. I followed <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">Google's TF tutorials</a> and got a validation accuracy of between 85% and 90% as they stated on the web page. However, LB is only 0.77. Could anyone provide some insights into the mismatch? Thanks.</p>\n\n<p>Another observation is that unknown should only take 9% but the output has ~ 30% unknown.</p>",
  "messages": [
    {
      "id": 248416,
      "postDate": "2017-11-26T01:33:38.957Z",
      "content": "<p>I just started this competition yesterday. I followed <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">Google's TF tutorials</a> and got a validation accuracy of between 85% and 90% as they stated on the web page. However, LB is only 0.77. Could anyone provide some insights into the mismatch? Thanks.</p>\n\n<p>Another observation is that unknown should only take 9% but the output has ~ 30% unknown.</p>",
      "rawMarkdown": "I just started this competition yesterday. I followed [Google's TF tutorials][1] and got a validation accuracy of between 85% and 90% as they stated on the web page. However, LB is only 0.77. Could anyone provide some insights into the mismatch? Thanks.\n\nAnother observation is that unknown should only take 9% but the output has ~ 30% unknown.\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition",
      "votes": 13
    },
    {
      "id": 252159,
      "postDate": "2017-12-02T11:37:52.013Z",
      "content": "<p>The solution I found to this problem is to split train/val by speaker and also re-sampling the datasets by class as to have a uniform distribution ( Otherwise the accuracy metric is over-optimistic as the unknown class is dominant) </p>",
      "rawMarkdown": "The solution I found to this problem is to split train/val by speaker and also re-sampling the datasets by class as to have a uniform distribution ( Otherwise the accuracy metric is over-optimistic as the unknown class is dominant) ",
      "votes": 7,
      "replies": [
        {
          "id": 252292,
          "postDate": "2017-12-02T17:33:17.323Z",
          "content": "<p>Thanks @CVxTz, I will take a try.</p>",
          "rawMarkdown": "Thanks @CVxTz, I will take a try.",
          "votes": 1
        },
        {
          "id": 261671,
          "postDate": "2017-12-23T14:34:49Z",
          "content": "<p>What do you mean by splitting train/val by speaker?</p>",
          "rawMarkdown": "What do you mean by splitting train/val by speaker?"
        },
        {
          "id": 261676,
          "postDate": "2017-12-23T14:50:02.203Z",
          "content": "<p>In the filenames there an Id specific to each user. I use it in order to avoid training on a word by speaker i and then evaluating on the same word by the same speaker i.</p>",
          "rawMarkdown": "In the filenames there an Id specific to each user. I use it in order to avoid training on a word by speaker i and then evaluating on the same word by the same speaker i."
        },
        {
          "id": 261874,
          "postDate": "2017-12-24T10:46:35.033Z",
          "content": "<p>Did the gap close completely? (If it closed completely, it means there's no other kind of mismatch, like acoustic environment mismatch.)</p>\n\n<p>P.S. You probably don't want to do this to the class distribution <em>during training</em>, since the private LB test set has many more unknowns than the public: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298</a></p>",
          "rawMarkdown": "Did the gap close completely? (If it closed completely, it means there's no other kind of mismatch, like acoustic environment mismatch.)\n\nP.S. You probably don't want to do this to the class distribution *during training*, since the private LB test set has many more unknowns than the public: https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298"
        }
      ]
    },
    {
      "id": 250094,
      "postDate": "2017-11-29T19:25:07.693Z",
      "content": "<p>They got 85% - 90% with 10% of train data (~6400). In the LB we use other dataset to test with 158538 examples...</p>",
      "rawMarkdown": "They got 85% - 90% with 10% of train data (~6400). In the LB we use other dataset to test with 158538 examples...",
      "votes": 5
    },
    {
      "id": 249672,
      "postDate": "2017-11-29T00:24:35.263Z",
      "rawMarkdown": "",
      "votes": 1,
      "replies": [
        {
          "id": 249676,
          "postDate": "2017-11-29T00:55:33.753Z",
          "content": "<p>I think that just means that your model is overfitting your training/validation/test data. Maybe you need to add some regularization?</p>",
          "rawMarkdown": "I think that just means that your model is overfitting your training/validation/test data. Maybe you need to add some regularization?",
          "votes": 1
        },
        {
          "id": 250117,
          "postDate": "2017-11-29T20:34:18.250Z",
          "rawMarkdown": ""
        },
        {
          "id": 250120,
          "postDate": "2017-11-29T20:42:37.437Z",
          "content": "<p>I'm curious what the distribution of the predictions is on your validation set versus the distribution of predictions on the test set. Are they similar?</p>",
          "rawMarkdown": "I'm curious what the distribution of the predictions is on your validation set versus the distribution of predictions on the test set. Are they similar?"
        },
        {
          "id": 251759,
          "postDate": "2017-12-01T17:07:00.243Z",
          "content": "<p>Deleted my posts because I thought I had made a mistake, turns out I hadn't..\nThey are similar yes. Which is of course bad as the LB has only 9% unknown while the local test/val has &gt;50%. Looking into how to solve that.</p>",
          "rawMarkdown": "Deleted my posts because I thought I had made a mistake, turns out I hadn't..\nThey are similar yes. Which is of course bad as the LB has only 9% unknown while the local test/val has &gt;50%. Looking into how to solve that.",
          "votes": 2
        }
      ]
    },
    {
      "id": 248573,
      "postDate": "2017-11-26T14:06:09.963Z",
      "content": "<p>If your validation score is lower than the LB score then it means that your validation set is not fully representative of the test set. This makes sense because the test set contains speakers that are not in your validation set, noisy data, and words that are not in the training data.</p>",
      "rawMarkdown": "If your validation score is lower than the LB score then it means that your validation set is not fully representative of the test set. This makes sense because the test set contains speakers that are not in your validation set, noisy data, and words that are not in the training data.",
      "votes": 1,
      "replies": [
        {
          "id": 248633,
          "postDate": "2017-11-26T16:33:06.507Z",
          "content": "<p>Thanks, @Human Analog. That helped me understand the problem.</p>",
          "rawMarkdown": "Thanks, @Human Analog. That helped me understand the problem.",
          "votes": 1
        }
      ]
    },
    {
      "id": 248553,
      "postDate": "2017-11-26T13:30:41.383Z",
      "content": "<p>From the train set there are samples from 30 keywords, of those 20 go into unknown. So it seems natural that unknown would be the largest class in the test set as well. </p>",
      "rawMarkdown": "From the train set there are samples from 30 keywords, of those 20 go into unknown. So it seems natural that unknown would be the largest class in the test set as well. ",
      "votes": 2,
      "replies": [
        {
          "id": 248632,
          "postDate": "2017-11-26T16:31:44.920Z",
          "content": "<p>Good to know that. Thanks, @Ben Lai.</p>",
          "rawMarkdown": "Good to know that. Thanks, @Ben Lai."
        },
        {
          "id": 248709,
          "postDate": "2017-11-26T20:09:27.687Z",
          "content": "<p>But if 0.09 is maximal constant prediction score, the balance is different in the test set.</p>",
          "rawMarkdown": "But if 0.09 is maximal constant prediction score, the balance is different in the test set.",
          "votes": 1
        },
        {
          "id": 248711,
          "postDate": "2017-11-26T20:12:40.653Z",
          "content": "<p>Not necessarily; the balance is different only in the part of the test set that is used to compute the LB score. Across the <em>entire</em> test set, including the examples that are ignored for scoring, the \"unknown\" class appears to be the largest by far.</p>",
          "rawMarkdown": "Not necessarily; the balance is different only in the part of the test set that is used to compute the LB score. Across the *entire* test set, including the examples that are ignored for scoring, the \"unknown\" class appears to be the largest by far.",
          "votes": 3
        },
        {
          "id": 249595,
          "postDate": "2017-11-28T19:46:02.373Z",
          "content": "<p>Right, thank you.</p>",
          "rawMarkdown": "Right, thank you."
        }
      ]
    },
    {
      "id": 251211,
      "postDate": "2017-11-30T20:44:47.097Z",
      "content": "<p>Your validation set probably has the same speakers from your training set.</p>\n\n<p>“the Speech Commands set has people repeating the same word multiple times. Each one of those repetitions is likely to be pretty close to the others, so if training was overfitting and memorizing one, it could perform unrealistically well when it saw a very similar copy in the test set. To avoid this danger, Speech Commands trys to ensure that all clips featuring the same word spoken by a single person are put into the same partition. Clips are assigned to training, test, or validation sets based on a hash of their filename, to ensure that the assignments remain steady even as new clips are added and avoid any training samples migrating into the other sets.”</p>",
      "rawMarkdown": "Your validation set probably has the same speakers from your training set.\n\n“the Speech Commands set has people repeating the same word multiple times. Each one of those repetitions is likely to be pretty close to the others, so if training was overfitting and memorizing one, it could perform unrealistically well when it saw a very similar copy in the test set. To avoid this danger, Speech Commands trys to ensure that all clips featuring the same word spoken by a single person are put into the same partition. Clips are assigned to training, test, or validation sets based on a hash of their filename, to ensure that the assignments remain steady even as new clips are added and avoid any training samples migrating into the other sets.”",
      "replies": [
        {
          "id": 251250,
          "postDate": "2017-11-30T21:37:38.727Z",
          "content": "<p>The provided val/test sets have no overlap in speakers with the training set.</p>",
          "rawMarkdown": "The provided val/test sets have no overlap in speakers with the training set.",
          "votes": 1
        },
        {
          "id": 251251,
          "postDate": "2017-11-30T21:39:42.537Z",
          "content": "<p>There's something wrong in your comment. They have provided only two sets: train and test. If the author of this thread created a validation set from the train set, then validation probably has the same speakers from train.</p>",
          "rawMarkdown": "There's something wrong in your comment. They have provided only two sets: train and test. If the author of this thread created a validation set from the train set, then validation probably has the same speakers from train.",
          "votes": 1
        },
        {
          "id": 251260,
          "postDate": "2017-11-30T21:54:57.430Z",
          "content": "<p>I think @RuAB refers to the suggested train/val/test split that is provided as part of the training set. If you make a random split then speakers will have overlap, but by using the provided split they won't.</p>",
          "rawMarkdown": "I think @RuAB refers to the suggested train/val/test split that is provided as part of the training set. If you make a random split then speakers will have overlap, but by using the provided split they won't.",
          "votes": 1
        }
      ]
    },
    {
      "id": 259409,
      "postDate": "2017-12-18T10:05:01.040Z",
      "content": "<p>I only got 0.71 on LB after following the tutorial and changing the  '_unknown_' and \"_silence_\".  Is there anyone same as me? Did you change any of the default parameters to get 0.77?</p>",
      "rawMarkdown": "I only got 0.71 on LB after following the tutorial and changing the  '\\_unknown\\_' and \"\\_silence\\_\".  Is there anyone same as me? Did you change any of the default parameters to get 0.77?",
      "replies": [
        {
          "id": 259546,
          "postDate": "2017-12-18T15:59:39.850Z",
          "content": "<p>0.77 - 0.78 are a normal score with this tutorial.</p>\n\n<p>So, check your code. </p>",
          "rawMarkdown": "0.77 - 0.78 are a normal score with this tutorial.\n\nSo, check your code. ",
          "votes": 2
        }
      ]
    },
    {
      "id": 249506,
      "postDate": "2017-11-28T16:35:10.263Z",
      "content": "<p>I believe you mean to say that 'silence' should take 9%, as evidenced by the 9% 'Silence is Golden' benchmark.</p>",
      "rawMarkdown": "I believe you mean to say that 'silence' should take 9%, as evidenced by the 9% 'Silence is Golden' benchmark.",
      "replies": [
        {
          "id": 249573,
          "postDate": "2017-11-28T19:00:09.823Z",
          "content": "<p>I also submited a file with 'unknown'. Same score with 'silence.'</p>",
          "rawMarkdown": "I also submited a file with 'unknown'. Same score with 'silence.'"
        },
        {
          "id": 249876,
          "postDate": "2017-11-29T12:11:38.477Z",
          "content": "<p>I also submit  \"unknown\" file, too! But if \"unknown\" appears most in test data, why take 9%, not more?</p>",
          "rawMarkdown": "I also submit  \"unknown\" file, too! But if \"unknown\" appears most in test data, why take 9%, not more?",
          "votes": 1
        },
        {
          "id": 249900,
          "postDate": "2017-11-29T13:26:10.150Z",
          "content": "<p>On the Data page it says: \"Not all of the files are evaluated for the leaderboard score.\"</p>",
          "rawMarkdown": "On the Data page it says: \"Not all of the files are evaluated for the leaderboard score.\"",
          "votes": 1
        },
        {
          "id": 250986,
          "postDate": "2017-11-30T15:10:39.600Z",
          "content": "<p>So the data for LB score , with 9% silence and 9% unknown. </p>",
          "rawMarkdown": "So the data for LB score , with 9% silence and 9% unknown. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 248742,
      "postDate": "2017-11-26T21:48:43.643Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 248745,
          "postDate": "2017-11-26T21:58:07.597Z",
          "content": "<p>Haha. Did you change ' _ unknown _ ' to 'unknown' and ' _ silence _ ' to 'silence'? Little trick.</p>",
          "rawMarkdown": "Haha. Did you change ' _ unknown _ ' to 'unknown' and ' _ silence _ ' to 'silence'? Little trick.",
          "votes": 7
        },
        {
          "id": 248758,
          "postDate": "2017-11-26T22:39:08.113Z",
          "rawMarkdown": "",
          "votes": 1,
          "isDeleted": true
        },
        {
          "id": 249527,
          "postDate": "2017-11-28T17:10:35.710Z",
          "content": "<p>OMG!\nI was stuck at .65 after days of tweaking the example code and then I see your post.\nAfter fixing this mistake, it's up to .81.</p>",
          "rawMarkdown": "OMG!\nI was stuck at .65 after days of tweaking the example code and then I see your post.\nAfter fixing this mistake, it's up to .81.",
          "votes": 4
        },
        {
          "id": 249575,
          "postDate": "2017-11-28T19:01:46.200Z",
          "content": "<p>@bsp2020, I should have kept it as a secret. Haha.</p>",
          "rawMarkdown": "@bsp2020, I should have kept it as a secret. Haha.",
          "votes": 2
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 252159,
      "author_name": "CVxTz",
      "author_url": "",
      "post_date": "2017-12-02T11:37:52.013000",
      "content": "<p>The solution I found to this problem is to split train/val by speaker and also re-sampling the datasets by class as to have a uniform distribution ( Otherwise the accuracy metric is over-optimistic as the unknown class is dominant) </p>",
      "votes": 7,
      "replies": [
        {
          "id": 252292,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2017-12-02T17:33:17.323000",
          "content": "<p>Thanks @CVxTz, I will take a try.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 261671,
          "author_name": "Lugi",
          "author_url": "",
          "post_date": "2017-12-23T14:34:49",
          "content": "<p>What do you mean by splitting train/val by speaker?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 261676,
          "author_name": "CVxTz",
          "author_url": "",
          "post_date": "2017-12-23T14:50:02.203000",
          "content": "<p>In the filenames there an Id specific to each user. I use it in order to avoid training on a word by speaker i and then evaluating on the same word by the same speaker i.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 261874,
          "author_name": "Aleksandr Dubinsky",
          "author_url": "",
          "post_date": "2017-12-24T10:46:35.033000",
          "content": "<p>Did the gap close completely? (If it closed completely, it means there's no other kind of mismatch, like acoustic environment mismatch.)</p>\n\n<p>P.S. You probably don't want to do this to the class distribution <em>during training</em>, since the private LB test set has many more unknowns than the public: <a href=\"https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298\">https://www.kaggle.com/c/tensorflow-speech-recognition-challenge/discussion/46298</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 250094,
      "author_name": "Dario Lopez Padial",
      "author_url": "",
      "post_date": "2017-11-29T19:25:07.693000",
      "content": "<p>They got 85% - 90% with 10% of train data (~6400). In the LB we use other dataset to test with 158538 examples...</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 249672,
      "author_name": "RuAB",
      "author_url": "",
      "post_date": "2017-11-29T00:24:35.263000",
      "content": "",
      "votes": 1,
      "replies": [
        {
          "id": 249676,
          "author_name": "bsp2020",
          "author_url": "",
          "post_date": "2017-11-29T00:55:33.753000",
          "content": "<p>I think that just means that your model is overfitting your training/validation/test data. Maybe you need to add some regularization?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 250117,
          "author_name": "RuAB",
          "author_url": "",
          "post_date": "2017-11-29T20:34:18.250000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 250120,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2017-11-29T20:42:37.437000",
          "content": "<p>I'm curious what the distribution of the predictions is on your validation set versus the distribution of predictions on the test set. Are they similar?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 251759,
          "author_name": "RuAB",
          "author_url": "",
          "post_date": "2017-12-01T17:07:00.243000",
          "content": "<p>Deleted my posts because I thought I had made a mistake, turns out I hadn't..\nThey are similar yes. Which is of course bad as the LB has only 9% unknown while the local test/val has &gt;50%. Looking into how to solve that.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 248573,
      "author_name": "Human Analog",
      "author_url": "",
      "post_date": "2017-11-26T14:06:09.963000",
      "content": "<p>If your validation score is lower than the LB score then it means that your validation set is not fully representative of the test set. This makes sense because the test set contains speakers that are not in your validation set, noisy data, and words that are not in the training data.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 248633,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2017-11-26T16:33:06.507000",
          "content": "<p>Thanks, @Human Analog. That helped me understand the problem.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 248553,
      "author_name": "Ben Lai",
      "author_url": "",
      "post_date": "2017-11-26T13:30:41.383000",
      "content": "<p>From the train set there are samples from 30 keywords, of those 20 go into unknown. So it seems natural that unknown would be the largest class in the test set as well. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 248632,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2017-11-26T16:31:44.920000",
          "content": "<p>Good to know that. Thanks, @Ben Lai.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 248709,
          "author_name": "ThomasV",
          "author_url": "",
          "post_date": "2017-11-26T20:09:27.687000",
          "content": "<p>But if 0.09 is maximal constant prediction score, the balance is different in the test set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 248711,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2017-11-26T20:12:40.653000",
          "content": "<p>Not necessarily; the balance is different only in the part of the test set that is used to compute the LB score. Across the <em>entire</em> test set, including the examples that are ignored for scoring, the \"unknown\" class appears to be the largest by far.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 249595,
          "author_name": "ThomasV",
          "author_url": "",
          "post_date": "2017-11-28T19:46:02.373000",
          "content": "<p>Right, thank you.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 251211,
      "author_name": "Rafael Barbolo",
      "author_url": "",
      "post_date": "2017-11-30T20:44:47.097000",
      "content": "<p>Your validation set probably has the same speakers from your training set.</p>\n\n<p>“the Speech Commands set has people repeating the same word multiple times. Each one of those repetitions is likely to be pretty close to the others, so if training was overfitting and memorizing one, it could perform unrealistically well when it saw a very similar copy in the test set. To avoid this danger, Speech Commands trys to ensure that all clips featuring the same word spoken by a single person are put into the same partition. Clips are assigned to training, test, or validation sets based on a hash of their filename, to ensure that the assignments remain steady even as new clips are added and avoid any training samples migrating into the other sets.”</p>",
      "votes": 0,
      "replies": [
        {
          "id": 251250,
          "author_name": "RuAB",
          "author_url": "",
          "post_date": "2017-11-30T21:37:38.727000",
          "content": "<p>The provided val/test sets have no overlap in speakers with the training set.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 251251,
          "author_name": "Rafael Barbolo",
          "author_url": "",
          "post_date": "2017-11-30T21:39:42.537000",
          "content": "<p>There's something wrong in your comment. They have provided only two sets: train and test. If the author of this thread created a validation set from the train set, then validation probably has the same speakers from train.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 251260,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2017-11-30T21:54:57.430000",
          "content": "<p>I think @RuAB refers to the suggested train/val/test split that is provided as part of the training set. If you make a random split then speakers will have overlap, but by using the provided split they won't.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 259409,
      "author_name": "Qixiu",
      "author_url": "",
      "post_date": "2017-12-18T10:05:01.040000",
      "content": "<p>I only got 0.71 on LB after following the tutorial and changing the  '_unknown_' and \"_silence_\".  Is there anyone same as me? Did you change any of the default parameters to get 0.77?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 259546,
          "author_name": "Dario Lopez Padial",
          "author_url": "",
          "post_date": "2017-12-18T15:59:39.850000",
          "content": "<p>0.77 - 0.78 are a normal score with this tutorial.</p>\n\n<p>So, check your code. </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 249506,
      "author_name": "rtmitch",
      "author_url": "",
      "post_date": "2017-11-28T16:35:10.263000",
      "content": "<p>I believe you mean to say that 'silence' should take 9%, as evidenced by the 9% 'Silence is Golden' benchmark.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 249573,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2017-11-28T19:00:09.823000",
          "content": "<p>I also submited a file with 'unknown'. Same score with 'silence.'</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 249876,
          "author_name": "whiteworld",
          "author_url": "",
          "post_date": "2017-11-29T12:11:38.477000",
          "content": "<p>I also submit  \"unknown\" file, too! But if \"unknown\" appears most in test data, why take 9%, not more?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 249900,
          "author_name": "Human Analog",
          "author_url": "",
          "post_date": "2017-11-29T13:26:10.150000",
          "content": "<p>On the Data page it says: \"Not all of the files are evaluated for the leaderboard score.\"</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 250986,
          "author_name": "whiteworld",
          "author_url": "",
          "post_date": "2017-11-30T15:10:39.600000",
          "content": "<p>So the data for LB score , with 9% silence and 9% unknown. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 248742,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-11-26T21:48:43.643000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 248745,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2017-11-26T21:58:07.597000",
          "content": "<p>Haha. Did you change ' _ unknown _ ' to 'unknown' and ' _ silence _ ' to 'silence'? Little trick.</p>",
          "votes": 7,
          "replies": []
        },
        {
          "id": 248758,
          "author_name": "",
          "author_url": "",
          "post_date": "2017-11-26T22:39:08.113000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 249527,
          "author_name": "bsp2020",
          "author_url": "",
          "post_date": "2017-11-28T17:10:35.710000",
          "content": "<p>OMG!\nI was stuck at .65 after days of tweaking the example code and then I see your post.\nAfter fixing this mistake, it's up to .81.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 249575,
          "author_name": "Shujian Liu",
          "author_url": "",
          "post_date": "2017-11-28T19:01:46.200000",
          "content": "<p>@bsp2020, I should have kept it as a secret. Haha.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "248416": "I just started this competition yesterday. I followed [Google's TF tutorials][1] and got a validation accuracy of between 85% and 90% as they stated on the web page. However, LB is only 0.77. Could anyone provide some insights into the mismatch? Thanks.\n\nAnother observation is that unknown should only take 9% but the output has ~ 30% unknown.\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition",
    "252159": "The solution I found to this problem is to split train/val by speaker and also re-sampling the datasets by class as to have a uniform distribution ( Otherwise the accuracy metric is over-optimistic as the unknown class is dominant) ",
    "250094": "They got 85% - 90% with 10% of train data (~6400). In the LB we use other dataset to test with 158538 examples...",
    "249672": "",
    "248573": "If your validation score is lower than the LB score then it means that your validation set is not fully representative of the test set. This makes sense because the test set contains speakers that are not in your validation set, noisy data, and words that are not in the training data.",
    "248553": "From the train set there are samples from 30 keywords, of those 20 go into unknown. So it seems natural that unknown would be the largest class in the test set as well. ",
    "251211": "Your validation set probably has the same speakers from your training set.\n\n“the Speech Commands set has people repeating the same word multiple times. Each one of those repetitions is likely to be pretty close to the others, so if training was overfitting and memorizing one, it could perform unrealistically well when it saw a very similar copy in the test set. To avoid this danger, Speech Commands trys to ensure that all clips featuring the same word spoken by a single person are put into the same partition. Clips are assigned to training, test, or validation sets based on a hash of their filename, to ensure that the assignments remain steady even as new clips are added and avoid any training samples migrating into the other sets.”",
    "259409": "I only got 0.71 on LB after following the tutorial and changing the  '\\_unknown\\_' and \"\\_silence\\_\".  Is there anyone same as me? Did you change any of the default parameters to get 0.77?",
    "249506": "I believe you mean to say that 'silence' should take 9%, as evidenced by the 9% 'Silence is Golden' benchmark.",
    "248742": ""
  }
}