{
  "id": 45642,
  "title": "A  0.09  submission , which part is wrong ?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45642",
  "author_name": "Steven Du",
  "post_date": "2017-12-14T02:44:22.346000",
  "votes": 0,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi,  my best submission is 0.09 on the leaderboard but my local cross validation  (0.88) is not that bad, I rebuild my system twice today and find no problems. So I listed what I had done, maybe someone could help to figure out which part is wrong.</p>\n\n<h1>Training</h1>\n\n<h2>Training X</h2>\n\n<p>Standard training data.  Silence files are filtered out by VAD.</p>\n\n<h2>Training Y</h2>\n\n<p>The training targets are 30 classes</p>\n\n<p><code>\n ['sheila', 'seven', 'right', 'house', 'dog', 'four', 'zero', 'go', 'yes', 'down', 'no', 'wow', 'six', 'three', 'bird', 'happy', 'marvin', 'stop', 'eight', 'two', 'one', 'on', 'off', 'tree', 'up', 'bed', 'cat', 'nine', 'five', 'left']\n</code></p>\n\n<h2>Cross Validation</h2>\n\n<h3>10 fold CV</h3>\n\n<p><a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html#sklearn.model_selection.GroupKFold\">10 fold CV  spitted</a>  by speaker label</p>\n\n<p>Then I build a model which gives avg 0.88 classification accuracy over 10-fold CV.</p>\n\n<h3>CV by validation_list.txt</h3>\n\n<p>Retrain the model and validate by <code>/train/validation_list.txt</code>, the score is also 0.88</p>\n\n<h1>Testing</h1>\n\n<p>The trained model predicts 30 class labels on testing set,  labels  not in\n<code>[yes, no, up, down, left, right, on, off, stop, go]</code>\nare marked as \"unknown\" and silence files are marked by VAD.</p>\n\n<p>And this approach will give 0.09 ...</p>\n\n<p>Please, tell me why ..</p>",
  "messages": [
    {
      "id": 257349,
      "postDate": "2017-12-14T04:36:52.377Z",
      "content": "<p>Hi Steven,</p>\n\n<p>Well that's very unlikely to happen. You can do these checks and see for any bugs:</p>\n\n<p>1) if you are normalizing training and validation sets, are you also normalizing test set? you need to normalize test set too.</p>\n\n<p>2) you are probably not writing the submission file correctly, i.e. the file name versus the label is not being printed properly. Most likely this is the issue.</p>\n\n<p>3) try printing all the labels of the test set. If all the labels are the same class (e.g. all labels are 'yes') or two classes (e.g. all labels are either 'yes' or 'no'), then try using batch normalization in your architecture. I encountered this issue for two other projects and batch normalization fixed it.</p>\n\n<p>Hope this helps.</p>",
      "rawMarkdown": "Hi Steven,\n\nWell that's very unlikely to happen. You can do these checks and see for any bugs:\n\n1) if you are normalizing training and validation sets, are you also normalizing test set? you need to normalize test set too.\n\n2) you are probably not writing the submission file correctly, i.e. the file name versus the label is not being printed properly. Most likely this is the issue.\n\n3) try printing all the labels of the test set. If all the labels are the same class (e.g. all labels are 'yes') or two classes (e.g. all labels are either 'yes' or 'no'), then try using batch normalization in your architecture. I encountered this issue for two other projects and batch normalization fixed it.\n\nHope this helps.",
      "votes": 2,
      "replies": [
        {
          "id": 257429,
          "postDate": "2017-12-14T08:24:28.653Z",
          "content": "<p>Wow, thanks,  I read the wrong index when writing the submission file.</p>",
          "rawMarkdown": "Wow, thanks,  I read the wrong index when writing the submission file."
        },
        {
          "id": 257441,
          "postDate": "2017-12-14T08:44:38.313Z",
          "content": "<p>awesome!</p>",
          "rawMarkdown": "awesome!"
        }
      ]
    },
    {
      "id": 260689,
      "postDate": "2017-12-20T18:26:39.367Z",
      "content": "<p>Do you have the test set only giving a list of all one class? 0.09 percent is the same percentage as a file submitted with only Silence for every file.</p>",
      "rawMarkdown": "Do you have the test set only giving a list of all one class? 0.09 percent is the same percentage as a file submitted with only Silence for every file."
    },
    {
      "id": 257353,
      "postDate": "2017-12-14T04:52:40.497Z",
      "rawMarkdown": ""
    },
    {
      "id": 257331,
      "postDate": "2017-12-14T02:44:22.347Z",
      "content": "<p>Hi,  my best submission is 0.09 on the leaderboard but my local cross validation  (0.88) is not that bad, I rebuild my system twice today and find no problems. So I listed what I had done, maybe someone could help to figure out which part is wrong.</p>\n\n<h1>Training</h1>\n\n<h2>Training X</h2>\n\n<p>Standard training data.  Silence files are filtered out by VAD.</p>\n\n<h2>Training Y</h2>\n\n<p>The training targets are 30 classes</p>\n\n<p><code>\n ['sheila', 'seven', 'right', 'house', 'dog', 'four', 'zero', 'go', 'yes', 'down', 'no', 'wow', 'six', 'three', 'bird', 'happy', 'marvin', 'stop', 'eight', 'two', 'one', 'on', 'off', 'tree', 'up', 'bed', 'cat', 'nine', 'five', 'left']\n</code></p>\n\n<h2>Cross Validation</h2>\n\n<h3>10 fold CV</h3>\n\n<p><a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html#sklearn.model_selection.GroupKFold\">10 fold CV  spitted</a>  by speaker label</p>\n\n<p>Then I build a model which gives avg 0.88 classification accuracy over 10-fold CV.</p>\n\n<h3>CV by validation_list.txt</h3>\n\n<p>Retrain the model and validate by <code>/train/validation_list.txt</code>, the score is also 0.88</p>\n\n<h1>Testing</h1>\n\n<p>The trained model predicts 30 class labels on testing set,  labels  not in\n<code>[yes, no, up, down, left, right, on, off, stop, go]</code>\nare marked as \"unknown\" and silence files are marked by VAD.</p>\n\n<p>And this approach will give 0.09 ...</p>\n\n<p>Please, tell me why ..</p>",
      "rawMarkdown": "Hi,  my best submission is 0.09 on the leaderboard but my local cross validation  (0.88) is not that bad, I rebuild my system twice today and find no problems. So I listed what I had done, maybe someone could help to figure out which part is wrong.\n\n# Training\n## Training X\nStandard training data.  Silence files are filtered out by VAD.\n## Training Y\nThe training targets are 30 classes\n\n```\n ['sheila', 'seven', 'right', 'house', 'dog', 'four', 'zero', 'go', 'yes', 'down', 'no', 'wow', 'six', 'three', 'bird', 'happy', 'marvin', 'stop', 'eight', 'two', 'one', 'on', 'off', 'tree', 'up', 'bed', 'cat', 'nine', 'five', 'left']\n```\n\n## Cross Validation  \n\n### 10 fold CV\n[10 fold CV  spitted](http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html#sklearn.model_selection.GroupKFold)  by speaker label\n\nThen I build a model which gives avg 0.88 classification accuracy over 10-fold CV.\n### CV by validation_list.txt\nRetrain the model and validate by ```/train/validation_list.txt ```, the score is also 0.88\n\n# Testing\n\nThe trained model predicts 30 class labels on testing set,  labels  not in\n```  [yes, no, up, down, left, right, on, off, stop, go] ```\nare marked as \"unknown\" and silence files are marked by VAD.\n\nAnd this approach will give 0.09 ...\n\nPlease, tell me why ..\n\n"
    },
    {
      "id": 259241,
      "postDate": "2017-12-18T01:37:20.100Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 257349,
      "author_name": "Sahil Rajesh Dhayalkar",
      "author_url": "",
      "post_date": "2017-12-14T04:36:52.377000",
      "content": "<p>Hi Steven,</p>\n\n<p>Well that's very unlikely to happen. You can do these checks and see for any bugs:</p>\n\n<p>1) if you are normalizing training and validation sets, are you also normalizing test set? you need to normalize test set too.</p>\n\n<p>2) you are probably not writing the submission file correctly, i.e. the file name versus the label is not being printed properly. Most likely this is the issue.</p>\n\n<p>3) try printing all the labels of the test set. If all the labels are the same class (e.g. all labels are 'yes') or two classes (e.g. all labels are either 'yes' or 'no'), then try using batch normalization in your architecture. I encountered this issue for two other projects and batch normalization fixed it.</p>\n\n<p>Hope this helps.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 257429,
          "author_name": "Steven Du",
          "author_url": "",
          "post_date": "2017-12-14T08:24:28.653000",
          "content": "<p>Wow, thanks,  I read the wrong index when writing the submission file.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 257441,
          "author_name": "Sahil Rajesh Dhayalkar",
          "author_url": "",
          "post_date": "2017-12-14T08:44:38.313000",
          "content": "<p>awesome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 260689,
      "author_name": "Altourus",
      "author_url": "",
      "post_date": "2017-12-20T18:26:39.367000",
      "content": "<p>Do you have the test set only giving a list of all one class? 0.09 percent is the same percentage as a file submitted with only Silence for every file.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 257353,
      "author_name": "Jeff",
      "author_url": "",
      "post_date": "2017-12-14T04:52:40.497000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 259241,
      "author_name": "",
      "author_url": "",
      "post_date": "2017-12-18T01:37:20.100000",
      "content": "",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "257349": "Hi Steven,\n\nWell that's very unlikely to happen. You can do these checks and see for any bugs:\n\n1) if you are normalizing training and validation sets, are you also normalizing test set? you need to normalize test set too.\n\n2) you are probably not writing the submission file correctly, i.e. the file name versus the label is not being printed properly. Most likely this is the issue.\n\n3) try printing all the labels of the test set. If all the labels are the same class (e.g. all labels are 'yes') or two classes (e.g. all labels are either 'yes' or 'no'), then try using batch normalization in your architecture. I encountered this issue for two other projects and batch normalization fixed it.\n\nHope this helps.",
    "260689": "Do you have the test set only giving a list of all one class? 0.09 percent is the same percentage as a file submitted with only Silence for every file.",
    "257353": "",
    "257331": "Hi,  my best submission is 0.09 on the leaderboard but my local cross validation  (0.88) is not that bad, I rebuild my system twice today and find no problems. So I listed what I had done, maybe someone could help to figure out which part is wrong.\n\n# Training\n## Training X\nStandard training data.  Silence files are filtered out by VAD.\n## Training Y\nThe training targets are 30 classes\n\n```\n ['sheila', 'seven', 'right', 'house', 'dog', 'four', 'zero', 'go', 'yes', 'down', 'no', 'wow', 'six', 'three', 'bird', 'happy', 'marvin', 'stop', 'eight', 'two', 'one', 'on', 'off', 'tree', 'up', 'bed', 'cat', 'nine', 'five', 'left']\n```\n\n## Cross Validation  \n\n### 10 fold CV\n[10 fold CV  spitted](http://scikit-learn.org/stable/modules/generated/sklearn.model_selection.GroupKFold.html#sklearn.model_selection.GroupKFold)  by speaker label\n\nThen I build a model which gives avg 0.88 classification accuracy over 10-fold CV.\n### CV by validation_list.txt\nRetrain the model and validate by ```/train/validation_list.txt ```, the score is also 0.88\n\n# Testing\n\nThe trained model predicts 30 class labels on testing set,  labels  not in\n```  [yes, no, up, down, left, right, on, off, stop, go] ```\nare marked as \"unknown\" and silence files are marked by VAD.\n\nAnd this approach will give 0.09 ...\n\nPlease, tell me why ..\n\n",
    "259241": ""
  }
}