{
  "id": 46229,
  "title": "Additional data (83,698 samples)",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46229",
  "author_name": "",
  "post_date": "2017-12-22T19:48:04.497535700Z",
  "votes": 5,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Dear all,</p>\n\n<p>I have uploaded to kaggle <a href=\"https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset\">https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset</a> a dataset with text-to-speech pronounced words, closely matching the data here.</p>\n\n<p>These data may be useful in the following ways:</p>\n\n<ul>\n<li>For this challenge, you could extend your training data to make your classifier more robust, or it might at least help you come up with better feature selection methods (e.g. spectrograms or other methods).</li>\n<li>As a robustness check, see if your method performs OK on this dataset too.</li>\n<li>Outside this challenge, I am interested whether synthetic data sufficient for classifying real datasets.</li>\n</ul>\n\n<p>More information can be found at <a href=\"https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset\">https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset</a> , with more details on how the data were created.</p>\n\n<p>As far as I understand, providing this data is within the challenge rules, because I am sharing the data with everyone. I also do not intend to submit a solution myself.</p>\n\n<p>I want to say thank you to Pete Warden and team, the volunteers and everyone submitting solutions for trying to improve the state of open source speech recognition.</p>\n\n<p>Cheers,\n                Johannes</p>",
  "messages": [
    {
      "id": "261494",
      "postDate": "12/22/2017 19:48:04",
      "content": "<p>Dear all,</p>\n\n<p>I have uploaded to kaggle <a href=\"https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset\">https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset</a> a dataset with text-to-speech pronounced words, closely matching the data here.</p>\n\n<p>These data may be useful in the following ways:</p>\n\n<ul>\n<li>For this challenge, you could extend your training data to make your classifier more robust, or it might at least help you come up with better feature selection methods (e.g. spectrograms or other methods).</li>\n<li>As a robustness check, see if your method performs OK on this dataset too.</li>\n<li>Outside this challenge, I am interested whether synthetic data sufficient for classifying real datasets.</li>\n</ul>\n\n<p>More information can be found at <a href=\"https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset\">https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset</a> , with more details on how the data were created.</p>\n\n<p>As far as I understand, providing this data is within the challenge rules, because I am sharing the data with everyone. I also do not intend to submit a solution myself.</p>\n\n<p>I want to say thank you to Pete Warden and team, the volunteers and everyone submitting solutions for trying to improve the state of open source speech recognition.</p>\n\n<p>Cheers,\n                Johannes</p>",
      "rawMarkdown": "Dear all,\n\nI have uploaded to kaggle https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset a dataset with text-to-speech pronounced words, closely matching the data here.\n\nThese data may be useful in the following ways:\n\n  - For this challenge, you could extend your training data to make your classifier more robust, or it might at least help you come up with better feature selection methods (e.g. spectrograms or other methods).\n  - As a robustness check, see if your method performs OK on this dataset too.\n  - Outside this challenge, I am interested whether synthetic data sufficient for classifying real datasets.\n\nMore information can be found at https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset , with more details on how the data were created.\n\nAs far as I understand, providing this data is within the challenge rules, because I am sharing the data with everyone. I also do not intend to submit a solution myself.\n\nI want to say thank you to Pete Warden and team, the volunteers and everyone submitting solutions for trying to improve the state of open source speech recognition.\n\nCheers,\n                Johannes",
      "votes": null
    },
    {
      "id": "261501",
      "postDate": "12/22/2017 20:04:46",
      "content": "<p>Not very sure about this, but my interpretation of the rule </p>\n\n<p>&gt; Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions.  </p>\n\n<p>suggests that you can't use them.</p>\n\n<p>EDIT: Oh, sorry, if you do not intend to submit your solution, I guess you could share... But we can't use them.</p>",
      "rawMarkdown": "Not very sure about this, but my interpretation of the rule \n\n&gt; Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions.  \n\nsuggests that you can't use them.\n\nEDIT: Oh, sorry, if you do not intend to submit your solution, I guess you could share... But we can't use them.",
      "votes": null
    },
    {
      "id": "261527",
      "postDate": "12/22/2017 20:47:27",
      "content": "<p>[edited to expand] Hmm, maybe the hosts can clarify. I think you can at least use it for feature selection, but maybe not for training. I mean if someone has previous knowledge from other works on speech (e.g. frequencies where human voice is), they can not avoid but use it. It is also not external \"data\" in the usual sense, and a \"model\" could include code to generate this from scratch as I did. I suspect their intent was to bar training against other coded human speech datasets (e.g. TED transcripts, audio books, switchboard dataset).</p>",
      "rawMarkdown": "[edited to expand] Hmm, maybe the hosts can clarify. I think you can at least use it for feature selection, but maybe not for training. I mean if someone has previous knowledge from other works on speech (e.g. frequencies where human voice is), they can not avoid but use it. It is also not external \"data\" in the usual sense, and a \"model\" could include code to generate this from scratch as I did. I suspect their intent was to bar training against other coded human speech datasets (e.g. TED transcripts, audio books, switchboard dataset).",
      "votes": null
    },
    {
      "id": "261658",
      "postDate": "12/23/2017 12:50:36",
      "content": "<p>I think it depends on how many parameters does the data augmentation step have, and how those parameters are chosen. A modern, learned TTS model has so many parameters that it, in part, has memorized the voices of many speakers. A traditional TTS model might actually use recorded samples. So I think your suggestion violates the intent of the rules. But if you, say, chose to use phonemes (as is common), you would also be relying on prior knowledge, but at least it would be a very small number of parameters.</p>",
      "rawMarkdown": "I think it depends on how many parameters does the data augmentation step have, and how those parameters are chosen. A modern, learned TTS model has so many parameters that it, in part, has memorized the voices of many speakers. A traditional TTS model might actually use recorded samples. So I think your suggestion violates the intent of the rules. But if you, say, chose to use phonemes (as is common), you would also be relying on prior knowledge, but at least it would be a very small number of parameters.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 261501,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "12/22/2017 20:04:46",
      "content": "<p>Not very sure about this, but my interpretation of the rule </p>\n\n<p>&gt; Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions.  </p>\n\n<p>suggests that you can't use them.</p>\n\n<p>EDIT: Oh, sorry, if you do not intend to submit your solution, I guess you could share... But we can't use them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 261527,
          "author_name": "jbuchner",
          "author_url": "",
          "post_date": "12/22/2017 20:47:27",
          "content": "<p>[edited to expand] Hmm, maybe the hosts can clarify. I think you can at least use it for feature selection, but maybe not for training. I mean if someone has previous knowledge from other works on speech (e.g. frequencies where human voice is), they can not avoid but use it. It is also not external \"data\" in the usual sense, and a \"model\" could include code to generate this from scratch as I did. I suspect their intent was to bar training against other coded human speech datasets (e.g. TED transcripts, audio books, switchboard dataset).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 261658,
          "author_name": "adubinsky",
          "author_url": "",
          "post_date": "12/23/2017 12:50:36",
          "content": "<p>I think it depends on how many parameters does the data augmentation step have, and how those parameters are chosen. A modern, learned TTS model has so many parameters that it, in part, has memorized the voices of many speakers. A traditional TTS model might actually use recorded samples. So I think your suggestion violates the intent of the rules. But if you, say, chose to use phonemes (as is common), you would also be relying on prior knowledge, but at least it would be a very small number of parameters.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "261494": "Dear all,\n\nI have uploaded to kaggle https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset a dataset with text-to-speech pronounced words, closely matching the data here.\n\nThese data may be useful in the following ways:\n\n  - For this challenge, you could extend your training data to make your classifier more robust, or it might at least help you come up with better feature selection methods (e.g. spectrograms or other methods).\n  - As a robustness check, see if your method performs OK on this dataset too.\n  - Outside this challenge, I am interested whether synthetic data sufficient for classifying real datasets.\n\nMore information can be found at https://www.kaggle.com/jbuchner/synthetic-speech-commands-dataset , with more details on how the data were created.\n\nAs far as I understand, providing this data is within the challenge rules, because I am sharing the data with everyone. I also do not intend to submit a solution myself.\n\nI want to say thank you to Pete Warden and team, the volunteers and everyone submitting solutions for trying to improve the state of open source speech recognition.\n\nCheers,\n                Johannes",
    "261501": "Not very sure about this, but my interpretation of the rule \n\n&gt; Unless otherwise expressly stated on the Competition Website, you may not use data other than the Competition Data to develop and test your models and Submissions.  \n\nsuggests that you can't use them.\n\nEDIT: Oh, sorry, if you do not intend to submit your solution, I guess you could share... But we can't use them.",
    "261527": "[edited to expand] Hmm, maybe the hosts can clarify. I think you can at least use it for feature selection, but maybe not for training. I mean if someone has previous knowledge from other works on speech (e.g. frequencies where human voice is), they can not avoid but use it. It is also not external \"data\" in the usual sense, and a \"model\" could include code to generate this from scratch as I did. I suspect their intent was to bar training against other coded human speech datasets (e.g. TED transcripts, audio books, switchboard dataset).",
    "261658": "I think it depends on how many parameters does the data augmentation step have, and how those parameters are chosen. A modern, learned TTS model has so many parameters that it, in part, has memorized the voices of many speakers. A traditional TTS model might actually use recorded samples. So I think your suggestion violates the intent of the rules. But if you, say, chose to use phonemes (as is common), you would also be relying on prior knowledge, but at least it would be a very small number of parameters."
  },
  "source": "meta"
}