{
  "id": 45324,
  "title": "Unknown or silence",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/45324",
  "author_name": "",
  "post_date": "2017-12-09T09:01:14.644843900Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>One question about what's rule to know if this is a \"silence\" or \"unknown\" wav file.</p>\n\n<p>If wav have background sound , white noise and no human speech , will those be need to label as silence ?\nIf wav have very low energy (hard to know if this is human speech or other animal or nature sound) , this should be a \"unknown\" case ?</p>",
  "messages": [
    {
      "id": "255507",
      "postDate": "12/09/2017 09:01:14",
      "content": "<p>One question about what's rule to know if this is a \"silence\" or \"unknown\" wav file.</p>\n\n<p>If wav have background sound , white noise and no human speech , will those be need to label as silence ?\nIf wav have very low energy (hard to know if this is human speech or other animal or nature sound) , this should be a \"unknown\" case ?</p>",
      "rawMarkdown": "One question about what's rule to know if this is a \"silence\" or \"unknown\" wav file.\n\nIf wav have background sound , white noise and no human speech , will those be need to label as silence ?\nIf wav have very low energy (hard to know if this is human speech or other animal or nature sound) , this should be a \"unknown\" case ?",
      "votes": null
    },
    {
      "id": "255535",
      "postDate": "12/09/2017 10:35:25",
      "content": "<p>In theory: Unknown is a recognizable <strong>word</strong> not in {yes, no, up, down, left, right, on, off, stop, go} and silence is background noise. In practice, there may be gray areas such as extremely low volume levels, simultaneous speakers or more than one word per sample. I guess gray areas are intentional and part of the game.</p>",
      "rawMarkdown": "In theory: Unknown is a recognizable **word** not in {yes, no, up, down, left, right, on, off, stop, go} and silence is background noise. In practice, there may be gray areas such as extremely low volume levels, simultaneous speakers or more than one word per sample. I guess gray areas are intentional and part of the game.",
      "votes": null
    },
    {
      "id": "255539",
      "postDate": "12/09/2017 10:42:15",
      "content": "<p>background noise (if no human speech) should be silence , make sense.  But if there is human speech detected with very low volume (or we can say 20-20khz ) also not in word list , those should be category as \"unknown\" case as well ?</p>",
      "rawMarkdown": "background noise (if no human speech) should be silence , make sense.  But if there is human speech detected with very low volume (or we can say 20-20khz ) also not in word list , those should be category as \"unknown\" case as well ?",
      "votes": null
    },
    {
      "id": "256450",
      "postDate": "12/12/2017 00:13:56",
      "content": "<p>I would assume that <em>voiced</em> audios should result in unknown... But that is definitely gray area, I agree</p>",
      "rawMarkdown": "I would assume that *voiced* audios should result in unknown... But that is definitely gray area, I agree",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 255535,
      "author_name": "batangas",
      "author_url": "",
      "post_date": "12/09/2017 10:35:25",
      "content": "<p>In theory: Unknown is a recognizable <strong>word</strong> not in {yes, no, up, down, left, right, on, off, stop, go} and silence is background noise. In practice, there may be gray areas such as extremely low volume levels, simultaneous speakers or more than one word per sample. I guess gray areas are intentional and part of the game.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 255539,
      "author_name": "chesterkuo",
      "author_url": "",
      "post_date": "12/09/2017 10:42:15",
      "content": "<p>background noise (if no human speech) should be silence , make sense.  But if there is human speech detected with very low volume (or we can say 20-20khz ) also not in word list , those should be category as \"unknown\" case as well ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 256450,
      "author_name": "douglasandrade",
      "author_url": "",
      "post_date": "12/12/2017 00:13:56",
      "content": "<p>I would assume that <em>voiced</em> audios should result in unknown... But that is definitely gray area, I agree</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "255507": "One question about what's rule to know if this is a \"silence\" or \"unknown\" wav file.\n\nIf wav have background sound , white noise and no human speech , will those be need to label as silence ?\nIf wav have very low energy (hard to know if this is human speech or other animal or nature sound) , this should be a \"unknown\" case ?",
    "255535": "In theory: Unknown is a recognizable **word** not in {yes, no, up, down, left, right, on, off, stop, go} and silence is background noise. In practice, there may be gray areas such as extremely low volume levels, simultaneous speakers or more than one word per sample. I guess gray areas are intentional and part of the game.",
    "255539": "background noise (if no human speech) should be silence , make sense.  But if there is human speech detected with very low volume (or we can say 20-20khz ) also not in word list , those should be category as \"unknown\" case as well ?",
    "256450": "I would assume that *voiced* audios should result in unknown... But that is definitely gray area, I agree"
  },
  "source": "meta"
}