{
  "id": 43696,
  "title": "How was silence data generated in test set?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43696",
  "author_name": "",
  "post_date": "2017-11-18T00:16:56.425626300Z",
  "votes": 7,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I may be missing something, but I can't find any information regarding how the silence data in the test set was generated.  Are the samples artificially generated similar to the data in the provided background noise samples, or are they non-speech portions of the audio collected from the speakers during data collection?</p>",
  "messages": [
    {
      "id": "245293",
      "postDate": "11/18/2017 00:16:56",
      "content": "<p>I may be missing something, but I can't find any information regarding how the silence data in the test set was generated.  Are the samples artificially generated similar to the data in the provided background noise samples, or are they non-speech portions of the audio collected from the speakers during data collection?</p>",
      "rawMarkdown": "I may be missing something, but I can't find any information regarding how the silence data in the test set was generated.  Are the samples artificially generated similar to the data in the provided background noise samples, or are they non-speech portions of the audio collected from the speakers during data collection?",
      "votes": null
    },
    {
      "id": "245300",
      "postDate": "11/18/2017 01:21:27",
      "content": "<p>The silence data was generated by taking similar .wavs to the ones in the background_noise folder of the released data set and extracting one-second clips with randomly-scaled volumes. The ones in the source background_noise folder are a mix of synthetic pink/white noise and field recordings of things like quiet rooms, some with fans or water backgrounds.</p>",
      "rawMarkdown": "The silence data was generated by taking similar .wavs to the ones in the background_noise folder of the released data set and extracting one-second clips with randomly-scaled volumes. The ones in the source background_noise folder are a mix of synthetic pink/white noise and field recordings of things like quiet rooms, some with fans or water backgrounds.",
      "votes": null
    },
    {
      "id": "245302",
      "postDate": "11/18/2017 01:23:58",
      "content": "<p>From the <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">tutorial</a> referenced in the overview the section on background noise and silence describes this, looks like they come from combinations of deliberate recordings of noise and generated white noise. I haven't read the <a href=\"http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf\">paper</a> yet - this may have more detail. I don't know how much the competition data differs from these.</p>",
      "rawMarkdown": "From the [tutorial][1] referenced in the overview the section on background noise and silence describes this, looks like they come from combinations of deliberate recordings of noise and generated white noise. I haven't read the [paper][2] yet - this may have more detail. I don't know how much the competition data differs from these.\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 245300,
      "author_name": "petewarden",
      "author_url": "",
      "post_date": "11/18/2017 01:21:27",
      "content": "<p>The silence data was generated by taking similar .wavs to the ones in the background_noise folder of the released data set and extracting one-second clips with randomly-scaled volumes. The ones in the source background_noise folder are a mix of synthetic pink/white noise and field recordings of things like quiet rooms, some with fans or water backgrounds.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 245302,
      "author_name": "stprior",
      "author_url": "",
      "post_date": "11/18/2017 01:23:58",
      "content": "<p>From the <a href=\"https://www.tensorflow.org/versions/master/tutorials/audio_recognition\">tutorial</a> referenced in the overview the section on background noise and silence describes this, looks like they come from combinations of deliberate recordings of noise and generated white noise. I haven't read the <a href=\"http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf\">paper</a> yet - this may have more detail. I don't know how much the competition data differs from these.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "245293": "I may be missing something, but I can't find any information regarding how the silence data in the test set was generated.  Are the samples artificially generated similar to the data in the provided background noise samples, or are they non-speech portions of the audio collected from the speakers during data collection?",
    "245300": "The silence data was generated by taking similar .wavs to the ones in the background_noise folder of the released data set and extracting one-second clips with randomly-scaled volumes. The ones in the source background_noise folder are a mix of synthetic pink/white noise and field recordings of things like quiet rooms, some with fans or water backgrounds.",
    "245302": "From the [tutorial][1] referenced in the overview the section on background noise and silence describes this, looks like they come from combinations of deliberate recordings of noise and generated white noise. I haven't read the [paper][2] yet - this may have more detail. I don't know how much the competition data differs from these.\n\n\n  [1]: https://www.tensorflow.org/versions/master/tutorials/audio_recognition\n  [2]: http://www.isca-speech.org/archive/interspeech_2015/papers/i15_1478.pdf"
  },
  "source": "meta"
}