{
  "id": 47026,
  "title": "what should be labeled to sounds like \"five....no\"",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47026",
  "author_name": "",
  "post_date": "2018-01-07T04:50:51.881225400Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I kept noticing some audio in the train and test set have two words spoken,like </p>\n\n<blockquote>\n  <p>train/audio/no/763188c4_nohash_4.wav</p>\n</blockquote>\n\n<p>contains \"five...no\" but was labeled \"no\" in the train set.\nSo should we treat it as the latter word like the sample above in the training set, or should we treat it as an \"unknown\"?  </p>",
  "messages": [
    {
      "id": "265914",
      "postDate": "01/07/2018 04:50:51",
      "content": "<p>I kept noticing some audio in the train and test set have two words spoken,like </p>\n\n<blockquote>\n  <p>train/audio/no/763188c4_nohash_4.wav</p>\n</blockquote>\n\n<p>contains \"five...no\" but was labeled \"no\" in the train set.\nSo should we treat it as the latter word like the sample above in the training set, or should we treat it as an \"unknown\"?  </p>",
      "rawMarkdown": "I kept noticing some audio in the train and test set have two words spoken,like \n\n&gt; train/audio/no/763188c4_nohash_4.wav\n\n\n\ncontains \"five...no\" but was labeled \"no\" in the train set.\nSo should we treat it as the latter word like the sample above in the training set, or should we treat it as an \"unknown\"?",
      "votes": null
    },
    {
      "id": "266593",
      "postDate": "01/09/2018 05:44:27",
      "content": "<p>I encountered this problem as well, I deal with this kind of files as non-unknown if any non-unknown word including, otherwise label it as unknown</p>",
      "rawMarkdown": "I encountered this problem as well, I deal with this kind of files as non-unknown if any non-unknown word including, otherwise label it as unknown",
      "votes": null
    },
    {
      "id": "266713",
      "postDate": "01/09/2018 14:26:26",
      "content": "<p>But How do your model predict it? is it five or no?</p>",
      "rawMarkdown": "But How do your model predict it? is it five or no?",
      "votes": null
    },
    {
      "id": "267118",
      "postDate": "01/10/2018 16:06:06",
      "content": "<p>It's according to the training data, because the numbers of 'five' and 'no' are nearly the same, the model would label this kinda files as 'five' or 'no' fairly.</p>",
      "rawMarkdown": "It's according to the training data, because the numbers of 'five' and 'no' are nearly the same, the model would label this kinda files as 'five' or 'no' fairly.",
      "votes": null
    },
    {
      "id": "267747",
      "postDate": "01/12/2018 07:15:03",
      "content": "<p>I just thought that how this wav was recorded. The former word mistakenly was recorded into the next word, so the labeling is for the latter word. Does it make sense?</p>",
      "rawMarkdown": "I just thought that how this wav was recorded. The former word mistakenly was recorded into the next word, so the labeling is for the latter word. Does it make sense?",
      "votes": null
    },
    {
      "id": "267752",
      "postDate": "01/12/2018 07:35:30",
      "content": "<p>I think so, we can try it</p>",
      "rawMarkdown": "I think so, we can try it",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 266593,
      "author_name": "rtygbwwwerr",
      "author_url": "",
      "post_date": "01/09/2018 05:44:27",
      "content": "<p>I encountered this problem as well, I deal with this kind of files as non-unknown if any non-unknown word including, otherwise label it as unknown</p>",
      "votes": null,
      "replies": [
        {
          "id": 266713,
          "author_name": "ildoonet",
          "author_url": "",
          "post_date": "01/09/2018 14:26:26",
          "content": "<p>But How do your model predict it? is it five or no?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 267118,
          "author_name": "rtygbwwwerr",
          "author_url": "",
          "post_date": "01/10/2018 16:06:06",
          "content": "<p>It's according to the training data, because the numbers of 'five' and 'no' are nearly the same, the model would label this kinda files as 'five' or 'no' fairly.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 267747,
      "author_name": "sort777",
      "author_url": "",
      "post_date": "01/12/2018 07:15:03",
      "content": "<p>I just thought that how this wav was recorded. The former word mistakenly was recorded into the next word, so the labeling is for the latter word. Does it make sense?</p>",
      "votes": null,
      "replies": [
        {
          "id": 267752,
          "author_name": "rtygbwwwerr",
          "author_url": "",
          "post_date": "01/12/2018 07:35:30",
          "content": "<p>I think so, we can try it</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "265914": "I kept noticing some audio in the train and test set have two words spoken,like \n\n&gt; train/audio/no/763188c4_nohash_4.wav\n\n\n\ncontains \"five...no\" but was labeled \"no\" in the train set.\nSo should we treat it as the latter word like the sample above in the training set, or should we treat it as an \"unknown\"?",
    "266593": "I encountered this problem as well, I deal with this kind of files as non-unknown if any non-unknown word including, otherwise label it as unknown",
    "266713": "But How do your model predict it? is it five or no?",
    "267118": "It's according to the training data, because the numbers of 'five' and 'no' are nearly the same, the model would label this kinda files as 'five' or 'no' fairly.",
    "267747": "I just thought that how this wav was recorded. The former word mistakenly was recorded into the next word, so the labeling is for the latter word. Does it make sense?",
    "267752": "I think so, we can try it"
  },
  "source": "meta"
}