{
  "id": 47312,
  "title": "how is flawed pronunciation samples scored?",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/47312",
  "author_name": "",
  "post_date": "2018-01-11T22:05:29.443227500Z",
  "votes": 1,
  "comment_count": 1,
  "views": 0,
  "content": "<p>take 'left' as an example, i noticed that my phoneme model identified bunch of words, the intention is pronounced as left,however, the t is not pronounced, so the pronunciation is sound like lef. i listened to some of the automatic detected cases and found majority of them should be correct identification. I tried to override them as unknown but the score in LB dropped. </p>\n\n<p>robustness on pronunciation flaw is good here?  anyone have similar observation?</p>",
  "messages": [
    {
      "id": "267616",
      "postDate": "01/11/2018 22:05:29",
      "content": "<p>take 'left' as an example, i noticed that my phoneme model identified bunch of words, the intention is pronounced as left,however, the t is not pronounced, so the pronunciation is sound like lef. i listened to some of the automatic detected cases and found majority of them should be correct identification. I tried to override them as unknown but the score in LB dropped. </p>\n\n<p>robustness on pronunciation flaw is good here?  anyone have similar observation?</p>",
      "rawMarkdown": "take 'left' as an example, i noticed that my phoneme model identified bunch of words, the intention is pronounced as left,however, the t is not pronounced, so the pronunciation is sound like lef. i listened to some of the automatic detected cases and found majority of them should be correct identification. I tried to override them as unknown but the score in LB dropped. \n\nrobustness on pronunciation flaw is good here?  anyone have similar observation?",
      "votes": null
    },
    {
      "id": "270125",
      "postDate": "01/17/2018 19:23:16",
      "content": "<p>Just try to revive this question. \nI tried a very traditional old approach that is using HTK to cold start a monophone HMM/GMM model, and first use it to force align predictions from deep learning word modeling, i found it do identify flawed words that was verified by my ear, but when i tried to relabel them as unknown or silence, the LB statistics drops. I was originally planning to train further HMM/NN model for phoneme verification but i just give up on this approach.  </p>\n\n<p>how is the ground truth labels generated? what is the rubric for correct/incorrect/unknown? is there an estimation of labeling errors using multiple labelers and calculate the agreement between them?</p>",
      "rawMarkdown": "Just try to revive this question. \nI tried a very traditional old approach that is using HTK to cold start a monophone HMM/GMM model, and first use it to force align predictions from deep learning word modeling, i found it do identify flawed words that was verified by my ear, but when i tried to relabel them as unknown or silence, the LB statistics drops. I was originally planning to train further HMM/NN model for phoneme verification but i just give up on this approach.  \n\nhow is the ground truth labels generated? what is the rubric for correct/incorrect/unknown? is there an estimation of labeling errors using multiple labelers and calculate the agreement between them?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 270125,
      "author_name": "tank8888",
      "author_url": "",
      "post_date": "01/17/2018 19:23:16",
      "content": "<p>Just try to revive this question. \nI tried a very traditional old approach that is using HTK to cold start a monophone HMM/GMM model, and first use it to force align predictions from deep learning word modeling, i found it do identify flawed words that was verified by my ear, but when i tried to relabel them as unknown or silence, the LB statistics drops. I was originally planning to train further HMM/NN model for phoneme verification but i just give up on this approach.  </p>\n\n<p>how is the ground truth labels generated? what is the rubric for correct/incorrect/unknown? is there an estimation of labeling errors using multiple labelers and calculate the agreement between them?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "267616": "take 'left' as an example, i noticed that my phoneme model identified bunch of words, the intention is pronounced as left,however, the t is not pronounced, so the pronunciation is sound like lef. i listened to some of the automatic detected cases and found majority of them should be correct identification. I tried to override them as unknown but the score in LB dropped. \n\nrobustness on pronunciation flaw is good here?  anyone have similar observation?",
    "270125": "Just try to revive this question. \nI tried a very traditional old approach that is using HTK to cold start a monophone HMM/GMM model, and first use it to force align predictions from deep learning word modeling, i found it do identify flawed words that was verified by my ear, but when i tried to relabel them as unknown or silence, the LB statistics drops. I was originally planning to train further HMM/NN model for phoneme verification but i just give up on this approach.  \n\nhow is the ground truth labels generated? what is the rubric for correct/incorrect/unknown? is there an estimation of labeling errors using multiple labelers and calculate the agreement between them?"
  },
  "source": "meta"
}