{
  "id": 46962,
  "title": "Post prediction logic",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46962",
  "author_name": "",
  "post_date": "2018-01-05T21:25:28.010571Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Besides using some neural network to make the predictions, I also implemented some logic that after the prediction looks at the probabilities for the known words and if they are low (let's say below a threshold of  0.5), decide if it is not better to label them as unknown or even silence instead.</p>\n\n<p>The idea behind this is that the new unknown words might confuse some of the learned features for the known words and this would show up as a lower probability in the predictions. </p>\n\n<p>However whenever I tried this strategy, the end result (PB score) was worse. Still believe this is a viable path, so I'm wondering if there is anyone that has had some success with this kind of extra step at the end to deal with the difference in distribution between train and PB?</p>",
  "messages": [
    {
      "id": "265563",
      "postDate": "01/05/2018 21:25:28",
      "content": "<p>Besides using some neural network to make the predictions, I also implemented some logic that after the prediction looks at the probabilities for the known words and if they are low (let's say below a threshold of  0.5), decide if it is not better to label them as unknown or even silence instead.</p>\n\n<p>The idea behind this is that the new unknown words might confuse some of the learned features for the known words and this would show up as a lower probability in the predictions. </p>\n\n<p>However whenever I tried this strategy, the end result (PB score) was worse. Still believe this is a viable path, so I'm wondering if there is anyone that has had some success with this kind of extra step at the end to deal with the difference in distribution between train and PB?</p>",
      "rawMarkdown": "Besides using some neural network to make the predictions, I also implemented some logic that after the prediction looks at the probabilities for the known words and if they are low (let's say below a threshold of  0.5), decide if it is not better to label them as unknown or even silence instead.\n\nThe idea behind this is that the new unknown words might confuse some of the learned features for the known words and this would show up as a lower probability in the predictions. \n\nHowever whenever I tried this strategy, the end result (PB score) was worse. Still believe this is a viable path, so I'm wondering if there is anyone that has had some success with this kind of extra step at the end to deal with the difference in distribution between train and PB?",
      "votes": null
    },
    {
      "id": "265573",
      "postDate": "01/05/2018 22:28:25",
      "content": "<p>I think thresholding at 0.5 is too much for these case. In hard samples like \"on\" &amp; \"off\" probabilities could be something like 0.45 \"on\" + 0.35 \"off\" +0.2 for rest.</p>",
      "rawMarkdown": "I think thresholding at 0.5 is too much for these case. In hard samples like \"on\" &amp; \"off\" probabilities could be something like 0.45 \"on\" + 0.35 \"off\" +0.2 for rest.",
      "votes": null
    },
    {
      "id": "265579",
      "postDate": "01/05/2018 23:02:54",
      "content": "<p>Actually, I did experiment with different values, but didn't make a difference. However this was still one value for all the words. So indeed I should try to establish a threshold per word (perhaps I can use the validation set results for this).</p>",
      "rawMarkdown": "Actually, I did experiment with different values, but didn't make a difference. However this was still one value for all the words. So indeed I should try to establish a threshold per word (perhaps I can use the validation set results for this).",
      "votes": null
    },
    {
      "id": "265580",
      "postDate": "01/05/2018 23:09:11",
      "content": "<p>I tried this approach, but no luck. I even tried differently thresholds, the precision is no better than without them.</p>",
      "rawMarkdown": "I tried this approach, but no luck. I even tried differently thresholds, the precision is no better than without them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 265573,
      "author_name": "raalsky",
      "author_url": "",
      "post_date": "01/05/2018 22:28:25",
      "content": "<p>I think thresholding at 0.5 is too much for these case. In hard samples like \"on\" &amp; \"off\" probabilities could be something like 0.45 \"on\" + 0.35 \"off\" +0.2 for rest.</p>",
      "votes": null,
      "replies": [
        {
          "id": 265579,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "01/05/2018 23:02:54",
          "content": "<p>Actually, I did experiment with different values, but didn't make a difference. However this was still one value for all the words. So indeed I should try to establish a threshold per word (perhaps I can use the validation set results for this).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 265580,
      "author_name": "abnerchou",
      "author_url": "",
      "post_date": "01/05/2018 23:09:11",
      "content": "<p>I tried this approach, but no luck. I even tried differently thresholds, the precision is no better than without them.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "265563": "Besides using some neural network to make the predictions, I also implemented some logic that after the prediction looks at the probabilities for the known words and if they are low (let's say below a threshold of  0.5), decide if it is not better to label them as unknown or even silence instead.\n\nThe idea behind this is that the new unknown words might confuse some of the learned features for the known words and this would show up as a lower probability in the predictions. \n\nHowever whenever I tried this strategy, the end result (PB score) was worse. Still believe this is a viable path, so I'm wondering if there is anyone that has had some success with this kind of extra step at the end to deal with the difference in distribution between train and PB?",
    "265573": "I think thresholding at 0.5 is too much for these case. In hard samples like \"on\" &amp; \"off\" probabilities could be something like 0.45 \"on\" + 0.35 \"off\" +0.2 for rest.",
    "265579": "Actually, I did experiment with different values, but didn't make a difference. However this was still one value for all the words. So indeed I should try to establish a threshold per word (perhaps I can use the validation set results for this).",
    "265580": "I tried this approach, but no luck. I even tried differently thresholds, the precision is no better than without them."
  },
  "source": "meta"
}