{
  "id": 43773,
  "title": "uncertainty about unknown",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/43773",
  "author_name": "",
  "post_date": "2017-11-19T09:55:27.721709Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Does anyone know if the distribution of the \"unknown\" target value is the same for the evaluation set as it is for the training data. </p>\n\n<p>So in other words are the \"unknown\" words used for testing the same unknown words that are also present in the training data (like sheila, house, marvin, ...).</p>\n\n<p>My first approach was to label all the words outside the 10 keywords as unknown and train on that. But if the test set has other words included that also need to be flagged as unknown, I might get better results looking at the probability for the 10 words.  </p>",
  "messages": [
    {
      "id": "245693",
      "postDate": "11/19/2017 09:55:27",
      "content": "<p>Does anyone know if the distribution of the \"unknown\" target value is the same for the evaluation set as it is for the training data. </p>\n\n<p>So in other words are the \"unknown\" words used for testing the same unknown words that are also present in the training data (like sheila, house, marvin, ...).</p>\n\n<p>My first approach was to label all the words outside the 10 keywords as unknown and train on that. But if the test set has other words included that also need to be flagged as unknown, I might get better results looking at the probability for the 10 words.  </p>",
      "rawMarkdown": "Does anyone know if the distribution of the \"unknown\" target value is the same for the evaluation set as it is for the training data. \n\nSo in other words are the \"unknown\" words used for testing the same unknown words that are also present in the training data (like sheila, house, marvin, ...).\n\nMy first approach was to label all the words outside the 10 keywords as unknown and train on that. But if the test set has other words included that also need to be flagged as unknown, I might get better results looking at the probability for the 10 words.",
      "votes": null
    },
    {
      "id": "245857",
      "postDate": "11/19/2017 21:58:26",
      "content": "<p>I'd like to know that too</p>",
      "rawMarkdown": "I'd like to know that too",
      "votes": null
    },
    {
      "id": "245898",
      "postDate": "11/20/2017 02:47:29",
      "content": "<p>The words \"backward\" and \"learn\", which are in the test set but not in the training set, do seem to confuse the model.</p>",
      "rawMarkdown": "The words \"backward\" and \"learn\", which are in the test set but not in the training set, do seem to confuse the model.",
      "votes": null
    },
    {
      "id": "245932",
      "postDate": "11/20/2017 05:42:32",
      "content": "<p>Your correct that the test set has more unknown words than the training one. We've been expanding the range of words we've been collecting since the original release, and the product goal with the data is that eventually <em>any</em> non-command word (even ones it hasn't seen before) should be classified as unknown. The idea is that the unknown category should drive the model to declare when it's unsure.</p>\n\n<p>I know this is a bit confusing though, does that explanation make sense?</p>",
      "rawMarkdown": "Your correct that the test set has more unknown words than the training one. We've been expanding the range of words we've been collecting since the original release, and the product goal with the data is that eventually *any* non-command word (even ones it hasn't seen before) should be classified as unknown. The idea is that the unknown category should drive the model to declare when it's unsure.\n\nI know this is a bit confusing though, does that explanation make sense?",
      "votes": null
    },
    {
      "id": "245957",
      "postDate": "11/20/2017 07:07:23",
      "content": "<p>Thanks for the explanation. I guess this partially contributes to the difference in accuracy on the cross validation and test set, the other part of course being me and my model ;)</p>\n\n<p>So I guess now it is up to us to come up with a good strategy, like taking more care of the probability/confidence levels when making the final prediction. </p>",
      "rawMarkdown": "Thanks for the explanation. I guess this partially contributes to the difference in accuracy on the cross validation and test set, the other part of course being me and my model ;)\n\nSo I guess now it is up to us to come up with a good strategy, like taking more care of the probability/confidence levels when making the final prediction.",
      "votes": null
    },
    {
      "id": "246018",
      "postDate": "11/20/2017 09:53:59",
      "content": "<p>Wouldn't it have been better to have more diversity in the train set rather than repeating the same few words ? It's already a small train set and now it has to be smaller because we have to remove loads of words . Why didn't you just take a random sample of \"unknown\"? It's like \"here's loads of specific words we're not interested in but you might as well as forget about them because we put totally different ones in the test set\" The train set is tiny now</p>",
      "rawMarkdown": "Wouldn't it have been better to have more diversity in the train set rather than repeating the same few words ? It's already a small train set and now it has to be smaller because we have to remove loads of words . Why didn't you just take a random sample of \"unknown\"? It's like \"here's loads of specific words we're not interested in but you might as well as forget about them because we put totally different ones in the test set\" The train set is tiny now",
      "votes": null
    },
    {
      "id": "246057",
      "postDate": "11/20/2017 13:16:53",
      "content": "<p>There are three kinds of words in the test set:</p>\n\n<ol>\n<li>Core words (yes, no, up, down)</li>\n<li>non-core words from train (marvin, sheila)</li>\n<li>Words not in train (learn, backward, visual)</li>\n</ol>\n\n<p>It's tempting to augment the training set with some hand-labeled examples of the new words, but I think that's against the rules. It seems like the model needs to get smart enough that it won't have many examples where it gives a high probability that a word is \"no\", but in fact it's \"learn\".</p>",
      "rawMarkdown": "There are three kinds of words in the test set:\n\n1. Core words (yes, no, up, down)\n2. non-core words from train (marvin, sheila)\n3. Words not in train (learn, backward, visual)\n\nIt's tempting to augment the training set with some hand-labeled examples of the new words, but I think that's against the rules. It seems like the model needs to get smart enough that it won't have many examples where it gives a high probability that a word is \"no\", but in fact it's \"learn\".",
      "votes": null
    },
    {
      "id": "246095",
      "postDate": "11/20/2017 15:04:02",
      "content": "<p>I agree that hand-labeling is against the rules.</p>\n\n<p>However I do think that pre-training your model using for example convolutional-autoencoders with the test data is acceptable.  That way at least the lower level conv-layers in your model better understand the relevant features in the test data, also those for those words not in train. </p>\n\n<p>And as a result your model might be able to better differentiate between known and unknown words.</p>",
      "rawMarkdown": "I agree that hand-labeling is against the rules.\n\nHowever I do think that pre-training your model using for example convolutional-autoencoders with the test data is acceptable.  That way at least the lower level conv-layers in your model better understand the relevant features in the test data, also those for those words not in train. \n\nAnd as a result your model might be able to better differentiate between known and unknown words.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 245857,
      "author_name": "domcastro",
      "author_url": "",
      "post_date": "11/19/2017 21:58:26",
      "content": "<p>I'd like to know that too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 245898,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/20/2017 02:47:29",
      "content": "<p>The words \"backward\" and \"learn\", which are in the test set but not in the training set, do seem to confuse the model.</p>",
      "votes": null,
      "replies": [
        {
          "id": 245932,
          "author_name": "petewarden",
          "author_url": "",
          "post_date": "11/20/2017 05:42:32",
          "content": "<p>Your correct that the test set has more unknown words than the training one. We've been expanding the range of words we've been collecting since the original release, and the product goal with the data is that eventually <em>any</em> non-command word (even ones it hasn't seen before) should be classified as unknown. The idea is that the unknown category should drive the model to declare when it's unsure.</p>\n\n<p>I know this is a bit confusing though, does that explanation make sense?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 245957,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "11/20/2017 07:07:23",
          "content": "<p>Thanks for the explanation. I guess this partially contributes to the difference in accuracy on the cross validation and test set, the other part of course being me and my model ;)</p>\n\n<p>So I guess now it is up to us to come up with a good strategy, like taking more care of the probability/confidence levels when making the final prediction. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 246018,
          "author_name": "domcastro",
          "author_url": "",
          "post_date": "11/20/2017 09:53:59",
          "content": "<p>Wouldn't it have been better to have more diversity in the train set rather than repeating the same few words ? It's already a small train set and now it has to be smaller because we have to remove loads of words . Why didn't you just take a random sample of \"unknown\"? It's like \"here's loads of specific words we're not interested in but you might as well as forget about them because we put totally different ones in the test set\" The train set is tiny now</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 246057,
          "author_name": "jfkingiii",
          "author_url": "",
          "post_date": "11/20/2017 13:16:53",
          "content": "<p>There are three kinds of words in the test set:</p>\n\n<ol>\n<li>Core words (yes, no, up, down)</li>\n<li>non-core words from train (marvin, sheila)</li>\n<li>Words not in train (learn, backward, visual)</li>\n</ol>\n\n<p>It's tempting to augment the training set with some hand-labeled examples of the new words, but I think that's against the rules. It seems like the model needs to get smart enough that it won't have many examples where it gives a high probability that a word is \"no\", but in fact it's \"learn\".</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 246095,
          "author_name": "peterdekkers101",
          "author_url": "",
          "post_date": "11/20/2017 15:04:02",
          "content": "<p>I agree that hand-labeling is against the rules.</p>\n\n<p>However I do think that pre-training your model using for example convolutional-autoencoders with the test data is acceptable.  That way at least the lower level conv-layers in your model better understand the relevant features in the test data, also those for those words not in train. </p>\n\n<p>And as a result your model might be able to better differentiate between known and unknown words.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "245693": "Does anyone know if the distribution of the \"unknown\" target value is the same for the evaluation set as it is for the training data. \n\nSo in other words are the \"unknown\" words used for testing the same unknown words that are also present in the training data (like sheila, house, marvin, ...).\n\nMy first approach was to label all the words outside the 10 keywords as unknown and train on that. But if the test set has other words included that also need to be flagged as unknown, I might get better results looking at the probability for the 10 words.",
    "245857": "I'd like to know that too",
    "245898": "The words \"backward\" and \"learn\", which are in the test set but not in the training set, do seem to confuse the model.",
    "245932": "Your correct that the test set has more unknown words than the training one. We've been expanding the range of words we've been collecting since the original release, and the product goal with the data is that eventually *any* non-command word (even ones it hasn't seen before) should be classified as unknown. The idea is that the unknown category should drive the model to declare when it's unsure.\n\nI know this is a bit confusing though, does that explanation make sense?",
    "245957": "Thanks for the explanation. I guess this partially contributes to the difference in accuracy on the cross validation and test set, the other part of course being me and my model ;)\n\nSo I guess now it is up to us to come up with a good strategy, like taking more care of the probability/confidence levels when making the final prediction.",
    "246018": "Wouldn't it have been better to have more diversity in the train set rather than repeating the same few words ? It's already a small train set and now it has to be smaller because we have to remove loads of words . Why didn't you just take a random sample of \"unknown\"? It's like \"here's loads of specific words we're not interested in but you might as well as forget about them because we put totally different ones in the test set\" The train set is tiny now",
    "246057": "There are three kinds of words in the test set:\n\n1. Core words (yes, no, up, down)\n2. non-core words from train (marvin, sheila)\n3. Words not in train (learn, backward, visual)\n\nIt's tempting to augment the training set with some hand-labeled examples of the new words, but I think that's against the rules. It seems like the model needs to get smart enough that it won't have many examples where it gives a high probability that a word is \"no\", but in fact it's \"learn\".",
    "246095": "I agree that hand-labeling is against the rules.\n\nHowever I do think that pre-training your model using for example convolutional-autoencoders with the test data is acceptable.  That way at least the lower level conv-layers in your model better understand the relevant features in the test data, also those for those words not in train. \n\nAnd as a result your model might be able to better differentiate between known and unknown words."
  },
  "source": "meta"
}