{
  "id": 46298,
  "title": "LB 0.80 while I predict 47% labels as \"unknown\"",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/46298",
  "author_name": "",
  "post_date": "2017-12-24T07:03:56.142319900Z",
  "votes": 14,
  "comment_count": 6,
  "views": 0,
  "content": "<p>I had assumed that public LB test portion is a random sample of overall test data. However, that does not seem to be the case. Here is the reason.</p>\n\n<p>The attached submission attains an LB score of 0.80. This score is, as we know, on the public LB. \nOut of 158539 predictions, this submission predicts \"unknown\" for 73880 samples, i.e. for 47% of the samples. However, we know that a submission which predicts \"unknown\" for all samples scores 11%.</p>\n\n<p>Thus, if public and private LB were having identical distribution of labels, at most 11% of my \"unknown\" prediction could have been correct, and thus, I would lose 47% - 11% = 36% points right there, and the LB score could be no more than 64%. However, it is 80%.</p>\n\n<p>This means one of the two things:\n(1) Private part of test cases have much higher proportion of unknowns\n(2) Private and public parts of test cases of identical distribution: my model just got lucky to have 80% in public part, and will be much poorer in private part. </p>\n\n<p>I think (1) above is more likely.</p>\n\n<p>Do others have similar observations?</p>\n\n<p>Thanks,\nVivek</p>",
  "messages": [
    {
      "id": "261845",
      "postDate": "12/24/2017 07:03:56",
      "content": "<p>I had assumed that public LB test portion is a random sample of overall test data. However, that does not seem to be the case. Here is the reason.</p>\n\n<p>The attached submission attains an LB score of 0.80. This score is, as we know, on the public LB. \nOut of 158539 predictions, this submission predicts \"unknown\" for 73880 samples, i.e. for 47% of the samples. However, we know that a submission which predicts \"unknown\" for all samples scores 11%.</p>\n\n<p>Thus, if public and private LB were having identical distribution of labels, at most 11% of my \"unknown\" prediction could have been correct, and thus, I would lose 47% - 11% = 36% points right there, and the LB score could be no more than 64%. However, it is 80%.</p>\n\n<p>This means one of the two things:\n(1) Private part of test cases have much higher proportion of unknowns\n(2) Private and public parts of test cases of identical distribution: my model just got lucky to have 80% in public part, and will be much poorer in private part. </p>\n\n<p>I think (1) above is more likely.</p>\n\n<p>Do others have similar observations?</p>\n\n<p>Thanks,\nVivek</p>",
      "rawMarkdown": "I had assumed that public LB test portion is a random sample of overall test data. However, that does not seem to be the case. Here is the reason.\n\nThe attached submission attains an LB score of 0.80. This score is, as we know, on the public LB. \nOut of 158539 predictions, this submission predicts \"unknown\" for 73880 samples, i.e. for 47% of the samples. However, we know that a submission which predicts \"unknown\" for all samples scores 11%.\n\nThus, if public and private LB were having identical distribution of labels, at most 11% of my \"unknown\" prediction could have been correct, and thus, I would lose 47% - 11% = 36% points right there, and the LB score could be no more than 64%. However, it is 80%.\n\nThis means one of the two things:\n(1) Private part of test cases have much higher proportion of unknowns\n(2) Private and public parts of test cases of identical distribution: my model just got lucky to have 80% in public part, and will be much poorer in private part. \n\nI think (1) above is more likely.\n\nDo others have similar observations?\n\nThanks,\nVivek",
      "votes": null
    },
    {
      "id": "261871",
      "postDate": "12/24/2017 10:19:27",
      "content": "<p>In the data description says the following:</p>\n\n<blockquote>\n  <p>Not all of the files are evaluated for the leaderboard score.</p>\n</blockquote>\n\n<p>So my hypothesis is that most of that unknown predictions are not being evaluated, they are given just to make the test set bigger and deter hand-labeling.</p>",
      "rawMarkdown": "In the data description says the following:\n\n&gt; Not all of the files are evaluated for the leaderboard score.\n\nSo my hypothesis is that most of that unknown predictions are not being evaluated, they are given just to make the test set bigger and deter hand-labeling.",
      "votes": null
    },
    {
      "id": "262014",
      "postDate": "12/24/2017 23:20:10",
      "content": "<p>I think that @ironbar is right. The distribution of train is ~36.(6)% for 10 commands and ~63,(3)% for unknown (plus 10% of silence). So if we support that test distribution is the same, <strong>but for evaluation they took only part of unknown to get 9%</strong>, so we may conclude that the perfect submition can (or should) contain not only 47%, but 63% of unknown words (or 54%, taking into account the silence).</p>",
      "rawMarkdown": "I think that @ironbar is right. The distribution of train is ~36.(6)% for 10 commands and ~63,(3)% for unknown (plus 10% of silence). So if we support that test distribution is the same, **but for evaluation they took only part of unknown to get 9%**, so we may conclude that the perfect submition can (or should) contain not only 47%, but 63% of unknown words (or 54%, taking into account the silence).",
      "votes": null
    },
    {
      "id": "262395",
      "postDate": "12/26/2017 12:27:53",
      "content": "<p>Thanks for the input. Yes, an alternative model of mine predicted more labels as unknown, and got a bit higher LB score.</p>",
      "rawMarkdown": "Thanks for the input. Yes, an alternative model of mine predicted more labels as unknown, and got a bit higher LB score.",
      "votes": null
    },
    {
      "id": "265081",
      "postDate": "01/04/2018 15:11:30",
      "content": "<p>What technique did you use to detect silence? It seems pretty accurate</p>",
      "rawMarkdown": "What technique did you use to detect silence? It seems pretty accurate",
      "votes": null
    },
    {
      "id": "265297",
      "postDate": "01/05/2018 04:36:56",
      "content": "<p>I have not used any different technique for silence. I treat it as yet another category.</p>",
      "rawMarkdown": "I have not used any different technique for silence. I treat it as yet another category.",
      "votes": null
    },
    {
      "id": "265317",
      "postDate": "01/05/2018 06:25:10",
      "content": "<p>thanks, it is important for me</p>",
      "rawMarkdown": "thanks, it is important for me",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 261871,
      "author_name": "ironbar",
      "author_url": "",
      "post_date": "12/24/2017 10:19:27",
      "content": "<p>In the data description says the following:</p>\n\n<blockquote>\n  <p>Not all of the files are evaluated for the leaderboard score.</p>\n</blockquote>\n\n<p>So my hypothesis is that most of that unknown predictions are not being evaluated, they are given just to make the test set bigger and deter hand-labeling.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 262014,
      "author_name": "mikhas",
      "author_url": "",
      "post_date": "12/24/2017 23:20:10",
      "content": "<p>I think that @ironbar is right. The distribution of train is ~36.(6)% for 10 commands and ~63,(3)% for unknown (plus 10% of silence). So if we support that test distribution is the same, <strong>but for evaluation they took only part of unknown to get 9%</strong>, so we may conclude that the perfect submition can (or should) contain not only 47%, but 63% of unknown words (or 54%, taking into account the silence).</p>",
      "votes": null,
      "replies": [
        {
          "id": 262395,
          "author_name": "thevivekpandey",
          "author_url": "",
          "post_date": "12/26/2017 12:27:53",
          "content": "<p>Thanks for the input. Yes, an alternative model of mine predicted more labels as unknown, and got a bit higher LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265081,
          "author_name": "mikhas",
          "author_url": "",
          "post_date": "01/04/2018 15:11:30",
          "content": "<p>What technique did you use to detect silence? It seems pretty accurate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265297,
          "author_name": "thevivekpandey",
          "author_url": "",
          "post_date": "01/05/2018 04:36:56",
          "content": "<p>I have not used any different technique for silence. I treat it as yet another category.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 265317,
          "author_name": "mikhas",
          "author_url": "",
          "post_date": "01/05/2018 06:25:10",
          "content": "<p>thanks, it is important for me</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "261845": "I had assumed that public LB test portion is a random sample of overall test data. However, that does not seem to be the case. Here is the reason.\n\nThe attached submission attains an LB score of 0.80. This score is, as we know, on the public LB. \nOut of 158539 predictions, this submission predicts \"unknown\" for 73880 samples, i.e. for 47% of the samples. However, we know that a submission which predicts \"unknown\" for all samples scores 11%.\n\nThus, if public and private LB were having identical distribution of labels, at most 11% of my \"unknown\" prediction could have been correct, and thus, I would lose 47% - 11% = 36% points right there, and the LB score could be no more than 64%. However, it is 80%.\n\nThis means one of the two things:\n(1) Private part of test cases have much higher proportion of unknowns\n(2) Private and public parts of test cases of identical distribution: my model just got lucky to have 80% in public part, and will be much poorer in private part. \n\nI think (1) above is more likely.\n\nDo others have similar observations?\n\nThanks,\nVivek",
    "261871": "In the data description says the following:\n\n&gt; Not all of the files are evaluated for the leaderboard score.\n\nSo my hypothesis is that most of that unknown predictions are not being evaluated, they are given just to make the test set bigger and deter hand-labeling.",
    "262014": "I think that @ironbar is right. The distribution of train is ~36.(6)% for 10 commands and ~63,(3)% for unknown (plus 10% of silence). So if we support that test distribution is the same, **but for evaluation they took only part of unknown to get 9%**, so we may conclude that the perfect submition can (or should) contain not only 47%, but 63% of unknown words (or 54%, taking into account the silence).",
    "262395": "Thanks for the input. Yes, an alternative model of mine predicted more labels as unknown, and got a bit higher LB score.",
    "265081": "What technique did you use to detect silence? It seems pretty accurate",
    "265297": "I have not used any different technique for silence. I treat it as yet another category.",
    "265317": "thanks, it is important for me"
  },
  "source": "meta"
}