{
  "id": 44234,
  "title": "Public/Private Leaderboard Split",
  "url": "/competitions/tensorflow-speech-recognition-challenge/discussion/44234",
  "author_name": "",
  "post_date": "2017-11-25T18:43:25.794985300Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Is the leaderboard split completely random/unbiased? </p>\n\n<p>Reason I ask. I've made a few submissions hacking around with a baseline starting point. Making some rather significant changes to training parameters, trying two different models, and altering training distribution of the unknown class especially, the LB scores have all been in .81-.82 range. </p>\n\n<p>Analyzing the submissions from the experiments there are significant differences in the prediction class counts, noteworthy differences in my local validation accuracy and loss, but very little change in the public leaderboard accuracy...</p>",
  "messages": [
    {
      "id": "248341",
      "postDate": "11/25/2017 18:43:25",
      "content": "<p>Is the leaderboard split completely random/unbiased? </p>\n\n<p>Reason I ask. I've made a few submissions hacking around with a baseline starting point. Making some rather significant changes to training parameters, trying two different models, and altering training distribution of the unknown class especially, the LB scores have all been in .81-.82 range. </p>\n\n<p>Analyzing the submissions from the experiments there are significant differences in the prediction class counts, noteworthy differences in my local validation accuracy and loss, but very little change in the public leaderboard accuracy...</p>",
      "rawMarkdown": "Is the leaderboard split completely random/unbiased? \n\nReason I ask. I've made a few submissions hacking around with a baseline starting point. Making some rather significant changes to training parameters, trying two different models, and altering training distribution of the unknown class especially, the LB scores have all been in .81-.82 range. \n\nAnalyzing the submissions from the experiments there are significant differences in the prediction class counts, noteworthy differences in my local validation accuracy and loss, but very little change in the public leaderboard accuracy...",
      "votes": null
    },
    {
      "id": "248359",
      "postDate": "11/25/2017 20:29:44",
      "content": "<p>The train/test split is definitely not random, since test contains words not in train. Is the public/private split of test random, or will we encounter new \"unknown\" words on the private LB?</p>",
      "rawMarkdown": "The train/test split is definitely not random, since test contains words not in train. Is the public/private split of test random, or will we encounter new \"unknown\" words on the private LB?",
      "votes": null
    },
    {
      "id": "248413",
      "postDate": "11/26/2017 01:18:02",
      "content": "<p>I was also trying to understand this issue. therefore I manually classified the first 200 test data only to see what are they made of. Here is my list:</p>\n\n<ul>\n<li>down     2.5 %</li>\n<li>left      5.5 %</li>\n<li>no       4 %  </li>\n<li>off       4 %  </li>\n<li>on        5 %  </li>\n<li>right    1 %  </li>\n<li>silence  8.5 %  </li>\n<li>unknown  45 % </li>\n<li>up       3.5 %  </li>\n<li>yes      2.5 %  </li>\n<li>unknown* 18.5 % (not available in train data)</li>\n</ul>",
      "rawMarkdown": "I was also trying to understand this issue. therefore I manually classified the first 200 test data only to see what are they made of. Here is my list:\n\n\n- down     2.5 %\n- left      5.5 %\n- no       4 %  \n- off       4 %  \n- on        5 %  \n- right    1 %  \n- silence  8.5 %  \n- unknown  45 % \n- up       3.5 %  \n- yes      2.5 %  \n- unknown* 18.5 % (not available in train data)",
      "votes": null
    },
    {
      "id": "248417",
      "postDate": "11/26/2017 01:41:11",
      "content": "<p>That's not an accurate sampling of the test distribution because they add synthetic data to prevent hand-labeling.  However, it's easy to get the class distribution of the public leaderboard by making submissions that predict a single class.  For example, the silence benchmark scores .09, revealing that the 9% of the public leaderboard is silence.  Whether this approximately holds for the private leaderboard depends on whether the public/private split is random.</p>",
      "rawMarkdown": "That's not an accurate sampling of the test distribution because they add synthetic data to prevent hand-labeling.  However, it's easy to get the class distribution of the public leaderboard by making submissions that predict a single class.  For example, the silence benchmark scores .09, revealing that the 9% of the public leaderboard is silence.  Whether this approximately holds for the private leaderboard depends on whether the public/private split is random.",
      "votes": null
    },
    {
      "id": "248418",
      "postDate": "11/26/2017 01:42:01",
      "content": "<p>I do think some, potentially a lot, of what I'm seeing is the fact that I haven't properly dealt with the 'unknown unknown' yet. That they are likely being distributed between the various classes differently in each training session, with a consistent % test data being correctly (easily?) classified regardless of my hyper params. But I was still expecting more variability in my leaderboard scores. My unknown class has ranged from 39% of results to 52%. Other classes have swung by up to 25%. Leaderboard has been in .81 to .82 in all cases.</p>\n\n<p>I guess it's time to do some feature analysis, clustering, and visualization :)  </p>",
      "rawMarkdown": "I do think some, potentially a lot, of what I'm seeing is the fact that I haven't properly dealt with the 'unknown unknown' yet. That they are likely being distributed between the various classes differently in each training session, with a consistent % test data being correctly (easily?) classified regardless of my hyper params. But I was still expecting more variability in my leaderboard scores. My unknown class has ranged from 39% of results to 52%. Other classes have swung by up to 25%. Leaderboard has been in .81 to .82 in all cases.\n\nI guess it's time to do some feature analysis, clustering, and visualization :)",
      "votes": null
    },
    {
      "id": "248419",
      "postDate": "11/26/2017 01:54:07",
      "content": "<p>@sjv I know. Like I said, I was only trying to understand what I might be missing. I didn't know about unknown-unknown issue entirely before doing this. </p>",
      "rawMarkdown": "sjv I know. Like I said, I was only trying to understand what I might be missing. I didn't know about unknown-unknown issue entirely before doing this.",
      "votes": null
    },
    {
      "id": "248574",
      "postDate": "11/26/2017 14:09:39",
      "content": "<p>If you submit just unknown then it scores .09 too, just like the silence benchmark. There is also a note that says \"Not all of the [test set] files are evaluated for the leaderboard score.\" So it seems that a lot of these unknown examples are getting ignored when computing the LB score.</p>",
      "rawMarkdown": "If you submit just unknown then it scores .09 too, just like the silence benchmark. There is also a note that says \"Not all of the [test set] files are evaluated for the leaderboard score.\" So it seems that a lot of these unknown examples are getting ignored when computing the LB score.",
      "votes": null
    },
    {
      "id": "255051",
      "postDate": "12/08/2017 06:09:38",
      "content": "<p>I'm also curious about this split, and I'm siding with @Fatihkurt.  It seems strange that they would give us a train distribution that doesn't match the test distribution.</p>",
      "rawMarkdown": "I'm also curious about this split, and I'm siding with @Fatihkurt.  It seems strange that they would give us a train distribution that doesn't match the test distribution.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 248359,
      "author_name": "jfkingiii",
      "author_url": "",
      "post_date": "11/25/2017 20:29:44",
      "content": "<p>The train/test split is definitely not random, since test contains words not in train. Is the public/private split of test random, or will we encounter new \"unknown\" words on the private LB?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 248413,
      "author_name": "ftkurt",
      "author_url": "",
      "post_date": "11/26/2017 01:18:02",
      "content": "<p>I was also trying to understand this issue. therefore I manually classified the first 200 test data only to see what are they made of. Here is my list:</p>\n\n<ul>\n<li>down     2.5 %</li>\n<li>left      5.5 %</li>\n<li>no       4 %  </li>\n<li>off       4 %  </li>\n<li>on        5 %  </li>\n<li>right    1 %  </li>\n<li>silence  8.5 %  </li>\n<li>unknown  45 % </li>\n<li>up       3.5 %  </li>\n<li>yes      2.5 %  </li>\n<li>unknown* 18.5 % (not available in train data)</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 248417,
          "author_name": "seanjv",
          "author_url": "",
          "post_date": "11/26/2017 01:41:11",
          "content": "<p>That's not an accurate sampling of the test distribution because they add synthetic data to prevent hand-labeling.  However, it's easy to get the class distribution of the public leaderboard by making submissions that predict a single class.  For example, the silence benchmark scores .09, revealing that the 9% of the public leaderboard is silence.  Whether this approximately holds for the private leaderboard depends on whether the public/private split is random.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 248418,
          "author_name": "rwightman",
          "author_url": "",
          "post_date": "11/26/2017 01:42:01",
          "content": "<p>I do think some, potentially a lot, of what I'm seeing is the fact that I haven't properly dealt with the 'unknown unknown' yet. That they are likely being distributed between the various classes differently in each training session, with a consistent % test data being correctly (easily?) classified regardless of my hyper params. But I was still expecting more variability in my leaderboard scores. My unknown class has ranged from 39% of results to 52%. Other classes have swung by up to 25%. Leaderboard has been in .81 to .82 in all cases.</p>\n\n<p>I guess it's time to do some feature analysis, clustering, and visualization :)  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 248419,
          "author_name": "ftkurt",
          "author_url": "",
          "post_date": "11/26/2017 01:54:07",
          "content": "<p>@sjv I know. Like I said, I was only trying to understand what I might be missing. I didn't know about unknown-unknown issue entirely before doing this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 248574,
          "author_name": "humananalog",
          "author_url": "",
          "post_date": "11/26/2017 14:09:39",
          "content": "<p>If you submit just unknown then it scores .09 too, just like the silence benchmark. There is also a note that says \"Not all of the [test set] files are evaluated for the leaderboard score.\" So it seems that a lot of these unknown examples are getting ignored when computing the LB score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 255051,
      "author_name": "tmulc18",
      "author_url": "",
      "post_date": "12/08/2017 06:09:38",
      "content": "<p>I'm also curious about this split, and I'm siding with @Fatihkurt.  It seems strange that they would give us a train distribution that doesn't match the test distribution.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "248341": "Is the leaderboard split completely random/unbiased? \n\nReason I ask. I've made a few submissions hacking around with a baseline starting point. Making some rather significant changes to training parameters, trying two different models, and altering training distribution of the unknown class especially, the LB scores have all been in .81-.82 range. \n\nAnalyzing the submissions from the experiments there are significant differences in the prediction class counts, noteworthy differences in my local validation accuracy and loss, but very little change in the public leaderboard accuracy...",
    "248359": "The train/test split is definitely not random, since test contains words not in train. Is the public/private split of test random, or will we encounter new \"unknown\" words on the private LB?",
    "248413": "I was also trying to understand this issue. therefore I manually classified the first 200 test data only to see what are they made of. Here is my list:\n\n\n- down     2.5 %\n- left      5.5 %\n- no       4 %  \n- off       4 %  \n- on        5 %  \n- right    1 %  \n- silence  8.5 %  \n- unknown  45 % \n- up       3.5 %  \n- yes      2.5 %  \n- unknown* 18.5 % (not available in train data)",
    "248417": "That's not an accurate sampling of the test distribution because they add synthetic data to prevent hand-labeling.  However, it's easy to get the class distribution of the public leaderboard by making submissions that predict a single class.  For example, the silence benchmark scores .09, revealing that the 9% of the public leaderboard is silence.  Whether this approximately holds for the private leaderboard depends on whether the public/private split is random.",
    "248418": "I do think some, potentially a lot, of what I'm seeing is the fact that I haven't properly dealt with the 'unknown unknown' yet. That they are likely being distributed between the various classes differently in each training session, with a consistent % test data being correctly (easily?) classified regardless of my hyper params. But I was still expecting more variability in my leaderboard scores. My unknown class has ranged from 39% of results to 52%. Other classes have swung by up to 25%. Leaderboard has been in .81 to .82 in all cases.\n\nI guess it's time to do some feature analysis, clustering, and visualization :)",
    "248419": "sjv I know. Like I said, I was only trying to understand what I might be missing. I didn't know about unknown-unknown issue entirely before doing this.",
    "248574": "If you submit just unknown then it scores .09 too, just like the silence benchmark. There is also a note that says \"Not all of the [test set] files are evaluated for the leaderboard score.\" So it seems that a lot of these unknown examples are getting ignored when computing the LB score.",
    "255051": "I'm also curious about this split, and I'm siding with @Fatihkurt.  It seems strange that they would give us a train distribution that doesn't match the test distribution."
  },
  "source": "meta"
}