{
  "id": 116890,
  "title": "Question about short answer evaluation",
  "url": "/competitions/tensorflow2-question-answering/discussion/116890",
  "author_name": "Cory Binnersley",
  "post_date": "2019-11-12T05:23:54.525000",
  "votes": 10,
  "comment_count": 12,
  "views": 0,
  "content": "<p>In the evaluation description, these lines exist:</p>\n\n<p>&gt; There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</p>\n\n<p>Does this mean that there are more than one possible <em>correct</em> solution per example for each short/long answer, and we're scored correctly for the row if we match <em>any</em> of the possible solutions?</p>\n\n<p>Also, I've noticed that the short answers don't necessarily match any of the <code>start:end</code> combinations in the <code>long_answer_candidates</code> array from the <code>annotations</code> field. It does say that all of the short answers will be a substring of a possible long answer candidate.</p>\n\n<p>&gt; A short answer might be a sentence or phrase, or even in some cases a YES/NO. The short answers are always contained within / a subset of one of the plausible long answers</p>\n\n<p>The part I find confusing is that in the evaluation section, it says the following:</p>\n\n<p>&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels </p>\n\n<p>Does this mean that we're to predict the <code>start:end</code> pair for a short answer, even if it doesn't exist as a long answer candidate? The above wording seems to imply that we're given the ground truth labels for short answers somewhere, but the only data that looks like a candidate are the long answer candidates.</p>\n\n<p>Would truly appreciate clarity on this.</p>\n\n<p>Thanks!</p>\n\n<p>EDIT:</p>\n\n<p>It appears that yes, for short answer you need to predict the indices yourself. They're only guaranteed to be within a long answer candidate.</p>",
  "messages": [
    {
      "id": 670973,
      "postDate": "2019-11-12T05:23:54.527Z",
      "content": "<p>In the evaluation description, these lines exist:</p>\n\n<p>&gt; There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</p>\n\n<p>Does this mean that there are more than one possible <em>correct</em> solution per example for each short/long answer, and we're scored correctly for the row if we match <em>any</em> of the possible solutions?</p>\n\n<p>Also, I've noticed that the short answers don't necessarily match any of the <code>start:end</code> combinations in the <code>long_answer_candidates</code> array from the <code>annotations</code> field. It does say that all of the short answers will be a substring of a possible long answer candidate.</p>\n\n<p>&gt; A short answer might be a sentence or phrase, or even in some cases a YES/NO. The short answers are always contained within / a subset of one of the plausible long answers</p>\n\n<p>The part I find confusing is that in the evaluation section, it says the following:</p>\n\n<p>&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels </p>\n\n<p>Does this mean that we're to predict the <code>start:end</code> pair for a short answer, even if it doesn't exist as a long answer candidate? The above wording seems to imply that we're given the ground truth labels for short answers somewhere, but the only data that looks like a candidate are the long answer candidates.</p>\n\n<p>Would truly appreciate clarity on this.</p>\n\n<p>Thanks!</p>\n\n<p>EDIT:</p>\n\n<p>It appears that yes, for short answer you need to predict the indices yourself. They're only guaranteed to be within a long answer candidate.</p>",
      "rawMarkdown": "In the evaluation description, these lines exist:\n\n&gt; There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\n\nDoes this mean that there are more than one possible _correct_ solution per example for each short/long answer, and we're scored correctly for the row if we match _any_ of the possible solutions?\n\nAlso, I've noticed that the short answers don't necessarily match any of the `start:end` combinations in the `long_answer_candidates` array from the `annotations` field. It does say that all of the short answers will be a substring of a possible long answer candidate.\n\n&gt; A short answer might be a sentence or phrase, or even in some cases a YES/NO. The short answers are always contained within / a subset of one of the plausible long answers\n\nThe part I find confusing is that in the evaluation section, it says the following:\n\n&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels \n\nDoes this mean that we're to predict the `start:end` pair for a short answer, even if it doesn't exist as a long answer candidate? The above wording seems to imply that we're given the ground truth labels for short answers somewhere, but the only data that looks like a candidate are the long answer candidates.\n\nWould truly appreciate clarity on this.\n\nThanks!\n\nEDIT:\n\nIt appears that yes, for short answer you need to predict the indices yourself. They're only guaranteed to be within a long answer candidate.\n\n",
      "votes": 10
    },
    {
      "id": 680130,
      "postDate": "2019-11-24T04:10:44.170Z",
      "content": "<blockquote>\n  <p>Does this mean that there are more than one possible correct solution per example for each short/long answer, and we're scored correctly for the row if we match any of the possible solutions?</p>\n</blockquote>\n\n<p>In the original Natural Questions dataset there can be multiple short and long answers and matching any of them is correct. It's documented in the <a href=\"https://github.com/google-research-datasets/natural-questions/blob/c2c9b2fd85b5b23ae5e313b0d1c2658d1c1cd387/nq_eval.py#L71\">evaluation script</a>. I presume the same is true here. Note though that as documented on the evaluation page the threshold calculation is not applied (in the original a confidence is submitted).</p>\n\n<blockquote>\n  <p>The public test data set does not give start/end for short answers or even annotate if there exists a short answer</p>\n</blockquote>\n\n<p>Yes it does, the <code>annotations</code> element has both a <code>yes_no_answer</code> (YES/NO/NONE) and a <code>short_answers</code> list. For instance:\n```</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>obj['example_id']\n      4035615966981436342\n      obj['annotations'][0]\n      [{'yes_no_answer': 'NONE',\n        'long_answer': {'start_token': 22, 'candidate_index': 0, 'end_token': 235},\n        'short_answers': [{'start_token': 204, 'end_token': 210},\n         {'start_token': 211, 'end_token': 217},\n         {'start_token': 218, 'end_token': 224},\n         {'start_token': 226, 'end_token': 233}],\n        'annotation_id': 12271164458330946017}]\n      ```\n      As you can see all the possible short answers are contained within the single long answer (I haven't actually checked this applies across the whole set, but it is supposed to).</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>In spite of what the evaluation page says about multiple long answers I find that the training set items never have more than one long answer. Across the training dataset the frequency of short answer counts are:\n<code>\n0    177140\n1     96499\n2      5543\n3      2021\n4      1027\n5       628\n6       352\n7       286\n8       206\n10      183\n9       159\n12       10\n11        3\n21        2\n13        2\n17        2\n16        1\n18        1\n25        1\n</code></p>",
      "rawMarkdown": "&gt; Does this mean that there are more than one possible correct solution per example for each short/long answer, and we're scored correctly for the row if we match any of the possible solutions?\n\nIn the original Natural Questions dataset there can be multiple short and long answers and matching any of them is correct. It's documented in the [evaluation script](https://github.com/google-research-datasets/natural-questions/blob/c2c9b2fd85b5b23ae5e313b0d1c2658d1c1cd387/nq_eval.py#L71). I presume the same is true here. Note though that as documented on the evaluation page the threshold calculation is not applied (in the original a confidence is submitted).\n\n&gt; The public test data set does not give start/end for short answers or even annotate if there exists a short answer\n\nYes it does, the `annotations` element has both a `yes_no_answer` (YES/NO/NONE) and a `short_answers` list. For instance:\n```\n&gt;&gt;&gt; obj['example_id']\n4035615966981436342\n&gt;&gt;&gt; obj['annotations'][0]\n[{'yes_no_answer': 'NONE',\n  'long_answer': {'start_token': 22, 'candidate_index': 0, 'end_token': 235},\n  'short_answers': [{'start_token': 204, 'end_token': 210},\n   {'start_token': 211, 'end_token': 217},\n   {'start_token': 218, 'end_token': 224},\n   {'start_token': 226, 'end_token': 233}],\n  'annotation_id': 12271164458330946017}]\n```\nAs you can see all the possible short answers are contained within the single long answer (I haven't actually checked this applies across the whole set, but it is supposed to).\n\nIn spite of what the evaluation page says about multiple long answers I find that the training set items never have more than one long answer. Across the training dataset the frequency of short answer counts are:\n```\n0    177140\n1     96499\n2      5543\n3      2021\n4      1027\n5       628\n6       352\n7       286\n8       206\n10      183\n9       159\n12       10\n11        3\n21        2\n13        2\n17        2\n16        1\n18        1\n25        1\n```",
      "replies": [
        {
          "id": 680331,
          "postDate": "2019-11-24T13:29:18.537Z",
          "content": "<p>Thanks. I just verified last night that all short answers are within the long answer. There's only one long answer per question. That's why I don't understand why <a href=\"/binnersley\">@binnersley</a> says \"five labels for long answers\"</p>\n\n<p>The ID you have noted (4035615966981436342) is actually from the training dataset. The test dataset has no annotations.</p>\n\n<p>For predicting multiple short answers, the model should predict all short answers. Consider example 5619629171535320779: \"order of harry potter movies first to last\". This has 8 short answers. My guess (hope someone can confirm) is that the submission must be as follows:\n<code>\n5619629171535320779_short,2234:2241\n5619629171535320779_short,2273:2280\n...\n</code></p>",
          "rawMarkdown": "Thanks. I just verified last night that all short answers are within the long answer. There's only one long answer per question. That's why I don't understand why @binnersley says \"five labels for long answers\"\n\nThe ID you have noted (4035615966981436342) is actually from the training dataset. The test dataset has no annotations.\n\nFor predicting multiple short answers, the model should predict all short answers. Consider example 5619629171535320779: \"order of harry potter movies first to last\". This has 8 short answers. My guess (hope someone can confirm) is that the submission must be as follows:\n```\n5619629171535320779_short,2234:2241\n5619629171535320779_short,2273:2280\n...\n```\n",
          "votes": 2
        },
        {
          "id": 680641,
          "postDate": "2019-11-25T02:30:05.623Z",
          "content": "<p>The five labels statement is from the evaluation page. Re-reading it, it is specifically referring to ground truth labels so it may be that there are multiple potential long answers but the training set only contains a single one. Though this seems a little odd, I would expect the training set to contain all possible answers, as it does for short answers.</p>\n\n<p>Yes, I had misread that, the test set will of course contain no ground truths as that's what you're predicting.</p>\n\n<p>I'm pretty sure you only predict one long and one short answer for each example, each potentially blank. This is the format of the sample submission and the example on the evaluation page.</p>",
          "rawMarkdown": "The five labels statement is from the evaluation page. Re-reading it, it is specifically referring to ground truth labels so it may be that there are multiple potential long answers but the training set only contains a single one. Though this seems a little odd, I would expect the training set to contain all possible answers, as it does for short answers.\n\nYes, I had misread that, the test set will of course contain no ground truths as that's what you're predicting.\n\nI'm pretty sure you only predict one long and one short answer for each example, each potentially blank. This is the format of the sample submission and the example on the evaluation page."
        },
        {
          "id": 680662,
          "postDate": "2019-11-25T03:19:29.860Z",
          "content": "<p>Thanks for clarifying about the long answer. Makes sense.\nNot so convinced about the short answers.</p>",
          "rawMarkdown": "Thanks for clarifying about the long answer. Makes sense.\nNot so convinced about the short answers."
        },
        {
          "id": 680691,
          "postDate": "2019-11-25T04:45:51.450Z",
          "content": "<p>The bert joint baseline predicts one long and one short answer per example, following the sample submission and gets a micro F1 score of 0.7 which I understand indicates an accuracy of 70% across both long and short answers. I don't think it could get 70% accuracy if it were only predicting one of multiple expected answers.</p>",
          "rawMarkdown": "The bert joint baseline predicts one long and one short answer per example, following the sample submission and gets a micro F1 score of 0.7 which I understand indicates an accuracy of 70% across both long and short answers. I don't think it could get 70% accuracy if it were only predicting one of multiple expected answers."
        }
      ]
    },
    {
      "id": 679766,
      "postDate": "2019-11-23T10:39:15.033Z",
      "content": "<p>&gt; There may be up to five labels for long answers, and more for short. </p>\n\n<p>Did you get any clarity on this?</p>",
      "rawMarkdown": "&gt; There may be up to five labels for long answers, and more for short. \n\nDid you get any clarity on this?"
    },
    {
      "id": 679763,
      "postDate": "2019-11-23T10:36:58.777Z",
      "content": "<p>I think you're right in saying that short answers don't have <code>start:end</code> candidates. Once model has predicted the long answer, the short answer has to be somewhere within it. Seems tough to predict the exact indices.</p>",
      "rawMarkdown": "I think you're right in saying that short answers don't have `start:end` candidates. Once model has predicted the long answer, the short answer has to be somewhere within it. Seems tough to predict the exact indices."
    },
    {
      "id": 676946,
      "postDate": "2019-11-19T16:34:59.920Z",
      "content": "<p>My question is: Can the short answer prediction to the test data also be slicing of the corpus, not only YES/NO/NONE answer?</p>",
      "rawMarkdown": "My question is: Can the short answer prediction to the test data also be slicing of the corpus, not only YES/NO/NONE answer?",
      "replies": [
        {
          "id": 676964,
          "postDate": "2019-11-19T16:55:34.310Z",
          "content": "<p>Yes. Take a look at some EDA or perform your own. </p>",
          "rawMarkdown": "Yes. Take a look at some EDA or perform your own. "
        }
      ]
    },
    {
      "id": 675683,
      "postDate": "2019-11-18T12:20:37.980Z",
      "content": "<p>Thanks for asking. I had a similar doubt. </p>\n\n<p>If you look at the same submission file, short answers are blank, YES or NO.</p>\n\n<p>The public test data set does not give start/end for short answers or even annotate if there exists a short answer or a YES/NO answer. I guess the model has to figure this out on its own.</p>\n\n<p>Also, \"The short answers are always contained within / a subset of one of the plausible long answers\": I don't find this in the description. Could you please share where you read this?</p>",
      "rawMarkdown": "Thanks for asking. I had a similar doubt. \n\nIf you look at the same submission file, short answers are blank, YES or NO.\n\nThe public test data set does not give start/end for short answers or even annotate if there exists a short answer or a YES/NO answer. I guess the model has to figure this out on its own.\n\nAlso, \"The short answers are always contained within / a subset of one of the plausible long answers\": I don't find this in the description. Could you please share where you read this?\n\n",
      "replies": [
        {
          "id": 676968,
          "postDate": "2019-11-19T17:00:31.523Z",
          "content": "<p>It’s on Data description page. </p>",
          "rawMarkdown": "It’s on Data description page. "
        }
      ]
    },
    {
      "id": 673694,
      "postDate": "2019-11-15T11:00:40.133Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 680130,
      "author_name": "Thomas Brandon",
      "author_url": "",
      "post_date": "2019-11-24T04:10:44.170000",
      "content": "<blockquote>\n  <p>Does this mean that there are more than one possible correct solution per example for each short/long answer, and we're scored correctly for the row if we match any of the possible solutions?</p>\n</blockquote>\n\n<p>In the original Natural Questions dataset there can be multiple short and long answers and matching any of them is correct. It's documented in the <a href=\"https://github.com/google-research-datasets/natural-questions/blob/c2c9b2fd85b5b23ae5e313b0d1c2658d1c1cd387/nq_eval.py#L71\">evaluation script</a>. I presume the same is true here. Note though that as documented on the evaluation page the threshold calculation is not applied (in the original a confidence is submitted).</p>\n\n<blockquote>\n  <p>The public test data set does not give start/end for short answers or even annotate if there exists a short answer</p>\n</blockquote>\n\n<p>Yes it does, the <code>annotations</code> element has both a <code>yes_no_answer</code> (YES/NO/NONE) and a <code>short_answers</code> list. For instance:\n```</p>\n\n<blockquote>\n  <blockquote>\n    <blockquote>\n      <p>obj['example_id']\n      4035615966981436342\n      obj['annotations'][0]\n      [{'yes_no_answer': 'NONE',\n        'long_answer': {'start_token': 22, 'candidate_index': 0, 'end_token': 235},\n        'short_answers': [{'start_token': 204, 'end_token': 210},\n         {'start_token': 211, 'end_token': 217},\n         {'start_token': 218, 'end_token': 224},\n         {'start_token': 226, 'end_token': 233}],\n        'annotation_id': 12271164458330946017}]\n      ```\n      As you can see all the possible short answers are contained within the single long answer (I haven't actually checked this applies across the whole set, but it is supposed to).</p>\n    </blockquote>\n  </blockquote>\n</blockquote>\n\n<p>In spite of what the evaluation page says about multiple long answers I find that the training set items never have more than one long answer. Across the training dataset the frequency of short answer counts are:\n<code>\n0    177140\n1     96499\n2      5543\n3      2021\n4      1027\n5       628\n6       352\n7       286\n8       206\n10      183\n9       159\n12       10\n11        3\n21        2\n13        2\n17        2\n16        1\n18        1\n25        1\n</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 680331,
          "author_name": "Arvind Padmanabhan",
          "author_url": "",
          "post_date": "2019-11-24T13:29:18.537000",
          "content": "<p>Thanks. I just verified last night that all short answers are within the long answer. There's only one long answer per question. That's why I don't understand why <a href=\"/binnersley\">@binnersley</a> says \"five labels for long answers\"</p>\n\n<p>The ID you have noted (4035615966981436342) is actually from the training dataset. The test dataset has no annotations.</p>\n\n<p>For predicting multiple short answers, the model should predict all short answers. Consider example 5619629171535320779: \"order of harry potter movies first to last\". This has 8 short answers. My guess (hope someone can confirm) is that the submission must be as follows:\n<code>\n5619629171535320779_short,2234:2241\n5619629171535320779_short,2273:2280\n...\n</code></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 680641,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2019-11-25T02:30:05.623000",
          "content": "<p>The five labels statement is from the evaluation page. Re-reading it, it is specifically referring to ground truth labels so it may be that there are multiple potential long answers but the training set only contains a single one. Though this seems a little odd, I would expect the training set to contain all possible answers, as it does for short answers.</p>\n\n<p>Yes, I had misread that, the test set will of course contain no ground truths as that's what you're predicting.</p>\n\n<p>I'm pretty sure you only predict one long and one short answer for each example, each potentially blank. This is the format of the sample submission and the example on the evaluation page.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 680662,
          "author_name": "Arvind Padmanabhan",
          "author_url": "",
          "post_date": "2019-11-25T03:19:29.860000",
          "content": "<p>Thanks for clarifying about the long answer. Makes sense.\nNot so convinced about the short answers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 680691,
          "author_name": "Thomas Brandon",
          "author_url": "",
          "post_date": "2019-11-25T04:45:51.450000",
          "content": "<p>The bert joint baseline predicts one long and one short answer per example, following the sample submission and gets a micro F1 score of 0.7 which I understand indicates an accuracy of 70% across both long and short answers. I don't think it could get 70% accuracy if it were only predicting one of multiple expected answers.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 679766,
      "author_name": "Arvind Padmanabhan",
      "author_url": "",
      "post_date": "2019-11-23T10:39:15.033000",
      "content": "<p>&gt; There may be up to five labels for long answers, and more for short. </p>\n\n<p>Did you get any clarity on this?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 679763,
      "author_name": "Arvind Padmanabhan",
      "author_url": "",
      "post_date": "2019-11-23T10:36:58.777000",
      "content": "<p>I think you're right in saying that short answers don't have <code>start:end</code> candidates. Once model has predicted the long answer, the short answer has to be somewhere within it. Seems tough to predict the exact indices.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 676946,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2019-11-19T16:34:59.920000",
      "content": "<p>My question is: Can the short answer prediction to the test data also be slicing of the corpus, not only YES/NO/NONE answer?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676964,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-19T16:55:34.310000",
          "content": "<p>Yes. Take a look at some EDA or perform your own. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 675683,
      "author_name": "Arvind Padmanabhan",
      "author_url": "",
      "post_date": "2019-11-18T12:20:37.980000",
      "content": "<p>Thanks for asking. I had a similar doubt. </p>\n\n<p>If you look at the same submission file, short answers are blank, YES or NO.</p>\n\n<p>The public test data set does not give start/end for short answers or even annotate if there exists a short answer or a YES/NO answer. I guess the model has to figure this out on its own.</p>\n\n<p>Also, \"The short answers are always contained within / a subset of one of the plausible long answers\": I don't find this in the description. Could you please share where you read this?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 676968,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-19T17:00:31.523000",
          "content": "<p>It’s on Data description page. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673694,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-15T11:00:40.133000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "670973": "In the evaluation description, these lines exist:\n\n&gt; There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\n\nDoes this mean that there are more than one possible _correct_ solution per example for each short/long answer, and we're scored correctly for the row if we match _any_ of the possible solutions?\n\nAlso, I've noticed that the short answers don't necessarily match any of the `start:end` combinations in the `long_answer_candidates` array from the `annotations` field. It does say that all of the short answers will be a substring of a possible long answer candidate.\n\n&gt; A short answer might be a sentence or phrase, or even in some cases a YES/NO. The short answers are always contained within / a subset of one of the plausible long answers\n\nThe part I find confusing is that in the evaluation section, it says the following:\n\n&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels \n\nDoes this mean that we're to predict the `start:end` pair for a short answer, even if it doesn't exist as a long answer candidate? The above wording seems to imply that we're given the ground truth labels for short answers somewhere, but the only data that looks like a candidate are the long answer candidates.\n\nWould truly appreciate clarity on this.\n\nThanks!\n\nEDIT:\n\nIt appears that yes, for short answer you need to predict the indices yourself. They're only guaranteed to be within a long answer candidate.\n\n",
    "680130": "&gt; Does this mean that there are more than one possible correct solution per example for each short/long answer, and we're scored correctly for the row if we match any of the possible solutions?\n\nIn the original Natural Questions dataset there can be multiple short and long answers and matching any of them is correct. It's documented in the [evaluation script](https://github.com/google-research-datasets/natural-questions/blob/c2c9b2fd85b5b23ae5e313b0d1c2658d1c1cd387/nq_eval.py#L71). I presume the same is true here. Note though that as documented on the evaluation page the threshold calculation is not applied (in the original a confidence is submitted).\n\n&gt; The public test data set does not give start/end for short answers or even annotate if there exists a short answer\n\nYes it does, the `annotations` element has both a `yes_no_answer` (YES/NO/NONE) and a `short_answers` list. For instance:\n```\n&gt;&gt;&gt; obj['example_id']\n4035615966981436342\n&gt;&gt;&gt; obj['annotations'][0]\n[{'yes_no_answer': 'NONE',\n  'long_answer': {'start_token': 22, 'candidate_index': 0, 'end_token': 235},\n  'short_answers': [{'start_token': 204, 'end_token': 210},\n   {'start_token': 211, 'end_token': 217},\n   {'start_token': 218, 'end_token': 224},\n   {'start_token': 226, 'end_token': 233}],\n  'annotation_id': 12271164458330946017}]\n```\nAs you can see all the possible short answers are contained within the single long answer (I haven't actually checked this applies across the whole set, but it is supposed to).\n\nIn spite of what the evaluation page says about multiple long answers I find that the training set items never have more than one long answer. Across the training dataset the frequency of short answer counts are:\n```\n0    177140\n1     96499\n2      5543\n3      2021\n4      1027\n5       628\n6       352\n7       286\n8       206\n10      183\n9       159\n12       10\n11        3\n21        2\n13        2\n17        2\n16        1\n18        1\n25        1\n```",
    "679766": "&gt; There may be up to five labels for long answers, and more for short. \n\nDid you get any clarity on this?",
    "679763": "I think you're right in saying that short answers don't have `start:end` candidates. Once model has predicted the long answer, the short answer has to be somewhere within it. Seems tough to predict the exact indices.",
    "676946": "My question is: Can the short answer prediction to the test data also be slicing of the corpus, not only YES/NO/NONE answer?",
    "675683": "Thanks for asking. I had a similar doubt. \n\nIf you look at the same submission file, short answers are blank, YES or NO.\n\nThe public test data set does not give start/end for short answers or even annotate if there exists a short answer or a YES/NO answer. I guess the model has to figure this out on its own.\n\nAlso, \"The short answers are always contained within / a subset of one of the plausible long answers\": I don't find this in the description. Could you please share where you read this?\n\n",
    "673694": ""
  }
}