{
  "id": 115715,
  "title": "top_level tag in long_answer_candidates json",
  "url": "/competitions/tensorflow2-question-answering/discussion/115715",
  "author_name": "Alex Federation",
  "post_date": "2019-11-04T21:38:18.652000",
  "votes": 5,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hi, \nI had a question about the training data - there's a flag in the column long_answer_candidates called 'top_level'. Can anyone clarify what this flag is indicating?</p>\n\n<p>Thanks!</p>",
  "messages": [
    {
      "id": 665351,
      "postDate": "2019-11-04T22:54:36.587Z",
      "content": "<p>from <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/README.md\">https://github.com/google-research-datasets/natural-questions/blob/master/README.md</a> <a href=\"/afederation\">@afederation</a> </p>\n\n<p><strong>Long Answer Candidates</strong>\nThe first task in Natural Questions is to identify the smallest HTML bounding box that contains all of the information required to infer the answer to a question. These long answers can be paragraphs, lists, list items, tables, or table rows. While the candidates can be inferred directly from the HTML or token sequence, we also include a list of long answer candidates for convenience. Each candidate is defined in terms of offsets into both the HTML and the document tokens. As with all other annotations, start offsets are inclusive and end offsets are exclusive.</p>\n\n<p><code>\n\"long_answer_candidates\": [\n  { \"start_byte\": 32, \"end_byte\": 106, \"start_token\": 5, \"end_token\": 22, \"top_level\": true },\n  { \"start_byte\": 65, \"end_byte\": 102, \"start_token\": 13, \"end_token\": 21, \"top_level\": false },\n</code>\nIn this example, <strong>you can see that the second long answer candidate is contained within the first. We do not disallow nested long answer candidates, we just ask annotators to find the smallest candidate containing all of the information required to infer the answer to the question</strong>. However, we do observe that 95% of all long answers (including all paragraph answers) are not nested below any other candidates. Since we believe that some users may want to start by only considering <strong>non-overlapping candidates, we include a boolean flag top_level that identifies whether a candidate is nested below another (top_level = False) or not (top_level = True)</strong>. Please be aware that this flag is only included for convenience and it is not related to the task definition in any way. For more information about the distribution of long answer types, please see the data statistics section below.</p>",
      "rawMarkdown": "from https://github.com/google-research-datasets/natural-questions/blob/master/README.md @afederation \n\n**Long Answer Candidates**\nThe first task in Natural Questions is to identify the smallest HTML bounding box that contains all of the information required to infer the answer to a question. These long answers can be paragraphs, lists, list items, tables, or table rows. While the candidates can be inferred directly from the HTML or token sequence, we also include a list of long answer candidates for convenience. Each candidate is defined in terms of offsets into both the HTML and the document tokens. As with all other annotations, start offsets are inclusive and end offsets are exclusive.\n\n```\n\"long_answer_candidates\": [\n  { \"start_byte\": 32, \"end_byte\": 106, \"start_token\": 5, \"end_token\": 22, \"top_level\": true },\n  { \"start_byte\": 65, \"end_byte\": 102, \"start_token\": 13, \"end_token\": 21, \"top_level\": false },\n```\nIn this example, **you can see that the second long answer candidate is contained within the first. We do not disallow nested long answer candidates, we just ask annotators to find the smallest candidate containing all of the information required to infer the answer to the question**. However, we do observe that 95% of all long answers (including all paragraph answers) are not nested below any other candidates. Since we believe that some users may want to start by only considering **non-overlapping candidates, we include a boolean flag top_level that identifies whether a candidate is nested below another (top_level = False) or not (top_level = True)**. Please be aware that this flag is only included for convenience and it is not related to the task definition in any way. For more information about the distribution of long answer types, please see the data statistics section below.\n",
      "votes": 12
    },
    {
      "id": 665314,
      "postDate": "2019-11-04T21:38:18.653Z",
      "content": "<p>Hi, \nI had a question about the training data - there's a flag in the column long_answer_candidates called 'top_level'. Can anyone clarify what this flag is indicating?</p>\n\n<p>Thanks!</p>",
      "rawMarkdown": "Hi, \nI had a question about the training data - there's a flag in the column long_answer_candidates called 'top_level'. Can anyone clarify what this flag is indicating?\n\nThanks!",
      "votes": 5
    },
    {
      "id": 685638,
      "postDate": "2019-12-02T03:18:06.923Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 665351,
      "author_name": "SeshuRaju 🧘‍♂️",
      "author_url": "",
      "post_date": "2019-11-04T22:54:36.587000",
      "content": "<p>from <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/README.md\">https://github.com/google-research-datasets/natural-questions/blob/master/README.md</a> <a href=\"/afederation\">@afederation</a> </p>\n\n<p><strong>Long Answer Candidates</strong>\nThe first task in Natural Questions is to identify the smallest HTML bounding box that contains all of the information required to infer the answer to a question. These long answers can be paragraphs, lists, list items, tables, or table rows. While the candidates can be inferred directly from the HTML or token sequence, we also include a list of long answer candidates for convenience. Each candidate is defined in terms of offsets into both the HTML and the document tokens. As with all other annotations, start offsets are inclusive and end offsets are exclusive.</p>\n\n<p><code>\n\"long_answer_candidates\": [\n  { \"start_byte\": 32, \"end_byte\": 106, \"start_token\": 5, \"end_token\": 22, \"top_level\": true },\n  { \"start_byte\": 65, \"end_byte\": 102, \"start_token\": 13, \"end_token\": 21, \"top_level\": false },\n</code>\nIn this example, <strong>you can see that the second long answer candidate is contained within the first. We do not disallow nested long answer candidates, we just ask annotators to find the smallest candidate containing all of the information required to infer the answer to the question</strong>. However, we do observe that 95% of all long answers (including all paragraph answers) are not nested below any other candidates. Since we believe that some users may want to start by only considering <strong>non-overlapping candidates, we include a boolean flag top_level that identifies whether a candidate is nested below another (top_level = False) or not (top_level = True)</strong>. Please be aware that this flag is only included for convenience and it is not related to the task definition in any way. For more information about the distribution of long answer types, please see the data statistics section below.</p>",
      "votes": 12,
      "replies": []
    },
    {
      "id": 685638,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-02T03:18:06.923000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "665351": "from https://github.com/google-research-datasets/natural-questions/blob/master/README.md @afederation \n\n**Long Answer Candidates**\nThe first task in Natural Questions is to identify the smallest HTML bounding box that contains all of the information required to infer the answer to a question. These long answers can be paragraphs, lists, list items, tables, or table rows. While the candidates can be inferred directly from the HTML or token sequence, we also include a list of long answer candidates for convenience. Each candidate is defined in terms of offsets into both the HTML and the document tokens. As with all other annotations, start offsets are inclusive and end offsets are exclusive.\n\n```\n\"long_answer_candidates\": [\n  { \"start_byte\": 32, \"end_byte\": 106, \"start_token\": 5, \"end_token\": 22, \"top_level\": true },\n  { \"start_byte\": 65, \"end_byte\": 102, \"start_token\": 13, \"end_token\": 21, \"top_level\": false },\n```\nIn this example, **you can see that the second long answer candidate is contained within the first. We do not disallow nested long answer candidates, we just ask annotators to find the smallest candidate containing all of the information required to infer the answer to the question**. However, we do observe that 95% of all long answers (including all paragraph answers) are not nested below any other candidates. Since we believe that some users may want to start by only considering **non-overlapping candidates, we include a boolean flag top_level that identifies whether a candidate is nested below another (top_level = False) or not (top_level = True)**. Please be aware that this flag is only included for convenience and it is not related to the task definition in any way. For more information about the distribution of long answer types, please see the data statistics section below.\n",
    "665314": "Hi, \nI had a question about the training data - there's a flag in the column long_answer_candidates called 'top_level'. Can anyone clarify what this flag is indicating?\n\nThanks!",
    "685638": ""
  }
}