{
  "id": 123266,
  "title": "Question about getting short_answer end_token",
  "url": "/competitions/tensorflow2-question-answering/discussion/123266",
  "author_name": "higepon",
  "post_date": "2019-12-26T08:32:59.002000",
  "votes": 3,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hi everyone,\nI'm writing my own PyTorch training kernel and reading through bert_util.py for reference.\nOne thing I don't understand is why get_first_annotation method choose end_token as follows.</p>\n\n<p>```\n  positive_annotations = sorted(\n      [a for a in e[\"annotations\"] if has_long_answer(a)],\n      key=lambda a: a[\"long_answer\"][\"candidate_index\"])</p>\n\n<p>for a in positive_annotations:\n    if a[\"short_answers\"]:\n      idx = a[\"long_answer\"][\"candidate_index\"]\n      start_token = a[\"short_answers\"][0][\"start_token\"]\n      end_token = a[\"short_answers\"][-1][\"end_token\"]\n      return a, idx, (token_to_char_offset(e, idx, start_token),\n                      token_to_char_offset(e, idx, end_token) - 1)\n```</p>\n\n<p>Why <code>end_token = a[\"short_answers\"][-1][\"end_token\"]</code>?</p>\n\n<ul>\n<li>Why they get end_token from other short_answer? (Specifically last short_answer?)</li>\n<li>Isn't this make the model harder to predict exact short_answer because it mixing multiple short_answers into one?</li>\n</ul>\n\n<p>thanks,\nTaro</p>",
  "messages": [
    {
      "id": 703480,
      "postDate": "2019-12-26T08:32:59.003Z",
      "content": "<p>Hi everyone,\nI'm writing my own PyTorch training kernel and reading through bert_util.py for reference.\nOne thing I don't understand is why get_first_annotation method choose end_token as follows.</p>\n\n<p>```\n  positive_annotations = sorted(\n      [a for a in e[\"annotations\"] if has_long_answer(a)],\n      key=lambda a: a[\"long_answer\"][\"candidate_index\"])</p>\n\n<p>for a in positive_annotations:\n    if a[\"short_answers\"]:\n      idx = a[\"long_answer\"][\"candidate_index\"]\n      start_token = a[\"short_answers\"][0][\"start_token\"]\n      end_token = a[\"short_answers\"][-1][\"end_token\"]\n      return a, idx, (token_to_char_offset(e, idx, start_token),\n                      token_to_char_offset(e, idx, end_token) - 1)\n```</p>\n\n<p>Why <code>end_token = a[\"short_answers\"][-1][\"end_token\"]</code>?</p>\n\n<ul>\n<li>Why they get end_token from other short_answer? (Specifically last short_answer?)</li>\n<li>Isn't this make the model harder to predict exact short_answer because it mixing multiple short_answers into one?</li>\n</ul>\n\n<p>thanks,\nTaro</p>",
      "rawMarkdown": "Hi everyone,\nI'm writing my own PyTorch training kernel and reading through bert\\_util.py for reference.\nOne thing I don't understand is why get\\_first\\_annotation method choose end_token as follows.\n\n```\n  positive_annotations = sorted(\n      [a for a in e[\"annotations\"] if has_long_answer(a)],\n      key=lambda a: a[\"long_answer\"][\"candidate_index\"])\n\n  for a in positive_annotations:\n    if a[\"short_answers\"]:\n      idx = a[\"long_answer\"][\"candidate_index\"]\n      start_token = a[\"short_answers\"][0][\"start_token\"]\n      end_token = a[\"short_answers\"][-1][\"end_token\"]\n      return a, idx, (token_to_char_offset(e, idx, start_token),\n                      token_to_char_offset(e, idx, end_token) - 1)\n```\n\nWhy ```end_token = a[\"short_answers\"][-1][\"end_token\"]```?\n\n- Why they get end_token from other short_answer? (Specifically last short_answer?)\n- Isn't this make the model harder to predict exact short\\_answer because it mixing multiple short\\_answers into one?\n\nthanks,\nTaro",
      "votes": 3
    },
    {
      "id": 706430,
      "postDate": "2019-12-30T11:09:12.273Z",
      "content": "<p>I think that there is only one short answer in the train set, but it may consist of multiple short answer spans. In the test set, there may be multiple short answers (from 5 different annotators), and each of them may consist of multiple spans.  One of the possible improvements mentioned by the authors of the baseline model is to account for multiple short answer spans. </p>",
      "rawMarkdown": "I think that there is only one short answer in the train set, but it may consist of multiple short answer spans. In the test set, there may be multiple short answers (from 5 different annotators), and each of them may consist of multiple spans.  One of the possible improvements mentioned by the authors of the baseline model is to account for multiple short answer spans. ",
      "votes": 1,
      "replies": [
        {
          "id": 706894,
          "postDate": "2019-12-31T00:37:23.147Z",
          "content": "<p>Thank you Darek for your reply!\nI see. Would you be able to give us an example or a link what you mentioned here? I'd like to understand it.</p>",
          "rawMarkdown": "Thank you Darek for your reply!\nI see. Would you be able to give us an example or a link what you mentioned here? I'd like to understand it."
        },
        {
          "id": 706937,
          "postDate": "2019-12-31T03:31:02.390Z",
          "content": "<p>This is an excerpt from the article ( <a href=\"https://ai.google/research/pubs/pub47761.pdf\">https://ai.google/research/pubs/pub47761.pdf</a> ): \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F49a2838297a5e56b8c6cb9b60195739b%2FScreen%20Shot%202019-12-31%20at%2004.24.35.png?generation=1577762756198901&amp;alt=media\" alt=\"\">\nSo if a short answer is multiple entities, they can be indicated by multiple spans within the long answer.\nHere's an example from the dev set (from the natural questions website, it should be similar to our test set I hope) - there are 4 short answers, and one of them consists of 2 spans:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F24d4e1454142eb84548314b77dee0346%2FScreen%20Shot%202019-12-31%20at%2004.19.22.png?generation=1577762953334486&amp;alt=media\" alt=\"\">\n Hope this helps :) </p>",
          "rawMarkdown": "This is an excerpt from the article ( https://ai.google/research/pubs/pub47761.pdf ): \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F49a2838297a5e56b8c6cb9b60195739b%2FScreen%20Shot%202019-12-31%20at%2004.24.35.png?generation=1577762756198901&amp;alt=media)\nSo if a short answer is multiple entities, they can be indicated by multiple spans within the long answer.\nHere's an example from the dev set (from the natural questions website, it should be similar to our test set I hope) - there are 4 short answers, and one of them consists of 2 spans:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F24d4e1454142eb84548314b77dee0346%2FScreen%20Shot%202019-12-31%20at%2004.19.22.png?generation=1577762953334486&amp;alt=media)\n Hope this helps :) ",
          "votes": 1
        },
        {
          "id": 707086,
          "postDate": "2019-12-31T08:40:39.590Z",
          "content": "<p>Thank you Darek! I extracted the corresponding data from the dev set.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2Fb81456ca366cde6f650f694d1649ea87%2F2019-12-31%2017.28.33.png?generation=1577780979671866&amp;alt=media\" alt=\"\"></p>\n\n<h2>Facts</h2>\n\n<ul>\n<li>There can be multiple annotations</li>\n<li>In an annotation, there is one short_answers field. (= list)</li>\n<li>There can be multiple (start, end) pairs in short_answers field.</li>\n</ul>\n\n<h2>Scoring assumption</h2>\n\n<ul>\n<li>We predict only one short_answer for each question.</li>\n<li>The answer is considered correct if it matches one of  the short_answers</li>\n</ul>\n\n<h2>How to train?</h2>\n\n<p>We can do either\n- Choose one of the short_answers when training\n- Create multiple training instance for the short_answers\nor\n- Concatenate them as one answer? (As the code we found?)</p>\n\n<p>I'm not very sure which works better or worth doing as there are very few.</p>\n\n<h2>How to predict?</h2>\n\n<p>We just pick the highest score one. Any idea?</p>",
          "rawMarkdown": "Thank you Darek! I extracted the corresponding data from the dev set.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2Fb81456ca366cde6f650f694d1649ea87%2F2019-12-31%2017.28.33.png?generation=1577780979671866&amp;alt=media)\n\n## Facts\n- There can be multiple annotations\n- In an annotation, there is one short_answers field. (= list)\n- There can be multiple (start, end) pairs in short_answers field.\n\n## Scoring assumption\n- We predict only one short_answer for each question.\n- The answer is considered correct if it matches one of  the short_answers\n\n## How to train?\nWe can do either\n- Choose one of the short_answers when training\n- Create multiple training instance for the short_answers\nor\n- Concatenate them as one answer? (As the code we found?)\n\nI'm not very sure which works better or worth doing as there are very few.\n\n## How to predict?\nWe just pick the highest score one. Any idea?",
          "votes": 1
        },
        {
          "id": 714339,
          "postDate": "2020-01-09T10:18:10.580Z",
          "content": "<p>I basically wrote a full blog post length \"question\" about this but I've lost faith that that's particularly productive. </p>\n\n<h2>My uninformed take on this</h2>\n\n<p>If there are multiple short answer spans from the annotator, a correct answer must be precisely one single span. </p>\n\n<p>I think bert-joint was incentivized to return the smallest span covering all annotators short answer span (I wouldn't call this concatenating them). In which case those <code>[-1]</code> 's should <code>[0]</code>'s. \n(In the<code>run_nq.py</code> these appear twice <a href=\"https://github.com/google-research/language/blob/89e5b53a0e7f9f3e2da25a5da71ce5bd466acabd/language/question_answering/bert_joint/run_nq.py#L379\">here</a> and <a href=\"https://github.com/google-research/language/blob/89e5b53a0e7f9f3e2da25a5da71ce5bd466acabd/language/question_answering/bert_joint/run_nq.py#L257\">here</a>)</p>\n\n<p>Thanks <a href=\"/higepon\">@higepon</a> for this thread (and actually quite a few other of your contributions.) </p>\n\n<h2>Examples:</h2>\n\n<p>From the simplified train set, of the first 100 examples, 4 have been annotated multiple short spans (precisely: #45, #56, #74, #93). \nHere are the question numbers, questions, the annotators token spans for short answers and short answer text selected by the <code>run_nq.py</code> script:</p>\n\n<p><strong>45</strong> <code>where are the upcoming olympics to be held</code>\n<strong>short answer spans</strong> <code>[{'start_token': 204, 'end_token': 210}, {'start_token': 211, 'end_token': 217}, {'start_token': 218, 'end_token': 224}, {'start_token': 226, 'end_token': 233}]</code>\n<strong>text selected</strong> <code>Tokyo for the 2020 Summer Olympics , Beijing for the 2022 Winter Olympics , Paris for the 2024 Summer Olympics , and Los Angeles for the 2028 Summer Olympics</code>\n<strong>56</strong> <code>name some components of the central nervous system (cns)</code>\n<strong>short answer spans</strong> <code>[{'start_token': 140, 'end_token': 141}, {'start_token': 142, 'end_token': 144}]</code>\n<strong>text selected</strong> <code>brain and spinal cord</code>\n<strong>74</strong> <code>what teams are in the fa cup final</code>\n<strong>short answer spans</strong> <code>[{'start_token': 208, 'end_token': 210}, {'start_token': 211, 'end_token': 212}]</code>\n<strong>text selected</strong> <code>Manchester United and Chelsea</code>\n<strong>93</strong> <code>where was life or something like it filmed</code>\n<strong>short answer spans</strong> <code>[{'start_token': 1137, 'end_token': 1140}, {'start_token': 1145, 'end_token': 1147}]</code>\n<strong>text selected</strong> <code>Seattle , Washington although portions were filmed in downtown Vancouver</code></p>\n\n<p>By my assertion, in this competition (and in the original eval scripts, although these differ) \nhad the text selected been used to submit a prediction, then these would have all scored 0. \nFor this kaggle competition, a correct set of predictions is (a non-unique suggestion)</p>\n\n<p><strong>45</strong> <code>Tokyo for the 2020 Summer Olympics</code>\n<strong>56</strong> <code>brain</code>\n<strong>74</strong> <code>Manchester United</code>\n<strong>93</strong> <code>Seattle , Washington</code></p>",
          "rawMarkdown": "I basically wrote a full blog post length \"question\" about this but I've lost faith that that's particularly productive. \n\n## My uninformed take on this\n\nIf there are multiple short answer spans from the annotator, a correct answer must be precisely one single span. \n\nI think bert-joint was incentivized to return the smallest span covering all annotators short answer span (I wouldn't call this concatenating them). In which case those `[-1]` 's should `[0]`'s. \n(In the`run_nq.py` these appear twice [here](https://github.com/google-research/language/blob/89e5b53a0e7f9f3e2da25a5da71ce5bd466acabd/language/question_answering/bert_joint/run_nq.py#L379) and [here](https://github.com/google-research/language/blob/89e5b53a0e7f9f3e2da25a5da71ce5bd466acabd/language/question_answering/bert_joint/run_nq.py#L257))\n\n\nThanks @higepon for this thread (and actually quite a few other of your contributions.) \n\n## Examples:\n \nFrom the simplified train set, of the first 100 examples, 4 have been annotated multiple short spans (precisely: #45, #56, #74, #93). \nHere are the question numbers, questions, the annotators token spans for short answers and short answer text selected by the `run_nq.py` script:\n\n**45** `where are the upcoming olympics to be held`\n**short answer spans** `[{'start_token': 204, 'end_token': 210}, {'start_token': 211, 'end_token': 217}, {'start_token': 218, 'end_token': 224}, {'start_token': 226, 'end_token': 233}]`\n**text selected** `Tokyo for the 2020 Summer Olympics , Beijing for the 2022 Winter Olympics , Paris for the 2024 Summer Olympics , and Los Angeles for the 2028 Summer Olympics`\n**56** `name some components of the central nervous system (cns)`\n**short answer spans** `[{'start_token': 140, 'end_token': 141}, {'start_token': 142, 'end_token': 144}]`\n**text selected** `brain and spinal cord`\n**74** `what teams are in the fa cup final`\n**short answer spans** `[{'start_token': 208, 'end_token': 210}, {'start_token': 211, 'end_token': 212}]`\n**text selected** `Manchester United and Chelsea`\n**93** `where was life or something like it filmed`\n**short answer spans** `[{'start_token': 1137, 'end_token': 1140}, {'start_token': 1145, 'end_token': 1147}]`\n**text selected** `Seattle , Washington although portions were filmed in downtown Vancouver`\n\nBy my assertion, in this competition (and in the original eval scripts, although these differ) \nhad the text selected been used to submit a prediction, then these would have all scored 0. \nFor this kaggle competition, a correct set of predictions is (a non-unique suggestion)\n\n**45** `Tokyo for the 2020 Summer Olympics`\n**56** `brain`\n**74** `Manchester United`\n**93** `Seattle , Washington`\n "
        },
        {
          "id": 722360,
          "postDate": "2020-01-18T13:19:00.933Z",
          "content": "<p>Thanks! So this confirms that for multi-span short answers, each individual span is considered correct, but their merged version (text selected) is not. Right?</p>",
          "rawMarkdown": "Thanks! So this confirms that for multi-span short answers, each individual span is considered correct, but their merged version (text selected) is not. Right?"
        },
        {
          "id": 722600,
          "postDate": "2020-01-18T19:31:51.600Z",
          "content": "<p>That would be my guess... I haven't checked it though. </p>",
          "rawMarkdown": "That would be my guess... I haven't checked it though. "
        }
      ]
    },
    {
      "id": 703496,
      "postDate": "2019-12-26T08:57:44.840Z",
      "content": "<p>In my opinion, it indeed is a bad mixing</p>",
      "rawMarkdown": "In my opinion, it indeed is a bad mixing",
      "votes": 1,
      "replies": [
        {
          "id": 703499,
          "postDate": "2019-12-26T09:04:36.930Z",
          "content": "<p><a href=\"/mikelkl\">@mikelkl</a> thank you for your reply!.\nI guessed so too. But it's surprising if this is actually wrong considering the model is performing very well.</p>",
          "rawMarkdown": "@mikelkl thank you for your reply!.\nI guessed so too. But it's surprising if this is actually wrong considering the model is performing very well."
        },
        {
          "id": 703541,
          "postDate": "2019-12-26T09:29:22.390Z",
          "content": "<p>actually, i think the model performance still hv a long way to go, coz the precision of short answer is quite low, so u know where to find reason</p>",
          "rawMarkdown": "actually, i think the model performance still hv a long way to go, coz the precision of short answer is quite low, so u know where to find reason",
          "votes": 2
        },
        {
          "id": 703575,
          "postDate": "2019-12-26T10:08:28.127Z",
          "content": "<p>Yes. Thank you so much for your help!</p>",
          "rawMarkdown": "Yes. Thank you so much for your help!"
        }
      ]
    },
    {
      "id": 703489,
      "postDate": "2019-12-26T08:46:23.990Z",
      "content": "<p>There is only one ”short_answer“ in socalled \"short_answers\", so a[\"short_answers\"][0] == a[\"short_answers\"][-1]</p>",
      "rawMarkdown": "There is only one ”short_answer“ in socalled \"short_answers\", so a[\"short_answers\"][0] == a[\"short_answers\"][-1]",
      "votes": 2,
      "replies": [
        {
          "id": 703497,
          "postDate": "2019-12-26T09:02:33.723Z",
          "content": "<p>Thank you <a href=\"/noxuslol\">@noxuslol</a>! But <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/README.md#annotations\">https://github.com/google-research-datasets/natural-questions/blob/master/README.md#annotations</a> says there can be multiple short answers.</p>",
          "rawMarkdown": "Thank you @noxuslol! But https://github.com/google-research-datasets/natural-questions/blob/master/README.md#annotations says there can be multiple short answers."
        },
        {
          "id": 703506,
          "postDate": "2019-12-26T09:08:42.023Z",
          "content": "<p>That is indeed the case in the test set, but you can check the training data(simplified-nq-train.jsonl) and there is really only one short answer...</p>",
          "rawMarkdown": "That is indeed the case in the test set, but you can check the training data(simplified-nq-train.jsonl) and there is really only one short answer...",
          "votes": 1
        },
        {
          "id": 703515,
          "postDate": "2019-12-26T09:15:12.940Z",
          "content": "<p>Oh. Okay. Then this code works okay for training data, but looks really weird :)</p>",
          "rawMarkdown": "Oh. Okay. Then this code works okay for training data, but looks really weird :)"
        },
        {
          "id": 703547,
          "postDate": "2019-12-26T09:33:29.980Z",
          "content": "<p>hi there, in my EDA, there is unignorable number of multi short answers in train data</p>",
          "rawMarkdown": "hi there, in my EDA, there is unignorable number of multi short answers in train data",
          "votes": 2
        },
        {
          "id": 703557,
          "postDate": "2019-12-26T09:43:53.173Z",
          "content": "<p>Are you sure there are multiple short answers, not multiple 'long_answer_candidates'...?\nIn the cell [9] of <a href=\"https://www.kaggle.com/ragnar123/exploratory-data-analysis-and-baseline\">this EDA kernel</a> we can see that there is only one short answer and one long answer...</p>",
          "rawMarkdown": "Are you sure there are multiple short answers, not multiple 'long_answer_candidates'...?\nIn the cell [9] of [this EDA kernel](https://www.kaggle.com/ragnar123/exploratory-data-analysis-and-baseline) we can see that there is only one short answer and one long answer...",
          "votes": 1
        },
        {
          "id": 703574,
          "postDate": "2019-12-26T10:08:03.473Z",
          "content": "<p>Confirmed there are multiple short_answer(s).\nFor example:\n<code>\n\"short_answers\": [{\"start_token\": 208, \"end_token\": 210}, {\"start_token\": 211, \"end_token\": 212}], \"annotation_id\": 12688287890235717147}], \"document_url\": \"https://en.wikipedia.org//w/index.php?title=2018_FA_Cup_Final&amp;amp;oldid=842824259\", \"example_id\": 6823139846656798944}\n</code></p>",
          "rawMarkdown": "Confirmed there are multiple short_answer(s).\nFor example:\n```\n\"short_answers\": [{\"start_token\": 208, \"end_token\": 210}, {\"start_token\": 211, \"end_token\": 212}], \"annotation_id\": 12688287890235717147}], \"document_url\": \"https://en.wikipedia.org//w/index.php?title=2018_FA_Cup_Final&amp;oldid=842824259\", \"example_id\": 6823139846656798944}\n```"
        },
        {
          "id": 703584,
          "postDate": "2019-12-26T10:20:55.783Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 703591,
          "postDate": "2019-12-26T10:28:18.043Z",
          "content": "<p>I'm wrong...there are multiple short_answers</p>",
          "rawMarkdown": "I'm wrong...there are multiple short_answers",
          "votes": 2
        },
        {
          "id": 703781,
          "postDate": "2019-12-26T16:02:50.667Z",
          "content": "<p>I counted short answer counts in <a href=\"https://www.kaggle.com/kentaronakanishi/eda-question-text-and-annotations-wip\">my dirty eda kernel version 5 (not latest)</a></p>\n\n<p><img src=\"https://www.kaggleusercontent.com/kf/23145600/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..2MrnQySMMLxc779gOfnBTw.cpyFoPQO-yM86B7TPdFJdMf-ApNESmYbVHQhlPYCsG3Z5NWIQZHxJxcg3TUrJKPOrPSwsKoVk0y1rwdd98sF54WJRQPmRHuQWUbCAgk-6nIOFLa0RI2kjf4giFc_IFcOrDRJi2uJZ6n9UpMD4dX2fK2JV0rNU-u7UEj1bktddnO40HEdOMnARtI4lnAjT1z2Y78VA-B0Nnz9wl-Eo7ky5A.ZiTyvPX4y4qgBypj8UNRZw/__results___files/__results___39_2.png\" alt=\"\"></p>\n\n<p>There are some records which have multiple short answers, and the max number is 25.</p>\n\n<p>In fact, most of records have only 0 or 1 short answer. So I also think what <a href=\"/higepon\">@higepon</a> mentioned is a bug, but not critical to scores.</p>",
          "rawMarkdown": "I counted short answer counts in [my dirty eda kernel version 5 (not latest)](https://www.kaggle.com/kentaronakanishi/eda-question-text-and-annotations-wip)\n\n![](https://www.kaggleusercontent.com/kf/23145600/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..2MrnQySMMLxc779gOfnBTw.cpyFoPQO-yM86B7TPdFJdMf-ApNESmYbVHQhlPYCsG3Z5NWIQZHxJxcg3TUrJKPOrPSwsKoVk0y1rwdd98sF54WJRQPmRHuQWUbCAgk-6nIOFLa0RI2kjf4giFc_IFcOrDRJi2uJZ6n9UpMD4dX2fK2JV0rNU-u7UEj1bktddnO40HEdOMnARtI4lnAjT1z2Y78VA-B0Nnz9wl-Eo7ky5A.ZiTyvPX4y4qgBypj8UNRZw/__results___files/__results___39_2.png)\n\nThere are some records which have multiple short answers, and the max number is 25.\n\nIn fact, most of records have only 0 or 1 short answer. So I also think what @higepon mentioned is a bug, but not critical to scores.",
          "votes": 4
        },
        {
          "id": 704023,
          "postDate": "2019-12-27T00:08:49.650Z",
          "content": "<p>Thanks for confirming <a href=\"/kentaronakanishi\">@kentaronakanishi</a> !</p>",
          "rawMarkdown": "Thanks for confirming @kentaronakanishi !"
        }
      ]
    },
    {
      "id": 703501,
      "postDate": "2019-12-26T09:06:56.987Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 706430,
      "author_name": "Darek Kłeczek",
      "author_url": "",
      "post_date": "2019-12-30T11:09:12.273000",
      "content": "<p>I think that there is only one short answer in the train set, but it may consist of multiple short answer spans. In the test set, there may be multiple short answers (from 5 different annotators), and each of them may consist of multiple spans.  One of the possible improvements mentioned by the authors of the baseline model is to account for multiple short answer spans. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 706894,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-31T00:37:23.147000",
          "content": "<p>Thank you Darek for your reply!\nI see. Would you be able to give us an example or a link what you mentioned here? I'd like to understand it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 706937,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2019-12-31T03:31:02.390000",
          "content": "<p>This is an excerpt from the article ( <a href=\"https://ai.google/research/pubs/pub47761.pdf\">https://ai.google/research/pubs/pub47761.pdf</a> ): \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F49a2838297a5e56b8c6cb9b60195739b%2FScreen%20Shot%202019-12-31%20at%2004.24.35.png?generation=1577762756198901&amp;alt=media\" alt=\"\">\nSo if a short answer is multiple entities, they can be indicated by multiple spans within the long answer.\nHere's an example from the dev set (from the natural questions website, it should be similar to our test set I hope) - there are 4 short answers, and one of them consists of 2 spans:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F24d4e1454142eb84548314b77dee0346%2FScreen%20Shot%202019-12-31%20at%2004.19.22.png?generation=1577762953334486&amp;alt=media\" alt=\"\">\n Hope this helps :) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 707086,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-31T08:40:39.590000",
          "content": "<p>Thank you Darek! I extracted the corresponding data from the dev set.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2Fb81456ca366cde6f650f694d1649ea87%2F2019-12-31%2017.28.33.png?generation=1577780979671866&amp;alt=media\" alt=\"\"></p>\n\n<h2>Facts</h2>\n\n<ul>\n<li>There can be multiple annotations</li>\n<li>In an annotation, there is one short_answers field. (= list)</li>\n<li>There can be multiple (start, end) pairs in short_answers field.</li>\n</ul>\n\n<h2>Scoring assumption</h2>\n\n<ul>\n<li>We predict only one short_answer for each question.</li>\n<li>The answer is considered correct if it matches one of  the short_answers</li>\n</ul>\n\n<h2>How to train?</h2>\n\n<p>We can do either\n- Choose one of the short_answers when training\n- Create multiple training instance for the short_answers\nor\n- Concatenate them as one answer? (As the code we found?)</p>\n\n<p>I'm not very sure which works better or worth doing as there are very few.</p>\n\n<h2>How to predict?</h2>\n\n<p>We just pick the highest score one. Any idea?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 714339,
          "author_name": "waalge",
          "author_url": "",
          "post_date": "2020-01-09T10:18:10.580000",
          "content": "<p>I basically wrote a full blog post length \"question\" about this but I've lost faith that that's particularly productive. </p>\n\n<h2>My uninformed take on this</h2>\n\n<p>If there are multiple short answer spans from the annotator, a correct answer must be precisely one single span. </p>\n\n<p>I think bert-joint was incentivized to return the smallest span covering all annotators short answer span (I wouldn't call this concatenating them). In which case those <code>[-1]</code> 's should <code>[0]</code>'s. \n(In the<code>run_nq.py</code> these appear twice <a href=\"https://github.com/google-research/language/blob/89e5b53a0e7f9f3e2da25a5da71ce5bd466acabd/language/question_answering/bert_joint/run_nq.py#L379\">here</a> and <a href=\"https://github.com/google-research/language/blob/89e5b53a0e7f9f3e2da25a5da71ce5bd466acabd/language/question_answering/bert_joint/run_nq.py#L257\">here</a>)</p>\n\n<p>Thanks <a href=\"/higepon\">@higepon</a> for this thread (and actually quite a few other of your contributions.) </p>\n\n<h2>Examples:</h2>\n\n<p>From the simplified train set, of the first 100 examples, 4 have been annotated multiple short spans (precisely: #45, #56, #74, #93). \nHere are the question numbers, questions, the annotators token spans for short answers and short answer text selected by the <code>run_nq.py</code> script:</p>\n\n<p><strong>45</strong> <code>where are the upcoming olympics to be held</code>\n<strong>short answer spans</strong> <code>[{'start_token': 204, 'end_token': 210}, {'start_token': 211, 'end_token': 217}, {'start_token': 218, 'end_token': 224}, {'start_token': 226, 'end_token': 233}]</code>\n<strong>text selected</strong> <code>Tokyo for the 2020 Summer Olympics , Beijing for the 2022 Winter Olympics , Paris for the 2024 Summer Olympics , and Los Angeles for the 2028 Summer Olympics</code>\n<strong>56</strong> <code>name some components of the central nervous system (cns)</code>\n<strong>short answer spans</strong> <code>[{'start_token': 140, 'end_token': 141}, {'start_token': 142, 'end_token': 144}]</code>\n<strong>text selected</strong> <code>brain and spinal cord</code>\n<strong>74</strong> <code>what teams are in the fa cup final</code>\n<strong>short answer spans</strong> <code>[{'start_token': 208, 'end_token': 210}, {'start_token': 211, 'end_token': 212}]</code>\n<strong>text selected</strong> <code>Manchester United and Chelsea</code>\n<strong>93</strong> <code>where was life or something like it filmed</code>\n<strong>short answer spans</strong> <code>[{'start_token': 1137, 'end_token': 1140}, {'start_token': 1145, 'end_token': 1147}]</code>\n<strong>text selected</strong> <code>Seattle , Washington although portions were filmed in downtown Vancouver</code></p>\n\n<p>By my assertion, in this competition (and in the original eval scripts, although these differ) \nhad the text selected been used to submit a prediction, then these would have all scored 0. \nFor this kaggle competition, a correct set of predictions is (a non-unique suggestion)</p>\n\n<p><strong>45</strong> <code>Tokyo for the 2020 Summer Olympics</code>\n<strong>56</strong> <code>brain</code>\n<strong>74</strong> <code>Manchester United</code>\n<strong>93</strong> <code>Seattle , Washington</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722360,
          "author_name": "lucaskg",
          "author_url": "",
          "post_date": "2020-01-18T13:19:00.933000",
          "content": "<p>Thanks! So this confirms that for multi-span short answers, each individual span is considered correct, but their merged version (text selected) is not. Right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 722600,
          "author_name": "Darek Kłeczek",
          "author_url": "",
          "post_date": "2020-01-18T19:31:51.600000",
          "content": "<p>That would be my guess... I haven't checked it though. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 703496,
      "author_name": "Luke",
      "author_url": "",
      "post_date": "2019-12-26T08:57:44.840000",
      "content": "<p>In my opinion, it indeed is a bad mixing</p>",
      "votes": 1,
      "replies": [
        {
          "id": 703499,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-26T09:04:36.930000",
          "content": "<p><a href=\"/mikelkl\">@mikelkl</a> thank you for your reply!.\nI guessed so too. But it's surprising if this is actually wrong considering the model is performing very well.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 703541,
          "author_name": "Luke",
          "author_url": "",
          "post_date": "2019-12-26T09:29:22.390000",
          "content": "<p>actually, i think the model performance still hv a long way to go, coz the precision of short answer is quite low, so u know where to find reason</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 703575,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-26T10:08:28.127000",
          "content": "<p>Yes. Thank you so much for your help!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 703489,
      "author_name": "HZD",
      "author_url": "",
      "post_date": "2019-12-26T08:46:23.990000",
      "content": "<p>There is only one ”short_answer“ in socalled \"short_answers\", so a[\"short_answers\"][0] == a[\"short_answers\"][-1]</p>",
      "votes": 2,
      "replies": [
        {
          "id": 703497,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-26T09:02:33.723000",
          "content": "<p>Thank you <a href=\"/noxuslol\">@noxuslol</a>! But <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/README.md#annotations\">https://github.com/google-research-datasets/natural-questions/blob/master/README.md#annotations</a> says there can be multiple short answers.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 703506,
          "author_name": "HZD",
          "author_url": "",
          "post_date": "2019-12-26T09:08:42.023000",
          "content": "<p>That is indeed the case in the test set, but you can check the training data(simplified-nq-train.jsonl) and there is really only one short answer...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 703515,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-26T09:15:12.940000",
          "content": "<p>Oh. Okay. Then this code works okay for training data, but looks really weird :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 703547,
          "author_name": "Luke",
          "author_url": "",
          "post_date": "2019-12-26T09:33:29.980000",
          "content": "<p>hi there, in my EDA, there is unignorable number of multi short answers in train data</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 703557,
          "author_name": "HZD",
          "author_url": "",
          "post_date": "2019-12-26T09:43:53.173000",
          "content": "<p>Are you sure there are multiple short answers, not multiple 'long_answer_candidates'...?\nIn the cell [9] of <a href=\"https://www.kaggle.com/ragnar123/exploratory-data-analysis-and-baseline\">this EDA kernel</a> we can see that there is only one short answer and one long answer...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 703574,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-26T10:08:03.473000",
          "content": "<p>Confirmed there are multiple short_answer(s).\nFor example:\n<code>\n\"short_answers\": [{\"start_token\": 208, \"end_token\": 210}, {\"start_token\": 211, \"end_token\": 212}], \"annotation_id\": 12688287890235717147}], \"document_url\": \"https://en.wikipedia.org//w/index.php?title=2018_FA_Cup_Final&amp;amp;oldid=842824259\", \"example_id\": 6823139846656798944}\n</code></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 703584,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-12-26T10:20:55.783000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 703591,
          "author_name": "HZD",
          "author_url": "",
          "post_date": "2019-12-26T10:28:18.043000",
          "content": "<p>I'm wrong...there are multiple short_answers</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 703781,
          "author_name": "cfiken",
          "author_url": "",
          "post_date": "2019-12-26T16:02:50.667000",
          "content": "<p>I counted short answer counts in <a href=\"https://www.kaggle.com/kentaronakanishi/eda-question-text-and-annotations-wip\">my dirty eda kernel version 5 (not latest)</a></p>\n\n<p><img src=\"https://www.kaggleusercontent.com/kf/23145600/eyJhbGciOiJkaXIiLCJlbmMiOiJBMTI4Q0JDLUhTMjU2In0..2MrnQySMMLxc779gOfnBTw.cpyFoPQO-yM86B7TPdFJdMf-ApNESmYbVHQhlPYCsG3Z5NWIQZHxJxcg3TUrJKPOrPSwsKoVk0y1rwdd98sF54WJRQPmRHuQWUbCAgk-6nIOFLa0RI2kjf4giFc_IFcOrDRJi2uJZ6n9UpMD4dX2fK2JV0rNU-u7UEj1bktddnO40HEdOMnARtI4lnAjT1z2Y78VA-B0Nnz9wl-Eo7ky5A.ZiTyvPX4y4qgBypj8UNRZw/__results___files/__results___39_2.png\" alt=\"\"></p>\n\n<p>There are some records which have multiple short answers, and the max number is 25.</p>\n\n<p>In fact, most of records have only 0 or 1 short answer. So I also think what <a href=\"/higepon\">@higepon</a> mentioned is a bug, but not critical to scores.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 704023,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "2019-12-27T00:08:49.650000",
          "content": "<p>Thanks for confirming <a href=\"/kentaronakanishi\">@kentaronakanishi</a> !</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 703501,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-26T09:06:56.987000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "703480": "Hi everyone,\nI'm writing my own PyTorch training kernel and reading through bert\\_util.py for reference.\nOne thing I don't understand is why get\\_first\\_annotation method choose end_token as follows.\n\n```\n  positive_annotations = sorted(\n      [a for a in e[\"annotations\"] if has_long_answer(a)],\n      key=lambda a: a[\"long_answer\"][\"candidate_index\"])\n\n  for a in positive_annotations:\n    if a[\"short_answers\"]:\n      idx = a[\"long_answer\"][\"candidate_index\"]\n      start_token = a[\"short_answers\"][0][\"start_token\"]\n      end_token = a[\"short_answers\"][-1][\"end_token\"]\n      return a, idx, (token_to_char_offset(e, idx, start_token),\n                      token_to_char_offset(e, idx, end_token) - 1)\n```\n\nWhy ```end_token = a[\"short_answers\"][-1][\"end_token\"]```?\n\n- Why they get end_token from other short_answer? (Specifically last short_answer?)\n- Isn't this make the model harder to predict exact short\\_answer because it mixing multiple short\\_answers into one?\n\nthanks,\nTaro",
    "706430": "I think that there is only one short answer in the train set, but it may consist of multiple short answer spans. In the test set, there may be multiple short answers (from 5 different annotators), and each of them may consist of multiple spans.  One of the possible improvements mentioned by the authors of the baseline model is to account for multiple short answer spans. ",
    "703496": "In my opinion, it indeed is a bad mixing",
    "703489": "There is only one ”short_answer“ in socalled \"short_answers\", so a[\"short_answers\"][0] == a[\"short_answers\"][-1]",
    "703501": ""
  }
}