{
  "id": 125675,
  "title": "Clarification on multispan short answers",
  "url": "/competitions/tensorflow2-question-answering/discussion/125675",
  "author_name": "",
  "post_date": "2020-01-12T18:32:55.331259700Z",
  "votes": 2,
  "comment_count": 8,
  "views": 0,
  "content": "<p>A similar question has been asked a couple of times here, but after reading all the answers I feel like a lot of people are confusing multiple short answers (which are present in the dev and test sets) and multiple spans within one short answer. I have not found any official answer regarding the latter, so I would be grateful if I could get some clarification about it.</p>\n\n<p>In the NQ dataset, one short answer (given by one annotator) can consist of multiple spans. Such answers are represented by several pairs of start/end indices. The paper about NQ dataset states that predictions for such answers will be considered correct only if all these spans are correctly identified. However, the possibility of multiple spans for short answers is never mentioned in the competition submission instructions.</p>\n\n<p>So, the question is: can we submit multiple spans as one short answer? If so, what is the correct format for such submission? If we cannot submit multiple spans, how will examples that have multispan as an answer be evaluated?</p>",
  "messages": [
    {
      "id": "717105",
      "postDate": "01/12/2020 18:32:55",
      "content": "<p>A similar question has been asked a couple of times here, but after reading all the answers I feel like a lot of people are confusing multiple short answers (which are present in the dev and test sets) and multiple spans within one short answer. I have not found any official answer regarding the latter, so I would be grateful if I could get some clarification about it.</p>\n\n<p>In the NQ dataset, one short answer (given by one annotator) can consist of multiple spans. Such answers are represented by several pairs of start/end indices. The paper about NQ dataset states that predictions for such answers will be considered correct only if all these spans are correctly identified. However, the possibility of multiple spans for short answers is never mentioned in the competition submission instructions.</p>\n\n<p>So, the question is: can we submit multiple spans as one short answer? If so, what is the correct format for such submission? If we cannot submit multiple spans, how will examples that have multispan as an answer be evaluated?</p>",
      "rawMarkdown": "A similar question has been asked a couple of times here, but after reading all the answers I feel like a lot of people are confusing multiple short answers (which are present in the dev and test sets) and multiple spans within one short answer. I have not found any official answer regarding the latter, so I would be grateful if I could get some clarification about it.\n\nIn the NQ dataset, one short answer (given by one annotator) can consist of multiple spans. Such answers are represented by several pairs of start/end indices. The paper about NQ dataset states that predictions for such answers will be considered correct only if all these spans are correctly identified. However, the possibility of multiple spans for short answers is never mentioned in the competition submission instructions.\n\nSo, the question is: can we submit multiple spans as one short answer? If so, what is the correct format for such submission? If we cannot submit multiple spans, how will examples that have multispan as an answer be evaluated?",
      "votes": null
    },
    {
      "id": "717475",
      "postDate": "01/13/2020 07:11:20",
      "content": "<p>This is a very good question, and I hope someone can give us a definitive answer. When I wrote an evaluation script atttempting to reproduce LB scores, I went by the description on the <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a> for this comp:\n&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels</p>\n\n<p>With short answers I made a flattened list of all annotations from all annotators. Then I checked to see if my prediction matched any of those. Using this method, I got an F1 for the dev set which was very close to my leaderboard score, so I think it is close to correct.</p>\n\n<p><code>\ndef short_annotations(example):\n    shorts = [('%s:%s' % (s['start_token'],s['end_token']))\n              for s in \n              sum([a['short_answers'] for a in example['annotations']], [])\n             ]\n    return shorts\n</code>\n<code>example</code> is <code>json.loads(line)</code> from a single line in the annotated dev set. </p>",
      "rawMarkdown": "This is a very good question, and I hope someone can give us a definitive answer. When I wrote an evaluation script atttempting to reproduce LB scores, I went by the description on the [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation) for this comp:\n&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels\n\nWith short answers I made a flattened list of all annotations from all annotators. Then I checked to see if my prediction matched any of those. Using this method, I got an F1 for the dev set which was very close to my leaderboard score, so I think it is close to correct.\n\n```\ndef short_annotations(example):\n    shorts = [('%s:%s' % (s['start_token'],s['end_token']))\n              for s in \n              sum([a['short_answers'] for a in example['annotations']], [])\n             ]\n    return shorts\n```\n`example` is `json.loads(line)` from a single line in the annotated dev set.",
      "votes": null
    },
    {
      "id": "717478",
      "postDate": "01/13/2020 07:16:06",
      "content": "<p>The sum of lists function <code>sum(list_of_lists, [])</code> flattens the short answer spans into a single list. I am still not sure that this is the right answer to your question, but its what I did and got similar score on the dev set to what the same model got on kaggle LB.</p>",
      "rawMarkdown": "The sum of lists function `sum(list_of_lists, [])` flattens the short answer spans into a single list. I am still not sure that this is the right answer to your question, but its what I did and got similar score on the dev set to what the same model got on kaggle LB.",
      "votes": null
    },
    {
      "id": "717479",
      "postDate": "01/13/2020 07:17:28",
      "content": "<p>If it would help, I can post a kernel of my evaluation script later today...</p>",
      "rawMarkdown": "If it would help, I can post a kernel of my evaluation script later today...",
      "votes": null
    },
    {
      "id": "717504",
      "postDate": "01/13/2020 08:01:24",
      "content": "<p>Would be very helpful <a href=\"/kenkrige\">@kenkrige</a>! Really grateful for the amount of help you provide.</p>",
      "rawMarkdown": "Would be very helpful @kenkrige! Really grateful for the amount of help you provide.",
      "votes": null
    },
    {
      "id": "717509",
      "postDate": "01/13/2020 08:11:58",
      "content": "<p>From the comment <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#681290\">here</a>, my understanding is that an answer is correct if it matches precisely one of any of the annotators short spans.</p>\n\n<p>If a short answer annotation is: \n\"10:20 31:34\"\nthen a correct answer \"10:20\" or \"31:34\", but \"10:34\" would be marked as wrong.</p>\n\n<p>I think this agrees with <a href=\"/kenkrige\">@kenkrige</a>. (?)</p>",
      "rawMarkdown": "From the comment [here](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#681290), my understanding is that an answer is correct if it matches precisely one of any of the annotators short spans.\n\nIf a short answer annotation is: \n\"10:20 31:34\"\nthen a correct answer \"10:20\" or \"31:34\", but \"10:34\" would be marked as wrong.\n\nI think this agrees with @kenkrige. (?)",
      "votes": null
    },
    {
      "id": "717577",
      "postDate": "01/13/2020 10:03:07",
      "content": "<p>Hope <a href=\"https://www.kaggle.com/kenkrige/possible-evaluation-metric\">this</a> helps. Let me know if you disagree with the metric algorithm.</p>",
      "rawMarkdown": "Hope [this](https://www.kaggle.com/kenkrige/possible-evaluation-metric) helps. Let me know if you disagree with the metric algorithm.",
      "votes": null
    },
    {
      "id": "717620",
      "postDate": "01/13/2020 11:11:25",
      "content": "<p>Hi! This is correct.</p>",
      "rawMarkdown": "Hi! This is correct.",
      "votes": null
    },
    {
      "id": "717627",
      "postDate": "01/13/2020 11:19:33",
      "content": "<p>In the discussion you've linked, it is unclear whether they are talking about multiple short answers from different annotators or one short answer that consists of multiple spans. I guess that the phrase \"multiple candidates for short answers\" is more likely to refer to multiple answers from different annotators, although I might be wrong. That's why it would be great if we could get official clarification about multispan short answers.</p>\n\n<p><a href=\"/philculliton\">@philculliton</a> could you help us, please?</p>\n\n<p>--\nNever mind, Phil answered while I was typing this. Thanks)</p>",
      "rawMarkdown": "In the discussion you've linked, it is unclear whether they are talking about multiple short answers from different annotators or one short answer that consists of multiple spans. I guess that the phrase \"multiple candidates for short answers\" is more likely to refer to multiple answers from different annotators, although I might be wrong. That's why it would be great if we could get official clarification about multispan short answers.\n\n@philculliton could you help us, please?\n\n\n--\nNever mind, Phil answered while I was typing this. Thanks)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 717475,
      "author_name": "kenkrige",
      "author_url": "",
      "post_date": "01/13/2020 07:11:20",
      "content": "<p>This is a very good question, and I hope someone can give us a definitive answer. When I wrote an evaluation script atttempting to reproduce LB scores, I went by the description on the <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a> for this comp:\n&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels</p>\n\n<p>With short answers I made a flattened list of all annotations from all annotators. Then I checked to see if my prediction matched any of those. Using this method, I got an F1 for the dev set which was very close to my leaderboard score, so I think it is close to correct.</p>\n\n<p><code>\ndef short_annotations(example):\n    shorts = [('%s:%s' % (s['start_token'],s['end_token']))\n              for s in \n              sum([a['short_answers'] for a in example['annotations']], [])\n             ]\n    return shorts\n</code>\n<code>example</code> is <code>json.loads(line)</code> from a single line in the annotated dev set. </p>",
      "votes": null,
      "replies": [
        {
          "id": 717478,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/13/2020 07:16:06",
          "content": "<p>The sum of lists function <code>sum(list_of_lists, [])</code> flattens the short answer spans into a single list. I am still not sure that this is the right answer to your question, but its what I did and got similar score on the dev set to what the same model got on kaggle LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 717479,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/13/2020 07:17:28",
          "content": "<p>If it would help, I can post a kernel of my evaluation script later today...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 717504,
          "author_name": "msheriey",
          "author_url": "",
          "post_date": "01/13/2020 08:01:24",
          "content": "<p>Would be very helpful <a href=\"/kenkrige\">@kenkrige</a>! Really grateful for the amount of help you provide.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 717577,
          "author_name": "kenkrige",
          "author_url": "",
          "post_date": "01/13/2020 10:03:07",
          "content": "<p>Hope <a href=\"https://www.kaggle.com/kenkrige/possible-evaluation-metric\">this</a> helps. Let me know if you disagree with the metric algorithm.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 717509,
      "author_name": "algywallis",
      "author_url": "",
      "post_date": "01/13/2020 08:11:58",
      "content": "<p>From the comment <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#681290\">here</a>, my understanding is that an answer is correct if it matches precisely one of any of the annotators short spans.</p>\n\n<p>If a short answer annotation is: \n\"10:20 31:34\"\nthen a correct answer \"10:20\" or \"31:34\", but \"10:34\" would be marked as wrong.</p>\n\n<p>I think this agrees with <a href=\"/kenkrige\">@kenkrige</a>. (?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 717620,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "01/13/2020 11:11:25",
          "content": "<p>Hi! This is correct.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 717627,
          "author_name": "olegplatonov",
          "author_url": "",
          "post_date": "01/13/2020 11:19:33",
          "content": "<p>In the discussion you've linked, it is unclear whether they are talking about multiple short answers from different annotators or one short answer that consists of multiple spans. I guess that the phrase \"multiple candidates for short answers\" is more likely to refer to multiple answers from different annotators, although I might be wrong. That's why it would be great if we could get official clarification about multispan short answers.</p>\n\n<p><a href=\"/philculliton\">@philculliton</a> could you help us, please?</p>\n\n<p>--\nNever mind, Phil answered while I was typing this. Thanks)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "717105": "A similar question has been asked a couple of times here, but after reading all the answers I feel like a lot of people are confusing multiple short answers (which are present in the dev and test sets) and multiple spans within one short answer. I have not found any official answer regarding the latter, so I would be grateful if I could get some clarification about it.\n\nIn the NQ dataset, one short answer (given by one annotator) can consist of multiple spans. Such answers are represented by several pairs of start/end indices. The paper about NQ dataset states that predictions for such answers will be considered correct only if all these spans are correctly identified. However, the possibility of multiple spans for short answers is never mentioned in the competition submission instructions.\n\nSo, the question is: can we submit multiple spans as one short answer? If so, what is the correct format for such submission? If we cannot submit multiple spans, how will examples that have multispan as an answer be evaluated?",
    "717475": "This is a very good question, and I hope someone can give us a definitive answer. When I wrote an evaluation script atttempting to reproduce LB scores, I went by the description on the [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation) for this comp:\n&gt; Predicted long and short answers must match exactly the token indices of one of the ground truth labels\n\nWith short answers I made a flattened list of all annotations from all annotators. Then I checked to see if my prediction matched any of those. Using this method, I got an F1 for the dev set which was very close to my leaderboard score, so I think it is close to correct.\n\n```\ndef short_annotations(example):\n    shorts = [('%s:%s' % (s['start_token'],s['end_token']))\n              for s in \n              sum([a['short_answers'] for a in example['annotations']], [])\n             ]\n    return shorts\n```\n`example` is `json.loads(line)` from a single line in the annotated dev set.",
    "717478": "The sum of lists function `sum(list_of_lists, [])` flattens the short answer spans into a single list. I am still not sure that this is the right answer to your question, but its what I did and got similar score on the dev set to what the same model got on kaggle LB.",
    "717479": "If it would help, I can post a kernel of my evaluation script later today...",
    "717504": "Would be very helpful @kenkrige! Really grateful for the amount of help you provide.",
    "717509": "From the comment [here](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#681290), my understanding is that an answer is correct if it matches precisely one of any of the annotators short spans.\n\nIf a short answer annotation is: \n\"10:20 31:34\"\nthen a correct answer \"10:20\" or \"31:34\", but \"10:34\" would be marked as wrong.\n\nI think this agrees with @kenkrige. (?)",
    "717577": "Hope [this](https://www.kaggle.com/kenkrige/possible-evaluation-metric) helps. Let me know if you disagree with the metric algorithm.",
    "717620": "Hi! This is correct.",
    "717627": "In the discussion you've linked, it is unclear whether they are talking about multiple short answers from different annotators or one short answer that consists of multiple spans. I guess that the phrase \"multiple candidates for short answers\" is more likely to refer to multiple answers from different annotators, although I might be wrong. That's why it would be great if we could get official clarification about multispan short answers.\n\n@philculliton could you help us, please?\n\n\n--\nNever mind, Phil answered while I was typing this. Thanks)"
  },
  "source": "meta"
}