{
  "id": 121666,
  "title": "still confused by the metric",
  "url": "/competitions/tensorflow2-question-answering/discussion/121666",
  "author_name": "Yih-Dar SHIEH",
  "post_date": "2019-12-14T16:53:08.802000",
  "votes": 4,
  "comment_count": 3,
  "views": 0,
  "content": "<p>In <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a> , it's written</p>\n\n<p><code>There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</code></p>\n\n<p>In <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030\">Dieter's implementation</a> or <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061\">a modified version by cfiken</a>, it looks like <code>long_label</code> is assumed to have ONLY one label, but <code>short_label</code> have multiple labels.</p>\n\n<p><strong>Question 1.</strong> What's the actual situation for long labels? <a href=\"/christofhenkel\">@christofhenkel</a> , <a href=\"/kentaronakanishi\">@kentaronakanishi</a> , <a href=\"/kashnitsky\">@kashnitsky</a> , <a href=\"/boliu0\">@boliu0</a>  , <a href=\"/zhaomeng1126\">@zhaomeng1126</a> , do you use multiple long labels for your local CV?</p>\n\n<p>Furthermore, when I use <a href=\"https://ai.google.com/research/NaturalQuestions/download\">v1.0-simplified_nq-dev-all.jsonl</a> -- in order to do CV, after using <code>simplify_nq_example</code> from <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/text_utils.py\">text_utils.py</a>, i see something like at the bottom of this post. There are multiple annotations, and each annotation has a single <code>long_answer</code> and multiple <code>short_answer</code>.</p>\n\n<p><strong>Question 2.</strong> Suppose we do have multiple long labels for this Kaggle competition. Are the <code>long_lables</code> and <code>short_labels</code> are somehow independent? Use the example below, suppose we give <code>long_prediction</code> as <code>22:341</code> which matches the <code>1st annotation</code>, but <code>short_prediction</code> as <code>351:353</code> which matches the 2nd <code>short_answer</code> in the <code>2nd annotation</code>, but not match the short_answer in the <code>1st annotation</code>. Does this count a correct prediction or not? </p>\n\n<p>I hope I don't make any stupid mistake here, but I am really confused by the metric. <a href=\"/philculliton\">@philculliton</a>, could you confirm clearly how is the metric calculated, including how annotations are converted to the actual labels used for this competition, please ....</p>\n\n<p>&gt;        \"question_text\": \"who wrote the song photograph by ringo starr\",\n        \"example_id\": -8366545547296627039, <br>\n        \"annotations\": [\n        {\n            \"annotation_id\": 647276088892962831,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 10083697748159579777,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 6354110109716954275,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 126,\n                    \"start_token\": 124\n                },\n                {\n                    \"end_token\": 129,\n                    \"start_token\": 127\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 7396078045607493655,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 9210770612718242890,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        }\n    ]</p>",
  "messages": [
    {
      "id": 695148,
      "postDate": "2019-12-14T16:53:08.803Z",
      "content": "<p>In <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation\">evaluation page</a> , it's written</p>\n\n<p><code>There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</code></p>\n\n<p>In <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030\">Dieter's implementation</a> or <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061\">a modified version by cfiken</a>, it looks like <code>long_label</code> is assumed to have ONLY one label, but <code>short_label</code> have multiple labels.</p>\n\n<p><strong>Question 1.</strong> What's the actual situation for long labels? <a href=\"/christofhenkel\">@christofhenkel</a> , <a href=\"/kentaronakanishi\">@kentaronakanishi</a> , <a href=\"/kashnitsky\">@kashnitsky</a> , <a href=\"/boliu0\">@boliu0</a>  , <a href=\"/zhaomeng1126\">@zhaomeng1126</a> , do you use multiple long labels for your local CV?</p>\n\n<p>Furthermore, when I use <a href=\"https://ai.google.com/research/NaturalQuestions/download\">v1.0-simplified_nq-dev-all.jsonl</a> -- in order to do CV, after using <code>simplify_nq_example</code> from <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/text_utils.py\">text_utils.py</a>, i see something like at the bottom of this post. There are multiple annotations, and each annotation has a single <code>long_answer</code> and multiple <code>short_answer</code>.</p>\n\n<p><strong>Question 2.</strong> Suppose we do have multiple long labels for this Kaggle competition. Are the <code>long_lables</code> and <code>short_labels</code> are somehow independent? Use the example below, suppose we give <code>long_prediction</code> as <code>22:341</code> which matches the <code>1st annotation</code>, but <code>short_prediction</code> as <code>351:353</code> which matches the 2nd <code>short_answer</code> in the <code>2nd annotation</code>, but not match the short_answer in the <code>1st annotation</code>. Does this count a correct prediction or not? </p>\n\n<p>I hope I don't make any stupid mistake here, but I am really confused by the metric. <a href=\"/philculliton\">@philculliton</a>, could you confirm clearly how is the metric calculated, including how annotations are converted to the actual labels used for this competition, please ....</p>\n\n<p>&gt;        \"question_text\": \"who wrote the song photograph by ringo starr\",\n        \"example_id\": -8366545547296627039, <br>\n        \"annotations\": [\n        {\n            \"annotation_id\": 647276088892962831,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 10083697748159579777,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 6354110109716954275,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 126,\n                    \"start_token\": 124\n                },\n                {\n                    \"end_token\": 129,\n                    \"start_token\": 127\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 7396078045607493655,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 9210770612718242890,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        }\n    ]</p>",
      "rawMarkdown": "In [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation) , it's written\n\n`There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.`\n\nIn [Dieter's implementation](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030) or [a modified version by cfiken](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061), it looks like `long_label` is assumed to have ONLY one label, but `short_label` have multiple labels.\n\n**Question 1.** What's the actual situation for long labels? @christofhenkel , @kentaronakanishi , @kashnitsky , @boliu0  , @zhaomeng1126 , do you use multiple long labels for your local CV?\n\nFurthermore, when I use [v1.0-simplified_nq-dev-all.jsonl](https://ai.google.com/research/NaturalQuestions/download) -- in order to do CV, after using `simplify_nq_example` from [text_utils.py](https://github.com/google-research-datasets/natural-questions/blob/master/text_utils.py), i see something like at the bottom of this post. There are multiple annotations, and each annotation has a single `long_answer` and multiple `short_answer`.\n\n**Question 2.** Suppose we do have multiple long labels for this Kaggle competition. Are the `long_lables` and `short_labels` are somehow independent? Use the example below, suppose we give `long_prediction` as `22:341` which matches the `1st annotation`, but `short_prediction` as `351:353` which matches the 2nd `short_answer` in the `2nd annotation`, but not match the short_answer in the `1st annotation`. Does this count a correct prediction or not? \n\nI hope I don't make any stupid mistake here, but I am really confused by the metric. @philculliton, could you confirm clearly how is the metric calculated, including how annotations are converted to the actual labels used for this competition, please ....\n\n&gt;        \"question_text\": \"who wrote the song photograph by ringo starr\",\n        \"example_id\": -8366545547296627039,  \n        \"annotations\": [\n        {\n            \"annotation_id\": 647276088892962831,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 10083697748159579777,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 6354110109716954275,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 126,\n                    \"start_token\": 124\n                },\n                {\n                    \"end_token\": 129,\n                    \"start_token\": 127\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 7396078045607493655,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 9210770612718242890,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        }\n    ]\n",
      "votes": 4
    },
    {
      "id": 695515,
      "postDate": "2019-12-15T09:18:18.293Z",
      "content": "<p>Question 1.\nThere is at most 1 long answer as for the training data, but more (at most five answers) for test data.\nSome of us are using the split of training data in order to compute CV, so the implementation you mentioned is assumed to have only 1 long answer (especially for my implementation).</p>\n\n<p>Multiple annotations for dev dataset is made by the process of Natural Questions. There are 5 annotator for a sample and each annotation by them is one of the <code>multiple annotation</code>.\nAbout Natural Questions, you can check <a href=\"https://ai.google/research/pubs/pub47761\">official paper</a> for detail.</p>\n\n<p>Question 2.\nI think it is computed independently but I don't have confidence.\n<code>independently</code> means that it is ok to say true positive if a predicted short answer is one of the gold short answer in all multiple annotations.</p>",
      "rawMarkdown": "Question 1.\nThere is at most 1 long answer as for the training data, but more (at most five answers) for test data.\nSome of us are using the split of training data in order to compute CV, so the implementation you mentioned is assumed to have only 1 long answer (especially for my implementation).\n\nMultiple annotations for dev dataset is made by the process of Natural Questions. There are 5 annotator for a sample and each annotation by them is one of the `multiple annotation`.\nAbout Natural Questions, you can check [official paper](https://ai.google/research/pubs/pub47761) for detail.\n\nQuestion 2.\nI think it is computed independently but I don't have confidence.\n`independently` means that it is ok to say true positive if a predicted short answer is one of the gold short answer in all multiple annotations.",
      "votes": 3,
      "replies": [
        {
          "id": 695565,
          "postDate": "2019-12-15T10:50:21.923Z",
          "content": "<p>Thanks for your clarification, especially for <code>Some of us are using the split of training data in order to compute CV</code></p>",
          "rawMarkdown": "Thanks for your clarification, especially for `Some of us are using the split of training data in order to compute CV`"
        }
      ]
    },
    {
      "id": 695429,
      "postDate": "2019-12-15T06:13:22.233Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 695515,
      "author_name": "cfiken",
      "author_url": "",
      "post_date": "2019-12-15T09:18:18.293000",
      "content": "<p>Question 1.\nThere is at most 1 long answer as for the training data, but more (at most five answers) for test data.\nSome of us are using the split of training data in order to compute CV, so the implementation you mentioned is assumed to have only 1 long answer (especially for my implementation).</p>\n\n<p>Multiple annotations for dev dataset is made by the process of Natural Questions. There are 5 annotator for a sample and each annotation by them is one of the <code>multiple annotation</code>.\nAbout Natural Questions, you can check <a href=\"https://ai.google/research/pubs/pub47761\">official paper</a> for detail.</p>\n\n<p>Question 2.\nI think it is computed independently but I don't have confidence.\n<code>independently</code> means that it is ok to say true positive if a predicted short answer is one of the gold short answer in all multiple annotations.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 695565,
          "author_name": "Yih-Dar SHIEH",
          "author_url": "",
          "post_date": "2019-12-15T10:50:21.923000",
          "content": "<p>Thanks for your clarification, especially for <code>Some of us are using the split of training data in order to compute CV</code></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 695429,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-15T06:13:22.233000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "695148": "In [evaluation page](https://www.kaggle.com/c/tensorflow2-question-answering/overview/evaluation) , it's written\n\n`There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.`\n\nIn [Dieter's implementation](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030) or [a modified version by cfiken](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120061), it looks like `long_label` is assumed to have ONLY one label, but `short_label` have multiple labels.\n\n**Question 1.** What's the actual situation for long labels? @christofhenkel , @kentaronakanishi , @kashnitsky , @boliu0  , @zhaomeng1126 , do you use multiple long labels for your local CV?\n\nFurthermore, when I use [v1.0-simplified_nq-dev-all.jsonl](https://ai.google.com/research/NaturalQuestions/download) -- in order to do CV, after using `simplify_nq_example` from [text_utils.py](https://github.com/google-research-datasets/natural-questions/blob/master/text_utils.py), i see something like at the bottom of this post. There are multiple annotations, and each annotation has a single `long_answer` and multiple `short_answer`.\n\n**Question 2.** Suppose we do have multiple long labels for this Kaggle competition. Are the `long_lables` and `short_labels` are somehow independent? Use the example below, suppose we give `long_prediction` as `22:341` which matches the `1st annotation`, but `short_prediction` as `351:353` which matches the 2nd `short_answer` in the `2nd annotation`, but not match the short_answer in the `1st annotation`. Does this count a correct prediction or not? \n\nI hope I don't make any stupid mistake here, but I am really confused by the metric. @philculliton, could you confirm clearly how is the metric calculated, including how annotations are converted to the actual labels used for this competition, please ....\n\n&gt;        \"question_text\": \"who wrote the song photograph by ringo starr\",\n        \"example_id\": -8366545547296627039,  \n        \"annotations\": [\n        {\n            \"annotation_id\": 647276088892962831,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 10083697748159579777,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 6354110109716954275,\n            \"long_answer\": {\n                \"candidate_index\": 0,\n                \"end_token\": 341,\n                \"start_token\": 22\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 126,\n                    \"start_token\": 124\n                },\n                {\n                    \"end_token\": 129,\n                    \"start_token\": 127\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 7396078045607493655,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [],\n            \"yes_no_answer\": \"NONE\"\n        },\n        {\n            \"annotation_id\": 9210770612718242890,\n            \"long_answer\": {\n                \"candidate_index\": 36,\n                \"end_token\": 478,\n                \"start_token\": 341\n            },\n            \"short_answers\": [\n                {\n                    \"end_token\": 353,\n                    \"start_token\": 351\n                },\n                {\n                    \"end_token\": 373,\n                    \"start_token\": 371\n                }\n            ],\n            \"yes_no_answer\": \"NONE\"\n        }\n    ]\n",
    "695515": "Question 1.\nThere is at most 1 long answer as for the training data, but more (at most five answers) for test data.\nSome of us are using the split of training data in order to compute CV, so the implementation you mentioned is assumed to have only 1 long answer (especially for my implementation).\n\nMultiple annotations for dev dataset is made by the process of Natural Questions. There are 5 annotator for a sample and each annotation by them is one of the `multiple annotation`.\nAbout Natural Questions, you can check [official paper](https://ai.google/research/pubs/pub47761) for detail.\n\nQuestion 2.\nI think it is computed independently but I don't have confidence.\n`independently` means that it is ok to say true positive if a predicted short answer is one of the gold short answer in all multiple annotations.",
    "695429": ""
  }
}