{
  "id": 126455,
  "title": "Question about long answer metrics",
  "url": "/competitions/tensorflow2-question-answering/discussion/126455",
  "author_name": "John cz",
  "post_date": "2020-01-17T15:15:20.092000",
  "votes": 1,
  "comment_count": 0,
  "views": 0,
  "content": "<p>hi, I am a little confused about the metrics here:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3272467%2F28069d8aaa38e46530205aa32f32f481%2F1.png?generation=1579273657819443&amp;alt=media\" alt=\"\"></p>\n\n<p>In original natural qa dataset, there are situations that some annotators think the question is answerable and give their answers while other annotators give null answer. In original nq-eval metrics, there is a setting that if &gt;= 2 of the annotators marked a non-null long answer, then the prediction must match any one of the non-null long answers to be considered correct. </p>\n\n<p>But in our metrics, do we have similar settings? Can I get score when I give null answer if just some of the annotator give the answer while others doesn't. Or should I make it blank if only one annotators answer the question.</p>\n\n<p>More specifically, how can I set the long answer and short answer thresholds when I am using Bert-joint baseline. I am a little confused about that. Thanks a lot for help.</p>",
  "messages": [
    {
      "id": 721649,
      "postDate": "2020-01-17T15:15:20.093Z",
      "content": "<p>hi, I am a little confused about the metrics here:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3272467%2F28069d8aaa38e46530205aa32f32f481%2F1.png?generation=1579273657819443&amp;alt=media\" alt=\"\"></p>\n\n<p>In original natural qa dataset, there are situations that some annotators think the question is answerable and give their answers while other annotators give null answer. In original nq-eval metrics, there is a setting that if &gt;= 2 of the annotators marked a non-null long answer, then the prediction must match any one of the non-null long answers to be considered correct. </p>\n\n<p>But in our metrics, do we have similar settings? Can I get score when I give null answer if just some of the annotator give the answer while others doesn't. Or should I make it blank if only one annotators answer the question.</p>\n\n<p>More specifically, how can I set the long answer and short answer thresholds when I am using Bert-joint baseline. I am a little confused about that. Thanks a lot for help.</p>",
      "rawMarkdown": "hi, I am a little confused about the metrics here:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3272467%2F28069d8aaa38e46530205aa32f32f481%2F1.png?generation=1579273657819443&amp;alt=media)\n\nIn original natural qa dataset, there are situations that some annotators think the question is answerable and give their answers while other annotators give null answer. In original nq-eval metrics, there is a setting that if &gt;= 2 of the annotators marked a non-null long answer, then the prediction must match any one of the non-null long answers to be considered correct. \n\nBut in our metrics, do we have similar settings? Can I get score when I give null answer if just some of the annotator give the answer while others doesn't. Or should I make it blank if only one annotators answer the question.\n\nMore specifically, how can I set the long answer and short answer thresholds when I am using Bert-joint baseline. I am a little confused about that. Thanks a lot for help.",
      "votes": 1
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "721649": "hi, I am a little confused about the metrics here:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3272467%2F28069d8aaa38e46530205aa32f32f481%2F1.png?generation=1579273657819443&amp;alt=media)\n\nIn original natural qa dataset, there are situations that some annotators think the question is answerable and give their answers while other annotators give null answer. In original nq-eval metrics, there is a setting that if &gt;= 2 of the annotators marked a non-null long answer, then the prediction must match any one of the non-null long answers to be considered correct. \n\nBut in our metrics, do we have similar settings? Can I get score when I give null answer if just some of the annotator give the answer while others doesn't. Or should I make it blank if only one annotators answer the question.\n\nMore specifically, how can I set the long answer and short answer thresholds when I am using Bert-joint baseline. I am a little confused about that. Thanks a lot for help."
  }
}