{
  "id": 130059,
  "title": "How to evaluate the predictions (metrics)",
  "url": "/competitions/tensorflow2-question-answering/discussion/130059",
  "author_name": "Malasana",
  "post_date": "2020-02-11T22:42:24.999000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>I'm a bit lost as I'm taking on the original Google QA challenge for my own learning and pleasure and I was able to generate predictions using the pre-trained baseline BERT model as described here <a href=\"https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/\">https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/</a>. However, I'm not sure how to evaluate the predictions in terms of computing the F1, precision/recall scores that the original challenge is using for metrics. </p>\n\n<p>I want to use the <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">nq_eval.py</a> script but it appears that one should be run on the dev-annotated data which is very different then the kaggle-test data provided in the kaggle competition. In fact, when I try to run it on test data, it doesn't work because the example ids in the gold path annotations and the prediction files do no match up. Is it possible at all to generate your own official metric scores without submitting to the competition on Kaggle / Google?</p>\n\n<p>Any guidance or clarification of my misunderstanding of the data would be greatly appreciated!</p>",
  "messages": [
    {
      "id": 743230,
      "postDate": "2020-02-11T22:42:25Z",
      "content": "<p>I'm a bit lost as I'm taking on the original Google QA challenge for my own learning and pleasure and I was able to generate predictions using the pre-trained baseline BERT model as described here <a href=\"https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/\">https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/</a>. However, I'm not sure how to evaluate the predictions in terms of computing the F1, precision/recall scores that the original challenge is using for metrics. </p>\n\n<p>I want to use the <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">nq_eval.py</a> script but it appears that one should be run on the dev-annotated data which is very different then the kaggle-test data provided in the kaggle competition. In fact, when I try to run it on test data, it doesn't work because the example ids in the gold path annotations and the prediction files do no match up. Is it possible at all to generate your own official metric scores without submitting to the competition on Kaggle / Google?</p>\n\n<p>Any guidance or clarification of my misunderstanding of the data would be greatly appreciated!</p>",
      "rawMarkdown": "I'm a bit lost as I'm taking on the original Google QA challenge for my own learning and pleasure and I was able to generate predictions using the pre-trained baseline BERT model as described here [https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/](https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/). However, I'm not sure how to evaluate the predictions in terms of computing the F1, precision/recall scores that the original challenge is using for metrics. \n\nI want to use the [nq_eval.py](https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py) script but it appears that one should be run on the dev-annotated data which is very different then the kaggle-test data provided in the kaggle competition. In fact, when I try to run it on test data, it doesn't work because the example ids in the gold path annotations and the prediction files do no match up. Is it possible at all to generate your own official metric scores without submitting to the competition on Kaggle / Google?\n\nAny guidance or clarification of my misunderstanding of the data would be greatly appreciated!\n"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "743230": "I'm a bit lost as I'm taking on the original Google QA challenge for my own learning and pleasure and I was able to generate predictions using the pre-trained baseline BERT model as described here [https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/](https://www.kaggle.com/abhinand05/bert-for-humans-tutorial-baseline/). However, I'm not sure how to evaluate the predictions in terms of computing the F1, precision/recall scores that the original challenge is using for metrics. \n\nI want to use the [nq_eval.py](https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py) script but it appears that one should be run on the dev-annotated data which is very different then the kaggle-test data provided in the kaggle competition. In fact, when I try to run it on test data, it doesn't work because the example ids in the gold path annotations and the prediction files do no match up. Is it possible at all to generate your own official metric scores without submitting to the competition on Kaggle / Google?\n\nAny guidance or clarification of my misunderstanding of the data would be greatly appreciated!\n"
  }
}