{
  "id": 118758,
  "title": "Please provide an official metric implementation or clarification",
  "url": "/competitions/tensorflow2-question-answering/discussion/118758",
  "author_name": "",
  "post_date": "2019-11-24T09:07:44.399247300Z",
  "votes": 21,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Dear organizers, could you please provide a detailed description of the metric, or an implementation of it? There is so much that is unclear about it IMO. I checked nq_eval but it's not clear which details are different compared to kaggle metric. Would be great to have an official clarification here, as metric affects model design.</p>\n\n<p>Evaluation page says \"There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\"\n- but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?\n- if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?\n- if we can predict multiple short answer spans, what is the format in submission file?</p>\n\n<p>If anyone has pointers to answers that I missed, please share.</p>",
  "messages": [
    {
      "id": "680224",
      "postDate": "11/24/2019 09:07:44",
      "content": "<p>Dear organizers, could you please provide a detailed description of the metric, or an implementation of it? There is so much that is unclear about it IMO. I checked nq_eval but it's not clear which details are different compared to kaggle metric. Would be great to have an official clarification here, as metric affects model design.</p>\n\n<p>Evaluation page says \"There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\"\n- but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?\n- if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?\n- if we can predict multiple short answer spans, what is the format in submission file?</p>\n\n<p>If anyone has pointers to answers that I missed, please share.</p>",
      "rawMarkdown": "Dear organizers, could you please provide a detailed description of the metric, or an implementation of it? There is so much that is unclear about it IMO. I checked nq_eval but it's not clear which details are different compared to kaggle metric. Would be great to have an official clarification here, as metric affects model design.\n\nEvaluation page says \"There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\"\n- but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?\n- if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?\n- if we can predict multiple short answer spans, what is the format in submission file?\n\nIf anyone has pointers to answers that I missed, please share.",
      "votes": null
    },
    {
      "id": "680229",
      "postDate": "11/24/2019 09:13:08",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a> I had same doubt as you for the third point <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#latest-680226\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#latest-680226</a></p>",
      "rawMarkdown": "lopuhin I had same doubt as you for the third point https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#latest-680226",
      "votes": null
    },
    {
      "id": "680447",
      "postDate": "11/24/2019 17:04:15",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a> just an update: I made a submission with multiple short answers <code>start_token:end_token</code> separated by space. For example, <code>start_token_c1:end_token_c1 start_token_c2:end_token_c2</code> for any example in the submission file which can have multiple candidates, it works. My kernel was evaluated and scored without the failure.</p>\n\n<p>I believe we can predict multiple short answer spans. </p>",
      "rawMarkdown": "lopuhin just an update: I made a submission with multiple short answers `start_token:end_token` separated by space. For example, `start_token_c1:end_token_c1 start_token_c2:end_token_c2` for any example in the submission file which can have multiple candidates, it works. My kernel was evaluated and scored without the failure.\n\nI believe we can predict multiple short answer spans.",
      "votes": null
    },
    {
      "id": "681289",
      "postDate": "11/26/2019 00:04:19",
      "content": "<p>Hi! Thanks for your questions.</p>\n\n<blockquote>\n  <p>but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?</p>\n</blockquote>\n\n<p>The test set has up to five correct long answers. You only need to identify one of them. If you correctly identify any of the possible answers, your answer is correct.</p>\n\n<blockquote>\n  <p>if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?</p>\n</blockquote>\n\n<p>You only need to predict a single short answer. In the same way as above, in the event there are multiple short answer ground truths, if you correctly identify one of them, you've got a correct answer.</p>\n\n<blockquote>\n  <p>if we can predict multiple short answer spans, what is the format in submission file?</p>\n</blockquote>\n\n<p>You cannot predict multiple short answer spans. Predicting multiple short answer spans will be scored as an incorrect answer.</p>",
      "rawMarkdown": "Hi! Thanks for your questions.\n\n&gt; but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?\n\nThe test set has up to five correct long answers. You only need to identify one of them. If you correctly identify any of the possible answers, your answer is correct.\n\n&gt; if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?\n\nYou only need to predict a single short answer. In the same way as above, in the event there are multiple short answer ground truths, if you correctly identify one of them, you've got a correct answer.\n\n&gt; if we can predict multiple short answer spans, what is the format in submission file?\n\nYou cannot predict multiple short answer spans. Predicting multiple short answer spans will be scored as an incorrect answer.",
      "votes": null
    },
    {
      "id": "681361",
      "postDate": "11/26/2019 02:29:58",
      "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , I have some questions too.</p>\n\n<p>Can you clarify these two points on the evaluation page?</p>\n\n<blockquote>\n  <p>The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.</p>\n</blockquote>\n\n<p>how is the F1 score defined? (how do you calculate accuracy and recall?) If the ground truth of a row is empty, and I made a prediction, is there penalty to the score? (seems like there isn't) \nhow are short and long answers combined?</p>",
      "rawMarkdown": "Hi @philculliton , I have some questions too.\n\nCan you clarify these two points on the evaluation page?\n&gt; The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.\n\nhow is the F1 score defined? (how do you calculate accuracy and recall?) If the ground truth of a row is empty, and I made a prediction, is there penalty to the score? (seems like there isn't) \nhow are short and long answers combined?",
      "votes": null
    },
    {
      "id": "681417",
      "postDate": "11/26/2019 04:58:38",
      "content": "<p>As explained by <a href=\"/philculliton\">@philculliton</a> I now understand why my submission scored less with multiple answers</p>",
      "rawMarkdown": "As explained by @philculliton I now understand why my submission scored less with multiple answers",
      "votes": null
    },
    {
      "id": "681440",
      "postDate": "11/26/2019 05:39:19",
      "content": "<p>Thanks for your answers, now metric looks clear.</p>",
      "rawMarkdown": "Thanks for your answers, now metric looks clear.",
      "votes": null
    },
    {
      "id": "681441",
      "postDate": "11/26/2019 05:46:35",
      "content": "<p><a href=\"/boliu0\">@boliu0</a> I think this reply answers your questions: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117182#673999\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117182#673999</a> - so F1 is calculated in a standard way, and if you made a prediction while the row is empty, there should be a penalty to the score. But rows where both prediction and ground truth are empty are excluded from the mean. Short and long answers are combined into one set of predictions for a row.</p>",
      "rawMarkdown": "boliu0 I think this reply answers your questions: https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117182#673999 - so F1 is calculated in a standard way, and if you made a prediction while the row is empty, there should be a penalty to the score. But rows where both prediction and ground truth are empty are excluded from the mean. Short and long answers are combined into one set of predictions for a row.",
      "votes": null
    },
    {
      "id": "681784",
      "postDate": "11/26/2019 14:47:00",
      "content": "<p>I don't think False Positive (gold is empty; made a prediction) has any penalty to the score. Through my submissions, I noticed that not setting any threshold (i.e. threshold=0) always gives better score than setting a threshold, for the same model.</p>\n\n<p>Also I just did this experiment: take my best model, \n- (A) set all long rows to be empty prediction;\n- (B) set all long rows to some clear wrong prediction <code>0:1</code></p>\n\n<p>It turns out (A) and (B) have exactly the same non-zero LB score, showing that all these <code>0:1</code> didn't have penalty on the score.</p>",
      "rawMarkdown": "I don't think False Positive (gold is empty; made a prediction) has any penalty to the score. Through my submissions, I noticed that not setting any threshold (i.e. threshold=0) always gives better score than setting a threshold, for the same model.\n\nAlso I just did this experiment: take my best model, \n- (A) set all long rows to be empty prediction;\n- (B) set all long rows to some clear wrong prediction `0:1`\n\nIt turns out (A) and (B) have exactly the same non-zero LB score, showing that all these `0:1` didn't have penalty on the score.",
      "votes": null
    },
    {
      "id": "682169",
      "postDate": "11/27/2019 03:08:48",
      "content": "<p>Hi <a href=\"/boliu0\">@boliu0</a> , thanks for your sharing. Is it possible that the test set only contain <strong>Answerable</strong> example ?  Beacause there is a gap between my CV and LB . </p>",
      "rawMarkdown": "Hi @boliu0 , thanks for your sharing. Is it possible that the test set only contain **Answerable** example ?  Beacause there is a gap between my CV and LB .",
      "votes": null
    },
    {
      "id": "682184",
      "postDate": "11/27/2019 03:49:56",
      "content": "<p>I thought about this possibility, but I don't think so. If all test set examples are answerable (i.e. no empty gold), and if I always make a prediction for every example, then <code>precision=recall=accuracy=F1</code> but the LB metric does not seem to be additive: overall score does not equal to short only score + long only score. If it's accuracy, it should be additive.</p>\n\n<p>P.S. I really hate these forensic investigation/LB probing/analysis just to understand the metric. I think Kaggle should either release evaluation code for all competitions (like many ML research competitions), or do a much better job describing the metrics on evaluation page. For simple ones like AUC or MSE, fine, everyone understands them, but complicated ones like this, it's really frustrating spending a lot of time guessing.</p>",
      "rawMarkdown": "I thought about this possibility, but I don't think so. If all test set examples are answerable (i.e. no empty gold), and if I always make a prediction for every example, then `precision=recall=accuracy=F1` but the LB metric does not seem to be additive: overall score does not equal to short only score + long only score. If it's accuracy, it should be additive.\n\nP.S. I really hate these forensic investigation/LB probing/analysis just to understand the metric. I think Kaggle should either release evaluation code for all competitions (like many ML research competitions), or do a much better job describing the metrics on evaluation page. For simple ones like AUC or MSE, fine, everyone understands them, but complicated ones like this, it's really frustrating spending a lot of time guessing.",
      "votes": null
    },
    {
      "id": "682210",
      "postDate": "11/27/2019 05:03:39",
      "content": "<p>I totally  agree. Its really frustrating to figure out how the metric is defined, and takes a lot of time that could have been spend to actually solve the problem at hand.</p>",
      "rawMarkdown": "I totally  agree. Its really frustrating to figure out how the metric is defined, and takes a lot of time that could have been spend to actually solve the problem at hand.",
      "votes": null
    },
    {
      "id": "682983",
      "postDate": "11/28/2019 01:54:30",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> I observed the same phenomenon <a href=\"/boliu0\">@boliu0</a> , Could you please clarify this issue ?</p>",
      "rawMarkdown": "philculliton I observed the same phenomenon @boliu0 , Could you please clarify this issue ?",
      "votes": null
    },
    {
      "id": "685963",
      "postDate": "12/02/2019 15:27:41",
      "content": "<p>Thanks for the feedback, <a href=\"/boliu0\">@boliu0</a> - I'm looking into this now.</p>",
      "rawMarkdown": "Thanks for the feedback, @boliu0 - I'm looking into this now.",
      "votes": null
    },
    {
      "id": "686854",
      "postDate": "12/03/2019 16:09:55",
      "content": "<p>Hi Phil, now that Dieter has reverse-engineered the LB metric, it confirmed my guess that FP is always 0. Please see my reply <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030#686853\">in the other thread</a></p>",
      "rawMarkdown": "Hi Phil, now that Dieter has reverse-engineered the LB metric, it confirmed my guess that FP is always 0. Please see my reply [in the other thread](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030#686853)",
      "votes": null
    },
    {
      "id": "686887",
      "postDate": "12/03/2019 16:41:39",
      "content": "<p>Saw that, thanks! Yes. I think I found an issue with the way we're processing submission files. I'll be working on that today and update soon.</p>",
      "rawMarkdown": "Saw that, thanks! Yes. I think I found an issue with the way we're processing submission files. I'll be working on that today and update soon.",
      "votes": null
    },
    {
      "id": "686916",
      "postDate": "12/03/2019 17:23:04",
      "content": "<p>Would be nice not to change metric, now that I spent two weeks to re-engineer it.</p>",
      "rawMarkdown": "Would be nice not to change metric, now that I spent two weeks to re-engineer it.",
      "votes": null
    },
    {
      "id": "689312",
      "postDate": "12/06/2019 19:10:02",
      "content": "<p>Hi Phil, is there a plan to fix the FP issue?</p>\n\n<p>Dieter, I think we should still use the correct metric. Your work on replicating the wrong metric shall not be wasted -- it has helped Phil pinpoint the problem.</p>\n\n<p>The reason I support the correct F1 metric is that, not only it's the official metric, but it's also more stable since it covers both empty and non-empty examples. In my local experiments so far, this is indeed the case. Also, with the current metric, all the empty examples are ignored.</p>",
      "rawMarkdown": "Hi Phil, is there a plan to fix the FP issue?\n\nDieter, I think we should still use the correct metric. Your work on replicating the wrong metric shall not be wasted -- it has helped Phil pinpoint the problem.\n\nThe reason I support the correct F1 metric is that, not only it's the official metric, but it's also more stable since it covers both empty and non-empty examples. In my local experiments so far, this is indeed the case. Also, with the current metric, all the empty examples are ignored.",
      "votes": null
    },
    {
      "id": "689334",
      "postDate": "12/06/2019 19:39:48",
      "content": "<p>On the other hand, the current metric has the advantage, that you don't need to spend time LB probing to find a good threshold for setting blanks</p>",
      "rawMarkdown": "On the other hand, the current metric has the advantage, that you don't need to spend time LB probing to find a good threshold for setting blanks",
      "votes": null
    },
    {
      "id": "689383",
      "postDate": "12/06/2019 21:43:23",
      "content": "<p>Hi Bo - yep. Sorry it's taken a bit of time. I've been working through it with our back-end engineers. We plan to finish prepping and rescore on Monday.</p>\n\n<p>Re: helping to pinpoint the problem, absolutely. It is very much appreciated.</p>",
      "rawMarkdown": "Hi Bo - yep. Sorry it's taken a bit of time. I've been working through it with our back-end engineers. We plan to finish prepping and rescore on Monday.\n\nRe: helping to pinpoint the problem, absolutely. It is very much appreciated.",
      "votes": null
    },
    {
      "id": "689386",
      "postDate": "12/06/2019 21:49:06",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> if you want to express your gratitude with some kaggle swag, I won't stop you. :D</p>",
      "rawMarkdown": "philculliton if you want to express your gratitude with some kaggle swag, I won't stop you. :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 680229,
      "author_name": "axel81",
      "author_url": "",
      "post_date": "11/24/2019 09:13:08",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a> I had same doubt as you for the third point <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#latest-680226\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#latest-680226</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 680447,
      "author_name": "axel81",
      "author_url": "",
      "post_date": "11/24/2019 17:04:15",
      "content": "<p><a href=\"/lopuhin\">@lopuhin</a> just an update: I made a submission with multiple short answers <code>start_token:end_token</code> separated by space. For example, <code>start_token_c1:end_token_c1 start_token_c2:end_token_c2</code> for any example in the submission file which can have multiple candidates, it works. My kernel was evaluated and scored without the failure.</p>\n\n<p>I believe we can predict multiple short answer spans. </p>",
      "votes": null,
      "replies": [
        {
          "id": 681417,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "11/26/2019 04:58:38",
          "content": "<p>As explained by <a href=\"/philculliton\">@philculliton</a> I now understand why my submission scored less with multiple answers</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 681289,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "11/26/2019 00:04:19",
      "content": "<p>Hi! Thanks for your questions.</p>\n\n<blockquote>\n  <p>but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?</p>\n</blockquote>\n\n<p>The test set has up to five correct long answers. You only need to identify one of them. If you correctly identify any of the possible answers, your answer is correct.</p>\n\n<blockquote>\n  <p>if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?</p>\n</blockquote>\n\n<p>You only need to predict a single short answer. In the same way as above, in the event there are multiple short answer ground truths, if you correctly identify one of them, you've got a correct answer.</p>\n\n<blockquote>\n  <p>if we can predict multiple short answer spans, what is the format in submission file?</p>\n</blockquote>\n\n<p>You cannot predict multiple short answer spans. Predicting multiple short answer spans will be scored as an incorrect answer.</p>",
      "votes": null,
      "replies": [
        {
          "id": 681361,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "11/26/2019 02:29:58",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a> , I have some questions too.</p>\n\n<p>Can you clarify these two points on the evaluation page?</p>\n\n<blockquote>\n  <p>The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.</p>\n</blockquote>\n\n<p>how is the F1 score defined? (how do you calculate accuracy and recall?) If the ground truth of a row is empty, and I made a prediction, is there penalty to the score? (seems like there isn't) \nhow are short and long answers combined?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681440,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "11/26/2019 05:39:19",
          "content": "<p>Thanks for your answers, now metric looks clear.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681441,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "11/26/2019 05:46:35",
          "content": "<p><a href=\"/boliu0\">@boliu0</a> I think this reply answers your questions: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117182#673999\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117182#673999</a> - so F1 is calculated in a standard way, and if you made a prediction while the row is empty, there should be a penalty to the score. But rows where both prediction and ground truth are empty are excluded from the mean. Short and long answers are combined into one set of predictions for a row.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 681784,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "11/26/2019 14:47:00",
          "content": "<p>I don't think False Positive (gold is empty; made a prediction) has any penalty to the score. Through my submissions, I noticed that not setting any threshold (i.e. threshold=0) always gives better score than setting a threshold, for the same model.</p>\n\n<p>Also I just did this experiment: take my best model, \n- (A) set all long rows to be empty prediction;\n- (B) set all long rows to some clear wrong prediction <code>0:1</code></p>\n\n<p>It turns out (A) and (B) have exactly the same non-zero LB score, showing that all these <code>0:1</code> didn't have penalty on the score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682169,
          "author_name": "wangyunazx",
          "author_url": "",
          "post_date": "11/27/2019 03:08:48",
          "content": "<p>Hi <a href=\"/boliu0\">@boliu0</a> , thanks for your sharing. Is it possible that the test set only contain <strong>Answerable</strong> example ?  Beacause there is a gap between my CV and LB . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682184,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "11/27/2019 03:49:56",
          "content": "<p>I thought about this possibility, but I don't think so. If all test set examples are answerable (i.e. no empty gold), and if I always make a prediction for every example, then <code>precision=recall=accuracy=F1</code> but the LB metric does not seem to be additive: overall score does not equal to short only score + long only score. If it's accuracy, it should be additive.</p>\n\n<p>P.S. I really hate these forensic investigation/LB probing/analysis just to understand the metric. I think Kaggle should either release evaluation code for all competitions (like many ML research competitions), or do a much better job describing the metrics on evaluation page. For simple ones like AUC or MSE, fine, everyone understands them, but complicated ones like this, it's really frustrating spending a lot of time guessing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682210,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "11/27/2019 05:03:39",
          "content": "<p>I totally  agree. Its really frustrating to figure out how the metric is defined, and takes a lot of time that could have been spend to actually solve the problem at hand.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 682983,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "11/28/2019 01:54:30",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> I observed the same phenomenon <a href=\"/boliu0\">@boliu0</a> , Could you please clarify this issue ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685963,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/02/2019 15:27:41",
          "content": "<p>Thanks for the feedback, <a href=\"/boliu0\">@boliu0</a> - I'm looking into this now.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686854,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "12/03/2019 16:09:55",
          "content": "<p>Hi Phil, now that Dieter has reverse-engineered the LB metric, it confirmed my guess that FP is always 0. Please see my reply <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030#686853\">in the other thread</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686887,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/03/2019 16:41:39",
          "content": "<p>Saw that, thanks! Yes. I think I found an issue with the way we're processing submission files. I'll be working on that today and update soon.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686916,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 17:23:04",
          "content": "<p>Would be nice not to change metric, now that I spent two weeks to re-engineer it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 689312,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "12/06/2019 19:10:02",
          "content": "<p>Hi Phil, is there a plan to fix the FP issue?</p>\n\n<p>Dieter, I think we should still use the correct metric. Your work on replicating the wrong metric shall not be wasted -- it has helped Phil pinpoint the problem.</p>\n\n<p>The reason I support the correct F1 metric is that, not only it's the official metric, but it's also more stable since it covers both empty and non-empty examples. In my local experiments so far, this is indeed the case. Also, with the current metric, all the empty examples are ignored.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 689334,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/06/2019 19:39:48",
          "content": "<p>On the other hand, the current metric has the advantage, that you don't need to spend time LB probing to find a good threshold for setting blanks</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 689383,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/06/2019 21:43:23",
          "content": "<p>Hi Bo - yep. Sorry it's taken a bit of time. I've been working through it with our back-end engineers. We plan to finish prepping and rescore on Monday.</p>\n\n<p>Re: helping to pinpoint the problem, absolutely. It is very much appreciated.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 689386,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/06/2019 21:49:06",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> if you want to express your gratitude with some kaggle swag, I won't stop you. :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "680224": "Dear organizers, could you please provide a detailed description of the metric, or an implementation of it? There is so much that is unclear about it IMO. I checked nq_eval but it's not clear which details are different compared to kaggle metric. Would be great to have an official clarification here, as metric affects model design.\n\nEvaluation page says \"There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\"\n- but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?\n- if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?\n- if we can predict multiple short answer spans, what is the format in submission file?\n\nIf anyone has pointers to answers that I missed, please share.",
    "680229": "lopuhin I had same doubt as you for the third point https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118759#latest-680226",
    "680447": "lopuhin just an update: I made a submission with multiple short answers `start_token:end_token` separated by space. For example, `start_token_c1:end_token_c1 start_token_c2:end_token_c2` for any example in the submission file which can have multiple candidates, it works. My kernel was evaluated and scored without the failure.\n\nI believe we can predict multiple short answer spans.",
    "681289": "Hi! Thanks for your questions.\n\n&gt; but in provided annotations I see only one label for long answer, is it different in the test set? If yes, how are they scored?\n\nThe test set has up to five correct long answers. You only need to identify one of them. If you correctly identify any of the possible answers, your answer is correct.\n\n&gt; if ground truth contains multiple short answer spans, do we need to predict multiple short answer spans? Do they need to match all, or a partial match counts, or it's enough to get a match of any of the spans?\n\nYou only need to predict a single short answer. In the same way as above, in the event there are multiple short answer ground truths, if you correctly identify one of them, you've got a correct answer.\n\n&gt; if we can predict multiple short answer spans, what is the format in submission file?\n\nYou cannot predict multiple short answer spans. Predicting multiple short answer spans will be scored as an incorrect answer.",
    "681361": "Hi @philculliton , I have some questions too.\n\nCan you clarify these two points on the evaluation page?\n&gt; The metric in this competition diverges from the original metric in two key respects: 1) short and long answer formats do not receive separate scores, but are instead combined into a micro F1 score across both formats, and 2) this competition's metric does not use confidence scores to find an optimal threshold for predictions.\n\nhow is the F1 score defined? (how do you calculate accuracy and recall?) If the ground truth of a row is empty, and I made a prediction, is there penalty to the score? (seems like there isn't) \nhow are short and long answers combined?",
    "681417": "As explained by @philculliton I now understand why my submission scored less with multiple answers",
    "681440": "Thanks for your answers, now metric looks clear.",
    "681441": "boliu0 I think this reply answers your questions: https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117182#673999 - so F1 is calculated in a standard way, and if you made a prediction while the row is empty, there should be a penalty to the score. But rows where both prediction and ground truth are empty are excluded from the mean. Short and long answers are combined into one set of predictions for a row.",
    "681784": "I don't think False Positive (gold is empty; made a prediction) has any penalty to the score. Through my submissions, I noticed that not setting any threshold (i.e. threshold=0) always gives better score than setting a threshold, for the same model.\n\nAlso I just did this experiment: take my best model, \n- (A) set all long rows to be empty prediction;\n- (B) set all long rows to some clear wrong prediction `0:1`\n\nIt turns out (A) and (B) have exactly the same non-zero LB score, showing that all these `0:1` didn't have penalty on the score.",
    "682169": "Hi @boliu0 , thanks for your sharing. Is it possible that the test set only contain **Answerable** example ?  Beacause there is a gap between my CV and LB .",
    "682184": "I thought about this possibility, but I don't think so. If all test set examples are answerable (i.e. no empty gold), and if I always make a prediction for every example, then `precision=recall=accuracy=F1` but the LB metric does not seem to be additive: overall score does not equal to short only score + long only score. If it's accuracy, it should be additive.\n\nP.S. I really hate these forensic investigation/LB probing/analysis just to understand the metric. I think Kaggle should either release evaluation code for all competitions (like many ML research competitions), or do a much better job describing the metrics on evaluation page. For simple ones like AUC or MSE, fine, everyone understands them, but complicated ones like this, it's really frustrating spending a lot of time guessing.",
    "682210": "I totally  agree. Its really frustrating to figure out how the metric is defined, and takes a lot of time that could have been spend to actually solve the problem at hand.",
    "682983": "philculliton I observed the same phenomenon @boliu0 , Could you please clarify this issue ?",
    "685963": "Thanks for the feedback, @boliu0 - I'm looking into this now.",
    "686854": "Hi Phil, now that Dieter has reverse-engineered the LB metric, it confirmed my guess that FP is always 0. Please see my reply [in the other thread](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120030#686853)",
    "686887": "Saw that, thanks! Yes. I think I found an issue with the way we're processing submission files. I'll be working on that today and update soon.",
    "686916": "Would be nice not to change metric, now that I spent two weeks to re-engineer it.",
    "689312": "Hi Phil, is there a plan to fix the FP issue?\n\nDieter, I think we should still use the correct metric. Your work on replicating the wrong metric shall not be wasted -- it has helped Phil pinpoint the problem.\n\nThe reason I support the correct F1 metric is that, not only it's the official metric, but it's also more stable since it covers both empty and non-empty examples. In my local experiments so far, this is indeed the case. Also, with the current metric, all the empty examples are ignored.",
    "689334": "On the other hand, the current metric has the advantage, that you don't need to spend time LB probing to find a good threshold for setting blanks",
    "689383": "Hi Bo - yep. Sorry it's taken a bit of time. I've been working through it with our back-end engineers. We plan to finish prepping and rescore on Monday.\n\nRe: helping to pinpoint the problem, absolutely. It is very much appreciated.",
    "689386": "philculliton if you want to express your gratitude with some kaggle swag, I won't stop you. :D"
  },
  "source": "meta"
}