{
  "id": 119606,
  "title": "What is the metric",
  "url": "/competitions/tensorflow2-question-answering/discussion/119606",
  "author_name": "",
  "post_date": "2019-11-29T23:34:28.656084900Z",
  "votes": 5,
  "comment_count": 15,
  "views": 0,
  "content": "<p>I'm considering entering this competition, and I am not sure I understand what we are supposed to optimize.</p>\n\n<p>The evaluation page points to Wiikipedia page on F1 score, but says \"Micro F1\" score.  Micro F1 score is not described in this page unless mistaken.</p>\n\n<p>Then another link points to <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py</a> , calling it the original metric, then says the metric here differs from the original metric.</p>\n\n<p>I see what the metric is not, but at no point do I see a description of what the metric is. </p>\n\n<p>Can someone point me to what the metric is?  I would expect a formula or an algorithm computing it.</p>\n\n<p>What am I missing?</p>",
  "messages": [
    {
      "id": "684508",
      "postDate": "11/29/2019 23:34:28",
      "content": "<p>I'm considering entering this competition, and I am not sure I understand what we are supposed to optimize.</p>\n\n<p>The evaluation page points to Wiikipedia page on F1 score, but says \"Micro F1\" score.  Micro F1 score is not described in this page unless mistaken.</p>\n\n<p>Then another link points to <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py</a> , calling it the original metric, then says the metric here differs from the original metric.</p>\n\n<p>I see what the metric is not, but at no point do I see a description of what the metric is. </p>\n\n<p>Can someone point me to what the metric is?  I would expect a formula or an algorithm computing it.</p>\n\n<p>What am I missing?</p>",
      "rawMarkdown": "I'm considering entering this competition, and I am not sure I understand what we are supposed to optimize.\n\nThe evaluation page points to Wiikipedia page on F1 score, but says \"Micro F1\" score.  Micro F1 score is not described in this page unless mistaken.\n\nThen another link points to https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py , calling it the original metric, then says the metric here differs from the original metric.\n\nI see what the metric is not, but at no point do I see a description of what the metric is. \n\nCan someone point me to what the metric is?  I would expect a formula or an algorithm computing it.\n\nWhat am I missing?",
      "votes": null
    },
    {
      "id": "684564",
      "postDate": "11/30/2019 02:41:29",
      "content": "<p>The metric is to be predicted also in this competition 😫</p>",
      "rawMarkdown": "The metric is to be predicted also in this competition 😫",
      "votes": null
    },
    {
      "id": "684765",
      "postDate": "11/30/2019 12:49:16",
      "content": "<p>The metric is indeed based on <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py</a> with 2 differences:\n - long and short answers are nor scored independently, instead, common F1 is calculated for both long and short answers\n - no optimal thresholds are found, so the code shall be simpler that those of <code>nq_eval</code></p>\n\n<p>I haven't yet implemented the metric (so far using <code>nq_eval</code> ) but Dieter <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119462#684706\">can see</a> correlation between his local validation and LB.  </p>\n\n<p>So either we wait till Phil shares the official metric implementation (it's in C# though) or we go on with reverse-engineering. </p>",
      "rawMarkdown": "The metric is indeed based on https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py with 2 differences:\n - long and short answers are nor scored independently, instead, common F1 is calculated for both long and short answers\n - no optimal thresholds are found, so the code shall be simpler that those of `nq_eval`\n\nI haven't yet implemented the metric (so far using `nq_eval` ) but Dieter [can see](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119462#684706) correlation between his local validation and LB.  \n\nSo either we wait till Phil shares the official metric implementation (it's in C# though) or we go on with reverse-engineering.",
      "votes": null
    },
    {
      "id": "684781",
      "postDate": "11/30/2019 13:07:46",
      "content": "<p>My suggestion would be to try some metric learning. Just kidding, hehe ;)</p>\n\n<p>Seriously, though, I'm confused as well, it seems that the metric understanding is probably even more tough task than the actual task! Hope Kaggle will release the metric soon...</p>",
      "rawMarkdown": "My suggestion would be to try some metric learning. Just kidding, hehe ;)\n\nSeriously, though, I'm confused as well, it seems that the metric understanding is probably even more tough task than the actual task! Hope Kaggle will release the metric soon...",
      "votes": null
    },
    {
      "id": "685236",
      "postDate": "12/01/2019 08:34:15",
      "content": "<p>Thanks, I read the description you quote, but it says what the metric is no, it does not say what the metric is.</p>\n\n<p>Anyway, I now see that I'm not the first to wonder about this, and I'm sorry people had to spend time like this: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119143\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119143</a></p>",
      "rawMarkdown": "Thanks, I read the description you quote, but it says what the metric is no, it does not say what the metric is.\n\nAnyway, I now see that I'm not the first to wonder about this, and I'm sorry people had to spend time like this: https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119143",
      "votes": null
    },
    {
      "id": "685237",
      "postDate": "12/01/2019 08:35:12",
      "content": "<p>Yes, not having  a metric description is totally unbelievable.  </p>",
      "rawMarkdown": "Yes, not having  a metric description is totally unbelievable.",
      "votes": null
    },
    {
      "id": "685361",
      "postDate": "12/01/2019 13:58:43",
      "content": "<p>Well, it also says smth about what it <em>is</em>, but indeed you need to study and adapt the source code of <code>nq_eval</code> and, moreover, analyze different peculiarities of the metric mentioned here on the forum.  </p>",
      "rawMarkdown": "Well, it also says smth about what it _is_, but indeed you need to study and adapt the source code of `nq_eval` and, moreover, analyze different peculiarities of the metric mentioned here on the forum.",
      "votes": null
    },
    {
      "id": "685451",
      "postDate": "12/01/2019 17:57:16",
      "content": "<p>I think the simple point is to use long_answer and short_answer to calculate the F-score together, and then do not set any answer to null. For me,  the cv is <code>0.672(for long)/0.548(for short)/0.623(for all)</code> with bert_joint. But there are still some bugs in my inference, so at this stage I can only say that these are some of my guesses😂</p>",
      "rawMarkdown": "I think the simple point is to use long_answer and short_answer to calculate the F-score together, and then do not set any answer to null. For me,  the cv is `0.672(for long)/0.548(for short)/0.623(for all)` with bert_joint. But there are still some bugs in my inference, so at this stage I can only say that these are some of my guesses😂",
      "votes": null
    },
    {
      "id": "685933",
      "postDate": "12/02/2019 14:44:41",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> - Thanks for the feedback! I've added a link to the sklearn page describing \"micro F1\".</p>\n\n<p>The <code>nq_eval.py</code> script is the metric, with the differences listed in the Evaluation page. Our C# metric is unfortunately not a good candidate for clearing things up (several of the operations are being performed outside of the main code).</p>",
      "rawMarkdown": "cpmpml - Thanks for the feedback! I've added a link to the sklearn page describing \"micro F1\".\n\nThe `nq_eval.py` script is the metric, with the differences listed in the Evaluation page. Our C# metric is unfortunately not a good candidate for clearing things up (several of the operations are being performed outside of the main code).",
      "votes": null
    },
    {
      "id": "685936",
      "postDate": "12/02/2019 14:46:18",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> - that's exactly correct.</p>",
      "rawMarkdown": "kashnitsky - that's exactly correct.",
      "votes": null
    },
    {
      "id": "685993",
      "postDate": "12/02/2019 16:04:45",
      "content": "<p>That is probably accurate, but I don't think it is correct.  Asking participants to devise the metric code is not very useful for advancing the quality of models.  It would have been much better to provide the modified python code directly.  Anyway, i'll do like every other, work first on a metric code instead of models.</p>",
      "rawMarkdown": "That is probably accurate, but I don't think it is correct.  Asking participants to devise the metric code is not very useful for advancing the quality of models.  It would have been much better to provide the modified python code directly.  Anyway, i'll do like every other, work first on a metric code instead of models.",
      "votes": null
    },
    {
      "id": "685995",
      "postDate": "12/02/2019 16:08:55",
      "content": "<p>Releasing even C# code would also suffice, there enough qualified people here to check whether the C# code is doing what's expected. </p>",
      "rawMarkdown": "Releasing even C# code would also suffice, there enough qualified people here to check whether the C# code is doing what's expected.",
      "votes": null
    },
    {
      "id": "685997",
      "postDate": "12/02/2019 16:10:35",
      "content": "<p>To speculate: either the evaluation code contains some obscure parts that are not intended to be made public. Or reverse-engineering the metric is considered to be a part of the competition.</p>",
      "rawMarkdown": "To speculate: either the evaluation code contains some obscure parts that are not intended to be made public. Or reverse-engineering the metric is considered to be a part of the competition.",
      "votes": null
    },
    {
      "id": "685999",
      "postDate": "12/02/2019 16:13:36",
      "content": "<p>Ah, ok, now I see Phil's comment below :) If the evaluation code is too involved, with source code spread over several files, then indeed I'd expect organizers to have clearly described the metric on Evaluation page. At least providing some toy examples covering all peculiarities of evaluation. </p>",
      "rawMarkdown": "Ah, ok, now I see Phil's comment below :) If the evaluation code is too involved, with source code spread over several files, then indeed I'd expect organizers to have clearly described the metric on Evaluation page. At least providing some toy examples covering all peculiarities of evaluation.",
      "votes": null
    },
    {
      "id": "686358",
      "postDate": "12/03/2019 04:05:42",
      "content": "<p>Thanks for all the feedback! I've updated the description to add more detail about the metric. Hopefully this helps. Sorry about that - I felt that <code>nq_eval.py</code> + addendums covered the metric fairly well (and also the odd corner cases thereof) but that was incorrect. Please let me know if additional detail or information is needed!</p>",
      "rawMarkdown": "Thanks for all the feedback! I've updated the description to add more detail about the metric. Hopefully this helps. Sorry about that - I felt that `nq_eval.py` + addendums covered the metric fairly well (and also the odd corner cases thereof) but that was incorrect. Please let me know if additional detail or information is needed!",
      "votes": null
    },
    {
      "id": "686518",
      "postDate": "12/03/2019 08:40:56",
      "content": "<p>Thanks, Phil! It's definitely better now. There're still some strange quirks about the metric (for instance, <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#681784\">this one</a> reported by Bo), and we have to spend time and submissions to understand why this or that happens.\nI'm almost fine with the idea of treating this as part of the competition :)\nBut no. In future, for other competitions, pls try to prepare official metric implementation, so that competitors spend time on what actually matters. </p>",
      "rawMarkdown": "Thanks, Phil! It's definitely better now. There're still some strange quirks about the metric (for instance, [this one](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#681784) reported by Bo), and we have to spend time and submissions to understand why this or that happens.\nI'm almost fine with the idea of treating this as part of the competition :)\nBut no. In future, for other competitions, pls try to prepare official metric implementation, so that competitors spend time on what actually matters.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 684564,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "11/30/2019 02:41:29",
      "content": "<p>The metric is to be predicted also in this competition 😫</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 684765,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "11/30/2019 12:49:16",
      "content": "<p>The metric is indeed based on <a href=\"https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py\">https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py</a> with 2 differences:\n - long and short answers are nor scored independently, instead, common F1 is calculated for both long and short answers\n - no optimal thresholds are found, so the code shall be simpler that those of <code>nq_eval</code></p>\n\n<p>I haven't yet implemented the metric (so far using <code>nq_eval</code> ) but Dieter <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119462#684706\">can see</a> correlation between his local validation and LB.  </p>\n\n<p>So either we wait till Phil shares the official metric implementation (it's in C# though) or we go on with reverse-engineering. </p>",
      "votes": null,
      "replies": [
        {
          "id": 685236,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/01/2019 08:34:15",
          "content": "<p>Thanks, I read the description you quote, but it says what the metric is no, it does not say what the metric is.</p>\n\n<p>Anyway, I now see that I'm not the first to wonder about this, and I'm sorry people had to spend time like this: <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119143\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119143</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685361,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/01/2019 13:58:43",
          "content": "<p>Well, it also says smth about what it <em>is</em>, but indeed you need to study and adapt the source code of <code>nq_eval</code> and, moreover, analyze different peculiarities of the metric mentioned here on the forum.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685451,
          "author_name": "wochidadonggua",
          "author_url": "",
          "post_date": "12/01/2019 17:57:16",
          "content": "<p>I think the simple point is to use long_answer and short_answer to calculate the F-score together, and then do not set any answer to null. For me,  the cv is <code>0.672(for long)/0.548(for short)/0.623(for all)</code> with bert_joint. But there are still some bugs in my inference, so at this stage I can only say that these are some of my guesses😂</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685936,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/02/2019 14:46:18",
          "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> - that's exactly correct.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685993,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/02/2019 16:04:45",
          "content": "<p>That is probably accurate, but I don't think it is correct.  Asking participants to devise the metric code is not very useful for advancing the quality of models.  It would have been much better to provide the modified python code directly.  Anyway, i'll do like every other, work first on a metric code instead of models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685995,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/02/2019 16:08:55",
          "content": "<p>Releasing even C# code would also suffice, there enough qualified people here to check whether the C# code is doing what's expected. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685997,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/02/2019 16:10:35",
          "content": "<p>To speculate: either the evaluation code contains some obscure parts that are not intended to be made public. Or reverse-engineering the metric is considered to be a part of the competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 685999,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/02/2019 16:13:36",
          "content": "<p>Ah, ok, now I see Phil's comment below :) If the evaluation code is too involved, with source code spread over several files, then indeed I'd expect organizers to have clearly described the metric on Evaluation page. At least providing some toy examples covering all peculiarities of evaluation. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686358,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "12/03/2019 04:05:42",
          "content": "<p>Thanks for all the feedback! I've updated the description to add more detail about the metric. Hopefully this helps. Sorry about that - I felt that <code>nq_eval.py</code> + addendums covered the metric fairly well (and also the odd corner cases thereof) but that was incorrect. Please let me know if additional detail or information is needed!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686518,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 08:40:56",
          "content": "<p>Thanks, Phil! It's definitely better now. There're still some strange quirks about the metric (for instance, <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#681784\">this one</a> reported by Bo), and we have to spend time and submissions to understand why this or that happens.\nI'm almost fine with the idea of treating this as part of the competition :)\nBut no. In future, for other competitions, pls try to prepare official metric implementation, so that competitors spend time on what actually matters. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 684781,
      "author_name": "ddanevskyi",
      "author_url": "",
      "post_date": "11/30/2019 13:07:46",
      "content": "<p>My suggestion would be to try some metric learning. Just kidding, hehe ;)</p>\n\n<p>Seriously, though, I'm confused as well, it seems that the metric understanding is probably even more tough task than the actual task! Hope Kaggle will release the metric soon...</p>",
      "votes": null,
      "replies": [
        {
          "id": 685237,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/01/2019 08:35:12",
          "content": "<p>Yes, not having  a metric description is totally unbelievable.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 685933,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "12/02/2019 14:44:41",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> - Thanks for the feedback! I've added a link to the sklearn page describing \"micro F1\".</p>\n\n<p>The <code>nq_eval.py</code> script is the metric, with the differences listed in the Evaluation page. Our C# metric is unfortunately not a good candidate for clearing things up (several of the operations are being performed outside of the main code).</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "684508": "I'm considering entering this competition, and I am not sure I understand what we are supposed to optimize.\n\nThe evaluation page points to Wiikipedia page on F1 score, but says \"Micro F1\" score.  Micro F1 score is not described in this page unless mistaken.\n\nThen another link points to https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py , calling it the original metric, then says the metric here differs from the original metric.\n\nI see what the metric is not, but at no point do I see a description of what the metric is. \n\nCan someone point me to what the metric is?  I would expect a formula or an algorithm computing it.\n\nWhat am I missing?",
    "684564": "The metric is to be predicted also in this competition 😫",
    "684765": "The metric is indeed based on https://github.com/google-research-datasets/natural-questions/blob/master/nq_eval.py with 2 differences:\n - long and short answers are nor scored independently, instead, common F1 is calculated for both long and short answers\n - no optimal thresholds are found, so the code shall be simpler that those of `nq_eval`\n\nI haven't yet implemented the metric (so far using `nq_eval` ) but Dieter [can see](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119462#684706) correlation between his local validation and LB.  \n\nSo either we wait till Phil shares the official metric implementation (it's in C# though) or we go on with reverse-engineering.",
    "684781": "My suggestion would be to try some metric learning. Just kidding, hehe ;)\n\nSeriously, though, I'm confused as well, it seems that the metric understanding is probably even more tough task than the actual task! Hope Kaggle will release the metric soon...",
    "685236": "Thanks, I read the description you quote, but it says what the metric is no, it does not say what the metric is.\n\nAnyway, I now see that I'm not the first to wonder about this, and I'm sorry people had to spend time like this: https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119143",
    "685237": "Yes, not having  a metric description is totally unbelievable.",
    "685361": "Well, it also says smth about what it _is_, but indeed you need to study and adapt the source code of `nq_eval` and, moreover, analyze different peculiarities of the metric mentioned here on the forum.",
    "685451": "I think the simple point is to use long_answer and short_answer to calculate the F-score together, and then do not set any answer to null. For me,  the cv is `0.672(for long)/0.548(for short)/0.623(for all)` with bert_joint. But there are still some bugs in my inference, so at this stage I can only say that these are some of my guesses😂",
    "685933": "cpmpml - Thanks for the feedback! I've added a link to the sklearn page describing \"micro F1\".\n\nThe `nq_eval.py` script is the metric, with the differences listed in the Evaluation page. Our C# metric is unfortunately not a good candidate for clearing things up (several of the operations are being performed outside of the main code).",
    "685936": "kashnitsky - that's exactly correct.",
    "685993": "That is probably accurate, but I don't think it is correct.  Asking participants to devise the metric code is not very useful for advancing the quality of models.  It would have been much better to provide the modified python code directly.  Anyway, i'll do like every other, work first on a metric code instead of models.",
    "685995": "Releasing even C# code would also suffice, there enough qualified people here to check whether the C# code is doing what's expected.",
    "685997": "To speculate: either the evaluation code contains some obscure parts that are not intended to be made public. Or reverse-engineering the metric is considered to be a part of the competition.",
    "685999": "Ah, ok, now I see Phil's comment below :) If the evaluation code is too involved, with source code spread over several files, then indeed I'd expect organizers to have clearly described the metric on Evaluation page. At least providing some toy examples covering all peculiarities of evaluation.",
    "686358": "Thanks for all the feedback! I've updated the description to add more detail about the metric. Hopefully this helps. Sorry about that - I felt that `nq_eval.py` + addendums covered the metric fairly well (and also the odd corner cases thereof) but that was incorrect. Please let me know if additional detail or information is needed!",
    "686518": "Thanks, Phil! It's definitely better now. There're still some strange quirks about the metric (for instance, [this one](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#681784) reported by Bo), and we have to spend time and submissions to understand why this or that happens.\nI'm almost fine with the idea of treating this as part of the competition :)\nBut no. In future, for other competitions, pls try to prepare official metric implementation, so that competitors spend time on what actually matters."
  },
  "source": "meta"
}