{
  "id": 120030,
  "title": "A bit confused, using the new metric, offline ~0.5 but online 0.7+",
  "url": "/competitions/tensorflow2-question-answering/discussion/120030",
  "author_name": "",
  "post_date": "2019-12-03T07:54:58.344227Z",
  "votes": 8,
  "comment_count": 27,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "686484",
      "postDate": "12/03/2019 07:54:58",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "686486",
      "postDate": "12/03/2019 08:00:48",
      "content": "<p>Try to get offline ~0.78, then probably you end up 0.98+ in private LB 😄 </p>",
      "rawMarkdown": "Try to get offline ~0.78, then probably you end up 0.98+ in private LB 😄",
      "votes": null
    },
    {
      "id": "686491",
      "postDate": "12/03/2019 08:09:38",
      "content": "<p>Can you post code for your implementation of the metric? Might be not implemented correctly</p>",
      "rawMarkdown": "Can you post code for your implementation of the metric? Might be not implemented correctly",
      "votes": null
    },
    {
      "id": "686492",
      "postDate": "12/03/2019 08:16:31",
      "content": "<p>Hi，Does your offline score match the LB score now？I remember your offline score is also around 0.5</p>",
      "rawMarkdown": "Hi，Does your offline score match the LB score now？I remember your offline score is also around 0.5",
      "votes": null
    },
    {
      "id": "686497",
      "postDate": "12/03/2019 08:21:58",
      "content": "<p>it matched 2 days ago. Now I am having some weird behavior, which I try to investigate</p>",
      "rawMarkdown": "it matched 2 days ago. Now I am having some weird behavior, which I try to investigate",
      "votes": null
    },
    {
      "id": "686536",
      "postDate": "12/03/2019 09:09:08",
      "content": "<p><code>python\n    tp, fp, fn = 0, 0, 0\n    for gg, pp in zip(g, p):\n        if gg == pp:\n            if gg.s == -1: # no ground truth exists\n                continue\n            else:\n                tp += 1\n        else:\n            if pp.s == -1: # no prediction exists\n                fn += 1\n            else:\n                fp += 1\n</code>\n`\nThis is to compute long anwser score, * pp.s *is start token of prediction .\n&gt; TP = the predicted indices match one of the possible ground truth indices\nFP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists\nFN = no prediction has been made where a ground truth exists</p>\n\n<p><a href=\"/christofhenkel\">@christofhenkel</a> Can you find some error in my code? Thank you very much</p>",
      "rawMarkdown": "```python\n    tp, fp, fn = 0, 0, 0\n    for gg, pp in zip(g, p):\n        if gg == pp:\n            if gg.s == -1: # no ground truth exists\n                continue\n            else:\n                tp += 1\n        else:\n            if pp.s == -1: # no prediction exists\n                fn += 1\n            else:\n                fp += 1\n```\n`\nThis is to compute long anwser score, * pp.s *is start token of prediction .\n&gt; TP = the predicted indices match one of the possible ground truth indices\nFP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists\nFN = no prediction has been made where a ground truth exists\n\n@christofhenkel Can you find some error in my code? Thank you very much",
      "votes": null
    },
    {
      "id": "686554",
      "postDate": "12/03/2019 09:38:58",
      "content": "<p>I had something similar before, but I also got f1 of 0.5 when LB was 0.7. Now I use:</p>\n\n<p>```\nfrom sklearn.metrics import f1_score</p>\n\n<p>def get_long_f1(answer_df,label_df):\n    long_label = (label_df['long_answer'] != '').astype(int)\n    long_predict = np.zeros(answer_df.shape[0])\n    long_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n    return f1_score(long_label.values,long_predict)\n```</p>\n\n<p>and get f1 of 0.7</p>",
      "rawMarkdown": "I had something similar before, but I also got f1 of 0.5 when LB was 0.7. Now I use:\n\n```\nfrom sklearn.metrics import f1_score\n\ndef get_long_f1(answer_df,label_df):\n    long_label = (label_df['long_answer'] != '').astype(int)\n    long_predict = np.zeros(answer_df.shape[0])\n    long_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n    return f1_score(long_label.values,long_predict)\n```\n\nand get f1 of 0.7",
      "votes": null
    },
    {
      "id": "686586",
      "postDate": "12/03/2019 10:13:29",
      "content": "<p>You are the winner of the competition of metric learning!\nI realize precision is always 1, and that is explain why  LB metric does not seem to be additive.   <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#682184\">as Bo </a></p>",
      "rawMarkdown": "You are the winner of the competition of metric learning!\nI realize precision is always 1, and that is explain why  LB metric does not seem to be additive.   [as Bo ](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#682184)",
      "votes": null
    },
    {
      "id": "686594",
      "postDate": "12/03/2019 10:21:09",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> thanks for sharing, but this is only long answer F1, which is indeed expected to be at around 0.7, while in the metric description we are supposed to evaluate both short and long answers, and take micro-average.</p>",
      "rawMarkdown": "christofhenkel thanks for sharing, but this is only long answer F1, which is indeed expected to be at around 0.7, while in the metric description we are supposed to evaluate both short and long answers, and take micro-average.",
      "votes": null
    },
    {
      "id": "686602",
      "postDate": "12/03/2019 10:27:19",
      "content": "<p>Imagine label_df has a column 'short_answers' which contains all possible short answers joined with space. Then I use the following to calculate the micro f1</p>\n\n<p>```</p>\n\n<p>def in_shorts(row):\n    return row['short_answer'] in row['short_answers']</p>\n\n<p>from sklearn.metrics import f1_score</p>\n\n<p>def get_f1(answer_df,label_df):\n    short_label =  (label_df['short_answers'] != '').astype(int)\n    long_label =  (label_df['long_answer'] != '').astype(int)</p>\n\n<pre><code>long_predict = np.zeros(answer_df.shape[0])\nlong_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n\nshort_predict = np.zeros(answer_df.shape[0])\na = pd.concat([answer_df[['short_answer']],label_df[['short_answers']]], axis = 1)\na['short_answers'] = a['short_answers'].apply(lambda x: x.split())\nshort_predict[a.apply(lambda x: in_shorts(x), axis = 1) &amp; (a['short_answer'] != '')] = 1\n\nlong_f1 = f1_score(long_label.values,long_predict)\nshort_f1 = f1_score(short_label.values,short_predict)\nmicro_f1 = f1_score(np.concatenate([long_label,short_label]),np.concatenate([long_predict,short_predict]))\nreturn micro_f1, long_f1, short_f1\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "Imagine label_df has a column 'short_answers' which contains all possible short answers joined with space. Then I use the following to calculate the micro f1\n\n```\n\ndef in_shorts(row):\n    return row['short_answer'] in row['short_answers']\n\n\nfrom sklearn.metrics import f1_score\n\n\ndef get_f1(answer_df,label_df):\n    short_label =  (label_df['short_answers'] != '').astype(int)\n    long_label =  (label_df['long_answer'] != '').astype(int)\n\n    long_predict = np.zeros(answer_df.shape[0])\n    long_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n\n    short_predict = np.zeros(answer_df.shape[0])\n    a = pd.concat([answer_df[['short_answer']],label_df[['short_answers']]], axis = 1)\n    a['short_answers'] = a['short_answers'].apply(lambda x: x.split())\n    short_predict[a.apply(lambda x: in_shorts(x), axis = 1) &amp; (a['short_answer'] != '')] = 1\n\n    long_f1 = f1_score(long_label.values,long_predict)\n    short_f1 = f1_score(short_label.values,short_predict)\n    micro_f1 = f1_score(np.concatenate([long_label,short_label]),np.concatenate([long_predict,short_predict]))\n    return micro_f1, long_f1, short_f1\n\n```",
      "votes": null
    },
    {
      "id": "686606",
      "postDate": "12/03/2019 10:30:22",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> Many thanks.It really helps, tp,fp,fn didn't work well is so confused😂 </p>",
      "rawMarkdown": "christofhenkel Many thanks.It really helps, tp,fp,fn didn't work well is so confused😂",
      "votes": null
    },
    {
      "id": "686611",
      "postDate": "12/03/2019 10:39:14",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> wow that's clear now, thanks for sharing! I think the main difference to my implementation is that I compute F1 per each example, including short and long answers (this is my understanding of micro F1), while you compute F1 on the concatenation of long and short predictions, which seems a little different. But I guess our goal is not to implement metric as described, but to implement it as implemented by Kaggle. EDIT ahh by bad, it's me completely misunderstanding what micro-F1 means, sorry.</p>",
      "rawMarkdown": "christofhenkel wow that's clear now, thanks for sharing! I think the main difference to my implementation is that I compute F1 per each example, including short and long answers (this is my understanding of micro F1), while you compute F1 on the concatenation of long and short predictions, which seems a little different. But I guess our goal is not to implement metric as described, but to implement it as implemented by Kaggle. EDIT ahh by bad, it's me completely misunderstanding what micro-F1 means, sorry.",
      "votes": null
    },
    {
      "id": "686628",
      "postDate": "12/03/2019 11:05:41",
      "content": "<p>Yes, now Phil indicates clearly that it shall be a concatenation of long and short predictions \n&gt;In \"micro\" F1, both long and short answers count toward an overall precision and recall, which is then used to calculate a single F1 value. (In contrast, in \"macro\" F1 a separate F1 score is calculated for each class / label and then averaged.)</p>\n\n<p>The <em>metric learning</em> prize indeed goes to Dieter! 🎉 \nWaiting for an efficient Numba implementation to be shared by someone. </p>",
      "rawMarkdown": "Yes, now Phil indicates clearly that it shall be a concatenation of long and short predictions \n&gt;In \"micro\" F1, both long and short answers count toward an overall precision and recall, which is then used to calculate a single F1 value. (In contrast, in \"macro\" F1 a separate F1 score is calculated for each class / label and then averaged.)\n\nThe *metric learning* prize indeed goes to Dieter! 🎉 \nWaiting for an efficient Numba implementation to be shared by someone.",
      "votes": null
    },
    {
      "id": "686636",
      "postDate": "12/03/2019 11:11:52",
      "content": "<p>Well... we still have to check first whether this implementation is consistent with all these peculiarities of the metric, quite a few are described in multiple posts and comments. </p>",
      "rawMarkdown": "Well... we still have to check first whether this implementation is consistent with all these peculiarities of the metric, quite a few are described in multiple posts and comments.",
      "votes": null
    },
    {
      "id": "686656",
      "postDate": "12/03/2019 11:34:01",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> is specialist for </p>\n\n<blockquote>\n  <p>efficient Numba implementation to be shared by someone</p>\n</blockquote>",
      "rawMarkdown": "cpmpml is specialist for \n&gt; efficient Numba implementation to be shared by someone",
      "votes": null
    },
    {
      "id": "686709",
      "postDate": "12/03/2019 12:34:03",
      "content": "<p>Thanks, <a href=\"/christofhenkel\">@christofhenkel</a> . I really think with a clear metric, this competition will attract more participants!</p>",
      "rawMarkdown": "Thanks, @christofhenkel . I really think with a clear metric, this competition will attract more participants!",
      "votes": null
    },
    {
      "id": "686710",
      "postDate": "12/03/2019 12:36:11",
      "content": "<p>Thats what I feared :P</p>",
      "rawMarkdown": "Thats what I feared :P",
      "votes": null
    },
    {
      "id": "686743",
      "postDate": "12/03/2019 13:28:26",
      "content": "<p>This competition is too hard for me as a novice to playing with too many kaggle masters and grandmasters,too sad...😭 </p>",
      "rawMarkdown": "This competition is too hard for me as a novice to playing with too many kaggle masters and grandmasters,too sad...😭",
      "votes": null
    },
    {
      "id": "686750",
      "postDate": "12/03/2019 13:38:45",
      "content": "<p>You are doing pretty well so far ;)</p>",
      "rawMarkdown": "You are doing pretty well so far ;)",
      "votes": null
    },
    {
      "id": "686853",
      "postDate": "12/03/2019 16:05:49",
      "content": "<p>Great job <a href=\"/christofhenkel\">@christofhenkel</a> reverse engineering the LB metric!  Your metric matches all my LB experiments. I'm very confident this is what is implemented in the LB (probably unintended).</p>\n\n<p>But, this deviates from the official evaluation and the eval description page in a big way: \nThe precision is always 1, because when <code>long_label</code> or <code>short_label</code> is 0 (i.e. no gold span for an example), the <code>long_predict</code> and <code>short_predict</code> can never be 1. </p>\n\n<p>In other words, there is never FP. So the \"F1\" is just harmonic average of 1 and the <code>recall</code>. So the best strategy is to always make a prediction (and never empty prediction) regardless of confidence. This is different from official evaluation or the eval description page <code>FP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists</code> where a wrong prediction actually have negative impact to the score, thus choosing a threshold is important.</p>\n\n<p>FYI <a href=\"/philculliton\">@philculliton</a>  </p>",
      "rawMarkdown": "Great job @christofhenkel reverse engineering the LB metric!  Your metric matches all my LB experiments. I'm very confident this is what is implemented in the LB (probably unintended).\n\nBut, this deviates from the official evaluation and the eval description page in a big way: \nThe precision is always 1, because when `long_label` or `short_label` is 0 (i.e. no gold span for an example), the `long_predict` and `short_predict` can never be 1. \n\nIn other words, there is never FP. So the \"F1\" is just harmonic average of 1 and the `recall`. So the best strategy is to always make a prediction (and never empty prediction) regardless of confidence. This is different from official evaluation or the eval description page ```FP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists``` where a wrong prediction actually have negative impact to the score, thus choosing a threshold is important.\n\nFYI @philculliton",
      "votes": null
    },
    {
      "id": "688692",
      "postDate": "12/05/2019 23:59:58",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> <a href=\"/juliaelliott\">@juliaelliott</a> can we get some clarification about FPs: If no gold answer exists and you predict one, does it count as a False Positive (decreasing precision) or not? Empirically, seems like it does not. If so, would be nice to correct the eval page.</p>",
      "rawMarkdown": "philculliton @juliaelliott can we get some clarification about FPs: If no gold answer exists and you predict one, does it count as a False Positive (decreasing precision) or not? Empirically, seems like it does not. If so, would be nice to correct the eval page.",
      "votes": null
    },
    {
      "id": "698351",
      "postDate": "12/19/2019 05:28:53",
      "content": "<p>Does this still work after the update?</p>",
      "rawMarkdown": "Does this still work after the update?",
      "votes": null
    },
    {
      "id": "698446",
      "postDate": "12/19/2019 08:45:37",
      "content": "<p>no</p>",
      "rawMarkdown": "no",
      "votes": null
    },
    {
      "id": "698828",
      "postDate": "12/19/2019 18:31:48",
      "content": "<p>If thats the case .. then there isn't a way I can run a CV.\nI dunno if this is right, but form the way my score in LB is changing</p>\n\n<p>I feel blanks preds have far lower weight than wrong preds.\nLike: Loss due to FP &lt; FN , but there is a case where it might not be true also.</p>\n\n<p>I cant wrap my head around improving score without knowing how it's calculated...</p>",
      "rawMarkdown": "If thats the case .. then there isn't a way I can run a CV.\nI dunno if this is right, but form the way my score in LB is changing\n\n I feel blanks preds have far lower weight than wrong preds.\nLike: Loss due to FP &lt; FN , but there is a case where it might not be true also.\n\nI cant wrap my head around improving score without knowing how it's calculated...",
      "votes": null
    },
    {
      "id": "710165",
      "postDate": "01/04/2020 11:47:52",
      "content": "<p>Hello, <a href=\"/boliu0\">@boliu0</a>, <a href=\"/alonbochman\">@alonbochman</a>!\nI coudln't find any information about the results of the LB and Evaluation Metric update.\nIs this true for today's evalutaion?</p>\n\n<blockquote>\n  <p>In other words, there is never FP</p>\n</blockquote>\n\n<p>Empty sample_submission.csv gives 0.00 on Public LB. If the long and short answers are just concatenated before evaluation, then it means, that we don't have any empty answers on Public Test. Is it true?</p>",
      "rawMarkdown": "Hello, @boliu0, @alonbochman!\nI coudln't find any information about the results of the LB and Evaluation Metric update.\nIs this true for today's evalutaion?\n&gt; In other words, there is never FP\n\nEmpty sample_submission.csv gives 0.00 on Public LB. If the long and short answers are just concatenated before evaluation, then it means, that we don't have any empty answers on Public Test. Is it true?",
      "votes": null
    },
    {
      "id": "710168",
      "postDate": "01/04/2020 11:54:17",
      "content": "<p><a href=\"/alex0404\">@alex0404</a> it has been fixed, you can find it here <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120931\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120931</a></p>",
      "rawMarkdown": "alex0404 it has been fixed, you can find it here https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120931",
      "votes": null
    },
    {
      "id": "710170",
      "postDate": "01/04/2020 11:57:41",
      "content": "<p><a href=\"/axel81\">@axel81</a> , thank you!\nIf I understood correctly, then we need to predict whether <strong>there actually is</strong> an answer or <strong>not</strong>.\nIs it true?</p>",
      "rawMarkdown": "axel81 , thank you!\nIf I understood correctly, then we need to predict whether **there actually is** an answer or **not**.\nIs it true?",
      "votes": null
    },
    {
      "id": "710643",
      "postDate": "01/05/2020 02:52:29",
      "content": "<p>Hi <a href=\"/alex0404\">@alex0404</a> yes you're correct.</p>",
      "rawMarkdown": "Hi @alex0404 yes you're correct.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 686486,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "12/03/2019 08:00:48",
      "content": "<p>Try to get offline ~0.78, then probably you end up 0.98+ in private LB 😄 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 686491,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "12/03/2019 08:09:38",
      "content": "<p>Can you post code for your implementation of the metric? Might be not implemented correctly</p>",
      "votes": null,
      "replies": [
        {
          "id": 686492,
          "author_name": "zhaomeng1126",
          "author_url": "",
          "post_date": "12/03/2019 08:16:31",
          "content": "<p>Hi，Does your offline score match the LB score now？I remember your offline score is also around 0.5</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686497,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 08:21:58",
          "content": "<p>it matched 2 days ago. Now I am having some weird behavior, which I try to investigate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686536,
          "author_name": "zhaomeng1126",
          "author_url": "",
          "post_date": "12/03/2019 09:09:08",
          "content": "<p><code>python\n    tp, fp, fn = 0, 0, 0\n    for gg, pp in zip(g, p):\n        if gg == pp:\n            if gg.s == -1: # no ground truth exists\n                continue\n            else:\n                tp += 1\n        else:\n            if pp.s == -1: # no prediction exists\n                fn += 1\n            else:\n                fp += 1\n</code>\n`\nThis is to compute long anwser score, * pp.s *is start token of prediction .\n&gt; TP = the predicted indices match one of the possible ground truth indices\nFP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists\nFN = no prediction has been made where a ground truth exists</p>\n\n<p><a href=\"/christofhenkel\">@christofhenkel</a> Can you find some error in my code? Thank you very much</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686554,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 09:38:58",
          "content": "<p>I had something similar before, but I also got f1 of 0.5 when LB was 0.7. Now I use:</p>\n\n<p>```\nfrom sklearn.metrics import f1_score</p>\n\n<p>def get_long_f1(answer_df,label_df):\n    long_label = (label_df['long_answer'] != '').astype(int)\n    long_predict = np.zeros(answer_df.shape[0])\n    long_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n    return f1_score(long_label.values,long_predict)\n```</p>\n\n<p>and get f1 of 0.7</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686586,
          "author_name": "wangyunazx",
          "author_url": "",
          "post_date": "12/03/2019 10:13:29",
          "content": "<p>You are the winner of the competition of metric learning!\nI realize precision is always 1, and that is explain why  LB metric does not seem to be additive.   <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#682184\">as Bo </a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686594,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "12/03/2019 10:21:09",
          "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> thanks for sharing, but this is only long answer F1, which is indeed expected to be at around 0.7, while in the metric description we are supposed to evaluate both short and long answers, and take micro-average.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686602,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 10:27:19",
          "content": "<p>Imagine label_df has a column 'short_answers' which contains all possible short answers joined with space. Then I use the following to calculate the micro f1</p>\n\n<p>```</p>\n\n<p>def in_shorts(row):\n    return row['short_answer'] in row['short_answers']</p>\n\n<p>from sklearn.metrics import f1_score</p>\n\n<p>def get_f1(answer_df,label_df):\n    short_label =  (label_df['short_answers'] != '').astype(int)\n    long_label =  (label_df['long_answer'] != '').astype(int)</p>\n\n<pre><code>long_predict = np.zeros(answer_df.shape[0])\nlong_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n\nshort_predict = np.zeros(answer_df.shape[0])\na = pd.concat([answer_df[['short_answer']],label_df[['short_answers']]], axis = 1)\na['short_answers'] = a['short_answers'].apply(lambda x: x.split())\nshort_predict[a.apply(lambda x: in_shorts(x), axis = 1) &amp; (a['short_answer'] != '')] = 1\n\nlong_f1 = f1_score(long_label.values,long_predict)\nshort_f1 = f1_score(short_label.values,short_predict)\nmicro_f1 = f1_score(np.concatenate([long_label,short_label]),np.concatenate([long_predict,short_predict]))\nreturn micro_f1, long_f1, short_f1\n</code></pre>\n\n<p>```</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686606,
          "author_name": "zhaomeng1126",
          "author_url": "",
          "post_date": "12/03/2019 10:30:22",
          "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> Many thanks.It really helps, tp,fp,fn didn't work well is so confused😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686611,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "12/03/2019 10:39:14",
          "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> wow that's clear now, thanks for sharing! I think the main difference to my implementation is that I compute F1 per each example, including short and long answers (this is my understanding of micro F1), while you compute F1 on the concatenation of long and short predictions, which seems a little different. But I guess our goal is not to implement metric as described, but to implement it as implemented by Kaggle. EDIT ahh by bad, it's me completely misunderstanding what micro-F1 means, sorry.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686628,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 11:05:41",
          "content": "<p>Yes, now Phil indicates clearly that it shall be a concatenation of long and short predictions \n&gt;In \"micro\" F1, both long and short answers count toward an overall precision and recall, which is then used to calculate a single F1 value. (In contrast, in \"macro\" F1 a separate F1 score is calculated for each class / label and then averaged.)</p>\n\n<p>The <em>metric learning</em> prize indeed goes to Dieter! 🎉 \nWaiting for an efficient Numba implementation to be shared by someone. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686636,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 11:11:52",
          "content": "<p>Well... we still have to check first whether this implementation is consistent with all these peculiarities of the metric, quite a few are described in multiple posts and comments. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686656,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 11:34:01",
          "content": "<p><a href=\"/cpmpml\">@cpmpml</a> is specialist for </p>\n\n<blockquote>\n  <p>efficient Numba implementation to be shared by someone</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686709,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/03/2019 12:34:03",
          "content": "<p>Thanks, <a href=\"/christofhenkel\">@christofhenkel</a> . I really think with a clear metric, this competition will attract more participants!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686710,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 12:36:11",
          "content": "<p>Thats what I feared :P</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686743,
          "author_name": "zhaomeng1126",
          "author_url": "",
          "post_date": "12/03/2019 13:28:26",
          "content": "<p>This competition is too hard for me as a novice to playing with too many kaggle masters and grandmasters,too sad...😭 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686750,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 13:38:45",
          "content": "<p>You are doing pretty well so far ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686853,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "12/03/2019 16:05:49",
          "content": "<p>Great job <a href=\"/christofhenkel\">@christofhenkel</a> reverse engineering the LB metric!  Your metric matches all my LB experiments. I'm very confident this is what is implemented in the LB (probably unintended).</p>\n\n<p>But, this deviates from the official evaluation and the eval description page in a big way: \nThe precision is always 1, because when <code>long_label</code> or <code>short_label</code> is 0 (i.e. no gold span for an example), the <code>long_predict</code> and <code>short_predict</code> can never be 1. </p>\n\n<p>In other words, there is never FP. So the \"F1\" is just harmonic average of 1 and the <code>recall</code>. So the best strategy is to always make a prediction (and never empty prediction) regardless of confidence. This is different from official evaluation or the eval description page <code>FP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists</code> where a wrong prediction actually have negative impact to the score, thus choosing a threshold is important.</p>\n\n<p>FYI <a href=\"/philculliton\">@philculliton</a>  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 688692,
          "author_name": "alonbochman",
          "author_url": "",
          "post_date": "12/05/2019 23:59:58",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> <a href=\"/juliaelliott\">@juliaelliott</a> can we get some clarification about FPs: If no gold answer exists and you predict one, does it count as a False Positive (decreasing precision) or not? Empirically, seems like it does not. If so, would be nice to correct the eval page.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710165,
          "author_name": "alex0404",
          "author_url": "",
          "post_date": "01/04/2020 11:47:52",
          "content": "<p>Hello, <a href=\"/boliu0\">@boliu0</a>, <a href=\"/alonbochman\">@alonbochman</a>!\nI coudln't find any information about the results of the LB and Evaluation Metric update.\nIs this true for today's evalutaion?</p>\n\n<blockquote>\n  <p>In other words, there is never FP</p>\n</blockquote>\n\n<p>Empty sample_submission.csv gives 0.00 on Public LB. If the long and short answers are just concatenated before evaluation, then it means, that we don't have any empty answers on Public Test. Is it true?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710168,
          "author_name": "axel81",
          "author_url": "",
          "post_date": "01/04/2020 11:54:17",
          "content": "<p><a href=\"/alex0404\">@alex0404</a> it has been fixed, you can find it here <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120931\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120931</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710170,
          "author_name": "alex0404",
          "author_url": "",
          "post_date": "01/04/2020 11:57:41",
          "content": "<p><a href=\"/axel81\">@axel81</a> , thank you!\nIf I understood correctly, then we need to predict whether <strong>there actually is</strong> an answer or <strong>not</strong>.\nIs it true?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 710643,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "01/05/2020 02:52:29",
          "content": "<p>Hi <a href=\"/alex0404\">@alex0404</a> yes you're correct.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 698351,
      "author_name": "rarun2596",
      "author_url": "",
      "post_date": "12/19/2019 05:28:53",
      "content": "<p>Does this still work after the update?</p>",
      "votes": null,
      "replies": [
        {
          "id": 698446,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/19/2019 08:45:37",
          "content": "<p>no</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 698828,
          "author_name": "rarun2596",
          "author_url": "",
          "post_date": "12/19/2019 18:31:48",
          "content": "<p>If thats the case .. then there isn't a way I can run a CV.\nI dunno if this is right, but form the way my score in LB is changing</p>\n\n<p>I feel blanks preds have far lower weight than wrong preds.\nLike: Loss due to FP &lt; FN , but there is a case where it might not be true also.</p>\n\n<p>I cant wrap my head around improving score without knowing how it's calculated...</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "686484": "",
    "686486": "Try to get offline ~0.78, then probably you end up 0.98+ in private LB 😄",
    "686491": "Can you post code for your implementation of the metric? Might be not implemented correctly",
    "686492": "Hi，Does your offline score match the LB score now？I remember your offline score is also around 0.5",
    "686497": "it matched 2 days ago. Now I am having some weird behavior, which I try to investigate",
    "686536": "```python\n    tp, fp, fn = 0, 0, 0\n    for gg, pp in zip(g, p):\n        if gg == pp:\n            if gg.s == -1: # no ground truth exists\n                continue\n            else:\n                tp += 1\n        else:\n            if pp.s == -1: # no prediction exists\n                fn += 1\n            else:\n                fp += 1\n```\n`\nThis is to compute long anwser score, * pp.s *is start token of prediction .\n&gt; TP = the predicted indices match one of the possible ground truth indices\nFP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists\nFN = no prediction has been made where a ground truth exists\n\n@christofhenkel Can you find some error in my code? Thank you very much",
    "686554": "I had something similar before, but I also got f1 of 0.5 when LB was 0.7. Now I use:\n\n```\nfrom sklearn.metrics import f1_score\n\ndef get_long_f1(answer_df,label_df):\n    long_label = (label_df['long_answer'] != '').astype(int)\n    long_predict = np.zeros(answer_df.shape[0])\n    long_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n    return f1_score(long_label.values,long_predict)\n```\n\nand get f1 of 0.7",
    "686586": "You are the winner of the competition of metric learning!\nI realize precision is always 1, and that is explain why  LB metric does not seem to be additive.   [as Bo ](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/118758#682184)",
    "686594": "christofhenkel thanks for sharing, but this is only long answer F1, which is indeed expected to be at around 0.7, while in the metric description we are supposed to evaluate both short and long answers, and take micro-average.",
    "686602": "Imagine label_df has a column 'short_answers' which contains all possible short answers joined with space. Then I use the following to calculate the micro f1\n\n```\n\ndef in_shorts(row):\n    return row['short_answer'] in row['short_answers']\n\n\nfrom sklearn.metrics import f1_score\n\n\ndef get_f1(answer_df,label_df):\n    short_label =  (label_df['short_answers'] != '').astype(int)\n    long_label =  (label_df['long_answer'] != '').astype(int)\n\n    long_predict = np.zeros(answer_df.shape[0])\n    long_predict[(answer_df['long_answer'] == label_df['long_answer']) &amp; (answer_df['long_answer'] != '')] = 1\n\n    short_predict = np.zeros(answer_df.shape[0])\n    a = pd.concat([answer_df[['short_answer']],label_df[['short_answers']]], axis = 1)\n    a['short_answers'] = a['short_answers'].apply(lambda x: x.split())\n    short_predict[a.apply(lambda x: in_shorts(x), axis = 1) &amp; (a['short_answer'] != '')] = 1\n\n    long_f1 = f1_score(long_label.values,long_predict)\n    short_f1 = f1_score(short_label.values,short_predict)\n    micro_f1 = f1_score(np.concatenate([long_label,short_label]),np.concatenate([long_predict,short_predict]))\n    return micro_f1, long_f1, short_f1\n\n```",
    "686606": "christofhenkel Many thanks.It really helps, tp,fp,fn didn't work well is so confused😂",
    "686611": "christofhenkel wow that's clear now, thanks for sharing! I think the main difference to my implementation is that I compute F1 per each example, including short and long answers (this is my understanding of micro F1), while you compute F1 on the concatenation of long and short predictions, which seems a little different. But I guess our goal is not to implement metric as described, but to implement it as implemented by Kaggle. EDIT ahh by bad, it's me completely misunderstanding what micro-F1 means, sorry.",
    "686628": "Yes, now Phil indicates clearly that it shall be a concatenation of long and short predictions \n&gt;In \"micro\" F1, both long and short answers count toward an overall precision and recall, which is then used to calculate a single F1 value. (In contrast, in \"macro\" F1 a separate F1 score is calculated for each class / label and then averaged.)\n\nThe *metric learning* prize indeed goes to Dieter! 🎉 \nWaiting for an efficient Numba implementation to be shared by someone.",
    "686636": "Well... we still have to check first whether this implementation is consistent with all these peculiarities of the metric, quite a few are described in multiple posts and comments.",
    "686656": "cpmpml is specialist for \n&gt; efficient Numba implementation to be shared by someone",
    "686709": "Thanks, @christofhenkel . I really think with a clear metric, this competition will attract more participants!",
    "686710": "Thats what I feared :P",
    "686743": "This competition is too hard for me as a novice to playing with too many kaggle masters and grandmasters,too sad...😭",
    "686750": "You are doing pretty well so far ;)",
    "686853": "Great job @christofhenkel reverse engineering the LB metric!  Your metric matches all my LB experiments. I'm very confident this is what is implemented in the LB (probably unintended).\n\nBut, this deviates from the official evaluation and the eval description page in a big way: \nThe precision is always 1, because when `long_label` or `short_label` is 0 (i.e. no gold span for an example), the `long_predict` and `short_predict` can never be 1. \n\nIn other words, there is never FP. So the \"F1\" is just harmonic average of 1 and the `recall`. So the best strategy is to always make a prediction (and never empty prediction) regardless of confidence. This is different from official evaluation or the eval description page ```FP = the predicted indices do NOT match one of the possible ground truth indices, OR a prediction has been made where no ground truth exists``` where a wrong prediction actually have negative impact to the score, thus choosing a threshold is important.\n\nFYI @philculliton",
    "688692": "philculliton @juliaelliott can we get some clarification about FPs: If no gold answer exists and you predict one, does it count as a False Positive (decreasing precision) or not? Empirically, seems like it does not. If so, would be nice to correct the eval page.",
    "698351": "Does this still work after the update?",
    "698446": "no",
    "698828": "If thats the case .. then there isn't a way I can run a CV.\nI dunno if this is right, but form the way my score in LB is changing\n\n I feel blanks preds have far lower weight than wrong preds.\nLike: Loss due to FP &lt; FN , but there is a case where it might not be true also.\n\nI cant wrap my head around improving score without knowing how it's calculated...",
    "710165": "Hello, @boliu0, @alonbochman!\nI coudln't find any information about the results of the LB and Evaluation Metric update.\nIs this true for today's evalutaion?\n&gt; In other words, there is never FP\n\nEmpty sample_submission.csv gives 0.00 on Public LB. If the long and short answers are just concatenated before evaluation, then it means, that we don't have any empty answers on Public Test. Is it true?",
    "710168": "alex0404 it has been fixed, you can find it here https://www.kaggle.com/c/tensorflow2-question-answering/discussion/120931",
    "710170": "axel81 , thank you!\nIf I understood correctly, then we need to predict whether **there actually is** an answer or **not**.\nIs it true?",
    "710643": "Hi @alex0404 yes you're correct."
  },
  "source": "meta"
}