{
  "id": 117182,
  "title": "Confused about the evaluation score",
  "url": "/competitions/tensorflow2-question-answering/discussion/117182",
  "author_name": "Alon Bochman",
  "post_date": "2019-11-13T19:29:13.261000",
  "votes": 17,
  "comment_count": 11,
  "views": 0,
  "content": "<p>I'm confused about the evaluation score.  Feel like I'm missing something very basic, so please forgive me. The competition page says:</p>\n\n<blockquote>\n  <p>Submissions are evaluated using micro F1 between the predicted and expected answers. Predicted long and short answers must match exactly the token indices of one of the ground truth labels ((or match YES/NO if the question has a yes/no short answer). There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</p>\n</blockquote>\n\n<p>I take this to mean you get a point (+1) for each correct answer (1 for long, 1 for short) and a zero (0) for incorrect answer. Then you compute the F1 score on top of those 1s and 0s, with long and short answer equally weighted. Great, but how do blanks figure into this? 50.5% of the questions in the training set have no answers. If you predict blanks for those, do you get the +1 (actually +2, blank for short and blank for long)? Or are blanks excluded from the F1 \"denominators\"?</p>\n\n<p>Let's assume our test set comes from the same population as the train set (otherwise why use the train set?) and also has 50% blank answers. The sample submission, which predicts all blanks, would be correct for about 50% of the questions. Why does it only score a 0?</p>",
  "messages": [
    {
      "id": 672337,
      "postDate": "2019-11-13T19:29:13.260Z",
      "content": "<p>I'm confused about the evaluation score.  Feel like I'm missing something very basic, so please forgive me. The competition page says:</p>\n\n<blockquote>\n  <p>Submissions are evaluated using micro F1 between the predicted and expected answers. Predicted long and short answers must match exactly the token indices of one of the ground truth labels ((or match YES/NO if the question has a yes/no short answer). There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.</p>\n</blockquote>\n\n<p>I take this to mean you get a point (+1) for each correct answer (1 for long, 1 for short) and a zero (0) for incorrect answer. Then you compute the F1 score on top of those 1s and 0s, with long and short answer equally weighted. Great, but how do blanks figure into this? 50.5% of the questions in the training set have no answers. If you predict blanks for those, do you get the +1 (actually +2, blank for short and blank for long)? Or are blanks excluded from the F1 \"denominators\"?</p>\n\n<p>Let's assume our test set comes from the same population as the train set (otherwise why use the train set?) and also has 50% blank answers. The sample submission, which predicts all blanks, would be correct for about 50% of the questions. Why does it only score a 0?</p>",
      "rawMarkdown": "I'm confused about the evaluation score.  Feel like I'm missing something very basic, so please forgive me. The competition page says:\n\n&gt; Submissions are evaluated using micro F1 between the predicted and expected answers. Predicted long and short answers must match exactly the token indices of one of the ground truth labels ((or match YES/NO if the question has a yes/no short answer). There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\n\nI take this to mean you get a point (+1) for each correct answer (1 for long, 1 for short) and a zero (0) for incorrect answer. Then you compute the F1 score on top of those 1s and 0s, with long and short answer equally weighted. Great, but how do blanks figure into this? 50.5% of the questions in the training set have no answers. If you predict blanks for those, do you get the +1 (actually +2, blank for short and blank for long)? Or are blanks excluded from the F1 \"denominators\"?\n\nLet's assume our test set comes from the same population as the train set (otherwise why use the train set?) and also has 50% blank answers. The sample submission, which predicts all blanks, would be correct for about 50% of the questions. Why does it only score a 0?",
      "votes": 17
    },
    {
      "id": 673845,
      "postDate": "2019-11-15T14:50:30.880Z",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> , <a href=\"/alonbochman\">@alonbochman</a> \nYou probably should look at it differently, this is a binary classification with a twist.\nlet's remember F1 - is for binary classification, in the F1 score only the true predication are taken into account, the False predictions aren't (correct and incorrect). The twist here is that in order to be true positive, you should also predict the answer indices (or yes/no) correctly.\nTwo other twists are:\n1. It is a multi-label classification - for every example you need to give prediction for short and for long.\n2. Some times there are a few correct answer.       </p>",
      "rawMarkdown": "@kashnitsky , @alonbochman \nYou probably should look at it differently, this is a binary classification with a twist.\nlet's remember F1 - is for binary classification, in the F1 score only the true predication are taken into account, the False predictions aren't (correct and incorrect). The twist here is that in order to be true positive, you should also predict the answer indices (or yes/no) correctly.\nTwo other twists are:\n1. It is a multi-label classification - for every example you need to give prediction for short and for long.\n2. Some times there are a few correct answer.       ",
      "votes": 2
    },
    {
      "id": 673900,
      "postDate": "2019-11-15T16:04:17.287Z",
      "content": "<p>The question is how are blanks scored in our F1? Perhaps an example will help. Consider the following table:</p>\n\n<p>Question |True Answer |Predicted Answer |Scoring\n1 --Yes-- Yes-- True Positive\n2 --No-- No-- True Negative\n3 --blank-- Yes-- False Positive?\n4 --blank --blank-- ??</p>\n\n<p>F1 is the harmonic mean of precision and recall. Precision is TP/(TP+FP). What's the precision for this table? If #4 counts as a TP, it's 2/3. If it counts as a TN, it's 1/3. If it's excluded because the true answer is blank, precision is 1/2.\nRecall is similarly affected.</p>\n\n<p><a href=\"/yuval6967\">@yuval6967</a> I think you are saying #4 is excluded, which would be consistent with the sample submission scoring 0 and precision above being 1/2.</p>",
      "rawMarkdown": "The question is how are blanks scored in our F1? Perhaps an example will help. Consider the following table:\n\nQuestion |True Answer |Predicted Answer |Scoring\n1 --Yes-- Yes-- True Positive\n2 --No-- No-- True Negative\n3 --blank-- Yes-- False Positive?\n4 --blank --blank-- ??\n\nF1 is the harmonic mean of precision and recall. Precision is TP/(TP+FP). What's the precision for this table? If #4 counts as a TP, it's 2/3. If it counts as a TN, it's 1/3. If it's excluded because the true answer is blank, precision is 1/2.\nRecall is similarly affected.\n\n@yuval6967 I think you are saying #4 is excluded, which would be consistent with the sample submission scoring 0 and precision above being 1/2.",
      "replies": [
        {
          "id": 673999,
          "postDate": "2019-11-15T19:18:20.880Z",
          "content": "<p><a href=\"/alonbochman\">@alonbochman</a> </p>\n\n<p>I'm saying:\nIn the real situation:\nThere is an answer = Positive\nThere is no answer = Negative</p>\n\n<p>In the prediction:\nLong/Short + indices or Yes, No = Positive\nLong/Short + Blank = Negative</p>\n\n<p>And your table become:</p>\n\n<p>(Real | Prediction)\n1. There is an answer  | Long/Short + indices or Yes, No (and indices are correct) - True Positive\n2. There is an answer | Long/Short + Blank (or indices are wrong)                           - False Negative\n3. There is no answer | Long/Short + indices or Yes, No                                             - False Positive\n4. There is no answer | Long/Short + Blank                                                                  - True Negative</p>",
          "rawMarkdown": "@alonbochman \n\nI'm saying:\nIn the real situation:\nThere is an answer = Positive\nThere is no answer = Negative\n\nIn the prediction:\nLong/Short + indices or Yes, No = Positive\nLong/Short + Blank = Negative\n\nAnd your table become:\n\n(Real | Prediction)\n1. There is an answer  | Long/Short + indices or Yes, No (and indices are correct) - True Positive\n2. There is an answer | Long/Short + Blank (or indices are wrong)                           - False Negative\n3. There is no answer | Long/Short + indices or Yes, No                                             - False Positive\n4. There is no answer | Long/Short + Blank                                                                  - True Negative",
          "votes": 9
        },
        {
          "id": 674000,
          "postDate": "2019-11-15T19:21:58.287Z",
          "content": "<p>thanks, <a href=\"/yuval6967\">@yuval6967</a> ! I'll analyze the evaluation script <code>nq_eval</code> and get back to your comments. Things get clearer.</p>",
          "rawMarkdown": "thanks, @yuval6967 ! I'll analyze the evaluation script `nq_eval` and get back to your comments. Things get clearer.\n"
        },
        {
          "id": 674024,
          "postDate": "2019-11-15T20:04:03.067Z",
          "content": "<p>Thanks <a href=\"/yuval6967\">@yuval6967</a> !</p>",
          "rawMarkdown": "Thanks @yuval6967 !"
        },
        {
          "id": 674344,
          "postDate": "2019-11-16T10:00:21.163Z",
          "content": "<p>Thanks again <a href=\"/yuval6967\">@yuval6967</a>, it actually helped to achieve 0.73</p>",
          "rawMarkdown": "Thanks again @yuval6967, it actually helped to achieve 0.73",
          "votes": 1
        },
        {
          "id": 674396,
          "postDate": "2019-11-16T12:26:55.533Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 674433,
          "postDate": "2019-11-16T13:59:53.073Z",
          "content": "<p>Hi,how evaluation script helped you to achieve 0.73 ? Do you mean thinking about the validation data f1 ?</p>",
          "rawMarkdown": "Hi,how evaluation script helped you to achieve 0.73 ? Do you mean thinking about the validation data f1 ?"
        },
        {
          "id": 674478,
          "postDate": "2019-11-16T14:54:13.790Z",
          "content": "<p>Well, it helped in an indirect way :) but definitely understanding the metric is crucial in a competition </p>",
          "rawMarkdown": "Well, it helped in an indirect way :) but definitely understanding the metric is crucial in a competition ",
          "votes": 1
        },
        {
          "id": 676643,
          "postDate": "2019-11-19T11:41:47.050Z",
          "content": "<p>It's frustrating that I write a script like this，but got 0.25 ,PB got 0.55 ，I don't know why</p>",
          "rawMarkdown": "It's frustrating that I write a script like this，but got 0.25 ,PB got 0.55 ，I don't know why"
        }
      ]
    },
    {
      "id": 673708,
      "postDate": "2019-11-15T11:16:47.360Z",
      "content": "<p>Yes, that's very strange and puzzles me, to be honest.</p>\n\n<p>So we have 50% of empty long answers and 63% of empty short answers. Micro-F1 is just accuracy in a multiclass setting. Then if test set distribution is ~ the same as train set distribution (at least concerning empty answers), we should expect 0.5 * (0.5 + 0.63) = 0.565 to be hit with just a sample submission file.</p>",
      "rawMarkdown": "Yes, that's very strange and puzzles me, to be honest.\n\nSo we have 50% of empty long answers and 63% of empty short answers. Micro-F1 is just accuracy in a multiclass setting. Then if test set distribution is ~ the same as train set distribution (at least concerning empty answers), we should expect 0.5 * (0.5 + 0.63) = 0.565 to be hit with just a sample submission file.\n"
    }
  ],
  "comments": [
    {
      "id": 673845,
      "author_name": "yuval reina",
      "author_url": "",
      "post_date": "2019-11-15T14:50:30.880000",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> , <a href=\"/alonbochman\">@alonbochman</a> \nYou probably should look at it differently, this is a binary classification with a twist.\nlet's remember F1 - is for binary classification, in the F1 score only the true predication are taken into account, the False predictions aren't (correct and incorrect). The twist here is that in order to be true positive, you should also predict the answer indices (or yes/no) correctly.\nTwo other twists are:\n1. It is a multi-label classification - for every example you need to give prediction for short and for long.\n2. Some times there are a few correct answer.       </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 673900,
      "author_name": "Alon Bochman",
      "author_url": "",
      "post_date": "2019-11-15T16:04:17.287000",
      "content": "<p>The question is how are blanks scored in our F1? Perhaps an example will help. Consider the following table:</p>\n\n<p>Question |True Answer |Predicted Answer |Scoring\n1 --Yes-- Yes-- True Positive\n2 --No-- No-- True Negative\n3 --blank-- Yes-- False Positive?\n4 --blank --blank-- ??</p>\n\n<p>F1 is the harmonic mean of precision and recall. Precision is TP/(TP+FP). What's the precision for this table? If #4 counts as a TP, it's 2/3. If it counts as a TN, it's 1/3. If it's excluded because the true answer is blank, precision is 1/2.\nRecall is similarly affected.</p>\n\n<p><a href=\"/yuval6967\">@yuval6967</a> I think you are saying #4 is excluded, which would be consistent with the sample submission scoring 0 and precision above being 1/2.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 673999,
          "author_name": "yuval reina",
          "author_url": "",
          "post_date": "2019-11-15T19:18:20.880000",
          "content": "<p><a href=\"/alonbochman\">@alonbochman</a> </p>\n\n<p>I'm saying:\nIn the real situation:\nThere is an answer = Positive\nThere is no answer = Negative</p>\n\n<p>In the prediction:\nLong/Short + indices or Yes, No = Positive\nLong/Short + Blank = Negative</p>\n\n<p>And your table become:</p>\n\n<p>(Real | Prediction)\n1. There is an answer  | Long/Short + indices or Yes, No (and indices are correct) - True Positive\n2. There is an answer | Long/Short + Blank (or indices are wrong)                           - False Negative\n3. There is no answer | Long/Short + indices or Yes, No                                             - False Positive\n4. There is no answer | Long/Short + Blank                                                                  - True Negative</p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 674000,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-15T19:21:58.287000",
          "content": "<p>thanks, <a href=\"/yuval6967\">@yuval6967</a> ! I'll analyze the evaluation script <code>nq_eval</code> and get back to your comments. Things get clearer.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674024,
          "author_name": "Alon Bochman",
          "author_url": "",
          "post_date": "2019-11-15T20:04:03.067000",
          "content": "<p>Thanks <a href=\"/yuval6967\">@yuval6967</a> !</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674344,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-16T10:00:21.163000",
          "content": "<p>Thanks again <a href=\"/yuval6967\">@yuval6967</a>, it actually helped to achieve 0.73</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 674396,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-11-16T12:26:55.533000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674433,
          "author_name": "sakuranew",
          "author_url": "",
          "post_date": "2019-11-16T13:59:53.073000",
          "content": "<p>Hi,how evaluation script helped you to achieve 0.73 ? Do you mean thinking about the validation data f1 ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 674478,
          "author_name": "Yury Kashnitsky",
          "author_url": "",
          "post_date": "2019-11-16T14:54:13.790000",
          "content": "<p>Well, it helped in an indirect way :) but definitely understanding the metric is crucial in a competition </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 676643,
          "author_name": "sakuranew",
          "author_url": "",
          "post_date": "2019-11-19T11:41:47.050000",
          "content": "<p>It's frustrating that I write a script like this，but got 0.25 ,PB got 0.55 ，I don't know why</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 673708,
      "author_name": "Yury Kashnitsky",
      "author_url": "",
      "post_date": "2019-11-15T11:16:47.360000",
      "content": "<p>Yes, that's very strange and puzzles me, to be honest.</p>\n\n<p>So we have 50% of empty long answers and 63% of empty short answers. Micro-F1 is just accuracy in a multiclass setting. Then if test set distribution is ~ the same as train set distribution (at least concerning empty answers), we should expect 0.5 * (0.5 + 0.63) = 0.565 to be hit with just a sample submission file.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "672337": "I'm confused about the evaluation score.  Feel like I'm missing something very basic, so please forgive me. The competition page says:\n\n&gt; Submissions are evaluated using micro F1 between the predicted and expected answers. Predicted long and short answers must match exactly the token indices of one of the ground truth labels ((or match YES/NO if the question has a yes/no short answer). There may be up to five labels for long answers, and more for short. If no answer applies, leave the prediction blank/null.\n\nI take this to mean you get a point (+1) for each correct answer (1 for long, 1 for short) and a zero (0) for incorrect answer. Then you compute the F1 score on top of those 1s and 0s, with long and short answer equally weighted. Great, but how do blanks figure into this? 50.5% of the questions in the training set have no answers. If you predict blanks for those, do you get the +1 (actually +2, blank for short and blank for long)? Or are blanks excluded from the F1 \"denominators\"?\n\nLet's assume our test set comes from the same population as the train set (otherwise why use the train set?) and also has 50% blank answers. The sample submission, which predicts all blanks, would be correct for about 50% of the questions. Why does it only score a 0?",
    "673845": "@kashnitsky , @alonbochman \nYou probably should look at it differently, this is a binary classification with a twist.\nlet's remember F1 - is for binary classification, in the F1 score only the true predication are taken into account, the False predictions aren't (correct and incorrect). The twist here is that in order to be true positive, you should also predict the answer indices (or yes/no) correctly.\nTwo other twists are:\n1. It is a multi-label classification - for every example you need to give prediction for short and for long.\n2. Some times there are a few correct answer.       ",
    "673900": "The question is how are blanks scored in our F1? Perhaps an example will help. Consider the following table:\n\nQuestion |True Answer |Predicted Answer |Scoring\n1 --Yes-- Yes-- True Positive\n2 --No-- No-- True Negative\n3 --blank-- Yes-- False Positive?\n4 --blank --blank-- ??\n\nF1 is the harmonic mean of precision and recall. Precision is TP/(TP+FP). What's the precision for this table? If #4 counts as a TP, it's 2/3. If it counts as a TN, it's 1/3. If it's excluded because the true answer is blank, precision is 1/2.\nRecall is similarly affected.\n\n@yuval6967 I think you are saying #4 is excluded, which would be consistent with the sample submission scoring 0 and precision above being 1/2.",
    "673708": "Yes, that's very strange and puzzles me, to be honest.\n\nSo we have 50% of empty long answers and 63% of empty short answers. Micro-F1 is just accuracy in a multiclass setting. Then if test set distribution is ~ the same as train set distribution (at least concerning empty answers), we should expect 0.5 * (0.5 + 0.63) = 0.565 to be hit with just a sample submission file.\n"
  }
}