{
  "id": 417478,
  "title": "[MATH] Why optimizing each question independently doesn't necessarily give good results",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/417478",
  "author_name": "",
  "post_date": "2023-06-16T00:06:23.115933900Z",
  "votes": 12,
  "comment_count": 8,
  "views": 0,
  "content": "<h2>Context</h2>\n<p>First, a shout out to <a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a> for his <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417359\" target=\"_blank\">great post</a> on question-based optimization. This is something that probably everyone tried and failed with. For those interested, here is some math behind the story.</p>\n<h2>Basic definitions</h2>\n<p>In this competition <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/evaluation\" target=\"_blank\">macro averaging</a> is used, hence, the final F1 score is a mean of F1 scores for two classes: \"questions answered correctly\" and \"questions answered incorrectly\", i.e.</p>\n<p>$$F1 = \\frac{F1^{\\mathrm{CORRECT}} + F1^{\\mathrm{INCORRECT}}}{2}$$</p>\n<p>By <a href=\"https://en.wikipedia.org/wiki/F-score\" target=\"_blank\">definition</a> \\( F1^{\\mathrm{CORRECT}} \\) and \\( F1^{\\mathrm{INCORRECT}} \\) can be expressed in terms of true positives (TP), false positives (FP) and false negatives (FN) as follows</p>\n<p>$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2TP^\\mathrm{(IN)CORRECT}}{2TP^\\mathrm{(IN)CORRECT} + FP^\\mathrm{(IN)CORRECT} + FN^\\mathrm{(IN)CORRECT}} $$</p>\n<p>These TP, FP and FN are nothing but the sums of the corresponding per-question values</p>\n<p>$$ TP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT} $$</p>\n<p>$$ FP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FP_Q^\\mathrm{(IN)CORRECT} $$</p>\n<p>$$ FN^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FN_Q^\\mathrm{(IN)CORRECT} $$</p>\n<p><br></p>\n<h2>Overall F1 vs per-question F1</h2>\n<p>Now we can express \\( F_1^{\\mathrm{CORRECT}} \\) and \\( F_1^{\\mathrm{INCORRECT}} \\) in terms of per-question TP, FP and FN</p>\n<p>$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$</p>\n<p>However, for each question considered separately the F1 score is<br>\n$$ F1_Q^{\\mathrm{(IN)CORRECT}} = \\frac{2TP_Q^\\mathrm{(IN)CORRECT}}{2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}} $$</p>\n<p>As you see, denominator in the above formula changes from question to question, so it is obvious that you can't easily sum up all the question scores to get the final one, i.e. in the general case<br>\n$$ F1^{\\mathrm{(IN)CORRECT}} \\neq \\sum_{Q = 1}^{18}F1_Q^{\\mathrm{(IN)CORRECT}} $$</p>\n<h2>Conclusion</h2>\n<p>From the math point of view it is obvious that optimizing F1 score for each question separately is different from optimizing the overall F1 score. So, there is no surprise this method also didn't work in practice.</p>",
  "messages": [
    {
      "id": "2304340",
      "postDate": "06/16/2023 00:06:23",
      "content": "<h2>Context</h2>\n<p>First, a shout out to <a href=\"https://www.kaggle.com/gehallak\" target=\"_blank\">@gehallak</a> for his <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417359\" target=\"_blank\">great post</a> on question-based optimization. This is something that probably everyone tried and failed with. For those interested, here is some math behind the story.</p>\n<h2>Basic definitions</h2>\n<p>In this competition <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/evaluation\" target=\"_blank\">macro averaging</a> is used, hence, the final F1 score is a mean of F1 scores for two classes: \"questions answered correctly\" and \"questions answered incorrectly\", i.e.</p>\n<p>$$F1 = \\frac{F1^{\\mathrm{CORRECT}} + F1^{\\mathrm{INCORRECT}}}{2}$$</p>\n<p>By <a href=\"https://en.wikipedia.org/wiki/F-score\" target=\"_blank\">definition</a> \\( F1^{\\mathrm{CORRECT}} \\) and \\( F1^{\\mathrm{INCORRECT}} \\) can be expressed in terms of true positives (TP), false positives (FP) and false negatives (FN) as follows</p>\n<p>$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2TP^\\mathrm{(IN)CORRECT}}{2TP^\\mathrm{(IN)CORRECT} + FP^\\mathrm{(IN)CORRECT} + FN^\\mathrm{(IN)CORRECT}} $$</p>\n<p>These TP, FP and FN are nothing but the sums of the corresponding per-question values</p>\n<p>$$ TP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT} $$</p>\n<p>$$ FP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FP_Q^\\mathrm{(IN)CORRECT} $$</p>\n<p>$$ FN^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FN_Q^\\mathrm{(IN)CORRECT} $$</p>\n<p><br></p>\n<h2>Overall F1 vs per-question F1</h2>\n<p>Now we can express \\( F_1^{\\mathrm{CORRECT}} \\) and \\( F_1^{\\mathrm{INCORRECT}} \\) in terms of per-question TP, FP and FN</p>\n<p>$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$</p>\n<p>However, for each question considered separately the F1 score is<br>\n$$ F1_Q^{\\mathrm{(IN)CORRECT}} = \\frac{2TP_Q^\\mathrm{(IN)CORRECT}}{2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}} $$</p>\n<p>As you see, denominator in the above formula changes from question to question, so it is obvious that you can't easily sum up all the question scores to get the final one, i.e. in the general case<br>\n$$ F1^{\\mathrm{(IN)CORRECT}} \\neq \\sum_{Q = 1}^{18}F1_Q^{\\mathrm{(IN)CORRECT}} $$</p>\n<h2>Conclusion</h2>\n<p>From the math point of view it is obvious that optimizing F1 score for each question separately is different from optimizing the overall F1 score. So, there is no surprise this method also didn't work in practice.</p>",
      "rawMarkdown": "##Context \n\nFirst, a shout out to @gehallak for his [great post](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417359) on question-based optimization. This is something that probably everyone tried and failed with. For those interested, here is some math behind the story.\n\n## Basic definitions\nIn this competition [macro averaging](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/evaluation) is used, hence, the final F1 score is a mean of F1 scores for two classes: \"questions answered correctly\" and \"questions answered incorrectly\", i.e.\n\n$$F1 = \\frac{F1^{\\mathrm{CORRECT}} + F1^{\\mathrm{INCORRECT}}}{2}$$\n\nBy [definition](https://en.wikipedia.org/wiki/F-score) \\\\( F1^{\\mathrm{CORRECT}} \\\\) and \\\\( F1^{\\mathrm{INCORRECT}} \\\\) can be expressed in terms of true positives (TP), false positives (FP) and false negatives (FN) as follows\n\n$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2TP^\\mathrm{(IN)CORRECT}}{2TP^\\mathrm{(IN)CORRECT} + FP^\\mathrm{(IN)CORRECT} + FN^\\mathrm{(IN)CORRECT}} $$\n\nThese TP, FP and FN are nothing but the sums of the corresponding per-question values\n\n$$ TP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT} $$\n\n$$ FP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FP_Q^\\mathrm{(IN)CORRECT} $$\n\n$$ FN^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FN_Q^\\mathrm{(IN)CORRECT} $$\n\n<br />\n\n## Overall F1 vs per-question F1\n\nNow we can express \\\\( F_1^{\\mathrm{CORRECT}} \\\\) and \\\\( F_1^{\\mathrm{INCORRECT}} \\\\) in terms of per-question TP, FP and FN\n\n$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$\n\nHowever, for each question considered separately the F1 score is\n$$ F1_Q^{\\mathrm{(IN)CORRECT}} = \\frac{2TP_Q^\\mathrm{(IN)CORRECT}}{2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}} $$\n\nAs you see, denominator in the above formula changes from question to question, so it is obvious that you can't easily sum up all the question scores to get the final one, i.e. in the general case\n$$ F1^{\\mathrm{(IN)CORRECT}} \\neq \\sum_{Q = 1}^{18}F1_Q^{\\mathrm{(IN)CORRECT}} $$\n\n## Conclusion \n\nFrom the math point of view it is obvious that optimizing F1 score for each question separately is different from optimizing the overall F1 score. So, there is no surprise this method also didn't work in practice.",
      "votes": null
    },
    {
      "id": "2304404",
      "postDate": "06/16/2023 01:18:26",
      "content": "<p>But with this method of calculating f1 metric, we may get away with predicting completely wrongly for some  questions and we can never know the model performance with respect to individual questions:  refer my post here:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417070#2304397\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417070#2304397</a></p>",
      "rawMarkdown": "But with this method of calculating f1 metric, we may get away with predicting completely wrongly for some  questions and we can never know the model performance with respect to individual questions:  refer my post here:\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417070#2304397",
      "votes": null
    },
    {
      "id": "2304534",
      "postDate": "06/16/2023 04:26:18",
      "content": "<p>I see your point, but that’s probably a topic for a separate discussion. Here I just wanted to bring to everyone’s attention some basic math, that is behind the competition’s metric, whatever this metric is…</p>",
      "rawMarkdown": "I see your point, but that’s probably a topic for a separate discussion. Here I just wanted to bring to everyone’s attention some basic math, that is behind the competition’s metric, whatever this metric is…",
      "votes": null
    },
    {
      "id": "2304540",
      "postDate": "06/16/2023 04:36:04",
      "content": "<p>fine, good work!</p>",
      "rawMarkdown": "fine, good work!",
      "votes": null
    },
    {
      "id": "2304551",
      "postDate": "06/16/2023 04:44:55",
      "content": "<p>I don’t really know why the metric is F1 macro, may be host only cares about the overall student performance, i.e. passed/not passed and some particular questions are not that so important.</p>",
      "rawMarkdown": "I don’t really know why the metric is F1 macro, may be host only cares about the overall student performance, i.e. passed/not passed and some particular questions are not that so important.",
      "votes": null
    },
    {
      "id": "2305880",
      "postDate": "06/17/2023 01:08:54",
      "content": "<p>Here is an interesting situation though hypothetical: A student does all navigations as usual but finally answers all the questions correctly -some kind of question-answer leak!! The question wise average f1 score will still give high f1 score - true evaluation of the model, but the binary macro f1 score will give very low  f1 score :-) My sample analysis shows that if 50% users were evaluated wrongly, the average f1 score can still be above 50% but binary macro f1 will be very low </p>",
      "rawMarkdown": "Here is an interesting situation though hypothetical: A student does all navigations as usual but finally answers all the questions correctly -some kind of question-answer leak!! The question wise average f1 score will still give high f1 score - true evaluation of the model, but the binary macro f1 score will give very low  f1 score :-) My sample analysis shows that if 50% users were evaluated wrongly, the average f1 score can still be above 50% but binary macro f1 will be very low",
      "votes": null
    },
    {
      "id": "2307721",
      "postDate": "06/18/2023 11:35:16",
      "content": "<p>good observation. thank you</p>",
      "rawMarkdown": "good observation. thank you",
      "votes": null
    },
    {
      "id": "2315405",
      "postDate": "06/24/2023 05:49:50",
      "content": "<blockquote>\n  <p>Now we can express \\( F_1^{\\mathrm{CORRECT}} \\) and \\( F_1^{\\mathrm{INCORRECT}} \\) in terms of per-question TP, FP and FN</p>\n  <p>$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{2\\sum_{Q = 1}^{18}\\left(TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$</p>\n</blockquote>\n<p>Thanks for the heads up. Just want to point out one tiny mistake. In the formula above, 2 in the denominator should be inside the summation. $$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$</p>",
      "rawMarkdown": "> Now we can express \\\\( F_1^{\\mathrm{CORRECT}} \\\\) and \\\\( F_1^{\\mathrm{INCORRECT}} \\\\) in terms of per-question TP, FP and FN\n> \n> $$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{2\\sum_{Q = 1}^{18}\\left(TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$\n\nThanks for the heads up. Just want to point out one tiny mistake. In the formula above, 2 in the denominator should be inside the summation. $$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$",
      "votes": null
    },
    {
      "id": "2315417",
      "postDate": "06/24/2023 05:57:18",
      "content": "<p>You are right, I’ve just fixed the typo. Thank you for pointing this out.</p>",
      "rawMarkdown": "You are right, I’ve just fixed the typo. Thank you for pointing this out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2304404,
      "author_name": "murugesann",
      "author_url": "",
      "post_date": "06/16/2023 01:18:26",
      "content": "<p>But with this method of calculating f1 metric, we may get away with predicting completely wrongly for some  questions and we can never know the model performance with respect to individual questions:  refer my post here:<br>\n<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417070#2304397\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417070#2304397</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2304534,
          "author_name": "kononenko",
          "author_url": "",
          "post_date": "06/16/2023 04:26:18",
          "content": "<p>I see your point, but that’s probably a topic for a separate discussion. Here I just wanted to bring to everyone’s attention some basic math, that is behind the competition’s metric, whatever this metric is…</p>",
          "votes": null,
          "replies": [
            {
              "id": 2304540,
              "author_name": "murugesann",
              "author_url": "",
              "post_date": "06/16/2023 04:36:04",
              "content": "<p>fine, good work!</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2304551,
                  "author_name": "kononenko",
                  "author_url": "",
                  "post_date": "06/16/2023 04:44:55",
                  "content": "<p>I don’t really know why the metric is F1 macro, may be host only cares about the overall student performance, i.e. passed/not passed and some particular questions are not that so important.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2305880,
                      "author_name": "murugesann",
                      "author_url": "",
                      "post_date": "06/17/2023 01:08:54",
                      "content": "<p>Here is an interesting situation though hypothetical: A student does all navigations as usual but finally answers all the questions correctly -some kind of question-answer leak!! The question wise average f1 score will still give high f1 score - true evaluation of the model, but the binary macro f1 score will give very low  f1 score :-) My sample analysis shows that if 50% users were evaluated wrongly, the average f1 score can still be above 50% but binary macro f1 will be very low </p>",
                      "votes": null,
                      "replies": []
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2307721,
      "author_name": "yassinezarwal",
      "author_url": "",
      "post_date": "06/18/2023 11:35:16",
      "content": "<p>good observation. thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2315405,
      "author_name": "fulinjiang",
      "author_url": "",
      "post_date": "06/24/2023 05:49:50",
      "content": "<blockquote>\n  <p>Now we can express \\( F_1^{\\mathrm{CORRECT}} \\) and \\( F_1^{\\mathrm{INCORRECT}} \\) in terms of per-question TP, FP and FN</p>\n  <p>$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{2\\sum_{Q = 1}^{18}\\left(TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$</p>\n</blockquote>\n<p>Thanks for the heads up. Just want to point out one tiny mistake. In the formula above, 2 in the denominator should be inside the summation. $$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$</p>",
      "votes": null,
      "replies": [
        {
          "id": 2315417,
          "author_name": "kononenko",
          "author_url": "",
          "post_date": "06/24/2023 05:57:18",
          "content": "<p>You are right, I’ve just fixed the typo. Thank you for pointing this out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2304340": "##Context \n\nFirst, a shout out to @gehallak for his [great post](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417359) on question-based optimization. This is something that probably everyone tried and failed with. For those interested, here is some math behind the story.\n\n## Basic definitions\nIn this competition [macro averaging](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/overview/evaluation) is used, hence, the final F1 score is a mean of F1 scores for two classes: \"questions answered correctly\" and \"questions answered incorrectly\", i.e.\n\n$$F1 = \\frac{F1^{\\mathrm{CORRECT}} + F1^{\\mathrm{INCORRECT}}}{2}$$\n\nBy [definition](https://en.wikipedia.org/wiki/F-score) \\\\( F1^{\\mathrm{CORRECT}} \\\\) and \\\\( F1^{\\mathrm{INCORRECT}} \\\\) can be expressed in terms of true positives (TP), false positives (FP) and false negatives (FN) as follows\n\n$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2TP^\\mathrm{(IN)CORRECT}}{2TP^\\mathrm{(IN)CORRECT} + FP^\\mathrm{(IN)CORRECT} + FN^\\mathrm{(IN)CORRECT}} $$\n\nThese TP, FP and FN are nothing but the sums of the corresponding per-question values\n\n$$ TP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT} $$\n\n$$ FP^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FP_Q^\\mathrm{(IN)CORRECT} $$\n\n$$ FN^\\mathrm{(IN)CORRECT} = \\sum_{Q = 1}^{18}FN_Q^\\mathrm{(IN)CORRECT} $$\n\n<br />\n\n## Overall F1 vs per-question F1\n\nNow we can express \\\\( F_1^{\\mathrm{CORRECT}} \\\\) and \\\\( F_1^{\\mathrm{INCORRECT}} \\\\) in terms of per-question TP, FP and FN\n\n$$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$\n\nHowever, for each question considered separately the F1 score is\n$$ F1_Q^{\\mathrm{(IN)CORRECT}} = \\frac{2TP_Q^\\mathrm{(IN)CORRECT}}{2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}} $$\n\nAs you see, denominator in the above formula changes from question to question, so it is obvious that you can't easily sum up all the question scores to get the final one, i.e. in the general case\n$$ F1^{\\mathrm{(IN)CORRECT}} \\neq \\sum_{Q = 1}^{18}F1_Q^{\\mathrm{(IN)CORRECT}} $$\n\n## Conclusion \n\nFrom the math point of view it is obvious that optimizing F1 score for each question separately is different from optimizing the overall F1 score. So, there is no surprise this method also didn't work in practice.",
    "2304404": "But with this method of calculating f1 metric, we may get away with predicting completely wrongly for some  questions and we can never know the model performance with respect to individual questions:  refer my post here:\nhttps://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/417070#2304397",
    "2304534": "I see your point, but that’s probably a topic for a separate discussion. Here I just wanted to bring to everyone’s attention some basic math, that is behind the competition’s metric, whatever this metric is…",
    "2304540": "fine, good work!",
    "2304551": "I don’t really know why the metric is F1 macro, may be host only cares about the overall student performance, i.e. passed/not passed and some particular questions are not that so important.",
    "2305880": "Here is an interesting situation though hypothetical: A student does all navigations as usual but finally answers all the questions correctly -some kind of question-answer leak!! The question wise average f1 score will still give high f1 score - true evaluation of the model, but the binary macro f1 score will give very low  f1 score :-) My sample analysis shows that if 50% users were evaluated wrongly, the average f1 score can still be above 50% but binary macro f1 will be very low",
    "2307721": "good observation. thank you",
    "2315405": "> Now we can express \\\\( F_1^{\\mathrm{CORRECT}} \\\\) and \\\\( F_1^{\\mathrm{INCORRECT}} \\\\) in terms of per-question TP, FP and FN\n> \n> $$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{2\\sum_{Q = 1}^{18}\\left(TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$\n\nThanks for the heads up. Just want to point out one tiny mistake. In the formula above, 2 in the denominator should be inside the summation. $$ F1^{\\mathrm{(IN)CORRECT}} = \\frac{2\\sum_{Q = 1}^{18}TP_Q^\\mathrm{(IN)CORRECT}}{\\sum_{Q = 1}^{18}\\left(2TP_Q^\\mathrm{(IN)CORRECT} + FP_Q^\\mathrm{(IN)CORRECT} + FN_Q^\\mathrm{(IN)CORRECT}\\right)} $$",
    "2315417": "You are right, I’ve just fixed the typo. Thank you for pointing this out."
  },
  "source": "meta"
}