{
  "id": 384315,
  "title": "Confused about competition metric",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/384315",
  "author_name": "Alvor",
  "post_date": "2023-02-07T13:15:45.747000",
  "votes": 27,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Submission with a constant value of <strong>1</strong> scores <strong>0.414</strong> on the public leaderboard and this confuses me a little. </p>\n<p>Let's say test set has <strong>S</strong> samples = <strong>P</strong> positives + <strong>N</strong> negatives. Then our constant submission has <br>\ntp = P, fp = N, fn = 0. Then<br>\n$$<br>\nF_1 = \\frac{2tp}{2tp+fp+fn} = \\frac{2P}{2P+N+0} = \\frac{2P}{P+S} = \\frac{2\\frac{P}{S}}{\\frac{P}{S}+1}=0.414,<br>\n$$<br>\nhence<br>\n$$<br>\n\\frac{P}{S}\\approx0.261<br>\n$$<br>\nwhich is much lower than the average proportion of positive samples in the training set. Where am I wrong?</p>\n<p><strong>UPD</strong> And the second strange thing for me. Submission with a constant value of <strong>0</strong> scores <strong>0.226</strong> on the LB.  Shouldn't the F1 metric equal zero when tp=0?</p>",
  "messages": [
    {
      "id": 2133525,
      "postDate": "2023-02-07T13:15:45.747Z",
      "content": "<p>Submission with a constant value of <strong>1</strong> scores <strong>0.414</strong> on the public leaderboard and this confuses me a little. </p>\n<p>Let's say test set has <strong>S</strong> samples = <strong>P</strong> positives + <strong>N</strong> negatives. Then our constant submission has <br>\ntp = P, fp = N, fn = 0. Then<br>\n$$<br>\nF_1 = \\frac{2tp}{2tp+fp+fn} = \\frac{2P}{2P+N+0} = \\frac{2P}{P+S} = \\frac{2\\frac{P}{S}}{\\frac{P}{S}+1}=0.414,<br>\n$$<br>\nhence<br>\n$$<br>\n\\frac{P}{S}\\approx0.261<br>\n$$<br>\nwhich is much lower than the average proportion of positive samples in the training set. Where am I wrong?</p>\n<p><strong>UPD</strong> And the second strange thing for me. Submission with a constant value of <strong>0</strong> scores <strong>0.226</strong> on the LB.  Shouldn't the F1 metric equal zero when tp=0?</p>",
      "rawMarkdown": "Submission with a constant value of **1** scores **0.414** on the public leaderboard and this confuses me a little. \n\nLet's say test set has **S** samples = **P** positives + **N** negatives. Then our constant submission has \ntp = P, fp = N, fn = 0. Then\n$$\nF_1 = \\frac{2tp}{2tp+fp+fn} = \\frac{2P}{2P+N+0} = \\frac{2P}{P+S} = \\frac{2\\frac{P}{S}}{\\frac{P}{S}+1}=0.414,\n$$\nhence\n$$\n\\frac{P}{S}\\approx0.261\n$$\nwhich is much lower than the average proportion of positive samples in the training set. Where am I wrong?\n\n**UPD** And the second strange thing for me. Submission with a constant value of **0** scores **0.226** on the LB.  Shouldn't the F1 metric equal zero when tp=0?",
      "votes": 27
    },
    {
      "id": 2133695,
      "postDate": "2023-02-07T15:02:12.713Z",
      "content": "<p>The competition metric is <code>sklearn.metrics.f1_score(true, pred, average='macro')</code>. We compute the <code>f1_score</code> for class <code>0</code> then compute the <code>f1_score</code> for class <code>1</code> then average the two equally. Blog explains this <a href=\"https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f#989c\" target=\"_blank\">here</a></p>",
      "rawMarkdown": "The competition metric is `sklearn.metrics.f1_score(true, pred, average='macro')`. We compute the `f1_score` for class `0` then compute the `f1_score` for class `1` then average the two equally. Blog explains this [here][1]\n\n[1]: https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f#989c",
      "votes": 25,
      "replies": [
        {
          "id": 2133710,
          "postDate": "2023-02-07T15:08:01.543Z",
          "content": "<p>Oh, thanks for clarifying <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I didn't find that in Competition description.</p>",
          "rawMarkdown": "Oh, thanks for clarifying @cdeotte I didn't find that in Competition description.",
          "votes": 3,
          "replies": [
            {
              "id": 2133714,
              "postDate": "2023-02-07T15:10:27.953Z",
              "content": "<p>I was confused too. I tried different <code>f1_score</code> metrics until I got validation score of all zeros to equal 0.226 and all ones to equal 0.414 to match LB results.</p>",
              "rawMarkdown": "I was confused too. I tried different `f1_score` metrics until I got validation score of all zeros to equal 0.226 and all ones to equal 0.414 to match LB results.",
              "votes": 11
            },
            {
              "id": 2134980,
              "postDate": "2023-02-08T11:24:23.207Z",
              "content": "<p>since f1_score macro is average of positive f1_score and negative f1_scores.</p>\n<ul>\n<li>For all one case, positive f1_score =&gt;  (1/2*2p)/(2p+n) = 0.414</li>\n<li>for all zero case, negative f1_score =&gt; (1/2*2n)/(2n+p) = 0.226<br>\nis my analysis wrong? since p/n is not getting same value from both submission scores</li>\n</ul>",
              "rawMarkdown": "since f1_score macro is average of positive f1_score and negative f1_scores.\n\n- For all one case, positive f1_score =>  (1/2*2p)/(2p+n) = 0.414\n- for all zero case, negative f1_score => (1/2*2n)/(2n+p) = 0.226\nis my analysis wrong? since p/n is not getting same value from both submission scores"
            },
            {
              "id": 2135241,
              "postDate": "2023-02-08T14:24:52.343Z",
              "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> LB truncates scores, so<br>\n$$<br>\n0.414\\leq\\frac{P}{2P+N}&lt;0.415<br>\n$$<br>\nand<br>\n$$<br>\n0.226\\leq\\frac{N}{2N+P}&lt;0.227<br>\n$$<br>\nhence<br>\n$$<br>\n2.406976&lt;\\frac{P}{N}&lt;2.441176<br>\n$$<br>\nand<br>\n$$<br>\n2.405286&lt;\\frac{P}{N}&lt;2.424778<br>\n$$</p>\n<p>Finally, <br>\n$$<br>\n2.406976&lt;\\frac{P}{N}&lt;2.424778<br>\n$$ </p>",
              "rawMarkdown": "@seshurajup LB truncates scores, so\n$$\n0.414\\leq\\frac{P}{2P+N}<0.415\n$$\nand\n$$\n0.226\\leq\\frac{N}{2N+P}<0.227\n$$\nhence\n$$\n2.406976<\\frac{P}{N}<2.441176\n$$\nand\n$$\n2.405286<\\frac{P}{N}<2.424778\n$$\n\nFinally, \n$$\n2.406976<\\frac{P}{N}<2.424778\n$$ ",
              "votes": 6
            },
            {
              "id": 2135365,
              "postDate": "2023-02-08T15:34:47.247Z",
              "content": "<p><a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a> Thanks for your explanation. Now it makes sense more clear.</p>",
              "rawMarkdown": "@allvor Thanks for your explanation. Now it makes sense more clear."
            }
          ]
        },
        {
          "id": 2134322,
          "postDate": "2023-02-07T22:26:34.480Z",
          "content": "<p>Great find Chris!</p>",
          "rawMarkdown": "Great find Chris!",
          "votes": 2
        },
        {
          "id": 2134551,
          "postDate": "2023-02-08T04:53:15.227Z",
          "content": "<p>great blog, thank for sharing!</p>",
          "rawMarkdown": "great blog, thank for sharing!",
          "votes": 2
        },
        {
          "id": 2146758,
          "postDate": "2023-02-16T05:42:35.743Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2140067,
      "postDate": "2023-02-11T12:34:24.667Z",
      "content": "<p>here is simple code to calculate this metric; it is a lot faster than sklearn.metrics.f1_score.<br>\ndef f1l(labels, preds): # my calc of macro F1 - way faster<br>\n     c11 = (preds + labels == 2).mean()<br>\n     c00 = (preds + labels == 0).mean()<br>\n     e = 1 - (1 - c00 - c11) / (1 - (c00 - c11)**2)<br>\n     return e</p>",
      "rawMarkdown": "here is simple code to calculate this metric; it is a lot faster than sklearn.metrics.f1_score.\ndef f1l(labels, preds): # my calc of macro F1 - way faster\n     c11 = (preds + labels == 2).mean()\n     c00 = (preds + labels == 0).mean()\n     e = 1 - (1 - c00 - c11) / (1 - (c00 - c11)**2)\n     return e",
      "votes": 5,
      "replies": [
        {
          "id": 2154230,
          "postDate": "2023-02-21T22:33:31.563Z",
          "content": "<p>Here is above code implemented as CatboostClassifier custom metric</p>\n<pre><code> (): \n    c11 = (preds + labels == ).mean()\n    c00 = (preds + labels == ).mean()\n    e =  - ( - c00 - c11) / ( - (c00 - c11)**)\n     e\n\n ():\n    threshold = \n     ():\n         error\n\n     ():\n         \n\n     () -&gt; [, ]:\n        approx = approxes[]\n        approx = [ / ( + np.exp(-x))  x  approx]\n        pred_labels = [  x &gt; threshold    x  approx]\n        f1 = f1_avg(pred_labels,target)\n         f1, \n</code></pre>",
          "rawMarkdown": "Here is above code implemented as CatboostClassifier custom metric\n\n\n```python\ndef f1_avg(preds, labels): \n    c11 = (preds + labels == 2).mean()\n    c00 = (preds + labels == 0).mean()\n    e = 1 - (1 - c00 - c11) / (1 - (c00 - c11)**2)\n    return e\n\nclass F1Metric(object):\n    threshold = 0.5\n    def get_final_error(self, error, weight):\n        return error\n\n    def is_max_optimal(self):\n        return True\n\n    def evaluate(self, approxes, target, weight) -> Tuple[float, float]:\n        approx = approxes[0]\n        approx = [1 / (1 + np.exp(-x)) for x in approx]\n        pred_labels = [1.0 if x > threshold else 0.0 for x in approx]\n        f1 = f1_avg(pred_labels,target)\n        return f1, 1.0\n\n```\n",
          "votes": 2,
          "replies": [
            {
              "id": 2154955,
              "postDate": "2023-02-22T10:49:58.187Z",
              "content": "<p>I am still struggling with numba warnings. I tried annotating it with <a href=\"https://www.kaggle.com/njit\" target=\"_blank\">@njit</a> but it keeps having some numba errors. If anybody finds a solution I will highly appreciate it.</p>",
              "rawMarkdown": "I am still struggling with numba warnings. I tried annotating it with @njit but it keeps having some numba errors. If anybody finds a solution I will highly appreciate it."
            }
          ]
        }
      ]
    },
    {
      "id": 2133687,
      "postDate": "2023-02-07T14:59:24.740Z",
      "content": "<p>Looks like a mood-enhancing bonus: score+=0.2 </p>",
      "rawMarkdown": "Looks like a mood-enhancing bonus: score+=0.2 ",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 2133695,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-02-07T15:02:12.713000",
      "content": "<p>The competition metric is <code>sklearn.metrics.f1_score(true, pred, average='macro')</code>. We compute the <code>f1_score</code> for class <code>0</code> then compute the <code>f1_score</code> for class <code>1</code> then average the two equally. Blog explains this <a href=\"https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f#989c\" target=\"_blank\">here</a></p>",
      "votes": 25,
      "replies": [
        {
          "id": 2133710,
          "author_name": "Alvor",
          "author_url": "",
          "post_date": "2023-02-07T15:08:01.543000",
          "content": "<p>Oh, thanks for clarifying <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> I didn't find that in Competition description.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2133714,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-02-07T15:10:27.953000",
              "content": "<p>I was confused too. I tried different <code>f1_score</code> metrics until I got validation score of all zeros to equal 0.226 and all ones to equal 0.414 to match LB results.</p>",
              "votes": 11,
              "replies": []
            },
            {
              "id": 2134980,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2023-02-08T11:24:23.207000",
              "content": "<p>since f1_score macro is average of positive f1_score and negative f1_scores.</p>\n<ul>\n<li>For all one case, positive f1_score =&gt;  (1/2*2p)/(2p+n) = 0.414</li>\n<li>for all zero case, negative f1_score =&gt; (1/2*2n)/(2n+p) = 0.226<br>\nis my analysis wrong? since p/n is not getting same value from both submission scores</li>\n</ul>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2135241,
              "author_name": "Alvor",
              "author_url": "",
              "post_date": "2023-02-08T14:24:52.343000",
              "content": "<p><a href=\"https://www.kaggle.com/seshurajup\" target=\"_blank\">@seshurajup</a> LB truncates scores, so<br>\n$$<br>\n0.414\\leq\\frac{P}{2P+N}&lt;0.415<br>\n$$<br>\nand<br>\n$$<br>\n0.226\\leq\\frac{N}{2N+P}&lt;0.227<br>\n$$<br>\nhence<br>\n$$<br>\n2.406976&lt;\\frac{P}{N}&lt;2.441176<br>\n$$<br>\nand<br>\n$$<br>\n2.405286&lt;\\frac{P}{N}&lt;2.424778<br>\n$$</p>\n<p>Finally, <br>\n$$<br>\n2.406976&lt;\\frac{P}{N}&lt;2.424778<br>\n$$ </p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 2135365,
              "author_name": "SeshuRaju 🧘‍♂️",
              "author_url": "",
              "post_date": "2023-02-08T15:34:47.247000",
              "content": "<p><a href=\"https://www.kaggle.com/allvor\" target=\"_blank\">@allvor</a> Thanks for your explanation. Now it makes sense more clear.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2134322,
          "author_name": "Mayukh Bhattacharyya",
          "author_url": "",
          "post_date": "2023-02-07T22:26:34.480000",
          "content": "<p>Great find Chris!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2134551,
          "author_name": "Tim Yee",
          "author_url": "",
          "post_date": "2023-02-08T04:53:15.227000",
          "content": "<p>great blog, thank for sharing!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 2146758,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-02-16T05:42:35.743000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2140067,
      "author_name": "Youri Matiounine",
      "author_url": "",
      "post_date": "2023-02-11T12:34:24.667000",
      "content": "<p>here is simple code to calculate this metric; it is a lot faster than sklearn.metrics.f1_score.<br>\ndef f1l(labels, preds): # my calc of macro F1 - way faster<br>\n     c11 = (preds + labels == 2).mean()<br>\n     c00 = (preds + labels == 0).mean()<br>\n     e = 1 - (1 - c00 - c11) / (1 - (c00 - c11)**2)<br>\n     return e</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2154230,
          "author_name": "Maciej Król",
          "author_url": "",
          "post_date": "2023-02-21T22:33:31.563000",
          "content": "<p>Here is above code implemented as CatboostClassifier custom metric</p>\n<pre><code> (): \n    c11 = (preds + labels == ).mean()\n    c00 = (preds + labels == ).mean()\n    e =  - ( - c00 - c11) / ( - (c00 - c11)**)\n     e\n\n ():\n    threshold = \n     ():\n         error\n\n     ():\n         \n\n     () -&gt; [, ]:\n        approx = approxes[]\n        approx = [ / ( + np.exp(-x))  x  approx]\n        pred_labels = [  x &gt; threshold    x  approx]\n        f1 = f1_avg(pred_labels,target)\n         f1, \n</code></pre>",
          "votes": 2,
          "replies": [
            {
              "id": 2154955,
              "author_name": "Maciej Król",
              "author_url": "",
              "post_date": "2023-02-22T10:49:58.187000",
              "content": "<p>I am still struggling with numba warnings. I tried annotating it with <a href=\"https://www.kaggle.com/njit\" target=\"_blank\">@njit</a> but it keeps having some numba errors. If anybody finds a solution I will highly appreciate it.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2133687,
      "author_name": "Wojtek Rosa",
      "author_url": "",
      "post_date": "2023-02-07T14:59:24.740000",
      "content": "<p>Looks like a mood-enhancing bonus: score+=0.2 </p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2133525": "Submission with a constant value of **1** scores **0.414** on the public leaderboard and this confuses me a little. \n\nLet's say test set has **S** samples = **P** positives + **N** negatives. Then our constant submission has \ntp = P, fp = N, fn = 0. Then\n$$\nF_1 = \\frac{2tp}{2tp+fp+fn} = \\frac{2P}{2P+N+0} = \\frac{2P}{P+S} = \\frac{2\\frac{P}{S}}{\\frac{P}{S}+1}=0.414,\n$$\nhence\n$$\n\\frac{P}{S}\\approx0.261\n$$\nwhich is much lower than the average proportion of positive samples in the training set. Where am I wrong?\n\n**UPD** And the second strange thing for me. Submission with a constant value of **0** scores **0.226** on the LB.  Shouldn't the F1 metric equal zero when tp=0?",
    "2133695": "The competition metric is `sklearn.metrics.f1_score(true, pred, average='macro')`. We compute the `f1_score` for class `0` then compute the `f1_score` for class `1` then average the two equally. Blog explains this [here][1]\n\n[1]: https://towardsdatascience.com/micro-macro-weighted-averages-of-f1-score-clearly-explained-b603420b292f#989c",
    "2140067": "here is simple code to calculate this metric; it is a lot faster than sklearn.metrics.f1_score.\ndef f1l(labels, preds): # my calc of macro F1 - way faster\n     c11 = (preds + labels == 2).mean()\n     c00 = (preds + labels == 0).mean()\n     e = 1 - (1 - c00 - c11) / (1 - (c00 - c11)**2)\n     return e",
    "2133687": "Looks like a mood-enhancing bonus: score+=0.2 "
  }
}