{
  "id": 384677,
  "title": "Using F1 Score as Metric for Early Stopping",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/384677",
  "author_name": "Mohamed Eltayeb",
  "post_date": "2023-02-08T22:50:26.248000",
  "votes": 3,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Hello <br>\nI tried to use f1 score as a metric when I set the early stopping of LGBM. I used the following function:</p>\n<pre><code> ():\n  preds = sigmoid(preds)\n  binary_preds = [(p&gt;)  p  preds]\n  y_true = lgbDataset.get_label()\n   , f1_score(y_true, binary_preds,  average = ), \n</code></pre>\n<p>But I found it to give much worse cv (0.65) instead of (0.671). Just want to know if I am missing something?</p>",
  "messages": [
    {
      "id": 2135879,
      "postDate": "2023-02-08T23:23:03.703Z",
      "content": "<p>Perhaps try a different threshold here</p>\n<pre><code>binary_preds = [int(p&gt;0.5) for p in preds]\n</code></pre>\n<p>A good threshold is 0.63</p>",
      "rawMarkdown": "Perhaps try a different threshold here\n\n    binary_preds = [int(p>0.5) for p in preds]\n\nA good threshold is 0.63",
      "votes": 3,
      "replies": [
        {
          "id": 2136331,
          "postDate": "2023-02-09T08:52:15.797Z",
          "content": "<p>Thanks. I think this was the case. I tried several thresholds and they improved the score but didn't go beyond 0.670. So, maybe not a good idea to use it anyway.</p>",
          "rawMarkdown": "Thanks. I think this was the case. I tried several thresholds and they improved the score but didn't go beyond 0.670. So, maybe not a good idea to use it anyway.",
          "votes": 1
        },
        {
          "id": 2136340,
          "postDate": "2023-02-09T09:07:45.290Z",
          "content": "<p>Do you think the threshold is model independant ? (for well calibrated models)<br>\nquestion independant ?</p>",
          "rawMarkdown": "Do you think the threshold is model independant ? (for well calibrated models)\nquestion independant ?",
          "votes": 1,
          "replies": [
            {
              "id": 2136916,
              "postDate": "2023-02-09T16:22:08.587Z",
              "content": "<p>It is question independent at least.</p>\n<p>Took me awhile to understand the multilabel 'macro' F1 score, eventually realized that the 'macro' averaging wasn't per question, but was for 1 as the 'positive' label (TP if ground truth and prediction both 1) averaged with 0 as the 'positive' label (TP if ground truth and prediction both 0). </p>\n<p>Anyways, point being I found that all questions are combined before calculating the F1 scores. So, logically, there's a single optimal threshold, which will vary by the global data distribution. The biggest single factor would be the global percentage of true 1s vs true 0s. I don't think(?) that would be the only factor, though, so a very confident model with really high precision on 1s but less precise on 0s would probably skew the optimal threshold</p>",
              "rawMarkdown": "It is question independent at least.\n\nTook me awhile to understand the multilabel 'macro' F1 score, eventually realized that the 'macro' averaging wasn't per question, but was for 1 as the 'positive' label (TP if ground truth and prediction both 1) averaged with 0 as the 'positive' label (TP if ground truth and prediction both 0). \n\nAnyways, point being I found that all questions are combined before calculating the F1 scores. So, logically, there's a single optimal threshold, which will vary by the global data distribution. The biggest single factor would be the global percentage of true 1s vs true 0s. I don't think(?) that would be the only factor, though, so a very confident model with really high precision on 1s but less precise on 0s would probably skew the optimal threshold",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 2135862,
      "postDate": "2023-02-08T22:50:26.250Z",
      "content": "<p>Hello <br>\nI tried to use f1 score as a metric when I set the early stopping of LGBM. I used the following function:</p>\n<pre><code> ():\n  preds = sigmoid(preds)\n  binary_preds = [(p&gt;)  p  preds]\n  y_true = lgbDataset.get_label()\n   , f1_score(y_true, binary_preds,  average = ), \n</code></pre>\n<p>But I found it to give much worse cv (0.65) instead of (0.671). Just want to know if I am missing something?</p>",
      "rawMarkdown": "Hello \nI tried to use f1 score as a metric when I set the early stopping of LGBM. I used the following function:\n```python\ndef lgb_f1_score(preds, lgbDataset):\n  preds = sigmoid(preds)\n  binary_preds = [int(p>0.5) for p in preds]\n  y_true = lgbDataset.get_label()\n  return 'f1', f1_score(y_true, binary_preds,  average = 'macro'), True\n```\n\nBut I found it to give much worse cv (0.65) instead of (0.671). Just want to know if I am missing something?",
      "votes": 3
    },
    {
      "id": 2136962,
      "postDate": "2023-02-09T16:38:41.480Z",
      "content": "<p>Another option besides the computationally efficient preset threshold suggested by Chris is to calculate the exact optimal threshold each time, inside the lgb_f1_score function itself. </p>\n<p>Found some hints about how to do it online, but didn't find anything that looked cut and paste. What I found related to using precision_recall_curve function. But I'm unsure it would be very efficient, and you'd have to call it twice, once with pos_label=0. </p>\n<p>If I end up implementing it myself, I'll share. Sounds like a fun mini-challenge. :)</p>\n<p>UPDATE: Implemented! See link in replies to this comment :)</p>",
      "rawMarkdown": "Another option besides the computationally efficient preset threshold suggested by Chris is to calculate the exact optimal threshold each time, inside the lgb_f1_score function itself. \n\nFound some hints about how to do it online, but didn't find anything that looked cut and paste. What I found related to using precision_recall_curve function. But I'm unsure it would be very efficient, and you'd have to call it twice, once with pos_label=0. \n\nIf I end up implementing it myself, I'll share. Sounds like a fun mini-challenge. :)\n\nUPDATE: Implemented! See link in replies to this comment :)",
      "votes": 1,
      "replies": [
        {
          "id": 2137116,
          "postDate": "2023-02-09T18:31:25.383Z",
          "content": "<p>I think I will try to implement this as well. Looks interesting! <br>\nJust could I know why we need to call it with pos_label=0? Not sure I understood this well.</p>",
          "rawMarkdown": "I think I will try to implement this as well. Looks interesting! \nJust could I know why we need to call it with pos_label=0? Not sure I understood this well.",
          "votes": 1,
          "replies": [
            {
              "id": 2137146,
              "postDate": "2023-02-09T18:47:34.073Z",
              "content": "<p>'macro' f1_score can be thought of as the mean of binary F1 score for pos_label=[each possible label]</p>\n<p>In our case, that's pos_label=0 and pos_label=1. The precision_recall_curve function is the same, but doesn't have the built in 'macro' option. And with pos_label=1 being the default, we just need to average the default with the os_label=0 case. </p>\n<p>The following code I got from <a href=\"https://stats.stackexchange.com/questions/518616/how-to-find-the-optimal-threshold-for-the-weighted-f1-score-in-a-binary-classifi\" target=\"_blank\">https://stats.stackexchange.com/questions/518616/how-to-find-the-optimal-threshold-for-the-weighted-f1-score-in-a-binary-classifi</a> , it can probably be adapted but with larger prediction sets calculating the entire curve (twice) is overkill and might be quite slow in terms of performance. </p>\n<pre><code>precision, recall, thresholds = precision_recall_curve(y_true, y_score)\nf1_scores = *recall*precision/(recall+precision)\n(, thresholds[np.argmax(f1_scores)])\n(, np.(f1_scores))\n</code></pre>",
              "rawMarkdown": "'macro' f1_score can be thought of as the mean of binary F1 score for pos_label=[each possible label]\n\nIn our case, that's pos_label=0 and pos_label=1. The precision_recall_curve function is the same, but doesn't have the built in 'macro' option. And with pos_label=1 being the default, we just need to average the default with the os_label=0 case. \n\nThe following code I got from https://stats.stackexchange.com/questions/518616/how-to-find-the-optimal-threshold-for-the-weighted-f1-score-in-a-binary-classifi , it can probably be adapted but with larger prediction sets calculating the entire curve (twice) is overkill and might be quite slow in terms of performance. \n\n```python\nprecision, recall, thresholds = precision_recall_curve(y_true, y_score)\nf1_scores = 2*recall*precision/(recall+precision)\nprint('Best threshold: ', thresholds[np.argmax(f1_scores)])\nprint('Best F1-Score: ', np.max(f1_scores))\n```",
              "votes": 3
            },
            {
              "id": 2137374,
              "postDate": "2023-02-09T23:24:05.463Z",
              "content": "<p>Implemented here! :)</p>\n<p><a href=\"https://www.kaggle.com/code/roberthatch/optimal-f1-score\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/optimal-f1-score</a></p>",
              "rawMarkdown": "Implemented here! :)\n\nhttps://www.kaggle.com/code/roberthatch/optimal-f1-score",
              "votes": 2
            },
            {
              "id": 2137876,
              "postDate": "2023-02-10T11:46:57.070Z",
              "content": "<p>Ok nice!<br>\nI gave it a try with the LGBM notebook. I received some errors. Firstly, I got a division by zero error, so I added a small number to the denominator of the f1 scores:<br>\n<code>f1_scores = 2*recall*precision/(recall+precision + 0.0001)</code>  <br>\nAlso, not sure why this bug happens but the negative indexing with slicing didn't work for me so I wrote this as follows:<br>\n<code>f1_scores = f1_scores[:f1_scores.shape[0]-b]</code></p>\n<p>Btw, in the cv I received an improvement (0.65 --&gt; 0.661), but it still below the logloss score (0.671).</p>",
              "rawMarkdown": "Ok nice!\nI gave it a try with the LGBM notebook. I received some errors. Firstly, I got a division by zero error, so I added a small number to the denominator of the f1 scores:\n`f1_scores = 2*recall*precision/(recall+precision + 0.0001)`  \nAlso, not sure why this bug happens but the negative indexing with slicing didn't work for me so I wrote this as follows:\n`f1_scores = f1_scores[:f1_scores.shape[0]-b]`\n\nBtw, in the cv I received an improvement (0.65 --> 0.661), but it still below the logloss score (0.671).",
              "votes": 1
            },
            {
              "id": 2138081,
              "postDate": "2023-02-10T14:37:27.840Z",
              "content": "<p>Interesting. Sometimes early stopping on logloss can just work better. Another possibility is just more variance, if you increase the early_stopping_rounds you can see if that helps. For example from 100 to 500. </p>",
              "rawMarkdown": "Interesting. Sometimes early stopping on logloss can just work better. Another possibility is just more variance, if you increase the early_stopping_rounds you can see if that helps. For example from 100 to 500. ",
              "votes": 1
            },
            {
              "id": 2138102,
              "postDate": "2023-02-10T14:57:43.683Z",
              "content": "<p>Note that optimizing each target F1 separately will produce worse results than optimizing the entire F1 together. So if you're training different models per target, there is no way to stop the first targets you train without having the later targets yet.</p>",
              "rawMarkdown": "Note that optimizing each target F1 separately will produce worse results than optimizing the entire F1 together. So if you're training different models per target, there is no way to stop the first targets you train without having the later targets yet.",
              "votes": 5
            },
            {
              "id": 2138115,
              "postDate": "2023-02-10T15:08:31.017Z",
              "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> I increased the rounds from 200 to 600, and the num_iteartions from 3000 to 6000. The improvement was from 0.661 --&gt; 0.667! Though it didn't surpass the logloss (0.671), I think it will be a very good component in ensembling later.</p>",
              "rawMarkdown": "@roberthatch I increased the rounds from 200 to 600, and the num_iteartions from 3000 to 6000. The improvement was from 0.661 --> 0.667! Though it didn't surpass the logloss (0.671), I think it will be a very good component in ensembling later.",
              "votes": 2
            },
            {
              "id": 2139549,
              "postDate": "2023-02-10T21:11:22.197Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2135879,
      "author_name": "Chris Deotte",
      "author_url": "",
      "post_date": "2023-02-08T23:23:03.703000",
      "content": "<p>Perhaps try a different threshold here</p>\n<pre><code>binary_preds = [int(p&gt;0.5) for p in preds]\n</code></pre>\n<p>A good threshold is 0.63</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2136331,
          "author_name": "Mohamed Eltayeb",
          "author_url": "",
          "post_date": "2023-02-09T08:52:15.797000",
          "content": "<p>Thanks. I think this was the case. I tried several thresholds and they improved the score but didn't go beyond 0.670. So, maybe not a good idea to use it anyway.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2136340,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2023-02-09T09:07:45.290000",
          "content": "<p>Do you think the threshold is model independant ? (for well calibrated models)<br>\nquestion independant ?</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2136916,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-09T16:22:08.587000",
              "content": "<p>It is question independent at least.</p>\n<p>Took me awhile to understand the multilabel 'macro' F1 score, eventually realized that the 'macro' averaging wasn't per question, but was for 1 as the 'positive' label (TP if ground truth and prediction both 1) averaged with 0 as the 'positive' label (TP if ground truth and prediction both 0). </p>\n<p>Anyways, point being I found that all questions are combined before calculating the F1 scores. So, logically, there's a single optimal threshold, which will vary by the global data distribution. The biggest single factor would be the global percentage of true 1s vs true 0s. I don't think(?) that would be the only factor, though, so a very confident model with really high precision on 1s but less precise on 0s would probably skew the optimal threshold</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2136962,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2023-02-09T16:38:41.480000",
      "content": "<p>Another option besides the computationally efficient preset threshold suggested by Chris is to calculate the exact optimal threshold each time, inside the lgb_f1_score function itself. </p>\n<p>Found some hints about how to do it online, but didn't find anything that looked cut and paste. What I found related to using precision_recall_curve function. But I'm unsure it would be very efficient, and you'd have to call it twice, once with pos_label=0. </p>\n<p>If I end up implementing it myself, I'll share. Sounds like a fun mini-challenge. :)</p>\n<p>UPDATE: Implemented! See link in replies to this comment :)</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2137116,
          "author_name": "Mohamed Eltayeb",
          "author_url": "",
          "post_date": "2023-02-09T18:31:25.383000",
          "content": "<p>I think I will try to implement this as well. Looks interesting! <br>\nJust could I know why we need to call it with pos_label=0? Not sure I understood this well.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2137146,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-09T18:47:34.073000",
              "content": "<p>'macro' f1_score can be thought of as the mean of binary F1 score for pos_label=[each possible label]</p>\n<p>In our case, that's pos_label=0 and pos_label=1. The precision_recall_curve function is the same, but doesn't have the built in 'macro' option. And with pos_label=1 being the default, we just need to average the default with the os_label=0 case. </p>\n<p>The following code I got from <a href=\"https://stats.stackexchange.com/questions/518616/how-to-find-the-optimal-threshold-for-the-weighted-f1-score-in-a-binary-classifi\" target=\"_blank\">https://stats.stackexchange.com/questions/518616/how-to-find-the-optimal-threshold-for-the-weighted-f1-score-in-a-binary-classifi</a> , it can probably be adapted but with larger prediction sets calculating the entire curve (twice) is overkill and might be quite slow in terms of performance. </p>\n<pre><code>precision, recall, thresholds = precision_recall_curve(y_true, y_score)\nf1_scores = *recall*precision/(recall+precision)\n(, thresholds[np.argmax(f1_scores)])\n(, np.(f1_scores))\n</code></pre>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2137374,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-09T23:24:05.463000",
              "content": "<p>Implemented here! :)</p>\n<p><a href=\"https://www.kaggle.com/code/roberthatch/optimal-f1-score\" target=\"_blank\">https://www.kaggle.com/code/roberthatch/optimal-f1-score</a></p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2137876,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2023-02-10T11:46:57.070000",
              "content": "<p>Ok nice!<br>\nI gave it a try with the LGBM notebook. I received some errors. Firstly, I got a division by zero error, so I added a small number to the denominator of the f1 scores:<br>\n<code>f1_scores = 2*recall*precision/(recall+precision + 0.0001)</code>  <br>\nAlso, not sure why this bug happens but the negative indexing with slicing didn't work for me so I wrote this as follows:<br>\n<code>f1_scores = f1_scores[:f1_scores.shape[0]-b]</code></p>\n<p>Btw, in the cv I received an improvement (0.65 --&gt; 0.661), but it still below the logloss score (0.671).</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2138081,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-02-10T14:37:27.840000",
              "content": "<p>Interesting. Sometimes early stopping on logloss can just work better. Another possibility is just more variance, if you increase the early_stopping_rounds you can see if that helps. For example from 100 to 500. </p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2138102,
              "author_name": "Chris Deotte",
              "author_url": "",
              "post_date": "2023-02-10T14:57:43.683000",
              "content": "<p>Note that optimizing each target F1 separately will produce worse results than optimizing the entire F1 together. So if you're training different models per target, there is no way to stop the first targets you train without having the later targets yet.</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2138115,
              "author_name": "Mohamed Eltayeb",
              "author_url": "",
              "post_date": "2023-02-10T15:08:31.017000",
              "content": "<p><a href=\"https://www.kaggle.com/roberthatch\" target=\"_blank\">@roberthatch</a> I increased the rounds from 200 to 600, and the num_iteartions from 3000 to 6000. The improvement was from 0.661 --&gt; 0.667! Though it didn't surpass the logloss (0.671), I think it will be a very good component in ensembling later.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2139549,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-02-10T21:11:22.197000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2135879": "Perhaps try a different threshold here\n\n    binary_preds = [int(p>0.5) for p in preds]\n\nA good threshold is 0.63",
    "2135862": "Hello \nI tried to use f1 score as a metric when I set the early stopping of LGBM. I used the following function:\n```python\ndef lgb_f1_score(preds, lgbDataset):\n  preds = sigmoid(preds)\n  binary_preds = [int(p>0.5) for p in preds]\n  y_true = lgbDataset.get_label()\n  return 'f1', f1_score(y_true, binary_preds,  average = 'macro'), True\n```\n\nBut I found it to give much worse cv (0.65) instead of (0.671). Just want to know if I am missing something?",
    "2136962": "Another option besides the computationally efficient preset threshold suggested by Chris is to calculate the exact optimal threshold each time, inside the lgb_f1_score function itself. \n\nFound some hints about how to do it online, but didn't find anything that looked cut and paste. What I found related to using precision_recall_curve function. But I'm unsure it would be very efficient, and you'd have to call it twice, once with pos_label=0. \n\nIf I end up implementing it myself, I'll share. Sounds like a fun mini-challenge. :)\n\nUPDATE: Implemented! See link in replies to this comment :)"
  }
}