{
  "id": 70841,
  "title": "Including F1 metric in Keras",
  "url": "/competitions/quora-insincere-questions-classification/discussion/70841",
  "author_name": "",
  "post_date": "2018-11-07T18:38:49.482409Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://stackoverflow.com/questions/43547402/how-to-calculate-f1-macro-in-keras\">https://stackoverflow.com/questions/43547402/how-to-calculate-f1-macro-in-keras</a></p>",
  "messages": [
    {
      "id": "417097",
      "postDate": "11/07/2018 18:38:49",
      "content": "<p><a href=\"https://stackoverflow.com/questions/43547402/how-to-calculate-f1-macro-in-keras\">https://stackoverflow.com/questions/43547402/how-to-calculate-f1-macro-in-keras</a></p>",
      "rawMarkdown": "https://stackoverflow.com/questions/43547402/how-to-calculate-f1-macro-in-keras",
      "votes": null
    },
    {
      "id": "417118",
      "postDate": "11/07/2018 19:34:11",
      "content": "<p>In short:</p>\n\n<pre><code>def competitionMetric(true,pred): #considering sigmoid activation, threshold = 0.5\n    pred = K.cast(K.greater(pred,0.5), K.floatx())\n\n    groundPositives = K.sum(true) + K.epsilon()\n    correctPositives = K.sum(true * pred) + K.epsilon()\n    predictedPositives = K.sum(pred) + K.epsilon()\n\n    precision = correctPositives / predictedPositives\n    recall = correctPositives / groundPositives\n\n    m = (2 * precision * recall) / (precision + recall)\n\n    return m\n</code></pre>\n\n<p><strong>Warning:</strong> this is a metric per batch that collapses the \"samples\" dimension. If you need a metric for the whole data, you will need to get metrics for \"ground positive sum\", \"correct positive sum\" and \"predicted sum\" separately and calculate the final metric in a <code>LambdaCallback(on_epoch_end:......)</code> with these metrics. </p>",
      "rawMarkdown": "In short:\n\n\tdef competitionMetric(true,pred): #considering sigmoid activation, threshold = 0.5\n\t\tpred = K.cast(K.greater(pred,0.5), K.floatx())\n\t\t\n\t\tgroundPositives = K.sum(true) + K.epsilon()\n\t\tcorrectPositives = K.sum(true * pred) + K.epsilon()\n\t\tpredictedPositives = K.sum(pred) + K.epsilon()\n\n\t\tprecision = correctPositives / predictedPositives\n\t\trecall = correctPositives / groundPositives\n\n\t\tm = (2 * precision * recall) / (precision + recall)\n\n\t\treturn m\n\n**Warning:** this is a metric per batch that collapses the \"samples\" dimension. If you need a metric for the whole data, you will need to get metrics for \"ground positive sum\", \"correct positive sum\" and \"predicted sum\" separately and calculate the final metric in a `LambdaCallback(on_epoch_end:......)` with these metrics.",
      "votes": null
    },
    {
      "id": "417224",
      "postDate": "11/08/2018 00:50:16",
      "content": "<p>In my experience if you want f1/metrics that don't work when averaging over multiple batches, you should implement your own training loop with Keras' train_on_batch function</p>",
      "rawMarkdown": "In my experience if you want f1/metrics that don't work when averaging over multiple batches, you should implement your own training loop with Keras' train_on_batch function",
      "votes": null
    },
    {
      "id": "417365",
      "postDate": "11/08/2018 06:38:57",
      "content": "<p>I tried using F1 score (from stackoverflow question you mentioned) instead of accuracy metric in Keras but I found that results are getting worse (in public Leaderboard at least). Do you think there is some reason for that?</p>",
      "rawMarkdown": "I tried using F1 score (from stackoverflow question you mentioned) instead of accuracy metric in Keras but I found that results are getting worse (in public Leaderboard at least). Do you think there is some reason for that?",
      "votes": null
    },
    {
      "id": "417380",
      "postDate": "11/08/2018 07:06:59",
      "content": "<p>I have used the same F1 metric implementation with bi-lstm Keras. Was able to get accuracy ~0.96 for validation data and F1 metric ~0.65.  Include both accuracy and f1 in metrics as below\nmetrics=['accuracy', f1].  Maybe you need to finetune your model and choose the threshold wisely(0.5 is not always the best threshold for binary classification tasks).</p>",
      "rawMarkdown": "I have used the same F1 metric implementation with bi-lstm Keras. Was able to get accuracy ~0.96 for validation data and F1 metric ~0.65.  Include both accuracy and f1 in metrics as below\nmetrics=['accuracy', f1].  Maybe you need to finetune your model and choose the threshold wisely(0.5 is not always the best threshold for binary classification tasks).",
      "votes": null
    },
    {
      "id": "418622",
      "postDate": "11/10/2018 09:23:52",
      "content": "<p>Threshold should be half of the F1 score, so in this competition it should be ~0.34.</p>",
      "rawMarkdown": "Threshold should be half of the F1 score, so in this competition it should be ~0.34.",
      "votes": null
    },
    {
      "id": "418717",
      "postDate": "11/10/2018 13:40:07",
      "content": "<p>I'm also facing some problems with these metrics. </p>\n\n<p>Different models with significantly different F1 scores are not performing proportionally in the leaderbord.</p>\n\n<p>Models with around .5 F1 are way better than models with .62, I'm not sure about what is going on here. </p>",
      "rawMarkdown": "I'm also facing some problems with these metrics. \n\nDifferent models with significantly different F1 scores are not performing proportionally in the leaderbord.\n\nModels with around .5 F1 are way better than models with .62, I'm not sure about what is going on here.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 417118,
      "author_name": "danmoller",
      "author_url": "",
      "post_date": "11/07/2018 19:34:11",
      "content": "<p>In short:</p>\n\n<pre><code>def competitionMetric(true,pred): #considering sigmoid activation, threshold = 0.5\n    pred = K.cast(K.greater(pred,0.5), K.floatx())\n\n    groundPositives = K.sum(true) + K.epsilon()\n    correctPositives = K.sum(true * pred) + K.epsilon()\n    predictedPositives = K.sum(pred) + K.epsilon()\n\n    precision = correctPositives / predictedPositives\n    recall = correctPositives / groundPositives\n\n    m = (2 * precision * recall) / (precision + recall)\n\n    return m\n</code></pre>\n\n<p><strong>Warning:</strong> this is a metric per batch that collapses the \"samples\" dimension. If you need a metric for the whole data, you will need to get metrics for \"ground positive sum\", \"correct positive sum\" and \"predicted sum\" separately and calculate the final metric in a <code>LambdaCallback(on_epoch_end:......)</code> with these metrics. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417224,
      "author_name": "msf908",
      "author_url": "",
      "post_date": "11/08/2018 00:50:16",
      "content": "<p>In my experience if you want f1/metrics that don't work when averaging over multiple batches, you should implement your own training loop with Keras' train_on_batch function</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 417365,
      "author_name": "demonplus",
      "author_url": "",
      "post_date": "11/08/2018 06:38:57",
      "content": "<p>I tried using F1 score (from stackoverflow question you mentioned) instead of accuracy metric in Keras but I found that results are getting worse (in public Leaderboard at least). Do you think there is some reason for that?</p>",
      "votes": null,
      "replies": [
        {
          "id": 417380,
          "author_name": "rajeshbhat",
          "author_url": "",
          "post_date": "11/08/2018 07:06:59",
          "content": "<p>I have used the same F1 metric implementation with bi-lstm Keras. Was able to get accuracy ~0.96 for validation data and F1 metric ~0.65.  Include both accuracy and f1 in metrics as below\nmetrics=['accuracy', f1].  Maybe you need to finetune your model and choose the threshold wisely(0.5 is not always the best threshold for binary classification tasks).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418622,
          "author_name": "ffedericoni",
          "author_url": "",
          "post_date": "11/10/2018 09:23:52",
          "content": "<p>Threshold should be half of the F1 score, so in this competition it should be ~0.34.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 418717,
          "author_name": "danmoller",
          "author_url": "",
          "post_date": "11/10/2018 13:40:07",
          "content": "<p>I'm also facing some problems with these metrics. </p>\n\n<p>Different models with significantly different F1 scores are not performing proportionally in the leaderbord.</p>\n\n<p>Models with around .5 F1 are way better than models with .62, I'm not sure about what is going on here. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "417097": "https://stackoverflow.com/questions/43547402/how-to-calculate-f1-macro-in-keras",
    "417118": "In short:\n\n\tdef competitionMetric(true,pred): #considering sigmoid activation, threshold = 0.5\n\t\tpred = K.cast(K.greater(pred,0.5), K.floatx())\n\t\t\n\t\tgroundPositives = K.sum(true) + K.epsilon()\n\t\tcorrectPositives = K.sum(true * pred) + K.epsilon()\n\t\tpredictedPositives = K.sum(pred) + K.epsilon()\n\n\t\tprecision = correctPositives / predictedPositives\n\t\trecall = correctPositives / groundPositives\n\n\t\tm = (2 * precision * recall) / (precision + recall)\n\n\t\treturn m\n\n**Warning:** this is a metric per batch that collapses the \"samples\" dimension. If you need a metric for the whole data, you will need to get metrics for \"ground positive sum\", \"correct positive sum\" and \"predicted sum\" separately and calculate the final metric in a `LambdaCallback(on_epoch_end:......)` with these metrics.",
    "417224": "In my experience if you want f1/metrics that don't work when averaging over multiple batches, you should implement your own training loop with Keras' train_on_batch function",
    "417365": "I tried using F1 score (from stackoverflow question you mentioned) instead of accuracy metric in Keras but I found that results are getting worse (in public Leaderboard at least). Do you think there is some reason for that?",
    "417380": "I have used the same F1 metric implementation with bi-lstm Keras. Was able to get accuracy ~0.96 for validation data and F1 metric ~0.65.  Include both accuracy and f1 in metrics as below\nmetrics=['accuracy', f1].  Maybe you need to finetune your model and choose the threshold wisely(0.5 is not always the best threshold for binary classification tasks).",
    "418622": "Threshold should be half of the F1 score, so in this competition it should be ~0.34.",
    "418717": "I'm also facing some problems with these metrics. \n\nDifferent models with significantly different F1 scores are not performing proportionally in the leaderbord.\n\nModels with around .5 F1 are way better than models with .62, I'm not sure about what is going on here."
  },
  "source": "meta"
}