{
  "id": 68242,
  "title": "Question about (macro) f1 metric",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/68242",
  "author_name": "",
  "post_date": "2018-10-10T18:04:01.584469Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<p>Currently, we're using the following function to track the macro f1 score of my Keras model:</p>\n\n<pre><code>from keras import backend as K\n\n    def f1(y_true, y_pred):\n        def recall(y_true, y_pred):\n              true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n              possible_positives = K.sum(K.round(K.clip(y_true, 0, 1)))\n              recall = true_positives / (possible_positives + K.epsilon())\n              return recall\n\n        def precision(y_true, y_pred):\n              true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n              predicted_positives = K.sum(K.round(K.clip(y_pred, 0, 1)))\n              precision = true_positives / (predicted_positives + K.epsilon())\n              return precision\n\n    precision = precision(y_true, y_pred)\n    recall = recall(y_true, y_pred)\n    return 2*((precision*recall)/(precision+recall+K.epsilon()))\n</code></pre>\n\n<p>However, the scores we get on the Kaggle leaderboard are much lower than the metric is reporting. Have we missed something and if so, how can we improve this metric to match Kaggle's metric?</p>",
  "messages": [
    {
      "id": "401793",
      "postDate": "10/10/2018 18:04:01",
      "content": "<p>Hi guys,</p>\n\n<p>Currently, we're using the following function to track the macro f1 score of my Keras model:</p>\n\n<pre><code>from keras import backend as K\n\n    def f1(y_true, y_pred):\n        def recall(y_true, y_pred):\n              true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n              possible_positives = K.sum(K.round(K.clip(y_true, 0, 1)))\n              recall = true_positives / (possible_positives + K.epsilon())\n              return recall\n\n        def precision(y_true, y_pred):\n              true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n              predicted_positives = K.sum(K.round(K.clip(y_pred, 0, 1)))\n              precision = true_positives / (predicted_positives + K.epsilon())\n              return precision\n\n    precision = precision(y_true, y_pred)\n    recall = recall(y_true, y_pred)\n    return 2*((precision*recall)/(precision+recall+K.epsilon()))\n</code></pre>\n\n<p>However, the scores we get on the Kaggle leaderboard are much lower than the metric is reporting. Have we missed something and if so, how can we improve this metric to match Kaggle's metric?</p>",
      "rawMarkdown": "Hi guys,\n\nCurrently, we're using the following function to track the macro f1 score of my Keras model:\n\n    from keras import backend as K\n    \n        def f1(y_true, y_pred):\n            def recall(y_true, y_pred):\n                  true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n                  possible_positives = K.sum(K.round(K.clip(y_true, 0, 1)))\n                  recall = true_positives / (possible_positives + K.epsilon())\n                  return recall\n        \n            def precision(y_true, y_pred):\n                  true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n                  predicted_positives = K.sum(K.round(K.clip(y_pred, 0, 1)))\n                  precision = true_positives / (predicted_positives + K.epsilon())\n                  return precision\n    \n        precision = precision(y_true, y_pred)\n        recall = recall(y_true, y_pred)\n        return 2*((precision*recall)/(precision+recall+K.epsilon()))\n\nHowever, the scores we get on the Kaggle leaderboard are much lower than the metric is reporting. Have we missed something and if so, how can we improve this metric to match Kaggle's metric?",
      "votes": null
    },
    {
      "id": "401837",
      "postDate": "10/10/2018 19:01:40",
      "content": "<p>I have the same result (local value bigger than LB), but not so much. You can compare your results with <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">scikit f1</a> and with this <a href=\"https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras\">variant</a>. Also, may be you have bad split strategy or something else. I don't see any errors.</p>",
      "rawMarkdown": "I have the same result (local value bigger than LB), but not so much. You can compare your results with [scikit f1][1] and with this [variant][2]. Also, may be you have bad split strategy or something else. I don't see any errors.\n\n\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\n  [2]: https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras",
      "votes": null
    },
    {
      "id": "401841",
      "postDate": "10/10/2018 19:03:55",
      "content": "<p>All global metrics are only approximated during batch-wise training, so you want to calculate them only at epoch's end.</p>\n\n<p><a href=\"https://bit.ly/2n8buDO\"><strong>Link 1</strong></a></p>\n\n<p><a href=\"https://github.com/keras-team/keras/issues/5794\"><strong>Link 2</strong></a></p>\n\n<p><a href=\"https://github.com/keras-team/keras/issues/6507\"><strong>Link 3</strong></a></p>\n\n<p>Yet another explanation for score discrepancy is that you could be overfitting.</p>",
      "rawMarkdown": "All global metrics are only approximated during batch-wise training, so you want to calculate them only at epoch's end.\n\n[__Link 1__](https://bit.ly/2n8buDO)\n\n[__Link 2__](https://github.com/keras-team/keras/issues/5794)\n\n[__Link 3__](https://github.com/keras-team/keras/issues/6507)\n\nYet another explanation for score discrepancy is that you could be overfitting.",
      "votes": null
    },
    {
      "id": "402584",
      "postDate": "10/12/2018 00:17:00",
      "content": "<p>Thanks for the information! I think overfitting is the problem.</p>",
      "rawMarkdown": "Thanks for the information! I think overfitting is the problem.",
      "votes": null
    },
    {
      "id": "403074",
      "postDate": "10/12/2018 20:26:40",
      "content": "<p>Thank you for your insights! I think the validation set I used was far too small.</p>",
      "rawMarkdown": "Thank you for your insights! I think the validation set I used was far too small.",
      "votes": null
    },
    {
      "id": "403128",
      "postDate": "10/12/2018 23:54:45",
      "content": "<p>I've also seen much lower scores in my hold out validation dataset and the scores from the test dataset.\nIt may be just that that the test dataset is a bit on the small side, especially since only 29% (out of curiosity, how did you come up with that number?) is visible right now.</p>",
      "rawMarkdown": "I've also seen much lower scores in my hold out validation dataset and the scores from the test dataset.\nIt may be just that that the test dataset is a bit on the small side, especially since only 29% (out of curiosity, how did you come up with that number?) is visible right now.",
      "votes": null
    },
    {
      "id": "403758",
      "postDate": "10/14/2018 14:42:17",
      "content": "<p>I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it. \nI'm using F1_score average=macro from scikit learn. Val set is 30% of given dataset.\nCreating validation set with either one of the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819\">two methods</a> of this kaggle discussion doesn't change a thing.. I'm lost :) if anybody feels in a helping mood and has any idea ... much appreciated </p>",
      "rawMarkdown": "I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it. \nI'm using F1_score average=macro from scikit learn. Val set is 30% of given dataset.\nCreating validation set with either one of the [two methods][1] of this kaggle discussion doesn't change a thing.. I'm lost :) if anybody feels in a helping mood and has any idea ... much appreciated \n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819",
      "votes": null
    },
    {
      "id": "403896",
      "postDate": "10/14/2018 20:44:25",
      "content": "<blockquote>\n  <p>I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it.</p>\n</blockquote>\n\n<p>Either your 30% validation set or 29% of test data scored on public leaderboard is not representative enough to match the overall distribution of test data. You don't have the control of the latter, but you can do a better job of the former. For example, do N-fold validation rather than a simple holdout, as that increases the likelihood that the averaged cross-validation will do better approximating the test score than a simple holdout.</p>\n\n<p>I am assuming that Kaggle is scoring the subset of test data that is representative of the whole. That may or may not be correct, but it is certainly not in client's interest if Kaggle has selected a subset of test data for public scoring that is deliberately misleading.</p>",
      "rawMarkdown": "&gt; I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it.\n\nEither your 30% validation set or 29% of test data scored on public leaderboard is not representative enough to match the overall distribution of test data. You don't have the control of the latter, but you can do a better job of the former. For example, do N-fold validation rather than a simple holdout, as that increases the likelihood that the averaged cross-validation will do better approximating the test score than a simple holdout.\n\nI am assuming that Kaggle is scoring the subset of test data that is representative of the whole. That may or may not be correct, but it is certainly not in client's interest if Kaggle has selected a subset of test data for public scoring that is deliberately misleading.",
      "votes": null
    },
    {
      "id": "404087",
      "postDate": "10/15/2018 08:19:29",
      "content": "<p>Thanks I can try this, right. </p>",
      "rawMarkdown": "Thanks I can try this, right.",
      "votes": null
    },
    {
      "id": "408939",
      "postDate": "10/23/2018 16:16:24",
      "content": "<p>Ok, my issue was definitely <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\">this</a></p>\n\n<p>The order of the ids in the submission matters. </p>",
      "rawMarkdown": "Ok, my issue was definitely [this](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366)\n\nThe order of the ids in the submission matters.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 401837,
      "author_name": "kwentar",
      "author_url": "",
      "post_date": "10/10/2018 19:01:40",
      "content": "<p>I have the same result (local value bigger than LB), but not so much. You can compare your results with <a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">scikit f1</a> and with this <a href=\"https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras\">variant</a>. Also, may be you have bad split strategy or something else. I don't see any errors.</p>",
      "votes": null,
      "replies": [
        {
          "id": 403074,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "10/12/2018 20:26:40",
          "content": "<p>Thank you for your insights! I think the validation set I used was far too small.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 403758,
          "author_name": "equiplane",
          "author_url": "",
          "post_date": "10/14/2018 14:42:17",
          "content": "<p>I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it. \nI'm using F1_score average=macro from scikit learn. Val set is 30% of given dataset.\nCreating validation set with either one of the <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819\">two methods</a> of this kaggle discussion doesn't change a thing.. I'm lost :) if anybody feels in a helping mood and has any idea ... much appreciated </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 403896,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "10/14/2018 20:44:25",
          "content": "<blockquote>\n  <p>I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it.</p>\n</blockquote>\n\n<p>Either your 30% validation set or 29% of test data scored on public leaderboard is not representative enough to match the overall distribution of test data. You don't have the control of the latter, but you can do a better job of the former. For example, do N-fold validation rather than a simple holdout, as that increases the likelihood that the averaged cross-validation will do better approximating the test score than a simple holdout.</p>\n\n<p>I am assuming that Kaggle is scoring the subset of test data that is representative of the whole. That may or may not be correct, but it is certainly not in client's interest if Kaggle has selected a subset of test data for public scoring that is deliberately misleading.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 404087,
          "author_name": "equiplane",
          "author_url": "",
          "post_date": "10/15/2018 08:19:29",
          "content": "<p>Thanks I can try this, right. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 401841,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "10/10/2018 19:03:55",
      "content": "<p>All global metrics are only approximated during batch-wise training, so you want to calculate them only at epoch's end.</p>\n\n<p><a href=\"https://bit.ly/2n8buDO\"><strong>Link 1</strong></a></p>\n\n<p><a href=\"https://github.com/keras-team/keras/issues/5794\"><strong>Link 2</strong></a></p>\n\n<p><a href=\"https://github.com/keras-team/keras/issues/6507\"><strong>Link 3</strong></a></p>\n\n<p>Yet another explanation for score discrepancy is that you could be overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 402584,
          "author_name": "carlolepelaars",
          "author_url": "",
          "post_date": "10/12/2018 00:17:00",
          "content": "<p>Thanks for the information! I think overfitting is the problem.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 403128,
      "author_name": "nwbrown",
      "author_url": "",
      "post_date": "10/12/2018 23:54:45",
      "content": "<p>I've also seen much lower scores in my hold out validation dataset and the scores from the test dataset.\nIt may be just that that the test dataset is a bit on the small side, especially since only 29% (out of curiosity, how did you come up with that number?) is visible right now.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 408939,
      "author_name": "danmoller",
      "author_url": "",
      "post_date": "10/23/2018 16:16:24",
      "content": "<p>Ok, my issue was definitely <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\">this</a></p>\n\n<p>The order of the ids in the submission matters. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "401793": "Hi guys,\n\nCurrently, we're using the following function to track the macro f1 score of my Keras model:\n\n    from keras import backend as K\n    \n        def f1(y_true, y_pred):\n            def recall(y_true, y_pred):\n                  true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n                  possible_positives = K.sum(K.round(K.clip(y_true, 0, 1)))\n                  recall = true_positives / (possible_positives + K.epsilon())\n                  return recall\n        \n            def precision(y_true, y_pred):\n                  true_positives = K.sum(K.round(K.clip(y_true * y_pred, 0, 1)))\n                  predicted_positives = K.sum(K.round(K.clip(y_pred, 0, 1)))\n                  precision = true_positives / (predicted_positives + K.epsilon())\n                  return precision\n    \n        precision = precision(y_true, y_pred)\n        recall = recall(y_true, y_pred)\n        return 2*((precision*recall)/(precision+recall+K.epsilon()))\n\nHowever, the scores we get on the Kaggle leaderboard are much lower than the metric is reporting. Have we missed something and if so, how can we improve this metric to match Kaggle's metric?",
    "401837": "I have the same result (local value bigger than LB), but not so much. You can compare your results with [scikit f1][1] and with this [variant][2]. Also, may be you have bad split strategy or something else. I don't see any errors.\n\n\n  [1]: http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\n  [2]: https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras",
    "401841": "All global metrics are only approximated during batch-wise training, so you want to calculate them only at epoch's end.\n\n[__Link 1__](https://bit.ly/2n8buDO)\n\n[__Link 2__](https://github.com/keras-team/keras/issues/5794)\n\n[__Link 3__](https://github.com/keras-team/keras/issues/6507)\n\nYet another explanation for score discrepancy is that you could be overfitting.",
    "402584": "Thanks for the information! I think overfitting is the problem.",
    "403074": "Thank you for your insights! I think the validation set I used was far too small.",
    "403128": "I've also seen much lower scores in my hold out validation dataset and the scores from the test dataset.\nIt may be just that that the test dataset is a bit on the small side, especially since only 29% (out of curiosity, how did you come up with that number?) is visible right now.",
    "403758": "I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it. \nI'm using F1_score average=macro from scikit learn. Val set is 30% of given dataset.\nCreating validation set with either one of the [two methods][1] of this kaggle discussion doesn't change a thing.. I'm lost :) if anybody feels in a helping mood and has any idea ... much appreciated \n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67819",
    "403896": "&gt; I pass from 0.16 in hold-out validation set to 0.075 in test LB. I just don't get it.\n\nEither your 30% validation set or 29% of test data scored on public leaderboard is not representative enough to match the overall distribution of test data. You don't have the control of the latter, but you can do a better job of the former. For example, do N-fold validation rather than a simple holdout, as that increases the likelihood that the averaged cross-validation will do better approximating the test score than a simple holdout.\n\nI am assuming that Kaggle is scoring the subset of test data that is representative of the whole. That may or may not be correct, but it is certainly not in client's interest if Kaggle has selected a subset of test data for public scoring that is deliberately misleading.",
    "404087": "Thanks I can try this, right.",
    "408939": "Ok, my issue was definitely [this](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366)\n\nThe order of the ids in the submission matters."
  },
  "source": "meta"
}