{
  "id": 75472,
  "title": "Question about implementing f1 macro at the end of each epoch",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/75472",
  "author_name": "",
  "post_date": "2018-12-22T06:23:51.099426400Z",
  "votes": 2,
  "comment_count": 12,
  "views": 0,
  "content": "<p>I was implementing the f1 \"micro\" but it seems to have little correlation to the LB score, so I wanted to implement proper f1_macro at the end of each epoch to help me select the best model.</p>\n\n<p>I followed <a href=\"https://medium.com/@thongonary/how-to-compute-f1-score-for-each-epoch-in-keras-a1acd17715a2\">this article</a>, which is basically:</p>\n\n<pre><code>class Metrics(Callback):\ndef on_train_begin(self, logs={}):\n self.val_f1s = []\n self.val_recalls = []\n self.val_precisions = []\n\ndef on_epoch_end(self, epoch, logs={}):\n val_predict = (np.asarray(self.model.predict(self.model.validation_data[0]))).round()\n val_targ = self.model.validation_data[1]\n _val_f1 = f1_score(val_targ, val_predict)\n _val_recall = recall_score(val_targ, val_predict)\n _val_precision = precision_score(val_targ, val_predict)\n self.val_f1s.append(_val_f1)\n self.val_recalls.append(_val_recall)\n self.val_precisions.append(_val_precision)\n print “ — val_f1: %f — val_precision: %f — val_recall %f” %(_val_f1, _val_precision, _val_recall)\n return\n\nmetrics = Metrics()\n</code></pre>\n\n<p>And to use:</p>\n\n<pre><code>model.fit(training_data, training_target, \n validation_data=(validation_data, validation_target),\n nb_epoch=10,\n batch_size=64,\n callbacks=[metrics])\n</code></pre>\n\n<p>I have two problems with this code:</p>\n\n<p>I have two problems with this code. One that is somewhat solvable is that I use validation generator as I don't have that much ram, so I don't have the (validation_data, validation_target) handy. I could swicth to pre loading it but its a bit awkward.</p>\n\n<p>The second is that the code seems to perform predict again, although Keras has already predicted on the valuation set to calculate the val_loss. This will practically double the validation time.</p>\n\n<p>Any help or other ways to implement this will be appreciated.</p>",
  "messages": [
    {
      "id": "443685",
      "postDate": "12/22/2018 06:23:51",
      "content": "<p>I was implementing the f1 \"micro\" but it seems to have little correlation to the LB score, so I wanted to implement proper f1_macro at the end of each epoch to help me select the best model.</p>\n\n<p>I followed <a href=\"https://medium.com/@thongonary/how-to-compute-f1-score-for-each-epoch-in-keras-a1acd17715a2\">this article</a>, which is basically:</p>\n\n<pre><code>class Metrics(Callback):\ndef on_train_begin(self, logs={}):\n self.val_f1s = []\n self.val_recalls = []\n self.val_precisions = []\n\ndef on_epoch_end(self, epoch, logs={}):\n val_predict = (np.asarray(self.model.predict(self.model.validation_data[0]))).round()\n val_targ = self.model.validation_data[1]\n _val_f1 = f1_score(val_targ, val_predict)\n _val_recall = recall_score(val_targ, val_predict)\n _val_precision = precision_score(val_targ, val_predict)\n self.val_f1s.append(_val_f1)\n self.val_recalls.append(_val_recall)\n self.val_precisions.append(_val_precision)\n print “ — val_f1: %f — val_precision: %f — val_recall %f” %(_val_f1, _val_precision, _val_recall)\n return\n\nmetrics = Metrics()\n</code></pre>\n\n<p>And to use:</p>\n\n<pre><code>model.fit(training_data, training_target, \n validation_data=(validation_data, validation_target),\n nb_epoch=10,\n batch_size=64,\n callbacks=[metrics])\n</code></pre>\n\n<p>I have two problems with this code:</p>\n\n<p>I have two problems with this code. One that is somewhat solvable is that I use validation generator as I don't have that much ram, so I don't have the (validation_data, validation_target) handy. I could swicth to pre loading it but its a bit awkward.</p>\n\n<p>The second is that the code seems to perform predict again, although Keras has already predicted on the valuation set to calculate the val_loss. This will practically double the validation time.</p>\n\n<p>Any help or other ways to implement this will be appreciated.</p>",
      "rawMarkdown": "I was implementing the f1 \"micro\" but it seems to have little correlation to the LB score, so I wanted to implement proper f1_macro at the end of each epoch to help me select the best model.\n\nI followed [this article][1], which is basically:\n\n    class Metrics(Callback):\n    def on_train_begin(self, logs={}):\n     self.val_f1s = []\n     self.val_recalls = []\n     self.val_precisions = []\n     \n    def on_epoch_end(self, epoch, logs={}):\n     val_predict = (np.asarray(self.model.predict(self.model.validation_data[0]))).round()\n     val_targ = self.model.validation_data[1]\n     _val_f1 = f1_score(val_targ, val_predict)\n     _val_recall = recall_score(val_targ, val_predict)\n     _val_precision = precision_score(val_targ, val_predict)\n     self.val_f1s.append(_val_f1)\n     self.val_recalls.append(_val_recall)\n     self.val_precisions.append(_val_precision)\n     print “ — val_f1: %f — val_precision: %f — val_recall %f” %(_val_f1, _val_precision, _val_recall)\n     return\n     \n    metrics = Metrics()\n\nAnd to use:\n\n    model.fit(training_data, training_target, \n     validation_data=(validation_data, validation_target),\n     nb_epoch=10,\n     batch_size=64,\n     callbacks=[metrics])\nI have two problems with this code:\n\nI have two problems with this code. One that is somewhat solvable is that I use validation generator as I don't have that much ram, so I don't have the (validation_data, validation_target) handy. I could swicth to pre loading it but its a bit awkward.\n\nThe second is that the code seems to perform predict again, although Keras has already predicted on the valuation set to calculate the val_loss. This will practically double the validation time.\n\nAny help or other ways to implement this will be appreciated.\n\n  [1]: https://medium.com/@thongonary/how-to-compute-f1-score-for-each-epoch-in-keras-a1acd17715a2",
      "votes": null
    },
    {
      "id": "443868",
      "postDate": "12/22/2018 16:17:21",
      "content": "<p>I think the most efficient way would be to accumulate the sums of the true positives, false positives, and false negatives after each batch, and then all you have to do at the end of the epoch is calculate f1 using those values. I'm not totally sure how to do that in Keras though, probably a custom callback with logic for <code>on_batch_end</code> and <code>on_epoch_end</code>.</p>",
      "rawMarkdown": "I think the most efficient way would be to accumulate the sums of the true positives, false positives, and false negatives after each batch, and then all you have to do at the end of the epoch is calculate f1 using those values. I'm not totally sure how to do that in Keras though, probably a custom callback with logic for `on_batch_end` and `on_epoch_end`.",
      "votes": null
    },
    {
      "id": "443977",
      "postDate": "12/22/2018 20:26:43",
      "content": "<p>Does anyone knows if there is a penalty for running the fit 1 epoch at the time and doing the validation myself? Is the optimizer state remain between calls to fit?\nThat I mean is doing something like:</p>\n\n<pre><code>for i in range(n_epochs):\n    model.fit(...., epochs=1)\n    model.predict(....)\n    calc_metrics(...)\n</code></pre>",
      "rawMarkdown": "Does anyone knows if there is a penalty for running the fit 1 epoch at the time and doing the validation myself? Is the optimizer state remain between calls to fit?\nThat I mean is doing something like:\n\n    for i in range(n_epochs):\n        model.fit(...., epochs=1)\n        model.predict(....)\n        calc_metrics(...)",
      "votes": null
    },
    {
      "id": "444022",
      "postDate": "12/23/2018 00:18:47",
      "content": "<p>You can do the following. Write a functor that is passed as a metric function to fit and create a corresponding callback that would reset the functor and print F1 score calculated for entire dataset. I provide a code that would be comparable with <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">my kernel</a> (Pytorch + fast.ai), but you can adjust it to your needs:</p>\n\n<pre><code>class F1:\n    __name__ = 'F1 macro'\n    def __init__(self,n=28):\n        self.n = n\n        self.TP = np.zeros(self.n)\n        self.FP = np.zeros(self.n)\n        self.FN = np.zeros(self.n)\n\n    def __call__(self,preds,targs,th=0.0):\n        preds = (preds &amp;gt; th).int()\n        targs = targs.int()\n        self.TP += (preds*targs).float().sum(dim=0)\n        self.FP += (preds &amp;gt; targs).float().sum(dim=0)\n        self.FN += (preds &amp;lt; targs).float().sum(dim=0)\n        score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6)).mean()\n        return score\n\n    def reset(self):\n        #macro F1 score\n        score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6))\n        print('F1 macro:',score.mean(),flush=True)\n        print('F1:',score)\n        self.TP = np.zeros(self.n)\n        self.FP = np.zeros(self.n)\n        self.FN = np.zeros(self.n)\n\nclass F1_callback(Callback):\n    def __init__(self, n=28):\n        self.f1 = F1(n)\n\n    def on_epoch_end(self, metrics):\n        self.f1.reset()\n</code></pre>\n\n<p>And usage in the code:</p>\n\n<pre><code>f1_callback = F1_callback()\nlearner.metrics = [f1_callback.f1]\n\nlearner.fit(lr,1,callbacks=[f1_callback])\n</code></pre>",
      "rawMarkdown": "You can do the following. Write a functor that is passed as a metric function to fit and create a corresponding callback that would reset the functor and print F1 score calculated for entire dataset. I provide a code that would be comparable with [my kernel][1] (Pytorch + fast.ai), but you can adjust it to your needs:\n\n    class F1:\n        __name__ = 'F1 macro'\n        def __init__(self,n=28):\n            self.n = n\n            self.TP = np.zeros(self.n)\n            self.FP = np.zeros(self.n)\n            self.FN = np.zeros(self.n)\n            \n        def __call__(self,preds,targs,th=0.0):\n            preds = (preds &gt; th).int()\n            targs = targs.int()\n            self.TP += (preds*targs).float().sum(dim=0)\n            self.FP += (preds &gt; targs).float().sum(dim=0)\n            self.FN += (preds &lt; targs).float().sum(dim=0)\n            score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6)).mean()\n            return score\n            \n        def reset(self):\n            #macro F1 score\n            score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6))\n            print('F1 macro:',score.mean(),flush=True)\n            print('F1:',score)\n            self.TP = np.zeros(self.n)\n            self.FP = np.zeros(self.n)\n            self.FN = np.zeros(self.n)\n    \n    class F1_callback(Callback):\n        def __init__(self, n=28):\n            self.f1 = F1(n)\n        \n        def on_epoch_end(self, metrics):\n            self.f1.reset()\n\nAnd usage in the code:\n\n    f1_callback = F1_callback()\n    learner.metrics = [f1_callback.f1]\n    \n    learner.fit(lr,1,callbacks=[f1_callback])\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
      "votes": null
    },
    {
      "id": "444033",
      "postDate": "12/23/2018 01:30:00",
      "content": "<p>Thank you! (as a side note, your kernal was really inspirational for me - I am tempted to switch to pytorch)</p>\n\n<p>Wouldn't calculating the F1 on all the dataset be a bit misleading? These metrics are mainly used to see that the kernel progress and does not overfit. Also, because I am not a very patient man, my epochs are shorter then the number of samples in the full (kaggle+external) HPA dataset (that has 100000k samples. In this case, calculating the F1 over part of the db+the full validation will also be unstable.</p>\n\n<p>I will try the primitive method of opening the loop and compare it to the usual method and see if it gets the same results.</p>",
      "rawMarkdown": "Thank you! (as a side note, your kernal was really inspirational for me - I am tempted to switch to pytorch)\n\nWouldn't calculating the F1 on all the dataset be a bit misleading? These metrics are mainly used to see that the kernel progress and does not overfit. Also, because I am not a very patient man, my epochs are shorter then the number of samples in the full (kaggle+external) HPA dataset (that has 100000k samples. In this case, calculating the F1 over part of the db+the full validation will also be unstable.\n\nI will try the primitive method of opening the loop and compare it to the usual method and see if it gets the same results.",
      "votes": null
    },
    {
      "id": "444036",
      "postDate": "12/23/2018 01:49:35",
      "content": "<p>I meant entire <strong>val</strong> dataset, thank you for pointing it out.</p>",
      "rawMarkdown": "I meant entire **val** dataset, thank you for pointing it out.",
      "votes": null
    },
    {
      "id": "444040",
      "postDate": "12/23/2018 01:58:34",
      "content": "<p>I need to check when these callbacks are called. As they ar attached to fit, I don't know if they are called at the beginning and end of val or train or both or how to know if i am now processing a val or train batch. This is all a bit vague in the docs. Will do some (lengthy) research and post here my findings</p>",
      "rawMarkdown": "I need to check when these callbacks are called. As they ar attached to fit, I don't know if they are called at the beginning and end of val or train or both or how to know if i am now processing a val or train batch. This is all a bit vague in the docs. Will do some (lengthy) research and post here my findings",
      "votes": null
    },
    {
      "id": "444048",
      "postDate": "12/23/2018 02:31:21",
      "content": "<p>Unfortunately I'm not very familiar with Keras, so I cannot help much in that( but as  @William Horton wrote, there should be an appropriate callback</p>",
      "rawMarkdown": "Unfortunately I'm not very familiar with Keras, so I cannot help much in that( but as  @William Horton wrote, there should be an appropriate callback",
      "votes": null
    },
    {
      "id": "444464",
      "postDate": "12/24/2018 04:28:32",
      "content": "<p>Here are my findings:</p>\n\n<ol>\n<li><p>The callbacks have no access to y_true and y_pred. It has been reported here (<a href=\"https://github.com/tensorflow/tensorflow/issues/21174\">https://github.com/tensorflow/tensorflow/issues/21174</a>) and assigned to Chollet François :-)</p></li>\n<li><p>The loss function is part of the tf graph and as such very limited in what it can do and what it can access. Even knowing if you are in training is complicated (there is a is_training_phase function but it is actually a branch in the graph and can't be used as just if this else that).</p></li>\n<li><p>Running single epochs one after the other instead of fitting for several epochs is not the same. Must be something with the optimizer or somesuch. It consistently give worse results.</p></li>\n</ol>\n\n<p>So, I guess I am stuck with re-predicting the validation with model.predict or just trusting the val_loss :(</p>",
      "rawMarkdown": "Here are my findings:\n\n1. The callbacks have no access to y_true and y_pred. It has been reported here (https://github.com/tensorflow/tensorflow/issues/21174) and assigned to Chollet François :-)\n\n2. The loss function is part of the tf graph and as such very limited in what it can do and what it can access. Even knowing if you are in training is complicated (there is a is_training_phase function but it is actually a branch in the graph and can't be used as just if this else that).\n\n3. Running single epochs one after the other instead of fitting for several epochs is not the same. Must be something with the optimizer or somesuch. It consistently give worse results.\n\nSo, I guess I am stuck with re-predicting the validation with model.predict or just trusting the val_loss :(",
      "votes": null
    },
    {
      "id": "444469",
      "postDate": "12/24/2018 04:37:14",
      "content": "<p>Do you have an option to use  customized <strong>metric</strong> (not loss function) in Tensorflow (Keras)? I'm just not very familiar with this platform. In this case the approach I describe will work for you. You pass the functor as a metric (which has access to pred and target values) and call reset in <code>on_epoch_end</code> callback. You only need to rewrite FP, FN, TP compute part to work with Tensorflow tensors rather than Pytorch.</p>",
      "rawMarkdown": "Do you have an option to use  customized **metric** (not loss function) in Tensorflow (Keras)? I'm just not very familiar with this platform. In this case the approach I describe will work for you. You pass the functor as a metric (which has access to pred and target values) and call reset in `on_epoch_end` callback. You only need to rewrite FP, FN, TP compute part to work with Tensorflow tensors rather than Pytorch.",
      "votes": null
    },
    {
      "id": "444495",
      "postDate": "12/24/2018 06:00:36",
      "content": "<p>Yes, its possible but is is very limited. see <a href=\"https://keras.io/metrics/#custom-metrics\">https://keras.io/metrics/#custom-metrics</a>\nits not a class with init etc, so it is very difficult to implement validation only metrics or even to accumulate over batches.</p>",
      "rawMarkdown": "Yes, its possible but is is very limited. see https://keras.io/metrics/#custom-metrics\nits not a class with init etc, so it is very difficult to implement validation only metrics or even to accumulate over batches.",
      "votes": null
    },
    {
      "id": "444504",
      "postDate": "12/24/2018 06:34:30",
      "content": "<p>Found this example (<a href=\"https://github.com/keras-team/keras/blob/master/tests/keras/metrics_test.py\">https://github.com/keras-team/keras/blob/master/tests/keras/metrics_test.py</a>) referenced from an issue of and issue... will try implementing your code as the BinaryTruePositives is implemented there. looks very promising</p>",
      "rawMarkdown": "Found this example (https://github.com/keras-team/keras/blob/master/tests/keras/metrics_test.py) referenced from an issue of and issue... will try implementing your code as the BinaryTruePositives is implemented there. looks very promising",
      "votes": null
    },
    {
      "id": "444588",
      "postDate": "12/24/2018 10:33:29",
      "content": "<p>Nope, it doesn't help... Still no way to know if you are processing training or validation. This just takes too much time... I'll go with the 2 predicts method.</p>\n\n<p>Thank you @lafoss for your efforts. Its close, but no cigar. </p>",
      "rawMarkdown": "Nope, it doesn't help... Still no way to know if you are processing training or validation. This just takes too much time... I'll go with the 2 predicts method.\n\nThank you @lafoss for your efforts. Its close, but no cigar.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 443868,
      "author_name": "hortonhearsafoo",
      "author_url": "",
      "post_date": "12/22/2018 16:17:21",
      "content": "<p>I think the most efficient way would be to accumulate the sums of the true positives, false positives, and false negatives after each batch, and then all you have to do at the end of the epoch is calculate f1 using those values. I'm not totally sure how to do that in Keras though, probably a custom callback with logic for <code>on_batch_end</code> and <code>on_epoch_end</code>.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 443977,
      "author_name": "moshel",
      "author_url": "",
      "post_date": "12/22/2018 20:26:43",
      "content": "<p>Does anyone knows if there is a penalty for running the fit 1 epoch at the time and doing the validation myself? Is the optimizer state remain between calls to fit?\nThat I mean is doing something like:</p>\n\n<pre><code>for i in range(n_epochs):\n    model.fit(...., epochs=1)\n    model.predict(....)\n    calc_metrics(...)\n</code></pre>",
      "votes": null,
      "replies": []
    },
    {
      "id": 444022,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "12/23/2018 00:18:47",
      "content": "<p>You can do the following. Write a functor that is passed as a metric function to fit and create a corresponding callback that would reset the functor and print F1 score calculated for entire dataset. I provide a code that would be comparable with <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb\">my kernel</a> (Pytorch + fast.ai), but you can adjust it to your needs:</p>\n\n<pre><code>class F1:\n    __name__ = 'F1 macro'\n    def __init__(self,n=28):\n        self.n = n\n        self.TP = np.zeros(self.n)\n        self.FP = np.zeros(self.n)\n        self.FN = np.zeros(self.n)\n\n    def __call__(self,preds,targs,th=0.0):\n        preds = (preds &amp;gt; th).int()\n        targs = targs.int()\n        self.TP += (preds*targs).float().sum(dim=0)\n        self.FP += (preds &amp;gt; targs).float().sum(dim=0)\n        self.FN += (preds &amp;lt; targs).float().sum(dim=0)\n        score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6)).mean()\n        return score\n\n    def reset(self):\n        #macro F1 score\n        score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6))\n        print('F1 macro:',score.mean(),flush=True)\n        print('F1:',score)\n        self.TP = np.zeros(self.n)\n        self.FP = np.zeros(self.n)\n        self.FN = np.zeros(self.n)\n\nclass F1_callback(Callback):\n    def __init__(self, n=28):\n        self.f1 = F1(n)\n\n    def on_epoch_end(self, metrics):\n        self.f1.reset()\n</code></pre>\n\n<p>And usage in the code:</p>\n\n<pre><code>f1_callback = F1_callback()\nlearner.metrics = [f1_callback.f1]\n\nlearner.fit(lr,1,callbacks=[f1_callback])\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 444033,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/23/2018 01:30:00",
          "content": "<p>Thank you! (as a side note, your kernal was really inspirational for me - I am tempted to switch to pytorch)</p>\n\n<p>Wouldn't calculating the F1 on all the dataset be a bit misleading? These metrics are mainly used to see that the kernel progress and does not overfit. Also, because I am not a very patient man, my epochs are shorter then the number of samples in the full (kaggle+external) HPA dataset (that has 100000k samples. In this case, calculating the F1 over part of the db+the full validation will also be unstable.</p>\n\n<p>I will try the primitive method of opening the loop and compare it to the usual method and see if it gets the same results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444036,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/23/2018 01:49:35",
          "content": "<p>I meant entire <strong>val</strong> dataset, thank you for pointing it out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444040,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/23/2018 01:58:34",
          "content": "<p>I need to check when these callbacks are called. As they ar attached to fit, I don't know if they are called at the beginning and end of val or train or both or how to know if i am now processing a val or train batch. This is all a bit vague in the docs. Will do some (lengthy) research and post here my findings</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444048,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/23/2018 02:31:21",
          "content": "<p>Unfortunately I'm not very familiar with Keras, so I cannot help much in that( but as  @William Horton wrote, there should be an appropriate callback</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444464,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/24/2018 04:28:32",
          "content": "<p>Here are my findings:</p>\n\n<ol>\n<li><p>The callbacks have no access to y_true and y_pred. It has been reported here (<a href=\"https://github.com/tensorflow/tensorflow/issues/21174\">https://github.com/tensorflow/tensorflow/issues/21174</a>) and assigned to Chollet François :-)</p></li>\n<li><p>The loss function is part of the tf graph and as such very limited in what it can do and what it can access. Even knowing if you are in training is complicated (there is a is_training_phase function but it is actually a branch in the graph and can't be used as just if this else that).</p></li>\n<li><p>Running single epochs one after the other instead of fitting for several epochs is not the same. Must be something with the optimizer or somesuch. It consistently give worse results.</p></li>\n</ol>\n\n<p>So, I guess I am stuck with re-predicting the validation with model.predict or just trusting the val_loss :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444469,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/24/2018 04:37:14",
          "content": "<p>Do you have an option to use  customized <strong>metric</strong> (not loss function) in Tensorflow (Keras)? I'm just not very familiar with this platform. In this case the approach I describe will work for you. You pass the functor as a metric (which has access to pred and target values) and call reset in <code>on_epoch_end</code> callback. You only need to rewrite FP, FN, TP compute part to work with Tensorflow tensors rather than Pytorch.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444495,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/24/2018 06:00:36",
          "content": "<p>Yes, its possible but is is very limited. see <a href=\"https://keras.io/metrics/#custom-metrics\">https://keras.io/metrics/#custom-metrics</a>\nits not a class with init etc, so it is very difficult to implement validation only metrics or even to accumulate over batches.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444504,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/24/2018 06:34:30",
          "content": "<p>Found this example (<a href=\"https://github.com/keras-team/keras/blob/master/tests/keras/metrics_test.py\">https://github.com/keras-team/keras/blob/master/tests/keras/metrics_test.py</a>) referenced from an issue of and issue... will try implementing your code as the BinaryTruePositives is implemented there. looks very promising</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 444588,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/24/2018 10:33:29",
          "content": "<p>Nope, it doesn't help... Still no way to know if you are processing training or validation. This just takes too much time... I'll go with the 2 predicts method.</p>\n\n<p>Thank you @lafoss for your efforts. Its close, but no cigar. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "443685": "I was implementing the f1 \"micro\" but it seems to have little correlation to the LB score, so I wanted to implement proper f1_macro at the end of each epoch to help me select the best model.\n\nI followed [this article][1], which is basically:\n\n    class Metrics(Callback):\n    def on_train_begin(self, logs={}):\n     self.val_f1s = []\n     self.val_recalls = []\n     self.val_precisions = []\n     \n    def on_epoch_end(self, epoch, logs={}):\n     val_predict = (np.asarray(self.model.predict(self.model.validation_data[0]))).round()\n     val_targ = self.model.validation_data[1]\n     _val_f1 = f1_score(val_targ, val_predict)\n     _val_recall = recall_score(val_targ, val_predict)\n     _val_precision = precision_score(val_targ, val_predict)\n     self.val_f1s.append(_val_f1)\n     self.val_recalls.append(_val_recall)\n     self.val_precisions.append(_val_precision)\n     print “ — val_f1: %f — val_precision: %f — val_recall %f” %(_val_f1, _val_precision, _val_recall)\n     return\n     \n    metrics = Metrics()\n\nAnd to use:\n\n    model.fit(training_data, training_target, \n     validation_data=(validation_data, validation_target),\n     nb_epoch=10,\n     batch_size=64,\n     callbacks=[metrics])\nI have two problems with this code:\n\nI have two problems with this code. One that is somewhat solvable is that I use validation generator as I don't have that much ram, so I don't have the (validation_data, validation_target) handy. I could swicth to pre loading it but its a bit awkward.\n\nThe second is that the code seems to perform predict again, although Keras has already predicted on the valuation set to calculate the val_loss. This will practically double the validation time.\n\nAny help or other ways to implement this will be appreciated.\n\n  [1]: https://medium.com/@thongonary/how-to-compute-f1-score-for-each-epoch-in-keras-a1acd17715a2",
    "443868": "I think the most efficient way would be to accumulate the sums of the true positives, false positives, and false negatives after each batch, and then all you have to do at the end of the epoch is calculate f1 using those values. I'm not totally sure how to do that in Keras though, probably a custom callback with logic for `on_batch_end` and `on_epoch_end`.",
    "443977": "Does anyone knows if there is a penalty for running the fit 1 epoch at the time and doing the validation myself? Is the optimizer state remain between calls to fit?\nThat I mean is doing something like:\n\n    for i in range(n_epochs):\n        model.fit(...., epochs=1)\n        model.predict(....)\n        calc_metrics(...)",
    "444022": "You can do the following. Write a functor that is passed as a metric function to fit and create a corresponding callback that would reset the functor and print F1 score calculated for entire dataset. I provide a code that would be comparable with [my kernel][1] (Pytorch + fast.ai), but you can adjust it to your needs:\n\n    class F1:\n        __name__ = 'F1 macro'\n        def __init__(self,n=28):\n            self.n = n\n            self.TP = np.zeros(self.n)\n            self.FP = np.zeros(self.n)\n            self.FN = np.zeros(self.n)\n            \n        def __call__(self,preds,targs,th=0.0):\n            preds = (preds &gt; th).int()\n            targs = targs.int()\n            self.TP += (preds*targs).float().sum(dim=0)\n            self.FP += (preds &gt; targs).float().sum(dim=0)\n            self.FN += (preds &lt; targs).float().sum(dim=0)\n            score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6)).mean()\n            return score\n            \n        def reset(self):\n            #macro F1 score\n            score = (2.0*self.TP/(2.0*self.TP + self.FP + self.FN + 1e-6))\n            print('F1 macro:',score.mean(),flush=True)\n            print('F1:',score)\n            self.TP = np.zeros(self.n)\n            self.FP = np.zeros(self.n)\n            self.FN = np.zeros(self.n)\n    \n    class F1_callback(Callback):\n        def __init__(self, n=28):\n            self.f1 = F1(n)\n        \n        def on_epoch_end(self, metrics):\n            self.f1.reset()\n\nAnd usage in the code:\n\n    f1_callback = F1_callback()\n    learner.metrics = [f1_callback.f1]\n    \n    learner.fit(lr,1,callbacks=[f1_callback])\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb",
    "444033": "Thank you! (as a side note, your kernal was really inspirational for me - I am tempted to switch to pytorch)\n\nWouldn't calculating the F1 on all the dataset be a bit misleading? These metrics are mainly used to see that the kernel progress and does not overfit. Also, because I am not a very patient man, my epochs are shorter then the number of samples in the full (kaggle+external) HPA dataset (that has 100000k samples. In this case, calculating the F1 over part of the db+the full validation will also be unstable.\n\nI will try the primitive method of opening the loop and compare it to the usual method and see if it gets the same results.",
    "444036": "I meant entire **val** dataset, thank you for pointing it out.",
    "444040": "I need to check when these callbacks are called. As they ar attached to fit, I don't know if they are called at the beginning and end of val or train or both or how to know if i am now processing a val or train batch. This is all a bit vague in the docs. Will do some (lengthy) research and post here my findings",
    "444048": "Unfortunately I'm not very familiar with Keras, so I cannot help much in that( but as  @William Horton wrote, there should be an appropriate callback",
    "444464": "Here are my findings:\n\n1. The callbacks have no access to y_true and y_pred. It has been reported here (https://github.com/tensorflow/tensorflow/issues/21174) and assigned to Chollet François :-)\n\n2. The loss function is part of the tf graph and as such very limited in what it can do and what it can access. Even knowing if you are in training is complicated (there is a is_training_phase function but it is actually a branch in the graph and can't be used as just if this else that).\n\n3. Running single epochs one after the other instead of fitting for several epochs is not the same. Must be something with the optimizer or somesuch. It consistently give worse results.\n\nSo, I guess I am stuck with re-predicting the validation with model.predict or just trusting the val_loss :(",
    "444469": "Do you have an option to use  customized **metric** (not loss function) in Tensorflow (Keras)? I'm just not very familiar with this platform. In this case the approach I describe will work for you. You pass the functor as a metric (which has access to pred and target values) and call reset in `on_epoch_end` callback. You only need to rewrite FP, FN, TP compute part to work with Tensorflow tensors rather than Pytorch.",
    "444495": "Yes, its possible but is is very limited. see https://keras.io/metrics/#custom-metrics\nits not a class with init etc, so it is very difficult to implement validation only metrics or even to accumulate over batches.",
    "444504": "Found this example (https://github.com/keras-team/keras/blob/master/tests/keras/metrics_test.py) referenced from an issue of and issue... will try implementing your code as the BinaryTruePositives is implemented there. looks very promising",
    "444588": "Nope, it doesn't help... Still no way to know if you are processing training or validation. This just takes too much time... I'll go with the 2 predicts method.\n\nThank you @lafoss for your efforts. Its close, but no cigar."
  },
  "source": "meta"
}