{
  "id": 73929,
  "title": "How to validate result a bit more properly in keras",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73929",
  "author_name": "",
  "post_date": "2018-12-06T19:15:38.282307Z",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<p>It is known that <em>F1</em> score is assumed to be 0 if class has no <em>true positive</em> samples. If we choose a small batch size then it is highly possible that rare classes will not be present in it and thus <em>F1</em> score will be 0 for them. That will affect <em>F1</em> macro metric computed as an average among all <em>F1</em> scores for all classes. You can see the demo in <strong><a href=\"https://www.kaggle.com/malyutins/f1-macro-score-on-small-batches\">that kernel</a></strong></p>\n\n<p>When we train a network using <em>keras</em> all the metrics are computed after each batch (<strong><code>on_batch_end</code></strong>) and get averaged and printed to console at the end of epoch (<strong><code>on_epoch_end</code></strong>). Due to the effect of small <em>F1</em> score for small batch even for the correct predictions, displayed score will always be lower than real.</p>\n\n<p>To avoid such behavior we can create a callback to compute score for the entire validation set as follows:</p>\n\n<pre><code>class CustomValidationCallback(Callback):\n\n    def on_epoch_end(self, epoch, logs={}):\n        threshold = 0.1\n        x_val = self.validation_data[0]\n        y_true = self.validation_data[1]\n        y_pred = self.model.predict(x_val)\n        f1 = f1_score(y_true, (y_pred &amp;gt; threshold).astype(int), average='macro')\n        logs['val_f1_custom'] = f1\n        print('Epoch %05d: custom f1 score %0.3f' % (epoch + 1, round(f1, 3))) \n</code></pre>\n\n<p>Create an instance and pass it to fit method:</p>\n\n<pre><code>custom_val_callback = CustomValidationCallback()\n\nmodel.fit_generator(... , callbacks=[custom_val_callback, ...], ...)\n</code></pre>",
  "messages": [
    {
      "id": "434679",
      "postDate": "12/06/2018 19:15:38",
      "content": "<p>It is known that <em>F1</em> score is assumed to be 0 if class has no <em>true positive</em> samples. If we choose a small batch size then it is highly possible that rare classes will not be present in it and thus <em>F1</em> score will be 0 for them. That will affect <em>F1</em> macro metric computed as an average among all <em>F1</em> scores for all classes. You can see the demo in <strong><a href=\"https://www.kaggle.com/malyutins/f1-macro-score-on-small-batches\">that kernel</a></strong></p>\n\n<p>When we train a network using <em>keras</em> all the metrics are computed after each batch (<strong><code>on_batch_end</code></strong>) and get averaged and printed to console at the end of epoch (<strong><code>on_epoch_end</code></strong>). Due to the effect of small <em>F1</em> score for small batch even for the correct predictions, displayed score will always be lower than real.</p>\n\n<p>To avoid such behavior we can create a callback to compute score for the entire validation set as follows:</p>\n\n<pre><code>class CustomValidationCallback(Callback):\n\n    def on_epoch_end(self, epoch, logs={}):\n        threshold = 0.1\n        x_val = self.validation_data[0]\n        y_true = self.validation_data[1]\n        y_pred = self.model.predict(x_val)\n        f1 = f1_score(y_true, (y_pred &amp;gt; threshold).astype(int), average='macro')\n        logs['val_f1_custom'] = f1\n        print('Epoch %05d: custom f1 score %0.3f' % (epoch + 1, round(f1, 3))) \n</code></pre>\n\n<p>Create an instance and pass it to fit method:</p>\n\n<pre><code>custom_val_callback = CustomValidationCallback()\n\nmodel.fit_generator(... , callbacks=[custom_val_callback, ...], ...)\n</code></pre>",
      "rawMarkdown": "It is known that *F1* score is assumed to be 0 if class has no *true positive* samples. If we choose a small batch size then it is highly possible that rare classes will not be present in it and thus *F1* score will be 0 for them. That will affect *F1* macro metric computed as an average among all *F1* scores for all classes. You can see the demo in **[that kernel][1]**\n\nWhen we train a network using *keras* all the metrics are computed after each batch (**`on_batch_end`**) and get averaged and printed to console at the end of epoch (**`on_epoch_end`**). Due to the effect of small *F1* score for small batch even for the correct predictions, displayed score will always be lower than real.\n\nTo avoid such behavior we can create a callback to compute score for the entire validation set as follows:\n\n    class CustomValidationCallback(Callback):\n    \n        def on_epoch_end(self, epoch, logs={}):\n            threshold = 0.1\n            x_val = self.validation_data[0]\n            y_true = self.validation_data[1]\n            y_pred = self.model.predict(x_val)\n            f1 = f1_score(y_true, (y_pred &gt; threshold).astype(int), average='macro')\n            logs['val_f1_custom'] = f1\n            print('Epoch %05d: custom f1 score %0.3f' % (epoch + 1, round(f1, 3))) \n\nCreate an instance and pass it to fit method:\n\n    custom_val_callback = CustomValidationCallback()\n    \n    model.fit_generator(... , callbacks=[custom_val_callback, ...], ...)\n\n\n  [1]: https://www.kaggle.com/malyutins/f1-macro-score-on-small-batches",
      "votes": null
    },
    {
      "id": "434680",
      "postDate": "12/06/2018 19:23:32",
      "content": "<p>I was doing something similar but ran into issues when using a generator for validation data.</p>",
      "rawMarkdown": "I was doing something similar but ran into issues when using a generator for validation data.",
      "votes": null
    },
    {
      "id": "434704",
      "postDate": "12/06/2018 20:23:41",
      "content": "<p>That's a good point. Some of it was discussed early on <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242\"><strong>here</strong></a>. Yet another solution is presented <a href=\"https://bit.ly/2n8buDO\"><strong>here</strong></a>.</p>\n\n<p>As <a href=\"/ldm314\">@ldm314</a> (Brian) pointed out, it is non-trivial to set this up when using generators.</p>",
      "rawMarkdown": "That's a good point. Some of it was discussed early on [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242). Yet another solution is presented [__here__](https://bit.ly/2n8buDO).\n\nAs @ldm314 (Brian) pointed out, it is non-trivial to set this up when using generators.",
      "votes": null
    },
    {
      "id": "434719",
      "postDate": "12/06/2018 20:58:19",
      "content": "<p>The 2nd link also has the same problem. I based my example on that one. Here is it with the modifications, but it doubles the validation time. It also isn't very clean since it uses the global 'valid_gen'. It worked ok, but I ended up just not using it because of the performance.</p>\n\n<pre>class Metrics(Callback):\n    def on_train_begin(self, logs={}):\n        self.val_f1s = []\n        self.val_recalls = []\n        self.val_precisions = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        t_y = None\n        t_x = None\n        preds_y = []\n        valid_y = []\n        valid_gen.batch_size = VALID_BATCH_SIZE\n        batches = valid_gen.n//VALID_BATCH_SIZE+1\n        i=1\n        for _, (t_x, t_names) in zip((range(valid_gen.n//VALID_BATCH_SIZE+1)),\n                                    valid_gen):\n            t_y = self.model.predict(t_x)\n            for c_id, c_score in zip(t_names, t_y):\n                valid_y.append(c_id)\n                preds_y.append(c_score)\n            print('batch %d/%d' % (i,batches))\n            i+=1\n\n        predicted = np.array(preds_y)\n        val_targ = np.array(valid_y)\n        max_val = np.max(predicted)\n        val_predict = predicted &gt; (0.65 * max_val)\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        self.val_f1s.append(_val_f1)\n        self.val_recalls.append(_val_recall)\n        self.val_precisions.append(_val_precision)\n        print(\"val_f1: %f — val_precision: %f — val_recall %f\" %(_val_f1, _val_precision, _val_recall))\n        return\n</pre>",
      "rawMarkdown": "The 2nd link also has the same problem. I based my example on that one. Here is it with the modifications, but it doubles the validation time. It also isn't very clean since it uses the global 'valid_gen'. It worked ok, but I ended up just not using it because of the performance.\n\n<pre>class Metrics(Callback):\n    def on_train_begin(self, logs={}):\n        self.val_f1s = []\n        self.val_recalls = []\n        self.val_precisions = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        t_y = None\n        t_x = None\n        preds_y = []\n        valid_y = []\n        valid_gen.batch_size = VALID_BATCH_SIZE\n        batches = valid_gen.n//VALID_BATCH_SIZE+1\n        i=1\n        for _, (t_x, t_names) in zip((range(valid_gen.n//VALID_BATCH_SIZE+1)),\n                                    valid_gen):\n            t_y = self.model.predict(t_x)\n            for c_id, c_score in zip(t_names, t_y):\n                valid_y.append(c_id)\n                preds_y.append(c_score)\n            print('batch %d/%d' % (i,batches))\n            i+=1\n\n        predicted = np.array(preds_y)\n        val_targ = np.array(valid_y)\n        max_val = np.max(predicted)\n        val_predict = predicted &gt; (0.65 * max_val)\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        self.val_f1s.append(_val_f1)\n        self.val_recalls.append(_val_recall)\n        self.val_precisions.append(_val_precision)\n        print(\"val_f1: %f — val_precision: %f — val_recall %f\" %(_val_f1, _val_precision, _val_recall))\n        return\n</pre>",
      "votes": null
    },
    {
      "id": "435115",
      "postDate": "12/07/2018 14:23:09",
      "content": "<p>Thank you</p>",
      "rawMarkdown": "Thank you",
      "votes": null
    },
    {
      "id": "435295",
      "postDate": "12/07/2018 20:08:12",
      "content": "<p>Hey Sergey, thanks again. I have one small question. Does the threshold change or should that be left as is? </p>",
      "rawMarkdown": "Hey Sergey, thanks again. I have one small question. Does the threshold change or should that be left as is?",
      "votes": null
    },
    {
      "id": "435532",
      "postDate": "12/08/2018 07:18:23",
      "content": "<p>Threshold selection is another non trivial problem :) It is better to try different options. Once you have <code>y_pred</code> you can call <code>f1_score</code> with different values of <code>threshold</code> then print and log all of them.</p>",
      "rawMarkdown": "Threshold selection is another non trivial problem :) It is better to try different options. Once you have `y_pred` you can call `f1_score` with different values of `threshold` then print and log all of them.",
      "votes": null
    },
    {
      "id": "435811",
      "postDate": "12/08/2018 20:22:07",
      "content": "<p>Hi, Brian\nI think you can avoid run another validation generator loop by using stateful metrics. Assuming tensorflow as a backend, it already have a 'micro' f1 score: <code>tf.contrib.metrics.f1_score</code> but <a href=\"http://vict0rsch.github.io/2018/06/06/tensorflow-streaming-multilabel-f1/\">this article</a> explains about streaming metrics and implements a macro version. I'm using this class as f1_metric:  </p>\n\n<pre><code>class Metrics(keras.layers.Layer):\n    def __init__(self, num_classes, threshold, **kwargs):\n        super(Metrics, self).__init__(**kwargs)\n        self.num_classes = num_classes\n        self.threshold = threshold\n        self.stateful = True\n        self.name = 'f1'\n\n    def reset_states(self):\n        K.get_session().run(tf.variables_initializer(self.local_variables))\n\n    def metric_variable(self, shape, dtype, validate_shape=True, name=None):\n        return tf.Variable(\n                np.zeros(shape),\n                dtype=dtype,\n                trainable=False,\n                collections=[tf.GraphKeys.LOCAL_VARIABLES],\n                validate_shape=validate_shape,\n                name=name,\n                )\n\n    def streaming_counts(self, y_true, y_pred):\n        self.tp_mac = self.metric_variable(\n            shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"tp_mac\"\n        )\n        self.fp_mac = self.metric_variable(\n            shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fp_mac\"\n        )\n        self.fn_mac = self.metric_variable(\n            shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fn_mac\"\n        )\n\n        up_tp_mac = tf.assign_add(self.tp_mac, tf.count_nonzero(y_pred * y_true, axis=0))\n        self.add_update(up_tp_mac)\n        up_fp_mac = tf.assign_add(self.fp_mac, tf.count_nonzero(y_pred * (y_true - 1), axis=0))\n        self.add_update(up_fp_mac)\n        up_fn_mac = tf.assign_add(self.fn_mac, tf.count_nonzero((y_pred - 1) * y_true, axis=0))\n        self.add_update(up_fn_mac)        \n\n        self.local_variables = tf.get_collection(tf.GraphKeys.LOCAL_VARIABLES)\n\n    def __call__(self, y_true, y_pred):\n            rounded_pred = K.cast(K.greater_equal(y_pred, self.threshold), 'float32')\n            self.streaming_counts(y_true, rounded_pred)\n            prec_mac = self.tp_mac / (self.tp_mac + self.fp_mac)\n            rec_mac = self.tp_mac / (self.tp_mac + self.fn_mac)\n            f1_mac = 2 * prec_mac * rec_mac / (prec_mac + rec_mac)\n            f1_mac = tf.reduce_mean(tf.where(tf.is_nan(f1_mac), tf.zeros_like(f1_mac),f1_mac))\n            return f1_mac\n</code></pre>",
      "rawMarkdown": "Hi, Brian\nI think you can avoid run another validation generator loop by using stateful metrics. Assuming tensorflow as a backend, it already have a 'micro' f1 score: `tf.contrib.metrics.f1_score` but [this article](http://vict0rsch.github.io/2018/06/06/tensorflow-streaming-multilabel-f1/) explains about streaming metrics and implements a macro version. I'm using this class as f1_metric:  \n\n    class Metrics(keras.layers.Layer):\n        def __init__(self, num_classes, threshold, **kwargs):\n            super(Metrics, self).__init__(**kwargs)\n            self.num_classes = num_classes\n            self.threshold = threshold\n            self.stateful = True\n            self.name = 'f1'\n    \n        def reset_states(self):\n            K.get_session().run(tf.variables_initializer(self.local_variables))\n            \n        def metric_variable(self, shape, dtype, validate_shape=True, name=None):\n            return tf.Variable(\n                    np.zeros(shape),\n                    dtype=dtype,\n                    trainable=False,\n                    collections=[tf.GraphKeys.LOCAL_VARIABLES],\n                    validate_shape=validate_shape,\n                    name=name,\n                    )\n            \n        def streaming_counts(self, y_true, y_pred):\n            self.tp_mac = self.metric_variable(\n                shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"tp_mac\"\n            )\n            self.fp_mac = self.metric_variable(\n                shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fp_mac\"\n            )\n            self.fn_mac = self.metric_variable(\n                shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fn_mac\"\n            )\n            \n            up_tp_mac = tf.assign_add(self.tp_mac, tf.count_nonzero(y_pred * y_true, axis=0))\n            self.add_update(up_tp_mac)\n            up_fp_mac = tf.assign_add(self.fp_mac, tf.count_nonzero(y_pred * (y_true - 1), axis=0))\n            self.add_update(up_fp_mac)\n            up_fn_mac = tf.assign_add(self.fn_mac, tf.count_nonzero((y_pred - 1) * y_true, axis=0))\n            self.add_update(up_fn_mac)        \n            \n            self.local_variables = tf.get_collection(tf.GraphKeys.LOCAL_VARIABLES)\n            \n        def __call__(self, y_true, y_pred):\n                rounded_pred = K.cast(K.greater_equal(y_pred, self.threshold), 'float32')\n                self.streaming_counts(y_true, rounded_pred)\n                prec_mac = self.tp_mac / (self.tp_mac + self.fp_mac)\n                rec_mac = self.tp_mac / (self.tp_mac + self.fn_mac)\n                f1_mac = 2 * prec_mac * rec_mac / (prec_mac + rec_mac)\n                f1_mac = tf.reduce_mean(tf.where(tf.is_nan(f1_mac), tf.zeros_like(f1_mac),f1_mac))\n                return f1_mac",
      "votes": null
    },
    {
      "id": "435849",
      "postDate": "12/08/2018 22:34:36",
      "content": "<p>I'm have to give that one a try. I am indeed using a tf backend. I'm training up a new model that this will be helpful for. </p>",
      "rawMarkdown": "I'm have to give that one a try. I am indeed using a tf backend. I'm training up a new model that this will be helpful for.",
      "votes": null
    },
    {
      "id": "436854",
      "postDate": "12/11/2018 03:08:24",
      "content": "<p>Hi Sergey,  I've tried your custom callback but got an error like this:\n'  x_val = self.validation_data[0]\nTypeError: 'NoneType' object is not subscriptable'</p>",
      "rawMarkdown": "Hi Sergey,  I've tried your custom callback but got an error like this:\n'  x_val = self.validation_data[0]\nTypeError: 'NoneType' object is not subscriptable'",
      "votes": null
    },
    {
      "id": "437146",
      "postDate": "12/11/2018 12:52:12",
      "content": "<p>It depends on keras version, try self.model.validationdata[0] </p>",
      "rawMarkdown": "It depends on keras version, try self.model.validationdata[0]",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 434680,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "12/06/2018 19:23:32",
      "content": "<p>I was doing something similar but ran into issues when using a generator for validation data.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 434704,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "12/06/2018 20:23:41",
      "content": "<p>That's a good point. Some of it was discussed early on <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242\"><strong>here</strong></a>. Yet another solution is presented <a href=\"https://bit.ly/2n8buDO\"><strong>here</strong></a>.</p>\n\n<p>As <a href=\"/ldm314\">@ldm314</a> (Brian) pointed out, it is non-trivial to set this up when using generators.</p>",
      "votes": null,
      "replies": [
        {
          "id": 434719,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/06/2018 20:58:19",
          "content": "<p>The 2nd link also has the same problem. I based my example on that one. Here is it with the modifications, but it doubles the validation time. It also isn't very clean since it uses the global 'valid_gen'. It worked ok, but I ended up just not using it because of the performance.</p>\n\n<pre>class Metrics(Callback):\n    def on_train_begin(self, logs={}):\n        self.val_f1s = []\n        self.val_recalls = []\n        self.val_precisions = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        t_y = None\n        t_x = None\n        preds_y = []\n        valid_y = []\n        valid_gen.batch_size = VALID_BATCH_SIZE\n        batches = valid_gen.n//VALID_BATCH_SIZE+1\n        i=1\n        for _, (t_x, t_names) in zip((range(valid_gen.n//VALID_BATCH_SIZE+1)),\n                                    valid_gen):\n            t_y = self.model.predict(t_x)\n            for c_id, c_score in zip(t_names, t_y):\n                valid_y.append(c_id)\n                preds_y.append(c_score)\n            print('batch %d/%d' % (i,batches))\n            i+=1\n\n        predicted = np.array(preds_y)\n        val_targ = np.array(valid_y)\n        max_val = np.max(predicted)\n        val_predict = predicted &gt; (0.65 * max_val)\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        self.val_f1s.append(_val_f1)\n        self.val_recalls.append(_val_recall)\n        self.val_precisions.append(_val_precision)\n        print(\"val_f1: %f — val_precision: %f — val_recall %f\" %(_val_f1, _val_precision, _val_recall))\n        return\n</pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435811,
          "author_name": "vhessel",
          "author_url": "",
          "post_date": "12/08/2018 20:22:07",
          "content": "<p>Hi, Brian\nI think you can avoid run another validation generator loop by using stateful metrics. Assuming tensorflow as a backend, it already have a 'micro' f1 score: <code>tf.contrib.metrics.f1_score</code> but <a href=\"http://vict0rsch.github.io/2018/06/06/tensorflow-streaming-multilabel-f1/\">this article</a> explains about streaming metrics and implements a macro version. I'm using this class as f1_metric:  </p>\n\n<pre><code>class Metrics(keras.layers.Layer):\n    def __init__(self, num_classes, threshold, **kwargs):\n        super(Metrics, self).__init__(**kwargs)\n        self.num_classes = num_classes\n        self.threshold = threshold\n        self.stateful = True\n        self.name = 'f1'\n\n    def reset_states(self):\n        K.get_session().run(tf.variables_initializer(self.local_variables))\n\n    def metric_variable(self, shape, dtype, validate_shape=True, name=None):\n        return tf.Variable(\n                np.zeros(shape),\n                dtype=dtype,\n                trainable=False,\n                collections=[tf.GraphKeys.LOCAL_VARIABLES],\n                validate_shape=validate_shape,\n                name=name,\n                )\n\n    def streaming_counts(self, y_true, y_pred):\n        self.tp_mac = self.metric_variable(\n            shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"tp_mac\"\n        )\n        self.fp_mac = self.metric_variable(\n            shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fp_mac\"\n        )\n        self.fn_mac = self.metric_variable(\n            shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fn_mac\"\n        )\n\n        up_tp_mac = tf.assign_add(self.tp_mac, tf.count_nonzero(y_pred * y_true, axis=0))\n        self.add_update(up_tp_mac)\n        up_fp_mac = tf.assign_add(self.fp_mac, tf.count_nonzero(y_pred * (y_true - 1), axis=0))\n        self.add_update(up_fp_mac)\n        up_fn_mac = tf.assign_add(self.fn_mac, tf.count_nonzero((y_pred - 1) * y_true, axis=0))\n        self.add_update(up_fn_mac)        \n\n        self.local_variables = tf.get_collection(tf.GraphKeys.LOCAL_VARIABLES)\n\n    def __call__(self, y_true, y_pred):\n            rounded_pred = K.cast(K.greater_equal(y_pred, self.threshold), 'float32')\n            self.streaming_counts(y_true, rounded_pred)\n            prec_mac = self.tp_mac / (self.tp_mac + self.fp_mac)\n            rec_mac = self.tp_mac / (self.tp_mac + self.fn_mac)\n            f1_mac = 2 * prec_mac * rec_mac / (prec_mac + rec_mac)\n            f1_mac = tf.reduce_mean(tf.where(tf.is_nan(f1_mac), tf.zeros_like(f1_mac),f1_mac))\n            return f1_mac\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 435849,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "12/08/2018 22:34:36",
          "content": "<p>I'm have to give that one a try. I am indeed using a tf backend. I'm training up a new model that this will be helpful for. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 435115,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "12/07/2018 14:23:09",
      "content": "<p>Thank you</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 435295,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "12/07/2018 20:08:12",
      "content": "<p>Hey Sergey, thanks again. I have one small question. Does the threshold change or should that be left as is? </p>",
      "votes": null,
      "replies": [
        {
          "id": 435532,
          "author_name": "malyutins",
          "author_url": "",
          "post_date": "12/08/2018 07:18:23",
          "content": "<p>Threshold selection is another non trivial problem :) It is better to try different options. Once you have <code>y_pred</code> you can call <code>f1_score</code> with different values of <code>threshold</code> then print and log all of them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436854,
      "author_name": "vikeezhou",
      "author_url": "",
      "post_date": "12/11/2018 03:08:24",
      "content": "<p>Hi Sergey,  I've tried your custom callback but got an error like this:\n'  x_val = self.validation_data[0]\nTypeError: 'NoneType' object is not subscriptable'</p>",
      "votes": null,
      "replies": [
        {
          "id": 437146,
          "author_name": "malyutins",
          "author_url": "",
          "post_date": "12/11/2018 12:52:12",
          "content": "<p>It depends on keras version, try self.model.validationdata[0] </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "434679": "It is known that *F1* score is assumed to be 0 if class has no *true positive* samples. If we choose a small batch size then it is highly possible that rare classes will not be present in it and thus *F1* score will be 0 for them. That will affect *F1* macro metric computed as an average among all *F1* scores for all classes. You can see the demo in **[that kernel][1]**\n\nWhen we train a network using *keras* all the metrics are computed after each batch (**`on_batch_end`**) and get averaged and printed to console at the end of epoch (**`on_epoch_end`**). Due to the effect of small *F1* score for small batch even for the correct predictions, displayed score will always be lower than real.\n\nTo avoid such behavior we can create a callback to compute score for the entire validation set as follows:\n\n    class CustomValidationCallback(Callback):\n    \n        def on_epoch_end(self, epoch, logs={}):\n            threshold = 0.1\n            x_val = self.validation_data[0]\n            y_true = self.validation_data[1]\n            y_pred = self.model.predict(x_val)\n            f1 = f1_score(y_true, (y_pred &gt; threshold).astype(int), average='macro')\n            logs['val_f1_custom'] = f1\n            print('Epoch %05d: custom f1 score %0.3f' % (epoch + 1, round(f1, 3))) \n\nCreate an instance and pass it to fit method:\n\n    custom_val_callback = CustomValidationCallback()\n    \n    model.fit_generator(... , callbacks=[custom_val_callback, ...], ...)\n\n\n  [1]: https://www.kaggle.com/malyutins/f1-macro-score-on-small-batches",
    "434680": "I was doing something similar but ran into issues when using a generator for validation data.",
    "434704": "That's a good point. Some of it was discussed early on [__here__](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242). Yet another solution is presented [__here__](https://bit.ly/2n8buDO).\n\nAs @ldm314 (Brian) pointed out, it is non-trivial to set this up when using generators.",
    "434719": "The 2nd link also has the same problem. I based my example on that one. Here is it with the modifications, but it doubles the validation time. It also isn't very clean since it uses the global 'valid_gen'. It worked ok, but I ended up just not using it because of the performance.\n\n<pre>class Metrics(Callback):\n    def on_train_begin(self, logs={}):\n        self.val_f1s = []\n        self.val_recalls = []\n        self.val_precisions = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        t_y = None\n        t_x = None\n        preds_y = []\n        valid_y = []\n        valid_gen.batch_size = VALID_BATCH_SIZE\n        batches = valid_gen.n//VALID_BATCH_SIZE+1\n        i=1\n        for _, (t_x, t_names) in zip((range(valid_gen.n//VALID_BATCH_SIZE+1)),\n                                    valid_gen):\n            t_y = self.model.predict(t_x)\n            for c_id, c_score in zip(t_names, t_y):\n                valid_y.append(c_id)\n                preds_y.append(c_score)\n            print('batch %d/%d' % (i,batches))\n            i+=1\n\n        predicted = np.array(preds_y)\n        val_targ = np.array(valid_y)\n        max_val = np.max(predicted)\n        val_predict = predicted &gt; (0.65 * max_val)\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        self.val_f1s.append(_val_f1)\n        self.val_recalls.append(_val_recall)\n        self.val_precisions.append(_val_precision)\n        print(\"val_f1: %f — val_precision: %f — val_recall %f\" %(_val_f1, _val_precision, _val_recall))\n        return\n</pre>",
    "435115": "Thank you",
    "435295": "Hey Sergey, thanks again. I have one small question. Does the threshold change or should that be left as is?",
    "435532": "Threshold selection is another non trivial problem :) It is better to try different options. Once you have `y_pred` you can call `f1_score` with different values of `threshold` then print and log all of them.",
    "435811": "Hi, Brian\nI think you can avoid run another validation generator loop by using stateful metrics. Assuming tensorflow as a backend, it already have a 'micro' f1 score: `tf.contrib.metrics.f1_score` but [this article](http://vict0rsch.github.io/2018/06/06/tensorflow-streaming-multilabel-f1/) explains about streaming metrics and implements a macro version. I'm using this class as f1_metric:  \n\n    class Metrics(keras.layers.Layer):\n        def __init__(self, num_classes, threshold, **kwargs):\n            super(Metrics, self).__init__(**kwargs)\n            self.num_classes = num_classes\n            self.threshold = threshold\n            self.stateful = True\n            self.name = 'f1'\n    \n        def reset_states(self):\n            K.get_session().run(tf.variables_initializer(self.local_variables))\n            \n        def metric_variable(self, shape, dtype, validate_shape=True, name=None):\n            return tf.Variable(\n                    np.zeros(shape),\n                    dtype=dtype,\n                    trainable=False,\n                    collections=[tf.GraphKeys.LOCAL_VARIABLES],\n                    validate_shape=validate_shape,\n                    name=name,\n                    )\n            \n        def streaming_counts(self, y_true, y_pred):\n            self.tp_mac = self.metric_variable(\n                shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"tp_mac\"\n            )\n            self.fp_mac = self.metric_variable(\n                shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fp_mac\"\n            )\n            self.fn_mac = self.metric_variable(\n                shape=[self.num_classes], dtype=tf.int64, validate_shape=False, name=\"fn_mac\"\n            )\n            \n            up_tp_mac = tf.assign_add(self.tp_mac, tf.count_nonzero(y_pred * y_true, axis=0))\n            self.add_update(up_tp_mac)\n            up_fp_mac = tf.assign_add(self.fp_mac, tf.count_nonzero(y_pred * (y_true - 1), axis=0))\n            self.add_update(up_fp_mac)\n            up_fn_mac = tf.assign_add(self.fn_mac, tf.count_nonzero((y_pred - 1) * y_true, axis=0))\n            self.add_update(up_fn_mac)        \n            \n            self.local_variables = tf.get_collection(tf.GraphKeys.LOCAL_VARIABLES)\n            \n        def __call__(self, y_true, y_pred):\n                rounded_pred = K.cast(K.greater_equal(y_pred, self.threshold), 'float32')\n                self.streaming_counts(y_true, rounded_pred)\n                prec_mac = self.tp_mac / (self.tp_mac + self.fp_mac)\n                rec_mac = self.tp_mac / (self.tp_mac + self.fn_mac)\n                f1_mac = 2 * prec_mac * rec_mac / (prec_mac + rec_mac)\n                f1_mac = tf.reduce_mean(tf.where(tf.is_nan(f1_mac), tf.zeros_like(f1_mac),f1_mac))\n                return f1_mac",
    "435849": "I'm have to give that one a try. I am indeed using a tf backend. I'm training up a new model that this will be helpful for.",
    "436854": "Hi Sergey,  I've tried your custom callback but got an error like this:\n'  x_val = self.validation_data[0]\nTypeError: 'NoneType' object is not subscriptable'",
    "437146": "It depends on keras version, try self.model.validationdata[0]"
  },
  "source": "meta"
}