{
  "id": 67835,
  "title": "Note about metric",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/67835",
  "author_name": "Sergey Bryansky",
  "post_date": "2018-10-06T00:12:06.391000",
  "votes": 15,
  "comment_count": 28,
  "views": 0,
  "content": "<p>In this competition metric is Macro F-Score which is really tunable metric. And for beginners, I just want to mention that this metric penalizes more if we can't good predict the truth labels, but less penalizes if we predict both the truth and wrong labels. So, there is the <strong>all labels benchmark, [.111]</strong>:</p>\n\n<pre><code>import pandas as pd\nsubmission = pd.read_csv(\"../input/sample_submission.csv\")\npred = \" \".join((str(i) for i in range(28)))  # n_labels = 28\nsubmission[\"Predicted\"] = [pred for _ in range(len(submission))]\nsubmission.to_csv(\"all_labels_benchmark.csv\", index=None)\n</code></pre>\n\n<p>Happy tuning!</p>",
  "messages": [
    {
      "id": 399510,
      "postDate": "2018-10-06T00:12:06.393Z",
      "content": "<p>In this competition metric is Macro F-Score which is really tunable metric. And for beginners, I just want to mention that this metric penalizes more if we can't good predict the truth labels, but less penalizes if we predict both the truth and wrong labels. So, there is the <strong>all labels benchmark, [.111]</strong>:</p>\n\n<pre><code>import pandas as pd\nsubmission = pd.read_csv(\"../input/sample_submission.csv\")\npred = \" \".join((str(i) for i in range(28)))  # n_labels = 28\nsubmission[\"Predicted\"] = [pred for _ in range(len(submission))]\nsubmission.to_csv(\"all_labels_benchmark.csv\", index=None)\n</code></pre>\n\n<p>Happy tuning!</p>",
      "rawMarkdown": "In this competition metric is Macro F-Score which is really tunable metric. And for beginners, I just want to mention that this metric penalizes more if we can't good predict the truth labels, but less penalizes if we predict both the truth and wrong labels. So, there is the **all labels benchmark, [.111]**:\n\n\n    import pandas as pd\n    submission = pd.read_csv(\"../input/sample_submission.csv\")\n    pred = \" \".join((str(i) for i in range(28)))  # n_labels = 28\n    submission[\"Predicted\"] = [pred for _ in range(len(submission))]\n    submission.to_csv(\"all_labels_benchmark.csv\", index=None)\n\nHappy tuning!",
      "votes": 15
    },
    {
      "id": 399803,
      "postDate": "2018-10-06T20:04:43.433Z",
      "content": "<p>Hello Sergey,</p>\n\n<p>If I use f1_score from sklearn.metrics and use 'macro' as average, is this the exact metric used by Kaggle?</p>\n\n<p>Documentation:\n<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html</a></p>",
      "rawMarkdown": "Hello Sergey,\n\nIf I use f1_score from sklearn.metrics and use 'macro' as average, is this the exact metric used by Kaggle?\n\nDocumentation:\nhttp://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\n",
      "votes": 7,
      "replies": [
        {
          "id": 399838,
          "postDate": "2018-10-06T22:27:37.953Z",
          "content": "<p>Hello. Yes, it is the same thing!</p>",
          "rawMarkdown": "Hello. Yes, it is the same thing!",
          "votes": 5
        },
        {
          "id": 400621,
          "postDate": "2018-10-08T16:26:52.827Z",
          "content": "<p>Awesome! Thanks!</p>",
          "rawMarkdown": "Awesome! Thanks!",
          "votes": 1
        }
      ]
    },
    {
      "id": 400555,
      "postDate": "2018-10-08T14:35:13.660Z",
      "content": "<p>Did anybody try to compare their validation score (calculated with sklearn.metrics f1_score) and the public LB score? Somehow I'm getting very different result and not sure if test set is completely different from train/validation, or  sklearn.metrics f1_score is not the metric used for score evaluation in the competition, or I have some bug in test data evaluation.</p>",
      "rawMarkdown": "Did anybody try to compare their validation score (calculated with sklearn.metrics f1_score) and the public LB score? Somehow I'm getting very different result and not sure if test set is completely different from train/validation, or  sklearn.metrics f1_score is not the metric used for score evaluation in the competition, or I have some bug in test data evaluation.",
      "votes": 3,
      "replies": [
        {
          "id": 400623,
          "postDate": "2018-10-08T16:33:00.183Z",
          "content": "<p>Did you specify average='macro' in your sklearn.metrics.f1_score call? That is the average used by Kaggle and necessary for multiclass classification.</p>\n\n<p>Your score might be calculated for 'binary' classification if you don't specify the average argument.</p>",
          "rawMarkdown": "Did you specify average='macro' in your sklearn.metrics.f1_score call? That is the average used by Kaggle and necessary for multiclass classification.\n\nYour score might be calculated for 'binary' classification if you don't specify the average argument.",
          "votes": 4
        },
        {
          "id": 400727,
          "postDate": "2018-10-08T19:20:09.907Z",
          "content": "<p>I did: f1_score(pred&gt;0.44, y, average='macro'), where pred and y are arrays containing prediction and target vectors in the following form [[0,1,1,0,0...],[1,0,1,0,0...],...]. And I used ~10% of the provided data for validation, so the discrepancy is unlikely be resulted by statistical noise.</p>",
          "rawMarkdown": "I did: f1_score(pred&gt;0.44, y, average='macro'), where pred and y are arrays containing prediction and target vectors in the following form [[0,1,1,0,0...],[1,0,1,0,0...],...]. And I used ~10% of the provided data for validation, so the discrepancy is unlikely be resulted by statistical noise.\n\n"
        },
        {
          "id": 401118,
          "postDate": "2018-10-09T13:12:18.703Z",
          "content": "<p>Yes, my own implementation is equal to sklearn F1 and very different from LB. </p>\n\n<p>Actually the public LB is completely different from my metric. </p>\n\n<p>By the way, this is what I thought was the \"macro\" average metric from sklearn's description and wikipedia:</p>\n\n<pre><code>def competitionMetric(true,pred):\n    pred = K.cast(K.greater(pred, 0.5), K.floatx())\n\n    #f1 per feature\n    groundPositives = K.sum(true, axis=0) + K.epsilon()\n    correctPositives = K.sum(true * pred, axis=0)\n    predictedPositives = K.sum(pred, axis=0) + K.epsilon()\n\n    precision = correctPositives / predictedPositives\n    recall = correctPositives / groundPositives\n\n    m = (2 * precision * recall) / (precision + recall + K.epsilon())\n\n    #macro average\n    return K.mean(m)\n</code></pre>",
          "rawMarkdown": "Yes, my own implementation is equal to sklearn F1 and very different from LB. \n\nActually the public LB is completely different from my metric. \n\nBy the way, this is what I thought was the \"macro\" average metric from sklearn's description and wikipedia:\n\n\tdef competitionMetric(true,pred):\n\t\tpred = K.cast(K.greater(pred, 0.5), K.floatx())\n\n\t\t#f1 per feature\n\t\tgroundPositives = K.sum(true, axis=0) + K.epsilon()\n\t\tcorrectPositives = K.sum(true * pred, axis=0)\n\t\tpredictedPositives = K.sum(pred, axis=0) + K.epsilon()\n\n\t\tprecision = correctPositives / predictedPositives\n\t\trecall = correctPositives / groundPositives\n\n\t\tm = (2 * precision * recall) / (precision + recall + K.epsilon())\n\n\t\t#macro average\n\t\treturn K.mean(m)",
          "votes": 3
        },
        {
          "id": 401160,
          "postDate": "2018-10-09T14:32:48.160Z",
          "content": "<p>It looks similar to that I read about the metric. Thank you for letting me know that you also have such discrepancy.  What is value of the F1 score you are getting in validation?</p>\n\n<p>I still cannot get what is the problem. When I run my model on the validation set, it gives 0.63 F1 score (with using sklearn: f1_score(pred&gt;0.44, y, average='macro')) and 0.972 accuracy. However, when I submit the prediction, the result looks like it contains only random indexes... And the worst thing is that I cannot do visual analysis of what is going on in the validation and test.</p>",
          "rawMarkdown": "It looks similar to that I read about the metric. Thank you for letting me know that you also have such discrepancy.  What is value of the F1 score you are getting in validation?\n\nI still cannot get what is the problem. When I run my model on the validation set, it gives 0.63 F1 score (with using sklearn: f1_score(pred&gt;0.44, y, average='macro')) and 0.972 accuracy. However, when I submit the prediction, the result looks like it contains only random indexes... And the worst thing is that I cannot do visual analysis of what is going on in the validation and test."
        },
        {
          "id": 401201,
          "postDate": "2018-10-09T15:48:17.130Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 401232,
          "postDate": "2018-10-09T16:44:19.560Z",
          "content": "<p>Similar case here....</p>\n\n<p>My model is not well trained or very good yet, but I get 0.22+ in F1, but 0.06 (which is less than the all ones benchmark = 0.111) in public LB. </p>\n\n<p>Tests :</p>\n\n<ul>\n<li>Compared my implementation with F1 from sklearn with macro: exactly the same    </li>\n<li>Compared metrics for the entire prediction with metrics by small batches: similar results   </li>\n</ul>\n\n<p>I'm tempted to say that Kaggle's metric is not equal to sklearn's f1 + macro, but I need to wait for my model to have better results first.</p>",
          "rawMarkdown": "Similar case here....\n\nMy model is not well trained or very good yet, but I get 0.22+ in F1, but 0.06 (which is less than the all ones benchmark = 0.111) in public LB. \n\nTests :\n\n- Compared my implementation with F1 from sklearn with macro: exactly the same    \n- Compared metrics for the entire prediction with metrics by small batches: similar results   \n\nI'm tempted to say that Kaggle's metric is not equal to sklearn's f1 + macro, but I need to wait for my model to have better results first."
        },
        {
          "id": 401352,
          "postDate": "2018-10-09T22:34:07.030Z",
          "content": "<p>I also see a difference between sklearn f1+macro and LB scores.</p>",
          "rawMarkdown": "I also see a difference between sklearn f1+macro and LB scores.",
          "votes": 1
        },
        {
          "id": 401928,
          "postDate": "2018-10-10T22:05:03.443Z",
          "content": "<p>@lafoss you called the function with predictions as the first argument followed by the ground truth.  In the documentation it is the other way around:\nf1_score(y_true, y_pred,.....</p>",
          "rawMarkdown": "@lafoss you called the function with predictions as the first argument followed by the ground truth.  In the documentation it is the other way around:\nf1_score(y_true, y_pred,.....",
          "votes": 1
        },
        {
          "id": 401989,
          "postDate": "2018-10-11T02:19:01.117Z",
          "content": "<p>It is a good catch,  but when I rerun my  evaluation kernel it gave exactly the same(</p>",
          "rawMarkdown": "It is a good catch,  but when I rerun my  evaluation kernel it gave exactly the same(",
          "votes": 1
        },
        {
          "id": 402377,
          "postDate": "2018-10-11T15:52:48.553Z",
          "content": "<p>For reference, here is how I am using it. At the end of each epoch it prints out f1, precision, recall for 0.5, 0.4 and 0.6 thresholds:</p>\n\n<pre>import numpy as np\nfrom keras.callbacks import Callback\nfrom sklearn.metrics import f1_score, precision_score, recall_score\nclass Metrics(Callback):\n    def on_train_begin(self, logs={}):\n        self.val_f1s = []\n        self.val_recalls = []\n        self.val_precisions = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        val_targ = self.validation_data[1]\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.5\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        self.val_f1s.append(_val_f1)\n        self.val_recalls.append(_val_recall)\n        self.val_precisions.append(_val_precision)\n        print(\"val_f1: %f — val_precision: %f — val_recall %f\" %(_val_f1, _val_precision, _val_recall))\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.4\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        print(\"val_f1.4: %f — val_precision.4: %f — val_recall.4 %f\" %(_val_f1, _val_precision, _val_recall))\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.6\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        print(\"val_f1.6: %f — val_precision.6: %f — val_recall.6 %f\" %(_val_f1, _val_precision, _val_recall))\n        return\n\nmetrics_cb = Metrics()\n</pre>",
          "rawMarkdown": "For reference, here is how I am using it. At the end of each epoch it prints out f1, precision, recall for 0.5, 0.4 and 0.6 thresholds:\n\n<pre>import numpy as np\nfrom keras.callbacks import Callback\nfrom sklearn.metrics import f1_score, precision_score, recall_score\nclass Metrics(Callback):\n    def on_train_begin(self, logs={}):\n        self.val_f1s = []\n        self.val_recalls = []\n        self.val_precisions = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        val_targ = self.validation_data[1]\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.5\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        self.val_f1s.append(_val_f1)\n        self.val_recalls.append(_val_recall)\n        self.val_precisions.append(_val_precision)\n        print(\"val_f1: %f — val_precision: %f — val_recall %f\" %(_val_f1, _val_precision, _val_recall))\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.4\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        print(\"val_f1.4: %f — val_precision.4: %f — val_recall.4 %f\" %(_val_f1, _val_precision, _val_recall))\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.6\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        print(\"val_f1.6: %f — val_precision.6: %f — val_recall.6 %f\" %(_val_f1, _val_precision, _val_recall))\n        return\n \nmetrics_cb = Metrics()\n</pre>",
          "votes": 2
        },
        {
          "id": 408941,
          "postDate": "2018-10-23T16:16:43.360Z",
          "content": "<h1>solved</h1>\n\n<p>Ok, my issue was definitely <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\">this</a></p>\n\n<p>The order of the ids in the submission matters. </p>",
          "rawMarkdown": "# solved\n\nOk, my issue was definitely [this](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366)\n\nThe order of the ids in the submission matters. ",
          "votes": 2
        },
        {
          "id": 409074,
          "postDate": "2018-10-23T19:53:48.003Z",
          "content": "<p>Checked my submissions, all were in the \"right\" order. I still seen a .1-.2 score difference from local LB to public. </p>",
          "rawMarkdown": "Checked my submissions, all were in the \"right\" order. I still seen a .1-.2 score difference from local LB to public. "
        },
        {
          "id": 409082,
          "postDate": "2018-10-23T20:01:08.723Z",
          "content": "<p>But 0.43 vs 0.65 is much better than 0.06 vs 0.65))) And the remaining discrepancy may be attributed to a mismatch of train and test sets, as I wrote in on of comments in <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai</a>. </p>",
          "rawMarkdown": "But 0.43 vs 0.65 is much better than 0.06 vs 0.65))) And the remaining discrepancy may be attributed to a mismatch of train and test sets, as I wrote in on of comments in https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai. "
        },
        {
          "id": 415576,
          "postDate": "2018-11-05T10:50:14.917Z",
          "content": "<p>Hi @Brian, I tried your code but ran into some errors like \n\"  predict = np.asarray(self.model.predict(self.validation_data[0]))\nTypeError: 'NoneType' object is not subscriptable\"</p>",
          "rawMarkdown": "Hi @Brian, I tried your code but ran into some errors like \n\"  predict = np.asarray(self.model.predict(self.validation_data[0]))\nTypeError: 'NoneType' object is not subscriptable\""
        },
        {
          "id": 415898,
          "postDate": "2018-11-05T21:55:52.740Z",
          "content": "<p>For this to work you have to provide validation data to your fit method,  and it can't be a generator.  I'm working on fixing this, I'm not using it right now because I had to switch to using generators for validation. </p>",
          "rawMarkdown": "For this to work you have to provide validation data to your fit method,  and it can't be a generator.  I'm working on fixing this, I'm not using it right now because I had to switch to using generators for validation. "
        },
        {
          "id": 415940,
          "postDate": "2018-11-05T23:52:55.710Z",
          "content": "<p>[Edit]</p>\n\n<p>I was mistaken, I think. I sometimes test with small data for quickness and got confused by that. </p>",
          "rawMarkdown": "[Edit]\n\nI was mistaken, I think. I sometimes test with small data for quickness and got confused by that. "
        },
        {
          "id": 415942,
          "postDate": "2018-11-06T00:06:38.613Z",
          "content": "<p>With the code I posted here, in on_epoch_end it is relying on the callback's self.validation_data to exist and have the data. It works using fit_generator with the validataion_data parameter something like like: validation_data=(valid_x,valid_y). Keras lets you use a generator here, but internally Keras only sets the callback's validation_data if a generator wasn't passed in: <a href=\"https://github.com/keras-team/keras/blob/master/keras/engine/training_generator.py#L150-L151\">https://github.com/keras-team/keras/blob/master/keras/engine/training_generator.py#L150-L151</a></p>",
          "rawMarkdown": "With the code I posted here, in on_epoch_end it is relying on the callback's self.validation_data to exist and have the data. It works using fit_generator with the validataion_data parameter something like like: validation_data=(valid_x,valid_y). Keras lets you use a generator here, but internally Keras only sets the callback's validation_data if a generator wasn't passed in: https://github.com/keras-team/keras/blob/master/keras/engine/training_generator.py#L150-L151"
        },
        {
          "id": 415946,
          "postDate": "2018-11-06T00:14:02.310Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 415956,
          "postDate": "2018-11-06T00:36:46.987Z",
          "content": "<p>you can also print out area under precision-recall curve (pr auc). This function is available in sklearn as well. this is independent of threshold and is a good metric to monitor</p>",
          "rawMarkdown": "you can also print out area under precision-recall curve (pr auc). This function is available in sklearn as well. this is independent of threshold and is a good metric to monitor",
          "votes": 1
        }
      ]
    },
    {
      "id": 401226,
      "postDate": "2018-10-09T16:28:19.383Z",
      "content": "<p><a href=\"https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras\">Here</a> there is my keras implementation of the Macro F1-Score that has very similar value to the LB scores and the sklearn 'macro' f1.</p>",
      "rawMarkdown": "[Here][1] there is my keras implementation of the Macro F1-Score that has very similar value to the LB scores and the sklearn 'macro' f1.\n\n\n  [1]: https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras",
      "votes": 1
    },
    {
      "id": 407947,
      "postDate": "2018-10-22T03:23:40.020Z",
      "content": "<p>In fact, this so-called \"benchmark\" can be surpassed by using solely 25 instead of all numbers.\nFor reference who got LB .222: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678\">LB probing</a></p>",
      "rawMarkdown": "In fact, this so-called \"benchmark\" can be surpassed by using solely 25 instead of all numbers.\nFor reference who got LB .222: [LB probing][1]\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678",
      "replies": [
        {
          "id": 407955,
          "postDate": "2018-10-22T04:02:22.970Z",
          "content": "<p>The last column is the fraction of samples of a particular class in the public LB. If you submit a file with 25 predicted for all images, the public LB score is 0.013 that gives 0.222493888 fraction of 25 class. You can use it to adjust thresholds in you model.</p>",
          "rawMarkdown": "The last column is the fraction of samples of a particular class in the public LB. If you submit a file with 25 predicted for all images, the public LB score is 0.013 that gives 0.222493888 fraction of 25 class. You can use it to adjust thresholds in you model."
        },
        {
          "id": 407957,
          "postDate": "2018-10-22T04:09:00.397Z",
          "content": "<p>Sorry for misinterpreting your result. (But LB 0.013 is still better than 0.111) XD</p>",
          "rawMarkdown": "Sorry for misinterpreting your result. (But LB 0.013 is still better than 0.111) XD"
        },
        {
          "id": 407959,
          "postDate": "2018-10-22T04:15:45.657Z",
          "content": "<p>Here is the code:</p>\n\n<pre><code>import pandas as pd\nsub = pd.read_csv(\"sample_submission.csv\")\npred = \" \".join(str(25))\nsub[\"Predicted\"] = [pred for _ in range(len(sub))]\nsub.to_csv(\"25_labels_benchmark.csv\", index=None)\n</code></pre>\n\n<p>But my submission got LB 0.010 which is different from your result @lafoss </p>\n\n<p>========================================================</p>\n\n<p>Edit: my bad. It is 0.013</p>\n\n<pre><code>&gt;&gt;&gt; import pandas as pd\n&gt;&gt;&gt; sub = pd.read_csv(\"sample_submission.csv\")\n&gt;&gt;&gt; pred = \"25\"\n&gt;&gt;&gt; sub[\"Predicted\"] = [pred for _ in range(len(sub))]\n&gt;&gt;&gt; sub.to_csv(\"25_labels_benchmark.csv\", index=None)\n&gt;&gt;&gt; exit()\n</code></pre>",
          "rawMarkdown": "Here is the code:\n\n    import pandas as pd\n    sub = pd.read_csv(\"sample_submission.csv\")\n    pred = \" \".join(str(25))\n    sub[\"Predicted\"] = [pred for _ in range(len(sub))]\n    sub.to_csv(\"25_labels_benchmark.csv\", index=None)\n\nBut my submission got LB 0.010 which is different from your result @lafoss \n\n\n========================================================\n\nEdit: my bad. It is 0.013\n\n    &gt;&gt;&gt; import pandas as pd\n    &gt;&gt;&gt; sub = pd.read_csv(\"sample_submission.csv\")\n    &gt;&gt;&gt; pred = \"25\"\n    &gt;&gt;&gt; sub[\"Predicted\"] = [pred for _ in range(len(sub))]\n    &gt;&gt;&gt; sub.to_csv(\"25_labels_benchmark.csv\", index=None)\n    &gt;&gt;&gt; exit()"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 399803,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2018-10-06T20:04:43.433000",
      "content": "<p>Hello Sergey,</p>\n\n<p>If I use f1_score from sklearn.metrics and use 'macro' as average, is this the exact metric used by Kaggle?</p>\n\n<p>Documentation:\n<a href=\"http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">http://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html</a></p>",
      "votes": 7,
      "replies": [
        {
          "id": 399838,
          "author_name": "Sergey Bryansky",
          "author_url": "",
          "post_date": "2018-10-06T22:27:37.953000",
          "content": "<p>Hello. Yes, it is the same thing!</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 400621,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2018-10-08T16:26:52.827000",
          "content": "<p>Awesome! Thanks!</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 400555,
      "author_name": "Iafoss",
      "author_url": "",
      "post_date": "2018-10-08T14:35:13.660000",
      "content": "<p>Did anybody try to compare their validation score (calculated with sklearn.metrics f1_score) and the public LB score? Somehow I'm getting very different result and not sure if test set is completely different from train/validation, or  sklearn.metrics f1_score is not the metric used for score evaluation in the competition, or I have some bug in test data evaluation.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 400623,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2018-10-08T16:33:00.183000",
          "content": "<p>Did you specify average='macro' in your sklearn.metrics.f1_score call? That is the average used by Kaggle and necessary for multiclass classification.</p>\n\n<p>Your score might be calculated for 'binary' classification if you don't specify the average argument.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 400727,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-10-08T19:20:09.907000",
          "content": "<p>I did: f1_score(pred&gt;0.44, y, average='macro'), where pred and y are arrays containing prediction and target vectors in the following form [[0,1,1,0,0...],[1,0,1,0,0...],...]. And I used ~10% of the provided data for validation, so the discrepancy is unlikely be resulted by statistical noise.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 401118,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-10-09T13:12:18.703000",
          "content": "<p>Yes, my own implementation is equal to sklearn F1 and very different from LB. </p>\n\n<p>Actually the public LB is completely different from my metric. </p>\n\n<p>By the way, this is what I thought was the \"macro\" average metric from sklearn's description and wikipedia:</p>\n\n<pre><code>def competitionMetric(true,pred):\n    pred = K.cast(K.greater(pred, 0.5), K.floatx())\n\n    #f1 per feature\n    groundPositives = K.sum(true, axis=0) + K.epsilon()\n    correctPositives = K.sum(true * pred, axis=0)\n    predictedPositives = K.sum(pred, axis=0) + K.epsilon()\n\n    precision = correctPositives / predictedPositives\n    recall = correctPositives / groundPositives\n\n    m = (2 * precision * recall) / (precision + recall + K.epsilon())\n\n    #macro average\n    return K.mean(m)\n</code></pre>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 401160,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-10-09T14:32:48.160000",
          "content": "<p>It looks similar to that I read about the metric. Thank you for letting me know that you also have such discrepancy.  What is value of the F1 score you are getting in validation?</p>\n\n<p>I still cannot get what is the problem. When I run my model on the validation set, it gives 0.63 F1 score (with using sklearn: f1_score(pred&gt;0.44, y, average='macro')) and 0.972 accuracy. However, when I submit the prediction, the result looks like it contains only random indexes... And the worst thing is that I cannot do visual analysis of what is going on in the validation and test.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 401201,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-10-09T15:48:17.130000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 401232,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-10-09T16:44:19.560000",
          "content": "<p>Similar case here....</p>\n\n<p>My model is not well trained or very good yet, but I get 0.22+ in F1, but 0.06 (which is less than the all ones benchmark = 0.111) in public LB. </p>\n\n<p>Tests :</p>\n\n<ul>\n<li>Compared my implementation with F1 from sklearn with macro: exactly the same    </li>\n<li>Compared metrics for the entire prediction with metrics by small batches: similar results   </li>\n</ul>\n\n<p>I'm tempted to say that Kaggle's metric is not equal to sklearn's f1 + macro, but I need to wait for my model to have better results first.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 401352,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-10-09T22:34:07.030000",
          "content": "<p>I also see a difference between sklearn f1+macro and LB scores.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 401928,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-10-10T22:05:03.443000",
          "content": "<p>@lafoss you called the function with predictions as the first argument followed by the ground truth.  In the documentation it is the other way around:\nf1_score(y_true, y_pred,.....</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 401989,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-10-11T02:19:01.117000",
          "content": "<p>It is a good catch,  but when I rerun my  evaluation kernel it gave exactly the same(</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 402377,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-10-11T15:52:48.553000",
          "content": "<p>For reference, here is how I am using it. At the end of each epoch it prints out f1, precision, recall for 0.5, 0.4 and 0.6 thresholds:</p>\n\n<pre>import numpy as np\nfrom keras.callbacks import Callback\nfrom sklearn.metrics import f1_score, precision_score, recall_score\nclass Metrics(Callback):\n    def on_train_begin(self, logs={}):\n        self.val_f1s = []\n        self.val_recalls = []\n        self.val_precisions = []\n\n    def on_epoch_end(self, epoch, logs={}):\n        val_targ = self.validation_data[1]\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.5\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        self.val_f1s.append(_val_f1)\n        self.val_recalls.append(_val_recall)\n        self.val_precisions.append(_val_precision)\n        print(\"val_f1: %f — val_precision: %f — val_recall %f\" %(_val_f1, _val_precision, _val_recall))\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.4\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        print(\"val_f1.4: %f — val_precision.4: %f — val_recall.4 %f\" %(_val_f1, _val_precision, _val_recall))\n        val_predict = (np.asarray((self.model.predict(self.validation_data[0])))) &gt; 0.6\n        _val_f1 = f1_score(val_targ, val_predict, average='macro')\n        _val_recall = recall_score(val_targ, val_predict, average='macro')\n        _val_precision = precision_score(val_targ, val_predict, average='macro')\n        print(\"val_f1.6: %f — val_precision.6: %f — val_recall.6 %f\" %(_val_f1, _val_precision, _val_recall))\n        return\n\nmetrics_cb = Metrics()\n</pre>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 408941,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-10-23T16:16:43.360000",
          "content": "<h1>solved</h1>\n\n<p>Ok, my issue was definitely <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69366\">this</a></p>\n\n<p>The order of the ids in the submission matters. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 409074,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-10-23T19:53:48.003000",
          "content": "<p>Checked my submissions, all were in the \"right\" order. I still seen a .1-.2 score difference from local LB to public. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 409082,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-10-23T20:01:08.723000",
          "content": "<p>But 0.43 vs 0.65 is much better than 0.06 vs 0.65))) And the remaining discrepancy may be attributed to a mismatch of train and test sets, as I wrote in on of comments in <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai</a>. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415576,
          "author_name": "vikeezhou",
          "author_url": "",
          "post_date": "2018-11-05T10:50:14.917000",
          "content": "<p>Hi @Brian, I tried your code but ran into some errors like \n\"  predict = np.asarray(self.model.predict(self.validation_data[0]))\nTypeError: 'NoneType' object is not subscriptable\"</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415898,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-05T21:55:52.740000",
          "content": "<p>For this to work you have to provide validation data to your fit method,  and it can't be a generator.  I'm working on fixing this, I'm not using it right now because I had to switch to using generators for validation. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415940,
          "author_name": "Daniel Möller",
          "author_url": "",
          "post_date": "2018-11-05T23:52:55.710000",
          "content": "<p>[Edit]</p>\n\n<p>I was mistaken, I think. I sometimes test with small data for quickness and got confused by that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415942,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-11-06T00:06:38.613000",
          "content": "<p>With the code I posted here, in on_epoch_end it is relying on the callback's self.validation_data to exist and have the data. It works using fit_generator with the validataion_data parameter something like like: validation_data=(valid_x,valid_y). Keras lets you use a generator here, but internally Keras only sets the callback's validation_data if a generator wasn't passed in: <a href=\"https://github.com/keras-team/keras/blob/master/keras/engine/training_generator.py#L150-L151\">https://github.com/keras-team/keras/blob/master/keras/engine/training_generator.py#L150-L151</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415946,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-11-06T00:14:02.310000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 415956,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2018-11-06T00:36:46.987000",
          "content": "<p>you can also print out area under precision-recall curve (pr auc). This function is available in sklearn as well. this is independent of threshold and is a good metric to monitor</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 401226,
      "author_name": "Guglielmo Camporese",
      "author_url": "",
      "post_date": "2018-10-09T16:28:19.383000",
      "content": "<p><a href=\"https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras\">Here</a> there is my keras implementation of the Macro F1-Score that has very similar value to the LB scores and the sklearn 'macro' f1.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 407947,
      "author_name": "Hanke Chen",
      "author_url": "",
      "post_date": "2018-10-22T03:23:40.020000",
      "content": "<p>In fact, this so-called \"benchmark\" can be surpassed by using solely 25 instead of all numbers.\nFor reference who got LB .222: <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678\">LB probing</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 407955,
          "author_name": "Iafoss",
          "author_url": "",
          "post_date": "2018-10-22T04:02:22.970000",
          "content": "<p>The last column is the fraction of samples of a particular class in the public LB. If you submit a file with 25 predicted for all images, the public LB score is 0.013 that gives 0.222493888 fraction of 25 class. You can use it to adjust thresholds in you model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 407957,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2018-10-22T04:09:00.397000",
          "content": "<p>Sorry for misinterpreting your result. (But LB 0.013 is still better than 0.111) XD</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 407959,
          "author_name": "Hanke Chen",
          "author_url": "",
          "post_date": "2018-10-22T04:15:45.657000",
          "content": "<p>Here is the code:</p>\n\n<pre><code>import pandas as pd\nsub = pd.read_csv(\"sample_submission.csv\")\npred = \" \".join(str(25))\nsub[\"Predicted\"] = [pred for _ in range(len(sub))]\nsub.to_csv(\"25_labels_benchmark.csv\", index=None)\n</code></pre>\n\n<p>But my submission got LB 0.010 which is different from your result @lafoss </p>\n\n<p>========================================================</p>\n\n<p>Edit: my bad. It is 0.013</p>\n\n<pre><code>&gt;&gt;&gt; import pandas as pd\n&gt;&gt;&gt; sub = pd.read_csv(\"sample_submission.csv\")\n&gt;&gt;&gt; pred = \"25\"\n&gt;&gt;&gt; sub[\"Predicted\"] = [pred for _ in range(len(sub))]\n&gt;&gt;&gt; sub.to_csv(\"25_labels_benchmark.csv\", index=None)\n&gt;&gt;&gt; exit()\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "399510": "In this competition metric is Macro F-Score which is really tunable metric. And for beginners, I just want to mention that this metric penalizes more if we can't good predict the truth labels, but less penalizes if we predict both the truth and wrong labels. So, there is the **all labels benchmark, [.111]**:\n\n\n    import pandas as pd\n    submission = pd.read_csv(\"../input/sample_submission.csv\")\n    pred = \" \".join((str(i) for i in range(28)))  # n_labels = 28\n    submission[\"Predicted\"] = [pred for _ in range(len(submission))]\n    submission.to_csv(\"all_labels_benchmark.csv\", index=None)\n\nHappy tuning!",
    "399803": "Hello Sergey,\n\nIf I use f1_score from sklearn.metrics and use 'macro' as average, is this the exact metric used by Kaggle?\n\nDocumentation:\nhttp://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\n",
    "400555": "Did anybody try to compare their validation score (calculated with sklearn.metrics f1_score) and the public LB score? Somehow I'm getting very different result and not sure if test set is completely different from train/validation, or  sklearn.metrics f1_score is not the metric used for score evaluation in the competition, or I have some bug in test data evaluation.",
    "401226": "[Here][1] there is my keras implementation of the Macro F1-Score that has very similar value to the LB scores and the sklearn 'macro' f1.\n\n\n  [1]: https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras",
    "407947": "In fact, this so-called \"benchmark\" can be surpassed by using solely 25 instead of all numbers.\nFor reference who got LB .222: [LB probing][1]\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678"
  }
}