{
  "id": 70379,
  "title": "Official confirmation on the F1 metric",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/70379",
  "author_name": "",
  "post_date": "2018-11-02T17:35:02.713234400Z",
  "votes": 10,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Sorry to start another thread on this and I know this has been asked a few times but for my own sanity I need to ask.</p>\n\n<p>After browsing the forums a bit I don't see anywhere where the <strong>competition host or kaggle</strong> have confirmed how they are calculating the final F1 metric (not what they say they are doing but what they are <strong>actually</strong> doing per their code), though there are many discussion on this. In particular most participants notice a large drop in local CV to public LB so I just wanted to double check this point. </p>\n\n<p>Can kaggle/the host confirm if either:</p>\n\n<ol>\n<li>We calculate macro f1_score over the whole [validation] set</li>\n<li>We calculate macro f1 score per image then average</li>\n</ol>\n\n<p>These are different and most participants are doing 1 but 2 has a <strong>much</strong> closer CV to LB relationship. I'm only banging this drum as I've never seen a competition which such a massive gap between CV and LB which is raising a few red flags for me.</p>\n\n<p>This quote from the host Emma Lundberg seems to indicate they at some point have calculated performance per image:</p>\n\n<blockquote>\n  <p>For this article we tested experts individually for their performance\n  on single images, which resulted in the macro-F1 of 0.71</p>\n</blockquote>\n\n<p>If they truly are doing 1 then this is one almighty switch-up from train to test or we are missing something.</p>\n\n<p>Related discussions for reference <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70053\">here</a>, <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67835\">here</a> and <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242\">here</a>.</p>\n\n<p>Thanks and apologies if I've missed something obvious,</p>\n\n<p>Mark</p>\n\n<p>P. S. Side point: I really don't understand why kaggle just don't provided a [numpy] version of the competition metric at the start to stop all this time-wasting. In TGS it was equally confusing.</p>",
  "messages": [
    {
      "id": "414404",
      "postDate": "11/02/2018 17:35:02",
      "content": "<p>Sorry to start another thread on this and I know this has been asked a few times but for my own sanity I need to ask.</p>\n\n<p>After browsing the forums a bit I don't see anywhere where the <strong>competition host or kaggle</strong> have confirmed how they are calculating the final F1 metric (not what they say they are doing but what they are <strong>actually</strong> doing per their code), though there are many discussion on this. In particular most participants notice a large drop in local CV to public LB so I just wanted to double check this point. </p>\n\n<p>Can kaggle/the host confirm if either:</p>\n\n<ol>\n<li>We calculate macro f1_score over the whole [validation] set</li>\n<li>We calculate macro f1 score per image then average</li>\n</ol>\n\n<p>These are different and most participants are doing 1 but 2 has a <strong>much</strong> closer CV to LB relationship. I'm only banging this drum as I've never seen a competition which such a massive gap between CV and LB which is raising a few red flags for me.</p>\n\n<p>This quote from the host Emma Lundberg seems to indicate they at some point have calculated performance per image:</p>\n\n<blockquote>\n  <p>For this article we tested experts individually for their performance\n  on single images, which resulted in the macro-F1 of 0.71</p>\n</blockquote>\n\n<p>If they truly are doing 1 then this is one almighty switch-up from train to test or we are missing something.</p>\n\n<p>Related discussions for reference <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70053\">here</a>, <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67835\">here</a> and <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242\">here</a>.</p>\n\n<p>Thanks and apologies if I've missed something obvious,</p>\n\n<p>Mark</p>\n\n<p>P. S. Side point: I really don't understand why kaggle just don't provided a [numpy] version of the competition metric at the start to stop all this time-wasting. In TGS it was equally confusing.</p>",
      "rawMarkdown": "Sorry to start another thread on this and I know this has been asked a few times but for my own sanity I need to ask.\n\nAfter browsing the forums a bit I don't see anywhere where the **competition host or kaggle** have confirmed how they are calculating the final F1 metric (not what they say they are doing but what they are **actually** doing per their code), though there are many discussion on this. In particular most participants notice a large drop in local CV to public LB so I just wanted to double check this point. \n\nCan kaggle/the host confirm if either:\n\n 1. We calculate macro f1_score over the whole [validation] set\n 2. We calculate macro f1 score per image then average\n\nThese are different and most participants are doing 1 but 2 has a **much** closer CV to LB relationship. I'm only banging this drum as I've never seen a competition which such a massive gap between CV and LB which is raising a few red flags for me.\n\nThis quote from the host Emma Lundberg seems to indicate they at some point have calculated performance per image:\n\n&gt; For this article we tested experts individually for their performance\n&gt; on single images, which resulted in the macro-F1 of 0.71\n\nIf they truly are doing 1 then this is one almighty switch-up from train to test or we are missing something.\n\nRelated discussions for reference [here][1], [here][2] and [here][3].\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70053\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67835\n  [3]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242\n\nThanks and apologies if I've missed something obvious,\n\nMark\n\nP. S. Side point: I really don't understand why kaggle just don't provided a [numpy] version of the competition metric at the start to stop all this time-wasting. In TGS it was equally confusing.",
      "votes": null
    },
    {
      "id": "414428",
      "postDate": "11/02/2018 18:57:43",
      "content": "<p>There's only one way to calculate F1 macro which is to calculate F1 for each sample and average afterwards. I assume you are describing F1 micro is your option #1. The large LB drop is due to the difference in train and test sets. It  could be due to different proportion of each class, different image statistics or something else. </p>",
      "rawMarkdown": "There's only one way to calculate F1 macro which is to calculate F1 for each sample and average afterwards. I assume you are describing F1 micro is your option #1. The large LB drop is due to the difference in train and test sets. It  could be due to different proportion of each class, different image statistics or something else.",
      "votes": null
    },
    {
      "id": "414430",
      "postDate": "11/02/2018 19:01:44",
      "content": "<p>Hi Mark,</p>\n\n<p>You can test your implementation by comparing your local score to the 'All class benchmark' on the public LB.\nCalculate your F1 on the training set (full set or randomly selected x labels). Use the train labels and the all-class predictions: <code>np.ones((x, 28))</code> Your local result should be close to the public LB ~0.111</p>\n\n<p>My local F1 is ~0.100, the difference could come from the slightly different class distribution between the train and test set. Or my F1 is wrong :)</p>\n\n<p>I haven't submitted anything yet, so I have no idea how much of a gap I have between local and public LB, but here is my implementation:</p>\n\n<p>```\ndef calc_macro_f1(predict, truth, threshold=0.5):</p>\n\n<pre><code>truth = truth &gt; 0.5\npredict = predict &gt; threshold\n\ntp = np.sum(truth &amp; predict, axis=0)\nfp = np.sum(~truth &amp; predict, axis=0)\nfn = np.sum(truth &amp; ~predict, axis=0)\n\nprecision = tp / (tp + fp + EPS)\nrecall = tp / (tp + fn + EPS)\n\nf1_scores = (2 * precision * recall) / (precision + recall + EPS)\n\nreturn np.mean(f1_scores)\n</code></pre>\n\n<p>```</p>\n\n<p><code>predict</code> and <code>truth</code> both have dimension of (#samples, #classes)</p>",
      "rawMarkdown": "Hi Mark,\n\nYou can test your implementation by comparing your local score to the 'All class benchmark' on the public LB.\nCalculate your F1 on the training set (full set or randomly selected x labels). Use the train labels and the all-class predictions: `np.ones((x, 28))` Your local result should be close to the public LB ~0.111\n\nMy local F1 is ~0.100, the difference could come from the slightly different class distribution between the train and test set. Or my F1 is wrong :)\n\nI haven't submitted anything yet, so I have no idea how much of a gap I have between local and public LB, but here is my implementation:\n\n```\ndef calc_macro_f1(predict, truth, threshold=0.5):\n\n    truth = truth &gt; 0.5\n    predict = predict &gt; threshold\n\n    tp = np.sum(truth &amp; predict, axis=0)\n    fp = np.sum(~truth &amp; predict, axis=0)\n    fn = np.sum(truth &amp; ~predict, axis=0)\n\n    precision = tp / (tp + fp + EPS)\n    recall = tp / (tp + fn + EPS)\n\n    f1_scores = (2 * precision * recall) / (precision + recall + EPS)\n\n    return np.mean(f1_scores)\n```\n\n`predict` and `truth` both have dimension of (#samples, #classes)",
      "votes": null
    },
    {
      "id": "414436",
      "postDate": "11/02/2018 19:06:04",
      "content": "<p><a href=\"/sakvaua\">@sakvaua</a>, you mean calculate F1 for each <strong>labels</strong>?</p>",
      "rawMarkdown": "sakvaua, you mean calculate F1 for each **labels**?",
      "votes": null
    },
    {
      "id": "414442",
      "postDate": "11/02/2018 19:14:59",
      "content": "<p>I might be missing something but:</p>\n\n<pre><code>from sklearn.metrics import f1_score, fbeta_score\nimport numpy as np\n\ndef f1_np(y_true, y_pred):\n    eps = 1e-8\n    y_pred = np.round(y_pred)\n    # This sums over all examples\n    tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n    fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n    fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)\n\n    p = tp / (tp + fp + eps)\n    r = tp / (tp + fn + eps)\n\n    f1 = 2*p*r / (p+r+eps)\n    f1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\n    return np.mean(f1)\n\n# Dummy data\ny_true = np.array([[1,0], [1,1], [0,1], [1,1]])\ny_pred = np.array([[1,0], [1,0], [0,1], [1,0]])\n\n# Standard sklearn with numpy (both agree) give f1 = 0.75\nf1_score(y_true, y_pred, average='macro')\nf1_np(y_true, y_pred)\n\n# f1 score per example then average\nres = 0\nfor i in range(y_pred.shape[0]):\n    res += f1_score(y_true[i,:], y_pred[i,:], average='macro')\n\nres /= y_pred.shape[0]\nres  # 0.6667\n</code></pre>\n\n<p>In other words if you calculate the metric on a mini-batch basis you will get a different answer.</p>\n\n<p>P. S. Thanks Peter - your implementation agrees with mine.</p>",
      "rawMarkdown": "I might be missing something but:\n\n    from sklearn.metrics import f1_score, fbeta_score\n    import numpy as np\n\n    def f1_np(y_true, y_pred):\n        eps = 1e-8\n        y_pred = np.round(y_pred)\n        # This sums over all examples\n        tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n        fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n        fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)\n    \n        p = tp / (tp + fp + eps)\n        r = tp / (tp + fn + eps)\n\n        f1 = 2*p*r / (p+r+eps)\n        f1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\n        return np.mean(f1)\n\n    # Dummy data\n    y_true = np.array([[1,0], [1,1], [0,1], [1,1]])\n    y_pred = np.array([[1,0], [1,0], [0,1], [1,0]])\n\n    # Standard sklearn with numpy (both agree) give f1 = 0.75\n    f1_score(y_true, y_pred, average='macro')\n    f1_np(y_true, y_pred)\n\n    # f1 score per example then average\n    res = 0\n    for i in range(y_pred.shape[0]):\n        res += f1_score(y_true[i,:], y_pred[i,:], average='macro')\n\n    res /= y_pred.shape[0]\n    res  # 0.6667\n\nIn other words if you calculate the metric on a mini-batch basis you will get a different answer.\n\nP. S. Thanks Peter - your implementation agrees with mine.",
      "votes": null
    },
    {
      "id": "414459",
      "postDate": "11/02/2018 20:04:33",
      "content": "<p>+1 on this. Some clarification would be nice. My public LB scores are still way off local using sklearn f1 macro.</p>",
      "rawMarkdown": "1 on this. Some clarification would be nice. My public LB scores are still way off local using sklearn f1 macro.",
      "votes": null
    },
    {
      "id": "414507",
      "postDate": "11/02/2018 22:40:00",
      "content": "<p>you can actually verify the correct metric by probing the LB.</p>\n\n<ol>\n<li><p>submit \"prediction= a single class A\" for all images</p></li>\n<li><p>submit \"prediction= a single class B\" for all images</p></li>\n<li><p>submit \"prediction= 2 classes A and B\" for all images</p></li>\n</ol>\n\n<p>if the F1 is averaged over the class, you can derived  3. from 1.,2.</p>\n\n<p>e.g. </p>\n\n<p>say for class-A: public LB = 0.01, then F1 for class A = 28*0.01</p>\n\n<p>and for class-B: public LB = 0.02, then F1 for class B = 28*0.02,</p>\n\n<p>then for class-A and B , i should public LB = ( (28*0.01) + (28*0.02) + 0 + ...0 )/28</p>",
      "rawMarkdown": "you can actually verify the correct metric by probing the LB.\n\n1. submit \"prediction= a single class A\" for all images\n\n2. submit \"prediction= a single class B\" for all images\n\n3. submit \"prediction= 2 classes A and B\" for all images\n\nif the F1 is averaged over the class, you can derived  3. from 1.,2.\n\ne.g. \n\nsay for class-A: public LB = 0.01, then F1 for class A = 28*0.01\n\nand for class-B: public LB = 0.02, then F1 for class B = 28*0.02,\n\nthen for class-A and B , i should public LB = ( (28*0.01) + (28*0.02) + 0 + ...0 )/28",
      "votes": null
    },
    {
      "id": "414517",
      "postDate": "11/02/2018 22:58:27",
      "content": "<p>regarding the difference in local and public LB, it could have been:</p>\n\n<ol>\n<li><p>train set A : train a model</p></li>\n<li><p>train set B: determine threshold to maximize F1</p></li>\n<li><p>validation set: measure F1</p></li>\n</ol>\n\n<p>you need to test the performance of your threshold on a separate set.</p>\n\n<p>lastly, there are duplicate train images. these should not be in both train and validation set.\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69979\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69979</a></p>\n\n<p>there are also \"near duplicate images\", this maybe reason why local LB score is high?\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/414517/10595/near_dulplicate.png\" alt=\"enter image description here\"></p>",
      "rawMarkdown": "regarding the difference in local and public LB, it could have been:\n\n1. train set A : train a model\n\n2. train set B: determine threshold to maximize F1\n\n3. validation set: measure F1\n\nyou need to test the performance of your threshold on a separate set.\n\nlastly, there are duplicate train images. these should not be in both train and validation set.\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69979\n\n\nthere are also \"near duplicate images\", this maybe reason why local LB score is high?\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/414517/10595/near_dulplicate.png",
      "votes": null
    },
    {
      "id": "414588",
      "postDate": "11/03/2018 05:06:19",
      "content": "<p>Since 7 classes are likely <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678\">missed in public LB</a> and get zero score when average is computed, for me it looks that</p>\n\n<pre><code>LB ~ val - 0.25 \n</code></pre>\n\n<p>SO, the ideal solution with 1.0 val can get only 0.75 in public LB, and the gap is just the result of missing classes.</p>",
      "rawMarkdown": "Since 7 classes are likely [missed in public LB][1] and get zero score when average is computed, for me it looks that\n\n    LB ~ val - 0.25 \n\nSO, the ideal solution with 1.0 val can get only 0.75 in public LB, and the gap is just the result of missing classes.\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678",
      "votes": null
    },
    {
      "id": "440487",
      "postDate": "12/17/2018 15:58:17",
      "content": "<p>Hi Mark,\nI am a newbie for this competition, I have some uncertainties about f1 score and I want to confirm with you.\n1. We should calculate macro f1_score over the whole [validation] set rather than per image or per batch.\n2. We should use a function like the following instead of f1_score in sklearn.metrics.\nI hope you can answer for me.</p>\n\n<p>def f1_np(y_true, y_pred):\n    eps = 1e-8\n    y_pred = np.round(y_pred)\n    # This sums over all examples\n    tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n    fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n    fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)</p>\n\n<pre><code>p = tp / (tp + fp + eps)\nr = tp / (tp + fn + eps)\n\nf1 = 2*p*r / (p+r+eps)\nf1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\nreturn np.mean(f1)\n</code></pre>",
      "rawMarkdown": "Hi Mark,\nI am a newbie for this competition, I have some uncertainties about f1 score and I want to confirm with you.\n1. We should calculate macro f1_score over the whole [validation] set rather than per image or per batch.\n2. We should use a function like the following instead of f1_score in sklearn.metrics.\nI hope you can answer for me.\n\ndef f1_np(y_true, y_pred):\n    eps = 1e-8\n    y_pred = np.round(y_pred)\n    # This sums over all examples\n    tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n    fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n    fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)\n\n    p = tp / (tp + fp + eps)\n    r = tp / (tp + fn + eps)\n\n    f1 = 2*p*r / (p+r+eps)\n    f1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\n    return np.mean(f1)",
      "votes": null
    },
    {
      "id": "440491",
      "postDate": "12/17/2018 16:02:03",
      "content": "<p>Hi Femi,</p>\n\n<p>You can use the sklearn version of f1_score, just set <code>average = 'macro'</code>.</p>\n\n<p>And yes, calculate over the entire validation set.</p>",
      "rawMarkdown": "Hi Femi,\n\nYou can use the sklearn version of f1_score, just set `average = 'macro'`.\n\nAnd yes, calculate over the entire validation set.",
      "votes": null
    },
    {
      "id": "440761",
      "postDate": "12/18/2018 00:48:44",
      "content": "<p>thank you very much for your help!</p>",
      "rawMarkdown": "thank you very much for your help!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 414428,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "11/02/2018 18:57:43",
      "content": "<p>There's only one way to calculate F1 macro which is to calculate F1 for each sample and average afterwards. I assume you are describing F1 micro is your option #1. The large LB drop is due to the difference in train and test sets. It  could be due to different proportion of each class, different image statistics or something else. </p>",
      "votes": null,
      "replies": [
        {
          "id": 414436,
          "author_name": "pestipeti",
          "author_url": "",
          "post_date": "11/02/2018 19:06:04",
          "content": "<p><a href=\"/sakvaua\">@sakvaua</a>, you mean calculate F1 for each <strong>labels</strong>?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 414442,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "11/02/2018 19:14:59",
          "content": "<p>I might be missing something but:</p>\n\n<pre><code>from sklearn.metrics import f1_score, fbeta_score\nimport numpy as np\n\ndef f1_np(y_true, y_pred):\n    eps = 1e-8\n    y_pred = np.round(y_pred)\n    # This sums over all examples\n    tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n    fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n    fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)\n\n    p = tp / (tp + fp + eps)\n    r = tp / (tp + fn + eps)\n\n    f1 = 2*p*r / (p+r+eps)\n    f1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\n    return np.mean(f1)\n\n# Dummy data\ny_true = np.array([[1,0], [1,1], [0,1], [1,1]])\ny_pred = np.array([[1,0], [1,0], [0,1], [1,0]])\n\n# Standard sklearn with numpy (both agree) give f1 = 0.75\nf1_score(y_true, y_pred, average='macro')\nf1_np(y_true, y_pred)\n\n# f1 score per example then average\nres = 0\nfor i in range(y_pred.shape[0]):\n    res += f1_score(y_true[i,:], y_pred[i,:], average='macro')\n\nres /= y_pred.shape[0]\nres  # 0.6667\n</code></pre>\n\n<p>In other words if you calculate the metric on a mini-batch basis you will get a different answer.</p>\n\n<p>P. S. Thanks Peter - your implementation agrees with mine.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 414430,
      "author_name": "pestipeti",
      "author_url": "",
      "post_date": "11/02/2018 19:01:44",
      "content": "<p>Hi Mark,</p>\n\n<p>You can test your implementation by comparing your local score to the 'All class benchmark' on the public LB.\nCalculate your F1 on the training set (full set or randomly selected x labels). Use the train labels and the all-class predictions: <code>np.ones((x, 28))</code> Your local result should be close to the public LB ~0.111</p>\n\n<p>My local F1 is ~0.100, the difference could come from the slightly different class distribution between the train and test set. Or my F1 is wrong :)</p>\n\n<p>I haven't submitted anything yet, so I have no idea how much of a gap I have between local and public LB, but here is my implementation:</p>\n\n<p>```\ndef calc_macro_f1(predict, truth, threshold=0.5):</p>\n\n<pre><code>truth = truth &gt; 0.5\npredict = predict &gt; threshold\n\ntp = np.sum(truth &amp; predict, axis=0)\nfp = np.sum(~truth &amp; predict, axis=0)\nfn = np.sum(truth &amp; ~predict, axis=0)\n\nprecision = tp / (tp + fp + EPS)\nrecall = tp / (tp + fn + EPS)\n\nf1_scores = (2 * precision * recall) / (precision + recall + EPS)\n\nreturn np.mean(f1_scores)\n</code></pre>\n\n<p>```</p>\n\n<p><code>predict</code> and <code>truth</code> both have dimension of (#samples, #classes)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414459,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "11/02/2018 20:04:33",
      "content": "<p>+1 on this. Some clarification would be nice. My public LB scores are still way off local using sklearn f1 macro.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414507,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/02/2018 22:40:00",
      "content": "<p>you can actually verify the correct metric by probing the LB.</p>\n\n<ol>\n<li><p>submit \"prediction= a single class A\" for all images</p></li>\n<li><p>submit \"prediction= a single class B\" for all images</p></li>\n<li><p>submit \"prediction= 2 classes A and B\" for all images</p></li>\n</ol>\n\n<p>if the F1 is averaged over the class, you can derived  3. from 1.,2.</p>\n\n<p>e.g. </p>\n\n<p>say for class-A: public LB = 0.01, then F1 for class A = 28*0.01</p>\n\n<p>and for class-B: public LB = 0.02, then F1 for class B = 28*0.02,</p>\n\n<p>then for class-A and B , i should public LB = ( (28*0.01) + (28*0.02) + 0 + ...0 )/28</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414517,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/02/2018 22:58:27",
      "content": "<p>regarding the difference in local and public LB, it could have been:</p>\n\n<ol>\n<li><p>train set A : train a model</p></li>\n<li><p>train set B: determine threshold to maximize F1</p></li>\n<li><p>validation set: measure F1</p></li>\n</ol>\n\n<p>you need to test the performance of your threshold on a separate set.</p>\n\n<p>lastly, there are duplicate train images. these should not be in both train and validation set.\n<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69979\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69979</a></p>\n\n<p>there are also \"near duplicate images\", this maybe reason why local LB score is high?\n  <img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/414517/10595/near_dulplicate.png\" alt=\"enter image description here\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 414588,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "11/03/2018 05:06:19",
      "content": "<p>Since 7 classes are likely <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678\">missed in public LB</a> and get zero score when average is computed, for me it looks that</p>\n\n<pre><code>LB ~ val - 0.25 \n</code></pre>\n\n<p>SO, the ideal solution with 1.0 val can get only 0.75 in public LB, and the gap is just the result of missing classes.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440487,
      "author_name": "femichen",
      "author_url": "",
      "post_date": "12/17/2018 15:58:17",
      "content": "<p>Hi Mark,\nI am a newbie for this competition, I have some uncertainties about f1 score and I want to confirm with you.\n1. We should calculate macro f1_score over the whole [validation] set rather than per image or per batch.\n2. We should use a function like the following instead of f1_score in sklearn.metrics.\nI hope you can answer for me.</p>\n\n<p>def f1_np(y_true, y_pred):\n    eps = 1e-8\n    y_pred = np.round(y_pred)\n    # This sums over all examples\n    tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n    fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n    fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)</p>\n\n<pre><code>p = tp / (tp + fp + eps)\nr = tp / (tp + fn + eps)\n\nf1 = 2*p*r / (p+r+eps)\nf1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\nreturn np.mean(f1)\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 440491,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "12/17/2018 16:02:03",
          "content": "<p>Hi Femi,</p>\n\n<p>You can use the sklearn version of f1_score, just set <code>average = 'macro'</code>.</p>\n\n<p>And yes, calculate over the entire validation set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440761,
          "author_name": "femichen",
          "author_url": "",
          "post_date": "12/18/2018 00:48:44",
          "content": "<p>thank you very much for your help!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "414404": "Sorry to start another thread on this and I know this has been asked a few times but for my own sanity I need to ask.\n\nAfter browsing the forums a bit I don't see anywhere where the **competition host or kaggle** have confirmed how they are calculating the final F1 metric (not what they say they are doing but what they are **actually** doing per their code), though there are many discussion on this. In particular most participants notice a large drop in local CV to public LB so I just wanted to double check this point. \n\nCan kaggle/the host confirm if either:\n\n 1. We calculate macro f1_score over the whole [validation] set\n 2. We calculate macro f1 score per image then average\n\nThese are different and most participants are doing 1 but 2 has a **much** closer CV to LB relationship. I'm only banging this drum as I've never seen a competition which such a massive gap between CV and LB which is raising a few red flags for me.\n\nThis quote from the host Emma Lundberg seems to indicate they at some point have calculated performance per image:\n\n&gt; For this article we tested experts individually for their performance\n&gt; on single images, which resulted in the macro-F1 of 0.71\n\nIf they truly are doing 1 then this is one almighty switch-up from train to test or we are missing something.\n\nRelated discussions for reference [here][1], [here][2] and [here][3].\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70053\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/67835\n  [3]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68242\n\nThanks and apologies if I've missed something obvious,\n\nMark\n\nP. S. Side point: I really don't understand why kaggle just don't provided a [numpy] version of the competition metric at the start to stop all this time-wasting. In TGS it was equally confusing.",
    "414428": "There's only one way to calculate F1 macro which is to calculate F1 for each sample and average afterwards. I assume you are describing F1 micro is your option #1. The large LB drop is due to the difference in train and test sets. It  could be due to different proportion of each class, different image statistics or something else.",
    "414430": "Hi Mark,\n\nYou can test your implementation by comparing your local score to the 'All class benchmark' on the public LB.\nCalculate your F1 on the training set (full set or randomly selected x labels). Use the train labels and the all-class predictions: `np.ones((x, 28))` Your local result should be close to the public LB ~0.111\n\nMy local F1 is ~0.100, the difference could come from the slightly different class distribution between the train and test set. Or my F1 is wrong :)\n\nI haven't submitted anything yet, so I have no idea how much of a gap I have between local and public LB, but here is my implementation:\n\n```\ndef calc_macro_f1(predict, truth, threshold=0.5):\n\n    truth = truth &gt; 0.5\n    predict = predict &gt; threshold\n\n    tp = np.sum(truth &amp; predict, axis=0)\n    fp = np.sum(~truth &amp; predict, axis=0)\n    fn = np.sum(truth &amp; ~predict, axis=0)\n\n    precision = tp / (tp + fp + EPS)\n    recall = tp / (tp + fn + EPS)\n\n    f1_scores = (2 * precision * recall) / (precision + recall + EPS)\n\n    return np.mean(f1_scores)\n```\n\n`predict` and `truth` both have dimension of (#samples, #classes)",
    "414436": "sakvaua, you mean calculate F1 for each **labels**?",
    "414442": "I might be missing something but:\n\n    from sklearn.metrics import f1_score, fbeta_score\n    import numpy as np\n\n    def f1_np(y_true, y_pred):\n        eps = 1e-8\n        y_pred = np.round(y_pred)\n        # This sums over all examples\n        tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n        fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n        fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)\n    \n        p = tp / (tp + fp + eps)\n        r = tp / (tp + fn + eps)\n\n        f1 = 2*p*r / (p+r+eps)\n        f1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\n        return np.mean(f1)\n\n    # Dummy data\n    y_true = np.array([[1,0], [1,1], [0,1], [1,1]])\n    y_pred = np.array([[1,0], [1,0], [0,1], [1,0]])\n\n    # Standard sklearn with numpy (both agree) give f1 = 0.75\n    f1_score(y_true, y_pred, average='macro')\n    f1_np(y_true, y_pred)\n\n    # f1 score per example then average\n    res = 0\n    for i in range(y_pred.shape[0]):\n        res += f1_score(y_true[i,:], y_pred[i,:], average='macro')\n\n    res /= y_pred.shape[0]\n    res  # 0.6667\n\nIn other words if you calculate the metric on a mini-batch basis you will get a different answer.\n\nP. S. Thanks Peter - your implementation agrees with mine.",
    "414459": "1 on this. Some clarification would be nice. My public LB scores are still way off local using sklearn f1 macro.",
    "414507": "you can actually verify the correct metric by probing the LB.\n\n1. submit \"prediction= a single class A\" for all images\n\n2. submit \"prediction= a single class B\" for all images\n\n3. submit \"prediction= 2 classes A and B\" for all images\n\nif the F1 is averaged over the class, you can derived  3. from 1.,2.\n\ne.g. \n\nsay for class-A: public LB = 0.01, then F1 for class A = 28*0.01\n\nand for class-B: public LB = 0.02, then F1 for class B = 28*0.02,\n\nthen for class-A and B , i should public LB = ( (28*0.01) + (28*0.02) + 0 + ...0 )/28",
    "414517": "regarding the difference in local and public LB, it could have been:\n\n1. train set A : train a model\n\n2. train set B: determine threshold to maximize F1\n\n3. validation set: measure F1\n\nyou need to test the performance of your threshold on a separate set.\n\nlastly, there are duplicate train images. these should not be in both train and validation set.\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69979\n\n\nthere are also \"near duplicate images\", this maybe reason why local LB score is high?\n  ![enter image description here][1]\n\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/414517/10595/near_dulplicate.png",
    "414588": "Since 7 classes are likely [missed in public LB][1] and get zero score when average is computed, for me it looks that\n\n    LB ~ val - 0.25 \n\nSO, the ideal solution with 1.0 val can get only 0.75 in public LB, and the gap is just the result of missing classes.\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678",
    "440487": "Hi Mark,\nI am a newbie for this competition, I have some uncertainties about f1 score and I want to confirm with you.\n1. We should calculate macro f1_score over the whole [validation] set rather than per image or per batch.\n2. We should use a function like the following instead of f1_score in sklearn.metrics.\nI hope you can answer for me.\n\ndef f1_np(y_true, y_pred):\n    eps = 1e-8\n    y_pred = np.round(y_pred)\n    # This sums over all examples\n    tp = np.sum((y_true*y_pred).astype('float'), axis=0)\n    fp = np.sum(((1-y_true)*y_pred).astype('float'), axis=0)\n    fn = np.sum((y_true*(1-y_pred)).astype('float'), axis=0)\n\n    p = tp / (tp + fp + eps)\n    r = tp / (tp + fn + eps)\n\n    f1 = 2*p*r / (p+r+eps)\n    f1 = np.where(np.isnan(f1), np.zeros_like(f1), f1)\n    return np.mean(f1)",
    "440491": "Hi Femi,\n\nYou can use the sklearn version of f1_score, just set `average = 'macro'`.\n\nAnd yes, calculate over the entire validation set.",
    "440761": "thank you very much for your help!"
  },
  "source": "meta"
}