{
  "id": 228422,
  "title": "[TO ORGANIZERS] Please clarify the evaluation metric",
  "url": "/competitions/plant-pathology-2021-fgvc8/discussion/228422",
  "author_name": "",
  "post_date": "2021-03-24T16:21:36.105633400Z",
  "votes": 18,
  "comment_count": 5,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/fruitpathology\" target=\"_blank\">@fruitpathology</a>, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, we need your help.</p>\n<p>From the recent comments on the <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227237\" target=\"_blank\">CV-LB Inconcistency</a></strong> topic it's still unclear whether the evaluation metric for the competition is a row-wise F1, i.e.:</p>\n<pre><code>score = f1_score(y_true, y_pred, average='micro')\n</code></pre>\n<p>or a column-wise (or label-wise) F1, i.e.:</p>\n<pre><code>score = f1_score(y_true, y_pred, average='macro')\n</code></pre>\n<p>Information on the <strong>Evaluation</strong> page is not full to make sure whether it is former or latter (and has a broken link). As there's up to a 10% difference between them clarification of the metric is also necessary for understanding the reasons for inconsistency.</p>",
  "messages": [
    {
      "id": "1251267",
      "postDate": "03/24/2021 16:21:36",
      "content": "<p><a href=\"https://www.kaggle.com/fruitpathology\" target=\"_blank\">@fruitpathology</a>, <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a>, we need your help.</p>\n<p>From the recent comments on the <strong><a href=\"https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227237\" target=\"_blank\">CV-LB Inconcistency</a></strong> topic it's still unclear whether the evaluation metric for the competition is a row-wise F1, i.e.:</p>\n<pre><code>score = f1_score(y_true, y_pred, average='micro')\n</code></pre>\n<p>or a column-wise (or label-wise) F1, i.e.:</p>\n<pre><code>score = f1_score(y_true, y_pred, average='macro')\n</code></pre>\n<p>Information on the <strong>Evaluation</strong> page is not full to make sure whether it is former or latter (and has a broken link). As there's up to a 10% difference between them clarification of the metric is also necessary for understanding the reasons for inconsistency.</p>",
      "rawMarkdown": "fruitpathology, @sohier, we need your help.\n\nFrom the recent comments on the **[CV-LB Inconcistency](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227237)** topic it's still unclear whether the evaluation metric for the competition is a row-wise F1, i.e.:\n```\nscore = f1_score(y_true, y_pred, average='micro')\n```\nor a column-wise (or label-wise) F1, i.e.:\n```\nscore = f1_score(y_true, y_pred, average='macro')\n```\nInformation on the **Evaluation** page is not full to make sure whether it is former or latter (and has a broken link). As there's up to a 10% difference between them clarification of the metric is also necessary for understanding the reasons for inconsistency.",
      "votes": null
    },
    {
      "id": "1253022",
      "postDate": "03/26/2021 09:50:10",
      "content": "<p>I believe it's 'macro' as in many of competitions, it's default.</p>",
      "rawMarkdown": "I believe it's 'macro' as in many of competitions, it's default.",
      "votes": null
    },
    {
      "id": "1253705",
      "postDate": "03/27/2021 01:09:55",
      "content": "<p>I agree, there is still and abnormally large gap between LB scores and CV scores. Fingers crossed that we hear back from the competition organizers!</p>",
      "rawMarkdown": "I agree, there is still and abnormally large gap between LB scores and CV scores. Fingers crossed that we hear back from the competition organizers!",
      "votes": null
    },
    {
      "id": "1264107",
      "postDate": "04/05/2021 22:01:03",
      "content": "<p>It's closest to sklearn's <code>average='samples'</code>, though there might be some small differences in our implementations.</p>",
      "rawMarkdown": "It's closest to sklearn's `average='samples'`, though there might be some small differences in our implementations.",
      "votes": null
    },
    {
      "id": "1264197",
      "postDate": "04/06/2021 01:25:13",
      "content": "<p>So strange</p>",
      "rawMarkdown": "So strange",
      "votes": null
    },
    {
      "id": "1264256",
      "postDate": "04/06/2021 02:50:45",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thank you for this clarification. Could you please specify the details of your implementation? For example, consider 3 images with labels <code>['scab', 'healthy', 'powdery_mildew complex']</code> and the predictions <code>['rust complex', 'healthy', 'powdery_mildew']</code>. What is f1 score for this case?</p>",
      "rawMarkdown": "Hi @sohier Thank you for this clarification. Could you please specify the details of your implementation? For example, consider 3 images with labels `['scab', 'healthy', 'powdery_mildew complex']` and the predictions `['rust complex', 'healthy', 'powdery_mildew']`. What is f1 score for this case?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1253022,
      "author_name": "rizdelhi",
      "author_url": "",
      "post_date": "03/26/2021 09:50:10",
      "content": "<p>I believe it's 'macro' as in many of competitions, it's default.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1253705,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "03/27/2021 01:09:55",
      "content": "<p>I agree, there is still and abnormally large gap between LB scores and CV scores. Fingers crossed that we hear back from the competition organizers!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1264107,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "04/05/2021 22:01:03",
      "content": "<p>It's closest to sklearn's <code>average='samples'</code>, though there might be some small differences in our implementations.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1264256,
          "author_name": "buinyi",
          "author_url": "",
          "post_date": "04/06/2021 02:50:45",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thank you for this clarification. Could you please specify the details of your implementation? For example, consider 3 images with labels <code>['scab', 'healthy', 'powdery_mildew complex']</code> and the predictions <code>['rust complex', 'healthy', 'powdery_mildew']</code>. What is f1 score for this case?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1264197,
      "author_name": "baodev",
      "author_url": "",
      "post_date": "04/06/2021 01:25:13",
      "content": "<p>So strange</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1251267": "fruitpathology, @sohier, we need your help.\n\nFrom the recent comments on the **[CV-LB Inconcistency](https://www.kaggle.com/c/plant-pathology-2021-fgvc8/discussion/227237)** topic it's still unclear whether the evaluation metric for the competition is a row-wise F1, i.e.:\n```\nscore = f1_score(y_true, y_pred, average='micro')\n```\nor a column-wise (or label-wise) F1, i.e.:\n```\nscore = f1_score(y_true, y_pred, average='macro')\n```\nInformation on the **Evaluation** page is not full to make sure whether it is former or latter (and has a broken link). As there's up to a 10% difference between them clarification of the metric is also necessary for understanding the reasons for inconsistency.",
    "1253022": "I believe it's 'macro' as in many of competitions, it's default.",
    "1253705": "I agree, there is still and abnormally large gap between LB scores and CV scores. Fingers crossed that we hear back from the competition organizers!",
    "1264107": "It's closest to sklearn's `average='samples'`, though there might be some small differences in our implementations.",
    "1264197": "So strange",
    "1264256": "Hi @sohier Thank you for this clarification. Could you please specify the details of your implementation? For example, consider 3 images with labels `['scab', 'healthy', 'powdery_mildew complex']` and the predictions `['rust complex', 'healthy', 'powdery_mildew']`. What is f1 score for this case?"
  },
  "source": "meta"
}