{
  "id": 309408,
  "title": "Macro F1 clarification",
  "url": "/competitions/birdclef-2022/discussion/309408",
  "author_name": "",
  "post_date": "2022-02-23T09:46:40.492653600Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>According to <code>sklearn</code> metric description:<br>\n<code>Calculate metrics for each label, and find their unweighted mean. This does not take label imbalance into account.</code></p>\n<p>So am I right that in our case labels are scored birds ?</p>\n<p>P.S.<br>\nIt is pretty obvious question but I want to ask it in order to make things absolutely clear</p>",
  "messages": [
    {
      "id": "1702068",
      "postDate": "02/23/2022 09:46:40",
      "content": "<p>According to <code>sklearn</code> metric description:<br>\n<code>Calculate metrics for each label, and find their unweighted mean. This does not take label imbalance into account.</code></p>\n<p>So am I right that in our case labels are scored birds ?</p>\n<p>P.S.<br>\nIt is pretty obvious question but I want to ask it in order to make things absolutely clear</p>",
      "rawMarkdown": "According to `sklearn` metric description:\n`Calculate metrics for each label, and find their unweighted mean. This does not take label imbalance into account.`\n\nSo am I right that in our case labels are scored birds ?\n\nP.S.\nIt is pretty obvious question but I want to ask it in order to make things absolutely clear",
      "votes": null
    },
    {
      "id": "1702123",
      "postDate": "02/23/2022 10:56:45",
      "content": "<blockquote>\n  <p>So am I right that in our case labels are scored birds ?</p>\n</blockquote>\n<p>That's my understanding, for what it is worth.</p>",
      "rawMarkdown": "> So am I right that in our case labels are scored birds ?\n\nThat's my understanding, for what it is worth.",
      "votes": null
    },
    {
      "id": "1702370",
      "postDate": "02/23/2022 14:55:54",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<p>Am I write if I use something like this </p>\n<pre><code>import numpy as np\nfrom typing import List\n\ndef f1_score(true_pos, false_neg, false_pos):\n    true_pos, false_neg, false_pos = float(true_pos), float(false_neg), float(false_pos)\n    if true_pos == 0 and (false_neg + false_pos) == 0:\n        return 1.0\n    else:\n        return 2 * true_pos / (2 * true_pos + false_neg + false_pos)\n\ndef macro_f1_similarity(\n    y_true_all_episodes: List[List[int]], y_pred_all_episodes: List[List[int]], unique_bird_ids: List[int]\n) -&gt; float:\n    assert len(y_true_all_episodes) == len(y_pred_all_episodes), \"different number of episodes for true and pred\"\n    n_episodes = len(y_true_all_episodes)\n\n    stats_dict = {k:{\"tp\":0, \"fn\":0, \"fp\":0} for k in unique_bird_ids}\n\n    for episode_idx in range(n_episodes):\n        for true_elem in y_true_all_episodes[episode_idx]:\n            if true_elem in unique_bird_ids:\n                if true_elem in y_pred_all_episodes[episode_idx]:\n                    stats_dict[true_elem][\"tp\"] += 1\n                else:\n                    stats_dict[true_elem][\"fn\"] += 1\n\n        for pred_el in y_pred_all_episodes[episode_idx]:\n            if pred_el in unique_bird_ids:\n                if pred_el not in y_true_all_episodes[episode_idx]:\n                    stats_dict[pred_el][\"fp\"] += 1\n\n\n    f1_similarity = np.mean([f1_score(\n        true_pos=item[\"tp\"], \n        false_neg=item[\"fn\"], \n        false_pos=item[\"fp\"]\n    ) for item in stats_dict.values()])\n\n\n    return f1_similarity\n</code></pre>\n<p>Here<br>\nepisode - audio chunk<br>\ny_true_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are presented in this episode <br>\ny_pred_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are predicted in this episode <br>\nunique_bird_ids - birds ids (or birds names) that will be included in metric computation. In our case it is <code>scored_birds.json</code></p>\n<p>Thanks in advance !!!</p>",
      "rawMarkdown": "cpmpml @stefankahl \n\nAm I write if I use something like this \n\n```\nimport numpy as np\nfrom typing import List\n\ndef f1_score(true_pos, false_neg, false_pos):\n    true_pos, false_neg, false_pos = float(true_pos), float(false_neg), float(false_pos)\n    if true_pos == 0 and (false_neg + false_pos) == 0:\n        return 1.0\n    else:\n        return 2 * true_pos / (2 * true_pos + false_neg + false_pos)\n\ndef macro_f1_similarity(\n    y_true_all_episodes: List[List[int]], y_pred_all_episodes: List[List[int]], unique_bird_ids: List[int]\n) -> float:\n    assert len(y_true_all_episodes) == len(y_pred_all_episodes), \"different number of episodes for true and pred\"\n    n_episodes = len(y_true_all_episodes)\n    \n    stats_dict = {k:{\"tp\":0, \"fn\":0, \"fp\":0} for k in unique_bird_ids}\n\n    for episode_idx in range(n_episodes):\n        for true_elem in y_true_all_episodes[episode_idx]:\n            if true_elem in unique_bird_ids:\n                if true_elem in y_pred_all_episodes[episode_idx]:\n                    stats_dict[true_elem][\"tp\"] += 1\n                else:\n                    stats_dict[true_elem][\"fn\"] += 1\n\n        for pred_el in y_pred_all_episodes[episode_idx]:\n            if pred_el in unique_bird_ids:\n                if pred_el not in y_true_all_episodes[episode_idx]:\n                    stats_dict[pred_el][\"fp\"] += 1\n\n    \n    f1_similarity = np.mean([f1_score(\n        true_pos=item[\"tp\"], \n        false_neg=item[\"fn\"], \n        false_pos=item[\"fp\"]\n    ) for item in stats_dict.values()])\n\n   \n    return f1_similarity\n```\n\nHere\nepisode - audio chunk\ny_true_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are presented in this episode \ny_pred_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are predicted in this episode \nunique_bird_ids - birds ids (or birds names) that will be included in metric computation. In our case it is `scored_birds.json`\n\nThanks in advance !!!",
      "votes": null
    },
    {
      "id": "1730937",
      "postDate": "03/21/2022 19:40:31",
      "content": "<p>In that case, could you just do the following?</p>\n<pre><code>from sklearn.metrics import f1_score\n\nmetric = f1_score(y_true, y_pred, average='macro')\n</code></pre>\n<p>where y_true and y_pred are List[List[int]] for ground truth and predictions of samples for scored birds</p>",
      "rawMarkdown": "In that case, could you just do the following?\n\n```\nfrom sklearn.metrics import f1_score\n\nmetric = f1_score(y_true, y_pred, average='macro')\n```\n\nwhere y_true and y_pred are List[List[int]] for ground truth and predictions of samples for scored birds",
      "votes": null
    },
    {
      "id": "1763184",
      "postDate": "04/21/2022 10:24:27",
      "content": "<p><a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> Hello. Congrats on the 1st place on the leaderboard. I recently started this competition. As far as I know, the metric is not exactly macro f1. Knowing the metric is crucial for finding a CV correlated with the leaderboard. I am wondering if you let me know which metric you finally found suitable, macro f1 or something else?</p>",
      "rawMarkdown": "vladimirsydor Hello. Congrats on the 1st place on the leaderboard. I recently started this competition. As far as I know, the metric is not exactly macro f1. Knowing the metric is crucial for finding a CV correlated with the leaderboard. I am wondering if you let me know which metric you finally found suitable, macro f1 or something else?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1702123,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/23/2022 10:56:45",
      "content": "<blockquote>\n  <p>So am I right that in our case labels are scored birds ?</p>\n</blockquote>\n<p>That's my understanding, for what it is worth.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1702370,
      "author_name": "vladimirsydor",
      "author_url": "",
      "post_date": "02/23/2022 14:55:54",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> </p>\n<p>Am I write if I use something like this </p>\n<pre><code>import numpy as np\nfrom typing import List\n\ndef f1_score(true_pos, false_neg, false_pos):\n    true_pos, false_neg, false_pos = float(true_pos), float(false_neg), float(false_pos)\n    if true_pos == 0 and (false_neg + false_pos) == 0:\n        return 1.0\n    else:\n        return 2 * true_pos / (2 * true_pos + false_neg + false_pos)\n\ndef macro_f1_similarity(\n    y_true_all_episodes: List[List[int]], y_pred_all_episodes: List[List[int]], unique_bird_ids: List[int]\n) -&gt; float:\n    assert len(y_true_all_episodes) == len(y_pred_all_episodes), \"different number of episodes for true and pred\"\n    n_episodes = len(y_true_all_episodes)\n\n    stats_dict = {k:{\"tp\":0, \"fn\":0, \"fp\":0} for k in unique_bird_ids}\n\n    for episode_idx in range(n_episodes):\n        for true_elem in y_true_all_episodes[episode_idx]:\n            if true_elem in unique_bird_ids:\n                if true_elem in y_pred_all_episodes[episode_idx]:\n                    stats_dict[true_elem][\"tp\"] += 1\n                else:\n                    stats_dict[true_elem][\"fn\"] += 1\n\n        for pred_el in y_pred_all_episodes[episode_idx]:\n            if pred_el in unique_bird_ids:\n                if pred_el not in y_true_all_episodes[episode_idx]:\n                    stats_dict[pred_el][\"fp\"] += 1\n\n\n    f1_similarity = np.mean([f1_score(\n        true_pos=item[\"tp\"], \n        false_neg=item[\"fn\"], \n        false_pos=item[\"fp\"]\n    ) for item in stats_dict.values()])\n\n\n    return f1_similarity\n</code></pre>\n<p>Here<br>\nepisode - audio chunk<br>\ny_true_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are presented in this episode <br>\ny_pred_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are predicted in this episode <br>\nunique_bird_ids - birds ids (or birds names) that will be included in metric computation. In our case it is <code>scored_birds.json</code></p>\n<p>Thanks in advance !!!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1730937,
          "author_name": "yousof9",
          "author_url": "",
          "post_date": "03/21/2022 19:40:31",
          "content": "<p>In that case, could you just do the following?</p>\n<pre><code>from sklearn.metrics import f1_score\n\nmetric = f1_score(y_true, y_pred, average='macro')\n</code></pre>\n<p>where y_true and y_pred are List[List[int]] for ground truth and predictions of samples for scored birds</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1763184,
      "author_name": "tjamali",
      "author_url": "",
      "post_date": "04/21/2022 10:24:27",
      "content": "<p><a href=\"https://www.kaggle.com/vladimirsydor\" target=\"_blank\">@vladimirsydor</a> Hello. Congrats on the 1st place on the leaderboard. I recently started this competition. As far as I know, the metric is not exactly macro f1. Knowing the metric is crucial for finding a CV correlated with the leaderboard. I am wondering if you let me know which metric you finally found suitable, macro f1 or something else?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1702068": "According to `sklearn` metric description:\n`Calculate metrics for each label, and find their unweighted mean. This does not take label imbalance into account.`\n\nSo am I right that in our case labels are scored birds ?\n\nP.S.\nIt is pretty obvious question but I want to ask it in order to make things absolutely clear",
    "1702123": "> So am I right that in our case labels are scored birds ?\n\nThat's my understanding, for what it is worth.",
    "1702370": "cpmpml @stefankahl \n\nAm I write if I use something like this \n\n```\nimport numpy as np\nfrom typing import List\n\ndef f1_score(true_pos, false_neg, false_pos):\n    true_pos, false_neg, false_pos = float(true_pos), float(false_neg), float(false_pos)\n    if true_pos == 0 and (false_neg + false_pos) == 0:\n        return 1.0\n    else:\n        return 2 * true_pos / (2 * true_pos + false_neg + false_pos)\n\ndef macro_f1_similarity(\n    y_true_all_episodes: List[List[int]], y_pred_all_episodes: List[List[int]], unique_bird_ids: List[int]\n) -> float:\n    assert len(y_true_all_episodes) == len(y_pred_all_episodes), \"different number of episodes for true and pred\"\n    n_episodes = len(y_true_all_episodes)\n    \n    stats_dict = {k:{\"tp\":0, \"fn\":0, \"fp\":0} for k in unique_bird_ids}\n\n    for episode_idx in range(n_episodes):\n        for true_elem in y_true_all_episodes[episode_idx]:\n            if true_elem in unique_bird_ids:\n                if true_elem in y_pred_all_episodes[episode_idx]:\n                    stats_dict[true_elem][\"tp\"] += 1\n                else:\n                    stats_dict[true_elem][\"fn\"] += 1\n\n        for pred_el in y_pred_all_episodes[episode_idx]:\n            if pred_el in unique_bird_ids:\n                if pred_el not in y_true_all_episodes[episode_idx]:\n                    stats_dict[pred_el][\"fp\"] += 1\n\n    \n    f1_similarity = np.mean([f1_score(\n        true_pos=item[\"tp\"], \n        false_neg=item[\"fn\"], \n        false_pos=item[\"fp\"]\n    ) for item in stats_dict.values()])\n\n   \n    return f1_similarity\n```\n\nHere\nepisode - audio chunk\ny_true_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are presented in this episode \ny_pred_all_episodes - contains List of episodes, where each episode is introduced by a List of birds ids (or bird names), which are predicted in this episode \nunique_bird_ids - birds ids (or birds names) that will be included in metric computation. In our case it is `scored_birds.json`\n\nThanks in advance !!!",
    "1730937": "In that case, could you just do the following?\n\n```\nfrom sklearn.metrics import f1_score\n\nmetric = f1_score(y_true, y_pred, average='macro')\n```\n\nwhere y_true and y_pred are List[List[int]] for ground truth and predictions of samples for scored birds",
    "1763184": "vladimirsydor Hello. Congrats on the 1st place on the leaderboard. I recently started this competition. As far as I know, the metric is not exactly macro f1. Knowing the metric is crucial for finding a CV correlated with the leaderboard. I am wondering if you let me know which metric you finally found suitable, macro f1 or something else?"
  },
  "source": "meta"
}