{
  "id": 311493,
  "title": "Any competition metric implementation ?",
  "url": "/competitions/birdclef-2022/discussion/311493",
  "author_name": "Ulrich G.",
  "post_date": "2022-03-07T08:27:51.756000",
  "votes": 27,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi everyone. I hope you are doing well. I'd like to ask if somebody has already implemented the competition metric. Honestly i'm struggling understing what they mean by:</p>\n<ul>\n<li><em>After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight</em></li>\n</ul>\n<p>To what i undestood, species will have the same global weights, but i'm confused by the weights of true negatives and true positives for species.</p>\n<p>Please <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/amandanavine\" target=\"_blank\">@amandanavine</a> <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> can you give more details ? By the way, any other explanation will be appreciated. </p>\n<p>I provide the OOF predictions of a simple CNN model <a href=\"https://www.kaggle.com/ulrich07/birdclef-oof\" target=\"_blank\">here</a> with the following variables:</p>\n<ul>\n<li>file_id : concatenation of primary label and audio name</li>\n<li>end_time</li>\n<li>prob : predicted probabilty of the targeted bird (after softmax)</li>\n<li>bird : the bird we want to test</li>\n<li>label : integer label (0:most frequent bird --&gt; 151: less frequent bird)</li>\n<li>bird_true : primary label</li>\n<li>target : prediction</li>\n<li>target_true : true target (True if bird = primary label)</li>\n</ul>\n<p>You can use these OOF predictions to test your metric evaluation fonction.</p>",
  "messages": [
    {
      "id": 1714675,
      "postDate": "2022-03-07T08:27:51.757Z",
      "content": "<p>Hi everyone. I hope you are doing well. I'd like to ask if somebody has already implemented the competition metric. Honestly i'm struggling understing what they mean by:</p>\n<ul>\n<li><em>After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight</em></li>\n</ul>\n<p>To what i undestood, species will have the same global weights, but i'm confused by the weights of true negatives and true positives for species.</p>\n<p>Please <a href=\"https://www.kaggle.com/stefankahl\" target=\"_blank\">@stefankahl</a> <a href=\"https://www.kaggle.com/amandanavine\" target=\"_blank\">@amandanavine</a> <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a> <a href=\"https://www.kaggle.com/holgerklinck\" target=\"_blank\">@holgerklinck</a> can you give more details ? By the way, any other explanation will be appreciated. </p>\n<p>I provide the OOF predictions of a simple CNN model <a href=\"https://www.kaggle.com/ulrich07/birdclef-oof\" target=\"_blank\">here</a> with the following variables:</p>\n<ul>\n<li>file_id : concatenation of primary label and audio name</li>\n<li>end_time</li>\n<li>prob : predicted probabilty of the targeted bird (after softmax)</li>\n<li>bird : the bird we want to test</li>\n<li>label : integer label (0:most frequent bird --&gt; 151: less frequent bird)</li>\n<li>bird_true : primary label</li>\n<li>target : prediction</li>\n<li>target_true : true target (True if bird = primary label)</li>\n</ul>\n<p>You can use these OOF predictions to test your metric evaluation fonction.</p>",
      "rawMarkdown": "Hi everyone. I hope you are doing well. I'd like to ask if somebody has already implemented the competition metric. Honestly i'm struggling understing what they mean by:\n* *After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight*\n\nTo what i undestood, species will have the same global weights, but i'm confused by the weights of true negatives and true positives for species.\n\nPlease @stefankahl @amandanavine @tomdenton @holgerklinck can you give more details ? By the way, any other explanation will be appreciated. \n\nI provide the OOF predictions of a simple CNN model [here](https://www.kaggle.com/ulrich07/birdclef-oof) with the following variables:\n* file_id : concatenation of primary label and audio name\n* end_time\n* prob : predicted probabilty of the targeted bird (after softmax)\n* bird : the bird we want to test\n* label : integer label (0:most frequent bird --> 151: less frequent bird)\n* bird_true : primary label\n* target : prediction\n* target_true : true target (True if bird = primary label)\n\nYou can use these OOF predictions to test your metric evaluation fonction.",
      "votes": 27
    },
    {
      "id": 1716290,
      "postDate": "2022-03-08T21:09:26.240Z",
      "content": "<p>Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)</p>",
      "rawMarkdown": "Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)",
      "votes": 9,
      "replies": [
        {
          "id": 1716310,
          "postDate": "2022-03-08T22:01:53.320Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>  i will release an attempt soon. So you will be able to confirm if my understanding is correct or wrong.</p>",
          "rawMarkdown": "Thanks @tomdenton  i will release an attempt soon. So you will be able to confirm if my understanding is correct or wrong.",
          "votes": 5
        },
        {
          "id": 1719814,
          "postDate": "2022-03-12T06:28:33.313Z",
          "content": "<p>\"we chose a weighting which equalizes the impact of the positive/negative labels for each species\" - Can you elaborate on exactly how this works? </p>\n<p>I've observed the leaderboard seems to be punishing false negatives far more heavily than false positives and seems to care more about recall than precision, but I'm still having some difficulty getting my local cross validation results to be correlated with the leaderboard. Changes that cause my leaderboard score to improve frequently cause my macro f1 scores to decline, and vice versa. This seems related to the positive/negative label weightings you're using.</p>",
          "rawMarkdown": "\"we chose a weighting which equalizes the impact of the positive/negative labels for each species\" - Can you elaborate on exactly how this works? \n\nI've observed the leaderboard seems to be punishing false negatives far more heavily than false positives and seems to care more about recall than precision, but I'm still having some difficulty getting my local cross validation results to be correlated with the leaderboard. Changes that cause my leaderboard score to improve frequently cause my macro f1 scores to decline, and vice versa. This seems related to the positive/negative label weightings you're using.\n",
          "votes": 4
        },
        {
          "id": 1719961,
          "postDate": "2022-03-12T10:07:05.480Z",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> for your observations. But can you share your local evaluation function ?</p>",
          "rawMarkdown": "Thanks @jsday96 for your observations. But can you share your local evaluation function ?",
          "votes": 3
        },
        {
          "id": 1720469,
          "postDate": "2022-03-12T20:32:20.377Z",
          "content": "<p>Sure. None of the metrics it calculates align well with the LB, but it is possible to predict the LB scores pretty accurately by training a linear regression model on the f1 scores, precisions, and recalls it outputs. r2 = .991 for 9 test submissions with LB scores ranging from 0.56 - 0.72. </p>\n<pre><code>true_positive_count = 0\nfalse_positive_count = 0\ntrue_negative_count = 0\nfalse_negative_count = 0\n\nbird_names_to_stats = {\n    bird_name : {\n        'true_positive_count' : 0,\n        'false_positive_count' : 0,\n        'true_negative_count' : 0,\n        'false_negative_count' : 0,\n    }\n    for bird_name in scored_birds_names\n}\n\nfor audio_file_path, ground_truth_class_name in paths_to_classes.items():\n\n    # Audio preprocessing and model inference here.\n\n    for segment_index, segment_scores in enumerate(all_chunks_scores):  \n        for candidate_bird_name in scored_birds_names:\n\n            # Score to threshold comparison here.\n\n            actually_matches = (candidate_bird_name == ground_truth_class_name)\n            if bird_detected and actually_matches:\n                true_positive_count += 1\n                bird_names_to_stats[candidate_bird_name]['true_positive_count'] += 1\n            elif bird_detected and not actually_matches:\n                false_positive_count += 1\n                bird_names_to_stats[candidate_bird_name]['false_positive_count'] += 1\n            elif not bird_detected and actually_matches:\n                false_negative_count += 1\n                bird_names_to_stats[candidate_bird_name]['false_negative_count'] += 1\n            else:\n                true_negative_count += 1\n                bird_names_to_stats[candidate_bird_name]['true_negative_count'] += 1\n\nall_f1s = [\n    stats['true_positive_count'] / (stats['true_positive_count'] + 0.5*(stats['false_positive_count'] + stats['false_negative_count']))\n    for stats in bird_names_to_stats.values()\n    if stats['true_positive_count'] + stats['false_positive_count'] + stats['false_negative_count'] &gt; 0\n]\nmacro_f1 = np.mean(all_f1s)\n\nall_precisions = [\n    stats['true_positive_count'] / (stats['true_positive_count'] + stats['false_positive_count'])\n    for stats in bird_names_to_stats.values()\n    if stats['true_positive_count'] + stats['false_positive_count'] &gt; 0\n]\nmacro_precision = np.mean(all_precisions)\n\nall_recalls = [\n    stats['true_positive_count'] / (stats['true_positive_count'] + stats['false_negative_count'])\n    for stats in bird_names_to_stats.values()\n    if stats['true_positive_count'] + stats['false_negative_count'] &gt; 0\n]\nmacro_recall = np.mean(all_recalls)\n\nprecision = true_positive_count / (true_positive_count + false_positive_count)\nrecall = true_positive_count / (true_positive_count + false_negative_count)\nf1 = 2 * (precision * recall) / (precision + recall)\n</code></pre>",
          "rawMarkdown": "Sure. None of the metrics it calculates align well with the LB, but it is possible to predict the LB scores pretty accurately by training a linear regression model on the f1 scores, precisions, and recalls it outputs. r2 = .991 for 9 test submissions with LB scores ranging from 0.56 - 0.72. \n\n```\ntrue_positive_count = 0\nfalse_positive_count = 0\ntrue_negative_count = 0\nfalse_negative_count = 0\n\nbird_names_to_stats = {\n\tbird_name : {\n\t\t'true_positive_count' : 0,\n\t\t'false_positive_count' : 0,\n\t\t'true_negative_count' : 0,\n\t\t'false_negative_count' : 0,\n\t}\n\tfor bird_name in scored_birds_names\n}\n\nfor audio_file_path, ground_truth_class_name in paths_to_classes.items():\n\n\t# Audio preprocessing and model inference here.\n\t\n\tfor segment_index, segment_scores in enumerate(all_chunks_scores):\t\n\t\tfor candidate_bird_name in scored_birds_names:\n\n\t\t\t# Score to threshold comparison here.\n\t\t\t\n\t\t\tactually_matches = (candidate_bird_name == ground_truth_class_name)\n\t\t\tif bird_detected and actually_matches:\n\t\t\t\ttrue_positive_count += 1\n\t\t\t\tbird_names_to_stats[candidate_bird_name]['true_positive_count'] += 1\n\t\t\telif bird_detected and not actually_matches:\n\t\t\t\tfalse_positive_count += 1\n\t\t\t\tbird_names_to_stats[candidate_bird_name]['false_positive_count'] += 1\n\t\t\telif not bird_detected and actually_matches:\n\t\t\t\tfalse_negative_count += 1\n\t\t\t\tbird_names_to_stats[candidate_bird_name]['false_negative_count'] += 1\n\t\t\telse:\n\t\t\t\ttrue_negative_count += 1\n\t\t\t\tbird_names_to_stats[candidate_bird_name]['true_negative_count'] += 1\n\nall_f1s = [\n\tstats['true_positive_count'] / (stats['true_positive_count'] + 0.5*(stats['false_positive_count'] + stats['false_negative_count']))\n\tfor stats in bird_names_to_stats.values()\n\tif stats['true_positive_count'] + stats['false_positive_count'] + stats['false_negative_count'] > 0\n]\nmacro_f1 = np.mean(all_f1s)\n\nall_precisions = [\n\tstats['true_positive_count'] / (stats['true_positive_count'] + stats['false_positive_count'])\n\tfor stats in bird_names_to_stats.values()\n\tif stats['true_positive_count'] + stats['false_positive_count'] > 0\n]\nmacro_precision = np.mean(all_precisions)\n\nall_recalls = [\n\tstats['true_positive_count'] / (stats['true_positive_count'] + stats['false_negative_count'])\n\tfor stats in bird_names_to_stats.values()\n\tif stats['true_positive_count'] + stats['false_negative_count'] > 0\n]\nmacro_recall = np.mean(all_recalls)\n\nprecision = true_positive_count / (true_positive_count + false_positive_count)\nrecall = true_positive_count / (true_positive_count + false_negative_count)\nf1 = 2 * (precision * recall) / (precision + recall)\n```",
          "votes": 10
        }
      ]
    },
    {
      "id": 1716134,
      "postDate": "2022-03-08T17:10:29.857Z",
      "content": "<p>Evaluation page says that the metric is most similar to macro-f1, but since the submission predicting all False scored 0.48x, it is not at all similar to macro-f1 as far as model confidence is low.<br>\nAs you point out, the statement \"weights of true negatives and true positives for species.\" is very unclear. I have no evidence, but at this point I guess that the metric is something like jaccard-score.</p>\n<pre><code>import numpy as np\n\nfrom sklearn.metrics import jaccard_score\n\n\ndef macro_jaccard_index(pred, target, threshold=0.5):\n    pred = pred &gt;= threshold\n    scores = []\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=1) for i in range(pred.shape[1])]\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=0) for i in range(pred.shape[1])]\n    return np.mean(scores)\n</code></pre>\n<p>I have currently no submission, so it is very likely to contain mistakes.</p>",
      "rawMarkdown": "Evaluation page says that the metric is most similar to macro-f1, but since the submission predicting all False scored 0.48x, it is not at all similar to macro-f1 as far as model confidence is low.\nAs you point out, the statement \"weights of true negatives and true positives for species.\" is very unclear. I have no evidence, but at this point I guess that the metric is something like jaccard-score.\n\n```\nimport numpy as np\n\nfrom sklearn.metrics import jaccard_score\n\n\ndef macro_jaccard_index(pred, target, threshold=0.5):\n    pred = pred >= threshold\n    scores = []\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=1) for i in range(pred.shape[1])]\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=0) for i in range(pred.shape[1])]\n    return np.mean(scores)\n```\n\nI have currently no submission, so it is very likely to contain mistakes.",
      "votes": 8,
      "replies": [
        {
          "id": 1716284,
          "postDate": "2022-03-08T20:58:46.683Z",
          "content": "<p>Thks <a href=\"https://www.kaggle.com/naka2ka\" target=\"_blank\">@naka2ka</a> . i will try it to see if my CV get closer to my LB. I tried something that i will make public in kernel shortly.</p>",
          "rawMarkdown": "Thks @naka2ka . i will try it to see if my CV get closer to my LB. I tried something that i will make public in kernel shortly.",
          "votes": 5
        }
      ]
    },
    {
      "id": 1734953,
      "postDate": "2022-03-25T18:26:49.563Z",
      "content": "<p>I tried to write out a whole post explaining things <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">here</a>.</p>\n<p>I think I have it right but who knows.</p>\n<p><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a></p>",
      "rawMarkdown": "I tried to write out a whole post explaining things [here](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999).\n\nI think I have it right but who knows.\n\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/314999",
      "votes": 2
    },
    {
      "id": 1734929,
      "postDate": "2022-03-25T18:04:21.703Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1716290,
      "author_name": "Tom Denton",
      "author_url": "",
      "post_date": "2022-03-08T21:09:26.240000",
      "content": "<p>Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)</p>",
      "votes": 9,
      "replies": [
        {
          "id": 1716310,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2022-03-08T22:01:53.320000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/tomdenton\" target=\"_blank\">@tomdenton</a>  i will release an attempt soon. So you will be able to confirm if my understanding is correct or wrong.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 1719814,
          "author_name": "James Day",
          "author_url": "",
          "post_date": "2022-03-12T06:28:33.313000",
          "content": "<p>\"we chose a weighting which equalizes the impact of the positive/negative labels for each species\" - Can you elaborate on exactly how this works? </p>\n<p>I've observed the leaderboard seems to be punishing false negatives far more heavily than false positives and seems to care more about recall than precision, but I'm still having some difficulty getting my local cross validation results to be correlated with the leaderboard. Changes that cause my leaderboard score to improve frequently cause my macro f1 scores to decline, and vice versa. This seems related to the positive/negative label weightings you're using.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 1719961,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2022-03-12T10:07:05.480000",
          "content": "<p>Thanks <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> for your observations. But can you share your local evaluation function ?</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1720469,
          "author_name": "James Day",
          "author_url": "",
          "post_date": "2022-03-12T20:32:20.377000",
          "content": "<p>Sure. None of the metrics it calculates align well with the LB, but it is possible to predict the LB scores pretty accurately by training a linear regression model on the f1 scores, precisions, and recalls it outputs. r2 = .991 for 9 test submissions with LB scores ranging from 0.56 - 0.72. </p>\n<pre><code>true_positive_count = 0\nfalse_positive_count = 0\ntrue_negative_count = 0\nfalse_negative_count = 0\n\nbird_names_to_stats = {\n    bird_name : {\n        'true_positive_count' : 0,\n        'false_positive_count' : 0,\n        'true_negative_count' : 0,\n        'false_negative_count' : 0,\n    }\n    for bird_name in scored_birds_names\n}\n\nfor audio_file_path, ground_truth_class_name in paths_to_classes.items():\n\n    # Audio preprocessing and model inference here.\n\n    for segment_index, segment_scores in enumerate(all_chunks_scores):  \n        for candidate_bird_name in scored_birds_names:\n\n            # Score to threshold comparison here.\n\n            actually_matches = (candidate_bird_name == ground_truth_class_name)\n            if bird_detected and actually_matches:\n                true_positive_count += 1\n                bird_names_to_stats[candidate_bird_name]['true_positive_count'] += 1\n            elif bird_detected and not actually_matches:\n                false_positive_count += 1\n                bird_names_to_stats[candidate_bird_name]['false_positive_count'] += 1\n            elif not bird_detected and actually_matches:\n                false_negative_count += 1\n                bird_names_to_stats[candidate_bird_name]['false_negative_count'] += 1\n            else:\n                true_negative_count += 1\n                bird_names_to_stats[candidate_bird_name]['true_negative_count'] += 1\n\nall_f1s = [\n    stats['true_positive_count'] / (stats['true_positive_count'] + 0.5*(stats['false_positive_count'] + stats['false_negative_count']))\n    for stats in bird_names_to_stats.values()\n    if stats['true_positive_count'] + stats['false_positive_count'] + stats['false_negative_count'] &gt; 0\n]\nmacro_f1 = np.mean(all_f1s)\n\nall_precisions = [\n    stats['true_positive_count'] / (stats['true_positive_count'] + stats['false_positive_count'])\n    for stats in bird_names_to_stats.values()\n    if stats['true_positive_count'] + stats['false_positive_count'] &gt; 0\n]\nmacro_precision = np.mean(all_precisions)\n\nall_recalls = [\n    stats['true_positive_count'] / (stats['true_positive_count'] + stats['false_negative_count'])\n    for stats in bird_names_to_stats.values()\n    if stats['true_positive_count'] + stats['false_negative_count'] &gt; 0\n]\nmacro_recall = np.mean(all_recalls)\n\nprecision = true_positive_count / (true_positive_count + false_positive_count)\nrecall = true_positive_count / (true_positive_count + false_negative_count)\nf1 = 2 * (precision * recall) / (precision + recall)\n</code></pre>",
          "votes": 10,
          "replies": []
        }
      ]
    },
    {
      "id": 1716134,
      "author_name": "ynktk",
      "author_url": "",
      "post_date": "2022-03-08T17:10:29.857000",
      "content": "<p>Evaluation page says that the metric is most similar to macro-f1, but since the submission predicting all False scored 0.48x, it is not at all similar to macro-f1 as far as model confidence is low.<br>\nAs you point out, the statement \"weights of true negatives and true positives for species.\" is very unclear. I have no evidence, but at this point I guess that the metric is something like jaccard-score.</p>\n<pre><code>import numpy as np\n\nfrom sklearn.metrics import jaccard_score\n\n\ndef macro_jaccard_index(pred, target, threshold=0.5):\n    pred = pred &gt;= threshold\n    scores = []\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=1) for i in range(pred.shape[1])]\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=0) for i in range(pred.shape[1])]\n    return np.mean(scores)\n</code></pre>\n<p>I have currently no submission, so it is very likely to contain mistakes.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 1716284,
          "author_name": "Ulrich G.",
          "author_url": "",
          "post_date": "2022-03-08T20:58:46.683000",
          "content": "<p>Thks <a href=\"https://www.kaggle.com/naka2ka\" target=\"_blank\">@naka2ka</a> . i will try it to see if my CV get closer to my LB. I tried something that i will make public in kernel shortly.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 1734953,
      "author_name": "Darien Schettler",
      "author_url": "",
      "post_date": "2022-03-25T18:26:49.563000",
      "content": "<p>I tried to write out a whole post explaining things <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">here</a>.</p>\n<p>I think I have it right but who knows.</p>\n<p><a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/314999\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/314999</a></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1734929,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-03-25T18:04:21.703000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1714675": "Hi everyone. I hope you are doing well. I'd like to ask if somebody has already implemented the competition metric. Honestly i'm struggling understing what they mean by:\n* *After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight*\n\nTo what i undestood, species will have the same global weights, but i'm confused by the weights of true negatives and true positives for species.\n\nPlease @stefankahl @amandanavine @tomdenton @holgerklinck can you give more details ? By the way, any other explanation will be appreciated. \n\nI provide the OOF predictions of a simple CNN model [here](https://www.kaggle.com/ulrich07/birdclef-oof) with the following variables:\n* file_id : concatenation of primary label and audio name\n* end_time\n* prob : predicted probabilty of the targeted bird (after softmax)\n* bird : the bird we want to test\n* label : integer label (0:most frequent bird --> 151: less frequent bird)\n* bird_true : primary label\n* target : prediction\n* target_true : true target (True if bird = primary label)\n\nYou can use these OOF predictions to test your metric evaluation fonction.",
    "1716290": "Hi! There's typically much more 'negative' audio for a given species than positive audio. So we chose a weighting which equalizes the impact of the positive/negative labels for each species. (and also equalizing the total weight for each species.)",
    "1716134": "Evaluation page says that the metric is most similar to macro-f1, but since the submission predicting all False scored 0.48x, it is not at all similar to macro-f1 as far as model confidence is low.\nAs you point out, the statement \"weights of true negatives and true positives for species.\" is very unclear. I have no evidence, but at this point I guess that the metric is something like jaccard-score.\n\n```\nimport numpy as np\n\nfrom sklearn.metrics import jaccard_score\n\n\ndef macro_jaccard_index(pred, target, threshold=0.5):\n    pred = pred >= threshold\n    scores = []\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=1) for i in range(pred.shape[1])]\n    scores += [jaccard_score(target[:, i], pred[:, i], pos_label=0) for i in range(pred.shape[1])]\n    return np.mean(scores)\n```\n\nI have currently no submission, so it is very likely to contain mistakes.",
    "1734953": "I tried to write out a whole post explaining things [here](https://www.kaggle.com/competitions/birdclef-2022/discussion/314999).\n\nI think I have it right but who knows.\n\nhttps://www.kaggle.com/competitions/birdclef-2022/discussion/314999",
    "1734929": ""
  }
}