{
  "id": 314999,
  "title": "If You Are Confused About the Evaluation Metric - Read This",
  "url": "/competitions/birdclef-2022/discussion/314999",
  "author_name": "Darien Schettler",
  "post_date": "2022-03-25T18:02:30.896000",
  "votes": 49,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi all, I have seen a fairly large amount of confusion regarding the evaluation metric. I recently implemented/understood it and wanted to share some things. </p>\n<hr>\n<p><strong>SIDE NOTE</strong></p>\n<ul>\n<li>The user <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">James Day</a> ( <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> ) provided one implementation/approach in this <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493\" target=\"_blank\">thread</a>.</li>\n<li>If I had seen his post it would have saved me quite a bit of time… hence why I wanted to make this easy to find.</li>\n<li>Please upvote his initial comment if possible as he deserves credit for his contribution.</li>\n</ul>\n<hr>\n<p><br></p>\n<p><strong>Step 1: Review and summarize the evaluation page details…</strong></p>\n<blockquote>\n  <p>\"Submissions are evaluated on a metric that is most similar to the macro F1 score. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a <strong>weighted classification accuracy</strong> with the weights set such that <strong>all of the species are assigned the same total weight</strong> and the <strong>true negatives and true positives for each species have the same weight</strong>. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.\"</p>\n</blockquote>\n<ul>\n<li>All scored species are weighted the same</li>\n<li>'Accuracy' is given as an equally weighted average between the <strong>True Positive</strong> score and <strong>True Negative</strong> score for a given species.</li>\n</ul>\n<p><br></p>\n<p><strong>Step 2: Identify if this obviously exists as a metric implementation</strong></p>\n<ul>\n<li><strong>True Positive</strong> Score is <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_score.html\" target=\"_blank\"><strong>Precision</strong></a> and is defined within the SKLearn library for easy use.<ul>\n<li><strong><code>TP Score = (TP)/(TP+FP)</code></strong></li></ul></li>\n<li><strong>True Negative</strong> Score is not defined within the SKLearn library. The closest definition for this term I could find was <a href=\"https://www.wikiwand.com/en/Positive_and_negative_predictive_values\" target=\"_blank\">here</a> <ul>\n<li><strong><code>TN Score = (TN)/(TN+FN)</code></strong></li></ul></li>\n</ul>\n<p><br></p>\n<p><strong>Step 3: Make Sure This Makes Sense…</strong></p>\n<ul>\n<li>We know that an all <strong><code>False</code></strong> submission yields a score of <strong>0.48</strong></li>\n<li>We know that there are 21 scored bird species</li>\n<li>We know there are 330,000 seconds of test audio broken into 5-second chunks --&gt; <strong>66,000 examples</strong> </li>\n<li>We aren't certain that there is an equal distribution of birds within the scored species</li>\n<li>For the sake of experiment we assume each bird is found in only 1% of the audio clips <ul>\n<li>--&gt; 660 examples should have a label of <strong><code>True</code></strong></li>\n<li>--&gt; 65,330 should have a label of <strong><code>False</code></strong></li></ul></li>\n</ul>\n<p>Let's now calculate the table that would be representative of an <strong>All <code>False</code></strong> submission given the above info.</p>\n<table>\n<thead>\n<tr>\n<th><strong>SPECIES</strong></th>\n<th><strong>TP COUNT</strong></th>\n<th><strong>TN COUNT</strong></th>\n<th><strong>FP COUNT</strong></th>\n<th><strong>FN COUNT</strong></th>\n<th><strong>TP SCORE (TP/TP+FP)</strong></th>\n<th><strong>TN SCORE (TN/TN+FN)</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>bird species 1</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 2</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 3</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 4</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 5</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 6</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 7</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 8</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 9</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 10</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 11</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 12</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 13</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 14</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 15</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 16</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 17</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 18</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 19</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 20</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 21</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>If we calculate the average across all bird species we get the following:<ul>\n<li><strong>Average TP Score of 0.0000</strong></li>\n<li><strong>Average FP Score of 0.9900</strong></li></ul></li>\n<li>If we then calculate the average BETWEEN TP and FP…<ul>\n<li>(0.00+0.99)/2 = <strong>0.495</strong></li></ul></li>\n</ul>\n<p><strong><em>So we can see with our little experiment that we get near the expected score. If we tweak the number for which each class occurs within the test set we can get the 0.48 number… ps the approximate percentage of TP would be --&gt; 4% &lt;--</em></strong></p>\n<p><br></p>\n<p><strong>Step 4: Develop an Implementation</strong></p>\n<p>Since most metric functions expect <strong><code>y_true</code></strong> and <strong><code>y_pred</code></strong> we setup our function similarly.</p>\n<ul>\n<li>Small note: We leverage which yields a list of confusion matrices each containing a 2x2 array where <ul>\n<li>the top left value is the True Negative Count, </li>\n<li>the top right value is the False Positive Count, </li>\n<li>the bottom left value is the False Negative Count, </li>\n<li>the bottom right value is the True Positive Count, </li></ul></li>\n</ul>\n<p>i.e.</p>\n<pre><code>|TN, FP|\n|FN, TP|\n</code></pre>\n<p><br></p>\n<hr>\n<p><strong>IMPLEMENTATION</strong></p>\n<hr>\n<pre><code>import numpy as np\nimport sklearn.metrics\n\ndef comp_metric(y_true, y_pred, epsilon=1e-9):\n    \"\"\" Function to calculate competition metric in an sklearn like fashion\n\n    Args:\n        y_true{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Ground truth (correct) target values.\n        y_pred{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Estimated targets as returned by a classifier.\n    Returns:\n        The single calculated score representative of this competitions evaluation\n    \"\"\"\n\n    # Get representative confusion matrices for each label\n    mlbl_cms = sklearn.metrics.multilabel_confusion_matrix(y_true, y_pred)\n\n    # Get two scores (TP and TN SCORES)\n    tp_scores = np.array([\n        mlbl_cm[1, 1]/(epsilon+mlbl_cm[:, 1].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n    tn_scores = np.array([\n        mlbl_cm[0, 0]/(epsilon+mlbl_cm[:, 0].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n\n    # Get average\n    tp_mean = tp_scores.mean()\n    tn_mean = tn_scores.mean()\n\n    return round((tp_mean+tn_mean)/2, 8)\n\n# Define some details of the data\nn_ex = 66000\ncls_perc = 0.04\nn_classes = 21\n\n# Generate a random (but representative) ground truth multilabel array\ngt_arr = np.zeros((n_ex, n_classes), np.uint8)\nfor i in range(n_classes):\n    random_idxs = np.random.choice(np.arange(n_ex), int(n_ex*cls_perc), replace=False, )\n    gt_arr[random_idxs, i] = 1\n\n# Define our all False prediction array\npred_arr = np.zeros_like(gt_arr)\n\n# Get metric\nprint(\"\\n... COMPETITION METRIC SCORE...\")\nprint(\"\\t--&gt;\", comp_metric(gt_arr, pred_arr))\n</code></pre>\n<p>Here's a link to a <a href=\"https://colab.research.google.com/drive/15p9ye1mZM1GFSbcGOAagVqbkVTTwg5P1?usp=sharing\" target=\"_blank\"><strong>colab</strong></a> where you can play with this implementation.</p>\n<hr>\n<p>Hope this helps and I'm not wrong! I think it makes sense, but feel free to correct me if I missed anything or if there are better ways to handle this.</p>",
  "messages": [
    {
      "id": 1734928,
      "postDate": "2022-03-25T18:02:30.897Z",
      "content": "<p>Hi all, I have seen a fairly large amount of confusion regarding the evaluation metric. I recently implemented/understood it and wanted to share some things. </p>\n<hr>\n<p><strong>SIDE NOTE</strong></p>\n<ul>\n<li>The user <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">James Day</a> ( <a href=\"https://www.kaggle.com/jsday96\" target=\"_blank\">@jsday96</a> ) provided one implementation/approach in this <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311493\" target=\"_blank\">thread</a>.</li>\n<li>If I had seen his post it would have saved me quite a bit of time… hence why I wanted to make this easy to find.</li>\n<li>Please upvote his initial comment if possible as he deserves credit for his contribution.</li>\n</ul>\n<hr>\n<p><br></p>\n<p><strong>Step 1: Review and summarize the evaluation page details…</strong></p>\n<blockquote>\n  <p>\"Submissions are evaluated on a metric that is most similar to the macro F1 score. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a <strong>weighted classification accuracy</strong> with the weights set such that <strong>all of the species are assigned the same total weight</strong> and the <strong>true negatives and true positives for each species have the same weight</strong>. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.\"</p>\n</blockquote>\n<ul>\n<li>All scored species are weighted the same</li>\n<li>'Accuracy' is given as an equally weighted average between the <strong>True Positive</strong> score and <strong>True Negative</strong> score for a given species.</li>\n</ul>\n<p><br></p>\n<p><strong>Step 2: Identify if this obviously exists as a metric implementation</strong></p>\n<ul>\n<li><strong>True Positive</strong> Score is <a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_score.html\" target=\"_blank\"><strong>Precision</strong></a> and is defined within the SKLearn library for easy use.<ul>\n<li><strong><code>TP Score = (TP)/(TP+FP)</code></strong></li></ul></li>\n<li><strong>True Negative</strong> Score is not defined within the SKLearn library. The closest definition for this term I could find was <a href=\"https://www.wikiwand.com/en/Positive_and_negative_predictive_values\" target=\"_blank\">here</a> <ul>\n<li><strong><code>TN Score = (TN)/(TN+FN)</code></strong></li></ul></li>\n</ul>\n<p><br></p>\n<p><strong>Step 3: Make Sure This Makes Sense…</strong></p>\n<ul>\n<li>We know that an all <strong><code>False</code></strong> submission yields a score of <strong>0.48</strong></li>\n<li>We know that there are 21 scored bird species</li>\n<li>We know there are 330,000 seconds of test audio broken into 5-second chunks --&gt; <strong>66,000 examples</strong> </li>\n<li>We aren't certain that there is an equal distribution of birds within the scored species</li>\n<li>For the sake of experiment we assume each bird is found in only 1% of the audio clips <ul>\n<li>--&gt; 660 examples should have a label of <strong><code>True</code></strong></li>\n<li>--&gt; 65,330 should have a label of <strong><code>False</code></strong></li></ul></li>\n</ul>\n<p>Let's now calculate the table that would be representative of an <strong>All <code>False</code></strong> submission given the above info.</p>\n<table>\n<thead>\n<tr>\n<th><strong>SPECIES</strong></th>\n<th><strong>TP COUNT</strong></th>\n<th><strong>TN COUNT</strong></th>\n<th><strong>FP COUNT</strong></th>\n<th><strong>FN COUNT</strong></th>\n<th><strong>TP SCORE (TP/TP+FP)</strong></th>\n<th><strong>TN SCORE (TN/TN+FN)</strong></th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td><strong>bird species 1</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 2</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 3</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 4</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 5</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 6</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 7</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 8</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 9</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 10</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 11</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 12</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 13</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 14</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 15</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 16</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 17</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 18</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 19</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 20</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n<tr>\n<td><strong>bird species 21</strong></td>\n<td>0</td>\n<td>65340</td>\n<td>0</td>\n<td>660</td>\n<td>0</td>\n<td>0.990000</td>\n</tr>\n</tbody>\n</table>\n<ul>\n<li>If we calculate the average across all bird species we get the following:<ul>\n<li><strong>Average TP Score of 0.0000</strong></li>\n<li><strong>Average FP Score of 0.9900</strong></li></ul></li>\n<li>If we then calculate the average BETWEEN TP and FP…<ul>\n<li>(0.00+0.99)/2 = <strong>0.495</strong></li></ul></li>\n</ul>\n<p><strong><em>So we can see with our little experiment that we get near the expected score. If we tweak the number for which each class occurs within the test set we can get the 0.48 number… ps the approximate percentage of TP would be --&gt; 4% &lt;--</em></strong></p>\n<p><br></p>\n<p><strong>Step 4: Develop an Implementation</strong></p>\n<p>Since most metric functions expect <strong><code>y_true</code></strong> and <strong><code>y_pred</code></strong> we setup our function similarly.</p>\n<ul>\n<li>Small note: We leverage which yields a list of confusion matrices each containing a 2x2 array where <ul>\n<li>the top left value is the True Negative Count, </li>\n<li>the top right value is the False Positive Count, </li>\n<li>the bottom left value is the False Negative Count, </li>\n<li>the bottom right value is the True Positive Count, </li></ul></li>\n</ul>\n<p>i.e.</p>\n<pre><code>|TN, FP|\n|FN, TP|\n</code></pre>\n<p><br></p>\n<hr>\n<p><strong>IMPLEMENTATION</strong></p>\n<hr>\n<pre><code>import numpy as np\nimport sklearn.metrics\n\ndef comp_metric(y_true, y_pred, epsilon=1e-9):\n    \"\"\" Function to calculate competition metric in an sklearn like fashion\n\n    Args:\n        y_true{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Ground truth (correct) target values.\n        y_pred{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Estimated targets as returned by a classifier.\n    Returns:\n        The single calculated score representative of this competitions evaluation\n    \"\"\"\n\n    # Get representative confusion matrices for each label\n    mlbl_cms = sklearn.metrics.multilabel_confusion_matrix(y_true, y_pred)\n\n    # Get two scores (TP and TN SCORES)\n    tp_scores = np.array([\n        mlbl_cm[1, 1]/(epsilon+mlbl_cm[:, 1].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n    tn_scores = np.array([\n        mlbl_cm[0, 0]/(epsilon+mlbl_cm[:, 0].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n\n    # Get average\n    tp_mean = tp_scores.mean()\n    tn_mean = tn_scores.mean()\n\n    return round((tp_mean+tn_mean)/2, 8)\n\n# Define some details of the data\nn_ex = 66000\ncls_perc = 0.04\nn_classes = 21\n\n# Generate a random (but representative) ground truth multilabel array\ngt_arr = np.zeros((n_ex, n_classes), np.uint8)\nfor i in range(n_classes):\n    random_idxs = np.random.choice(np.arange(n_ex), int(n_ex*cls_perc), replace=False, )\n    gt_arr[random_idxs, i] = 1\n\n# Define our all False prediction array\npred_arr = np.zeros_like(gt_arr)\n\n# Get metric\nprint(\"\\n... COMPETITION METRIC SCORE...\")\nprint(\"\\t--&gt;\", comp_metric(gt_arr, pred_arr))\n</code></pre>\n<p>Here's a link to a <a href=\"https://colab.research.google.com/drive/15p9ye1mZM1GFSbcGOAagVqbkVTTwg5P1?usp=sharing\" target=\"_blank\"><strong>colab</strong></a> where you can play with this implementation.</p>\n<hr>\n<p>Hope this helps and I'm not wrong! I think it makes sense, but feel free to correct me if I missed anything or if there are better ways to handle this.</p>",
      "rawMarkdown": "Hi all, I have seen a fairly large amount of confusion regarding the evaluation metric. I recently implemented/understood it and wanted to share some things. \n\n---\n\n**SIDE NOTE**\n* The user [James Day](https://www.kaggle.com/jsday96) ( @jsday96 ) provided one implementation/approach in this [thread](https://www.kaggle.com/competitions/birdclef-2022/discussion/311493).\n* If I had seen his post it would have saved me quite a bit of time... hence why I wanted to make this easy to find.\n* Please upvote his initial comment if possible as he deserves credit for his contribution.\n\n---\n\n<br>\n\n**Step 1: Review and summarize the evaluation page details...**\n\n> \"Submissions are evaluated on a metric that is most similar to the macro F1 score. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a **weighted classification accuracy** with the weights set such that **all of the species are assigned the same total weight** and the **true negatives and true positives for each species have the same weight**. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.\"\n\n* All scored species are weighted the same\n* 'Accuracy' is given as an equally weighted average between the **True Positive** score and **True Negative** score for a given species.\n\n<br>\n\n**Step 2: Identify if this obviously exists as a metric implementation**\n\n* **True Positive** Score is [**Precision**](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_score.html) and is defined within the SKLearn library for easy use.\n  * **`TP Score = (TP)/(TP+FP)`**\n* **True Negative** Score is not defined within the SKLearn library. The closest definition for this term I could find was [here](https://www.wikiwand.com/en/Positive_and_negative_predictive_values) \n  * **`TN Score = (TN)/(TN+FN)`**\n\n<br>\n\n**Step 3: Make Sure This Makes Sense...**\n\n* We know that an all **`False`** submission yields a score of **0.48**\n* We know that there are 21 scored bird species\n* We know there are 330,000 seconds of test audio broken into 5-second chunks --> **66,000 examples** \n* We aren't certain that there is an equal distribution of birds within the scored species\n* For the sake of experiment we assume each bird is found in only 1% of the audio clips \n  * --> 660 examples should have a label of **`True`**\n  * --> 65,330 should have a label of **`False`**\n\nLet's now calculate the table that would be representative of an **All `False`** submission given the above info.\n\n|     **SPECIES**     | **TP COUNT** | **TN COUNT** | **FP COUNT** | **FN COUNT** | **TP SCORE (TP/TP+FP)** | **TN SCORE (TN/TN+FN)** |\n|:-------------------:|:------------:|:------------:|:------------:|:------------:|:-----------------------:|:-----------------------:|\n|  **bird species 1** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 2** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 3** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 4** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 5** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 6** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 7** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 8** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 9** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 10** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 11** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 12** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 13** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 14** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 15** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 16** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 17** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 18** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 19** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 20** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 21** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n\n* If we calculate the average across all bird species we get the following:\n  * **Average TP Score of 0.0000**\n  * **Average FP Score of 0.9900**\n* If we then calculate the average BETWEEN TP and FP...\n  * (0.00+0.99)/2 = **0.495**\n\n***So we can see with our little experiment that we get near the expected score. If we tweak the number for which each class occurs within the test set we can get the 0.48 number... ps the approximate percentage of TP would be --> 4% <--***\n\n<br>\n\n**Step 4: Develop an Implementation**\n\nSince most metric functions expect **`y_true`** and **`y_pred`** we setup our function similarly.\n* Small note: We leverage which yields a list of confusion matrices each containing a 2x2 array where \n  * the top left value is the True Negative Count, \n  * the top right value is the False Positive Count, \n  * the bottom left value is the False Negative Count, \n  * the bottom right value is the True Positive Count, \n\ni.e.\n```\n|TN, FP|\n|FN, TP|\n```\n\n<br>\n\n---\n\n**IMPLEMENTATION**\n\n---\n\n```\nimport numpy as np\nimport sklearn.metrics\n\ndef comp_metric(y_true, y_pred, epsilon=1e-9):\n    \"\"\" Function to calculate competition metric in an sklearn like fashion\n\n    Args:\n        y_true{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Ground truth (correct) target values.\n        y_pred{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Estimated targets as returned by a classifier.\n    Returns:\n        The single calculated score representative of this competitions evaluation\n    \"\"\"\n    \n    # Get representative confusion matrices for each label\n    mlbl_cms = sklearn.metrics.multilabel_confusion_matrix(y_true, y_pred)\n\n    # Get two scores (TP and TN SCORES)\n    tp_scores = np.array([\n        mlbl_cm[1, 1]/(epsilon+mlbl_cm[:, 1].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n    tn_scores = np.array([\n        mlbl_cm[0, 0]/(epsilon+mlbl_cm[:, 0].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n\n    # Get average\n    tp_mean = tp_scores.mean()\n    tn_mean = tn_scores.mean()\n\n    return round((tp_mean+tn_mean)/2, 8)\n\n# Define some details of the data\nn_ex = 66000\ncls_perc = 0.04\nn_classes = 21\n\n# Generate a random (but representative) ground truth multilabel array\ngt_arr = np.zeros((n_ex, n_classes), np.uint8)\nfor i in range(n_classes):\n    random_idxs = np.random.choice(np.arange(n_ex), int(n_ex*cls_perc), replace=False, )\n    gt_arr[random_idxs, i] = 1\n\n# Define our all False prediction array\npred_arr = np.zeros_like(gt_arr)\n\n# Get metric\nprint(\"\\n... COMPETITION METRIC SCORE...\")\nprint(\"\\t-->\", comp_metric(gt_arr, pred_arr))\n```\n\nHere's a link to a [**colab**](https://colab.research.google.com/drive/15p9ye1mZM1GFSbcGOAagVqbkVTTwg5P1?usp=sharing) where you can play with this implementation.\n\n---\n\nHope this helps and I'm not wrong! I think it makes sense, but feel free to correct me if I missed anything or if there are better ways to handle this.",
      "votes": 46
    },
    {
      "id": 1735156,
      "postDate": "2022-03-25T23:27:40.023Z",
      "content": "<p>Suppose mean GTR = 4% and TP/N and TN/N are equally weighted, scores in assumptions are:</p>\n<ul>\n<li>all-true submission: 0.02 (=(0.04 + 0) / 2)</li>\n<li>all-false submission: 0.48 (=(0+ 0.96) / 2)</li>\n<li>random submission: 0.25 (=(0.02 + 0.48) / 2)</li>\n</ul>\n<p>However, the actual scores are:</p>\n<ul>\n<li>all-true submission: 0.51</li>\n<li>all-false submission: 0.48</li>\n<li>random submission (seed=123): 0.50</li>\n</ul>\n<p>I think TP/N and TN/N are not equally weighted (if supposed LB = w_p * TP/N + w_n * TN/N).</p>",
      "rawMarkdown": "Suppose mean GTR = 4% and TP/N and TN/N are equally weighted, scores in assumptions are:\n\n* all-true submission: 0.02 (=(0.04 + 0) / 2)\n* all-false submission: 0.48 (=(0+ 0.96) / 2)\n* random submission: 0.25 (=(0.02 + 0.48) / 2)\n\nHowever, the actual scores are:\n\n* all-true submission: 0.51\n* all-false submission: 0.48\n* random submission (seed=123): 0.50\n\nI think TP/N and TN/N are not equally weighted (if supposed LB = w_p * TP/N + w_n * TN/N).",
      "votes": 1,
      "replies": [
        {
          "id": 1735177,
          "postDate": "2022-03-26T00:21:56.170Z",
          "content": "<p>This is a modified version of <em>unequal</em> weight. The assumed scores are close to the actual scores (thought differ in |error| &lt; 0.02).</p>\n<hr>\n<p>Assume the score is calculated by the below equation:</p>\n<p>$$S = w_p * \\text{TPR} + w_n * \\text{TNR}$$</p>\n<p>and TNR and FNR are <em>evenly</em> counted,</p>\n<p>$$w_p = 0.5 * \\frac{N}{GT}, w_n = 0.5 * \\frac{N}{N - GT}$$</p>\n<p>Concisely, </p>\n<p>$$S = 0.5 * \\left(\\frac{TP}{GT} + \\frac{TN}{N - GT}\\right) \\tag{1} = 0.5 * (\\text{TPR} + \\text{TNR})$$</p>\n<p>Thus, the assumed scores are:</p>\n<p>$$S_{\\text{True}} = 0.5 * \\left(\\frac{GT}{GT} + \\frac{0}{N - GT}\\right) = 0.5$$<br>\n$$S_{\\text{False}} = 0.5 * \\left(\\frac{0}{GT} + \\frac{N - GT}{N - GT}\\right) = 0.5$$<br>\n$$S_{\\text{Random}} = 0.5 * \\left(\\frac{GT/2}{GT} + \\frac{(N - GT)/2}{N - GT}\\right) = 0.5$$<br>\n$$S_{\\text{Perfect}} = 0.5 * \\left(\\frac{GT}{GT} + \\frac{N - GT}{N - GT}\\right) = 1.0$$</p>\n<hr>\n<p>Note: the weights (w_p and w_n) might be variable w.r.t. target species (or it might be fixed).</p>",
          "rawMarkdown": "This is a modified version of *unequal* weight. The assumed scores are close to the actual scores (thought differ in |error| < 0.02).\n\n---\nAssume the score is calculated by the below equation:\n\n$$S = w_p * \\text{TPR} + w_n * \\text{TNR}$$\n\nand TNR and FNR are *evenly* counted,\n\n$$w_p = 0.5 * \\frac{N}{GT}, w_n = 0.5 * \\frac{N}{N - GT}$$\n\nConcisely, \n\n$$S = 0.5 * \\left(\\frac{TP}{GT} + \\frac{TN}{N - GT}\\right) \\tag{1} = 0.5 * (\\text{TPR} + \\text{TNR})$$\n\nThus, the assumed scores are:\n\n$$S_{\\text{True}} = 0.5 * \\left(\\frac{GT}{GT} + \\frac{0}{N - GT}\\right) = 0.5$$\n$$S_{\\text{False}} = 0.5 * \\left(\\frac{0}{GT} + \\frac{N - GT}{N - GT}\\right) = 0.5$$\n$$S_{\\text{Random}} = 0.5 * \\left(\\frac{GT/2}{GT} + \\frac{(N - GT)/2}{N - GT}\\right) = 0.5$$\n$$S_{\\text{Perfect}} = 0.5 * \\left(\\frac{GT}{GT} + \\frac{N - GT}{N - GT}\\right) = 1.0$$\n\n---\n\nNote: the weights (w_p and w_n) might be variable w.r.t. target species (or it might be fixed).",
          "votes": 2
        },
        {
          "id": 1735187,
          "postDate": "2022-03-26T00:53:23.977Z",
          "content": "<p>You have this line…</p>\n<blockquote>\n  <p>\"all-true submission: 0.02 (=(0.04 + 1) / 2)\"</p>\n</blockquote>\n<p>But shouldn't it be…</p>\n<pre><code>x = (0.04+1.00)/2\nx = (1.04/2)\nx ~= 0.52\n</code></pre>\n<p>Which is close enough to the public LB score of 0.51 that it would match?</p>\n<hr>\n<p>That being said, I think you just typo'd… your logic and math make sense.</p>",
          "rawMarkdown": "You have this line...\n\n> \"all-true submission: 0.02 (=(0.04 + 1) / 2)\"\n\nBut shouldn't it be...\n\n```\nx = (0.04+1.00)/2\nx = (1.04/2)\nx ~= 0.52\n```\n\nWhich is close enough to the public LB score of 0.51 that it would match?\n\n---\n\nThat being said, I think you just typo'd... your logic and math make sense."
        },
        {
          "id": 1735191,
          "postDate": "2022-03-26T01:11:32.563Z",
          "content": "<blockquote>\n  <p>all-true submission: 0.02 (=(0.04 + 1) / 2)</p>\n</blockquote>\n<p>Oops. It’s typo as you said.</p>\n<p>Correct: all-true submission: 0.02 (=(0.04 + <strong>0</strong>) / 2)</p>\n<p></p>\n<p>Note: edited original comment: TPR, TNR -&gt; TP/N, TN/N</p>",
          "rawMarkdown": "> all-true submission: 0.02 (=(0.04 + 1) / 2)\n\nOops. It’s typo as you said.\n\nCorrect: all-true submission: 0.02 (=(0.04 + **0**) / 2)\n\n~~Maybe some other minor miscalculation exists in the original assumption because I used TP/N and FP/N instead of TPR and TNR.~~\n\nNote: edited original comment: TPR, TNR -> TP/N, TN/N"
        },
        {
          "id": 1771128,
          "postDate": "2022-04-28T23:44:55.610Z",
          "content": "<p>FYI: The equalized version is generally called <em>balanced accuracy</em>.<br>\n<a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\" target=\"_blank\">https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score</a></p>",
          "rawMarkdown": "FYI: The equalized version is generally called *balanced accuracy*.\nhttps://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score",
          "votes": 1
        }
      ]
    },
    {
      "id": 1735143,
      "postDate": "2022-03-25T22:45:45.917Z",
      "content": "<p>Good observation. Thank you for sharing.</p>\n<p>I think below sentence is mistaken:</p>\n<blockquote>\n  <p>ps the approximate percentage of TN would be --&gt; 4% &lt;--</p>\n</blockquote>\n<p>-&gt; Correct: the approximate of <strong>FN</strong> (I.e. GT) would be 4% (= 1 - 0.48 * 2).</p>",
      "rawMarkdown": "Good observation. Thank you for sharing.\n\nI think below sentence is mistaken:\n\n> ps the approximate percentage of TN would be --> 4% <--\n\n-> Correct: the approximate of **FN** (I.e. GT) would be 4% (= 1 - 0.48 * 2).\n",
      "replies": [
        {
          "id": 1735147,
          "postDate": "2022-03-25T23:07:27.660Z",
          "content": "<p>Yes you’re right. I meant to write TP which would have the same rate as FN.</p>\n<p>Thanks for the catch. I’ll update it now.</p>",
          "rawMarkdown": "Yes you’re right. I meant to write TP which would have the same rate as FN.\n\nThanks for the catch. I’ll update it now."
        }
      ]
    },
    {
      "id": 1801676,
      "postDate": "2022-05-26T03:16:29.897Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1735156,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-03-25T23:27:40.023000",
      "content": "<p>Suppose mean GTR = 4% and TP/N and TN/N are equally weighted, scores in assumptions are:</p>\n<ul>\n<li>all-true submission: 0.02 (=(0.04 + 0) / 2)</li>\n<li>all-false submission: 0.48 (=(0+ 0.96) / 2)</li>\n<li>random submission: 0.25 (=(0.02 + 0.48) / 2)</li>\n</ul>\n<p>However, the actual scores are:</p>\n<ul>\n<li>all-true submission: 0.51</li>\n<li>all-false submission: 0.48</li>\n<li>random submission (seed=123): 0.50</li>\n</ul>\n<p>I think TP/N and TN/N are not equally weighted (if supposed LB = w_p * TP/N + w_n * TN/N).</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1735177,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-03-26T00:21:56.170000",
          "content": "<p>This is a modified version of <em>unequal</em> weight. The assumed scores are close to the actual scores (thought differ in |error| &lt; 0.02).</p>\n<hr>\n<p>Assume the score is calculated by the below equation:</p>\n<p>$$S = w_p * \\text{TPR} + w_n * \\text{TNR}$$</p>\n<p>and TNR and FNR are <em>evenly</em> counted,</p>\n<p>$$w_p = 0.5 * \\frac{N}{GT}, w_n = 0.5 * \\frac{N}{N - GT}$$</p>\n<p>Concisely, </p>\n<p>$$S = 0.5 * \\left(\\frac{TP}{GT} + \\frac{TN}{N - GT}\\right) \\tag{1} = 0.5 * (\\text{TPR} + \\text{TNR})$$</p>\n<p>Thus, the assumed scores are:</p>\n<p>$$S_{\\text{True}} = 0.5 * \\left(\\frac{GT}{GT} + \\frac{0}{N - GT}\\right) = 0.5$$<br>\n$$S_{\\text{False}} = 0.5 * \\left(\\frac{0}{GT} + \\frac{N - GT}{N - GT}\\right) = 0.5$$<br>\n$$S_{\\text{Random}} = 0.5 * \\left(\\frac{GT/2}{GT} + \\frac{(N - GT)/2}{N - GT}\\right) = 0.5$$<br>\n$$S_{\\text{Perfect}} = 0.5 * \\left(\\frac{GT}{GT} + \\frac{N - GT}{N - GT}\\right) = 1.0$$</p>\n<hr>\n<p>Note: the weights (w_p and w_n) might be variable w.r.t. target species (or it might be fixed).</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 1735187,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-03-26T00:53:23.977000",
          "content": "<p>You have this line…</p>\n<blockquote>\n  <p>\"all-true submission: 0.02 (=(0.04 + 1) / 2)\"</p>\n</blockquote>\n<p>But shouldn't it be…</p>\n<pre><code>x = (0.04+1.00)/2\nx = (1.04/2)\nx ~= 0.52\n</code></pre>\n<p>Which is close enough to the public LB score of 0.51 that it would match?</p>\n<hr>\n<p>That being said, I think you just typo'd… your logic and math make sense.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1735191,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-03-26T01:11:32.563000",
          "content": "<blockquote>\n  <p>all-true submission: 0.02 (=(0.04 + 1) / 2)</p>\n</blockquote>\n<p>Oops. It’s typo as you said.</p>\n<p>Correct: all-true submission: 0.02 (=(0.04 + <strong>0</strong>) / 2)</p>\n<p></p>\n<p>Note: edited original comment: TPR, TNR -&gt; TP/N, TN/N</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1771128,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-04-28T23:44:55.610000",
          "content": "<p>FYI: The equalized version is generally called <em>balanced accuracy</em>.<br>\n<a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score\" target=\"_blank\">https://scikit-learn.org/stable/modules/model_evaluation.html#balanced-accuracy-score</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1735143,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-03-25T22:45:45.917000",
      "content": "<p>Good observation. Thank you for sharing.</p>\n<p>I think below sentence is mistaken:</p>\n<blockquote>\n  <p>ps the approximate percentage of TN would be --&gt; 4% &lt;--</p>\n</blockquote>\n<p>-&gt; Correct: the approximate of <strong>FN</strong> (I.e. GT) would be 4% (= 1 - 0.48 * 2).</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1735147,
          "author_name": "Darien Schettler",
          "author_url": "",
          "post_date": "2022-03-25T23:07:27.660000",
          "content": "<p>Yes you’re right. I meant to write TP which would have the same rate as FN.</p>\n<p>Thanks for the catch. I’ll update it now.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1801676,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-05-26T03:16:29.897000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1734928": "Hi all, I have seen a fairly large amount of confusion regarding the evaluation metric. I recently implemented/understood it and wanted to share some things. \n\n---\n\n**SIDE NOTE**\n* The user [James Day](https://www.kaggle.com/jsday96) ( @jsday96 ) provided one implementation/approach in this [thread](https://www.kaggle.com/competitions/birdclef-2022/discussion/311493).\n* If I had seen his post it would have saved me quite a bit of time... hence why I wanted to make this easy to find.\n* Please upvote his initial comment if possible as he deserves credit for his contribution.\n\n---\n\n<br>\n\n**Step 1: Review and summarize the evaluation page details...**\n\n> \"Submissions are evaluated on a metric that is most similar to the macro F1 score. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a **weighted classification accuracy** with the weights set such that **all of the species are assigned the same total weight** and the **true negatives and true positives for each species have the same weight**. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.\"\n\n* All scored species are weighted the same\n* 'Accuracy' is given as an equally weighted average between the **True Positive** score and **True Negative** score for a given species.\n\n<br>\n\n**Step 2: Identify if this obviously exists as a metric implementation**\n\n* **True Positive** Score is [**Precision**](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.precision_score.html) and is defined within the SKLearn library for easy use.\n  * **`TP Score = (TP)/(TP+FP)`**\n* **True Negative** Score is not defined within the SKLearn library. The closest definition for this term I could find was [here](https://www.wikiwand.com/en/Positive_and_negative_predictive_values) \n  * **`TN Score = (TN)/(TN+FN)`**\n\n<br>\n\n**Step 3: Make Sure This Makes Sense...**\n\n* We know that an all **`False`** submission yields a score of **0.48**\n* We know that there are 21 scored bird species\n* We know there are 330,000 seconds of test audio broken into 5-second chunks --> **66,000 examples** \n* We aren't certain that there is an equal distribution of birds within the scored species\n* For the sake of experiment we assume each bird is found in only 1% of the audio clips \n  * --> 660 examples should have a label of **`True`**\n  * --> 65,330 should have a label of **`False`**\n\nLet's now calculate the table that would be representative of an **All `False`** submission given the above info.\n\n|     **SPECIES**     | **TP COUNT** | **TN COUNT** | **FP COUNT** | **FN COUNT** | **TP SCORE (TP/TP+FP)** | **TN SCORE (TN/TN+FN)** |\n|:-------------------:|:------------:|:------------:|:------------:|:------------:|:-----------------------:|:-----------------------:|\n|  **bird species 1** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 2** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 3** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 4** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 5** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 6** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 7** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 8** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n|  **bird species 9** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 10** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 11** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 12** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 13** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 14** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 15** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 16** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 17** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 18** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 19** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 20** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n| **bird species 21** |       0      |     65340    |       0      |      660     |            0            |         0.990000        |\n\n* If we calculate the average across all bird species we get the following:\n  * **Average TP Score of 0.0000**\n  * **Average FP Score of 0.9900**\n* If we then calculate the average BETWEEN TP and FP...\n  * (0.00+0.99)/2 = **0.495**\n\n***So we can see with our little experiment that we get near the expected score. If we tweak the number for which each class occurs within the test set we can get the 0.48 number... ps the approximate percentage of TP would be --> 4% <--***\n\n<br>\n\n**Step 4: Develop an Implementation**\n\nSince most metric functions expect **`y_true`** and **`y_pred`** we setup our function similarly.\n* Small note: We leverage which yields a list of confusion matrices each containing a 2x2 array where \n  * the top left value is the True Negative Count, \n  * the top right value is the False Positive Count, \n  * the bottom left value is the False Negative Count, \n  * the bottom right value is the True Positive Count, \n\ni.e.\n```\n|TN, FP|\n|FN, TP|\n```\n\n<br>\n\n---\n\n**IMPLEMENTATION**\n\n---\n\n```\nimport numpy as np\nimport sklearn.metrics\n\ndef comp_metric(y_true, y_pred, epsilon=1e-9):\n    \"\"\" Function to calculate competition metric in an sklearn like fashion\n\n    Args:\n        y_true{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Ground truth (correct) target values.\n        y_pred{array-like, sparse matrix} of shape (n_samples, n_outputs)\n            - Estimated targets as returned by a classifier.\n    Returns:\n        The single calculated score representative of this competitions evaluation\n    \"\"\"\n    \n    # Get representative confusion matrices for each label\n    mlbl_cms = sklearn.metrics.multilabel_confusion_matrix(y_true, y_pred)\n\n    # Get two scores (TP and TN SCORES)\n    tp_scores = np.array([\n        mlbl_cm[1, 1]/(epsilon+mlbl_cm[:, 1].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n    tn_scores = np.array([\n        mlbl_cm[0, 0]/(epsilon+mlbl_cm[:, 0].sum()) \\\n        for mlbl_cm in mlbl_cms\n        ])\n\n    # Get average\n    tp_mean = tp_scores.mean()\n    tn_mean = tn_scores.mean()\n\n    return round((tp_mean+tn_mean)/2, 8)\n\n# Define some details of the data\nn_ex = 66000\ncls_perc = 0.04\nn_classes = 21\n\n# Generate a random (but representative) ground truth multilabel array\ngt_arr = np.zeros((n_ex, n_classes), np.uint8)\nfor i in range(n_classes):\n    random_idxs = np.random.choice(np.arange(n_ex), int(n_ex*cls_perc), replace=False, )\n    gt_arr[random_idxs, i] = 1\n\n# Define our all False prediction array\npred_arr = np.zeros_like(gt_arr)\n\n# Get metric\nprint(\"\\n... COMPETITION METRIC SCORE...\")\nprint(\"\\t-->\", comp_metric(gt_arr, pred_arr))\n```\n\nHere's a link to a [**colab**](https://colab.research.google.com/drive/15p9ye1mZM1GFSbcGOAagVqbkVTTwg5P1?usp=sharing) where you can play with this implementation.\n\n---\n\nHope this helps and I'm not wrong! I think it makes sense, but feel free to correct me if I missed anything or if there are better ways to handle this.",
    "1735156": "Suppose mean GTR = 4% and TP/N and TN/N are equally weighted, scores in assumptions are:\n\n* all-true submission: 0.02 (=(0.04 + 0) / 2)\n* all-false submission: 0.48 (=(0+ 0.96) / 2)\n* random submission: 0.25 (=(0.02 + 0.48) / 2)\n\nHowever, the actual scores are:\n\n* all-true submission: 0.51\n* all-false submission: 0.48\n* random submission (seed=123): 0.50\n\nI think TP/N and TN/N are not equally weighted (if supposed LB = w_p * TP/N + w_n * TN/N).",
    "1735143": "Good observation. Thank you for sharing.\n\nI think below sentence is mistaken:\n\n> ps the approximate percentage of TN would be --> 4% <--\n\n-> Correct: the approximate of **FN** (I.e. GT) would be 4% (= 1 - 0.48 * 2).\n",
    "1801676": ""
  }
}