{
  "id": 200757,
  "title": "Understanding the Evaluation Metric: Label Weighted LRAP",
  "url": "/competitions/rfcx-species-audio-detection/discussion/200757",
  "author_name": "",
  "post_date": "2020-12-01T18:43:32.317185100Z",
  "votes": 13,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Before diving deep into the metric, let's understand what we need to submit, it's vector dimensions and formats.</p>\n<h1>Submission Formats</h1>\n<p>We have got 24 classes and some 9000 training samples in total.  No, this much data is not sufficient to understand the <strong>Label weighted LRAP (label ranking average precision)</strong> metric. First, have a look at the below data chunk.</p>\n<table>\n<thead>\n<tr>\n<th>---</th>\n<th>recording_id</th>\n<th>species_id</th>\n<th>songtype_id</th>\n<th>t_min</th>\n<th>f_min</th>\n<th>t_max</th>\n<th>f_max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>00204008d</td>\n<td>21</td>\n<td>1</td>\n<td>13.8400</td>\n<td>3281.2500</td>\n<td>14.9333</td>\n<td>4125.0000</td>\n</tr>\n<tr>\n<td>1</td>\n<td>00204008d</td>\n<td>8</td>\n<td>1</td>\n<td>24.4960</td>\n<td>3750.0000</td>\n<td>28.6187</td>\n<td>5531.2500</td>\n</tr>\n<tr>\n<td>2</td>\n<td>00204008d</td>\n<td>4</td>\n<td>1</td>\n<td>15.0027</td>\n<td>2343.7500</td>\n<td>16.8587</td>\n<td>4218.7500</td>\n</tr>\n</tbody>\n</table>\n<p>What did you see? Well, if it didn't catch your attention, Let me explain as per my understanding. For single <code>recording id</code> <strong><em>00204008d</em></strong>, we have got 3 species ids in that, i.e. 21, 8, and 4. What does that mean? The recording, which means the audio sample having this mentioned id, has audio of species 24, species 8, and species 4. Hence, our ground truth for this data sample would be:</p>\n<blockquote>\n  <p>[0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1]. </p>\n</blockquote>\n<p>But what we need to predict? We need to predict the probability of the presence of each class in this audio sample. Hence our (hypothetical) predicted vector for this sample would be:</p>\n<blockquote>\n  <p>[0.01, 0.01, 0.01, 0.84, 0.01, 0.01, 0.01, 0.91, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.78]. </p>\n</blockquote>\n<p>So, yeah, it's a multilabel classification problem. </p>\n<p>Now quoting the submission format from the competition page:</p>\n<pre><code>recording_id,s0,...,s23\n000316da7,0.1,....,0.3\n003bc2cb2,0.0,...,0.8\n...\n</code></pre>\n<p>And we need to do this for what, 1992 samples? Okay, that being said, let's worry about the evaluation metric.</p>\n<p>PS: I will update this thread soon. Please correct me if I am wrong anywhere, thanks :)</p>",
  "messages": [
    {
      "id": "1098636",
      "postDate": "12/01/2020 18:43:32",
      "content": "<p>Before diving deep into the metric, let's understand what we need to submit, it's vector dimensions and formats.</p>\n<h1>Submission Formats</h1>\n<p>We have got 24 classes and some 9000 training samples in total.  No, this much data is not sufficient to understand the <strong>Label weighted LRAP (label ranking average precision)</strong> metric. First, have a look at the below data chunk.</p>\n<table>\n<thead>\n<tr>\n<th>---</th>\n<th>recording_id</th>\n<th>species_id</th>\n<th>songtype_id</th>\n<th>t_min</th>\n<th>f_min</th>\n<th>t_max</th>\n<th>f_max</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0</td>\n<td>00204008d</td>\n<td>21</td>\n<td>1</td>\n<td>13.8400</td>\n<td>3281.2500</td>\n<td>14.9333</td>\n<td>4125.0000</td>\n</tr>\n<tr>\n<td>1</td>\n<td>00204008d</td>\n<td>8</td>\n<td>1</td>\n<td>24.4960</td>\n<td>3750.0000</td>\n<td>28.6187</td>\n<td>5531.2500</td>\n</tr>\n<tr>\n<td>2</td>\n<td>00204008d</td>\n<td>4</td>\n<td>1</td>\n<td>15.0027</td>\n<td>2343.7500</td>\n<td>16.8587</td>\n<td>4218.7500</td>\n</tr>\n</tbody>\n</table>\n<p>What did you see? Well, if it didn't catch your attention, Let me explain as per my understanding. For single <code>recording id</code> <strong><em>00204008d</em></strong>, we have got 3 species ids in that, i.e. 21, 8, and 4. What does that mean? The recording, which means the audio sample having this mentioned id, has audio of species 24, species 8, and species 4. Hence, our ground truth for this data sample would be:</p>\n<blockquote>\n  <p>[0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1]. </p>\n</blockquote>\n<p>But what we need to predict? We need to predict the probability of the presence of each class in this audio sample. Hence our (hypothetical) predicted vector for this sample would be:</p>\n<blockquote>\n  <p>[0.01, 0.01, 0.01, 0.84, 0.01, 0.01, 0.01, 0.91, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.78]. </p>\n</blockquote>\n<p>So, yeah, it's a multilabel classification problem. </p>\n<p>Now quoting the submission format from the competition page:</p>\n<pre><code>recording_id,s0,...,s23\n000316da7,0.1,....,0.3\n003bc2cb2,0.0,...,0.8\n...\n</code></pre>\n<p>And we need to do this for what, 1992 samples? Okay, that being said, let's worry about the evaluation metric.</p>\n<p>PS: I will update this thread soon. Please correct me if I am wrong anywhere, thanks :)</p>",
      "rawMarkdown": "Before diving deep into the metric, let's understand what we need to submit, it's vector dimensions and formats.\n# Submission Formats\n We have got 24 classes and some 9000 training samples in total.  No, this much data is not sufficient to understand the **Label weighted LRAP (label ranking average precision)** metric. First, have a look at the below data chunk.\n\n|---|\trecording_id  |\tspecies_id  | songtype_id | t_min | f_min | t_max | f_max|\n|---|-------------- |------------|-------------- |-------|------ |--------|-------| \n| 0 | 00204008d | 21 | 1 | 13.8400 | 3281.2500 | 14.9333 | 4125.0000 |\n|1 | 00204008d| 8\t|1\t|24.4960\t|3750.0000 |28.6187 | 5531.2500\n|2 | 00204008d| 4 | 1| 15.0027  | 2343.7500 | 16.8587 |  4218.7500  |  \n\nWhat did you see? Well, if it didn't catch your attention, Let me explain as per my understanding. For single `recording id` ***00204008d***, we have got 3 species ids in that, i.e. 21, 8, and 4. What does that mean? The recording, which means the audio sample having this mentioned id, has audio of species 24, species 8, and species 4. Hence, our ground truth for this data sample would be:\n>  [0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1]. \n\nBut what we need to predict? We need to predict the probability of the presence of each class in this audio sample. Hence our (hypothetical) predicted vector for this sample would be:\n> [0.01, 0.01, 0.01, 0.84, 0.01, 0.01, 0.01, 0.91, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.78]. \n\nSo, yeah, it's a multilabel classification problem. \n\nNow quoting the submission format from the competition page:\n```\nrecording_id,s0,...,s23\n000316da7,0.1,....,0.3\n003bc2cb2,0.0,...,0.8\n...\n```\n\nAnd we need to do this for what, 1992 samples? Okay, that being said, let's worry about the evaluation metric.\n\nPS: I will update this thread soon. Please correct me if I am wrong anywhere, thanks :)",
      "votes": null
    },
    {
      "id": "1106991",
      "postDate": "12/09/2020 09:23:14",
      "content": "<p>This seems right <a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> in my opinion too. Instead of hard thresholding of each output neuron we just have to submit the probability score. </p>",
      "rawMarkdown": "This seems right @mrutyunjaybiswal in my opinion too. Instead of hard thresholding of each output neuron we just have to submit the probability score.",
      "votes": null
    },
    {
      "id": "1150094",
      "postDate": "01/12/2021 11:18:03",
      "content": "<p>That is very usefull thread. Looking forward for updates/discussions.. </p>\n<p>BTW you may want to correct a typo to avoid any confusion</p>\n<blockquote>\n  <p>has audio of species 24, species 8, and species 4. </p>\n</blockquote>\n<p>first one should be 21 (according to first line of data table)<br>\nsimilarly you need to correct the <code>ground truth</code> vector </p>",
      "rawMarkdown": "That is very usefull thread. Looking forward for updates/discussions.. \n\nBTW you may want to correct a typo to avoid any confusion\n> has audio of species 24, species 8, and species 4. \n\nfirst one should be 21 (according to first line of data table)\nsimilarly you need to correct the `ground truth` vector",
      "votes": null
    },
    {
      "id": "1163103",
      "postDate": "01/21/2021 14:29:05",
      "content": "<p>Cool!</p>\n<p>Waiting for the update  <a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> .</p>",
      "rawMarkdown": "Cool!\n\nWaiting for the update  @mrutyunjaybiswal .",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1106991,
      "author_name": "ayuraj",
      "author_url": "",
      "post_date": "12/09/2020 09:23:14",
      "content": "<p>This seems right <a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> in my opinion too. Instead of hard thresholding of each output neuron we just have to submit the probability score. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1150094,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "01/12/2021 11:18:03",
      "content": "<p>That is very usefull thread. Looking forward for updates/discussions.. </p>\n<p>BTW you may want to correct a typo to avoid any confusion</p>\n<blockquote>\n  <p>has audio of species 24, species 8, and species 4. </p>\n</blockquote>\n<p>first one should be 21 (according to first line of data table)<br>\nsimilarly you need to correct the <code>ground truth</code> vector </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1163103,
      "author_name": "joshi98kishan",
      "author_url": "",
      "post_date": "01/21/2021 14:29:05",
      "content": "<p>Cool!</p>\n<p>Waiting for the update  <a href=\"https://www.kaggle.com/mrutyunjaybiswal\" target=\"_blank\">@mrutyunjaybiswal</a> .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1098636": "Before diving deep into the metric, let's understand what we need to submit, it's vector dimensions and formats.\n# Submission Formats\n We have got 24 classes and some 9000 training samples in total.  No, this much data is not sufficient to understand the **Label weighted LRAP (label ranking average precision)** metric. First, have a look at the below data chunk.\n\n|---|\trecording_id  |\tspecies_id  | songtype_id | t_min | f_min | t_max | f_max|\n|---|-------------- |------------|-------------- |-------|------ |--------|-------| \n| 0 | 00204008d | 21 | 1 | 13.8400 | 3281.2500 | 14.9333 | 4125.0000 |\n|1 | 00204008d| 8\t|1\t|24.4960\t|3750.0000 |28.6187 | 5531.2500\n|2 | 00204008d| 4 | 1| 15.0027  | 2343.7500 | 16.8587 |  4218.7500  |  \n\nWhat did you see? Well, if it didn't catch your attention, Let me explain as per my understanding. For single `recording id` ***00204008d***, we have got 3 species ids in that, i.e. 21, 8, and 4. What does that mean? The recording, which means the audio sample having this mentioned id, has audio of species 24, species 8, and species 4. Hence, our ground truth for this data sample would be:\n>  [0, 0, 0, 1, 0, 0, 0, 1, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 0, 1]. \n\nBut what we need to predict? We need to predict the probability of the presence of each class in this audio sample. Hence our (hypothetical) predicted vector for this sample would be:\n> [0.01, 0.01, 0.01, 0.84, 0.01, 0.01, 0.01, 0.91, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.01, 0.78]. \n\nSo, yeah, it's a multilabel classification problem. \n\nNow quoting the submission format from the competition page:\n```\nrecording_id,s0,...,s23\n000316da7,0.1,....,0.3\n003bc2cb2,0.0,...,0.8\n...\n```\n\nAnd we need to do this for what, 1992 samples? Okay, that being said, let's worry about the evaluation metric.\n\nPS: I will update this thread soon. Please correct me if I am wrong anywhere, thanks :)",
    "1106991": "This seems right @mrutyunjaybiswal in my opinion too. Instead of hard thresholding of each output neuron we just have to submit the probability score.",
    "1150094": "That is very usefull thread. Looking forward for updates/discussions.. \n\nBTW you may want to correct a typo to avoid any confusion\n> has audio of species 24, species 8, and species 4. \n\nfirst one should be 21 (according to first line of data table)\nsimilarly you need to correct the `ground truth` vector",
    "1163103": "Cool!\n\nWaiting for the update  @mrutyunjaybiswal ."
  },
  "source": "meta"
}