{
  "id": 492876,
  "title": "Confusion regarding labeling",
  "url": "/competitions/birdclef-2024/discussion/492876",
  "author_name": "",
  "post_date": "2024-04-11T08:33:56.401279Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I have some confusions regarding the labelling of the train and test data. <br>\nIf I understand it correctly each recording we are given for training is assigned one label even though there are multiple 5 second windows. But I do not understand what the secondary labels are meant to mean. Is it that the recording contains sounds from multiple species? If so in which segments of the recording do these species show up?</p>",
  "messages": [
    {
      "id": "2746474",
      "postDate": "04/11/2024 08:33:56",
      "content": "<p>I have some confusions regarding the labelling of the train and test data. <br>\nIf I understand it correctly each recording we are given for training is assigned one label even though there are multiple 5 second windows. But I do not understand what the secondary labels are meant to mean. Is it that the recording contains sounds from multiple species? If so in which segments of the recording do these species show up?</p>",
      "rawMarkdown": "I have some confusions regarding the labelling of the train and test data. \nIf I understand it correctly each recording we are given for training is assigned one label even though there are multiple 5 second windows. But I do not understand what the secondary labels are meant to mean. Is it that the recording contains sounds from multiple species? If so in which segments of the recording do these species show up?",
      "votes": null
    },
    {
      "id": "2746481",
      "postDate": "04/11/2024 08:43:53",
      "content": "<p>This is the challenge of the competition, the answer to your question is that we don't know. The goal is to make strong labels from soft ones. It is up to you to find a way of training your model to confuse it as little as possible knowing those labels don't always appear in 5 second segments.</p>",
      "rawMarkdown": "This is the challenge of the competition, the answer to your question is that we don't know. The goal is to make strong labels from soft ones. It is up to you to find a way of training your model to confuse it as little as possible knowing those labels don't always appear in 5 second segments.",
      "votes": null
    },
    {
      "id": "2746484",
      "postDate": "04/11/2024 08:46:28",
      "content": "<p>The labelling is a major challenge in this competition, you only know that a certain set of species is present in a recording, but you do not know where in the recordings.</p>\n<p>So you could have a 2 minute recording with a single 1 second bird call.<br>\nWith multiple bird species in a single recording it gets tricky to assign the label to a random 5 second window.</p>\n<p>One approach would be to exclude the samples with multiple species and try to filter 5 second windows with low intensity, implying the absence of a bird call.</p>",
      "rawMarkdown": "The labelling is a major challenge in this competition, you only know that a certain set of species is present in a recording, but you do not know where in the recordings.\n\nSo you could have a 2 minute recording with a single 1 second bird call.\nWith multiple bird species in a single recording it gets tricky to assign the label to a random 5 second window.\n\nOne approach would be to exclude the samples with multiple species and try to filter 5 second windows with low intensity, implying the absence of a bird call.",
      "votes": null
    },
    {
      "id": "2746785",
      "postDate": "04/11/2024 13:14:11",
      "content": "<p>Well. Let's clarify a thing about labels. For example we have audio for \"Black-rumped Flameback\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Fa2803c4aa887228b63a5897eaa047df4%2FDinopium_benghalense.jpeg?generation=1712840084436385&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F2e5bb7a60593c9e5f6061c1ef60ad19e%2FScreenshot%202024-04-11%20at%2014.52.12.png?generation=1712839971577524&amp;alt=media\"></p>\n<p>Sound could be found in train /kaggle/input/birdclef-2024/train_audio/bkrfla1/XC142801.ogg</p>\n<p>It is 27 seconds record. Full spectrogram looks like </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F33db1efbb0bbe6229f7e3863de13db65%2FScreenshot%202024-04-11%20at%2014.58.51.png?generation=1712840363266721&amp;alt=media\"></p>\n<p>Whole record has 1 primary label - <strong>bkrfla1</strong></p>\n<p>But look what will happens if we cut it to 5 second specs:</p>\n<p>0-5 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F865745062f2e1ed2a47cce57c010f311%2FXC142801_05.png?generation=1712840492234676&amp;alt=media\"></p>\n<p>6-10 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F1e64f87d66423feca5969f30433cf335%2FXC142801_10.png?generation=1712840502819808&amp;alt=media\"></p>\n<p>11-15 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff1479bd0f947f43216303c94ee2e8e12%2FXC142801_15.png?generation=1712840510823608&amp;alt=media\"></p>\n<p>15-20 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff6a9913a420627949c788f5667ca87d9%2FXC142801_25.png?generation=1712840520406191&amp;alt=media\"></p>\n<p>These cuts is what your model should identify correctly. What would you and organizers expect from your model predict output for each of the 5 seconds cut? </p>\n<ul>\n<li>Is there a Black-rumped Flameback voice at 0-5 seconds ? What is the probability the bird is there?</li>\n<li>Is there a Black-rumped Flameback voice at 6-11 seconds? What is the probability the bird is there?</li>\n</ul>\n<p>Same questions for all other species should be answered for each 5 second split for each of the record model will got as an input. </p>\n<p>Personaly I'm currious now on the data we should train the model:</p>\n<ol>\n<li>Should we train the model(s) on silences, noises, and other \"empty' pieces of the recordings?</li>\n<li>Should we show whole 27 second record to model or it should be cleaned and trimed before?</li>\n</ol>",
      "rawMarkdown": "Well. Let's clarify a thing about labels. For example we have audio for \"Black-rumped Flameback\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Fa2803c4aa887228b63a5897eaa047df4%2FDinopium_benghalense.jpeg?generation=1712840084436385&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F2e5bb7a60593c9e5f6061c1ef60ad19e%2FScreenshot%202024-04-11%20at%2014.52.12.png?generation=1712839971577524&alt=media)\n\nSound could be found in train /kaggle/input/birdclef-2024/train_audio/bkrfla1/XC142801.ogg\n\nIt is 27 seconds record. Full spectrogram looks like \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F33db1efbb0bbe6229f7e3863de13db65%2FScreenshot%202024-04-11%20at%2014.58.51.png?generation=1712840363266721&alt=media)\n\nWhole record has 1 primary label - **bkrfla1**\n\nBut look what will happens if we cut it to 5 second specs:\n\n0-5 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F865745062f2e1ed2a47cce57c010f311%2FXC142801_05.png?generation=1712840492234676&alt=media)\n\n6-10 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F1e64f87d66423feca5969f30433cf335%2FXC142801_10.png?generation=1712840502819808&alt=media)\n\n11-15 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff1479bd0f947f43216303c94ee2e8e12%2FXC142801_15.png?generation=1712840510823608&alt=media)\n\n15-20 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff6a9913a420627949c788f5667ca87d9%2FXC142801_25.png?generation=1712840520406191&alt=media)\n\nThese cuts is what your model should identify correctly. What would you and organizers expect from your model predict output for each of the 5 seconds cut? \n\n- Is there a Black-rumped Flameback voice at 0-5 seconds ? What is the probability the bird is there?\n- Is there a Black-rumped Flameback voice at 6-11 seconds? What is the probability the bird is there?\n\nSame questions for all other species should be answered for each 5 second split for each of the record model will got as an input. \n\nPersonaly I'm currious now on the data we should train the model:\n1. Should we train the model(s) on silences, noises, and other \"empty' pieces of the recordings?\n2. Should we show whole 27 second record to model or it should be cleaned and trimed before?",
      "votes": null
    },
    {
      "id": "2760939",
      "postDate": "04/19/2024 15:41:21",
      "content": "<p>Hi! Yes, the recording can contain multiple species - the secondary label implies that there is another species that had been identified while labeling/annotating the data. </p>",
      "rawMarkdown": "Hi! Yes, the recording can contain multiple species - the secondary label implies that there is another species that had been identified while labeling/annotating the data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2746481,
      "author_name": "janmpia",
      "author_url": "",
      "post_date": "04/11/2024 08:43:53",
      "content": "<p>This is the challenge of the competition, the answer to your question is that we don't know. The goal is to make strong labels from soft ones. It is up to you to find a way of training your model to confuse it as little as possible knowing those labels don't always appear in 5 second segments.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2746484,
      "author_name": "markwijkhuizen",
      "author_url": "",
      "post_date": "04/11/2024 08:46:28",
      "content": "<p>The labelling is a major challenge in this competition, you only know that a certain set of species is present in a recording, but you do not know where in the recordings.</p>\n<p>So you could have a 2 minute recording with a single 1 second bird call.<br>\nWith multiple bird species in a single recording it gets tricky to assign the label to a random 5 second window.</p>\n<p>One approach would be to exclude the samples with multiple species and try to filter 5 second windows with low intensity, implying the absence of a bird call.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2746785,
      "author_name": "samvelkoch",
      "author_url": "",
      "post_date": "04/11/2024 13:14:11",
      "content": "<p>Well. Let's clarify a thing about labels. For example we have audio for \"Black-rumped Flameback\"<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Fa2803c4aa887228b63a5897eaa047df4%2FDinopium_benghalense.jpeg?generation=1712840084436385&amp;alt=media\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F2e5bb7a60593c9e5f6061c1ef60ad19e%2FScreenshot%202024-04-11%20at%2014.52.12.png?generation=1712839971577524&amp;alt=media\"></p>\n<p>Sound could be found in train /kaggle/input/birdclef-2024/train_audio/bkrfla1/XC142801.ogg</p>\n<p>It is 27 seconds record. Full spectrogram looks like </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F33db1efbb0bbe6229f7e3863de13db65%2FScreenshot%202024-04-11%20at%2014.58.51.png?generation=1712840363266721&amp;alt=media\"></p>\n<p>Whole record has 1 primary label - <strong>bkrfla1</strong></p>\n<p>But look what will happens if we cut it to 5 second specs:</p>\n<p>0-5 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F865745062f2e1ed2a47cce57c010f311%2FXC142801_05.png?generation=1712840492234676&amp;alt=media\"></p>\n<p>6-10 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F1e64f87d66423feca5969f30433cf335%2FXC142801_10.png?generation=1712840502819808&amp;alt=media\"></p>\n<p>11-15 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff1479bd0f947f43216303c94ee2e8e12%2FXC142801_15.png?generation=1712840510823608&amp;alt=media\"></p>\n<p>15-20 sec<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff6a9913a420627949c788f5667ca87d9%2FXC142801_25.png?generation=1712840520406191&amp;alt=media\"></p>\n<p>These cuts is what your model should identify correctly. What would you and organizers expect from your model predict output for each of the 5 seconds cut? </p>\n<ul>\n<li>Is there a Black-rumped Flameback voice at 0-5 seconds ? What is the probability the bird is there?</li>\n<li>Is there a Black-rumped Flameback voice at 6-11 seconds? What is the probability the bird is there?</li>\n</ul>\n<p>Same questions for all other species should be answered for each 5 second split for each of the record model will got as an input. </p>\n<p>Personaly I'm currious now on the data we should train the model:</p>\n<ol>\n<li>Should we train the model(s) on silences, noises, and other \"empty' pieces of the recordings?</li>\n<li>Should we show whole 27 second record to model or it should be cleaned and trimed before?</li>\n</ol>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2760939,
      "author_name": "wgbirds",
      "author_url": "",
      "post_date": "04/19/2024 15:41:21",
      "content": "<p>Hi! Yes, the recording can contain multiple species - the secondary label implies that there is another species that had been identified while labeling/annotating the data. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2746474": "I have some confusions regarding the labelling of the train and test data. \nIf I understand it correctly each recording we are given for training is assigned one label even though there are multiple 5 second windows. But I do not understand what the secondary labels are meant to mean. Is it that the recording contains sounds from multiple species? If so in which segments of the recording do these species show up?",
    "2746481": "This is the challenge of the competition, the answer to your question is that we don't know. The goal is to make strong labels from soft ones. It is up to you to find a way of training your model to confuse it as little as possible knowing those labels don't always appear in 5 second segments.",
    "2746484": "The labelling is a major challenge in this competition, you only know that a certain set of species is present in a recording, but you do not know where in the recordings.\n\nSo you could have a 2 minute recording with a single 1 second bird call.\nWith multiple bird species in a single recording it gets tricky to assign the label to a random 5 second window.\n\nOne approach would be to exclude the samples with multiple species and try to filter 5 second windows with low intensity, implying the absence of a bird call.",
    "2746785": "Well. Let's clarify a thing about labels. For example we have audio for \"Black-rumped Flameback\"\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Fa2803c4aa887228b63a5897eaa047df4%2FDinopium_benghalense.jpeg?generation=1712840084436385&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F2e5bb7a60593c9e5f6061c1ef60ad19e%2FScreenshot%202024-04-11%20at%2014.52.12.png?generation=1712839971577524&alt=media)\n\nSound could be found in train /kaggle/input/birdclef-2024/train_audio/bkrfla1/XC142801.ogg\n\nIt is 27 seconds record. Full spectrogram looks like \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F33db1efbb0bbe6229f7e3863de13db65%2FScreenshot%202024-04-11%20at%2014.58.51.png?generation=1712840363266721&alt=media)\n\nWhole record has 1 primary label - **bkrfla1**\n\nBut look what will happens if we cut it to 5 second specs:\n\n0-5 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F865745062f2e1ed2a47cce57c010f311%2FXC142801_05.png?generation=1712840492234676&alt=media)\n\n6-10 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2F1e64f87d66423feca5969f30433cf335%2FXC142801_10.png?generation=1712840502819808&alt=media)\n\n11-15 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff1479bd0f947f43216303c94ee2e8e12%2FXC142801_15.png?generation=1712840510823608&alt=media)\n\n15-20 sec\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10356799%2Ff6a9913a420627949c788f5667ca87d9%2FXC142801_25.png?generation=1712840520406191&alt=media)\n\nThese cuts is what your model should identify correctly. What would you and organizers expect from your model predict output for each of the 5 seconds cut? \n\n- Is there a Black-rumped Flameback voice at 0-5 seconds ? What is the probability the bird is there?\n- Is there a Black-rumped Flameback voice at 6-11 seconds? What is the probability the bird is there?\n\nSame questions for all other species should be answered for each 5 second split for each of the record model will got as an input. \n\nPersonaly I'm currious now on the data we should train the model:\n1. Should we train the model(s) on silences, noises, and other \"empty' pieces of the recordings?\n2. Should we show whole 27 second record to model or it should be cleaned and trimed before?",
    "2760939": "Hi! Yes, the recording can contain multiple species - the secondary label implies that there is another species that had been identified while labeling/annotating the data."
  },
  "source": "meta"
}