{
  "id": 176199,
  "title": "Dealing with secondary labels",
  "url": "/competitions/birdsong-recognition/discussion/176199",
  "author_name": "",
  "post_date": "2020-08-20T21:41:02.503772400Z",
  "votes": 9,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Given that the test data must be labelled in 5-second increments with multiple classifications, it seems important to take the <code>secondary_labels</code> into consideration when training. </p>\n<p>I had proposed this pipeline to myself:</p>\n<ol>\n<li>Load a training file.</li>\n<li>Pad the file to satisfy (length &gt;= 0)</li>\n<li>Randomly select 5 seconds of resulting file</li>\n<li>Augment data as desired</li>\n</ol>\n<p>Then, I realized that the <code>secondary_labels</code> associated with the audio file don't have any sort of timestamp (as discussed <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/164741\" target=\"_blank\">here</a>). And, of course, neither does the primary call. </p>\n<p>So how does one guarantee the proper mix of primary and secondary bird calls are in a random 5-second clip matches the primary and secondary labels? </p>\n<p>I thought to use <code>librosa.split</code> and <code>librosa.remix</code> to cut out the noise, but it may be that this noise contains the information matching the secondary labels. </p>\n<p>The other option I thought of was to pad all files to a uniform size (say of the longest training file), and use the entire files as training data. But this ignores the fact that the test data must be labelled in 5 second increments. </p>\n<p>Finally, I thought of this:</p>\n<ol>\n<li>Grab a 5-second clip with primary label \"bluebird\" and secondary label \"crow\".</li>\n<li>Hope there is actually a bluebird in that 5 seconds.</li>\n<li>Grab a random 5 second clip with primary label \"crow\"</li>\n<li>Hope there is actually a crow in the second clip.</li>\n<li>Mix the two.</li>\n<li>Hope that there is both a bluebird and a crow in the final result.</li>\n</ol>\n<p>This involves a whole lot of hope. </p>\n<p><strong>So none of my ideas seem to be good ones.</strong></p>\n<p>What are your thoughts?</p>",
  "messages": [
    {
      "id": "979476",
      "postDate": "08/20/2020 21:41:02",
      "content": "<p>Given that the test data must be labelled in 5-second increments with multiple classifications, it seems important to take the <code>secondary_labels</code> into consideration when training. </p>\n<p>I had proposed this pipeline to myself:</p>\n<ol>\n<li>Load a training file.</li>\n<li>Pad the file to satisfy (length &gt;= 0)</li>\n<li>Randomly select 5 seconds of resulting file</li>\n<li>Augment data as desired</li>\n</ol>\n<p>Then, I realized that the <code>secondary_labels</code> associated with the audio file don't have any sort of timestamp (as discussed <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/164741\" target=\"_blank\">here</a>). And, of course, neither does the primary call. </p>\n<p>So how does one guarantee the proper mix of primary and secondary bird calls are in a random 5-second clip matches the primary and secondary labels? </p>\n<p>I thought to use <code>librosa.split</code> and <code>librosa.remix</code> to cut out the noise, but it may be that this noise contains the information matching the secondary labels. </p>\n<p>The other option I thought of was to pad all files to a uniform size (say of the longest training file), and use the entire files as training data. But this ignores the fact that the test data must be labelled in 5 second increments. </p>\n<p>Finally, I thought of this:</p>\n<ol>\n<li>Grab a 5-second clip with primary label \"bluebird\" and secondary label \"crow\".</li>\n<li>Hope there is actually a bluebird in that 5 seconds.</li>\n<li>Grab a random 5 second clip with primary label \"crow\"</li>\n<li>Hope there is actually a crow in the second clip.</li>\n<li>Mix the two.</li>\n<li>Hope that there is both a bluebird and a crow in the final result.</li>\n</ol>\n<p>This involves a whole lot of hope. </p>\n<p><strong>So none of my ideas seem to be good ones.</strong></p>\n<p>What are your thoughts?</p>",
      "rawMarkdown": "Given that the test data must be labelled in 5-second increments with multiple classifications, it seems important to take the `secondary_labels` into consideration when training. \n\nI had proposed this pipeline to myself:\n1. Load a training file.\n2. Pad the file to satisfy (length >= 0)\n3. Randomly select 5 seconds of resulting file\n4. Augment data as desired\n\nThen, I realized that the `secondary_labels` associated with the audio file don't have any sort of timestamp (as discussed [here](https://www.kaggle.com/c/birdsong-recognition/discussion/164741)). And, of course, neither does the primary call. \n\nSo how does one guarantee the proper mix of primary and secondary bird calls are in a random 5-second clip matches the primary and secondary labels? \n\nI thought to use `librosa.split` and `librosa.remix` to cut out the noise, but it may be that this noise contains the information matching the secondary labels. \n\nThe other option I thought of was to pad all files to a uniform size (say of the longest training file), and use the entire files as training data. But this ignores the fact that the test data must be labelled in 5 second increments. \n\nFinally, I thought of this:\n\n1. Grab a 5-second clip with primary label \"bluebird\" and secondary label \"crow\".\n2. Hope there is actually a bluebird in that 5 seconds.\n3. Grab a random 5 second clip with primary label \"crow\"\n4. Hope there is actually a crow in the second clip.\n5. Mix the two.\n6. Hope that there is both a bluebird and a crow in the final result.\n\nThis involves a whole lot of hope. \n\n**So none of my ideas seem to be good ones.**\n\nWhat are your thoughts?",
      "votes": null
    },
    {
      "id": "979523",
      "postDate": "08/20/2020 23:08:10",
      "content": "<p>Now, I was trying mix two loss by first label * 0.7 + secondary label * 0.3.<br>\nThe reason of the weight rate is first label emerge more in audio.</p>\n<p>Network is Densenet161, and my local fold-0 f1 score is:<br>\nFIrst label only: 0.6770335821<br>\nUse secondary label: 0.6464293118</p>\n<p>It seems not good, but I try more.</p>",
      "rawMarkdown": "Now, I was trying mix two loss by first label * 0.7 + secondary label * 0.3.\nThe reason of the weight rate is first label emerge more in audio.\n\nNetwork is Densenet161, and my local fold-0 f1 score is:\nFIrst label only: 0.6770335821\nUse secondary label: 0.6464293118\n\nIt seems not good, but I try more.",
      "votes": null
    },
    {
      "id": "979676",
      "postDate": "08/21/2020 03:54:44",
      "content": "<p>I wasted some time on 'garbage in, garbage out' models - where input quality was not good enough ('hoping' there is a bluebird, crow etc). I don't know what the smartest way is to get to cleanly/better labeled input data for the model, but I didn't feel like I made much progress feeding data in without inspection.</p>",
      "rawMarkdown": "I wasted some time on 'garbage in, garbage out' models - where input quality was not good enough ('hoping' there is a bluebird, crow etc). I don't know what the smartest way is to get to cleanly/better labeled input data for the model, but I didn't feel like I made much progress feeding data in without inspection.",
      "votes": null
    },
    {
      "id": "979819",
      "postDate": "08/21/2020 06:14:09",
      "content": "<p>I prepared a small pseudo test set of examples from train that had secondary labels (mapped to ebird codes via scientific name) with 5 sec audio  site 1/2/3 similar to the birdcall_check used in the inference notebooks which uses the first 15 rows of aldfly. Then could use these to check if the model predicted correctly any of the secondary labels so they could be used for training since the 5 second segment was known. Even with the aldfly birdcall_check test set it contains some secondary labels and can predict some of them correctly.  <br>\nBut since the hidden test data is soundscapes and different to what is in train, not really sure it helps for the real predictions. And the labels themselves can be noisy.  Maybe SED could give you some other ideas - <br>\n<a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a></p>",
      "rawMarkdown": "I prepared a small pseudo test set of examples from train that had secondary labels (mapped to ebird codes via scientific name) with 5 sec audio  site 1/2/3 similar to the birdcall_check used in the inference notebooks which uses the first 15 rows of aldfly. Then could use these to check if the model predicted correctly any of the secondary labels so they could be used for training since the 5 second segment was known. Even with the aldfly birdcall_check test set it contains some secondary labels and can predict some of them correctly.  \nBut since the hidden test data is soundscapes and different to what is in train, not really sure it helps for the real predictions. And the labels themselves can be noisy.  Maybe SED could give you some other ideas - \nhttps://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection",
      "votes": null
    },
    {
      "id": "979844",
      "postDate": "08/21/2020 06:35:55",
      "content": "<p>i am trying to use self-supervised method to learn representation (e.g. SWav and wave2vec). Then i would hand-label a small number of train images for transfer learning.</p>\n<p>here are some examples of results of hand labels. It explain the amount of noise in the train set from <a href=\"https://www.xeno-canto.org/\" target=\"_blank\">https://www.xeno-canto.org/</a> and why your methods will work or not.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3b298bcc3a0958d55eb47e01b8961fe0%2FSelection_027.png?generation=1597991675813645&amp;alt=media\" alt=\"\"></p>\n<p>SED gives pretty good clip level prediction. It can ensure that most confidence frame level (time-wise) prediction is correct but it cannot guarantee good results on \"all\"  frame level.</p>\n<p>if i can just annotate  a part of the train audio clip and some algorithm cam automatically determine the \"same bird\" in the clip, that will be useful.</p>",
      "rawMarkdown": "i am trying to use self-supervised method to learn representation (e.g. SWav and wave2vec). Then i would hand-label a small number of train images for transfer learning.\n\nhere are some examples of results of hand labels. It explain the amount of noise in the train set from https://www.xeno-canto.org/ and why your methods will work or not.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3b298bcc3a0958d55eb47e01b8961fe0%2FSelection_027.png?generation=1597991675813645&alt=media)\n\nSED gives pretty good clip level prediction. It can ensure that most confidence frame level (time-wise) prediction is correct but it cannot guarantee good results on \"all\"  frame level.\n\nif i can just annotate  a part of the train audio clip and some algorithm cam automatically determine the \"same bird\" in the clip, that will be useful.",
      "votes": null
    },
    {
      "id": "980884",
      "postDate": "08/22/2020 01:00:28",
      "content": "<p>Apparently, there is an area of research called \"sound event detection\" and people are exploring ideas of this field.<br>\nI found this starter kernel on the topic, hope it helps 😃: <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a></p>",
      "rawMarkdown": "Apparently, there is an area of research called \"sound event detection\" and people are exploring ideas of this field.\nI found this starter kernel on the topic, hope it helps 😃: https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection",
      "votes": null
    },
    {
      "id": "980888",
      "postDate": "08/22/2020 01:07:09",
      "content": "<p>Hye! Nice solution there!<br>\nHave you heard of wav2vec* 2.0? Perhaps you can use it! <br>\nPaper <a href=\"https://arxiv.org/abs/2006.11477\" target=\"_blank\">https://arxiv.org/abs/2006.11477</a></p>\n<p>Note: I think they did not release the code in a public repository as I was not able to find it, and pytorch is still working on implementing the method, but perhaps keep your radar on 😁</p>\n<ul>\n<li>edit1: swapped words here</li>\n</ul>",
      "rawMarkdown": "Hye! Nice solution there!\nHave you heard of wav2vec* 2.0? Perhaps you can use it! \nPaper https://arxiv.org/abs/2006.11477\n\nNote: I think they did not release the code in a public repository as I was not able to find it, and pytorch is still working on implementing the method, but perhaps keep your radar on 😁\n\n* edit1: swapped words here",
      "votes": null
    },
    {
      "id": "980903",
      "postDate": "08/22/2020 01:54:24",
      "content": "<p>word2vec is a good pretraining model. there is one paper applied it for spectrogram (instead of wave).<br>\n<a href=\"https://openreview.net/forum?id=r1xBk6NKwr\" target=\"_blank\">https://openreview.net/forum?id=r1xBk6NKwr</a></p>\n<p>\"We propose Audio2Vec, a self-supervised learning task that is inspired by Word2Vec,<br>\nbut applied to audio spectrograms\"</p>",
      "rawMarkdown": "word2vec is a good pretraining model. there is one paper applied it for spectrogram (instead of wave).\nhttps://openreview.net/forum?id=r1xBk6NKwr\n\n\"We propose Audio2Vec, a self-supervised learning task that is inspired by Word2Vec,\nbut applied to audio spectrograms\"",
      "votes": null
    },
    {
      "id": "980904",
      "postDate": "08/22/2020 01:57:32",
      "content": "<p>it is important to note that SED  (instance based learning) is not the only way to learn the frame-wise labels without strong labels.</p>\n<p>you can be creative and try other methods too, e.g. just training with data \"without secondary birds\" using softmax and then retrieve the CAM (class activation map), etc</p>",
      "rawMarkdown": "it is important to note that SED  (instance based learning) is not the only way to learn the frame-wise labels without strong labels.\n\nyou can be creative and try other methods too, e.g. just training with data \"without secondary birds\" using softmax and then retrieve the CAM (class activation map), etc",
      "votes": null
    },
    {
      "id": "980906",
      "postDate": "08/22/2020 02:00:10",
      "content": "<p>Oh, I swapped words as I was working with transformer models earlier hahaha</p>\n<p>I meant wav2vec 2.0 </p>",
      "rawMarkdown": "Oh, I swapped words as I was working with transformer models earlier hahaha\n\nI meant wav2vec 2.0",
      "votes": null
    },
    {
      "id": "981042",
      "postDate": "08/22/2020 06:03:52",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I guess you would be interested in <a href=\"https://arxiv.org/abs/1807.03748\" target=\"_blank\">CPC from DeepMind</a>.<br>\nRecently main author made very nice presentation at ICML2020 virtually, you will like this video:<br>\n<a href=\"https://slideslive.com/38930729/contrastive-learning-in-audio\" target=\"_blank\">https://slideslive.com/38930729/contrastive-learning-in-audio</a><br>\nI've been looking into representation learning these days, CPC is used by many audio researchers as far as I've checked…</p>",
      "rawMarkdown": "hengck23 I guess you would be interested in [CPC from DeepMind](https://arxiv.org/abs/1807.03748).\nRecently main author made very nice presentation at ICML2020 virtually, you will like this video:\nhttps://slideslive.com/38930729/contrastive-learning-in-audio\nI've been looking into representation learning these days, CPC is used by many audio researchers as far as I've checked...",
      "votes": null
    },
    {
      "id": "981069",
      "postDate": "08/22/2020 06:35:09",
      "content": "<p><a href=\"https://www.kaggle.com/daisukelab\" target=\"_blank\">@daisukelab</a> <br>\nthanks for the link. i will definitely talk a look into it.</p>\n<p>I think large scale self-supervised pre-learning is the future. an example is GPT-3. I hope these methods will be introduced to kaggle solutions. i have been looking at constrastive learning for image in the past, but i think  all text, audio, image, video will converge to some common method </p>\n<p>it will be interesting if we can learn domain shift  from noiseless to noisy sound environment using such methods.</p>",
      "rawMarkdown": "daisukelab \nthanks for the link. i will definitely talk a look into it.\n\nI think large scale self-supervised pre-learning is the future. an example is GPT-3. I hope these methods will be introduced to kaggle solutions. i have been looking at constrastive learning for image in the past, but i think  all text, audio, image, video will converge to some common method \n\nit will be interesting if we can learn domain shift  from noiseless to noisy sound environment using such methods.",
      "votes": null
    },
    {
      "id": "998594",
      "postDate": "09/04/2020 21:10:40",
      "content": "<p>I tried <code>secondary_labels</code> as additional information for the training but it did not result in better LB scores, my CV got also worse. I dont see the point in using <code>secondary_labels</code> especially with cutting out random 5 seconds of each audio-file.</p>",
      "rawMarkdown": "I tried `secondary_labels` as additional information for the training but it did not result in better LB scores, my CV got also worse. I dont see the point in using `secondary_labels` especially with cutting out random 5 seconds of each audio-file.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 979523,
      "author_name": "takamichitoda",
      "author_url": "",
      "post_date": "08/20/2020 23:08:10",
      "content": "<p>Now, I was trying mix two loss by first label * 0.7 + secondary label * 0.3.<br>\nThe reason of the weight rate is first label emerge more in audio.</p>\n<p>Network is Densenet161, and my local fold-0 f1 score is:<br>\nFIrst label only: 0.6770335821<br>\nUse secondary label: 0.6464293118</p>\n<p>It seems not good, but I try more.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 979676,
      "author_name": "davidedwards1",
      "author_url": "",
      "post_date": "08/21/2020 03:54:44",
      "content": "<p>I wasted some time on 'garbage in, garbage out' models - where input quality was not good enough ('hoping' there is a bluebird, crow etc). I don't know what the smartest way is to get to cleanly/better labeled input data for the model, but I didn't feel like I made much progress feeding data in without inspection.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 979819,
      "author_name": "something4kag",
      "author_url": "",
      "post_date": "08/21/2020 06:14:09",
      "content": "<p>I prepared a small pseudo test set of examples from train that had secondary labels (mapped to ebird codes via scientific name) with 5 sec audio  site 1/2/3 similar to the birdcall_check used in the inference notebooks which uses the first 15 rows of aldfly. Then could use these to check if the model predicted correctly any of the secondary labels so they could be used for training since the 5 second segment was known. Even with the aldfly birdcall_check test set it contains some secondary labels and can predict some of them correctly.  <br>\nBut since the hidden test data is soundscapes and different to what is in train, not really sure it helps for the real predictions. And the labels themselves can be noisy.  Maybe SED could give you some other ideas - <br>\n<a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 979844,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "08/21/2020 06:35:55",
      "content": "<p>i am trying to use self-supervised method to learn representation (e.g. SWav and wave2vec). Then i would hand-label a small number of train images for transfer learning.</p>\n<p>here are some examples of results of hand labels. It explain the amount of noise in the train set from <a href=\"https://www.xeno-canto.org/\" target=\"_blank\">https://www.xeno-canto.org/</a> and why your methods will work or not.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3b298bcc3a0958d55eb47e01b8961fe0%2FSelection_027.png?generation=1597991675813645&amp;alt=media\" alt=\"\"></p>\n<p>SED gives pretty good clip level prediction. It can ensure that most confidence frame level (time-wise) prediction is correct but it cannot guarantee good results on \"all\"  frame level.</p>\n<p>if i can just annotate  a part of the train audio clip and some algorithm cam automatically determine the \"same bird\" in the clip, that will be useful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 980888,
          "author_name": "ricafernandes",
          "author_url": "",
          "post_date": "08/22/2020 01:07:09",
          "content": "<p>Hye! Nice solution there!<br>\nHave you heard of wav2vec* 2.0? Perhaps you can use it! <br>\nPaper <a href=\"https://arxiv.org/abs/2006.11477\" target=\"_blank\">https://arxiv.org/abs/2006.11477</a></p>\n<p>Note: I think they did not release the code in a public repository as I was not able to find it, and pytorch is still working on implementing the method, but perhaps keep your radar on 😁</p>\n<ul>\n<li>edit1: swapped words here</li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980903,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/22/2020 01:54:24",
          "content": "<p>word2vec is a good pretraining model. there is one paper applied it for spectrogram (instead of wave).<br>\n<a href=\"https://openreview.net/forum?id=r1xBk6NKwr\" target=\"_blank\">https://openreview.net/forum?id=r1xBk6NKwr</a></p>\n<p>\"We propose Audio2Vec, a self-supervised learning task that is inspired by Word2Vec,<br>\nbut applied to audio spectrograms\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980904,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/22/2020 01:57:32",
          "content": "<p>it is important to note that SED  (instance based learning) is not the only way to learn the frame-wise labels without strong labels.</p>\n<p>you can be creative and try other methods too, e.g. just training with data \"without secondary birds\" using softmax and then retrieve the CAM (class activation map), etc</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 980906,
          "author_name": "ricafernandes",
          "author_url": "",
          "post_date": "08/22/2020 02:00:10",
          "content": "<p>Oh, I swapped words as I was working with transformer models earlier hahaha</p>\n<p>I meant wav2vec 2.0 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 981042,
          "author_name": "daisukelab",
          "author_url": "",
          "post_date": "08/22/2020 06:03:52",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> I guess you would be interested in <a href=\"https://arxiv.org/abs/1807.03748\" target=\"_blank\">CPC from DeepMind</a>.<br>\nRecently main author made very nice presentation at ICML2020 virtually, you will like this video:<br>\n<a href=\"https://slideslive.com/38930729/contrastive-learning-in-audio\" target=\"_blank\">https://slideslive.com/38930729/contrastive-learning-in-audio</a><br>\nI've been looking into representation learning these days, CPC is used by many audio researchers as far as I've checked…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 981069,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "08/22/2020 06:35:09",
          "content": "<p><a href=\"https://www.kaggle.com/daisukelab\" target=\"_blank\">@daisukelab</a> <br>\nthanks for the link. i will definitely talk a look into it.</p>\n<p>I think large scale self-supervised pre-learning is the future. an example is GPT-3. I hope these methods will be introduced to kaggle solutions. i have been looking at constrastive learning for image in the past, but i think  all text, audio, image, video will converge to some common method </p>\n<p>it will be interesting if we can learn domain shift  from noiseless to noisy sound environment using such methods.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 980884,
      "author_name": "ricafernandes",
      "author_url": "",
      "post_date": "08/22/2020 01:00:28",
      "content": "<p>Apparently, there is an area of research called \"sound event detection\" and people are exploring ideas of this field.<br>\nI found this starter kernel on the topic, hope it helps 😃: <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection\" target=\"_blank\">https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 998594,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "09/04/2020 21:10:40",
      "content": "<p>I tried <code>secondary_labels</code> as additional information for the training but it did not result in better LB scores, my CV got also worse. I dont see the point in using <code>secondary_labels</code> especially with cutting out random 5 seconds of each audio-file.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "979476": "Given that the test data must be labelled in 5-second increments with multiple classifications, it seems important to take the `secondary_labels` into consideration when training. \n\nI had proposed this pipeline to myself:\n1. Load a training file.\n2. Pad the file to satisfy (length >= 0)\n3. Randomly select 5 seconds of resulting file\n4. Augment data as desired\n\nThen, I realized that the `secondary_labels` associated with the audio file don't have any sort of timestamp (as discussed [here](https://www.kaggle.com/c/birdsong-recognition/discussion/164741)). And, of course, neither does the primary call. \n\nSo how does one guarantee the proper mix of primary and secondary bird calls are in a random 5-second clip matches the primary and secondary labels? \n\nI thought to use `librosa.split` and `librosa.remix` to cut out the noise, but it may be that this noise contains the information matching the secondary labels. \n\nThe other option I thought of was to pad all files to a uniform size (say of the longest training file), and use the entire files as training data. But this ignores the fact that the test data must be labelled in 5 second increments. \n\nFinally, I thought of this:\n\n1. Grab a 5-second clip with primary label \"bluebird\" and secondary label \"crow\".\n2. Hope there is actually a bluebird in that 5 seconds.\n3. Grab a random 5 second clip with primary label \"crow\"\n4. Hope there is actually a crow in the second clip.\n5. Mix the two.\n6. Hope that there is both a bluebird and a crow in the final result.\n\nThis involves a whole lot of hope. \n\n**So none of my ideas seem to be good ones.**\n\nWhat are your thoughts?",
    "979523": "Now, I was trying mix two loss by first label * 0.7 + secondary label * 0.3.\nThe reason of the weight rate is first label emerge more in audio.\n\nNetwork is Densenet161, and my local fold-0 f1 score is:\nFIrst label only: 0.6770335821\nUse secondary label: 0.6464293118\n\nIt seems not good, but I try more.",
    "979676": "I wasted some time on 'garbage in, garbage out' models - where input quality was not good enough ('hoping' there is a bluebird, crow etc). I don't know what the smartest way is to get to cleanly/better labeled input data for the model, but I didn't feel like I made much progress feeding data in without inspection.",
    "979819": "I prepared a small pseudo test set of examples from train that had secondary labels (mapped to ebird codes via scientific name) with 5 sec audio  site 1/2/3 similar to the birdcall_check used in the inference notebooks which uses the first 15 rows of aldfly. Then could use these to check if the model predicted correctly any of the secondary labels so they could be used for training since the 5 second segment was known. Even with the aldfly birdcall_check test set it contains some secondary labels and can predict some of them correctly.  \nBut since the hidden test data is soundscapes and different to what is in train, not really sure it helps for the real predictions. And the labels themselves can be noisy.  Maybe SED could give you some other ideas - \nhttps://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection",
    "979844": "i am trying to use self-supervised method to learn representation (e.g. SWav and wave2vec). Then i would hand-label a small number of train images for transfer learning.\n\nhere are some examples of results of hand labels. It explain the amount of noise in the train set from https://www.xeno-canto.org/ and why your methods will work or not.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F3b298bcc3a0958d55eb47e01b8961fe0%2FSelection_027.png?generation=1597991675813645&alt=media)\n\nSED gives pretty good clip level prediction. It can ensure that most confidence frame level (time-wise) prediction is correct but it cannot guarantee good results on \"all\"  frame level.\n\nif i can just annotate  a part of the train audio clip and some algorithm cam automatically determine the \"same bird\" in the clip, that will be useful.",
    "980884": "Apparently, there is an area of research called \"sound event detection\" and people are exploring ideas of this field.\nI found this starter kernel on the topic, hope it helps 😃: https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection",
    "980888": "Hye! Nice solution there!\nHave you heard of wav2vec* 2.0? Perhaps you can use it! \nPaper https://arxiv.org/abs/2006.11477\n\nNote: I think they did not release the code in a public repository as I was not able to find it, and pytorch is still working on implementing the method, but perhaps keep your radar on 😁\n\n* edit1: swapped words here",
    "980903": "word2vec is a good pretraining model. there is one paper applied it for spectrogram (instead of wave).\nhttps://openreview.net/forum?id=r1xBk6NKwr\n\n\"We propose Audio2Vec, a self-supervised learning task that is inspired by Word2Vec,\nbut applied to audio spectrograms\"",
    "980904": "it is important to note that SED  (instance based learning) is not the only way to learn the frame-wise labels without strong labels.\n\nyou can be creative and try other methods too, e.g. just training with data \"without secondary birds\" using softmax and then retrieve the CAM (class activation map), etc",
    "980906": "Oh, I swapped words as I was working with transformer models earlier hahaha\n\nI meant wav2vec 2.0",
    "981042": "hengck23 I guess you would be interested in [CPC from DeepMind](https://arxiv.org/abs/1807.03748).\nRecently main author made very nice presentation at ICML2020 virtually, you will like this video:\nhttps://slideslive.com/38930729/contrastive-learning-in-audio\nI've been looking into representation learning these days, CPC is used by many audio researchers as far as I've checked...",
    "981069": "daisukelab \nthanks for the link. i will definitely talk a look into it.\n\nI think large scale self-supervised pre-learning is the future. an example is GPT-3. I hope these methods will be introduced to kaggle solutions. i have been looking at constrastive learning for image in the past, but i think  all text, audio, image, video will converge to some common method \n\nit will be interesting if we can learn domain shift  from noiseless to noisy sound environment using such methods.",
    "998594": "I tried `secondary_labels` as additional information for the training but it did not result in better LB scores, my CV got also worse. I dont see the point in using `secondary_labels` especially with cutting out random 5 seconds of each audio-file."
  },
  "source": "meta"
}