{
  "id": 171247,
  "title": "What's your current validation scheme?",
  "url": "/competitions/birdsong-recognition/discussion/171247",
  "author_name": "Hidehisa Arai",
  "post_date": "2020-07-31T04:01:55.699000",
  "votes": 33,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Many of us know good validation matters a lot and some say it is the most important thing.\nHowever for this competition, creating good validation strategy is particularly difficult, since the train dataset and test dataset are totally different: and that is what the host meant to be. It's super-challenging, and very interesting research topic.</p>\n\n<p>For the time being, I've been using oof mAP score on train dataset for model selection but this is definitely not a good validation strategy. It got stuck around 0.70 and I think this comes from label noise and class imbalance.</p>\n\n<p>I've been trying to create event-level / segment-level labels as suggested in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170959#951943\">https://www.kaggle.com/c/birdsong-recognition/discussion/170959#951943</a> , using Weakly-supervised SED model, but so far I haven't got a luck. I think the direction I'm in is right: if we get event level labels, we can do </p>\n\n<ul>\n<li>negative mining (<code>nocall</code> class)</li>\n<li>cut out the birdcall events and mix them with arbitrary background sounds without being influenced by label noise: I think this is quite important because in this way maybe we can synthesize good validation set</li>\n</ul>\n\n<p>I'm wondering how are you guys tackling on this. Have you already get stable validation? What metrics are you using so far and around what value your model scores on? What do you think of the idea of creating event-level labels?</p>",
  "messages": [
    {
      "id": 952577,
      "postDate": "2020-07-31T04:01:55.700Z",
      "content": "<p>Many of us know good validation matters a lot and some say it is the most important thing.\nHowever for this competition, creating good validation strategy is particularly difficult, since the train dataset and test dataset are totally different: and that is what the host meant to be. It's super-challenging, and very interesting research topic.</p>\n\n<p>For the time being, I've been using oof mAP score on train dataset for model selection but this is definitely not a good validation strategy. It got stuck around 0.70 and I think this comes from label noise and class imbalance.</p>\n\n<p>I've been trying to create event-level / segment-level labels as suggested in <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/170959#951943\">https://www.kaggle.com/c/birdsong-recognition/discussion/170959#951943</a> , using Weakly-supervised SED model, but so far I haven't got a luck. I think the direction I'm in is right: if we get event level labels, we can do </p>\n\n<ul>\n<li>negative mining (<code>nocall</code> class)</li>\n<li>cut out the birdcall events and mix them with arbitrary background sounds without being influenced by label noise: I think this is quite important because in this way maybe we can synthesize good validation set</li>\n</ul>\n\n<p>I'm wondering how are you guys tackling on this. Have you already get stable validation? What metrics are you using so far and around what value your model scores on? What do you think of the idea of creating event-level labels?</p>",
      "rawMarkdown": "Many of us know good validation matters a lot and some say it is the most important thing.\nHowever for this competition, creating good validation strategy is particularly difficult, since the train dataset and test dataset are totally different: and that is what the host meant to be. It's super-challenging, and very interesting research topic.\n\nFor the time being, I've been using oof mAP score on train dataset for model selection but this is definitely not a good validation strategy. It got stuck around 0.70 and I think this comes from label noise and class imbalance.\n\nI've been trying to create event-level / segment-level labels as suggested in https://www.kaggle.com/c/birdsong-recognition/discussion/170959#951943 , using Weakly-supervised SED model, but so far I haven't got a luck. I think the direction I'm in is right: if we get event level labels, we can do \n\n* negative mining (`nocall` class)\n* cut out the birdcall events and mix them with arbitrary background sounds without being influenced by label noise: I think this is quite important because in this way maybe we can synthesize good validation set\n\nI'm wondering how are you guys tackling on this. Have you already get stable validation? What metrics are you using so far and around what value your model scores on? What do you think of the idea of creating event-level labels?",
      "votes": 32
    },
    {
      "id": 954409,
      "postDate": "2020-08-01T17:53:50.740Z",
      "content": "<p>preliminary results i have:</p>\n\n<p>1.create validation set:\n- fake event labels by using signal-to-noise ratio:\n- create PCEN mel-spectrum from input wave\n- estimate noise = media value of spectrum\n- signal(t) = max of spectrum at time t\n- if signal(t)/noise is higher than threshold, event has occurred\n- select some clip. Sample  5-sec intervals in which event has occurred as validation set</p>\n\n<hr>\n\n<p>2.train set = random 5-sec intervals\n- validation F1 =0.68</p>\n\n<p>3.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold)\n- validation F1 =0.72</p>\n\n<p>4.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold) + hand label(~1%)\n- validation F1 =0.75</p>",
      "rawMarkdown": "preliminary results i have:\n\n1.create validation set:\n- fake event labels by using signal-to-noise ratio:\n- create PCEN mel-spectrum from input wave\n- estimate noise = media value of spectrum\n- signal(t) = max of spectrum at time t\n- if signal(t)/noise is higher than threshold, event has occurred\n- select some clip. Sample  5-sec intervals in which event has occurred as validation set\n\n----\n\n2.train set = random 5-sec intervals\n- validation F1 =0.68\n\n3.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold)\n- validation F1 =0.72\n\n4.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold) + hand label(~1%)\n- validation F1 =0.75",
      "votes": 11,
      "replies": [
        {
          "id": 954671,
          "postDate": "2020-08-02T01:47:21.560Z",
          "content": "<p>Thanks for sharing! This is quite interesting.</p>\n\n<blockquote>\n  <p>fake event labels by using signal-to-noise ratio:</p>\n</blockquote>\n\n<p>Do you label detected events with primary label or do you have some other techniques to automatically distinguish the difference between those events and somehow label it also with secondary labels?</p>\n\n<blockquote>\n  <p>select some clip</p>\n</blockquote>\n\n<p>Are the scores below calculated on this selected subset?</p>\n\n<blockquote>\n  <p>.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold)\n  validation F1 =0.72</p>\n</blockquote>\n\n<p>Is this chunk-level F1, which is the main metric of this challenge?</p>",
          "rawMarkdown": "Thanks for sharing! This is quite interesting.\n\n&gt; fake event labels by using signal-to-noise ratio:\n\nDo you label detected events with primary label or do you have some other techniques to automatically distinguish the difference between those events and somehow label it also with secondary labels?\n\n&gt; select some clip\n\nAre the scores below calculated on this selected subset?\n\n&gt; .train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold)\n&gt; validation F1 =0.72\n\nIs this chunk-level F1, which is the main metric of this challenge?",
          "votes": 1
        },
        {
          "id": 955279,
          "postDate": "2020-08-02T14:11:28.817Z",
          "content": "<p>\"distinguish the difference between those events and somehow label it also with secondary labels?\"</p>\n\n<p>for the fake labels, i do not distinguish them. I just select high signal-to-noise clips. Hence this is noisy labels. it doesn't matter for the initial training because i am doing label clean up latter. </p>\n\n<p>\"Are the scores below calculated on this selected subset?\"\nyes. It is  F1 for 5-sec chunk</p>\n\n<p>for file based F1 (whole file), it is about 0.85.</p>",
          "rawMarkdown": "\"distinguish the difference between those events and somehow label it also with secondary labels?\"\n\n for the fake labels, i do not distinguish them. I just select high signal-to-noise clips. Hence this is noisy labels. it doesn't matter for the initial training because i am doing label clean up latter. \n\n\"Are the scores below calculated on this selected subset?\"\nyes. It is  F1 for 5-sec chunk\n\nfor file based F1 (whole file), it is about 0.85.",
          "votes": 2
        },
        {
          "id": 955306,
          "postDate": "2020-08-02T14:38:52.623Z",
          "content": "<p>Seems event-detection based on SNR somewhat works, thank you for sharing!</p>",
          "rawMarkdown": "Seems event-detection based on SNR somewhat works, thank you for sharing!",
          "votes": 2
        },
        {
          "id": 960383,
          "postDate": "2020-08-06T10:58:32.067Z",
          "content": "<p>Hello <a href=\"/hengck23\">@hengck23</a> and <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> ,</p>\n\n<blockquote>\n  <p>fake event labels by using signal-to-noise ratio</p>\n</blockquote>\n\n<p>What is the event in these fake event labels , I am assuming it is the label of which bird has called? </p>\n\n<blockquote>\n  <p>select some clip. Sample 5-sec intervals in which event has occurred as validation set</p>\n</blockquote>\n\n<p>Can you shed some light on this</p>",
          "rawMarkdown": "Hello @hengck23 and @hidehisaarai1213 ,\n\n&gt; fake event labels by using signal-to-noise ratio\n\nWhat is the event in these fake event labels , I am assuming it is the label of which bird has called? \n\n&gt; select some clip. Sample 5-sec intervals in which event has occurred as validation set\n\nCan you shed some light on this",
          "votes": 1
        },
        {
          "id": 961380,
          "postDate": "2020-08-07T06:12:11.060Z",
          "content": "<blockquote>\n  <p>What is the event in these fake event labels </p>\n</blockquote>\n\n<p>It's based on SNR, so event just means high SNR time span. As <a href=\"/hengck23\">@hengck23</a> describe, this can be treated as call event of the bird in primary label (but sometimes this assumption may be wrong).</p>\n\n<blockquote>\n  <blockquote>\n    <p>select some clip. Sample 5-sec intervals in which event has occurred as validation set</p>\n  </blockquote>\n  \n  <p>Can you shed some light on this</p>\n</blockquote>\n\n<p>I'm also interested in this.\n<a href=\"/hengck23\">@hengck23</a> How do you select clips? Just random?</p>",
          "rawMarkdown": "&gt; What is the event in these fake event labels \n\nIt's based on SNR, so event just means high SNR time span. As @hengck23 describe, this can be treated as call event of the bird in primary label (but sometimes this assumption may be wrong).\n\n&gt;&gt; select some clip. Sample 5-sec intervals in which event has occurred as validation set\n\n&gt; Can you shed some light on this\n\nI'm also interested in this.\n@hengck23 How do you select clips? Just random?",
          "votes": 2
        }
      ]
    },
    {
      "id": 952677,
      "postDate": "2020-07-31T06:22:39.767Z",
      "content": "<p>Hi <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> , thank you very much for sharing. Could you share more about what is SED model and how it is used for labels propagation? I think the direction you are going is definitely worth trying, but am wondering about how accurate we can label audio clips using weakly supervised learning.</p>",
      "rawMarkdown": "Hi @hidehisaarai1213 , thank you very much for sharing. Could you share more about what is SED model and how it is used for labels propagation? I think the direction you are going is definitely worth trying, but am wondering about how accurate we can label audio clips using weakly supervised learning.",
      "votes": 3,
      "replies": [
        {
          "id": 952820,
          "postDate": "2020-07-31T08:40:06.163Z",
          "content": "<p>Sure, SED is Sound Event Detection and it outputs probability of each class existence for each time segment. You can check the basic idea from PANNs repository, which is the one I'm also using in this competition.</p>\n\n<p><a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\">https://github.com/qiuqiangkong/audioset_tagging_cnn</a></p>\n\n<p>Basic idea of SED model is simple:\nthe output of CNN feature extractor still contains information about frequency and time(it should be 4 dimensional if that model relies on spectrogram based features: (batch size, channels, frequency, time)), so if we aggregate it only in frequency axis we can preserve time information on that feature map. That feature map has information about which time segment has what sound event. </p>\n\n<p>In weakly-supervised setting, we only have clip-level annotation, therefore we also need to aggregate that in time axis (In PANNs, they tried attention aggregation, MaxPooling aggregation and AvgPooling aggregation). Therefore we at first put classifier that outputs class existence probability for each time step just after the feature extractor and then aggregate the output of the classifier result in time axis.\nIn this way we can get both clip-level prediction and segment-level prediction (if the time resolution is high, it can be treated as event-level prediction). Then we train it normally by using BCE loss with clip-level prediction and clip-level annotation.</p>\n\n<p>I will show the prediction result example later on the data of this competition.</p>",
          "rawMarkdown": "Sure, SED is Sound Event Detection and it outputs probability of each class existence for each time segment. You can check the basic idea from PANNs repository, which is the one I'm also using in this competition.\n\nhttps://github.com/qiuqiangkong/audioset_tagging_cnn\n\nBasic idea of SED model is simple:\nthe output of CNN feature extractor still contains information about frequency and time(it should be 4 dimensional if that model relies on spectrogram based features: (batch size, channels, frequency, time)), so if we aggregate it only in frequency axis we can preserve time information on that feature map. That feature map has information about which time segment has what sound event. \n\nIn weakly-supervised setting, we only have clip-level annotation, therefore we also need to aggregate that in time axis (In PANNs, they tried attention aggregation, MaxPooling aggregation and AvgPooling aggregation). Therefore we at first put classifier that outputs class existence probability for each time step just after the feature extractor and then aggregate the output of the classifier result in time axis.\nIn this way we can get both clip-level prediction and segment-level prediction (if the time resolution is high, it can be treated as event-level prediction). Then we train it normally by using BCE loss with clip-level prediction and clip-level annotation.\n\nI will show the prediction result example later on the data of this competition.",
          "votes": 13
        },
        {
          "id": 953806,
          "postDate": "2020-08-01T05:48:50.350Z",
          "content": "<p>Thanks for the reply! Always a joy to learn from you :)</p>",
          "rawMarkdown": "Thanks for the reply! Always a joy to learn from you :)",
          "votes": 1
        },
        {
          "id": 962798,
          "postDate": "2020-08-08T12:53:49.743Z",
          "content": "<p>I just shared the <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection?scriptVersionId=40372870\">result example notebook</a>.</p>\n\n<p>Note that I didn't share any trained weight nor training method I used.</p>",
          "rawMarkdown": "I just shared the [result example notebook](https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection?scriptVersionId=40372870).\n\nNote that I didn't share any trained weight nor training method I used.",
          "votes": 2
        },
        {
          "id": 962877,
          "postDate": "2020-08-08T14:10:44.780Z",
          "content": "<p>Thanks for all of your sharings.</p>",
          "rawMarkdown": "Thanks for all of your sharings."
        }
      ]
    },
    {
      "id": 952611,
      "postDate": "2020-07-31T04:52:09.193Z",
      "content": "<p>currently i am not having any special validation strategy. My thought is to first crack the problem with data augmentation, domine adaptation(as we all know test data is in a different form), and a better feature extractor robust enough for domine changes.</p>\n\n<p>Now my validation loop consist of calculating F1 score for the validation split and i am using only the primary label for that . approx i get 0.6 - 0.7 .but, they don't have a lot of effect on the LB</p>",
      "rawMarkdown": "currently i am not having any special validation strategy. My thought is to first crack the problem with data augmentation, domine adaptation(as we all know test data is in a different form), and a better feature extractor robust enough for domine changes.\n\nNow my validation loop consist of calculating F1 score for the validation split and i am using only the primary label for that . approx i get 0.6 - 0.7 .but, they don't have a lot of effect on the LB",
      "votes": 4,
      "replies": [
        {
          "id": 952823,
          "postDate": "2020-07-31T08:46:33.223Z",
          "content": "<p>Thank you for sharing!</p>\n\n<blockquote>\n  <p>i am using only the primary label for that </p>\n</blockquote>\n\n<p>do you have any specific reason for this?</p>\n\n<blockquote>\n  <p>approx i get 0.6 - 0.7</p>\n</blockquote>\n\n<p>I also got around this score on macro-F1 when I tried training only with primary label, so I guess we don't have so much difference so far.</p>",
          "rawMarkdown": "Thank you for sharing!\n\n&gt; i am using only the primary label for that \n\ndo you have any specific reason for this?\n\n&gt; approx i get 0.6 - 0.7\n\nI also got around this score on macro-F1 when I tried training only with primary label, so I guess we don't have so much difference so far.",
          "votes": 2
        },
        {
          "id": 952854,
          "postDate": "2020-07-31T09:14:17.100Z",
          "content": "<p>My idea behind using primary label as the target over secondary labels.\n1. i did a run with both primary and secondary labels and it did not show any improvement in LB.\n2. compared to secondary label, primary label should be stronger.\n3. i found few BG species which are not belonging to the classes provided to use in the competition. i am not sure how to handle it as of now.</p>\n\n<p>i am pretty sure at some point i have to incorporate information from BG species as you shared in other posts . but, as of now i am not sure which method is best suited to crack this problem. one that is solved then i will start with enriching the data.</p>\n\n<p>i am considering SED using weak labels after reading you post. will try this over this weekend.</p>",
          "rawMarkdown": "My idea behind using primary label as the target over secondary labels.\n1. i did a run with both primary and secondary labels and it did not show any improvement in LB.\n2. compared to secondary label, primary label should be stronger.\n3. i found few BG species which are not belonging to the classes provided to use in the competition. i am not sure how to handle it as of now.\n\ni am pretty sure at some point i have to incorporate information from BG species as you shared in other posts . but, as of now i am not sure which method is best suited to crack this problem. one that is solved then i will start with enriching the data.\n\ni am considering SED using weak labels after reading you post. will try this over this weekend.",
          "votes": 3
        },
        {
          "id": 953642,
          "postDate": "2020-08-01T00:17:20.250Z",
          "content": "<p>Thanks! I got it.</p>\n\n<blockquote>\n  <p>i found few BG species which are not belonging to the classes provided to use in the competition. i am not sure how to handle it as of now.</p>\n</blockquote>\n\n<p>I think this should be <code>nocall</code>. This is also asked in kaggle.com/c/birdsong-recognition/discussion/171004 .</p>\n\n<blockquote>\n  <p>i am considering SED using weak labels after reading you post.</p>\n</blockquote>\n\n<p>Nice :)</p>",
          "rawMarkdown": "Thanks! I got it.\n\n&gt; i found few BG species which are not belonging to the classes provided to use in the competition. i am not sure how to handle it as of now.\n\nI think this should be `nocall`. This is also asked in kaggle.com/c/birdsong-recognition/discussion/171004 .\n\n&gt; i am considering SED using weak labels after reading you post.\n\nNice :)",
          "votes": 3
        }
      ]
    },
    {
      "id": 954338,
      "postDate": "2020-08-01T16:38:21.257Z",
      "content": "<p>Hey <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> Thanks for so nicely explaining SED . I am very new here , It would be highly helpful if you clear my following doubts : -\n* How are you splitting the dataset for a relieable CV , like what I am having a hard time with is the difference between train and test . If we just do stratified K-fold , it won't be correct as the test set contains 5 sec predictions from single file . So what should I do , should I also make chunks of single audio file and then use it as cv like alex has created for custom check? If you can explain it in depth\n* Is F1-score a good way to save models?\n* Are you using only mel spectograms rn?</p>\n\n<p>Thanks</p>",
      "rawMarkdown": "Hey @hidehisaarai1213 Thanks for so nicely explaining SED . I am very new here , It would be highly helpful if you clear my following doubts : -\n* How are you splitting the dataset for a relieable CV , like what I am having a hard time with is the difference between train and test . If we just do stratified K-fold , it won't be correct as the test set contains 5 sec predictions from single file . So what should I do , should I also make chunks of single audio file and then use it as cv like alex has created for custom check? If you can explain it in depth\n* Is F1-score a good way to save models?\n* Are you using only mel spectograms rn?\n\nThanks",
      "votes": 2,
      "replies": [
        {
          "id": 954656,
          "postDate": "2020-08-02T01:02:16.880Z",
          "content": "<blockquote>\n  <p>How are you splitting the dataset for a relieable CV , like what I am having a hard time with is the difference between train and test . If we just do stratified K-fold , it won't be correct as the test set contains 5 sec predictions from single file . </p>\n</blockquote>\n\n<p>I'm currently using Stratified KFold + 5sec random crop but I don't think this is reliable method. I also tried using <code>type</code> information (type of the call like <code>song</code>, <code>flight call</code>, <code>call</code>, etc.) and did 781 class classification (like <code>aldfly_song</code>, <code>btbwar_call</code>, etc.). At this time, I used Multilabel Stratified Kfold on that calltype label but it seems it's also not a very good choice: both CV and LB went down.</p>\n\n<blockquote>\n  <p>So what should I do , should I also make chunks of single audio file and then use it as cv like alex has created for custom check? If you can explain it in depth</p>\n</blockquote>\n\n<p>Sad to say I'm not sure about this. If we find out something stable, it would be a great advantage.</p>\n\n<blockquote>\n  <p>Is F1-score a good way to save models?</p>\n</blockquote>\n\n<p>I must be clear on this: F1-score on clip and F1-score on chunk is different. We usually don't have event level label so we can't use F1-score on chunk (I think this is what we need to maximize). F1-score on clip is therefore different from what we need to maximize, so I think it's not good to rely on it in the long run, but in the mean time we need to use F1-score on clip / mAP / lwlrap etc.</p>\n\n<blockquote>\n  <p>Are you using only mel spectograms rn?</p>\n</blockquote>\n\n<p>Yes. Technically log-melspectrogram.</p>",
          "rawMarkdown": "&gt; How are you splitting the dataset for a relieable CV , like what I am having a hard time with is the difference between train and test . If we just do stratified K-fold , it won't be correct as the test set contains 5 sec predictions from single file . \n\nI'm currently using Stratified KFold + 5sec random crop but I don't think this is reliable method. I also tried using `type` information (type of the call like `song`, `flight call`, `call`, etc.) and did 781 class classification (like `aldfly_song`, `btbwar_call`, etc.). At this time, I used Multilabel Stratified Kfold on that calltype label but it seems it's also not a very good choice: both CV and LB went down.\n\n&gt; So what should I do , should I also make chunks of single audio file and then use it as cv like alex has created for custom check? If you can explain it in depth\n\nSad to say I'm not sure about this. If we find out something stable, it would be a great advantage.\n\n&gt; Is F1-score a good way to save models?\n\nI must be clear on this: F1-score on clip and F1-score on chunk is different. We usually don't have event level label so we can't use F1-score on chunk (I think this is what we need to maximize). F1-score on clip is therefore different from what we need to maximize, so I think it's not good to rely on it in the long run, but in the mean time we need to use F1-score on clip / mAP / lwlrap etc.\n\n&gt; Are you using only mel spectograms rn?\n\nYes. Technically log-melspectrogram.",
          "votes": 2
        }
      ]
    },
    {
      "id": 966710,
      "postDate": "2020-08-11T16:10:13.910Z",
      "content": "<p>If I train with mixup, but validate with raw data. Can the validation set pick the best model? Why?</p>",
      "rawMarkdown": "If I train with mixup, but validate with raw data. Can the validation set pick the best model? Why?"
    },
    {
      "id": 959520,
      "postDate": "2020-08-05T16:48:09.017Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 954409,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2020-08-01T17:53:50.740000",
      "content": "<p>preliminary results i have:</p>\n\n<p>1.create validation set:\n- fake event labels by using signal-to-noise ratio:\n- create PCEN mel-spectrum from input wave\n- estimate noise = media value of spectrum\n- signal(t) = max of spectrum at time t\n- if signal(t)/noise is higher than threshold, event has occurred\n- select some clip. Sample  5-sec intervals in which event has occurred as validation set</p>\n\n<hr>\n\n<p>2.train set = random 5-sec intervals\n- validation F1 =0.68</p>\n\n<p>3.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold)\n- validation F1 =0.72</p>\n\n<p>4.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold) + hand label(~1%)\n- validation F1 =0.75</p>",
      "votes": 11,
      "replies": [
        {
          "id": 954671,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-02T01:47:21.560000",
          "content": "<p>Thanks for sharing! This is quite interesting.</p>\n\n<blockquote>\n  <p>fake event labels by using signal-to-noise ratio:</p>\n</blockquote>\n\n<p>Do you label detected events with primary label or do you have some other techniques to automatically distinguish the difference between those events and somehow label it also with secondary labels?</p>\n\n<blockquote>\n  <p>select some clip</p>\n</blockquote>\n\n<p>Are the scores below calculated on this selected subset?</p>\n\n<blockquote>\n  <p>.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold)\n  validation F1 =0.72</p>\n</blockquote>\n\n<p>Is this chunk-level F1, which is the main metric of this challenge?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 955279,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2020-08-02T14:11:28.817000",
          "content": "<p>\"distinguish the difference between those events and somehow label it also with secondary labels?\"</p>\n\n<p>for the fake labels, i do not distinguish them. I just select high signal-to-noise clips. Hence this is noisy labels. it doesn't matter for the initial training because i am doing label clean up latter. </p>\n\n<p>\"Are the scores below calculated on this selected subset?\"\nyes. It is  F1 for 5-sec chunk</p>\n\n<p>for file based F1 (whole file), it is about 0.85.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 955306,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-02T14:38:52.623000",
          "content": "<p>Seems event-detection based on SNR somewhat works, thank you for sharing!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 960383,
          "author_name": "Mr_KnowNothing",
          "author_url": "",
          "post_date": "2020-08-06T10:58:32.067000",
          "content": "<p>Hello <a href=\"/hengck23\">@hengck23</a> and <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> ,</p>\n\n<blockquote>\n  <p>fake event labels by using signal-to-noise ratio</p>\n</blockquote>\n\n<p>What is the event in these fake event labels , I am assuming it is the label of which bird has called? </p>\n\n<blockquote>\n  <p>select some clip. Sample 5-sec intervals in which event has occurred as validation set</p>\n</blockquote>\n\n<p>Can you shed some light on this</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 961380,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-07T06:12:11.060000",
          "content": "<blockquote>\n  <p>What is the event in these fake event labels </p>\n</blockquote>\n\n<p>It's based on SNR, so event just means high SNR time span. As <a href=\"/hengck23\">@hengck23</a> describe, this can be treated as call event of the bird in primary label (but sometimes this assumption may be wrong).</p>\n\n<blockquote>\n  <blockquote>\n    <p>select some clip. Sample 5-sec intervals in which event has occurred as validation set</p>\n  </blockquote>\n  \n  <p>Can you shed some light on this</p>\n</blockquote>\n\n<p>I'm also interested in this.\n<a href=\"/hengck23\">@hengck23</a> How do you select clips? Just random?</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 952677,
      "author_name": "Alan Choon",
      "author_url": "",
      "post_date": "2020-07-31T06:22:39.767000",
      "content": "<p>Hi <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> , thank you very much for sharing. Could you share more about what is SED model and how it is used for labels propagation? I think the direction you are going is definitely worth trying, but am wondering about how accurate we can label audio clips using weakly supervised learning.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 952820,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-07-31T08:40:06.163000",
          "content": "<p>Sure, SED is Sound Event Detection and it outputs probability of each class existence for each time segment. You can check the basic idea from PANNs repository, which is the one I'm also using in this competition.</p>\n\n<p><a href=\"https://github.com/qiuqiangkong/audioset_tagging_cnn\">https://github.com/qiuqiangkong/audioset_tagging_cnn</a></p>\n\n<p>Basic idea of SED model is simple:\nthe output of CNN feature extractor still contains information about frequency and time(it should be 4 dimensional if that model relies on spectrogram based features: (batch size, channels, frequency, time)), so if we aggregate it only in frequency axis we can preserve time information on that feature map. That feature map has information about which time segment has what sound event. </p>\n\n<p>In weakly-supervised setting, we only have clip-level annotation, therefore we also need to aggregate that in time axis (In PANNs, they tried attention aggregation, MaxPooling aggregation and AvgPooling aggregation). Therefore we at first put classifier that outputs class existence probability for each time step just after the feature extractor and then aggregate the output of the classifier result in time axis.\nIn this way we can get both clip-level prediction and segment-level prediction (if the time resolution is high, it can be treated as event-level prediction). Then we train it normally by using BCE loss with clip-level prediction and clip-level annotation.</p>\n\n<p>I will show the prediction result example later on the data of this competition.</p>",
          "votes": 13,
          "replies": []
        },
        {
          "id": 953806,
          "author_name": "Alan Choon",
          "author_url": "",
          "post_date": "2020-08-01T05:48:50.350000",
          "content": "<p>Thanks for the reply! Always a joy to learn from you :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 962798,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-08T12:53:49.743000",
          "content": "<p>I just shared the <a href=\"https://www.kaggle.com/hidehisaarai1213/introduction-to-sound-event-detection?scriptVersionId=40372870\">result example notebook</a>.</p>\n\n<p>Note that I didn't share any trained weight nor training method I used.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 962877,
          "author_name": "Yu Kang",
          "author_url": "",
          "post_date": "2020-08-08T14:10:44.780000",
          "content": "<p>Thanks for all of your sharings.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 952611,
      "author_name": "yuvaramsingh",
      "author_url": "",
      "post_date": "2020-07-31T04:52:09.193000",
      "content": "<p>currently i am not having any special validation strategy. My thought is to first crack the problem with data augmentation, domine adaptation(as we all know test data is in a different form), and a better feature extractor robust enough for domine changes.</p>\n\n<p>Now my validation loop consist of calculating F1 score for the validation split and i am using only the primary label for that . approx i get 0.6 - 0.7 .but, they don't have a lot of effect on the LB</p>",
      "votes": 4,
      "replies": [
        {
          "id": 952823,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-07-31T08:46:33.223000",
          "content": "<p>Thank you for sharing!</p>\n\n<blockquote>\n  <p>i am using only the primary label for that </p>\n</blockquote>\n\n<p>do you have any specific reason for this?</p>\n\n<blockquote>\n  <p>approx i get 0.6 - 0.7</p>\n</blockquote>\n\n<p>I also got around this score on macro-F1 when I tried training only with primary label, so I guess we don't have so much difference so far.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 952854,
          "author_name": "yuvaramsingh",
          "author_url": "",
          "post_date": "2020-07-31T09:14:17.100000",
          "content": "<p>My idea behind using primary label as the target over secondary labels.\n1. i did a run with both primary and secondary labels and it did not show any improvement in LB.\n2. compared to secondary label, primary label should be stronger.\n3. i found few BG species which are not belonging to the classes provided to use in the competition. i am not sure how to handle it as of now.</p>\n\n<p>i am pretty sure at some point i have to incorporate information from BG species as you shared in other posts . but, as of now i am not sure which method is best suited to crack this problem. one that is solved then i will start with enriching the data.</p>\n\n<p>i am considering SED using weak labels after reading you post. will try this over this weekend.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 953642,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-01T00:17:20.250000",
          "content": "<p>Thanks! I got it.</p>\n\n<blockquote>\n  <p>i found few BG species which are not belonging to the classes provided to use in the competition. i am not sure how to handle it as of now.</p>\n</blockquote>\n\n<p>I think this should be <code>nocall</code>. This is also asked in kaggle.com/c/birdsong-recognition/discussion/171004 .</p>\n\n<blockquote>\n  <p>i am considering SED using weak labels after reading you post.</p>\n</blockquote>\n\n<p>Nice :)</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 954338,
      "author_name": "Mr_KnowNothing",
      "author_url": "",
      "post_date": "2020-08-01T16:38:21.257000",
      "content": "<p>Hey <a href=\"/hidehisaarai1213\">@hidehisaarai1213</a> Thanks for so nicely explaining SED . I am very new here , It would be highly helpful if you clear my following doubts : -\n* How are you splitting the dataset for a relieable CV , like what I am having a hard time with is the difference between train and test . If we just do stratified K-fold , it won't be correct as the test set contains 5 sec predictions from single file . So what should I do , should I also make chunks of single audio file and then use it as cv like alex has created for custom check? If you can explain it in depth\n* Is F1-score a good way to save models?\n* Are you using only mel spectograms rn?</p>\n\n<p>Thanks</p>",
      "votes": 2,
      "replies": [
        {
          "id": 954656,
          "author_name": "Hidehisa Arai",
          "author_url": "",
          "post_date": "2020-08-02T01:02:16.880000",
          "content": "<blockquote>\n  <p>How are you splitting the dataset for a relieable CV , like what I am having a hard time with is the difference between train and test . If we just do stratified K-fold , it won't be correct as the test set contains 5 sec predictions from single file . </p>\n</blockquote>\n\n<p>I'm currently using Stratified KFold + 5sec random crop but I don't think this is reliable method. I also tried using <code>type</code> information (type of the call like <code>song</code>, <code>flight call</code>, <code>call</code>, etc.) and did 781 class classification (like <code>aldfly_song</code>, <code>btbwar_call</code>, etc.). At this time, I used Multilabel Stratified Kfold on that calltype label but it seems it's also not a very good choice: both CV and LB went down.</p>\n\n<blockquote>\n  <p>So what should I do , should I also make chunks of single audio file and then use it as cv like alex has created for custom check? If you can explain it in depth</p>\n</blockquote>\n\n<p>Sad to say I'm not sure about this. If we find out something stable, it would be a great advantage.</p>\n\n<blockquote>\n  <p>Is F1-score a good way to save models?</p>\n</blockquote>\n\n<p>I must be clear on this: F1-score on clip and F1-score on chunk is different. We usually don't have event level label so we can't use F1-score on chunk (I think this is what we need to maximize). F1-score on clip is therefore different from what we need to maximize, so I think it's not good to rely on it in the long run, but in the mean time we need to use F1-score on clip / mAP / lwlrap etc.</p>\n\n<blockquote>\n  <p>Are you using only mel spectograms rn?</p>\n</blockquote>\n\n<p>Yes. Technically log-melspectrogram.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 966710,
      "author_name": "Yu Kang",
      "author_url": "",
      "post_date": "2020-08-11T16:10:13.910000",
      "content": "<p>If I train with mixup, but validate with raw data. Can the validation set pick the best model? Why?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 959520,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-08-05T16:48:09.017000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "952577": "Many of us know good validation matters a lot and some say it is the most important thing.\nHowever for this competition, creating good validation strategy is particularly difficult, since the train dataset and test dataset are totally different: and that is what the host meant to be. It's super-challenging, and very interesting research topic.\n\nFor the time being, I've been using oof mAP score on train dataset for model selection but this is definitely not a good validation strategy. It got stuck around 0.70 and I think this comes from label noise and class imbalance.\n\nI've been trying to create event-level / segment-level labels as suggested in https://www.kaggle.com/c/birdsong-recognition/discussion/170959#951943 , using Weakly-supervised SED model, but so far I haven't got a luck. I think the direction I'm in is right: if we get event level labels, we can do \n\n* negative mining (`nocall` class)\n* cut out the birdcall events and mix them with arbitrary background sounds without being influenced by label noise: I think this is quite important because in this way maybe we can synthesize good validation set\n\nI'm wondering how are you guys tackling on this. Have you already get stable validation? What metrics are you using so far and around what value your model scores on? What do you think of the idea of creating event-level labels?",
    "954409": "preliminary results i have:\n\n1.create validation set:\n- fake event labels by using signal-to-noise ratio:\n- create PCEN mel-spectrum from input wave\n- estimate noise = media value of spectrum\n- signal(t) = max of spectrum at time t\n- if signal(t)/noise is higher than threshold, event has occurred\n- select some clip. Sample  5-sec intervals in which event has occurred as validation set\n\n----\n\n2.train set = random 5-sec intervals\n- validation F1 =0.68\n\n3.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold)\n- validation F1 =0.72\n\n4.train set = 5-sec intervals with fake event (i.e. signal(t)/noise is higher than threshold) + hand label(~1%)\n- validation F1 =0.75",
    "952677": "Hi @hidehisaarai1213 , thank you very much for sharing. Could you share more about what is SED model and how it is used for labels propagation? I think the direction you are going is definitely worth trying, but am wondering about how accurate we can label audio clips using weakly supervised learning.",
    "952611": "currently i am not having any special validation strategy. My thought is to first crack the problem with data augmentation, domine adaptation(as we all know test data is in a different form), and a better feature extractor robust enough for domine changes.\n\nNow my validation loop consist of calculating F1 score for the validation split and i am using only the primary label for that . approx i get 0.6 - 0.7 .but, they don't have a lot of effect on the LB",
    "954338": "Hey @hidehisaarai1213 Thanks for so nicely explaining SED . I am very new here , It would be highly helpful if you clear my following doubts : -\n* How are you splitting the dataset for a relieable CV , like what I am having a hard time with is the difference between train and test . If we just do stratified K-fold , it won't be correct as the test set contains 5 sec predictions from single file . So what should I do , should I also make chunks of single audio file and then use it as cv like alex has created for custom check? If you can explain it in depth\n* Is F1-score a good way to save models?\n* Are you using only mel spectograms rn?\n\nThanks",
    "966710": "If I train with mixup, but validate with raw data. Can the validation set pick the best model? Why?",
    "959520": ""
  }
}