{
  "id": 167202,
  "title": "Very challenging points about CBR",
  "url": "/competitions/birdsong-recognition/discussion/167202",
  "author_name": "",
  "post_date": "2020-07-15T15:41:33.548391600Z",
  "votes": 43,
  "comment_count": 20,
  "views": 0,
  "content": "<p>Three weeks now since the start of this CBR competition and we could hardly overtake the <strong>sample_submission</strong> ! During my experiments, I come up with these challenging points which require specific attention:</p>\n\n<ul>\n<li>Training set is somehow clean but test set is very noisy</li>\n<li>There is a lot of \"<strong>nocall</strong>\" in the test set (~46%)</li>\n<li><strong>Nocall</strong> prediction by thresholding is somehow subjective (need a cross-val there but ...)</li>\n<li>Testing is not obvious as we have no significant sample of the test set</li>\n<li>Training set contains one bird at a time while test set does contain many</li>\n<li>....</li>\n</ul>\n\n<p>Those points make me believing that preprocessing will play a huge role during this CBR competition.</p>\n\n<h1>Update 1 (<a href=\"/hidehisaarai1213\">@hidehisaarai1213</a>)</h1>\n\n<ul>\n<li>There are many types of song / calls even in a single species</li>\n<li>Annotation to the training set is actually not enough\n<ul><li>Some of them have songs/calls of multiple species in fact, but not annotated</li>\n<li>Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training</li></ul></li>\n<li>We are not sure about the annotation level in test set\n<ul><li>Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.</li></ul></li>\n<li>Some calls of different species are very similar and not so distinguishable.</li>\n</ul>\n\n<p><strong>PS</strong>: Don't mind adding your challenging points in the comment section, I will manage to report them all !</p>",
  "messages": [
    {
      "id": "930613",
      "postDate": "07/15/2020 15:41:33",
      "content": "<p>Three weeks now since the start of this CBR competition and we could hardly overtake the <strong>sample_submission</strong> ! During my experiments, I come up with these challenging points which require specific attention:</p>\n\n<ul>\n<li>Training set is somehow clean but test set is very noisy</li>\n<li>There is a lot of \"<strong>nocall</strong>\" in the test set (~46%)</li>\n<li><strong>Nocall</strong> prediction by thresholding is somehow subjective (need a cross-val there but ...)</li>\n<li>Testing is not obvious as we have no significant sample of the test set</li>\n<li>Training set contains one bird at a time while test set does contain many</li>\n<li>....</li>\n</ul>\n\n<p>Those points make me believing that preprocessing will play a huge role during this CBR competition.</p>\n\n<h1>Update 1 (<a href=\"/hidehisaarai1213\">@hidehisaarai1213</a>)</h1>\n\n<ul>\n<li>There are many types of song / calls even in a single species</li>\n<li>Annotation to the training set is actually not enough\n<ul><li>Some of them have songs/calls of multiple species in fact, but not annotated</li>\n<li>Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training</li></ul></li>\n<li>We are not sure about the annotation level in test set\n<ul><li>Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.</li></ul></li>\n<li>Some calls of different species are very similar and not so distinguishable.</li>\n</ul>\n\n<p><strong>PS</strong>: Don't mind adding your challenging points in the comment section, I will manage to report them all !</p>",
      "rawMarkdown": "Three weeks now since the start of this CBR competition and we could hardly overtake the **sample_submission** ! During my experiments, I come up with these challenging points which require specific attention:\n\n* Training set is somehow clean but test set is very noisy\n*  There is a lot of \"**nocall**\" in the test set (~46%)\n* **Nocall** prediction by thresholding is somehow subjective (need a cross-val there but ...)\n* Testing is not obvious as we have no significant sample of the test set\n* Training set contains one bird at a time while test set does contain many\n* ....\n\n\nThose points make me believing that preprocessing will play a huge role during this CBR competition.\n\n#Update 1 (@hidehisaarai1213)\n* There are many types of song / calls even in a single species\n* Annotation to the training set is actually not enough\n  * Some of them have songs/calls of multiple species in fact, but not annotated\n  * Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training\n* We are not sure about the annotation level in test set\n   * Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.\n* Some calls of different species are very similar and not so distinguishable.\n\n**PS**: Don't mind adding your challenging points in the comment section, I will manage to report them all !",
      "votes": null
    },
    {
      "id": "930939",
      "postDate": "07/15/2020 20:29:39",
      "content": "<p>Thanks for summarization!</p>\n\n<p>These are difficult but interesting points of this competition. </p>",
      "rawMarkdown": "Thanks for summarization!\n\nThese are difficult but interesting points of this competition.",
      "votes": null
    },
    {
      "id": "931016",
      "postDate": "07/15/2020 22:37:42",
      "content": "<p>Yes...very challenging and of course interesting 👍 </p>",
      "rawMarkdown": "Yes...very challenging and of course interesting 👍",
      "votes": null
    },
    {
      "id": "931032",
      "postDate": "07/15/2020 22:56:50",
      "content": "<p>Here are some others I have found</p>\n\n<ul>\n<li>There are many types of song / calls even in a single species</li>\n<li>Annotation to the training set is actually not enough\n<ul><li>Some of them have songs/calls of multiple species in fact, but not annotated</li>\n<li>Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training</li></ul></li>\n<li>We are not sure about the annotation level in test set\n<ul><li>Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.</li></ul></li>\n<li>Some calls of different species are very similar and not so distinguishable.</li>\n</ul>\n\n<p>The host is also sharing some of the challenges <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/163899#914790\">here</a>.</p>",
      "rawMarkdown": "Here are some others I have found\n\n* There are many types of song / calls even in a single species\n* Annotation to the training set is actually not enough\n  - Some of them have songs/calls of multiple species in fact, but not annotated\n  - Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training\n* We are not sure about the annotation level in test set\n  - Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.\n* Some calls of different species are very similar and not so distinguishable.\n\nThe host is also sharing some of the challenges [here](https://www.kaggle.com/c/birdsong-recognition/discussion/163899#914790).",
      "votes": null
    },
    {
      "id": "931128",
      "postDate": "07/16/2020 02:19:09",
      "content": "<p>Thanks for pointing out those points. I've reported them ! </p>",
      "rawMarkdown": "Thanks for pointing out those points. I've reported them !",
      "votes": null
    },
    {
      "id": "931299",
      "postDate": "07/16/2020 05:34:59",
      "content": "<p>My biggest challenge is the lack of proper validation set. Sound-scape recordings could have completely different characteristics than the training set.</p>\n\n<p>I think this competition requires relatively complex pipeline (audio preprocessing, call/not call classifier, augmentation, CNN, prediction post-processing) compared to the typical kaggle competitions. It might also explain the slow progress.</p>",
      "rawMarkdown": "My biggest challenge is the lack of proper validation set. Sound-scape recordings could have completely different characteristics than the training set.\n\nI think this competition requires relatively complex pipeline (audio preprocessing, call/not call classifier, augmentation, CNN, prediction post-processing) compared to the typical kaggle competitions. It might also explain the slow progress.",
      "votes": null
    },
    {
      "id": "931356",
      "postDate": "07/16/2020 06:31:09",
      "content": "<p>totally agree</p>",
      "rawMarkdown": "totally agree",
      "votes": null
    },
    {
      "id": "931557",
      "postDate": "07/16/2020 09:09:07",
      "content": "<p>Preprocessing and augmentation would give for sure a significant uplift during this CBR competition. If we had a serious test set, it could guide us in building such preprocessing  and augmentation pipelines. While I can understand the concern about leakage issues, I still think that hiding all the test set from Kagglers during such  a tough competition isn't necessarily optimal for both organizers and kagglers : <strong>garbage in, garbage out</strong> !</p>",
      "rawMarkdown": "Preprocessing and augmentation would give for sure a significant uplift during this CBR competition. If we had a serious test set, it could guide us in building such preprocessing  and augmentation pipelines. While I can understand the concern about leakage issues, I still think that hiding all the test set from Kagglers during such  a tough competition isn't necessarily optimal for both organizers and kagglers : **garbage in, garbage out** !",
      "votes": null
    },
    {
      "id": "931751",
      "postDate": "07/16/2020 12:25:42",
      "content": "<p>Data quality - some mp3 files are not decoded properly however they have labels. Published kernels assume empty signals however we know nothing about the test test, i.e .whether we had a decoding error or not.</p>",
      "rawMarkdown": "Data quality - some mp3 files are not decoded properly however they have labels. Published kernels assume empty signals however we know nothing about the test test, i.e .whether we had a decoding error or not.",
      "votes": null
    },
    {
      "id": "932804",
      "postDate": "07/17/2020 09:48:49",
      "content": "<p>It could be useful to look at the secondary labels and consider this as multi-class.  Started to create a pseudo testset like these and found it would sometimes predict out of the secondary label correctly. <br>\nFor no call it may be an idea to augment with other sounds like from some of the Freesound competitions or some Urban Sounds databases.  And maybe thresholds need to be different per species or groups of species. Something like a crow is maybe easier to recognise than songbirds with varied vocalisations.  </p>",
      "rawMarkdown": "It could be useful to look at the secondary labels and consider this as multi-class.  Started to create a pseudo testset like these and found it would sometimes predict out of the secondary label correctly.  \nFor no call it may be an idea to augment with other sounds like from some of the Freesound competitions or some Urban Sounds databases.  And maybe thresholds need to be different per species or groups of species. Something like a crow is maybe easier to recognise than songbirds with varied vocalisations.",
      "votes": null
    },
    {
      "id": "932877",
      "postDate": "07/17/2020 11:12:10",
      "content": "<p>Those ideas are absolutely great ! It would be interesting to know if someone have successfully tried some of them. I was doing some preprocessing to split audios into call/nocall before classification but I can't tell If I was successful in it.</p>",
      "rawMarkdown": "Those ideas are absolutely great ! It would be interesting to know if someone have successfully tried some of them. I was doing some preprocessing to split audios into call/nocall before classification but I can't tell If I was successful in it.",
      "votes": null
    },
    {
      "id": "932988",
      "postDate": "07/17/2020 12:07:24",
      "content": "<p>I've actually tried multi-label and pseudo labeling on oof (to get strong labels on train set)</p>\n\n<ol>\n<li>multi-label: worth doing</li>\n<li>pseudo labeling: we'll need good model otherwise we'll face many false positives which may corrupt the dataset.</li>\n</ol>\n\n<p>I think that once we get strong labels on train set, it would be much easier for us to mix train data with some background sound, so I'm still trying to do pseudo labeling on train set.</p>",
      "rawMarkdown": "I've actually tried multi-label and pseudo labeling on oof (to get strong labels on train set)\n\n1. multi-label: worth doing\n2. pseudo labeling: we'll need good model otherwise we'll face many false positives which may corrupt the dataset.\n\nI think that once we get strong labels on train set, it would be much easier for us to mix train data with some background sound, so I'm still trying to do pseudo labeling on train set.",
      "votes": null
    },
    {
      "id": "933706",
      "postDate": "07/17/2020 23:53:29",
      "content": "<p>Nice to know that pseudo labeling is just working fine … finally something joyceful among all those \"this thing is not wroking\" comments xD </p>",
      "rawMarkdown": "Nice to know that pseudo labeling is just working fine ... finally something joyceful among all those \"this thing is not wroking\" comments xD",
      "votes": null
    },
    {
      "id": "933708",
      "postDate": "07/17/2020 23:55:54",
      "content": "<p>For the multi label thing  <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , are you using the <strong>species</strong> column in the train meta data ?</p>",
      "rawMarkdown": "For the multi label thing  @hidehisaarai1213 , are you using the **species** column in the train meta data ?",
      "votes": null
    },
    {
      "id": "933728",
      "postDate": "07/18/2020 00:58:56",
      "content": "<p>I'm using <code>secondary_labels</code></p>",
      "rawMarkdown": "I'm using `secondary_labels`",
      "votes": null
    },
    {
      "id": "933750",
      "postDate": "07/18/2020 01:42:35",
      "content": "<blockquote>\n  <p>Nice to know that pseudo labeling is just working fine … </p>\n</blockquote>\n\n<p>What I meant was that it's not working fine so far... but I believe it's good to pursue on this line.</p>",
      "rawMarkdown": "&gt; Nice to know that pseudo labeling is just working fine … \n\nWhat I meant was that it's not working fine so far... but I believe it's good to pursue on this line.",
      "votes": null
    },
    {
      "id": "933955",
      "postDate": "07/18/2020 06:50:35",
      "content": "<blockquote>\n  <p>What I meant was that it's not working fine so far…</p>\n</blockquote>\n<p>smh 🙁</p>",
      "rawMarkdown": "&gt; What I meant was that it's not working fine so far…\n\nsmh 🙁",
      "votes": null
    },
    {
      "id": "934585",
      "postDate": "07/18/2020 15:03:48",
      "content": "<p>very helpfull</p>",
      "rawMarkdown": "very helpfull",
      "votes": null
    },
    {
      "id": "942441",
      "postDate": "07/23/2020 18:37:30",
      "content": "<p>Nice Insight. Thks <a href=\"/kneroma\">@kneroma</a> </p>",
      "rawMarkdown": "Nice Insight. Thks @kneroma",
      "votes": null
    },
    {
      "id": "969460",
      "postDate": "08/13/2020 17:43:29",
      "content": "<blockquote>\n  <p>There is a lot of \"nocall\" in the test set (~46%)</p>\n</blockquote>\n<p>How do you know this?</p>",
      "rawMarkdown": "> There is a lot of \"nocall\" in the test set (~46%)\n\nHow do you know this?",
      "votes": null
    },
    {
      "id": "981073",
      "postDate": "08/22/2020 06:38:45",
      "content": "<p>Yeah, <a href=\"https://www.kaggle.com/kkillamsetti\" target=\"_blank\">@kkillamsetti</a>, what makes you say that nocall in the test dataset is ~46%? I made a rough estimation based on an all-nocall submission and I came up with between 57% and 63% but maybe my math is off. </p>",
      "rawMarkdown": "Yeah, @kkillamsetti, what makes you say that nocall in the test dataset is ~46%? I made a rough estimation based on an all-nocall submission and I came up with between 57% and 63% but maybe my math is off.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 934585,
      "author_name": "chaitanya565656",
      "author_url": "",
      "post_date": "07/18/2020 15:03:48",
      "content": "<p>very helpfull</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 969460,
      "author_name": "marcogorelli",
      "author_url": "",
      "post_date": "08/13/2020 17:43:29",
      "content": "<blockquote>\n  <p>There is a lot of \"nocall\" in the test set (~46%)</p>\n</blockquote>\n<p>How do you know this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 981073,
          "author_name": "lewington",
          "author_url": "",
          "post_date": "08/22/2020 06:38:45",
          "content": "<p>Yeah, <a href=\"https://www.kaggle.com/kkillamsetti\" target=\"_blank\">@kkillamsetti</a>, what makes you say that nocall in the test dataset is ~46%? I made a rough estimation based on an all-nocall submission and I came up with between 57% and 63% but maybe my math is off. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 930939,
      "author_name": "ttahara",
      "author_url": "",
      "post_date": "07/15/2020 20:29:39",
      "content": "<p>Thanks for summarization!</p>\n\n<p>These are difficult but interesting points of this competition. </p>",
      "votes": null,
      "replies": [
        {
          "id": 931016,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "07/15/2020 22:37:42",
          "content": "<p>Yes...very challenging and of course interesting 👍 </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 931032,
      "author_name": "hidehisaarai1213",
      "author_url": "",
      "post_date": "07/15/2020 22:56:50",
      "content": "<p>Here are some others I have found</p>\n\n<ul>\n<li>There are many types of song / calls even in a single species</li>\n<li>Annotation to the training set is actually not enough\n<ul><li>Some of them have songs/calls of multiple species in fact, but not annotated</li>\n<li>Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training</li></ul></li>\n<li>We are not sure about the annotation level in test set\n<ul><li>Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.</li></ul></li>\n<li>Some calls of different species are very similar and not so distinguishable.</li>\n</ul>\n\n<p>The host is also sharing some of the challenges <a href=\"https://www.kaggle.com/c/birdsong-recognition/discussion/163899#914790\">here</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 931128,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "07/16/2020 02:19:09",
          "content": "<p>Thanks for pointing out those points. I've reported them ! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 931299,
      "author_name": "gaborfodor",
      "author_url": "",
      "post_date": "07/16/2020 05:34:59",
      "content": "<p>My biggest challenge is the lack of proper validation set. Sound-scape recordings could have completely different characteristics than the training set.</p>\n\n<p>I think this competition requires relatively complex pipeline (audio preprocessing, call/not call classifier, augmentation, CNN, prediction post-processing) compared to the typical kaggle competitions. It might also explain the slow progress.</p>",
      "votes": null,
      "replies": [
        {
          "id": 931356,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/16/2020 06:31:09",
          "content": "<p>totally agree</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 931557,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "07/16/2020 09:09:07",
          "content": "<p>Preprocessing and augmentation would give for sure a significant uplift during this CBR competition. If we had a serious test set, it could guide us in building such preprocessing  and augmentation pipelines. While I can understand the concern about leakage issues, I still think that hiding all the test set from Kagglers during such  a tough competition isn't necessarily optimal for both organizers and kagglers : <strong>garbage in, garbage out</strong> !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932804,
          "author_name": "something4kag",
          "author_url": "",
          "post_date": "07/17/2020 09:48:49",
          "content": "<p>It could be useful to look at the secondary labels and consider this as multi-class.  Started to create a pseudo testset like these and found it would sometimes predict out of the secondary label correctly. <br>\nFor no call it may be an idea to augment with other sounds like from some of the Freesound competitions or some Urban Sounds databases.  And maybe thresholds need to be different per species or groups of species. Something like a crow is maybe easier to recognise than songbirds with varied vocalisations.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932877,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "07/17/2020 11:12:10",
          "content": "<p>Those ideas are absolutely great ! It would be interesting to know if someone have successfully tried some of them. I was doing some preprocessing to split audios into call/nocall before classification but I can't tell If I was successful in it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 932988,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/17/2020 12:07:24",
          "content": "<p>I've actually tried multi-label and pseudo labeling on oof (to get strong labels on train set)</p>\n\n<ol>\n<li>multi-label: worth doing</li>\n<li>pseudo labeling: we'll need good model otherwise we'll face many false positives which may corrupt the dataset.</li>\n</ol>\n\n<p>I think that once we get strong labels on train set, it would be much easier for us to mix train data with some background sound, so I'm still trying to do pseudo labeling on train set.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933706,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "07/17/2020 23:53:29",
          "content": "<p>Nice to know that pseudo labeling is just working fine … finally something joyceful among all those \"this thing is not wroking\" comments xD </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933708,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "07/17/2020 23:55:54",
          "content": "<p>For the multi label thing  <a href=\"https://www.kaggle.com/hidehisaarai1213\" target=\"_blank\">@hidehisaarai1213</a> , are you using the <strong>species</strong> column in the train meta data ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933728,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/18/2020 00:58:56",
          "content": "<p>I'm using <code>secondary_labels</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933750,
          "author_name": "hidehisaarai1213",
          "author_url": "",
          "post_date": "07/18/2020 01:42:35",
          "content": "<blockquote>\n  <p>Nice to know that pseudo labeling is just working fine … </p>\n</blockquote>\n\n<p>What I meant was that it's not working fine so far... but I believe it's good to pursue on this line.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 933955,
          "author_name": "kneroma",
          "author_url": "",
          "post_date": "07/18/2020 06:50:35",
          "content": "<blockquote>\n  <p>What I meant was that it's not working fine so far…</p>\n</blockquote>\n<p>smh 🙁</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 931751,
      "author_name": "snovik1975",
      "author_url": "",
      "post_date": "07/16/2020 12:25:42",
      "content": "<p>Data quality - some mp3 files are not decoded properly however they have labels. Published kernels assume empty signals however we know nothing about the test test, i.e .whether we had a decoding error or not.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 942441,
      "author_name": "ulrich07",
      "author_url": "",
      "post_date": "07/23/2020 18:37:30",
      "content": "<p>Nice Insight. Thks <a href=\"/kneroma\">@kneroma</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "930613": "Three weeks now since the start of this CBR competition and we could hardly overtake the **sample_submission** ! During my experiments, I come up with these challenging points which require specific attention:\n\n* Training set is somehow clean but test set is very noisy\n*  There is a lot of \"**nocall**\" in the test set (~46%)\n* **Nocall** prediction by thresholding is somehow subjective (need a cross-val there but ...)\n* Testing is not obvious as we have no significant sample of the test set\n* Training set contains one bird at a time while test set does contain many\n* ....\n\n\nThose points make me believing that preprocessing will play a huge role during this CBR competition.\n\n#Update 1 (@hidehisaarai1213)\n* There are many types of song / calls even in a single species\n* Annotation to the training set is actually not enough\n  * Some of them have songs/calls of multiple species in fact, but not annotated\n  * Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training\n* We are not sure about the annotation level in test set\n   * Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.\n* Some calls of different species are very similar and not so distinguishable.\n\n**PS**: Don't mind adding your challenging points in the comment section, I will manage to report them all !",
    "930939": "Thanks for summarization!\n\nThese are difficult but interesting points of this competition.",
    "931016": "Yes...very challenging and of course interesting 👍",
    "931032": "Here are some others I have found\n\n* There are many types of song / calls even in a single species\n* Annotation to the training set is actually not enough\n  - Some of them have songs/calls of multiple species in fact, but not annotated\n  - Many part of the training set is background, so we may label non-call events as call events if we crop some short period out of the original clip in training\n* We are not sure about the annotation level in test set\n  - Whether they put a label even on a call event of small volume or not. CNN is quite powerful that we can detect small-volume-events.\n* Some calls of different species are very similar and not so distinguishable.\n\nThe host is also sharing some of the challenges [here](https://www.kaggle.com/c/birdsong-recognition/discussion/163899#914790).",
    "931128": "Thanks for pointing out those points. I've reported them !",
    "931299": "My biggest challenge is the lack of proper validation set. Sound-scape recordings could have completely different characteristics than the training set.\n\nI think this competition requires relatively complex pipeline (audio preprocessing, call/not call classifier, augmentation, CNN, prediction post-processing) compared to the typical kaggle competitions. It might also explain the slow progress.",
    "931356": "totally agree",
    "931557": "Preprocessing and augmentation would give for sure a significant uplift during this CBR competition. If we had a serious test set, it could guide us in building such preprocessing  and augmentation pipelines. While I can understand the concern about leakage issues, I still think that hiding all the test set from Kagglers during such  a tough competition isn't necessarily optimal for both organizers and kagglers : **garbage in, garbage out** !",
    "931751": "Data quality - some mp3 files are not decoded properly however they have labels. Published kernels assume empty signals however we know nothing about the test test, i.e .whether we had a decoding error or not.",
    "932804": "It could be useful to look at the secondary labels and consider this as multi-class.  Started to create a pseudo testset like these and found it would sometimes predict out of the secondary label correctly.  \nFor no call it may be an idea to augment with other sounds like from some of the Freesound competitions or some Urban Sounds databases.  And maybe thresholds need to be different per species or groups of species. Something like a crow is maybe easier to recognise than songbirds with varied vocalisations.",
    "932877": "Those ideas are absolutely great ! It would be interesting to know if someone have successfully tried some of them. I was doing some preprocessing to split audios into call/nocall before classification but I can't tell If I was successful in it.",
    "932988": "I've actually tried multi-label and pseudo labeling on oof (to get strong labels on train set)\n\n1. multi-label: worth doing\n2. pseudo labeling: we'll need good model otherwise we'll face many false positives which may corrupt the dataset.\n\nI think that once we get strong labels on train set, it would be much easier for us to mix train data with some background sound, so I'm still trying to do pseudo labeling on train set.",
    "933706": "Nice to know that pseudo labeling is just working fine ... finally something joyceful among all those \"this thing is not wroking\" comments xD",
    "933708": "For the multi label thing  @hidehisaarai1213 , are you using the **species** column in the train meta data ?",
    "933728": "I'm using `secondary_labels`",
    "933750": "&gt; Nice to know that pseudo labeling is just working fine … \n\nWhat I meant was that it's not working fine so far... but I believe it's good to pursue on this line.",
    "933955": "&gt; What I meant was that it's not working fine so far…\n\nsmh 🙁",
    "934585": "very helpfull",
    "942441": "Nice Insight. Thks @kneroma",
    "969460": "> There is a lot of \"nocall\" in the test set (~46%)\n\nHow do you know this?",
    "981073": "Yeah, @kkillamsetti, what makes you say that nocall in the test dataset is ~46%? I made a rough estimation based on an all-nocall submission and I came up with between 57% and 63% but maybe my math is off."
  },
  "source": "meta"
}