{
  "id": 322317,
  "title": "Idea: using 2021 soundscape as a local evaluation dataset",
  "url": "/competitions/birdclef-2022/discussion/322317",
  "author_name": "",
  "post_date": "2022-05-01T11:57:04.627689400Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I will share my own CV strategy that I am currently considering.</p>\n<p>The base model is a model trained embedding by prototypical network[1-2]. The advantage of this model trained by meta learning is that we can classify species not used during training without training a new model.</p>\n<p>Thus, we can use this model to achieve local validation with the 2021 soundscape.</p>\n<p>We can achieve local validation by the following procedure:</p>\n<ol>\n<li>for the 48 bird species that appear in the 2021 training soundscape, randomly select 21 species (this step may be omitted).</li>\n<li>for short clips of 21 bird calls, use a binary discriminator to extract clips that contain the calls</li>\n<li>input the clips extracted in step 2 into a prototype trained encoder and generate embedding of the support set</li>\n<li>input 5-second clips from the 2021 training soundscape to the binary discriminator to discriminate call/nocall; for clips determined to be call, use the prototype encoder to classify and compute a score.</li>\n</ol>\n<p>The code for submission is also done by the same process as for CV generation. This time, however, the support set is generated using recordings of the 21 bird species to be evaluated in 2022.</p>\n<p>These processes are somewhat more complex, but they do address the issue of the lack of soundscapes in 2022.</p>\n<p><a href=\"https://ibb.co/fq5R7Xm\"><img src=\"https://i.ibb.co/175Vx0H/Screen-Shot-2022-05-01-at-21-20-05.png\" alt=\"Screen-Shot-2022-05-01-at-21-20-05\"></a></p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/319603\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/319603</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311005\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311005</a></li>\n</ul>",
  "messages": [
    {
      "id": "1773725",
      "postDate": "05/01/2022 11:57:04",
      "content": "<p>I will share my own CV strategy that I am currently considering.</p>\n<p>The base model is a model trained embedding by prototypical network[1-2]. The advantage of this model trained by meta learning is that we can classify species not used during training without training a new model.</p>\n<p>Thus, we can use this model to achieve local validation with the 2021 soundscape.</p>\n<p>We can achieve local validation by the following procedure:</p>\n<ol>\n<li>for the 48 bird species that appear in the 2021 training soundscape, randomly select 21 species (this step may be omitted).</li>\n<li>for short clips of 21 bird calls, use a binary discriminator to extract clips that contain the calls</li>\n<li>input the clips extracted in step 2 into a prototype trained encoder and generate embedding of the support set</li>\n<li>input 5-second clips from the 2021 training soundscape to the binary discriminator to discriminate call/nocall; for clips determined to be call, use the prototype encoder to classify and compute a score.</li>\n</ol>\n<p>The code for submission is also done by the same process as for CV generation. This time, however, the support set is generated using recordings of the 21 bird species to be evaluated in 2022.</p>\n<p>These processes are somewhat more complex, but they do address the issue of the lack of soundscapes in 2022.</p>\n<p><a href=\"https://ibb.co/fq5R7Xm\"><img src=\"https://i.ibb.co/175Vx0H/Screen-Shot-2022-05-01-at-21-20-05.png\" alt=\"Screen-Shot-2022-05-01-at-21-20-05\"></a></p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/319603\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/319603</a></li>\n<li>[2] <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/311005\" target=\"_blank\">https://www.kaggle.com/competitions/birdclef-2022/discussion/311005</a></li>\n</ul>",
      "rawMarkdown": "I will share my own CV strategy that I am currently considering.\n\nThe base model is a model trained embedding by prototypical network[1-2]. The advantage of this model trained by meta learning is that we can classify species not used during training without training a new model.\n\nThus, we can use this model to achieve local validation with the 2021 soundscape.\n\nWe can achieve local validation by the following procedure:\n\n1. for the 48 bird species that appear in the 2021 training soundscape, randomly select 21 species (this step may be omitted).\n2. for short clips of 21 bird calls, use a binary discriminator to extract clips that contain the calls\n3. input the clips extracted in step 2 into a prototype trained encoder and generate embedding of the support set\n4. input 5-second clips from the 2021 training soundscape to the binary discriminator to discriminate call/nocall; for clips determined to be call, use the prototype encoder to classify and compute a score.\n\nThe code for submission is also done by the same process as for CV generation. This time, however, the support set is generated using recordings of the 21 bird species to be evaluated in 2022.\n\nThese processes are somewhat more complex, but they do address the issue of the lack of soundscapes in 2022.\n\n<a href=\"https://ibb.co/fq5R7Xm\"><img src=\"https://i.ibb.co/175Vx0H/Screen-Shot-2022-05-01-at-21-20-05.png\" alt=\"Screen-Shot-2022-05-01-at-21-20-05\" border=\"0\"></a>\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/319603\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/311005",
      "votes": null
    },
    {
      "id": "1773729",
      "postDate": "05/01/2022 11:58:37",
      "content": "<p>This is an idea that has not been fully explored. I will share it when I see the prospect of its realization.</p>",
      "rawMarkdown": "This is an idea that has not been fully explored. I will share it when I see the prospect of its realization.",
      "votes": null
    },
    {
      "id": "1774260",
      "postDate": "05/02/2022 02:59:53",
      "content": "<p>I also used 2021 soundscape as a local evaluation dataset.</p>\n<p>In my case, I used this dataset for threshhold optimization.<br>\nBut it didn't work. Because 2021 soundscape has no positive data in <strong>birdclef-2022.</strong></p>",
      "rawMarkdown": "I also used 2021 soundscape as a local evaluation dataset.\n\nIn my case, I used this dataset for threshhold optimization.\nBut it didn't work. Because 2021 soundscape has no positive data in **birdclef-2022.**",
      "votes": null
    },
    {
      "id": "1774285",
      "postDate": "05/02/2022 03:52:48",
      "content": "<p>Unfortunately, you are right. The above method can only be used for meta learning methods such as prototypical networks.</p>\n<p>However, the 2021 soundscape can still be used to evaluate binary discriminant models.</p>",
      "rawMarkdown": "Unfortunately, you are right. The above method can only be used for meta learning methods such as prototypical networks.\n\nHowever, the 2021 soundscape can still be used to evaluate binary discriminant models.",
      "votes": null
    },
    {
      "id": "1774288",
      "postDate": "05/02/2022 03:55:23",
      "content": "<p>Another way I came up with was to artificially create evaluation data by embedding positive data for the year 2022 in an arbitrary no-call soundscape.<br>\nThis one requires human engineering to create strong labels.</p>",
      "rawMarkdown": "Another way I came up with was to artificially create evaluation data by embedding positive data for the year 2022 in an arbitrary no-call soundscape.\nThis one requires human engineering to create strong labels.",
      "votes": null
    },
    {
      "id": "1774465",
      "postDate": "05/02/2022 07:29:44",
      "content": "<p>This is a great idea - I'm looking forward to seeing the results!</p>",
      "rawMarkdown": "This is a great idea - I'm looking forward to seeing the results!",
      "votes": null
    },
    {
      "id": "1775386",
      "postDate": "05/03/2022 01:07:04",
      "content": "<p>The challenge in this approach is to extract recordings that contain only the target species as a support set.</p>\n<p>This is because if the background contains recordings other than the target species, the embedding of the recordings will move away from the true center of the embedding of the target species.</p>\n<p>The same applies to no-calls, but it can be selected by a binary classifier that determines call/no call.</p>\n<p>Methods for extracting recordings of only the target species are as follows</p>\n<ol>\n<li>using a source separation method such as MixIT[1], and a human discriminate the target sound source and others</li>\n<li>detecting time regions containing the target species by SED model</li>\n<li>human review of the audio and spectrogram to extract the time region containing the target species.</li>\n</ol>\n<p>1, 2 would require training an additional neuron network, and 3 would consume human resources.</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2006.12701\" target=\"_blank\">https://arxiv.org/abs/2006.12701</a></li>\n</ul>",
      "rawMarkdown": "The challenge in this approach is to extract recordings that contain only the target species as a support set.\n\nThis is because if the background contains recordings other than the target species, the embedding of the recordings will move away from the true center of the embedding of the target species.\n\nThe same applies to no-calls, but it can be selected by a binary classifier that determines call/no call.\n\nMethods for extracting recordings of only the target species are as follows\n1. using a source separation method such as MixIT[1], and a human discriminate the target sound source and others\n2. detecting time regions containing the target species by SED model\n3. human review of the audio and spectrogram to extract the time region containing the target species.\n\n1, 2 would require training an additional neuron network, and 3 would consume human resources.\n\n## Reference\n\n- [1] https://arxiv.org/abs/2006.12701",
      "votes": null
    },
    {
      "id": "1775392",
      "postDate": "05/03/2022 01:27:55",
      "content": "<p>To explain a little more neatly, the 21 species to be selected in the above process are different from the 21 species to be evaluated in 2022.<br>\nSince the prototypical network is a method of learning adaptability to an unknown sample, it is not required that the target species be exactly the same (although of course it is desirable that they be the same).</p>",
      "rawMarkdown": "To explain a little more neatly, the 21 species to be selected in the above process are different from the 21 species to be evaluated in 2022.\nSince the prototypical network is a method of learning adaptability to an unknown sample, it is not required that the target species be exactly the same (although of course it is desirable that they be the same).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1773729,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/01/2022 11:58:37",
      "content": "<p>This is an idea that has not been fully explored. I will share it when I see the prospect of its realization.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1774260,
      "author_name": "shinmurashinmura",
      "author_url": "",
      "post_date": "05/02/2022 02:59:53",
      "content": "<p>I also used 2021 soundscape as a local evaluation dataset.</p>\n<p>In my case, I used this dataset for threshhold optimization.<br>\nBut it didn't work. Because 2021 soundscape has no positive data in <strong>birdclef-2022.</strong></p>",
      "votes": null,
      "replies": [
        {
          "id": 1774285,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/02/2022 03:52:48",
          "content": "<p>Unfortunately, you are right. The above method can only be used for meta learning methods such as prototypical networks.</p>\n<p>However, the 2021 soundscape can still be used to evaluate binary discriminant models.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1774288,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/02/2022 03:55:23",
          "content": "<p>Another way I came up with was to artificially create evaluation data by embedding positive data for the year 2022 in an arbitrary no-call soundscape.<br>\nThis one requires human engineering to create strong labels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1775392,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "05/03/2022 01:27:55",
          "content": "<p>To explain a little more neatly, the 21 species to be selected in the above process are different from the 21 species to be evaluated in 2022.<br>\nSince the prototypical network is a method of learning adaptability to an unknown sample, it is not required that the target species be exactly the same (although of course it is desirable that they be the same).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1774465,
      "author_name": "",
      "author_url": "",
      "post_date": "05/02/2022 07:29:44",
      "content": "<p>This is a great idea - I'm looking forward to seeing the results!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1775386,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "05/03/2022 01:07:04",
      "content": "<p>The challenge in this approach is to extract recordings that contain only the target species as a support set.</p>\n<p>This is because if the background contains recordings other than the target species, the embedding of the recordings will move away from the true center of the embedding of the target species.</p>\n<p>The same applies to no-calls, but it can be selected by a binary classifier that determines call/no call.</p>\n<p>Methods for extracting recordings of only the target species are as follows</p>\n<ol>\n<li>using a source separation method such as MixIT[1], and a human discriminate the target sound source and others</li>\n<li>detecting time regions containing the target species by SED model</li>\n<li>human review of the audio and spectrogram to extract the time region containing the target species.</li>\n</ol>\n<p>1, 2 would require training an additional neuron network, and 3 would consume human resources.</p>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/2006.12701\" target=\"_blank\">https://arxiv.org/abs/2006.12701</a></li>\n</ul>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1773725": "I will share my own CV strategy that I am currently considering.\n\nThe base model is a model trained embedding by prototypical network[1-2]. The advantage of this model trained by meta learning is that we can classify species not used during training without training a new model.\n\nThus, we can use this model to achieve local validation with the 2021 soundscape.\n\nWe can achieve local validation by the following procedure:\n\n1. for the 48 bird species that appear in the 2021 training soundscape, randomly select 21 species (this step may be omitted).\n2. for short clips of 21 bird calls, use a binary discriminator to extract clips that contain the calls\n3. input the clips extracted in step 2 into a prototype trained encoder and generate embedding of the support set\n4. input 5-second clips from the 2021 training soundscape to the binary discriminator to discriminate call/nocall; for clips determined to be call, use the prototype encoder to classify and compute a score.\n\nThe code for submission is also done by the same process as for CV generation. This time, however, the support set is generated using recordings of the 21 bird species to be evaluated in 2022.\n\nThese processes are somewhat more complex, but they do address the issue of the lack of soundscapes in 2022.\n\n<a href=\"https://ibb.co/fq5R7Xm\"><img src=\"https://i.ibb.co/175Vx0H/Screen-Shot-2022-05-01-at-21-20-05.png\" alt=\"Screen-Shot-2022-05-01-at-21-20-05\" border=\"0\"></a>\n\n## Reference\n\n- [1] https://www.kaggle.com/competitions/birdclef-2022/discussion/319603\n- [2] https://www.kaggle.com/competitions/birdclef-2022/discussion/311005",
    "1773729": "This is an idea that has not been fully explored. I will share it when I see the prospect of its realization.",
    "1774260": "I also used 2021 soundscape as a local evaluation dataset.\n\nIn my case, I used this dataset for threshhold optimization.\nBut it didn't work. Because 2021 soundscape has no positive data in **birdclef-2022.**",
    "1774285": "Unfortunately, you are right. The above method can only be used for meta learning methods such as prototypical networks.\n\nHowever, the 2021 soundscape can still be used to evaluate binary discriminant models.",
    "1774288": "Another way I came up with was to artificially create evaluation data by embedding positive data for the year 2022 in an arbitrary no-call soundscape.\nThis one requires human engineering to create strong labels.",
    "1774465": "This is a great idea - I'm looking forward to seeing the results!",
    "1775386": "The challenge in this approach is to extract recordings that contain only the target species as a support set.\n\nThis is because if the background contains recordings other than the target species, the embedding of the recordings will move away from the true center of the embedding of the target species.\n\nThe same applies to no-calls, but it can be selected by a binary classifier that determines call/no call.\n\nMethods for extracting recordings of only the target species are as follows\n1. using a source separation method such as MixIT[1], and a human discriminate the target sound source and others\n2. detecting time regions containing the target species by SED model\n3. human review of the audio and spectrogram to extract the time region containing the target species.\n\n1, 2 would require training an additional neuron network, and 3 would consume human resources.\n\n## Reference\n\n- [1] https://arxiv.org/abs/2006.12701",
    "1775392": "To explain a little more neatly, the 21 species to be selected in the above process are different from the 21 species to be evaluated in 2022.\nSince the prototypical network is a method of learning adaptability to an unknown sample, it is not required that the target species be exactly the same (although of course it is desirable that they be the same)."
  },
  "source": "meta"
}