{
  "id": 319603,
  "title": "few-shot learning demo - prevent overfitting to few sample",
  "url": "/competitions/birdclef-2022/discussion/319603",
  "author_name": "",
  "post_date": "2022-04-18T06:14:16.274956Z",
  "votes": 15,
  "comment_count": 1,
  "views": 0,
  "content": "<h1>TL;DR</h1>\n<p>I released <a href=\"https://www.kaggle.com/code/tatamikenn/birdcleff22-a-working-demo-for-few-shot-learning?scriptVersionId=93308494\" target=\"_blank\">a notebook</a> to train audio classifier using few-shot learning.</p>\n<p>In short, few-shot learning is a training technique to learn a classifier without overfitting to train data even if training samples per class are extremely few (1-10 samples/class is basic settings).</p>\n<p>As you know, part of the scored species are rare and there are literally a few samples in this dataset (e.g. <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species\" target=\"_blank\">see this EDA notebook</a>). So, few-shot learning is suitable to the situation.</p>\n<h2>What this notebook provides</h2>\n<ol>\n<li>a working demo of few-shot learning on BirdCLEF 2022 dataset (Prototypical network[1])</li>\n<li>a learned model of good classification accuracy. The condition and performance of the learned model is:<ul>\n<li>5-way(5 classes), 5-shot(5 samples per class) training, evaluating for new class (not appeared on train set)</li>\n<li>using small model (~0.31M trainable parameters)</li>\n<li>train 10 epochs (10,000 episodes)</li>\n<li>the mean classification accuracy are ~76%</li></ul></li>\n</ol>\n<h2>What this notebook doesn't provide</h2>\n<ul>\n<li>call/no call classifier<ul>\n<li>since the models are trained only by positive (bird call) samples, it doesn't support distinguishing background from bird call (; maybe you need another classifier, or you need to explicitly feed background samples to the model).</li></ul></li>\n<li>multi-label classification<ul>\n<li>it only classifies most probable 1 class with inputting 5 second audio frame</li>\n<li><code>secondary_labels</code> are completely ignored</li></ul></li>\n</ul>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/1703.05175\" target=\"_blank\">https://arxiv.org/abs/1703.05175</a> - Prototypical network</li>\n<li>[2] <a href=\"https://github.com/Frankluox/LightningFSL\" target=\"_blank\">https://github.com/Frankluox/LightningFSL</a> - the training code is mainly adopted from the repository</li>\n</ul>",
  "messages": [
    {
      "id": "1758895",
      "postDate": "04/18/2022 06:14:16",
      "content": "<h1>TL;DR</h1>\n<p>I released <a href=\"https://www.kaggle.com/code/tatamikenn/birdcleff22-a-working-demo-for-few-shot-learning?scriptVersionId=93308494\" target=\"_blank\">a notebook</a> to train audio classifier using few-shot learning.</p>\n<p>In short, few-shot learning is a training technique to learn a classifier without overfitting to train data even if training samples per class are extremely few (1-10 samples/class is basic settings).</p>\n<p>As you know, part of the scored species are rare and there are literally a few samples in this dataset (e.g. <a href=\"https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species\" target=\"_blank\">see this EDA notebook</a>). So, few-shot learning is suitable to the situation.</p>\n<h2>What this notebook provides</h2>\n<ol>\n<li>a working demo of few-shot learning on BirdCLEF 2022 dataset (Prototypical network[1])</li>\n<li>a learned model of good classification accuracy. The condition and performance of the learned model is:<ul>\n<li>5-way(5 classes), 5-shot(5 samples per class) training, evaluating for new class (not appeared on train set)</li>\n<li>using small model (~0.31M trainable parameters)</li>\n<li>train 10 epochs (10,000 episodes)</li>\n<li>the mean classification accuracy are ~76%</li></ul></li>\n</ol>\n<h2>What this notebook doesn't provide</h2>\n<ul>\n<li>call/no call classifier<ul>\n<li>since the models are trained only by positive (bird call) samples, it doesn't support distinguishing background from bird call (; maybe you need another classifier, or you need to explicitly feed background samples to the model).</li></ul></li>\n<li>multi-label classification<ul>\n<li>it only classifies most probable 1 class with inputting 5 second audio frame</li>\n<li><code>secondary_labels</code> are completely ignored</li></ul></li>\n</ul>\n<h1>Reference</h1>\n<ul>\n<li>[1] <a href=\"https://arxiv.org/abs/1703.05175\" target=\"_blank\">https://arxiv.org/abs/1703.05175</a> - Prototypical network</li>\n<li>[2] <a href=\"https://github.com/Frankluox/LightningFSL\" target=\"_blank\">https://github.com/Frankluox/LightningFSL</a> - the training code is mainly adopted from the repository</li>\n</ul>",
      "rawMarkdown": "# TL;DR\n\nI released [a notebook] to train audio classifier using few-shot learning.\n\n[a notebook]: https://www.kaggle.com/code/tatamikenn/birdcleff22-a-working-demo-for-few-shot-learning?scriptVersionId=93308494\n\nIn short, few-shot learning is a training technique to learn a classifier without overfitting to train data even if training samples per class are extremely few (1-10 samples/class is basic settings).\n\nAs you know, part of the scored species are rare and there are literally a few samples in this dataset (e.g. [see this EDA notebook](https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species)). So, few-shot learning is suitable to the situation.\n\n## What this notebook provides\n\n1. a working demo of few-shot learning on BirdCLEF 2022 dataset (Prototypical network[1])\n2. a learned model of good classification accuracy. The condition and performance of the learned model is:\n    - 5-way(5 classes), 5-shot(5 samples per class) training, evaluating for new class (not appeared on train set)\n    - using small model (~0.31M trainable parameters)\n    - train 10 epochs (10,000 episodes)\n    - the mean classification accuracy are ~76%\n\n## What this notebook doesn't provide\n\n- call/no call classifier\n    - since the models are trained only by positive (bird call) samples, it doesn't support distinguishing background from bird call (; maybe you need another classifier, or you need to explicitly feed background samples to the model).\n- multi-label classification\n    - it only classifies most probable 1 class with inputting 5 second audio frame\n    - `secondary_labels` are completely ignored\n\n# Reference\n\n* [1] https://arxiv.org/abs/1703.05175 - Prototypical network\n* [2] https://github.com/Frankluox/LightningFSL - the training code is mainly adopted from the repository",
      "votes": null
    },
    {
      "id": "3282533",
      "postDate": "09/06/2025 12:37:48",
      "content": "<p>Thank you ! You really helped me</p>",
      "rawMarkdown": "Thank you ! You really helped me",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3282533,
      "author_name": "cisko6koman",
      "author_url": "",
      "post_date": "09/06/2025 12:37:48",
      "content": "<p>Thank you ! You really helped me</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1758895": "# TL;DR\n\nI released [a notebook] to train audio classifier using few-shot learning.\n\n[a notebook]: https://www.kaggle.com/code/tatamikenn/birdcleff22-a-working-demo-for-few-shot-learning?scriptVersionId=93308494\n\nIn short, few-shot learning is a training technique to learn a classifier without overfitting to train data even if training samples per class are extremely few (1-10 samples/class is basic settings).\n\nAs you know, part of the scored species are rare and there are literally a few samples in this dataset (e.g. [see this EDA notebook](https://www.kaggle.com/code/tatamikenn/birdclef22-eda-on-the-scored-species)). So, few-shot learning is suitable to the situation.\n\n## What this notebook provides\n\n1. a working demo of few-shot learning on BirdCLEF 2022 dataset (Prototypical network[1])\n2. a learned model of good classification accuracy. The condition and performance of the learned model is:\n    - 5-way(5 classes), 5-shot(5 samples per class) training, evaluating for new class (not appeared on train set)\n    - using small model (~0.31M trainable parameters)\n    - train 10 epochs (10,000 episodes)\n    - the mean classification accuracy are ~76%\n\n## What this notebook doesn't provide\n\n- call/no call classifier\n    - since the models are trained only by positive (bird call) samples, it doesn't support distinguishing background from bird call (; maybe you need another classifier, or you need to explicitly feed background samples to the model).\n- multi-label classification\n    - it only classifies most probable 1 class with inputting 5 second audio frame\n    - `secondary_labels` are completely ignored\n\n# Reference\n\n* [1] https://arxiv.org/abs/1703.05175 - Prototypical network\n* [2] https://github.com/Frankluox/LightningFSL - the training code is mainly adopted from the repository",
    "3282533": "Thank you ! You really helped me"
  },
  "source": "meta"
}