{
  "id": 413383,
  "title": "What did not work for me",
  "url": "/competitions/birdclef-2023/discussion/413383",
  "author_name": "",
  "post_date": "2023-05-28T12:03:35.331021500Z",
  "votes": 8,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello Kagglers,</p>\n<p>Big thanks to the hosts once again for resuming the bird series - with even a higher prize this year! Congratulations to my last year's teammate Volodymyr for his well-deserved and strong finish.</p>\n<p>I'd like to share my negative learnings about this year's competition. Primarily, this year I had no desire to replicate the methodology of the previous year and aimed to attempt something entirely different. The most popular approach to audio classification nowadays using a mel-spectrogram (which is an intermediate, not-learned representation). A pre-trained image backbone is fine-tuned using mel representation and a new downstream task (such as bird classification) is learned. The desire was to entirely move away from mel-spectrograms.</p>\n<p>I was in luck! <a href=\"https://ai.honu.io/papers/encodec/samples.html\" target=\"_blank\">EnCodec</a>, a relatively recent method, provides an alternative to long-standing mel-spectrograms. The approach is quite popular in audio research recently and used by works like MusicLM, VALL-E, Diffsound, and many others. EnCodec is a residual vector quantizer. It transforms a continuous waveform [channel, time] to a discrete representation of [number_of_residuals, frames]. This highly compressed representation allows us to transform 5 seconds of audio into 375 discrete tokens (frames) with a vocabulary of 1024. The discrete audio representation enables something entirely new, which is casting audio classification similar to the text classification (as opposed to image classification). </p>\n<p>My baseline was the combination of <a href=\"https://github.com/karpathy/nanoGPT\" target=\"_blank\">nanoGPT</a> and <a href=\"https://github.com/facebookresearch/encodec\" target=\"_blank\">EnCodec</a>. The plan was to quantize bird audios into sequences of discrete tokens using EnCodec. Later, fine-tuning a Transformer model. In this case,  I used a pre-trained GPT-2 with several modifications such as turning causal attention to full attention and replacing embedding and linear layers to adopt the new downstream task. </p>\n<h3>What didn't work?</h3>\n<p>Obviously, the idea wasn't competitive with brilliant Kagglers' work. The biggest stopper for my research was CPU runtime. Although I optimised both the Transformer and EnCodec for the CPU runtime (onnx, quantisation, other tricks) it still took around 35-40 seconds to predict 10 minutes of audio, without optimisations it took around a minute and 20 seconds. Ensembling wasn't too practical after these timings. Another issue was not being able to fit a pre-trained network as it is. For example, a regular GPT-2 (12 layers) + EnCodec results in a time-out. Therefore, I ended up going with much smaller Transformer sizes (i.e., by pruning the last 6-8 layers). Unfortunately, small data and small model approaches seem to be challenging with Transformers. </p>\n<p>Finally, although the approach sounds promising, using CNN + mel-spectrogram would be the strategically correct decision for this competition. However, for other scales such as GPU-enabled runtime and 100x more data, the proposed approach could become relevant and perhaps scale better.</p>\n<p>I am hoping to investigate the idea further and hopefully turn it into a paper submission. Negative learnings from BirdCLEF2023 should have it's own section😄.</p>",
  "messages": [
    {
      "id": "2278162",
      "postDate": "05/28/2023 12:03:35",
      "content": "<p>Hello Kagglers,</p>\n<p>Big thanks to the hosts once again for resuming the bird series - with even a higher prize this year! Congratulations to my last year's teammate Volodymyr for his well-deserved and strong finish.</p>\n<p>I'd like to share my negative learnings about this year's competition. Primarily, this year I had no desire to replicate the methodology of the previous year and aimed to attempt something entirely different. The most popular approach to audio classification nowadays using a mel-spectrogram (which is an intermediate, not-learned representation). A pre-trained image backbone is fine-tuned using mel representation and a new downstream task (such as bird classification) is learned. The desire was to entirely move away from mel-spectrograms.</p>\n<p>I was in luck! <a href=\"https://ai.honu.io/papers/encodec/samples.html\" target=\"_blank\">EnCodec</a>, a relatively recent method, provides an alternative to long-standing mel-spectrograms. The approach is quite popular in audio research recently and used by works like MusicLM, VALL-E, Diffsound, and many others. EnCodec is a residual vector quantizer. It transforms a continuous waveform [channel, time] to a discrete representation of [number_of_residuals, frames]. This highly compressed representation allows us to transform 5 seconds of audio into 375 discrete tokens (frames) with a vocabulary of 1024. The discrete audio representation enables something entirely new, which is casting audio classification similar to the text classification (as opposed to image classification). </p>\n<p>My baseline was the combination of <a href=\"https://github.com/karpathy/nanoGPT\" target=\"_blank\">nanoGPT</a> and <a href=\"https://github.com/facebookresearch/encodec\" target=\"_blank\">EnCodec</a>. The plan was to quantize bird audios into sequences of discrete tokens using EnCodec. Later, fine-tuning a Transformer model. In this case,  I used a pre-trained GPT-2 with several modifications such as turning causal attention to full attention and replacing embedding and linear layers to adopt the new downstream task. </p>\n<h3>What didn't work?</h3>\n<p>Obviously, the idea wasn't competitive with brilliant Kagglers' work. The biggest stopper for my research was CPU runtime. Although I optimised both the Transformer and EnCodec for the CPU runtime (onnx, quantisation, other tricks) it still took around 35-40 seconds to predict 10 minutes of audio, without optimisations it took around a minute and 20 seconds. Ensembling wasn't too practical after these timings. Another issue was not being able to fit a pre-trained network as it is. For example, a regular GPT-2 (12 layers) + EnCodec results in a time-out. Therefore, I ended up going with much smaller Transformer sizes (i.e., by pruning the last 6-8 layers). Unfortunately, small data and small model approaches seem to be challenging with Transformers. </p>\n<p>Finally, although the approach sounds promising, using CNN + mel-spectrogram would be the strategically correct decision for this competition. However, for other scales such as GPU-enabled runtime and 100x more data, the proposed approach could become relevant and perhaps scale better.</p>\n<p>I am hoping to investigate the idea further and hopefully turn it into a paper submission. Negative learnings from BirdCLEF2023 should have it's own section😄.</p>",
      "rawMarkdown": "Hello Kagglers,\n\nBig thanks to the hosts once again for resuming the bird series - with even a higher prize this year! Congratulations to my last year's teammate Volodymyr for his well-deserved and strong finish.\n\nI'd like to share my negative learnings about this year's competition. Primarily, this year I had no desire to replicate the methodology of the previous year and aimed to attempt something entirely different. The most popular approach to audio classification nowadays using a mel-spectrogram (which is an intermediate, not-learned representation). A pre-trained image backbone is fine-tuned using mel representation and a new downstream task (such as bird classification) is learned. The desire was to entirely move away from mel-spectrograms.\n\nI was in luck! [EnCodec](https://ai.honu.io/papers/encodec/samples.html), a relatively recent method, provides an alternative to long-standing mel-spectrograms. The approach is quite popular in audio research recently and used by works like MusicLM, VALL-E, Diffsound, and many others. EnCodec is a residual vector quantizer. It transforms a continuous waveform [channel, time] to a discrete representation of [number_of_residuals, frames]. This highly compressed representation allows us to transform 5 seconds of audio into 375 discrete tokens (frames) with a vocabulary of 1024. The discrete audio representation enables something entirely new, which is casting audio classification similar to the text classification (as opposed to image classification). \n\nMy baseline was the combination of [nanoGPT](https://github.com/karpathy/nanoGPT) and [EnCodec](https://github.com/facebookresearch/encodec). The plan was to quantize bird audios into sequences of discrete tokens using EnCodec. Later, fine-tuning a Transformer model. In this case,  I used a pre-trained GPT-2 with several modifications such as turning causal attention to full attention and replacing embedding and linear layers to adopt the new downstream task. \n\n### What didn't work?\n\nObviously, the idea wasn't competitive with brilliant Kagglers' work. The biggest stopper for my research was CPU runtime. Although I optimised both the Transformer and EnCodec for the CPU runtime (onnx, quantisation, other tricks) it still took around 35-40 seconds to predict 10 minutes of audio, without optimisations it took around a minute and 20 seconds. Ensembling wasn't too practical after these timings. Another issue was not being able to fit a pre-trained network as it is. For example, a regular GPT-2 (12 layers) + EnCodec results in a time-out. Therefore, I ended up going with much smaller Transformer sizes (i.e., by pruning the last 6-8 layers). Unfortunately, small data and small model approaches seem to be challenging with Transformers. \n\nFinally, although the approach sounds promising, using CNN + mel-spectrogram would be the strategically correct decision for this competition. However, for other scales such as GPU-enabled runtime and 100x more data, the proposed approach could become relevant and perhaps scale better.\n\nI am hoping to investigate the idea further and hopefully turn it into a paper submission. Negative learnings from BirdCLEF2023 should have it's own section😄.",
      "votes": null
    },
    {
      "id": "2280479",
      "postDate": "05/30/2023 06:07:54",
      "content": "<p>That's a really cool idea! Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ? <br>\nI wanted to know whether given the massive compute that you need, the score will indeed improve. </p>",
      "rawMarkdown": "That's a really cool idea! Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ? \nI wanted to know whether given the massive compute that you need, the score will indeed improve.",
      "votes": null
    },
    {
      "id": "2281195",
      "postDate": "05/30/2023 16:51:12",
      "content": "<p>Thanks for sharing! </p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "2291718",
      "postDate": "06/07/2023 19:14:01",
      "content": "<p>Hey, </p>\n<p>Thank you for your message.</p>\n<blockquote>\n  <p>Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ?</p>\n</blockquote>\n<p>Not yet, but also curious about this. I will get to it as soon as I find some time 😅</p>",
      "rawMarkdown": "Hey, \n\nThank you for your message.\n\n>  Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ?\n\nNot yet, but also curious about this. I will get to it as soon as I find some time 😅",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2280479,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "05/30/2023 06:07:54",
      "content": "<p>That's a really cool idea! Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ? <br>\nI wanted to know whether given the massive compute that you need, the score will indeed improve. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2291718,
          "author_name": "realsleim",
          "author_url": "",
          "post_date": "06/07/2023 19:14:01",
          "content": "<p>Hey, </p>\n<p>Thank you for your message.</p>\n<blockquote>\n  <p>Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ?</p>\n</blockquote>\n<p>Not yet, but also curious about this. I will get to it as soon as I find some time 😅</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2281195,
      "author_name": "sohier",
      "author_url": "",
      "post_date": "05/30/2023 16:51:12",
      "content": "<p>Thanks for sharing! </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2278162": "Hello Kagglers,\n\nBig thanks to the hosts once again for resuming the bird series - with even a higher prize this year! Congratulations to my last year's teammate Volodymyr for his well-deserved and strong finish.\n\nI'd like to share my negative learnings about this year's competition. Primarily, this year I had no desire to replicate the methodology of the previous year and aimed to attempt something entirely different. The most popular approach to audio classification nowadays using a mel-spectrogram (which is an intermediate, not-learned representation). A pre-trained image backbone is fine-tuned using mel representation and a new downstream task (such as bird classification) is learned. The desire was to entirely move away from mel-spectrograms.\n\nI was in luck! [EnCodec](https://ai.honu.io/papers/encodec/samples.html), a relatively recent method, provides an alternative to long-standing mel-spectrograms. The approach is quite popular in audio research recently and used by works like MusicLM, VALL-E, Diffsound, and many others. EnCodec is a residual vector quantizer. It transforms a continuous waveform [channel, time] to a discrete representation of [number_of_residuals, frames]. This highly compressed representation allows us to transform 5 seconds of audio into 375 discrete tokens (frames) with a vocabulary of 1024. The discrete audio representation enables something entirely new, which is casting audio classification similar to the text classification (as opposed to image classification). \n\nMy baseline was the combination of [nanoGPT](https://github.com/karpathy/nanoGPT) and [EnCodec](https://github.com/facebookresearch/encodec). The plan was to quantize bird audios into sequences of discrete tokens using EnCodec. Later, fine-tuning a Transformer model. In this case,  I used a pre-trained GPT-2 with several modifications such as turning causal attention to full attention and replacing embedding and linear layers to adopt the new downstream task. \n\n### What didn't work?\n\nObviously, the idea wasn't competitive with brilliant Kagglers' work. The biggest stopper for my research was CPU runtime. Although I optimised both the Transformer and EnCodec for the CPU runtime (onnx, quantisation, other tricks) it still took around 35-40 seconds to predict 10 minutes of audio, without optimisations it took around a minute and 20 seconds. Ensembling wasn't too practical after these timings. Another issue was not being able to fit a pre-trained network as it is. For example, a regular GPT-2 (12 layers) + EnCodec results in a time-out. Therefore, I ended up going with much smaller Transformer sizes (i.e., by pruning the last 6-8 layers). Unfortunately, small data and small model approaches seem to be challenging with Transformers. \n\nFinally, although the approach sounds promising, using CNN + mel-spectrogram would be the strategically correct decision for this competition. However, for other scales such as GPU-enabled runtime and 100x more data, the proposed approach could become relevant and perhaps scale better.\n\nI am hoping to investigate the idea further and hopefully turn it into a paper submission. Negative learnings from BirdCLEF2023 should have it's own section😄.",
    "2280479": "That's a really cool idea! Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ? \nI wanted to know whether given the massive compute that you need, the score will indeed improve.",
    "2281195": "Thanks for sharing!",
    "2291718": "Hey, \n\nThank you for your message.\n\n>  Did you have CV score comparison between yours and other winning method, or perhaps the score comparison with your last year winning solution ?\n\nNot yet, but also curious about this. I will get to it as soon as I find some time 😅"
  },
  "source": "meta"
}